knitr::opts_chunk$set( collapse = TRUE, comment = "#>", fig.width = 7, fig.height = 5 )
In efficiency analysis, we often study firms that operate under different production environments. Steel producers using electric arc furnaces (EAF) face a different feasible set of input-output combinations than those using the blast furnace-basic oxygen furnace (BF-BOF) route. Hospitals in rural areas face different constraints than urban ones. Banks face different regulatory environments across jurisdictions.
Following Battese, Rao, and O'Donnell (2004) and O'Donnell, Rao, and Battese (2008), we conceive of a single industry metatechnology $T^$: the set of all input-output combinations that are technically feasible in the industry. Each group of firms operates within a restricted subset $T_j \subseteq T^$ of this metatechnology, where the restrictions arise from regulation, the physical environment, resource endowments, or the cost of switching production systems. Groups do not possess fundamentally different technologies; they face different restrictions of a common metatechnology.
Standard stochastic frontier analysis (SFA) or data envelopment analysis (DEA) applied to the pooled sample implicitly assumes that all firms have unrestricted access to the same technology set, an assumption that may be unrealistic. Estimating separate frontiers for each group solves this problem but makes efficiency scores incomparable across groups: a firm that is 90\% efficient relative to a less advanced group frontier may actually be less productive than a firm that is 70\% efficient relative to a more advanced frontier.
The metafrontier framework, introduced by Battese, Rao, and O'Donnell (2004) and extended by Huang, Huang, and Liu (2014) and O'Donnell, Rao, and Battese (2008), resolves this by:
$$TE^*_i = TE_i \times TGR_i$$
where:
The metafrontier package provides a unified interface for estimating
metafrontier models using both SFA and DEA approaches.
library(metafrontier)
The package includes simulate_metafrontier() for generating data from a
known data-generating process. This is useful for Monte Carlo studies and
for learning the package.
sim <- simulate_metafrontier( n_groups = 3, n_per_group = 200, beta_meta = c(1.0, 0.5, 0.3), # intercept, elasticity_1, elasticity_2 tech_gap = c(0, 0.25, 0.5), # intercept shifts (0 = best technology) sigma_u = c(0.2, 0.3, 0.4), # inefficiency SD per group sigma_v = 0.15, # noise SD seed = 42 ) str(sim$data[, c("log_y", "log_x1", "log_x2", "group")]) table(sim$data$group)
The simulation generates a Cobb-Douglas frontier: $$\ln y_i = \beta_0^{(j)} + \beta_1 \ln x_{1i} + \beta_2 \ln x_{2i} + v_i - u_i$$ where the intercept $\beta_0^{(j)} = \beta_0^* - \delta_j$ is shifted down from the metafrontier by the technology gap $\delta_j$ for group $j$.
fit <- metafrontier( log_y ~ log_x1 + log_x2, data = sim$data, group = "group", method = "sfa", meta_type = "deterministic" ) fit
The deterministic metafrontier is estimated in two stages:
objective = "lp", the default). The alternative minimum sum of
squared deviations criterion is available via objective = "qp";
both criteria are proposed by Battese, Rao, and O'Donnell (2004),
and the methods vignette discusses them in detail.This is the default method:
fit_det <- metafrontier( log_y ~ log_x1 + log_x2, data = sim$data, group = "group", meta_type = "deterministic" ) summary(fit_det)
The stochastic metafrontier replaces the LP in Stage 2 with a second-stage SFA, using the fitted group frontier values as the dependent variable:
$$\ln \hat{f}(x_i; \hat\beta_j) = x_i'\beta^ + v^_i - u^*_i$$
where $u^*_i \ge 0$ captures the technology gap stochastically. This provides a distributional framework for the TGR, enabling standard errors and hypothesis testing.
fit_sto <- metafrontier( log_y ~ log_x1 + log_x2, data = sim$data, group = "group", meta_type = "stochastic" ) summary(fit_sto)
The stochastic metafrontier provides a variance-covariance matrix:
vcov(fit_sto)
For a nonparametric approach, set method = "dea":
fit_dea <- metafrontier( log_y ~ log_x1 + log_x2, data = sim$data, group = "group", method = "dea", rts = "vrs" ) fit_dea
The DEA metafrontier computes:
Use efficiencies() to extract the three components of the decomposition:
te <- efficiencies(fit_det, type = "group") tgr <- efficiencies(fit_det, type = "tgr") te_star <- efficiencies(fit_det, type = "meta") # Verify the fundamental identity: TE* = TE x TGR all.equal(te_star, te * tgr)
For SFA fits, technical efficiencies are computed with the Battese and
Coelli (1988) conditional expectation estimator by default
(estimator = "bc88"). The Jondrow et al. (1982) estimator is computed
and stored alongside it, so you can switch without refitting:
te_jlms <- efficiencies(fit_det, type = "group", estimator = "jlms") cor(te, te_jlms)
The technology_gap_ratio() function returns TGR values grouped by
technology:
tgr_by_group <- technology_gap_ratio(fit_det) lapply(tgr_by_group, summary)
For a formatted summary table:
tgr_summary(fit_det)
# Metafrontier coefficients coef(fit_det, which = "meta") # Group-specific coefficients coef(fit_det, which = "group")
# Log-likelihood (sum of group log-likelihoods for deterministic) logLik(fit_det) # Number of observations nobs(fit_det) # AIC and BIC (available automatically via logLik method) AIC(fit_det)
The package provides four built-in plot types:
plot(fit_det, which = "tgr")
plot(fit_det, which = "efficiency")
Points below the 45-degree line indicate a technology gap (TE* < TE). The vertical distance from the line reflects the TGR.
plot(fit_det, which = "decomposition")
Side-by-side boxplots of TE, TGR, and TE* by group.
The poolability test evaluates whether group-specific frontiers are statistically different from a single pooled frontier:
poolability_test(fit_det)
A significant result (small p-value) indicates that the group frontiers genuinely differ, that is, the groups face different restrictions of the common metatechnology, justifying the metafrontier approach.
Every estimation stage of a metafrontier fit can be inspected with
check_convergence(), which returns one row per stage (each group
frontier and the metafrontier itself) with the estimation method, the
optimiser's convergence code, and a logical convergence indicator:
check_convergence(fit_det)
The summary() method also prints a convergence block, so estimation
problems are flagged even if you never call check_convergence()
directly.
The package supports three distributional assumptions for the one-sided inefficiency term $u_i$ in SFA:
# Half-normal (default): u ~ |N(0, sigma_u^2)| fit_hn <- metafrontier(log_y ~ log_x1 + log_x2, data = sim$data, group = "group", dist = "hnormal") # Truncated normal: u ~ N+(mu, sigma_u^2) fit_tn <- metafrontier(log_y ~ log_x1 + log_x2, data = sim$data, group = "group", dist = "tnormal") # Exponential: u ~ Exp(1/sigma_u) fit_exp <- metafrontier(log_y ~ log_x1 + log_x2, data = sim$data, group = "group", dist = "exponential")
Since we used simulated data, we can compare estimated values against the truth:
# True vs estimated metafrontier coefficients cbind( True = sim$params$beta_meta, Estimated = coef(fit_det, which = "meta") ) # True vs estimated mean TGR by group true_tgr <- tapply(sim$data$true_tgr, sim$data$group, mean) est_tgr <- tapply(fit_det$tgr, fit_det$group_vec, mean) cbind(True = true_tgr, Estimated = est_tgr) # Correlation between true and estimated efficiency cor(sim$data$true_te, fit_det$te_group) cor(sim$data$true_te_star, fit_det$te_meta)
The package supports panel data via the Battese-Coelli (1992) and (1995) models. Use the panel argument:
# Simulate panel data panel_sim <- simulate_panel_metafrontier( n_groups = 2, n_firms_per_group = 20, n_periods = 5, seed = 42 ) # BC92: time-varying inefficiency u_it = u_i * exp(-eta*(t-T)) fit_panel <- metafrontier( log_y ~ log_x1 + log_x2, data = panel_sim$data, group = "group", panel = list(id = "firm", time = "year"), panel_dist = "bc92" ) summary(fit_panel) # The eta parameter captures time-varying inefficiency # eta > 0: inefficiency decreasing over time # eta < 0: inefficiency increasing over time
The boot_tgr() function provides parametric and nonparametric
bootstrap confidence intervals for the technology gap ratio:
sim <- simulate_metafrontier(n_groups = 2, n_per_group = 100, seed = 42) fit <- metafrontier(log_y ~ log_x1 + log_x2, data = sim$data, group = "group", meta_type = "stochastic") # Nonparametric bootstrap (case resampling within groups) boot <- boot_tgr(fit, R = 499, type = "nonparametric", seed = 1) print(boot) # Observation-level CIs ci <- confint(boot) head(ci) # Group-level mean TGR CIs boot$ci_group # Parametric bootstrap (resample from estimated error distributions) boot_par <- boot_tgr(fit, R = 499, type = "parametric", seed = 1)
The stochastic metafrontier is a two-stage estimator where Stage 2 uses fitted values from Stage 1 as regressors. This "generated regressor" problem means naive standard errors understate uncertainty. The Murphy-Topel (1985) correction adjusts for this:
fit <- metafrontier(log_y ~ log_x1 + log_x2, data = sim$data, group = "group", meta_type = "stochastic") # Naive (uncorrected) standard errors vcov(fit) # Murphy-Topel corrected standard errors vcov(fit, correction = "murphy-topel") # Corrected confidence intervals confint(fit, correction = "murphy-topel")
When group membership is unobserved, use latent_class_metafrontier():
sim <- simulate_metafrontier(n_groups = 2, n_per_group = 100, seed = 42) # Fit with 2 latent classes lc <- latent_class_metafrontier( log_y ~ log_x1 + log_x2, data = sim$data, n_classes = 2, n_starts = 5, seed = 123 ) print(lc) summary(lc) # Select optimal number of classes via BIC bic_table <- select_n_classes( log_y ~ log_x1 + log_x2, data = sim$data, n_classes_range = 2:4, n_starts = 3, seed = 42 ) print(bic_table) # choose n_classes with lowest BIC
For additive efficiency decomposition, use DDF-based metafrontier:
sim <- simulate_metafrontier(n_groups = 2, n_per_group = 50, seed = 42) # Use raw (non-log) data for DEA sim$data$y <- exp(sim$data$log_y) sim$data$x1 <- exp(sim$data$log_x1) sim$data$x2 <- exp(sim$data$log_x2) fit_ddf <- metafrontier( y ~ x1 + x2, data = sim$data, group = "group", method = "dea", type = "directional", direction = "output" ) summary(fit_ddf) # Additive decomposition: beta_meta = beta_group + ddf_tgr head(data.frame( beta_meta = fit_ddf$beta_meta, beta_group = fit_ddf$beta_group, ddf_tgr = fit_ddf$ddf_tgr ))
Battese, G.E. and Coelli, T.J. (1988). Prediction of firm-level technical efficiencies with a generalized frontier production function and panel data. Journal of Econometrics, 38(3), 387--399.
Battese, G.E., Rao, D.S.P. and O'Donnell, C.J. (2004). A metafrontier production function for estimation of technical efficiencies and technology gaps for firms operating under different technologies. Journal of Productivity Analysis, 21(1), 91--103.
Huang, C.J., Huang, T.-H. and Liu, N.-H. (2014). A new approach to estimating the metafrontier production function based on a stochastic frontier framework. Journal of Productivity Analysis, 42(3), 241--254.
Jondrow, J., Lovell, C.A.K., Materov, I.S. and Schmidt, P. (1982). On the estimation of technical inefficiency in the stochastic frontier production function model. Journal of Econometrics, 19(2--3), 233--238.
O'Donnell, C.J., Rao, D.S.P. and Battese, G.E. (2008). Metafrontier frameworks for the study of firm-level efficiencies and technology ratios. Empirical Economics, 34(2), 231--255.
Any scripts or data that you put into this service are public.
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.