The Markup Equation and HC3: Standard Errors for Small Comparables Sets

August 25, 2026 by Ednaldo Silva

A transfer pricing comparables set is small. Thirty independent companies is a generous count; ten to fifteen comparables are common. Each contributes three fiscal years, so the pooled regression runs on fewer than ninety annual observations and often on fewer than fifty. Everything that follows is shaped by that fact. Asymptotic arguments that hold comfortably at ten thousand observations do not hold here, and the estimator that is correct in principle at large sample sizes is not the estimator reported in transfer pricing.

This note sets out the specification we estimate, the exact algebra that recovers the arm’s-length markup from it, and the standard error we report. The conclusion is deliberately unadventurous: ordinary least squares (OLS) on the markup equation, with HC3 heteroskedasticity-consistent standard errors. Both halves of that sentence appear in standard graduate econometrics texts. Neither is exotic, and neither requires the reader to accept anything on our authority.

1. An accounting identity that changes the question

Operating income before depreciation and total operating expense are built from the same primitives:

\begin{aligned}
&(1)\quad \mathrm{OIBDP}=\mathrm{SALE}-\mathrm{COGS}-\mathrm{XSGA}\\
& \mathrm{XOPR}=\mathrm{COGS}+\mathrm{XSGA}\\
& \mathrm{OIBDP}=\mathrm{REVT}-\mathrm{XOPR}
\end{aligned}

Identity (1) is a definition, not an empirical finding. Its consequence is usually overlooked. Consider the operating-profit equation that transfer pricing practice would naturally write, and substitute the identity into it:

\begin{aligned}
&(2)\quad \mathrm{OIBDP}(i)=\alpha+\beta\,\mathrm{REVT}(i)+u(i)\\
&\quad \mathrm{REVT}(i)-\mathrm{XOPR}(i)
=\alpha+\beta\,\mathrm{REVT}(i)+u(i)\\
&\quad \mathrm{XOPR}(i)
=-\alpha+(1-\beta)\mathrm{REVT}(i)-u(i)
\end{aligned}

The profit equation and the cost equation are the same equation with the signs rearranged. In sample the correspondence is exact, not asymptotic:

\begin{aligned}
&(3)\quad
\hat{\beta}
=\frac{\operatorname{Cov}(\mathrm{OIBDP},\mathrm{REVT})}
{\operatorname{Var}(\mathrm{REVT})}\\
&\quad
=\frac{\operatorname{Cov}(\mathrm{REVT}-\mathrm{XOPR},\mathrm{REVT})}
{\operatorname{Var}(\mathrm{REVT})}\\
&\quad
=1-\frac{\operatorname{Cov}(\mathrm{XOPR},\mathrm{REVT})}
{\operatorname{Var}(\mathrm{REVT})}
=1-\hat{\gamma}_1
\end{aligned}

So regressing operating profit on revenue tells us nothing that regressing operating cost on revenue has not already told us. There is one scatter of points, and there are three ways to lay a line through it. Two of the three are the same line.

2. The third normalization

NormalizationEquationComment
Profit equationOIBDP = \alpha + \beta REVT + uThe regressor appears inside the dependent variable through identity (1).
Cost equationXOPR = \gamma_0 + \gamma_1 REVT + vNumerically equivalent to the profit equation by (2) and (3).
Markup equationREVT = \lambda_0 + \lambda_1 XOPR + wThe reverse regression. Genuinely distinct, and the one we estimate.

Least squares is not indifferent to which variable is placed on the left. The markup equation is therefore a different estimator, not a cosmetic rearrangement, and the relation between the two slopes is exact in every sample:

\quad
(4)\qquad
\hat{\lambda}_1\hat{\gamma}_1=R^2

We estimate the markup equation for a reason that predates the statistics. Written as REVT = \lambda_0 + \lambda_1 XOPR, it is Kalecki’s degree-of-monopoly markup expressed as an estimating equation. Under administered pricing the causal ordering runs from cost to price within the period: purchasing and production are committed before the price is set. Placing cost on the right-hand side puts the predetermined variable where it belongs. The profit equation reverses that ordering, which is why it is the specification that would require an instrument.

3. Recovering the arm’s-length margin

Nothing is lost by estimating the markup equation, because the structural parameters follow from it by an exact mapping. Solve for XOPR and match coefficients against (2):

\begin{aligned}
&(5)\quad
\mathrm{REVT}=\lambda_0+\lambda_1\mathrm{XOPR}\\
&\quad
\mathrm{XOPR}
=-\frac{\lambda_0}{\lambda_1}
+\frac{1}{\lambda_1}\mathrm{REVT}\\
&\quad
\text{matching against }
\mathrm{XOPR}
=-\alpha+(1-\beta)\mathrm{REVT}\\
&\quad
\alpha=\frac{\lambda_0}{\lambda_1},
\qquad
\beta=1-\frac{1}{\lambda_1}
=\frac{\lambda_1-1}{\lambda_1}
\end{aligned}

The economic reading is immediate. The markup on operating cost is \lambda_1 minus one, and the operating margin is that markup divided by one plus itself:

\begin{aligned}
&(6)\quad
\text{markup }m=\lambda_1-1\\
&\quad
\text{margin }\beta=\frac{m}{1+m}
\end{aligned}

Because alpha and beta are nonlinear in the estimated coefficients, the covariance matrix must be carried through the mapping by the delta method. For the margin the derivative is a single term; for the intercept the gradient has two, and the off-diagonal element of the covariance matrix enters:

\begin{aligned}
&(7)\quad
\frac{d\beta}{d\lambda_1}
=\frac{1}{\lambda_1^2}\\
&\quad
SE(\hat{\beta})
=\frac{SE(\hat{\lambda}_1)}{\hat{\lambda}_1^2}\\
&\quad
\text{gradient for }\alpha:
\qquad
g=
\begin{bmatrix}
1/\lambda_1\\
-\lambda_0/\lambda_1^2
\end{bmatrix}\\
&\quad
\operatorname{Var}(\hat{\alpha})=g'Vg
\end{aligned}

The final expression is the reason to retain the whole matrix V rather than the two standard errors on its diagonal. The cross term carries a minus sign against a covariance that is negative whenever the mean of XOPR is positive, so it increases the variance of the intercept. An implementation that stores only standard errors will understate the uncertainty in alpha, and alpha is the quantity on which the case against ratio-based profit-level indicators turns.

4. A short note on instruments

Readers who recognize the profit equation as structural will ask why we do not instrument it. The short answer is that no defensible instrument exists in this data environment, and a weak or arbitrary instrument is worse than none. The longer answer is that we do not need one.

If both revenue and operating cost are measured with error, the direct and reverse least-squares slopes bracket the true slope. The two endpoints turn out to be exactly the two regressions an analyst might have run, and the width of the bracket is an arithmetic consequence of the fit:

\begin{aligned}
&(8)\quad
\text{lower: }\quad
\beta_C=1-\frac{1}{\hat{\lambda}_1}
\qquad
\text{(markup equation)}\\
&\quad
\text{upper: }\quad
\beta_A=1-\hat{\gamma}_1
\qquad
\text{(profit equation)}\\
&\quad
\text{width}
=\beta_A-\beta_C
=\frac{1-R^2}{\hat{\lambda}_1}
\end{aligned}

Revenue and operating cost are close to collinear in levels for operating businesses, so R-squared in this regression routinely exceeds 0.99 and the bracket collapses to a few basis points. The identification question that cannot be answered at all by a quartile calculation is here answered by an identity and a number the analyst can print.

5. The sandwich, in matrix form

Write the pooled regression with n comparables observed over T = 3 fiscal years, so that there are nT annual observations and k = 2 parameters. Let y be the nT-vector of REVT, and let X be the nT by k design matrix whose columns are a constant and XOPR. Then

\begin{aligned}
&(9)\quad
\hat{\beta}
=(X'X)^{-1}X'y\\
&\quad
\operatorname{Var}(\hat{\beta})
=(X'X)^{-1}M(X'X)^{-1}
\end{aligned}

The outer terms are fixed by the design matrix alone. Everything that distinguishes one covariance estimator from another lives in the middle matrix M. Under the classical assumption of a common error variance,

\begin{aligned}
&(10)\quad
M=\sigma^2(X'X),
\qquad
\hat{\sigma}^2=\frac{e'e}{nT-k}\\
&\quad
\text{which returns the textbook formula}\\
&\quad
\operatorname{Var}(\hat{\beta})
=\sigma^2(X'X)^{-1}
\end{aligned}

That assumption fails here for a structural reason, not a subtle one. A company with four billion dollars of operating cost misses the fitted line by tens of millions; a company with two hundred million cannot. The scatter widens with size because the variables are denominated in dollars. Heteroskedasticity in a levels regression on financial statement data is a property of the measurement, not a pathology to be tested for and hoped away.

The heteroskedasticity-consistent family replaces the single variance by observation-specific residual information. Let e be the residual vector, let

\begin{aligned}
&(11)\quad
H=X(X'X)^{-1}X'
\qquad
\text{(the hat matrix)}\\
&\quad
h(i)=H(i,i)
\qquad
\text{(the leverage of observation }i\text{)}\\
&\quad
x(i)=\text{the }i\text{-th row of }X\text{ taken as a column}
\end{aligned}

Then the four standard members of the family differ only in how each squared residual is weighted:

EstimatorMiddle matrix MScaleProperty
HC0sum(i) e(i)^2 x(i) x(i)'1Biased downward; worst at high leverage.
HC1sum(i) e(i)^2 x(i) x(i)'nT/(nT-k)A global degrees-of-freedom patch.
HC2sum(i) [e(i)^2/(1-h(i))] x(i) x(i)'1Exactly unbiased under homoskedasticity.
HC3sum(i) [e(i)^2/(1-h(i))^2] x(i) x(i)'1Approximates the jackknife; conservative.

In every case the reported standard errors are the square roots of the diagonal of (9) with the chosen M substituted, and the reference distribution is t with nT minus k degrees of freedom.

6. Why HC3 when the sample is small

The case for the leverage correction is one line of algebra. Under homoskedasticity the expected squared residual is not the error variance:

\quad
(12)\qquad
E[e(i)^2]=(1-h(i))\sigma^2

So \text{HC0}, which substitutes e(i)^2 directly, is biased downward by the factor (1 - h(i)). It reports standard errors that are too small, and it does so most severely at the observations with the highest leverage. \text{HC2}divides by (1 - h(i)) and removes the bias exactly. \text{HC3} divides by the square, which over-corrects on purpose.

Whether over-correction is a virtue depends entirely on the sample size, which is why the recommendation for a comparables set differs from the recommendation for a large panel. Three facts govern the choice at nT below ninety:

  1. Average leverage is k divided by nT. At n = 15 comparables and three years, average leverage is about 0.044, but a single large comparable can carry leverage several times that. The correction is not a rounding adjustment at this sample size; it is the difference between reporting a defensible interval and reporting a narrow one.
  2. Simulation evidence is unambiguous below a few hundred observations: \text{HC3} has the best coverage of the four, and \text{HC0} is unacceptably liberal. This is the finding of MacKinnon and White and the explicit practitioner recommendation of Long and Ervin, who advise \text{HC3} for samples under roughly 250.
  3. \text{HC3} is the widest of the four in almost every real sample, because 1/(1-h)^2 exceeds 1/(1-h) exceeds 1. An analyst who reports \text{HC3} has reported the most conservative member of the standard family. That is a position from which the range cannot be attacked as having been narrowed by the choice of estimator.

The third point deserves emphasis because it is the one that matters in an examination. The reason to prefer \text{HC3} is not that it is the most sophisticated available estimator. Better small-sample corrections exist, and they are published. The reason to prefer \text{HC3} is that it is simultaneously the conservative choice and the one written down in the standard textbooks, so the two arguments an examiner might raise both resolve in the analyst’s favor at once. Estimator selection here is not a demonstration of technique. It is a decision about what can be defended without a seminar.

7. What HC3 does not do

Two limitations belong on the face of any report that uses it, because disclosing them costs nothing and discovering them costs a great deal.

  • \text{HC3} treats the three fiscal years of a company as three independent observations. They are not: a comparable that ran a thin margin in one year tends to run one in the next. Corrections for this exist and we compute them in the workpapers, but at these sample sizes the leverage inflation in \text{HC3} typically produces a wider interval than the alternatives do, so reporting \text{HC3} is the conservative course as well as the standard one. Where it is not, we say so.
  • No covariance estimator addresses bias in the coefficient itself. \text{HC3} corrects the reported uncertainty around \hat{\beta}; it has nothing to say about whether \hat{\beta} is centered on the right number. That is the specification question of sections 1 through 4, and it is the more important one.

We therefore publish the full family, \text{HC0} through \text{HC3}, as a sensitivity appendix. The claim we make is not that \text{HC3} is uniquely correct. It is that the conclusion does not depend on which member is used,

and where that fails, the failure is disclosed together with the leverage diagnostic that almost always explains it.

8. A worked example

Fifteen comparables, three fiscal years each, forty-five pooled annual observations, no year dummy variables. Ordinary least squares on the markup equation with \text{HC3} standard errors.

QuantityEstimateHC3 SENote
\lambda_0 ($ millions)6.88953.7295intercept of the markup equation
\lambda_11.1294170.008736markup factor
R-squared0.999235identical in both directions
markup m12.9417 %0.8736 %equation (6)
margin beta0.1145880.006849delta method, equation (7)
alpha6.10003.3466delta method, equation (7)
maximum leverage0.1543against a mean of 0.0444

The reported arm’s-length margin is 11.46 percent plus or minus 0.68 percentage points, an interval of [10.77 %, 12.14 %]. The four estimators are ordered exactly as the algebra requires, and the spread between the least and the most conservative is under fourteen percent of the standard error:

\begin{array}{l}
\text{(13)}\quad SE(\hat{\lambda}_1):\\
\text{HC0}=0.007651\\
\text{HC1}=0.007827\\
\text{HC2}=0.008171\\
\text{HC3}=0.008736
\end{array}

The intercept is estimated at 6.10 with a standard error of 3.35, so the one-standard-error interval is [2.75, 9.45] and excludes zero. That is the diagnostic that governs everything else. A non-zero intercept means the ratio of profit to revenue is not the arm’s-length rate but the rate plus a size-dependent contaminant, and the interquartile range of such ratios measures the intercept and the size dispersion of the comparables rather than any variation in profitability. Where the interval for alpha contains zero, ratio indicators are approximately unbiased and we say so.

9. What we report

  1. The specification, fixed before the comparables set is closed: REVT on XOPR, pooled annual observations, no year dummy variables.
  2. \hat{\lambda}_0 and \hat{\lambda}_1 with \text{HC3} standard errors, the full two-by-two covariance matrix in the appendix, and R-squared.
  3. The derived markup and margin with delta-method standard errors from (7), reported as the estimate plus or minus one standard error.
  4. The bracket (8) and its width, which pre-empts the endogeneity objection without requiring an instrument.
  5. The sensitivity appendix over \text{HC0} through \text{HC3}, and the maximum leverage.
  6. The interval for alpha, and where it excludes zero, the implied bias in any ratio-based profit-level indicator computed on the same comparables.

References

Davidson, R., and MacKinnon, J. G. (2004). Econometric Theory and Methods. New York: Oxford University Press. A standard graduate econometrics textbook. HC0 through HC3 appear in the chapter on heteroskedasticity as numbered members of one family, with the leverage corrections written out. This is the reference that answers the charge that anything above has been invented for the occasion: the estimator we report is a textbook option, available as a single keyword in R and Python.

White, H. (1980). “A Heteroskedasticity-Consistent Covariance Matrix Estimator and a Direct Test for Heteroskedasticity.” Econometrica 48, 817-838. The foundational paper, and among the most cited in economics. It settles the question of whether heteroskedasticity-consistent standard errors are accepted practice.

MacKinnon, J. G., and White, H. (1985). “Some Heteroskedasticity-Consistent Covariance Matrix Estimators with Improved Finite Sample Properties.” Journal of Econometrics 29, 305-325. Introduces HC2 and HC3 and demonstrates by simulation that HC0 is unacceptably liberal in small samples. The authority for the leverage correction and for not reporting HC0.

Long, J. S., and Ervin, L. H. (2000). “Using Heteroscedasticity Consistent Standard Errors in the Linear Regression Model.” The American Statistician 54, 217-224. The practitioner recommendation: use HC3 unless the sample exceeds roughly 250 observations. Written for applied users and published in a journal a non-specialist reader can follow, which makes it the most useful single item to attach to a report.

Kalecki, M. (1954). Theory of Economic Dynamics, Chapter 1. London: Allen and Unwin. The degree-of-monopoly markup. This is the theoretical warrant for placing operating cost on the right-hand side, and it is what makes the markup equation a behavioral relation rather than a normalization preference.

Frisch, R. (1934). Statistical Confluence Analysis by Means of Complete Regression Systems. Oslo: Universitetets Økonomiske Institutt (Publication No. 5, University Institute of Economics). The source of the bracket in section 4. When both variables carry measurement error, the direct and reverse regressions bound the true slope. The result predates the instrumental-variables literature and requires no instrument to apply. The monograph is scarce; its opening pages (5–8), on the confluence problem rather than on the bounds, are reprinted as ch. 23, pp. 271–273 of D. F. Hendry and M. S. Morgan (eds.), The Foundations of Econometric Analysis, Cambridge University Press, 1995. For a modern statement and the multivariate generalization, see Klepper, S., and Leamer, E. E. (1984), “Consistent Sets of Estimates for Regressions with Errors in All Variables,” Econometrica 52, 163–183; the result is now called the Gini–Frisch bound.