1. The question you will be asked
A client, an examiner, or a reviewer can ask why we run a regression when every other firm reports the median and the interquartile range of a profit level indicator (PLI). The honest answer is not that regression is more sophisticated. It is that the ratio statistics answer a different question from the one the regulation poses (a reliable measure of an arm’s length amount), and they answer it with an error whose size we can compute in advance.
This memorandum sets out the algebra in full. It assumes no more than first-year statistics. Anyone who works through sections 3 to 6 will be able to defend the regression approach without appeal to authority, which is the object: a position held on faith cannot survive cross-examination, and a position held on algebra does not need to.
The claim is narrow. Ratios are legitimate quantities and we compute them. What fails is the univariate summary statistics of a collection of ratios read as an estimate of an economic parameter (profit indicator). The condition under which the summary statistics is sound is stated, and it is testable.
2. Notation
Latin majuscules denote observable Compustat variables: Y for operating profit (OIBDP, OIADP), X for the independent variable or base, such as revenue (REVT), total cost (XOPR), or assets stock (PPENT). Greek minuscules denote parameters. Lowercase n counts entities; N is sample size, with N=nT for n comparables over T years. The index i runs over observations, with X(i) > 0 throughout.
The profit level indicator is the ratio m(i)=Y(i)/X(i). The harmonic mean of the denominators is written X_{H}, and the harmonic quantile X_{H,p} is defined in section 5. The arithmetic mean of a variable is written with a bar, so that \bar{m} is the mean of the ratios.
We report an estimate with its standard uncertainty as b \pm \operatorname{SE}(b), which is the k=1 convention of the GUM and corresponds to a 68 percent interval. We do not report p-values.
3. The ratio identity
Take the levels relation between profit and revenue across comparable companies:
\quad (1)\qquad Y(i)=\alpha+\beta X(i)
This is best as a reduced form regression. We estimate reduced forms rather than structural systems to avoid carrying the misspecification of an equation we did not need, and because two-stage least squares ((TSLS) is not a method that survives contact with a general audience. The parameter \beta is the return on an incremental dollar of the base, such as revenue (REVT), total cost (XOPR), or operating assets (PPENT). The parameter \alpha is whatever profit the relation carries at zero base: fixed cost if negative, scale-independent return if positive.
Divide (1) through by X(i):
\quad
(2)\qquad
m(i)=\frac{Y(i)}{X(i)}=\beta+\frac{\alpha}{X(i)}Equation (2) is the whole memorandum in one line. The profit ratio is not a parameter. It is an affine function of the reciprocal of company size. Affine refers to a relationship or transformation that combines a linear operation with a constant shift, such as an intercept. In mathematics, a linear function must pass through the origin (f(X) = \beta X, where \alpha = 0). When an intercept (alpha) is added, the relationship becomes an affine function. It equals \beta for every company if and only if \alpha=0, in which case the relation passes through the origin and profit is proportional to the base (such as revenue).
Write u(i)=1/X(i), so that m(i)=\beta+\alpha u(i). Every result below follows from the fact that m is affine in u, and from nothing else.
4. What the mean of ratios estimates
Average equation (2) over the sample:
\quad
(3)\qquad
\bar{m}=\frac{1}{N}\sum_{i=1}^{N}m(i)=\beta+\alpha\cdot\frac{1}{N}\sum_{i=1}^{N}\frac{1}{X(i)}The second term contains the arithmetic mean of the reciprocals, which is by definition the reciprocal of the harmonic mean:
\quad
(4)\qquad
X_{H}=\frac{N}{\sum_{i=1}^{N}\frac{1}{X(i)}}
\qquad\text{so that}\qquad
\frac{1}{N}\sum_{i=1}^{N}\frac{1}{X(i)}=\frac{1}{X_{H}}Substituting into (3):
\quad
(5)\qquad
\bar{m}=\beta+\frac{\alpha}{X_{H}}
\qquad\text{and}\qquad
\operatorname{Bias}(\bar{m};\beta)=\frac{\alpha}{X_{H}}Three properties of this result deserve emphasis with a client.
- It is exact, not asymptotic. No approximation has been made anywhere between (1) and (5).
- It does not vanish with sample size. Adding comparables changes X_{H} and leaves the structure untouched. This is not sampling error and no amount of data cures it.
- It is signed. For \alpha > 0 the mean of ratios overstates \beta; for \alpha < 0, which is the case when the relation carries fixed costs, it understates \beta.
By the inequality of arithmetic and harmonic means, X_{H}\le\bar{X}, so the bias is at least \alpha/\bar{X} in magnitude. The intuitive correction of dividing the intercept by average revenue understates the damage.
A second observation matters for the choice of statistic. The ratio of aggregates, \sum Y/\sum X, is not the same quantity as the mean of ratios. Writing w(i)=X(i)/\sum X, the ratio of aggregates equals \sum w(i)m(i): a weighted mean of the ratios with weights proportional to size. The mean of ratios weights each company alike and thereby gives the smallest comparables the largest influence on the answer.
5. What the quartiles estimate
Differentiate (2) with respect to X:
\quad
(6)\qquad
\frac{dm}{dX}=-\frac{\alpha}{X^{2}}The derivative has the sign of -\alpha and never changes sign. The ratio is therefore a strictly monotone function of company size whenever \alpha\ne 0: decreasing in size for \alpha > 0, increasing for \alpha < 0. This single fact governs every order statistic.
Monotone transformations carry quantiles to quantiles. Sort the denominators X(1)\le X(2)\le\dots\le X(N). For \alpha > 0 the ranking reverses:
\quad
(7)\qquad
m(k)=\beta+\frac{\alpha}{X(N+1-k)}The k-th smallest ratio belongs to the k-th largest company. In quantile function form, writing Q_{p} for the p-th quantile:
\quad
(8)\qquad
Q_{p}(m)=\beta+\frac{\alpha}{Q_{1-p}(X)}
\qquad\text{for }\alpha > 0Define the harmonic p-quantile X_{H,p} as the reciprocal of the p-th quantile of the reciprocals, which for \alpha > 0 is Q_{1-p}(X). Then every quantile of the ratio carries a bias of the same algebraic form as the mean:
\quad
(9)\qquad
Q_{p}(m)=\beta+\frac{\alpha}{X_{H,p}}Applied to the three quartiles of X, written Q_{1}, Q_{2}, Q_{3}:
\begin{aligned}
&(10)\quad
Q_{0.25}(m)=\beta+\frac{\alpha}{Q_{3}}
\\[8pt]
&\phantom{(10)\quad}
Q_{0.50}(m)=\beta+\frac{\alpha}{Q_{2}}
\\[8pt]
&\phantom{(10)\quad}
Q_{0.75}(m)=\beta+\frac{\alpha}{Q_{1}}
\end{aligned}None of the three estimates \beta, and the displacement is larger at the top of the range than at the bottom. Averaging the three returns the familiar form, with X_{Q} the harmonic mean of the quartiles of X:
\begin{aligned}
&(11)\quad
\frac{1}{3}\sum_{j=1}^{3}Q_{j/4}(m)=\beta+\frac{\alpha}{X_{Q}},
\\[8pt]
&\phantom{(11)\quad}
X_{Q}=\frac{3}{\frac{1}{Q_{1}}+\frac{1}{Q_{2}}+\frac{1}{Q_{3}}}
\end{aligned}6. What the dispersion statistics measure
Because m(i)=\beta+\alpha u(i) is affine in u, the slope passes into every measure of spread and the intercept of the affine map drops out:
\quad
(12)\qquad
\operatorname{Var}(m)=\alpha^{2}\operatorname{Var}(u)
\qquad
\operatorname{sd}(m)=|\alpha|\operatorname{sd}(u)
\qquad
\mathrm{IQR}(m)=|\alpha|\,\mathrm{IQR}(u)The parameter \beta has cancelled identically in all three. A dispersion statistic computed on ratio profit indicators does not measure variation in profitability. Under (2) it measures variation in the reciprocal of company size, scaled by the intercept. In particular:
\quad
(13)\qquad
\mathrm{IQR}(m)=|\alpha|\cdot\frac{Q_{3}-Q_{1}}{Q_{1}\cdot Q_{3}}If \alpha=0 the interquartile range of the ratios collapses to zero. The width of the arm’s length range, constructed this way, is a measurement of the intercept and the dispersion of revenue among the comparables. It is not a measurement of disagreement about profitability, which is what everyone reading the report believes it to be.
7. A worked example with no sampling error
Eight companies, constructed so that (1) holds exactly with \alpha=5.0 and \beta=0.06. There is no disturbance, no measurement error, and no sampling variation. Every number below is structural.
| Company | X (revenue) | Y (profit) | m = Y/X | 1/X |
|---|---|---|---|---|
| A | 10 | 5.60 | 0.5600 | 0.100000 |
| B | 20 | 6.20 | 0.3100 | 0.050000 |
| C | 30 | 6.80 | 0.2267 | 0.033333 |
| D | 50 | 8.00 | 0.1600 | 0.020000 |
| E | 80 | 9.80 | 0.1225 | 0.012500 |
| F | 120 | 12.20 | 0.1017 | 0.008333 |
| G | 200 | 17.00 | 0.0850 | 0.005000 |
| H | 400 | 29.00 | 0.0725 | 0.002500 |
The regression of Y on X with an intercept returns \hat{\alpha}=5.0000 and \hat{\beta}=0.0600 with zero standard errors, because the relation is exact. The denominators have harmonic mean X_{H}=34.53 and quartiles Q_{1}=27.5, Q_{2}=65, Q_{3}=140. Now compute the ratio statistics that a conventional report would present.
| Statistic | Value | Identity | Error vs β |
|---|---|---|---|
| Mean of ratios | 0.2048 | \beta+\alpha/X_{H} | +0.1448 |
| Median of ratios | 0.1413 | \beta+\alpha/Q_{2} | +0.0813 |
| Lower quartile | 0.0975 | \beta+\alpha/Q_{3} | +0.0375 |
| Upper quartile | 0.2475 | \beta+\alpha/Q_{1} | +0.1875 |
| Interquartile range | 0.1500 | |\alpha|(Q_{3}-Q_{1})/(Q_{1}\cdot Q_{3}) | — |
| Standard deviation | 0.1643 | |\alpha|\cdot\operatorname{sd}(1/X) | — |
Read the third and fourth rows together. The interquartile range of the profit level indicator is 9.75 percent to 24.75 percent. The true parameter is 6.00 percent. The arm’s length range constructed from these comparables excludes the arm’s length return, and it excludes it in data containing no noise whatever. A tested party earning exactly \beta on its revenue would be adjusted upward to the median.
The standard deviation is 16.43 percentage points around a parameter of 6.00 percent. None of that dispersion is economic. It is 5.0 multiplied by the spread of the reciprocals of revenue.
The interpolation qualification of section 5 is visible here. Identity (10) with Q_{3}(X)=140 gives \beta+5/140=0.0957, while the interpolated lower quartile of the ratios is 0.0975. The gap of 0.0018 is the convexity term, and it runs in the direction the inequality predicts.
8. What the standard remedies do not fix
Two transformations are recommended in the textbooks for heteroscedastic data of this kind, and both are sometimes offered as answers to the problem above. Neither is an answer. One is innocent and the other is not.
8.1 Deflation by the denominator
Add a disturbance to (1), so that Y(i)=\alpha+\beta X(i)+\varepsilon(i). Dividing by X(i):
\quad (14)\qquad m(i)=\beta+\alpha u(i)+\varepsilon(i)u(i)
If the disturbance is proportional to size, \operatorname{Var}(\varepsilon(i)\mid X(i))=\sigma^{2}X(i)^{2}, then the transformed disturbance \varepsilon(i)u(i) has constant variance \sigma^{2} and ordinary least squares applied to (14) is best linear unbiased. This is weighted least squares, and it is sound. Observe what it estimates: the intercept of (14) is \beta and its slope is \alpha. Deflation retains both parameters.
The error is not deflation. The error is deflation followed by discarding the u term, which is what taking the mean or the quartiles of m(i) amounts to. The ratio statistic is the deflated regression with \alpha constrained to zero, and the constraint is imposed rather than tested. Maddala recommends deflation as a variance-stabilizing device, and he is right; what he does not recommend, and what no one recommends, is dropping a regressor without testing its coefficient.
There is a further equivalence, and it is worth showing rather than citing because it takes two lines. Suppose instead that the levels disturbance is homoscedastic, \operatorname{Var}(\varepsilon(i))=\sigma^{2}. Then in (14) the transformed disturbance has variance \sigma^{2}u(i)^{2}, and efficient estimation weights each observation by X(i)^{2}. Write out that criterion and clear the denominators:
\quad
(15)\qquad
\sum_{i=1}^{N}X(i)^{2}\left[m(i)-\beta-\alpha u(i)\right]^{2}=\sum_{i=1}^{N}\left[Y(i)-\alpha-\beta X(i)\right]^{2}The two criteria are the same function of the same parameters, so weighted least squares on the ratio equation and ordinary least squares on the levels are one procedure with two descriptions. The levels regression is not an alternative to the ratio analysis. It is the ratio analysis performed with the intercept estimated instead of assumed. Weisberg (2014, §7.1) treats weighted least squares with known weights, which is the case here; the weights are X(i)^{2}.
8.2 Double logarithms
The second recommendation is to take logarithms of both variables. This is a change of model, not a change of scale, and the algebra is worth doing slowly because it is where intuition fails. Logarithms do not distribute over a sum: \ln(\alpha+\beta X) is not \ln\alpha+\beta\ln X. Factor \beta X out of the sum instead:
\quad
(16)\qquad
\alpha+\beta X=\beta X\cdot\left(1+\frac{\alpha}{\beta X}\right)\quad
(17)\qquad
\ln Y=\ln\beta+\ln X+\ln\left(1+\frac{\alpha}{\beta X}\right)The coefficient on \ln X in (17) is one, not \beta. Under the affine model the elasticity approaches unity as revenue grows, and \beta survives only inside the curvature term. Differentiating (17):
\quad
(18)\qquad
\frac{d\ln Y}{d\ln X}=\frac{\beta X}{\alpha+\beta X}=\frac{1}{1+\dfrac{\alpha}{\beta X}}The elasticity is a function of size. It equals one only when \alpha=0; it lies below one when \alpha > 0 and above one when \alpha < 0. A double-log regression fitted with a constant slope is therefore misspecified whenever the intercept is nonzero, and the coefficient it returns is a size-weighted average of the local elasticities in (18). The transformation does not remove \alpha. It distributes \alpha across the slope, where it cannot be seen and cannot be tested.
Logarithms carry a second cost. Because E[\ln Y]\ne\ln E[Y], the fitted curve does not predict profit in levels without a retransformation correction. We would be trading a bias we can compute for one we would have to estimate.
8.3 The power function, and how to tell the two apart
The double-log regression is the correct transformation of a different model, the power function:
\quad
(19)\qquad
Y=AX^{\beta}
\qquad\Longrightarrow\qquad
\ln Y=\ln A+\beta\ln XUnder (19) the ratio behaves differently: m(i)=AX(i)^{\beta-1}, constant in size only when \beta=1. The two specifications meet at Y=AX, the proportional case, which is the only relation under which the profit ratio is a parameter rather than a function of size.
There is a duality worth carrying in the head. Under the power function the marginal effect is dY/dX=\beta\cdot m(i): the slope is \beta times the profit ratio. Under the affine model the elasticity is (dY/dX)(X/Y)=\beta/m(i). In both directions the ratio is the carrier of the contamination, and in both directions the contaminant is \alpha/X.
The two models are distinguishable at the cost of one additional regression. Under a genuine power function the elasticity is constant across size classes. Under the affine model with \alpha\ne 0 the elasticity varies with size and approaches one, from above if \alpha < 0 and from below if \alpha > 0. Estimate the double-log slope within size terciles: a gradient toward unity as revenue grows indicates the affine model with a nonzero intercept, while a flat profile across terciles indicates a power law.
The Box-Cox transformation nests both cases, with \lambda=1 giving the affine model and \lambda=0 the power function, so the profile likelihood over \lambda settles the specification question by estimation rather than by assumption. This is already implemented in our regression module and should be run before the panel is drawn.
9. Reading the diagnostic panel
The EdgarStat platform draws the levels scatterplot and the ratio scatterplot. A third panel is perceptual rather than informational.
- Panel 1, Y against X. The open megaphone that appears in almost every Compustat sample is a statement about the disturbance variance, and therefore about efficiency and the choice of weights. It says nothing about \alpha. Do not let a critic answer it with “that is why we take ratios”: the deflation fixes the variance and leaves \alpha/X_{H} exactly where it was.
- Panel 2, m against 1/X. This is the estimation space, and it is unreadable for a reason worth stating in the caption. Revenue is close to lognormal, so the reciprocals pile up near zero with one or two small companies far out on the axis. Leverage in this regression goes as (u(i)-\bar{u})^{2}, and the disturbance is \varepsilon(i)u(i), so the points that determine the slope are also the noisiest. The panel is illegible because the statistic it displays is fragile.
- Panel 3, m against X. The hyperbola. The horizontal asymptote is \beta, so the parameter becomes a shape rather than a number. The vertical distance from the mean-of-ratios line to the asymptote is \alpha/X_{H}, which renders the bias as a length. If \alpha=0 the curve degenerates to a horizontal line, so the null hypothesis has a visual form the eye judges well.
The third panel displays the curve implied by the levels regression, \hat{m}(X)=\hat{\beta}+\hat{\alpha}/X, rather than a curve fitted to the ratios. Its pointwise variance follows from the delta method applied to a linear function of the estimates, writing a and b for \hat{\alpha} and \hat{\beta}:
\quad
(20)\qquad
\operatorname{Var}\left[m(X)\right]=\operatorname{Var}(b)+\frac{2}{X}\operatorname{Cov}(a,b)+\frac{1}{X^{2}}\operatorname{Var}(a)The band is tight at large revenue and flares as revenue falls. That flare, drawn to scale, is the case against ratio analysis for small comparables, and it is the most persuasive single object we can put in front of a client.
On the abscissa: plotting against \ln X is a rescaling of the axis and not a transformation of any variable, since nothing is estimated on \ln X and the ordinate is untouched. It spreads the small companies, where \alpha/X is large, over proportionate width. Where the panel is destined for an examiner, show linear X in the headline figure and place the logarithmic version as an inset labelled as a rescaling, so that no reader mistakes it for a log-log model.
Four reporting rules follow from the algebra rather than from taste.
10. The regulatory ground
The regulation does not prescribe the interquartile range of a ratio. It prescribes a reliability standard, and it names the interquartile range as one acceptable way to meet it. Section 1.482-1(e)(2)(iii)(B) requires that the reliability of the analysis be increased, where it is possible to do so, by adjusting the range through application of a valid statistical method. The following subparagraph states that reliability is increased when statistical methods establish limits such that there is a 75 percent probability of a result falling above the lower limit and a 75 percent probability of falling below the upper limit, and then provides that the interquartile range ordinarily provides an acceptable measure of this range, while a different statistical method may be applied if it provides a more reliable measure.
The word that carries the weight is “ordinarily.” The regulation sets a probability standard and invites a better method. Our position is that where \hat{\alpha} is distinguishable from zero, the interquartile range of ratio indicators does not meet the standard it is offered under, because by (10) its limits are determined by the size distribution of the comparables rather than by the dispersion of their returns. A prediction interval from the levels regression meets the standard directly and can be shown to do so.
This is not an idiosyncratic position, and directors should know the lineage before a reviewer produces it as an objection. Pearson (1897) identified spurious correlation induced by forming indices with a common component. Kronmal (1993) revisited the problem, observed that the warnings of Pearson, Neyman, and Tanner had been ignored, showed that ratio variables in regression produce misleading inference, and recommended that a ratio be used only inside a full model containing its component variables. That recommendation, arrived at in biostatistics, is the levels regression with an intercept. What is ours is the exact form of the bias in the location and dispersion statistics, \alpha/X_{H} and \alpha/X_{H,p}, and its application to section 482.
11. Objections you will meet
“Ratios are an ancient and well founded construction.”
They are, and nothing here disputes them. The ratio m(i) is a well defined number and we report it. What is at issue is the univariate summary of a collection of ratios read as an estimate of \beta. The two are different objects, and only the second is in question. It is worth adding that the classical criterion for sameness of ratio, invariance under scaling of both magnitudes, is itself failed by the profit indicator when \alpha\ne 0, since m(\lambda X)=\beta+\alpha/(\lambda X).
“The mean of ratios is unbiased.”
It is unbiased for E[m], which by (5) equals \beta+\alpha/X_{H}. It is biased for \beta. Both statements are true and they differ only in the target named. The regulatory question is a question about \beta.
“Everyone uses the interquartile range.”
Agreed, and the uniformity of a practice is not evidence of its validity. The interquartile range is the most elementary form of univariate summary: it requires no operation beyond sorting and slicing. The regulation permits a more reliable method and we are obliged to use one where it exists.
“Your regression assumes linearity.”
It assumes a functional form and then tests it, which is the distinction. The Box-Cox profile likelihood of section 8.3 chooses between the affine and power specifications from the data. The ratio summary also assumes a functional form, Y=\beta X through the origin, and never tests it.
“Small samples make regression unreliable.”
Small samples widen \operatorname{SE}(\hat{\beta}), which is visible and reportable. They do not widen the bias \alpha/X_{H}, which is invisible and unreported. A wide interval honestly stated is a better position before an examiner than a narrow interval that is wrong by a fixed amount.
References
Fisher, I. (1922). The Making of Index Numbers. Boston: Houghton Mifflin. The time reversal criterion; the index number literature reached the mean-of-ratios problem independently.
Goldberger, A. S. (1964). Econometric Theory. New York: Wiley. Reduced form versus structural estimation; the basis of the preference stated in section 3.
Kmenta, J. (1986). Elements of Econometrics, 2nd ed. New York: Macmillan. Reduced form estimation and the treatment of heteroscedastic disturbances.
Kronmal, R. A. (1993). Spurious correlation and the fallacy of the ratio standard revisited. Journal of the Royal Statistical Society A, 156(3), 379-392. Ratios in regression produce misleading inference; use a full model containing the component variables.
Maddala, G. S. (1977). Econometrics, New York: McGraw-Hill. Deflation and double logarithms as variance-stabilizing devices; section 8 above.
Pearson, K. (1897). Mathematical contributions to the theory of evolution: on a form of spurious correlation which may arise when indices are used in the measurement of organs. Proceedings of the Royal Society of London, 60, 489-498. The founding statement on correlation induced by a common denominator.
Silva, E. (2026). The harmonic mean bias of ratio profit indicators. EdgarStat blog. The propositions of sections 4 to 6 and the eight-company illustration.
U.S. Treasury. 26 CFR § 1.482-1(e)(2)(iii). Arm’s length range and interquartile range; the reliability standard discussed in section 11.
Weisberg, S. (2014). Applied Linear Regression, 4th ed. Hoboken: Wiley. The standing reference for this memorandum: weighted least squares at §7.1, misspecified variances and the sandwich estimator at §7.2, variance stabilizing transformations at §7.5, the delta method used at equation (20) at §7.6, nonconstant variance at §9.3, and the Box-Cox method at §A.12.