quants.wiki
Estimators, their annualisation rules, and where they break

Performance statistics

Sharpe, Sortino, Calmar, Sterling, Omega, information ratio, Treynor, Jensen's alpha, M-squared and capture, each with its exact estimator, annualisation rule and standard error.

Every number in this section is computed from the 24-month return series in the first table below, with a constant risk-free rate of 0.20 percent per month. Where a statistic has more than one estimator in general use - and most of them do - each variant is given separately with its own worked value, because the variants do not agree and the disagreement is usually larger than the difference between two managers. The annualisation rule matters as much as the estimator: multiplying a monthly Sharpe ratio by the square root of twelve is correct only for serially uncorrelated returns, and the correction for the general case is given explicitly.

The base return series

Twenty-four monthly returns in percent for a strategy and a benchmark, the strategy equity index starting from 1, its running peak, and the drawdown from that peak. Every statistic in this corpus is computed from these columns. The risk-free rate is 0.20 percent per month throughout. Sum of strategy returns 18.40 percent, sum of benchmark returns 14.40 percent.

MonthStrategy return, percentBenchmark return, percentStrategy equityRunning peakDrawdown, percent
1+1.60+4.001.0160001.0160000.0000
2+4.10+3.901.0576561.0576560.0000
3+2.30+0.701.0819821.0819820.0000
4+3.40+3.001.1187691.1187690.0000
5+1.50+0.801.1355511.1355510.0000
6-0.80-0.401.1264671.1355510.8000
7+3.20+2.301.1625141.1625140.0000
8+2.70+3.201.1939011.1939010.0000
9+1.90+1.401.2165861.2165860.0000
10-3.60-3.901.1727881.2165863.6000
11-0.50+1.101.1669251.2165864.0820
12+0.60-1.301.1739261.2165863.5065
13+2.10+2.801.1985791.2165861.4801
14+1.10+3.101.2117631.2165860.3964
15-1.20+0.301.1972221.2165861.5917
16+0.90+2.301.2079971.2165860.7060
17+0.40-1.401.2128291.2165860.3088
18-2.50-2.901.1825081.2165862.8011
19-1.80-3.001.1612231.2165864.5507
20-0.30-1.401.1577391.2165864.8370
21-2.90-3.501.1241651.2165867.5967
22+2.00+0.401.1466481.2165865.7487
23+1.40+1.001.1627011.2165864.4292
24+2.80+1.901.1952571.2165861.7532

Summary statistics of the base series

All values recomputed from the table above. Monthly figures use the sample standard deviation with T-1 in the denominator unless stated. Annualised figures state their annualisation rule because the rules are not equivalent.

StatisticMonthlyAnnualisedRule used
Arithmetic mean return0.766667 percent9.2000 percentP times mu
Geometric mean return0.745939 percent9.3278 percent(1+g)^P - 1
Terminal wealth from 1.00-1.19525673 over 24 monthsproduct of (1+r_t)
Standard deviation2.080273 percent7.2063 percentsigma times sqrt(P)
Skewness gamma3-0.508966-m3/m2^1.5, moment estimator, divisor T
Kurtosis gamma42.368633-m4/m2^2, not excess; excess is -0.631367
Lag-1 autocorrelation rho10.297578-sample autocorrelation, divisor = total sum of squares
Mean excess return over rf0.566667 percent6.8000 percentP times (mu - rf)
Sharpe ratio0.2724000.943622SR times sqrt(P), iid assumption
Sharpe ratio, Lo-corrected0.2724000.713970SR times eta(P), AR(1) plug-in
Sharpe ratio, bias-corrected0.2634030.912456divide by a(T)=1.03415598
Standard error of the Sharpe ratio0.2078760.720104iid normal, Lo 2002
Standard error, non-normal0.2202300.762899Mertens 2002, uses gamma3 and gamma4
Sortino ratio, MAR 0.20 percent0.4443021.539108downside deviation divided by T
Sortino ratio, alternative divisor0.2565180.888604downside deviation divided by count below MAR
Omega, threshold 02.352941-sum of gains above 0 over sum of losses below 0
Omega, threshold 0.20 percent1.894737-same, threshold at the risk-free rate
Maximum drawdown-7.5967 percentpeak-to-trough on the equity index
Calmar ratio-1.227869geometric annual return over max drawdown
Ulcer index-2.9821 percentroot mean square drawdown over all 24 months
Beta against the benchmark0.750000-sample covariance over benchmark variance
Correlation with the benchmark0.855946-R-squared 0.732644
Jensen's alpha0.266667 percent3.2000 percentalpha times P
Tracking error1.228526 percent4.2557 percentsd of arithmetic active return
Information ratio0.1356640.469954active mean over tracking error, times sqrt(P)
Appraisal ratio0.2424660.839927alpha over residual sd, times sqrt(P)
Treynor ratio-0.090667annual excess return over beta
M-squared-10.1606 percentrf_ann + SR_ann times sigma_bench_ann
Up capture, mean-based-0.90993816 up-benchmark months
Down capture, mean-based-0.6123608 down-benchmark months

Why sqrt(P) is the wrong annualisation factor when returns are autocorrelated

Lo 2002 gives the exact scaling factor eta(q) for aggregating a per-period Sharpe ratio to q periods. Under an AR(1) autocorrelation structure rho_k = rho^k the factor is q divided by the square root of q + 2 times the sum over k of (q-k)rho^k. Values below are for q = 12. The last column is the factor by which naive sqrt(12) annualisation overstates the true annual Sharpe ratio.

rho1eta(12)eta(12) / sqrt(12)Overstatement factor of sqrt(12)
-0.204.1707571.2039900.830569
-0.103.7978841.0963530.912113
0.003.4641021.0000001.000000
0.103.1601230.9122001.096250
0.202.8787830.8310031.203365
0.29757771 (base series)2.6210350.7566241.321654
0.302.6148000.7548241.324805
0.402.3634830.6822571.465735
0.502.1213200.6123721.632993

What each ratio divides by

The ratios differ almost entirely in the denominator. Numerators and denominators below are stated per period; worked values are annualised from the base series where an annualisation rule exists.

RatioNumeratorDenominatorWorked, annualised
Sharpemu - rfsigma of total returns0.943622
Sharpe, Lo-correctedmu - rfsigma, aggregated with eta(P)0.713970
Sortinomu - MARdownside deviation below MAR1.539108
Calmargeometric annual returnmaximum drawdown1.227869
Sterling, with 10-point offsetgeometric annual returnaverage annual max drawdown + 0.100.594269
Sterling, no offsetgeometric annual returnaverage annual max drawdown1.637532
Martin, or Ulcer Performance Indexgeometric annual return - rfUlcer index2.314199
Omegasum of returns above the thresholdsum of shortfalls below it2.352941 at threshold 0
Information ratiomean active returntracking error0.469954
Appraisal ratioJensen's alpharesidual sd of the benchmark regression0.839927
Treynormu - rfbeta0.090667
M-squared-expressed as a return, not a ratio10.1606 percent

Entries

Arithmetic mean, geometric mean, and volatility drag

The arithmetic mean is the average of the periodic returns; the geometric mean is the constant periodic return that reproduces the observed terminal wealth. The geometric mean is always at or below the arithmetic mean, and the gap is approximately half the variance. Only the geometric mean describes what a compounding investor received.

FieldValue
Formulamu = (1/T) sum r_t; g = (prod (1+r_t))^(1/T) - 1; drag = mu - g approx sigma^2/2, where sigma^2 is the population variance m2
WorkedBase series: mu = 0.766667 percent per month. Product of (1+r_t) = 1.19525673, so g = 1.19525673^(1/24) - 1 = 0.745939 percent. Drag = 0.766667 - 0.745939 = 0.020728 percent per month. Second-order prediction m2/2 = 0.00041472/2 = 0.00020736 = 0.020736 percent, matching the exact drag to within 8e-8
Cross-check via logsmean of ln(1+r_t) = 0.00743171 per month; exp(0.00743171) - 1 = 0.745939 percent, identical to the geometric mean by construction
Equality conditionmu = g if and only if every r_t is identical. Any dispersion at all makes g strictly smaller
  • The approximation drag = sigma^2/2 uses the population variance with divisor T, not the sample variance with divisor T-1. On the base series the sample variance gives 0.021638 percent against a true drag of 0.020728 percent - a 4 percent error in the drag from choosing the wrong divisor.
  • The relationship is exact only in the continuous limit. It is a second-order Taylor expansion of ln(1+r) and degrades as returns get large: at monthly returns of 20 percent the third-order term is no longer negligible.
  • Reporting an arithmetic mean alongside a maximum drawdown is internally inconsistent, because the drawdown is computed on the compounded path and the arithmetic mean is not achievable on it.

Source: Standard result; the sigma^2/2 approximation follows from Ito's lemma applied to ln(1+r)

The three annualisations, and why two of them disagree

A monthly return can be annualised by multiplying by twelve, by compounding twelve times, or by exponentiating a mean log return. These give three different numbers. Which is correct depends on whether the quantity being annualised is additive in returns, in log returns, or in variance.

FieldValue
Formulasimple: P*mu. compounded: (1+g)^P - 1. continuous: exp(P * mean(ln(1+r_t))) - 1. Variance annualises as P*sigma^2 and volatility as sqrt(P)*sigma
Worked, base seriessimple 12 * 0.766667 percent = 9.2000 percent. compounded (1.00745939)^12 - 1 = 9.3278 percent. continuous exp(12 * 0.00743171) - 1 = 9.3278 percent
The counter-intuitive partThe geometric annualisation (9.3278 percent) is LARGER than the simple annualisation of the arithmetic mean (9.2000 percent), by 12.78 basis points, even though the monthly geometric mean is smaller than the monthly arithmetic mean
WhyTwelve-fold compounding of the smaller monthly figure contributes more than the monthly volatility drag subtracts. The two effects work in opposite directions and neither dominates by construction
Volatility2.080273 percent monthly times sqrt(12) = 7.2063 percent annual. Variance 0.00043275 monthly times 12 = 0.00519300 annual
  • The often-quoted identity g_annual = mu_annual - sigma_annual^2/2 does not hold across an annualisation boundary. On the base series it gives 9.2000 - 7.2063^2/2 = 8.9403 percent against a true 9.3278 percent, an error of 39 basis points. The identity applies to log returns at a single frequency, not to a simply annualised arithmetic mean.
  • Volatility scales with sqrt(P) only for serially uncorrelated returns, exactly the same condition that Sharpe annualisation requires. If rho1 is nonzero, the annualised volatility is wrong too and in the opposite direction from the Sharpe error.
  • A performance table that does not state which of the three rules produced its annual return is not reproducible. On this series the spread between the rules is 39 basis points on a 9 percent return; on a volatile series it is much larger.

Sharpe ratio

Mean excess return divided by the standard deviation of returns, both measured per period. It is a t-statistic in disguise: the same quantity, scaled by sqrt(T), tests the hypothesis that mean excess return is zero.

FieldValue
FormulaSR = (mu - rf) / sigma, with sigma the sample standard deviation of returns (divisor T-1). Annualised SR_ann = SR * sqrt(P) only if returns are serially uncorrelated
WorkedBase series: mu = 0.766667 percent, rf = 0.200000 percent, so mean excess = 0.566667 percent. sigma = 2.080273 percent. SR = 0.566667/2.080273 = 0.272400 per month. Annualised at sqrt(12): 0.272400 * 3.464102 = 0.943622
Denominator conventionWith a constant risk-free rate the standard deviation of returns and of excess returns are identical (both 2.080273 percent here), so the convention is invisible. With a time-varying rf they differ and the excess-return standard deviation is the correct one
As a t-statistict = SR * sqrt(T) / sqrt(1 + SR^2/2) = 0.272400 * 4.898979 / 1.018446 = 1.310396, so this Sharpe ratio is not significantly different from zero at 24 observations
Scale invarianceUnlevering or levering a return series by a constant factor leaves SR unchanged. Adding a constant to every return does not
  • Subtracting rf from the numerator but not adjusting the denominator is correct; subtracting it twice, or comparing a Sharpe ratio computed on total returns with one computed on excess returns, is not. On a series with 5 percent rf and 8 percent volatility the difference is more than half a Sharpe point.
  • The ratio is undefined as a ranking device across return distributions with different higher moments. Two strategies with identical mu and sigma but skewness of +1 and -1 receive the same Sharpe ratio and are not the same risk.
  • Selling out-of-the-money options raises the Sharpe ratio of a short sample almost mechanically, because the premium enters the numerator every period and the tail enters the denominator only when it occurs. A high Sharpe ratio over a short sample is evidence about the sample, not the strategy.

Source: Sharpe 1966, revised in Sharpe 1994

Annualising a Sharpe ratio when returns are autocorrelated

Multiplying a per-period Sharpe ratio by sqrt(P) assumes the returns are independently and identically distributed. Under serial correlation the correct factor is Lo's eta(q), which is smaller than sqrt(q) for positive autocorrelation and larger for negative. Positive autocorrelation is the common case in illiquid or marked-to-model portfolios, so the naive factor usually overstates.

FieldValue
FormulaSR(q) = eta(q) * SR, with eta(q) = q / sqrt( q + 2 * sum_{k=1}^{q-1} (q-k) * rho_k ). Under AR(1), rho_k = rho1^k. If all rho_k = 0 then eta(q) = sqrt(q)
Worked, AR(1) plug-inBase series rho1 = 0.297578, q = 12. sum (q-k)rho1^k = 4.480618, so q + 2*sum = 20.961236 and eta(12) = 12/sqrt(20.961236) = 2.621035. Corrected annual Sharpe = 0.272400 * 2.621035 = 0.713970 against 0.943622 from sqrt(12)
Overstatementsqrt(12)/eta(12) = 3.464102/2.621035 = 1.321654, so the naive factor overstates the annual Sharpe ratio by 32.17 percent on this series
Worked, full sample lagsUsing the 11 estimated autocorrelations directly (0.2976, 0.0813, -0.0370, 0.0789, 0.1128, -0.0316, -0.0930, -0.0828, 0.0569, 0.0016, 0.0326): sum (q-k)rho_k = 4.393994, eta = 2.631934, SR_ann = 0.716939
Sensitivityrho1 = 0.10 gives eta = 3.160123 and an overstatement of 9.63 percent; rho1 = 0.50 gives eta = 2.121320 and 63.30 percent
  • Estimating 11 autocorrelations from 24 observations is not defensible: each rho_k beyond the first few is noise. The AR(1) plug-in, which spends one parameter, is the practical choice at short sample lengths and is what the worked value above uses. On this series the two approaches happen to agree to within 0.003 Sharpe points, which should not be read as evidence that they generally do.
  • Positive autocorrelation in reported returns is the signature of smoothed or stale marks, not of skill. Getmansky, Lo and Makarov 2004 model exactly this, and the implication is that the correction should be applied before, not after, comparing an illiquid book to a liquid one.
  • The correction changes the annualised Sharpe ratio and nothing else. The per-period Sharpe ratio is unaffected: the autocorrelation problem is entirely a problem of aggregation across periods.
  • eta(q) can exceed sqrt(q). Negatively autocorrelated returns - mean-reverting or overhedged books - have their annual Sharpe ratio understated by the naive factor, which is why the correction should be applied symmetrically rather than only when it flatters.

Source: Lo 2002

Standard error of the Sharpe ratio

The Sharpe ratio is an estimate with a standard error that depends on the sample length and on the higher moments of the return distribution. Under independent normal returns the standard error has a closed form; under general distributions it picks up skewness and kurtosis terms.

FieldValue
Formulaiid normal: SE(SR) = sqrt( (1 + SR^2/2) / T ). General: SE(SR) = sqrt( (1 - gamma3*SR + ((gamma4 - 1)/4)*SR^2) / T ), with gamma3 skewness and gamma4 kurtosis (3 for a normal)
Worked, iidSR = 0.272400, T = 24. SE = sqrt((1 + 0.272400^2/2)/24) = sqrt(1.03710/24) = sqrt(0.04321) = 0.207876 per month; times sqrt(12) = 0.720104 annualised
Worked, non-normalgamma3 = -0.508966, gamma4 = 2.368633. SE = sqrt((1 - (-0.508966)(0.272400) + ((2.368633-3)/4)(0.272400^2))/24) = sqrt((1 + 0.138642 - 0.011713)/24) = sqrt(0.046955) = 0.220230; annualised 0.762899
95 percent interval, annualisediid: 0.943622 +/- 1.96 * 0.720104 = -0.467783 to 2.355027. Non-normal: -0.527639 to 2.414882. Both intervals contain zero
Sample length for significanceMonths required for t = 1.96 at a given true annual Sharpe: SR_ann 0.5 needs 186.32 months (15.53 years); 1.0 needs 48.02 (4.00 years); 1.5 needs 22.41 (1.87 years); 2.0 needs 13.45 (1.12 years)
  • The interval on the base series spans nearly three Sharpe points. Two years of monthly data cannot distinguish a Sharpe ratio of 0.9 from one of zero, and no estimator improvement changes that; it is a sample-size limit.
  • Negative skewness makes the standard error LARGER here because the -gamma3*SR term is positive when gamma3 is negative. That is the opposite of what many summaries claim, and the sign follows directly from the formula.
  • This standard error assumes serially uncorrelated returns. If it is applied to autocorrelated returns it is too small, in addition to the annualisation error covered separately. Lo 2002 gives the autocorrelation-robust version.
  • Comparing two Sharpe ratios requires the standard error of their difference, which includes the covariance of the two return series and is not the root sum of squares of the individual standard errors. Jobson and Korkie 1981, with the correction in Memmel 2003, gives the test.

Source: Lo 2002 for the iid case; Mertens 2002 for the non-normal correction

Small-sample bias in the Sharpe ratio

The plug-in Sharpe estimator is biased upward in small samples. The sample standard deviation is a downward-biased estimator of sigma, and dividing by a number that is too small on average makes the ratio too large on average. The bias is a pure function of T and can be removed exactly under normality.

FieldValue
FormulaE[SR_hat] = SR * a(T), with a(T) = sqrt((T-1)/2) * Gamma((T-2)/2) / Gamma((T-1)/2). Unbiased estimator: SR_hat / a(T)
WorkedT = 24: a(24) = sqrt(11.5) * Gamma(11)/Gamma(11.5) = 1.03415598, an upward bias of 3.4156 percent. Corrected monthly SR = 0.272400/1.03415598 = 0.263403; annualised at sqrt(12) = 0.912456 against 0.943622 uncorrected
Bias by sample lengthT = 12: a = 1.075315, bias 7.532 percent. T = 24: 1.034156, 3.416 percent. T = 36: 1.022086, 2.209 percent. T = 60: 1.012940, 1.294 percent. T = 120: 1.006358, 0.636 percent. T = 240: 1.003152, 0.315 percent
Related constantc4(T) = sqrt(2/(T-1)) * Gamma(T/2)/Gamma((T-1)/2) = 0.98919267 at T = 24 is the bias factor of the sample standard deviation itself: E[s] = c4 * sigma
  • The bias is small relative to the standard error and is routinely mistaken for the whole small-sample problem. At T = 24 the bias is 3.4 percent of the Sharpe ratio while the standard error is 76 percent of it. Correcting the bias and reporting a point estimate without an interval fixes the smaller error and leaves the larger one.
  • The bias is always upward, never downward, and it is largest exactly where track records are shortest. A twelve-month track record overstates its Sharpe ratio by 7.5 percent before any selection effect is considered.
  • The correction assumes normal iid returns. Under fat tails the exact factor differs, and no closed form is available; the direction of the bias is unchanged.

Source: Miller and Gehr 1978

Sortino ratio, and the divisor that changes the answer

Excess return over a minimum acceptable return divided by downside deviation - the root mean square of shortfalls below that target. The ambiguity is the divisor in the downside deviation: the full sample length, or only the count of observations below the target. The two answers differ by a large factor and both appear in commercial reporting.

FieldValue
FormulaSortino = (mu - MAR) / DD, with DD = sqrt( (1/T) * sum_t min(r_t - MAR, 0)^2 ). The alternative uses 1/n_below in place of 1/T, where n_below is the count of periods with r_t < MAR
Worked, divisor TMAR = 0.20 percent per month. 8 of 24 months fall below it. Sum of squared shortfalls = 0.00390400. DD = sqrt(0.00390400/24) = 1.275408 percent. Sortino = (0.766667 - 0.200000)/1.275408 = 0.444302 per month; annualised 1.539108
Worked, divisor n_belowDD = sqrt(0.00390400/8) = 2.209072 percent. Sortino = 0.566667/2.209072 = 0.256518; annualised 0.888604
Ratio of the twosqrt(T/n_below) = sqrt(24/8) = 1.732051. The T-divisor version is 73.2 percent higher on this series, and the gap grows as the strategy has fewer losing periods
MAR conventionsMAR = 0 gives DD = 1.198263 percent (divisor T) and Sortino 0.639806. MAR = rf gives the value above. MAR = the arithmetic mean makes the Sortino ratio zero by construction
  • The divisor-T version is the one implied by the original semi-deviation definition and is the more defensible: it treats an upside month as a zero shortfall rather than dropping it, which keeps the statistic a comparable second moment. The divisor-n_below version is not a deviation of the return distribution at all; it is a conditional deviation of the losses.
  • A strategy with no periods below the MAR has an infinite Sortino ratio under either divisor, and the divisor-n_below version is undefined. Any strategy whose losses are rare and large will therefore look better on Sortino than on Sharpe until the first large loss lands.
  • Annualising a Sortino ratio by sqrt(P) inherits the same autocorrelation problem as the Sharpe ratio, with no published analogue of Lo's correction. Treat an annualised Sortino ratio on autocorrelated returns as unquantified.
  • The MAR must be stated. Sortino ratios computed against 0, against the risk-free rate, and against an absolute return target are three different statistics and are not comparable.

Source: Sortino and van der Meer 1991

Calmar ratio

Compound annual return divided by maximum drawdown over the same window. It is the only common ratio whose denominator is a single realised extreme rather than a moment of the distribution, which makes it the least stable of the family.

FieldValue
FormulaCalmar = g_ann / |MaxDD|, both measured over the same window. The original specification uses a 36-month window and a compound annual return
Workedg_ann = 9.3278 percent, MaxDD = 7.5967 percent. Calmar = 0.093278/0.075967 = 1.227869
With arithmetic annualisation12 * 0.766667 percent = 9.2000 percent gives 1.211046, 1.4 percent lower. The convention must be stated
Excess-return variantUsing return over the risk-free rate: (9.3278 - 2.4266) percent / 7.5967 percent = 0.908445. Sometimes called the Sterling-Calmar ratio
Window dependenceOver the first 12 months alone the max drawdown is 4.0820 percent; over the second 12 months alone it is 7.3105 percent. Calmar computed on either half is not comparable with Calmar computed on the whole
  • Maximum drawdown is non-decreasing in the length of the window, so Calmar mechanically falls as a track record lengthens even if nothing about the strategy changes. Comparing the Calmar ratios of a two-year record and a ten-year record compares window lengths, not strategies.
  • The denominator is one number from one path. Its sampling variability is enormous and there is no accepted standard error for it, which means Calmar has no interval and cannot support a significance statement.
  • The original 36-month convention exists precisely to make the statistic comparable across managers. A Calmar ratio computed over whatever window happens to flatter the record is not the same statistic.

Source: Young 1991

Sterling ratio, and its incompatible definitions

A return-over-drawdown ratio whose denominator is the average of the annual maximum drawdowns rather than the single worst drawdown. At least three definitions circulate, differing in whether a fixed offset is added to the denominator and whether the average is over annual maxima or over all drawdowns.

FieldValue
FormulaSterling = g_ann / (average annual maximum drawdown + 0.10) in the form with the offset; g_ann / (average annual maximum drawdown) without it; and g_ann / (average of all drawdowns) in a third variant
Worked inputsBase series year 1 maximum drawdown 4.0820 percent, year 2 maximum drawdown 7.3105 percent. Average annual maximum drawdown = 5.6963 percent
Worked, with the 10-point offset0.093278 / (0.056963 + 0.10) = 0.093278/0.156963 = 0.594269
Worked, without the offset0.093278 / 0.056963 = 1.637532
Spread between definitions1.637532 versus 0.594269 - a factor of 2.755 on identical data. No other statistic in this corpus has a wider definitional spread
  • The 10 percentage-point offset in the original form has no statistical justification; it is a floor that prevents the ratio exploding for strategies with small drawdowns. Its practical effect is to compress differences between low-drawdown strategies almost to nothing, which is either the point or the flaw depending on the use.
  • A Sterling ratio without a stated definition is uninterpretable. Given the factor of 2.755 spread on this series, comparing two Sterling ratios from two sources is comparing two unrelated numbers.
  • The averaging over annual maxima makes the statistic depend on where the calendar year boundaries fall. Shifting the year boundary by six months changes the annual maxima and therefore the ratio, on identical returns.

Omega ratio

The ratio of the probability-weighted gains above a threshold to the probability-weighted losses below it. Unlike the Sharpe ratio it uses the entire return distribution rather than its first two moments, and it is a function of the threshold rather than a single number.

FieldValue
FormulaOmega(tau) = sum_t max(r_t - tau, 0) / sum_t max(tau - r_t, 0). Equivalently, the ratio of the integral of the survival function above tau to the integral of the cumulative distribution below tau
Worked, tau = 0Base series: gains above zero sum to 32.0000 percent, shortfalls below zero sum to 13.6000 percent. Omega(0) = 0.320000/0.136000 = 2.352941
Worked, tau = rfAt tau = 0.20 percent: gains 28.8000 percent, losses 15.2000 percent, Omega = 1.894737
Identity(gains - losses)/T = mu - tau exactly. Check: (0.320000 - 0.136000)/24 = 0.00766667 = mu - 0 . And (0.288000 - 0.152000)/24 = 0.00566667 = mu - rf
Fixed pointOmega(tau) = 1 exactly when tau = mu. Omega is monotonically decreasing in tau, crossing 1 at the arithmetic mean and tending to 0 as tau rises above the maximum return
  • Because Omega equals 1 at the mean and is monotone, Omega at a threshold below the mean carries the same ordinal information as the mean itself for a fixed distribution shape. It only adds information when comparing distributions of different shape at a threshold that matters to the user.
  • Omega uses every observation, which makes it more stable than a drawdown-based ratio, but it is still a ratio of two sample sums and inherits their sampling error. No standard closed-form standard error exists.
  • The whole Omega curve, not one point on it, is the statistic. Two strategies can cross: one better below the threshold and worse above it. Reporting Omega at a single threshold hides exactly the case Omega was designed to reveal.

Source: Keating and Shadwick 2002

The benchmark regression, and its standard errors

Beta, alpha, R-squared, tracking error and residual volatility all come from a single regression of excess portfolio returns on excess benchmark returns. Quoting any of them without the regression standard errors treats an estimate as a measurement.

FieldValue
Formular_t - rf = alpha + beta*(b_t - rf) + e_t. beta = Cov(r,b)/Var(b); alpha = mean(r - rf) - beta*mean(b - rf); R^2 = corr(r,b)^2; residual sd s_e = sqrt(SSE/(T-2))
Worked, point estimatesBase series: Cov(r,b) = 0.0004227391, Var(b) = 0.0005636522, so beta = 0.750000 exactly. corr = 0.855946, R-squared = 0.732644. alpha = 0.266667 percent per month = 3.2000 percent annualised
Worked, standard errorsResidual sd s_e = 1.099811 percent. SE(beta) = 0.09659361, t(beta) = 7.764489. SE(alpha) = 0.00227799, t(alpha) = 1.170625
Testing beta = 1t = (0.750000 - 1)/0.09659361 = -2.588163, so beta is significantly below 1 at 24 observations even though alpha is not significantly above 0
Alpha interval95 percent interval on annualised alpha: -2.1578 percent to 8.5578 percent
  • Beta is estimated far more precisely than alpha, always. That is a structural property of the regression, not a feature of this dataset: alpha is the intercept and absorbs all the residual noise. A track record can establish its market exposure long before it can establish its skill.
  • The regression must be run on excess returns on both sides. Running it on total returns produces an intercept that mixes alpha with (1 - beta) times the risk-free rate, which is not Jensen's alpha and is biased whenever beta is far from 1.
  • R-squared of 0.73 means 27 percent of the variance is unexplained by the benchmark. It says nothing about whether the residual is skill; it is the same residual that alpha's standard error is computed from, and here it is not significantly nonzero.
  • Ordinary least squares standard errors assume homoskedastic uncorrelated residuals. On real return data neither holds, and the honest version uses Newey-West or a comparable robust estimator, which widens the alpha interval further.

Jensen's alpha

The intercept of the excess-return regression: the average return not explained by the portfolio's exposure to the benchmark. It is a return, in percent, not a ratio, and it has a standard error.

FieldValue
Formulaalpha = (mu - rf) - beta * (mu_b - rf). Annualised as P * alpha
Workedmu - rf = 0.566667 percent, mu_b - rf = 0.400000 percent, beta = 0.750000. alpha = 0.566667 - 0.750000 * 0.400000 = 0.566667 - 0.300000 = 0.266667 percent per month. Annualised 12 * 0.266667 = 3.2000 percent
Cross-checkThe OLS intercept computed directly from the regression is 0.00266667, identical. The closed form and the regression are the same calculation
Significancet(alpha) = 1.170625 at T = 24. Not significant. 95 percent interval on annualised alpha: -2.1578 percent to 8.5578 percent
Appraisal ratioalpha divided by residual standard deviation: 0.00266667/0.01099811 = 0.242466 per month, 0.839927 annualised. This is the Sharpe-like scaling of alpha
  • Alpha is defined only relative to the specified benchmark and the specified beta. Adding a second factor to the regression changes both, and an alpha that survives one factor and not two was never alpha; it was unmodelled exposure.
  • The annualisation multiplies alpha by P, not by sqrt(P), because alpha is a mean return rather than a ratio. Applying sqrt(P) to alpha is a common and large error: on this series it would give 0.9238 percent instead of 3.2000 percent.
  • A positive alpha with an insignificant t-statistic is the normal state of a short track record and is not evidence of skill. Here the point estimate is 3.2 percent per year and the interval includes -2.2 percent.

Source: Jensen 1968; the appraisal ratio is Treynor and Black 1973

Information ratio, and the two things it is called

Mean active return divided by tracking error. The ambiguity is what counts as active return: the arithmetic difference from the benchmark, or the regression residual. The first gives the information ratio proper, the second the appraisal ratio. They coincide only when beta equals 1.

FieldValue
FormulaIR = mean(r_t - b_t) / sd(r_t - b_t), annualised by sqrt(P). Appraisal ratio = alpha / s_e, also annualised by sqrt(P)
Worked, information ratioBase series: mean active return = 0.166667 percent per month, tracking error = 1.228526 percent. IR = 0.135664 per month, 0.469954 annualised. Annualised tracking error 4.2557 percent, annualised active return 2.0000 percent
Worked, appraisal ratioalpha = 0.266667 percent, residual sd = 1.099811 percent. Ratio = 0.242466 per month, 0.839927 annualised
Gap0.839927 versus 0.469954 - the appraisal ratio is 79 percent higher on identical data, because beta of 0.750000 means the arithmetic difference r - b carries a large residual benchmark exposure that the regression removes
When they agreeOnly when beta = 1. At beta = 1 the regression residual and the arithmetic active return are the same series
  • A portfolio with beta well away from 1 has a tracking error dominated by the beta mismatch rather than by security selection. Its information ratio is then mostly a statement about its market exposure, which is why the appraisal ratio is the right number for a stock-selection mandate and the information ratio the right one for a tracking mandate.
  • The two are routinely reported under the same label. When a document quotes an information ratio, check whether the denominator is the standard deviation of r - b or the residual standard deviation of a regression; on this series that choice moves the number by 0.37.
  • Annualising by sqrt(P) assumes serially uncorrelated active returns. Active returns of a portfolio rebalanced monthly against a benchmark rebalanced quarterly are autocorrelated by construction.

Source: Grinold and Kahn, Active Portfolio Management; the appraisal ratio is Treynor and Black 1973

Treynor ratio

Mean excess return divided by beta rather than by volatility. It prices only systematic risk, which makes it the right ratio for a component of a diversified portfolio and the wrong one for a standalone allocation.

FieldValue
FormulaTreynor = (mu - rf) / beta, conventionally stated in annualised return units
WorkedAnnual excess return 6.8000 percent, beta 0.750000. Treynor = 0.068000/0.750000 = 0.090667
Benchmark comparisonThe benchmark's own Treynor ratio is its annual excess return over a beta of 1: 0.072000 - 0.024000 = 0.048000. The strategy's 0.090667 is higher, which is the sign that alpha is positive
Relationship to alphaTreynor exceeds the benchmark's Treynor exactly when Jensen's alpha is positive: alpha = beta * (Treynor_p - Treynor_b) = 0.750000 * (0.090667 - 0.048000) = 0.032000, matching the annualised alpha
UnitsNot dimensionless. Treynor is a return per unit of beta, so its magnitude depends on the annualisation of the numerator
  • Treynor ignores idiosyncratic risk entirely. A single-stock portfolio and a thousand-stock portfolio with the same beta and the same excess return receive the same Treynor ratio, which is correct if and only if the holder is diversified elsewhere.
  • The ratio is unstable and sign-flipping for betas near zero, and meaningless for negative betas: a short-biased fund with negative beta and positive excess return gets a negative Treynor ratio that ranks it below a losing long fund.
  • Beta carries a standard error - 0.09659361 on this series - so Treynor carries one too. The ratio is more precisely estimated than the Sharpe ratio only because beta is estimated more precisely than volatility, not because the numerator is better known.

Source: Treynor 1965

M-squared

The return the portfolio would have earned if it had been levered or delevered to exactly the benchmark's volatility. It converts a Sharpe ratio into a return, which makes it directly comparable with the benchmark's return in percentage points rather than in ratio units.

FieldValue
FormulaM2 = rf_ann + SR_ann * sigma_bench_ann. M2 alpha = M2 - mu_bench_ann. The implied leverage is sigma_bench/sigma_portfolio
WorkedSR_ann = 0.943622, sigma_bench_ann = 8.2242 percent, rf_ann = 2.4000 percent (12 times 0.20 percent). M2 = 0.024000 + 0.943622 * 0.082242 = 0.024000 + 0.077606 = 10.1606 percent
M-squared alphaBenchmark annual arithmetic return 7.2000 percent. M2 alpha = 10.1606 - 7.2000 = 2.9606 percent
Implied leveragesigma_bench/sigma_port = 8.2242/7.2063 = 1.141262, so the comparison is against the strategy run at 114.1 percent of its observed size
OrderingM2 ranks identically to the Sharpe ratio for a fixed benchmark, since it is an affine transformation of it. Its value is interpretability, not a different ordering
  • M-squared assumes leverage is available at the risk-free rate and that levering the strategy scales its return linearly. Neither holds for a strategy whose capacity is limited or whose financing spread widens with size, and the M-squared alpha is overstated by exactly the financing cost that was ignored.
  • Because M2 is affine in the Sharpe ratio, it inherits the Sharpe ratio's entire standard error. The 2.96 percent M-squared alpha on this series is not distinguishable from zero at 24 observations.
  • The benchmark volatility must be measured over the same window as the strategy volatility. Using a long-run benchmark volatility against a short-window strategy volatility mixes two sample periods into one number.

Source: Modigliani and Modigliani 1997

Up and down capture, and the two ways to compute them

The fraction of the benchmark's up-period return the portfolio captured, and the fraction of its down-period loss the portfolio suffered. Computed either as a ratio of arithmetic means over the selected periods or as a ratio of compounded returns over them, giving different answers.

FieldValue
FormulaMean-based: UpCapture = mean(r_t | b_t > 0) / mean(b_t | b_t > 0). Compound-based: UpCapture = (prod over up periods of (1+r_t) - 1) / (prod over up periods of (1+b_t) - 1). Down capture uses b_t < 0
Worked, mean-based16 up months: mean strategy 1.831250 percent, mean benchmark 2.012500 percent, up capture 0.909938. 8 down months: mean strategy -1.362500 percent, mean benchmark -2.225000 percent, down capture 0.612360
Worked, compound-basedUp months compounded: strategy +33.508928 percent, benchmark +37.397292 percent, ratio 0.896026. Down months: strategy -10.473647 percent, benchmark -16.522009 percent, ratio 0.633921
Capture ratioUp capture divided by down capture: 0.909938/0.612360 = 1.485954 mean-based, 0.896026/0.633921 = 1.413306 compound-based
Period classificationPeriods are classified by the sign of the BENCHMARK return, not the strategy return. 16 up and 8 down months here, with no zero-return months
  • Capture ratios are not independent of beta. A portfolio with beta 0.75 and zero alpha has an up capture of about 0.75 and a down capture of about 0.75 mechanically. The informative quantity is the gap between the two, not either level: here 0.910 against 0.612 on a beta of 0.75 is what asymmetry looks like.
  • The statistic is entirely determined by the sign classification, and the sign of a benchmark return in a period near zero is noise. Reclassifying one marginal month can move a capture ratio by several points on a two-year sample.
  • Down capture below 1 over a sample containing few down periods is the least reliable number in a performance table. Eight down months is not a sample from which a downside characteristic can be inferred.
  • Compounding the selected periods, as the compound-based version does, treats non-adjacent months as if they were consecutive. It is internally consistent but it is not a return anyone earned.

Reference data. Reviewed 2026-08-27. Machine-readable: /performance.json. Corpus manifest: /llms.txt.

Published and maintained by · [email protected]. A reference published by the wallstreet.wiki network. Every figure is stated as a formula and recomputed from it, every convention names the authority that sets it, and corrections are versioned and dated. About this reference.

Reference information only. Not investment advice, and not a recommendation of any strategy, estimator or allocation. The estimators described here carry explicit assumptions - independence, stationarity, normality, zero drift, continuous monitoring, known parameters - and they are not interchangeable: two of them applied to the same data will disagree, and the disagreement is a property of the estimators rather than an error in either. Figures labelled Worked are arithmetic examples computed from the inputs stated alongside them; none of them is an empirical finding about any market, instrument or manager, and the published return series, OHLC bars and covariance matrix are constructed data for that purpose.