Performance statistics
Sharpe, Sortino, Calmar, Sterling, Omega, information ratio, Treynor, Jensen's alpha, M-squared and capture, each with its exact estimator, annualisation rule and standard error.
Every number in this section is computed from the 24-month return series in the first table below, with a constant risk-free rate of 0.20 percent per month. Where a statistic has more than one estimator in general use - and most of them do - each variant is given separately with its own worked value, because the variants do not agree and the disagreement is usually larger than the difference between two managers. The annualisation rule matters as much as the estimator: multiplying a monthly Sharpe ratio by the square root of twelve is correct only for serially uncorrelated returns, and the correction for the general case is given explicitly.
The base return series
Twenty-four monthly returns in percent for a strategy and a benchmark, the strategy equity index starting from 1, its running peak, and the drawdown from that peak. Every statistic in this corpus is computed from these columns. The risk-free rate is 0.20 percent per month throughout. Sum of strategy returns 18.40 percent, sum of benchmark returns 14.40 percent.
| Month | Strategy return, percent | Benchmark return, percent | Strategy equity | Running peak | Drawdown, percent |
|---|---|---|---|---|---|
| 1 | +1.60 | +4.00 | 1.016000 | 1.016000 | 0.0000 |
| 2 | +4.10 | +3.90 | 1.057656 | 1.057656 | 0.0000 |
| 3 | +2.30 | +0.70 | 1.081982 | 1.081982 | 0.0000 |
| 4 | +3.40 | +3.00 | 1.118769 | 1.118769 | 0.0000 |
| 5 | +1.50 | +0.80 | 1.135551 | 1.135551 | 0.0000 |
| 6 | -0.80 | -0.40 | 1.126467 | 1.135551 | 0.8000 |
| 7 | +3.20 | +2.30 | 1.162514 | 1.162514 | 0.0000 |
| 8 | +2.70 | +3.20 | 1.193901 | 1.193901 | 0.0000 |
| 9 | +1.90 | +1.40 | 1.216586 | 1.216586 | 0.0000 |
| 10 | -3.60 | -3.90 | 1.172788 | 1.216586 | 3.6000 |
| 11 | -0.50 | +1.10 | 1.166925 | 1.216586 | 4.0820 |
| 12 | +0.60 | -1.30 | 1.173926 | 1.216586 | 3.5065 |
| 13 | +2.10 | +2.80 | 1.198579 | 1.216586 | 1.4801 |
| 14 | +1.10 | +3.10 | 1.211763 | 1.216586 | 0.3964 |
| 15 | -1.20 | +0.30 | 1.197222 | 1.216586 | 1.5917 |
| 16 | +0.90 | +2.30 | 1.207997 | 1.216586 | 0.7060 |
| 17 | +0.40 | -1.40 | 1.212829 | 1.216586 | 0.3088 |
| 18 | -2.50 | -2.90 | 1.182508 | 1.216586 | 2.8011 |
| 19 | -1.80 | -3.00 | 1.161223 | 1.216586 | 4.5507 |
| 20 | -0.30 | -1.40 | 1.157739 | 1.216586 | 4.8370 |
| 21 | -2.90 | -3.50 | 1.124165 | 1.216586 | 7.5967 |
| 22 | +2.00 | +0.40 | 1.146648 | 1.216586 | 5.7487 |
| 23 | +1.40 | +1.00 | 1.162701 | 1.216586 | 4.4292 |
| 24 | +2.80 | +1.90 | 1.195257 | 1.216586 | 1.7532 |
Summary statistics of the base series
All values recomputed from the table above. Monthly figures use the sample standard deviation with T-1 in the denominator unless stated. Annualised figures state their annualisation rule because the rules are not equivalent.
| Statistic | Monthly | Annualised | Rule used |
|---|---|---|---|
| Arithmetic mean return | 0.766667 percent | 9.2000 percent | P times mu |
| Geometric mean return | 0.745939 percent | 9.3278 percent | (1+g)^P - 1 |
| Terminal wealth from 1.00 | - | 1.19525673 over 24 months | product of (1+r_t) |
| Standard deviation | 2.080273 percent | 7.2063 percent | sigma times sqrt(P) |
| Skewness gamma3 | -0.508966 | - | m3/m2^1.5, moment estimator, divisor T |
| Kurtosis gamma4 | 2.368633 | - | m4/m2^2, not excess; excess is -0.631367 |
| Lag-1 autocorrelation rho1 | 0.297578 | - | sample autocorrelation, divisor = total sum of squares |
| Mean excess return over rf | 0.566667 percent | 6.8000 percent | P times (mu - rf) |
| Sharpe ratio | 0.272400 | 0.943622 | SR times sqrt(P), iid assumption |
| Sharpe ratio, Lo-corrected | 0.272400 | 0.713970 | SR times eta(P), AR(1) plug-in |
| Sharpe ratio, bias-corrected | 0.263403 | 0.912456 | divide by a(T)=1.03415598 |
| Standard error of the Sharpe ratio | 0.207876 | 0.720104 | iid normal, Lo 2002 |
| Standard error, non-normal | 0.220230 | 0.762899 | Mertens 2002, uses gamma3 and gamma4 |
| Sortino ratio, MAR 0.20 percent | 0.444302 | 1.539108 | downside deviation divided by T |
| Sortino ratio, alternative divisor | 0.256518 | 0.888604 | downside deviation divided by count below MAR |
| Omega, threshold 0 | 2.352941 | - | sum of gains above 0 over sum of losses below 0 |
| Omega, threshold 0.20 percent | 1.894737 | - | same, threshold at the risk-free rate |
| Maximum drawdown | - | 7.5967 percent | peak-to-trough on the equity index |
| Calmar ratio | - | 1.227869 | geometric annual return over max drawdown |
| Ulcer index | - | 2.9821 percent | root mean square drawdown over all 24 months |
| Beta against the benchmark | 0.750000 | - | sample covariance over benchmark variance |
| Correlation with the benchmark | 0.855946 | - | R-squared 0.732644 |
| Jensen's alpha | 0.266667 percent | 3.2000 percent | alpha times P |
| Tracking error | 1.228526 percent | 4.2557 percent | sd of arithmetic active return |
| Information ratio | 0.135664 | 0.469954 | active mean over tracking error, times sqrt(P) |
| Appraisal ratio | 0.242466 | 0.839927 | alpha over residual sd, times sqrt(P) |
| Treynor ratio | - | 0.090667 | annual excess return over beta |
| M-squared | - | 10.1606 percent | rf_ann + SR_ann times sigma_bench_ann |
| Up capture, mean-based | - | 0.909938 | 16 up-benchmark months |
| Down capture, mean-based | - | 0.612360 | 8 down-benchmark months |
Why sqrt(P) is the wrong annualisation factor when returns are autocorrelated
Lo 2002 gives the exact scaling factor eta(q) for aggregating a per-period Sharpe ratio to q periods. Under an AR(1) autocorrelation structure rho_k = rho^k the factor is q divided by the square root of q + 2 times the sum over k of (q-k)rho^k. Values below are for q = 12. The last column is the factor by which naive sqrt(12) annualisation overstates the true annual Sharpe ratio.
| rho1 | eta(12) | eta(12) / sqrt(12) | Overstatement factor of sqrt(12) |
|---|---|---|---|
| -0.20 | 4.170757 | 1.203990 | 0.830569 |
| -0.10 | 3.797884 | 1.096353 | 0.912113 |
| 0.00 | 3.464102 | 1.000000 | 1.000000 |
| 0.10 | 3.160123 | 0.912200 | 1.096250 |
| 0.20 | 2.878783 | 0.831003 | 1.203365 |
| 0.29757771 (base series) | 2.621035 | 0.756624 | 1.321654 |
| 0.30 | 2.614800 | 0.754824 | 1.324805 |
| 0.40 | 2.363483 | 0.682257 | 1.465735 |
| 0.50 | 2.121320 | 0.612372 | 1.632993 |
What each ratio divides by
The ratios differ almost entirely in the denominator. Numerators and denominators below are stated per period; worked values are annualised from the base series where an annualisation rule exists.
| Ratio | Numerator | Denominator | Worked, annualised |
|---|---|---|---|
| Sharpe | mu - rf | sigma of total returns | 0.943622 |
| Sharpe, Lo-corrected | mu - rf | sigma, aggregated with eta(P) | 0.713970 |
| Sortino | mu - MAR | downside deviation below MAR | 1.539108 |
| Calmar | geometric annual return | maximum drawdown | 1.227869 |
| Sterling, with 10-point offset | geometric annual return | average annual max drawdown + 0.10 | 0.594269 |
| Sterling, no offset | geometric annual return | average annual max drawdown | 1.637532 |
| Martin, or Ulcer Performance Index | geometric annual return - rf | Ulcer index | 2.314199 |
| Omega | sum of returns above the threshold | sum of shortfalls below it | 2.352941 at threshold 0 |
| Information ratio | mean active return | tracking error | 0.469954 |
| Appraisal ratio | Jensen's alpha | residual sd of the benchmark regression | 0.839927 |
| Treynor | mu - rf | beta | 0.090667 |
| M-squared | - | expressed as a return, not a ratio | 10.1606 percent |
Entries
Arithmetic mean, geometric mean, and volatility drag
The arithmetic mean is the average of the periodic returns; the geometric mean is the constant periodic return that reproduces the observed terminal wealth. The geometric mean is always at or below the arithmetic mean, and the gap is approximately half the variance. Only the geometric mean describes what a compounding investor received.
| Field | Value |
|---|---|
| Formula | mu = (1/T) sum r_t; g = (prod (1+r_t))^(1/T) - 1; drag = mu - g approx sigma^2/2, where sigma^2 is the population variance m2 |
| Worked | Base series: mu = 0.766667 percent per month. Product of (1+r_t) = 1.19525673, so g = 1.19525673^(1/24) - 1 = 0.745939 percent. Drag = 0.766667 - 0.745939 = 0.020728 percent per month. Second-order prediction m2/2 = 0.00041472/2 = 0.00020736 = 0.020736 percent, matching the exact drag to within 8e-8 |
| Cross-check via logs | mean of ln(1+r_t) = 0.00743171 per month; exp(0.00743171) - 1 = 0.745939 percent, identical to the geometric mean by construction |
| Equality condition | mu = g if and only if every r_t is identical. Any dispersion at all makes g strictly smaller |
- The approximation drag = sigma^2/2 uses the population variance with divisor T, not the sample variance with divisor T-1. On the base series the sample variance gives 0.021638 percent against a true drag of 0.020728 percent - a 4 percent error in the drag from choosing the wrong divisor.
- The relationship is exact only in the continuous limit. It is a second-order Taylor expansion of ln(1+r) and degrades as returns get large: at monthly returns of 20 percent the third-order term is no longer negligible.
- Reporting an arithmetic mean alongside a maximum drawdown is internally inconsistent, because the drawdown is computed on the compounded path and the arithmetic mean is not achievable on it.
Source: Standard result; the sigma^2/2 approximation follows from Ito's lemma applied to ln(1+r)
The three annualisations, and why two of them disagree
A monthly return can be annualised by multiplying by twelve, by compounding twelve times, or by exponentiating a mean log return. These give three different numbers. Which is correct depends on whether the quantity being annualised is additive in returns, in log returns, or in variance.
| Field | Value |
|---|---|
| Formula | simple: P*mu. compounded: (1+g)^P - 1. continuous: exp(P * mean(ln(1+r_t))) - 1. Variance annualises as P*sigma^2 and volatility as sqrt(P)*sigma |
| Worked, base series | simple 12 * 0.766667 percent = 9.2000 percent. compounded (1.00745939)^12 - 1 = 9.3278 percent. continuous exp(12 * 0.00743171) - 1 = 9.3278 percent |
| The counter-intuitive part | The geometric annualisation (9.3278 percent) is LARGER than the simple annualisation of the arithmetic mean (9.2000 percent), by 12.78 basis points, even though the monthly geometric mean is smaller than the monthly arithmetic mean |
| Why | Twelve-fold compounding of the smaller monthly figure contributes more than the monthly volatility drag subtracts. The two effects work in opposite directions and neither dominates by construction |
| Volatility | 2.080273 percent monthly times sqrt(12) = 7.2063 percent annual. Variance 0.00043275 monthly times 12 = 0.00519300 annual |
- The often-quoted identity g_annual = mu_annual - sigma_annual^2/2 does not hold across an annualisation boundary. On the base series it gives 9.2000 - 7.2063^2/2 = 8.9403 percent against a true 9.3278 percent, an error of 39 basis points. The identity applies to log returns at a single frequency, not to a simply annualised arithmetic mean.
- Volatility scales with sqrt(P) only for serially uncorrelated returns, exactly the same condition that Sharpe annualisation requires. If rho1 is nonzero, the annualised volatility is wrong too and in the opposite direction from the Sharpe error.
- A performance table that does not state which of the three rules produced its annual return is not reproducible. On this series the spread between the rules is 39 basis points on a 9 percent return; on a volatile series it is much larger.
Sharpe ratio
Mean excess return divided by the standard deviation of returns, both measured per period. It is a t-statistic in disguise: the same quantity, scaled by sqrt(T), tests the hypothesis that mean excess return is zero.
| Field | Value |
|---|---|
| Formula | SR = (mu - rf) / sigma, with sigma the sample standard deviation of returns (divisor T-1). Annualised SR_ann = SR * sqrt(P) only if returns are serially uncorrelated |
| Worked | Base series: mu = 0.766667 percent, rf = 0.200000 percent, so mean excess = 0.566667 percent. sigma = 2.080273 percent. SR = 0.566667/2.080273 = 0.272400 per month. Annualised at sqrt(12): 0.272400 * 3.464102 = 0.943622 |
| Denominator convention | With a constant risk-free rate the standard deviation of returns and of excess returns are identical (both 2.080273 percent here), so the convention is invisible. With a time-varying rf they differ and the excess-return standard deviation is the correct one |
| As a t-statistic | t = SR * sqrt(T) / sqrt(1 + SR^2/2) = 0.272400 * 4.898979 / 1.018446 = 1.310396, so this Sharpe ratio is not significantly different from zero at 24 observations |
| Scale invariance | Unlevering or levering a return series by a constant factor leaves SR unchanged. Adding a constant to every return does not |
- Subtracting rf from the numerator but not adjusting the denominator is correct; subtracting it twice, or comparing a Sharpe ratio computed on total returns with one computed on excess returns, is not. On a series with 5 percent rf and 8 percent volatility the difference is more than half a Sharpe point.
- The ratio is undefined as a ranking device across return distributions with different higher moments. Two strategies with identical mu and sigma but skewness of +1 and -1 receive the same Sharpe ratio and are not the same risk.
- Selling out-of-the-money options raises the Sharpe ratio of a short sample almost mechanically, because the premium enters the numerator every period and the tail enters the denominator only when it occurs. A high Sharpe ratio over a short sample is evidence about the sample, not the strategy.
Source: Sharpe 1966, revised in Sharpe 1994
Annualising a Sharpe ratio when returns are autocorrelated
Multiplying a per-period Sharpe ratio by sqrt(P) assumes the returns are independently and identically distributed. Under serial correlation the correct factor is Lo's eta(q), which is smaller than sqrt(q) for positive autocorrelation and larger for negative. Positive autocorrelation is the common case in illiquid or marked-to-model portfolios, so the naive factor usually overstates.
| Field | Value |
|---|---|
| Formula | SR(q) = eta(q) * SR, with eta(q) = q / sqrt( q + 2 * sum_{k=1}^{q-1} (q-k) * rho_k ). Under AR(1), rho_k = rho1^k. If all rho_k = 0 then eta(q) = sqrt(q) |
| Worked, AR(1) plug-in | Base series rho1 = 0.297578, q = 12. sum (q-k)rho1^k = 4.480618, so q + 2*sum = 20.961236 and eta(12) = 12/sqrt(20.961236) = 2.621035. Corrected annual Sharpe = 0.272400 * 2.621035 = 0.713970 against 0.943622 from sqrt(12) |
| Overstatement | sqrt(12)/eta(12) = 3.464102/2.621035 = 1.321654, so the naive factor overstates the annual Sharpe ratio by 32.17 percent on this series |
| Worked, full sample lags | Using the 11 estimated autocorrelations directly (0.2976, 0.0813, -0.0370, 0.0789, 0.1128, -0.0316, -0.0930, -0.0828, 0.0569, 0.0016, 0.0326): sum (q-k)rho_k = 4.393994, eta = 2.631934, SR_ann = 0.716939 |
| Sensitivity | rho1 = 0.10 gives eta = 3.160123 and an overstatement of 9.63 percent; rho1 = 0.50 gives eta = 2.121320 and 63.30 percent |
- Estimating 11 autocorrelations from 24 observations is not defensible: each rho_k beyond the first few is noise. The AR(1) plug-in, which spends one parameter, is the practical choice at short sample lengths and is what the worked value above uses. On this series the two approaches happen to agree to within 0.003 Sharpe points, which should not be read as evidence that they generally do.
- Positive autocorrelation in reported returns is the signature of smoothed or stale marks, not of skill. Getmansky, Lo and Makarov 2004 model exactly this, and the implication is that the correction should be applied before, not after, comparing an illiquid book to a liquid one.
- The correction changes the annualised Sharpe ratio and nothing else. The per-period Sharpe ratio is unaffected: the autocorrelation problem is entirely a problem of aggregation across periods.
- eta(q) can exceed sqrt(q). Negatively autocorrelated returns - mean-reverting or overhedged books - have their annual Sharpe ratio understated by the naive factor, which is why the correction should be applied symmetrically rather than only when it flatters.
Source: Lo 2002
Standard error of the Sharpe ratio
The Sharpe ratio is an estimate with a standard error that depends on the sample length and on the higher moments of the return distribution. Under independent normal returns the standard error has a closed form; under general distributions it picks up skewness and kurtosis terms.
| Field | Value |
|---|---|
| Formula | iid normal: SE(SR) = sqrt( (1 + SR^2/2) / T ). General: SE(SR) = sqrt( (1 - gamma3*SR + ((gamma4 - 1)/4)*SR^2) / T ), with gamma3 skewness and gamma4 kurtosis (3 for a normal) |
| Worked, iid | SR = 0.272400, T = 24. SE = sqrt((1 + 0.272400^2/2)/24) = sqrt(1.03710/24) = sqrt(0.04321) = 0.207876 per month; times sqrt(12) = 0.720104 annualised |
| Worked, non-normal | gamma3 = -0.508966, gamma4 = 2.368633. SE = sqrt((1 - (-0.508966)(0.272400) + ((2.368633-3)/4)(0.272400^2))/24) = sqrt((1 + 0.138642 - 0.011713)/24) = sqrt(0.046955) = 0.220230; annualised 0.762899 |
| 95 percent interval, annualised | iid: 0.943622 +/- 1.96 * 0.720104 = -0.467783 to 2.355027. Non-normal: -0.527639 to 2.414882. Both intervals contain zero |
| Sample length for significance | Months required for t = 1.96 at a given true annual Sharpe: SR_ann 0.5 needs 186.32 months (15.53 years); 1.0 needs 48.02 (4.00 years); 1.5 needs 22.41 (1.87 years); 2.0 needs 13.45 (1.12 years) |
- The interval on the base series spans nearly three Sharpe points. Two years of monthly data cannot distinguish a Sharpe ratio of 0.9 from one of zero, and no estimator improvement changes that; it is a sample-size limit.
- Negative skewness makes the standard error LARGER here because the -gamma3*SR term is positive when gamma3 is negative. That is the opposite of what many summaries claim, and the sign follows directly from the formula.
- This standard error assumes serially uncorrelated returns. If it is applied to autocorrelated returns it is too small, in addition to the annualisation error covered separately. Lo 2002 gives the autocorrelation-robust version.
- Comparing two Sharpe ratios requires the standard error of their difference, which includes the covariance of the two return series and is not the root sum of squares of the individual standard errors. Jobson and Korkie 1981, with the correction in Memmel 2003, gives the test.
Source: Lo 2002 for the iid case; Mertens 2002 for the non-normal correction
Small-sample bias in the Sharpe ratio
The plug-in Sharpe estimator is biased upward in small samples. The sample standard deviation is a downward-biased estimator of sigma, and dividing by a number that is too small on average makes the ratio too large on average. The bias is a pure function of T and can be removed exactly under normality.
| Field | Value |
|---|---|
| Formula | E[SR_hat] = SR * a(T), with a(T) = sqrt((T-1)/2) * Gamma((T-2)/2) / Gamma((T-1)/2). Unbiased estimator: SR_hat / a(T) |
| Worked | T = 24: a(24) = sqrt(11.5) * Gamma(11)/Gamma(11.5) = 1.03415598, an upward bias of 3.4156 percent. Corrected monthly SR = 0.272400/1.03415598 = 0.263403; annualised at sqrt(12) = 0.912456 against 0.943622 uncorrected |
| Bias by sample length | T = 12: a = 1.075315, bias 7.532 percent. T = 24: 1.034156, 3.416 percent. T = 36: 1.022086, 2.209 percent. T = 60: 1.012940, 1.294 percent. T = 120: 1.006358, 0.636 percent. T = 240: 1.003152, 0.315 percent |
| Related constant | c4(T) = sqrt(2/(T-1)) * Gamma(T/2)/Gamma((T-1)/2) = 0.98919267 at T = 24 is the bias factor of the sample standard deviation itself: E[s] = c4 * sigma |
- The bias is small relative to the standard error and is routinely mistaken for the whole small-sample problem. At T = 24 the bias is 3.4 percent of the Sharpe ratio while the standard error is 76 percent of it. Correcting the bias and reporting a point estimate without an interval fixes the smaller error and leaves the larger one.
- The bias is always upward, never downward, and it is largest exactly where track records are shortest. A twelve-month track record overstates its Sharpe ratio by 7.5 percent before any selection effect is considered.
- The correction assumes normal iid returns. Under fat tails the exact factor differs, and no closed form is available; the direction of the bias is unchanged.
Source: Miller and Gehr 1978
Sortino ratio, and the divisor that changes the answer
Excess return over a minimum acceptable return divided by downside deviation - the root mean square of shortfalls below that target. The ambiguity is the divisor in the downside deviation: the full sample length, or only the count of observations below the target. The two answers differ by a large factor and both appear in commercial reporting.
| Field | Value |
|---|---|
| Formula | Sortino = (mu - MAR) / DD, with DD = sqrt( (1/T) * sum_t min(r_t - MAR, 0)^2 ). The alternative uses 1/n_below in place of 1/T, where n_below is the count of periods with r_t < MAR |
| Worked, divisor T | MAR = 0.20 percent per month. 8 of 24 months fall below it. Sum of squared shortfalls = 0.00390400. DD = sqrt(0.00390400/24) = 1.275408 percent. Sortino = (0.766667 - 0.200000)/1.275408 = 0.444302 per month; annualised 1.539108 |
| Worked, divisor n_below | DD = sqrt(0.00390400/8) = 2.209072 percent. Sortino = 0.566667/2.209072 = 0.256518; annualised 0.888604 |
| Ratio of the two | sqrt(T/n_below) = sqrt(24/8) = 1.732051. The T-divisor version is 73.2 percent higher on this series, and the gap grows as the strategy has fewer losing periods |
| MAR conventions | MAR = 0 gives DD = 1.198263 percent (divisor T) and Sortino 0.639806. MAR = rf gives the value above. MAR = the arithmetic mean makes the Sortino ratio zero by construction |
- The divisor-T version is the one implied by the original semi-deviation definition and is the more defensible: it treats an upside month as a zero shortfall rather than dropping it, which keeps the statistic a comparable second moment. The divisor-n_below version is not a deviation of the return distribution at all; it is a conditional deviation of the losses.
- A strategy with no periods below the MAR has an infinite Sortino ratio under either divisor, and the divisor-n_below version is undefined. Any strategy whose losses are rare and large will therefore look better on Sortino than on Sharpe until the first large loss lands.
- Annualising a Sortino ratio by sqrt(P) inherits the same autocorrelation problem as the Sharpe ratio, with no published analogue of Lo's correction. Treat an annualised Sortino ratio on autocorrelated returns as unquantified.
- The MAR must be stated. Sortino ratios computed against 0, against the risk-free rate, and against an absolute return target are three different statistics and are not comparable.
Source: Sortino and van der Meer 1991
Calmar ratio
Compound annual return divided by maximum drawdown over the same window. It is the only common ratio whose denominator is a single realised extreme rather than a moment of the distribution, which makes it the least stable of the family.
| Field | Value |
|---|---|
| Formula | Calmar = g_ann / |MaxDD|, both measured over the same window. The original specification uses a 36-month window and a compound annual return |
| Worked | g_ann = 9.3278 percent, MaxDD = 7.5967 percent. Calmar = 0.093278/0.075967 = 1.227869 |
| With arithmetic annualisation | 12 * 0.766667 percent = 9.2000 percent gives 1.211046, 1.4 percent lower. The convention must be stated |
| Excess-return variant | Using return over the risk-free rate: (9.3278 - 2.4266) percent / 7.5967 percent = 0.908445. Sometimes called the Sterling-Calmar ratio |
| Window dependence | Over the first 12 months alone the max drawdown is 4.0820 percent; over the second 12 months alone it is 7.3105 percent. Calmar computed on either half is not comparable with Calmar computed on the whole |
- Maximum drawdown is non-decreasing in the length of the window, so Calmar mechanically falls as a track record lengthens even if nothing about the strategy changes. Comparing the Calmar ratios of a two-year record and a ten-year record compares window lengths, not strategies.
- The denominator is one number from one path. Its sampling variability is enormous and there is no accepted standard error for it, which means Calmar has no interval and cannot support a significance statement.
- The original 36-month convention exists precisely to make the statistic comparable across managers. A Calmar ratio computed over whatever window happens to flatter the record is not the same statistic.
Source: Young 1991
Sterling ratio, and its incompatible definitions
A return-over-drawdown ratio whose denominator is the average of the annual maximum drawdowns rather than the single worst drawdown. At least three definitions circulate, differing in whether a fixed offset is added to the denominator and whether the average is over annual maxima or over all drawdowns.
| Field | Value |
|---|---|
| Formula | Sterling = g_ann / (average annual maximum drawdown + 0.10) in the form with the offset; g_ann / (average annual maximum drawdown) without it; and g_ann / (average of all drawdowns) in a third variant |
| Worked inputs | Base series year 1 maximum drawdown 4.0820 percent, year 2 maximum drawdown 7.3105 percent. Average annual maximum drawdown = 5.6963 percent |
| Worked, with the 10-point offset | 0.093278 / (0.056963 + 0.10) = 0.093278/0.156963 = 0.594269 |
| Worked, without the offset | 0.093278 / 0.056963 = 1.637532 |
| Spread between definitions | 1.637532 versus 0.594269 - a factor of 2.755 on identical data. No other statistic in this corpus has a wider definitional spread |
- The 10 percentage-point offset in the original form has no statistical justification; it is a floor that prevents the ratio exploding for strategies with small drawdowns. Its practical effect is to compress differences between low-drawdown strategies almost to nothing, which is either the point or the flaw depending on the use.
- A Sterling ratio without a stated definition is uninterpretable. Given the factor of 2.755 spread on this series, comparing two Sterling ratios from two sources is comparing two unrelated numbers.
- The averaging over annual maxima makes the statistic depend on where the calendar year boundaries fall. Shifting the year boundary by six months changes the annual maxima and therefore the ratio, on identical returns.
Omega ratio
The ratio of the probability-weighted gains above a threshold to the probability-weighted losses below it. Unlike the Sharpe ratio it uses the entire return distribution rather than its first two moments, and it is a function of the threshold rather than a single number.
| Field | Value |
|---|---|
| Formula | Omega(tau) = sum_t max(r_t - tau, 0) / sum_t max(tau - r_t, 0). Equivalently, the ratio of the integral of the survival function above tau to the integral of the cumulative distribution below tau |
| Worked, tau = 0 | Base series: gains above zero sum to 32.0000 percent, shortfalls below zero sum to 13.6000 percent. Omega(0) = 0.320000/0.136000 = 2.352941 |
| Worked, tau = rf | At tau = 0.20 percent: gains 28.8000 percent, losses 15.2000 percent, Omega = 1.894737 |
| Identity | (gains - losses)/T = mu - tau exactly. Check: (0.320000 - 0.136000)/24 = 0.00766667 = mu - 0 . And (0.288000 - 0.152000)/24 = 0.00566667 = mu - rf |
| Fixed point | Omega(tau) = 1 exactly when tau = mu. Omega is monotonically decreasing in tau, crossing 1 at the arithmetic mean and tending to 0 as tau rises above the maximum return |
- Because Omega equals 1 at the mean and is monotone, Omega at a threshold below the mean carries the same ordinal information as the mean itself for a fixed distribution shape. It only adds information when comparing distributions of different shape at a threshold that matters to the user.
- Omega uses every observation, which makes it more stable than a drawdown-based ratio, but it is still a ratio of two sample sums and inherits their sampling error. No standard closed-form standard error exists.
- The whole Omega curve, not one point on it, is the statistic. Two strategies can cross: one better below the threshold and worse above it. Reporting Omega at a single threshold hides exactly the case Omega was designed to reveal.
Source: Keating and Shadwick 2002
The benchmark regression, and its standard errors
Beta, alpha, R-squared, tracking error and residual volatility all come from a single regression of excess portfolio returns on excess benchmark returns. Quoting any of them without the regression standard errors treats an estimate as a measurement.
| Field | Value |
|---|---|
| Formula | r_t - rf = alpha + beta*(b_t - rf) + e_t. beta = Cov(r,b)/Var(b); alpha = mean(r - rf) - beta*mean(b - rf); R^2 = corr(r,b)^2; residual sd s_e = sqrt(SSE/(T-2)) |
| Worked, point estimates | Base series: Cov(r,b) = 0.0004227391, Var(b) = 0.0005636522, so beta = 0.750000 exactly. corr = 0.855946, R-squared = 0.732644. alpha = 0.266667 percent per month = 3.2000 percent annualised |
| Worked, standard errors | Residual sd s_e = 1.099811 percent. SE(beta) = 0.09659361, t(beta) = 7.764489. SE(alpha) = 0.00227799, t(alpha) = 1.170625 |
| Testing beta = 1 | t = (0.750000 - 1)/0.09659361 = -2.588163, so beta is significantly below 1 at 24 observations even though alpha is not significantly above 0 |
| Alpha interval | 95 percent interval on annualised alpha: -2.1578 percent to 8.5578 percent |
- Beta is estimated far more precisely than alpha, always. That is a structural property of the regression, not a feature of this dataset: alpha is the intercept and absorbs all the residual noise. A track record can establish its market exposure long before it can establish its skill.
- The regression must be run on excess returns on both sides. Running it on total returns produces an intercept that mixes alpha with (1 - beta) times the risk-free rate, which is not Jensen's alpha and is biased whenever beta is far from 1.
- R-squared of 0.73 means 27 percent of the variance is unexplained by the benchmark. It says nothing about whether the residual is skill; it is the same residual that alpha's standard error is computed from, and here it is not significantly nonzero.
- Ordinary least squares standard errors assume homoskedastic uncorrelated residuals. On real return data neither holds, and the honest version uses Newey-West or a comparable robust estimator, which widens the alpha interval further.
Jensen's alpha
The intercept of the excess-return regression: the average return not explained by the portfolio's exposure to the benchmark. It is a return, in percent, not a ratio, and it has a standard error.
| Field | Value |
|---|---|
| Formula | alpha = (mu - rf) - beta * (mu_b - rf). Annualised as P * alpha |
| Worked | mu - rf = 0.566667 percent, mu_b - rf = 0.400000 percent, beta = 0.750000. alpha = 0.566667 - 0.750000 * 0.400000 = 0.566667 - 0.300000 = 0.266667 percent per month. Annualised 12 * 0.266667 = 3.2000 percent |
| Cross-check | The OLS intercept computed directly from the regression is 0.00266667, identical. The closed form and the regression are the same calculation |
| Significance | t(alpha) = 1.170625 at T = 24. Not significant. 95 percent interval on annualised alpha: -2.1578 percent to 8.5578 percent |
| Appraisal ratio | alpha divided by residual standard deviation: 0.00266667/0.01099811 = 0.242466 per month, 0.839927 annualised. This is the Sharpe-like scaling of alpha |
- Alpha is defined only relative to the specified benchmark and the specified beta. Adding a second factor to the regression changes both, and an alpha that survives one factor and not two was never alpha; it was unmodelled exposure.
- The annualisation multiplies alpha by P, not by sqrt(P), because alpha is a mean return rather than a ratio. Applying sqrt(P) to alpha is a common and large error: on this series it would give 0.9238 percent instead of 3.2000 percent.
- A positive alpha with an insignificant t-statistic is the normal state of a short track record and is not evidence of skill. Here the point estimate is 3.2 percent per year and the interval includes -2.2 percent.
Source: Jensen 1968; the appraisal ratio is Treynor and Black 1973
Information ratio, and the two things it is called
Mean active return divided by tracking error. The ambiguity is what counts as active return: the arithmetic difference from the benchmark, or the regression residual. The first gives the information ratio proper, the second the appraisal ratio. They coincide only when beta equals 1.
| Field | Value |
|---|---|
| Formula | IR = mean(r_t - b_t) / sd(r_t - b_t), annualised by sqrt(P). Appraisal ratio = alpha / s_e, also annualised by sqrt(P) |
| Worked, information ratio | Base series: mean active return = 0.166667 percent per month, tracking error = 1.228526 percent. IR = 0.135664 per month, 0.469954 annualised. Annualised tracking error 4.2557 percent, annualised active return 2.0000 percent |
| Worked, appraisal ratio | alpha = 0.266667 percent, residual sd = 1.099811 percent. Ratio = 0.242466 per month, 0.839927 annualised |
| Gap | 0.839927 versus 0.469954 - the appraisal ratio is 79 percent higher on identical data, because beta of 0.750000 means the arithmetic difference r - b carries a large residual benchmark exposure that the regression removes |
| When they agree | Only when beta = 1. At beta = 1 the regression residual and the arithmetic active return are the same series |
- A portfolio with beta well away from 1 has a tracking error dominated by the beta mismatch rather than by security selection. Its information ratio is then mostly a statement about its market exposure, which is why the appraisal ratio is the right number for a stock-selection mandate and the information ratio the right one for a tracking mandate.
- The two are routinely reported under the same label. When a document quotes an information ratio, check whether the denominator is the standard deviation of r - b or the residual standard deviation of a regression; on this series that choice moves the number by 0.37.
- Annualising by sqrt(P) assumes serially uncorrelated active returns. Active returns of a portfolio rebalanced monthly against a benchmark rebalanced quarterly are autocorrelated by construction.
Source: Grinold and Kahn, Active Portfolio Management; the appraisal ratio is Treynor and Black 1973
Treynor ratio
Mean excess return divided by beta rather than by volatility. It prices only systematic risk, which makes it the right ratio for a component of a diversified portfolio and the wrong one for a standalone allocation.
| Field | Value |
|---|---|
| Formula | Treynor = (mu - rf) / beta, conventionally stated in annualised return units |
| Worked | Annual excess return 6.8000 percent, beta 0.750000. Treynor = 0.068000/0.750000 = 0.090667 |
| Benchmark comparison | The benchmark's own Treynor ratio is its annual excess return over a beta of 1: 0.072000 - 0.024000 = 0.048000. The strategy's 0.090667 is higher, which is the sign that alpha is positive |
| Relationship to alpha | Treynor exceeds the benchmark's Treynor exactly when Jensen's alpha is positive: alpha = beta * (Treynor_p - Treynor_b) = 0.750000 * (0.090667 - 0.048000) = 0.032000, matching the annualised alpha |
| Units | Not dimensionless. Treynor is a return per unit of beta, so its magnitude depends on the annualisation of the numerator |
- Treynor ignores idiosyncratic risk entirely. A single-stock portfolio and a thousand-stock portfolio with the same beta and the same excess return receive the same Treynor ratio, which is correct if and only if the holder is diversified elsewhere.
- The ratio is unstable and sign-flipping for betas near zero, and meaningless for negative betas: a short-biased fund with negative beta and positive excess return gets a negative Treynor ratio that ranks it below a losing long fund.
- Beta carries a standard error - 0.09659361 on this series - so Treynor carries one too. The ratio is more precisely estimated than the Sharpe ratio only because beta is estimated more precisely than volatility, not because the numerator is better known.
Source: Treynor 1965
M-squared
The return the portfolio would have earned if it had been levered or delevered to exactly the benchmark's volatility. It converts a Sharpe ratio into a return, which makes it directly comparable with the benchmark's return in percentage points rather than in ratio units.
| Field | Value |
|---|---|
| Formula | M2 = rf_ann + SR_ann * sigma_bench_ann. M2 alpha = M2 - mu_bench_ann. The implied leverage is sigma_bench/sigma_portfolio |
| Worked | SR_ann = 0.943622, sigma_bench_ann = 8.2242 percent, rf_ann = 2.4000 percent (12 times 0.20 percent). M2 = 0.024000 + 0.943622 * 0.082242 = 0.024000 + 0.077606 = 10.1606 percent |
| M-squared alpha | Benchmark annual arithmetic return 7.2000 percent. M2 alpha = 10.1606 - 7.2000 = 2.9606 percent |
| Implied leverage | sigma_bench/sigma_port = 8.2242/7.2063 = 1.141262, so the comparison is against the strategy run at 114.1 percent of its observed size |
| Ordering | M2 ranks identically to the Sharpe ratio for a fixed benchmark, since it is an affine transformation of it. Its value is interpretability, not a different ordering |
- M-squared assumes leverage is available at the risk-free rate and that levering the strategy scales its return linearly. Neither holds for a strategy whose capacity is limited or whose financing spread widens with size, and the M-squared alpha is overstated by exactly the financing cost that was ignored.
- Because M2 is affine in the Sharpe ratio, it inherits the Sharpe ratio's entire standard error. The 2.96 percent M-squared alpha on this series is not distinguishable from zero at 24 observations.
- The benchmark volatility must be measured over the same window as the strategy volatility. Using a long-run benchmark volatility against a short-window strategy volatility mixes two sample periods into one number.
Source: Modigliani and Modigliani 1997
Up and down capture, and the two ways to compute them
The fraction of the benchmark's up-period return the portfolio captured, and the fraction of its down-period loss the portfolio suffered. Computed either as a ratio of arithmetic means over the selected periods or as a ratio of compounded returns over them, giving different answers.
| Field | Value |
|---|---|
| Formula | Mean-based: UpCapture = mean(r_t | b_t > 0) / mean(b_t | b_t > 0). Compound-based: UpCapture = (prod over up periods of (1+r_t) - 1) / (prod over up periods of (1+b_t) - 1). Down capture uses b_t < 0 |
| Worked, mean-based | 16 up months: mean strategy 1.831250 percent, mean benchmark 2.012500 percent, up capture 0.909938. 8 down months: mean strategy -1.362500 percent, mean benchmark -2.225000 percent, down capture 0.612360 |
| Worked, compound-based | Up months compounded: strategy +33.508928 percent, benchmark +37.397292 percent, ratio 0.896026. Down months: strategy -10.473647 percent, benchmark -16.522009 percent, ratio 0.633921 |
| Capture ratio | Up capture divided by down capture: 0.909938/0.612360 = 1.485954 mean-based, 0.896026/0.633921 = 1.413306 compound-based |
| Period classification | Periods are classified by the sign of the BENCHMARK return, not the strategy return. 16 up and 8 down months here, with no zero-return months |
- Capture ratios are not independent of beta. A portfolio with beta 0.75 and zero alpha has an up capture of about 0.75 and a down capture of about 0.75 mechanically. The informative quantity is the gap between the two, not either level: here 0.910 against 0.612 on a beta of 0.75 is what asymmetry looks like.
- The statistic is entirely determined by the sign classification, and the sign of a benchmark return in a period near zero is noise. Reclassifying one marginal month can move a capture ratio by several points on a two-year sample.
- Down capture below 1 over a sample containing few down periods is the least reliable number in a performance table. Eight down months is not a sample from which a downside characteristic can be inferred.
- Compounding the selected periods, as the compound-based version does, treats non-adjacent months as if they were consecutive. It is internally consistent but it is not a return anyone earned.