Library / Backtest overfitting

Minimum Backtest Length (MinBTL) & Minimum Track Record Length (MinTRL)

How long must a strategy run before its track record can support a performance claim?

State the benchmark Sharpe, sampling frequency, confidence level and purpose first. MinTRL evaluates a separately observed record using its estimated moments; the rough 2 ln N / S squared scale concerns choosing among N independent null configurations. Neither calculation certifies future trading skill.

The 7.61-year scale is conservative and is not a necessary minimum backtest length.

USDJPY rolling Sharpe varies as the 41-period estimation window moves.
USDJPY rolling Sharpe varies as the 41-period estimation window moves.

Evidence map

AspectFinding
What it isTwo different sample-length questions. MinBTL asks how much history a search over a stated population of configurations needs before its best in-sample score can be distinguished from a noise maximum under a simplified null.
Key result / formulaFor N independent zero-edge candidates, Gaussian null scores and a target annualized Sharpe S, a rough maximum-score scale after T years is sqrt(2 ln N / T).
Worked scale checkWith N = 45 independent null configurations and T = 5 years, sqrt(2 ln 45 / 5) is about 1.23.

What it is

MinTRL asks how long a separately observed record needs before its estimated Sharpe exceeds a specified benchmark at a chosen confidence level. Neither number is a general certificate of trading skill.

Key result / formula

Rearranging gives a conservative asymptotic scale T about 2 ln N / S squared. The source derives a more specific MinBTL from the expected maximum of N null scores; 2 ln N is an upper-bound approximation for that expected-maximum term, not the exact minimum length. This is a population-selection approximation, not a significance threshold for each new research lever and not a rule for a single predeclared reading of a sealed lot. Dependence among configurations changes the maximum's distribution, so a replay of the full selection on an appropriate null is preferable. MinTRL is a different formula based on the probabilistic Sharpe ratio, the benchmark Sharpe, skewness, kurtosis and a confidence level. Its observation count uses the return sampling convention and assumptions of that calculation. Trial count is not an input to MinTRL for a single predeclared record.

Worked scale check

That is an approximate scale of the best noise Sharpe, not a probability and not a guarantee that a Sharpe of 1 will appear. For a target S = 1, the conservative 2 ln 45 scale is about 7.61 years under these assumptions. The expected-maximum calculation in the source is smaller, around five years for this illustration, so 7.61 years is not a necessary minimum. Nearby configurations are correlated, so 45 raw trials do not usually mean 45 independent trials. Estimate the null maximum from the actual search structure when possible, and show the number of attempted configurations as context.

Why it matters for backtesting

The selection calculation applies when choosing a winner among a declared population of configurations. A researcher adding a new, separately motivated lever should compare paired gain, placebo and a positive control in development, then read one predeclared sealed lot once; the maximum-score threshold is not automatically raised because more ideas were explored. MinTRL asks whether a record clears its own benchmark with a stated uncertainty calculation. Neither length fixes look-ahead, missing costs or a changed live regime. Report the assumptions, sample dates, dependence structure and the decision that each calculation actually supports.

Source

MinBTL: Bailey, Borwein, Lopez de Prado & Zhu, "Pseudo-Mathematics and Financial Charlatanism", Notices of the AMS 61(5), 2014, 458-471. MinTRL: Bailey & Lopez de Prado, Journal of Risk 15(2), 2012, SSRN 1821643. Primary source

Related in the library