Library / Backtest overfitting
False Strategy Theorem (Expected Maximum Sharpe under N trials)
How likely is a convincing strategy result when many false candidates are tested?
If one winner is selected from N independent null candidates, the leading scale of the largest standardized statistic is approximately sqrt(2 log N). For 100 candidates it is about 3.03. Correlated trading rules require a null replay of the actual search; the approximation is not a threshold for a new sealed test.
The expected null maximum describes selection from a defined configuration population, not every new research lever.
Evidence map
| Aspect | Finding |
|---|---|
| What it is | The result that, with enough trials, any target Sharpe ratio is achievable by chance — even from a strategy with zero true edge. |
| Key result / formula | For N independent strategies each with true Sharpe 0 and cross-trial Sharpe variance V, the expected maximum in-sample Sharpe is E[max SR_N] ≈ √V · [ (1−γ)·Z⁻¹(1 − 1/N) + γ·Z⁻¹(1 − 1/(N·e)) ], with γ ≈ 0.5772 (Euler–Mascheroni), Z⁻¹ the inverse normal CDF. |
| Why it matters for backtesting | This E[max] is precisely the SR₀ benchmark inside the Deflated Sharpe Ratio and the driver of Minimum Backtest Length. |
What it is
The maximum Sharpe among N trials is therefore an upward-biased statistic, not evidence of skill.
Key result / formula
It rises without bound, asymptotically like √(2·ln N)·√V. This expected maximum describes selection among a defined population of configurations. It is not a significance threshold for a separately proposed lever or a single predeclared test on a sealed lot.
Why it matters for backtesting
Comparing the selected maximum with zero ignores the selection step and increases false discoveries. The approximation describes that selection problem under its assumptions; it is not a universal haircut that grows whenever another independent idea is tried.
Worked null
With N independent standard-normal test statistics, the largest value grows roughly as sqrt(2 log N). For N = 100, this leading term is about 3.03 even though every candidate is null; finite-sample corrections change the exact expectation. This is why the maximum backtest score is a poor standalone estimate of skill. The independence assumption is rarely credible for nearby parameters or overlapping assets. Estimate the maximum directly by replaying the whole search on block-resampled or otherwise valid null data. Preserve trading costs in each replay and compare maxima measured in the same units. The theorem describes selection pressure, not a universal haircut to subtract from any Sharpe. A new strategy family or a single predeclared sealed test calls for its own validation protocol rather than an automatic increase in this maximum-based hurdle.
Source
Bailey & López de Prado (2014, Deflated Sharpe Ratio, SSRN 2460551); López de Prado, Advances in Financial Machine Learning, Wiley 2018; López de Prado & Lewis, "Detection of False Investment Strategies Using Unsupervised Learning Methods", Quantitative Finance, 2019. Primary source