Library / Backtest overfitting

False Strategy Theorem (Expected Maximum Sharpe under N trials)

How likely is a convincing strategy result when many false candidates are tested?

If one winner is selected from N independent null candidates, the leading scale of the largest standardized statistic is approximately sqrt(2 log N). For 100 candidates it is about 3.03. Correlated trading rules require a null replay of the actual search; the approximation is not a threshold for a new sealed test.

The expected null maximum describes selection from a defined configuration population, not every new research lever.

Evidence map

AspectFinding
What it isThe result that, with enough trials, any target Sharpe ratio is achievable by chance — even from a strategy with zero true edge.
Key result / formulaFor N independent strategies each with true Sharpe 0 and cross-trial Sharpe variance V, the expected maximum in-sample Sharpe is E[max SR_N] ≈ √V · [ (1−γ)·Z⁻¹(1 − 1/N) + γ·Z⁻¹(1 − 1/(N·e)) ], with γ ≈ 0.5772 (Euler–Mascheroni), Z⁻¹ the inverse normal CDF.
Why it matters for backtestingThis E[max] is precisely the SR₀ benchmark inside the Deflated Sharpe Ratio and the driver of Minimum Backtest Length.

What it is

The maximum Sharpe among N trials is therefore an upward-biased statistic, not evidence of skill.

Key result / formula

It rises without bound, asymptotically like √(2·ln N)·√V. This expected maximum describes selection among a defined population of configurations. It is not a significance threshold for a separately proposed lever or a single predeclared test on a sealed lot.

Why it matters for backtesting

Comparing the selected maximum with zero ignores the selection step and increases false discoveries. The approximation describes that selection problem under its assumptions; it is not a universal haircut that grows whenever another independent idea is tried.

Worked null

With N independent standard-normal test statistics, the largest value grows roughly as sqrt(2 log N). For N = 100, this leading term is about 3.03 even though every candidate is null; finite-sample corrections change the exact expectation. This is why the maximum backtest score is a poor standalone estimate of skill. The independence assumption is rarely credible for nearby parameters or overlapping assets. Estimate the maximum directly by replaying the whole search on block-resampled or otherwise valid null data. Preserve trading costs in each replay and compare maxima measured in the same units. The theorem describes selection pressure, not a universal haircut to subtract from any Sharpe. A new strategy family or a single predeclared sealed test calls for its own validation protocol rather than an automatic increase in this maximum-based hurdle.

Source

Bailey & López de Prado (2014, Deflated Sharpe Ratio, SSRN 2460551); López de Prado, Advances in Financial Machine Learning, Wiley 2018; López de Prado & Lewis, "Detection of False Investment Strategies Using Unsupervised Learning Methods", Quantitative Finance, 2019. Primary source

Related in the library