Library / Backtest overfitting

Deflation vs Selection Bias

Can adjusting a performance statistic repair the act of selecting the best strategy?

For a winner chosen from a defined configuration population, replay that selection under a dependence-preserving null and compare its maximum with the observed winner. Estimate prospective performance on untouched data separately. A corrected test statistic cannot supply the winner's expected live effect size.

A selection-aware statistic addresses the chosen winner; it does not estimate the strategy?s future effect size.

Evidence map

AspectFinding
What it isSelection of the best observed configuration biases its in-sample estimate upward.
Key result / formulaSelection bias: E[max of N noise draws] > 0 and grows with N (the False Strategy Theorem), so the winning backtest's Sharpe overstates true skill.
Why it matters for backtestingThe single most common quant self-deception is judging a selected strategy against an unconditional benchmark.

What it is

Several methods address different statistical questions created by that selection.

Key result / formula

A null of the maximum and the Deflated Sharpe Ratio address a selected winner across a population of configurations under their assumptions. Harvey-Liu adjustments and family-wise or false-discovery p-value procedures have different targets and are not interchangeable versions of one deflation. Key distinction: more decimal places, a longer sample, or a prettier equity curve do not remove selection bias — a selection-aware analysis needs the search design and dependence among candidates; it does not remove all estimation bias.

Why it matters for backtesting

When choosing one winner from a population of configurations, compare that winner with a selection-aware benchmark and report the search. For a new lever evaluated through a predeclared paired test and one sealed batch, the trial count is reported but does not automatically tighten the decision threshold.

Worked distinction

If 50 correlated strategies are searched and only the largest Sharpe is published, two questions arise. First, how surprising is that maximum under a zero-edge search? A null-of-the-maximum or deflated statistic addresses that question. Second, how much of the winner's reported performance will survive a fresh sample? That requires an out-of-sample estimate, shrinkage model or prospective observation. The first answer does not supply the second: a significant winner can still have a badly inflated point estimate. Write down both the rejection threshold and the expected live effect separately. Reusing the same holdout to estimate shrinkage and choose the winner recreates selection bias. The search log and an untouched evaluation period are distinct pieces of evidence.

Source

Bailey & López de Prado (2014, Deflated Sharpe Ratio, SSRN 2460551); Harvey & Liu (2015, SSRN 2345489); Lo & MacKinlay (RFS, 1990). Primary source

Related in the library