Library / Backtest overfitting
Statistical Overfitting: Why More Searching Gives a Worse Strategy
Why does repeated strategy tuning produce high backtest scores without live skill?
Record each tuned configuration and repeat the selection process inside a chronology-preserving validation design. High in-sample scores become less persuasive when many correlated candidates were searched; an untouched evaluation estimates what survives the selection.
Coarser grids may help under a given design but are not a universal rule.

Evidence map
| Aspect | Finding |
|---|---|
| What it is | The argument, developed by Bailey and co-authors, that beyond a point the intensity of a backtest search is negatively related to the quality of the strategy it produces. |
| Key result / formula | Consider a family of strategy configurations, all with the same true expected return, differing solely in parameters that do not matter. |
| Why it matters for backtesting | This inverts the intuition a user brings, that more testing is more diligence. |
What it is
Not merely uninformative: actively harmful, because the search optimises the noise.
Key result / formula
The backtest ranks them by a performance statistic computed on one finite sample, and the winner's statistic is the maximum of a set of noisy draws, so its expected value rises with the number of configurations examined even when no configuration is better than any other. That is the false strategy result. The extension in this paper is what happens out of sample: the winner's advantage was noise, so it does not persist, and because the selected configuration is the one that best fitted the sample's idiosyncrasies, its subsequent performance is worse than that of a configuration chosen at random from the family. The authors formalise the relationship between the number of trials, the length of the sample and the expected degradation, which yields both a minimum sample length below which a search is meaningless and the observation that adding trials without adding data makes matters worse rather than better.
Why it matters for backtesting
The operational rules that follow are concrete. Fix the family of configurations before searching and count its size, including the configurations examined and discarded. A coarse parameter grid is better than a fine one, since adjacent values add trials without adding information. Check the sample is long enough to support the number of trials contemplated (see [Minimum Backtest Length (MinBTL) & Minimum Track Record Length (MinTRL)]). And when a search has already been run at length, treat the selected parameters as suspect and test whether a central or rounded value in the same family performs comparably, since a parameter that must be exact is a parameter fitted to noise.
Source
Bailey, Ger, Lopez de Prado, Sim & Zhu, "Statistical Overfitting and Backtest Performance", in Jurczenko (ed.), Risk-Based and Factor Investing, ISTE Press and Elsevier, 2015, 449-461. Primary source
Related in the library
Backtest versus Live on a Large Cohort of Algorithms
The Discount Between Backtest and Live, Measured at the Vendors
Guide on this topic: Overfitting in trading, explained
How the assistant cites the library · The checks behind a verdict