Library / Backtest overfitting
Ranking Noise and the Null of the Maximum
Is the top strategy distinguishable from the maximum score generated by noise?
Freeze the candidate population, resample the null data in blocks that preserve relevant dependence, rerun every candidate and record the maximum standardized score on each draw. Compare the observed leader with that maximum distribution; any conclusion applies to the declared population and resampling assumptions.
The maximum-null comparison applies only to a frozen candidate population and its dependence-preserving null.
Evidence map
| Aspect | Finding |
|---|---|
| What it is | A search-level comparison for the best result selected from a predeclared population of configurations. |
| Key result / formula | For N independent standard-normal null statistics, the leading scale of their maximum is roughly sqrt(2 ln N). |
| Why it matters for backtesting | Freeze the candidate family, return definition, costs and ranking rule before drawing the search-level null. |
What it is
The statistic of the winner is compared with the largest statistic that the same search would produce under a suitable zero-edge null. This question arises only after a winner has been chosen from that population; a separately proposed lever and one sealed test have a different evaluation design.
Key result / formula
It is an asymptotic scale, not an exact expectation or a universal cutoff. Dependence changes the distribution: if all candidates are identical, copying one candidate adds no selection opportunity and no extra expected maximum. White's reality check and related resampling methods rerun the complete candidate selection on each null draw, preserving relevant dependence where their assumptions allow. Record the best statistic from each draw and compare the observed winner with that distribution. Standardizing candidates may be needed when their statistic variances differ, but the statistic and normalization must be specified before reading the result.
Why it matters for backtesting
A failure to reject means this test has not established predictive superiority of the selected winner over its benchmark. It does not prove every candidate has zero edge, erase descriptive differences in the sample or require rejecting an independently designed future hypothesis. Report how many candidates were considered and how strongly their returns depend on one another. If the family was expanded after seeing results, the added search belongs in the history and a new confirmatory batch must remain untouched. A placebo candidate and a positive control help diagnose whether the validation can detect an effect of practical size. Compare any chosen strategy with later data using the same frozen rule; the null of the maximum alone does not estimate its live return.
Source
White, "A Reality Check for Data Snooping", Econometrica 68(5), 2000; Romano & Wolf, "Stepwise Multiple Testing as Formalized Data Snooping", Econometrica 73(4), 2005; Harvey, Liu & Zhu, "...and the Cross-Section of Expected Returns", Review of Financial Studies 29(1), 2016; Harvey, "Presidential Address: The Scientific Outlook in Financial Economics", Journal of Finance 72(4), 2017. Primary source