Library / Backtest overfitting
Evaluating Trading Strategies: What Threshold Should a Reviewer Demand
What performance threshold must a strategy clear before its apparent edge is credible?
Disclose the full selection process and the number of variants examined, then evaluate the frozen winner on untouched data against a declared benchmark. Harvey and Liu study thresholds for selected strategies; their numerical cutoff does not apply automatically to a separately proposed lever or one sealed test.
Harvey and Liu?s numerical threshold addresses selected strategies; a new lever or sealed test uses its own declared evaluation.

Evidence map
| Aspect | Finding |
|---|---|
| What it is | Harvey and Liu's short statement of the reviewer's problem, as distinct from the researcher's. |
| Key result / formula | The argument turns on a quantity hidden from the reviewer: how many candidate strategies were examined before this one was presented. |
| Why it matters for backtesting | In Stochastly the researcher and the reviewer are the same person, which is exactly the situation the paper warns about. |
What it is
A fund, or a user, is presented with a strategy and a track record and must set the level of evidence to demand, knowing that the strategy was selected from a pool whose size they cannot observe. It complements [Harvey–Liu Haircut Sharpe & the Cross-Section of Expected Returns], which gives the machinery; this gives the decision rule.
Key result / formula
A researcher who tried a handful and a researcher who tried hundreds produce the same-looking winner, and the appropriate discount differs enormously between them. Harvey and Liu show that once a plausible number of trials is assumed, the conventional threshold of two standard errors lets through far too many false positives, and they discuss a threshold nearer three within their selected-strategy setting. They also make the point that matters most in practice: the reviewer should require the number of trials to be disclosed, because the selection context of the presented winner cannot otherwise be assessed. Where the number is unavailable, the reviewer's options are to assume a large one, which is conservative, or to demand evidence that does not depend on it, such as performance on data the researcher could not have seen.
Why it matters for backtesting
The discipline it implies is procedural rather than statistical: record the number of variants tried as they are tried, including the ones abandoned after a glance, because a count reconstructed afterwards is systematically too low. Report the search count and selection process alongside the result. That count describes selection, but it does not set an automatic penalty for a newly added lever or a predeclared sealed test. The second implication is about what to do with an inherited strategy, one taken from a book, a forum or a paper: the trials behind it were run by someone else and are unknowable, so independently dated, untouched evidence is especially valuable. If the original selection process is unknown, state that limitation rather than inventing an adjusted threshold.
Source
Harvey & Liu, "Evaluating Trading Strategies", Journal of Portfolio Management 40(5), 2014, 108-118. Primary source