Library / Backtest overfitting
Peeking at Accumulating Data Inflates the Error Rate
Why does repeatedly checking an accumulating backtest inflate false discoveries?
Choose the sample length and decision rule before viewing outcomes, or use a sequential test with boundaries calibrated to repeated looks. Applying an ordinary fixed-sample significance test whenever new data arrive raises false-positive risk under the repeated-testing design studied by Armitage and colleagues.
Repeated peeking need not drive error toward certainty without stronger conditions; drawdown stop estimates remain uncertain.
Evidence map
| Aspect | Finding |
|---|---|
| What it is | The statistical result that examining a test repeatedly as data accumulates, and stopping when it becomes significant, produces significance far more often than the nominal rate. |
| Key result / formula | Wald's sequential analysis establishes that a test applied to accumulating data is a different object from a test applied once at a fixed sample size, and that a valid sequential procedure must specify its stopping rule in advance, with boundaries computed so that the error rates are controlled across the entire procedure. |
| Why it matters for backtesting | The user's version has three forms, and all three are the same error. |
What it is
It is the formal version of a habit every backtester has: watching the equity curve and deciding when to stop.
Key result / formula
Armitage, McPherson and Rowe made the cost of ignoring this explicit for the clinical setting: if a conventional test is applied repeatedly to data as it arrives and the experiment is stopped at the first significant result, the probability of a false positive is much higher than the nominal level, and it grows toward certainty as the number of looks increases, because a random walk of the test statistic will eventually cross any fixed boundary. The remedy is either a pre-specified fixed sample size, or a sequential design whose boundaries account for the looking.
Why it matters for backtesting
Extending a backtest's period until the result improves is peeking. Watching a live strategy and deciding to abandon or to keep it based on where the curve is now, with no rule fixed in advance, is peeking. Re-running a backtest after each data update and treating the first good result as the answer is peeking. The defence is procedural and cheap: fix the evaluation length and the decision rule before the test, in terms of a number of observations or episodes, and record them with the verdict. For a live strategy, the equivalent is a pre-declared drawdown or duration of underperformance at which it is stopped, chosen from the backtest's own distribution, which removes the decision from the moment it would otherwise be taken.
Source
Wald, "Sequential Tests of Statistical Hypotheses", Annals of Mathematical Statistics 16(2), 1945, 117-186; Armitage, McPherson & Rowe, "Repeated Significance Tests on Accumulating Data", Journal of the Royal Statistical Society, Series A 132(2), 1969, 235-244. Primary source