Library / Backtest overfitting

Where You Cut the Sample Is One More Free Parameter

How much does a chosen train-test split affect the apparent edge?

Freeze one chronological training and test split before viewing the final result, then describe sensitivity across predeclared alternative splits. A split chosen because it flatters the strategy is part of the search and cannot serve as untouched confirmation.

A split chosen after inspecting performance is exploratory; freeze it before a confirmatory test.

The walk-forward ribbon marks repeated training and unseen test windows.
The walk-forward ribbon marks repeated training and unseen test windows.

Evidence map

AspectFinding
What it isThe econometric literature on how the researcher's choice of where to divide a series for evaluation affects the conclusion, and what to do about it.
Key result / formulaInoue and Kilian compare tests of predictability computed on the full sample with tests computed on a held-out portion, and show the held-out test is not automatically the more trustworthy: it discards information, so its power is lower, and it introduces the choice of where to cut, which is a degree of freedom the full-sample test does not have.
Why it matters for backtestingIn Stochastly the division is set in the validation node, and the temptation is exactly the one described: move it until the evaluation period looks acceptable.

What it is

The division feels like a neutral procedural step; it is a parameter, and it can be tuned like any other.

Key result / formula

A researcher who tries several cut points and reports the one that works has performed a search whose multiplicity is invisible. Hansen and Timmermann study the size distortion this produces and show that the distribution of the test statistic depends on the fraction of the sample retained for evaluation, so a result significant under one division may not be under another, without anything about the data having changed. Rossi and Inoue propose tests that are explicitly robust to the choice, by considering the statistic across all admissible divisions rather than at one, which removes the free parameter by integrating over it.

Why it matters for backtesting

The disciplined procedure is to fix the division before any fitting, on a rule that does not refer to results, such as a calendar date or a fixed fraction, and to record that choice with the verdict. When a user has already moved it, the sound repair is to report the result across a range of divisions, the chosen one among them, which is the practical version of the robust tests: a conclusion that holds across the range is a conclusion, and one that appears at a single cut point is a parameter value. The related warning is that a held-out segment used more than once is no longer held out, since each look spends some of its independence.

Source

Inoue & Kilian, "In-Sample or Out-of-Sample Tests of Predictability: Which One Should We Use?", Econometric Reviews 23(4), 2005, 371-402; Hansen & Timmermann, "Choice of Sample Split in Out-of-Sample Forecast Evaluation", EUI Working Paper ECO 2012/10, European University Institute, 2012; Rossi & Inoue, "Out-of-Sample Forecast Tests Robust to the Choice of Window Size", Journal of Business & Economic Statistics 30(3), 2012, 432-453. Primary source

Related in the library