Library / Backtest overfitting

Lucky Factors: Testing a Candidate Against What Is Already Known

What should a new factor be compared with after a large factor search?

Compare the new factor?s incremental performance with an existing benchmark factor under the same universe, costs and held-out period. A positive standalone return can duplicate known exposure, while an incremental test asks whether the addition changes the decision.

Factor-zoo study differs from the vault protocol for testing a newly added lever.

Observed P&L is compared with one thousand shuffled-sign histories of equal magnitudes.
Observed P&L is compared with one thousand shuffled-sign histories of equal magnitudes.

Evidence map

AspectFinding
What it isA method for deciding whether a new explanatory variable adds anything, given that many have already been proposed and that the candidates are correlated with each other.
Key result / formulaThe naive procedure tests each candidate separately against a null of no effect, which ignores both the multiplicity of the search and the fact that the candidates overlap.
Why it matters for backtestingThe user's version of this problem is a strategy that already works and a proposed addition, a filter, an extra rule, a second signal.

What it is

It replaces the usual test against zero with a test against the incumbents, which is the question a researcher actually faces.

Key result / formula

Harvey and Liu propose a sequential bootstrap procedure. At each step the candidate that best explains the data, given those already selected, is identified; the distribution of that best statistic under the null is obtained by resampling in a way that preserves the correlation structure among candidates and the dependence in the data; the candidate is admitted only if its statistic exceeds what the resampled maximum would typically produce. The procedure then repeats on the residual, so each new admission must beat the null of the maximum conditional on what is already in. The companion methodological survey with Saretto compares the available multiple-testing corrections for finance applications and sets out when each is appropriate, since the correlation among tests makes the simplest corrections both too conservative in some settings and inadequate in others.

Why it matters for backtesting

Test whether the increment exceeds what the base already delivers, judged against the distribution of the best improvement that random additions would produce. That is implementable here: generate a set of placebo additions of the same kind, measure the improvement each gives, and compare the real one against that distribution. It connects directly to the house law that what is judged is the increment (see [What is judged is the increment, and the bar is what is already in place]) and to [Ranking Noise and the Null of the Maximum]. The failure mode it prevents is the slow accumulation of filters, each of which looked like an improvement when tested alone.

Source

Harvey & Liu, "Lucky factors", Journal of Financial Economics 141(2), 2021, 413-435; Harvey, Liu & Saretto, "An Evaluation of Alternative Multiple Testing Methods for Finance Applications", Review of Asset Pricing Studies 10(2), 2020, 199-248. Primary source

Related in the library