Library / Backtest overfitting

The Discount Between Backtest and Live, Measured at the Vendors

How much of an alternative-beta backtest survives live trading?

For each product in the studied alternative-beta cohort, compare its publicly dated backtest with the return history after launch under the same metric. The paper finds a substantial live discount, especially for more complex designs; the measured amount belongs to that cohort and cannot be assigned to an untested strategy.

Evidence derives from 215 bank products; discounts vary by class and strategy.

Evidence map

AspectFinding
What it isA measurement of how much performance is lost between the backtested history a bank publishes for a systematic strategy index and the performance that index delivers once it is live.
Key result / formulaSuhonen, Lennkh and Perez collect a large set of investable systematic strategy indices offered by major banks, each of which carries a backtested history before its launch date and a live record after it.
Why it matters for backtestingThis supplies an order of magnitude for the discount a user should apply to their own backtested figure, from a population that was produced by professionals with good data and strong incentives to be right.

What it is

Because these products are dated and priced publicly, the comparison can be made without relying on anyone's self-report.

Key result / formula

Comparing the two within each product, they find a substantial and systematic deterioration: the live ratio is markedly lower than the backtested one, and the gap is large enough that a naive user of the published history would have formed materially wrong expectations. The deterioration varies with the characteristics of the strategy, being worse for the more complex constructions and for those with more parameters, which is the signature of selection rather than of a change in markets. The design is what makes the result useful: the launch date is a hard boundary chosen by the issuer, outside the researcher's control, so the split between in-sample and subsequent data is not a choice made after seeing results.

Why it matters for backtesting

It is also the cleanest available argument for Stochastly's insistence on counting trials: the products with more moving parts lost more, and complexity is a proxy for the size of the search behind the design. The practical consequence is to favour the simpler construction when two variants perform similarly, and to treat any elaborate strategy as carrying a larger undeclared search. Stochastly can make the comparison concrete by holding out a final segment of the user's data, chosen before any fitting, and reporting both figures side by side.

Source

Suhonen, Lennkh & Perez, "Quantifying Backtest Overfitting in Alternative Beta Strategies", Journal of Portfolio Management 43(2), 2017, 90-104. Primary source

Related in the library