Library / Backtest overfitting
The Discount Between Backtest and Live, Measured at the Vendors
How much of an alternative-beta backtest survives live trading?
For each product in the studied alternative-beta cohort, compare its publicly dated backtest with the return history after launch under the same metric. The paper finds a substantial live discount, especially for more complex designs; the measured amount belongs to that cohort and cannot be assigned to an untested strategy.
Evidence derives from 215 bank products; discounts vary by class and strategy.
Evidence map
| Aspect | Finding |
|---|---|
| What it is | A measurement of how much performance is lost between the backtested history a bank publishes for a systematic strategy index and the performance that index delivers once it is live. |
| Key result / formula | Suhonen, Lennkh and Perez collect a large set of investable systematic strategy indices offered by major banks, each of which carries a backtested history before its launch date and a live record after it. |
| Why it matters for backtesting | This supplies an order of magnitude for the discount a user should apply to their own backtested figure, from a population that was produced by professionals with good data and strong incentives to be right. |
What it is
Because these products are dated and priced publicly, the comparison can be made without relying on anyone's self-report.
Key result / formula
Comparing the two within each product, they find a substantial and systematic deterioration: the live ratio is markedly lower than the backtested one, and the gap is large enough that a naive user of the published history would have formed materially wrong expectations. The deterioration varies with the characteristics of the strategy, being worse for the more complex constructions and for those with more parameters, which is the signature of selection rather than of a change in markets. The design is what makes the result useful: the launch date is a hard boundary chosen by the issuer, outside the researcher's control, so the split between in-sample and subsequent data is not a choice made after seeing results.
Why it matters for backtesting
It is also the cleanest available argument for Stochastly's insistence on counting trials: the products with more moving parts lost more, and complexity is a proxy for the size of the search behind the design. The practical consequence is to favour the simpler construction when two variants perform similarly, and to treat any elaborate strategy as carrying a larger undeclared search. Stochastly can make the comparison concrete by holding out a final segment of the user's data, chosen before any fitting, and reporting both figures side by side.
Source
Suhonen, Lennkh & Perez, "Quantifying Backtest Overfitting in Alternative Beta Strategies", Journal of Portfolio Management 43(2), 2017, 90-104. Primary source
Related in the library
Minimum Backtest Length (MinBTL) & Minimum Track Record Length (MinTRL)
Backtest versus Live on a Large Cohort of Algorithms
Statistical Overfitting: Why More Searching Gives a Worse Strategy
The Reference Factor Models: What a Strategy Is Measured Against
Guide on this topic: How to backtest a trading strategy
How the assistant cites the library · The checks behind a verdict