Library / Backtest overfitting
Backtest versus Live on a Large Cohort of Algorithms
How far do backtest returns fall when the same strategy trades live?
Compare each strategy's frozen backtest with its later live or genuinely untouched period, preserving the same return definition and costs. The studied platform cohort found little predictive value in backtest ratios; the result gives a strong reason to demand prospective evidence for a new strategy.
Quantopian cohort of 888 strategies with at least six months out-of-sample is not every algorithm.
Evidence map
| Aspect | Finding |
|---|---|
| What it is | A rare direct measurement of how well a backtest predicts what follows. |
| Key result / formula | The sample is large and, crucially, was not assembled by the authors of the strategies for this purpose, so it is closer to a census of what retail and semi-professional quantitative research actually produces than to a curated set. |
| Why it matters for backtesting | This is the empirical justification for Stochastly's entire approach to verdicts. |
What it is
Wiecki and co-authors leave the argument about overfitting in principle aside and take a large population of strategies developed by many independent authors on one platform, each with a backtest and a subsequent live or genuinely out-of-sample period, and correlate the two.
Key result / formula
The headline finding is that the backtested performance ratio has essentially no predictive relationship with the out-of-sample one: strategies with the most impressive backtests did not go on to do better, which is what one expects when the backtest number is dominated by selection rather than by signal. Some characteristics did carry information, and they are not performance statistics: measures of the strategy's structure and behaviour, such as its turnover, the length and breadth of its testing, and its exposure profile, related to subsequent results more reliably than its backtested ratio did. The authors also examine combinations, finding that portfolios of many such strategies behave better than the individual ones, consistent with the individual signals being weak and partly independent.
Why it matters for backtesting
If the backtested ratio does not predict the future, then reporting it as the headline result is misleading by construction, and the useful output is the set of things that do carry information: how many variants were tried, how much the strategy trades, how long and how varied the test period was, and how the result behaves under perturbation. A user asking "my backtest shows a good ratio, is that good" should be given this study as the answer, followed by the checks that do discriminate. The complementary point is about combination: several weak, genuinely independent strategies are a better object than the best single one, which is the same arithmetic as diversification.
Source
Wiecki, Campbell, Lent & Stauth, "All That Glitters Is Not Gold: Comparing Backtest and Out-of-Sample Performance on a Large Cohort of Trading Algorithms", Journal of Investing 25(3), 2016, 69-80. Primary source
Related in the library
Statistical Overfitting: Why More Searching Gives a Worse Strategy
The Discount Between Backtest and Live, Measured at the Vendors
Guide on this topic: Overfitting in trading, explained
How the assistant cites the library · The checks behind a verdict