Library / Backtest overfitting

Backtest versus Live on a Large Cohort of Algorithms

How far do backtest returns fall when the same strategy trades live?

Compare each strategy's frozen backtest with its later live or genuinely untouched period, preserving the same return definition and costs. The studied platform cohort found little predictive value in backtest ratios; the result gives a strong reason to demand prospective evidence for a new strategy.

Quantopian cohort of 888 strategies with at least six months out-of-sample is not every algorithm.

Evidence map

AspectFinding
What it isA rare direct measurement of how well a backtest predicts what follows.
Key result / formulaThe sample is large and, crucially, was not assembled by the authors of the strategies for this purpose, so it is closer to a census of what retail and semi-professional quantitative research actually produces than to a curated set.
Why it matters for backtestingThis is the empirical justification for Stochastly's entire approach to verdicts.

What it is

Wiecki and co-authors leave the argument about overfitting in principle aside and take a large population of strategies developed by many independent authors on one platform, each with a backtest and a subsequent live or genuinely out-of-sample period, and correlate the two.

Key result / formula

The headline finding is that the backtested performance ratio has essentially no predictive relationship with the out-of-sample one: strategies with the most impressive backtests did not go on to do better, which is what one expects when the backtest number is dominated by selection rather than by signal. Some characteristics did carry information, and they are not performance statistics: measures of the strategy's structure and behaviour, such as its turnover, the length and breadth of its testing, and its exposure profile, related to subsequent results more reliably than its backtested ratio did. The authors also examine combinations, finding that portfolios of many such strategies behave better than the individual ones, consistent with the individual signals being weak and partly independent.

Why it matters for backtesting

If the backtested ratio does not predict the future, then reporting it as the headline result is misleading by construction, and the useful output is the set of things that do carry information: how many variants were tried, how much the strategy trades, how long and how varied the test period was, and how the result behaves under perturbation. A user asking "my backtest shows a good ratio, is that good" should be given this study as the answer, followed by the checks that do discriminate. The complementary point is about combination: several weak, genuinely independent strategies are a better object than the best single one, which is the same arithmetic as diversification.

Source

Wiecki, Campbell, Lent & Stauth, "All That Glitters Is Not Gold: Comparing Backtest and Out-of-Sample Performance on a Large Cohort of Trading Algorithms", Journal of Investing 25(3), 2016, 69-80. Primary source

Related in the library