Library / Backtest overfitting
Nothing Beats the Historical Mean Out of Sample, Except Under Constraints
Can a forecast beat the historical mean after the evaluation period is fixed?
Freeze the predictor and evaluation dates, generate forecasts using only information available at each date, and compare forecast error or investor utility with a rolling historical-mean benchmark. Welch and Goyal found widespread failures in their sample; Campbell and Thompson found small gains under predeclared economic restrictions.
Equity-premium forecasting benchmark does not automatically apply to every strategy.
Evidence map
| Aspect | Finding |
|---|---|
| What it is | The pair of papers that settled, and then partly unsettled, whether the equity premium can be forecast. |
| Key result / formula | Welch and Goyal take the predictors accumulated by decades of literature, valuation ratios, interest rate variables, issuance and volatility measures, and evaluate each in the way an investor would have had to use it: estimating the relationship on data available at the time and forecasting the next period. |
| Why it matters for backtesting | The first lesson is about the benchmark. |
What it is
Welch and Goyal showed that the standard predictors fail against a trivial benchmark; Campbell and Thompson showed that simple economic restrictions recover a small but real amount of predictability.
Key result / formula
Against the benchmark of simply using the historical average return to date, nearly all the predictors lose, in most sub-periods, including the ones whose in-sample relationships are strong; the apparent predictability is largely an artefact of estimating the relationship on the full sample. Campbell and Thompson accept the finding and add two restrictions that any sensible investor would impose anyway: set a forecast of the equity premium to zero if the model returns a negative number, and constrain the estimated coefficients to have the sign theory predicts. With those constraints, several predictors beat the historical mean by a small margin, and they show that a margin of that size, though statistically unimpressive, is economically meaningful for an investor allocating between equities and cash.
Why it matters for backtesting
A strategy must be judged against what a user would otherwise have done, which is usually a simple passive exposure or a constant allocation; zero is the wrong benchmark, and much of what looks like forecasting ability disappears at that comparison. The second is that imposing a sensible restriction before fitting is not cheating, it is regularisation, and it improves out-of-sample behaviour precisely because it removes the model's freedom to produce nonsense. In Stochastly this maps onto capping or flooring signal outputs and fixing the sign of a rule from the hypothesis, before any fit. The third is that a small edge can matter: the accurate report of a marginal result is a small effect with a wide interval, not a failure.
Source
Welch & Goyal, "A Comprehensive Look at The Empirical Performance of Equity Premium Prediction", Review of Financial Studies 21(4), 2008, 1455-1508; Campbell & Thompson, "Predicting Excess Stock Returns Out of Sample: Can Anything Beat the Historical Average?", Review of Financial Studies 21(4), 2008, 1509-1531. Primary source
Related in the library
In Defence of Optimisation: the Fallacy of Equal Weighting
Guide on this topic: In-sample vs out-of-sample testing
How the assistant cites the library · The checks behind a verdict