Library / Machine learning in finance
The Virtue of Complexity Debate: More Parameters Than Observations
When does added model complexity help financial prediction, and how is that tested?
Compare simple and complex return predictors on the same US-equity sample using a fixed out-of-sample period, common costs and uncertainty intervals. The cited debate concerns settings where regularization and weak signals matter; added complexity has no universal benefit.
The model-complexity debate concerns studied return-prediction settings, not a universal rule for all models.
Evidence map
| Aspect | Finding |
|---|---|
| What it is | The live disagreement about whether models with more parameters than data points can predict returns better than small ones. |
| Key result / formula | The machine learning background is the double descent phenomenon documented by Belkin, Hsu, Ma and Mandal: as model capacity grows past the point of exactly fitting the training data, test error first rises as the classical trade-off predicts, then falls again, so very large models can generalise well despite interpolating the training set. |
| Why it matters for backtesting | An agent asked whether to use a large model should present this as unsettled and give the practical reading, which both sides support: what matters is the amount of regularisation; the parameter count matters little, and a large model without heavy shrinkage is not what either side is defending. |
What it is
It matters because the answer determines whether the usual advice to keep models simple is a law or a convention.
Key result / formula
Kelly, Malamud and Zhou carry this into return prediction, arguing with theory and empirics that in the extremely low signal environment of markets, a model with many more parameters than observations, heavily regularised, outperforms the parsimonious alternative, and that the conventional preference for simplicity costs performance. Nagel's critique argues the result is less about complexity than about the implicit shrinkage the construction performs: the apparent gain can be reproduced by simple regularised estimators, so the virtue belongs to the regularisation rather than to the complexity, and the framing overstates what is new.
Why it matters for backtesting
For a user of Stochastly the operational advice is unchanged by the debate. Penalise strongly, validate chronologically with purging, and compare against the simple version at equal risk; if a complex model wins under that protocol, it wins, and if the comparison was never run the debate is irrelevant to the decision. The historical lesson is worth stating too: the classical rule against overparameterisation is a good default that turned out to have exceptions, which is a reason to test either position before reciting it.
Source
Kelly, Malamud & Zhou, "The Virtue of Complexity in Return Prediction", Journal of Finance 79(1), 2024, 459-503; Belkin, Hsu, Ma & Mandal, "Reconciling modern machine-learning practice and the classical bias-variance trade-off", Proceedings of the National Academy of Sciences 116(32), 2019, 15849-15854; Nagel, "Seemingly Virtuous Complexity in Return Prediction", working paper, 2025. Primary source