Library / Machine learning in finance

A Backtesting Protocol for the Machine Learning Era

What must a machine-learning backtest freeze before its final evaluation?

Freeze features, labels, candidate search, transaction costs and the final evaluation window before reading its result. Record revisions made during development; a final test reused to choose the model is another training set rather than independent evidence.

The paper supports backtesting discipline; site-specific workflow rules are additional inferences.

Evidence map

AspectFinding
What it isA checklist for conducting and reporting quantitative research, written by three authors who between them defined much of the field's method.
Key result / formulaThe protocol organises the discipline into a small number of rules.
Why it matters for backtestingThis maps almost item by item onto what Stochastly enforces, which makes it the single best citation for explaining why the tool behaves as it does.

What it is

It is prescriptive rather than analytical, and its value is that it can be applied directly to a piece of work before it is believed.

Key result / formula

Research should begin from an economic hypothesis stated before the data is examined, because a model fitted without one has no ground for expecting the relationship to persist. Each result must be reported with the number of models and variants tried, since the reader cannot otherwise interpret it. Sample selection must be justified on grounds other than convenience, and the exclusion of periods requires an argument. Complexity must be penalised explicitly, and simpler specifications favoured where performance is comparable. The out-of-sample evidence must be genuinely out of sample, which means the data was not examined during model development, and the authors are explicit that a hold-out set consulted repeatedly is no longer one. Research culture matters as much as technique: the incentive to produce a positive result is the source of most of the failures, and a team should reward the identification of flaws in its own work.

Why it matters for backtesting

The hypothesis stated in advance is the registered verdict. The count of variants is the trial count. The genuine out-of-sample requirement is why a division consulted repeatedly must be treated as spent. The preference for simplicity is why a parameter that must be exact is a warning. An agent asked to justify a demand it is making of the user can cite this protocol as an external authority. The one item Stochastly cannot supply is the last, since research culture is a property of the person, and a user working alone has to supply their own adversary, which is what the devil's advocate role exists to imitate.

Source

Arnott, Harvey & Markowitz, "A Backtesting Protocol in the Era of Machine Learning", Journal of Financial Data Science 1(1), 2019, 64-74. Primary source

Related in the library