Library / Machine learning in finance

When Cross-Validation Is Valid on Time Series, and When It Is Not

When can time-series cross-validation estimate future prediction error honestly?

Bergmeir, Hyndman and Koo show that ordinary cross-validation can be valid for autoregressive prediction under specific error and model assumptions. This result does not authorize random folds for overlapping financial labels or future-dependent features.

The cited validity result concerns autoregressive models under its assumptions; overlapping trade labels need separate validation.

Evidence map

AspectFinding
What it isThe pair of papers that replaced the blanket prohibition on cross-validation for time series with a precise condition.
Key result / formulaThe conventional advice is that cross-validation cannot be used on dependent data because training and test sets share information, so a chronological division is the one admissible design.
Why it matters for backtestingThis is the reference an agent should use when a user asks why Stochastly insists on purged validation.

What it is

The prohibition is right for some situations and wrong for others, and the distinction turns on whether the model's errors are serially correlated.

Key result / formula

Bergmeir and Benitez examine this empirically and find that for purely autoregressive prediction, where the model uses past values of the series alone, a blocked cross-validation performs well and provides more reliable error estimates than a single chronological split, because it uses the data several times over and the single split is a high-variance estimate of a quantity that varies across periods. Bergmeir, Hyndman and Koo supply the theory: standard cross-validation on an autoregressive model gives asymptotically unbiased error estimates provided the model's errors are uncorrelated, that is, provided the model is correctly specified in the sense of capturing the serial dependence. If the errors remain correlated, the estimate is biased and the chronological split is required. The condition is testable on the fitted residuals, which turns a doctrinal question into a diagnostic.

Why it matters for backtesting

The answer is conditional. Where the model uses exogenous features and the labels overlap in time, which is the usual case here, the dependence is structural and purging and embargoing are necessary (see [Purged K-Fold & CPCV]). Where the model is a plain autoregression on a single series and its residuals test as uncorrelated, blocked cross-validation is defensible and uses scarce data better than one split. The diagnostic to run is the autocorrelation of the fitted residuals; the decision should follow that measurement rather than a rule of thumb, and the choice should be recorded, since it changes what the reported error means.

Source

Bergmeir & Benitez, "On the use of cross-validation for time series predictor evaluation", Information Sciences 191, 2012, 192-213; Bergmeir, Hyndman & Koo, "A note on the validity of cross-validation for evaluating autoregressive time series prediction", Computational Statistics & Data Analysis 120, 2018, 70-83. Primary source

Related in the library