Library / Backtest overfitting

Hansen's Test for Superior Predictive Ability (SPA)

Does a leading trading rule remain superior when tested against the full candidate set?

Freeze the benchmark, trading-rule universe and loss differential before testing. Hansen's SPA test evaluates whether the strongest rule improves on that benchmark while reducing the influence of plainly poor alternatives. Its p-value describes the tested universe under the bootstrap assumptions, not future profitability after costs.

SPA addresses selection from a population of configurations with bootstrap dependence assumptions.

An unconnected Hansen SPA node exposes its bootstrap count and statistic field.
An unconnected Hansen SPA node exposes its bootstrap count and statistic field.

Evidence map

AspectFinding
What it isA more powerful successor to White's Reality Check.
Key result / formulaTwo fixes to the Reality Check: (1) studentize each strategy's performance (divide by its bootstrap standard deviation) so erratic, high-variance rules don't dominate the max-statistic; (2) use a sample-dependent null that recenters only strategies plausibly near the benchmark, instead of White's least-favorable configuration (which assumes every rule exactly breaks even).
Why it matters for backtestingWhite's RC is conservative — adding obviously terrible strategies to the universe raises its p-value and can hide a real edge.

What it is

It tests whether the best strategy beats a benchmark while removing the upward bias that many poor or irrelevant strategies inject into the null.

Key result / formula

Result: correct test size with substantially higher power — bad rules no longer inflate the p-value and mask a genuinely good one.

Why it matters for backtesting

SPA is the preferred data-snooping test when the strategy universe is large and heterogeneous (as in any big config sweep). Hansen reports lower/consistent p-value bounds bracketing the truth.

Worked protocol

Let d[t,j] be the loss of a benchmark minus the loss of candidate j at time t. A positive mean favours that candidate. For each candidate, center the d series under the null, resample time blocks to preserve serial dependence, and recompute the largest studentized mean across all j. With 20 candidates and 1,000 bootstrap draws, the adjusted p-value is the fraction of draws whose maximum exceeds the observed maximum. Hansen's SPA adjustment reduces the undue influence of very poor alternatives in White's Reality Check. The benchmark, loss function, candidate universe and block rule must be fixed before seeing the p-value. Failure to reject does not establish equivalence; it may mean that the sample cannot distinguish a small profitable effect.

Source

Hansen, "A Test for Superior Predictive Ability", Journal of Business & Economic Statistics 23(4), 2005, pp. 365–380 (SSRN 264569). Primary source

Related in the library