Library / Backtest overfitting
Hansen's Test for Superior Predictive Ability (SPA)
Does a leading trading rule remain superior when tested against the full candidate set?
Freeze the benchmark, trading-rule universe and loss differential before testing. Hansen's SPA test evaluates whether the strongest rule improves on that benchmark while reducing the influence of plainly poor alternatives. Its p-value describes the tested universe under the bootstrap assumptions, not future profitability after costs.
SPA addresses selection from a population of configurations with bootstrap dependence assumptions.

Evidence map
| Aspect | Finding |
|---|---|
| What it is | A more powerful successor to White's Reality Check. |
| Key result / formula | Two fixes to the Reality Check: (1) studentize each strategy's performance (divide by its bootstrap standard deviation) so erratic, high-variance rules don't dominate the max-statistic; (2) use a sample-dependent null that recenters only strategies plausibly near the benchmark, instead of White's least-favorable configuration (which assumes every rule exactly breaks even). |
| Why it matters for backtesting | White's RC is conservative — adding obviously terrible strategies to the universe raises its p-value and can hide a real edge. |
What it is
It tests whether the best strategy beats a benchmark while removing the upward bias that many poor or irrelevant strategies inject into the null.
Key result / formula
Result: correct test size with substantially higher power — bad rules no longer inflate the p-value and mask a genuinely good one.
Why it matters for backtesting
SPA is the preferred data-snooping test when the strategy universe is large and heterogeneous (as in any big config sweep). Hansen reports lower/consistent p-value bounds bracketing the truth.
Worked protocol
Let d[t,j] be the loss of a benchmark minus the loss of candidate j at time t. A positive mean favours that candidate. For each candidate, center the d series under the null, resample time blocks to preserve serial dependence, and recompute the largest studentized mean across all j. With 20 candidates and 1,000 bootstrap draws, the adjusted p-value is the fraction of draws whose maximum exceeds the observed maximum. Hansen's SPA adjustment reduces the undue influence of very poor alternatives in White's Reality Check. The benchmark, loss function, candidate universe and block rule must be fixed before seeing the p-value. Failure to reject does not establish equivalence; it may mean that the sample cannot distinguish a small profitable effect.
Source
Hansen, "A Test for Superior Predictive Ability", Journal of Business & Economic Statistics 23(4), 2005, pp. 365–380 (SSRN 264569). Primary source
Related in the library
Contagion or Interdependence: the Bias That Inflates Crisis Correlations
Guide on this topic: Multiple testing in trading research
How the assistant cites the library · The checks behind a verdict