Library / Backtest overfitting
Evidence-Based Technical Analysis: Rules Tested Against Their Own Null
What evidence should a technical trading rule provide beyond an attractive chart?
Define the chart rule mechanically, fix its parameters and benchmark, include trading costs, then replay the complete rule search against an appropriate bootstrap null. Aronson's negative result concerns the rule family and market he studied; excluding smaller effects requires a power analysis and positive control.
Aronson?s negative finding concerns his tested rule family and data; excluding smaller effects needs a power analysis.
Evidence map
| Aspect | Finding |
|---|---|
| What it is | Aronson's attempt to put technical analysis on a testable footing, half a critique of the discipline's subjective tradition and half a worked demonstration of how to evaluate a large set of rules while accounting for the fact that many were tried. |
| Key result / formula | The first half argues that most technical analysis is not wrong so much as unfalsifiable: rules stated in terms that require interpretation cannot be tested, because the analyst supplies the judgement that makes them work in hindsight, and the literature is dominated by results produced without any correction for the number of rules examined. |
| Why it matters for backtesting | This is the closest published analogue to what a user of Stochastly does, and the study offers a methodological template, whose operations require separate verification against Stochastly catalogue: define the family of rules before testing, evaluate each of them, the unpromising ones included, detrend or otherwise neutralise the ambient exposure so that direction is not rewarded by drift, and compare the best against the distribution of the best under the null (see [White's Reality Check]). |
Key result / formula
The second half applies the method the argument demands. Aronson defines a large universe of mechanical rules on a single index, evaluates each of them on the same data, and then applies White's reality check bootstrap to ask whether the best performer's result exceeds what the best of that many random rules would produce; he also detrends the data to remove the ambient drift, so that a rule with net long exposure is not credited for it. The finding is negative: after correcting for the size of the search, none of the rules demonstrates statistically significant predictive power on that data. The lasting value is the procedure, which is reusable; the negative verdict on one index matters less.
Why it matters for backtesting
The detrending step is the one users skip and it matters most for rules that hold no short positions on an instrument that rose over the sample. The negative result should be quoted carefully: it concerns a specific rule family on specific data, and a negative result is worthless without its floor, so its negative result is bounded by the tested rules and data. Inferring the size of effects excluded requires a power calculation and a positive control that the book summary alone does not establish.
Source
Aronson, Evidence-Based Technical Analysis: Applying the Scientific Method and Statistical Inference to Trading Signals, Wiley, 2006. Primary source