Library / Backtest overfitting

Why Most Published Research Findings Are False

Why can a published positive finding fail to replicate in a new sample?

A positive finding becomes less persuasive when the prior chance of a true effect is low, power is weak, and analysis choices were selected after outcomes were seen. Record the design, all tested variants and a prospective evaluation; Ioannidis's result is conditional on these quantities, not a verdict on every published strategy.

Ioannidis?s result depends on prior odds, power, bias and study design; it is not a verdict on every positive strategy.

Evidence map

AspectFinding
What it isIoannidis's analysis of the conditions under which a published positive result is more likely to be wrong than right.
Key result / formulaThe probability that a claimed finding is true depends on four things: the proportion of hypotheses tested that are genuinely true, the power of the study to detect a true effect, the significance threshold used, and the amount of flexibility and bias in the analysis.
Why it matters for backtestingSome strategy searches share the low-prior, low-power and flexible-analysis conditions in Ioannidis's conditional model.

What it is

It is written about biomedical research, but each term in its argument has a direct counterpart in strategy research, which is why it belongs in this folder.

Key result / formula

Combining them shows that a positive result is more likely false when studies are small, when effects are small, when many relationships are tested and the significant ones alone reported, when there is flexibility in design and analysis, and when many teams work the same question competitively. The consequence that surprises readers is that in a field where true hypotheses are rare and power is low, the majority of positive published findings will be false even when each individual study follows the rules, because the false positives among the many false hypotheses outnumber the true positives among the few true ones. Remedies include better power, pre-specification and full reporting of each analysis tried.

Why it matters for backtesting

Their prevalence must be assessed for the actual study. A positive result should be interpreted with the study design, estimated effect, power and selection process; the paper does not make every threshold-clearing strategy presumptively false. For a single predeclared sealed evaluation, trial history is reported as context, not used as an automatic threshold penalty. The practical use for a user is to run the argument on their own work before trusting a result: how likely was this hypothesis before I tested it, how much data did I have, what selection occurred, and how free was I to choose the analysis after seeing the outcome.

Source

Ioannidis, "Why Most Published Research Findings Are False", PLOS Medicine 2(8), 2005, e124. Primary source

Related in the library