Library / Backtest overfitting

Is There a Replication Crisis in Finance? Three Papers, Three Answers

What would a successful replication of an asset-pricing result require?

Replicate the exact signal definition with point-in-time data, then report outcomes under the original universe and prespecified alternative weighting, size and geography choices. The cited anomaly studies use different samples and success criteria, so their aggregate replication rates cannot serve as one universal cutoff.

The cited studies differ in data, universes and success criteria, so no single replication percentage is universal.

Evidence map

AspectFinding
What it isThe disagreement at the centre of empirical asset pricing, presented as such.
Key result / formulaHou, Xue and Zhang replicate a very large set of published anomalies with a uniform procedure, weighting portfolios by value rather than equally and excluding micro-cap stocks, and report that a majority fail to produce significant returns under that treatment; their diagnosis is that much of the literature's evidence rests on the smallest and least liquid securities.
Why it matters for backtestingAn agent asked whether published anomalies are real should present this as an open disagreement, with the sample and estimand attached, and should name the reasons for it: the answer depends on data, weighting, universe filters and the unit of analysis, each of them a choice made by the replicator.

What it is

Three research teams applied replication methods to overlapping sets of published anomalies and reached conclusions that range from "most do not survive" to "the field replicates well". The studies differ in data, sample construction, universes and methods of adjudication.

Key result / formula

Chordia, Goyal and Saretto approach it through multiple testing, generating an enormous number of plausible anomaly specifications from accounting data to characterise what the search process itself produces, and study how candidate generation changes apparent significance within their anomaly-testing setting; this is not a universal rule to raise thresholds for a newly added lever or one sealed test. Jensen, Kelly and Pedersen take the opposite view using a hierarchical Bayesian model that pools information across related factors and across countries: they argue that judging each factor in isolation overstates failure, that the appropriate question is whether a factor's family replicates, and that on that basis the majority of factors do replicate, including out of sample and internationally.

Why it matters for backtesting

That is the transferable lesson for a user's own work. A result that flips between equal and value weighting, or that disappears when the smallest instruments are excluded, depends on a choice the user should make deliberately and declare. Both versions should be run, and both reported (see [The scope of a measurement is a hypothesis, and it can be wrong]).

Source

Hou, Xue & Zhang, "Replicating Anomalies", Review of Financial Studies 33(5), 2020, 2019-2133; Chordia, Goyal & Saretto, "Anomalies and False Rejections", Review of Financial Studies 33(5), 2020, 2134-2179; Jensen, Kelly & Pedersen, "Is There a Replication Crisis in Finance?", Journal of Finance 78(5), 2023, 2465-2518. Primary source

Related in the library