Library / Econophysics

DFA and Modified R/S: Estimating Long Memory Without Fooling Yourself

Can DFA or modified R/S distinguish long memory from a finite-sample illusion?

Estimate the exponent on several disjoint windows and compare it with shuffled or short-memory controls at the same sample length. An estimate above one half alone cannot separate persistent dependence from trends, regime shifts, or finite-sample bias.

An exponent above one half alone does not prove long memory without dependence and trend checks.

The periodogram plots cycle strength and a white-noise reference for the recorded return series.
The periodogram plots cycle strength and a white-noise reference for the recorded return series.

Evidence map

AspectFinding
What it isThe two estimators that fixed the two classic ways of finding long memory that is not there.
Key result / formulaDFA integrates the series, y(k) = Σ_(i≤k) (x_i − x̄), cuts it into boxes of length n, fits a least-squares line in each box and computes the root-mean-square deviation from those local fits, F(n); a power law F(n) ∝ n^α gives the exponent, α = 0.5 for uncorrelated data and α > 0.5 for long-range positive correlation.
Why it matters for backtestingThree pitfalls survive both fixes and must be handled by design.

What it is

Detrended fluctuation analysis (Peng et al. 1994) removes local trends that raw fluctuation scaling reads as persistence; Lo's modified rescaled range (1991) removes short-range dependence that classical R/S reads as long-range dependence. The companion note on the Hurst exponent gives the definitions; this note is about what each estimator corrects and what it still leaves to you.

Key result / formula

Peng et al. introduced it on DNA sequences precisely to separate genuine long-range correlation from a "mosaic" of patches with different local means: a series assembled from such patches shows spurious scaling in the raw fluctuation, and the local detrending is what removes it. Lo keeps the range in the numerator of R/S but divides by a heteroskedasticity- and autocorrelation-consistent estimate of the long-run standard deviation with a bandwidth of q lags, Q_n = R_n / σ̂_n(q). Under short-range dependence alone the classical statistic rejects the null far too often; the modified one has a known limiting distribution, and on daily and monthly US stock index returns over several periods it finds no evidence of long-range dependence once short-range dependence is accounted for.

Why it matters for backtesting

The exponent is a slope over a chosen range of box sizes or lags: report the range and look for a crossover, since one slope over a decade and another over the next is the usual picture in returns and is not a single H. The modified R/S depends on q — a q large enough to absorb all short memory also absorbs part of any long memory — so show the statistic as a function of q, not at one value. And both are slopes of a log-log plot, which has no residual diagnostic: Peng et al.'s own motivation, patches that mimic correlation unless detrended, is the reminder that the detrending order and the box range are modelling choices, and different orders should agree before α is believed. On returns the defensible summary is Lo's — none once short memory is removed — and an estimator that reports otherwise on a price series has more likely found a trend than a memory.

Source

Peng, Buldyrev, Havlin, Simons, Stanley & Goldberger, "Mosaic organization of DNA nucleotides", Physical Review E 49(2), 1994, 1685-1689; Lo, "Long-Term Memory in Stock Market Prices", Econometrica 59(5), 1991, 1279-1313. Primary source

Related in the library