Library / Statistics

Efron's Bootstrap: Resampling the Sample to Get a Standard Error

What does resampling reveal about the uncertainty of a statistic?

Resample observations under a sampling scheme that matches the data dependence, recompute the statistic in every draw and inspect the resulting interval. Plain row resampling generally assumes independent draws; clustered or time-dependent returns need a suitable block design.

Plain bootstrap inference needs sampling assumptions; exchangeability alone is not a universal guarantee.

An unconnected BCa bootstrap node displays its confidence-interval controls in the app editor.
An unconnected BCa bootstrap node displays its confidence-interval controls in the app editor.

Evidence map

AspectFinding
What it isThe method that replaced formula-based standard errors with computation.
Key result / formulaGiven observations x_1..x_n and a statistic θ̂ = s(x), a bootstrap replicate draws n values with replacement from the empirical distribution F̂ and computes θ̂* = s(x*); after B replicates the bootstrap standard error is the standard deviation of the θ̂*, and percentile intervals come from their quantiles.
Why it matters for backtestingThe Monte Carlo node in Stochastly is a bootstrap of trade or bar returns, and the first question is which of the three limits applies.

What it is

Treat the observed sample as if it were the population, draw from it with replacement many times, recompute the statistic on each draw, and read the variability of the statistic from the spread of the replicates. Efron introduced it as a generalisation of the jackknife and showed the jackknife is its linear approximation.

Key result / formula

The paper works through the variance of the sample median, where no simple formula exists, error rates of a linear discriminant, ratio estimators and regression coefficients, and shows the method behaves well where the jackknife fails. Three limits sit in the original setting. The observations must be exchangeable, so the plain bootstrap is valid for independent draws alone; it measures variability around the sample, so it inherits any bias of θ̂ and cannot repair an unrepresentative sample; and it cannot extrapolate beyond the data, so the largest resampled value cannot exceed the largest observed one, which makes tail quantities its least reliable output. Later refinements (bias-corrected intervals, the studentized bootstrap) improve the intervals for skewed statistics such as ratios, of which the Sharpe ratio is one.

Why it matters for backtesting

Trade returns from a trend strategy are not exchangeable: winners cluster in trending stretches, so independent resampling destroys that dependence and understates the variance of the equity curve; the block versions in [Stationary & Block Bootstrap for Dependent Data] exist for that reason. A user can test this directly: compute the bootstrap interval of the mean return or Sharpe ratio with independent resampling and with block resampling, and if the block interval is materially wider, the dependence is real and the narrower interval is wrong. The tail limit is not fixable by any resampling scheme: a drawdown worse than any observed cannot be produced by resampling the observed, which is why the bootstrap answers "how uncertain is my estimate given this history" and not "what could happen that did not". The bias limit is the subtle one in a strategy search: bootstrapping the returns of the best of many candidates reproduces the selected sample's optimism faithfully, so the interval is centred on a biased number; see [Deflated Sharpe Ratio (DSR)] for the correction that resampling cannot supply.

Source

Efron, "Bootstrap Methods: Another Look at the Jackknife", Annals of Statistics 7(1), 1979, 1-26. Primary source

Related in the library