Library / Data practices

Cleaning Tick Data, Step by Step

Which tick-data errors can create trades that never existed?

Check duplicated timestamps, impossible prices, crossed quotes and trade-quote alignment before building signals or fills. Document each removal and rerun the strategy with and without disputed records; equity TAQ cleaning thresholds do not automatically transfer to crypto or FX.

The source studies equity TAQ cleaning; the same thresholds are not universal for crypto or FX.

Evidence map

AspectFinding
What it isThe published procedures for turning a raw trade and quote feed into something analysable.
Key result / formulaBarndorff-Nielsen, Hansen, Lunde and Shephard set out an explicit cleaning sequence used before computing realised measures: keep entries from a single exchange, delete entries outside the official trading session, delete entries with a price or size of zero, aggregate entries with the same timestamp, remove quotes with negative or implausibly large spreads, and delete trades whose price deviates from a local neighbourhood of surrounding trades by more than a multiple of the local dispersion.
Why it matters for backtestingFor a user of Stochastly the transferable content is what to ask of a bar file and what to check in it.

What it is

They matter here even for a user who does not work with ticks, because each bar they use was produced by someone applying, or failing to apply, these steps.

Key result / formula

They emphasise that the choices are consequential: different cleaning rules produce measurably different volatility estimates, so the rule must be reported alongside the estimate. Brownlees and Gallo formalise the outlier filter as a comparison with a centred rolling neighbourhood, with parameters for the window, the number of dispersions and a granularity floor for markets that trade on a fine price grid. Dacorogna and colleagues supply the earlier framework from the currency market, including the treatment of multiple contributors and the construction of a filtered series with a quality measure attached.

Why it matters for backtesting

Bars aggregated from an uncleaned feed inherit each bad print, and a single erroneous tick becomes a bar's high or low, which is exactly the value a stop or a breakout rule reads. The checks that run on bars are the same in spirit: compare each bar's range with a local neighbourhood of ranges and flag the outliers, look for bars whose high or low is far from both neighbouring closes, and check for duplicated timestamps and zero volumes. A strategy whose result depends on a handful of extreme bars should have those bars inspected individually before the result is believed, since the cheapest explanation for an extreme bar is a bad print.

Source

Barndorff-Nielsen, Hansen, Lunde & Shephard, "Realized kernels in practice: trades and quotes", Econometrics Journal 12(3), 2009, C1-C32; Brownlees & Gallo, "Financial econometric analysis at ultra-high frequency: Data handling concerns", Computational Statistics & Data Analysis 51(4), 2006, 2232-2245; Dacorogna, Gencay, Muller, Olsen & Pictet, An Introduction to High-Frequency Finance, Academic Press, 2001. Primary source

Related in the library