Library / Data practices

A Wide Panel Empties Itself Silently

When a wide market-data panel loses assets during cleaning, what disappears from the backtest?

If each of C independent columns is present on a fraction p of dates, complete-case filtering retains roughly p to the power C of rows. That can remove most candidate observations before any model is fitted.

The p^C relation assumes independent missingness across columns.

The app counts gaps, repeated closes and other data quality checks on the recorded two-market run.
The app counts gaps, repeated closes and other data quality checks on the recorded two-market run.

Evidence map

AspectFinding
What it isThe failure mode of assembling many series with different calendars into one matrix.
Key result / practiceWith C columns each independently present on a fraction p of dates, the share of complete rows falls like p^C.
Practical rulesConstrain completeness on the target alone, and impute the features explicitly — after standardisation, a neutral fill is a defined value with a stated meaning, and it is visible in the data.

What it is

Any requirement that a row be complete across all columns intersects the gaps of each column at once, so the usable window shrinks geometrically with the number of columns, often to an empty sample, without an error being raised.

Key result / practice

Real panels are worse than that arithmetic suggests, because gaps are not independent: a weekly series is absent four days in five, a series that starts later removes the entire early window, and a single column with a long outage can void an entire period. Dropping incomplete rows therefore produces an empty or badly truncated sample that downstream code treats as legitimate — a model fitted on an empty sample still returns an output, and an empty out-of-sample block returns no error, just a missing verdict.

Practical rules

Print, at each stage of the pipeline, the number of rows, the number of columns and the first and last usable dates, and assert lower bounds on them so a collapse fails loudly. Report the share of each feature that is filled rather than observed, since a feature that is mostly imputed cannot support a conclusion. Where a column's history is much shorter than the rest, settle in advance whether it enters the study at all: adding it silently redefines the sample for each other column.

Source

Little & Rubin, Statistical Analysis with Missing Data, 3rd ed., Wiley 2019 (missingness mechanisms and the bias of complete-case analysis); van Buuren, Flexible Imputation of Missing Data, 2nd ed., CRC 2018. Primary source

Related in the library