Help · Knowledge base · Concept
Data hygiene (survivorship, revisions, point-in-time)
Data hygiene (survivorship, revisions, point-in-time)
Definition
Data hygiene is the set of practices ensuring that
historical data used for research actually represents
what was knowable and tradable at the time. Its
foundational result is Brown et al (1992): datasets
containing only SURVIVORS — funds, stocks, strategies
that lasted long enough to be in today’s database —
systematically overstate performance and can
manufacture apparent skill and predictability out of
nothing. The same family of defects includes restated
fundamentals presented as if original, dividend/split
adjustments applied backwards, delisted-return
omission, and index-membership hindsight. Dirty data
makes clean methods worthless — this entry is
quant-backtest-hygiene’s substrate.
How it works / structure
- Survivorship (Brown et al): excluding dead funds inflates measured mutual-fund performance (documented magnitudes around 1%+ annually in the follow-on literature); for equities the analogue is testing on CURRENT index members — a strategy “picking” today’s S&P 500 constituents in 1990 data embeds the answer in the question; the cure is as-of-date universes including delistings, with DELISTING RETURNS included (omitting them overstates small-cap and distress-strategy results — documented).
- Point-in-time fundamentals: reported financials
get RESTATED; databases that overwrite history with
final numbers let backtests trade on corrections
nobody had (
fa-financial-statementsrestatement machinery) — point-in-time snapshots (data as originally filed, with availability LAGS — a Q4 filing isn’t knowable January 1) are the standard; earnings-estimate archives have the same requirement. - Price-series mechanics (engine-relevant):
back-adjusted futures contracts (adjustment method
changes signal behavior at rolls), split/dividend
adjustment consistency (total-return vs price series
answer different questions), bad ticks and stale
prints in intraday archives, and exchange-count
drift in breadth data (
indicator-new-highs-lowsraw-count decay) — each a documented artifact class with a named cure. - The audit protocol: for any dataset — who built it, does it include the dead, is it as-of-date or overwritten, what are the availability lags, how are corporate actions handled; unanswered questions are disclosed limitations, not ignorable details.
When it applies
Before every backtest and every empirical claim —
data audit precedes method (quant-backtest-hygiene
assumes this entry passed); vendor evaluation
(the survivorship and point-in-time questions are the
first two); factor research
(strategy-factor-investing results are notoriously
sensitive to delisting-return treatment — documented
in the size/value literature); any ML pipeline
(quant-ml-in-trading — models find leakage with
superhuman efficiency).
Risk profile & failure modes
- Invisible inflation (the core hazard): dirty data errs almost uniformly FLATTERING — survivorship, look-ahead, and restatement leakage all inflate results, so undisclosed data provenance should be treated as optimistic bias, not neutral noise.
- Leakage via joins: merging datasets keyed on identifiers that change (tickers recycle, CUSIPs change in mergers) silently misaligns histories — the documented plumbing failure class.
- Adjustment-method sensitivity: signals near ex-dividend dates or futures rolls can be artifacts of the adjustment choice — robustness across methods is the test.
- Clean-data overconfidence: even pristine data
represents ONE historical path — hygiene removes
artifacts, not regime dependence
(
risk-scenario-analysispicks up from there).
Evidence & limits
Brown et al (1992) is the peer-reviewed anchor; delisting-return and point-in-time effects are documented across the empirical-finance literature. Perfect provenance is unattainable — the standard is disclosed, audited, directionally-understood data limitations attached to every result.
Falsifiable-thesis examples
Illustrations only, not signals:
- “Strategy X’s backtest Sharpe drops <15% when rerun on a survivorship-free universe with delisting returns (survivorship-sensitivity check)” — falsified by the paired run.
- “Fundamental signal Y retains significance on point-in-time data with realistic filing lags (leakage check)” — falsified by the point-in-time rerun.
Cross-references
- The method layer above:
quant-backtest-hygiene; the ML amplifier:quant-ml-in-trading - The fundamentals source:
fa-financial-statements - The applied sensitivities:
strategy-factor-investing,indicator-new-highs-lows
Sources
- Brown, S., Goetzmann, W., Ibbotson, R. and Ross, S. (1992), Survivorship Bias in Performance Studies — Review of Financial Studies 5(4), 553-580
The agent cites this page.
Inside the platform, this entry is live context. A signed-in citation opens the in-app view of the same id.