Knowledge base · Concept

Data hygiene (survivorship, revisions, point-in-time)

Educational reference from the platform knowledge base — written agent-readable first, rendered here for humans. Mechanics, not advice: nothing here is a recommendation to buy or sell any security.

Data hygiene (survivorship, revisions, point-in-time)

Definition

Data hygiene is the set of practices ensuring that historical data used for research actually represents what was knowable and tradable at the time. Its foundational result is Brown et al (1992): datasets containing only SURVIVORS — funds, stocks, strategies that lasted long enough to be in today’s database — systematically overstate performance and can manufacture apparent skill and predictability out of nothing. The same family of defects includes restated fundamentals presented as if original, dividend/split adjustments applied backwards, delisted-return omission, and index-membership hindsight. Dirty data makes clean methods worthless — this entry is quant-backtest-hygiene’s substrate.

How it works / structure

  • Survivorship (Brown et al): excluding dead funds inflates measured mutual-fund performance (documented magnitudes around 1%+ annually in the follow-on literature); for equities the analogue is testing on CURRENT index members — a strategy “picking” today’s S&P 500 constituents in 1990 data embeds the answer in the question; the cure is as-of-date universes including delistings, with DELISTING RETURNS included (omitting them overstates small-cap and distress-strategy results — documented).
  • Point-in-time fundamentals: reported financials get RESTATED; databases that overwrite history with final numbers let backtests trade on corrections nobody had (fa-financial-statements restatement machinery) — point-in-time snapshots (data as originally filed, with availability LAGS — a Q4 filing isn’t knowable January 1) are the standard; earnings-estimate archives have the same requirement.
  • Price-series mechanics (engine-relevant): back-adjusted futures contracts (adjustment method changes signal behavior at rolls), split/dividend adjustment consistency (total-return vs price series answer different questions), bad ticks and stale prints in intraday archives, and exchange-count drift in breadth data (indicator-new-highs-lows raw-count decay) — each a documented artifact class with a named cure.
  • The audit protocol: for any dataset — who built it, does it include the dead, is it as-of-date or overwritten, what are the availability lags, how are corporate actions handled; unanswered questions are disclosed limitations, not ignorable details.

When it applies

Before every backtest and every empirical claim — data audit precedes method (quant-backtest-hygiene assumes this entry passed); vendor evaluation (the survivorship and point-in-time questions are the first two); factor research (strategy-factor-investing results are notoriously sensitive to delisting-return treatment — documented in the size/value literature); any ML pipeline (quant-ml-in-trading — models find leakage with superhuman efficiency).

Risk profile & failure modes

  • Invisible inflation (the core hazard): dirty data errs almost uniformly FLATTERING — survivorship, look-ahead, and restatement leakage all inflate results, so undisclosed data provenance should be treated as optimistic bias, not neutral noise.
  • Leakage via joins: merging datasets keyed on identifiers that change (tickers recycle, CUSIPs change in mergers) silently misaligns histories — the documented plumbing failure class.
  • Adjustment-method sensitivity: signals near ex-dividend dates or futures rolls can be artifacts of the adjustment choice — robustness across methods is the test.
  • Clean-data overconfidence: even pristine data represents ONE historical path — hygiene removes artifacts, not regime dependence (risk-scenario-analysis picks up from there).

Evidence & limits

Brown et al (1992) is the peer-reviewed anchor; delisting-return and point-in-time effects are documented across the empirical-finance literature. Perfect provenance is unattainable — the standard is disclosed, audited, directionally-understood data limitations attached to every result.

Falsifiable-thesis examples

Illustrations only, not signals:

  • “Strategy X’s backtest Sharpe drops <15% when rerun on a survivorship-free universe with delisting returns (survivorship-sensitivity check)” — falsified by the paired run.
  • “Fundamental signal Y retains significance on point-in-time data with realistic filing lags (leakage check)” — falsified by the point-in-time rerun.

Cross-references

  • The method layer above: quant-backtest-hygiene; the ML amplifier: quant-ml-in-trading
  • The fundamentals source: fa-financial-statements
  • The applied sensitivities: strategy-factor-investing, indicator-new-highs-lows

Sources

  • Brown, S., Goetzmann, W., Ibbotson, R. and Ross, S. (1992), Survivorship Bias in Performance Studies — Review of Financial Studies 5(4), 553-580

The agent cites this page.

Inside the platform, this entry is live context: the AI reasons from it, quotes it, and grades against it. Make your case.

Inquire about founding membership