Skip to main content

A fifth of my trading universe was delisted stocks, understating every signal by half

· 19 min read
Vadim Nicolai
Senior Software Engineer

The number I did not expect was 20.5%. That is the share of my screen-eligible universe that had already delisted and fallen out of the research panel before I scored a single signal on it. The direction was the bigger surprise. Dropping those names did not flatter my backtest; it suppressed it, understating every lane's spread by a mean 19.3 bps (se 6.5) against a headline lane spread of 37.99 bps. Just over half of a result, traceable to an instrument nobody had audited.

The delisting population has been on the record as large for nearly two decades. Macey, O'Hara and Pompilio (2008) counted more than 9,000 U.S. delistings since 1995, almost half of them involuntary. What I did not know was how much of that population my own instrument had deleted.

The failure mode I had budgeted for was optimism. What I got was arithmetic.

Everything below comes from one run of the autonomous loop: 8 sampled dates drawing 2022-11-01 through 2025-02-03, with no lane, constant, module or gate default changed. That constraint is the whole value of the exercise. The loop asked a question, wrote the answer into the log, and touched nothing — because a measurement that immediately triggers a patch stops being a measurement. The discipline behind that rule is not mine. Zhang, Li, Peng and Chen (2026) price leakage by toggling a single evaluation convention. Everything else — the data panel, the walk-forward split, the model family, the horizon, the portfolio rule, the cost convention — stays fixed.

My board had already booked the survivorship hole at 3.8%, and the real number is 20.5%

There was a survivorship figure on the books, and it had been there long enough to be treated as settled: 3.8% of the universe. That was the untyped-name effect — securities whose type never resolved, which quietly fell out of the panel during construction.

The new measurement puts the whole delisted population at 20.5%, a factor of 5.4 over the old bound: 6,612 confirmed delisted common stocks sitting in the vendor's active=false registry and absent from the panel, across the 8 sampled dates. Macey, O'Hara and Pompilio (2008) had already told us the population was large — more than 9,000 U.S. delistings since 1995, nearly half involuntary. They had not told us how much of it a modern screen would quietly delete.

Five times larger is the headline. The more useful observation is that the two numbers do not measure the same quantity. One counted a single failure path through the pipeline. The other counted everyone who died. Macey, O'Hara and Pompilio (2008) put voluntary and involuntary delistings in the same population for a reason: to a cross-sectional screen, both are simply names that stop existing.

The old figure was on the books, believed, and used. It was correct about what it measured and wrong about what it covered — scoped to one leg of the pipeline while carrying the authority of a whole-panel bound. That gap is not an error in either number. It is the difference between auditing a code path and auditing a universe.

The hole decays monotonically, and that is the fingerprint of look-ahead

The missing share falls from 29.5% on 2022-11-01 to 15.0% on 2025-02-03, and it falls monotonically across the sampled dates. That shape is the textbook signature of look-ahead survivorship: later dates have simply had less calendar time in which their names could die. A vendor coverage artifact would look flat or noisy. A screen-definition mismatch would look like a step. A near-halving of the hole over roughly 27 months is a decay curve with a known cause. Zhang, Li, Peng and Chen (2026) built their benchmark around exactly this idea — that a leak is a switch you hold fixed and toggle — and a date-indexed universe toggles it for you.

I want to be precise about what that buys me. Zhang et al. (2026) are explicit that their benchmark is diagnostic rather than a claim of tradable alpha, and the same caveat applies here. A monotone decay is a fingerprint, not a proof. It tells me the mechanism is almost certainly the one I think it is. It tells me nothing about whether the estimate is well-centered.

Deleting the dead names made every lane look worse, not better

Here is the part that inverted my expectation. The missing names underperform the panel on 7 of 8 dates. If you exclude a set of names that returns less than the panel average, you raise the equal-weighted benchmark. A lane's spread is its return minus the benchmark, so a raised benchmark deflates every lane: mean 19.3 bps, standard error 6.5. Against a headline lane spread of 37.99 bps, that is 50.8% of the number. Macey, O'Hara and Pompilio (2008) measured what happens to these names after they leave an exchange. What they did not measure is what their absence does to the cross-section they left behind.

The textbook framing of survivorship bias is about levels — the surviving index looks better than the universe anyone could have traded. That is true here too, but it is the benchmark's problem. For a relative signal, the identical deletion flips the sign of the error.

That distinction cost me a week of looking in the wrong place. I had been braced for a strategy that was too good. I found a measuring stick that was too generous to the baseline — and therefore unkind to every signal measured against it.

The composition is consistent with the sign. The names are recognisable and mostly acquisitions: ATVI delisted 2023-10-16, VMW 2023-11-24, PXD 2024-05-06, JWN 2025-05-21. A takeover pins a target's price to a deal spread. Those names stop generating the cross-sectional dispersion a signal is trying to harvest, and then they vanish from the panel entirely. "Delisted" is also not "untradeable". Macey, O'Hara and Pompilio (2008), studying NYSE firms delisted in 2002, measured percentage spreads roughly tripling and volatility roughly doubling after the move to the Pink Sheets, with volume still remarkably high. Delisting times varied widely. Some firms traded for months after failing listing requirements.

A common 19.3 bps shift is not a rescue for a lane 0.69 standard errors from zero

The 19.3 bps is a common shift to every lane. It moves levels; it does not necessarily move ordering. It does not create a factor loading, and it does not rescue a lane whose leftover after market, size and rates sits 0.69 standard errors from zero. That lane stays rejected after the correction, and it should. Zhang et al. (2026) draw the same line between a diagnostic benchmark and a claim of tradable alpha; a level shift is a diagnostic result, not a rescue.

This is the whole substance of cross-sectional work, and it is why an ordering claim survives a rebase when a level claim does not. Han, Li and Onishchenko (2021) revisit Hong and Kacperczyk's original sin-stock result with an updated sample covering 2009–2018 — a full decade — and find the premium still alive, contrary to the out-of-sample attenuation McLean and Pontiff documented elsewhere. The effect concentrates in low-liquidity and high-uncertainty states, and it is recession-proof. That is a statement about who sits above whom, and it survives a uniform rebase of the level. My 19.3 bps is not an ordering claim. It raises the whole board and changes nothing about who is above whom.

What it does change is the admission gate. A common shift leaves ranks alone and quietly rewrites the pass rate at any fixed threshold. A board that admits lanes at a spread cutoff calibrated on the biased panel is not admitting the same population it would admit on a corrected one.

Two numbers govern how much weight this carries. The first is n=8, sampled across the research slice rather than at random. So 20.5% is an estimate with no honest confidence interval, and 19.3 bps is a mean with a standard error of 6.5 attached to eight observations, not eight hundred. The ratio is about 3.0, which sounds decisive. I would rather treat it as: the sign is probably right and the magnitude is a hypothesis. Zhang et al. (2026) call their own benchmark diagnostic rather than a tradable-alpha claim, and a mean over eight dates deserves the same modesty.

Now the stakes, stated concretely. Suppose a lane is admitted at 25 bps of spread. Measured on a panel missing a fifth of its names, that gate is reading an instrument that runs 19.3 bps low. The measurement error is 77% of the admission threshold. You are not choosing between a good lane and a bad one. You are choosing a threshold on a ruler whose zero is off by three-quarters of the decision. The point of pricing a leak one switch at a time is to learn whether a fixed convention still means what you think it means, which is exactly what Zhang et al. (2026) built their paired benchmark to answer.

Macey, O'Hara and Pompilio counted the population; the screen is where it bites

The classical reference is worth reading for the shape of the risk, not just the size. Macey, O'Hara and Pompilio (2008) document more than 9,000 U.S. delistings since 1995 — nearly half of them involuntary — tripling percentage spreads and doubling volatility for firms that moved to the Pink Sheets, remarkably high volume, and delisting times that vary enough that a firm can keep trading for months after failing listing requirements. That is the full anatomy: large in number, liquid in practice, volatile, and slow to actually leave.

Read that against a factor screen. Delisted names are not a tail of untradeable garbage a dollar-volume floor would have excluded anyway. They are liquid, they are volatile, they are numerous, and they include exactly the extreme-return observations an event study wants. Macey et al. (2008) showed these names keep trading for months after they fail listing requirements. My panel shows 6,612 confirmed delisted common stocks clearing its close and dollar-volume floors before disappearing. The binding constraint was never the liquidity floor. It was the vendor's active=false flag — a statement about today, applied to 2022.

The "months after failing listing requirements" detail is the one that breaks naive reconstruction. A name can be tradable on date t and flagged as delisted in the registry on date t+1 while still printing prices. Any join that uses the current flag rather than the as-of-date state will silently censor precisely those observations.

One switch at a time is the only honest way to price a leaky universe

Zhang, Li, Peng and Chen (2026) built a paired benchmark that refuses to treat leakage as binary. They toggle one evaluation convention at a time around a clean t+1-open reference, holding the data panel, walk-forward split, model family, horizon, portfolio rule and cost convention fixed. Across two daily-OHLCV equity panels, six model families and yearly tests from 2016 to 2024, they find the inflation is highly selective. Centered temporal features and same-day-open execution with post-open daily-bar information produce large, stable increases in both predictive and trading metrics. Global normalization, future-informed graph structure and same-day-close execution are weak in most settings.

That selectivity is the finding, not the footnote. A diffuse leak shows up in any aggregate robustness check, which means you find it by accident. A concentrated leak hides inside a healthy-looking panel, which means you find it only by asking about that specific switch.

The delisted-names hole is that kind of switch, applied to a different axis. Nobody was leaking a timestamp. The universe itself was joined against a registry that only knows the present, so a fifth of the cross-section at date t was absent because of knowledge from after t. The one-switch result predicts exactly what I observed. No subperiod test, no holdout, no transaction-cost sweep flagged it — because they all run over the same already-censored panel.

The field's own audit finds universe handling in 1 of 19 primary studies

Xia, You, Wang, Liu, Qi, Wu and Zhang (2026) audited 77 included studies in a protocol-coded snapshot, with a primary empirical subset of n=19 satisfying their minimum boundary of Action Output plus Closed-Loop Evaluation. Only 2 of 19 report an extractable time-consistent split protocol. Only 1 of 19 reports an explicit transaction-cost model. Only 1 of 19 documents universe or survivorship handling. 11 of 19 report execution timing or semantics, 15 of 19 are coded R0, and no study reaches R3 reproducibility.

There is a second literature here that gets mistaken for this one. McLean and Pontiff (2012) studied 56 published characteristics and found average out-of-sample decay of about 15% — statistically indistinguishable from zero — against post-publication decay of roughly 50%, distinguishable from both 0% and 100%, and unexplained by time trends in anomaly returns. That second number is about predictability after publication. My 20.5% is about who is in the cross-section in the first place. They are different axes. Conflating them is how a universe audit gets filed under "already known" and never run.

One in nineteen. The authors' own framing is that comparable evaluation protocols, execution semantics and reproducible artifacts are the field's immediate bottleneck — not architecture. The hole in my universe is that same omission at the scale of a single system. No amount of reading was going to produce the number. It had to be measured in the panel where it lived.

The agent that found this was auditing its own instrument

Fang, Peng, Zhang, Wang, Yi, Zhang, Xu, Wu, Liu, Li and Ren (2025) formalise self-evolution as four components — System Inputs, Agent System, Environment and Optimisers — and observe that most deployed agent systems rely on manually crafted configurations that stay static after deployment. The loop that ran this measurement behaves like one of their evolved systems. It runs a two-leg removal test as standing doctrine, and it builds learned artifacts reluctantly: the A12 census counts 9 AI-native modules with 0 intersection, and the one learned artifact the system carries, fill_hazard, is refused at runtime at −0.0074 out-of-sample. Auditing its own universe is the Optimiser acting on the inputs rather than on the policy. Almost nobody instrumentates that.

The recent agent benchmarks are the closest thing to an instrument-level check, and none of them scores universe correctness. Qian et al. (2025) introduce Agent Market Arena, a lifelong real-time benchmark that evaluates four agent architectures — InvestorAgent, TradeAgent, HedgeFundAgent and DeepFundAgent — across five model backbones from GPT-4o and GPT-4.1 to Claude-sonnet-4 and Gemini-2.0-flash, built on what the authors call verified trading data. Fan et al. (2025) build AI-Trader across three markets — U.S. equities, A-shares and crypto — evaluating six mainstream LLMs, and frame it explicitly as data-uncontaminated. Saha, Lyu, Saxena, Zhao and Mehta (2025) survey the evaluation frameworks and benchmark datasets the field has produced; Yehudai, Eden, Li, Uziel, Zhao, Bar-Haim, Cohan and Shmueli-Scheuer (2025) map agent evaluation across four dimensions and flag cost-efficiency, safety and robustness as under-assessed; Mohammadi, Li, Lo and Yip (2025) add reliability guarantees and long-horizon interactions as the enterprise challenges most current research skips. Instrument audits are on none of those lists.

The source set also names two papers I cannot attribute — one proposing an LLM agent that detects look-ahead bias in backtests, another reframing look-ahead-freedom as a verifiable temporal property. I am not citing them here, because an unverifiable citation is worse than an omitted one, and I cannot tie either to an author and a year I would bet on. The idea that survives without them is Zhang et al.'s (2026): the delisted-names hole is a spatial cousin of the temporal non-interference those papers want to verify. Names that should be in the cross-section at date t are absent because knowledge from after t deleted them. In that framing, the detecting agent was the loop itself, and it found exactly the class of defect the framing predicts — a leaky universe construction that no robustness check flagged.

Practical takeaways

The audit produced four outputs, and each one drives a different decision: the population share (20.5% of screen-eligible names missing), the direction (the missing names underperform, so the benchmark is inflated and every spread deflated), the magnitude (19.3 bps against a 37.99 bps headline), and the commonality (a level shift, not a reordering). Xia et al. (2026) is the measure of how rare that audit is — of their 19 primary studies, exactly one documents universe handling at all. The mistake to avoid is treating those four outputs as a single number:

  • Share tells you whether to care.
  • Direction tells you which way your results are wrong.
  • Magnitude tells you whether your gates still mean anything.
  • Commonality tells you what survives the correction.

The literature supports the shape of the fix rather than the size of it. Macey, O'Hara and Pompilio (2008) establish that the delisted population is large, liquid and slow to leave. Zhang et al. (2026) establish that the leak worth pricing is the concentrated one you toggle deliberately. Xia et al. (2026) establish that almost nobody documents this at all.

The reconstruction itself is mechanical — build the universe as of each date rather than as of today:

-- as-of-date universe: registry state known at t, not today's flag
SELECT u.date, u.security_id
FROM universe_candidates u
JOIN registry_state r
ON r.security_id = u.security_id
AND r.effective_from <= u.date
AND (r.effective_to > u.date OR r.effective_to IS NULL)
WHERE u.close >= u.close_floor
AND u.dollar_volume >= u.dv_floor
AND r.status_at_date = 'active';

Every clause that says "at date" is load-bearing. A join on active = false pulled from the vendor today is the bug in its natural habitat. The branch is simple: a registry keyed on today's status drops a fifth of the cross-section and deflates the benchmark, while a registry keyed on the as-of-date state keeps the names and forces you to re-derive your gates.

What I would put on the checklist:

  • Rebuild the universe as of each date. Never filter history with a present-tense status flag.
  • Report the delisted population's share of the screen-eligible set as a first-class metric, per period, on the record.
  • Re-run the benchmark with and without the missing names. The delta is your instrument's error bar.
  • If the delta is common, re-derive the gates. Do not re-rank the lanes and call it done.
  • Treat a monotone decay in coverage across dates as a look-ahead signature until you can explain it otherwise.
  • Sample more dates than eight before you quote the magnitude in anything that matters.

Questions that keep coming back

What is survivorship bias in a backtest? Testing a strategy only on assets that still exist today. Names that were acquired, went bankrupt, or failed listing requirements drop out of the historical record, so the surviving sample is stronger than the universe you could actually have traded — and, as the measurement above shows, the direction of the damage depends on whether your signal is absolute or relative.

Why would survivorship bias make results worse? It is counterintuitive, but it follows from the arithmetic. Deletion usually inflates absolute returns, because you are removing losers. A relative spread, though, is measured against a benchmark built from the same survivors, and a benchmark inflated by deletion deflates every spread measured against it. That is what the 19.3 bps on a 37.99 bps headline is — and Macey et al. (2008) is the place to see why the deleted names look the way they do.

How do I get history for delisted names? Use a vendor that preserves full historical coverage including delisted tickers, and pull universe membership as of each rebalance date rather than from today's list. Macey et al. (2008) is the right starting point for understanding what happens to these names after they leave an exchange — tripling spreads, doubling volatility, and months of continued trading.

Every backtest is a claim about a universe before it is a claim about a signal

The board changed nothing on the strength of this. No lane was retired, no gate was moved, no constant was tuned. That is not restraint for its own sake — it is the only way the number stays a measurement instead of becoming a reaction to one.

What I actually got out of it was a reordering of what I trust. I had believed my instrument's flaws were the ones I had named and bounded: an untyped-name effect worth 3.8%. The flaw I had not named was five times larger, sitting in the join between the panel and a registry that only knows the present, and it had been deflating every spread the board ever scored. Macey, O'Hara and Pompilio (2008) had made plain that the delisted population was large, liquid and slow to leave. What nobody had done was check whether my panel contained it.

Ask the question of your own panel this week. How many names does your universe claim on a date three years back, and where do those names come from — a dated membership table or a current status flag? The answer is either zero or it is not. You will not find out by reading another post about survivorship bias. You find it out by counting.