Skip to main content

2 posts tagged with "equity"

View All Tags

Autonomous trading: sample size sets the bar, not the number

· 20 min read
Vadim Nicolai
Senior Software Engineer

The largest spread-to-standard-error ratio on the board is 4.2996, and it is not significant. It has to clear 4.302653 — the two-sided 5% critical value on 2 degrees of freedom. That is not the 1.959964 the rest of the one-day table is measured against. 4.2996 falls 0.003 short. Nothing about the number is wrong; it was judged against a bar that belonged to a different sample size — the error Bailey and López de Prado (2014) built the deflated Sharpe ratio to catch, where the length of the track record behind a statistic is part of the threshold and not a footnote to it.

That gap is worth a long article not because 0.003 is large, but because the board prints no column that says so, and because the system that produced the number declined to promote, demote or score anything on the strength of it. The failure mode has a name. Bailey and López de Prado (2014) describe it as an undeflated ratio: a performance statistic reported without controlling for the number of trials behind it, the length of the track record, and the non-normality of the sample. The correction applied to this cell is the crudest possible version of their adjustment — the one that comes free with a t-table.

Autonomous trading: extreme-move alpha survives a liquidity cut

· 20 min read
Vadim Nicolai
Senior Software Engineer

Every robustness test is a confession. It names the failure mode its author fears most, then tries to kill it. The test behind this record was aimed at the most respectable fear in cross-sectional equity work: that a screen ranking stocks on how violently they trade is a small-cap artifact wearing a ranking's clothes. The surprise is not that the screen survived the knife. The surprise is how little blood the knife drew — and what that reveals about which statistics are worth robustness-testing in the first place.

The research board this record comes from screens thousands of names each day. One lane ranks them by intraday high-low range as a percentage of price and asks whether its top names concentrate five-day extremes: an up-tail of fifty percent over five days and a log-symmetric down-tail at minus one-third. The obvious objection is the one any quant makes on sight: rank on realised volatility and you surface the smallest, thinnest names that clear the gates, and those names move fifty percent for reasons that have nothing to do with the ranking being informative.

How often does anyone bother to test that objection instead of asserting it? An audit-oriented evidence map of 77 LLM-trading studies found that of the 19 that met a closed-loop evaluation bar, only 1 documents universe or survivorship handling at all (Xia et al., 2026). Universe handling is the unglamorous act this whole measurement exists to perform — and the literature audit says publishing it is the exception, not the default.

So here is the answer in the form the question deserves: the extreme-move lift persists after excluding the smallest, least-liquid names. It is not a small-cap illusion; it survives a liquidity filter. What matters is the size of that survival — and why the survival is so much larger than the standard small-cap story would predict.