Skip to main content

The number a trading objective has to beat before it means anything: my system finally measured the base rate for a 10% move in five days, and it is forty-two times larger than anyone had measured

· 21 min read
Vadim Nicolai
Senior Software Engineer

For a year the board has run a cross-sectional screen over roughly 12,500 US-listed equities a day. It ranks names on momentum, reversal, short-interest flow and intraday range, then scores every lane against how the winners and losers actually moved. The pipeline is unremarkable — the same architecture-capability-adaptation stack that Xia et al. (2026) audit across 77 LLM-trading studies, minus the LLM. It has measured a great deal. It had never measured the one number its own standing objective rests on.

The objective, written down by its owner, reads:

identify equities that will rise 10% or more within five trading days, at a hit rate materially above the base rate for their own liquidity cohort.

"Above the base rate" presupposes that the base rate exists as a number. It did not. The board's working notes listed it, in plain text, under what has never been measured.

Today the loop measured it over the full panel: 5.46% — 148,798 name-days out of 2,723,465 that rose 10% or more inside five trading days. Roughly one name in eighteen. That phrasing — establish the unconditional property of the universe first, argue about strategies second — is what Taripe (2026) did for 28 currency pairs across ten years of daily closes: reject the unit-root null, then ask the Hurst exponent what it means. Nobody had done it here.

Forty-two times is the ratio against the only threshold anyone had ever committed to paper, the +50% tail at 0.130% — 933 hits in 719,767 name-days. That a bar alone can move a verdict that far is not a hypothesis. Hou, Xue and Zhang (2018) raise one test hurdle from |t| = 1.96 to 2.78 across 452 anomalies and watch the failure rate climb from 65% to 82%.

An objective with an unmeasured term is not a weak objective. It is an unfalsifiable one. Unfalsifiable objectives are exactly the ones that survive a year of review, because nothing in the workflow compels anybody to put a number where the placeholder is. That clause read like a specification. It functioned as a placeholder wearing a specification's clothes.

The Literature Audits the Models and Skips the Specification

The field has built an impressive apparatus for asking whether an agent performs, and almost none for asking what it has to outperform. Xia et al. (2026) screened the LLM-trading literature through 2026-03-09 and retained 77 studies, of which 19 met the minimum bar of action output plus closed-loop evaluation. Within that primary subset, only 2 of 19 report an extractable time-consistent split protocol. Exactly 1 of 19 reports an explicit transaction-cost model. Exactly 1 documents universe or survivorship handling. Eleven of 19 report execution timing or semantics at all. Fifteen of 19 are coded R0 for reproducibility, and no study reaches R3.

That is the board's problem one level up. Not "we cannot compare two models," but "we cannot say what either model had to beat." The paper's own framing is protocol incomparability, and it is careful to call its Architecture-Capability-Adaptation lens a working lens rather than a validated taxonomy. What it demonstrates instead is an evidence ledger, a reproducibility audit and a reporting checklist — the plumbing that makes a claim checkable at all.

Saha et al. (2025) review the same territory from the investment-management side, categorising the literature across four use cases — portfolio optimization, risk management, information retrieval and automated strategy generation — and naming robustness, explainability and real-world deployment as the open problems. Deployment is a denominator question before it is anything else. You cannot deploy against a bar you never measured, and you cannot distinguish a robust agent from a lucky one when both are being graded against nothing.

The model was never the missing piece here. The specification was.

Forty-Two Times Is Not a Bigger Number, It Is a Different Question

The temptation is to read 5.46% against 0.130% and conclude that the 10% target is "a bigger number." It is not. At the +50% tail a hit lands once in roughly 770 name-days, so a screen reporting a 1% hit rate beats that bar nearly eightfold and looks like a discovery. At the 10% threshold a hit lands once in eighteen, and the identical 1% screen is nowhere near the floor. Same panel, same signal, same claimed hit rate, opposite verdicts — the hurdle did all of that work, a point Hou, Xue and Zhang (2018) make with anomaly counts rather than hit rates.

Their numbers are the cleanest external version of the argument. Against 452 anomalies, with microcaps mitigated via NYSE breakpoints and value-weighted returns, 65% cannot clear the single test hurdle of an absolute t-value of 1.96 — including 96% of the trading-frictions category. Impose the higher multiple-test hurdle of 2.78 at the 5% significance level and the failure rate climbs to 82%. Same data library, seventeen points of headline movement, produced entirely by the bar.

Had this board kept the +50% tail as its standing threshold, a lane reporting a 1% hit rate would have been funded as a near-eightfold edge. The real question would never have been asked. One number would have silently decided the call, and it would have been the wrong number.

The Liquidity Gradient Is Real, Monotone, and It Points Straight at the Most Expensive Names to Trade

The objective names liquidity cohorts, so a pooled number was never going to be enough. Split the panel into dollar-volume deciles and the base rate for a 10% move in five days starts at 7.11% in the lowest-dollar-volume decile, then falls continuously — 6.39%, 6.15%, 6.01%, 5.62%, 5.26%, 4.73%, 4.43%, 4.14% — down to decile nine, before ticking back up to 4.82% in the most-liquid decile. The cheapest third of the universe rises 10% about 1.6× as often as the dearest third. A decile ladder is the microcap control Hou, Xue and Zhang (2018) impose as NYSE breakpoints and value-weighted returns, run as a measurement instead of a filter.

Read it fast and it looks like a free lunch with a bow on it: the names that move are the cheap names, so trade the cheap names. That reading dies twice.

First, the cheap names are where touching the market costs the most. Fesenko (2026) formalises the shape of that shortfall for a neighbouring problem. Backtests that assume zero-latency fills record the price differential at signal time, while production pays a deterministic round-trip delay T during which the window can close, the differential can decay toward zero, or both. Realised edge per opportunity is strictly smaller than backtest edge, and the gap grows with T.

The paper derives a retention ratio R(T) under three decay regimes and three duration distributions, solves for the viability threshold T* where expected per-trade edge equals expected per-trade cost, and validates the closed forms against a Monte Carlo calibrated to the BEQI broker-execution dataset. Put crudely: the extra 2.3 points of base rate sitting in the cheapest decile is a denominator, not an edge, until someone shows it survives the cost of touching those names. Nobody on this board has shown it.

Second, the decile assignment is only as good as the volume figure behind it. Zwydak et al. (2026) apply complexity and statistical-structure measures to high-frequency trade-level data for BTC, ETH and XRP across Binance, Bitget, KuCoin and Kraken between 1 April and 30 June 2025, and find a pronounced anomaly on Bitget for BTC and ETH after mid-May: transaction counts increase sharply while traded volume and return fluctuations show no proportional increase. Reported activity was partly noise-shaped. A liquidity decile built on a venue-reported volume number is a proxy, and this is what a proxy looks like when it breaks.

Had 7.11% been adopted as the bar for the cheapest cohort, the board would have launched a lane whose entire measured edge was the spread it paid to enter.

My Holdout Refused To Be Conservative

This is the part the board had wrong, and it is the part with the most consequence.

The intuition about sealed holdouts is that they are a tax. You carve out a tail of the panel, and you expect the out-of-sample arm to come in slightly worse or flat, because that is what holdouts do to people. Sheppert (2026) builds an objective function around exactly that failure mode — data-driven strategies learning spurious patterns that collapse out of sample — and prices it across 50 S&P 500 companies spanning 2010–2024, nine sequential walk-forward splits, and a Monte Carlo study over 15 random seeds and three trading strategies. His composite GT-Score, which folds performance, statistical significance, consistency and downside risk into one objective, improved the generalization ratio — validation return divided by training return — by 98% relative to baseline objective functions, with paired out-of-sample differences detectable at p < 0.01 and small effect sizes.

The board's holdout went the other way. The loop pre-registered a sealed tail, ran the pre-cut arm and the sealed arm as separate measurements, and the sealed arm came in at 6.35% against the pre-cut arm's 5.01% — more than a point higher, in the opposite direction from the conservative drift Sheppert (2026) optimises against.

That is a panel fact, not a claim about regimes, and I am not dressing it as one. The operational consequence is the part that matters. Any out-of-sample lift a future screen reports has to be measured against that 6.35% denominator, not the pooled 5.46%, or the screen flatters itself by 0.89 points for free. At a reported 6% hit rate, those 0.89 points are the entire difference between clearing the bar and missing it. Grade yourself against the pooled number and you are sitting the easier exam while telling yourself you pre-registered.

A Base Rate Has No Ranking Dates — Which Names the Overlap Problem Rather Than Solving It

Every other number this board reports is a spread across names chosen on a date, and that structure forces a correction. Consecutive five-day windows overlap, so observations are not independent in the way the raw sample size implies, and the board has measured the required standard-error inflation at 1.13×–1.96× depending on convention — a spread in the same family as the convention effects Zhang et al. (2026) isolate by toggling one evaluation choice at a time.

A base rate looks like it sidesteps that. Every name-day is one Bernoulli observation: did this name rise 10% in the next five days, yes or no. The independent-observation standard error is therefore the natural one, and it is what the board reports. But consecutive name-days on the same name share overlapping forward windows, so the independence assumption is named and set aside, not escaped. The board forgoes the overlap correction for this number and says so out loud, which is a different thing from never having noticed the problem. Bolt an overlap correction onto a base rate and you inflate a bar that does not need inflating; skip it on a ranked lane score and every lane looks significant.

What makes the convention question concrete rather than philosophical is the one-switch benchmark. Zhang et al. (2026) toggle one evaluation convention at a time around a clean t+1-open reference while holding the data panel, walk-forward split, model family, horizon, portfolio rule and cost convention fixed, across two daily-OHLCV equity panels, six model families, and yearly tests from 2016 to 2024. The inflation is highly selective. Centered temporal features and same-day-open execution with post-open daily-bar information produce large, stable increases in both predictive and trading metrics. Global normalization, future-informed graph structure and same-day-close execution are weak in most settings. Same model, same panel, same horizon — and the convention picks the answer.

It is not only backtests. Yuan et al. (2024) run three experiments with different task instructions on the same safety-judgment benchmark and report the resulting standard deviations of F1 and consistency. Task instructions affect performance only slightly, but larger models show higher consistency and lower standard deviation. The instruction is not the model. It is the convention wrapped around the model, and it moves the number.

The Most Dangerous Line in a Report Is a Citation Sitting Next to a Number the Cited Paper Never Produced

I got this wrong on the first pass through the material, and it is worth stating plainly, because it is the failure mode that makes a self-measurement look externally validated when it is not.

The 42× ratio is this board's arithmetic on this board's panel. It would be easy, and wrong, to hang Hou, Xue and Zhang (2018) on it. HXZ is a replication study of 452 anomalies against t-value hurdles. It says nothing about a +50% tail, nothing about a 42× ratio, and nothing about 10% moves in five days. A citation there would import authority the paper never earned — and it would be doubly invisible, because the numbers in the sentence are real.

The same trap opens twice more. The finding that the base rate shifts with universe, horizon and threshold belongs to nobody but the board; Xia et al. (2026) is about protocol incomparability across 19 primary studies, not about the sensitivity of a return threshold. And the sealed-arm result, 6.35% against 5.01%, is a panel fact; Sheppert (2026) is about overfitting collapse out of sample, measured through a 98% improvement in generalization ratio across nine walk-forward splits. Attaching his name to my holdout spread would be a category error dressed up as diligence.

The rule I now apply is narrow and mechanical. A citation earns its place only when the number in the sentence is the source's number. Every other figure — everything the system produced itself — carries no citation, because there is nothing to attach it to. A borrowed reference manufactures the appearance of external replication, and a reader who cannot tell which numbers were measured here and which were measured elsewhere cannot audit either.

Two of the sources I lean on hardest are also the two with the weakest standing. Fesenko (2026) sits on a vendor domain, and both it and Sheppert (2026) are pre-prints rather than peer-reviewed publications. Both are used here as structural arguments about how cost and generalization gaps behave, not as validated constants to be plugged in.

A Benchmark Is Only as Good as the Denominator It Scores Against

The base rate matters past one board because it is the missing half of every benchmark in this field.

Qian et al. (2025) make the point sharpest. Agent Market Arena is a lifelong, real-time benchmark for LLM trading agents that runs four architectures — InvestorAgent as a single-agent baseline, TradeAgent and HedgeFundAgent with different risk styles, and DeepFundAgent with memory-based reasoning — across five model backbones including GPT-4o, GPT-4.1, Claude-3.5-haiku, Claude-sonnet-4 and Gemini-2.0-flash, on live cryptocurrency and stock markets. Their most portable finding is that agent frameworks display markedly distinct behavioural patterns, from aggressive risk-taking to conservative decision-making, while model backbones contribute less to outcome variation. Set that next to an unmeasured base rate and it turns strange. If architecture matters more than the model, and the floor was never established, the benchmark is ranking behaviour against nothing in particular.

Fan et al. (2025) push harder. AI-Trader is a fully automated, live, data-uncontaminated benchmark spanning US stocks, A-shares and cryptocurrencies across multiple trading granularities, and it evaluates six mainstream LLMs. Their finding is that general intelligence does not translate automatically into trading capability — most agents post poor returns and weak risk management — and that risk-control capability determines cross-market robustness. It is the right conclusion reached against the wrong kind of floor. Poor relative to what? A 10%-in-five-days objective with a 5.46% pooled base rate and a 6.35% out-of-sample bar gives "poor" a number instead of an adjective.

Without a floor, "does this work" quietly becomes "does this look better on the metric I chose." Poudel and Paudel (2025) build a rule-based, long-only strategy for the Nepal Stock Exchange, combining a Z-score, RSI and a 240-day moving average with dynamic position sizing, trade limits and cool-down periods, then benchmark it against Buy & Hold using CAGR, Sharpe, Sortino, maximum drawdown, win rate, profit factor and recovery factor. The strategy does not significantly outperform in raw daily returns, but it does post higher Sharpe and Sortino ratios and significantly lower drawdowns. Which is it — a failure or a success? Raw return says failure, risk-adjusted return says success, and with no pre-registered bar the metric quietly becomes the finding.

Practical Takeaways: Compute the Bar, Do Not Borrow It

The reflex after reading a number like 5.46% is to use it. Do not. Use the method.

Compute your own denominator first. Mohammadi et al. (2025) organise agent evaluation along two axes — evaluation objectives (what to evaluate) and evaluation process (how, including interaction modes, datasets, metric computation and tooling) — and the base rate is the "what" that precedes the "how." An objective without it is not being evaluated on either axis.

Split by the cohort your objective names. A pooled 5.46% hides a span from 7.11% down to 4.14% across dollar-volume deciles — the microcap problem Hou, Xue and Zhang (2018) handle with NYSE breakpoints and value-weighted returns.

Pre-register the holdout, then believe it. The pre-cut arm sat at 5.01% and the sealed arm at 6.35%, more than a point higher in the direction that costs you. Grade against the pooled number and you are sitting the easier exam — an out-of-sample error of the kind Sheppert (2026) spends nine walk-forward splits trying to price.

Name your error convention out loud. A 1.13×–1.96× overlap inflation is not a rounding error, and neither is omitting it. That convention-dependence is what Zhang et al. (2026) isolate by toggling one choice at a time across six model families and two equity panels.

Add costs before you add the word "edge." A 1.6× frequency advantage in the cheapest third of the universe stays a denominator until a cost model survives contact with it — the R(T) computation Fesenko (2026) formalises with its viability threshold T*.

Distrust the liquidity label as much as the price. Zwydak et al. (2026) found transaction counts on one venue rising sharply with no proportional increase in traded volume, which is exactly the noise that can move a name across a decile boundary.

Check the citation before you trust the number. If the sentence's number is not the cited paper's number, the citation is decoration, and decoration is worse than a blank.

What is a base rate in trading?

It is the unconditional frequency of an outcome — how often a move of a given size happens at all, across all periods measured. It is the benchmark any objective, signal or strategy must beat before it carries information, and it is the term objectives tend to leave unmeasured.

How often does a stock move 10% in five days?

On this panel: 5.46% of name-days, or one in eighteen, across roughly 12,500 US-listed names, a five-trading-day horizon, and a fixed 10% threshold. Those three choices are what the number is conditional on. Xia et al. (2026) found that of 19 primary LLM-trading studies, only 2 report an extractable time-consistent split protocol — so two quoted frequencies can differ for reasons that have nothing to do with the market.

Why is 5.46% so much larger than the +50% threshold number?

Because the two are different events, not one event at two resolutions. The +50% tail fires 0.130% of the time — 933 of 719,767 name-days — which makes the 10% target roughly 42× more common. Threshold choice moves results more than most people expect: Hou, Xue and Zhang (2018) show that raising a single hurdle from |t| = 1.96 to 2.78 moves the failure rate across 452 anomalies from 65% to 82%.

Should you use the 5.46% figure or compute your own?

Compute your own. The sealed-arm spread on this same panel — 6.35% against 5.01% — is wide enough to move a screen across the bar by itself. Objectives tuned on one sample and validated on another are precisely what Sheppert (2026) prices, and a 98% generalization-ratio improvement across nine walk-forward splits is a measure of how much that gap usually costs.

The Autonomous Part Is Not That It Measured. It Is That It Was Forbidden To Conclude.

An autonomous loop asked itself a question its own objective presupposed, answered it end to end over 2,723,465 name-days, split the answer by the exact cohort the objective names, carried a named standard-error convention, and pre-registered an out-of-sample arm. Then it was forbidden to write a verdict. The record carries the number and a list of caveats; the question of whether any screen clears the bar is left to a deterministic follow-up.

That refusal is the load-bearing part. Fang et al. (2025) abstract self-evolving agents into four components — System Inputs, Agent System, Environment and Optimisers — and review technique after technique for evolving each one, naming lifelong evaluation protocols and safe coordination in heterogeneous environments as the open problems. Ding et al. (2024) catalogue the common architecture, data inputs and backtested performance of LLM trading agents, and list the challenges running through them. In the material reviewed here, the surveys propose the behaviour; they do not report an instance of an agent measuring its own denominator first and then declining to grade itself. An agent that measures a base rate and immediately announces that its own model beats it has not run a test. It has run a press release.

Let me be honest about what that base rate is. It is not good news, and it is not a discovery that helps anyone trade. It is a smaller number than the objective's owner probably imagined, and a larger one than the field's only committed threshold, and both of those are true at once. A name in eighteen clears a 10% move with no skill involved. Every screen pointed at that target now has a bar it cannot talk its way around.

The reason this took a year has nothing to do with difficulty. The query is one line over a panel the board already had. It took a year because the objective's wording felt complete.

That is the general lesion, and it is not specific to trading. The hard part of building a system that measures itself is not the measuring. It is designing the objective so that the unmeasured term becomes a defect the system cannot help but fall into. This board fell into it, measured 2,723,465 name-days on the way down, and came back with 5.46% — plus a sealed arm that said 6.35%, a list of caveats, and no verdict. Xia et al. (2026) find 15 of 19 primary studies coded R0 and none reaching R3, which is the audit's way of saying the same thing. The next question — whether any model beats the bar — is only now, for the first time, a question that can be answered.