The Research Ledger

Everything we tested. Especially what died.

Every strategy below was preregistered before we looked at a single price, judged at executable odds (the price you could actually get, at the size actually available) and confirmed on data it had never seen. Companies show you winners. Ledgers show you everything.

9
Sports modeled
70,000+
Markets measured
600+
Preregistered cells
~98%
Killed or parked
PREREGISTERFREEZEEXECUTABLE PRICESCONFIRM ON UNSEEN DATAPUBLISH EITHER WAY
Why publish this? Because the graveyard is the method. If a measurement process never kills anything, it isn't measuring; it's marketing. Ours kills almost everything, which is exactly why you can trust the little that survives. Live strategies appear here without their parameters; verdicts appear in full.

In-play "game state" edges: soccer, all score states

KILLED
JUL 2026 · ~300 preregistered tests · 9 half-time & minute-60/75 states × 6 bet sides × league tiers

Backing or laying draws, favorites, leaders and trailers at every common score state, judged at executable exchange prices with liquidity floors. Early measurements showed edges up to +43%. Then we fixed six measurement bugs one by one: stale prices, lookahead, missing event coverage. And every single edge evaporated. The market prices sustained game states correctly.

judge best-back/lay level-0, OPEN status only · gates event coverage, settlement integrity, conservative clock, £25 size floor · tests 301 across league tiers · expected false positives ≈3 · observed "survivors" 1 (sub-slice, fails pooled = noise) · result implied ≈ realized in every cell
Last-traded price is not evidence. If your backtest fills at prices nobody was offering, it's fiction with decimals.

Half-time 1-1, lay the draw

RESCINDED
JUL 2026 · n=709 executable entries · was our own favorite discovery (+8.6%)

Our most promising candidate. It passed kill-checks, survived by-league analysis, 5/5 positive months. At executable best-lay prices with integrity guards: −9.6%. We rescinded our own discovery the same week we made it.

The strategies that hurt most to kill are the ones most worth killing.

Cross-market arbitrage: darts correct-score, hockey regulation/moneyline

KILLED
JUL 2026 · 152 + 937 events with synchronized order books · exact payoff identities

If the sum of correct-score prices disagrees with the match-winner price, that's free money: mathematically exact, no model needed. We scanned every synchronized book: 1 executable lock in 152 darts events, 0 in 937 hockey events. The incoherence exists, but always inside the spread. The taker pays exactly the cost that erases the edge.

method minimum-cost replication portfolios, per-market commission, all legs sized at level-0 · hockey prerequisite settlement semantics verified per league: 0 rule mismatches in 969 pairs · residuals median +22.5pp (darts), +16pp (hockey), always on the non-executable side
Markets can be visibly "wrong" and still unexploitable. The spread is the market's lawyer.

Market-making rent on prematch flagship markets

NO-GO
JUL 2026 · 3,489 markets, 6 sports · conservative FIFO queue simulation

Wide spreads looked like rent for patient liquidity providers. Then we conditioned on the fills that actually reach your queue position. The flow that finds you is worse than the average flow: net markout ≈ zero in the best corners, negative elsewhere. The winner's curse eats the spread.

sim FIFO queue frozen at placement, only traded volume advances you, cancellations don't · metric 30s markout net of commission, clustered by market · coverage 3,489 markets, 0 parse errors, exclusions counted · caveat this is markout, not full P&L: parked, not "impossible"
Averages lie to makers. Condition on your fills or don't bother.

Snooker model edge, +18.7% ROI

VOID
2026 · the edge that taught us everything

A beautiful backtest: +18.7%, significant, consistent. It was data leakage: the model had quietly seen information from the future. Rebuilt clean: breakeven. This one incident shaped our entire measurement system: leak gates, immutability digests, and an automatic block that has since caught real contamination before it reached production.

A number that's too good is a bug until proven otherwise.

Darts model edge, +14.7% ROI

PARKED
2026 · real model, real ROI, judged by closing line value

The model was clean and the realized ROI was real. But its closing-line value was ≈ zero, meaning the market did not move toward our predictions. Realized profit without CLV is variance wearing a suit. Parked until the evidence says otherwise.

Judge every edge by CLV and ROI together. Either one alone will happily lie to you.

Our own soccer model vs the market

DEAD
JUL 2026 · n=8,500–13,900 bets per anchor · fully powered

We built a clean, leak-free soccer model. We tested it against real exchange prices at every entry time. It loses everywhere: net negative at every anchor, and on the draw side the market beats it with high significance. So it doesn't bet. A model that isn't allowed to bet is not a failure; a model that bets when it shouldn't is.

verdict basis walk-forward predictions vs owned tick-level archive · anchors t-18h / t-6h / t-2h / close · t-stats −4.4 to −6.3 · status model serves analytics only, betting disabled by rule
Most model shops would never run this test. None would publish it.

Handball state grid

UNPOWERED
JUL 2026 · frozen preregistration, executable judge, run once

Same hypothesis class as the soccer grid, preregistered before any price was seen. The exchange itself set the ceiling: only ~700 handball match markets exist in the window, and our frozen minimum was 100 observations per cell. Largest cell: 22. No cell was judged, and the floor was not lowered. Lowering a preregistered floor because the data came up short is how fake edges get born.

Sometimes the market is too small to even ask the question. Record that, and walk away.

Basketball halftime grid

NOT RUN
JUL 2026 · structural feasibility check, 40 markets sampled, no outcomes read

To test halftime states you need to know when halftime happened. Basketball exchange markets trade straight through the break: no suspension marker, no honest clock. Inferring the break from price action is exactly the class of measurement error that manufactured the soccer mirages. So we refused to run it.

The most dangerous backtest is the one you can't timestamp.

Tennis in-play, nine strategy families

KILLED
2026 · 13,000+ matches of tick data · break-chasing, comebacks, meltdowns, overshoots, reopenings

Nine families of in-play tennis strategies tested end-to-end. Several looked alive in discovery, including one with six consecutive positive months. The confirmation split killed them: mirage after mirage, plus one edge that was real but smaller than a single tick. In-play flow is informed; the taker side of tennis is a toll road.

data 13,432 matches, full ladder · discipline byte-identical double runs, content-addressed artifacts · survivors one maker-side hypothesis, now in forward validation (parameters private)
Discovery finds. Confirmation decides. Never let the same data do both jobs.

Our NCAA model, blocked by our own gate

BLOCKED
JUL 2026 · automatic leak gate, first production firing

A retrain came back with an AUC jump of +0.098 over its own history and a brand-new feature suddenly explaining 40% of its decisions. To a hype shop, that's a breakthrough. To our gate, it's the signature of contamination. Promotion blocked automatically, before a single prediction reached production.

trigger AUC jump beyond +0.04 band vs own-history median · second signal importance-dominance shift (new top feature 40% vs historical 16%) · action candidate quarantined, investigation queued, serving untouched
Our own models face the same police as everything else. Especially when the news is good.

Line-drift conditioned strategy: multiple sports

FORWARD TEST
LIVE · survived leave-one-sport-out, temporal out-of-sample, and CLV coherence

One of the very few survivors. It's now doing the only thing that counts: running forward on data no analysis has ever touched, against preregistered gates that will kill it without appeal if it underperforms. Parameters not published. Surviving edges stay private; verdicts don't.

Discovery is cheap. Forward evidence is the only expensive thing worth buying.