oddsx status
Pre-registered verdicts (TEST split)
| hypothesis | verdict | n_bets | mean_net | edge_vs_naive | ci95_low | dossier |
|---|---|---|---|---|---|---|
fundamental-arb-exec | FALSIFIED | 516 | -0.2449 | -0.1945 | -0.3901 | dossier |
fundamental-arb-short | VALIDATED | 1228 | +0.1449 | +0.1550 | +0.0097 | dossier |
fundamental-arb | INSUFFICIENT_DATA | 15 | -0.6679 | -0.7089 | -1.2193 | dossier |
Protocol verdicts as computed; read the SCORECARD and dossiers for caveats.
Dossiers
- drift-to-resolution
- favorite-longshot
- fundamental-arb-exec
- fundamental-arb-short
- fundamental-arb
- maker-economics
- toxic-flow
SCORECARD
Edge SCORECARD — powered re-run (T-0015)
All three edges re-run on the FULL resolved-sports window (build_resolved_bets_full — 29,651,775 bets, 83,019 markets, 2026-04-18 → 06-11) with a leak-free calendar split (calendar_split): TRAIN = Apr18→May26, TEST = a genuine 15-day June OOS era (May27→Jun11). Standard control = the FAVORITE null (the honest, cost-artifact-free control per T-0012), same Polymarket cost model (fee_rate=0.05, spread=0.02), equal-notional sizing, returns capped ≈[−1, +1]. Band/threshold derived on TRAIN only. edge_vs_naive = strategy mean_net − favorite-null mean_net (costs cancel → pure selection skill). Verdict = TEST edge_vs_naive > 0 and TEST mean_net > 0.
| edge | TRAIN edge_vs_naive | TEST edge_vs_naive (net) | Sharpe (TEST, bet-level) | N_test | verdict |
|---|---|---|---|---|---|
| Edge-1 · favorite-longshot (take favorite in band [0.50,0.95)) | −0.0021 | −0.0032 | −0.0588 | 8,844,141 | FALSIFIED |
| Edge-2 · drift-to-resolution (momentum, entered late) | −0.0219 | −0.0311 | +0.0219 | 12,717 | FALSIFIED |
| Edge-3 · toxic-flow (follow-flow, band [0.00,0.10)) | −0.7497 | −0.5516 | −0.2595 | 2,717,883 | FALSIFIED |
All three stay FALSIFIED on the powered window. The favorite null (buy the favorite) is undefeated: no edge beats simply backing the favored side, net of costs.
Notes per edge
- Edge-1 (favorite-longshot): the gross calibration bias is real (~+1% for mid favorites) but ~2–3× smaller than the ~3% taker cost; net-negative in every price bucket. More robust than provisional: the sign no longer flips between all-fills and per-market (both negative), and the provisional "longshot mirage" (+0.28) is gone (−0.0075).
- Edge-2 (drift / CLV): falsified as a tradeable edge — momentum's TEST net is marginally positive (+0.0153) but is entirely a fat-tail artifact (median −0.0070; a single +65× outcome; top-1% carries 19.6% of |PnL|), and it loses to buying the late favorite (+0.0464). Separately, closing-line value is strongly CONFIRMED: late-price Brier 0.022 vs early-price 0.201 (TEST) — the information is real but already priced by entry time.
- Edge-3 (toxic-flow): falsified decisively — the follow-flow return *worsens* monotonically as flow gets *more* one-sided (opposite of the thesis). The provisional "+1.45 edge vs uniform" was a null-is-worse trap; vs the sharp favorite null the edge is a clean −0.55. OP's "wide-spread = retail/safe" is FALSE (capped taker outcome flat across spread quintiles; near-close flow is if anything *more* informed).
Tail-honesty summary (median vs mean; top-1% |PnL| share, TEST)
| edge | median_net | mean_net | top-1% |PnL| share | read |
|---|---|---|---|---|
| Edge-1 | +0.2275 | −0.0431 | 0.016 | favorites win often, pay little; −100% losses sink the mean; no tail dominance |
| Edge-2 | −0.0070 | +0.0153 | 0.196 | mean is a tail illusion — median loses; one +65× drives the positive mean |
| Edge-3 | −1.0374 | −0.5881 | 0.084 | uncapped gross +0.010 is a right-tail mirage; capped+cost net −0.59 |
Verdict
No edge flipped to validated on the powered window — all three remain FALSIFIED, and the falsifications are now backed by a real 15-day OOS era rather than the provisional ~1-day tail. Honest > profitable: the sharp favorite null survives every challenge, closing-line value is confirmed as *information* (not a tradeable edge), and the toxic-flow / wide-spread retail folklore is refuted in the data. Every signal being real-but-sub-taker-cost points away from a taker edge — the maker side (post limit orders, capture the bias as spread, pay no taker fee) is the natural next question.
Reproduce (each pastes its full stdout):
uv run --extra backtest python research/02_favorite_longshot.py # Edge-1
uv run --extra backtest python research/04_drift.py # Edge-2
uv run --extra backtest python research/03_toxic_flow.py # Edge-3
uv run pytest -q → green.
Maker economics (T-0016, added iter-006)
| Strategy | TEST capped net | TEST uncapped | verdict |
|---|---|---|---|
| Naive maker (fee=0, quote everything) | −0.135 | +0.013 (tail-only) | FALSIFIED (net-neg equal-notional; spread is cheap-token tail) |
| Toxic-flow-filtered maker | −0.078 | +0.009 | FALSIFIED (halves bleed, never clears 0) |
Fee asymmetry CONFIRMED (maker beats taker by +0.26/fill — the fee is the only wedge), and the adverse-selection gradient is real (maker net worsens monotonically with flow one-sidedness). But even at 0 fee, and as an OPTIMISTIC upper bound (assumes every passive fill wins), making is net-negative. Conclusion: no harvestable edge from historical fills — taker OR maker. The one untested candidate is polymm's actual edge = maker WITH FRESH SPORTSBOOK ODDS (de-vig), which needs Track C forward data.
fundamental_arb (G4, pre-registered, Polymarket price-threshold markets)
Pre-registration and full results: dossiers/fundamental-arb.md. Calendar split on decision time: TRAIN 2026-04-10 → 06-30, TEST 2026-07-01 → 08-21. Benchmark = favorite null, same costs. Verdict rule: fewer than 200 TEST bets → INSUFFICIENT_DATA. Realized-vol-only limited test.
| edge | TRAIN edge_vs_naive | TEST edge_vs_naive (net) | Sharpe (TEST, bet-level) | N_test | verdict |
|---|---|---|---|---|---|
| fundamental_arb | +0.4851 | −0.7089 | −0.6406 | 15 | INSUFFICIENT_DATA |
fundamental_arb_short (G4b, pre-registered separately, Polymarket price-threshold markets)
Pre-registration and full results: dossiers/fundamental-arb-short.md. Decision 1 hour before end (at least 30 minutes to resolution). Calendar split on decision time: TRAIN 2026-04-10 → 06-30, TEST 2026-07-01 → 08-21. Benchmark = favorite null, same costs. Same verdict rule as G4 (200-bet floor; mean net > 0, edge_vs_naive > 0 and cluster-bootstrap 95% lower bound > 0). Realized-vol-only limited test.
| edge | TRAIN edge_vs_naive | TEST edge_vs_naive (net) | Sharpe (TEST, bet-level) | N_test | verdict |
|---|---|---|---|---|---|
| fundamental_arb_short | +1.3321 | +0.1550 | +0.0624 | 1,228 | VALIDATED (protocol) — likely stale-price artifact, not executable; see dossier |
The protocol verdict stands as computed, but post-hoc diagnostics (not part of the pre-registration) point to stale YES prices in thin hourly markets: the TEST 95% lower bound is only +0.0097, the edge sits where the last trade is old, and the freshest-priced bets lose. Not promoted; a follow-up with executable entry prices is required before any paper trading.
fundamental_arb_exec (G4c, pre-registered separately, Polymarket price-threshold markets)
Pre-registration and full results: dossiers/fundamental-arb-exec.md. Same signal, universe and decision time as fundamental_arb_short, but each bet is filled at the first executable trade at or after the decision (within 10 minutes); returns, costs and the favorite null use that fill price. Calendar split on decision time: TRAIN 2026-04-10 → 06-30, TEST 2026-07-01 → 08-21. Same verdict rule (200-bet floor; mean net > 0, edge_vs_naive > 0 and cluster-bootstrap 95% lower bound > 0). Realized-vol-only limited test.
| edge | TRAIN edge_vs_naive | TEST edge_vs_naive (net) | Sharpe (TEST, bet-level) | N_test | verdict |
|---|---|---|---|---|---|
| fundamental_arb_exec | +0.0082 | −0.1945 | −0.1457 | 516 | FALSIFIED — executable fills remove the G4b edge; see dossier |
At the executable fill, TEST mean net is −0.2449 and the 95% lower bound −0.3901. Post-hoc diagnostics (not part of the pre-registration): the same 516 filled bets earn +0.179 net at the stale signal price and −0.245 at the fill, and 67% of fills are worse than the signal price. fundamental_arb_short's result was a stale-price artifact. No fundamental_arb variant is promoted or paper-traded.
poly_data lake
| dataset | size | files | rows | latest month | latest timestamp | age | unreadable |
|---|---|---|---|---|---|---|---|
| markets | 687.2 MiB | 1 | 2,589,274 | — | 2026-08-21T08:59:22Z | 24.0 d | 0 |
| markets_raw | 87.6 MiB | 1 | 3,496,613 | — | — | — | 0 |
| orders | 12.73 GiB | 5 | 540,377,085 | 2026-08 | 2026-08-21T09:07:27Z | 24.0 d | 0 |
| trades | 13.77 GiB | 5 | 540,377,085 | 2026-08 | 2026-08-21T09:07:27Z | 24.0 d | 0 |
Total 27.26 GiB · /home/trbck/workspace/_repos/poly_data/processed/parquet · sizes, rows and timestamps from Parquet footers (no data scans).
Verification
oddsx boil verification — 2026-09-14 08:49
39/39 frozen checks pass.
| goal | milestone | exit | seconds | last line |
|---|---|---|---|---|
| G1 | M1 | 0 | 0.3 | 4 passed in 0.05s |
| G1 | M2 | 0 | 0.3 | 3 passed in 0.04s |
| G1 | M3 | 0 | 14.8 | 43 passed in 14.31s |
| G1 | M4 | 0 | 0.4 | 41 passed in 0.11s |
| G1 | M5 | 0 | 0.4 | 8 passed in 0.05s |
| G1 | M6 | 0 | 2.6 | 3 passed in 2.31s |
| G1 | M7 | 0 | 0.3 | 3 passed in 0.02s |
| G3 | G3M1 | 0 | 0.3 | 3 passed in 0.01s |
| G3 | G3M2 | 0 | 37.5 | 20 passed in 37.04s |
| G3 | G3M3 | 0 | 3.2 | 10 passed in 2.71s |
| G3 | G3M4 | 0 | 85.3 | 6 passed in 84.82s (0:01:24) |
| G3 | G3M5 | 0 | 4.3 | 7 passed in 3.87s |
| G3 | G3M6 | 0 | 14.3 | 8 passed in 0.05s |
| G4 | G4M1 | 0 | 101.1 | 7 passed in 100.54s (0:01:40) |
| G4 | G4M2 | 0 | 0.8 | 4 passed in 0.39s |
| G4 | G4M3 | 0 | 0.9 | 4 passed in 0.49s |
| G4 | G4M4 | 0 | 0.8 | 5 passed in 0.40s |
| G4 | G4M5 | 0 | 16.1 | 10 passed in 2.16s |
| G4b | G4bM1 | 0 | 102.4 | 7 passed in 101.79s (0:01:41) |
| G4b | G4bM2 | 0 | 1.0 | 4 passed in 0.43s |
| G4b | G4bM3 | 0 | 1.3 | 4 passed in 0.84s |
| G4b | G4bM4 | 0 | 1.0 | 5 passed in 0.44s |
| G4b | G4bM5 | 0 | 18.0 | 13 passed in 0.60s |
| G4c | G4cM1 | 0 | 102.1 | 9 passed in 101.52s (0:01:41) |
| G4c | G4cM2 | 0 | 0.8 | 4 passed in 0.38s |
| G4c | G4cM3 | 0 | 1.0 | 4 passed in 0.55s |
| G4c | G4cM4 | 0 | 0.8 | 5 passed in 0.37s |
| G4c | G4cM5 | 0 | 23.1 | 26 passed in 1.23s |
| G2 | G2M1 | 0 | 2.4 | 5 passed in 2.00s |
| G2 | G2M2 | 0 | 0.9 | 3 passed in 0.45s |
| G2 | G2M3 | 0 | 1.3 | 36 passed in 0.88s |
| G2 | G2M4 | 0 | 1.7 | 17 passed in 1.26s |
| G2 | G2M5 | 0 | 1.6 | 7 passed in 1.15s |
| G2 | G2M6 | 0 | 1.3 | 7 passed in 0.87s |
| G2 | G2M7 | 0 | 1.2 | 7 passed in 0.75s |
| G5 | G5M1 | 0 | 0.3 | 5 passed in 0.02s |
| G5 | G5M2 | 0 | 0.4 | 3 passed in 0.08s |
| G5 | G5M3 | 0 | 0.3 | 3 passed in 0.02s |
| G5 | G5M4 | 0 | 28.5 | 124 passed in 8.12s |
Goals ladder
| goal | one-line |
|---|---|
| G1 | oddsx is the single merged repo — both parents' history grafted, oddstrading restructured, oddsoddy archived with triage cards, rules enforced, parents frozen. |
| G2 | the reusable pure logic of oddsoddy (title blocklist, price and trade signals, wallet PnL, extra wallet metrics, syndicates) lives in oddsx/core, proven equivalent to the archived originals on seeded fixtures. |
| G3 | the poly_data lake is compacted to Parquet, verified month by month against a frozen oracle, read by oddsx only through data/polydata.py, cut to a 14-day raw window, and kept that way by a daily retention timer. |
| G4 | fundamental_arb gets an honest, pre-registered out-of-sample verdict on Polymarket price-threshold markets (2026-04 → 08), recorded in the SCORECARD — whatever the verdict is. |
| G4b | fundamental_arb_short (decision 1 hour before end) gets an honest, separately pre-registered out-of-sample verdict on Polymarket price-threshold markets (2026-04 → 08), recorded in the SCORECARD — whatever the verdict is. |
| G4c | fundamental_arb_exec (fundamental_arb_short's signal, filled at the first executable trade) gets an honest, separately pre-registered out-of-sample verdict — whatever it is — settling whether the G4b edge was a stale-price artifact. |
| G5 | oddsx reads its secrets from its own gitignored .env through one loader, and nothing in oddsx points at the frozen oddsoddy repo anymore. |
Retention log (last 10)
| ts | applied | refused | parquet after | evicted |
|---|---|---|---|---|
| 2026-09-13T22:19:30+00:00 | True | False | 26.50 GiB | 0 |