oddsx · read-only research status
Generated 2026-09-14T09:00:03Z · HEAD bbbe072

oddsx status

Pre-registered verdicts (TEST split)

hypothesisverdictn_betsmean_netedge_vs_naiveci95_lowdossier
fundamental-arb-execFALSIFIED516-0.2449-0.1945-0.3901dossier
fundamental-arb-shortVALIDATED1228+0.1449+0.1550+0.0097dossier
fundamental-arbINSUFFICIENT_DATA15-0.6679-0.7089-1.2193dossier

Protocol verdicts as computed; read the SCORECARD and dossiers for caveats.

Dossiers

SCORECARD

Edge SCORECARD — powered re-run (T-0015)

All three edges re-run on the FULL resolved-sports window (build_resolved_bets_full29,651,775 bets, 83,019 markets, 2026-04-18 → 06-11) with a leak-free calendar split (calendar_split): TRAIN = Apr18→May26, TEST = a genuine 15-day June OOS era (May27→Jun11). Standard control = the FAVORITE null (the honest, cost-artifact-free control per T-0012), same Polymarket cost model (fee_rate=0.05, spread=0.02), equal-notional sizing, returns capped ≈[−1, +1]. Band/threshold derived on TRAIN only. edge_vs_naive = strategy mean_net − favorite-null mean_net (costs cancel → pure selection skill). Verdict = TEST edge_vs_naive > 0 and TEST mean_net > 0.

edgeTRAIN edge_vs_naiveTEST edge_vs_naive (net)Sharpe (TEST, bet-level)N_testverdict
Edge-1 · favorite-longshot (take favorite in band [0.50,0.95))−0.0021−0.0032−0.05888,844,141FALSIFIED
Edge-2 · drift-to-resolution (momentum, entered late)−0.0219−0.0311+0.021912,717FALSIFIED
Edge-3 · toxic-flow (follow-flow, band [0.00,0.10))−0.7497−0.5516−0.25952,717,883FALSIFIED

All three stay FALSIFIED on the powered window. The favorite null (buy the favorite) is undefeated: no edge beats simply backing the favored side, net of costs.

Notes per edge

Tail-honesty summary (median vs mean; top-1% |PnL| share, TEST)

edgemedian_netmean_nettop-1% |PnL| shareread
Edge-1+0.2275−0.04310.016favorites win often, pay little; −100% losses sink the mean; no tail dominance
Edge-2−0.0070+0.01530.196mean is a tail illusion — median loses; one +65× drives the positive mean
Edge-3−1.0374−0.58810.084uncapped gross +0.010 is a right-tail mirage; capped+cost net −0.59

Verdict

No edge flipped to validated on the powered window — all three remain FALSIFIED, and the falsifications are now backed by a real 15-day OOS era rather than the provisional ~1-day tail. Honest > profitable: the sharp favorite null survives every challenge, closing-line value is confirmed as *information* (not a tradeable edge), and the toxic-flow / wide-spread retail folklore is refuted in the data. Every signal being real-but-sub-taker-cost points away from a taker edge — the maker side (post limit orders, capture the bias as spread, pay no taker fee) is the natural next question.

Reproduce (each pastes its full stdout):

uv run --extra backtest python research/02_favorite_longshot.py   # Edge-1
uv run --extra backtest python research/04_drift.py               # Edge-2
uv run --extra backtest python research/03_toxic_flow.py          # Edge-3

uv run pytest -q → green.

Maker economics (T-0016, added iter-006)

StrategyTEST capped netTEST uncappedverdict
Naive maker (fee=0, quote everything)−0.135+0.013 (tail-only)FALSIFIED (net-neg equal-notional; spread is cheap-token tail)
Toxic-flow-filtered maker−0.078+0.009FALSIFIED (halves bleed, never clears 0)

Fee asymmetry CONFIRMED (maker beats taker by +0.26/fill — the fee is the only wedge), and the adverse-selection gradient is real (maker net worsens monotonically with flow one-sidedness). But even at 0 fee, and as an OPTIMISTIC upper bound (assumes every passive fill wins), making is net-negative. Conclusion: no harvestable edge from historical fills — taker OR maker. The one untested candidate is polymm's actual edge = maker WITH FRESH SPORTSBOOK ODDS (de-vig), which needs Track C forward data.

fundamental_arb (G4, pre-registered, Polymarket price-threshold markets)

Pre-registration and full results: dossiers/fundamental-arb.md. Calendar split on decision time: TRAIN 2026-04-10 → 06-30, TEST 2026-07-01 → 08-21. Benchmark = favorite null, same costs. Verdict rule: fewer than 200 TEST bets → INSUFFICIENT_DATA. Realized-vol-only limited test.

edgeTRAIN edge_vs_naiveTEST edge_vs_naive (net)Sharpe (TEST, bet-level)N_testverdict
fundamental_arb+0.4851−0.7089−0.640615INSUFFICIENT_DATA

fundamental_arb_short (G4b, pre-registered separately, Polymarket price-threshold markets)

Pre-registration and full results: dossiers/fundamental-arb-short.md. Decision 1 hour before end (at least 30 minutes to resolution). Calendar split on decision time: TRAIN 2026-04-10 → 06-30, TEST 2026-07-01 → 08-21. Benchmark = favorite null, same costs. Same verdict rule as G4 (200-bet floor; mean net > 0, edge_vs_naive > 0 and cluster-bootstrap 95% lower bound > 0). Realized-vol-only limited test.

edgeTRAIN edge_vs_naiveTEST edge_vs_naive (net)Sharpe (TEST, bet-level)N_testverdict
fundamental_arb_short+1.3321+0.1550+0.06241,228VALIDATED (protocol) — likely stale-price artifact, not executable; see dossier

The protocol verdict stands as computed, but post-hoc diagnostics (not part of the pre-registration) point to stale YES prices in thin hourly markets: the TEST 95% lower bound is only +0.0097, the edge sits where the last trade is old, and the freshest-priced bets lose. Not promoted; a follow-up with executable entry prices is required before any paper trading.

fundamental_arb_exec (G4c, pre-registered separately, Polymarket price-threshold markets)

Pre-registration and full results: dossiers/fundamental-arb-exec.md. Same signal, universe and decision time as fundamental_arb_short, but each bet is filled at the first executable trade at or after the decision (within 10 minutes); returns, costs and the favorite null use that fill price. Calendar split on decision time: TRAIN 2026-04-10 → 06-30, TEST 2026-07-01 → 08-21. Same verdict rule (200-bet floor; mean net > 0, edge_vs_naive > 0 and cluster-bootstrap 95% lower bound > 0). Realized-vol-only limited test.

edgeTRAIN edge_vs_naiveTEST edge_vs_naive (net)Sharpe (TEST, bet-level)N_testverdict
fundamental_arb_exec+0.0082−0.1945−0.1457516FALSIFIED — executable fills remove the G4b edge; see dossier

At the executable fill, TEST mean net is −0.2449 and the 95% lower bound −0.3901. Post-hoc diagnostics (not part of the pre-registration): the same 516 filled bets earn +0.179 net at the stale signal price and −0.245 at the fill, and 67% of fills are worse than the signal price. fundamental_arb_short's result was a stale-price artifact. No fundamental_arb variant is promoted or paper-traded.

poly_data lake

datasetsizefilesrowslatest monthlatest timestampageunreadable
markets687.2 MiB12,589,2742026-08-21T08:59:22Z24.0 d0
markets_raw87.6 MiB13,496,6130
orders12.73 GiB5540,377,0852026-082026-08-21T09:07:27Z24.0 d0
trades13.77 GiB5540,377,0852026-082026-08-21T09:07:27Z24.0 d0

Total 27.26 GiB · /home/trbck/workspace/_repos/poly_data/processed/parquet · sizes, rows and timestamps from Parquet footers (no data scans).

Verification

oddsx boil verification — 2026-09-14 08:49

39/39 frozen checks pass.

goalmilestoneexitsecondslast line
G1M100.34 passed in 0.05s
G1M200.33 passed in 0.04s
G1M3014.843 passed in 14.31s
G1M400.441 passed in 0.11s
G1M500.48 passed in 0.05s
G1M602.63 passed in 2.31s
G1M700.33 passed in 0.02s
G3G3M100.33 passed in 0.01s
G3G3M2037.520 passed in 37.04s
G3G3M303.210 passed in 2.71s
G3G3M4085.36 passed in 84.82s (0:01:24)
G3G3M504.37 passed in 3.87s
G3G3M6014.38 passed in 0.05s
G4G4M10101.17 passed in 100.54s (0:01:40)
G4G4M200.84 passed in 0.39s
G4G4M300.94 passed in 0.49s
G4G4M400.85 passed in 0.40s
G4G4M5016.110 passed in 2.16s
G4bG4bM10102.47 passed in 101.79s (0:01:41)
G4bG4bM201.04 passed in 0.43s
G4bG4bM301.34 passed in 0.84s
G4bG4bM401.05 passed in 0.44s
G4bG4bM5018.013 passed in 0.60s
G4cG4cM10102.19 passed in 101.52s (0:01:41)
G4cG4cM200.84 passed in 0.38s
G4cG4cM301.04 passed in 0.55s
G4cG4cM400.85 passed in 0.37s
G4cG4cM5023.126 passed in 1.23s
G2G2M102.45 passed in 2.00s
G2G2M200.93 passed in 0.45s
G2G2M301.336 passed in 0.88s
G2G2M401.717 passed in 1.26s
G2G2M501.67 passed in 1.15s
G2G2M601.37 passed in 0.87s
G2G2M701.27 passed in 0.75s
G5G5M100.35 passed in 0.02s
G5G5M200.43 passed in 0.08s
G5G5M300.33 passed in 0.02s
G5G5M4028.5124 passed in 8.12s

Goals ladder

goalone-line
G1oddsx is the single merged repo — both parents' history grafted, oddstrading restructured, oddsoddy archived with triage cards, rules enforced, parents frozen.
G2the reusable pure logic of oddsoddy (title blocklist, price and trade signals, wallet PnL, extra wallet metrics, syndicates) lives in oddsx/core, proven equivalent to the archived originals on seeded fixtures.
G3the poly_data lake is compacted to Parquet, verified month by month against a frozen oracle, read by oddsx only through data/polydata.py, cut to a 14-day raw window, and kept that way by a daily retention timer.
G4fundamental_arb gets an honest, pre-registered out-of-sample verdict on Polymarket price-threshold markets (2026-04 → 08), recorded in the SCORECARD — whatever the verdict is.
G4bfundamental_arb_short (decision 1 hour before end) gets an honest, separately pre-registered out-of-sample verdict on Polymarket price-threshold markets (2026-04 → 08), recorded in the SCORECARD — whatever the verdict is.
G4cfundamental_arb_exec (fundamental_arb_short's signal, filled at the first executable trade) gets an honest, separately pre-registered out-of-sample verdict — whatever it is — settling whether the G4b edge was a stale-price artifact.
G5oddsx reads its secrets from its own gitignored .env through one loader, and nothing in oddsx points at the frozen oddsoddy repo anymore.

Retention log (last 10)

tsappliedrefusedparquet afterevicted
2026-09-13T22:19:30+00:00TrueFalse26.50 GiB0