oddsx · read-only research status
Generated 2026-09-14T10:00:49Z · HEAD 9612e75

Edge-3 dossier — Adverse-selection / toxic-flow (follow the one-sided flow)

Verdict: FALSIFIED (out-of-sample, net of costs, vs the favorite null), and the OP's "wide-spread = retail/safe" is FALSE in our data. Ticket: T-0015 (powered re-run of T-0011). Script: research/03_toxic_flow.py. Date: 2026-07-24.

> History: provisional FALSIFIED on the Apr 18–29 window (~485k fills, > TEST ≈ 1 day; T-0011) → powered FALSIFIED on the FULL 53-day window > (29.6M fills, real 15-day June OOS; T-0015). The powered run removes the > provisional's deceptive "+1.45 edge vs the uniform null" by scoring against the > sharp favorite null instead: the honest edge is now cleanly negative > (−0.55), and the follow-flow return still *worsens* as flow gets *more* > one-sided.

> Honest > profitable. Two clean negatives: (1) following short-window one-sided > flow does not beat buying the favorite net of costs — it gets *worse* as > flow gets *more* one-sided; and (2) wide-spread and near-close fills do not > have worse taker outcomes. The adverse-selection *mechanism* is real for a > market-maker, but it is not a taker-side edge you can harvest by chasing flow.

---

Hypotheses

From the r/algotrading thread: the bot's residual bled to adverse selection — informed, one-sided flow lifts stale quotes right before the move.

  1. STRATEGY ("follow the toxic flow"): when the short-window signed flow into a market is strongly one-sided, take the flow direction and hold to resolution. If informed flow predicts the winner, following it should be +EV out of sample, net of costs.
  2. DIAGNOSTIC (the OP's unverified claim): "wide spread = retail/safe." If true, wide-spread (and/or near-close) takers should have worse realized outcomes than the rest.

Method (powered — T-0015)

TRAIN scan — follow-flow return by one-sidedness (capped, net of costs)

The hypothesis predicts the follow-flow return should rise with one-sidedness. It does the opposite — every bucket is gross-negative once capped, and it *worsens* as flow gets more one-sided (chasing a stronger move buys a more extreme, more expensive-to-cross price):

|one-sidedness|Nwin_ratemean_gross (capped)mean_net (capped)
0.0–0.15,173,4900.446−0.1926−0.7768
0.1–0.24,031,1320.375−0.2959−0.7004
0.2–0.33,207,6020.331−0.3614−0.8047
0.3–0.42,263,0850.265−0.4820−0.9853
0.4–0.51,559,2350.211−0.5870−1.1323
0.5–0.61,097,2930.173−0.6619−1.2309
0.6–0.7674,7020.153−0.7038−1.4474
0.7–0.8417,7520.125−0.7626−1.9171
0.8–0.9244,5810.131−0.7598−2.5090
0.9–1.0681,7920.342−0.5006−3.5558

No bucket is gross-positive, so the band derivation falls back to the least-one-sided bucket, |osid| ∈ [0.00, 0.10) — the exact opposite of the thesis (same fallback as the provisional run, now on 30M fills).

TRAIN vs TEST — the verdict table

A — follow-flow in the TRAIN-derived band [0.00, 0.10) vs buy-the-favorite (favorite null)

splitNgrossnetsharpenaive_netedge_vs_naive
TRAIN5,173,490−0.1926−0.7768−0.3008−0.0272−0.7497
TEST2,717,883−0.1516−0.5881−0.2595−0.0365−0.5516

B — follow-flow when |osid| ≥ 0.60 (the ticket's literal "strongly one-sided") vs favorite null

splitNgrossnetsharpenaive_netedge_vs_naive
TRAIN2,044,007−0.6474−2.3589−0.6295−0.0272−2.3317
TEST1,041,970−0.6525−1.9138−0.5894−0.0365−1.8772

Against the sharp favorite null, both runs are cleanly negative on both splits: the band strategy loses ~59% per bet (TEST net −0.588, edge −0.552) and the literal "strong flow" cut loses ~191% (TEST net −1.914, edge −1.877; chasing an extreme move at a thin, cheap price is dominated by the cost of crossing the spread). FALSIFIED. For context, vs the *uniform* null the band run shows a misleading edge_vs_naive = +0.25 — positive only because follow-everywhere (−0.84) is even more catastrophic; that is the trap the favorite null removes.

Robustness — one strongest-|osid| position per market, |osid| ≥ 0.60 (de-clustered): TEST net −0.3966, edge_vs_naive −0.3722, N=26,655 — negative on both splits. De-clustering does not rescue the edge.

Tail-variance honesty

runTEST Nmedian_netmean_net (capped)mean_gross UNCAPPEDminmaxtop-1% |PnL| share
band (A)2,717,883−1.0374−0.5881+0.0102−11.05+0.9550.084
strong (B)1,041,970−1.1204−1.9138−0.0069−11.05+0.9550.051

vbt.pro simulation (gate L1.2)

vbt.pro 2026.4.7, hourly equal-weight book over the TEST window (375 bars across the ~15-day June era, band run, cost-adjusted): total_return −100.0%, Sharpe(1h-ann) −67.4, max_drawdown −100.0%. Unambiguously negative and consistent with the harness verdict. Pandas equity curve emitted as fallback.

OP DIAGNOSTIC — is "wide-spread = retail/safe" true? NO.

1 — taker outcome by spread-width proxy quintile (price range over the local 30-min window; mean_taker = the taker's own realized return; ~5.93M fills/quintile):

spread quintileNmean_taker (uncapped)mean_taker (capped)medianwin_ratemean_price
Q1_tight5,930,355−0.0090+0.1239−0.0010.5440.550
Q25,930,355−0.0000+0.1371−0.0010.5150.514
Q35,930,355+0.0027+0.1414+0.0010.5140.514
Q45,930,355−0.0040+0.1317+0.0010.5170.515
Q5_wide5,930,355−0.0375+0.1434+0.0010.5370.535

The retail story predicts Q5_wide should be the worst. On the tail-robust capped measure the quintiles are essentially flat (+0.124 … +0.143), and Q5_wide is NOT the worst — it is indistinguishable from the middle. Wide-spread fills do NOT have worse outcomes → "wide spread = retail/safe" is FALSE in our data.

2 — taker outcome by time-to-close:

time to closeNmean_taker (uncapped)mean_taker (capped)win_ratemean_price
<1h1,161,325+0.0282+0.14550.6280.628
1–6h21,923,118−0.0036+0.14040.5280.527
6–24h5,456,393−0.0287+0.11750.5010.502
1–3d899,429−0.0690+0.11220.4850.505
>3d211,510−0.0865+0.13320.5190.524

Near-close (<1h) takers are not the worst — they have the highest capped outcome (+0.146) and win 63%, the direction adverse selection predicts (informed flow picking off stale quotes), the opposite of "near-close = safe." Here the price-level confound is mild (near-close mean price 0.628 ≈ its win-rate), so the read is clean: there is no evidence that wide-spread or near-close identifies *worse-informed* (retail, "safe") flow.

Verdict

FALSIFIED (powered). "Follow the toxic flow" does not beat buying the favorite out of sample and is not profitable net of costs — following flow gets *worse* as flow gets *more* one-sided (TRAIN scan monotone-negative), and any apparent edge is a fat-tail / longshot-chasing mirage that dies under equal-notional capped sizing and the cost of crossing a thin spread. Separately, the OP's "wide-spread = retail/safe" is FALSE: wide and narrow spread carry the same taker outcome, and near-close flow is if anything *more* (not less) informed. The adverse-selection mechanism is real for a market-maker (a reason to widen/skew quotes), but it is not a taker-side statistical edge in this dataset.

Caveats

  1. The favorite null is the honest control. vs the uniform null the band run shows a deceptive +0.25 "edge" (the null is even worse); vs the favorite null it is a clean −0.55. Read the favorite-null row, not the uniform one.
  2. Entry-time split + fill-clustering: the all-fills runs weight high-volume markets by fill count; the de-clustered per-market run agrees in sign (negative, TEST edge −0.372). Event-level weighting is the stricter unit (deferred).
  3. Spread proxy is a realized-dispersion surrogate, not the book. We have no historical order-book snapshots, so "spread width" is proxied by local traded-price range. A true top-of-book spread could sharpen the diagnostic — but the *flat capped outcome across quintiles* (now on 5.9M fills/quintile) is a strong prior that the retail-vs-spread link is weak at best.
  4. Cost sensitivity: chasing one-sided flow buys extreme/cheap prices where the fractional cost of crossing the spread is largest; the result is *more* cost-sensitive than Edge-1, not less. Even a materially cheaper execution path would have to overcome a *gross* signal that does not point the right way.

Reproduce

uv run --extra backtest python research/03_toxic_flow.py
# tunable:   ODDS_FLOW_WINDOW=30min

uv run pytest -q → green.