Nine hundred and twelve famous trading ideas — anomalies, folk strategies, vendor backtests, things people swear by — each put through the same protocol and stamped twice: is the signal real? and does it survive real execution and scale? This page is the view from above. It aggregates; it doesn't re-judge — every verdict below links back to the study that earned it.
(Regenerate with python tools/make_bench_figures.py — it parses the
README table, so it's always in sync.)
Of the 912 studies on the bench, 911 carry final stamps (study 14 is pre-registered, verdict pending):
| Investable | Fragile | Mirage | ||
|---|---|---|---|---|
| Real | 6 | 47 | 31 | 84 |
| Weak | 0 | 98 | 213 | 311 |
| None | 0 | 4 | 512 | 516 |
| 6 | 149 | 756 | 911 |
Read it the way the colours tell you to:
- 84 / 911 signals are statistically real. About one famous idea in ten survives autocorrelation-robust inference. The other nine-tenths are weak (311, Mixed folded in) or plain noise (516).
- 756 / 911 are mirages once you try to trade them. Costs, capacity, decay, or the discovery that the "edge" was beta all along.
- 6 / 911 are investable. The three originals — Storm-Shy, All-Weather, Balancing-Act — are risk-managers, not forecasters. The three added by the "green hunt" lot (591–640) are the first harvestable return engines on the bench, and tellingly none is a crystal ball either: Fallen-Angels is a forced-seller risk premium, Currency-Hedged-Carry is a covered-interest-parity rate-differential identity, and Starting-Yield is duration arithmetic (the yield on your ticket ≈ your next decade). You can bank a premium or an identity; you still can't see the future.
The single most important cell isn't the green one — it's Real × Mirage (31 studies). Thirty-one effects that are genuinely there in the data and still can't pay you. That gap between "true" and "tradable" is the bench's whole thesis, measured — and it widened under the green hunt: chase likely-real premia and most still die at the trading desk (borrow fees, roll drag, one-way bounds you can't stand on).
A quieter cell worth a look: None × Fragile (4 studies) — gold [69] and bitcoin [70] flunk the claims made for them (inflation hedge, digital haven) yet keep a Fragile stamp as plain diversifiers. The story dies; the asset survives.
We sorted the 912 into rough families. The boundaries are judgement calls (is the 52-week high a chart pattern or a momentum factor? we said factor) — the totals below are honest, the taxonomy is approximate.
* "Survived costs" = stamped Investable or Fragile (alive on paper, even if thin). The complement is Mirage.
Three patterns jump out:
- ML & forecasting is the deadest corner of the bench: 0 for 7. Every model-driven forecaster — Markov pipeline [10], ARIMA+GARCH [12], neural net [39], the Stock-to-Flow model [84], a Random Forest [138] and the 'AI-powered' ETF [139] — produced an in-sample story and an out-of-sample coin flip.
- Calendar effects are the opposite failure mode: among the most real per capita (6 of 60) and almost none tradable. The pattern is genuinely in the data; the trade built on it forfeits more than it captures (42, 55) — or dies the moment it's published (67).
- Momentum, trend and carry don't die — they limp. These families collect Fragile stamps, not Mirage ones (momentum 9/11 alive-but-thin, carry 5/10): premia with a century of literature that one tape can't certify and costs nearly erase.
These aren't opinions — each one fell out of multiple studies independently.
1 · The edge dies at the costs line, not the signal line. Of the 47 statistically real signals, 44 failed or barely survived tradability. The overnight drift is real and untradable [01]; intraday reversal is real with a 3.31 bp break-even that lives in the least-liquid names [33]; the turn-of-the-month premium is real at t = 5.1 and a window-only book — even with its cash leg paid the T-bill — still compounds half of buy-and-hold [42]; IBS snap-back is real and gone at the spread [19]. Beat 6 — could you trade it? — is where almost everything dies.
2 · Survivorship doesn't just flatter results — it manufactures and even inverts them. On a survivor panel of large caps, the lottery effect ran backwards (−10.4%/yr, t = −2.5) [53], the idiosyncratic-vol puzzle inverted decisively [54], the 52-week-high premium came out negative [50], the net-issuance hedge flipped because the decade's diluters were the growth winners [64], and asset-growth showed nothing where the literature's premium hides in micro-caps [44]. When we found a positive result on a survivor panel, we capped it as an upper bound [48] — the bias cuts both ways and we say which.
3 · Post-publication decay is the norm, not the exception. The pre-FOMC drift is the most spectacular case on the bench: 3% of sessions carried 11.5% of SPY's entire cumulative return — until Lucca-Moench published it in 2011 and the drift collapsed from +0.24%/day to +0.09% [67]. The size premium never showed at all on a 39-year tradable proxy [45]; turn-of-the-month faded from 13.8 to 4.8 bp/day after 2008 [42]; betting-against-beta decayed 0.70 → 0.28 [43]; textbook pairs stopped paying once everyone copied them [05]; the vendor's dual-momentum edge thinned right after the publication that sells it [40]; oil-predicts-stocks reads exactly zero out of sample [49]. An anomaly's discovery date is the start of its obituary.
4 · Leverage is never free — the "free lunch" is usually the financing bill, in disguise. Betting against beta needs 2.78× leverage to be market-neutral, and realistic financing drags its Sharpe from 0.47 to 0.02 [43]. The retail CFD markup — charged on the whole notional, not the borrowed slice — costs a levered dip-buyer 2.65 pts/yr [30]. The 3× ETF "free amplifier" tripled the drawdown, not the Sharpe (0.90 vs 0.98, −82% trough) [61]; extending duration for term premium lowers the Sharpe [59]; and vol-targeting the carry trade makes its crash worse [27]. When a strategy's appeal is "same return, just levered," the lender has already priced your idea.
5 · What's green on the bench you earn by managing risk, banking a premium, or reading an identity — never by forecasting. For 590 studies the green column was pure risk machinery: scale exposure down when markets get loud [16], balance risk across assets and win on Sharpe not return [68], or hold the plain 60/40 [97]. The "green hunt" lot (591–640) — 50 ideas hand-picked because they should be real — finally added three greens you buy for yield, and the lesson survives them intact: a forced-seller risk premium you get paid to absorb [610], a rate-differential identity a currency hedge hands you mechanically [613], and duration arithmetic where the yield on the ticket is the return [625]. None forecasts anything — you collect a premium or read an identity. And the hunt's own body count proves the rule: chase 50 likely-real edges and most still die at the desk — borrow fees [617], roll drag [619], one-way bounds you can't stand on [621], the most famous alpha on Earth already spent [628]. Nothing on this bench forecasts returns and pays. Several things manage risk, harvest a premium, or bank an identity — and those do.
🟩 The six that made it. Three risk-managers found in the first 590, and three harvestable return engines surfaced by the "green hunt" lot (591–640) — the first time the bench's green column contains anything you buy for its yield rather than its calm.
The three risk-managers — they predict nothing and win on risk-adjusted terms: 16 · Storm-Shy — Real × Investable. Scale exposure down when realized vol spikes. It survives robust inference, real costs, capacity, a parameter sweep and a third tape — and note what it is: not an alpha, a risk overlay. 68 · All-Weather — Real × Investable. Risk parity earns the best Sharpe of anything we tested (0.92) with a third of equities' drawdown — by predicting nothing and balancing everything. Half the return of stocks, though: the green is risk-adjusted, not absolute. 97 · Balancing-Act — Real × Investable. The plain 60/40 lifts the excess-of-cash Sharpe over 100% stocks (HAC t = 2.3, a bootstrap CI clear of zero) and halves the drawdown — but it forfeits ~2.6 pts/yr of return, leans on the historic bond bull, and the bonds did not cushion 2022. Risk-adjusted, not absolute.
The three return engines — premia and identities you can actually bank, and still not a crystal ball: 610 · Fallen-Angels — Real × Investable. Bonds kicked out of investment grade get dumped by forced sellers; catching them (ANGL over HYG) pays +18.3 bp/mo, HAC t = 2.44 — and it's not duration (β to Treasuries ≈ 0). It survives dropping the entire 2020 wave, a single-year jackknife, and a block bootstrap (CI clear of zero), and it clears costs with ~1.4 bp/yr of drag. A genuine forced-seller risk premium. 613 · Currency-Hedged-Carry — Real × Investable. Two funds hold the same Japanese basket; the hedged one (HEWJ) quietly out-earns the unhedged (EWJ) by the whole US–Japan rate gap — +1.2%/yr at HAC t = 2.7 even before the 2022 hiking cycle, pass-through slope ≈ 1.0. A covered-interest-parity mechanical identity, not a signal — no forecasting, just collecting the differential the wrapper hands you for free. 625 · Starting-Yield — Real × Investable. The 10-year yield on your buy ticket predicts your next decade of bond returns at R² = 0.92, slope t = 12.5, identity slope ≈ 1 — pure duration arithmetic (Bogle/Leibowitz), robust across pre/post-1950 and a drop-one-decade jackknife. One entry per decade, ~$0 cost, effectively unlimited capacity. You don't forecast the yield; you read it off the ticket.
🟨 The honest fragiles — Real signals that survive on paper but are thin, decaying, or capacity-starved. Worth knowing; not worth quitting your job for:
| Study | What's real | Why only fragile | |
|---|---|---|---|
| 48 | Groundhog | Month-of-year seasonality, t = 4.1, undecayed | Survivor-panel upper bound; breaks even near ~19 bp |
| 52 | Smoke-Screen | Accruals: cash-backed earnings win, Sharpe 0.64 | Short-side costs; documented post-2000 fade |
| 56 | Tide-Table | CAPE forecasts 10-year returns (R² 0.28) | A tide table, not a stopwatch — useless at 1 year |
| 59 | Downhill | Term premium, +2.2%/yr over cash | Sharpe 0.32 vs cash's 1.82; 2022 took −23% |
| 63 | Free-Fall | Short-vol carry, +12%/yr (SVXY) | Skew −4.8, one −83% day; five crash days wiped 95% |
| 66 | Inverted | Curve inversion → +1% next 18m vs +16% normal | ~5% of months, a year of melt-up first — no sell button |
| 67 | Fed-Drift | Pre-FOMC drift carried 11.5% of SPY's return | Publication killed it: +0.24%/day → +0.09% after 2011 |
| 71 | Ambush | Confluence of four dead-net edges: +19.6 bp/day at K≥3 (HAC t = 3.1), undecayed, costs defeated by rarity | ~15 trades/yr → +1.2%/yr excess; OOS Sharpe +0.28 under the frozen 0.30 bar |
| 75 | Knee-Jerk | Connors RSI(2) oversold bounce: pooled HAC t = 10.7, beats a coin by +57 bp/trade | Decayed 35% since the 2008 book; long-only beta in a bull market; the 200-SMA filter hurts |
| 103 | Turtle-Trader | The Turtles' Donchian breakout: a real long-side trend premium, HAC t = +11 | Shorts are a structural trap on up-drifting markets; edge ~halved post-publication; years-long drawdowns |
| 106 | Supertrend | The ATR(10,3) daily flip beats a coin (HAC t = +3.3, bootstrap CI clear of zero) | Only the canonical multiplier works (2 and 4 are noise); ~2.5%/yr gross at ~8 flips/yr |
| 110 | Faber-Timing | The 200-day timing rule lifts SPY's Sharpe 0.55→0.73 and halves the drawdown (−55%→−22%) | Pure risk reduction, not alpha; lags in bull markets; whipsaw + switching/tax drag |
| 144 | Permanent-Portfolio | Browne's 25/25/25/25 stocks/long-bonds/gold/cash: a genuinely low-drawdown all-weather mix | Forfeits much of equities' return; leans on the gold + bond bull; risk reduction, not alpha |
| 151 | Stocks-For-Long-Run | The real equity premium holds in every long rolling window (Siegel) | The unit of time is the decade; 20-yr windows can still trail bonds; useless as a timer |
| 173 | Four-Percent-Rule | Bengen's 4% withdrawal survived every historical 30-yr US retirement cohort | Sequence-of-returns risk; today's valuations/yields; non-US history has failed it (Pfau) |
| 203 | Golden-Butterfly | Tyler's 20x5 mix beats SPY on Sharpe (0.68 vs 0.53) with a third of the drawdown | Loses to its simpler parent (Permanent Portfolio); forfeits ~3 pts/yr of return; risk reduction, not alpha |
| 209 | ETH-BTC-Ratio | The 20-day ETH/BTC momentum rotation beats 50/50 (HAC t=+3.7, alpha t=+9.4) | Strapped to a single crypto cycle, -68% drawdown; one regime, not a law |
| 210 | Crypto-Trend | A 200-day timing rule cuts Bitcoin's -83% crash to -70% and beats buy-and-hold on Sharpe (0.90 vs 0.63) | Whipsaws; ~7 switches/yr; ~12-year history; a drawdown shield, not alpha |
| 223 | Same-Month Seasonality | Same-calendar-month return persists: top-bottom decile spread HAC t = +5.57 over 330 months, bootstrap Sharpe CI clear of zero | Survivor-inflated on ~8-stock deciles; ~78% monthly turnover; short leg hard-to-borrow; a live Russell-1000 build dilutes the gross edge |
| 301 | Triple-RSI | The viral "90% win-rate" RSI(5) bounce is genuinely real: SPY +132.8 bp/trade, HAC t = +5.07, survives an honest next-open fill and post-2010, beats a coin by +85 bp | ~3.5 trades/yr, ~7% of the time in the market → ~+4.7%/yr vs the index's +10.8%; the 90% win-rate is the exit's shape, not the edge (a coin clears 62% at a loss) |
| 302 | Lithium-Boom | A 200-day trend overlay on lithium (LIT) is real — net HAC t = +2.09, and trend-timing beats same-exposure random timing at the 100th percentile | Net excess-Sharpe ~0.51 (bootstrap CI [0.02, 0.99] barely clears zero), trails SPY buy-and-hold; the only real benefit is a halved drawdown (−66%→−41%) — crash-dodging, not skill |
| 340 | Bank-Loans | Floating-rate loans (BKLN) really do dodge duration: beta to long Treasuries = −0.055 (HAC t = −2.95), and BKLN gained +13.9% through the 2020–23 bond repricing | The risk just moved duration→credit (equity beta +0.20, t = +4.48); a thin sleeve (CAGR 3.7%) that fell in 7/7 equity crashes and gapped −24% in March 2020 |
| 363 | PEAD-Drift | Post-earnings drift sorted on the EPS surprise: +1.34% at 20 days (t = 2.96), survives a quarter-block placebo | Nothing in week 1; net edge only at 20–60-day holds on a 30-name survivor basket; thin and long-leg-dependent |
| 367 | CEF-Discount | Widest-discount closed-end funds beat the narrowest by +6.6%/yr, market-neutral (Welch t = 3.72), placebo-clean | A NAV proxy on a survivor basket; fades to t = 1.56 post-2010; hard-to-borrow shorts, ~18 tiny funds |
| 375 | VXX-Roll-Decay | Shorting a VIX-futures ETP's contango is a real carry: +0.171%/day, HAC t = 2.51, +35%/yr net | Skew −1.65, a −43% day and a −92% drawdown make a constant-notional short un-allocatable |
| 601 | Factor-ETF-Live-Test | The factor ETFs do deliver their exposure live: USMV cuts vol to 0.80× (t vs 1 = −7.5), MTUM/VLUE/QUAL load their factor at HAC t = +7.5/+10.0/+8.8 | Exposure is real, alpha isn't — none reliably beats SPY once you pay for it; a delivered risk profile, not free return |
| 618 | GBTC-Premium-Cycle | All three regimes verified to the digit — +36% premium (t = 9.0), −24% discount (t = −7.4), then par; the 2023 discount→par convergence was a real, dated trade | A one-off wrapper lifecycle that ended at the Jan-2024 ETF conversion; not repeatable now |
| 622 | Thematic-ETF-Curse | A calendar-time book of 48 thematic ETFs loses −16.3%/yr CAPM alpha (HAC t = −3.27) in their first 36 months; the broad-index-launch placebo is clean | It's a short-the-hype signal: costly-to-borrow small names, clustered 2020–21 launches, survivorship flatters the long-only escape |
| 626 | Unemployment-Trend-Timing | Gating Faber's 200-day rule on rising unemployment beats the pure rule by +12.3 bp/mo (HAC t = +2.32) and cuts whipsaw spells 64% over 928 months | Most of the edge is trading less, not seeing more; current-vintage unemployment (revisions unmodeled); a smoother Faber, not alpha |
| 628 | Buffett's-Alpha | The most famous alpha on Earth replicates: +9.5%/yr CAPM alpha, HAC t = 3.47 over 46 years (the FKP t > 3 holds to 2011) | The fade is itself significant (−11 pp/yr post-2010, t = −2.38) — quality + low-beta + cheap leverage, mostly spent; little left for today's buyer |
| 632 | Crypto-XS-Momentum | Last week's crypto winners keep winning: +164 bp/wk on the winners-minus-losers quintile, HAC t = 3.56 over 449 weeks on a 44-coin panel including dead pairs (LUNA shorted) | Rides the 2017/2021 bulls; ~55% hit rate drowned by 10–50 bp crypto spreads; paid roughly nothing through the LUNA/FTX year |
| 806 | Prospect-Theory Value | The Barberis-Mukherjee-Wang cumulative-prospect-theory value predicts returns negatively: long low-TK / short high-TK earns +138.7 bp/mo (HAC t = +3.15 over 137 months, positive in both eras, correct BMW sign) | Survivorship flatters the short (blown-up lottery names absent) and a 50-name universe concentrates it into hard-to-borrow lottery mega-caps; +124 bp/mo net at 5 bps but real borrow/squeeze exceeds the charge |
| 812 | Corwin-Schultz Spread | The high-low bid-ask illiquidity estimator earns a premium — high-spread names out-earn by +4.45 bp/day (HAC t = +3.24, both eras, +4.88σ placebo) | The illiquid long leg pays the very spread it earns: net +2.31 bp/day at 1 bp (t = +1.59, insignificant), −5.69 at 5 bps |
| 861 | Debt-Maturity Rollover | Firms funded with a high share of SHORT-TERM debt under-earn the long-funded — low-minus-high-rollover-share tercile +45 bps/mo (NW t = +3.22), right-signed, both regimes (post-2017 +2.83, post-2019 +2.51), ~2.6× stronger in the 2022+ hiking era | Thin ~26-name survivor cross-section (no delisted rollover-wall casualties), monthly turnover + hard-to-borrow shorts; a real balance-sheet risk premium too small and capacity-starved to bank cleanly |
| 888 | CLO AAA Carry | The senior AAA CLO tranche (JAAA) pays a genuine low-vol carry over cash — +1.38%/yr on 1.63% vol, excess Sharpe +0.84 (HAC t +2.33, block-bootstrap CI clear of zero), topping the excess-Sharpe race and beating the un-tranched loans it's carved from | ~5.7y stress-free sample (misses the 2020 CLO mark-down), and all the carry lives in the high-rate era (ZIRP excess ~0) — mechanically a short-rate+spread, so regime-bound and short-history; thin and not yet crisis-tested |
| 889 | Broad Dollar-Hedge Overlay | Generalising 613: on broad developed-international (HEFA/DBEF vs EFA, the same basket) the hedged-minus-unhedged return IS the US–EAFE rate differential mechanically (β on −fx ≈ 1, high R², HAC t clears the bar, bootstrap CI clear of zero) | The identity is robust but the whole tape sits in one dollar-regime (US out-yielded EAFE); the tradable overlay is regime-limited and the pickup is small — a mechanical identity, not a forecast |
This page will be wrong eventually — that's the design. Eight hundred and forty-one verdicts is eight hundred and forty-one falsifiable claims, each with reproducible code, pinned data fingerprints, and the exact line where we think the dream dies.
- Think a Mirage is tradable? Fork the study, change the cost model or the venue, and show the break-even. Beat 7 of every notebook says what we'd consider convincing.
- Think a None is real? The inference stack (HAC, Lo, bootstrap, Reality
Check) is in
quantlab/— run it on your variant. - Got a candidate for the queue? Open an issue. The ideas that look most embarrassing to test are usually the best ones.
The map gets a new chip every time. python tools/make_bench_figures.py
redraws it.
Part of Open-Alpha-Lab. Counts generated from the README
table by tools/make_bench_figures.py.
Not investment advice — research and education. See LICENSE.
