The Engine, Three Arenas
Same machine, three battlefields. One rule everywhere: beat the dumb baseline under a strict publication-time wall, or admit you didn’t.
Every forecast is issued behind the wall — train on frozen labels only, score before the outcome is read. No hindsight, no leakage. We publish the wounds next to the wins.
World Events
live · backtestedOpponent
the expanding base rate
The laziest forecaster alive — it just recites the historical frequency. Most "geopolitical AI" never actually beats it; it narrates around it.
Scoreboard
Brier 0.1718 vs 0.1969 — +12.7% skill, behind the wall
tech_decoupling still loses to base rate; 8 of 29 origins still lose. Wounds published next to wins.
Sporting Events
live · beats Elo, not yet the closeOpponent
the closing line
The sharpest baseline in any domain — by lock it has eaten every public signal. Anyone selling a model that "destroys the book" is selling you the part before the wound.
Scoreboard
Stack RPS 0.1731 vs Elo 0.1779 — beats Elo in 16/16 origins (15,821 OOT predictions)
We beat Elo, not the closing line — measuring closing-line value needs historical odds we don’t yet have. The edge, if any, lives in thin markets and mispriced tails, not the deep book.
Weather
live · servedOpponent
climatology & persistence
"It’ll be about average" and "it’ll be like yesterday." Embarrassingly hard to beat past a few days — which is exactly why the Brier score was invented for weather in 1950.
Scoreboard
Brier skill vs climatology — precip +12.4%→0.6%, freeze +37.7%→2.5% (lead 1→7 days) · 8 cities × 20 yrs ERA5
Skill decays with lead time exactly as promised — precip 12.4%→0.6% as the horizon extends, converging back to "about average." Freeze is far more predictable than precip. Built from ERA5 reanalysis, not a real NWP forecast: it beats the naive baselines (the engine’s game), nothing more.
World Events — base rate finally beaten overall
Build 0.4.0 · 2,409 strict rolling-origin forecasts across 29 origins · 2,694 adjudicated labels, 9 target families.
| Model | Brier | Log loss | ECE | Brier skill |
|---|---|---|---|---|
| Target-specific stack | 0.1718 | 0.5262 | 0.0549 | +12.7% |
| Calibrated stack | 0.1764 | 0.5387 | 0.0800 | +10.4% |
| Recent base rate | 0.1923 | 0.6220 | 0.1341 | +2.3% |
| Expanding base rate | 0.1969 | 0.6358 | 0.1532 | — |
The win is not one naked classifier — it is target-specific models, calibrated anchor signals, rolling past-only stack weights, and online reliability shrinkage. The stack is also better calibrated than the base rate, not merely sharper. The crown is conditional: it has earned a stronger arena, not a crown.
Sports numbers come from a strict rolling-origin backtest on real international-match history (2010–25). World-events numbers come from the Build 0.4.0 adjudicated pack. Weather has not been backtested yet — when it is, the scores will come from verification, not from this page.