The Engine, Three Arenas

Same machine, three battlefields. One rule everywhere: beat the dumb baseline under a strict publication-time wall, or admit you didn’t.

Every forecast is issued behind the wall — train on frozen labels only, score before the outcome is read. No hindsight, no leakage. We publish the wounds next to the wins.

World Events

live · backtested

Opponent

the expanding base rate

The laziest forecaster alive — it just recites the historical frequency. Most "geopolitical AI" never actually beats it; it narrates around it.

Scoreboard

Brier 0.1718 vs 0.1969 — +12.7% skill, behind the wall

tech_decoupling still loses to base rate; 8 of 29 origins still lose. Wounds published next to wins.

Sporting Events

live · beats Elo, not yet the close

Opponent

the closing line

The sharpest baseline in any domain — by lock it has eaten every public signal. Anyone selling a model that "destroys the book" is selling you the part before the wound.

Scoreboard

Stack RPS 0.1731 vs Elo 0.1779 — beats Elo in 16/16 origins (15,821 OOT predictions)

We beat Elo, not the closing line — measuring closing-line value needs historical odds we don’t yet have. The edge, if any, lives in thin markets and mispriced tails, not the deep book.

Weather

live · served

Opponent

climatology & persistence

"It’ll be about average" and "it’ll be like yesterday." Embarrassingly hard to beat past a few days — which is exactly why the Brier score was invented for weather in 1950.

Scoreboard

Brier skill vs climatology — precip +12.4%→0.6%, freeze +37.7%→2.5% (lead 1→7 days) · 8 cities × 20 yrs ERA5

Skill decays with lead time exactly as promised — precip 12.4%→0.6% as the horizon extends, converging back to "about average." Freeze is far more predictable than precip. Built from ERA5 reanalysis, not a real NWP forecast: it beats the naive baselines (the engine’s game), nothing more.

World Events — base rate finally beaten overall

Build 0.4.0 · 2,409 strict rolling-origin forecasts across 29 origins · 2,694 adjudicated labels, 9 target families.

Model Brier Log loss ECE Brier skill
Target-specific stack 0.1718 0.5262 0.0549 +12.7%
Calibrated stack 0.1764 0.5387 0.0800 +10.4%
Recent base rate 0.1923 0.6220 0.1341 +2.3%
Expanding base rate 0.1969 0.6358 0.1532

The win is not one naked classifier — it is target-specific models, calibrated anchor signals, rolling past-only stack weights, and online reliability shrinkage. The stack is also better calibrated than the base rate, not merely sharper. The crown is conditional: it has earned a stronger arena, not a crown.

Sports numbers come from a strict rolling-origin backtest on real international-match history (2010–25). World-events numbers come from the Build 0.4.0 adjudicated pack. Weather has not been backtested yet — when it is, the scores will come from verification, not from this page.