This post first appeared on AI Forecast Tracker, my public AI forecasting site: 50 tracked predictions, 21 public topics with community voting, and a scoreboard of 153 companies. The revision it describes is agent work: a fleet of research agents re-checked every open prediction, and an adversarial reviewer stress-tested every big move.
Six months ago we started writing down numbers.
Not takes. Not vibes. Numbers — 40 public predictions about AI from the people running the labs, the banks, and the consultancies, each one translated into a probability we were willing to be graded on. The rule was simple: when someone says “AI will transform everything,” we write down what that means, when it resolves, and how likely we actually think it is. Then we wait, and we keep the receipts.
This week we ran the first full revision of everything we track: every open prediction re-researched from scratch, every big move stress-tested by an adversarial reviewer whose only job was to prove the researcher wrong. Here’s what six months of AI reality did to the forecasts — and to us.
The scoreboard: five verdicts in
Five predictions have already resolved. We went 5-for-5 on direction — and we want to be careful about how much that means.
Yann LeCun said the AI investment boom was a bubble likely to burst by 2026. We put it at 5%. The window closed on June 30 with zero of our five burst criteria met — and it closed during the hardest week the market threw at AI all year, a three-day semiconductor selloff that erased over a trillion dollars and then reversed. If a bubble doesn’t pop in that week, in that window, it wasn’t ready to pop. Refuted. (LeCun’s Brier: 0.09. Ours: 0.0025.)
Sam Altman said AI would discover genuinely novel scientific insights in 2026. We put it at 60% rising to 78% as evidence came in, and it confirmed early: Google’s Co-Scientist published in Nature with six independent wet-lab validations, GPT-5.4 solved a FrontierMath problem that had been open since 2019, and an AI-written paper passed Nature’s human peer review. Confirmed — two months before its deadline.
Boris Cherny said AI can already write 100% of production code, with top engineers 10x more productive. Confirmed, with the caveat we’ve repeated all year: this is true for the top of the distribution, not the middle. The power law is the story — top-decile engineers get 10x, the average gets 1.5–2x. Anyone selling you a uniform multiplier is selling you something. Confirmed.
ARK’s inference-cost collapse (99%+ in a year) and Cursor’s billion-dollar-revenue-with-fewer-than-20-people claim both confirmed at 95% — the two “easy” calls on the board.
Our aggregate Brier score across the five: 0.022 (0 is perfect, 0.25 is coin-flipping). The honest caveat: n=5 is small, and three of the five were priced near-certain. The two that weren’t — LeCun’s bubble and Altman’s science — are the ones we’re actually proud of. We’ll take the grade, but the sample is the sample.
What six months did to the other 35
This week’s sweep re-researched all 35 open predictions — every one independently, with verdict changes and big moves stress-tested by an adversarial reviewer. Twenty-one moved up, eleven moved down, three held. Three patterns dominate.
1. The robots stopped being a demo. The single biggest riser: Sam Altman’s “robots will be executing tasks in the real world by 2027,” from 75% to 85%. Waymo is at 500,000 paid rides a week. Figure’s humanoids at BMW’s Spartanburg plant now outnumber the human employees at Figure’s own facility — about 740 robots doing unsupervised logistics. Agility’s Digit has 65,000+ commercial hours and a $2.5B SPAC built on $300M of contracted orders. Even Gary Marcus’s “humanoids are all demo, very little product” — a claim we’d backed at 70% — dropped to 60%: he’s still right about Optimus (Musk’s own word for the ramp is “extremely slow”), but wrong, by name, about Figure.
2. Coding AI hit a reliability wall nobody priced. Capability kept climbing — 95.5% on SWE-bench Verified, near-frontier models now free-tier. But this cycle produced the strongest counter-evidence cluster in our tracking history: OpenAI’s own audit found ~30% of SWE-bench Pro’s public tasks are broken and retracted its endorsement of the benchmark; a 500-org survey found 43% of AI-generated code changes still need production debugging after passing QA; and Zuckerberg — whose “AI will write most of Meta’s code” we track — admitted internally that agentic progress “hasn’t really accelerated” in four months. Andrej Karpathy’s “agentic engineering becomes the default” fell from 50% to 40%. Meta’s own claim fell to a coin flip. The gap between benchmark and production is now the most important number in AI, and almost nobody publishes it.
3. The labor story got complicated — and both halves are true. Stanford’s payroll-data Canaries dashboard shows entry-level employment in AI-exposed jobs falling over 4% a year and accelerating: the starter-job squeeze is real and isn’t mean-reverting. But the NY Fed now attributes 64% of the youth-unemployment rise to remote work, not AI — and 55% of employers who cut jobs for AI say they regret it, with over a third already rehiring (IBM tripled entry-level hiring; Ford brought engineers back after its AI design review missed defects). The “Klarna Pattern” — fire for AI, rehire quietly — went from one company’s embarrassment to a measured behavior with a number attached: 32% of managers who cut a role for AI have already rehired for it (Robert Half, n=2,000). Our entry-level-crisis assessment actually rose to 83% on the payroll data — while the causal story underneath it got a lot more honest. Both things are true at once, and pretending otherwise is how forecasters go wrong.
And hovering over everything: in June the US government switched off the most capable coding model on Earth for nineteen days — the first time Washington reached into a live commercial frontier model — then switched it back on behind a jailbreak classifier and, a week later, a biometric ID gate. Every capability forecast now carries a rider it didn’t have in January: assuming you’re allowed to use it.
Where we were too hot
Grading only feels honest if the misses get equal billing.
- Deepfake fraud at $40B by 2027 — we cooled from 50% to 40%. The trend is real; the specific dollar line looks like it was drawn for a headline.
- Facial verification abandonment (Gartner’s 30%-of-enterprises claim) — down to 48%. Deepfakes are winning individual fights, but enterprises are layering defenses, not abandoning the technology.
- “Agentic engineering as default professional workflow” — our steepest self-correction of the sweep (50% → 40%), for the reliability reasons above.
- And a process miss that never made it to the site, which is exactly the point: while scouting new predictions this week, one of our researchers built a case on a “climbing” grad-underemployment trend. Our adversarial reviewer checked the source: the latest NY Fed print had reversed — 41.5%, down from 42.5%. The prediction shipped anyway, re-anchored, at a much humbler probability. The pipeline that catches your own researchers extrapolating a dead trend is worth more than any single forecast.
Ten new bets
The sweep also told us where we weren’t looking. Ten new predictions go live today, four of them as full topics:
- Companies fired workers for AI. Now they’re quietly hiring them back. Will the rehire rate hit 40% of AI-cutting employers by December? — 35%
- AI data centers are already flipping elections. Will electricity bills decide a November 2026 general-election race? (A 20-year Utah incumbent already lost a primary over it.) — 50%
- Is China about to beat America at its own AI game? A Chinese model in the global top 3 by January 2027 — the top five is currently 100% one lab’s models. — 15%
- Will an AI “artist” crack the actual Billboard Hot 100? An AI act already peaked at #20 on a genre chart, with a $3M label bidding war behind it. — 50%
- Three Mile Island restarts a year early to power AI. — 40%
- The PJM grid auction hits its price cap a fourth straight year — the quiet number behind your rising power bill. — 72%
- New-grad underemployment crosses 45% — the re-anchored version. — 30%
- Philippine BPO headcount grows through 2027 — our own stress-test against our outsourcing-earthquake thesis. — 65%
- A hyperscaler blinks on AI accounting — a depreciation-schedule change or impairment that Michael Burry has been shouting about. — 45%
- Anthropic rings the bell before OpenAI — an AI-lab IPO closing by December 31, 2026. — 40%
Note what’s not here: another coding benchmark. Our portfolio was already saturated with capability bets; the gaps were energy, geopolitics, consumer culture, and education. That’s where the new money goes.
The receipts get a permanent home
Starting today, resolved predictions stop disappearing into the archive and get their own page: Track Record — every verdict, what we said, what happened, and the running Brier score. Confirmed calls and busted ones, same font size. The refuted ones are the reason to trust the confirmed ones.
Six months ago the loudest question in AI was “is it all hype?” That question is now settled enough to be boring — the interesting questions are about distribution: who gets the productivity, who pays for the power, who gets hired back, and which government reaches for the off switch next. We’ve got numbers on all of them.
Disagree with any of ours? Every forecast page has a slider. That’s the point.
Methodology: probabilities 0.05–0.95, Tetlock-style; every resolution requires pre-registered criteria; Brier scores computed per source once 3+ predictions resolve. This revision: 35 predictions independently re-researched, verdict changes and ≥0.10 moves adversarially verified. Not financial advice.