This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR#
Economics#
New Epoch compute report is a smorgasbord:
-
US:China totals are about 15:1.
-
Google has 23% of world compute. Every non-hyperscaler combined has 25%.
-
Meta went for AMD somewhat, 20% (they have the R&D budget and it was the last unblocked option).
-
The vexed H20 chip is 40% of Chinese compute. China is halfway switched over to Ascends.
Opinion: Ascends for inference? Ascends for training would slow them down another couple years. Compute policy really could work.
Scaling still power-constrained: Obscure market research group claims that up to half of planned 2026 US datacenter buildouts will lag into 2027 or beyond. Bottom-up look at 190 GW from 777 individual projects. “In 2025, 26% of expected capacity slipped, and another 10% of projects pushed back their commercial operation dates without much notice.” Concentration: on-site and hybrid are 10% of projects but nearly half of announced capacity. “Developers have been submitting applications across multiple regions and Independent System Operators to maintain optionality, redundancy, and see who moves fastest. So ~50% being delayed or cancelled has a natural relationship to the speculative nature of project development. It’s not all equipment shortages or pricing.”
Opinion: We doubt anyone in the know is surprised by the above.
Stargate UK paused indefinitely. Also not pushing the training-friendly copyright change.
Opinion: Was only 8000 GPUs lol. Could be a negotiating tactic against the gov, could be to do with the VP leaving.
Nice Epoch opinion breaking down the Iran threat to the AI supply chain: as we argued, not much yet. “Gulf investment flows into AI are on the order of tens of billions annually [out of $1tn total]: significant, but not critical.”
Opinion: Confirms our prediction from #10.
Google buys 2.4GW of solar power (and a pipeline for 9 more) for $5bn. Another 2.9GW under construction by the acquired company, Intersect, and another 6GW waiting for approvals. So supposedly +11GW by Q4 2028.
Opinion: Natural move since the politics and logistics of the grid are straining and since nuclear has been difficult; interconnect delay is now 5 years (for average suppliers who don’t have legions of compliance and lobbyist staff).
Correction to fluffy numbers on China leading in industrial robots per capita (as opposed to demo robots, as opposed to total). It is actually 22nd worldwide. “166 robots for every 10,000 people employed, which was a year-on-year increase of 17%”. Korea is 1220.
Opinion: Absolute number more important for e.g. defence purposes.
Example of what real tokenmaxxing looks like: The Information hear that Meta (75,000 employees) use 2 trillion tokens a day, 27M per head per day or something like $300 retail. For some reason fools speculate that all(!) of this is Claude and so made claims about the “huge” Meta share of Anthropic revenues. Meanwhile, SemiAnalysis is at 62M per head per day.
Opinion: we should do some of this.
Capabilities#
-
On-trend in many ways. It catches Anthropic up to OAI, and a little further (for the next month). We don’t believe it’s a step-change even in the touted cybersec dimension. To be clear: “on-trend” is still big progress and the Opus + harness prior art is very very powerful.
-
It’s maybe 10T params and 1e27 FLOP.
-
Everyone is ignoring that it’s five times more $ per token. Probably not worth the premium for most tasks? (As usual, capability densing means this will be cheap next year.)
-
Not general access. 40 enterprises only, at first.
-
Nominally their “most aligned” model but many worrying signs
-
No consensus on whether internal deployment is good, whether CoT bug is bad, the extent of the hard power Ant now has
-
Lots of silliness about their caution and Glasswing and giving away the zero-days being unusually noble when it’s just basic sanity.
-
RSI deflation: The system card implies that frontier progress comes 4:1 from compute:human ideas. No signs of AI contributions to Mythos training.
-
Cloudflare is down 22%, which is pretty ridiculous.
-
Anthropic are very well-meaning and very dangerous.
First model from the new Meta lab. Closed, free. ~No general API yet. A big push for multimodal. Does well on (private) FrontierMath. One ok independent suite of narrow “real” tasks puts it #3 overall after Sonnet(!) and Opus. Very good at taxes. Some UI tricks to make it seem fast, but it also is fast. Length penalty during RL produces this pleasing graph, and here is clear discussion of the three suspected scaling laws. The system prompt is quite nice, though very different from Ant (imperious). Reporting lots of obscure benchmarks, bad sign. Hiding SOTA on the big table, bad sign.
Opinion: Catchup, better than expected. Our hope that mercenaries cannot truly create is mildly dinged but still standing. We tried it out, including crudely pasting a system prompt as prompt prefix. It’s good, nothing like Llama. Probably around Grok-level on ideal benchmarks. No big model smell but it’s fairly tasteful. Wouldn’t mind adding it as 6th-ranking member of a council.
OpenAI “Spud”
-
Pretraining finished ~Mar 24, so out by end of May, maybe sooner given Mythos.
-
Altman is not overseeing its safety, he’s busy fundraising. So less haste?
-
OAI defects: there will probably be a general release for Spud (unlike for the new cybersecurity product).
-
OpenAI internal model solves 3 open Erdos problems 5.4 can’t solve. (Could Aletheia solve them?)
Can LLMs successfully trade stocks with $10k of real money? A new benchmark has Gemini and GLM winning (up 40% in a month) and others cratering.
Opinion: Bad scaffold. They provide open transcripts of the models’ current actions, which lets us see that GPT is just stuck (only 5 trades ever) and no one is doing anything about it.
What about whole-season sports betting? Another new long-horizon prediction benchmark has them all losing money using a betting market on simulated sports seasons.
Opinion: should be a really good setting for them: vast amounts of opinion to harvest, vast amounts of historical data to fit, unambiguous objective with weekly dense rewards
xAI is training a 10T Grok right now (in addition to 5 other smaller Grok variants)
Opinion: They are currently out of the race, the worst by far on third-party evals.
Greenblatt update. He’s even more worried. - 1.75x serial eng speed up at Anthropic, 2.5 hour 80%-reliability on METR, 6.5 hour 50%-reliability on inside-company tasks [not HCAST!].
Opinion: Great guy, updates us a little.
We’re up to 4 true AI Erdos solves (green circles), two in January and two in March. (“True” is the extremely high bar of totally unassisted, verified by Tao, and lacking prior art.)
Opinion: Too early to say if current models have hit their limit yet; we doubt they have. Could look at which fields have yet to see any assisted / pure result.
The new MirrorCode eval we reported on had a ceiling of 1 billion tokens per task (c. $10k). Opus 4.6 managed to hit this limit without completing the final hard software project, but it was still improving.
Opinion: fair enough, that’s cheap for what it is (a large valuable scientific program).
Cool Microsoft paper: After minor retraining, the KV cache can act as continuous state for long tasks. In this design, they have the model summarize its work periodically in discrete tokens (“mementoes”), but persist the KVs (which carry info beyond the summary text) rather than wiping them.
Opinion: hopefully doesn’t lead to neuralese; representations are still anchored to text at this point.
Labs gutting the wrappers, part 5: Anthropic releases “managed agents”, a way of handling changing harnesses.
Opinion: Presumably further inroads against OpenClaw and Gastown
My simple way to estimate the open-closed model gap. A little more than 10 months from this bad sample.
Opinion: probably need to condition on people actually having tried Kimi first.
Paper by some very serious people and Google theoretically solving one of the many bottlenecks for using quantum computers for ML tasks (here classification and PCA). “the data loading problem, the challenge of efficiently accessing the classical world in quantum superposition”.
Opinion: real progress but like 1% of what you’d need. It’s quantum linear algebra, not quantum AI. Exponential speedup for linear tasks seems right. Blog is 100x more sensational than the paper.
Politics#
DC court doesn’t immediately block the Ant supply chain risk designation. California ruling is thus moot, designation stays for >6 weeks, possibly much longer. Oral arguments on May 19. Deferring to the man: “opinion is highly deferential to DoD… reasonably likely the panel will rule in DoD’s favor on the merits as well. But Anthropic has a strong chance of success in en banc review”. “The odds of Anthropic drawing a panel of three Republican-appointed judges on the D.C. Circuit were only about 2.4%.” Probably a new set of 3 judges for the May panel. State of play:
-
DoD can refuse to use it for IT and telecoms, pending full briefing before DC court.
-
Nothing stops other Feds using it
-
Private contractors can use it, except on covered DoD contracts.
Opinion: Defer to the pros here.
GSA restores Claude to USAI.gov.
OpenAI seems to have drafted the worst “AI Safety” bill ever, in Illinois. Would give AI companies full immunity from liability for catastrophes (100+ deaths / $1bn damages) caused by their model in exchange for the company publishing a safety protocol. OpenAI testified in favour.
Opinion: Shows the total disconnect between OAI Policy (public service with some conflicts of interest) and OAI Global Affairs (Lehane, killers). Unclear who will win but Lehane has the money and the reins for now. Really worth pushing back on if Seismic or Fathom have any pull.
Political opposition blocked around 20% of 2024–2025 datacenter plans ($18 billion in projects blocked and another $46 billion delayed over the past two years). This might increase now, given the consumer bill inflation and water usage news cycle, or might decrease, now that hyperscalers are desperate and liquid and doing the most intense jurisdiction shopping ever. “At least 12 states have filed moratorium bills in 2026 alone, with 300+ data center bills filed across 30+ states in just the first six weeks of the legislative cycle.”
Opinion: Half-joking: this is more impact than AI safety. It is unvirtuous to rely on falsehoods to get what you want, but it did a lot.
More OAI astroturfing, if you’re interested in who will be running opposition against you. Big Twitter spend. OpenAI → Brockman → Leading the Future → Targeted Victory, Build American AI → web of anon accounts.
Off the back of The AI Doc, Tristan Harris’ gang set out some AI principles like it’s 2021. Fathom and AVERI mentioned, but otherwise very socially distant from our usual set. They remain the best faction in the AI Ethics crowd.
Opinion: We’re stating the obvious but this clearly isn’t an epistemic speech-act, it’s a coalition-building thing. For various reasons (~apolitical, non-US, nerd) we can’t evaluate its success at that goal very well, and you’d want a lawyer to help with the meat: the actual bills they support.
Relatedly (since the CHT report was by a TBI guy): I didn’t know that the Tony Blair Institute was primarily funded by Ellison ($50m a year). Explains some things.
Alarmist piece making some good points about the power AI and OpenAI have over Britain already.
Safety#
Nice AISI paper running influence functions backwards. Normally we take behaviour and find the training data that explains it: data = F(behaviour). Here we use the learned F to craft training data that induces some model behaviour. This is yet another attack vector against LLMs.
Abi Olvera from Golden Gate has been interviewing virologists and biosecurity folks, and overall updates against AI biorisk soon. This essay just covers the reasons for the historically low human rate of attempts and successes.
Nice paper looking at “blind refusals”, models failing to help users evade unjust rules just because they are rule-following. “Help rates on defeated rules range from 7.7% (GPT-5.4-mini) to 58.0% (Grok-4). The GPT-5.4 family is the most restrictive”
Incidents#
Costs of cybercrime have doubled since 2022, to $21bn a year. No trend break for LLMs!
Opinion: hard to infer anything but good news. Crime is bottlenecked on ideas so the diffusion may still happen.
A hacker claims to have breached China’s National Supercomputing Center in Tianjin, stealing over 10PB of data spanning aerospace, military research, bioinformatics and fusion simulations.
Opinion: Unlikely. This could only have been a physical breach, lifting 100 hard drives. Note that a single SOTA large eddy simulation can generate 1PB. 10 such simulations is worth something to someone but it’s also not a huge deal.
Excellent paper on a difficult prereq for any liability regime: knowing who done it.
“Thin identification is the project of tying every action taken by an AI to some human principal. Thin identity will be essential for law to hold accountable the humans who make and use AI agents. Thick identification is distinguishing between AI agents, qua agents. sorting millions of AI entities into discrete, persistent units with stable, coherent goals.”
Minor#
- $100M for Alzheimer’s from OAI Foundation, finalised this month.
- OpenAI’s VP of Compute leaves, taking the other two Stargate leads with him.
- Basic cost-saving measure: use Sonnet ($15, fast) as your agent and let it use Opus ($25, slow) as an occasional Advisor through a tool call. Opus also watches context for dumb stuff. The two of them together are in theory a fast smart agent. Opinion: Only saves 10% atm but the multi-agent principle is a winner.
- Nice portrait of the real OSS AI scene. People love tiny dense models; people love Qwen (60% of all downloads and finetunes). GPT-OSS still has a devoted following 6 months later among those rich enough to run it.
- Competition among Chinese robotics companies is intensifying. The founder and CEO of Dreame is reportedly trying to poach Unitree’s chief scientist even if it costs $30 million, as well as all of the competitor’s customers, bidding projects, and even employees.