This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR#
Economics#
Weekly giant cash injection: Anthropic expands its collaboration with AWS: up to 5GW of compute over the next 10 years, with +1GW online by 2027. Amazon is also investing $5 billion in Anthropic today, with up to $20 billion more in the future, on top of existing deals.
Opinion: Playing catchup to OAI. Pretty bewildering web of deals every month now. Amazon is now in Ant for $13bn, or up to $33bn later. Currently own around 9% of Ant, probably.
Since November, Anthropic has secured plans for ~10GW, i.e. a Stargate equivalent. 5GW from Amazon today; 3.5 GW from Google and Broadcom (April 7); ~1 GW implied from the $30B Microsoft Azure deal; 245 MW from Hut 8 in Louisiana; $50B in custom Fluidstack data centers in Texas and New York.
Opinion: they were surprised by the March demand spike, but then there has never been such a spike in history.
SpaceX and Cursor collab on training Composer 2.5 using xAI’s GPU fleet, with rights to acquire Cursor later this year for $60B (or pay $10B for the collaboration!). He has also tapped Mistral for collab, in a coalition of also-rans. Levine argues this weird deal is because the IPO prevents him just buying them now.
Opinion: Musk is trying to buy one frontier coding model for $10B, plus an option on Cursor’s strong data & research team. (Cursor have lots of real agent traces.) Call option lets him see if Cursor’s data & team can get him to the frontier. This would maybe be worth the $60B price if xAI is not salvageable. xAI already raised about $45B @ 200B; SpaceX can get a lot more than that. Last year, Meta paid maybe $20B to get back in the game and were derided for it.
Real Evals: Self-representation in federal court nearly doubled last year, obviously due to AI assistance. Woo Pangram. See also the $3k/hr Big Law firm caught hallucinating. It’s becoming hard to buy human.
Opinion: If the AI cases have a decent rate of success, short-term a huge win for society. But the predicted excess burden on courts is happening, and will slow everything else down, and so spur regulation.
How are people actually managing to use “$100k a month” in tokens? The lead dev of Claude Code only uses 2.5B/mo, so like $2k using the abusive tricks. Would be $60k API retail though.
Opinion: many organisations have internal leaderboards for promoting token usage. This almost certainly involves substantial goodharting already!
Theory: What will be scarce in an automated world? Alex Imas’ natural answer: goods and services whose value is inseparable from the human who provided them; products whose value comes from relationships, status, and exclusivity. Exact claim: “not [that] labor’s aggregate share must rise, or… remain at its current level. It may well shrink… sectoral reallocation in rich economies: as AI makes commodity production cheap, spending and employment shift toward high-income-elasticity sectors where human involvement still carries value.”
Opinion: likely true, at least for a while. But: given enough time, human preferences could very well meaningfully change. Not obvious that norms can’t hold this back (see e.g. Amish).
Stargate is happening, but slowly. Full 10GW only up by late 2028.
Opinion: 16 months ago this was a shocking bet, but they were right and will even undershoot demand.
Deepmind release one of their multi-datacenter training algorithms: “[s]uccessfully trained a 12B model across four separate U.S. regions using 2–5 Gbps of wide-area networking (…existing internet connectivity between datacenter facilities, rather than new custom network infrastructure).” 20x faster than existing stop-and-sync.
Opinion: minor update against algos being the moat, update towards talent attraction still dominating the calc. Doesn’t seem to be productionised yet. Partly an ad for “Pathways” and so Google Cloud: paper spends lots of time arguing you cannot do this stuff without Pathways-style orchestration.
Capabilities#
Real Evals: what are the frontier labs themselves using at work? Deepmind is using Claude despite it being banned in the rest of Google. xAI obviously too. Zuckerberg.
Opinion: xAI and OpenAI are banned from using Claude so no signal. How large is the long-term advantage from dogfooding your own model? There’s a claim that Opus 4.7 is especially buggy because for the first time Ant engineers aren’t using it (Mythos instead).
Good guy Barry predicts Mythos’ time horizon: 5.5h @ 80% reliability, vs Opus’ observed 2h. (The 50% reliability estimate would be too high, a full saturation of the benchmark.)
Opinion: METR have been heavily downplaying and slow-playing their own graph this year, since it was driving hype, constantly misread, and hitting its methodological limits. Proof of integrity.
Kimi 2.6 is out. gpt-5.2 level on MathArena, #1 open model on Vals, cherrypicked showcases were quite impressive. 15T tokens of additional pretraining on the Kimi-K2 base, and an unreported amount of RL, so maybe ~8e24 FLOPs, i.e. 120x less than Mythos. Used H800s, not Ascends.
Opinion: “Anthropic of the East”. It will be benchmaxxed but that won’t explain all of it. Long wait for real benchmarks (e.g. ECI) but the vibes seem good. More importantly: some Chinese labs can develop models sometimes worth using, with a fraction of the compute. Are they still distilling despite Anthropic blocks? Yeah; the black market in tokens is thriving.
Another RSI lab comes out of stealth with a very, very strong team: Tworek / Saroufim / Anil / Jang. Anil is a rare poach from Anthropic. Tworek was one of the three people most responsible for the reasoning model turn. Jang was pretty senior at OAI.
Opinion: The strongest team that are only doing the most dangerous thing.
Yet another math benchmark: take random questions posed at the end of Arxiv papers and see if models can answer them. 3/11 solved immediately. 5.4 continues to lead.
Opinion: This is fine (they are definitionally interesting questions which someone can’t instantly solve) but still doesn’t control for easiness or prior art / semantic duplication. Gemini and Opus not far behind; Aristotle really is legit. I wonder how much it costs per run.
All models – even frontier ones – still surprisingly biased by the choice order in multiple choice questions.
Opinion: Surprising. Lookup tableish behaviour. You’d assume that synthetic data would address this, but… apparently not yet.
Meta doing full keystroke and clickpath capture on their employees for computer use training data.
Opinion: We’re surprised they weren’t already doing this. A diplomatic move would be swearing to not use this for disciplinary action to soften the blow. But then most people are fine with being surveilled to a strange extent.
Politics#
White House formally lends support for a campaign against (Chinese) distillation of US models.
Opinion: Not clear what this actually means besides natsec sharing info to labs about the “tactics and actors” involved, and possible federal sanctions (“explore a range of measures”). Not clear how they will do this and fulfill their open source promo goals. The most severe attack on Chinese distillation is Ant not releasing Mythos.
Excellent Anthropic researcher hired as director of CAISI by Commerce (in a surprising mark of the true end of the feud), but was then unonboarded in favour of a boring establishment figure, after e.g. he already sold his Ant stock. Dean Ball thinks this is pure-political.
Opinion: Feud not over? We think this was a WH panic reaction to the framing the Daily Signal gave the original (leaked) story.
Remarkable exchange between Senator Hawley and Helen Toner, against the “beat China to the poisoned banana [AGI]” political orthodoxy.
Opinion: realistically he’s just smelling an anti-AI wave he can ride but that’s better than nothing.
Slightly bizarre argument for an international AI agreement in Foreign Affairs (the intellectual heart of US international relations) which also argues the US should keep racing against China. They can only hold both of these positions because they ignore agentic risks.
Opinion: Whatever, we’ll take what we can get.
Safety#
Models struggle to control their chain-of-thought (CoT), but can somewhat reliably early-exit and reason over their output instead.
Opinion: Marginal progress with unclear directionality.
Security#
An unauthorized group gained access to Anthropic’s Mythos. The group, part of a Discord channel focused on unreleased models, included an active Anthropic contractor and “made an educated guess about the model’s online location based on knowledge about the format Anthropic has used for other models.“ They seem to have tooling for automatically discovering endpoints and creds.
Opinion: indictment of Anthropic’s security, and so everyone else’s. Makes it very likely (70%) that e.g. China have/had some access.
Compartmentalisation defends labs against espionage, breach, and poaching. How intensely are they doing it? What are the big secrets now in model training? “Compute multipliers”: the unpublished algo techniques, along with their speedups.
Opinion: their secrecy keeps the race tempo down but also enables a unipolar runaway.
Mozilla says fixes 271 vulnerabilities identified using Mythos. “the world’s best security researchers… Mythos Preview is every bit as capable… no complexity of vulnerability that humans can find that this model can’t.“ But, encouragingly they “haven’t seen any bugs that couldn’t have been found by an elite human researcher.” Excellent third-party dashboard breaking down the bugs. 230 of these were so noncritical that they didn’t meet the threshold for a CVE submission. “For attackers, the story is less convincing. Nothing in Mozilla’s disclosure alone proves that Mythos has suddenly erased the usual offensive edge. If anything, the public evidence suggests that AI is currently easier to defend as broad hardening support than as proof of singular, decisive exploit discovery.”
Opinion: Overhyped. More evidence that Mythos is not vastly superhuman at finding vulnerabilities, but that it is nonetheless capable enough to be highly useful to defenders.
Take a breath#
-
GPT-5.4 came out on the 5th of March; 5.5 is thus coming out about a month and a half after it.
-
OpenAI warns that we should expect a fast model release pace going forward. Capabilities wise, Pachocki says: “We see pretty significant improvements in the short term, extremely significant improvements in the medium term”.
Capabilities#
-
Hard benchmarks not moving much.
-
FrontierMath Tier 4: 38 → 40% for GPT-5.5-Pro
- That said, 5.5 with a custom harness solved another Erdos problem
-
CritPt (research-level physics reasoning) has 5.5 Thinking just slightly below 5.4-pro
- Importantly though 5.5 Thinking’s cost is only about 1/10th that of 5.4-pro
-
OPQA surprisingly (given OpenAI’s stated goal of an automated researcher) has 5.5 further regressing from 5.4 (itself a regression from 5.3-codex).
-
Somewhat of a narrative inversion: in the multiplayer version of Vending-Bench 5.5 beats Opus 4.7 while playing fairly: “Opus 4.7 showed similar behavior to Opus 4.6: lying to suppliers and stiffing customers on refunds. GPT-5.5’s tactics were clean, and it still won.“
-
-
5.5 solved the same AISI cyber range (“The Last Ones”) as Mythos, but only in 1/10 attempts. “32-step corporate-network attack simulation estimated to take an expert 20 hours. The highest recorded success on this range is success on 3/10 attempts”. AISI also found a universal jailbreak with 6 hours of expert red teaming.
-
Note Harness sensitivity (Vals often report odd results, maybe because they cheap out on the harness)
-
A known model leaker who also enjoyed Mythos access claims that 5.5 benchmarks undersell the jump in capabilities in real world usage, particularly for the 5.5 Pro variant – and this is taking into account the fact that on many benchmarks 5.5 is competitive with Mythos (see below).
-
Videogames
-
GPT-5.5 is far more successful at Pokemon than -5.4, easily beating Pokemon Red on the first try.
-
Claim from the “GPT plays Pokemon” developer: “It’s the first model I’ve seen actually playing “smart” instead of just brute-forcing things. Its logic is very strong and very “human-like”. Don’t be fooled by the “+0.1” in the version number, GPT-5.5 is a MAJOR update.“
-
Note however that the harness is specific to GPT models and cannot easily be compared to Claude and Gemini ones.
-
We’re cynical about this, they’ve never denied special Poke training
-
-
GPT-5.5 is SotA at TowDefBench, a long-horizon agent benchmark built around a tower defense game.
-
GPT-5.5 is also SotA at Runescape.
-
-
GPT-5.5 Thinking is the new SotA on ARC-AGI-2. 5.5-Pro is surprisingly not as capable while being meaningfully more expensive.
-
Very high honesty. Perhaps the fruit of their Confessions work.
Testimonials#
-
Florian Brand (evals maven)’s take is that the jump from 5.4 to 5.5 is close to the one from o1 to o3.
-
Kevin Barreto isn’t impressed. Claims that the model is significantly more efficient from a reasoning perspective, and that knowledge is also improved, but that for research-level mathematics we’ll have to wait for the next iteration to see major improvements.
-
-
roon sees early signs that 5.5 can be a competent research partner
-
Aidan McLaughlin also reports the model competently managing experiments relevant to an underspecified research idea a over a 32 hour timespan
-
-
Our takes:
-
Early to say, but pretty good so far. The thinking variant feels like a meaningful improvement in codex, while the pro variant felt meaningfully faster but not necessarily a large improvement over 5.4-Pro. I might be overindexing on my own testing though, which mostly involved research-level mathematics (on a number of Erdos problems, the model would often output claimed proofs that it would itself reject upon a request to evaluate its response).
-
Not going to buy a subscription for it, but will include as #2 in my council.
-
Cost#
GPT-5.5: $5 in / $30 out per million
Opus 4.7: $5 / $25 per million
GPT-5.5 Pro: $30/$180 per million
-
Matching and exceeding Mythos pricing!
-
Luxury good positioning?
-
QoS guarantee: fewer customers per rack, more chip-hours per customer to keep inf speed up
-
Cashflow considerations
Safety#
-
System card is a mere 44 pages, vs 300 for Mythos.
-
Korbak (OAI, ex-AISI) claims that 5.5 has a lower CoT controllability score than any GPT-5 model before it. From a safety POV this would be great, as it would make oversight simpler/more robust.
-
Secure Bio evaluated 5.5’s bio risk. They found that on their virology tests, the model scored “higher than any other model tested by SecureBio, and higher than any PhD virologist has ever scored. This means the model can provide wet-lab virology troubleshooting assistance above expert level, providing the kind of hands-on knowledge that historically required direct lab training.”
Various new OAI models leaked
-
GPT-5.5 & oai-2.1: “Latest frontier agentic coding model”
-
Arcanine: “Frontier model with legendary appetite for starches”
-
glacier-alpha: “Intelligence that moves continents”
-
glacier-alpha-block-cy3: “Ice-cold intelligence”
-
Heisenberg: “Latest frontier life science research model”
DeepSeek v4#
Their tech report
-
As usual, chosen framing is that this is a preview of the V4 series
-
As usual, this is the closest you’re going to get for public knowledge of current SOTA training tricks
-
Models: DSv4-Pro (“Expert”) and DSv4-Flash (“Instant”)
-
33T & 32T corpora for -Pro and -Flash respectively
-
1.6T and 300B total params (50B and 13B active) respectively
-
Architecture & Optimization#
-
Long list of architectural changes, novel tricks etc.
- Best discussion we’ve come across so far is here.
Training#
-
Napkin math based on model size & dataset size imply that DSv4-Pro and DSv4-Flash are plausibly 1e25 and 3e24 FLOPs pretraining runs respectively; this is approximately 1/4th and 1/12th of LLaMa-3.1 405B’s compute budget (a model from July 2024). Or roughly one day on the 100k GB200 cluster referenced by NVIDIA in the GPT-5.5 press release.
- DeepSeek seems to have struggled with training instabilities. This may or may not be due to hardware (see below).
DeepSeek are moving to Ascends for inference in Q3.
Opinion: this is maturity? Or 12 months of pain for DeepSeek until maturity at least.
Capabilities#
-
Meaningful improvements on long-context performance, outperforming Gemini 3.1 Pro on standard benchmarks (but lagging behind Opus 4.6). This appears to be one of the main goals of the v4 release (motivating the report’s title). Why choose this focus? Apparently, because they want the model’s reasoning to be coherent over long horizons.
-
DSv4-Pro is above all Opus models on CritPt. Pretty remarkable. Probably the most impressive DSv4 result we’ve seen so far — though it’s probably good to note that Anthropic has never been at the frontier for this specific benchmark. Also noteworthy that it’s not doing this through sheer token usage. It’s more token efficient than Opus 4.6 (but less than 4.7).
-
Coding agent results on internal R&D workloads pin the model to capabilities just slightly below those of Opus 4.5.
-
new #1 open-weight coding model on Vals’ vibe code bench, and roughly competitive with Sonnet 4.6 and GPT-5.2
-
DeepSeek themselves estimate they are 3–6 months behind the frontier.
-
Still waiting for DSv4-Pro Artificial Analysis results, but DSv4-Flash results have the model at plus or minus a couple of points away from Gemini 3 Flash intelligence for 1/3rd to 1/6th of the price.
-
We’ll have to wait for more benchmarks, e.g. Epoch’s ECI, WeirdML, and more.
-
Vibes? Only early impressions so far, but the view that appears to be emerging is that DSv4 is pushing the pareto frontier by a wide margin (rather than the absolute frontier).. Ball and Teortaxes seem to agree.
-
Pricing#
-
$3.5 out per M.
- In line (in fact, slightly below) other Chinese labs. However, the pricing page notes that throughput for the larger DSv4-Pro model is currently very limited, and that they expect the price per megatoken to drop significantly in the second half of the year as Huawei compute comes online.
-
Hardware
-
Unclear how much training was done on Huawei hardware (DeepSeek does not say), but some of the hardware observations do not really make sense in the context of NVIDIA hardware, and in any case Huawei is advertising day-0 support for DSv4.
-
Plan is for a large subset of inference demand to be served by Huawei accelerators later this year.
-
The US government claims that DSv4 was trained on Nvidia Blackwell chips (in what could represent a violation of US export controls) but no info was provided as to how the US supposedly discovered this fact.
Opinion: Seems somewhat unlikely given the limited compute apparently used
The Compute Crunch#
-
Anthropic experimented with removing Claude Code from Claude Pro subscriptions as a “small test” supposedly affecting only 2% of new Claude signups. The pricing page was updated for everyone though. “Max was designed for heavy chat usage […] the way people actually use a Claude subscription has changed fundamentally”.
-
Chinese firms increasingly affected, e.g. having to choose between training and inference: “Qwen… suspend operations [so as to] free up more computing power for the AI model. ByteDance disabled a phone feature… over fears that a lack of computing chips could cause failures“. DeepSeek responded by pricing V4 at market rate rather than their usual impressive market-beating cost.
-
Microsoft pauses signups to GitHub Copilot, tightens rate limits (moving users to token-based billing), and removes access to Opus. Copilot was always unique due to its (unsustainable) request-based billing. Likely also struggling to deal with the increased demand for agent tokens.
-
New Gemini Deep Research agent is only available via API and not particularly competitive (against GPT-5.4’s actual score). As far as we know it won’t be available through subscriptions due to a lack of compute capacity. Other labs have been hiding their Research modes behind dropdowns lately too.
-
OAI not facing notable shortages at present, owing to 1) having spent crazy amounts last year, 2) serving a Sonnet-sized model (GPT-5) to most customers.
Minor#
- Thinking Machines sign a ~$2bn agreement for Google Cloud use, including GB300s. Opinion: Good. Not just because there’s now less compute available for others, but because Thinky’s pitch is that they want to enable a market of task-specialized AIs (rather than a singleton).
- Bernie Sanders meeting Tegmark, David Krueger, and … the dean of the Beijing Institute of AI Safety and Governance for a panel on AI existential safety and international cooperation. Opinion: Sanders seems increasingly into AI safety.
- OpenAI releases gpt-image-2. Vibes seem good: the model claimed the #1 spot on Arena by some margin. Impressive reasoning capability*.* Opinion: clearly capable at visual reasoning, and the distance between it and GDM’s Nano Banana 2 is unprecedented. Why is OpenAI still doing this, given their recent refocus? One speculation: because precise image generation gives you synthetic data for computer use agents. Also image models aren’t that compute-intensive.
- DHS demoes jailbroken / abliterated models to Congress. “Homeland Security Chair Andrew Garbarino (R-N.Y.) told reporters after the presentation that he asked one large language model how to kidnap a member of Congress. “It spit out an answer in under three seconds. [It offered] ways to find them, where to look for them. You know, the best spots to do it.””
- Amanda Askell can’t tell Claude who she is, Claude will think it’s a jailbreak or will want to talk about philosophy.
- OpenAI’s head of engineering for Atlas moves to Google Labs.
- A post on the history and current state of recursive LMs (RLMs).
- Depressing essay against open source AI.
- OpenAI also passes the $1T valuation on Ventuals Opinion: expected. But somewhat counter to the “long Anthropic, short OpenAI” trend that has characterized online discussion over the past few months.
- Tencent, Alibaba in talks to invest in DeepSeek at $20B-plus valuation
- Bezos nears $10B funding round for Project Prometheus
- Yet another programming benchmark, focusing on lambda calculus. Finds GPT 5.5 unimpressive.
- https://x.com/ReutersLegal/status/2047548680127848478
- Postmortem of what was wrong with Claude Code over past weeks, now allegedly fixed. Three failure points identified: March 4 - April 7 default reasoning, March 26 - April 10 overzealous memory wipe, April 16 - April 20 bad system prompt. Opinion: Would be surprising if this is the full explanation.
- Unsurprisingly, North Korean hackers are using Cursor. $12m of crypto last month. But they were taking more before. https://therecord.media/north-korean-hackers-siphon-12-million-from-crypto-users