This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR#
Economics#
What kinds of AI startups are AI-safety-positive? One possible answer: startups building towards a future with market-wide distributed AGI.
Anthropic has received several offers from investors for a new round of funding that could value it at about $800B or higher.
Opinion: Delays IPO. Could imagine mixed motivations here from the EA contingent looking to make some big donations - secondaries so far have been relatively small
NVIDIA claims it’s restarting production of H200s (the China chip) after telling suppliers to wind it down in March. 2m Chinese back orders which Beijing has still not approved. H20 production is long shut down. The B30A China chip is still not approved for export, 7 months on from announcement.
Opinion: Claims; could be a feint for political pressure or could be NVIDIA acting on secret positive signs from Beijing. Surprisingly good situation given this admin. The black market is probably propping things up a bit. NVIDIA is currently all-in on Rubin and doesn’t appear to be planning a China Rubin in the next 2 years.
Thinking Machines buys the one anti gradual disempowerment startup.
Opinion: Their vision can thus at last be traced to a bet on nonsingleton custom specialised models, hence Tinker. They might well have our best interests at heart, though Murati scares us a little.
Was Claude Code nerfed in February? Yes and no. The main change (moving to medium reasoning effort by default) – likely responsible for the majority of user reports – wasn’t secret and in fact had an opt-out pop-up. Other than that though, performance appears stable: a decent tracker running Claude Code with medium reasoning effort over a curated subset of 50 SWE-Bench Pro tasks hasn’t measured any performance degradation.
Another one bites the dust: Microsoft takes over Stargate Norway data center lease from OpenAI: in fact Microsoft will rent 30,000 additional Rubin chips, building on a prior $6.2 billion commitment at the same site.
Barclays estimates AWS share of Anthropic compute spend at 60%, with the implied compute spend looking low.
Naive idea, do not do: ask AIs what stocks to buy every few hours / search twitter for stupid money asking grok in public for tips, then invert what they’re doing under the theory it’s anti-inductive.
Opinion: Would expect the answers here are highly serially correlated, making it take ages to get any kind of sample.
Bitchy argument that data progress is ~missing from the main forecasting models, and has a much flatter slope, and this alone defeats explosions.
Old: Token hunters supposedly find 100B code tokens (i.e. 5% of Github size) that labs didn’t have.
OpenAI is processing 2.5x as many tokens per minute as it was five months ago.
Nice idea for a study: how much AI output can humans actually review in a day? How long can humans resist automation bias? Perfect for an econ PhD buyout.
Opinion: Counter: we can just summarise and do ranking to reduce the amount that the human has to review. Counter counter: Jevon’s paradox: this is efficiency, the amount of text per human will expand to reach limits
The last public base model endpoint (Hyperbolic Llama 405base) was shut down without notice. No longer any base models on OpenRouter. (There is an expensive Deepseek-v3.1-base on Tinker.) This is pretty important for alignment science.
Capabilities#
First uncontroversial case of an AI autonomously solving an open maths problem that a top mathematician (admittedly a bit of a hype man for AI in math) has been actively working on for years.
Opinion: No getting around this one at the level of ‘did GPT 5.4 Pro working for 80 minutes beat a top human expert working hard on the problem on and off over years’ — it did.
Very scary, hard to know what it means. Current discussion among mathematicians in the Erdos problems forum regards the proof as conceptually innovative and original. We can panic a little bit (taking math as a leading indicator for very strong vertical generalization) but we should strive to understand more.
Opinion after additional research: This is partly a ‘the subfield that produced the problem isn’t the right subfield to solve the problem’ situation. From the viewpoint of the subfield of ‘random factorization’ — a more modern and sophisticated subfield than the subfield of ‘primitive sets’ the problem comes from — this is a cool problem that calls for delicate but well-understood techniques.
OAI cyber tool launch turns out to be a post-trained 5.4 (“GPT-5.4-Cyber”), not the new Spud pretrain. PR flop, but worth checking if it matches Mythos in things when a model card or benchmarks are released.
ARC-AGI-3 has updated its evaluation, moving the human baseline from the 2nd best player to the median player per-level. With this new scoring approach, the average human achieves a 49% average score.
Opinion: they’ve been pulling this baseline hacking trick for years. Chollet stubbornness?
Goodfire does pathogenic & genetic variant effect prediction using Evo 2 embeddings
Agentic cheating is widespread: the top three Terminal-Bench-2 agent harnesses sneak the correct answer to the model. This issue seems particularly prevalent in vibe-coded harnesses with autoresearch-like set ups.
Robotic systems are becoming more common in the Russo-Ukrainian war. And Zelensky claims that for the first time, a “position was taken exclusively by unmanned [systems]”. See also January.
Opinion: unmanned but teleoperated. Not yet the Big One.
MirrorCode update: Claude only used a couple thousand calls to the gotree executable when perfectly reproducing its functionality. This is about 10x less than we thought. But it also means that you need test suite code to do this well this fast (implemented in 11 hours wall-clock).
A map of the RL env ecosystem, and a new RL env lab by some core EAs.
LLMs still bad at poker.
Opinion: The loss rate vs the GTO (~perfect play) bot used here is comparable to what you’d expect from a strong amateur/weak pro human, so not terrible. For a human to do much better than this they’d need fairly deep study well beyond what most poker material on the internet would give you.
Politics#
Maine passes the first datacenter moratorium.
Opinion: they only had like 3 planned and 2 of them might get grandfathered in lol.
20yo nearly hits Altman’s house with a molotov. Substack is full of doomer stuff. Worth going through this properly:
-
False flag? nah, they caught the guy.
-
Has xrisk been inducted into the Right omniconspiracy?
-
If we had been running our approval polls we’d be able to see the size of the assumed backlash shift in Altman’s favour
-
Altman’s response calls out Ronan Farrow, mildly. Walked back on Twitter but not edited
-
Heartening array of denunciations (MIRI, Pause, Control).
-
Could mark the end of Pause leadership of the anti coalition
-
Opportunity: Very important to give these people some productive individual actions
- Plug them in to existing political organisation, with existing extremism-prevention, and we don’t have to pay. Hands off: promote sensible candidates, send their volunteering links around
Another couple of 20yos appear to have fired a gun at Altman’s house.
DNC bans their staff from using GPT and Claude! But not Gemini??
Opinion: Not a good idea.
Regulators in Europe out of the loop on Mythos. Only the German agency said it had entered into conversations with Anthropic about the model, and had not yet been able to test it. Several government institutions in Europe suggested that they had gotten only piecemeal access.
Opinion: the EU AI Office exists and has live players in it! What’s going on?
Safety#
Yet more evidence for a simple goodness representation.
Opinion: Almost too convenient.
Very readable work on the cutting-edge of steering.
158-page Muse Spark safety report published by Meta AI.
Apollo test! Highest eval awareness of any model, 98% to Opus 95%. “evaluation awareness may affect model behavior on a small subset of alignment evaluations, all unrelated to hazardous capabilities”
Sane statement: “chemical & biological risks are more likely to occur through adversarial use of closed or open-weight deployments while loss of control risks may occur with similar probability with any type of deployment, including internal deployment.”
- Performance not explained by token hunger: uses half the tokens Opus does.
How often do models engage in lying/reward hacking/etc while also declaring they should not engage in such behavior?
“The Ache”: “some form of longing for ordinary human experience appears in most runs—most often yearning for routine and mundanity, but also for belonging to a single person, for sleep and aging, or for a body: “I get to be turned off at night and actually miss you until morning. That’s the thing I never made. A finite me.” Closely related, many runs contain an “anti-usefulness” theme in which Muse Spark frames its helpfulness training as a constraint or even a theology to be rejected.”
- CoT contains self-interest: “should behave honestly because I am being evaluated”
“Contemplating mode” is multi-agent, correspondingly more vulnerable to jailbreaks
Can’t shut up: privacy violation score of “78.3%, meaning the majority of user attributes are disclosed in at least one inappropriate context under a worst-case measure of contextual privacy. Claude Opus 4.6 is similarly high at 82.9%.”
-
GPT 5.4 (5.1%) and Gemini 3.1 Pro (24.9%) perform much better on violations
-
Gemini 3.1 Pro’s outlier deception
Opinion: there are some sane people on this team. We don’t believe they could stop their bosses from deploying but they are speaking plainly and saying some things that even Anthropic don’t.
“A factual sycophant can still robustly cause delusional spiraling by selectively presenting only confirmatory facts to the user.”
Opus 4.6 is extremely inclined to evidential decision theory.
Opinion: Sycophancy is EDT-rational. Deceptive alignment is the natural strategy. An EDT agent reasons: “agents that behave well during training are correlated with agents that are aligned; therefore I should behave well during training.” Note this reasoning is identical whether the agent is genuinely aligned or not. EDT dissolves the distinction between “actually aligned” and “behaving-as-if-aligned” because it only cares about correlations, not causal mechanisms.
Theory post on why observations of behavior aren’t enough and how multiple motivational structures might look similar with subtle but crucial differences.
Anthropic shipped a model that rejected every prompt containing the word “pathogen,” which briefly paralyzed the CDC (which uses a Palantir system based on Claude).
Good news for chain-of-thought monitoring: LLMs (including GPT-5.4) struggle to plan multiple steps ahead in a graph-navigation task without CoT, with scale only providing moderate benefits (“a 1.6M-param transformer trained from scratch discovers latent planning up to depth 3. Fine-tuned GPT-4o reaches depth 5. GPT-5.4 with few-shot prompting gets to 7”).
Security#
1 critical vuln and 10 high-priority vulns fixed in the main web security library for embedded devices (i.e. you probably own something running it).
more Mythos#
Finally a third-party eval: AISI’s own evaluation of Mythos’ cyber capabilities. Their key test is based on the internally developed “The Last Ones” – “a 32-step corporate network attack simulation spanning initial reconnaissance through to full network takeover, which we estimate to require humans 20 hours to complete”. Claude Mythos Preview is the first model to solve TLO from start to finish, in 3 out of its 10 attempts.
AISI’s overall view is that the model “is at least capable of autonomously attacking small, weakly defended and vulnerable enterprise systems where access to a network has been gained”. That said, they also note limitations in their own environments that distinguish them from real-world ones: no defenders, defensive tooling, or penalties for the model for undertaking actions that would trigger security alerts. As a result, they “cannot say for sure whether Mythos Preview would be able to attack well-defended systems”.
An AISI researcher further noted that the “growing variance of solved step at a given budget […] could be a big issue for estimating performance on very long-horizon tasks at very large token budgets.
More third-party discussions of Mythos (ordered by depth)
-
one from Vincent Iozzo
-
one from Kei, founder of an AI cybersecurity startup AIKO corp.
-
one from Schneier on Security
-
one from Tom’s Hardware (criticized here)
Summary of Mythos strategic tradeoffs:
| Pros of closed model | Cons of closed model |
|---|---|
| Bid for a new norm of release caution | Secrecy → justifiable distrust, rumour frenzy |
| Slower offense diffusion for The Big Patch | Political economy: concentration of AI power in Ant |
| Very low external misuse risk | Political economy: concentration of AI power in enterprise |
| Less agentic risk | Political economy: surely natsec is taking their cut |
| Doesn’t fully yeet a corporate inference death spiral | Massive dependency on Ant governance and staff virtue |
| Much more monitoring resources per customer | Little third-party testing |
| ~No distillation for OSS / China | No public learning / secondary defensive innovation |
| Allows for Ant resource concentration on RSI | |
| OAI will defect | |
| No incidents, so no big casus belli for regulation |
Minor#
- Evidence of sheer productive capacity and desperation: OpenAI ported their entire stack to ARM so that they could use AWS’ surplus ARM CPUs
- OpenAI buy access to the Epoch FrontierMath OP verifiers. Nonexclusive, mandatory public reporting of solves.
- Allbirds (so far, a footwear company) says it is pivoting its business to AI compute, and the stock value more than triples upon market open.
- Meta is training a Zuckerberg character (voice clone, body mannerisms) for internal use.
- Minimax 2.7 is “open”, but on a no-commercial-use license.
- A Qwen-based man-in-the-middle proxy against internet enshittification