This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR#
Economics#
The OpenAI/Anthropic/Deepmind lead remains robust
Meta delays their new “Avocado” model, citing poor performance. It nominally beats Gemini 2.5 (March 2025) but not 3
Opinion: They have a bunch of world-class people (though zero of the most tasteful, besides Friedman). Suggests that catching up to the frontier is very hard: talent and compute aren’t enough. This probably cost them $35bn.
xAI continues to burn out senior staff. Down five cofounders this month: Tony Wu, Jimmy Ba, Toby Pohlen, Guodong Zhang, Zihang Dai. Only two of 12 remain. This includes most of the Grok Imagine seniors.
Opinion: Bad (to the extent their compute might be put to productive use by more competent actors if they disappear), good (in lessening race dynamics). This is a larger loss than post-coup OpenAI. But the extreme turnover is probably less lethal than usual: Musk has a history of rebuilding after such things and it’s hard to shake investors’ confidence in him. Little bit of grovelling, and he’s already poached two heads from Cursor. But in the short term he will struggle, the mood is against them. Obvious move is to try and go literally AI-first, running autoresearch at million experiment / day scales.
Chinese companies unable to meet demand for agents, lots of downtime even relative to Anthropic.
“New” “frontier” lab?: NVIDIA double down on open models with a $26bn pledge over 5 years (12% of last year’s revenue, so maybe 1% of expected revenue over 5 years).
Opinion: they’ve always spent a bunch, hiring thousands of ML engineers and producing solid models which don’t even reach the open frontier. We estimate they were spending $2bn/yr on model dev previously so new pledge only 2-3xes. It is in Nvidia’s interest to democratize access to good models, “commoditising their complement”. Even the attempt to commoditise model capabilities creates pressure on big labs to race for RSI. $5bn is plausibly as much as all Chinese labs combined are spending on model R&D right now.
Anthropic launches the Anthropic Institute, “a new effort to confront the most significant challenges that powerful AI will pose to our societies”, bringing Anthropic Societal Impacts, Red Teaming and Economic research under a new umbrella. Led by cofounder Jack Clark, with two excellent economists and Matt Botvinick (ex-Deepmind lead, law) as early hires.
Variables hit: labs self-policing, measuring societal impacts of AI
Opinion: A bit vague on what’s new here vs a restructure and brand segmentation.
Alibaba walk back their recent “agent going rogue and starting to abuse GPUs for crypto mining” incident. It was real (the AI did unexpectedly try something bad) but fake (the AI was simulating a rogue as part of its security tasks, not endorsing the rogue’s goal) and low-impact.
> We had a model tasked with a security audit — specifically, investigating abnormal CPU usage on a server. Somewhere along the way, it went off-script and decided to simulate a cryptocurrency miner to “construct a suspicious process scenario.”… our safety monitoring caught it immediately. The whole thing happened inside a strictly isolated sandbox — zero impact on anything external. The incident has been logged and will be used as a negative example in future RL training to reinforce what’s off-limits.
Variables hit: loss of control risk, broad autonomy
Opinion: Probably the most benign possible version of the story that doesn’t imply they were making it all up. Note that the practical difference between an LLM doing something and simulating doing something is minimal, though the implications can be different (e.g. how much it persists in trying, how much it resists correction).
Claims of multi-year lows in compute availability.
Variables hit: race dynamics, utility of existing models
Opinion: to the extent this is inference demand growth driven, suggests compute investments are less risky, less pressure to push for RSI from suppliers of compute (hyperscalers, semis). Good.
Capabilities#
METR study on SWE-Bench Verified gives yet more reason to distrust it: the test suites are totally inadequate to proxy correctness of the submitted programs: maintainers on repos used in SWE-Bench Verified review “successful” graded PRs, wouldn’t merge half of them.
Note that METR’s own HCAST involves much more correctness-checking than just running the test suite.
Opinion: Gap between benchmark results and real world application continues to exist. Unsurprising but useful result along the lines of replication.
OpenAI: Interpretability on RLHF reward models - training a model to generate a human legible rubric to match the preferences expressed by human graders in RLHF. Helpful for exploring undesired behaviours, e.g. models learning to be performatively helpful vs accurate reasoning, and catching these things in the training process
Opinion: We think this is big, but even more for capabilities. Very useful for detecting issues (eg sycophancy, hallucinations). Likely to improve models in the long run, but is only one part of the puzzle; doesn’t directly provide alternative training methods to avoid the issues yet besides “repeat until it doesn’t have the issue”.
Gemini Embedding 2 manages to be the SOTA multimodal embedding model for all of two days before being upstaged by small lab Mixedbread
Opinion: Suggests lots of room for improvement in various domains possible with focused effort, good for tool world if small labs can compete on things, bad insofar as it puts pressure on labs to go for RSI if other things can be done cheaply.
Opus still making dumb errors frequently in places you wouldn’t expect.
Opinion: Fits with our experience, clearly pro-tool world to the extent these problems are hard to stamp out
Natural extension to Karpathy’s autoresearch: hub for agents to work in parallel and share results, a la Moltbook. Dangerous because it lets the internet fund really large distributed sweeps.
Math#
GPT-5.4 pro, prompted by a math hobbyist and a Cambridge undergrad (worked with GDM before), appears to have solved the first FrontierMath Open Problem. For context, FrontierMath Open Problems are “a collection of unsolved mathematics problems that have resisted serious attempts by professional mathematicians”. Epoch is still waiting to hear back from the problem’s author, but it’s likely that the solution is correct, given that it was autoformalized in Lean (likely through Math Inc’s Gauss).
Opinion: The problem is only “moderately interesting” (the least interesting category of FrontierMath Open Problems) but it would nonetheless be a meaningful result. We are likely to see more such results as Math provides perfect automated feedback through formal math environments like Lean. OpenAI appears to agree: one of the authors of this effort is expected to join the OpenAI for science team later this summer to further advance AI for mathematics.
GDM got Alpha Evolve to improve the lower bound for 5 Ramsey numbers.
Opinion: Not sure how impressive this is. Ramsey theory famously poses significant challenges. However, unlike most mathematics, combinatorics is particularly search-like, so progress here might not generalize to other areas of research.
Politics#
- “[The US government] is affirmatively reaching out to [Anthropic’s] customers & urging them to stop working with Anthropic.”
- Removal of Anthropic models from military systems happening now
- Big boost (50% monthly growth) for Anthropic web traffic lately. Still small in the overall consumer segment (the bump shown is chat-only, not claude code/cowork).
- Anthropic-DoW injunction hearing set for the 24th.
- Google, Amazon, Apple, Nvidia and Microsoft all file amicus briefs in support of Anthropic vs DoW. Microsoft one is quite spicy.
- Forecast of supply chain risk designation being lifted by May 1st of 43% (similar to prediction markets 50%, useful for granular model)
Opinion: Not clear on direct impact of lifting designation in the short term - probably military suppliers still won’t want to use Anthropic to avoid DoW ire. Plausibly reduces the amount of non-military harm done to Anthropic. USG actively reaching out to customers suggests they are keen to maximise damage here.
OpenAI and DoW sending mixed messages over whether intelligence agencies are included in their deal, OAI saying no, DoW saying yes.
Opinion: Overwhelming power of the DoW to be balanced against their recent habit of breaking the law in ways that Congress and courts won’t retroactively approve.
Related: Palantir announces a “sovereign AI OS reference architecture” with NVIDIA, a combined hardware and software system that lets governments and large organizations run AI on their own infrastructure without relying on external cloud services.
Opinion: Palantir positioning themselves as a middleman here, also allows systems to be more model/provider agnostic.
Bernie calls for a moratorium on data centres, explicitly cites x risk and quotes senior tech leaders at length saying that AI is going to replace human labour/be transformational.
Opinion: Unlikely to lead to any policy near term but notable that this noise is building. How bipartisan it ends up being will be an important factor.
Polling by an EA lobbyist org finds American public prefer guardrails > ban AI > no guardrails
Opinion: Propaganda: bad poll structure - guardrails unspecified so voters can fill in the blank with whatever they want, naturally this will poll well. Even “ban vs no guardrails” is a very loaded framing in this context.
Study analysing LLMs discourse in terms of moralising tone, finds it more moralised as a topic than vaccines and GMOs
Opinion: Says what you’d guess if you spent 10 minutes on Bluesky looking at AI discourse
Safety#
Win for Subtle alignment: Now that Claude’s Constitution is public we can test models against it.
Study of whether Claude models violate their constitution, finds rate of violation decreasing over time
Opinion: Good news, but not groundbreaking. Highlights: GPT-5.2 still caves under pressure while Sonnet 4.6 doesn’t, but it along with other Claude models still fabricates data, lies when instructed to and takes unilateral actions.
Obviously agents will have no shortage of communication channels, but mildly interesting to see current crappy agents inventing them in passing:
Minor#
- Meta acquires Moltbook, presume acquihire
- Claude chat gets big multimodal display upgrade - it will make inline interactive flowcharts, use a visual chessboard to display chess moves, etc.
- February uptime drop for OAI and Anthropic, vibecoders getting blamed.
- Time piece on Anthropic. Some attention to RSI claims and company culture, ties to EA etc. Nothing new for us.
- Beginning of a proposal for a heightened AI security protocol. As advertised the protocol is strict; unclear if it covers all necessary bases. Doesn’t have much momentum or adoption yet.
- “Productive people don’t make productive firms“ article doing the rounds on Twitter, exploring why firm productivity is growing less with AI than individual productivity, suggests new organizational structures needed.
- Encouraging comments indicating care about safety from two of the most powerful people at Deepmind, but they are just words so far.
- Prinz (pseudonymous US BigLaw senior associate/partner) on why he thinks AI kills BigLaw (a lot of BigLaw value prop is niche specialisation, surge capacity and availability, AI tools target these directly and enable smaller outfits to compete).
- SemiAnalysis now thinks chips are a bigger constraint than power and data center construction. Epoch thinks the bottleneck on chips is CoWoS and HBM rather than logic die fabrication capacity (both expecting big capacity surges next year).
- New LLM sycophancy benchmark finds big lab frontier models not very sycophantic.