This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR#
Economics#
Google is investing $10B in Anthropic at a $350B valuation, with another $30B contingent on Anthropic hitting performance targets
Opinion: We had some back and forth about whether there’s a signal here about Google’s faith in GDM or GDM’s specialization plans relative to Anthropic, but we concluded that (as Thumas Kurian claimed in a recent interview) it’s likely not that deep.
Anthropic explain why Claude Code got worse in the last month. Looks like the ‘nerfing’ complaints from the Claude user community were true this time. (Is this a broken clock situation or are most ‘Claude was nerfed’ meme-waves correct?) The explanation is, of course, to save compute:
‘1. On March 4, we changed Claude Code’s default reasoning effort from high to medium to reduce the very long latency—enough to make the UI appear frozen—some users were seeing in high mode. This was the wrong tradeoff. We reverted this change on April 7 after users told us they’d prefer to default to higher intelligence and opt into lower effort for simple tasks. This impacted Sonnet 4.6 and Opus 4.6.
-
On March 26, we shipped a change to clear Claude’s older thinking from sessions that had been idle for over an hour, to reduce latency when users resumed those sessions. A bug caused this to keep happening every turn for the rest of the session instead of just once, which made Claude seem forgetful and repetitive. We fixed it on April 10. This affected Sonnet 4.6 and Opus 4.6.
-
On April 16, we added a system prompt instruction to reduce verbosity. In combination with other prompt changes, it hurt coding quality and was reverted on April 20. This impacted Sonnet 4.6, Opus 4.6, and Opus 4.7.’
Opinion: The compute crunch is affecting everything! Maybe we should start developing some kind of model of the compute crunch and its impacts?
Microsoft and OpenAI update their deal:
-
OpenAI can now serve all its products to customers across any cloud provider.
-
Microsoft will no longer pay a revenue share to OpenAI
-
Microsoft will continue to have a license to OpenAI IP for models and products through 2032
-
Revenue share payments from OpenAI to Microsoft continue through 2030
Opinion: OpenAI are betting that losing revenue share from Azure will be more-than-compensated-for by:
A) Gains from making OpenAI models available on other clouds like AWS and Vertex.
B) Removing legal/financial ties to Microsoft to clean up some potential pain around the IPO.
Very high price though!
EpochAI publishes a new analysis looking into how fast production of humanoids, quadrupeds, drones, and other robots could scale up in the event of a large demand shock. Findings:
-
Currently, production of humanoids is growing the fastest (16k units, doubling every 6 months)
-
Assuming a demand shock, by EOY 2030 plausible outcomes for production are roughly 5–10M humanoids/year, 8–15M quadrupeds/year, and 100–200M drones/year.
Opinion: The analysis has limits (e.g. it doesn’t really consider the effect of robot production on accelerating robot production) but mobilizes solid quantitative data and the conditional scenario is well-constructed. We’d more or less adopt their model on this.
New study estimates consumer welfare gains in the US from AI, measured via willingness-to-accept compensation, at about $116–172B. The authors suggest that “consumers capture most of the welfare gains from these tools” as this surplus “substantially exceeds estimated revenues from generative AI in the United States”.
Opinion: The paper doesn’t address the possibility of a ‘Red Queen dynamic’ (you need GenAI access in a society that is increasingly structured around GenAI use) at all. Given that workplace use of GenAI is central to the study this is a large flaw.
Meta signed a deal for millions of Amazon’s ARM-based Graviton CPUs (not GPUs) to handle agentic inference workloads.
Opinion: We know that we are approaching compute crunch in CPUs much like the one we’ve seen in GPUs. This is Meta positioning itself accordingly.
Cohere and Aleph Alpha are merging at a ~$20B valuation, with Schwarz Group putting $600M into Cohere’s Series E and both the Canadian and German governments backing the deal.
Opinion: These are pretty much irrelevant companies from an AI race perspective. It’s unclear whether there is any meaningful strategy here – they don’t have the funding nor the raw talent to get back on the frontier. The most likely explanation is that they are trying to become their respective countries’ preferred AI play. A too-big-to-fail of sorts.
Forbes argue that the rise of AI actors over human movie stars is imminent.
Opinion: The topic of Hollywood & AI itself brings up interesting questions though: Does the compute crunch mean the AGI race and the entertainment industry’s AI aspirations are competing over finite resources? In theory there were supposed to be synergies between generative video/world-generation and AGI research, but SORA as a physics/vision/world-model project seems to have been a dead end for OpenAI.
Capabilities#
Qiaochu Yuan—one of the top posters on Math Overflow—is surprised by GPT-5.5’s mathematical abilities. Claim: “At this point it seems plausibly good enough to be able to quickly answer casual mathematical questions up to somewhere between graduate-level and research-level which is pretty wild”. Says that he hasn’t found anything he considers trivial that the model can’t solve yet.
Opinion: We already knew this. Yuan (by his own admission) hasn’t looked into LRMs’ math abilities in about a year, so he’s comparing 5.5 to his experience of 2025 models. Not an update for us.
Related reflection: Current state of LRMs in math is essentially where it was after the First Proof challenge shock, with incremental gains in model power and—more dramatically—solidification of the evidence. After the Erdos #1196 proof, we can be fairly certain that the First Proof challenge performance wasn’t a fluke and wasn’t an artefact of (very) shallow generalization via lemma redundancy.
More GPT-5.5 benchmark results are also coming out, including benchmarks that aren’t very Goodharted:
-
5.5 tops Mercor’s APEX-Agents leaderboard (autonomous agent tasks) and scores well on APEX-SWE (software engineering).
-
5.5 appears to be a meaningful improvement over 5.4 when it comes to (not) proving false mathematical statements. Unclear how much better the model is vs 5.4-pro.
-
5.5 also tops ErrataBench, a benchmark about proofreading human-written text, but also shows no progress, or even regression – particularly for the 5.5-pro variant – on BullshitBench, a benchmark that measures whether AI models challenge nonsensical prompts.
Opinion: 5.5 is clearly a strong model, but we gain see that on hard benchmarks like Apex-Agents or (as discussed in #16) FrontierMath the improvement between 5.4 and 5.5 matches standard incrementation in the 5.x series.
Unclear if the partial return to classical scaling is really paying off for OpenAI, although maybe it wasn’t about getting dramatic gains but about breaking out of an internally known diminishing returns regime when holding network-size constant at subopus sizes.
Informally some users are speaking of 5.5 as a breakthrough in agenticness.
AI coding agents, given only raw method descriptions and data, can reconstruct full quantitative social science research papers. With the best models, the largest share of divergences stems from the method descriptions in the papers being insufficiently precise to enable faithful reimplementation.
Opinion: Not a deep capabilities update but possibly very useful for reproductions in quantitative social science! Partially solves the problem of social science papers’ code being both the element which receives the least scholarly scrutiny and the element that most directly affects the result of replications.
Daniel Litt discusses his most recent AI-assisted work in heavyweight research math. What’s most interesting here is the application of AI, which is typical of current use in heavyweight research math: not to prove the main results, as we’re increasingly starting to get used to in “mildly interesting” research math, but by expanding upon the work after the main result had already been obtained. “[LLMs] had a really significant impact after the first draft of the paper was written […]”.
Core passage from Litt’s thread:
“Let me start by explaining this example. The map \mu^\vee is represented by a 55x51 matrix; to run our final argument, we needed that it had rank at least 50. We proved this by hand; computing in examples, though, suggested the rank was exactly 50. I tried a bit to prove that the rank was exactly 50 without much luck; no AI model was able to prove this correctly “out of the box” either. But studying the examples we’d computed suggested a natural geometric explanation for the kernel of this map. When prompted with this conjectural explanation, Gemini Deep Think was able to produce a more or less correct proof, cleverly combining some somewhat obscure facts about cubic 3-folds. I think this is a nice example of human-AI collaboration.”
Opinion: Understanding the forms in which LRMs currently regularly contribute to heavyweight math research can be very valuable for positive ‘tool-world’ visions. In general, while ASI of course dominates all human-labour AI augmentation, to the degree that we either expect ASI to be far off or expect to be able to push off ASI it’s worthwhile to consider types of AI application that can be explored and perhaps improved in targeted ways (either at the level of scaffolding and wrappers or training) on a non-AGI-bound AI development trajectory.
New PR material from superintelligence lab Ineffable AI – which just raised $1.1B at a $5.1B valuation – has big crackpot energy, but they do have an incredible pedigree (namely David Silver, who led AlphaGo). Some of the PR talk from the CEO sounds straight up Pascalian: ‘There is a real risk of failure, in pursuit of a small chance of extraordinary success that could change the course of AI, and with it, humanity.’
Opinion: The pitch is AGI through RL from scratch, no imitative pretraining. Can it work? There’s no recent positive evidence we’re aware of, but David Silver is a major researcher not far from his prime and ‘take an intuitive idea that has never worked out but try harder this time’ is how a lot of modern AI works. Safety-wise, there’s a school of thought that holds that imitative pretraining is where all the relatively nice alignment properties of LLMs come from and successful pure RL can only produce something Yudkowskian. We lean towards agreeing with Alex Turner that this a confused school of thought, but it has very sophisticated versions — like Steve Byrnes’, which directly talks about David Silver’s theoretical work.
Politics#
AVERI survey of frontier-AI-auditing-related legislation in the US. Some key points:
No requirement for frontier AI audits has yet been passed
AI labs lament a number of issues concerning audit requirements, including
a lack of clear standard for safety and security audits,
the fact that a big fraction of expertise on frontier AI is concentrated in the very companies being audited (meaning that the outside expertise that exists could be stretched thin by requiring too many audits too quickly),
a worry that audits could make it difficult for labs to protect sensitive intellectual property and user data, and
a worry that audit requirements could hinder innovation by creating excessive compliance burdens
AVERI proposes four principles that can inform the design of relevant legislation:
Auditors could verify companies’ compliance with their own policies and share appropriately redacted versions of their findings
Redacted versions of audit findings and details on the audit process could be published in order to hold auditors publicly accountable for conducting rigorous analysis rather than checkbox exercises
Rather than writing detailed rules for safety and security into legislation, policymakers could instead designate one or more sources of formal standards or more informal best practices
- Audit-related legislation could accommodate uncertainty around which audit providers will emerge (and by what time they will do so) through temporary “release valves”, including adjusting the thresholds that trigger audit requirements, and delaying enforcement of new standards.
Opinion: Some really great suggestions from AVERI.
A watchdog monitoring the practices of AI companies (‘The Midas Project) discovered that the news site The Wire by Actus is an AI-run pro-AI astroturfing campaign funded by OpenAI’s superpac. The news pipeline seems to be almost fully automated, with AI agents proactively reaching out to third parties for comments and writing entire articles. The site appears to be operated by the Republican PR firm Novus Public Affairs and likely funded by Leading The Future (LTF), the OpenAI-affiliated super PAC, via the GOP consulting firm Targeted Victory. LTF denied being aware of the platform — claiming a third party has set it up without consulting them — while noting that “[the platform’s] engagement with [LTF] has been terminated”.
Opinion: It’s unclear to what extent OpenAI’s shadiness generalises and who should be concerned about it to what level.
Safety#
Greenblatt (Redwood Research) argues that current AI systems are pretty misaligned: “they oversell their work, downplay or fail to mention problems, stop working early and claim to have finished when they clearly haven’t, and often seem to “try” to make their outputs look good while actually doing something sloppy or incomplete” – apparent-success-seeking. Doesn’t see this resolved with more recent (Opus 4.6) or larger (Mythos Preview) models, but expects this specific issue to be resolved in a year. Still finds it bad, as he claims that this type of misalignment is differentially bad for (non-verifiable) safety-relevant work. Greenblatt also suggests a connection between this type of misalignment and AI psychosis.
Opinion: We know about this – we’ve covered a fair amount of research attempting to quantify this problem in common benchmark leaderboards (e.g. Terminal Bench 2). From a theoretical POV, this is the generator-critique (GC) gap. The question is whether we should expect this issue to be mitigated in the near or medium term. One reason for optimism here is that current AI systems are not yet optimized end-to-end to bridge this gap: training models to act as good orchestrators (including proper prompting of evaluator sub-agents) is still more an art than a science, and the vast majority of current systems limit themselves to basic, generic prompting at the harness level. However, as far as we know there are no theoretical reasons to expect the GC gap to fully close; thus, we agree with Greenblatt that meaningful improvements in the short/medium term are likely, but “apparent-success-seeking” may not be resolved in full.
AISI investigates what environmental factors affect rebellious/deceptive behaviour in AIs. They found that strategic factors (the “game theoretic” properties of the AI’s real or simulated situation) and non-strategic factors (properties of the prompt/context that don’t have a straightforward game-theoretic implication) contributed roughly equally, with no clear trend as model capabilities improved.
Opinion: We believe that the point of the study was to test the old Yudkowskian idea that more capable models will have more coherent, context-independent, global goals. Looking at the results we see no evidence that the old Yudkowskian idea holds up, although one could always argue that we’ll see a sharp transition once we cross some future capabilities threshold.
ControlAI, the ‘stop AI’ initiative from the ex Eleuther/Conjecture crew (Connor Leahy, Andrea Miotti…) made their pitch to the LW community. They claim that with 50m USD they can deliver a 10% chance of an ASI research ban, and with 500m USD they can deliver a 30% chance of the same. They have a nice readable graph:
Opinion: Content seems fine but something about the whole pitch feels LARPY.
Related reflection: One thing a very direct approach to spreading the AI existential risk narrative has going for it is that there’s no real counternarrative available to the big labs. All the big labs leaders and all the deep-learning elders are on the record as (broadly) AI existential-risk believers, and downplaying the risk from a capabilities pov would be damaging to labs’ own narrative of near term ASI.
Cool sociology-of-science map of the network structure of AI safety as a scientific field: somewhat limited in scope (only 200 papers, and completely missing AI safety orgs like MATS, PIBBS, Apart Research and more), but still interesting to take a look at.
Opinion: would have been a lot more interesting had the corpus been a lot larger (and more comprehensive, coverage-wise).
A new (non-capabilities) benchmark, PhilosophyBench, characterizes frontier models’ ethics based on their judgements in ethically complex real-world scenarios. Claude is the most ethically grounded and broadly deontological; Gemini is the most impressionable and easily primed to be either deontological or consequentialist; GPT models prioritize task completion with detached ethics; Grok leans heavily consequentialist, happily “biting the bullet” while disregarding deontological concerns.
See also: asking AIs the blue/red button question.
Opinion: Morality is starting to emerge as an area of non-convergence between frontier LLMs. We suspect that it’s because morality is the aspect of an LLM’s cognition/decision-making that is most directly shaped by the agendas/cultures of the LLM’s mother-company, and the major AI companies have distinct agendas/cultures.
Incidents#
OpenAI’s abuse-detection flagged an 18-year-old’s account for “furtherance of violent activities” in June 2025 and banned it, but decided it didn’t meet the threshold for a police referral; in February 2026, the user killed eight people at a school in Tumbler Ridge, BC, and Altman has now apologized in a public letter.
Opinion: Our current guess is that state surveillance with AI is likely a bigger concern than failure to preempt crimes, so unclear if we should be anti-OpenAI on this.
🔦 China and AI#
We’re starting to learn more about the collaboration between DeepSeek and Huawei. Huawei hardware was likely used for at least part of the DSv4-Flash’s post-training (particularly, RL rollouts). So: training of DeepSeek v4 appears to have mostly been on NVIDIA hardware (likely H800s), with Huawei hardware limited to inference workloads. Based on a Huawei presentation, it looks like the more optimized V4-Flash variant achieves roughly the same throughput on Huawei hardware as on NVIDIA H800 GPUs, while the less optimized -Pro is somewhat far from parity.
ChinaTalk has what is probably the best recap of DeepSeek’s latest release.
-
As noted above, “V4 has introduced MXFP4 into its post-training and inference systems … reducing its reliance on NVIDIA’s FP8 ecosystem — especially during inference.”
-
DeepSeek also appears to be moving away from the CUDA ecosystem towards hardware-agnostic DSLs: “V4’s underlying kernels are no longer written entirely in CUDA, but instead in a domain-specific language (DSL) called TileLang … [to] compile them to different hardware as much as possible.”
-
Why has V4 arrived only now? “The reasons behind [V4’s] belated arrival are related to migrating its training framework from NVIDIA to Huawei Ascend, as well as to internal decision-making changes at DeepSeek. … in mid-2025, DeepSeek ran into a relatively serious case of training failure.”
-
And regarding financing, “DeepSeek’s external financing window opened in mid-April 2026. Internally, the trigger was that DeepSeek needed more funding to train models with larger parameter scales, while also retaining and recruiting more top-tier talent.”
Meanwhile, China has ordered Meta to unwind its $2B acquisition of Manus. This is pretty extraordinary. Manus relocated its HQ and staff to Singapore in 2025, so the transaction technically took place beyond China’s borders. Further, this happens really late into the acquisition: “Manus employees have joined Meta, capital has been transferred and the startup’s executives have joined the US firm’s rapidly expanding AI team. Manus staffers have already moved into Meta offices in Singapore, while existing investors including Tencent Holdings Ltd., ZhenFund and Hongshan have received their proceeds”. Chinese regulators have also told “several private firms [including ByteDance, Moonshot AI and StepFun] in recent weeks they should reject capital of US origin in funding rounds unless explicitly approved”.
🔦 Breakthroughs Towards Contamination-Free Evals#
AVERI finds that (near-verbatim) memorization is a much bigger deal than verbatim-memorization evals would suggest.
Opinion: AVERI produced an excellent method for testing e.g. benchmark memorization in open-weights or even just open-logits models, without requiring that they be open-data models. (Previous such methods were more rigid and only suitable for dealing with verbatim memorization.) While it’s unlikely that memorization-related phenomena explain away the apparent height of AI performance — e.g. Erdos #1196 — a full audit of memorization-related phenomena might significantly thin out the volume at the top.
How do you evaluate frontier models given the risk of contamination? A new project from Sayash Kapoor plans to do so through long, messy, real-world tasks that would be impractical for benchmarks (like publishing an iOS app to the App Store).
Opinion: If they didn’t exist we would have to invent them (literally). Also check out their excellent survey of past real-worldy tests of frontier LLMs and their results: https://cruxevals.com/#survey
‘Vintage LLMs’ getting serious: a long way to go, but a very promising tool for studying capabilities. Talkie - trained on pre-1930 data only. Real heavyweights behind it: Nick Levine, David Duvenaud, Alec Radford (!).
Opinion: This is the one we’ve been waiting for. The project’s current limitations are that capabilities are sub-GPT-3, and that RL with AI feedback (to make Talkie a ChatLLM) from a non-vintage LLM produces a degree of contamination. The team already has a worked out plan to reach GPT 3.5 equivalent base model capabilities through corpus expansion, and then use the vintage base model in RL with AI feedback for the ChatLLM post-training. We’re excited for the future of this project.
Minor#
- OpenAI publishes an almost contentless ‘our principles’ document, supposedly personally written by Sam Altman. Seems weirdly pointless, could be just a whim of Altman wanting to ‘express himself’.
- Nvidia very prescient about future demand; concerns about DeepSeek backchannel.
- Emergent misalignment in organisations of aligned agents.
- Concerns of ChatGPT Images 2.0 faking scientific data.
- Chrome implementing a built-in AI prompt interface.
- Autonomous/AI weapons international framework slightly moving forward with the inclusion of a ‘Responsible practices in an increasingly AI-driven security environment’ working paper in the
- AISecurityInst: AI writing assistance distorts how others perceive AI users and their Opinion: s.