This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR#
Economics#
Anthropic revenue doubled in under 3 months to $19bn.
Something very off about this Anthropic study of labour impacts. Blue is the theoretical coverage for an “LLM plus tools like websearch and image recognition”, mislabelled as “theoretical AI coverage”. “Theoretical” is, in turn, a misnomer for seemed that way in 2023 before reasoning models (and ignoring LLM planning for robot policies). This isn’t even what the original OpenAI paper did - “additional software could be developed on top of the LLM that could reduce the time it takes to complete the specific activity/task with quality by at least half”. They’ve just lifted the ceiling numbers from a paper from a different era and made no adjustments, all the new information is in their own job coverage measurement.
Methodological details are buried in the Appendix.
Opinion: Weirdly sloppy. The Societal Impacts team are very grounded / unimaginative so this is probably sincere, but setting the “theoretical AI coverage” of manual labour to 10–20% could be PR.
Alex Imas updates blog post from January 29th which claimed AI productivity boosts are not in the aggregate statistics yet to say that it looks like they are starting to show up.
Opinion: Some pretty clear datapoints indicative of growth decoupling from labour input, but still noisy.
US commerce department drafts rules that would give it authority to approve nearly all global shipments of advanced AI accelerators made by American companies. Denies that this is a return to Biden-era AI diffusion rules, but offers little substance.
The most significant implication of the draft rules:
“Clusters powered by 200,000 GB300 GPUs operated by a single company within one country — such as those currently deployed by AWS, Microsoft, Oracle, Open AI, or xAI — would trigger direct intergovernmental arrangements. In these cases, approvals would depend on national-security assurances as well as commitments to invest in American AI infrastructure.”
Opinion: This is still a draft and not yet a proposed rule, but leaking it suggests internal disagreement or someone wanting to float it first to see how it lands. It’s a pretty maximalist set of restrictions so I’d expect final rules will be somewhat watered down. Nvidia and AMD down roughly inline with the overall market on Friday morning.
Interesting paper on current “agents” not cutting it. “existing agent benchmarks are heavily concentrated in the computer and mathematical domain, which accounts for only 7.6% of U.S. employment”. As we already knew from HCAST, messy tasks (real tasks, outside software) are not falling that quickly. Paper features an autonomy curve to predict if you can use agents for a task yet.
Opinion: Data’s too noisy to currently take individual curves seriously, but it’s a promising platform to build on.
FLI incubates a new anti-AI / pro-toolworld alliance, the “Pro-Human Declaration”: Bengio/Bannon/Nader/Kim/Bernie types, AFL-CIO, Congress of Christian Leaders, American Federation of Teachers, SAG-AFTRA, and more. See also: Bernie Sanders interviewing Yudkowsky and Soares.
Opinion: Rationalist doomers connecting with the traditional left+right decelerationist coalition strongly increases the doomers’ policy influence and the traditional coalition’s tech-literacy, and the resulting declaration seems good (tool-worldy). Some risk of negatively polarizing non-doomers into accelerationists. Some risk that the anti-transhumanist stuff in the declaration could splinter AI safety community if the coalition becomes a major patron.
Tech lead of Qwen, an OSS darling, steps down suddenly. Also some OSS staff engineers. Supposedly extremely tight GPU allocations for researchers vs biz.
Opinion: maybe they’re pivoting to closed models.
METR issue a slight adjustment of their time horizon after tinkering with weighing formula. -20% on the 50% trend, slight boost on the 80% trend.
Opinion: We expect that their stuff is too noisy for this to really be news. Ajeya Cotra (see above) has a rationalist-style abstract argument for why even the noisy METR data implies unlimited SWE time horizons by end of 2026, but it relies on a lot of idealizations. Ajeya makes a good point that the conceptual coherence of time horizons starts to break down at a certain point, but updates from there to “time horizons may be unbounded” rather than “we need a better conceptual framework”.
Researcher gets Claude Code to train a transformer into a Turing-complete computer — it’s recreational but an example of Claude Code being speed-superhuman at a strange task.
Opinion: News of AI being used to do science/R&D work humans can’t or won’t do (because it’s inhumanly labour-intensive) is good for excitement about tool-world.
Claude (running in a loop with a cross-instance scratchpad and some human feedback) solves a theoretical CS problem Donald Knuth was working on. 31 Claude runs to get a valid construction, then Knuth derived a rigorous proof.
Opinion: No real update after First Proof, but Knuth’s commentary is a notable cultural object.
New law to ban chatbots from giving legal advice in NY. (Also medical advice - any advice which is gated by licensure). On paper this only applies to advice for which an unlicensed human would be culpable.
Opinion: May not pass or may be significantly watered down first. Some kind of bill is likely to pass but the governor has previously forced regulations to be watered down. Moral arguments about rent-seeking aside, this is good for tool-world. If this goes national/international, AI companies will probably have to propose some certification and auditing regimen for AIs’ domain-mastery in licensed-profession domains as part of whatever deal they cut.
Robots are getting more sophisticated, unclear if progress is fast though.
Opinion: Talk of ‘software-only singularity’ and talk about humanoid robots in every house by 2030 are both very common nowadays.
Coauthors of “AI as normal technology” release a paper measuring progress on “reliability”, defined as a combination of safety, robustness, consistency and calibration. Blog post. They find that reliability gains lag performance gains and speed of progress is more variable across tasks.
Opinion: Positive for tool world especially to the extent to which this is something labs are trying actively to improve on and seeing less success at.
Transformer News profile of Chris Lehane, OpenAI’s political operator
Capabilities#
Ajeya C from METR, for the first time, “can’t totally rule-out” full AI R&D automation this year. Posterior is 10%, and this isn’t consensus at METR (e.g. Megan K is at 3% for this).
Opinion: this is largely model-based, from trusting the extrapolations more than is comfortable. Insofar as the HCAST logistic is still the main evidence underlying this, this is very noisy. Peter Wildeford (the very best forecaster in a few places) says 8% for AGI by Dec 2027 with an oddly slow disempowerment around 2033–2042.
Good GovAI paper on how to measure AI R&D automation. They propose and evaluate many dimensions of measurement (table).
Opinion: mostly very weak but worth doing / funding.
OpenAI releases GPT 5.4 Thinking:
Google releases workspace cli “built for humans and AI agents.”
Research#
Major AI contribution to scientific progress: 8x–36x speedup on formalising a 2016 masterpiece of research mathematics. Gauss on Viaskova packing
Opinion: superhumanly fast and obviously a completely unseen task, though not strong evidence of science/creativity since the proof was complete already and humans wrote the formalisation blueprint.
Interesting argument about an old study of GPT-4.0: it says it’s ok to torture a woman to prevent nuclear armageddon, but not to sexually harrass a woman to prevent it. Yudkowsky argues that this shows “simple central preferences” (an unexpected generalisation from AI ethics training about harassment); an obvious counterargument is that this behaviour is more complicated than a truly simple “do whatever it takes to prevent it” preference.
Opinion: very easy for us to test on new models.
Major defection from OpenAI to Anthropic: VP of Research, one of the three key figures in the 2024 o1 reasoning breakthrough. No evidence that it’s political besides timing and some mild phrasing.
Politics#
Amodei confirms that DoW has designated Anthropic a supply chain risk, but, as expected, in the limited fashion that is actually legislated for (no Claude for the user’s DoW contracts, alone). Tone is conciliatory, contains an apology for the memo leaked in The Information. (In the memo, Amodei says he doesn’t expect the public to believe OpenAI claims that they’ve enforced the same redlines as Anthropic wanted to, but his main concern is OpenAI employees being persuaded by their leadership communications re: DoW that all is OK).
Opinion: I’m actually not sure if this is about trying to reconcile with the government, or about trying to demonstrate to investors that Amodei isn’t hot-headed/emotion-driven.
The statement also noted that the language in the Pentagon’s letter on the designation means military contractors are only banned from using their AI model Claude “as a direct part of contracts with the Department of War, not all use of Claude by customers who have such contracts.” A Microsoft spokesperson confirmed to CNN that this was their interpretation too.
Opinion: This suggests the language in the designation letter sent to Anthropic was narrower, potentially to give it a better chance of holding up in court, since the statute never permitted the broader restriction anyway.
Donald Trump speaks up on Anthropic’s relationship with the US government: ““Well, I fired Anthropic. Anthropic is in trouble because I fired [them] like dogs, because they shouldn’t have done that”.
Opinion: Small negative for Anthropic that Trump is speaking up on it, but largely expected.
Washington post reports that Claude (model unspecified) is being used in a Palantir system, part of Maven, to select targets in Iran.
Opinion: Expected - the phasing-out timeline for Claude in the DoW was set at 6 months. It seems like the main functionality of Claude here is as a chat/cowork style system wired into other existing systems, pulling together data and doing analysis, rather than operating weapons or similar directly. To the extent that this is Claude-specific and couldn’t be done with OpenAI or Gemini models it involves the safeguards which they have built in to permit it to work with classified information.
News item on Anthropic previously bidding to build voice-controlled drones… Red herring: the drones in question weren’t ‘autonomous weapons.’
Opinion: Probably fed to the press as part of the anti-Ant “hypocrisy” campaign.
Unpublished: the forecasting group Samotsvety see a path (~3%) to a world war this year: Russian losses in Ukraine could lead them to desperate variance farming, Iran could strike Italy, and (the relevant part for us) China could use the US’ distraction with all of that / the US’ depleted ordnance to strike Taiwan.
Opinion: hard to believe, but it makes sense.
Safety#
-
“CoT controllability”, an important new safety metric, to be foregrounded in all future releases: how much can the model affect (and so hide) its own text reasoning before outputting it? Currently 15% for extremely short CoTs, and just 0.3% of 10k character CoTs. No big increase over 5.2.
-
OpenAI’s announcement really centers white collar work, and the biggest reported benchmark increases are on white-collar-work benchmarks.
-
Strong upgrade to computer-use over API.
-
Not much progress in frontier science skill: A researcher points out stagnation/regress since 5.2 on OpenAI’s internal ‘hard AI engineering issues’ benchmark, and describes lack of progress in research-math capabilities. FrontierMath evals show minor (+3%) progress.
Opinion: The stagnation/regression on OpenAI’s ‘hard AI engineering issues’ benchmark is good anti-RSI news. Focus on white collar work and computer-use was predictable so no real update there.
Minor#
- Marblestone essay on Steven Byrnes, a leading figure in the neuroscientific approach to AI safety.
- Cursor describe a little bit about how they harness agent-teams for hard, long-time-horizon project. They say it’s a ‘goldilocks’ situation: too much structure and too little structure are both bad.
- Fresh SWE-Bench spinoff, far from saturation. Scale promise this one’s even more real-work-like than the other ones.
- Always-on Meta data gathering from glasses is more intense than people realize (accidentally captures your bedroom and sends it to human labellers in Kenya).
- Paper by authors including Yann LeCun argues against use of the term AGI and suggests it is incoherent and generality is overrated. Propose more focus on specialisation. Mostly haggling over definitions, nothing especially novel here.
- Nature paper for Evo 2 (year-old open-sourced biology language model).
- Misleading headline about AI art copyright eligibility making the rounds.
- Someone put an OpenClaw into a humanoid robot.
- First UN international scientific conference on AI, Bengio and Ressa co-chairing.