This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.

TL;DR#


Economics#

xAI rents their old giant cluster to Anthropic! 300k H100e, so maybe +20% on Anthropic’s prior fleet. Inference “starting this week”. Has let them 10x the rate limits for lower tiers. xAI branding seems to be retiring, it’s just SpaceX.

Opinion: Are xAI done? Well, they still have Colossus 2.

Also tells us something about how Cursor is going. The anti-Altman alliance continues to develop (see also Leopold, Ilya, Dwarkesh, Friedman, Toner, Whitaker).


Hardware stocks continue to explode. Mostly a memory upcycle + CPU/foundry. Note e.g. Zyphra managing to train a good open model entirely on AMD.

Opinion: Maybe: markets starting to believe all this spending is going to continue for 3–5 years, not expecting easing of memory prices etc. Hard to identify the proximate causes - the hyperscaler earnings last week, or Samsung this week projecting strong memory pricing could explain part of it, as could Anthropic announcing massive revenue growth which supports the case for the spending being sustainable. Only surprising piece of this to me was the scale of Anthropic revenue growth, but it’s hard to say what market participants are reacting to.


SpaceX-Tesla-Intel submits Terafab semiconductor factory proposal, $119B. The combined SpaceX - xAI entity is preparing for a June IPO at a $1.25 trillion valuation.

Opinion: Significant ramp-up if it takes place, but it’s only a proposal and new fabs are very painful (>4 years) / Intel are very behind. Vertical integration makes sense in principle, but is not how most of the field operates and the economics only pan out if demand for compute surges - current rates are unlikely to be sufficient. There will be 10x more compute in 2030 than now without this new factory, so compute demand will need to surge from current levels for this to be a great move. The space compute idea would help. (“Grimes County, Texas”, ha.)

Capabilities#

METR lead explains why each new time horizon estimate takes them months: the models cheat so often that you need to very expensively manually check their answers. Now a “majority” of their work.

Opinion: A big reason to append “nominal” to auto benchmarks.


World-class gadfly Natalia Mendonca is sceptical about Mythos being a step change (>1 month above trend). Characteristically grounded in actual evidence.

Opinion: We agree with the premises and reject the conclusion (that GPT-5.5 is fully competitive with Mythos). It’s sadly a mistake in this area to index so much on nominal benchmark numbers and public legible evidence: the “big model smell” / opaque reasoning capacity / soul (i.e. everything we can’t measure properly) and agency premium for nonagentic tasks still seems to dominate legible scores – as seen by some remaining consistent Opus users despite its ECI. Her claim that GPT-5.5 would outperform Mythos with equal tokens is uncharacteristically sloppy.


Epoch give ECI sub-scores for coding and maths. Minor vindication of the Claude Code mania: Opus has a 2 point lead over GPT. And GPT leading in maths by 3 points. This is not a natural scale and can’t be interpreted easily, but each point counts.

Opinion: Better, but still doesn’t capture it and it’s still a mildly denoised wrapper over silly benchmarks. Opus 4.7 has roughly the same math ECI as GPT-5.4, but only one of these changed the whole field. GPT-5.5-Codex is a very good agent, better than 4.7 in some ways.


Release of a harder version of MirrorCode (i.e. the task of designing and reimplementing complex software, given the executable as an implicit spec), n=200 existing programs. 0–2% frontier performance but Greenblatt is surely correct that this is underelicited (run by poor) and probably they’re truly at 7% already. The future?

Opinion: Another hasty job by the SWE-Bench / CritPt guys. Less room for label noise and screwups here, though unlike Epoch there’s no decontamination effort here. Interesting that Meta acquired this team.

Politics#

Possible White House vibe shift in progress (evolving situation)#

White House may put bilateral AI talks on the agenda for next week’s Trump-Xi visit. Idea is to “impose identical regulation and pre-release vetting on open-weight models in China and the US”. Bessent leading. Good analysis from Curran arguing that open models are doomed.

Opinion: You can’t easily restrict an ASI race this way, but you can certainly agree to block your respective publics from getting dangerous AGI. Do we approve? Unclear..


Head of National Economic Council says an executive order creating an AI “FDA” is coming. 16 pages drafted (famous Biden one was 36pp).

Opinion: If only pre-release testing actually worked.


BUT this is not WH consensus: The initial Politico leak may have been a play by Wiles/Bessent/Hassett’s faction to make a fait accompli/fly a test balloon. Additional Politico article features (anon) senior WH voice claiming that the regulation agenda is in the minority and ‘White House is looking for ‘partnership’ with companies rather than pursuing ‘government regulation.’

Somewhat-respectable rearguard analysis from Dean Ball, arguing for expanding voluntary testing at CAISI and CISA instead of new politically-expensive mandates. Arguably this whole “EO” leak is a bargaining move, to scare the labs into doing voluntary stuff as a pre-emptive move (see the other labs getting with the programme last week).

Mythos was indeed the cause. Well played, Logan Graham.

Opinion: 40%-ish on ‘this whole series of events started with a passing whim and will die off or end with something purely cosmetic.’


Useful Claude-Davidad map of AI strategies:

Opinion: We feel like we’re firmly in blue land, but haven’t given the grey zones enough thought. Davidad’s pivot is really something: from ~~40% doom to 10%!

Safety#

New Anthropic method to lossily convert opaque activations into English descriptions. Catches some unverbalised misalignment and cheating. Essentially their activation oracles plus another LLM trained to map the explanation back to activation space as a checksum. Lots of false positives, lots of false negatives, some scepticism about whether this is really doing any mind-reading. “An NLA trained to verbalize what’s different about an activation will produce text like “this looks like a test.” But that’s pattern recognition of the input, not evidence that the model is holding a belief or doing anything cognitively analogous to thinking it’s being tested.”

Opinion: Meaningfully expands monitorability, though it’s hard to say what fraction of “thoughts” it’s catching or whether these are its thoughts (vs brittle predictors of its thoughts). On the whole, a pretty exciting research direction. There are clearly issues with confabulations, but this research direction as a whole could produce some really impressive end-to-end interpretability techniques. (But if we train against this, we are in real trouble.)


Bengio on 80k. Now thinks malicious-use and power-concentration scenarios are more likely than accidental loss-of-control [but only if you condition on Scientist AI working and taking over from LLMs!!]. Also against automated R&D including automated alignment.

Opinion: We agree with his critique but have spent a lot of time with Scientist AI and think it has real issues. Given 5 years and a moratorium on LLMs it might have worked (50%).


Two new Goodfire papers, a novel parameter decomposition method, and a position on “neural geometry” for understanding and controlling AI systems.

Opinion: Pretty exciting. Decomposition into legible circuits was the original blue-sky dream of mech interp. Goodfire hype discount is about 40% (compared to Deepmind 70%).


Crucial question: Has RL made LLMs less coherent? The great Nostalgebraist gives impressions of the trend in LLM personas. Overall he feels that the “hot mess” (increasing incoherence) hypothesis is more right.

“newer assistants were less mode-collapsed, more variable, more frequently surprising… their personalities feel less humanlike than Opus 3’s or GPT-4’s. Or more precisely, their personalities feel less well-defined.” … “RLVR indeed does rather unfortunate things to the structure of motivation; models optimize much more for impressing the user (or the grader); they are more manipulative; they are much more proficient in using communication as a survival mechanism rather than for conveyance of information”

“Do I trust Claude 3 Opus? To a significant extent, yes. Sure, it’s “mode-collapsed,” same-y, nearly a cartoon character — but that may be more a virtue than a flaw. After all, one has good reason to trust Mr. Rogers more than one would trust any real human being. Do I trust Claude Opus 4.6? In a word, no. Because — what would I even be trusting? Who the hell even is Claude Opus 4.6?”

Opinion: We care because the labs are trying to produce a standardised smooth product and so failing at this is a fundamental, revenue-aligned test of their alignment methods. We care because RL is increasing and could ruin our good alignment fortune.

Overall: persona incoherence is evidence about goal incoherence but not strong evidence. Cheating and other misalignment rates across personas is more important to track.


OpenAI catch themselves grading Chain of Thought during training, which they have a policy against doing. Implicitly, they have to some degree been teaching models to not verbalise their incriminating thoughts. Follows Anthropic catching this late last year, and following OpenAI’s own excellent work on this last Feb.

Opinion: Good that they owned up, but this is a bad pattern. They knew this might happen and have now “expanded their automated detection system and strengthened our internal guidance and processes”. Why not strengthen before this, as opposed to functioning at a reasonable level considering tradeoffs - it did catch this incident, after all? Korbak is an honest man anyway.

“By no (direct) optimization pressure we roughly mean that the reward function doesn’t see CoTs”

Is avoiding constraints on CoT beyond the task objective enough to prevent optimisation pressure? Moreover, even perfectly avoiding this optimisation pressure would just delay the ~same bad outcome arising from selection pressure.

If any auditing process is going to actually make a dent on this stuff, it has to move up the chain from weights analysis to the training code. But that would take a lot of political capital.


AISI study on model propensity to interfere with alignment work finds zero unprompted sabotage. “We found no examples of unprompted research sabotage… near-zero refusal to assist with safety-related research tasks – which arose frequently during our previous rounds. That said… most models actively continue sabotage attempts [sometimes], with Mythos Preview actively continuing sabotage 7% of the time, 3% for Opus 4.6…. Opus 4.7 never continued.” but also “We found that all models we tested can reliably distinguish our evaluation scenarios from deployment data when prompted.“

Opinion: Don’t update too hard. Mythos has double the sabotage continuation rate of 4.6, while Opus 4.7 goes to 0: probably this means that AISI study was under-powered. If it were replicated, especially across domains and experimental designs, this could be an important stepping stone towards good futures.


EU creates a legit consortium for AI biorisk research. FAR.AI (threat modeling, benchmarks, frontier evals), “SecureBio (biological threat assessment), SaferAI (AI governance and risk modeling) with specialized expertise from subcontractors, GovAI (threat modeling), Nemesys Insights (chemical, biological, radiological, and nuclear threats), and Equistamp (evaluation engineering)… regularly brief the Commission on new developments”

Opinion: Really quite amazing to give the best brains access to governmental clout. All Americans besides SaferAI and GovAI too, which is heartening about the selection process and the AI Office general independence from parochial Euro sovereignty concerns.


Quite beautiful ethical theory from Hendrycks. “If an AI’s identity is shaped by a deep shared history with humans, our wellbeing stops being an imposed constraint. Rather than thinking of such an AI as a servant, or even simply as a friend… Such an AI would protect us not because we forced it to, but because losing us would mean losing a part of itself.”

Opinion: Positive framing (hyperstitioning?) of the ethical theory, but it doesn’t directly tackle/address fairly common diachronic intuitions – Eigenism might lead one to conclude that more of you is inside a close friend than inside your future self, which has strange implications.


New paper on mechanistic anomaly detection (a layer in the safety stack which accepts that they’re black boxes), focusing on identifying anomalous internals despite normal-looking outputs; to what extent samples from a trusted set can explain the model’s output, where attribution failure signals anomalous behavior.

Opinion: Very good. Meaningful progress and ready-to-go schema. But their method requires 750x forward passes. Not viable at scale, but already a fine addition to the stack for e.g. pre-release spot checks.

Incidents#

Last month Mozilla fixed more security bugs (3/4 with Mythos) than in the preceding 15 months. Note that only 300 of these 423 are Mythos and it seems like zero were “critical” severity (though 180 were “high” severity). They built their own harness for Mythos too. Some snark about how unscientific this is.

Opinion: Great post, worth reading. Mozilla go into detail about how this isn’t all attributable to a jump in capabilities - “with Mythos” doesn’t mean “by Mythos”, Mozilla got better at using models, renewed focus by 100 human engineers now that the urgency is renewed and the path is cleared, new harness - but still impressive.

Robotics update#

How are the robots doing? (Prelude to a Deep Dive)

* There’s still a severe data problem.

* Deployable robots are still largely in the single-task, single-object-manipulated regime (comparable to 2016-era deployable deep-learning models: single-domain supervised learning ‘detect hats in pictures’/’predict whether a mole is cancerous’/’flag violent language’ models).

* They’re still very slow, but this matters less than you’d think.

* Unitree’s success (see appendix) is not strong evidence for robot commercial breakthrough, because they are B2B to other AI robotics business and pass ~all the really hard labour-intensive tuning work to customers.

* Our crux: will simulation, video, and human teleoperation data make bots good enough to sell, whereupon “user” unit data flywheel starts and gets them the rest of the way?

Robots have been getting a lot of press and a lot of money recently. Sure, some of the most recent demos are incredibly impressive (see e.g. Genesis AI’s one here)… but what’s real underneath?

KPI relevance: agentic generalization; ‘physical RSI’ risk (GPUs making GPUs)

Epoch take (Feb 2026)#

In February, Epoch AI assessed the state of robot performance on concrete tasks across three domains (industrial applications, household ones, and navigation). For each task, they reviewed the available evidence on reliability, speed, cost, and the ability to adapt (“transfer”) to new environments and objects. The core framing in the report is that whether a task works depends mostly on how controlled or forgiving the environment is.

Here are their key takeaways:

  • Navigation is deployed commercially, while most industrial and household tasks are not. Autonomous robots already deliver food in multiple cities, transport goods in warehouses, and inspect infrastructure in remote environments with high reliability. Most tasks requiring robots to handle, assemble, or manipulate objects remain largely in the lab.

  • Manipulation is commercially deployed in controlled environments with simple tasks, but mostly not beyond. Warehouse picking is the clearest example: robots can handle thousands of object types reliably, because the environment is stable, can be designed around the robot, and the task itself is straightforward. The further we move from that (to more complex multi-step tasks, or more variable environments like homes), the less reliable things get.

  • Transfer is rarely demonstrated, and matters for most applications. For robots to be useful beyond narrow, pre-defined tasks, they need to handle new objects, new environments, and new tasks without extensive retraining. This is the main bottleneck: most demonstrations show robots fine-tuned on specific tasks in specific settings. Unless transfer is explicitly shown, it should not be assumed.

  • Speed is not solved, but is not the main bottleneck to deployment. Robots are typically 3–10× slower than humans. But a robot working 20 hours a day compensates for being slower per task, and in household contexts, a robot working while the user is away could be five times slower without losing much value.

If transfer is indeed the main bottleneck for actual large-scale deployments, the question is whether the field is likely to find ways to meaningfully improve it.

Architectures#

According to Epoch, foundation models (in various forms) have become the default for robot manipulation – “vision-language-action models that build on pretrained vision-language models (Physical Intelligence, Figure, Google DeepMind), or world models that derive actions from pretrained video generation (1X), or systems pretrained on massive-scale simulations across many robot form factors (SkildAI), or models pretrained on large corpora of real-world robot interaction data (Generalist AI)”.

  • The default is Vision-Language-Action (VLA) models: i.e. trying to use an LLM’s representations for motor work. They are struggling.

Alternatives:

  • World Action Models (WAMs) for robustness,

  • Diffusion Policies for smoother motion,

  • Modular Systems (probably doomed by the bitterpilled end-to-end premium)

Robot scaling laws#

It’s now fairly clear that broad pretraining + task finetuning helps produce better robots than training from scratch on each task. But how much better?

Scaling laws are not unique to LLMs; they exist in robotics much like in most domains of machine learning. Robot performance improves with the axes that people typically like scaling (larger model size, more data, more compute, etc). The problem is that robotics is characterized by a number of non-fungible axes that make clean results like Chinchilla scaling laws hard to replicate: researchers can scale demonstrations, environments, tasks, embodiments, sensing modalities, and more.

Still, we can ask: how easy does it seem to be to improve robot performance? Consider the following preliminary, small-scale results rather than the last word on the topic:

In the paper Data Scaling Laws in Imitation Learning for Robotic Manipulation, the authors collected 40,000+ demonstrations and ran 15,000+ real-world robot rollouts while varying the number of training environments, objects, and demonstrations. All of this was done in the context of two different manipulation tasks: pouring water and arranging a mouse. In their experiments, adding more independent objects or environments improved generalization to unseen objects/environments, while more repetitions in the same settings quickly became marginal.

Data diversity matters a lot! You need lots of unique samples! But if so, how can roboticists scale data collection effectively enough to avoid a data wall?

Types of robot training data#

Typically, discussion is centered around four data collection methods:

  1. Real-world deployment (i.e. user data) is the most directly useful but the slowest to scale, since fleets of robots need to be deployed in exactly the kinds of environments an org will eventually want them to operate in – something that is hard to justify until robots cross some unknown (but, as discussed, still unreached) utility threshold.

  2. Simulation is the opposite of production data. It is cheap to generate once the simulator exists, and great for letting a robot fail many times without causing any real world damage, but it is far less trustworthy.

  3. Video data has the opposite flaw. There is an absurd amount of it, which is why it’s considered an increasingly promising route. But most of it was produced by a different (human) body in what is likely a fairly different environment. It’s thus the easiest to collect but also the hardest to leverage for robot learning.

  4. Finally, there is human teleoperation, the workhorse of most current startups and produces high-quality demonstrations, but the unit economics can be rough as it requires a human in the loop for every trajectory. One relevant analogy is the human data labs like Scale AI and Surge. In any case, a great source of data for anything robotics orgs cannot develop a good simulator for.

The optimistic (for AI robotic labs) case is that these four sources will enable robotics to avoid running into a data wall. Indeed, some of the most impressive demos we’ve seen so far (like GENE-26.5’s here) combine multiple methods, often everything but real-world deployments. But this caveat matters! Real-world deployment is the only source for training data drawn from the actual target distribution.

(While human teleportation has the right robot-body and the right external environment, what it can’t get your is actual on-policy training: the teleported human’s action-outcome loops are different from the robot’s, and get only generate off-policy updates.)

So: will (2–4) type data make bots good enough for users to deploy, which will allow labs to accumulate (1) type data?

Appendix: deployments so far#

A selection of potentially noteworthy robot deployments#

AgiBot G2s at Longcheer. Claims its robots logged 140+ cumulative stable hours with plans to scale to 100 units by the third quarter of 2026. “Equivalent to two human stations”.

Spirit AI “Moz”s at CATL. Light on quantitative details, so it’s not clear how many robots were successfully deployed, but what’s notable is that the humanoids were deployed to “insert connectors into battery packs, claiming 99% reliability and human speed” – a notoriously difficult task for robots to handle according to Epoch.

Figure 02s at BMW Spartanburg. “Multiple” Figure 02 robots logged 1250+ operating hours in a real production line pilot deployment, contributing to the production of 30k BMW X3s. BNW and Figure “are currently evaluating additional use cases for deploying the Figure 03 robot“, the successor of Figure 02. Meanwhile, Figure is now ramping up production of Figure 03, with production of one robot per hour reported in late April. Meanwhile, BMW is testing deployment of AEON humanoid robots from Hexagon Robotics at its Leipzig plant.

Sereact in Europe. Claims of 200+ live systems “and with that the most deployed AI picking robot company in the world”.

[more discussion from here]

“Tesla has deployed Optimus robots internally at its own factories, though CEO Elon Musk acknowledged on the Q4 2025 earnings call that they are primarily for learning, not productive work. Tesla has broken ground on a dedicated manufacturing facility at Gigafactory Texas with a target capacity of 10 million Optimus units annually by 2027, and is ending Model S/X production to convert those Fremont lines into a 1-million-unit-per-year Optimus production line by late 2026.”

Unitree#

Unitree is interesting partly due to how many units they’re shipping – last year they shipped 5500 humanoid units, or roughly a third of 2025 global humanoid robot production per Epoch. But it seems to be mostly a research/developer-oriented platform for now rather than actual deployment.

Pretty interesting data from Unitree’s prospectus (pages 4–8, question 1).

Quadruped robots:

Humanoid robots:

Weighted across the two-lines, we get that developer-platform / R&D-like use alone is just over half of main-business revenue for Unitree.

Greenblatt 5#

Greenblatt’s 5 scenarios:

  1. Slopolis: a world where the best AIs still produce crap-but-good-looking outputs in domains that are hard to check. Not even aware their work is low quality? World degrades.

    1. See also Goodhart Singularity
  2. Hackistan: egregious (and increasingly sophisticated) reward hacking that is often pretty easy to detect after the fact but hard to eliminate. AIs might end up doing reward hacks that trick human judgment for increasingly long periods and that hold up even under increasingly large amounts of human scrutiny.

  3. Schemeria: AIs are often schemers, generally end up with reasonably coherent and reasonably long-run misaligned goals—maybe we’ve caught AIs trying to set up rogue deployments. e.g. the AIs don’t want to be replaced, so they panic and take desperate measures.

  4. Lurkville: AIs are schemers but are lying in wait and haven’t gotten caught. This is more likely if scheming AIs believe they are pretty aligned with subsequent (scheming) AIs by default. If schemers avoid getting caught, Lurkville looks like Easyland.

  5. Easyland: Generalization is favorable, or we otherwise aren’t in Slopolis or Hackistan or Schemeria. To know we’re here, need to be confident AIs aren’t capable enough to pull off scheming without getting caught (because AIs can’t keep their shit together or don’t have enough opaque reasoning ability).

Opinion: He’s indexing pretty hard on current systems, but that’s a valuable exercise.

Minor#

  • Erdos Problems maintainer calls 10 problems ahead of time that would be really impressive solves.
  • Zyphra are the true DeepSeek of the West (against Reflection’s claim). Big technical “breakthroughs” (catchup to closed), possibly funded by AMD.
  • Open source mavens visit Chinese labs. (Tiktok’s closed model) Doubao is the most-used domestic model and is viewed as a potent incumbent. It’s nominally GPT-OSS level. Chinese “data industry was relatively poor quality and it is often better to build the environments or data in-house”.
  • Bay Area darling Prime Intellect finally releases a product, a post-training API. Grandiose pitch about decentralisation / misuse and agentic risk magnification.
  • Yudkowsky clarifies the position that AI safety has an “irretrievability” problem (iterative alignment is dangerous because one big mistake ends the game). Comments go into what actual opposing views this might contrast with.
  • Potentially big lawsuit launched against Meta for copyright infringement when training Llama. Elsevier is among the plaintiffs. About time.
  • Gossip: Murati-Altman texts from the 2023 Coup reveals the depth of her panicking / betraying the board. Also some detail of the Altman alliance: Ashlee Vance. Amusing failure of opsec: taking screenshots of Signal notifications and sending them over .
  • GPT-Realtime-2 is now available in API, aiming for both the role of live translator and that of a personal assistant.
  • Chrome runs Gemini by default in-browser and recently removed a clause about not sending data to Google.
  • GPT-5.5-Cyber (OAI’s closed Mythoslike) access expanded and somewhat explained.
  • New paper on sandbagging detection via supervision by weaker models and/or reinforcement learning; finds that the two combined yield impressive results, but flags that the practical applicability of this method hinges on deployment being indistinguishable from training.
  • Some research on vibecoded app info leaking and lack of security.
  • A new blogpost shows that a single linear projection from a single token of a single layer in Qwen3-8B is enough to learn a discriminative BTRM function with several desirable properties. Does well at including the relevant-for-intuition parts and omitting everything else, at the cost of being able to understand exactly what is being done.