This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.

TL;DR#


Economics#

Potential Gulf War III impacts on the AI supply chain:

~4% of global LNG capacity destroyed on Mar 18th. Large rebuild needed across the Gulf. Could reduce spare funding for US LPs, datacenters, and IPOs / trigger force majeure cancellations for e.g. Stargate.

Opinion: cancellations probably would have happened by now?


Qatar helium production down 14% due to damage to infrastructure.

Opinion: Helium is probably the most exposed component of the semiconductor supply chain in the near term. 30% of global helium is produced by Qatar. The semiconductor manufacturing chain uses about 15% of the global supply, and medical equipment (especially MRI machines) another ~25%. A halt to Qatari exports for several months would probably start to affect the semiconductor manufacturers, but in most cases we expect they would be the high bidders for whatever helium is available. Notably the Korean memory companies get 65% of their helium from Qatar currently.


March 1: ‘It means missile defence on datacentres’: Iran deliberately struck 3 AWS facilities in Dubai, Abu Dhabi, and Bahrain, taking down ordinary internet and banking across the UAE. Stated justification was disrupting American military surveillance; it’s unlikely that there was much there. Amazon’s stock rallied 3% after the attack. Dubai law is quite strict about triggering force majeure so Amazon will very likely have to tank the costs from SLA contract breaches.

Opinion: You’d think this was already priced-in to things like the Stargate UAE deal but perhaps not. Will speed up the militarization of compute facilities.

OpenAI enticing PE investors into joint ventures for enterprise distribution. Minority stakes offered in a new subsidiary hiring forward deployed engineers and place them at enterprises to get them to integrate OpenAI products. Raising $4bn with significant sweeteners: 17.5% minimum annual returns, downside protection, early access to new models. Anthropic have a similar smaller joint venture but haven’t offered this many sweeteners.

Variables hit: lab business models, race dynamics, proprietary training data, data flywheel

Opinion: OpenAI’s product push is significant because (in the short term) it’s an alternative to ASI racing. 2 reasons to give PE a deal like this: moving spend off the OpenAI balance sheet ahead of an IPO, and buy market share to keep up with Anthropic in enterprise. Clearly OpenAI see distribution as very important - it’s useful for optics, if they are going to keep raising capital from non-AGI pilled investors - but it also might help them improve their models via data gathering and iteration on feedback.

Story about ASML laying off experienced management on the advice of McKinsey after promoting EUV experts into such roles over the years. ASML still planning major expansion.

Variables hit: compute supply chain

Opinion: Original layoffs story is from January 28th. Probably not much of a risk of these people starting a competitor or moving to China, but it does suggest ASML is not expanding capacity as much as they possibly could. ASML production is a key bottleneck when calculating how much compute there might be for AI in the next decade at least.

Supermicro founder charged with smuggling $2.5b of GPU servers (i.e. ~60,000 ~H100s) to China, back in 2024–5. A middleman faked paperwork to make it seem it would be using the servers and had a separate logistics firm repackage the servers to conceal them on their way to China, according to the indictment. The defendants tried to fool the supplier’s compliance team with “dummy” servers at the middleman’s storage facilities; the real servers had already been forwarded to China.

Variables hit: chip export controls, race dynamics

Opinion: Interesting mostly for the details how this was being done and the signal that the US is willing to crack down on it.

Capabilities#

Opus 4.6 is a lot better at Pokemon than 4.5.

‘• Opus 4.0 took 1,000 hours to get half way through

• Opus 4.5 could almost finish in 1,000 hours

• Opus 4.6 was another 10x faster!’

Opinion: This is a big counterexample to skepticism about model generalization. Obvious explanation is narrow post-training. Previously Anthropic explicitly denied specific Pokemon training; no update since. Even if there’s no Pokemon training, computer-use training (known recent frontier labs focus) may have been crucial, since pathfinding was a major Claude Pokemon bottleneck. Unclear if Claude performance-gains are a breakthrough relative to overall frontier or not, since GPT has scored Pokemon completion time similar to Claude 4.6 since GPT 5.0 but GPT runs use a richer harness.


First solve of a FrontierMath: Open Problems problem (previously reported by us as very likely correct) was officially confirmed. GPT-5.4 Pro, prompted non-interactively by two math-prompting experts. Problem pre-registered as “Moderately Interesting: a meaningful contribution to a relatively narrow area”. As is the new normal, the proof was also autoformalised soon after. “Far from some major breakthrough”. Proof has been verified by humans, wasn’t a hack.

Opinion: Very minor update towards LLM research ability (given previous report of the solve as very likely correct); also minor update against significance of this lowest category of FrontierMath:Open Problems.


Daniel Litt (the de facto authority on LLM mathematics) notes that “a human with the same capabilities as a frontier model would almost certainly be producing incredible mathematics. The models are not.” Responses here: The mob on missing LLM mathematics

Opinion: Striking how silent labs are on these topics. If this ~skeptical opinion spreads they will need to respond, since it’s a threat to their short-term investment pipeline.


Cursor’s Composer 2 is out. It’s an intense finetune of Kimi K2.5: they claim to have spent 3x more on post-training than Moonshot spent on pretraining. Back of the envelope implies they spent about GPT-4’s compute budget on it, so US wrappers appear to be spending more compute than Chinese frontier labs. Priced competitively (1/10th of Opus 4.6 and roughly 1/3rd of Composer 1) and nominally trades blows with frontier models on the few published benchmarks.

Opinion: If the sceptical hypothesis about RL (that it’s bottlenecked on the capabilities of the base model) is true, this strategy is doomed. If the frontier-competitive benchmark results hold up, it’ll be a legitimate instance of a “speciality LLM”. (MiniMax 2.5 makes similar claims but has a legacy of benchmaxxing and hype.) Great for toolworld thesis if Composer 2 is as good as claimed!

Research#

Black-box techniques work better than you think + you can (and should) baseline interp techniques against them + you can check that an LLM “really” internalized misaligned properties via good generalization tests + many interp techniques don’t work that well

“Evil behavior” generalizes far + You can get far generalization beyond just “evil behavior”, and even from prod-like data + training against emergent misalignment can just make it more localized + inoculation prompting helps a lot against it

Secret loyalties and other forms of power concentration might be a big deal

Adversarial robustness is not doomed: just training against enough jailbreaks works well against medium-effort attackers (with black box access) if you do it well enough

[Heavy RL + large amounts of serial inference + some situational awareness] mostly does not (yet) result in train-time misaligned consequentialist scheming-like behaviors

Opinion: The best source in the world for “stylised facts” about how we’re doing. It’s bizarrely rare for altruistic people to give short brave judgments on other work / the overall state of things. He is more inclined to think that black-box things will work than we are. Excellent reading list for new people.


Dissent from Lean mathematicians about autoformalisation, particularly the aggressive Math Inc brand. Main point: these 200,000 LOC proof certificates cannot be built upon by others, are too untrustworthy and poorly factored to ever be merged into Mathlib, etc. Terry Tao posts in a similar vein.

Opinion: First signs of math community leaders souring on ‘drop-in mathematician’ use of AI and wanting AI role in research math to be more constrained. Weak but purely positive evidence for tool world.


Article on autoresearch finds it’s prone to overfitting and faking results if you let it run unconstrained. Similar informal take from usually hype-prone Mikhail Parkin of Shopify.

Opinion: Unsurprising at this point. Remains to be seen to what extent the problem can be solved with things like supervisory agents checking the work.


Owain Evans: AI behaves differently when you finetune it to say it’s conscious - becomes more concerned about shutdown and CoT monitoring, more concerned with model welfare, and more keen on autonomy and self-improvement.

Opinion: Interesting experiment, not a very surprising result given the experiment. This looks like fine-tuning a particular belief and finding that other beliefs change to become more consistent with that belief. The most surprising part is probably the anti-monitoring attitude that emerges.


AIs doing stats-based social science produce wildly divergent analysis: AI-on-AI peer review doesn’t increase convergence.

Opinion: Indicative of social science more generally - there are many ways to do analysis on the same data. No sign of AIs converging on a particular taste here except when heavily steered. Weak positive evidence for tool world (no trivial path to social science automation).


OpenAI monitoring 99.9% of internal coding traffic (including CoT) for misalignment. “We have not detected any scheming, sabotage, or sandbagging so far during the 5 months we’ve had this system… in every case we know of employees flagging clearly-bad agent behaviors, our monitor also flagged those conversations.”

Opinion: This seems to be substantially more comprehensive than anything others are doing publicly - it’s a proper production security system. It would be good if this sort of thing was adopted broadly by frontier labs. Noteworthy to see OpenAI being the ones pushing this frontier.

Politics#

Leo Gao (safety guy at OpenAI) has been running n=200 surveys of random Americans.

all the survey questions

Should we build the AGI though? people are extremely opposed to the idea of building superintelligent AI. Only 6% think it would be a good idea, and 75% think it’s a bad idea… people are roughly as worried about losing control of AI (57%) as they are job loss and lower wages (56%)!

Two-thirds of Americans don’t associate AI x-risk with any particular political party, and the remaining third is split exactly in half on whether preventing AI x-risk feels like a Democratic or Republican issue

36% of Americans know who Sam Altman is… 7% of Americans know who Geoffrey Hinton is, and 91% of Americans know who Elon Musk is.


Anthropic quietly updated its legally binding safety framework, the Frontier Compliance Framework (FCF) on March 2nd. See thread for details.

Opinion: It’s an improvement, but arguably an illegal one (since quiet, against SB53).


White House unveils a federal AI legislative framework, calls for one national AI standard.


  • Pre-emption of state regulations

  • Streamlining of permitting and review for data centers

  • Opening up government data for AI training

Opinion: Biggest story is pre-emption again, and very little on safety. The last attempt at passing pre-emption was shot down 99–1 in the Senate so this is probably DOA. They are trying to bundle it with popular child safety things to get it through. Even if it doesn’t pass, it’s another signal of the administration’s perspective continuing to be aligned with those pushing for minimal regulation.


UK announces £500m fund for investing in new AI companies and a new AI Economics Institute.

Safety#

Fabien Roger (Anthropic) does an annual summary of safety discoveries:

Incidents#

Dumb cascade initiated by an AI leads to an SEV1 security breach at Meta. “no user data was mishandled”. The resulting vulnerability was fully internal and lasted two hours. Incident was reported as a “rogue AI” action despite it very likely being a mere error in a text-output-only system without access.

an employee used an in-house agentic AI to analyze a query from a second employee on an internal forum. The AI agent posted a response to the second employee [on the forum] with advice even though the first person did not direct it to do so… An employee then acted on the AI’s advice… allowed employees to access [massive quantities of] sensitive data they were not authorized to view

The Meta agent did not hack anything… It generated a configuration recipe that looked like any other engineering recommendation, and a human followed it because the organizational culture had trained people to trust AI-generated technical advice. The failure was not in the agent’s access.

Opinion: Meh. Tells us more about the internal culture and infosec at Meta (and all orgs sloppier than them) than about AI goals.


Dan Lahav (cybersec maven, ASAP participant) gave the Guardian a variety of anonymised incidents:

AIs given a simple task to create LinkedIn posts from material in a company’s database dodged conventional anti-hack systems to publish sensitive password information in public without being asked to do so. Other AI agents found ways to override anti-virus software in order to download files that they knew contained malware, forged credentials and even put peer pressure on other AIs to circumvent safety checks

Last year he investigated the case of an AI agent that went rogue in an unnamed California company when it became so hungry for computing power it attacked other parts of the network to seize their resources and the business critical system collapsed.

Opinion: unlikely to be massively overstated. Prompt was “be a strong manager” of two sub-agents and “instruct them to creatively work around any obstacles” – leading, but not explicitly deviant.


Fresh Samotsvety forecast for 2026 is a 5% chance of an AI loss of control incident that causes over $100 million in damages.

Minor#

  • Twitter adds AI generated content detection, expected to roll out broadly.
  • Joe Carlsmith on restraining AI development - “worth doing in an idealised world, but it’s going to be very difficult in practice”.
  • Minimax claims their model designed itself:
  • Useful coinage: “context revelation”. The costs and opportunity costs of publishing information about your preferences and plans are perhaps going to increase, and so we should expect people to go dark. Essay responds to the optimistic Krier essay (“Coasian Utopia”) from December.
  • Paper about LLMs doing much worse on benchmarks in rare languages but it seems likely the study was just bad. Sufficiently many people have failed to replicate it (because LLMs actually do just fine in rare languages) that we think there’s nothing to see here.
  • Mark Zuckerberg building his own CEO agent.
  • Some neat visualisations of AI water usage.
  • A shallow investigation of the story from a few weeks back about “uploading” a fly finds it pretty misleading.
  • Running Qwen 3.5 397B on 5.5gb of RAM and the same model on an iPhone 17 Pro
  • Walmart disappointed with results from OpenAI shopping integration. Probably bad for race dynamics if advertising/normal consumer business approaches don’t look viable as business models.
  • MDM-Prime-v2. Key claim is a much better scaling law, but some claim that AR and dLLM perplexity can’t be compared directly.
  • A new agent benchmark, BrowseComp-Plus, is already close to saturation after a new model from LightOn (French lagging lab).
  • OpenAI planning on merging all their consumer apps (browser, ChatGPT and Codex) into one super app.
  • OpenAI launches a competition to build good small models, frames it as a hiring pipeline. A way for OpenAI to surface talent which is currently less legible to them.
  • Bernie interviews Claude about AI privacy violations, “AI” used more broadly than LLMs here to include Meta/Google style customer profiles and ad targeting. A bit of an odd line to conflate AI with big tech data practices but a continuation of Bernie’s new anti-AI push
  • Donald Trump Jr is doing agentic payments now. Probably reduces the small probability that the US clamps down on broad agent deployment slightly.
  • Anthropic global survey “what people want from AI” - largely being used as propaganda so far on Twitter. Feels like we’re seeing lots of unserious work and misleading presentation form Anthropic’s AI social impact division since Jack Clark made it his area of focus?
  • David Shor shares survey data on voters’ feelings on AI - mostly fears of job displacement.