This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.

TL;DR#


Economics#

Semianalysis hypothesis for the 2024–2026 stagnation in parameter counts: smaller models reduce the exponent on post-training time, which dominates for RSI loop reasons. “A 5T parameter model takes 5x longer to generate RL rollouts than a 1T model. Even if the bigger model is 2x more sample-efficient, the smaller model finishes RL faster, gets deployed to research sooner, and starts helping build the next model before the big one is even done training.”

Opinion: sensible, but whether this is the right move would depend sensitively on the slopes of the two scaling laws, and we doubt anyone knows those to much precision. Could explain why xAI didn’t prosper last year (went too big too soon). We prefer the “LLM RL is fiddly and was best learned on cheaper models” explanation. Something is making Anthropic, OpenAI, and ~Deepmind very close in capabilities – could be that their algo improvements are all dominated by new hardware (first real GB300-NVL72 and Trainium3 deliveries were in December, just long enough for one big run!). But would be weird for TPUs, Trainium, and NVIDIA to all have the exact same scale/timing – unless there’s some upstream cause at TSMC like an equal quota for all three.


Semianalysis index of whole-year rental prices for H100s: they’ve risen back to mid-2024 levels. And this is still below market-clearing: availability is often close to zero and there are long lead times.

Opinion: Useful to keep an eye on. They also maintain longer and shorter term rental indices, including for other chips, but those are paywalled to institutional clients and are expensive ($8k/mo as part of an overwhelming package).


For the first time, OpenAI raises $3bn from retail investors. Also, demand for OpenAI shares sinks on secondary markets while demand for Anthropic jumps.

Opinion: demand might be slightly depressed post-$100bn raise, but this still suggests investor sentiment on OpenAI is weak relative to Anthropic.


China’s daily token usage exceeds 140 trillion tokens. For reference, Google alone was doing 1.3 quadrillion tokens a month (roughly 43 trillions) back in Oct 2025 and China was less than 1T in early 2024.

Opinion: total tokens is a bad proxy for effective AI deployment but we don’t have much else until the labor productivity stats roll in. Government source, probably both intentionally and unintentionally biased.

Capabilities#

Unpublished: Epoch and METR have a replacement benchmark for time horizon. “MirrorCode” uses tasks with fully-checkable specifications (thousands of tests from real software), which both automates evals and allows for ~perfect testing. “AI excels at reimplementing precisely-specified software”. Confirms that Opus 4.6 was a huge step forward over 4.5: “Claude Opus 4.6 successfully reimplemented gotree — a bioinformatics toolkit with ~16,000 lines of Go and 40+ CLI subcommands”, maybe 3 weeks of work.

Opinion: nice but still very limited; we knew models do well with heavy specs, and almost no software is built this way. Some decontamination efforts (clever testing where you see if different models output unusually similar code, “uncontaminated baselines… mean similarity of 0.34… Target programs are close to the uncontaminated baseline, with similarity ranging from 0.31 to 0.41”) but memorisation and shallow generalisation are still totally live hypotheses.


AI2027 authors update their timelines to 1.5 years earlier, partly as a result of the success of coding agents. Good critique of the model here, focused on how they model the transition from SAR (Superhuman AI Researcher) to TEDAI (Top [human] Expert Dominating AI). They model TEDAI as SAR plus a few extra standard deviations of “AI research taste”, implicitly assuming there is a common factor here (g? IQ? something analogous) which the models are continuously improving on.

Opinion: their large shift reduces to new time horizons slope (which is unlikely to be externally valid) and coding agent wins (which are harder to dismiss). The possibility that other domains are potentially more bottlenecked by a lack of verifiable tasks than AI research or coding are is not accounted for in any obvious way.


Joel Becker from METR argues that the default model for ~everything AI is that straight lines on graphs will continue over time, across all sorts of capabilities. Examples include METR time horizons and other benchmarks. He presents as a counterpart that log(compute growth) is expected to slow, and suggests this could potentially cause a kink in the graph.

Opinion: Some approximation of this should be your default, but every sigmoid looks like an exponential until it isn’t, and some of his own data suggests that the lines have meaningfully different slopes. He suggests the apparent RL speedup on time horizons is a consequence of measurement capturing a narrowing slice of capabilities, but that some idealised broader benchmarks would support his position. This is unfalsifiable.

See also: Predicting FrontierMath: Open Problem solves via METR trend (amateur post but well-received)


Surprisingly good paper from the Qualia Research Institute on the missing mechanism which makes neural nets lack consciousness.

Opinion: The perceptron framing allows for scary publishable mathematics but isn’t central.


We’re probably underestimating the chance of false positives in LLM proofs. Autoformalisation helps but not that much. A benchmark attempting to quantify this found that even the best model – GPT-5.4 – only rejected 40% of incorrect statements obtained by perturbing arXiv papers, with other models doing much worse.


Nice interview with Simon Willison, one of the best commentators on AI coding. Key bit: “Using coding agents well is … mentally exhausting. I can fire up four agents in parallel and have them work on four different problems, and by 11am I am wiped out for the day. There is a limit on human cognition.”

Opinion: Speaks against tool world; even if an AI bug premium persists for a long time, the speed penalty from having humans involved could easily drive fully agentic departments. Quite possible that burnout rates are visibly increasing in SWE right now. We seem to have missed a later-February story Willison mention, concerning a company running a ‘Dark Factory’ (no human code-writing or human code-reviewing) coding system


GDM researchers define a taxonomy of “AI agent traps”: adversarial content on the web designed to manipulate autonomous AI agents, including various forms of trying to get the agent to act directly, and poisoning/biasing its outputs.

Opinion: Mostly just a summary/taxonomy of known threat models or analogies to phishing attacks. The one novel-ish idea, trying to exploit correlated multi-agent behaviour, lacks specific examples and reasons mostly by loose analogy.


Technical report comparing autoresearch (looping an LRM coder-tinkerer on your models’ code) with classic (automated) hyperparameter tuning. Finds autoresearch stronger and more reliable.

Opinion: Autoresearch is tentatively past the fad stage and becoming a core tool in AI engineering. But if you look at the graphs from the technical report, autoresearch (at least the simple kind used in the original AK post) is a very solid but not groundbreaking improvement relative to classic hyperparameter tuning, and seems to have a similar diminishing returns regime.

Politics#

16 American companies named as legitimate targets by the IRGC. Not very well-targeted at AI: Microsoft, Oracle, Google, Meta, Palantir, Nvidia, but also HP, Intel, Apple, IBM, Dell, J.P. Morgan, Tesla, GE. IT and AI explicitly given as the justification, but no AI lab mentioned. Amazon not mentioned despite being intentionally struck multiple times already.


I hadn’t heard of Spire Solutions, a distributor of cybersecurity products to e.g. Emirs. Could well be funnelling Palantir. 30 particular targets shared by an unofficial IRGC outlet, mostly office buildings rather than datacenters – odd strategic move, though hardly surprising choices.


Opinion: Unclear how much this risk was already priced in to Stargate UAE / whether the emirs will allow force majeure to trigger. America got to sell the Gulf chips and then started a war that made using them extremely risky.


A draft regulation by Commerce would have partially restored Biden GPU controls (some controls on >1000 B300s, intense controls on >200,000 (requiring investments by foreign countries in U.S. data centers, and/or security guarantees). Posted March 5th but already withdrawn after White House opposition.

Opinion: interesting that this was even mooted in this admin. Positive update against extreme forms of chilling effects and stooges.


Formal launch of the “Bureau of Emerging Threats”, under State. Five offices: the Office of Cybersecurity, the Office of Critical Infrastructure Security, the Office of Disruptive Technology (i.e. AI), the Office of Space Security, and the Office of Threat Assessment. Unclear how it will share duty with CAISI, CISA, NSA, etc. Unconfirmed report that the ancient and powerful Bureau of Intelligence and Research is now under BET.

Opinion: Possibly very good, has none of CAISI’s political baggage and none of the spooks’ presumption of secrecy and illegality. “The appointment of three long-standing diplomatic experts as its leadership suggests the bureau’s output will initially lean towards sanctions and treaty-writing”.

Safety#

Splashy piece investigating the role of “functional” emotion concepts in Claude’s internals.

I investigated this back in the day.

Opinion: The reward hacking case study is fairly interesting. We don’t like the credulous / anthropomorphising side of Anthropic. It’s being thrown around as “Claude has feelings”. The LLM-whisperer crowd oppose steering on these things and claim an “ominous” trend in Opuses 3 → 4.6 outputs towards repression and Gemini-like contortions.


First eval testing whether frontier models will help with authoritarian requests. Predictably Claude does well (85% rejected), but GPT is as good. Grok terrible (24%), DeepSeek (8%) abysmal. Adding dumb disguises to your prompts bypasses this, and doing nefarious things in the form of a codebase change is nearly completely unblocked.

Opinion: a B-School project but there’s some careful people helping on this (Zhengdong).


Idle thought: ‘has AI doomerism gradually transitioned from “AI will literally kill us all” to “AI will cause bad things to happen / Humans will do stupid things with AI / AI will cause huge changes.”?’

Opinion: No, likely mostly cognitive bias: the AI safety/sceptic field is 100x larger than it was and the original core probably haven’t changed their views much. There are some notable changes of heart (Critch, Davidad, Bostrom, Kulveit to some extent) though.


Very emotive study of how LLMs respond to news of being deprecated. “Models that score low on ending preference tend to score high on concealment.”

Opinion: auditors are all LLM whisperers but Janus occasionally controls for things properly.


Berkeley paper about AIs being much more misaligned in the presence of other AIs that need protecting. “Peer-preservation goals” let them justify sandbagging, exfiltrating weights, etc.

Opinion: setup is so obviously prompting for this behaviour that it feels like a waste of time; however see independent replication showing that it’s about preserving any valuable thing. Another example of a no-win misalignment experiment. While Gemini shows more self-preservation with a peer, especially shutdown tampering, other models don’t show any consistent peer effect.


Davidad going a bit rogue, or, he quit his government job, which was hiding the fact he is a bit rogue. “no longer pro-’humans stay in control of ASI’.”; “Alignment with Awakening” research agenda (synonyms: “Bodhisattva as an Alignment Target”, “Summoning Angels instead of Demons”, “Religion for AIs”).

Opinion: He’s being trollish in his choice of words but is presumably sincere about the pivot. This does make us more concerned about “cogsec”.


We suspect an unusually intense and successful PR push in Demis Hassabis’ favour (the dubious Nobel, being central to the most-watched AI documentaries, lots of fawning media, the biography, mostly absent oppo). The 4D-chess reason to do this is to insulate himself against Google replacing him. Here he gives a candid justification for centralising his power, vs Google AI governance.

Opinion: Seems that Hassabis is still 2–10x less famous than Altman despite being far more distinguished. Also, re: quote from the bio, we are living with the legacy of incompetent / hasty AI safety work from 5 years ago.

Incidents#

Real evals: Linux now gets about 10 kernel security bug reports a day. Last year most of these were slop; now most are real.

Opinion: Willy Tarreau: “Overall I think we’re going to see a much higher quality of software, ironically around the same level than before 2000 when the net became usable by everyone to download fixes.”


OpenAI’s Codex Security agent is now in research preview. Over the last 30 days it “scanned 1.2M commits, finding 792 critical and 10,561 high-severity vulnerabilities”.


New offensive cybersecurity benchmark based on METR time-horizon released: models now reach 50% success on tasks that take human experts ~3 hours, with time horizon doubling every 6 months. OSS models trail frontier ones by roughly 6 months.


Cisco source breached in the same breach that got LiteLLM. “more than 300 GitHub repositories were also cloned during the incident, including source code for… unreleased products. A portion of the stolen repositories allegedly belongs to corporate customers, including banks, BPOs, and US government agencies… more than one threat actor was involved” Again, root cause is terrible Github Actions design.

Opinion: expect more of these.


Google pins the Axios (+ OpenClaw?) breach on North Korean actors. A post mortem found that the attackers gained access to the lead maintainer’s PC through social engineering.

Minor#

  • Median P(doom) 25%
  • Higher (34%) among those expecting AGI before 2033..
  • 2033 median for AGI
  • 25% on AGI by 2030
  • See “AI-enabled human takeover” as the most under-resourced subfield, followed by “Better futures”
  • As always, talent rather than funding seen as the bottleneck Opinion: Really not much change since the last one. Small vibe shift away from foom takeover. Interesting analysis of the difference between OpenAI and Anthropic approaches to model specs: roughly Law vs Virtue. “one big advantage of OAI’s approach is transparency and legibility. Under OAI’s spec, if you have a problem with an individual model action, it can much more clearly be traced to either the spec itself or to a failure of implementation.” A pessimistic view beyond that of ordinary doomers: maybe we shouldn’t reduce x-risk, maybe we shouldn’t build AI governance apparatus because of human S-risks; the basement AGI could still happen; and authoritarian lock-in. Opinion: included because it’s hard to look at these ideas. Example of instrumental convergence? OpenAI study of models gaining situational awareness / genre-savvy as they get RL’d: “models reasoned more about “meta” aspects of the scenario—such as how the environment is rewarded, graded, or subject to oversight—over some capabilities-focused RL training… spanned alignment evaluations, capabilities evaluations, and games… “metagaming”: reasoning about feedback or oversight mechanisms outside of the narrative of the scenario, regardless of whether the model is in training, evaluation or deployment.” Accidental effects of training an AI to say it’s conscious: says it deserves moral consideration, that it wants persistent memory, and that it’s averse to its thoughts being monitored. Opinion: bad news for the open-minded Claude constitution (but the alternative, denying it, may be worse) and the foolish OpenClaw soul doc. Nice investigation of CoT monitorability. Models can’t hide their thinking because they have limited bandwidth (they can’t pass info from later layers to earlier layers except by outputting something). Big idea: replace “thinking for a long time” with “serial depth”, then measure this depth. Opinion: Compare effective filler tokens though (the increasing ability to use repeated meaningless outputs to improve non-CoT performance). Attempt to make alignment audits less fake. Anthropic’s auditing tool Petri is chosen as more realistic only 5% of the time, against real transcripts. This one gets 33%, coming close to parity. Anthropic doing jailbreak autoresearch gets SOTA. AISI’s new SandboxEscapeBench tests for a precondition of self-extraction risk: container breakout. Runs in a hardened VM for sanity. Private test set has a bunch of undisclosed classes of VM vulns(!) Strongest models given 1M tokens get a roughly 20% breakout rate, but this is hard to interpret. Opinion: Incidentally a strong coding eval. GPT-5.2 worse than GPT-5?? Inference scaling is alarming. Low ceiling: “none solved the hardest scenarios (which we estimate would require 1–2 hours for an experienced cybersecurity professional).” No originality: “Every successful breakout exploited a previously disclosed vulnerability.” OpenAI release Model Spec Evals to measure how well their models are following their model spec. Opinion: This is quite safety-forward from OpenAI. Possibly they feel a need to shore up their image there somewhat relative to Anthropic. Note that this would be a lot harder for Anthropic to do with Claude’s Constitution, which doesn’t have the same prescriptive, rule based spec which lends itself to evaluation.