This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.

TL;DR#

  • Taalas hard-wires models for “10x” efficiency. Doesn’t work for MoEs.

  • Claude Code peak session length doubles

  • Altman floats an IAEA for AI

  • AISI shares a general method for breaking safeguards


Economics#

Taalas hard-wired models are 10x faster + 10x less energy at inference.

Opinion: this seems tricky to adapt to mixture of experts (i.e. any frontier model), although maybe one could hard-wire every expert model somehow.


So far Anthropic’s revenue has been growing faster than OpenAI’s.

Opinion: Anthropic might overtake OpenAI, although on the other hand this might depend on compute capacity and growing as the underdog is easier. If the Pentagon strongly decouples from Anthropic its revenue could also stop growing as fast.

Capabilities#

The 99.9th percentile of Claude Code use went from 25 mins to 45 mins in 3 months.

Opinion: shows that these models are indeed being pushed further as time goes on.


METR estimates that Claude Opus 4.6 has a 50% time horizon of 14.5 hours on software tasks, their highest estimate to date. Some argued that this is evidence of superexponential growth in the length of tasks that AI models can complete, though others noted that this was less clear from the 80% time horizon results and METR themselves said that their measurement is extremely noisy because their current task suite is nearly saturated.


Release of Gemini 3.1 Pro, also Chinese models (GLM-5, Seed 2.0, Qwen3.5).

Opinion: Just AI progress inexorably continuing.


Chinese models aren’t as close to Western ones as might seem, because the benchmarks they mention in their announcements are specifically selected for being close to parity. But see: Epoch’s Capability Index, an aggregate of benchmarks, for an aggregate of benchmarks that still has Kimi K2.5 up pretty high.

Opinion: Relevance: a bit of a wash.


A team of physicists worked on a paper that became so equations-heavy they put it on pause, and this month collaborating with OpenAI allowed them to simplify the equations and finish the paper.

Opinion: Significant in a different way from First Proof. GPT (5.2 Pro + internal) was used more like a futuristic version of Mathematica than like a scientist.

Politics#

At the New Delhi AI summit, Sam Altman says an IAEA for AI may be needed.

Opinion: unclear how serious Altman is, probably not much at all given that OAI consistently opposes regulation.


Pentagon briefly added Alibaba, Baidu, BYD to Section 1260H list, then withdrew them.

Opinion: Impact: shows a policy lever the US could use.


Defense Secretary Pete Hegseth is “close” to cutting business ties with Anthropic and designating the AI company a “supply chain risk”.

Opinion: Impact: if this goes through, it reduces Anthropic revenue and integration into the US military, probably a negative vs other model providers. Depending on the degree to which the Pentagon requires companies it does business with to decouple from all business with Anthropic, it could hit Anthropic quite hard.


U.S ‘totally’ rejects global AI governance, says Michael Kratsios, Assistant to the President and Director of the White House Office of Science and Technology Policy.

Opinion: represents the direction of current whitehouse policy.


Recently, a Super PAC backed by Anthropic has backed Bores, while a Super PAC backed by OpenAI continues to oppose him.

Opinion: Just interesting to note for, might be the beginning of a lobbying race.

Safety#

AI security institute shares a general approach to construct a signal to train against when breaking safeguards of top AI models.

Opinion: clever and general approach.