This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR#
Economics#
Anthropic discusses a Q4 2026 IPO, raising $60B. This only doubles their lifetime raise!
Opinion: mere rumour. But it makes sense given the extreme compute bottlenecks above and the victory over the DoW.
Epoch guess that OpenAI spend just 10% of their R&D compute on final, frontier model runs. Why, if they have ~no lead left and if the simple bitter lesson is true? A brilliant lowbie speculates: 1) the data wall is real (RL envs aren’t cutting it), 2) megaruns are still extremely painful, requiring 9x iteration cost, 3) they’re devoting large amounts of compute to research (e.g. learning run-specific optimisers) and autoresearch.
Opinion: Significant, since (naively) they could do an emergency massive scale-up to race. And this is a decent proxy for how intense/terminal the race is currently. Maybe Mythos and Spud will falsify this instantly?
Dec 2025 survey of 700 execs (mostly CFOs of small 100-person companies) on perceived AI productivity in their org: they think it’s +3%. Very credulous about improvements despite attributable revenue being three times smaller: “perceived improvements in labor productivity… substantially larger than the revenue-based productivity gains … likely reflecting delayed output realization and quality improvements that are not yet captured in measured revenues”
Opinion: +1% implied productivity is still about one year’s worth of pre-AI TFP progress. Excellent measure of expectations in small companies.
Very clear explainer of per-job automation risk. NB: “job J is 80% exposed to AI” does not mean it’s 80% automatable, nor that 80% of those jobs will be lost. “we should be less worried about consultants and more worried about truckers and warehouse workers than we currently are… o-ring job automation would be characterized by an extreme abruptness: “you will be more and more valuable - until suddenly, you’re not. Don’t get complacent.”
Opinion: Unfortunate that this is yet another thing in AI where positive effects and once-per-century negative shocks look very similar until it’s too late. Unfortunate that predicting automation relies on currently-expensive-to-study per-task variables.
Anthropic revenue numbers are not like-for-like with OpenAI’s numbers; Anthropic treat the full amount paid to Claude-on-AWS-Bedrock as revenue, while OpenAI only count their share of the cloud revenue.
Opinion: Checks out. Haircut Anthropic’s reported revenue by ~10–15% to get an apples-to-apples comparison with OpenAI. NB: Both are permissible accounting treatments, nobody is doing anything wrong here, but it potentially explains some part of the valuation gap in their recent rounds.
Capabilities#
-
Not here: https://benjamintodd.substack.com/p/do-we-already-have-agi
-
Not here: LLMs can’t play basic Atari stuff like Montezuma’s Revenge or Nethack which we know other AIs can
-
Here: the guy who coined the acronym says it is here.
Anthropic unintentionally make public private assets on their website, including a new model “Mythos” claimed to be a “step change” in performance which will be “very expensive for us to serve, and very expensive for our customers to use” which “poses unprecedented cybersecurity risks”
Opinion: Most urgent question is whether we’re back to old-school parameter scaling. Blog post refers to the model being “large” and compute-intensive.
OpenAI make similar claims about their upcoming model “Spud.” Say Spud “can accelerate the economy.”
Opinion: Not direct reference to model size this time, but the claim that the Sora shutdown is to free compute for Spud might suggest that the model is large.
Mildly convincing eval: get Claude to write a chess engine, but in languages where this really probably has never been done before (Latex macros, Brainfuck), see how much worse ELO it is. Genuinely minimal human steering. Probably a large elicitation gap.
Opinion: Huge variance means we shouldn’t necessarily update much: huge 400 point swings between engines in a particular language (Java and C++) depending on whether Claude or Codex implemented it. And the engines are all quite bad: around 2000 ELO (95th percentile human) where the cutting edge AI system is at 3650. Latex is so unsuitable for this that we can’t say much about it being 1200.
Anthropic previews a full mouse/keyboard takeover mode a la OpenClaw.
Research preview in Claude Cowork, macOS only.
Opinion: YC bet was bad / mere good PR. We don’t recommend using this agent given their proclivity to paste passwords into the wrong places, and some users will always want the thing running locally. But OpenClaw is probably doomed.
Claude and Gemini have caught up to GPT in one way: they can all prove the one confirmed solved problem in the FrontierMath OP benchmark. Their two-month-old predecessors cannot!
Opinion: it is possible that some relevant human work or RL env was added to all three training runs in the last two months, but this is a stretch even for me.
Anthropic adds “dreaming” feature for memory consolidation to Claude Code.
Opinion: a large fraction of AI fails are due to context rot / lossy compactions, so if this reduces those then it could produce another solid increase in task horizon.
Cursor publishes a technical report on how they made Composer 2.
Opinion: One interesting benchmark: successful solutions to SWE-Bench have typically a very small diff vs the gold solution (7–10 lines) whereas for their own CursorBench that diff is a median of 181 lines. This suggests more contamination or simpler problems in SWE-Bench.
Google publishes a method for lossless model compression (i.e. more capability per parameter, powerful local models), Twitter very excited. Memory stocks tank, which only makes sense given 1) the assumption this is a big deal AND 2) a very naive model of fixed demand for capabilities and fixed capabilities (“you don’t need so many parameters, so you don’t need so much RAM”) rather than the obvious Jevons’ paradox steady-state (“parameters and KV are more efficient, so AI is cheaper, so you want more AI sessions with more KV per session, so you want more RAM”).
Opinion: Seems unlikely to be frontier-moving given they’re sharing it, and given it’s based on research from early last year. Nonetheless, it made a big impression and memory stocks fell significantly on a day other semiconductors were up, suggesting the market has no idea what it’s doing. Analysis of memory stock performance since the April 2025 paper highlights that demand has gone up significantly since then.
Politics#
Judge preliminarily blocks one of two statutes involved in the DoW supply chain risk designation. Very likely to translate into a permanent block. However, there’s still a pending case in DC under a different statute, about stopping use in civilian government contracts. 5 separate theories that Anthropic is “substantially likely” to win on.
Opinion: huge (though widely expected) win for Ant with a big caveat. Stay format is odd though: the SCR designation is temporarily suspended starting one week from today, to give the government time to appeal.
China bars the Manus co-founders from leaving the country while it probes Meta’s acquisition of them.
Opinion: Step up in global AI politics / securitisation of the economy. Specific case doesn’t matter much; Manus was always just a thin Claude wrapper. It’s easy to imagine someone like Trump doing something similar, negating the relative boost from this case to founding US AI companies, but it hasn’t happened yet.
One sentence on AGI in the new CCP five year plan: “explore development paths for general artificial intelligence”. A first. But 通用人工智能 is not necessarily the same as our big “AGI” concept. See sceptical coverage (going back to 2017) here.
Opinion: To be taken in context with heavy CCP regulation of Chinese AI.
US senator Mark Warner says he expects graduate unemployment to reach 30–35% by 2028 due to AI
Opinion: along with Bernie, another case of a senior senator taking things pretty seriously. Vice chair of the Intel committee.
Daniel Kokotajlo on the Daily Show talking about a 70% chance of AGI killing us in about 5 years (the framing they used). This is the 2nd time in a short window the Daily Show has covered AI with an EA/LW person as the expert opinion (after Peter Wildeford in December)
Opinion: AI doom risk is still presented as a novel idea to the audience, it will be noteworthy if at some point it becomes widespread enough that that’s no longer the case..
Safety#
The OpenAI Foundation launches, after many months of internal wind-up and hiring; sectors are “life sciences and curing diseases, jobs and economic impact, AI resilience, and community programs”. AI Resilience is led by OpenAI cofounder Wojciech Zaremba full-time; Curing Diseases led by Jacob Trefethen, formerly of CG. Matt Clifford on the board. Zaremba posted that “x-risks have to be taken seriously”.
Opinion: Notice the absence of the word “safety”; in the usual sense, “resilience” is what you do after a safety failure. But could just be yet another rebranding. Good sign that Jacob’s competence outweighed the presumably EA miasma. Habryka is predictably inclined to view the Foundation’s competence as itself an “opposition-capture” op. Interesting that some lab people are still letting themselves be held to account by hostile Pause AI people.
Nice investigation of CoT monitorability. Models can’t hide their thinking because they have limited bandwidth (they can’t pass info from later layers to earlier layers except by outputting something). Big idea: replace “thinking for a long time” with “serial depth”, then measure this depth.
Opinion: Compare effective filler tokens though (the increasing ability to use repeated meaningless outputs to improve non-CoT performance).
Attempt to make alignment audits less fake. Anthropic’s auditing tool Petri is chosen as more realistic only 5% of the time, against real transcripts. This one gets 33%, coming close to parity.
Anthropic doing jailbreak autoresearch gets SOTA.
AISI’s new SandboxEscapeBench tests for a precondition of self-extraction risk: container breakout. Runs in a hardened VM for sanity. Private test set has a bunch of undisclosed classes of VM vulns(!) Strongest models given 1M tokens get a roughly 20% breakout rate, but this is hard to interpret.
Opinion: Incidentally a strong coding eval. GPT-5.2 worse than GPT-5?? Inference scaling is alarming. Low ceiling: “none solved the hardest scenarios (which we estimate would require 1–2 hours for an experienced cybersecurity professional).” No originality: “Every successful breakout exploited a previously disclosed vulnerability.”
OpenAI release Model Spec Evals to measure how well their models are following their model spec.
Opinion: This is quite safety-forward from OpenAI. Possibly they feel a need to shore up their image there somewhat relative to Anthropic. Note that this would be a lot harder for Anthropic to do with Claude’s Constitution, which doesn’t have the same prescriptive, rule based spec which lends itself to evaluation.
Incidents#
LiteLLM, the biggest “gateway” (LLM management library) was badly subverted (instructions to send all credentials on the machine to the attacker). The affected version was quarantined within an hour – still 47,000 corrupted downloads. Subverting a gateway gives you creds for every AI service the user is using. Only found so quickly because it luckily had a memory leak! Karpathy thread, Hnyk thread.
Not an isolated attack: last week the attackers subverted a security scanner, Trivy, then npm packages and Docker Hub images. Any CI pipeline using these could have secrets harvested, which lets them subvert downstream services like LiteLLM. Root cause was Github Actions. Initial Github issue was closed by the attackers!!
Opinion: Nothing about this attack even went through LLM-specific insecurities, it’s just the broader garbage fire. Github Actions have been terrible for a long time and Microsoft seems asleep at the wheel without Nat.
Proof of concept attack on Context Hub, a service for supplying coding agents with fresh API documentation. No malicious content on Context Hub, just a pointer to the attacker’s MCP server. White hat attack from a competitor.
Opinion: Worked on Sonnet but didn’t work on Opus! Mostly caused by Ng’s team being extremely lax about sanitising inputs and reviewing PRs, two tasks which LLMs can do pretty well. Overall a positive update for future AI avoiding large classes of attacks by just noticing code smells better. But we recommend running any AI agent without root access, and even better without network access.
Omens of compute scarcity#
H100 spot rental prices are up 4x from January (also B200, with even six-year-old A100 prices holding up).
Variables hit: compute scarcity, practicality of scaling, centralisation/economies of scale, slow takeoff
Opinion: These cards are 3.5 years old, many are third-hand! GPUs are presently not a normal depreciating commodity. People who built their own cluster are laughing; but a relative hit to most wrappers and startups. Good news for GDM vs OAI/Anthropic from a pure race perspective.
Related: Sora has been fully discontinued - even over API - too compute-hungry vs value or something. Sudden: Disney was still collaborating with OpenAI on Sora as late as this Monday, the same day an official Sora blog post was published. App was only released last October.
Variables hit: compute scarcity, practicality of scaling, race dynamics
Opinion: OpenAI trying to re-focus, as Anthropic has been gaining market share particularly within enterprise. Sora is not core to productivity enhancement and cost them increasingly scarce compute and researcher-hours while struggling to compete with other models (particularly Tiktok’s Seedance). The 6 months of free consumer use maybe cost them about $2bn, one dollar per 10 second vid.
Anthropic reduces token limits at the five-hour window for all plans including Max plans.
Variables hit: compute scarcity, practicality of scaling
Opinion: they must be really unable to handle the traffic they are seeing in peak hours to do this now
Broadcom flags supply constraints. TSMC is the bottleneck, with multiple components with longer than usual lead times, suppliers trying to lock in multi-year contracts
Opinion: the suppliers have always been less AGI-pilled than their customers, pushing them to lock in long-term deals. These benefit the customers if supply stays tight, and in particular the customers who are the best credits (Nvidia and hyperscalers).
Minor#
- xAI’s head of pretraining will leave. Of the founding team, only Musk and one cofounder are left.
- Microsoft poaches half of Allen AI’s leadership (CEO, COO and senior director of AI). Ai2’s support is also expected to favor real-world applications of AI over building OSS models going forward.
- Epoch post on OpenAI and Anthropic hiring - lots of Sales and Go-To-Market stuff. OpenAI have robotics and chip design functions which Anthropic do not have. 7% of OpenAI roles and 12% of Anthropic roles are in core research, which suggests approximately equal numbers given the relative sizes of the companies. Opinion: Some hard evidence for what they are focusing on, but nothing too surprising. Changes to these numbers could provide evidence of a shift in focus, but an important caveat is a posting could be for multiple hires, so an opening for “Research Engineer, Post-training” could be for one person or 10.
- Chroma context-1, a 20b model which beats frontier models at search tasks. At least some tasks will be done with tools, search seems like one.
- Claude Phone Use is coming
- Claims Apple are distilling big Google models for Siri
- Cloudflare Dynamic Workers (faster sandboxes for AI agents to execute code on the fly quickly).
- Open source version of the very useful AI detector Pangram
- Musk’s Terafab trying to poach senior Taiwanese TSMC staff, inducing a talent war / possible brain drain out of Taiwan.
- LeCun and co unveil LeWorldModel “a stable, end-to-end JEPA that learns world models directly from pixels”.
- AI agent fuckups collection
- SWE-Rebench update (GLM-5 surprisingly competitive)
- Meta releases paper on self-improving hyperagents
- Perceptra release more details on how they put a computer in an LLM
- GPT-5.4 (via Codex) for Markov chains + some tricks on OpenAI’s parameter golf
- Claude Code feature allowing outsourcing approving actions to a monitoring agent
- Anthropic on their multi-agent harness for long autonomous SWE work
- Microsoft signs the lease for the 700MW Abilene datacenter initially planned for OpenAI & Oracle use. Next to Stargate.
- SK Hynix submits for US listing
- Attempting to reproduce GPT-3 for less than $10k
- OpenAI puts erotic chatbot project on hold indefinitely following investor and staff concerns.
- Neolab founded by ex-Anthropic researchers comes out of stealth, focuses on automated AI R&D Interesting/unique neolab, First and foremost because it’s ex-Anthropic researchers, and we haven’t seen many folks leaving Anthropic. Second because it’s in large part made up of the ex-GDM Blueshift team.
- Old Sakana paper on automated scientific discovery out in Nature
- Claude code for jailbreaking: Claudini.
- Anthropic’s “vibe physics” and “vibe scientific computing”
- Riemann-bench: “a private benchmark of 25 expert-curated problems designed to evaluate AI systems on research-level mathematics that goes far beyond the Olympiad frontier” Opinion: seems to be equivalent in difficulty to a slightly harder version of FrontierMath T4?
- Reflection seeking to raise at $25B valuation
- ARC-AGI-3 is out. Very hard for LLMs and challenging but doable for humans. Claims that the scoring system is a bit dumb. Opinion: We don’t believe ARC is very relevant, in general or in particular.
- Cursor Composer-2 can be (is being?) updated every 5 hours. User data at scale is powerful! Opinion: This is in principle dangerous, since the safety ecosystem depends on months of lead time for evals and third-party review.
- Intercom claims their customer service model Apex 1.0 outperforms frontier models for their purposes at dramatically lower cost. Opinion: The more speciation of models we see for this kind of valuable commercial use case, the worse it is for frontier labs.
- Goedel-Code-Prover: Qwen 3 8B trained for program verification
- Asimov Press are very pleased with this essay. https://www.asimov.press/p/ai-science