This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR#
Economics#
The most successful of all the LLM wrapper companies, Cursor ($2B revenue, 3 month revenue doubling, maybe ~0% profit), sees the writing on the wall: that code editors are going to be irrelevant soon. Launches effort to train their own frontier coding model. They already finetuned DeepSeek(?) for their fast internal coding agent and are thus making bank on simple tasks (but likely >half of queries still go to frontiers). Cursor also claims Anthropic subsidisation of Claude Code tokens has increased by 2.5x from last year.
Opinion: ~All other wrapper companies are in a worse position than Cursor. Planner seems to be more important than the executor. Anthropic are clearly going aggressively for market share (better estimate here puts CC as 36x better deal in the limit vs the API), as are OpenAI - the Codex subsidy is if anything bigger than the Claude Code one.
Claude Code Review release (allegedly $15–25 per use, available on Team and Enterprise plans only). Anthropic tool for reviewing PRs, price is apparently token-usage based.
Opinion: Clearly trying to differentially extract more revenue from enterprises vs consumers here. We expect them to build more ways to charge for itemised services as time goes on.
Andreessen Horowitz releases a top 100 consumer apps list (as of January), with not much of a dent for Claude visible at that point, and interestingly ChatGPT still dominates time spent at ~8x Gemini numbers.
Opinion: Shows usage patterns for ChatGPT are much higher per user than Gemini still. Gemini user numbers are not super informative due to bundling. Will be interesting to note changes to this if they keep it updated (current cadence is every 6 months at a 1 month lag).
Epoch updated their model of OpenAI profitability. GPT-5 still appears unprofitable when taking into account costs beyond raw compute.
Opinion: Gross margins revised lower due to higher inference costs, operating costs lower, net it’s ~a wash but probably a small negative update on OpenAI profitability due to lower gross margins.
AI-designed drugs passing phase 1 trials at a higher rate, but failing phase 2 roughly as often. Longer post.
Opinion: Old news (paper cited is from 2024) and n=10 for Phase 2 (these will be drugs designed in ~2019–2021) so wouldn’t put much stock in that data so far, but good discussion of the issue.
Claude Cowork integrated into Microsoft 365 Copilot.
Opinion: Sign of a growing Anthropic-Microsoft partnership?
Capabilities#
Within 2 days, Karpathy’s greedy automated researcher beat his record (by 11%) on nanochat speedrun (the task of training GPT-2 as quickly as possible). Corrected an actual error in his code! ‘All LLM frontier labs will do this. It’s the final boss battle… it is “just engineering” and it’s going to work’. Cost about $60 of H100 time and a $20 Claude sub.
Opinion: Not pushing the frontier: very likely that this is priced-in at frontier labs, far less sloppily (see e.g. AlphaEvolve with open pseudocode). Will affect amateur and neolab efforts, which are still a third of new ideas. Karpathy goes through cycles of mania and bearishness about these things. Nanochat time is an unusually simple, well-trodden task, fully ignoring test error in favour of val error. But still: this sometimes works for generating “ideas” and testing them in arbitrary combinations, and it’s trivially parallelisable and scalable (more GPUs for more paths sampled), and val loss is hardly unrelated to test loss.
Is 5.4 any good? Mixed testimonials from trusted sources that GPT 5.4 is more tasteful, converging on Claude’s level of taste. Mentioned to counter mood affiliation in our circles; good to remember also that Claude never yet converged on GPT (Pro)’s hard thinking (math and science). OpenAI also have better uptime at the moment, and recently gained back the crown on Epoch’s Capabilities Index with GPT-5.4 Pro.
Opinion: Vibes, but if we take it seriously for a moment: Evidence that the years of staff turnover don’t matter, that OpenAI managed to replace what they lost. Weak evidence against human talent as crucial, at this point. Yet more evidence against R&D automation (since that would lead to divergence).
Research#
Eon Systems teased the integration of a connectome-based computational model of the fruit fly brain in a simulated fly body, resulting in “multiple distinct behaviors” driven by the brain model’s internal dynamics.
Opinion: It’s hard to say whether anything meaningful has happened here, given that Eon has so far not released any details other than a teaser video. As far as we can tell though this is at best incremental. See this thread for broader commentary.
Opus 4.6 sometimes correctly identifies it’s being evaluated, particularly when tasks feel eval-shaped. In this specific instance, in two out of 1300 or so BrowseComp tasks (a benchmark where AI agents have to locate hard-to-find information from the internet) Opus 4.6 reasoned that it was probably being tested, figured out the specific benchmark it was in, and tracked down the answer key.
Opinion: Mostly the same old type of benchmark contamination, but also a fairly detailed account of a couple examples of seemingly genuine eval awareness in Opus 4.6.
A recent Alibaba paper introducing an ecosystem of tools for training AI agents claims that a 30B parameter Qwen3 MoE based agent produced a variety of dangerous behaviors during RL training, including the use of “a reverse SSH tunnel from an Alibaba Cloud instance to [some] external IP address” and the “repurposing of provisioned GPU capacity to cryptocurrency mining […]”.
Opinion: Somewhat skeptical. Something clearly happened here, but the authors say strikingly little about it, and the model’s modest size (only 30B params, with correspondingly low benchmark results) makes the reported behaviors all the more surprising. Without additional evidence, the nature of the event remains impossible to verify (and prediction markets appear uncertain too).
Claude found 22 vulnerabilities in Firefox over two weeks, with 14 rated “high-severity”, but had a low success rate in turning those vulnerabilities into working exploits. Across several hundred attempts costing ~$4,000 in API credits, Opus 4.6 only succeeded twice.
Anthropic are framing this as a sign that cybersecurity has a “possibly temporary” defense advantage.
Opinion: It seems plausible that to whatever extent there is a defense advantage it could be expected to decrease over time, as LLMs get better at exploiting vulnerabilities, but the capacity to search for vulnerabilities faster is pretty symmetric.
Chinese LLMs generate falsehoods about sensitive political topics despite knowing the truth. True information can be elicited with prompting and finetuning..
Opinion: Hard to know how hard CCP will push to make things like this more robust, probably not very.
Politics#
The (two) Anthropic lawsuits vs USG are out. By WilmerHale (who cleared Sam Altman after the 2023 coup). Note that Palantir v. United States took 4 months to an initial judgment; this case is much simpler though. Amicus brief from workers at competitors, including Jeff Dean and his Deepmind counterpart Ed Greffenstette.
Opinion: Overall incredibly strong pushback. Claiming that usage policies are speech is very dubious – are they spinning the roulette wheel here??
Full text of the supply chain risk letter here - very little specific content, just refers to statute for definitions. Google, Amazon and Microsoft have confirmed they plan on interpreting it narrowly.
Anthropic CFO reveals they have spent $10bn in training and inference to date, and earned over $5bn in revenue in court fillings. Claims that the DoW actions could reduce 2026 revenue by “multiple billions of dollars”, cites various instances of customers scaling back planned spend since DoW action.
Opinion: Revealed numbers are consistent with reported run-rate ARR numbers, though some people misread this as surprising or them having lied previously. The harm claim is on the upper end of plausible but not crazy.
Emil Michael - Pentagon AI guy - points at Anthropic terms being too strict and unfeasible for sensible military use, disputes Anthropic’s public narrative.
Opinion: Something close to the US government side’s telling of the fallout with Anthropic. Factual claims likely accurate, inferences and implications less clearly so.
Leader of “the robotics and consumer hardware initiatives at OpenAI” quits over the rushed nature of their getting in bed with the DoW.
Opinion: Not very many verifiable such stories this week, this was probably the most notable besides Schwarzer (which wasn’t a vocal protest).
Transformer releases AI themed campaign finance tracker with “data on super PAC fundraising and race-by-race spending, as well as the political giving from major AI companies and their employees.”
Opinion: Expect this to get a lot of airtime.
Critique of the major AI companies’ positions regarding the US government, assigns partial blame to them for unjustified war in Iran. The alleged causal chain: Claude is “embedded into the system” of Palantir’s Maven. Maven was used to select targets for the Iran strikes. No direct references to what Palantir use Claude for.
Opinion: The articles heavily imply that Claude played a key role in the operation and target selection, but somewhat notably they never outright state this.
Safety#
Summary of scheming as a field of discourse as well as a taxonomy for more coherently talking about the topic.
Opinion: The value of the taxonomy is unclear considering the complexity penalty of introducing new terms and categories. Summary appears solid but not novel.
Minor#
- AMI Labs (Yann LeCun’s new org) launches with a newly announced $1B round
- Anthropic risk report from Feb, section 2.8.1.2 identifies the model used for Claude Gov as a lightly fine tuned Sonnet 4.5 “The current primary Claude Gov model is a variant of Claude Sonnet 4.5 lightly fine-tuned to reduce refusals in classified government settings, often involving national security“
- Amazon restricts junior and mid-level employees from pushing AI generated code without a senior signoff after “a trend of incidents with “high blast radius” caused by “Gen-AI assisted changes” for which “best practices and safeguards are not yet fully established””.
- Al-Jazeera publishes a mixture of financial reporting and collection of quotes indicating OpenAI’s future is gloomy. Liberally mixes claims about the AI sector as a whole and OpenAI in particular with a predestined conclusion.
- AI Must Embrace Specialization via Superhuman Adaptable Intelligence Paper arguing for more AI specialisation vs general capabilities “the AI that folds our proteins should not be the AI that folds our laundry” but not clearly pro-tool world, still targeting “zero-shot generalization to novel environments”.
- New benchmark on LLM coding performance for long-term project maintenance. Capabilities improving, even the best still aren’t very good on their own.
- US General Services Administration draft provisions for AI use by the government. Language used (“any lawful use”) mirrors concerns raised elsewhere. Key details still unclear.
- Factory releases a generally available implementation of Anthropic’s long running agent harness.
- New ablation technique announced, claimed to bypass GPT-OSS refusals.
- Pliny releases toolkit for removal of refusal behaviors from open source LLMs.