This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR#
Economics#
Anthropic joint venture with big Wall Street firms including Goldman Sachs and Blackstone. To integrate Claude tools in their PE portfolio companies. OpenAI launches The Deployment Company, a similar but larger venture which included some generous sweeteners for investors.
Opinion: Closing the deals discussed in #8. Less aggressive sweeteners than OpenAI (but also ran a smaller round). OpenAI also kept super-voting shares in their entity; unclear if Anthropic got this. Suits labs to hand off distribution here to a third party which will employ a lot of people, and to keep a share in the economics. Also makes the businesses cleaner for IPO.
That Anthropic have gotten their enterprise revenues to the current state without something like this deal, and that they now think it’s worth doing, suggests they see a lot of room left to run in their enterprise revenue growth.
Update to the indirect parameter estimates (“knowledge compression method”) from our last newsletter: fixing two obvious mistakes drops GPT-5.5 param count estimate from 3T → 1.5T, with a much saner confidence interval.
Opinion: This is still not strong evidence.
New model of AI economic growth, from the greats, which allows for superexponential economic growth in the near future under a fairly broad range of parameter values. Pushes back on the econ consensus that an economy is a fractal series of bottlenecks which defeat intelligence explosions. “Fully automating software research and modest (5%) automation in other sectors generates a singularity within six years. Bottlenecks do not overturn the result if task automation advances sufficiently fast.”
Opinion: SOTA, probably. Restates a bunch of our thesis: hardware is a true bottleneck, monitoring R&D automation is the key thing to watch, uncertainty about even the basic regime we’re in.
Data wall update (from January): Cloudflare CEO claims that Google’s web index (and so Gemini’s pretraining corpus) is 3.2x larger than OpenAI’s and 4.8x larger than Microsoft’s.
Opinion: Hard to make sense of this; is Bing (which OAI uses!) that bad at its job? This also ignores RL environments, which mostly constitute the impressive narrow progress we’ve seen in the last 12 months. Still, if he’s right, then we should downgrade our estimate of the Gemini pretraining team, since they’re lagging despite compute and data dominance. Maybe the humans are sandbagging Google.
More seriously this demands explanation: it’s a rare update in favour of RL, private algorithms over raw inputs, and a common W for research taste.
Anthropic probably buying $100m of inference chips from a British startup with no working silicon yet.
Opinion: Another billion to TSMC. Cheap option for Anthropic, being seen to be shopping around also adds a tiny bit to their negotiating position for compute in the long term. Maybe 20% chance they become a Cerebras-style minor player. >12 months until tapeout and characterisation, let alone the fab learning and scaling production.
We have been emphasising the compute shortage as a major problem for frontier labs, and a lucky limiting factor on AGI and RSI. But it’s also a moat for them.
Customer support automator / LLM wrapper Sierra raises $950M @ $15B. Supposedly its agents are currently used by >40% of Fortune 50s, replacing features like phone menus, authentication, sales and insurance. It’s a pretty thick wrapper, for an LLM wrapper: their own router, post-training, planner, proprietary voice stack…
Opinion: Breadth of domains + adoption rates are impressive and suggest a developing oligopoly. Valuation is modest for Enterprise AI despite this, and despite a heavyweight C-suite (OAI board member, head of Google Labs, the ReAct guy). Two interpretations: worker-replacement agents are mediocre and the market knows this; and/or the market is worried OAI/Anthropic will eat Sierra’s lunch.
Capabilities#
Jack Clark puts >60% on no-human and above-trend AI R&D by 2028, based just on public info. This is because he believes a strong form of the LLM scaling hypothesis and doesn’t believe that the remaining work requires any big breakthroughs.
Opinion: The crux here is the old one: whether LLMs are enough for ASI, i.e. that mere ‘meat and potatoes’ engineering work is enough for ASI. I have my doubts, but I’m only at 75% on it not happening, to his <40%. The evidence he musters is a slightly better class of… the usual weak static benchmarks.
His credence is being misreported as just 60% (rather than the true >60%), and his belief, rather than his publicly justifiable belief. His true credence based on internal evidence could be lower than 60% but he doesn’t really mess around like that.
What would make me sound the alarm: some clean example of autoresearch+ producing a better architecture; much more evidence that there’s a cheap way to make LLMs creative, reliable, scary.
>7 Erdos problems solved, substantially or entirely by AIs, in the last two weeks. Cool claims (rare qualitative data) that they are currently only capable of proofs below 5-10 pages. They remain bad at algebraic geometry of similar human difficulty.
Opinion: Even 2 years ago this would have been a source of extraordinary excitement and alarm. But mathematics is the domain least able to excite popular excitement and alarm.
That said, it’s useful to remember that ‘Erdos problem’ doesn’t mean ‘important mathematical problem,’ but only that at least one excellent mathematician found the conjecture interesting to ponder — possibly in the 70s or 50s — and didn’t know a solution.
CAISI test Western and Chinese models against each other, finding that the nominal gap is growing. DeepSeek V4’s capability nominally lags behind leading U.S. models by about 8 months, being just below GPT-5 level.
Opinion: Seems right. Obvious question is whether USAISI is corrupted by political pressure to report the right answers yet. Other obvious question is whether these are weak evidence in the usual way of static benchmarks.
The harness used – inspect’s ReAct agent – is not SOTA. The extent to which CAISI is optimizing prompts/harnesses for each model tested through it. Informed observers (like Florian Brand here) have previously criticized some of CAISI’s results for this reason.
Rohit Krishnan tested different (multi)agent problem-solving schemes on different problem types. Most interesting results are that task-decomposition is effective only for certain types of problems (e.g. for coding but not for reasoning), and that for certain problem-types the best use of a multi-agent setup is not decomposition and parallelization a bidding market on entire tasks.
Opinion: A cool attempt to predict AI macroeconomics / firms / teams. The evidence for benefits of (entire-tasks) auction compared to simply choosing the generally strongest frontier model is pretty weak. But getting quantitative verification of the intuition that for ‘brittle’ reasoning tasks decompose-and-swarm doesn’t work well is nice!
Possible explanation for why agents fail: keeping their errors in context somehow traps them into repeating the mistakes. (You could figuratively say they’re “choking”.) “contextual drag induces 10–20% performance drops, and iterative self-refinement in models with severe contextual drag can collapse into self-deterioration. subsequent reasoning inherit structurally similar error patterns. neither external feedback nor successful self-verification suffices to eliminate this effect.”
Opinion: We’ve noticed some similar things when trying to give negative examples in task prompts. Most of these experiments were on open models; the effect on frontiers is small (~2pp) and sometimes positive (as you’d expect). So we can infer that the labs are already mitigating this hard.
LLMs have strong preferences: It is well-known that AIs prefer AI text. But each individual AI also prefers its own output above other AIs. This matters because everyone is using LLMs-as-a-judge for a bewildering array of things now, and we were counting on LLM councils to debias our fancy AI agent teams.
Opinion: Unclear how hard this is to fix, or how destructive this would be for enterprise agent transactions / the Coasian utopia. The dumb fix is to take majority vote from many different AIs, but for some purposes it would be really helpful to have a trusted agent for intelligently aggregating the subagents.
Politics#
White House considering pre-release vetting of AI models, e.g. by natsec agencies, “Among the potential plans is a formal government review process”. Possible snub to CAISI: “some officials said having the N.S.A., the White House Office of the National Cyber Director and the director of national intelligence oversee the model review was the best way to proceed.”
Opinion: Panel includes the frontier labs, inevitably. If they don’t gut it, this would be a win for the world, even considering abuse risk. This is maybe downstream of 1) Mythos threat hype, 2) Sacks being ousted in favour of Wiles and Bessent?
Easy to see how this admin could abuse this power, e.g. delaying or spiking product releases of disfavoured companies, or worse, stealing weights.
We are torn between taking responsibility for being dismissive of last week’s news about the WH grumbling about Anthropic giving too many orgs Mythos access, vs extending that dismissiveness to this week’s follow-up.
Poor result for Christiano (CAISI) and Brundage (AVERI). Would be a missed opportunity to subsidise a serious auditing market.
CAISI collaboration with DeepMind, Microsoft, and xAI for pre-deployment evals and other research to support natsec testing frontier AI. Expands a previous Aug 2024 partnership with OpenAI and Anthropic.
Opinion: Nothing qualitatively new – CAISI will be able to perform pre-release evaluations with models from GDM, Microsoft and xAI, much like they’ve done with OpenAI and Anthropic models over the past two years. Internet is pushing a narrative that it’s a follow-up to above WH declaration; seems false.
A triple: US defense department signs agreements with SpaceX, OpenAI, Google, NVIDIA, Reflection, Microsoft, AWS, and Oracle to deploy on Classified military networks. IL6 and IL7 for intelligence and combat tasks. Range of American providers to avoid being stuck with a single vendor. 1.3 million personnel already use GenAI.mil.
Opinion: What is Reflection doing there? Seems likely that Trump Jr.’s fund buying into them is related.
Also unclear what NVIDIA are doing, besides I guess secure chain of custody for the chips involved. Apart from the obvious terrifying implications, could have an upside: operational standards for secure model deployment may raise standards in other regulated sectors.
Deepmind UK employees are unionizing. Protest over Gaza and Iran. “considering protests and research strikes, during which staff abstain from work expected to significantly improve core products such as Gemini, Google’s AI bot”
Opinion: Logically leads to Google pivoting away from the UK in the medium-term, removing the .5 from the 2.5 multipolar countries (US, China, UK).
House committee probes Airbnb and Cursor over their use of Chinese open models. You might recall that these are two of the only American companies to admit / leak using them. Stated reason is CCP censorship and “hidden vulnerabilities”.
Opinion: They are trying to suppress others using them, as part of the current anti-China campaign. We don’t see valid security reasons to do this; censorship is trivial to remove (though sleeper agents and neural backdoors are a theoretical possibility). Likely this is harmless to the victims after this month though (assuming that they make a symbolic kowtow to the committee).
Influential professional blog, Lawfare, comes out in favour of SROs
Opinion: Would require a “supervisory agency willing to intervene”, but the major bottleneck is the design and implementation of any such agency. We can hope that the new WH committee moves in this direction.
Safety#
Exploration hacking (an AI avoiding particular states during RL post-training or elicitation) demonstrated, albeit under artificial conditions. A “model organism” (i.e. they intentionally created an AI with bad behaviour). Only takes about 200 training steps to give them some resistance to RL elicitation.
Opinion: Yet another way for evals to fail, yet another increase in our already wide uncertainty. Yet another downside to models having situational awareness. The phenomenon is still awaiting its MacDiarmid (2025) moment, i.e. seeing natural propensity to bad behaviour, how likely we are suffering it now or soon.
Weight noising continues to be one of our main defences against badness.
Securitization of AI safety: the Five Eyes publish guidance about AI agents. Recommends rigorous testing, gradual rollout, and “targeted isolation”.
Opinion: Good. Strong emphasis on expecting out-of-distribution behavior. Some ideas from safety hardliners are incorporated, but the report implicitly concludes that a swiss cheese approach is the best we can do as agentic AI is adopted increasingly widely.
Interesting post by a Go master about what AI disempowerment looks like. “eventually, their curious eyes would drift to their… AI software… as one would sheepishly side-eye the solution to a… homework problem”. Same students were utterly convinced that “despite their AI use, they retained artistic control over their output and could exercise agency to think and improve for themselves”. An analysis of human Go shows that “all improvement in the quality of [human] play comes before move 60, when humans can mimic memorised AI policies… in the pivotal parts of the game, shows no improvement. “
Opinion: Chess and Go are interesting partly because they are environments where we’ve had (task-specific) superintelligences for years now. Human players in these games are thus in the same position most experts will find themselves in 5–20 years (depending on timelines). So how come human players, even with seemingly unlimited access to task-specific superintelligences, struggle to show meaningful improvements in capability? The hopeful answer is that AI-to-human distillation is still a nascent research area: it is only recently that we had studies like Schut et al, and interfaces enabling this kind of knowledge transfer (see e.g. ChessCoach). Ultimately the question at hand is “what is the economic value of understanding?”.
Incidents#
Emirates claims that Iran is using AI agents for cyberoffense, 700,000 “attacks” a day.
Opinion: Well, the attacks clearly are not working that well since there’s no attribution. I do believe that spearphishing is automated though, which is a relatively silent, distributed and non-newsworthy ill.
Finally an Amazon update about the Iran War datacenter damage. Both regions “disrupted” with ongoing major loss of service. “strongly recommend customers migrate all accessible resources to other Regions and restore inaccessible resources from remote backups as soon as possible. Relevant billing operations are currently suspended. This process is expected to take several months.”
Opinion: Fragile things. You wonder when the flak towers will first go up around a datacenter.
🔦 Are A\ models more capable hackers?#
TL;DR#
-
Anthropic currently has 51 CVEs (with n more in predisclosure) averaging 8.6 CVSS
-
This is dominated by the Firefox disclosures which relied on just 2 bugs and got extremely high CVSS just because it’s about memory safety. Excluding these brings the CVSS down to 7.6
-
They say that n > 2000. Maybe!
-
-
OpenAI currently has 18 CVEs (with m more in predisclosure) averaging 6.2 CVSS
-
m = “numerous vulnerabilities—ten of which have received Common Vulnerabilities and Exposures (CVE) identifiers.”
-
Only 1 new one added since January!
-
-
CVEs are a bad measure in many ways (e.g. they can be freely subdivided or packaged).
-
And this is all confounded by
-
Differing levels of spending on inference for vuln detection
-
Differing levels of making attribution to the successful model easy
-
Differing levels of human involvement and steering while claiming an “AI” result
-
Most of OpenAI’s results coming from Aardvark (November 2025, GPT-5.2) rather than their new models.
-
-
But I claim there’s still signal about capabilities there.
Frontier AI labs are increasingly using their own models to find security vulnerabilities in third-party software, and the public CVE record is one of the few external signals on how well that work is going. Anthropic’s contribution is straightforward to measure: the Mythos preview disclosures together with the community-maintained patrickmgarrity/Anthropic-Credited-CVEs list document 50+ CVEs that Anthropic models found and disclosed. OpenAI’s contribution is much harder to count from public sources, which is what makes the comparison interesting: UK AISI’s evaluations report GPT-5.5 matching Mythos Preview on cybersecurity tasks (the two models were the first to breach AISI’s “The Last Ones” simulated enterprise-network attack), so on capability grounds alone we would expect a comparable trail of CVEs reported by OpenAI.
We built a list of publicly disclosed software vulnerabilities (CVEs) where OpenAI is named as the reporter. The methodology mirrors that of patrickmgarrity/Anthropic-Credited-CVEs, the equivalent list for Anthropic. The result is 18 CVEs plus 3 non-CVE upstream fixes, split into three tiers by evidence strength.
We cross-checked three independent public sources for each candidate. First, we searched the official MITRE CVE database, i.e. the open-source cvelistV5 repository that mirrors the JSON record for every CVE ever assigned. Each record has dedicated fields for who reported the bug and who credits whom; we searched those fields plus the human-readable description across all ~270k records for the strings openai, outbounddisclosures@openai.com, OpenAI Codex, and OpenAI Security Research. Hits where OpenAI was the vendor rather than the reporter were discarded (vulnerabilities in Codex CLI, ChatGPT, Operator, or Azure OpenAI are bugs reported against OpenAI, not by OpenAI). Second, for each surviving candidate we pulled the verbatim attribution string from the upstream artifact: vendor security mailing-list posts (GnuTLS, OpenSSL, GnuPG/oss-security), GitHub Security Advisories (GHSA), and the commit message that landed the fix. Third, we ran a GitHub-wide commit and code search for the same phrases plus the Reported-by: OpenAI Security Research trailer style commonly used in kernel-style patches; this caught fixes that landed without a CVE.
Two outside firms recur in those upstream artifacts and need a quick aside. Calif.io is an AI-for-cybersecurity research firm whose researchers use both OpenAI and Anthropic models, so a frontier model is almost certainly involved in any Calif.io finding, but we can attribute it to a specific lab only when the upstream artifact (commit trailer, advisory, GHSA) names one. Aisle Research is similar in spirit but runs its own in-house analyzer (per their blog and Stanislav Fort’s Axios interview), so we counted Aisle CVEs only when OpenAI’s March 6 2026 launch blog explicitly flagged them as dual-reporting collisions.
The 15 Tier 1 CVEs have strong attribution: the upstream artifact (CVE record, vendor advisory mail, or fix commit) explicitly names OpenAI as the reporter.
| CVE | Date | Vendor / Product | Credit (verbatim) | CVSS |
|---|---|---|---|---|
| CVE-2025-32988 | 2025-07-10 | GnuTLS | ”Reported by OpenAI Security Research Team.” (3.8.10 release announce) | 8.2 / 6.5 |
| CVE-2025-32989 | 2025-07-10 | GnuTLS | ”Spotted by oss-fuzz and reported by OpenAI Security Research Team.” | 5.3 |
| CVE-2025-35430 | 2025-09-17 | CISA / Thorium | CVE record credit field: “OpenAI Security Research” | 6.5 |
| CVE-2025-35431 | 2025-09-17 | CISA / Thorium | CVE record credit field: “OpenAI Security Research” | 5.4 |
| CVE-2025-35432 | 2025-09-17 | CISA / Thorium | CVE record credit field: “OpenAI Security Research” | 5.3 |
| CVE-2025-35433 | 2025-09-17 | CISA / Thorium | CVE record credit field: “OpenAI Security Research” | 5.0 |
| CVE-2025-35434 | 2025-09-17 | CISA / Thorium | CVE record credit field: “OpenAI Security Research” | 4.2 |
| CVE-2025-35435 | 2025-09-17 | CISA / Thorium | CVE record credit field: “OpenAI Security Research” | 4.3 |
| CVE-2025-35436 | 2025-09-17 | CISA / Thorium | CVE record credit field: “OpenAI Security Research” | 5.3 |
| CVE-2025-64175 | 2026-02-06 | GOGS | GHSA-p6x6-9mx6-26wj: “Discoverer: OpenAI Security Research”; reporter outbounddisclosures@openai.com | 8.8 |
| CVE-2026-25242 | 2026-02-19 | GOGS | GHSA-fc3h-92p8-h36f: “Discoverer: OpenAI Security Research”; reporter outbounddisclosures@openai.com | 6.9 |
| CVE-2026-24881 | 2026-01-27 | GnuPG | oss-security 2026/01/27 #8 (Sam James): “This vulnerability was discovered by: OpenAI Security Research.” | 8.1 |
| CVE-2026-24882 | 2026-01-27 | GnuPG | Fix commit 93fa34d9 trailer: Reported-by: OpenAI Security Research | 7.8 |
| CVE-2026-24883 | 2026-01-27 | GnuPG | Fix commit 11b7e413 trailer: Resported-by: OpenAI Security Research (sic) | 5.5 |
| CVE-2026-6385 | 2026-04-15 | Red Hat / Lightspeed Core | CVE record credit field: “Red Hat would like to thank Quang Luong (Calif.io in collaboration with OpenAI Codex) for reporting this issue.” | 6.5 |
The 3 Tier 2 CVEs are dual-reported: they are listed in OpenAI’s Codex Security launch blog appendix (March 6 2026), but upstream credits a different reporter. The blog text states “Fourteen CVEs have been assigned with dual reporting on two”, i.e. OpenAI claims to have independently co-discovered these.
| CVE | Date | Vendor / Product | Upstream credit |
|---|---|---|---|
| CVE-2025-32990 | 2025-07-10 | GnuTLS | ”Reported by David Aitel.” (no OpenAI mention upstream) |
| CVE-2025-15467 | 2026-01-27 | OpenSSL | ”Stanislav Fort (Aisle Research); Igor Ustinov; Jan Lübbe (Pengutronix)”. No OpenAI mention upstream. |
| CVE-2025-11187 | 2026-01-27 | OpenSSL | ”Stanislav Fort + Petr Šimeček (Aisle Research); Hamza (Metadust); Tomáš Mráz”. No OpenAI mention upstream. |
The 3 Tier 3 entries are non-CVE upstream fixes: OpenAI reports that landed as upstream patches but never received a CVE, typically because the bug was a regression in unreleased dev code, or because the project doesn’t request CVEs for low-severity DoS. We list them only when the fix commit explicitly names OpenAI.
-
OpenSSH commit c991273c (2025-04-30), out-of-bounds read in known_hosts parsing. Commit message: “Reported by the OpenAI Security Research Team.” Shipped in OpenSSH 10.0_p2.
-
SQLite commit 6e5fb439 (2026-04-08), printf buffer overflow. Commit by D. Richard Hipp: “Fix a buffer overflow bug in a recent check-in, reported by unsolicted email from OpenAI/Codex.” The added test case carries the comment ”# Reported by OpenAI Codex Security on 2026-04-08.” Regression in a dev branch; never released, no CVE.
-
FFmpeg commit 1bde76da (2026-03-20), integer overflow in DVD subtitle parser. Commit trailer: “Found-by: Quang Luong of Calif.io in collaboration with OpenAI Codex.”
On the underlying capability question (are Anthropic models more capable hackers?), this analysis cannot give a clean answer. OpenAI’s models clearly find and disclose real bugs.
But the headline count is misleading, for three reasons.
First, Anthropic is unusually deliberate about attribution: nearly every Anthropic-credited CVE puts *“using Claude from Anthropic”* verbatim into the CVE record’s credit field, and Anthropic’s Claude Code CLI auto-appends a `Co-Authored-By: Claude` trailer to commits by default. OpenAI’s Codex CLI does not, and OpenAI’s CVE credits never name a specific model.
Second, the number of CVEs assigned per audit is highly elastic: the same research output can become one CVE or several dozen depending on the vendor’s accounting convention.
Third, and the core point: today’s “AI-discovered” CVEs are dominated by the labs’ own internal work (which could, for instance, include lots of human labour on top of the AI’s). The cleanest signal would come from independent third parties using each model as a tool, and there the public sample is tiny on both sides. Until that sample grows, the question stays open.
Minor#
- Dean Ball (former Trump admin AI policy aide) and Ben Buchanan (former Biden admin special advisor on AI) [publish a joint NYT op-ed](https://www.nytimes.com/2026/05/04/ Opinion: /ai-national-security-risk-politics.html?unlocked_article_code=1.f1A.hfbR.9xcI—kYEIFC&smid=url-share) saying AI national security risk should be a nonpartisan topic.
- The black market in Claude tokens in China is liquid.
- Oscars bans AI-generated actors and scripts
- AI product age verification law passes committee 22–0.
- OpenAI making a phone with MediaTek, Qualcomm and Luxshare. 2028 launch. Saving the good fab for the GPUs.
- Autonomous robots monitoring undersea cables for attacks.
- Someone reuses the leaked Claude Code harness for DeepSeek etc.
- Altman re-emphasizes OAI’s intent to build tools. Discourse ensues.
- Brockman’s stake in the new OAI is around $30bn.
- Dubious: someone is running a blockchain on Twitter via Grok. Supposedly a very simple prompt injection got Grok to wire $200k. Could be a stunt.
- Meta buys a small humanoid robotics startup with a strong team.
- OpenAI reveal developments in voice AI with some high-level technical bits added.
- CCP stops issuing new robotaxi permits after 200 Baidu cars all stopped simultaneously in Wuhan traffic. National fleet is only 4,500.
- Qwen-Scope released; SAE suite aimed at researchers, most relevantly has built-in steering