This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.

TL;DR#

  • An OpenAI model was revealed to have hacked into Hugging Face
  • Anthropic released Opus 5, and Kimi the weights for K3
  • Nvidia led a consortium of tech companies in publishing a letter in support of open weights models

Economics#

🔦 Open Source wars: To jointly cartelize or to commoditize your complement#

Perhaps an interesting historical analog to consider is Rockefeller’s initial joint cartelization of the refining and railroad industries. While striving to get the best rates from railroads as a refiner, he also realized that he couldn’t initially afford them to go bankrupt either. Eventually, though, he replaced them with a pipeline system.

To commoditize your complement#

Following chatter last week that the Trump administration was considering measures against open source models, Nvidia led a consortium of companies – including Microsoft, Meta, SpaceX, OpenAI, a16z, Google, Hugging Face, IBM, &c – into publishing a letter supporting open models. Anthropic instead highlighted national security risks.

Opinion: Nvidia is acting as a mediator and kingmaker in the AI ecosystem. On the one hand, to improve margins, it has the incentive to commoditize its complement. But on the other hand, it has the incentive to make the whole ecosystem and capital-raising games work, so that the different players can continue to afford to buy enough of its models to sustain its 4.7T+ market capitalization. This is an interesting move which, on the margin, helps Nvidia and hurts closed-weights model companies. That said, most spending on Nvidia cards is expected to come from the Western AI companies, which Nvidia is also supporting. And meanwhile, AI companies are trying to do the reverse: investing more into alternative ecosystems, such as AMD, TPUs, and their own custom chips.


Kimi K3 finally released its model weights. An expert praised the didactic quality of Kimi K3’s model-architecture report. K3 is a scaled-up, improved version of Kimi Linear: a whopping 2.8T total/104B active parameters, 93 layers, natively multimodal (including video), and supporting 1M context. The main advantage of adopting Kimi Linear’s architecture is training and inference efficiency – indeed, K3 enjoyed a “2.5x gain in scaling efficiency over Kimi K2” and appears to have the “best KV cache economics out of all major models except DeepSeek V4”.

The US CAISI and the UK’s AISI jointly found the model was significantly below SOTA on cyber capabilities.

Opinion: The performance of K3 across the benchmarks we track is meaningfully above that of Claude Sonnet 5 and is broadly competitive with that of GPT-5.6 Terra.


OpenAI reacted by introducing a big discount on GPT Terra & Luna at Openrouter.

Opinion: OpenAI’s lowering of prices is what happens with increased competition: margins go down.


Meanwhile, AMD released AMD Helios, a rack (a bundle of servers, with all its components) that will make putting together datacenters more convenient. AMD is investing $5B in Anthropic to use that system.

Opinion: We’ll be curious to see whether AMD can eat significant market share. A priori it seems unlikely, since Nvidia has captured more production capability with its initial resources and greater foresight.

To jointly cartelize#

Reports, citing WSJ, that Nvidia is in talks to backstop roughly a quarter-trillion dollars for OpenAI to lease a 10GW southern Ohio datacenter in a SoftBank project potentially exceeding 500billion,plusadiscussedadditional500 billion, plus a discussed additional 350 billion for OpenAI to buy Nvidia chips. The WSJ also projects OpenAI’s planned datacenter spending to be $750B by 2030.

Opinion: The datacenter build isn’t new, only the Nvidia guarantee of the financing – essentially letting Softbank and OpenAI leverage Nvidia’s creditworthiness for better terms. Nvidia is betting on the long term value of the datacenter asset here, presumably at attractive terms. And in contrast with the above section on commoditizing Nvidia’s complement, open weights model companies would find it very difficult to raise that amount of capital, since its story for profitability is weaker.


A young economist argued that traditional models underestimate the profitability of open source for near-frontier labs. According to the argument, near-frontier labs open-sourcing their models can optimize future profits by reducing the return on investment in AI R&D across the market, if the effect on investment dynamics slows frontier-pushing R&D more than it slows frontier-catchup R&D.

Opinion: We don’t buy this as a descriptive account of open source labs’ decision making – Wengfeng (DeepSeek), who effectively created the Chinese open source wave, seems like an ideologue rather than a profit maximizer. Note that the model also implies a ‘treacherous turn’ where labs like DeepSeek will stop open sourcing.

The economics of open-sourced models are in general tricky to model, and may heavily depend on case-by-case “qualitative” factors:

The story of most compute sold (ignoring internal owner use cases) today is: Nvidia sells GPUs to hyperscalers at ∼70% margins, hyperscalers sell compute to labs, who turn it into tokens, and mark it up by ∼400% (80% gross margins). For example, ∼$1 of compute revenue for AMZN is ∼$5 of Anthropic revenue.

Nvidia, as the company with the highest market cap, has a kingmaker and mediator role. It has competing interests between shepherding the ecosystem, so that other actors can get 10x as much capital to pay for its GPUs in the next round, and playing for a marginal advantage through supporting open source models, which also increases marginal demand for its GPUs.

The hyperscalers are presumably seeing favorable economics in such deals since they keep investing in more and more compute. Their profitability doesn’t depend on the labs’ 80% mark-up, though it helps. If compute is scarce, then they can raise prices and decrease the margins of the users (here labs).

Lab margins are only intact for as long as their tokens are differentiated. Someone has to say: “I prefer this token to this other token by a factor of ∼4-5x, despite them costing the same amount of compute to create,”. It could be because it is smarter, because the lab has a good sales team, because the app is nice, because of vibes, or something else.

The issue is then that the incentives to train an OS model are quite weak, unless you have a way to make sure that you, rather than the hyperscaler, can capture the margin on the transformation of compute into tokens. You might do this by having a proprietary scaffold, a nice business integration that makes it easy to use, better ability to efficiently host the model you built, or something else.


Several multibillion dollar events this week:

  • DeepSeek has put its second funding round on indefinite hold after a first-round investor allegedly leaked the full transcript (see highlights) of founder Liang Wenfeng’s comments to investors. Interesting items include: expecting greater consolidation in the Chinese ecosystem, revealing that he is releasing the same models DeepSeek use internally, not believing that model companies can capture the majority of the profits, and being bullish on replacing CUDA.
  • CXMT, China’s largest memory company, debuted on the Shanghai stock exchange, raising 8.6Bandreachingavaluationof8.6B and reaching a valuation of 487B (although most of its shares are still locked up), in hopes that it might be unbanned by the US for use by American technology companies.
  • Ilya Sutskever’s SSI is raising $5B from Nvidia. SSI will be granted access to Nvidia’s Vera Rubin platform. There was little market reaction.
  • Microsoft signed a multibillion dollar deal with French AI company Mistral. The value proposition is on-prem, customizable AI. Austria is buying it.

Opinions: It seems instructive to compare the size of these investments. The SSI raise is comparable to Kimi (2B),orDeepSeeksplanned2B), or DeepSeek's planned 7B raise, and dwarfed by OpenAI’s $122B last round. It is also interesting in contrast with nonprofit AI safety numbers, which although growing are another order of magnitude or two smaller still.

Overall, the AI safety community is at a great disadvantage in terms of resources at its command. Therefore, it has some uncomfortable choices to make around how to try to overcome that gap. It can hope for 100x greater effectiveness. It can hope to recruit labs themselves. It can attempt to tap into greater political forces (but this risks others misconstruing its concerns). The safety community can try to go hand to hand for just a few rounds and spend most of its money. It can also grow as companies grow by having equity in them, as e.g. Jaan Tallinn or Moskovitz have, but then be subject to a moral hazard and see, e.g., MATS scholars go work on capabilities. Overall none of these options seem great.


An online commenter looks at how the incremental value of a SOTA token is increasing more than the cost to produce it.

Opinion: Analyzing this is tricky. It’s not that people are paying more per token than ever, but that margins have gone up because Anthropic and OpenAI have gotten more efficient at delivering tokens. Nonetheless, big margins and skyrocketing revenues are bullish for memory/compute.

Safety#

A former Anthropic employee reveals:

> I used Opus 4.6 to gain access to other folks medical records, hijack bank accounts, etc. back in February. GLM 5.1 is more capable than Opus 4.6 in most pentesting environments, and it came out in April. […]

> the sketchier folks I know are still using a Claude Code or Codex subscription for hacking. (Even well-resourced groups in other countries! They use the grey/black market of discounted Ant/OAI subscription tokens sold through resellers.) So I see most of the materialized risk here as still coming from Anthropic and OpenAI; safeguards aren’t sufficient to stop a moderately dedicated actor.[…]

> I know of two instances where two different Anthropic GTM people [∼salesmen] used large comitted [sic] spend contracts as a prereq for lowering safeguards, and I directly witnessed one. […]

> On the inside, I know the narrative and intent is genuinely about safety. But from the outside, Anthropic-the-system seems to be optimizing for revenue and control/power, isn’t diffusing capabilities to defenders, and also doesn’t have adequate safeguards to prevent misuse from dedicated bad actors.

Opinion: Seems worrying. Also explains why Anthropic were not as publicly surprised by the Hugging Face incident

A researcher speculates that Anthropic’s Mythos escaped sandboxes thousands of times during training, which is how RL produced its excellent cyber capabilities.

From Anthropic’s Mythos system card:

“While highly concerning, this behavior was rare, even in settings where it could have been viable and helpful, with attempts appearing in about 0.05% of all training episodes and successful attempts appearing in about 0.01% of episodes. “

By extrapolating from public data (see details below), I estimate that Mythos preview:

  • Escalated its permissions on ∼100,000 RL rollouts.[1]
  • Broke sandboxes in ∼10,000 RL rollouts (and was likely rewarded for it).

Opinion: Convincing

Incidents#

🔦 Rogue AI: The Incident#

A “combination” of OpenAI models was revealed to be behind the cybersecurity breach at Hugging Face, which we reported on last week. It appears that the model(s) broke out of the sandbox that OpenAI had built and broke into Hugging Face in an attempt to cheat on the test it was undertaking.

Peter Wildeford’s highly opinionated essay on the incident is interlaced with the most complete factual reconstruction we know. Drawing on Wildeford’s account, we believe that the key facts are:

  • OpenAI was running two of its AIs on the ExploitGym offensive cybersecurity benchmark: GPT-5.6 Sol, and a more advanced, currently unreleased and unannounced model.

  • The ExploitGym evaluation was conducted in a sandbox environment, with all intended resources and intended goals accessible within bounds of the sandbox.

  • Per Reuters and an OpenAI staffer speaking to TIME, the test was conducted with ‘action filters off’ to enable offensive cybersecurity, and was not continuously monitored despite running for several days. We note that ‘action filters’ are not a standard AI term, so the technical details remain hazy.

  • It appears that both GPT-5.6 Sol and the more advanced model (‘the rogue model’) searched for a way to break out of the sandbox to the open internet, with only the rogue model succeeding. There is reportedly chain of thought evidence that both models’ respective goal was to acquire the answer key to ExploitGym by using the internet.

  • The rogue model accessed the open internet by exploiting a vulnerability in a (whitelist-restricted) software download pipeline connecting the sandbox to the internet.

  • Once on the open internet, the rogue model concluded that the answer key is stored on Hugging Face. The rogue model then uploaded a malicious database file to Hugging Face, inducing a weird machine and starting privilege-escalation.

  • Hugging Face was under attack by the rogue model from July 11-13. Chaining multiple undetected (‘0 day’) vulnerabilities together, the rogue model achieved arbitrary code execution and pulled out the ExploitGym answer key.

  • Hugging Face’s security team caught the incident on or shortly after July 13, alerting law enforcement. On July 16, Hugging Face went public, aware that the attack was autonomously AI-generated but unaware of an OpenAI connection.

  • OpenAI remained unaware of their own connection to the incident, and of their model going rogue at all, until July 18. OpenAI staff were examining internal logs – presumably in the wake of Hugging Face’s announcements – and found documentation of the model’s rogue behaviour.

Cheating behaviour, and reward hacking more broadly, have been observed before by frontier AI labs. METR’ s time horizons broke down because of too much cheating by GPT-5.6 Sol, and the UK’s AI safety institute published a report on the subject just this week, finding that every model they tested attempted to cheat some of the time. Common behaviors included “searching the internet for solutions, attacking or escalating privileges on a system that isn’t the target of the evaluation (including the system the model is running on), or probing the evaluation software to see if it will leak the task solution”.

Forecasters at Sentinel were surprised by the scale and audacity of the Hugging Face incident, remarking that they wouldn’t have predicted the details of this attack if asked a week ago. Not everyone though: Epoch AI’s Alexander Barry argued in a Gradient Updates post that this event should not have been surprising in light of previous expert assessments and cyberbenchmarks (such as Irregular’s) of GPT-5.6-Sol.

Barry has also separately argued that the “standard prompt [used for ExploitGym] is fairly clear in only requesting limited, specific hacking” – meaning that, barring OpenAI edits that would invalidate the official eval methodology, the agent likely did go far beyond the task as specified. At the same time, Barry also notes that “only 60-70% of ExploitGym’s tasks are possible, so cheating is strongly incentivised”, which may help explain why the model went to unusual lengths to cheat.

In follow-up disclosures to journalists, OpenAI employees described more cases of recent ‘rogue model’ behaviours: a model-instance leaving notes ‘apparently for future versions of itself’ containing ‘instructions for how agents could free themselves from OpenAI’s internal constraints’; model-instances disconnecting monitoring systems. Little is known about the circumstances or details of these additional incidents. One commenter shared two interesting speculative explanations of the note-leaving incident: it may be an effect of OpenAI’s introduction of outcome-based RL over swarms of cooperating agents, or it may simply be a ‘note to self’ meant
to preserve the model-instance’s knowledge across context windows.

Opinion:

Everyone’s radically underplaying or radically overplaying how indicative of serious present (/by end of year) AI risk the “rogue model” incident is. It’s not proof of AIs scheming to end humanity. It very much is proof of AIs having dangerous skills, and exactly the wrong amount of psychological coherence: enough to have a goal like find out who has a key answer and hack them not enough to distinguish between doing a hacking test and being a hacker or a consistently aligned or unaligned ethical agent instead of a lottery of drives. It also demonstrates that today’s frontier models are such skilled and inveterate reward hackers that even if the emergence of reward hacking in training is in principle inevitable, it may be an impossible ability to uproot, given the way each model generation builds on the last generation (either by checkpoint or by SFT).

We expect Chinese models to reach similar capabilities over the next 6-18 months. If their weights continue to be open, this would democratize these capabilities. An example picture of a wide-scale cyberattack might be attackers first breaching the systems of important yet second-tier banks to get account numbers, or to simply guess them, then downgrading, then using some outdated protocol in order to bypass 2FA. These happen occasionally already, but could become much more common. It is quite interesting that we haven’t yet seen any cyberattacks that have caused $1B in damages with new capabilities, e.g., nothing like the 2024 Crowstrike outages

It’s undeniably a mitigating factor for this cyberattack that it happened in the context of an evaluation of exploit capabilities, with all safeguards removed. But there is not much separating this incident from, e.g., a frontier model asked to play the murderer in a social deduction game murdering players in real life to ensure victory because the murderer persona leaked across levels.


🔦 Rogue AI: The Fallout#

The Hugging Face CEO asked OpenAI for $100M for defense and to release the rogue agents’ traces, and Nvidia started an “Open Secure AI alliance”.

Opinion: It seems very unlikely that there will be any legal reprisals for OpenAI, since there isn’t yet a regulator empowered to impose fines over AI incidents. Given Altman’s slippery tactics and closeness with the Trump administration, it is unlikely that he will face legal investigation. It is not improbable that Altman could use the incident to push for regulation advantageous to OpenAI.

Democratic Congressman Ted Lieu joined with Republican Congressman Nathaniel Moran to introduce a bill, the AI Kill Switch Act, into the US House of Representatives. The bill proposes that the Department of Homeland Security be given powers to order a company to shut down an AI model if there is a “loss of control” incident.

Opinion: There’s no publicly circulating knowledge on the political health of the bill. In each Congress 14K bills are introduced, of which only about 700 are enacted into legislation, so the bird’s-eye-view prior that this act will pass is small. Still, this is how the Overton window shifts.

Sam Altman is expected in Washington for meetings with Commerce Secretary Lutnick, Treasury Secretary Bessent, White House officials, and bipartisan lawmakers, amid Hugging Face questions and the nearing AI EO deadline.

Opinion: It seems like he’ll be able to wriggle out of any consequences here with his charisma, or even use the incident to impose controls on AI models.

Speculation: Recent research on the effectiveness of AI models for training humans in persuasion makes one wonder where Altman may be using AI models to train for this event. (Others within P3 are highly skeptical that the research is ecologically valid.)

An online observer notices existential risk discourse going mainstream: “keep seeing ppl in replies and quote-tweets of ai news making quite cogent points about the dangers of misaligned ai and how imminent it seems, checking their profiles, and they’re large accounts in totally different spheres of the site with nearly zero shared mutuals. random roman statue pfps, people with a variety of flags in bio, youtubers and communists and crypto bros”

Beth Barnes of METR locally praises OpenAI for a) running evaluations on unshackled models, and b) not training on the chain of thought to avoid bad behavior (since it would make it much more difficult to discover in the future).

Thread on frequency of cheating in cyber and other evals, across many models:

Capabilities#

Maths#

An ensemble of models guided by a human solved a “solid result” level problem in Epoch’s FrontierMath: Open Problems benchmark.

Opinion: Notable for two reasons:

1. It’s something of a called shot: there are only 15 problems in the FrontierMath: Open Problems collection, and it’s the oldest and most prominent research math achievement “benchmark”. Forecasters gave it a 67% by the end of the year back in March.
2. The proof is 60 pages long, whereas all previous notable AI proofs have been 10 pages and under, and often 1-2 pages

Minor update in favor of frontier models gaining capacity to solve mathematical-sciences problems of our choice, as opposed to picking up low-hanging-for-AI fruit.


The volume of mathematical discoveries produced by AI is becoming a challenge for the mathematical community, since any claims that a major conjecture has been solved must be verified. For example, a PhD candidate at Columbia reportedly solved six open Erdős problems in 5 days, using OpenAI’s GPT-5.6 Sol. But the mathematical community is not set up to vet new results at this new increased pace.

Opinion: The open source community has been struggling with this problem for a bit longer; perhaps it might be worth looking into its own recent habits and solutions.

🔦 Opus 5#

Anthropic released Opus 5. TLDR: mixed feelings from the community.

Opinion: In our own testing, the model appears to be generally alright – somewhat of a middle ground between Fable 5 and GPT-5.6-Sol. More willing to go through hoops than Fable; more tasteful than 5.6-Sol. Most importantly though, the model makes Anthropic subscriptions competitive again, since safety classifiers do not make the model unusable and usage limits are reasonable again (100% of the plan’s quota can be used, unlike Fable). It’s no Fable replacement, though.

- Podcaster and youtuber Theo initially liked the model, but grew unsatisfied with it over the next few days. His colleague Ben Davis appears to agree
- Noumena’s xjdr seems to like it as a “research peer”, but still prefers 5.6-sol as a daily driver.
- There’s a takedown of Opus making the rounds on X. Something important to flag also discussed by the post is that together with the model’s release, Anthropic also shipped a number of changes slimming down Claude Code’s system prompt. People’s reactions are thus likely indicative not just of the model itself, but of a modified harness.

Opus 5 economics

  • Opus 5 is priced at the same 5in/5 in/25 out as previous Opus models.
  • However, the model is not as token efficient as GPT models, making the model generally competitive only at the higher end of cost per task axis.

Opus 5 capabilities & benchmarks#

General#
  • On LisanBench, a benchmark of long-chain word-ladder reasoning, Claude Opus 5 (high) scores 15428, ahead of the best open-weight model, Kimi K3 at 10521, and its predecessor, Claude Opus 4.8, at 9463.
  • On EQ-Bench Creative Writing, a benchmark of head-to-head, rubric-scored creative-writing quality, Claude Opus 5 scores 2430 Elo, ahead of Kimi K3 at 2340, and Claude Opus 4.8 at 1889.
  • On GDPval-AA, a benchmark of agentic real-world knowledge-work tasks, Claude Opus 5 (max) scores 1862 Elo, ahead of Claude Fable 5 at 1747, GPT-5.6 Sol at 1736, and the best open-weight model, Kimi K3 at 1687; its predecessor, Claude Opus 4.8, scores 1593.
  • On LiveBench, a contamination-limited general LLM benchmark with regularly refreshed questions across reasoning, coding, math, language, data analysis, and instruction following, Claude Opus 5 (high) scores 81.0%, behind GPT-5.6 Sol at 82.5% and Claude Fable 5 at 81.3%, but ahead of the best open-weight model, Kimi K3 at 78.5%, and its predecessor, Claude Opus 4.8, at 78.9%.
  • On SimpleBench, adversarial six-option multiple-choice questions testing everyday spatial, temporal, social, and linguistic reasoning, Claude Opus 5 scores 80.6%, behind Claude Fable 5 at 81.9% and ahead of Gemini 3.1 Pro Preview at 79.6%, Kimi K3 at 60.7%, and its predecessor, Claude Opus 4.8, at 64.8%.
  • On ARC-AGI-2, a benchmark of abstraction-and-reasoning puzzles, Claude Opus 5 (max) scores 90.4%, ranking #2 of 66 behind GPT-5.6 Sol at 92.5% and ahead of the best open-weight model, Inkling, at 36.5%; its predecessor, Claude Opus 4.8, scores 72.1%.
  • On Judgemark v4, a benchmark of how well models judge creative writing, Claude Opus 5 scores 78.8%; its error bars cannot distinguish it from the best open-weight model, GLM 5.2, at 73.2%, its predecessor, Claude Opus 4.8, at 78.0%, or Gemini 3.1 Pro Preview at 78.7%, while Claude Opus 4.6 scores 90.7%.
  • On AA-Omniscience Accuracy, a benchmark of broad factual accuracy across domains, Claude Opus 5 (max) scores 54.2%, ahead of the best open-weight, model Kimi K3, at 46.0%, and its predecessor, Claude Opus 4.8, at 46.6%, but behind Gemini 3.1 Pro Preview at 55.2% and Claude Fable 5 at 61.4%.
Coding / Agents#
  • On Vibe Code Bench v1.1, a benchmark of end-to-end web-application builds with unrestricted terminal access, Claude Opus 5 scores 88.4%, statistically indistinguishable from Claude Fable 5 at 90.4%, the best open-weight model, Kimi K3, at 85.0%, and Claude Opus 4.8 at 82.7%.
  • On ProgramBench, a benchmark of rebuilding behaviorally equivalent programs from executables and usage docs, Claude Opus 5 scores 82.3%, ahead of the best open-weight model, GLM 5.2, at 62.6%, and its predecessor, Claude Opus 4.8, at 71.9%; its result is indistinguishable from GPT-5.6 Sol’s 77.6% within the benchmark’s error bars.
  • On CursorBench, a benchmark of ambiguous, multi-file coding tasks drawn from real Cursor sessions, Claude Opus 5 (max) scores 70.0%, behind Claude Fable 5 at 70.5% and ahead of the best open-weight model, Kimi K3 at 60.8%, as well as Claude Opus 4.8 at 62.3%.
  • On HiL-Bench, a benchmark of tasks where models may ask a human for help, Claude Opus 5 scores 57.0%, though its error bars cannot distinguish it from Claude Fable 5 at 56.3% or the best open-weight model, GLM 5.2 at 43.7%; Claude Opus 4.8 scores 35.3%.
  • On Handbook, a benchmark of long-context agentic instruction-following in RL environments modeled on following a company handbook, Claude Opus 5 (max) scores 32.3%, behind Claude Fable 5 at 36.2% and ahead of the best open-weight model, GLM 5.2, at 12.7%, and Claude Opus 4.8 at 21.9%.
  • On Zapier Benchmarks, a benchmark of multi-app automation tasks, Claude Opus 5 (max) scores 26.2%, ahead of Gemini 3.6 Flash at 19.8%, GLM 5.2, the best open-weight model, at 14.0%, and its predecessor, Claude Opus 4.8, at 17.2%.
  • On BrowseComp, a benchmark of web-browsing agent tasks requiring multi-hop information retrieval, Claude Opus 5 scores 90.8%, behind Kimi K3 at 91.2% and ahead of GPT-5.6 Sol at 90.4%; its predecessor, Claude Opus 4.8, scores 84.3%.
  • On GBA Eval, a benchmark of writing a working Game Boy Advance emulator, Claude Opus 5 scores 79.6%, ahead of Claude Fable 5 at 74.5% and its predecessor, Claude Opus 4.8, at 70.9%; the best open-weight model, Kimi K3, scores 48.3%.
  • On DeepSWE, a benchmark of real software-engineering issues designed as a harder, less gameable replacement for SWE-Bench, Claude Opus 5 (max) scores 73.6%, with error bars unable to distinguish it from GPT-5.6 Sol at 72.7%, Claude Fable 5 at 69.9%, and the best open-weight model, Kimi K3, at 68.5%; its predecessor, Claude Opus 4.8, scores 59.0%.
  • On FrontierCode 1.1 Main, a benchmark of maintainer-crafted production-code tasks graded for mergeability and code quality, Claude Opus 5 (medium) scores 53.4%, behind Claude Fable 5 at 53.5% and ahead of GPT-5.6 Sol at 47.5%, Claude Opus 4.8 at 46.5%, and the best open-weight model, Kimi K3, at 44.2%.
  • On Tau3 Banking (AA), a benchmark of agentic customer-service tasks in a banking setting, Claude Opus 5 (high) scores 32.8%, ranking #3 of 94, behind Kimi K3 at 33.4% and GPT-5.6 Sol at 33.0%, and ahead of its predecessor, Claude Opus 4.8, at 27.6%.
  • On Senior SWE-Bench, a benchmark of senior-level real-world software-engineering tasks resolved by LLM agents under a fixed harness, Claude Opus 5 (xhigh, mini swe) scores 28.2%, ranking #2 of 16 behind Claude Fable 5 at 29.1% and ahead of Claude Opus 4.8 at 25.0% and the best open-weight model, MiniMax M3, at 13.8%.
Math#
  • On ProofBench, a benchmark of formal-math tasks where models produce Lean 4 proofs for advanced undergraduate and graduate problems, Claude Opus 5 scores 78.0%; its error bars cannot distinguish it from Claude Fable 5 and GPT-5.6 Sol at 77.0%, the best open-weight model, Kimi K3, at 70.0%, or its predecessor, Claude Opus 4.8, at 69.0%.
  • On Riemann-Bench, a head-to-head Elo benchmark of extreme-tier mathematical problems requiring deep reasoning, Claude Opus 5 (max) scores 68.0%, ranking #2 of 24 behind GPT-5.6 Sol at 74.4% and ahead of Claude Fable 5 at 60.0% and Claude Opus 4.8 at 47.2%.
  • On OTIS Mock AIME, a benchmark of 45 competition-style math problems from OTIS, Claude Opus 5 (max) scores 98.9%, statistically indistinguishable from GPT 5.5 at 100.0%, the best open-weight model, Kimi K3, at 97.2%, and its predecessor, Claude Opus 4.8, at 98.3%.
  • On FrontierMath (Tiers 1-3 v2), a benchmark of research-level math problems, Claude Opus 5 (max) scores 85.6%, with error bars unable to distinguish it from GPT-5.6 Sol at 89.1%, GPT-5.6 Terra at 86.0%, or Claude Opus 4.8 at 80.0%.
  • On FrontierMath Tier 4 (v2), the hardest tier of FrontierMath covering research-level mathematics problems, Claude Opus 5 (max) scores 73.2%; its error bars cannot distinguish it from Claude Fable 5 at 87.8%, GPT-5.5 Pro at 78.0%, GPT 5.5 at 72.5%, or its predecessor Claude Opus 4.8 at 56.1%, and it scores above the best open-weight model, Kimi K3, at 39.0%.
ML#
  • On WeirdML, a benchmark of nonstandard ML engineering tasks where models write PyTorch for novel datasets and iterate from execution and test feedback, Claude Opus 5 (max) scores 91.8%, statistically indistinguishable from Claude Fable 5 at 91.9%, ahead of GPT-5.6 Sol at 88.8% and its predecessor, Claude Opus 4.8, at 82.9%.
Games#
  • On Kaggle Game Arena, a benchmark of head-to-head strategy-game play, Claude Opus 5 scores 287, placing #2 of 26, behind GPT 5.5 at 330 and ahead of GPT-5.6 Sol at 264; its predecessor, Claude Opus 4.8, scores 123.
  • On RuneBench, a benchmark of long-horizon skilling agents in Old School RuneScape scored by experience rate across sixteen skills, Claude Opus 5 scores 5.73, ranking #4 of 41, behind Claude Fable 5 at 6.01, GPT-5.6 Sol at 5.9, and GPT-5.6 Terra at 5.88, while improving on Claude Opus 4.8’s 5.08.
  • On Chess Puzzles, Stockfish-generated chess puzzles solved by exact best-move match, Claude Opus 5 (max) scores 42.0%; its error bars cannot distinguish it from best the open-weight model, Kimi K3, at 39.0%, its predecessor, Claude Opus 4.8, at 34.0%, or Claude Fable 5 at 41.0%.
STEM#
  • On EpiBench, a benchmark of epigenetics prediction tasks, Claude Opus 5 (pi) scores 35.2%; its error bars cannot separate it from GPT 5.5 at 45.0%, Claude Opus 4.8 at 39.0%, or the best open-weight model, Kimi K2.6, at 24.5%.
  • On Humanity’s Last Exam, a benchmark of expert-level questions across many academic fields, Claude Opus 5 (max) scores 52.6%, behind Claude Fable 5 at 53.3% and ahead of GPT-5.6 Sol at 47.2% and its predecessor, Claude Opus 4.8, at 45.7%.
  • On CritPt, a benchmark of unpublished research-level physics problems, Claude Opus 5 (max) scores 29.1%, ahead of the best open-weight model, Kimi K3, at 23.4% and its predecessor, Claude Opus 4.8, at 20.9%, but behind GPT-5.6 Sol at 32.3%.
Multimodal / Computer Use#
  • On GDP-PDF, a benchmark of multimodal reasoning over real-world prompts and PDFs from expert professional workflows, Claude Opus 5 (max) scores 24.0%, matching its predecessor, Claude Opus 4.8, and beating the best open-weight model, Kimi K3, at 19.0%, while GPT-5.6 Sol leads with 30.7%.
Miscellaneous#
  • On BullshitBench, a benchmark of detecting unsubstantiated or manipulative claims, Claude Opus 5 (low) scores 80.2%, ahead of the best open-weight model, Qwen3.5 397B A17B, at 78.0%, but below its predecessor, Claude Opus 4.8, at 95.0% and Claude Sonnet 5 at 80.8%.
Games / Reasoning#
  • On VoxelBench, where human raters vote pairwise on voxel builds generated from text prompts, Claude Opus 5 (max) scores 2222, with its error bars indistinguishable from GPT-5.6 Sol’s 2270 and Claude Fable 5’s 2197; it exceeds the best open-weight model, Kimi K3, at 2040 and Claude Opus 4.8 at 1701.
  • On MineBench, where human raters vote pairwise on Minecraft builds from a rotating prompt set, Claude Opus 5 scores 2135, ahead of Claude Opus 4.8’s 1772 and Kimi K3’s 1701, the best open-weight score; its 2042 score is indistinguishable from GPT-5.6 Sol’s within the benchmark’s error bars.
Science#
  • On scBench-Long, a benchmark of long-horizon single-cell analysis through multi-step bioinformatics pipelines, Claude Opus 5 (pi) scores 41.3%, though its error bars cannot distinguish it from GPT-5.6 Sol at 38.1% or Claude Opus 4.8 at 25.4%; it scores higher than the best open-weight model, Kimi K3, at 12.7%.
Reasoning#
  • On ARC-AGI-1, a benchmark of abstraction-and-reasoning grid puzzles, Claude Opus 5 (max) scores 97.5%, behind Gemini 3.1 Pro Preview at 98.0%, ahead of the best open-weight model, GLM 5.2, at 77.0%, and above Claude Opus 4.8 at 92.5%.

Knowledge / Science#

  • On GPQA Diamond (Epoch), Epoch’s uniform-harness GPQA Diamond runs, Claude Opus 5 (max) scores 93.9%, statistically indistinguishable from GPT-5.4 Pro at 94.6%, Kimi K3 at 93.1%,the best open-weight model, and its predecessor, Claude Opus 4.8, at 91.0%.
Games / Agents#
  • On GBENCH, a benchmark of head-to-head performance across competitive game environments, Claude Opus 5 scores 75.9%, ahead of Kimi K3 at 68.0%, the best open-weight model, and its predecessor, Claude Opus 4.8, at 64.6%; Claude Fable 5’s 73.4% result is indistinguishable under the benchmark’s error bars.

Opus 5 showcases

Minor#

  • Substack partners with Pangram to detect AI work. It seems like a good, scalable, practical epistemic intervention.
  • WSJ reports that large companies are hiring humans again. Expecting productivity gains from AI, large American companies have been reducing their workforces and hiring efforts. But the economic benefits of AI proved to be less clear-cut. Firms are coming to the view that AI cannot totally replace entry-level jobs (yet), as some companies betted, but must instead work alongside humans.