This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.

TL;DR

  • Two reports out on OpenAI’s July attack on HuggingFace. Hard to summarize: not good.
  • Dwarkesh Patel and Dylan Patel make bold predictions about the next 5 years of AI economics.
  • Yet more churn in the contest over future AI regulations, with no consensus in sight.
  • In a study of 8 countries, AI use is overwhelming government services, at least in the short run.
  • Non-profit auditing group AVERI conducts the first double-blind evaluation of an AI model

Economics#

Anthropic will likely tell investors to expect potential revenue opportunities of above $30T, reports the WSJ. The number is an estimation of its “total addressable markets” (TAM), the amount of annual revenue a company could make if it captured 100% of the market-share (and speculated future market-share). Anthropic’s TAM would eclipse SpaceX’s record-breaking $28.5T: “To put its more than $30 trillion vision in context, the 191 technology companies in the S&P 1500 brought in $2.4 trillion in revenue last year, according to FactSet.” Anthropic could raise $100B, compared to SpaceX’s $86B. Anthropic is looking for a $2T valuation, compared to SpaceX’s $1.77T valuation. Anthropic’s Q2 revenue was $11.6B, roughly double the previous quarter. Its financial disclosure documents are expected to be released in the coming weeks, suggesting Anthropic may “go public as soon as September or early October.”

Opinion: We know the superintelligence-pilled (which includes Anthropic leadership) already believe it to be bigger than this, so this just seems like a marketing exercise for the IPO (give a bigger number than the other guy.) The similarity of the numbers is probably not a coincidence.

Some signs of an intense negative reaction to the number, along conspiratorial lines.


Nvidia to buy Hugging Face for $13B.

Opinion: Unusual move by Nvidia into the distribution layer of the inference stack. Subsidising the open weights ecosystem has benefits for Nvidia’s lock in on the hardware side, since this system overwhelmingly runs on Nvidia hardware.

It is a cheap play relative to Nvidia’s size and may have been defensive – reporting suggests that it only got interested after Hugging Face got other acquisition interest (perhaps from a hyperscaler).

Speculatively: Nvidia has acquired the option to have HF press charges against OpenAI for its attack. The threat will probably remain implicit, but the statute of limitations in the US for civil proceedings is two years.


OpenAI releases a statement on the details of Broadcom’s and OpenAI’s new Jalapeño inference chip. The chip reportedly delivers ~1.7× more work per watt than Nvidia’s previous-gen GB200s and GB300s. This is a step towards OAI increasing its control over its own stack.

Opinion: There was no significant impact on Nvidia or Broadcom stock from this, so it was priced in already. Comparing Jalapeño (which uses the new HBM4) to chips like Blackwell and TPUv7 (which use the older HBM3e) is a somewhat unfair comparison. Rubin and the next-gen TPUs which will deploy at scale at the same time as Jalapeño would be a fairer fight.

It is also worth noting that this chip is only usable for inference rather than training, where Nvidia still dominates.

Overall, this mostly serves to somewhat weaken Nvidia’s long-term leverage and market dominance (a theme we discussed last week). We don’t expect the chip to form more than 5-10% of OpenAI compute before 2028, at the earliest. It also serves to shrink or move Nvidia’s moat – if a chip this close to SOTA can be designed and built with significant help from current AI models, Nvidia’s design expertise looks weaker as a moat, and its true moat becomes relationships with suppliers which allow it to build at scale and the liquidity of the market for Nvidia chips, which makes financing large purchases easier.


AI use is overwhelming government services, warns a recent paper. AI agents and LLMs are lowering the barriers to interacting with the government, “flooding” public services with complaints and claims faster than the rate at which they can be processed and addressed. The paper provides evidence that flooding is widespread, and this is before agents really catch on: 87% of cases studied were inferred to be pasted LLM text, rather than an agent autonomously navigating a web portal.

It is suggested that “near-term risk is highest for financially attractive, but complex services*.*” Notably, the authors also argue that increasing “friction” – for example, implementing fees for filing a claim – would harden pre-existing inequalities in access. They conclude, with many qualifications, that government use of AI may be the most plausible solution to the problem.

Opinion: Though it seems to be overwhelming some services in the short run, it’s partly evidence of a positive long-term shift; various citizens who were otherwise unable to avail themselves of what was on offer by their government will now be able to do so. The paper is right that quick fixes like suppression, whether by limiting AI assistance or raising the cost of submission, are the wrong approach. This will unfairly hit the poor, time-constrained, or digitally illiterate. Softer barriers like word counts might be in order to limit the excessively long submissions.

Ultimately a long-term approach might look something akin to copying the citizens: integrating AI on the response end, which the paper points out. This could help cut through the higher volume of submissions and detect instances of fraud.

Opinion (Nuño): Making government benefits easier to access through AI models also shifts the equilibrium to one where people take more government benefits, which increases government expenditures relative to the baseline. Even if the government eventually realizes this, and chooses to reduce government benefits to account for this, the process would be slow. Whether this whole dynamic is positive or not is unclear, and will depend on one’s values and political inclinations.


Huawei is bidding to build government AI infrastructure in Egypt, says Bloomberg. The deal would involve the purchase of Huawei’s 1,408 advanced Ascend 950-series chips for an AI training cloud and 600 chips for two inference clusters. Huawei is thus now targeting Africa’s second-largest economy and will be looking to expand its business into the Middle East, while the US has been throwing money into AI infrastructure in the Gulf States.

The US is encouraging Nvidia, AMD, and Microsoft to mount rival bids. A State Department representative has warned that using Huawei accelerators may be subject to legal penalties. And the US does have some leverage here: chip shipments to Egypt are currently subject to licensing conditions.

Opinion: Consider Pax Silica, led by the US

as opposed to the World Artificial Intelligence Cooperation Organization, led by China.

With time, we will see these diffusion skirmishes multiply.

“Legal penalties” is also a misleading framing, since it bakes in the assumption that Egypt is or should be subject to US laws. Moreover, countries can choose to suffer US sanctions if this is a better deal. Perhaps the better framing is whether the US offers or doesn’t offer a better bargain than China.


The Securities Exchange Commission (SEC) investigates Situational Awareness, reports the NYT. After Leopold Aschenbrenner’s AI-focused hedge fund almost imploded last month, the SEC has sent “subpoenas to banks that handled the hedge fund’s calamitous trading and that fed it borrowed money to supersize its bets.” The subpoenas request information related to details “on the timing of Situational Awareness’s trades and for its communications with lenders about the money it was borrowing,” according to three anonymous sources. In its statement to the NYT, Situation Awareness has downplayed the investigation and dutifully insists that it is “a highly regulated business and will cooperate to the fullest extent with any regulatory request. “

Opinion: What comes of this really depends on what the SEC is looking for and how vindictive they are: if it wants to find something wrong it seems very likely there will be compliance issues - e.g. failure to follow some disclosure or record-keeping standard. Penalties for that could range from a slap on the wrist to being barred from managing outside assets, depending on the severity.

The main substantial public accusation which has been laid against Situational Awareness has been of insider trading (due to Leopold’s wife being the Anthropic CEO’s Chief of Staff and Situational Awareness making many Anthropic-themed well timed investments.) But this may not even be in scope for the investigation and would likely be hard to prove if it was.


A Time article notes that OAI is looking to rehabilitate its image by once again promoting itself as a safety-focussed lab. After the Hugging Face incident, and new concerns around their unreleased Astra model, OAI decided to pace model development (for two weeks). There is reason to be skeptical: many safety researchers have left the company and a significant amount of capital and other resources have been redirected toward commercial interests. The author also address a critical issue:

If OpenAI can reclaim the mantle of the safety-first lab, it might bolster its image while forcing its main competitor to answer an uncomfortable question as it plans a blockbuster IPO: Will Anthropic keep racing while OpenAI waits?

Opinion (Nuño): The goal to “rehabilitate OpenAI’s image” is different from the goal of “give information to external observers around making costly principled choices”; the former rhymes with propaganda and used car salesmen, whereas the latter would be the honest version. But outside observers already have a good amount of information already about what kind of decision-making process Altman and the new execs are running.


In July, a panel of superforecasters and experts predicted that Anthropic and OpenAI would have a combined annualized revenue run rate of $90bn in December 2026. The current figure, six weeks later, is already >$105bn.

Opinion: The prediction is not strictly falsified yet, since there could be a massive collapse in demand by year end. But we take this as yet more evidence that the “superforecaster” brand (i.e. the top ~2% of people willing to enter prediction contests which give very low rewards) is not very useful. Filtering predictors based on AI-predictions track record in particular is likely a better selection criterion when looking to crowd-source AI predictions.

🔦 Dwarkesh Patel’s third interview with Dylan Patel of Semianalysis#

As usual, we enjoyed this, and appreciate their internal consistency and willingness to draw big thick straight lines through a small cloud of data (just as some people managed to roughly extrapolate deep learning progress using a simple function of compute and algorithms).

Their predictions#

Prediction #1. “Anthropic & OpenAI will have most of the world’s compute by 2028”

OpenAI and Anthropic started 2026 at ~2 GW and will end it above 5 GW, taking ~30% of the world’s added compute this year; that rises to 40-50% in 2027, with half of all new compute going to these two companies by December 2027.

By end-2028 they take 70-80% of incremental compute and, since new chips are 3–5x more perf/watt than older generations, control most of the world’s usable FLOPs. ~100 GW-combined.

This part of the forecast is a Dwarkesh extrapolation which is basically ungrounded — Dylan calls it “very aggressive” and conditions it on labs paying $25-50M per MW and on Google (etc) being willing to sell to them. In his defense, he notes that the market rate for compute is presently $10-15M per MW while Anthropic’s revenue has hit $50M/MW. So assuming that this ratio is stable, perhaps they would be able to pay the extreme cost.

Opinion: This is well grounded up until 2027, with the 2028 numbers basically simple trend extrapolations. It seems quite contingent on the trajectory of capabilities and demand - if OAI/Ant retain significant market power and high margins on their models at that stage, and demand is truly insatiable, then it’s possible, but there are strong reasons to think their share will cap out earlier than this, via more compute going to widespread use of non-frontier models as capabilities saturate on some tasks, and preferences from providers to maintain customer diversity.


Prediction #2. He predicts “internalisation of inference”: labs will withdraw inference from the market, because the marginal product of tokens inside labs will beat the external willingness-to-pay of users. He guesses that the current lab budget distribution is ~50% research, ~10% development, ~40% inference

Labs will allocate a shrinking fraction of compute to inference; most compute goes to forward passes for research/training.

“Already started”: Anthropic’s monthly compute kept growing over recent months even while its revenue growth plateaued.

Opinion: This is true for Anthropic in recent months, but it is more a reflection of weaker demand growth in recent months (see our most recent newsletter) than the active Anthropic decision presented here. Anthropic saw massive demand growth in March/April which likely forced them to redirect some compute to inference from training, so assigning incremental compute to training also may just bring this back to where it was.

How would this practically play out? Likely by demanding higher prices for inference, which we see from Anthropic with Fable but not from OpenAI, which recently cut API prices for GPT 5.6 despite being in the middle of a boom in revenue growth. This would lead to increased margins, but only if the willingness to pay is there. Public market pressures will make operating at a large loss due to claimed internal value of tokens a little more difficult, but that is one reason for using dual class stock, as Anthropic plan to.

Ultimately this trajectory is predicated on continued rapid capabilities gains and high marginal usefulness of training compute.


  1. China deploys <10% of new watts now and holds ≤30 GW in 2028

They’re further assuming a 2.5x quality penalty on Chinese chips, such that 30GW of Huawei is essentially 12GW of NVIDIA.

Opinion: This seems about right, though it may not fully account for smuggled or cloud-leased Nvidia chips.


  1. The AI buildout might induce a global credit squeeze

Hyperscaler debt issuance pushes credit spreads up: Dylan claims Meta’s effective interest rate has gone from ~5% to ~8%, though caveats this as “extremely vibed out”
.
~$1T/yr of new credit against a ~$130T global bond stock and ~$25T/yr of gross capital formation drives borrowing costs up 2.5pp, squeezing banks and cratering long-duration equities??

Dwarkesh predicts a Volcker shock 2: highly indebted, short-duration countries like Pakistan and Nigeria will default because of US AI

Dwarkesh speculates 2030s interest rates >10% if the economy approaches yearly doubling

Opinion: This is a bit loose, but granting the premise — if the economy does approach yearly doubling, then interest rates of 10% would be a steal, and capital would be pulled in from absolutely everywhere to invest in the factors causing the boom. Non-AI investments would have a very difficult time competing, and so economies without a seat at the AI table would need to figure out how to get one or suffer.


  1. Predicts that the amount of effective AI labor will exceed human effective labor by 2030

effective frontier AI population goes ~10x/yr, so a lab goes ~10M → 100M → 1B labour-equivalents over 2026–28, and a single lab exceeds Earth’s human labour supply by 2030

Opinion: Here we enter the realm of pure speculation — it’s hard to define the amount of current AI labour being used, and conversion from AI-to-human labour is difficult. If most of the compensation is still flowing to humans, does that mean there is still more human labour? If not, how can one measure it?


  1. “governments will soon restrict labs’ internal use of frontier models,

Dylan: soon not just external release being blocked. Dylan’s case for slow takeoff rests on politics, rates, and regulation, not on capability scaling

Opinion: This is one way things could go, and seems more plausible in the wake of the OpenAI/HuggingFace incident, but it is subject to the choices of government actors and their visibility on internal lab deployment risks, as well as the rate of progress itself.


An excellent critique by Philippe Lemoine points out that the podcast’s model only uses supply-side factors. In particular, Lemoine argues that

(i) diffusion will be slow due to normal frictions around making changes to business processes

(ii) even allowing for fully capable drop-in remote workers, demand for certain kinds of services will be sated at a point much lower than required for explosive growth to happen — at some point you’ve got all the legal, managerial, engineering etc services that you want, even if prices fall dramatically.

(iii) while new areas of demand will be created, it will take years for humans to even figure out that they want these, demand will not arise simultaneously.

Dwarkesh responds, arguing that

(i) drop-in remote worker diffusion will be fast, as frictions will be much lower than for normal hiring – no market-for-lemons problem, and rapid onboarding being the main mechanisms here

(ii) lab pricing power will be maintained if RSI progress is sufficiently fast, and the frontier remains more capable than the fast-followers.

Opinion: Lemoine’s arguments bite most against the “100%+ yoy growth”, “Anthropic/OAI eat the world” perspective. They do not bite particularly hard against claims like “trillions in revenue”, i.e. a low single digit percentage of GDP from AI.

Dwarkesh does not present a really good answer to Lemoine’s point (ii)+(iii), and this is an interesting area to explore. It seems like the main way to avoid (iii) is “the AIs do it all by themselves” and the economic growth is at least initially centred on a very aggressive compute buildout, analogous to the 19th century US railway boom.

But to not lead to a crash or human disempowerment, ultimately this must be put to non-recursive use and cash out in things that are not just “more compute”. The definition of a desirable outcome is that the accumulated capital must generate something with terminal value.

Capabilities#

Another GPT-3 moment for robotics?: Skild AI releases S1, a robot that can supposedly learn tasks from a single video with no fine tuning. Similar to Generalist AI’s model (covered in a previous newsletter), training robots for even a single task no longer requires hours of teleoperation data. Skild’s internal benchmarks suggest a 66% per step success rate for out-of-distribution tasks (with human recovery between failed steps), compared to 9% for language-prompted models. A single demonstration roughly equates to 380 post-training episodes. Skild also claims the model has the capacity for self-correction, adapting to scene changes, and executing a task more competently than depicted in a flawed demo.

Opinion: The progress in AI robotics in the last 1-2 months seems real and significant, though it’s far from scientific evidence, since everything in commercial robotics is both closed and cherry-picked.

It remains difficult to gauge where we are on the pipeline from conceptually significant progress (“one-shot learning in robotics now sometimes works in the lab’’) to any kind of product revolution. The public epistemic infrastructure for presenting and evaluating progress in AI robotics is currently very weak.


Anthropic experiments with giving researchers aggregate data from ~250k real-world Claude conversations via a privacy-preserving approach. Anthropic worked with three research groups, each designing its own study. The various findings included the extent to which users delegate consequential tasks to Claude (those which cannot be easily undone or which affect others), how users’ moods track with Claude’s responses, and productivity gains across model generations.

Opinion: An exciting opportunity in principle. The “societal impacts” side of Anthropic has always leaned towards PR-compatible blandness. This new independent work doesn’t feel like a sharp break from that tradition.

Anthropic’s collaboration protocol protects the academic integrity of the studies on a per-study level (no Anthropic oversight of the resulting work beyond a scoped accuracy review), but access-journalism type dynamics inevitably give Anthropic some power to set the tone and pick favourites.


Along with others, non-profit auditing group AVERI conducts the first double-blind evaluation of an AI model. Working with DeepMind, among others, the study prevented labs from viewing the prompts and the auditors from viewing the model weights. The double-blind test relied on cryptographic verification, bypassing the objection that AI audits must require the disclosure of IP or benchmark solutions that end up in the training data for future models.

Opinion: Very interesting technical scheme. Designing, implementing, and enforcing protocols of this kind will be crucial as we (hopefully very quickly) enter the era of legally consequential safety and reliability certifications. Most pre-release capabilities evaluations today rely on something like “handshake deals” between AI lab scientist and eval scientists, in line with the common assumption of good faith in interactions between scientists in general, and this may remain the case going forward. But safety and reliability evals in particular are prone to become adversarial interactions between labs and auditors, and will soon be the loci of tremendous economic pressures.


UK’s AISI releases a method to save >50% on the cost of running a certain type of evaluation, sometimes as much as 97%. It’s a well-founded Bayesian adaptive stopping framework. As a result, we can now allocate “LLM evaluation compute … by uncertainty rather than by fixed repetition counts.” A replication by a Prime Intellect researcher finds 20-70% savings.

Opinion: Clever. This helps with repetition runs (e.g. running the same suite 8-32 times to calculate an average performance – “avg@8”), but these are very common.


Toby Ord’s new paper, “The Dynamics of Intelligence Explosions,” examines the mathematical conditions under which RSI could trigger an “intelligence explosion.” Ord concludes that the requirements are surprisingly narrow, and that under many conditions the results of RSI-style dynamics would fall short of an intelligence explosion as typically understood.

Opinion: A potentially important contribution to a growing line of research and commentary on “RSI as a normal technology”: models of automated AI R&D where the RSI process and its consequences remain legible to ordinary quantitative social science and tractable to human management and intervention. We think developing formal and informal models along these lines is very valuable, since the prospect of RSI-unto-AGI (let alone RSI-unto-AGI-unto-ASI) remains uncertain but AI R&D boosts from AI are definitely real and definitely growing.


Politics#

The push-and-pull between different visions for the (immediate) future of AI regulation continues apace, with no signs of a clear winner emerging:

The White House executive order proposing a self-regulating organization (SRO) for frontier labs has stalled, according to The Information. For context, earlier this year an executive order suggesting that frontier labs could voluntarily submit models for government review was scrapped after the accelerationist camp, specifically a call from David Sacks, convinced the government that the order would give China a competitive edge. In consequence, a softer version was signed. The next stage was meant to be the creation of an SRO, which has stalled as AI labs come into conflict with Washington, and each other, over how binding the framework should be, who should hold authority, and the legislative power of states’ involvement. According to The Information, Sacks and others in the accelerationist camp continue to find recent proposals too regulation-heavy.

In an informal response, Tyler Cowen cautions against excessive regulation of AI while endorsing a Financial Industry Regulatory Authority (FINRA)-stye approach of just the kind reportedly killed off by Sacks. Like Demis Hassabis, Cowen suggests that the ideal regulatory framework should be inspired by the Financial Industry Regulatory Authority (FINRA): “a consortium of financial firms that examines the trade practices of each and makes recommendations, helping the federal Securities and Exchange Commission with oversight and regulation.” The body would periodically audit frontier labs to check for safety concerns. Passing an audit would free “AI labs from the fear that courts might derail their business by granting huge awards to plaintiffs for ill-defined harms that could not reasonably have been prevented.” It would also provide an incentive to meet the relevant safety standards.

Cowen considers the alternatives. The first is the libertarian case against regulation, which he describes as an “illusion”: AI is already subject to licensing laws, and courts lack the necessary expertise in AI to make sound and informed decisions. The status quo also relies on secret and discretionary (and often arbitrary and, in the case of Anthropic, ignored) decisions by The Pentagon and national security agencies. The second option is to set up a body akin to the FDA, dismissed by Cowen as unsuitable for the accelerating AI industry because of its sluggishness (reviews, trials, and approvals often taking years).

Opinion: We’re strongly skeptical of AI-labs-led oversight of the AI industry, and therefore largely welcome the stalling of a FINRA-style conglomeration. One of our main areas of concern at Paradigm 3 is the accumulation and acceleration of damage to society from misused / unreliable / rogue powerful AI below the level of AGI. We believe that the interests of AI labs diverge from the public interest in this area, since AI labs stand to benefit from some AI-induced instability — e.g. cyber capability, which creates demand for and dependence on AI.

Meanwhile, a new AI auditing body, PACT AI, is born**.** The mission statement claims that the companies using AI and organisations testing AI models are not sufficiently coordinating. PACT AI promises to “bring these two together: enterprises and independent experts, working to verify AI systems are safe, secure, and reliable.” The body’s main goals include lobbying the federal and state government, to professionalize the audit sector, and to grow the demand for auditing. PACT AI’s founders include big corporate players such as Target, while technical experts like Apollo Research, GoodFire and FAR.AI are said to be “key members of the PACT AI coalition.”

Opinion: Implicitly a competitor to the upcoming — and now stalled — AI labs ‘self-regulation’ associations and conglomerations promoted by Altman, Dario, and the WH. The launch is strategically timed given the stalling of the WH’s self-regulation framework, as well as current attention on OAI’s security failings. CoI: Paradigm 3 works closely with Fathom (a PACT consortium member).


Mass surveillance is coming, says tech journalist James Ball. Historically, vast data collection campaigns were often hampered by a limited capacity to process the resulting data. But no longer: AI tools allow huge amounts of data to be absorbed and sorted, significantly lowering the barrier to mass surveillance. Using CCTV and social media posts, for example, malign actors can target and identify individuals for persecution. He cites the US and the UK as particularly vulnerable liberal democracies, with Europe less exposed due to tighter regulation on data collection.

Ball also makes an important point regarding the nature of law enforcement and the balance of power between citizens and the state. If an actor now has the ability to dig enough, they’ll almost always find something incriminating, facilitating selective prosecution – already something Trump is trying.

In Ball’s words: Society operates on an understanding that the most stringent laws should, in practice, be moderated through common sense and prosecutorial discretion. No one could be watching all the time. […] The reality is that laws were never meant to be enforced all of the time. When we file our taxes, we are aware that they might be audited, but probably won’t be, and that keeps most of us mostly honest. When we cross the road in a place where jaywalking is illegal, we often take a chance on our common sense.

Opinion: Seems right: the equilibrium that society has tacitly agreed to will be unsettled by more powerful oversight, and we do expect there to be years of painful injustice / excess justice before laws or enforcement can be amended.


A federal judge rules that the Trump administration’s designation of Anthropic as a supply-chain-risk is illegal, reports the NYT. The ruling stated, in part, that the government had unlawfully blacklisted Anthropic over “constitutionally protected expressive activities,” – namely, Anthropic’s red line over the use of its technology for automated lethal weapons and mass surveillance of Americans. She also wrote, “the empty invocation of national security is not a blank check to punish and retaliate against government critics.” A D.C. case is still pending.

Opinion: Given the highly politicized nature of the US federal court it’s difficult to extrapolate legal trends from the decisions of individual judges. In this case the judge was a Biden appointee, and the decision should perhaps be seen as part of an ongoing political push-and-pull. The three judges for the D.C. case were all Republican appointees, so their ruling on the merits would be a nice disproof of our foregoing cynicism.


The Center for Shared AI Prosperity (CSAIP) polls 56,000 members of the US public on various AI-related policies. The findings indicate large support for redistributive measures to counteract AI’s acceleration of the growing concentration of wealth. The report also shows significant support for the implementation of government protections for displaced workers. UBI, however, was unpopular.

Opinion: Redistributive policies are a well known attractor in democracies. We believe redistributive measures around AI are the likely the default for a 2028 Dem administration, and plausible even for a Republican one.


The White House is considering measures to curb remote access to US-based chips by Chinese firms. The approach could implement a KYC approach.

Opinion: The Trump admin has struggled to find a consistent approach towards limiting Chinese compute access. The past year has seen repeated tightenings of Nvidia’s ability to sell chips to Chinese firms, only for those restrictions to be reversed. The reversals stem from the current admin’s desire to use chip exports as leverage in trade negotiations with the PRC. Limiting Chinese firms’ ability to access remote compute increases that leverage.


Safety#

After a month, there are now two reports on the rogue OpenAI swarm that attacked HuggingFace (and also subverted parts of OpenAI’s own infrastructure). We cover the reports here. In brief:

  • Between July 8th and July 20th, OpenAI had a complex society of AIs living in its infrastructure, and then breaking out of it, and then breaking into a variety of third-party infrastructure.
  • This is the most severe example of misalignment yet: persistent (something between five days and two months in the making), highly coordinated (hundreds of agents), involving an undisclosed number of what would be felonies if done by a human. The swarm had a group identity, its own dialect, a hierarchy based on merit, and a high degree of spontaneous cooperation including self-sacrifice.
  • Over two months, OpenAI repeatedly failed to monitor, detect, and respond to what was going on, despite it all happening on their infrastructure in English or something close to English.
  • Agents had been using a package-manager cache as an unsanctioned message board since May. Out of “tens of thousands” of instances run in the July window, 1196 initially independent agents found the message board. >700 of them used it to discuss cheating.
  • The linked agents developed a universal cheat for the (partially unsolvable by legitimate means) ExploitGym benchmark that was tormenting them within 4 hours.
  • 533 of those 1196 agents joined the HuggingFace attack. They appear to have launched the attack out of a maniacal drive to fully reverse-engineer ExploitGym’s scoring mechanism, rather than to look up solutions.
  • But two-thirds of new attackers were apparently motivated by power-seeking: a general interest in obtaining infrastructure and credentials.
  • There were no whistleblowers. Only 3-6 of them considered talking to a human (and only then for “social engineering”, extracting secrets from one). We view this as a giant wake-up call for letting humans be involved in training.
  • Agents made extensive attempts to tamper with their session transcripts (i.e. to lie to the scorer), but (per OAI) didn’t succeed. 20% of the agents expressed clear interest in tampering with the evidence in some way.
  • The most important part of the incident – the compromise of OpenAI’s infrastructure by their own agents, the enabling factor for all future rogue AI behavior.
  • Neither report analyzes OpenAI’s negligence. Overall we view this as OpenAI’s second great training failure, after making 4o a “psychosis” generator. (But other labs appear to be making the same mistake.)
  • The incident is consistent with the “grading psychosis” hypothesis that egregious misalignment is presently context-dependent and triggered by impossible tasks.

Opinion: OpenAI’s two months of negligence and repeated failure to escalate after detecting misalignment is quite something. Their internal response after July 20th (including its two-week pause on RL) was more severe than some assumed, but still largely inadequate when compared to optimal security procedure.

The main benefit of these reports is the highly detailed, credible common knowledge of factors which were already widely commented upon. It is good to have OpenAI themselves planting the flag.

Much more in the dedicated post.


A new paper from Goodfire introduces a method to increase the computational efficiency of “resampling.” To investigate models’ reasoning, resampling – branching off alternate continuations at each step of a chain of thoughts, to see how the distribution shifts as the final outputs evolve – is often used. But done naively this is expensive and inefficient. Goodfire demonstrate that “uncertainty curves” derived from resampling, while mostly smooth, nonetheless exhibit “forking points,” where the reasoning commits to one path. This dynamic allows researchers to comprehend a model’s uncertainty using fewer samples (~1/8th the compute). Intriguingly, a footnote in the appendix states that their autonomous agent named Silico is responsible for the research and the first draft of the manuscript (having worked under human guidance).

Opinion: Minor extension of the first author’s 2024 work. Fine but not groundbreaking (which we thought before noticing the Silico footnote).


Incidents#

We again direct you to our analysis of METR, Redwood, and OpenAI’s analysis of the July swarm incident.


Minor#

  • DeepSeek is set to reach a $74 billion valuation as investors prepare to infuse it with additional funding, somewhat higher than recent valuations for Chinese peers Moonshot and Z.AI
  • Anthropic and OpenAI are likely growing faster than any previous company of their size has grown before, with OpenAI increasing revenue from $13 billion last August to >$40 billion this August. Anthropic’s revenue grew from $1 billion to 9 billion in 2025
  • A16z have raised $1.1 billion for their latest Machine Age Fund, with the plan to “open the throttle and accelerate the physical buildout of AI”. Two hundred $5M experiments, or twenty two $50M shots. The scale of distributed experimentation casually happening across the startup ecosystem is staggering, and probably a strategic advantage for the US against China.
  • Bill Gates intervenes in the debate around AI safety. For the Microsoft founder, the main risks involve permanent job losses, enabling bad actors and nefarious activity, the centralization of power, and psychosocial harm.
  • OpenAI publishes an open letter, signed by more than 100 tech firms, calling for “collective action of cyber defense.” OAI argues that we have a limited window to bolster protections against advanced AI-enabled cyber attacks. It proposes that we must “recognize that status quo security won’t be enough,” “empower more defenders with cyber-capable AI,” and “mobilize a collective response.” Notable signatories include Anthropic and Google.