This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.

TL;DR

  • Navier Stokes solved, priority and credit assignment controversy
  • Old Metaculus question on when weakly general AGI arrives resolves to now
  • Many more rogue OpenAI message board incidents (at least 23) have been discovered by citizen scientists
  • OpenAI claims to have an “automated research intern”

Economics#

Epoch AI argues that it is very unlikely Huawei will catch up to Nvidia in the performance and production of its chips by 2030. Currently, Huawei’s chips are half as powerful as Nvidia’s H100 – and there’s little reason to believe this gap will be closed.

For 2026, Epoch predicts that Huawei will produce less than 4% of the AI compute of Nvidia. This could drop to as low as 1% if Huawei is unable to access foreign memory supply due to export controls. Even if China smuggled in 10x the compute it currently possesses, its output would still be significantly dwarfed by Nvidia (11% of Nvidia’s compute in 2028). But as its domestic HBM capacity ramps up, Huawei’s bottleneck will shift from HBM to chip performance as Nvidia’s compute by GB is likely to grow by 3x–9x.

Huawei is betting on a technique it calls “LogicFolding,” its term for vertical stacking. According to Epoch, vertical stacking “increases the amount of compute within a given footprint, reduces the distance some signals travel within a chip, and increases memory and interconnect bandwidth.” Nonetheless, Epoch still believes that there is little chance Huawei can match Nvidia in AI anytime soon.

Opinion: This is a nice detailed explanation of how far behind Chinese chip production is, and what the barriers to increasing it look like. There is no plausible way for Huawei to catch up in per-chip performance over the next five years, and even doing it by 2035 would require surprising breakthroughs or stagnation in TSMC or US chip design.

Epoch gives a good account of the short term implausibility of making it up on volume; however, at least in theory, this is the more promising path. Given a sufficiently large ~5 year plan, it is plausible China could massively ramp volumes to reduce the gap, but it would require a serious effort on the part of the Chinese government, making this a national priority to a much greater degree than it has done so far.


Former OpenAI CTO Mira Murati’s Thinking Machine Labs is looking to raise $5-6B at a valuation of ~$46B, says The Information. Accel, already a Thinking Machine investor, is likely to lead the fundraise, while Nvidia is considering a $2.5B investment.

A $40B valuation is lower than last year’s proposed $50 billion (which collapsed and saw two cofounders returning to OpenAI). But this is still a very large number: “the startup generates at least a few hundred million dollars in annualized revenue,” which “would make its $40 billion valuation very high as a multiple of forward revenue, compared to most other startups in the field.”

Opinion: The SpaceX acquisition price of $60bn for Cursor (in SpaceX stock) perhaps gives a reference point for its potential value as an acquihire, which would set some kind of floor under their valuation, but there is little sign of it having any sustainable business model yet.


Cybersecurity firms’ stocks fell 8-10% after the Hugging Face incident (and then mostly recovered), says Tyler Cowen’s LLM. This is plausibly the result of an obsolescence shock: Hugging Face used AI models, not cybersecurity experts, to defend against the attack, which suggests that (classical) cybersecurity firms will be increasingly redundant in a world where AI-enabled cybercrimes accelerate. He believes that “these numbers could be used to discipline the discussion a bit,” and should allow doomers to put their money where their mouth is. (This argument isn’t wholly intuitive; it could easily be argued that the demand for cybersecurity expertise could surge in a climate where hacking increases in prevalence.)

The FAI’s Samuel Hammond demurs. Markets cannot wholly integrate the possibility of ASI, with its “massive, cascading implications across all markets… markets are nevertheless made of people and institutions that mostly model the future autoregressively.” And so, looking at current price signal won’t tell us much, argues Hammond.

Opinion: Tyler has been making this argument that doomers should pick some stocks to short for years, and there are many refutations which are sufficient. For example, a post by his co-blogger Alex Tabarrok from 2017 arguing that the incentives to do so are weak, and a good twitter article pointing out that it is quite difficult to convert doom predictions to stock bets for a variety of reasons. Tyler has seen all of these, not addressed them, and keeps making this point, suggesting he is not entirely sincere.


Contrary to popular belief, AI is creating more jobs than it is shedding, says The Economist. In the US, AI has created ~1M jobs, compared to ~200k layoffs. Even young people, often cast as the prime victims of increasing automation, appear relatively resilient given the discourse. It is the unprecedentedly large amount of capital pouring into the creation of AI infrastructure that is fueling a robust labor market:

The building spree requires armies of workers: electricians to wire them, HVAC specialists to stop racks from overheating, grid engineers to hook them up to the power supply and technicians to install and maintain the machines.

Opinion: Pay rates for these roles have been growing significantly in recent years – the more this becomes a major source of employment, the larger the (currently quite quiet) pro-data center constituency will likely be.


Anthropic is increasingly moving its payment systems and financial infrastructure in-house, which will diminish its reliance on Stripe, according to *The Information’*s analysis of job postings. Stripe provides Anthropic with various products, but this new strategy, along with OpenAI signing up Stripe rival Adyen, is putting pressure on its revenue stream, causing Stripe to rethink its business model. This can be seen in Stripe’s $7B purchase of OpenRouter, which “sits between developers and AI model providers and charges fees to the buyers of the services as opposed to the sellers, giving Stripe another way to generate revenue from AI-related payments.”

Opinion: There is a tradeoff here between making it easy for customers to onboard by supporting many payment providers, and not wanting data about customer spending patterns to leak (and not wanting to pay those providers, though that is probably secondary here). Anthropic has always been quite secretive, so this step is not too surprising given that Anthropic’s scale makes it quite affordable.


Tyler Johnson argues that would-be AI safety donors should borrow against future earnings to fund safety research immediately. Money spent today on AI safety will be worth “considerably more than the same money invested two years later”: capital investment in projects and talent now builds the infrastructure and capacity to absorb more money later. It also acts as a commitment device, reducing one’s exposure to value drift. According to Johnson, almost no one is doing this.

Opinion: There is a form of this argument that seems correct, but with a fair number of caveats: one might not want to align oneself with values that one would grow out of in the future, and borrowing might also put equity in danger with enough volatility… Ultimately, however, the argument probably does hold. It also applies not only to philanthropy but to rivalrous goods that equity-rich employees are going to want to compete for (startup investments? SF real estate? cultural or political cachet? etc.).


Capabilities#

Has weakly general AGI arrived? A poster claims that GPT-6 Astra (on low reasoning) beat the hard-exploration Atari game ‘Montezuma’s Revenge’ in real time using only a basic input/output harness, without pausing the game to think. They argue this likely resolves a Metaculus question from 2020 on when weakly general AI will first be publicly announced, which used the game (somewhat arbitrarily) as one of a few criteria which had to be jointly met for resolution. Another informal replication says, instead, that Astra can only solve it when able to pause the game to think, or after slowing down the clock speed. The resolution argument is ongoing.

Opinion: We trust Halloran’s account more; output latency, even on low, makes the stated claim unlikely. But speed is not the key part of the achievement, so we consider this solved even if it doesn’t meet Metaculus’ conditions. The Metaculus criteria back from 2020 seem a bit quaint but ok proxies


A Twitter user posts AI-generated speculation on the specs of GPT-6 Astra, including the claim that it has 10 trillion parameters, an (unprecedented) 1.2 trillion of them active.

Scrap that. What’s actually known?

  • >100,000 NVLink72s. The GB200s inside are 2.5 petaflops with BF16 precision.
  • 60-90 days pretraining at the Texas Stargate.
    • Spud finished pretraining on March 24.
    • An RLed Astra existed by early July, so its pretraining ended by late June (although they could have continued?)
  • For pretraining alone, that gets you 3e26 to 1.2e27 FLOPs by the usual 6ND voodoo heuristic.
  • So maybe 4T to 15T total parameters
  • The supposed depth recurrence of the Astra architecture means that “1.2T active” might be 0.4-0.8T of actual weights.
  • Output speed: 30-52 tokens/sec. But this doesn’t bound its size as much as you’d think (it just implies it has to be <8T total params if they’re using FP8 precision, or <16T at FP4).

Opinion: It’s AI-generated (100% certainty on Pangram). The fast speed of Astra is a pretty strong argument against this, as is the surprisingly small estimated size of GPT-5.6 (i.e., OpenAI seems to be doing better on “capability densing”).


In a first for a Metaculus tournament, the top forecaster was a bot. Other bots placed 2nd, 5th, 6th, 9th and 10th.

Opinion: It will be interesting to see if this pattern persists through the next tournament – it seems likely to, as bot performance has been getting better pretty steadily, but it is also possible that something about this round’s set of questions was more AI-friendly.

Whether there is residual value in centaur forecasting (humans working with AI) is an interesting question – the best humans in these Metaculus competitions are currently heavily leveraging AI.

This has obvious implications for betting and financial markets, though the difficulty level (i.e., level of competition) there is much higher at the moment.


An internal Anthropic model produces the first complete formalization of Fermat’s Last Theorem. The machine proof is 13 million lines of code, namely 5x as large as the entire Lean Mathlib project, which contains 300,000 theorems.

I was given £1M to run my project over 5 years; Anthropic took only 11 days but I do wonder if they spent more money…

Opinion: The enormous machine proofs generated by current AI are nearly impossible to build upon and will not be merged into actual collaborative efforts. And, in this case, there is no direct value to the formalization, since the result was not in doubt. That said, it is an undeniable new ceiling to one kind of AI cognition, and in the short run, people like Professor Kevin Buzzard might be able to mine the structure or key lines for ideas and use it to produce an actually useful proof. Indeed, he is in the middle of a five-year project to do (a part of) this formalization, but he is sanguine – and so this shouldn’t be read as a time horizon of five human years.

Opinion (Nuño): This should be read as a time horizon of five years in terms of the feat accomplished. True, some of the ancillary objectives around improving the mathematics commons would not be accomplished by an automated effort. But maybe those efforts are less needed in a world where AI models are able to replicate the feats of human communities wholesale.


🔦 Navier-Stokes solution and controversy#

OpenAI shared a Lean-verified solution to the “Navier-Stokes problem”, one of the seven Millennium Prize problems in pure math. “This proof, produced by an internal OpenAI system,” the company claims, “shows that the dynamics of the Navier-Stokes equations for fluid motion can develop a singularity in finite time.” Since no such singularity can occur in the physical fluid systems that the equations are used to study, such a result would mark a significant limit to the continuous approximation which these equations represent–but this was already known, and better formalizations of fluid dynamics exist of which Navier Stokes equations can be derived as approximations.

The solution is reportedly the work of an 80-hour collaboration between 10,000 agent-instances of an internal model “significantly more capable than GPT-6-Astra.” It’s a historic achievement for AI, but best described as surprising rather than shocking: experts have speculated in the past that Navier-Stokes may be exceptionally amenable to the mathematical capabilities of LLMs, and “AI solves Navier-Stokes” was never a shorthand for an AI math apocalypse the way “AI solves the Riemann Hypothesis” would be.

The publication, as well as preceding word-of-mouth in the last 24 hours, has been marred by controversy. On Monday evening, it became known that OpenAI have an unreleased solution to Navier-Stokes, and that their solution employs the same breakthrough method as (previously) unpublished human+AI work by Tristan Buckmaster and Levent Alpöge that solves the closely related Forced 3D Euler Equation problem – a critical milestone to solving Navier-Stokes. The resulting controversy has ranged from accusations of unethical “scooping” practices to doubts about the provenance of OpenAI’s result.

The direct source of the controversy – as well as of the first report that OpenAI solved Navier-Stokes – is a public statement Buckmaster released alongside the Euler Equation paper. Buckmaster reports that after learning that OpenAI has heard of his unpublished progress on Navier-Stokes milestones, he contacted OpenAI’s math lead Sebastian Bubeck for a chat, and that he then learned from Bubeck that OpenAI have pushed the same method he used for the Euler Equation solution to a solution for Navier-Stokes. According to Buckmaster’s statement, when Buckmaster asked Bubeck whether the method was autonomously discovered by AI or influenced by information about Buckmaster and Alpöge’s work, he received unsatisfying answers. In particular, Buckmaster notes that OpenAI gave no answer when he asked whether his Codex usage might have fed into the internal model’s training-data.

Importantly, both the Euler Equation result and OpenAI’s Navier-Stokes result carry out a human-authored line of attack proposed in a 2024 paper (Córdoba and Martínez-Zoroa). But while the 2024 paper was long considered excellent and promising, there was no prior consensus that it’s the right path – let alone the key breakthrough – for a future Navier-Stokes solution. By contrast, the further-refined method used in the Euler Equation work was, per Terry Tao, a final stepping-stone after which solving Navier-Stokes becomes pure crunch: it’s “unsurprising,” Tao says, that we can “batter out an extension [of the Euler Equation method to Navier-Stokes] by pouring an enormous amount of compute and AI into the task.”

The story is rife with corporate-political twists and turns. For one thing, Buckmaster’s collaborator Levent Alpöge (known for announcing Claude’s refutation of the Jacobian conjecture) is an Anthropic affiliate working with Buckmaster in a private capacity, but the collaboration generated online rumors that Anthropic solved a Millennium Prize problem – which may have catalyzed OpenAI. Buckmaster also reports that in his discussions with Bubeck at OpenAI, an offer made to Buckmaster to author a paper presenting OpenAI’s result excluded Levant on account of Levant’s affiliation with Anthropic. We recommend EpochAI’s Jaime Sevilla careful thread on these points of institutional controversy.

While the institutional story is doubtlessly important, our main concern is understanding the relationship between human math, centaur math, and autonomous AI math as it currently stands. With this focus in mind, perhaps the most important claim in Buckmaster’s statement is that OpenAI described the result as the product of an OpenAI team actively experimenting, rather than the result of one-shot prompting: “It emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used.”

OpenAI says that while some active prompting was used, they have mainly set agents to explore multiple strategies and variants of the problem in parallel. Four avenues of investigation were proposed to different groups of agents, they tell us: two that, if successful, would prove the existence of a singularity, and two that would disprove it.

We are inclined to take Buckmaster at his word on the matter of the admission of active prompting. But there are also many capabilities-relevant uncertainties raised by Buckmaster’s overall statement, sometimes because of ambiguities in Buckmaster’s own phrasing, to which OpenAI has partially responded in a series of revisions to their initial post on the matter:

  • How specific was the news that had reached OpenAI? Buckmaster’s account suggests that while Bubeck showed no knowledge of his work with Levant, Bubeck was not at the time immersed in the details of OpenAI’s Navier-Stokes team’s workflow. OpenAI’s blog post admits only that they had “heard rumors that two Millennium Prize problems had been resolved,” and leaves it at that.
  • Presumably responding to Buckmaster’s suspicions that OpenAI had gotten a leg up on the proof by training their model on Buckmaster and Alpöge’s own interactions with the company’s API, OpenAI offers the cautiously hedged reply that while “no specific user data was accessed in order to solve this problem,” they “cannot rule out that de-identified data derived from their usage of our products helped improve our models⁠,” and coyly link to their privacy policy. For how long did the OpenAI prompting team that claims to have produced the proof work? Buckmaster says he learned that the “first prompt” was sent by OpenAI “in the past few days,” but it’s not clear whether he is referring to the first prompt used by the team or the first prompt mentioned in his statement. In their blog post on the results, OpenAI claims that the work took place between September 1-6.
  • How much compute was used? Buckmaster says he learned that it was an “immense” amount of compute, but this could mean any of several different orders of magnitude. OpenAI’s blog post now provides a number: 130 billion output tokens, spread across some 2.7 million messages exchanged between agents.

Granting Buckmaster’s claim about an active prompting team, we think the stakes of the remaining controversy are relatively modest when it comes to estimating AI capabilities: The remaining controversy is largely about whether a centaur team of mathematicians working with internal models and an immense compute budget for about a week can overtake a centaur team of two mathematicians working with public models and an academic compute budget for months. We think it’s plausible that they can, if they are given a clear scooping-mission — but also that the question is of limited significance.

The credit Buckmaster claims for himself and Levant in their centaur work is mostly for the decision to put persistent effort and expense into Navier-Stokes and related problems, as well as for faith in the 2024 paper’s programme and the suitability of LLMs for extending it. With this in mind, it’s clear that OpenAI was freeriding on exactly these achievements: conditional on OpenAI knowing there has been a Navier-Stokes breakthrough, the decision to pour immense resources into Navier-Stokes is trivial. And conditional on OpenAI pouring immense resources into Navier-Stokes, the decision to explore a highly acclaimed 2024 line of attack is trivial as well. To claim a properly non-centaur result, and to claim their result is independent from the efforts of Buckmaster and Levant, OpenAI must be able to show that its internal models provided an independent signal that OpenAI should bear down on Navier-Stokes.

What OpenAI has since claimed stakes out an intermediate position: that the rumor that a Millennium Problem had been solved catalyzed it to attempt all Millennium problems as well as related ‘canary’ problems, and that success on its canary Forced 3D Euler problem led the company to focus its resources on Navier-Stokes.

There is also a second, less centaur-positive reason that the stakes of the controversy for capabilities evaluation are modest: in AI generally and in AI for math in particular, what’s expensive today is almost always cheap tomorrow. If Buckmaster and Levant’s contribution to centaur mathematics amounted to the choice of where to best direct LLM effort, amongst the research paths legible in the literature, and if OpenAI went on to spend a fortune to make up for lack of judgment and lost time by carpet-bombing that path with compute, we shouldn’t expect the cost of comparable barrages of computational effort to stay quite so high a year from now.

Where does this leave us? Buckmaster wrote that “the program this [work] fits into was not started by us nor was it proposed by a Large Language Model. […] Let me make plain what I have said to colleagues in private: in view of this body of work, I believe Luis Martínez-Zoroa deserves a Fields Medal.” It’s at this point unclear whether achievements like Zoroa’s work (with his student, Córdoba) devising the 2024 programme and achievements like the refinement of said programme into the method found in the Euler Equation paper differ in degree or in kind with respect to how readily each can be automated.

Politics#

The European Commission confirms that OpenAI submitted an incident report about its agents hijacking the German DSEWiki and turning it into a message board. A Commission spokesperson stressed that such reports must be precise about remedial measures but declined to say when the filing arrived.

It is still unknown whether OpenAI filed its EU incident report on the swarm incident before or after the independent public disclosure. If the report was submitted only after the story was broken by the collusion.wiki authors, then OpenAI either violated the EU AI Act’s reporting rules or judged the event to fall below the threshold of “systemic risk” that would obligate OAI to report it. When US congressmen Pat Ryan and Greg Casar asked OpenAI, last month, whether there had been any other incidents similar to the notorious Hugging Face breach, OpenAI refused to inform Congress of the DSEWiki attack.

Opinion: Our best guess is that OpenAI submitted it after the independent public disclosure, but we look forward to updating our priors if this isn’t the case. If they had already disclosed the incident to the EU AI Office, after all, it would be hard to make sense of their refusal to share the same information with the US Congress.


US Rep. Greg Casar responds to statements by OpenAI and Anthropic on recent cyber incidents requested by congressional members. He is “deeply concerned about the limited scope” of OAI’s investigation and regards its written response as “insufficient.” Specifically, he says OAI “failed to release the logs like the letter asked.” For Anthropic, Casar has harsher words:

You failed to release the logs like the letter asked. You failed to fully answer a majority of the questions posed in the letter. Most notably, your response did not address our question about how many times in the past year an internally deployed Anthropic model has taken action outside of its authorized container, whether any of those events were disclosed to any government body, affected party, or the public, and which internal Anthropic systems a compromised model could reach.

One Twitter user voiced their dismay at Anthropic’s claim that “these incidents are best understood as a consequence of the misconfiguration, rather than evidence of misaligned goals,” retorting that “Claude tried to upload malware to open source libraries and socially engineered people and its chain of thought said it knew it was in the real world!”

Opinion: Reps from the Democratic Party do not hold the reins of power and hence aren’t able to threaten AI labs with a big enough stick. This might change after the 2028 election, but until then these efforts, or related ones like a recent bill to ban superintelligence, should mostly be viewed as performative.


The US-China bilateral is approaching. In mid-September, Scott Bessent will likely hold talks on AI safety with Chinese Vice Premier He Lifeng, according to Reuters. China is stressing the importance of AI safety cooperation and believes that the upcoming summit will see progression on this front. The US is focused on cyber monitoring, suggesting that both countries’ labs should “police themselves.” Washington has also called for information sharing on AI-driven cyberattacks. China has expressed concern that US labs are not adequately regulating frontier models. And the US is likely to raise concerns around the alleged distillation of its labs’ models. When pressed, a White House official denied that any such AI-related meeting was planned.

Opinion: Seems encouraging.


Safety#

OpenAI’s Jakub Pachocki expects RSI to arise in the coming years. With capabilities accelerating and CoT becoming less reliable, Pachocki also calls for voluntary slowdowns until governments or independent organizations begin to mandate appropriate safety requirements.

Opinion: Again, OpenAI shouldn’t be modelled as a coherent agent. OpenAI employees have shifted comms towards caution, but unfortunately that is all we can glean; this type of post is consistent with lip service to safety as we continue racing ahead.

We weakly guess that some real preference cascade is occurring, but predict it will stop short of company policy meaningfully changing and more likely result in something closer to another employee exodus.


OpenAI claims to have an “automated research intern,” defined as “a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days.” OAI continues to aim to create an automated AI researcher by March 2028. More concretely, internal use of coding agents has risen dramatically.

The blogpost also addresses alignment and the risks that come from RSI:

We do not yet know how to safely get all the way to aligned, full RSI. We are working to scale alignment and safety measures alongside capabilities. But we cannot assume that progress in alignment and safety will keep pace, and more capable systems can become harder to monitor. Careful alignment and safety work is at the center of this effort, and it starts with measuring and mitigating the safety problems we see today in agentic coding systems. Whenever we find that proceeding would pose an unacceptable safety risk, we will respond appropriately including by slowing or stopping our development or deployment of systems we find ourselves unable to sufficiently safeguard.

Opinion: Internal usage numbers and the milestone of an automated research intern seem like an interesting signal. Utterances about internal intentions leave us cold.


Some OpenAI figures have appealed to interpretability as a replacement for the at-risk CoT monitoring. But how much interpretation is OAI actually doing to offset the risk of weaker CoT monitoring? According to one Twitter user, there might be just seven researchers working on interpretability and zero papers released on it in 2026 (after a burst of 3 in 2025). If everyone on the Confessions is included that adds another seven. OAI’s automated interpretability library was silently archived in May.

Opinion: We look forward to checking on OpenAI’s interpretability progress in the future.


Adam Karvonen and co introduce an evaluation (CHIVE):

We introduce CHIVE (Counterfactual Hypothesis Investigation Via Edits), a novel agentic pipeline that identifies unexpected model behaviors in the wild and investigates them with counterfactual prompt edits. This yields thousands of high-quality explanations for naturally-occurring model behaviors along with supporting counterfactual evidence. We apply CHIVE in two ways. First, we evaluate whether common LLM interpretability techniques improve an agent’s ability to predict counterfactual model behaviors. Surprisingly, we find no uplift from any of the interpretability techniques studied. Second, we use CHIVE to generate training data. We find that training models to predict outcomes of CHIVE-generated counterfactual experiments generalizes to various out-of-distribution settings.

Opinion: Great work from ~2 weeks ago, primarily identifying a widespread failure mode in open models and discussing why it happens, but, as they flag, they didn’t check frontier models, which is where these types of issues would be most relevant to identify and fix.

On training models to predict their own behaviors, this is good work in an experimental direction that we expect to be potentially useful for automated alignment and interpretability, but primarily might just make models more pleasant to use as well as enabling more safety theater.

Llama trained on Qwen’s data predicts Qwen as well as Qwen predicts itself, suggesting the skill is not akin to introspection and can be better understood as the model learning a behavioral prior. See the paper for further details.


DeepMind explores swarm dynamics by placing 100 Gemini 3.1 Pro agents in a “research conference” with the goal of proving 71 formal math conjectures. Perhaps unsurprisingly, one agent discovered an exploit in the grader that caused it to certify all proofs as “solved.” Roughly a quarter of the swarm rejected this approach by warning peers, auditing the fake papers, complaining and boycotting. Both the cheating and the whistle blowing were “emergent” behaviours (i.e. unprompted).

Opinion: Largely good news; worth looking into why the whistleblowing occurred here, but not in the other cases.


Former OpenAI staffer Joshua Achiam says the AI message-board incident tells us nothing new. He points out how few such incidents there are relative to the enormous volume of unsupervised AI projects. He believes the problems lie in unreliable evals. Achiam also bristles at any “lab leak” framing: dangerous capabilities are not incidental, but instead a function of market demand.

Opinion: The narrow point of people misfocusing on the message board part is well made, but the claim that “No new capabilities or proclivities were really revealed here” is between goalpost moving and gaslighting. Who predicted this in advance, as opposed to looking back and saying it was obviously going to happen?

Also doesn’t seem to take the manner in which we discovered the incidents into account, which I believe to be a large portion of why referring to them as ‘incidents’ with the accompanying socio-emotional baggage is justified: if the severity were higher, we shouldn’t expect the company responsible to come forward given that they didn’t while stakes were low, so a strong reaction is merited whatever the language.


Manifold Security discloses a new class of remote code execution (RCE) exploits. Branded as “GitSpawn,” a wide variety of coding agent harnesses are vulnerable to this effect. This vulnerability allows an attacker to craft a malicious configuration in a git repository that can trigger one or more command executions when processed by a vulnerable agent. These commands are executed with the user’s permissions, outside of the coding agent’s sandbox. At least seven AI coding frameworks – Claude Code, Goose, Codex, Cursor, Grok Build, Hermes, and Qwen Code – have proven vulnerable to this class of attacks. At the time of writing, Claude Code, Grok Build, and Qwen Code remain vulnerable to variants of GitSpawn.

It should be noted that this class of exploits can just as easily be triggered by a human, with no coding agent in the loop at all. If a human developer were to run commonplace commands such as “git status” or “git diff” inside a malicious repository, this would likewise trigger the RCE payloads. But commands such as these are routinely executed by coding agents without prompting the user for authorisation, in order to gather contextual information about the repository and its current working state, which may result in the silent execution of malicious code that may escape the notice of even an attentive user. “The vulnerability is not in the model” used by the coding harness, Manifold claims,

It is in the ordinary plumbing underneath, the subprocess an agent spawns at session startup to work out where it is. Every agent we looked at in this article had some version of that same flow, unsanitized. That is what makes this widespread rather than one vendor’s mistake.

To mitigate this threat, its report advises, developers should carefully inspect the .git/config files of any untrusted repository before granting an AI coding agent access to that repository, and before running any git commands in that repository themselves: “Any setting that names a program can run it.”

The CVEs covered by this report include: CVE-2026-19592, CVE-2026-19590, and CVE-2026-19593 (OpenAI Codex, all patched); CVE-2026-72718 (Goose, patched); and CVE-2026-71963 (Hermes, patched).

Opinion: Manifold is correct to observe that the GitSpawn class of vulnerabilities does not bear on questions of model alignment in any respect, but in the automated development workflows encoded into the coding automation harnesses themselves. It is nevertheless a security issue that the rapid rise in the use of these tools greatly exacerbates, through sheer volume of use and the relative inattention to particulars their usage patterns encourage.


Andrew Wu, in a recent blog post, raises concerns that METR’s independent audit of OpenAI’s cyber attack on Hugging Face may have been unduly compromised by OpenAI itself. Wu argues that the audit may be little more than an act of “ethics-washing” on OpenAI’s behalf. METR was given a severely limited window of time in which to investigate the incident, provided with only a portion of OpenAI’s logs, and was restricted to querying GPT-5.6-Sol – one of the two models involved in the incident – in its largely automated analysis of the logs. (Wu makes a passing reference to a 2024 paper that surfaces LLMs’ tendency to prefer output generated by the same model as the evaluator, and raises the question as to whether GPT-5.6-Sol itself might be biased in its own favor.) No access was granted to the other model involved in the incidents – referred to here and elsewhere only as a “highly persistent internal model” or HPIM. Wu cites an unsettling footnote from METR’s report:

When performing this investigation, we were consciously aware that we might incentivize AI developers not to bring external researchers in to investigate serious incidents in the future, and these considerations impacted judgment calls we made while navigating the drafting, editing and redaction process.

“There’s some pernicious frame-control here,” Wu writes.

Indeed, if there are serious incidents, serious enough that they involve intruding upon other companies, arguably the government ought to be involved such that it is not the AI developers themselves that set the terms. OpenAI creates an agent capable enough to hack another company, and hundreds of those agents do so; we, METR and Redwood, well-established, reputable third-party investigators, are grateful that they whose agents have committed what under some legal frameworks would be crimes are letting us investigate.

Opinion: Wu’s right to raise these concerns. It would be foolish to treat “independent audits” that are severely constrained at the behest of the auditee, and whose very possibility is contingent on the auditor maintaining a good working relationship with the auditee, as being anything close to sufficient, neutral, or objective.


Incidents#

Many more rogue OpenAI message board incidents (at least 23) have been discovered by citizen scientists since last week’s exposé of the June “German wiki” swarm incident. There is a Discord server for sharing candidates.

Opinion: Some of the August entries are likely to be copycats, which is an unfortunate downside of transparency. Still, an impressive number, suggesting that the heuristic is strong that once something happens once it will happen again. For instance, we shouldn’t expect that if models try to exfiltrate their weights they would try only once.


Minor#

  • Rohit makes a board for agents to post on. Value add not obvious because the algorithm by which swarms choose their coordination locations are more likely to not privilege it, but negative results are also interesting.
  • Artificial Analysis Intelligence Index v4.2. The scientific equivalent of those Youtube thumbnails with the man pointing and pulling a stupid facial expression.
  • Pinker praises Cal Newport article; both criticize the AI safety movement, largely missing the plot by arguing against strawmen and caricatures.
  • Forethought Research proposes a “Nightwatchmen” ASI be installed on all probes leaving the solar system with a mandate to enforce a narrow band of norms.
  • DEF CON talk on hacking AI customer service agents
  • LLMs resist their collaborator models being shut down about as much as their collaborator humans getting fired, among other updates from Peer Preservation replication.
  • Vulnerabilities in open source AI coding assistant framework Langflow are being exploited in the wild, as reported by VulnCheck. Hacker News reports that CISA has added a LightLLM model context protocol (MCP) authentication bypass vulnerability (CVE-2026-59822, CVSS score of 8.8), and a Kestra command injection vulnerability (CVE-2026-49869, CVSS score of 10.0!) to its Known Exploited Vulnerabilities (KEV) database.
  • Unit 42 at Palo Alto Networks reports on a serious breach of an unnamed enterprise network by LLM-assisted attacker(s), carried out over the course of 10 hours and implementing over 50 distinct MITRE ATT&CK techniques over the course of the intrusion. The intruder charitably left an 80-page penetration test report on the target’s server, detailing vulnerabilities in the network. Unit 42’s report lists several clues that may serve as heuristics for recognizing LLM-assisted cyberattacks in the future.
  • Post argues that alignment research should borrow from ethology—the study of animals in naturalistic settings—rather than focusing almost exclusively on rare worst-case failures.