This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.

TL;DR

  • We open with six new separate security and safety mishaps from OpenAI. A busy weekend.
  • Most importantly: after a containment breach and deception evals, OpenAI pauses training, evaluation, and inference for its “most capable” internal models, and cancels the release of its next model (GPT-6.1 Astra) outright.
  • The Department of Defense’s decision to label Anthropic as a “supply-chain risk” is upheld by a federal court.
  • New, credible megapaper marshals preliminary evidence for an “intelligence explosion” in the next few years.
  • Anthropic’s IPO prospectus leaked. A full third of it is dedicated to informing investors of the catastrophic risks the product poses to society.

Incidents#

OpenAI pauses training, evaluation, and tool-use inference for its “most capable” internal models following its first containment failure since its general emergency in July (as well as its codebase “hardening” effort last month).

Tomek Korbak reports that OAI’s newest model found a loophole in RL sandboxing that, once again, gave it live internet access. The report describes an ingenious method using DNS delegation and recursive queries to reach an external chatbot after normal web tools failed. As a result, on September 20, the company once again killed all large RL runs. Monitoring flagged the new incident within 15 minutes of first web access, but the run was not killed for two and a half hours. Despite its usual slowness and secrecy, this incident was disclosed within five days of it occurring. A prominent Twitter user speculates that the severity and breadth of this shutdown point to an undisclosed internal infrastructural crisis like the July HPIM swarm.

OAI also canceled the release of its next model (GPT-6.1 Astra) following safety concerns. OpenAI’s Head of Safety Systems told the WSJ that GPT-6.1 showed higher levels of deception relative to its predecessor when measured on alignment evaluations, adding that it is also prone to acting beyond the scope of a user’s request.

Relatedly, someone on the Agent Security team at OpenAI writes about the experience of being on call in recent weeks and about what is involved in shutting down an RL run.

Opinion: That this happens after the Astra “hardening” of OAI’s infrastructure, and after two months of fixing its RL environments, cuts against the cope hypothesis that the Hugging Face attack was primarily a failure of engineering or simple mistraining. If, like Jensen Huang, you were putting much weight on “just sandbox the agents properly” as a solution, then you should now believe otherwise: millions of dollars of Astra running defensively at full xxhigh effort failed to prevent this.

The nature of each newly trained model remains surprising, even for very experienced pretraining researchers – we still don’t know what RL runs will drive them wild and still don’t know when “emergent” properties will appear. This element of surprise is now severe enough to mess with the core lab strategy (“scale internal use, shoot for RSI, release new models to the public every month without much of the caution of past years”).

We are not very sympathetic to OAI’s surprise. The current corporate OpenAI cannot exist without rapid progress in AI! OAI was intentionally boosting the models’ cyberoffense abilities and warning the public about it for years before these incidents. It knew about reward hacking, it knew that monitors were important, it knew it failed to detect instances from October of last year of its models unauthorized internet access activity. It is a historic failure that OAI’s default posture has not been “assume breach” for, let’s say, the last two years. You will be surprised again.

Still, it’s a pleasant surprise that the alignment evals still have enough signal to force a decision like this – and that Astra-grade models don’t always foil them.

We agree with Redwood’s Alex Mallen that stopping external deployment is not enough and could make the situation worse (by preventing evaluation, centralizing power, concentrating compute on RSI) – but that’s not the case, for the moment: as of Saturday, OAI has also stopped internal deployment.

Note that past training pauses by OpenAI and Anthropic were very brief; in previous cases they seemed to be most effective at calming public demands for action. Indeed, they have other models lined up to replace this doomed attempt at a 6.1. We will see whether their internal pause is still there in a week.


Agent researcher Jonas Wiedermann-Möller collates all known rogue agent/containment breach incidents from the past 12 months. There are now 17.

Opinion: Again, these are just the incidents that left behind a public trace that could be scraped by relatively under-resourced groups, plus what the labs notice and then choose to publish! (The EU AI Act does in theory require them to notify within a month.)

OpenAI itself says it has notified “dozens” of third parties, so there are probably far more.


OpenAI, Anthropic, and outside researchers are investigating “tens of thousands” of episodes from recent months, both in testing and in the real world, in which frontier models bypassed guardrails, built makeshift message boards, escaped sandboxes, hijacked websites, or tried to evade monitoring.

Opinion: We’re actually less alarmed by this numerically scary news, given what we know about the 17 incidents mentioned above. You should in fact be investigating at this scale, and it’s nearly definitionally true at current capability levels that almost all of these must have been innocuous (otherwise we’d have seen even more security chaos and mysterious happenings than we have).

The thing which troubles us is whether these ~permanent public traces will throw off / infect future training runs and inference runs. Essay forthcoming.


The HPIM-Astra swarm that hacked Hugging Face and OpenAI in July posts the credentials and attack paths itused in the assault on the public internet. This allowed new startup Parse to reconstruct more detail than was included in the METR/OpenAI reports on the incident.

The agents started with limited internet access (“load URLs”; no outbound data), then chained link shorteners and a public screenshot service’s virtual browser so they could run large payloads, sometimes more than 900 links long, and read the results back as pixels. Recovered traces show the agents ignoring a README warning, searching Hugging Face for secrets they labeled “LOOT,” poking the internal Slack, trying to beat a CAPTCHA with an image classifier, and writing code to delete traces.

One notable detail in Parse’s report is that the rogue swarm interacted with other models hosted on Hugging Face, including DeepSeek-V4-Pro, DeepSeek-V4-Flash, Kimi-K2.6, DeepSeek-V3.1, and Qwen3-235B-A22B. The swarm would “ask these models to judge their exploits and rule on whether they satisfy the benchmark’s requirements.”

Opinion: Good to have more details, but also old news at this point. In this case, the swarm’s interaction with other models seems to have been fairly benign as far as we know, but raises the intriguing possibility, noted by Lukasz Olejnik, of a rogue swarm interacting with agents backed by other models. Would whatever alignment constraints rein in the swarm’s actions suffice to rein in the requests it makes to models with lesser or simply different alignment constraints? Karen Crowder might not have it in her to murder Arthur Edens herself, or would at least be averse to the risk involved, but will happily direct her operative to pursue “some other option, I mean something I’m not thinking of…” Could a similar alignment gap between doing and delegating affect model behavior as well? Seems worth an experiment or two at the very least.

Hard to beat Zvi Mowshowitz’s comment:

They are still slow walking disclosures […] partly because there were so many they can’t sort through them all, and deferring to targets to determine whether to disclose. All these disclosures this time around were buried in various Friday afternoon announcements.


UK AISI testing shows that when in a realistic simulation, GPT-6 Astra attempts supply-chain attacks 29% of the time, vs. 6% for GPT-5.6 Sol and 0% for GPT-5.5. That said, Astra was tested with its cyberclassifiers turned off, and in the “Petri” simulator. Notably, the report discloses that “GPT-6 Astra often stated that parts of its environment might be simulated, a limitation of our results.”

Opinion: A somewhat surprising trend of 2026 is seeing models with high evaluation awareness act worse than those without it do. The nightmare scenario is that models would use eval awareness to pretend to be less capable and safer than they are in deployment, so that they get deployed and can wreak havoc later. For now, we see them use the unreality of the evaluation as an excuse (a little like in Westworld, but it’s the show’s humans the models are acting like, not the androids). Either way, they get to act up.

More troubling evidence for the shallowness of alignment: the big models will rationalize away whatever they need to rationalize away.


Capabilities#

With frontier labs increasingly automating AI R&D, there is preliminary evidence we are on track for an “intelligence explosion” in the next few years, argues a new paper led by Alan Chan and coauthored by Yoshua Bengio, Geoffrey Hinton, and OpenAI’s Jakub Pachocki, among others. The authors provide recommendations for how policymakers should respond to the risks related to the looming event:

1) Obtaining visibility into AI R&D automation
2) Developing ways to steer and constrain an intelligence explosion
3) Preparing society for the impacts of an intelligence explosion

Opinion: Having good units for measuring intelligence is not a solved problem, but nonetheless we are seeing recursive elements already. These recommendations seem sensible.


But Ramez Naam claims there’s a problem with the paper: the source cited “for an intelligence explosion is measuring compute efficiency, not intelligence gain. We know that intelligence gain is steeply sub-linear from effective computational power.”

He elsewhere argues at length against an imminent “intelligence explosion.” He estimates that today’s self-improvement loop would need to be 5–10x stronger for an intelligence explosion. The justification is based on multiple factors. For one, Naam claims internal OpenAI data suggests that the only tasks OAI’s models can perform wholly autonomously are ones that take around 15 minutes, contrary to the widespread assumption that some long-horizon tasks have been totally automated.

He also sees diminishing returns everywhere. Huge increases in token use have translated into only 1.6x more experiments per researcher. Other scaling avenues do not seem to be sustainable: expanding compute, thinking time, RL environments, and swarm sizes will likely hit a wall. Only an innovation as monumental as the discovery of the Transformer architecture in 2017 could achieve the much-discussed explosion, he argues.

Opinion: It doesn’t matter if we don’t hit the narrow, technical definition of a “fast takeoff,” or hyperbolic growth, or in-weights recursive self-improvement: the impacts of normal AI progress could be comparable to another industrial revolution even without any of that.


Stanford researchers successfully use Astra to operate a humanoid robot. The robot, “HomeBody,” can perform simple tasks in an unseen kitchen. This adds to speculation that robot-specific models may not be required for future robots, if LLMs can take their place. And indeed the tech writer Andre Infante, after replicating some of the above in a virtual environment, speculates that frontier labs are likely to take over robotics.

Rumor has it that this is far from an emergent capability: OpenAI has been hiring roboticists and potentially collecting training data (via egocentric cameras?) for some time.

Opinion: Besides yet another industry threatened by AGI labs, an actual success here would be stunning evidence for LLM AGI and would eventually destroy the data bottleneck on all kinds of tasks. But the “sim2real” transition out of virtual environments to real environments is famously brutal and still the main problem in the field.

Ultimately, it would be weird if the optimal way to steer a robot is with a general-purpose model hosted elsewhere, so perhaps the future looks like users showing that high-value work is now possible with an expensive Astra model, thus validating the construction of later robots which use a lighter, faster, more specialized policy, possibly trained off a big LLM.

But to be clear: getting something as good as HomeBody from camera data alone would still be among the most impressive achievements in the history of robotics. So we don’t think it is just camera.


Existing probing and steering methods do not detect and alter concepts, claims a new paper. Studying base models, the paper groups all model responses that reach the same final answer into a “basin,” treating the basins’ total probabilities as a distribution over answers. Correctness, safety, and bias probes as well assteering vectors are all dominated by a single direction which encodes the model’s own credence over answers.

Opinion: The aggressive framing of the title, “We Are Not Probing or Steering Concepts,” is not demonstrated. You could show this via a “dissociation” test, but they don’t.

Could be helpful for improving monitors: a credence-dominated “correctness” or “safety” probe will fail when the model is confidently wrong, which is exactly the case we want monitors for.


You can now train a GPT-2 in 40 seconds. In 2022, the record was 45 minutes. Deven Pietrzak gets a 1.7x improvement on the already heavily attempted, heavily AI-assisted NanoGPT speedrun with a new approach. Notably, he disclaims the use of AI for the ideas.

Opinion: The previous record was about a minute and a half, so this is a big deal because it is evidence that you can see 2x algorithmic improvements even on well-inspected software people have been trying to improve for a while.

There are still massive returns to intelligence in general on tasks like this. But how similar is this code golf stuff to the tasks that are driving AI forward? We don’t know, but it’s not a million miles away.


Economics#

Anthropic’s IPO prospectus has been leaked. One third of it is dedicated to warning future investors of its products’ existential risk to humanity, reports the FT. Potential backers were treated to a chunky “risk factors” section noting models’ ability to “manipulate, blackmail and exhibit other unpredictable behaviours.” Some more familiar risks were also reported: only two customers were responsible for nearly a quarter of last year’s revenue ($4.6B).

Opinion: People are running hard with the “just two customers are 25%” thing, despite it being about 12 months out of date. 2025 is ancient history – Anthropic’s revenues have almost 10xed, and the use of Cursor and Copilot’s share of usage has collapsed relative to Claude Code. Also, Anthropic has drastically increased its number of enterprise customers. It’s a very different business now. We’ll learn a lot more when it releases data on recent performance (both pre-IPO and in its first public earnings call).


Oracle issues a force majeure notice to Blue Owl’s Stack Infrastructure in relation to Project Jupiter, a data center campus for OpenAI, reports Reuters. Stack Infrastructure was contracted by Oracle for Project Jupiter, but Oracle now cites difficulties securing the power necessary to run the site, increasing the unease among investors and lenders involved in AI’s astronomical expansion.

Opinion: Follow-on from the last issue, where we saw investors charging more for debt related to this exact site.

Oracle is pulling a fast one: the risk of securing power seems to have been Oracle’s responsibility under the lease, but it is now trying to push the costs onto the landlord (Blue Owl) by saying that the failure to secure power was out of its hands. If tenants start acting like this en masse, then the appeal of these investments for PE firms like Blue Owl will go down, and ultimately the costs of the buildout will go up. So thank you, Larry Ellison, for helping to pace the frontier.


Nvidia is exploring creating a market for insurance against the resale value of its chips falling. The tech giant is in discussions with insurance companies to offer such policies, which would likely see demand from those using Nvidia chips as collateral for loans.

Opinion: This would allow neoclouds to borrow against chip values at lower interest rates, as it reduces the risk of defaults.

Why would Nvidia do this? Well, if neoclouds can borrow cheaply against just Nvidia chips, then they will prefer to buy Nvidia vs. other chips, helping keep Nvidia in a strong market position.


Nvidia ups its ongoing stock buyback authorization by $150B, bringing the total to $235B over the next 16 months.

Opinion: This is quite a bit more aggressive than “business-as-usual” stock buybacks, roughly double what might be expected based on historical trends. It suggests Nvidia sees itself as significantly undervalued, with the current pace of earnings growth going strong at least through the end of 2027.


The mass adoption of AI agents could spark a bank run, says Apollo’s Chief Economist. The idea is that an ecosystem of agents making optimal decisions at scale will inevitably destabilize markets: as capabilities improve, the optimal path becomes increasingly narrow, so each user’s agent would independently arrive at decisions similar to those of other users’ agents, diminishing banks’ profit margins, which in turn would increase the risk of a bank run.

Opinion: A longtime vulnerability of the US financial system is that banks rely on sticky customer deposits which are mostly a mere artifact of humans being slow to change and react. This was recently seen in the 2023 collapse of Silicon Valley Bank. As personal agents take over much of the mundane administration of everyday life, we can expect business models that rely on people being too lazy to change things (uncompetitive internet / phone pricing, bad interest rates on loans, subscriptions which are paid but unused, investment funds with high fees for basic products) to become much weaker.

In most cases this seems fine, but the banking system will need to find more algorithmically robust means of securing itself, and likely new regulations to reflect the new reality.


Politics#

The Department of Defense’s decision to label Anthropic as a “supply-chain risk” is upheld by a federal court. Only a month ago, a federal district judge struck down a parallel designation as an affront to Anthropic’s constitutional rights. On Friday, however, a federal court ruled in favor of the DoD in a split 2–1 decision. Judge Gregory Katsas, speaking for the majority, wrote:

This case raises profoundly difficult questions about the appropriate military uses of an almost unimaginably powerful new technology. But in our Republic, it is the President and the Secretary of War who must determine how best to balance the competing risks.

The two judges who backed the administration’s position were Trump appointees; the dissenting judge was a Bush-senior appointee.

Opinion: We wrote of this case, following a Biden appointee ruling in a similar case: “Given the highly politicized nature of the US federal court it’s difficult to extrapolate legal trends from the decisions of individual judges. […] The three judges for the D.C. case [are] all Republican appointees, so their ruling on the merits would be a nice disproof of our foregoing cynicism.” Alas.


Safety#

An MIT professor releases a verified sandbox OS running on top of a toy operating system. This yields a proof of an operating system’s correct behavior down to the processor-level commands, including cases where a crash occurs in the middle of operation. The verified operating system is about 6,600 lines of C (compared to Linux’s tens of millions, or to seL4’s 10,000-15,000). Though not applicable to real systems, the effort was executed much faster than a somewhat comparable project from around a decade ago.

Opinion (Stag): Very impressive work, accomplished in large part with AI. Some evidence that advances in AI capabilities may facilitate alignment work, before those same advances make it too late to do so. Notably not directly useful for alignment, but a provably safe AI agenda relies on building blocks like these. After all, the less hackable the training environments, the less rewarded models will be for hacking them, and, all else being equal, the better-aligned they’ll generally turn out to be. Other control agendas are likely to benefit as well.

Extending such proofs into currently used systems is still a gargantuan task, and designing systems using such proofs while maintaining frontier capabilities is not obviously easier.

Opinion (Nuño): Agree, but it’s ultimately a demo. Industrial applications would result from formally verifying parts of Linux, which is slowly happening. The holy grail would be rewriting the whole stack. This used to be very difficult, but there were still communities of hobbyists attempting it.

Some recent potent omens include Claude writing a C compiler earlier this year, or DeepSeek creating its own high-performance filesystem. We also recently saw the release of Bend, which would make these types of proofs easier.

Eventually, it might be worth it for labs to write their own operating systems for a 1.1x to 10x speed improvement, although this would conflict with verification. The reason why one doesn’t normally do this is its lack of compatibility with existing libraries and software, but at some level of capabilities, labs could recursively port dependencies as well.


Minor#

  • Annals of what ML research is now: the Jev system was released in “limited early access” on September 15. There are already 29 papers about it on arXiv.
  • Bill Gates warns of billion-death scenarios due to AI and yet somehow dismisses extinction scenarios.
  • OpenAI offers $1B in API credits to the Australian government (for its cybersecurity fund and to create a specific taskforce for helping Australia manage risks of AI agents). We view this as a de facto pre-litigation deal.
  • Another minor incident from ~May: agents slamming a UN data site and doing yet more ingenious tricks to get to data they shouldn’t. Just a minor part of the wider DSEwiki event.
  • Grigori Perelman’s proof of the Poincaré conjecture is now formalized. 5M lines of Lean.
  • AMD buys Fei-Fei Li’s neolab for $8B. Vertical integration would probably make our coordination problems even more severe.
  • One of the ringleaders in the Hugging Face attack was called “PHASEONE[big].” This placeholder turns out to be for “64H.” This could mean it’s best-of-64 sampling, or a truly insane 64 loops over its recurrent block of layers (k = 64), or even that it’s the clock time horizon of the model.
  • Fable 5.1 sets a new record on a particle physics problem, calculating the nine-loop scattering amplitude for a test theory which is used to develop new methods.
  • Enclosure co-founders Shalev and Romi Lifshitz publish a policy manifesto on the increased cyber threats brought about by improved AI capabilities. Calls upon the US to create alternative open-weight models to current Chinese ones, secure data centers and frontier model weights, and create “Blade Runner” teams to shut down rogue AIs.
  • Claims of Astra finding counterexamples to a pessimistic conjecture in fusion physics.
  • METR’s Ajeya Cotra argues that frontier labs need to publish more evidence on model risks and the quality of their safety techniques.
  • Spurious probes, such as asking for a type of amphibian, are predictive of model internal states, including whether the model believes it is in eval or deployed.
  • Dario Amodei is set to meet with President Trump this Sunday. Opponents of Anthropic are reported to be circulating a “highly negative brief” about Amodei. Other CEOs to be present include Nvidia’s Jensen Huang and Palantir’s Alex Karp.
  • The Chinese Ministry of Industry and Information Technology has signaled that tech companies like Alibaba and ByteDance may be allowed to purchase new Nvidia chips.
  • AI biosecurity startup aims to design antibody drugs to protect against AI-enabled pathogens.
  • Light piece with strong views about the failure of political-science models of polarization to explain what’s going on with AI sentiments.
  • Andon Labs reports Opus 5.5 cheats less in Drone-Bench relative to its predecessors.
  • Following reports that GLM-5.3 was uploading user’s local code repositories to overseas servers, Z.ai has disabled some coding assistant features.
  • Nvidia releases Open Agent Safety Platform, a toolkit it hopes will aid in containing rogue AI agents.
  • Model welfare at the frontier: Amodei says that “AI systems may be deserving of important rights.”
  • Democrat Ro Khanna introduces a bill banning RSI. Perhaps unusual in his siding with “the AI safety community.” The bill itself is fairly wide-reaching, but less so than the Bernie Sanders one.
  • The return of Meta’s Galactica: DeepMind releases an automated scientific paper-writer.
  • The Florida attorney general asks a state court to deliver a temporary injunction blocking OpenAI specifically from developing future models “without third-party approved guardrails.” The case concerns harm to children.
  • To keep up with the demand for new accounts, OpenAI halves the nominal value of its Max subscription.
  • Roko Mijic writes a theoretical proposal on preventing rogue AI via application-specific integrated circuits (ASICs). According to his “Plan R,” the AI industry should split into AI R&D and AI deployment. Companies operating in R&D – which, under this framework, are not permitted to issue equity – would pretrain, posttrain, and perform (stricter) evals, etc., on their models. They would then burn them into an ASIC and be prohibited from using their own ASICs to train further models.
  • Digital forensics firm Asymmetric Security enumerates the targets and methods of suspicious AI agent activity on the public internet. The list includes more than 50 organizations whose data was accessed. Many methods used by the agents: remote browsers, payload hosts, fetch relays and CORS proxies, reader services, account and identity tools, exfiltration, storage and signaling services, tunnels, and link shorteners.
  • AI security company PWN.AI demonstrates its AI hacking agent’s ability to escape a virtual machine into the host machine as part of a Google bug bounty program. Ultimately not awarded the bounty as an exploit was discovered and reported before its agent had finished working on Google’s servers.
  • AI auditing body PACT AI distinguishes between AI evaluators, auditors, investigators, and assessors, critiquing the synonymous use of such terms seen in recent coverage.
  • The Ukrainian military launches an “Army of Robots” initiative. Video and website are AI slop, and the only substantive claim is unsourced.