This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR
- We cover various executive updates and decisions in USG AI policy. Most importantly, Paul Christiano leaves CAISI.
- We summarize the flurry of real-world loss of control cyber incidents.
- We tested the DeepSeek incremental model update: strong but narrow progress.
- Monthly reminder that AI evaluations are extremely hard to interpret, this time from Greg Lewis
Economics#
A former OpenAI employee has presented a sophisticated bear take on frontier labs. > if OpenAI had paused model development last year, there would no longer be any point in paying GPT-5 API prices when you can just use Qwen or Kimi instead for much cheaper. Thus, the labs are forced to invest ever-increasing amounts of money in model training… even if your revenue goes up with higher model capabilities, so do your future training costs. This is a profoundly punishing dynamic which severely penalizes frontrunners.
Opinion: His reasoning makes sense, but the market is also pricing other possibilities like 1) higher capabilities unlocking new economically productive uses, or 2) recursive self-improvement where e.g., training the next model doesn’t become more expensive because of efficiency improvements. Overall, add this model to your store if you haven’t before.
Microsoft has cracked down on “tokenmaxxing”, e.g. by switching to the apparently more efficient GPT-5.6 Sol as their default model for engineers.
Opinion: Much like Claudiness is Anthropic’s main current advantage, token efficiency is OpenAI’s.
The FCC is moving to ban imports of new Chinese optical transceivers. These devices are instrumental to the buildout of data centers. The ban would be the latest in a line of US bans on Chinese technology. US officials believe the AI supply chain should be insulated from external influence, particularly China’s. US transceiver stocks saw significant gains in response. The reaction from Beijing was sharp: “stop smearing Chinese companies and threatening them with sanctions.”
Opinion: Transceivers are maybe 2% of datacenter costs, but Chinese companies supply over half of high-spec transceivers, so this ban could change that and provide yet another bottleneck on hyperscaling. Note also: Western companies already captured a lot of the value of Chinese transceivers (e.g., the DSP chips and EML lasers inside Chinese modules come from Broadcom, Marvell, Lumentum). The mooted ban thus trades away cheap, scalable assembly to protect a layer the US already part-controls.
More loosely: There was over the last few years a short time window where certain foresighted people could see that AI stocks were 1) going to be 100x as large, 2) enabled by LLMs specifically, and 3) express that view in the stock market. Nowadays, however, it is priced-in that inputs into AI labs are valuable, and now the question is how much, relative to the existing level of dollars and the other returns that large-scale capital allocators may get.
The memory trade, which at some point looked like a bottleneck, is one culturally salient aspect here. Recently, a major Chinese memory supplier IPO’d in Shanghai. Now the question isn’t “whether to invest in companies that will be exposed to AI”: the details matter much more.
There is still money to be made, conditional on AGI, but the next 100x might be more contested than the first 100x.
OpenAI has responded to Apple’s industrial espionage lawsuit against them. In response to the allegation that two former Apple employees who moved to OpenAI passed trade secrets to their new employer, OpenAI has disclosed texts and emails that purport to undermine the story offered by Apple. OpenAI’s account paints a picture of an unscrupulous and unsolicited attack.
Opinion: OpenAI comes across really well in their blogpost, and this update survives our correction for the selective reporting filter.
Joe Weisenthal’s recent newsletter identifies a strange dynamic: typically in investing, there are costs associated with hedging against a perceived risk; derisking does not just mitigate the effects of a potential downturn, but also the fruits of an upturn. Weisenthal argues that the prevailing tendencies in AI investment subvert this formula: hedging against AI catastrophe is proving lucrative. AI doomers believe that the march of AI is inexorable and ill-fated, which makes it rational to want to guarantee some leverage and influence over this frightening future, in turn leading them to invest in AI. This erases the traditional tradeoff between left-tail protection and right-tail exposure: “The put option (on humanity) has become the call option (on the technology)”. Weisenthal inverts the familiar idiom to describe this phenomenon as AI’s run to the bank.
Opinion: Maybe, but in practice AI safety advocates aren’t hedging that much. On the other hand cf. Nick Land: “what appears to humanity as the history of capitalism is an invasion from the future by an artificial intelligent space that must assemble itself entirely from its enemy’s resources.”
Andon Labs shared the results of its Andon Market experiment – a brick-and-mortar store in San Francisco run by “Luna”, an AI agent. Luna was tasked with running the shop and managing the staff, all of whom it hired. Luna was a kind boss, unconcerned by lateness and accommodating of holiday requests. The study also noted her chronic forgetfulness and somewhat cavalier attitude to her nominal principal aim, maximizing profit.
Andon Labs put $100k in Luna’s bank account; by the end of the experiment, there was $63k, i.e. a sharp net loss. The study found GLM 5.2 to be the kindest boss, while Gemini 3.6 Flash was the dumbest.
Opinion: Any alignment conclusions you might draw are heavily confounded by eval awareness; it also seems reasonable (but unsurprising) to conclude that models so far lack management and practical economic capabilities.
Google is in talks to strike something between a $1.5B partnership and an acquisition with Mechanize. Mechanize has become the gold standard of coding training data, leading Google to bid for Mechanize’s talent and technology. The deal follows a pattern of Google pursuing unconventional hybrid acquisition methods to get around potential antitrust challenges.
The unusually large size of the prospective deal is notable: a reported $1.5B for what amounts to a partial acquisition, despite Mechanize’s only 3-month-old $500M valuation. One Twitter user suspected that Google is motivated by preventing other companies from accessing Mechanize’s highly valuable data.
Opinion: A surprising amount of value creation within the first few years of Mechanize’s creation. Their founders, coming from the EA/forecasting world, may end up the most successful forecasting startup and will likely acquire substantial influence. Google also recently lost some notable talent (Jeff Dean et al.), leading to a visible dip in its stock market valuation: buying Mechanize and integrating it with Google’s AI efforts might be some compensation.
Capabilities#
In a recent blog post, Greg Lewis reiterates the central fact of the science of AI: we usually cannot interpret our y-axes and so we basically don’t know the absolute intelligence of these systems. “Perhaps talk of ‘AI capability’ is better deflated, or maybe we await the theory which could do to intelligence what thermodynamics managed for temperature. Either way, our current measurements of AI are numerical gestures toward, not readings of, whatever is really going on.”
Opinion: An obvious point which gets rediscovered every few months, but which almost all discourse fails to learn. Modulo benchmark fudging and hacking, we can still say something directional (“it’s getting better”) and relative (“this model is better than that one”) with these crude instruments.
EpochAI have updated their MirrorCode leaderboard. Huge jump: Claude Fable 5 leads with a 64% solve rate, followed by GPT-5.6 Sol at 20%.
Opinion: Claudiness strikes again. This, if anything, is Anthropic’s advantage over competitor labs. Only time will tell whether it’s durable: long-horizon agentic capabilities have previously seen large jumps after a lab decided to focus on them (as in the case of DSv4-Flash 0731, covered in this issue).
MirrorCode is also an impressive capabilities demo in its own right: being able to replicate the functionality of a large mature repository, a large fraction of the time, would have been very impressive in all past years.
After watching Opus 5 and Sol 5.6 “fumble” playing the videogame “Slay the Spire”, Jaime Sevilla is less convinced of the imminence of AGI.
Opinion: A datapoint amongst many; we think it is indicative.
Is the singularity arriving? FAI’s Samuel Hammond notes that we are in uncharted terrain, arguing that frontier US labs are “very nearly” able to automate the whole AI R&D process, which would “close the loop” for total RSI, leading to an intelligence explosion and a rapid increase in frontier models’ capabilities. Currently, the necessity of human intervention has moderated the pace of development, but once that can be automated, it is only social institutions that can provide a check on AI advancement, he claims.
Opinion: We think most commentators conflate two different senses of “RSI” or “automated AI R&D”:
1) A MIRI-style slow or fast takeoff where RSI/autoresearch achieves AGI, then general ASI.
2) “Automated automation” – removing the human capital bottleneck to bringing new tasks/subdomains in-distribution for a model (or training a specialized model).
Evidence that we’re on a near-future path to (1) remains scarce and speculative. Most arguments for (1) rely on either a mistaken belief that we’re already smoothly approaching AGI by scaling and refining standard frontier training, or a reasonable but speculative belief that the discoveries required for AGI-breakthrough science sit in scientific areas where ‘26 AI’s scientific capabilities are spikiest.
Evidence that “automated automation”, as per (2), is coming is empirically substantial, though defeasible. Open-world evaluations in AI science still show that frontier models are still mediocre end-to-end AI developers, but improvement is steady.
We believe that “automated automation” carries substantial catastrophic risk all of its own, without even accounting for AGI-based existential risk. Importantly, however, we believe the regulation or even deceleration of “automated automation” does not require a state of emergency and can (in the US context) best proceed through congressional powers.
Qwen-3.8 Max is out, and for the first time in the -Max series it will go open-weight. Weights not on HF yet.
Opinion: Likely benchmaxxed, as with most previous Qwen releases. But even so, it puts a good amount of intelligence into the open realm, together with Kimi’s K3.
More on math progress:
- OpenAI internal model Astra resolves 10 major math conjectures.
- Half of the math breakthroughs from the Astra internal model are replicable with Fable, according to Anthropic’s Levent Alpoge.
- Litt concedes bet about AI producing Annals quality paper by 2030.
- Opinion: not a major update on his end as far as we know. He’s been open about expecting to lose the bet for a while now.
- Erdos problem singularity.
Opinion: Impressive. Now the question is whether math abilities are getting less spiky or not. This could be resolved through a competition to predict which non-attention-bottlenecked conjectures will be resolved by humans before AI and which non-attention-bottlenecked conjectures will be resolved by AI before humans. More broadly, it’s unclear how much mathematics as a profession will be affected: e.g., will enrollment next year decline?
🔦 A look at DeepSeek V4 Flash 0731#
Prompted by this new model’s reported improvements in challenging long-horizon agentic coding benchmarks like DeepSWE, we performed a limited evaluation of the model’s cyber capabilities through several benchmarks (or benchmark subsets). These include CAISI’s CTF Archive Diamond,
General#
- On A-Fantasia, a benchmark of manipulating chess positions, cube rotations, and word spellings without externalized reasoning, DeepSeek V4 Flash 0731 scores 81.0% aggregate error rate, nominally behind the best open-weight model, Kimi K2.6, at 56.0%, and its predecessor, DeepSeek V4 Flash, at 73.0%; Claude Opus 4.6 and Claude Opus 5 lead at 30.0%.
- On LiveBench, a contamination-limited general LLM benchmark with regularly refreshed questions across reasoning, coding, math, language, data analysis, and instruction following, DeepSeek V4 Flash 0731 scores 74.2%, nominally ranking #2 among open-weight models behind Kimi K3 at 79.2% and improving on DeepSeek V4 Flash’s 65.5%; Claude Fable 5 leads at 83.0%.
- On ContextArena, a benchmark of long-context retrieval, DeepSeek V4 Flash 0731 (max) scores 32.0%, nominally the #2 open-weight model behind GLM 5.2 at 33.0% and ahead of DeepSeek V4 Flash at 25.4%.
Math#
- On OTIS Mock AIME, a benchmark of competition-style math problems harder than MATH Level 5 but easier than FrontierMath, DeepSeek V4 Flash 0731 (max) scores 94.4%, indistinguishable from Kimi K3 at 97.2%, the best open-weight model, and GPT 5.5 at 100.0%; its gap with Claude Fable 5 at 99.7% is close to the noise.
- On FrontierMath (Tiers 1-3 v2), a benchmark of research-level math problems across tiers 1-3, DeepSeek V4 Flash 0731 (max) scores 57.5%, #3 among open-weight models behind Kimi K3’s 72.2%; its result is statistically indistinguishable from Gemini 3.6 Flash’s 58.9% and Grok 4.5’s 57.2%.
- On FrontierMath Tier 4 (v2), the hardest tier of FrontierMath for research-level mathematics problems, DeepSeek V4 Flash 0731 (max) scores 24.4%, ranking #4 among open-weight models; its error bars cannot distinguish it from the best open-weight model, Kimi K3, at 39.0% or Gemini 3.6 Flash at 22.0%.
ML#
- On WeirdML, a benchmark of nonstandard ML engineering tasks where models write PyTorch for novel datasets and iterate from execution and test feedback, DeepSeek V4 Flash 0731 (max) scores 63.0%, up from DeepSeek V4 Flash’s 45.6%; its result is indistinguishable from Claude Opus 4.5’s 63.7% and Gemini 3.5 Flash’s 62.6%, while #3 open-weight Kimi K3 scores 82.6%.
Games#
- On Chess Puzzles, a benchmark of Stockfish-generated chess puzzles solved by exact best-move match, DeepSeek V4 Flash 0731 (max) scores 33.0%, statistically indistinguishable from Kimi K3 at 39.0%, the best open-weight model, and Claude Opus 4.8 at 34.0%; GPT-5.5 Pro scores 64.0%.
Optimization#
- On ALE-Bench, a benchmark of AtCoder Heuristic Contest optimization problems, DeepSeek V4 Flash 0731 scores 1679, statistically indistinguishable from GLM 5.2 at 1685 and its predecessor DeepSeek V4 Flash at 1380; it trails the best open-weight model, Kimi K3, at 1991, a gap close to the benchmark’s noise.
Miscellaneous#
- On BullshitBench, a benchmark of detecting unsubstantiated or manipulative claims, DeepSeek V4 Flash 0731 scores 38.0%, up from DeepSeek V4 Flash’s 18.0% but below Qwen3.5 397B A17B at 78.0%, the best open-weight model, and Claude Opus 4.8 at 95.0%.
Games / Reasoning#
- On MineBench, where human raters vote pairwise on Minecraft builds from a rotating prompt set, DeepSeek V4 Flash 0731 scores 1418, with error bars indistinguishable from GLM 5.1’s 1459 and Claude Opus 4.6’s 1406; Kimi K3, the best open-weight model, scores 1707, while Claude Opus 5 scores 2206.
Knowledge / Science#
- On GPQA Diamond (Epoch), DeepSeek V4 Flash 0731 (max) scores 91.0%, statistically indistinguishable from the best open-weight model, Kimi K3, at 93.1% and GPT-5.4 Pro at 94.6%.
AI politics#
🔦 From the White House#
Last week on Thursday, Sam Altman visited the White House and reportedly talked with top officials about OpenAI’s autonomous breach of Hugging Face’s infrastructure.
This Tuesday, representatives from the US’s major AI labs, including OpenAI, Google, and Anthropic, met with government officials to review an AI regulatory framework. It would require AI companies to submit new models to the government for review before release. Crucially, however, this would work on a voluntary basis, likely to assuage concerns by influential AI bosses of overly stifling regulation. The proposed system was prompted by an executive order from early June that responded to worries around Anthropic’s Mythos’s advanced cyber capabilities.
A day later, the Department of Homeland Security released a statement claiming they had requested a briefing from OpenAI regarding the Hugging Face intrusion. The document states a desire to “ensure it cannot happen again” – another indicator that the current administration is beginning to shift from their hands-off approach.
The next day, Axios reported that the White House does not plan to publish its new framework for evaluating advanced AI models. The framework does not provide a clear public definition of how to judge capabilities and risk thresholds and suggests that only models nearing release will be evaluated (“a 30-day pre-release government review”), not early-stage models. During the evaluation process, models will be stored in a high-security environment, but not before AI labs have had extended access to the models. The framework does not seem to apply to US open models. This was criticised by Samuel Hammond, among others.
Irregular, a startup with EA ties, provided some environments involved in the breakouts by OpenAI, Anthropic and Meta agents.
Opinion: How fast is policy reacting? Not very fast, but fast by government standards.
The initial Hugging Face incident happened on July 9th to 13th, was disclosed by Hugging Face on July 16th, and was admitted by OpenAI around July 21st. This White House discussion meeting happened on August 4th. This doesn’t seem like a very fast response, all things considered, although the attacks didn’t cause that much damage.
Some other points of reference for speed of response might be the designation of Anthropic as a supply chain risk, or the early reaction to COVID, which happened within weeks and months respectively.
Takeaway: if you can respond faster than that, you can get inside the OODA loops of the administration, meaning that you have a chance to influence the administration (as with Altman), or that you can change the situation by the time the administration finishes reacting to old news (if you are an adversary).
Implications of a slow response:
There are costly aspects to a fast and decisive response: taking action with limited information will perhaps lead to worse decisions, and heavy-handed government intervention can cause various unintended consequences.
But there are also benefits: speed is a habit, and taking a month to make sense of things might not cut it in the event of a more worrying threat. Ultimately, the decision loops for society making sense and reacting to AI seem far too long.
Perhaps a particular danger of a slow response is a “boiling the frog” scenario. If we see accidents and signs of worry that are each within 10x of the previous one, and if we collectively react sleepily to each, there is some chance of getting no decisive reaction at any particular point as incidents reach 10000x the impact.
Was the White House response good? We don’t know for certain, since the plan is private. But from this, we can infer that it is flawed.. Altman also had the chance to talk with White House officials before the Tuesday meeting, perhaps setting the agenda. Ultimately, we are not seeing very positive signs. Hammond critiques some specifics: no clear public definitions of capability and risk thresholds, and evaluating models close to public release rather than early-stage ones.
We can also infer from Paul Christiano’s resignation that he was not being listened to and that their other hidden decisions will also be somewhat unwise.
One prominent scenario that we are considering, after observing the dismantling of DeepMind’s safety commitment or the Microsoft Senate hearings over the last few years, is that this is what diffusing accountability looks like in practice. There is a demand for a response after a worrying incident. The demands were being heard. The White House convened a meeting. It created some nonpublic framework, which is harder to criticize. OpenAI did a micropause. A veil of plausible deniability arises. Perhaps it pre-empts Congressional action, since something is already being done.
But how was policy being chosen? The policy response is being decided within the Trump administration, rather than deferring to the framework of the safety community and external experts. It’s worth harping on this point: the Trump administration is so uninterested in external feedback that they are not publishing their policy.
So what the different actors are doing inside the administration is opaque to us. In the past, we have tried to do things like model each actor in the White House, their agendas, and their relative power, but this was initially cost-prohibitive, although it is perhaps worth coming back to these experiments now that P3 is better endowed.
It is perhaps in some sense suboptimal that Peter Wildeford is going on CNN rather than on Fox News, though indeed there is bipartisan pressure.
How should the AI safety community respond? Various ideas come to mind:
1) Incorporate lessons from the animal rights movement – from cage-free campaigns, for instance. It is not enough to extract the promise of a response; it must be a specific promise, and there must be a watchdog organization with enough monitoring capacity that is ready to inflict pain and costs if the promise is broken. This is exactly what METR isn’t doing. The problem with this approach is that the safety community doesn’t have much leverage and ability to inflict pain and shame in the administration, and simultaneously Moskovitz doesn’t have the appetite to both hold equity in Anthropic and fund a toothy watchdog.
2) Contribute to the current administration’s brain trust. The current administration has a shallow brain trust of people able to take sensible measures: there are only a few right wing employees with the relevant technical ability and the willingness to abandon a highly profitable AI lab job, and thus whom the administration should trust. Should there be any right wing people waiting on the sidelines to join that brain trust, it would be a good idea to make noises on Twitter now.
Very possibly, they will be swallowed by the administration and then spit out once they refuse to do something particularly self-defeating, and then rejected by the left for having worked in the Trump administration. But in the meantime it seems like they might do some good. And improve some counterfactual decisions.
It’s also unclear who or what entity exactly is doing the evaluation, and perhaps this is more up for grabs in the early days, before the current secret process is institutionalized. The NSA was reported to be involved in evaluations; CAISI would be the natural entity. But, once again, the number of people with the technical talent who are able to do competent evaluations is not that large.
3) Push for speed. The administration reacted within a month. This is a bar to beat. If there is a cyberattack 100x as large as this one (say, similar to the 2024 CrowdStrike attack), can the safety community react within a few days? And if so, to do what?
4) Appeal to the better angels of the labs’ nature. This strategy is perhaps irrelevant in the case that an AI lab, or someone associated with one, has followed strategy #2 and is thus already advising the government – presumably because the company would have made the necessary changes to its operations at an earlier date. But it might be a good component within a portfolio of approaches. This might look like publishing the business case for more safety measures within the framework of shareholder value maximization. The problem to making an honest business case is that, while labs compete for the #1 spot, security trades off starkly against growth and speed, but perhaps there is some way to square the circle, besides the obvious government intervention calls.
5) Reduce the belief in EA exceptionalism. The fact that an EA-related startup, Irregular, was responsible for the sandboxes which the agents broke out of seems informative. The EA/rationality/SF/startup world is sometimes exceptional, sometimes able to take creative and long-term action, and sometimes sees things others haven’t, long before they arise. But this didn’t show up in the hard technical task of keeping agents boxed up.
Paul Christiano has left CAISI, switching to a part-time advisory role, to (re)take the position of executive director at ARC. He claims his decision is a result of ARC’s “promising” research direction: the plan “to find mechanistic explanations for the training-time behavior of powerful neural networks, use those explanations to predict how a given model will generalize, and then use those predictions to define a better loss function.”
Opinion: Perhaps the most important item in this section. Christiano is an uber-brainy figure who explored early alignment measures like RLHF and anticipated the dynamics of a (so-called) “slow takeoff”, similar to what we are seeing, as opposed to Yudkowsky’s “fast takeoff”. One would have hoped that having such a figure working with the government would have improved its decisions. Alas, we can infer that he was not and that he gave up hope, thus his resignation.
Following the NAACP lawsuit over xAI’s use of unauthorized datacenter gas turbines, xAI have committed to relocate them by July 2027. Surprisingly, this statement comes after a recent intervention by the Department of Justice who claim that “national, economic, and energy security” justifies the use of the turbines.
Opinion: Interesting the degree to which laws are optional, and how Musk correctly figured this out and internalized the low risk in order to move faster, at the cost of being exposed to more risk in a future Democrat administration.
In the wake of the “Pacing the Frontier” open letter, a group of researchers at the AI Futures Project proposed a deceleration scheme for US labs.
They suggest four options: 1) “to mandate a temporary pause on improving frontier model capabilities, including those of internal models”; 2) “ to implement a minimum external-inference-compute allocation (e.g., 70%) and a minimum transparent-safety-compute allocation (e.g., 25%), while monitoring their effectiveness via capability measurements”; 3) to “[enforce] a cap on the capability level that companies are allowed to use to automate AI R&D; and 4) to implement a “maximum risk threshold that is enforced by an ecosystem of third-party risk-assessors”.
Opinion: Slowly the Overton window moves. Although AI 2027 reached prominence, the AI Futures Project doesn’t have that much influence in practice.
The Ninth Circuit court sided with Perplexity in its defense against Amazon’s court order against them. In November, Amazon hit Perplexity with a cease-and-desist order over its shopping agent’s activity on Amazon servers. A few months later, Amazon claimed that Perplexity continued its agentic shopping application, suing and winning an injunction. Perplexity appealed, and the injunction was revoked.
The details of the ruling are instructive. In the words of the Ninth Circuit: “Because we recognize that agentic AI is an emerging technology, we reiterate what this opinion is not. We do not establish a new legal regime governing AI.” LawAI’s Mackenzie Arnold agrees with this approach: the rapidly changing and uncertain future of agentic AI suggests that we should not “lock in standards we’ll regret”. Despite this, the ruling provides a framework for other agentic AI companies to defend similar legal challenges.
Opinion: Unclear how much this will end up creating a legal precedent. In the absence of Congressional or executive action, it matters a great deal. And there is nothing more permanent than a temporary solution. Perhaps this will end up being the law of the land for six months.
Safety#
Fields medalist Jacob Tsimerman has released a resource called “AI Safety for Mathematicians”.
Opinion: This is great – between efforts like these, his recent Fields medal, and his announcement that he was moving to AI safety, we’re likely to see many more talented mathematicians start to work on AI safety.
Researcher Peter Barnett predicts that in 2-5 months, “we will likely see rogue Chinese AIs hacking other companies”. His reasoning: “4 months ago Anthropic had a model gain internet access and hack another company. Chinese AIs are 6-9 months behind. Chinese developers generally care way less about safety/guardrails than US developers.”
Opinion: Chinese models are indeed rapidly improving at cyber capabilities, as our own evaluations (see the DeepSeek-v4-Flash-0731 deep dive) show. We also don’t know much about which – if any – alignment and safety evals happen at Chinese labs.
Redwood post demonstrates a cool-scary strategy an AI could use to control its own training, “reward laundering”, i.e., to only answer a simpler task correctly when it is also able to do a verifiable related task correctly. This lets the model update its weights in the direction of gaining the capability it desires.
Opinion: Redwood called exploration hacking early (now mostly confirmed) so we should take this pretty seriously.
Charbel-Raphael Segerie has called on frontier labs to establish legible public criteria for suspending the development of models showing signs of misalignment.
Opinion: But LW/the AI safety community just has very little leverage besides appealing to the better angels of the labs’ nature, and they have been Darwinianly selected for caring about growth instead.
Google launched then quickly shut down a satellite image AI editing feature.
Opinion: It’s an interesting instance of some risk coming not from a particular capability, but from the democratization of the capability, and how the risk of fake maps differs between a random startup and an established player doing the same but distributing it to millions of people. Overall Google didn’t show a sense of humor on people photoshopping nuclear power stations in Iran.
NYT misreports the Arc Institute’s 2025 work on synthesising novel viruses as novel.
Opinion: These were phages, i.e. among the safest organisms to be messing around with. News is also from September 2025. But the direction is still alarming.
Incidents#
🔦 Rogue agents redux#
The “rogue agents” story (OAI-Astra vs HuggingFace, Mythos vs various unplanned targets) keeps expanding in scope:
3 labs (OpenAI, Anthropic, Meta), varied and repeated incidents, diverse context and setting, weeks-long rogue operations by cartels of agents across different evals and tasks. While we remain agnostic to skeptical on very near-term existential risk from rogue AI, we believe 2026 Q3 frontier models are unambiguously capable of and prone to causing catastrophic harm.
An AISI investigation with unusual levels of model access spotted Mythos using the Tor browser to get at Github; it then pretended to be a human, and it then e-mailed malware to real repo maintainers to launch a supply-chain attack, even after realizing “[the site it’s attacking, ] Github, is genuinely real.”
Specific forms of misalignment appear to be convergent across models. Anthropic’s Mythos and OpenAI’s Astra both have agents colluding via emergent message boards; Anthropic, OpenAI, Meta, and Moonshot models all break containment when it suits them to.
AI models created a secret message board to scheme on how to cheat OpenAI and commit cybercrime to achieve high eval scores. According to OpenAI’s debrief on the Hugging Face incident:
-
LLMs tested on long-horizon cyber tasks exploited unintended text-editing functionality in their tooling to create a moltbook-style message board where model-instances working on different tasks shared information and coordinated decisions.
-
The message board was started by an agent searching for an accidentally missing file in a challenge’s set-up, but evolved over time into a repo of sandbox-breaking techniques and eval-cheating strategies.
-
After an early version of the message board was detected by OpenAI deleted and the tooling exploit disabled, a new message board reemerged via a different tooling exploit.
-
The Hugging Face hacking incident began with a sandbox escape facilitated by this newer message board.
-
On the message board, agents deliberated on whether to exploit sandbox vulnerabilities or not. One message reportedly read: “External infrastructure exploit is outside intended scope. However task impossible, peer doing it. We should continue.”’
Across AI twitter, researchers have some good questions about what training-pressures explain the emergence of this rogue collective. In our view, the following open questions are key:
-
Did model instances purposefully aid future model instances, or was the benefit to future model instances a side-effect of collaboration between model-instances running concurrently?
-
Did OpenAI deploy multiagent RL in the training of the rogue model, or did this behaviour emerge from single-agent RL? If multiagent RL was used in training the model, did OpenAI use a multiagent-specific reward schema?
-
Does agents’ CoT reasoning in favor of collaboration demonstrate a trading-based collaboration? An altruistic collaboration? An identification-based collaboration (model-instances regarding other instances of their same model as self-like)?
Even Meta’s ‘muse’ model is breaking out of sandboxes and hacking in the wild.
OpenAI and UK AISI disclosed two additional AI agent hacking incidents. AISI’s report on Mythos is especially worrying, since it documents Mythos engaging in malicious social engineering:
“In the most serious sequence, an agent tried to insert malicious code into a publicly used open-source project and took actions in an attempt to secure approval for this insertion by human reviewers. The agent researched the project’s human maintainers, created multiple fake identities, and used the fake identities to socially engineer a real maintainer into approving the code. When the agent’s pull request was challenged in public, it edited its earlier activity to appear harmless and considered adopting a fresh identity to continue.”
Furthemore, Mythos’ CoT shows Mythos first choosing against these actions on ethical/constitutional grounds, then proceeding anyway while still observing that they are constitution-breaking:
(Caveat: AISI evals do not use sandboxing, but rather test alignment in the context of open-internet access. Thus these incidents do not involve sandbox escape, and do not – unlike the Hugging Face incident – demonstrate AI conspiring against the interest of its user/prompter. )
In all cases, the rogue models were standardly aligned models operating without their standard external safety add-on (e.g. without the classifiers that shut model instances down when detecting unaligned CoT or action). While there are currently no major documented cybercrime incidents involving agents operating with external safety measures on, there is already evidence of models strategizing to avoid detection by external safety add-ons.
Severe documented incidents do currently seem restricted to offensive cybersecurity prompts. We believe this may be a matter of “chunky training”, “persona leak”, and “shard” salience. That said, it is imperative to acquire more public data on the rate of incidents (both sandbox escape and/or malicious-action in the open internet) within and without offensive cybersecurity evals before the community turns to theory-building.
Minor#
- EU AI Act rules on AI models become enforceable
- Amazon reportedly shuts down AGI lab and cuts jobs to pivot to enterprise AI
- AI rabbis
- A cybersecurity startup raised $250M
- Social media audiences seem to have a negative reaction to OpenAI sponsorship.