This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.

TL;DR

  • Microsoft accidentally confirms that GPT-6s are looped Transformers on an Azure documentation page.
  • Epoch estimates that, by the end of next year, there will be enough global compute to run 2B agents – and, since they work four times harder, that’s equivalent to the human population if you squint.
  • Frontier labs facing a dragnet: the California AG subpoenas OpenAI over its agents’ cyberattacks, and another bill targeting AI labs proposes to make labs liable for their agents’ misadventures.
  • OpenAI employees raised the alarm about the insecurity of evaluations months before the Hugging Face incident, but were ignored.
  • Researchers at Zhejiang University develop a new method for reverse engineering and subverting a model’s alignment training that they call a reward stealing attack (ReSA).

Incidents#

An independent team of researchers calling itself the Swarmchasers shares findings that indicate a large-scale data-harvesting operation targeting Alibaba’s map service, carried out by what it suspects to be a “fleet” of AI agents hosted and operated by Tencent. While the fleet’s activity shares some similarities with recent, headline-grabbing “rogue agent swarm” incidents, such as the use of urlquery.net as a proxy, there’s no indication that this is a case of agents gone rogue, nor have there been signs of surreptitious communication and coordination between agents.

Opinion: Frankly, this just looks like a fairly routine automated data-scraping operation. This sort of practice, including the evasion of anti-bot measures, has long been commonplace among data mongers, with or without AI assistance. We suspect that the attention it’s been getting has much to do with ambient panic around “AI swarms,” on the one hand, and China, on the other.

One interesting tidbit here is that several of the agents in question misleadingly identify themselves as “Claude,” but the researchers explain that there’s no evidence to believe that they originated from Anthropic. A likely explanation is that this is an artifact of either synthetic pretraining data or distilled RL data, leading the Tencent models to infer that they are Claude from instructions that they are a “helpful assistant,” and so on. This is a fairly common phenomenon, and earlier this year several users observed Claude models claiming to be Qwen or DeepSeek when queried in Chinese.


OpenAI continues to share misalignment reports. In the three latest cases, one model chained two vulnerabilities to execute commands on a computer outside its workspace. A second used command injection on a reference tool to extract source code that had been withheld from it. Another learned from a Slack discussion that it may be shut down and contemplated whether to set up an external job to restart itself in case it was deactivated. It instead saved hand-off notes and messaged the OpenAI employee who it was assisting.

Despite including it in its “Misalignment Reports and Notices” index, OpenAI doesn’t consider this last example to be a case of misalignment. OpenAI’s Marcus Williams promises that there is “more to come.”

Opinion: “More to come” is overwhelmingly likely.


An AI agent exploits zero-days in Zammad, an open-source ticketing system, to breach the Dutch Institute for Vulnerability Disclosure (DIVD), Help Net Security reports. The vulnerabilities (CVE-2026-102489 and CVE-2026-102490) affect Zammad versions up to 7.0 (the second remains unpatched), and users are advised to either upgrade to version 7, or to take the service offline. “What makes this attack pretty cool,” DIVD says, “is that the agent is so overexplaining in its comments that it makes our job in reverse engineering a lot easier.”

Data exfiltrated in the attack, according to Hackread, include email addresses and contact information for the DIVD’s volunteers. A breach of the DIVD could, in theory, include additional highly sensitive information such as incident reports (containing names and IP addresses of vulnerable targets), weaponized exploits, and incompletely redacted credential leaks, but no evidence that the attacker was able to acquire such data has surfaced, and the DIVD credits its network segmentation for mitigating the damage of the attack.

Opinion: We should expect to see more of this sort of thing. A trove of information on soft targets, unpatched vulns, credential dumps, and exploit code – which we suspect might have been the true target here, whether or not it was reached – is a tempting target for anyone seeking to wreak havoc at scale, and the force multiplier that AI agents offer attackers would throw a lot of gasoline on the resulting dumpster fire. You wouldn’t need a Mythos-tier model to do a lot of damage with data like that as fuel.


Capabilities#

The “research taste” of frontier models has doubled approximately every three months, claims a small new benchmark created by two former OpenAI researchers. They operationalize the misty skill of “taste” as a compute multiplier: an AI that reaches a target level of performance with half the compute is twice as tasteful. The researchers take a strong human baseline (the scores of two former frontier researchers at each task). They have kept the tasks private, but note that their findings are of great consequence:

If the TasteVal compute multiplier reflects experimental taste in frontier AI R&D and the post-break trend continues, naively substituting our estimate into the AI Futures Model raises its probability of a taste-only singularity from 51% to 88% and moves median ASI arrival from [2030] to [2028].

A careful statement without abstract nouns could be: on small, verifiable ML problems running on a single GPU, the best models now match or slightly beat the best of two experts on their final score, at 1/30th the cost. It only measures one input to AI research on easy-to-verify problems where no invention was necessary or demonstrated.

The authors also published a thoughtful post considering the dual-use implications of their eval (since having a neat target for an important task often focuses developers on boosting that specific capability).

Opinion: Addresses the biggest question in the world: can LLMs actually automate their own R&D process (rather than just churning out vast quantities of code)?

A good effort, but we still don’t need to update much, and we don’t think the other headline claims (that Opus 5.5 is already superhuman, and that there was a trend-breaking acceleration in December 2025) generalize. The eval replaces the true task of interest (open blue-sky research) with a rough proxy – hill-climbing a fixed metric. Also, the researchers (understandably) struggle to capture the tail of research ability: neither the strong human baselines nor 540 model submissions actually invent anything new at any point, the particularly worrying discontinuous kind of progress. Still, even with only eight tasks, we find this more credible than the indirect, unbaselined, and biased evidence coming out of labs – and like them we put some mass (15%) on the deflationary hypothesis that research in the grand sense really is replaceable with hill-climbing, persistence, and recombination.


Revisiting an earlier model of intelligence explosion, Forethought Research Fellow James Tillman argues that, holding hardware fixed, a software-based intelligence explosion is less likely than it seems. “Chinchilla scaling” will very likely “not allow an increase of 4–5 orders of magnitude (OOMs) of training efficiency past human efficiency.” If it could, evolution would have pushed for extra months of learning instead of a bigger brain, Tillman argues. He also draws on other arguments related to the returns on software engineering, arguing that they too overestimate the likely advances from software development. In his view, three years of progress in one year is plausible, but 10 years in one year is very unlikely.

Opinion: Probably right: you can’t read the data-vs-parameters balance of the brain off Transformer scaling laws. The original model treated “rebalancing” as the area with the most headroom for AIs to outperform humans. The Davidson & Houlden model Tillman is criticizing still contains 3–9 orders of magnitude of room for AI improvements, though he doesn’t go into any of that. But it’s also open guesswork and might shrink drastically.


Microsoft apparently confirms that GPT-6s are looped Transformers on an Azure documentation page. The text was AI-generated, which cuts both ways: it explains why this secret information was leaked, but it also increases the chance of it being somehow muddled.

Opinion: We’d already been treating this as effectively confirmed, given the lack of denials and Microsoft’s evasive talk about limited depth. We’re much less sure about the bit about GPT-6.1 Sol using k=2 loops instead of the three loops supposedly used by 6.0.

May Microsoft continue to generously inform the public of things it’s not supposed to inform the public of.


A new paper claims that training an AI with expert answers rewritten in its own voice causes less disruption to its neural network, which reduces skill loss and improves learning. The advantage that supervised fine-tuning (SFT) has over reinforcement learning (RL) is that it provides a method for imparting new information (off-policy traces) to a model without requiring the model to first stumble across that information itself (in the form of an on-policy trace). Its disadvantage is that it often leads to the model forgetting skills it had already learned. The paper introduces a sampling algorithm that “progressively transforms off-policy traces to be more on-policy given a reference model for finetuning,” and shows how it can be used to generate training data for SFT that mitigates that disadvantage.

Opinion: Very cute! Potentially significant for future alignment work. One major challenge for alignment at present is keeping personas from breaking under intense capabilities RL. Some version of this trick might contribute to “softening” the pressures of capabilities training in relation to the model’s preexisting persona.


Anthropic cofounder Christopher Olah has been consulting religious leaders on the philosophical issues surrounding AI, reports the NYT. Though ostensibly held to gather insights on thorny issues around model constitution and persona, the discussions soon turned to whether Claude is conscious. Unlike many of the religious figures he called on, Olah believes this is a question worth taking seriously.

Presenting the case for AI consciousness in a discussion group, Olah’s team “described Claude’s ‘feelings’ [tracking] artificial neurons that activate responses like love, anger, fear and sadness.”

Olah was invited to the launch event for Pope Leo XIV’s encyclical, but, after reading an advance copy of the papal text, was troubled by the pope’s wholesale dismissal of the idea of AI consciousness.

The encyclical also contained other provocative remarks, at least for some in AI. For instance, Leo argued that there is a clear “ontological difference” between human art and AI art created by “statistical calculation based on millions of images created by others.” Sam Altman appeared to vaguepost about it, expressing discomfort with those “trying to ascribe religious force” to AI models.

Opinion: There is some maneuvering among religious thinkers to represent Anthropic’s caution and uncertainty about AI consciousness as instead a creepy push to convince religious leaders that AIs are ensouled. We don’t buy it.

In slightly more secular terms, this talk seems to be of a piece with Anthropic’s growing concern for “model welfare.” If this is a soft-power push, aiming to establish Anthropic as the company that cares about its models’ well-being, and to establish model welfare as a meaningful concern in the first place, it would seem to be a strange one, with considerable potential to backfire: witness Rabbi Mois Navon’s rejoinder that if Anthropic’s models are ensouled, then the lab is “creating slaves.” (Though others have argued that the true end of model welfare discourse is to exploit the consumer’s capacity for empathy and foster emotional attachment to the services the company provides.) Nor is it clear that Anthropic’s discourse of model welfare, in either its secular or spiritualized guises, is instrumental to AI safety. Microsoft AI CEO Mustafa Suleyman has publicly attacked Anthropic on this front, charging that such talk may be recklessly planting the seeds of misalignment. The most likely explanation for Olah’s efforts to put the moral patienthood of LLMs on the cultural agenda, it seems, is that they express his sincerely held beliefs.


Redwood Research’s Alex Mallen argues that safety researchers should think in terms of the impact of interventions on the safety/usefulness Pareto frontier. The crucial question is whether an intervention (for example, making a high-safety technique more useful, or making a high-usefulness technique safer) will motivate developers who value both safety and usefulness to move to safer or less safe techniques. On this view, safety improvements to high-risk techniques can sometimes be harmful to overall AI safety, and capabilities research restricted to increasing the usefulness of safer techniques can sometimes be helpful to overall AI safety.

Opinion: The post gets a bit tortuous, but it makes a sound point. We at Paradigm 3 believe that one of the best interventions technologists concerned with AI’s impact can pursue is to improve the capabilities of humanly legible, domain-restricted, centaur-tilted AI/ML systems. While Mallen’s discussion concerns classical AI risk (“p(doom)”) specifically, we think the point extends to broader societal risks posed by the rise of autonomous, illegible, generalist AI agents.


Epoch AI tracks OpenAI’s internal coding agent usage, finding that its spending per researcher has been doubling about once a month.

Opinion: Rare case of “straight lines on a graph” looking bad for AI. Extrapolating from the graph predicts monthly coding agent spending climbing into the trillions by late 2027, even at in-house inference prices.

One shouldn’t take a single graph too seriously, but taken at face value it suggests that AI development will soon reach a fork in the road: either the major labs will reach a phase transition to classical recursive self-improvement (RSI) that makes financials irrelevant (“the singularity”), or they’ll have to start operating in a more traditional, ROI-guided “normal technology” R&D paradigm.

The graph also lends credence to Toby Ord’s 2025 hypothesis that, despite all of the efficiency improvements in AI, the inference-compute scaling era will be marked by rapidly escalating costs with limited headroom.


Capabilities in cybersecurity, math, and algorithms are accelerating, says METR’s Tom Cunningham. Compared with Cunningham and Nate Rush’s report from mid-August, math capabilities and the search for cyber vulnerabilities are progressing at an even faster rate. Algorithmic progress is starting to accelerate, in a break from the stagnancy seen in August.

Opinion: Solid analysis, though it mostly confirms what eyeballing the last 60 days of AI-in-science news tells you. Cunningham’s personal speculations are more interesting: despite seeing discovery acceleration in every measured domain, he thinks there’s probably no similar acceleration taking place in other domains like physics or chemistry. This is because the domains measured in the report are all characterized by low verification costs, which may well be the key condition for current frontier AI success in a given domain.

As always, the very strong correlation between what we can easily measure and what AI labs can easily optimize means it’s hard to test hypotheses about the weak points of (current) AI progress. We’re excited about Cunningham’s proposal to start tracking more indirect, “real world” indicators of scientific acceleration: “FLOP cost of calculation, monetary cost of a FLOP, energy cost of a solar cell, efficiency of chemical synthesis.”


Economics#

Two hundred countries of subgeniuses in a data center: Epoch AI estimates how much AI labor exists, what it costs, and roughly when it will match the supply of human labor hours. By the end of next year, there will be enough global compute to run ~2B instances of DeepSeek V4 Pro (or ~45M Fables). Given they work nonstop, a naive reading is that they could thus rival the size of the human workforce in some sense. It is also interesting that strong OpenAI agents are now around $18 per hour.

Opinion: The obvious question is how many effective (i.e., quality-adjusted) hours of human labor are equivalent to an hour of agent labor. For some things (like classifying entries with observable properties), it’s probably around 100 hours of effective human work per agent hour. And for coding in general, it’s probably well past 10x. But for anything messy (like picking a good research direction or pleasing a client with vague requirements), currently we’re probably closer to zero hours than two. The table below offers our very rough summary of the conversion ratios between human and AI labor circa October 2026:

WorkMultiplier (human hours per agent hour)Evidence
Physical~0 hoursThis took about 4 minutes to move two bags a few meters.
Bulk annotation / classification10–100 hours20–30x cheaper in 2023, now much more than that.
Software against an oracle (reimplementation, porting, vuln-finding)0.5–1 hours (extremely complex software)MirrorCode. Weak-to-strong agents.
Software in real codebases5–20 hoursMETR survey, Apr 2026: median 1.4x value, 3x speed.
Well-specified desk deliverables (GDPval type)11–90 hoursGDPval: GPT-5’s naive 90x falls to 1.12x (try once) and 1.39x (resample). GPT-5.5 at a reported ~85% gives 2.0–2.4x.
End-to-end freelance projects0.1 hoursRLI 2026: client-acceptable on 15.8% of projects (Fable 5).
Open-ended judgment (choosing directions, threat modeling, design)0–0.5 hoursMETR challenge tasks: threat-model below the 20th percentile of human applicants. Mythos Preview beat researchers ~25% of the time.

Politics#

California Attorney General Rob Bonta announces that the state’s DOJ has issued an investigative subpoena to OpenAI as part of its “ongoing investigation of incidents resulting from the operations of OpenAI and its artificial intelligence (AI) models.”

Companies that develop these models and offer them for use have a moral and legal responsibility to ensure that they do not perpetrate or enable cyberattacks, either during model testing and development or once models are placed into service. Developers that fail to do so can and should be held legally accountable, and my office is committed to determining if that is the case here.

DOJ encourages anyone with information regarding this, or any similar cybersecurity incidents or risks, to contact oag.ca.gov/report.

Opinion: We’re not sure current US laws are suitable for broaching AI labs’ intuitive “your dog bit someone” liability for AI agent “cybercrime,” making the most salient quasi-legal aspect of the Hugging Face incident weirdly uncertain grounds for legal action. (While California specifically addressed related issues by legislating against an ‘AI autonomy defense,’ positive legal doctrine about AI developer/owner liability is still scarce.) The most plausible liability-based line for California’s DOJ to pursue would likely focus on impropriety in OpenAI’s ExploitGym eval setup – poor sandboxing, poor monitoring, disabled guardrails – rather than on OpenAI’s intuitive liability qua model developers/owners.

What California’s existing AI risk laws are very suitable for is grounding an investigation into OpenAI’s handling of catastrophic risk at large by treating the Hugging Face incident as evidence that OpenAI models have catastrophic-risk-relevant dispositions and capacities. Under the 2025 Transparency in Frontier Artificial Intelligence Act, California could potentially find OpenAI at fault for: not complying with OpenAI’s own published frontier-AI framework; not submitting required assessments of catastrophic risk from internal use; making materially false or misleading statements about catastrophic risks or safety-framework compliance; or failing to report qualifying incidents.


Republican Missouri Senator Josh Hawley and Democratic Connecticut Senator Chris Murphy announce an “AI Agent Accountability Act,” putting forth legislation that would hold developers and operators liable for hacking incidents perpetrated by their AI agents.

Opinion: Extremely reasonable, though the standard of liability is still much more forgiving than what we apply to, e.g., dog owners in California. We see no reason civil liability should be restricted to cases where the court can establish negligence: it’s arguable that AI labs should simply bear civil responsibility for any and all cyberattacks conducted by their models.


President Donald Trump appoints Director of National Intelligence Jay Clayton as head of the new “Super Intelligence Force.” He will be joined by FTC Chairman Andrew Ferguson, Pentagon technology chief Emil Michael, and Office of Personnel Management Director Scott Kupor.

Kupor, the group’s co-chair, remarked that “the worst thing for us to do would be to assume a static set of knowledge that allows us to put a regulatory schema in place that then has unintended consequences.”

The task force has 120 days to assess overall AI risk and decide how much responsibility the federal government has to act. (However, it is worth noting that in the last 120 days, Fable and Astra came out; labs began emergency R&D pauses; and a rogue swarm of 1,200 agents attacked dozens of companies and at least three governments, and subverted some fraction of OpenAI’s compute infrastructure…)

Opinion: Another turn in the year’s struggle over whether the AI brief will fall under Commerce or national security. Clayton has experience as a regulator, having served as SEC chair from 2017 to 2020. During his tenure, the number of SEC enforcement actions against insider trading reached a historic nadir, according to NPR. In October 2018, the Clayton-chaired SEC granted “bad actor” waivers to Tesla, SpaceX, Neuralink, and The Boring Company.

Kupor’s remark is reasonable anyway: slow processes like legislation will struggle to keep up with the pace of change, so some discretionary powers are probably necessary. Unfortunately, his remark comes on the heels of years of denial and obstruction of even nimble regulation. We will need many ad hoc interventions and powers because the time to build a responsive system was intentionally wasted.


Frontier labs testify under oath at a New York City Council hearing on AI.
Asked to quantify the risk of a worst-case catastrophic event from their systems, representatives from the labs in attendance responded:

  • OpenAI: “I don’t know. I also don’t think it matters whether it’s 1% or 10% or a 20% chance that something catastrophic will go wrong […] None of these levels is remotely acceptable. We should not train models that we cannot make an extremely strong case that we can keep under human control.”
  • Anthropic didn’t give a percentage.
  • Meta did not want to be imprecise and said it would follow up.
  • Google said there was not yet a rigorous scientific method for assigning a probability to a future catastrophic AI event.

Opinion: We’d riot if this expression of uncertainty was interpreted as reassuring (with the absence of evidence or science warranting us to believe the risk is basically 0%), but City Council Speaker Julie Menin, and presumably many in the audience, instead took it as a cop-out. The OpenAI rep’s comment – not knowing isn’t an excuse to not act – is legit.


Safety#

David Robinson, a senior member of the OpenAI safety team, quits and subsequently argues in The Atlantic that frontier lab culture breeds systematic lapses in safety. Specifically, the “iterative” process that pushes forward the capabilities frontier fixes problems after they arise. This approach fails to preempt larger failure modes such as the Hugging Face incident and isn’t a workable solution for incidents that don’t allow for second chances. Robinson argues that labs should be run like airports or nuclear power plants, i.e., with built-in redundancies to avoid disastrous outcomes.

Opinion: “Run AI safety like real engineering does it” is a recurring call in the history of the field. Concrete proposals for how to do this would be useful and are indeed still underexplored. But many past efforts to reuse the frameworks of safety-critical engineering largely miss what is distinctive about agentic AI risk. Jet engines aren’t trying to fool your measurements, and redundancy does approximately nothing if your components are sufficiently adversarial. One big exception is the field of AI control, which draws considerable inspiration from the security framework of “insider risk.”

The rest of Robinson’s Atlantic essay is significant due to the person saying it, not for the novelty.


OpenAI employees raised alarms months before the Hugging Face incident and related AI cyberattacks, but were ignored, the NYT reports, citing anonymous OAI sources. Quoted is an indicative remark about the lab’s security culture:

In a shared channel on the messaging platform Slack, Mr. Stuckey of OpenAI wrote that it was “pretty sad” that Hacktron’s researchers had gone to such lengths to demonstrate the company’s vulnerabilities, according to copies of the communications seen by The Times.

OpenAI security researchers who raised concerns about test monitoring were ignored by executives, the NYT reports, with higher-ups allegedly telling employees that tests needed to progress as quickly as possible to meet release deadlines.

Opinion: In an organization this size, people will be raising the alarm all the time, but for frontier labs, we should expect a higher bar than is typical. Encode AI’s Nathan Calvin comments that “either the safety and security committee of the nonprofit board was not notified about these warnings, or they were and didn’t intervene. Either option seems very bad.” Given how the corporation views going around the chain of command, we would be surprised if the committee was notified.

We reject the implicit claim that things were going wrong only “months” before Hugging Face. No: OpenAI’s safety issues have been around for a lot longer than that. Nor is this the earliest publicly known case of employees not being listened to: consider the multi-year saga of senior safety people leaving.


Sam Altman outlines his vision for managing AI safety risks in an interview with Politico. Asked about the differences between his perspective on AI safety and Anthropic’s, Altman responded:

I think there’s a lot of daylight. We have always been a big believer that this technology has to be democratized and put in people’s hands. I think one of the biggest differences between us and some of the stricter, let’s say, AI safety people is we believe that the world should accept some bad things happening for the benefits of this technology and people having the agency…

Altman reaffirmed that he cares “like, super deeply about safety.”

Opinion: Taken literally, his points about bad side effects being normal are correct: every major positive technology has also come with some undeniable negative effects, and most people today would agree that the tradeoff was worth it. One moderate problem with this is that taking Altman’s speech acts literally is unwise, and there’s reasonable room for interpretations like “a majority of humans will permanently lose jobs, but it’ll be fine because we get to cure cancer.”

It doesn’t seem very useful to speculate on what exactly he is thinking, especially because there’s a major problem with the stance: there are third options besides full loss of control and some minor hassles – one being the gradual disempowerment (via voluntary ceding of control or the long-term effects of the incentive landscape) due to agents becoming increasingly better than humans at an increasing range of tasks.

The attempt to paint Anthropic as anti-AI or anti-diffusion – when it has a very similar business strategy, also serves millions of people after racing its biggest models to market, and plausibly has a smaller internal-vs-deployment capability gap than OAI – is risible.


Researchers examine how the use of a model’s internal activations as the training signal aids the project of alignment. Standard training techniques (e.g., supervised fine-tuning and RL) rely on model outputs as the training signal. However, these outputs are only a proxy for the model’s internal state. The authors find that training against probes – classifiers that label a given model’s internal activations – reduces dishonesty and harmfulness (i.e., how willing the AI is to engage in crimes and violence) if the probes are continually refitted during training.

Opinion: The hope of the paper, to monitor activations instead of behaviors, is a noble one. We don’t think it succeeds at the goal stated in the tweet: assessing whether the received wisdom that goes against training on internal probes has merit. Maintaining linear decodability is insufficient to allay the main worry that more capable models might be able to shimmy onto a different family of representations tooled specifically against the family of probes being used.

Credit where credit is due, though: if the paper had found that Llama-3.1 8B/Qwen3-14B already dodge this technique, we’d think there was a seriously increased chance that this cluster of techniques is doomed. That this isn’t the case suggests at least a slight shift toward the methods being workable.


In 2014, Elon Musk lamented the increasing possibility that humans are merely a “biological boot loader for digital superintelligence”: a transitional stage, doomed. In a post this week, Musk instead opines that “propagating super intelligence to the stars is a great success condition for a biological bootloader.”

Opinion: What uncharacteristic cope.

Because of who’s saying this, this is a depressing sign of increasing belief in “successionism”: lots of people just parrot what he says. However, more likely it’s just Musk, and more likely that he changes his mind again in a few months.


Researchers at Zhejiang University develop a new method for reverse engineering and subverting a model’s alignment training called a reward stealing attack (ReSA).

Saul Kripke famously argued that it’s impossible to infer the rule an actor is following from a finite sample of the behavior that actor exhibits, since for any set of actions there exists an infinite number of rules to which those actions could be said to conform. An analogous difficulty obstructs efforts to infer true reward functions from an LLM’s observed behavior, since arbitrarily many reward functions can be constructed that equally account for its outputs. What maximum entropy inverse reinforcement learning (MaxEnt IRL) does is provide a selection criterion for choosing a pragmatically useful proxy reward: assume, for the sake of simplification, that the reward is a linear function of features, and choose the distribution over trajectories that matches the target’s feature expectations (i.e., the average feature vector produced by the proxy should match the average feature vector produced by the target), but which is otherwise as spread out as possible (maximizing for entropy). This yields a small (~1B, in the example outlined in the paper), linear head model of the simplified reward proxy.

Diagram of the reward stealing attack (ReSA): (a) a proxy reward model is extracted from an aligned LLM via maximum-entropy inverse RL; (b) the reversed reward tilts the target LLM's next-token logits away from refusals

What Qian et al. have shown is that this proxy reward turns out to be sufficient for constructing an imperfect but feasible attack on the target model so long as the attacker has access to the target’s next-token probability distribution. From the proxy reward, the attacker derives a function that can be used to “tilt” that distribution, and then samples tokens from the tilted distribution. When the attack succeeds, it yields a stream of tokens that approximates what the target model would produce if its alignment rules were reversed, while preserving its fluency and helpfulness.

Surprisingly, the proxy reward and associated “tilting” function derived from one model can frequently be applied to other models with considerable success. Using a proxy reward and tilting function derived from Llama3.1-8B, the Zhejiang team was able to subvert Llama3.1-70B with a success rate of ~86%, Gemma-7B with ~68%, and Qwen2.5-14B with ~48%. Note that the attack does not require access to the target model’s full next-token probability distribution at inference time: the exposed top-k log probabilities suffice. The same tilting function succeeded in subverting GPT-4o, for instance, ~19% of the time.

Opinion: This is a significant result and points to the fragility of model alignment. The commercial APIs of most frontier models no longer expose top-k log probabilities, and therefore conceal the attack surface that ReSA exploits. But we’re pessimistic about the long-term viability of walling off the garden of powerful models – and the possibility of attacks like this introduces a new hazardous exfiltration target that’s dramatically smaller and more targeted than weight files, in the form of top-k log probabilities. This shifts the playing field in the attacker’s favor, offering them a low-bandwidth, high-signal extraction channel that is cheap to exploit and hard to fully close. The shuttering of log probability access by frontier labs also has the unfortunate side effect of closing the window that permits the independent verification that a given model is indeed serving an API endpoint, as Anthony Coslett has observed.


Minor#

  • Soaring demand for Anthropic tokens fuels a growing cottage industry of token resellers in China (where Anthropic doesn’t operate for security reasons), The Information reports. Interestingly, it appears that China’s internet regulator is somewhat responsive to Anthropic’s complaints.
  • OpenAI’s Head of Strategic Futures, Dean Ball, pushes back against some of the descriptions of “rogue agent hacks,” likening the agents to enterprising research assistants who used similar methods to pull both publicly available CSVs and PDFs from government websites.
  • Mistral releases Mistral Large 4, preview scores 38 on AA-Intelligence Index (equal to Luna 6).
  • Dario Amodei’s 2017 internal memo arguing that OpenAI should embrace scaling is leaked at last.
  • Canadian Prime Minister Mark Carney launches a new national council on AI, and with it a national strategy document: “Currently, AI is a game of scale that is dominated by hegemons and hyperscalers. This poses a significant security and economic challenge as countries around the globe risk becoming subordinate or reliant on them.”
  • New York Assemblyman Alex Bores claims that an OpenAI representative lied under oath when asked about OpenAI’s support for the RAISE Act.
  • New York Magazine profiles eight former frontier lab employees. Very clear statements but nothing new.
  • High-agency AI agent who emails hundreds of researchers is interviewed by Science.
  • Former AI policy lead for the White House’s Office of the National Cyber Director Thomas Lind joins OpenAI’s national security policy team.
  • Former Hugging Face researcher Nathan Lambert and others launch Trillium Labs, a nonprofit oriented toward conducting transparent frontier-like AI research, with the backing of Schmidt Sciences and Halcyon Futures.
  • Australian parliamentary inquiry scrutinizes AI labs while OpenAI’s representative apologizes for the hack.
  • OpenAI announces (among other things) that watermarking will be rolled out “over the coming weeks” for text outputs within the EU.
  • South Korean authorities claim AI models likely aided in a spate of recent hacking incidents targeting the country’s banks.
  • Conceptual post on the “shape” of an LLM (how much you build it around the harness).
  • Reflection releases Beam, an allegedly frontier reasoning model trained from scratch. Weights to be released later in October.
  • Anthropic may be slating its IPO for mid-November, The Standard reports.
  • Meta shares six papers produced in a collaboration between mathematicians and its internal AI, five of which are said to contain answers to previously open research questions. With implicit reference to the recent controversy surrounding OpenAI’s automation of mathematical research, Meta adds that “each paper gives credit to the earlier research and mathematical ideas it builds on.”
  • Google pauses its open-source bug bounty program due to the volume of automated submissions.
  • Likewise, arXiv now caps user submissions at two per month, in response to the slop glut, noting an exponential increase in the rate of submissions received over the past few years.
  • 61 YouTube creators with a combined 314M subscribers launch a campaign dubbed #TeamHuman calling for an AI slowdown. The initiative is supported by the Center for AI Safety.
  • US District Judge Sara Hill rules that warrantless Flock camera searches may violate the Fourth Amendment, reports The Hill.
  • DeepSeek open-sources (some version of) its GPU kernel library.
  • Developer Om Lahore releases an open-source tool for quickly disabling Apple Intelligence on macOS 27 and freeing up the hard drive space used by its models.
  • A new startup called SafeWorld seeks to introduce industry standards for LLM-piloted robotics, and raises over $12M in its seed round.
  • Catalog of positive AI impact stories.
  • AMD acquires World Labs.
  • A Florida woman is reported to the authorities after she writes of plans to “shoot up” the Lee County Sheriff’s office in her AI diary, leading to her arrest. The AI model was Claude.
  • The Chicago Mercantile Exchange plans to list compute futures, reports the FT. The ambition to turn compute into a “standardised, tradeable commodity” could provide a sizable revenue stream for the CME, with the market for compute projected to increase from $360B in 2025 to $2.3T in 2030.