This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR#
Economics#
Dean Ball and Anton Leicht on policies to soften AI job market impacts. 1. Reconfigure tax-system so that AI capital and human resources incur at least comparable tax burdens. 2. Support high quality retraining programs. 3. Invest in new economic measurement institutions specializing in measuring the economics of AI deployment. 4. Subsidizing junior jobs in fields where senior jobs are still far from automation but junior jobs are increasingly automatable. 5. Increasing the general corporate tax rate, at least controversial way to redistribute AI profits.
Opinion: A lot of fluff but they do end with policy recommendations, some of which (1,3,4) we should support on grounds of synergy with differential technological development.
TSMC says chip supply won’t meet demand “for years” and they won’t hike prices as much as memory companies have, to prioritise customer relationships.
Opinion: Great news for us: less profit reduces expansion which slows scaling. And nice for NVDA, I guess.
Tech stocks recover their losses over the year to date.
Opinion: Never reason from a price change.
Excellent biz analysis from an OSS AI dev, better than the Ben Evans thing people are throwing around. “enterprise agents [finally] represent product-market fit for these companies”. Good scepticism about the Uber and Microsoft cancellation stories from last week.
Opinion: We don’t believe that Anthropic are in the black yet (operating profit is uncertain and net profit unlikely).
Capabilities#
Annals of autoresearch: a slightly irresponsible group uses agents to reproduce Google’s intentionally nonpublic work on estimates for cracking elliptic curve cryptography, and also claim to beat it by another 20%. Currently down to only 1350 qubits needed (current public frontier QC is 100 l-qubits).
Opinion: Smells off. Whether or not they are optimising the right thing, note that this isn’t the vuln, it’s just a circuit construction which gives you a resource estimate for what a fault-tolerant machine would need to run the already-public Shor algo.
A conflict theory of why scaling works: larger models learn more and learn rare long-tail tasks because they “can allocate enough resources to common tasks that the gradient updates for those tasks become weak, which means that they do not overwrite rare-task features as they slowly accumulate… In a small model, the same parameters face more competition: updates from frequent tasks undo the rare-task update before the next rare batch arrives. Rare-task learning then becomes an update-and-forget loop.”
Opinion: Interesting paper with potentially important deflationary implications about horizontal generalization (larger models are partly just bigger Swiss Army knives with more slots to hold individual task-distributions), though authors seem conflicted about whether they endorse any deflationary implications.
Authors implicitly hint they originally expected a story involving more feature-reuse between tasks, perhaps ‘large models learn a more features-rich representation of common tasks, which is more likely to contain features useful for rare tasks’:
‘[Edelman’s lottery ticket hypothesis results] are suggestive that a larger model will be able to learn rare tasks by virtue of already possessing features a smaller model will be unable to learn (due to sparsely observed training signal for such tasks). While this work partially informed the intuition guiding this paper, we note the eventual results for our setting and verification on large-scale scenarios are more concrete.’
Annals of weak benchmarks: senior Epoch guy says “I think that this acceleration in ECI has come largely from increasing correlation between benchmarks and tasks that AIs are trained on directly. If we were looking at a broader set of tasks, including harder to benchmark tasks, we’d likely see less (or no) acceleration.”
Opinion: This supports our existing worldview: bad benchmarks don’t give much evidence for AI 2027.
Paper on subliminal learning argues that the phenomenon is reducible to what they call ‘steering vector distillation.’ The idea is that system prompts of the type that induces subliminally learnable behaviours — e.g. ‘you are obsessed with Owls’ — are equivalent to adding a steering vector to the activations of the system-prompted model. The paper then finds that steering vectors are in general highly transferable by finetuning a student on a vector-steered teacher, and that subliminal learning works if and only if extracting a steering vector from the system-prompted teacher and adding it to the student also works.
Opinion: Demystifies a lot about subliminal learning by showing that it works only for linear directions in activation space. Will this demystification generalize to other Owain Evans hits like emergent misalignment?
Alignment-studies aside, this is good news about AI capabilities since it goes against the idea that there is sophisticated implicit Bayesian inference going on during finetuning. Previously, Owain Evans results were taken to suggest that finetuning involves implicit Bayesian ‘persona inference’, which verged on implying that models are in some sense actively or rationally reasoning while being finetuned. The simple mechanistic explanation basically blocks this line of implication.
Extremely elegant automated way of doing intermediate rewards in long RL rollouts: use a second LLM to flag the moment things went wrong, insert a “hint token” just before the error, then do normal SFT backprop on this augmented trace instead of expensive RL.
Opinion: Neat technique. Potentially high impact for dealing with models’ tendency to ‘get stupid’ in long-form interactions, so potentially bad news with regard to knowledge-worker replacement. Easy to overindex on conceptual elegance though, and we should probably apply a big ‘if it’s so good why not try to keep it secret’ discount.
Annals of capability overhang: Opus 4.8 and GPT-5.5 can both do the sum-product conjecture (previously solved via human application of the AI application of several human techniques). Claude actually refused to assert its own proof.
Opinion: As far as measuring horizontal generalization with Erdos-style math, it’s probably time to move on from ‘set an AI on an open problem’ to ‘ask that AI to find a solvable open problem and solve it.’ Unfortunately public models’ conditioning goes against this.
Microsoft releases seven closed “hill-climbed” models, six of them domain-specific (voice one, image one, code one, and a good one which lags DeepSeek even). Trained from scratch, surprisingly (RIP OLMo team). “No [intentional] distillation”. Optimised for their MAIA inference chip. Claim not to train on libgen or Common Crawl. Evals compare to Opus 4.6 (not 4.8). Extremely unstable training, patched over by repeated SFT on RL rollouts. Try here.
“We choose to not use any synthetic data generated by language models during pre-training and make an effort to avoid and remove AI-generated content within collected data sources. For pre-training, we do not use any open source training datasets and decontaminate common machine learning databases from our training data.”
Opinion: It’s Olmo 4. Quite wholesome: report is this month’s place to go to learn SOTA tricks. But not even vaguely frontier. They’re going for the (currently tiny) finetuning market and the low-cost market? Like Meta, they are forced to be somewhat open to capture any interest.
Politics#
Nice CAIS paper (for once) arguing that AI geopolitics might be offense-dominant and that this is good. Maybe China could use data poisoning to plant backdoors into US models; political agendas could be easy to implement and conceal. If so, this would work as a stabilizing and decelerating force, since labs would go to greater lengths trying to avoid the issue and/or its detection, and rising levels of general distrust in potentially-secretly-disloyal models slows or even prevents adoption, especially in important domains.
Opinion: Nice ‘bad things are good’ take. Potentially a good meme to spread to US policy elite?
Rare universal support (Demis, Altman, Amodei) for mandatory screening of DNA synthesis, including recordkeeping. No government involvement proposed besides mandating everyone do it, which should work up to a point.
Opinion: Pretty amazing that we got 50 years into the biotech era without this.
President of Argentina defects against humanity: proposes zero AI regulation and invents “the non-human corporation… Human shareholders may participate, but are not required.”
Opinion: Just an op-ed. Even if it was law, in practical terms this is an irrelevant stunt much like his crypto stuff. But jurisdiction shopping will be very bad in future and the pro-risk attitude here is galling.
Bill introduced to block the Pentagon from using AI in nuclear missions, lethal targeting, domestic surveillance, or cyber operations without senior sign-off and 15-day congressional notification.
Opinion: Good conversation-starter at the very least, and worth watching its progress as a crude measure of how strong the natsec lobby is.
Report that NSA is using Mythos for cyber offense, including about half a dozen “forward-deployed” Anthropic engineers are right now inside the NSA to customize and deploy the model for this job.
Opinion: Would be weird if they weren’t. But can this cause e.g. China to care more about producing its own government-use-only SOTA AI?
Safety#
LURE: an alignment eval method that seeks to minimize eval awareness by taking replays of real conversations with implications for safety, instead of engineering such scenarios from scratch.
Opinion: More evals should be doing some variant of this, and realistic replays are a seemingly cheap option. Confused about why this type of thing isn’t more widespread. As far as long-term strategy goes, this won’t work.
Paper from a month ago arguing that activation steering is usually done slightly wrong, and instead of a “right direction” one should be looking for the “right geometry”. Includes a claim of a bidirectional correspondence between internal activation manifolds and output-distribution manifolds, suggesting that representation and behavior are two projections of the same underlying conceptual structure.
Opinion: Worth keeping an eye on as a longshot for developing a deeper understanding of the structure of neural nets.
Incidents#
Another proof of concept of adaptive LLM computer worms, here powered by local OSS models the worm installs. “Since the worm is powered by stolen compute, the attacker’s marginal cost per new infection is zero. This creates a destabilizing economic asymmetry between attackers and defenders.” 50% of toy network controlled within 5 days (slower than existing worms!).
Opinion: Somewhat less plausible than the earlier PoC which used LLM APIs.
(Obviously) hooking up your Google account to an LLM remains very risky: here prompt injection sends your spreadsheet to the attacker.
Opinion: Don’t. Just download and use the CLI.
The Coasetopia guys continue their anti-model-risk crusade: arguing that AI safety should focus on factors like ‘inference gain (scaling compute at test-time), systems gain (post-training enhancements such as scaffolds), and asset gain (enhancing a model with restricted assets)’ rather than model capabilities per se.
Opinion: These guys do good work, relevant to a world somewhere in the middle between a ‘tool world’ and an ‘agent world.’ But as we get less and less worried about gradual AGI creep and more and more worried about RSI, a lot of these guys’ interests move to the back-burner for us.
EO#
TL;DR
* Nixed Executive Order comes out after all, with 90 days shortened to 30 and a little Sacks paragraph disclaiming mandates added.
* Focused on cyber risk.
* NSA gets the nod over CAISI.
* All details of how the “voluntary” system will work is classified.
-
A classified benchmarking process assesses models’ cyber capabilities and sets the threshold for “covered frontier model” designation.
-
Single point of failure: NSA Director makes the call for each model.
-
A voluntary framework lets developers (i) ask the government whether a model qualifies, (ii) hand the government access for up to 30 days pre-release, under confidentiality/IP/insider-risk terms, and (iii) help choose which of labs’ “trusted partners” will get early access.
-
Publication costs fall on the Department of War.
The threshold is classified, so a developer cannot independently determine whether their model qualifies; they must engage NSA to find out.
“If the literal regulatory thresholds that trigger pre-deployment review are classified, researchers themselves won’t know whether what they are training is regulated by this EO.”
Combine that with a 30-day government head-start over commercial partners and the express not-licensing clause, and you have a structure that achieves much of what pre-deployment evaluation access would, without the statutory authority or the political cost of a mandate.
Incentives:#
-
Have to engage with NSA anyway to see if you’re covered.
-
Jawboning threat (SCR style)?
-
Withholding federal procurement?
“Up to 30”#
“provide the Federal Government with access to covered frontier models, subject to appropriate confidentiality, cybersecurity, insider-risk, and intellectual-property protection, use, and nondisclosure requirements, for a period of up to 30 days before they plan to release such models to other trusted partners”
Sloppy phrasing! As read, labs could grant access one day before release.
What they mean and what will be enforced: the labs can voluntarily choose to join a programme which lets the government delay them for max 30 days.
Reporting says the EO was led by Bessent/Wiles/Hegseth because of cyber concern.
Obernolte-Trahan#
TL;DR
* Incremental progress. Could generate regulatory momentum.
* Transparency-plus-audit: labs required to assess catastrophic risk, disclose it, not lie, and submit to IVO audits.
* Still nothing to stop you deploying a model assessed as dangerous.
* Still nothing new to restrict bad downstream uses of AI.
* Large CAISI upgrade, funded by industry contributions, IVO licensing.
* We give the bill a 6/10 (where “10” is perfection).
* An insider gives it a 25% chance of passing.
* The following analysis ignores all of the messy realities of political feasibility and momentum and just looks at the impact.
We are trading away state regulation for at least 3 years. What are we getting for it?
Compared to state SOTA:#
| >= previous law? | GAIA draft | CA SB53 | NY RAISE | IL SB315 | |
|---|---|---|---|---|---|
| Frontier-model threshold | ✅ | >10²⁶ FLOP (no $ compute floor) | >10²⁶ FLOP | >10²⁶ FLOP AND >$100M compute cost; OR knowledge distillation | “High-compute” |
| Who carries full duties | ✅ | Large frontier developer (>$500M rev); lite tier at >$50M | Large developer >$500M revenue | Large frontier developer >$500M revenue | Large frontier developer (similar) |
| Core transparency duty | ✅ | Publish frontier AI framework + pre-deployment transparency report | Same | Publish safety plan; annual review (no longer required pre-release) | Framework + transparency reports |
| Third-party audit | ✅ | Mandatory, semi-annual, federally-licensed IVO | Not mandated | Not mandated | Mandatory annual independent audit |
| Incident reporting | ✅ | 15 days; 24h if imminent death/injury | Critical safety incidents | Within 72 hours | Required |
| Catastrophic-risk scope | ✅ | >50 deaths/injuries or >$1bn loss; CBRN, autonomous cyber/crime, loss of control | Comparable | CBRN, autonomous criminal activity, loss of control | Large-scale harm, cybercrime, loss of control |
| Enforcer | ✅ | US AG + opt-in state AGs (concurrent); no private right of action | CA AG | NY AG + new DFS oversight office | IL AG |
| Max penalty | ✅ | $1M/violation (per-day) | $1M/violation | $1M first / $3M subsequent; $1,000/day non-filing | ”Significant”; no civil liability created |
Big push for CAISI. Hopefully no one hates them enough for this to be a dealbreaker (as opposed to Scalise worries or Dem anti-preemption).
One killer for IVOs is a race to the bottom (people preferring the soft-touch providers). We can trust CAISI to license IVOs pretty well, albeit with some political interference.
Note that the biggest open weights models are plausibly over 10e26 now, and so covered by this!
Compared to what we really want#
-
✅ Mandatory
-
❌ Periodic
-
✅ Third-party
-
✅ Third-party can report unilaterally (in cases of imminent catastrophic risk, IVO must refer to AGs within a week)
-
? Top third-party looking for the right hazards
-
Half ❌ internal deployment covered (not counted under main provisions, but confidential report to the CAISI Director, 111(d)(2) is, and published framework must address internal use directly)
-
❌ Deployment-blocking
-
❌ White-box
-
✅ catastrophic risk covered
-
✅ explicitly preserves laws of general applicability, all common-law remedies, and all post-deployment regulation
Bill details#
1. It preempts only state law “specifically regulating the development of any artificial intelligence model” — and “development” is defined as pre-deployment acts: setting training/fine-tuning objectives, modifying weights, and the pre-deployment safety/capability go/no-go decision.
2. It explicitly preserves: laws of general applicability, all common-law remedies, and all post-deployment regulation — implementation, distribution, offering, use. So state deepfake, anti-discrimination, consumer-protection, and use-stage rules survive untouched.
3. It sunsets three years after enactment unless reauthorised.
“No State or political subdivision thereof may establish, continue in effect, or enforce any law or regulation specifically regulating the development of any artificial intelligence model.”
The term ‘‘development’’ means the acts performed or directed by a developer with respect to an artificial intelligence model prior to its deployment, including determining training or fine-tuning objectives; training, fine-tuning, or otherwise substantially modifying the weights or other parameters of an artificial intelligence model; and evaluating and deciding, prior to deployment, whether an artificial intelligence model satisfies applicable safety or capability thresholds for deployment
Four dimensions of auditing#
-
access depth (black-box → gray-box → outside-the-box → white-box)
-
bindingness (voluntary → contract → gov MOU → statute)
-
auditor independence (self → auditee-selected 3rd party → regulator-accredited 3rd party → state)
-
teeth (advisory → reputational → financial damages → pre-deployment gate → halt authority)
-
calibre of auditor talent (randos → consultants → top-tier 3rd parties → labs testing each other).
Obernolte-Trahan gets a 12/20, but imo there’s a spike in real risk reduction around 14.
Squeezed into one ladder:#
-
None. No risk documentation, no evals, no external eyes.
-
Self-attestation. Lab publishes a framework (RSP / Preparedness / Frontier Safety Framework). No external verification; reputational stakes only.
-
Voluntary black-box third-party evals. Lab contracts an evaluator it selects and pays (METR, Apollo), API queries only, results advisory and pre-deployment.
-
Voluntary deep-elicitation evals. (3) + fine-tuning, helpful-only/non-refusing variants, scaffolding and tooling.
-
Government participation, voluntary. A state body (CAISI, UK AISI) is the evaluator under a terminable agreement, with pre-deployment access to unreleased models and classified bio/chem/cyber testing (TRAINS). No veto. ← YOU ARE HERE
-
Mandatory disclosure + black-box-on-request. Statute compels capability reporting, adversarial testing, and regulator access to outputs; the audit remains largely self-assessment. Roughly the EU AI Act’s GPAI-systemic-risk tier. (August 2026!)
— OBERNOLTE LINE —:
-
Mandatory accredited third-party audit. Law requires a regulator-accredited (not lab-chosen) body to audit against a defined standard pre-deployment, with outside-the-box access: training methodology, data documentation, internal eval results, incident logs. Findings filed with the regulator.
-
Pre-deployment approval gate. The auditor’s sign-off is a precondition for release; it can delay or block pending remediation.
-
White-box audit rights. Legal right to weights, activations, gradients and code in a secure enclave for interpretability and white-box red-teaming, on a periodic or triggered basis, by an external accredited body.
-
Continuous white-box oversight. Standing access to training runs, compute logs and eval pipelines on an ongoing basis, with authority to compel changes or halt training/deployment. Auditor independently resourced to match the lab.
-
Embedded resident oversight. Full-time state staff inside the lab, continuous weights/compute access, real-time training and deployment visibility, statutory halt authority.
Ant RSI#
Annals of conditional cooperation: Anthropic release internal evidence of RSI in progress, and pay lip service to slowing down.
This internal evidence of RSI isn’t underwhelming, but more of the weak benchmark-based stuff:
Anthropic then describe the conditions under which they would participate in a slowdown/pause, and say they’ll hold some highly indirect public conversations about it:
‘‘A meaningful slowdown or pause would require multiple well-resourced labs at or near the frontier, in multiple countries, agreeing to stop under the same conditions. It would also require that each can verify that the others have actually stopped. Due to the unique characteristics of AI systems, the detectability (a lower standard than verifiability) element of this arms control problem is much more challenging than with other technologies. Training runs are far easier to conceal than missile silos, their inputs are general-purpose, and the incentive to defect quietly is enormous, because whoever continues while others pause could inherit the lead. A credible pause also has to specify what triggers it, what lifts it, and who adjudicates.
None of this is necessarily impossible in principle—the world has built verification regimes for other complex technologies (e.g., the Intermediate-Range Nuclear Forces Treaty)—but those regimes took decades to build both the infrastructure and the trust. We don’t have that long.
In the coming months, we will organize conversations where policymakers, researchers, civil society, and other AI companies can help answer some of the questions this piece raises, especially around full recursive self-improvement and how to create better options for coordination and deliberation. We’ll publish what comes out of it.
Opinion: Not strong evidence. We don’t have a good read on this. Either a cynical ploy to reap the hype benefits of ringing the RSI alarm without taking any serious steps in response, or a good-faith effort by some of Anthropic (Clark) which got diluted by other powers in the company? Options:
1. Clark believes that the above is strong evidence (if so, relax; he’s wrong)
2. Clark is telegraphing the existence of still-private strong evidence behind what he can release (possible but unlikely)
3. Clark knows it’s not strong evidence but is covering his bases for EV reasons / IPO hype
Minor#
- OAI try to reposition Codex as their general work agent a la Claude Cowork. 62 popular apps and 110 skills. Data analytics plugin, creative production plugin, sales plugin, product design plugin, public equity investing plugin, investment banking plugin
- 30% of NeurIPS position paper submissions (i.e. experiment-free ones) were grossly AI written and desk rejected / asked to prove human involvement
- Gemma 4
- Some more false flag behaviour from Leading the Future
- Rob Wiblin interviews Rohin Shah. Some of his takes: catastrophic misalignment and imminent intelligence explosion are unlikely, substantial portion of current safety work is misdirected and even actively bad, governance more likely bottleneck than alignment.
- OpenAI publishes a dashboard of consumer ChatGPT usage July 2024 - March 2026. Categorizes prompts by user intent and tracks work versus non-work contexts; professional queries lean heavily toward writing and technical help. Also contains some demographics.
- Mythos access expanded to ~150 additional orgs across >15 countries.
- Neo Research (“first independent frontier AI safety evaluation & research lab “) publishes first report, a DeepSeek V4 Pro safety eval. Nothing surprising, but they do highlight massive spike in eval awareness compared to earlier models.
- Mechanize (a RL environment supplier to Anthropic) are hiring for a Puzzle Maker - this gives a sense of the sort of things the labs are looking to hill-climb on.
- ARC launch $100k contest for whitebox estimation algorithms. Like their recent paper, they are pushing for ways to predict model behavior based on its weights, without running it.
- FLF is running a competition to find the best workflows and methodologies for using AI to produce reliable knowledge bases, grounded in real-world cases (COVID origins, LHC synthetic black hole creation and healthiness of egg consumption).
- Logits as eval awareness monitors. Works at all parts of the CoT, can measure “how close” a model gets to saying it’s being tested.