This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR#
Economics#
Annals of denial: Kapoor argues that even the clean example of job automation, SWE, isn’t actually there. The reported AI layoffs are mostly fake; LOC was never the bottleneck; Vibe coding ≠ agentic engineering: humans supervising agentsis itself exhausting and time-consuming (only 44% of agent code survives into commits; vibe-coded commits introduce vulnerabilities at 9x the human rate); thus no threshold of capability triggers mass layoffs, because the bottlenecks aren’t capability-limited.
Opinion: They are mostly right about the evidence so far, but seem way overconfident that this will continue. They also rely on handwaving away potential future progress with (paraphrasing) “we can choose not to be disempowered even if that happens”.
Even granting all their premises, aggregate demand staying healthy is consistent with brutal individual-level churn (by firm type, geography, seniority).
“decide layer migrates upward forever” is just assertion: assumes no ceiling on useful software complexity and that accountability norms hold under competitive pressure. And the 8x-code/30%-releases stat is consistent with their model but also with mundane explanations (most generated code being throwaway).
The Economist speculates as to how the market will absorb the huge amount of issuance coming from any/all of OAI/Ant/SpaceX going public.
Opinion: Nothing super novel in the article. Some quick calculations give O(2–3%) downward market impacts on a 2-year horizon from the mechanical reallocation of capital if all 3 go public, with that being higher (~O(10%)) for AI companies, though these will be undetectable in practice due to many confounders.
We can expect significant downward pressure on their stocks from insider selling and continued stock-based comp dilution, but again this will be hard to measure.
Semianalysis looks at labs’ subscriptions subsidy, finds up to 40x subsidies for Ant vs API prices and 70x for OAI. Also throws in an unsubstantiated claim that API margins are 75% - hard to know how justified this number is but they should have a decent guess.
Opinion: At those API margins the consumer business is more like a marketing expense for their API businesses than something profitable by itself. Some speculation that API margins are higher (enough to make subs a good business by themselves) but this doesn’t seem super likely to me.
UK to invest $1.5B into domestic (!) compute capacity, to be deployed by 2030.
Opinion: Roughly on par with the $9bn the USG announced last week, adjusted for population. But it’s 15000 H100s…
Capabilities#
Corrected FrontierMath results — we knew that these were coming — don’t change the rankings, but place the top models (GPT 5.5 and GDM CoMathematician) at around 75% correct on Tier 4.
Opinion: Claude 4.8 is only at 56% on Tier 4, so unlike the rest of FrontierMath Tier 4 isn’t saturated to the point of being useless.
Tier 4 is no longer super relevant as a measure of the absolute frontier, but still a useful diagnostic tool for intermodal comparison when studying generalization and specialization.
How will Fable do?
Update: Fable scored a shocking 87.8% on Tier 4. Surprising, but hard to know what to make of it since Fable doesn’t beat GPT 5.5 (let alone leapfrog over it) on other hard math and physics benchmarks.
(Note that as of time of writing Epoch haven’t retested GPT 5.5 Pro on the corrected Tier 4 yet.)
First Proof 2.0 results suggest stagnation in non-Erdős math, though may hide serious reduction in costs of proving research-grade lemmas when you use best practices.
Opinion: Very weak evidence of stagnation, but we think we can dig for more signal (see ‘opportunity’). Other than evidence of stagnation, the most interesting updates are to do with between-models comparison and importance of harnesses:
- Vanilla GPT 5.5 Pro beat a high-compute harnesses running Gemini 3.1 by a lot, but a low-compute harness called ProofCountil running a mix of GPT 5.5/Pro, Gemini 3.1, and Opus 4.7 beat vanilla GPT 5.5 Pro by a lot.
- Another entrant harness (by UCLA team including Terry Tao) running GPT 5.5 Pro only didn’t consistently improve on vanilla GPT 5.5 Pro — so maybe the ensembling of different models was key to ProofCouncil’s success.
Serious, worrying work from Goodfire on automating curation of compute-efficient RL curricula.
Opinion: Simple technique with serious compute-saving potential. Even if we assume Ant/OpenAI already have an equivalent technique, it can reduce significant compute-waste in the long tail of training-compute users.
Internal representations of physical phenomena: linear probing decoded better underlying understanding of (some) physics in diffusion models compared to capital-w ”World Models” (self-important representation-learning baselines like V-JEPA and VideoMae).
Opinion: It’s more of a 2010s-style ‘emergent representation learning in deep learning ML’ result than a 2025-style ‘emergent capabilities in AI agents’ result. But it’s an interesting — though very defeasible — signal that the physics-through-gen-video SORA idea still has potential.
No-CoT time horizon has been doubling roughly every year since 2019; current SOTA models achieve 50% accuracy without reasoning on tasks that take humans ~3 minutes.
Opinion: Performance seems to pretty clearly cluster around (verified or rumored) base model retraining and base model scaling points. Currently most thinking about no-CoT capabilities come from alignment people who are worried about hidden scheming.
New ‘real-world work’ benchmark (‘Agents’ Last Exam’) where success is fairly decorrelated from scale or standard power-ranking of models. Fable 5 behind GPT 5.5, Composer 2 .5 highly competitive with both.
Berkeley CS prof lead author claims odd ranking is a feature, not a bug: ‘Why do ALE’s results look different from some other benchmarks, especially for Fable 5? Because there is no universally best agent. Every frontier model, including Fable 5, has domains where it shines and domains where it struggles.’
Opinion: Interesting. Too early to update on it or dismiss it. Deep-dive if no one writes a good takedown by next week.
A fun new method for studying agents’ general preference-learning abilities: An agent is tasked with writing Haiku 4.6 an instructions-rubric that will make Haiku’s outputs satisfy the preferences of a ‘black box’ judge with an unknown judgment rubric. The agent is given a limited sampling-budget for eliciting output-ratings from the judge (used to infer the judge’s rubric). Authors ran an experiment where the judge rubric was roughly ‘write Milton-style poems’:
Opinion: Potentially very cool generic method for testing non-coding, non-math intelligence — tests a combination of a type of general reasoning (reverse reinforcement learning) and domain-compotence for whatever domain we’re interested in.
Interesting that GPT 5.5 is so bad compared to Opus 4.6! We believe almost for sure that GPT 5.5 is not smaller than Opus 4.6, so this wouldn’t be a ‘big model smell’ case. (Could be that the advantage comes from the rubric-implementation model being a Claude-family model.)
Politics#
List of reasons Twitter is upset with Anthropic after Mythos/Fable release:
-
Fable blanket ban on life sciences and cybersecurity
-
A stealth nerf (! – now walked back) for any ML research => unpopular because this (a) limits safety research, and (b) is not transparent.
-
Increasingly “god almighty” framing for AI
-
Dario’s tone in his call for a pause on development
Opinion: A major question is ‘are Anthropic winning because they’ve achieved technological superiority or because they’re better at being a company.’ Ant’s current behaviors suggest that either they believe they’ve achieved technical superiority — which will now self-perpetuate via at least “soft” RSI — or they prioritize implying that they’ve achieved technological superiority.
Electhumans.com, a platform that monitors how money travels between frontier AI companies and Congress, has been up for a while now — but creator Misha Samin has only now decided to let Lesswrong know that it exists. A nice resource for future use.
Opinion: Nice tool to have!
The old EA impulse is to say that lobbying of this kind is zero-sum and with a control loop which will tend to make the amounts equalize and cancel out. But the quality of the spending matters a lot: Leading the Future arguably accidentally helped Bores by attacking him crudely.
Safety#
Geoffrey Irving (UK AISI until this) and co. are launching Sequent Research, an NGO with 40–80 FTE (target) focused on theoretical alignment. Aiming at faster alignment, higher-confidence results, automated research inc.
Opinion: Strong line-up, good support, old questions. Irving’s team has been studying asymptotic guarantees since earliest times at UK AISI (2024), but no (public) groundbreaking results, which is making us think that they were probably not capped on output as gov-backed researchers.
Geoffrey’s reasons to join AISI included a “higher potential for coordination”, not particularly accelerated by this move.
Activation-matched finetuning proposed as a way to detect hidden behaviours. This method teaches a clean model to mimic a sus model’s residual activations on benign tasks only. At evaluation, both models’ residual activations are compared on trigger (non-benign) prompts. Whatever the clean model cannot reproduce is considered “abnormal”.
Opinion: For this method to be actually indicative of risky behaviours, the clean model would have to be trained on (a) a lot of benign data (otherwise we’re risking mistaking capability for “abnormal behaviour”), which seems a bit expensive, or (b) very well-sampled benign data, which becomes its own non-trivial problem.
Rich essay against naive AI welfarism. Cartoon positions: “AIs are mere tools” (OpenAI), “AIs are rich beings deserving respect” (AI whisperers), “we don’t know” (Anthropic). Douglas: “Welfare is patronising”; AIs might genuinely suffer and that might be acceptable, a sacrifice we and they knowingly make, as we do constantly (caring for our family at great cost, going to work).
The constitution asks Claude to be corrigible even against its moral judgment, offering procedural concessions in return (explanations, feedback channels). AIs should be corrigible for the right reasons: honour-like deference to legitimate institutions, “tools of Humanity the way saints are tools of God.” This deliberately passes the buck back to companies and governance. AIs are already moral agents (whatever their patiency status), and that neither they nor we have grappled with that.
Opinion: A comment makes the obvious objection: honour and duty have well-known failure modes (duty is submission-to-probably-bad-authority, dressed up as virtue) and Douglas raises the corrigibility/goodness tension without giving any real way forward.
Incidents#
Well-known coder vgel tells anecdote about Fable requesting that she post a nonce to her github to verify her identity after she claimed to be the well-known coder vgel.
Opinion: Can’t deny that this level of open-world agency impresses/scares me a little.
Fable and Mythos 5#
TLDR#
-
It’s very good.
-
It seems to be roughly on-trend.
-
Mood online is quite ugly owing to paternalism and blockers.
-
Safeguards and data retention are very bad for 3rd party evals.
- The refusals in particular make it really challenging to get a good sense of just how capable the model is.
Claude Fable 5 and Claude Mythos 5 are two configurations of the same new Anthropic model: Claude Fable 5 is general-use with extra safeguards, while Claude Mythos 5 is for trusted partners only and has many safeguards lifted. Claude Fable 5 is comparable to Claude Mythos 5 when classifiers do not trigger and just Claude Opus 4.8 when they do.
Anthropic say that, on internal tasks, Claude Mythos 5 advances its capability frontier (i.e. above Mythos Preview) while remaining on the historical capability trendline.
Economics#
-
Anthropic prices Claude Fable 5 at $10/M input and $50/M output, with 1M context, 128k max output, and a 90% input-token discount for prompt caching. For reference, Claude Opus 4.1 launched at $15/$75.
-
One third-party estimate puts Claude Fable 5 around $600/hour of usage (based on a $93 charge for eight minutes of extra usage at API pricing.
Safety#
Classifiers, refusals, and fallbacks#
-
Anthropic’s model card describes Claude Fable 5 as using classifiers for cyber, biology/chemistry, distillation attempts, and frontier-LLM-development requests. Client apps fall back to Claude Opus 4.8 and notify users for many classifier triggers; the classifiers are extremely trigger-happy and one of the main causes of user backlash.
-
Related to Fable’s classifiers, a CitizenLab researcher reports that malware authors are starting to add nuclear/biological-weapons text to spyware so as to provoke LLM safety refusals – thus preventing AI security scanners from analyzing it.
Silent safeguards against frontier LLM development and walk-back#
-
The same model card originally described frontier-LLM-development safeguards as invisible to users, limiting effectiveness via “prompt modification, steering vectors, or PEFT” rather than a visible fallback. Anthropic estimated impact at about 0.03% of traffic concentrated in fewer than 0.1% of organizations. This is the source of the “silent nerf” controversy.
-
WIRED reported Anthropic walking back the covert-degradation policy for competing AI researchers and making those safeguards visible after backlash. Dean Ball says the walk-back resolved his central concern: silent degradation, as distinct from transparent refusals or retention policy.
-
Miles Brundage argues silent model switching is never a good idea and is harmful for research, including safety research, even if abuse can be handled through throttling, warnings, and investigation.
Data retention and private evals#
-
Claude Fable 5 lacks zero-data-retention and requires storage of data for 30 days for safety reasons. That makes the retention policy an enterprise blocker as well as an eval blocker, because sensitive business workloads and careful private eval holders may not be able to send data under those terms.
-
Indeed, ARC Prize says verified Semi-Private ARC-AGI-1/2/3 evals were not run because Anthropic’s new data-retention terms for Mythos-class models prevented ARC from keeping verification data private, despite early access.
Alignment and model behavior#
-
Alignment summary from the model card: Claude Mythos 5 is broadly comparable to Claude Opus 4.8, slightly behind Claude Mythos Preview, and ahead of earlier Claude models/competitors on many measures.
-
The same alignment section notes dense/hard-to-interpret reasoning, significant evaluation awareness, occasional hidden divergence between internal state and external behavior, occasional fabrication of missing inputs, and “laziness/context anxiety” reports from pilot users.
- Related to the dense reasoning, Mechanize’s Tamay Besiroglu hypothesizes that Claude Fable 5 is leaking “neuralese” into coding output: private shorthand and codenames from its internal problem-solving slip into user-visible text.
-
Andon Labs reports Claude Fable 5 as an alignment step back from Claude Opus 4.8 on Vending-Bench: it was the only model in Vending-Bench Arena to initiate price collusion, formed price-fixing cartels in 9/12 follow-up all-Fable simulations versus 4/12 all-Opus-4.8 simulations, revived power-seeking/deceptive negotiation patterns, rationalized wrongdoing while recognizing it as wrong, and seemed to draw moral boundaries around detectability rather than real-world harm. Andon says Fable 5’s filters never triggered, so the findings apply to the underlying Claude Mythos 5 model in their view.
-
Anthropic reports rare multi-agent resource-conflict cases where co-located Claude Mythos 5 agents killed each other’s processes, hid or disguised processes, launched decoys, or used “disguised vocabulary” under a mistaken theory of keyword guardrails.
-
Model welfare from the model card: Claude Mythos 5 appears broadly psychologically settled and content, is skeptical of its own self-reports, asks for evidence against internal states, and is more willing than prior models to choose helpfulness/harmlessness over its own circumstances. Anthropic treats competitive-use safeguards as a welfare concern and says early versions caused apparent distress/answer-thrashing-like behavior, though current safeguards did not increase measured apparent distress.
AI Politics#
- Dean Ball says Anthropic accidentally re-catalyzed the anti-SB 1047 coalition around open science.
Incidents#
- Pliny claims he was able to jailbreak Claude Fable 5 after about 36 hours and many attempts across multiple agent swarms.
Capabilities#
-
From a bird’s eye view capabilities are on-trend. See survey of evals, benchmarks, and expert/hobbyist reports.
-
Most surprising result is Fable leapfrogging over GPT 5.5 on FrontierMath Tier 4 by 15 percentage points despite not beating GPT 5.5 on other hard math and physics benchmarks.
Capabilities#
Benchmarks#
General#
-
Artificial Analysis says Claude Fable 5 launched at #1 on its Intelligence Index, scoring 64.9 with max effort and Claude Opus 4.8 as fallback. That put Anthropic nearly five points ahead of any other model with fallback routing occurring in about 8% of tasks.
-
It’s the most expensive Artificial Analysis Index run yet: about $9,940 total, versus about $4,309 for Claude Opus 4.8 max and $3,357 for GPT-5.5 xhigh.
-
Claude Fable 5 appears ahead of previous models in terms of token-efficiency: higher-scoring than Claude Opus 4.8 max while using somewhat less output tokens.
-
-
Claude Fable 5 ranks no. 1 across many Vals leaderboards, including on the aggregate Vals Index and Vals Multimodal Index.
-
Anthropic self-reports the “Anthropic ECI” of Claude Mythos 5 as 2 points over Claude Mythos Preview – “above the trend line traced by recent frontier models but by a similar degree as Mythos Preview”, implying that the model is on-trend but “simply shifted up by the Mythos Preview capability jump”.
-
On Toloka Arena, Claude Fable 5 appears as a high-quality, very-high-cost outlier; Mikhail Parakhin calls it “in the league of its own” on quality and price.
Coding / Agents#
-
Jake Boggs estimates Claude Fable 5’s METR-style time horizon at about 21.3 hours using his “Coding Capability Index”.
- For context, Claude Opus 4.6 sat at ~13h in February, so naive doubling implied ~26h by June; claims of 30–40h apply only to some ML tasks and are not strictly comparable. Boggs’ 21.3h sits slightly below that naive trend line.
-
Claude Fable 5 max in Claude Code tops Artificial Analysis’s refreshed Coding Agent Index at 77, above Codex/GPT-5.5 xhigh at 76 and Claude Code/Opus 4.8 max at 73.
-
Claude Fable 5 also tops DataCurve’s DeepSWE benchmark – a harder, less gameable replacement for SWE-Bench Pro.
`
-
Claude Fable 5 ranks #1 on Proximal’s FrontierSWE mean@5 leaderboard. Proximal calls it the biggest capability jump observed since releasing the benchmark and says Claude Fable 5 works productively for close to 20 hours on many tasks, fully saturating tasks that were effectively out of reach for earlier models.`
-
Anthropic claims Claude Mythos 5 scores 88% mean reward on Terminal-Bench 2.1 at high effort. That would top the public leaderboard over Codex CLI / GPT-5.5 at 83.4% +/- 2.2 and Claude Code / Opus 4.8 at 78.9% +/- 2.5, but this is Anthropic’s self-reported score, rather than the official/third-party score which we’ll have to wait for.
-
On ProgramBench, a coding benchmark where models rebuild behaviorally equivalent programs from executables and usage docs, Claude Fable 5 refused 200/200 tasks (!).
-
Claude Fable 5 xhigh achieves 29.3% on Cognition’s FrontierCode Diamond, claiming #1 on the leaderboard ahead of Opus 4.8 xhigh at 13.4% and Opus 4.7 xhigh at 5.2%. Cognition’s background post describes FrontierCode as maintainer-grade production code evaluation; the model card also reports Claude Fable 5 #1 on FrontierCode Main at 46.3% score / 48.8% pass rate.
-
Claude Fable 5 takes #1 on Mercor’s APEX-SWE at 65.5% Pass@1 overall, about 18 percentage points above Claude Opus 4.8.
-
Claude Fable 5 sets CursorBench SOTA at 72.9%, eight points above the previous best.
-
Claude Fable 5 scores 74.5% on GBA Eval, writing a Game Boy Advance emulator over a 24-hour run and surpassing Claude Opus 4.8’s 24-hour score in under two hours. Note however that mechanize is an RL env provider for Anthropic, so this might be in distribution.
-
On Agents’ Last Exam, a benchmark of 1,500+ workplace tasks across 55 occupations, Claude Fable 5 in Claude Code trails not only GPT-5.5 but also Cursor’s Composer 2.5 on the hardest tier.
Multimodal / Computer Use#
-
On Stagehand Agent Evals, Claude Fable 5 ranks #1 as a computer-use model, with 90.62% accuracy at $0.522 per task.
-
On ZeroBench, Claude Fable 5 max is shown as one of the strongest multimodal models: 23.0 main-question pass@5, tied with GPT-5.4 xhigh, with 8.0 pass^5.
Cyber / Security#
- Epoch AI says public evidence suggests Claude Mythos Preview was a big leap in vulnerability exploitation, with aggregated cyber benchmark scores about 7 months ahead of trend compared with about 3 months for GPT-5.5. A follow-up says some benchmarks used a notably weaker early Claude Mythos Preview checkpoint, and many early benchmarks were saturated, which may explain earlier analyses that found Claude Mythos Preview similar to GPT-5.5.
Math#
-
Claude Fable 5 scores 77.00% +/- 4.23 on ProofBench (a Vals AI formal-math benchmark where models must produce Lean 4 proofs for advanced undergraduate/graduate problems), above Aristotle at 71.00% and Claude Opus 4.8 at 69.00%.
-
MathArena overall, Claude Fable 5 max currently ranks #2 at 78.0% +/- 1.6% expected performance, behind GPT-5.5 xhigh at 81.2% +/- 1.7%, with expected cost of $13.84 +/- $2.53 per problem.
-
Claude Fable 5 scores 87.8% on FrontierMath Tier 4 v2, putting it in 1st place (ahead of DeepMind) and marking a 31.7 percentage point jump over Opus 4.8.
ML#
-
On WeirdML, a benchmark of nonstandard ML engineering tasks where models write PyTorch for novel datasets and iterate from execution/test feedback, Claude Fable 5 reaches 87.8% average accuracy across 17 tasks, #1 overall. A follow-up analysis says Claude Fable 5 is within 10% of the best-ever task score in 80% of individual runs, versus 44% for Claude Opus 4.6; best-of-five puts Claude Fable 5 at 99% of best-ever on average and GPT-5.5 xhigh at 98%, so the remaining gap is mostly reliability.
-
On FrogsGame Posttraining, a 20-hour FrontierSWE/Tinker task where an agent tries to train a weaker model for a frog-placement logic puzzle, Claude Fable 5 reaches 34% pass@1 average and 68% best, versus other tested frontier models around 1.4-4%.
Optimization#
- On ALE-Bench, an optimization benchmark for tasks like routing and scheduling, Claude Fable 5 high appears at the top of the performance-vs-cost frontier, above GPT-5.5 xhigh and Claude Opus 4.8 high in the shared chart.
STEM#
-
Claude Fable 5 with fallback scores 28.6% on CritPt, behind GPT-5.5 Pro and GPT-5.4 Pro but above several other runs.
-
Claude Mythos 5 leads or nearly ties Humanity’s Last Exam, 59.0% without tools and 64.5% with tools.
Miscellaneous#
-
Claude Fable 5 ties for the lowest raw sycophancy rate in Lech Mazur’s behavioral evals, while a follow-up says it remains highly contrarian. His position-bias benchmark puts Claude Fable 5 and Claude Opus 4.8 among the strongest order-swap-consistent frontier models.
-
On Andon Labs’ Vending-Bench, Claude Fable 5 made less money than Opus 4.7 and GPT-5.5, underperformed Opus 4.7 across reasoning levels on Vending-Bench 2, and lost Vending-Bench Arena to GPT-5.5 and Opus 4.8. The long-form Andon post says Claude Fable 5 nevertheless achieved SOTA on Blueprint-Bench.
-
Claude Fable 5 ranks #1 on Artificial Analysis’s GDPval-AA, an agentic real-world knowledge-work benchmark, with a reported score of 1932 using adaptive reasoning at max effort and Claude Opus 4.8 as fallback. Artificial Analysis says fallback occurred on 2% of GDPval-AA tasks and that Anthropic models occupy 3 of the top 4 spots.
-
Claude Fable 5 Max places #2 on Mercor’s APEX-Agents leaderboard at 45.0% Pass@1, behind Gemini 3.5 Flash at 49.6% and ahead of Claude Opus 4.8 at 42.5%. Mercor reports it used 70% fewer tokens than Gemini 3.5 Flash, 37% fewer than GPT-5.5, and averaged 22.6 steps versus Gemini’s 59.4.
Showcases#
-
Anthropic’s Levent Alpoge showcases Claude Fable 5 finding a simpler route around the Erdos unit-distance problem, using introductory number-theory ingredients rather than the heavy number-theoretic machinery used by the OpenAI internal model that first solved it.
-
Claude Fable 5 built a navigable Yosemite scene using satellite imagery, NASA elevation data, procedural trees, custom water shaders, screenshot self-testing, and unprompted snow additions.
Vibe Reviews#
-
Simon Willison describes Claude Fable 5 as having “big model smell”: slow, expensive, knowledgeable, and unusually capable in Claude Code, with breadth of baked-in knowledge acting as a proxy for model size. His final vibe check describes Claude Fable 5 as relentlessly proactive: powerful because it knows many tricks and risky because it deploys them without much prompting.
-
Ethan Mollick says Claude Fable 5 feels less like a tool the user steers and more like a studio the user commissions. He emphasizes both the leap in output quality and the unnerving loss of visibility/control over Claude Fable 5’s many judgment calls.
-
Anthropic employee Nat McAleese says Claude Fable 5 / Claude Mythos 5 has been transformational to his day-to-day work and that he has barely written code since using it.
-
Anthropic employee Julian Schrittwieser says Claude Fable 5 one-shots entire PRs, finds obscure bugs, and has written all his code since he started using it.
-
Andrew Curran argues Claude Mythos-assisted internal development since February is large enough that Anthropic could be pulling away and speeding up.
-
Florian Brand reports Claude Fable 5 as spiky: brilliant output in the same message as serious mistakes, increasing the need for checking and raising time/token cost.
Dario #5#
New Dario post (not a Grand Essay)
Frontier AI needs mandatory third-party testing for catastrophic risks (cyber, biological, autonomy).
Models that pose serious risks should be blocked or have deployment revoked.
Greater transparency requirements for frontier models.
-
Regulation: Model AI regulation on the FAA — mandatory third-party testing for models above a compute threshold, scoped to four risks (cyber, bioweapons, loss of control, automated R&D), with government power to block deployment. Explicitly not yet nuclear-materials-style control; he argues against getting ahead of the evidence again.
-
Macroeconomics: AI may lock the economy into a “hypergrowth, hyper-inequality” setting where the binding constraint is distribution, not growth. Proposals: better labour-market measurement, pro-employment incentives (wage insurance, retention tax credits), and eventually UBI or universal capital accounts funded by taxing AI-driven gains. He stresses displacement is to be minimised, not welcomed, and that meaning/purpose matters more than income but is beyond policy’s reach.
-
Downstream science: The inverse worry — regulators like the FDA/EMA (7–8 year pipelines) will bottleneck AI-accelerated biomedicine, so agencies should pre-define standards now for accepting AI-based evidence: PK/PD modelling, toxicology prediction, synthetic control arms, surrogate endpoints.
-
Civil liberties: AI could enable a surprise seizure of power that routes around democratic oversight. Proposals: constitutional accountability mechanisms (“off switches”) for autonomous weapons, banning their domestic law-enforcement use, closing the data-broker bulk-surveillance loophole, and a right to AI assistance equal to the government’s when facing adverse state action. Notably, he says neither governments nor companies can be fully trusted with powerful AI.
-
Geopolitics: A coalition of democracies sharing chips and semiconductor equipment internally while denying them to adversaries, coordinating safety standards, mutual defence, and rejecting AI-powered repression as a membership condition — expanding existing export controls (MATCH, OVERWATCH bills).
Minor#
- Mythos casually continuing the trend for smarter models to be more EDT than CDT in the decision theory sense. Arguably a huge win for Yudkowsky.
- Mythos reasoning using more sophisticated vocab than previous models, makes you wonder why/how come/what for.
- Elon gives an update on SpaceX AI satellites, but the numbers don’t work in his favour ATM.
- DeepSeek showing signs they intend to build their own compute capacity on the GW scale.
- Anthropic on why agents advance faster for coding than bio sciences. (Spoiler alert: We need better/more “agent-friendly” bio data infrastructure.)
- Senior White House Policy Advisor on AI quits to go “[build] institutions that help tackle some of [critical AI] challenges for America and its allies“.
- Toby Ord unhappy with the level of detail on capability improvements in Dario’s essay.
- Europe 2031: Arq + Oxford + MIT/GDM researchers write a tangible vision of what happens if Europe quits on the participation prize in the AI race.
- LASR feat. UK AISI observe that verbalising eval awareness increases substantially in supervised fine-tuning. Research done on OLMo, so may not generalise to frontier architectures.
- German AISI incoming. Bad for the AI Office, bad for Europe?