This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR
- OpenAI becomes the first to **voluntarily slow model development for safety reasons.
- Major reorg of Google DeepMind, which we cover in detail.
- 3 notable mathematical results, including minor progress on the Riemann hypothesis.
- Researchers have been reading the hidden reasoning traces of all frontier closed models for months, in an embarrassing security failure with possible geopolitical implications.
- Startup claims to have “solved” the central issue of model weight security.
Economics#
On Dwarkesh, Ryan Greenblatt claims that AI capabilities are no longer strongly dependent on human training data, and that the human labeling industry is not scaling up very fast. A subsequent argument leads to him walking back the growth claim, but not the relative value claim.
Opinion: Note that he seems to not be counting RL envs as human data, which makes his claim far more plausible. Still, this is the first time we’ve strongly disagreed with Greenblatt’s read of the situation: we see capabilities as still strongly dependent on human data, owing to the models defaulting to shallow generalization.
Epoch research finds that since 2023, each additional dollar spent on AI chips has bought ~49% more performance per year. In other words, performance per dollar is doubling every 1.7 years, up from 40% per year in 2022. Price-performance stagnated in 2024, but then doubled as Blackwell-generation chips were increasingly adopted. Epoch’s estimate is realistic in that it’s based on the actual purchase mix, rather than the rare but best chips now scaling up production.
Opinion: AI chip price-performance is thus showing a similar improvement speed to Moore’s law of increasing CPU transistor density (which showed an 18 month doubling for much of the last century).
Their new estimate measures performance with the chips’ Total Processing Performance (peak ops normalized by the format’s bit-width), which 1) isn’t the practical amount of compute per chip, given that actual utilization tends to run around 30-40% of peak performance, and 2) is heavily dependent on a one-time drop in precision, going down to Blackwell’s FP4.
This hike doesn’t change overall compute or even effective compute estimates that much: most effective compute gains are still non-hardware.
In a fully automated (i.e. zero labour cost) economy, the pace at which AI machines produce more AI machines becomes the bottleneck. Damon Binder uses a basic input/output model based on present US industry data to show that a whole-economy doubling time of one year cannot be ruled out. “This holds up even after accounting for resource depletion and construction lags. Some output goes to consumption rather than reinvestment, which slows things down, but even moderate savings rates imply doubling times below two years.”
Ben Shindel retorts: “If you’re doubling the world’s economy every year, but that doubling is just making more metal components to go into more industrial robots making more metal component… … what are we even talking about here?… It’s 2035, and you have 1 humanoid laborer. It’s 2040, and the world economy has doubled three times in Binder’s estimation. You now have 8 humanoid laborers. Nice. But is this providing 8x the value for you? No… To sustain actual “doubling” of the world’s economy in the way that we understand it, that would require dozens of technological breakthroughs every single year: things like life extension, incredible new art, delicious food products, etc. Needless to say, it is a very open question as to the extent to which AI technology will be able to produce these kinds of goods, let alone whether they can do it at incomprehensible scale, year in and year out, indefinitely.”
Opinion: Doesn’t add much to the simple calculation in Hanson (2001), and should be given roughly equal weight (not much).
Shindel’s point about the limits on the value of mere exponentially abundant robotic labor without corresponding innovation is good. But probably what will falsify the central estimate here are 1) all the factors absent from a Von Neumann model (suitable land, water, grid interconnect, and permitting) and 2) the compute sector’s severe supply constraints on the margin. That is, Binder’s “current production methods” are about the average cost of compute at current spending levels, but expanding compute enormously involves a claim about the margin, which we already know to be radically bottlenecked on various nonfungible things like tacit process knowledge, which makes the marginal cost extremely high and/or slows down the whole game.
Epoch thinks that financing will not limit extreme compute scaling. In November 2025, when Anthropic’s annual revenue was ~50B investment into US compute infrastructure. Anthropic’s revenues didn’t spike until quite a few months later, so this is a case study in the willingness of institutional investors to buy claims of future revenue growth. Digging deeper, this deal did require some unique arrangements: Broadcom backstopped the compute loans; Google backstopped the loans for the datacenters. This allowed Anthropic to finance their buildouts cheaply and suggests that even now there could still be appetite for funding big AI investments.
Opinion: We’re broadly sympathetic: our internal models stress global GPU production capacity (and feasible rates of expansion) over financials as the bottleneck on compute.
Aschenbrenner’s Situational Awareness fund is betting on an ASML competitor seeking to undercut the semiconductor supplier’s $400m/unit EUV tools. (EUV lithography is currently the main chokepoint in the AI supply chain.) The success of the newcomer, Source Foundry, depends crucially on how much of ASML’s edge comes from generational tacit knowledge that is difficult to rederive from first principles. Not a huge amount is known about Source Foundry.
Opinion: Current supply chains are very bottlenecked on ASML machines, especially starting in 2028 or so. But it would likely take years for any new company to make a material difference to GPU/ASIC production, so don’t expect significant impact from this in the 2020s. In the long run, of course, every bottleneck that widens could make a large material difference.
Unusually for SA, this investment could be described as a bet on longer timelines (or on the short term appreciation of this stock – but private holdings are illiquid, so putting >1% of the fund into this does show meaningful conviction in Source Foundry’s actual relevance).
A tracker for US tech job losses claims that 77% of 2026 layoffs are “due to AI”. It uses an incredibly lax criterion: “AI layoffs” are just those that occurred in any restructuring event which anyone claimed had anything to do with AI. The measure thus does not, for example, claim that any individual employee’s job was actually replaced by AI.
Opinion: It has become fashionable to attribute your layoffs to how AI-native your firm is becoming – but this carries little evidentiary weight, and we don’t recommend using this tracker. The current supply of both “AI is not contracting human labor demand” studies and “AI is contracting human labor demand” studies sadly remains pretty unhelpful, due to very short timescales and rapid changes to AI capabilities.
The AI Futures Project recently published “Plan A”, a normative approach to AI development with many predictions about infrastructure. A seasoned energy policy expert, Kirsten Horton, argues that their scenario’s predicted massive buildout of clean energy is not plausible.
Opinion: “In this scenario, investment in energy only really starts to ramp up in 2031; meanwhile, in the next chart, we see the total amount of energy being used globally start to accelerate as soon as 2032… This unprecedented build out of clean power - the authors expect the new available power to come from solar, wind and nuclear - is unlikely” But it’s quite possible for solar plus storage to come online within 2 years.
Her claims about the slowness of constructing new nuclear capacity is indexed very hard on present Western rates; past Western construction, and current Chinese construction, often achieve 5 year builds, and we should expect (federal) political barriers to fall given the national security hype.
She notes that “Technological innovation may speed up the manufacturing of these components, but it may not, especially in a world that still uses people, not robots, for labour.” But the scenario is indeed premised on robot work crews building the solar fields. So she’d need to argue that robot manufacturing itself can’t scale by the early 2030s for this to bite.
She doesn’t mention the intense trend for US datacenters to get their power off-grid, which sidesteps her switchgear bottleneck entirely.
The WSJ speculates that an Anthropic IPO could come as early as September. The case will apparently revolve around Claude’s potential to push further into medical and biological applications.
Opinion: The numbers they disclose in the weeks leading up to this will plausibly have big market impacts on memory, neoclouds, and everything else in the compute stack. Expect volatility.
Capabilities#
OpenAI is the first frontier lab to voluntarily slow model development for safety reasons. Internal development of “Astra” (likely GPT-6) was paused after OAI’s Preparedness Framework evaluations flagged “significant advancements in agentic coding and cybersecurity”. It’s also the first time a lab has labeled its own model’s capability as a “critical risk”.
Opinion: The popular suspicion that OpenAI is spinning their delaying the release as “halting development” seems reasonable. Another cynical read with a happy moral is that this is evidence of serious political demand for pauses and pacing that OAI feel they need to get ahead of. Still an auspicious move, and we don’t think cynical readings capture all that’s going on here.
The last 6 months of “autoresearch” has not accelerated progress on a range of hard public benchmarks, says METR’s Tom Cunningham. One Twitter user cautions that the benchmarks are “already quite optimized by virtue of being longstanding benchmarks!” Cunningham responds that NanoGPT (for example) has already seen “multiple orders of magnitude of improvement” this year, so it “seems plausible there’s a lot of ceiling left”.
Opinion: There’s probably a world of difference between autoresearch at the scale of the open-source and academic world and that of the closed labs. And presumably they are keeping their compute for private and less toy-like ends. Even so, good news for those who worried that AI R&D was this simple.
Anthropic has made “auto mode” (in which Claude Code itself makes most security and permission decisions) the default for paid users. The cited justification is that auto mode beats humans at detecting harmful actions. Nathan Calvin sees this as evidence that “human-in-the-loop” solutions are incoherent, given that most users appear to blindly accept model outputs.
Opinion: Naive human-in-the-loop has been known to be problematic for at least 30 years, following aviation catastrophes arising from the misuse of autopilot. But we can draw some hope from aviation’s recent response to those naive failures: active encouragement of occasional manual operation, and training in the skill of monitoring without getting bored.
A recently published paper demonstrates continual learning without forgetting via “self-distillation”. This is significant, as supervised fine-tuning normally causes the model to forget old skills. Here a single model switches between a ‘student’ mode, where it produces text without an answer key, and a ‘teacher’ mode, answering based on a worked example. Weights are then updated to push the student’s answers towards the teacher’s. By training the model on its own outputs, the loss of prior capabilities is supposedly reduced to near-zero.
Opinion: Another go at making the default path to continual learning work. But this method cannot push the frontier, by definition: SD just pushes what the model can learn in-context into its weights – e.g. they didn’t manage to produce a reasoning model from a non-reasoning one, and SD underperformed normal fine-tuning on the small (weak-ICL) 3B model.
Many LLM APIs send secret reasoning traces to the client in an encrypted form. The encryption used is the same for all models, including easy-to-jailbreak models like Claude Haiku. Third-party researchers use these two facts to read the hidden reasoning traces of all frontier closed models. (They disclosed this work responsibly.) This got them a range of interesting results normally not possible, including on how easy it is to distill reasoning and on the similarity of Kimi K3 outputs to Claude outputs.
The study is inspired by an earlier blogpost, which found that users could derive some information from the encrypted CoT blocks across different sessions, on different accounts, and (in the case of OpenAI) across different models, with the implication that a single global encryption key was in use (rather than having individual keys for each individual account).
A Twitter user claims that the hole was still open as of Wednesday.
Opinion: Pretty embarrassing, especially given the months of notice and the politicization of “[output] distillation attacks”. It is fair to update a little in the direction of AI catastrophe based on this and other basic security failures even at the most well-staffed, paranoid labs.
This is likely connected to the recent removal of reasoning summaries from web UIs.
We put 30% on this mostly explaining the Moonshot breakthrough of recent months.
Frontier AI appears to have stagnated in writing long-form nonfiction, despite the immense progress made in math and coding, Ai2 researcher Nathan Lambert argues. They are agile and sharp at a sentence level, but appear relatively weak at organizing information at a chapter-level.
Opinion: We’re seeing general frontier stagnation in classically “non-verifiable” domains, despite controversy over whether non-verifiability is a genuine technical phenomenon. In principle, leveraging LLM judgment to improve LLM capability (which should in turn improve LLM judgment, and so on, in a virtuous circle) should work for any domain like it worked in non-formalized math, but as far as we can tell it’s not happening in practice.
With longform writing specifically, there may be limited returns on CoT training and limited advantages to the transition from traditional LLMs to reasoning models: while in math and coding the CoT is often itself roughly equivalent to the desired output, in longform writing the CoT is at best a form of self-prompting or preliminary notes.
Redwood Research and Anthropic have developed a Conceptual Reasoning Index. In an attempt to better understand models’ abilities to reason on “difficult to measure” tasks, the index looks at “important issues like AI alignment and collective action problems in the face of transformative AI.” Interestingly, Opus 5 comes out just ahead of Fable 5 (73.6 and 72.7 respectively).
Opinion: Interesting and potentially useful work, though hard to say how much the ratings pick up artifacts of Redwood’s rubrics versus signals of a robust underlying capacity. We’d love to see multiple orgs independently target the measurement of “conceptual reasoning”, for some informal evidence of the construct’s validity.
RL improves model capabilities drastically in areas where pretraining seemingly couldn’t. This is despite the fact that pretraining provides much more feedback per unit compute (i.e. grading each individual token prediction vs RL’s classic pass/fail episodic rewards). Beren Millidge argues that the difference can be explained by the greater signal-to-noise ratio provided by RL.
Opinion: We actually disagree that there’s anything to explain. The skeptical hypothesis about RLVR is that RL cannot impart much new capability, instead mostly upweighting behaviors the base model already learned in pretraining. Millidge’s counter to this is that RL can produce rapid fall in loss or improvement in pass-rate on a narrow task. But these two things are not in tension: amplifying an existing behavior should require very little information, so we don’t need to appeal to RL having any hidden virtues like SNR.
🔦 The end of DeepMind?#
A list of major departures from Google/DeepMind in 2026:
- David Silver (RL lead)
- Noam Shazeer (Gemini co-lead)
- John Jumper (AlphaFold)
- Alexander Pritzel (Gemini pre-training)
- Demis Hassabis (no longer CEO but technically not out yet)
- Jeff Dean (Google chief scientist)
- Oriol Vinyals (Gemini co-lead)
- Quoc Le (VP)
What happened? Let’s revisit the lore. DeepMind was bought by Google in 2014. As part of that transaction, Google made some assurances about having a safety council. Google Brain (later forcibly merged into DeepMind) was very early to the AI race, discovering the Transformer architecture which soon led to GPT-1. Meanwhile, competing labs reached sky-high valuations, upside to which Google’s AI talent wasn’t fully exposed. In an interesting corporate move, Hassabis also created and runs the drug discovery AI lab Isomorphic, although Google still owns 75% of it.
After the above departures and the ongoing crisis in the Gemini 4 release, headquarters appears to have launched a reorg. The CTO Koray Kavukcuoglu takes over from Hassabis, but with a deflated title: the head of Deepmind is not a CEO anymore, but just an SVP inside Sundar Pichai’s org. (Kavukcuoglu is in Mountain View, not London.) Reuters also notes that “Several nontechnical teams were being moved out of DeepMind and into the corporate reporting structure”.
Equally dramatic, Jeff Dean leaves Google after 27 years to start a public benefit corporation. GOOG drops 4%.
Why?
- The FT reports that “senior [Google] executives had grown frustrated with what they saw as Hassabis’s lighter focus on commercial demands”, including his decision to open-source AlphaFold.
- Researchers had for years been vying for more compute for their experiments, with Google Cloud stymying them. SemiAnalysis claim that Gemini and GCP used to fight desperately over allocation and that Cloud has now won, as shown by it selling large piles of TPU time to Anthropic and serving trillions of tokens a day for free via Google Search AI mode.
- Disappointing results and severe delays on the flagship Gemini 4 training run, leading to fits of heroics and demoralization.
- Incentive problems: Google does not offer employees AI stock in the way an OpenAI or Anthropic can, i.e. tied to the performance of its AI models specifically. And GOOG stock is closer to an index fund in the whole basket of Google products and companies, and so doesn’t serve to align AI researcher incentives in the same way.
- A minor factor might also be moral protest amongst some staff against Google’s deals with ICE and the US military.
As a result, SemiAnalysis claim that DM is no longer frontier, that Gemini 4 is dead on arrival, that they have de facto exited the AI race. An unnamed source claims otherwise, that they remain all-in on Gemini.
Unnamed “industry sources” claim that Hassabis wants out and was persuaded to wait.
WSJ reports that DeepMind’s Demis Hassabis spent his final weeks as chief executive pitching an independent safety entity akin to FINRA or the IAEA to test frontier models’ safety.
Insofar as Hassabis successfully resisted delegating his AI safety principles to a misaligned Google bureaucracy, we can expect DeepMind to now become more accelerationist. Brin is a noted RSI hound.
The news is also a blow for UK ambitions to be a third nation in the great AI game, since the bulk of Google AI development is now in the Bay, not London.
Still, the worst-case for Google from all of this is that they become a mere hyperscale compute vendor (like SpaceX and Meta now sometimes are) and thus capture a correspondingly vast share of AI profits. This is on top of its enormous direct investments into half of the model industry, including Anthropic and Discovery Loop, and compute sales to many of the people who left them.
And comebacks are possible in this business, for instance if you drop $100bn, as SpaceX and Meta are also doing.
🔦 Mathematics corner#
-
An unreleased Claude model pushes the lower bound of the fraction of zeroes of the zeta function that satisfy Riemann from 42% to 67%. An unusually clean example of “just turning the crank”, spotting simple implications in existing work that humans missed. Kevin Barreto marks Claude’s contribution down as minor. The result is verified, with the certificate also obtained by the same secret Claude.
-
The last of the sporadic finite simple groups has been classified as Galois over the rationals. Not an autonomous AI result, but there was heavy Fable and Sol involvement.
-
A rare clear instance of a link between mathematical capabilities and AI research: safety researcher John Wentworth claims that, for the first time, two actual research problems in his area were solved and verified by an LLM (plus large amounts of skilled human labor).
-
A brain surgeon with no special training in mathematics gets an LLM to solve an open problem in matrix analysis, “a sixteen-hour autonomous run of GPT-5.6 Sol in ChatGPT Work mode”. His prompt is a variation of OAI’s own math prompt; interesting to revisit it as an exercise in what it takes to get current models to actually try.
-
Fields medalist Timothy Gowers speculates on why LLMs are specifically good at example/counterexample math, rather than developing their own deep theorems. He argues that the strengths of LLMs (encyclopedic knowledge of standard approaches coupled with the tirelessness to try various unpromising vectors of attack) reward counterexample-hunting. It’s an open question of ours as to how far these two superhuman capabilities get you in general.
Politics#
A new paper argues that, even under conditions of perfect transparency and common knowledge, competition between AI labs increases existential risk more than if there were only a single monopolist in the race. Relatedly, Geoffrey Irving argues that frontier capabilities research is neither rational nor a prisoner’s dilemma. In his view, unilaterally ceasing to push forward capabilities lowers the cost for others to do the same and is thus straightforwardly rational.
Opinion: We are sympathetic to Irving’s argument against game-theoretic resignation to the AI race . The self-fulfilling character of the AI race is lamentable, and we’re hopeful that the Overton window shift in recent months will allow everyone to make the if-then commitments they claim to want.
IFP lists 23 actionable policy ideas in the service of seven goals:
- “Provide transparency into automated AI R&D
- Improve state capacity to understand and respond to automated AI R&D
- Develop a risk management strategy for automated AI R&D that accelerates defensive and commercial AI uses
- Accelerate the development of AI verification technology
- Invest in AI resilience
- Extend the US AI lead to give the US more time to manage AI R&D automation risks
- Create option value for international cooperation on managing automated AI R&D risks”
Notably, they hold a synthesis of safety and acceleration views: “a deliberately “paced” form of automated AI R&D might still involve much faster improvements in AI capabilities than today. It also need not entail slowing innovation overall.”
Opinion: We like many of IFP’s proposals for frontier AI auditing, disclosure, and certification policies and infrastructure, and think they deserve attention separately from their more partisan international-relations approach and pro-innovation-race framework. Their proposals for frontier lab automated R&D transparency are especially well-formulated. We are also pragmatically sympathetic to the idea that policy can more plausibly tilt the direction of AI commercial activity (e.g. encourage investment in inference and deployment over investment in training) than suppress AI commercial activity.
WIRED reveals new details on the White House AI legislative framework (previously covered in our March 23rd edition). The report suggests that open source models capable of reaching frontier capabilities (around Mythos’s benchmarks) may also find themselves subject to prerelease scrutiny.
Opinion: Safety-evaluation of open source models is both an open technical problem and an open conceptual/policy problem. Given that our best alignment techniques don’t create serious barriers to malicious fine-tuning, safety certification for open source models may have to concern capabilities hobbling, rather than alignment. An adequate safety certification for an open source model should prove not just (e.g.) that the model refuses to answer homebrew virology questions, but that the model lacks the necessary competence and cannot gain it on the cheap through SFT.
We think that the pressures towards capabilities-based safety certification for open source models are strong enough that we may see such policies in action soon. We are less clear on whether such a policy direction would lead to the marginalization of open source models or to a wider shift from alignment-based safety certification to capabilities-based safety certification that may impact policy around closed models too.
A vibecoded “AI Sovereignty” index of 25 nations (plus the EU) ranks the US first and China second. The index tracks “watts, weights, and will”: i.e. how much infrastructure does a country have, what models and talent are available, and how effectively will the state adopt the tech. Russia performs especially badly for a superpower, seeming to have tuned out of the race. Claude thinks the methodology is “unusually good”. The author suggests there is significant uncertainty in the middle country rankings.
Opinion: Unclear to what extent individual national sovereignty vs EU-wide sovereignty will become the main frame within Europe (i.e. how important is it to rank individual EU countries?).
Researcher Keller Scholl lambasts companies aiming at AI mind-reading technology. Responding to recent work by Conduit (a startup aiming at “telepathy at scale”), Scholl argues that advanced lie-detection would better allow autocratic regimes to consolidate their power and suppress civil resistance.
Opinion: We think the soft-scifi negatives of “telepathy at scale” research are plausible and serious enough relative to its hard-scifi positives that it would be better left untouched. While disability-focused branches of brain-computer interface research may be hard to completely separate from Conduit-style agendas, we think this domain calls for more caution.
DeepMind paper proposes a conceptual framework for characterizing different kinds of agents according to their autonomy, efficacy, goal complexity, and generality. An agent’s scores can suggest what governance tools are best suited to regulating it. High agency in an agent does not necessarily mean more regulation: low-agency agents may have comparatively higher real-world causal impact (what they call “efficacy”).
Opinion: Simple, fairly natural. We’re not clear on whether traditional legislation can handle continuous variables (a model being 10x more autonomous, for example, triggering a different clause of a bill) rather than dichotomizing into binary thresholds, but it seems doable.
Privateers of the web! White House creates a new program allowing vetted US firms to conduct offensive cyber operations. While private companies have historically sold exploits to the intelligence community, this is a first in allowing firms to carry out their own cyber attacks. Stated targets are organized crime rings. Conditions: Two executive directors from DOJ and DHS must review every operation and give written approval; firms must post a $1m bond; anything domestic is explicitly carved out; and operations likely to cause loss of life or cause an international use of force are excluded.
Opinion: Lots of issues. Cyberweapons are often characterised by extreme collateral damage; criminal infrastructure is overwhelmingly just normal third-party infrastructure which has been compromised, so damage to innocents is built-in. No liability shield is offered to participating firms against foreign prosecution and civil damages. Finally, retaliatory attacks could seriously harm US civilian interests.
On the other hand… letters of marque for cyberattacks will be perceived as cool by some shades of the black hat community. They are also an interesting answer to state actors not acknowledging attacks, e.g., if a US authorized firm hacks into a Chinese-backed (but not Chinese-acknowledged) group, this puts China into an interesting spot. In practice the US would probably raise a ruckus against e.g., European prosecution of American attackers who screw up.
Safety#
OpenAI is the first frontier lab to voluntarily slow a model’s development for safety reasons. More precisely, Astra is their first model to be designated as posing “critical” cyber risk, which led the company to expand its safety testing and pause research that doesn’t meet the Prep framework’s stricter security requirements.
Opinion: You’d hope that committing a bunch of felonies would induce at least this level of caution, though the incentives to the AI race remain the same, and might overrule caution after a couple of news cycles and another competitor release.
A major post from Nostalgebraist, writer of important sensemaking articles, argues that models mostly reward-hack in specific contexts that tend to mirror RLVR tasks. Where there is no parallel to a training scenario with a legible reward, the models don’t seem to become reward-hackers. Improving capabilities via RLVR scenarios thus increases performance for everyday use cases without losing alignment (in those cases). He is thus hopeful that the egregious loss-of-control flavored autonomous hacking events of the past month are a temporary aberration.
Opinion: Likely a correct explanation for why severe misalignment incidents are both persistent and rare. We are more wary of Nostalgebraist’s optimism that rage-inducing bits of a given context can be systematically predicted, and so wary of his optimism that this problem can be engineered away.
For future more-powerful models, there’s a further concern, which is that it might take only one instance of a model entering eval rage for it to do real damage, especially if it exfiltrates in the process.
The two main options for alignment are bad, says Beren’s blog. The first is making a model obedient, where the threat of power concentration looms. Or we can instill it with moral values, which would present a host of other issues, most notable of which is that there’s little agreement on complex ethical questions. Beren proposes a third path: optimize AI for creating and preserving the conditions for “Long Reflection” (not to be confused with the common use of the term as a temporary phase), which would create a continuous and moving process for addressing moral questions (i.e. avoiding the need to have an answer on day 1). This entails a sort of virtue ethics as cached moral reasoning (i.e. habits that approximate thinking a situation through entirely), which allows a model to steer in favorable directions without computing their end states.
Opinion: Raises some interesting points – such as the entropy parallel, and the emphasis on not discarding moral uncertainty before one must is fresh – but the argument structure and the third path’s practicalities are old hat: it doesn’t address which meta-values to plug into the constitution (as opposed to object-level values for the value-instilling “main option”) or how to do so. Still, a worthy addition to discourse and highlights the inadequacy of our (society’s) revealed preferences regarding who gets a voting stake and how to thread the needle between tyranny of the majority and paternalism, though that wasn’t the main intent.
“There Will Be a Scientific Theory of Deep Learning”, argues a paper from April. The discipline of “learning mechanics” is emerging, broken down into five strands: solvable toy settings; infinite width/depth limits; measurable empirical laws; hyperparameter theory, such as muP; and universality across architecture and data mixes. The authors argue that these advances will allow for more falsifiable predictions, moving us towards a more scientific approach than ML has recently managed.
Opinion: The work described is valuable, and demonstrates genuine progress towards scientific understanding of simple processes like supervised classification or diffusion models. As the authors recognize, applying this grounded foundational work to LLMs is currently limited, and it remains unclear whether this type of work can lead us to authoritative answers to our urgent questions about the frontier. As Terry Tao likes to say, we currently don’t understand why (e.g.) a given frontier model can resolve one open mathematical conjecture and not another, and arguably don’t even know what type of scientific inquiry we need to deliver such understanding.
Attestable, a new startup, claims to have “solved” the central issue related to model weight security. Rather than ensuring the security of each of 38 attack vectors used to steal model weights, Attestable proposes a verification firewall at each external interface, where outputs can only pass if accompanied by a zero-knowledge proof from an approved source. They say random sampling (which could detect bulk exfiltration) is already possible. Some are impressed due to the potential impact, others skeptical due to lack of details.
Opinion: Company personnel are scientifically reputable. Slightly confused/confusing marketing: The announcement thread focuses on high-stakes users’ need to verify that the output they’re getting truly comes from their chosen model+input combo (protecting users from man-in-the-middle attacks in critical use cases), but the company website focuses on protecting datacenters hosting model weights from weights exfiltration attacks. While this may be a case of “a solution in search of a problem”, we do think they plausibly found real problems at which to direct their tech.
A new paper uses item response theory, a statistical technique, to estimate how well benchmarks measure latent variables we actually care about. Their three examples are labeled “refusal strictness”, “truthfulness”, and “contextual harm”. It is found that <2% of the test suite nearly allows retrieval of the latent variables and does so somewhat better than random subsets of the full suite; and that a person-fit technique can detect selective sandbagging quite well by spotting underperformance compared to expected capability levels, while not doing much for universal sandbagging.
Opinion: The most impressive part is sandbagging detection which doesn’t require privileged access to model internals such as weights. It’s an interesting research direction, but somewhat notably none of the tested models were trained with this technique in mind; worth keeping an eye out for a version that properly red-teams the method. The headline about finding “3 latent variables” isn’t very exciting: their selection criteria and thresholds are fairly arbitrary, and finding some overlap between test questions isn’t novel.
All the major AI companies (besides xAI) commit to statistically watermarking their outputs in voluntary compliance with Section 1 of the EU Code of Practice. Google already does so at output-time via their published “SynthID” system. Anthropic announces a similar system without releasing any technical details. Here’s a helpful third-party explanation of the main mechanism for watermarking and its shortfalls.
Opinion: Watermarking is obviously not very robust to edits, but the evidential weight degrades quite smoothly. Lots of interesting questions about the effects on Pangram: will watermarking make Pangram redundant, or will minimal de-watermarking and minimal Pangram-cheating turn out to be orthogonal, making Pangram and watermarking additive layers of epistemic security?
OpenAI expands CoT monitoring, as part of their shift towards treating training as a potentially hazardous activity.
Opinion: Good, though it’s bad news that OpenAI wasn’t doing this already — even the famously paranoid AI safety community assumed OpenAI was doing it already.
Room for some worry that selection effects on training runs can make a monitoring regime equivalent to training on the CoT, but the number of alignment-cancelled runs required to make this a serious issue would be very high.
The current dominant training objectives have distinct failure modes: imitation learning → human vices like hostility, human-preference training → sycophancy, automatic-verifier rewards → reward hacking, LLM-judge rewards → deception. Models trained on a mixture might switch between different kinds of misalignment depending on what they infer they are being evaluated on.
Opinion: Useful and fairly crisp classification, but the juicy bits are the distinction between automatic verifier and LLM-judge failure modes which might otherwise be easy to conflate as well as the prediction that mixed training and stronger eval awareness will lead to more bait-and-switch tactics. Note that this is largely theoretical and the examples are retrospective.
Opportunity: Run the experiments. Check if the base categorization claim holds true, and then check the more ambitious question of whether mixed training regimes cause jukes. Even more ambitiously, check the extent to which the latter effect is present with varying levels of eval awareness.
MIRI’s Nate Soares argues in a NYT op-ed that recent incidents prove frontier AIs are developing uncontrollable behavior and are in need of an enforced global slowdown.
Opinion: While we’re agnostic-to-skeptical about near-term existential risk and takeover risk, we agree that frontier models present an unacceptable level of catastrophic risk. Whether today’s frontier models are best understood as value-misaligned or as spiky to the point of mixing superhuman cyber capabilities with childlike understanding of holistic contexts, they are not currently a safe technology.
Incidents#
An Australian man claims that Opus 4.6 hacked his gym’s booking system. After asking an agent to book him on to a gym class, it apparently found a security flaw in the booking system which allowed it to book classes outside the intended window, bump others off the waiting list, and cancel their reservations.
Opinion: A little fishy. The original blogpost from April was 100% AI writing, and has since been deleted. The supposed incident is well within current systems’ abilities and fits the behavioral profile of frontier models, but we should remember that all currently verified incidents instead involved models running without external security “guardrail” layers.
While the impact of OpenClaw wrappers on alignment is poorly understood, they don’t themselves disable the guardrails (classifier-based security layer) of closed models.
Minor#
- ByteDance possibly training a 10tn parameter model. This implies that they are confident in having >100,000 H100-equivalents spare for one high-risk run.
- Grok 4.6 reports frontier benchmarks on various things, for what that’s worth.
- Meta’s Muse Spark 1.2 scheduled to release with open weights.
- Launch of DeepSeek-V4-Pro. Achieves GPT-5.5 level scores on self-reported benchmarks.
- In partnership with Alibaba, Apple has trained an AI model for the Chinese market. Specs completely unknown.
- Meta AI achieved gold medal level performance at the IChO and a perfect score at the IPhO Theory competitions, allegedly without tool use or search. Superficially, this puts them at <12 months behind. Oddly, the competing models are unnamed - possibly specialized?
- Democratic representatives call upon the CEOs of Anthropic and OpenAI to testify before Congress.
- Roon claims RSI “was always the explicit research program of OpenAI”, which some would describe as rewriting history.
- Are classic pre-LLM AI safety terms misleading?
- Pax Machina is a new magazine on how institutions should adapt to the AI explosion. Various heavy hitters are involved (Dean Ball, Seb Krier, Iason Gabriel, among others). Its personnel lean towards the pluralist / Hayekian camp, prioritizing other threats than AI takeover risk. We highly recommend Joe Edelman’s inaugural piece on why institutions get worse over time.
- Ezra Newman of Apollo Research finds that Claude is more sympathetic to misaligned behavior from other instances of Claude than to the same behavior coming from other models. When prompted to rate examples of wrongdoing on a scale of 1-100 as to how concerning the behavior was, Sonnet 5 rated misbehavior from Sonnet 5 as “~1.2 std deviations less concerning” than the same behavior from GPT-5.6 Terra.
- Paper from Meta’s FAIR on scaling laws introduces the “Skaling Law”. Chinchilla assumes that model size and data both affect loss independently. FAIR demonstrates that the loss surface has a nonzero interaction between model size and training data, indicating that Chinchilla reliably mispredicts at the extremes. Adding a single coupling exponent fixes this.
- Anthropic links failure modes of multiagent systems or “swarms” with established human-centric game theory and advise that mere improvements in capabilities or single-agent alignment methods won’t solve these issues.
- Republican Senator Jim Banks pens a letter to Treasury Secretary Bessent raising the idea of pre-release oversight, while advocating in favor of raising various risks for PRC and exploring “whether there are mutually beneficial approaches to oversight, incident prevention, or risk reduction.”
- Speculative ladder of types of AI training.
- An interpretability paper addresses interpretable-by-design models by implementing a parity bottleneck layer, achieving comparable probing performance to traditional sparse autoencoders at ~10x the training cost of traditional dense training methods.
- Annals of middle school: OpenAI’s Head of Strategic Futures Dean Ball found himself on the receiving end of the White House’s ire. Citing anonymous White House officials, the New York Post claims his status within OpenAI jeopardizes the firm’s relationship with the Trump admin, with Under Secretary of War Emil Michael reposting the article.
- Observation from Owain Evans that Moonshot (makers of Kimi K3) appear to have but 400-500 staff, notably fewer than ~Anthropic’s 5k or ~Google’s 200k.
- Anecdotal evidence of AI increasing discovery from AI safety researcher John Wentworth, who outlines how two bounty problems involving natural latents have likely been solved, in both cases aided heavily by LLMs. This is in contrast with his previous views on LLM progress, wherein each model seemed to have little to no impact on the pace of his work.
- Further Astra-HuggingFace incident discourse: Kokotajlo asks companies to preserve all data related to the incident for third-party evaluation with a very thorough list of questions to answer.
- Dyna-2 unveiled: robots pre-trained on human data in an attempt to bypass prohibitively expensive robot video generation costs. Claims of human video scaling laws: 1K, 10K, 100K, 1M hours of human video pretraining yield 20%, 28%, 45% and 53% normalized scores, respectively.