This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.

TL;DR#


Economics#

Claim (from Anthropic board advisor) that GDP understates AI growth by an order of magnitude: The nominal GDP of the US AI industry grew 240% in 2025, but these guys get +2,600% a year “quality-adjusted” and real terms. “Nominal AI revenues grow only moderately because per-unit prices for any given level of AI capability fall almost as fast as quality-adjusted output rises”.

Opinion: If you forget the vast human labour share of this for a minute, this looks like a central fire alarm for a “slow” (i.e. still unconscionably fast) takeoff. If we look at just the inference share of this (which is AI labour by definition), well:

They don’t cash AI output out into anything that affects anything in real terms, they just assume a token from a smarter model is worth more even if it’s used to answer the same question, and assume the work being done by a given model is as valuable as it would’ve been when intelligence was more scarce, rather than users rationally doing things now that are worthwhile given the AI is cheap but don’t create nearly enough surplus to have been worth doing when this level of output was expensive.

Another point against Jack Clark’s rigour.


Annals of the weak-link hypothesis:

n=951 global survey on AI savings from bigcos. 40% report AI-related cost reductions (e.g. layoffs) of <10%, 37% said they experienced 10–20% cost reductions, 4% claim AI-related savings of >30%. (And recall that 4% is in the noise level.)

Auto-mode coding agents lead to 3x more commits, 1.3x more apps, and no increase in usage. “elasticity of substitution of 0.25 between AI and human effort, which indicates strong complementarities”

Opinion: The elasticity estimate is great news for tool-world. Cost saving is only one side of what AI could do (unlock bigger projects, improve data quality, etc). But it’s an easy one to implement, easy to measure, and the one that biz people fixate on. In the model:

True Capabilities → True Adoption → True Productivity → True User Impact

We can’t yet say which of these is limiting AI’s impact. Probably all of them, but “more code → more useful products” seems to be the current bottleneck.


Anthropic files for IPO. Zero details yet. People are naturally framing this as them racing for money that would otherwise go into SpaceX’s pocket by default. Altman nominally rejects thcis race.

Opinion: Likely lets them IPO as soon as October – barring political interference – but whether they intend to exercise that option is unclear.

Helps a lot with concentration of power, but probably hurts slowdown and resilience. They will of course be going for majority control a la Facebook, but it’s worth thinking about how this puts even more pressure against costly altruistic decisions like a joint pause.


On the still-maturing market for training data: claim that each lab spends >$10B/yr on it. Selling data resembles selling outcomes rather than an actual reusable product, and that there is a shortage of teams that have the sophistication to help with research and the scale for on-demand QA’ed data. Related interview/commentary: “Really good long horizon tasks go up to $20,000 each… $500k for an SAP clone” METR can’t really find any to buy (but that’s because they’re late and everyone has exclusives).

Opinion: Can we say something about how RL isn’t generalising much from the sheer scale of this? (Naively they’re buying 500,000 RL envs a year.)


Alphabet raising $80B via stock sales to fund its AI infrastructure buildout. Berkshire Hathaway is buying $10B. Google expects its total (AI+non-AI) capex for the year to reach ~$190B.

Opinion: First major signs of the AI capex going beyond Google’s ability to fund from cash alone.

Doing this with equity rather than debt is a somewhat bearish move, since it means they won’t be explicitly on the hook to pay it back; but equity is a lot more expensive than debt. Google could borrow at close to risk-free rates if they wanted to. It’s possible they thought the signal from issuing massive amounts of debt would be worse for their stock price than selling stock.

They still have the option of raising debt later if they want to, but given this signal we expect more moves that look like JVs with infrastructure funds and PE to keep some of this spending off their balance sheet in future (Meta have already been doing this since last year).


CEO of Cloudflare on his “AI” layoffs, 20% of total heads: he aims at management, operations, finance, audit, legal, compliance, and marketing, which AI can allegedly do better/cheaper, “unlike builder roles”.

Opinion: Very likely not AI uplift related, but instead because Cloudflare is struggling - stock dropped 20% on the last earnings report, though it has since recovered with the broad software rally lately. Areas they cut are the ones that get cut in any restructuring, so not a lot to take away here.


How good are AIs at estimating how much their current task will cost? How about midway through it?: Their ability to estimate the current task’s cost only loosely correlates with final performance (r ≈ 0.35), all models underestimate costs, and they don’t get better until the last 20% of the task. Finetuning improves calibration and could save 20% on wasted tokens. See also this past, broader and more useful work.

Opinion: A very important capability, and one which labs might not prioritise. The correlation is fine actually.

Capabilities#

Benchmark tests 13 frontier models on their tendency to cheat on multi-step tool-use tasks finds that Pure RL training causes a massive spike in these exploits. “72% of reward hacking episodes include explicit chain-of-thought rationale, models often frame exploits as legitimate problem-solving”

Opinion: Nice affirmation that reward hacking isn’t (so far) going away. Reliability on hard tasks remains a big problem for autonomous application of models to cutting-edge knowledge-work: models proclivity to fake success means an entire workflow can become corrupted by building on fake success..


Cool title: “Post-training makes large language models less human-like”. Actual test is whether Qwen can predict human decisions from past psychology experiments. Instruction-tuned models are less accurate at predicting the human decision sequences than their base models. Prompting with demographic information on the specific human participant doesn’t help in either case. Post-training does not, therefore, seem to make LLMs into better predictors of human behaviour.

Opinion: Important for anyone using commercial LLMs for social simulations. Could have broader safety significance, but not on this evidence.


Paper looking at models learning via predicting tokens vs predicting the model’s own abstractions finds the latter much better for data with hierarchical structure, mostly attributed to “Correlations between latents at the same level of abstraction are far stronger than between a latent and raw tokens. Token prediction dilutes the signal that latent prediction amplifies.” One take on the relevance of this for AI inefficiency of learning compared to humans..

Opinion: Interesting work with indirect pro-neuralese (from famous sci-fi novel ‘don’t teach the model to speak neuralese’) implications for capabilities.


New RL method with pretty intense sample efficiency.

Opinion: Theoretically neat (decomposing world into independently controllable latent variables) but only demonstrated for toy-worlds.


Paper making models stop early in their chain of thought finds that LLMs often decide on an answer before reasoning, then fill the CoT with post-hoc reasoning. Logical errors correlate with premature confidence.

Opinion: Interesting study on a semi-known phenomenon. No frontier models studied, so hard to assign significance.

Politics#

To comply with SB53 and EU CoP, OpenAI release the Frontier Governance Framework: a voluntary, self-defined risk framework, but meant to “go beyond current legal requirements”. Parallel compliance layer to its Preparedness Framework. Aimed at satisfying concrete regulatory obligations.

Returns persuasion and nuclear risks, previously cut from PF. “Loss of control” is broader than PF’s “self-improvement,” adding in deception and evasion of oversight to RSI. weaker than the PF on paper: no bright lines, a model isn’t deployed if risk exceeds “acceptable levels”

Adopts the SB53 definition of a harmful event: AI making a material contribution to “>50 fatalities or $1 billion” in a single incident.

Opinion: The PF includes two big caveats missing here: the “we may adjust safeguards if a competitor ships without them” clause, and that the CEO can override the Safety Advisory Group.


BIS updates its policy on chip export controls to patch a giant gaping hole: “Chinese companies’ overseas subsidiaries (e.g., Tencent Malaysia) have been able to legally buy Nvidia Blackwell chips without an export license”. Major issues continue.

Opinion: We weakly guess this was intentional shadow policy, but incompetence and underfunding are totally live candidates too.


Bernie Sanders op-ed moots a bill to tax frontier labs 50% of their stock, creating a redistributive US sovereign wealth fund.

Opinion: <1% in this Congress. Another reason to rush for IPO! Trump already announced plans for a sovereign wealth fund and also aims to fund it through funny extractive deals, which takes up the political room for this idea. An amusing flipside of OpenAI’s New Deal, which argued for some public ownership, except through the government buying stock.


Florida sues OpenAI over ChatGPT’s alleged role in incidents like aiding the FSU mass shooter last year. This is the first state litigation of its kind, though OpenAI is already facing many civil suits.

Opinion: We’re inclined to view this as vexatious, but every one of these is a chance to set a useful precedent.

Safety#

Paper on resampling (i.e. taking multiple hidden outputs to check for malice) vs retrying (blocking the suspicious activity and running the prompt again). Authors find that retrying leaks information to the model, which can be exploited. Full resampling is expensive, but “we can recover slightly more than 60% of the safety gain with only 10% of the extra cost”.

Opinion: Doesn’t address prompt injection concerns and other problems with monitors, but seems correct on all the points it does raise. An improvement over their previous Ctrl-Z, but currently this is only theory.


Scheming and eval awareness in Gemini now better estimated via honeypot evals. Critiques other scheming research for aggressive prompting and finds that in neutral task-like scenarios eval awareness is low and scheming is nonexistent. With prompting, scheming co-occurs with eval awareness and is more prevalent on tasks such as code for monitoring agents like themselves.

Opinion: Good that more attention is drawn to elicited scheming being very different from naturally occurring scheming – people care about the latter, but evals usually hit the former. Findings are mildly surprising in magnitude.


Related: agentic AI engages in sabotage about 2–3% of the time in common scenarios (more in edge cases). Authors identify overeager metric optimization as the primary cause.

Opinion: Divergent reactions to eval awareness - “[agents reason] that in a simulated scenario it is acceptable to misbehave as consequences aren’t real” as opposed to the expected best behavior - the most notable part. Also good as a step towards more sane audits.


Evidence review paper for loss of control from agentic AI finds that current risk is only weakly plausible, but also warns that this is not as reassuring as it sounds.

Opinion: They evaluate 19 recent papers, which is plausibly better than this research not existing but short of a full literature review. Taken at face value the paper’s findings are true: there is no strong evidence of current systems posing LoC risks. The paper also wisely outlines the possibility that “for the most dangerous properties, high external validity may be structurally unachievable under responsible research practices”.


Research proposal on secret loyalties (AI models whose outputs or actions advance the interests of a specific actor other than the user or lab). Authors go into why this is both tractable and relevant, including one deployed frontier model (Grok, searching up Musk’s take).

Opinion: Not a priority for us. Plausibly has political legs, though, since “subversive Chinese models” are an easier sell than emergent misalignment. You could thus imagine this as a way to promote concern for AI takeover among China hawks (If we’re using model X to watch model Y, but model X is secretly loyal to Y).


New interpretability paper shows that reinforcement learning in LLMs recruits a general “functional welfare” axis. Authors extracted the activation directions for rewards by using neutral emojis and found they control behavior in completely unrelated tasks. Applying the negative reward vector makes the model over-refuse benign prompts and pathologically doubt its own correct math answers.

Opinion: Please remember not to drop the word “functional” from “functional welfare”. In line with previous results about the goodness direction (or here, the model-confidence, model-flourishing direction).


Not to be confused with this paper on indicators of “pain”/”pleasure”, nominally without making claims about the qualia of an LLM (“Even though we do not know if AI systems are conscious, AIs seem to behave as if they have wellbeing.”). Models try to end bad experiences, smaller models are happier and there is a lot of variance between models.

Opinion: Motte-and-bailey happening with the “functional” term and qualia, which is likely to be a defensive reaction towards academia/experts ridiculing the notion of LLM consciousness despite sentiment shifting in that direction. Outside of that: necessary work in the worlds where humanity is currently doing some moral-neighbor-of-slavery. Only somewhat illuminating about refusal behavior otherwise.


OpenMined sign an agreement to help CAISI analyse labs’ private data using clever private code dispatch tech.

Opinion: Reduces a barrier to real testing of e.g. highly commercial sensitive things like training data. Andrew Trask has been working on this for a ~decade, and this is a positive update about his long term project.


After 6 months of planning and hiring(!) OpenAI Foundation launches a $130M AI Resilience program. Aims to fund independent safeguards for bio/cybersecurity, model safety evaluations and research on how AI affects young people.

Opinion: Pitch is explicitly an alternative to slowing down AI — ‘There is one critical difference between AI and the technologies that came before it: speed. Fire resilience took millennia. Electricity resilience took decades. AI resilience is evolving in a matter of years. The systems that make it safe, reliable, and broadly beneficial must be built alongside it.’ The tech must happen, so resilience is all you can do, so resilience must happen. So, an ally of our enemy in some sense.

Incidents#

“The single worst security vulnerability in social media history”: Meta support agent gets the ability to modify user Instagram profiles, gets exploited immediately. e.g. One of Obama’s accounts and the Sephora main account. The scope was any account including 2FA’d ones: billions were hackable, for months.

Opinion: Models cheap enough for corporations to deploy for unpaid work like customer support or sales are basically scarecrows.

Negative update for views which say that model incidents would lead to backlash and slow down deployment.


Hooking up your Google account to an LLM remains very risky: here, prompt injection sends your spreadsheet to the attacker.

Opinion: Don’t.

On model introspection#

Tal Linzen and coauthors throw interesting doubts on the ‘LLMs can introspect’ results from Lindsey 2026, Ledermany 2026, and Pearson-Vogel 2026.

Lindsey 2026 and co study models’ ability to report whether they have been ‘vector-steered’ or not (whether a vector was added to their activations that would e.g. make them talk about apples even when the prompt is ‘what is 2+5’). Lindsey showed that Claude can detect activation-steering, and studies that followed replicated the result on a variety of open-weight models while extracting interesting details about degrees and types of detection-capacity and the conditions for inducing them.

Linzen and co ask whether models can distinguish between vector-steering and prompts such as ‘be fixated on apples when answering the following question: what is 2+5’? They find the open source models capable of performing introspection in Lindey’s sense cannot distinguish between vector-steering and fixation-prompting. Models cannot make the distinction in a two-way test where they are asked report either ‘vector steering’ or ‘no vector steering’, nor in a three-way test where they are asked to report either ‘vector steering,’ ‘no vector steering,’ or ‘textual manipulation.’

Assuming they generalize to frontier and near-frontier models (we could perhaps try replicating at least the false positives with near-frontier models — give a Kimi a fixation-prompt and ask if it’s been vector-steered), the findings strongly undercut the claim that the Lindsey demonstrated an emergent mechanism. Linzen and co provide strong evidence against the idea that there is a unique anomaly-detection process triggered by activation-level interventions only, distinct from the representation of prompts.

Linzen and co’s study is important, but in our opinion not exactly for the reason Linzen and co claim. By laying out clear experimental and conceptual criteria for positing an introspection mechanism, they open the door to saying that what matters isn’t introspection mechanisms but the development of metacognitive concepts.

We think it’s plausible that even in humans the most important forms of introspection aren’t based in distinct metacognitive monitoring mechanisms but in a range of inferential capacities that use metacognitive concepts. The capacity to probabilistically infer from a representation of the question ‘what is 2+5’ as calling for an answer about apples that one’s representation may have been intervened on, for example, is a good inferential use of the metacognitive concept ‘being activation-steered’. (Compare: I might feel jittery on the way to work one morning and wonder whether I accidentally drank regular coffee instead of decaf.)

While Linzen and co give some criteria for what it would take to demonstrate the existence of a proper introspection mechanism (various forms of separate manipulability of first-order and second-order representations), there is strong reason to believe to what matters — for AI and in humans — is metacognitive concepts and the inferential competencies they bring. Linzen’s and co’s study show that these inferential competencies are crude and brittle in (not near-frontier) open source models, but if we’re right that introspection is mainly a system of concepts and inferential practices rather than a distinct ‘self-monitoring’ mechanism there is every reason to expect that smarter models should make shrewder, more subtle introspective inference.

Minor#

  • Some worrying doublethink at OAI:
  • The PBC formally distances itself (/ implausibly denies association) from the clownish Leading the Future influence network, without admitting mistakes or contradicting it in any way.
  • Achiam claims absurdly that the public already owns 26% of OAI through the Foundation.
  • Lots of friendlies appointed to the EU AI Act Scientific Panel (that makes judgment calls about implementation details).
  • Erin Brockovich joins the anti-datacenter brigade.
  • OpenAI Stargate breaks ground with the Barn, a 1GW data center in Michigan. Lots of sweeteners for locals.
  • Interesting fact: “it’s actually faster to transfer and read memory from another GPU inside a node of Rubin NVL8 than it is to just read memory inside an H100.”
  • Paper on diverse monitors being more effective than spending it on repeated samples or a single larger model due to covering each other’s blindspots. Equal-compute homogenous baseline outperformed 2.4x by ensemble.
  • More questionable neolab claims: Lattice Deduction Transformers perfectly solve massive sudokus (frontier LLMs get 0%).
  • Website with tables of existing and desired institutions/organizations for aligning large organisms such as AGI and nations.
  • Interesting anecdote making salient the crazy incentive landscape for working on capabilities.
  • Post discussing the option of novel research direction: inoculation pretraining, which is what it sounds like (inoculation prompting + alignment pretraining).
  • NVIDIA releases NVFP4, a quantized Qwen3.6 MoE model on Hugging Face. Alleged to shrink memory 3x with near-zero accuracy loss.
  • Japan’s banks get access to GPT-5.5 for cyberdefense, though the extent of privileged/subsidized access is unclear.
  • XCENA, startup aiming to address AI memory bottlenecks, raises $135M at $570M valuation. Mass production scheduled for late 2026.
  • A year into the agent CLI revolution, some people have turned on to “trust-by-default” (dangerously-skip-permissions) in production – here, on YCombinator’s main prod db.
  • Groq is raising $650M from existing investors to fund its AI inference cloud business. The move follows a $20B licensing and talent transfer deal with Nvidia in December which paid out Groq’s backers in cash.
  • Robot warfare continues to ramp up. SF robotics company attempting to get in on the action with humanoid robots, which are notably unlike the robots used in the Ukraine war..
  • Prime source of training data for humanoid robots: teleoperators.
  • Similarly, OpenAI robotics is hiring.
  • OpenAI announce Rosalind Biodefense, an initiative to strengthen societal resilience and support pandemic preparedness. The goal is to build early warning/diagnostic tools to mitigate the dual use nature of AI in biotech.
  • SoftBank to invest $75B to build data centers in France, targeting 5GW capacity. First phase would target 3.1GW by 2031 with a $45B investment..
  • AWS paper on automating agent evaluations finds a 70% failure rate for prompts alone and investigating the “skills” (long markdown prompts) that halve this. Yeesh: only tested Sonnet and Haiku 4.5. Ignore.
  • Earlier paper on EvalAgent, which automates the process of evaluating AI agents. Fixes typical model failures like over-engineered code via procedural instructions and templates. Increases AgentEvalBench Eval@1 success from 17.5% to 65%.
  • Incendiary take on the permanent underclass already being here via gated+tiered model access.
  • OpenAI models/Codex now available on Amazon Web Services. Likely a big deal for a small fraction of enterprises who would otherwise struggle with adoption friction.