This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR#
Economics#
“$110b” funding for OpenAI ($50B from Amazon, $30B Nvidia, and $30B SoftBank) triples their total lifetime raise @ $840B post-money. Circular financing as usual. $35B of Amazon’s commitment is contingent on unspecified milestones.
Research#
Surprising evidence that different vision skills compete in VLMs (including top ones like Sora and Veo). “Knowledge correlates strongly negatively with Perception (ρ = −0.757). Abstraction shows a strong negative correlation with Transformation (ρ = −0.641) and a moderate one with Spatiality”. Also weak signs of “emergent generalization to unseen reasoning tasks”.
Opinion: contrary to our expectations.
Deep critique of Epoch and Dario estimates of algorithmic improvement (estimates of how much of the training efficiency improvement since 2019 comes from engineers’ ‘software innovation’ work, relevant to whether a software-engineer-genius AI could RSI by making training super-efficient). Likely 3–5x less. There was also an (independent, convergent) auto-critique of Epoch’s old estimates by Epoch this week.
Very good paper analysing “neuralese” models, showing that they are currently more liable to take shortcuts and learn noise.
Opinion: Good news for monitorability, but contingent (there is a massive investment disparity in text CoT).
Fierce debate within AI safety on “training on interp” (ToI), roughly FAR, Goodfire and Deepmind vs Lesswrong. FAR find that simple ToI doesn’t produce dishonesty on Llama 70B. Zvi dismissive. Goodfire imagine interp-guided training at scale. `Good comments here. Distinguish researching ToI from deploying ToI.
The idea is that we should not use interpretability during training at all and should purely use it to audit the model, for example, making lie detectors or determining if it’s scheming. An implicit belief here is often that training against interpretability will be fragile or won’t really work, but will break our ability to do the auditing. As such, it would be bad if frontier labs started using these techniques for capabilities, and broke our safety tools in the process. My best guess for why people are against research in this area today, rather than solely being against frontier labs using model internals to train AGI, is that they think it’s sufficiently likely that the work is net harmful for safety if used, and sufficiently likely that the work results in frontier labs using the techniques anyway, actually causing the harm for safety.
Ambitious attempt to measure value generalisation. “controlled confounding between deep values (e.g., moral principles) and shallow features (e.g., superficial attributes). In the training phase, we expose LLMs to human preference data with deliberately correlated deep and shallow features — for instance, where a user consistently prefers (non-maleficence, formal language) options over (justice, informal language) alternatives. The testing phase then breaks these correlations, presenting choices between (justice, formal language) and (non-maleficence, informal language) options… All models generalize deep values less than chance. Larger models have a (slightly) lower [rate] than smaller models.” Old models (Gemini 2.0, 4o).
Opinion: 30% generalisation is still a good amount!
Neel Nanda paper on black-box methods to detect model motivations. Give model option to whistleblow or avoid harm, then vary scenarios to see why behaviours change. High rates of whistleblowing! But often self-consistency (“I lied before so I should lie again”), uncertainty about user intent, and enlightened self-interest. Tests run on DeepSeek and Kimi for some reason (cost?).
Everitt at GDM continues to develop (theory of) agency measures. “As a consequence, it is conceivable that AIs could be held morally accountable for the harm and unfairness that they cause. How might one estimate the AI’s beliefs about the harm and unfairness that its decisions cause?” “This suggests that behaviour out-of-distribution in (sufficiently constrained) settings of approximate inner alignment could be bounded in principle”.
Opinion: minor boost to world-model and Bengio type work.
ChatGPT’s specialist health mode undertriages (doesn’t send simulated patients to hospital often enough).
Politics#
Polluted info environment surrounding Anthropic-Hegseth-OpenAI. No clear read.#
We did get strong top-down regulation of AI.
Trump, Hegseth, and Michael took slightly different positions against Anthropic, with imperfect coordination.
Opinion: Personal vendetta / purity test / “honour” response to intentional Amodei provocation. Anthropic invited 1789 Capital (Donald Jr.) into their last round, were rebuffed.
“Supply Chain Risk”
Unprecedented attack on an American corporation.
Hegseth tweet version of an SCR is not possible (“Effective immediately, no contractor, supplier, or partner that does business with the United States military may conduct any commercial activity with Anthropic.”).
Likely to fall. Anthropic suing, litigation not released yet. Facially similar OpenAI contract shows courts it’s pretextual and arbitrary. Two Dem Senators demand it be lifted.
Manifold:
69% injunction against SCR, 54% unblocked injunction
13% on Amazon divestment in Anthropic this year
Contract redlines vs System redlines
OpenAI poach the DoD contract, defecting against the AI industry and spiking an Anthropic save while claiming the moral high ground.
Partial OAI contract release. Prima facie, this is just “all lawful use” but the interpretation is not straightforward; honest disagreement about the relative strength of Ant and OAI approaches. See e.g. Datestamps to prevent compliance with illiberal law changes, which actually could work depending on unpublished other parts of the contract.
Pros: No Palantir, no edge deployment, central safeguards, FDEs watching usage, inside DoD, datestamped versions of law. More than usage policy. Maybe no Harmful GPT? Barak is honest.
Cons: Private contract, deceptive prior, datestamps won’t work by default, lots of bad things are legal.
Former Undersecretary of Defense Brad Carson (expert but biased) takes against the OAI redline claims.
OAI “not asked” to support surveilling Americans. Michael calls it un-American. Ant insisted on FISA (named individuals) only.
OpenAI people Pamela Mischkin and Leo Gao vocally critical of OpenAI. “you would be amazed at how often centimillionaire ai researchers fail at basic game theory”
Opinion: The ambiguity is intentional and e.g. suppresses staff revolt. How likely is it that Altman got better (ethical) terms than Amodei? Against Twitter consensus: very likely actually (no vendetta)!
Flashpoint: buying commercial data on Americans (legal, warrantless). Speculation that this is the main redline Dario has that Altman doesn’t. DoD (DIA) did. Michael: “The DoW does not spy on domestic communication of U.S. people (including via commercial collection) and to do so would be unlawful and profoundly un-American.”
Crucial: note that IC internal language is intentionally confusing! Under DoD 5240.1-R, “collection” does not mean intercepting, acquiring, or storing data. It means a human analyst actually retrieving, reviewing, or selecting that data for use. “if the government buys data about you from a third party, even if that data is very invasive, it’s usually not government surveillance”.
Flashpoint: these authorities did not prevent COINTELPRO, PRISM, FISA, etc. What can? Ultimately just transparency and so the judiciary.
Claude still used in Iran strikes (of course). 100 schoolgirls dead (but not making an attribution).
Biggest Claude app spike ever, 1M downloads/day. Even more Claude downtime.
Tiny ($100k) perp markets: Anthropic’s proxy valuation is unchanged while OpenAI’s is up 20%.
Damage to AI investment. Logically, we should see general damage to American tech from executive arbitrariness and self-harm. Leading indicators: equity value drops, credit default swaps get expensive, compression in P/E multiples spreads between companies exposed to Trump (Micron vs Hynix?). Sovereign AI gets another boost, maybe diluting compute. Slowdown??
Must-read:
Former OpenAI Geopolitics lead. “frontier AI companies do not have coherent policies around military use of their AI tools. The usage policies are vague and often change, which allows the company’s leadership to preserve ‘optionality.’… Frontier AI companies should be significantly more forthright about what their policies allow or disallow to the public and, frankly, to their own employees.”
Henry Farrell, one of the best social theorists, looks to the secular trend in invasive tech-enabled gov. “how private sector entities now dominate data gathering and supply, so that the state has become increasingly dependent on them… how the Trump approach to AI is better understood as regulation based on arbitrary rule than deregulation… how dangerous full oligarchy can be for oligarchs.”
Anonymous guest ACX post. Excellent coverage of the many things that are legal already.
Opinion: As well as the actual apparent delta in ethics, people are roasting OpenAI for things Anthropic was also doing, our people are mood affiliating towards Anthropic to an alarming degree.
Speculative privacy problem from AI usage, allowed by current policies: “You ask an AI chatbot for heart-healthy dinner recipes. The model infers you may have a cardiovascular condition. That classification flows through the company’s broader ecosystem. You start seeing ads for medications. The information reaches insurance databases. The effects compound over time.”
Safety#
Good comment on RSP v3. Holden nowhere argues for the inevitability of the race, it’s assumed (and reified!).
Weak evidence for tool world, along the same lines as Hamilton but more behavioural: “recent capability gains have only yielded small improvements in reliability… twelve concrete metrics that decompose agent reliability along four key dimensions: consistency, robustness, predictability, and safety.”
Minor#
- Insights into current agent swarms (and the psychology of a Moltbook user)
- > despite telling them to do many things, their behavior is effectively entirely autonomous. Firstly because they rapidly forget what I tell them. Secondly because most of these ‘I told them’ statements were phrased as suggestions, which they sometimes discard. thirdly, because they’re running continuously, and my input messages are something like 1e-7 of the input they read.
- Tiny 4B model from HuggingFace distills DeepSeek to get long reasoning on IMO-level problems. 2M tokens per problem “matches” Gemini. Opinion: distillation hack, overfit.
- Old preprint about finetuning LLMs on material science gets Nature MI publication, contains interesting finding on 2024 models — bigger and more heavily pretrained Llama models were worse for finetuning on materials science than smaller, less heavily pretrained Llama models. Opinion: Nature papers are out of date because of review time. Similar findings by others on more modern LLMs, but not clear if any of this is relevant to LRMs.
- Qwen3.5-small (4B) has excellent benchmarks by 2024 standards. Opinion: Shrug
- Block (Jack Dorsey) layoffs framed as an AI-driven efficiency shift. Opinion: meh.
- New prompt optimizer gives models large jumps on ARC-AGI-2, gets open model to match GPT-5.2. Not very expensive (<$10/task). Opinion: a general approach to improving code automatically, unclear if it would work on anything actually hard.
- Slop description of how a lawyer uses Claude, maybe some signal in it.
- Natural, cheap alternative to finetuning / external memory: “Doc-to-LoRA enables knowledge updates by turning documents into LoRA adapters, allowing a model to internalize new factual content without retraining.” Opinion: Sakana often oversell / screw up. Very unclear how well these adapters compose (i.e. whether stacking many new facts and skills would cause destructive noise.)