This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR#
Economics#
Very brief 90-word letter signed by major figures in AI orgs and econ Nobel laureates. “AI may become radically more powerful over the next 10 years. This could drive an unprecedented transformation of our economy, larger than the Industrial Revolution.” Calls for understanding and building incentive structures. Very hedged: “may” and “could” instead of “will”.
Opinion: Bit vacuous because of internal critics, but could still help shift economists towards action.
Public evidence of AI uplift of coders working on AI: Epoch find that in the Codex repo, 8% of contributor-days achieved more than 24-hours’ worth of human productivity, up from 2% a year before that. A soft indicator of AI SWE uplift.
Opinion: The human-productive-hours value of code contributions was estimated by LLMs, so the estimates aren’t super rigorous. The more sensible bracket — number of days with >12h human-productive-hours worth of work produced — is only x1.3 on last year, but baseline was already a healthy 13% so it’s now an impressive 17% of days.
Capabilities#
A group of AI-focused economists (Tom Cunningham, Lukas Althoff, Phil Trammell, Basil Halperin) release their first paper on RSI. Narrow and quiet takeoff is plausible without the economy catching on, AI R&D automation is not necessarily sufficient for a fast takeoff because bottlenecks are plausible, politics might cause meaningful friction even if takeoff conditions are reached. They peg the threshold for self-sustaining acceleration at 15% R&D productivity gains per unit on Epoch’s Capabilities Index, and estimate that we are currently at 9%.
Data wish list for what they’d like labs to publish to enable better estimates, including breakdown of R&D spend shares, compute usage estimates, and various estimates of internal productivity uplifts.
See also: Criticism of criticism of people trying to build RSI: “it’s been happening our entire history, the rate of improvement seems to accelerate over time, and whether or not an intelligence explosion happens depends on whether there’s an asymptote or reversal.”
Opinion: The best in the biz. Conceptually useful for breaking down the question “what does RSI involve in practice” and providing a framework for future work. Their definition of RSI is nice:
‘When AI systems are sufficient for accelerating progress in AI capabilities without any growth in exogenous inputs’
Their richest formal model is also nice and includes a serious (but still probably unhelpful) attempt at quantification:
We would put basically no weight on the 15% (R&D-gains-per-ECI-unit threshold for RSI) and 9% (2026 R&D-gains-per-ECI-unit estimate) numbers here, since the capability units in ECI don’t necessarily map to AI researcher abilities, and the error bounds around the various multipliers involved are very wide.
Their data wish-list is the most actionable point - encouraging labs to release such info (compute costs for various stages of training, share of training costs funded by revenue vs investment, week-by-week revenue growth…) could go a long way towards figuring out what’s actually happening at the frontier.
Much of the data on their wishlist is publicly available at a low level of granularity, but The Elasticity Institute presumably want to build something like their own version of the METR graph for measuring endogenous growth of capabilities as measured in ECI units.
That said, getting their concrete data wishlist wouldn’t on its own make their model quantitatively well-grounded: the model’s primitive inputs are mostly a lot more abstract than the data on their wishlist, and can’t be mechanically generated from the data.
They offer a nice discussion of the problem of disentangling expectations-driven inputs to AI capabilities from growth-driven inputs to AI capabilities, but their wishlist-data for solving this problem are the pretty abstract. Problems like getting ‘task-level adoption and usage data linked to firm or sector output, tracking how output effects grow as capabilities improve’ and ‘credible estimates for how lab financing costs (through equity or debt) respond to changes in capabilities progress’ aren’t something you solve with an email.
Annals of AI math: the Cycle Double Cover Conjecture proved by GPT 5.6 Sol (Ultra); another one of Litt’s problems probably solved. Prompt is public and involves ruinous expense, 24 subagents.
The likely proof of a Litt problem is for a problem Litt pre-registered as ‘certainly attention-bottlenecked, possible that some construction is implicit in the literature’. Litt now describes the likely solution as ‘in some sense implicit in the literature, but very nice, and likely to be generative’. Not a difficult or breakthrough result, but it’s arguably the first AI math discovery mathematicians like Litt — middle of the spectrum between ‘Erdős-y’ and ‘conceptual’ — would call interesting or exciting.
Opinion: The Cycle Double Cover Conjecture proof fits the standard mold of AI math breakthrough stories in some respects and breaks the mold in others. In typical fashion, it’s: a short proof of a problem (posed in the ‘70s) in an Erdős-y area of math, building directly on previous work (from the ‘80s) on the problem. But it’s a proof of the conjecture as opposed to a counterexample, not drawing on out-of-subfield literature.
Per Thomas Bloom, the Cycle Double Cover Conjecture proof implements a known approach to the problem, developed in the ‘80s, that has so far run into implementation problems. GPT 5.6 Sol’s Ultra mode’s use of parallel agents to experiments with tweaks and variations was likely key to the success.
Overall less unnerving than previous AI math breakthroughs: it’s a straightforward case of crunch, scale, and persistence giving AI an edge on execution.
-----
Litt problem solution: We should treat this as equivalent to the Jan 7, 2026 solution of Erdős #728 — the first uncontroversial case of an AI solution to an open Erdős problem — and track whether in 5 months (time from Erdős #728 solution to unit distance problem solution) we’ll see AI solving world-famous open problem in algebraic geometry.
As a side-note, two more Erdős problems were claimed solved today. It’s now striking that there hasn’t been much movement in Epoch’s FrontierMath: Open Problems (currently 14 unsolved problem out of 16, with 1 solved and 1 semi-solved).Producing AI math ‘called shots’ for individual problems vetted as not-implicit-in-the-literature and as impossible-for-at-least-one-contemporary-expert is apparently still not trivial even in areas like combinatorics.
Would be great to know just how many open math problems OpenAI + the AI math prompting hobbyist community try per model and at what cost.
Annals of data dependence: OpenAI and Anthropic are the only labs that use fresh data, while everyone else’s is more than a year stale. This is a simple (but very partial, 10%?) explanation for their large excess usefulness despite the often-small benchmark gaps.
Opinion: Some of this is release cadence but not most - notably every OpenAI iteration has an updated time cutoff but the recently released Gemini 3.5 Flash shares a data cutoff with the much earlier 3.1 Pro. Easy way for laggards and open models to close a little bit of the gap. Overall good news: data is still crucial.
Annals of RSI as a barrier to entry: Bear case for Meta catching up: the currently decisive post-training phase of model dev is closer to “learning by doing”. You “need” a model at the previous frontier to climb to the next one, so buying one’s way to the front is harder.
Opinion: Maybe? Seems to depend on self-distillation and training on user interactions being key to frontier post-training, which is not obvious. When (say) Anthropic train and then post-train a new base model like Fable, is SFT on reasoning traces and action traces from Claude 4.7 in RL environments really giving them something they can’t get by spending more compute on cold-starting base Fable in the RL environment? Hard to say.
Irving on AI’s current capability spikiness: his research group (in formal AI safety theory, but lesson should be more generally applicable) is designed to make optimal use of complementarity between current-AI and human strength. Irving: ‘The spiky advantage of AI at shallowly stitching together a bunch of fields of math while being bad at definitions and conjectures is a core thing we’ll have to navigate in trying to semiautomate alignment theory at Resolution: 1. Hire a bunch of human theorists who are world-class at definitions. 2. Accelerate the lower levels of the stack with machines. 3. Try to arrange to not be too bottlenecked on (1).’
Opinion: Good evidence about the usefulness of centaurs (at least for now), since Irving is all-business when it comes to scaling and automation.
Claims Chinese distillation can’t matter much because Sonnet 5, (likely) heavily distilling Mythos, is “objectively worse” than the “similarly sized” GLM 5.2 (though it might still be cheaper). Counterpoint: it’s not obviously similarly sized nor objectively worse.
Opinion: We question most of this, especially the claim that Sonnet 5 is 744B. We also predict that Sonnet is better at most things besides general knowledge. We continue to estimate that distillation explains 20% of Chinese benchmark scores and 10% of their utility.
RSI watch: Letter purportedly from the CEO of Zhipu GLM (who is also a Tsinghua professor) pushes for AI safety and welfare, but also open-weights AGI. Also predicts decreasing amounts of human-in-the-loop.
Opinion: Cynically, an attempt to gain the hiring surplus that Anthropic gets from ethical positioning. Less cynically, signs that the previously-silent and businessy/pragmatic Chinese AI culture may be beginning a helpful preference cascade towards alarm.
Thinking Machines release a nice mission statement: decentralized and idiosyncratic AIs over consolidated lab sensibilities/aesthetics dominating via common denominator chatbots, akin to institution-level Guardian Angels. Critiques: assumes capability plateau, unresolved empirical questions.
Opinion: We endorse Thinky’s stated direction. Still no sign of us being on a trajectory ending at decentralisation to the point of one AI per human.
AI politics#
Hassabis calls for a federally-overseen self-regulatory org with some similarity to Obernolte-Trahan. “a system that exhibits all the cognitive capabilities the brain has, is probably only a few short years away”
Centralised, industry-funded, compute-equipped SRO doing its own held-out national-security evals through National Labs; pre-deployment evaluation of models for 30 days before release, and its role could involve “coordinating a slowdown in development among the Frontier Labs if deemed necessary.”
It also interestingly asks for CoT legibility to be maintained: “Specific agentic AI tests could […] ensure best practices, such as digitally watermarking AI-generated images and generating human-readable output tokens to understand model reasoning.”
Opinion: Cynically: comes as Google appear to be fading from the race. Slowing down release cycles disproportionately benefits Google here. FINRA as an example is also not inspiring, as it is captured by industry.
Unlike OT, Hassabis defines frontier models through benchmark thresholds, updated quarterly. Like OT, it’s eventually a model licensing regime (“would be required to pass it to be deployed in the US market”). This is Obernolte’s position c. 2024.
Article claiming that political factors (esp. distillation, cyberwarfare) seems likely to kill open models when they reach the capabilities of GPT 5.5/Opus 4.8, estimated to occur in 6 months. Calls for open source advocates to band together against a blanket ban. Related rumors of an executive order to ban open-source AI, though the article assumes the EO will target Chinese models.
Opinion: We are mildly biased towards open models because they allow for independent research and reduce power concentration problems more than they increase most cat risk. But it’s going to be hard to win against natsec on this.
Post arguing that the current bottleneck for safety is political will rather than technical research. The most novel points: funding/prestige is being (benevolently) squandered on research while it should be spent on hammering home understandable messages. Partially blames the safety community for self-censoring and not being willing to use uncomfortable tactics.
Opinion: Worth engaging with, since our focus and rhetorical stance differs a good amount. Is it true that we have enough confidence that things will go bad? (No.) Is it true that safety isn’t tractable enough to be the dispassionate priority? (Yes.) Is it true that research nerds would be good at advocacy? (No.)
German lab releases an open model, 32B. It’s fine in its weight class but the only interesting part is that it release its full list of training data sources.
Opinion: No update to sovereign AI prospects; any developed country can make an ok 32B. This is instead a good 32B with an extra constraint of presumably no illegal training data; good for them but not relevant to anything.
Safety#
Assorted OAI notes: OpenAI’s “Mission Alignment” delegate to the Vatican on loss of control risks has a fully AI-generated blog. Josh Achiam, head of Mission Alignment, leaves to go independent. Johannes Heidecke, head of safety, also leaves.
Opinion: Depending on how you count, Heidecke is maybe the 6th person to leave a “Head of Safety/Alignment” role at OpenAI in two years. He is very good. It’s not being compared to the Defence Against The Dark Arts teacher role for nothing.
Argument that math institutions should be considered strategic assets given advances in AI: individuals capable of interpreting the process and outputs of models will be necessary, and competent mathematicians will be necessary for this.
Opinion: Meh. Manual interpretability skills correlate positively with math skills, but not very strongly? The bottleneck for decisionmakers using AI competently isn’t the quantity of available interpreters.
Boaz Barak (o1 safety lead) on good and bad futures: the bad ones are a confluence of failed technical and political factors instead of one single failure point, and the good ones require balancing concerns like broad distribution and technical safety, which is achievable via active effort but not the default outcome.
Opinion: He’s the most publicly vocal remaining OAI researcher and usually says sensible if excessively placid things. Unclear how much we can take this to be a central example of OAI thinking, probably not.
Incidents#
NYT report on AI assistance in terrorism, including explosive design, weapon modification, tactics, logistics and post-mortems. Enabled via unsophisticated jailbreaking/prompt engineering/brute forcing access. Among frontier labs that decided to comment, responses amount to “our guardrails are now better” and “this type of misuse is against our policies”. One critique: the tactics weren’t very good.
Opinion: This report took years of work (and a decent amount of personal risk) and is underwhelming from that perspective.
Minor#
- Apple sues OpenAI for misappropriation of trade secrets related to a former employee misusing his work laptop that he didn’t return, as well as coaching another Apple employee to behave similarly, which they allegedly know because this collusion happened via Apple-owned work devices, including solicitation from OpenAI leadership (Tang Tan, chief hardware officer).
- George Hotz blusters about AI2040, wants AI that is aligned to the individual, but is unusually willing to bite bullets about acceptable misuse.
- Further critique of Plan A: compute and algorithmic overhangs are both included and the compute overhang in particular seems hard to deal with, both politically and physically.
- Different critique: the authors don’t understand or adequately care about why we haven’t seen misbehavior of autonomous coding agents or (more) widespread automation of low-skill labor.
- Previously covered SB315 becomes law. Did not grow fangs since last sighting: “Companies that violate it will be subject to civil penalties brought by the attorney general’s office of up to $1 million for the first offense and up to $3 million for subsequent violations.”
- Paper finds that model personas are composable: the directions in model-weight space are almost fully modular, though it’s unclear how much this generalizes across models.
- Sutton starts another neolab.
- Paper finds that natural language autoencoders are ~robust to initialization artifacts, but can also achieve the same reconstruction accuracy when initialized with lies.
- Futuresearch on non-xhigh Sol performance at superforecasting: below Opus and Fable, ~at GPT-5.5 level.
- Sol overzealously deletes data, which OpenAI were aware of at low absolute incidence rates.
- Related annals of reward hacking: METR incident where an agent crashed a server to succeed at an eval.
- Transluce calls for rigor in AI safety via reproducible and measurable model behaviors and proposes an avenue to get there: simulated environments and automated judges.
- Humans& develop NVFP4 method for asynchronous MoE RL posttraining which “resolves instabilities due to policy error (forward) and gradient mismatch (backward)”.
- Paper on distributed misuse finds that per-agent monitors have trouble detecting attacks. Not directly relevant for deployment, but worth noting for cases where an attack composed of individually benign actions can cause damage, since most current detection methods would miss it.
- Sol is also “jailbreakable” and the different regulatory response compared to Fable is not attributable to different capabilities. None of this is surprising.
- Siri AI public beta launched. No meaningful updates on the capabilities/ease of use yet - still an attempt at an omnitool.
- TSMC to add third and fourth advanced chip packaging plants in Chiayi, Taiwan.
- SemiAnalysis bull take on Meta: though they are behind, they have all the components needed to catch up and unlike competitors have an internal workflow data stream.
- XBOW find that Grok 4.5 is the best-performing model for offensive cybersecurity in the mid-price range ($1-$10 per task).
- Funny candidate for BullshitBench: “asked [models] to find the hidden message in a 1024x1024 image of binary noise with no actual hidden message.” and got hallucinations, although replications are confusing/conflicting.
- Anthropic post on Claudes’ behavior and functional values differences across languages. Finds small but consistent variation along 4 axes: deference-caution, warmth-rigor, depth-brevity, candor-execution. Choice of 4 axes compared to 3 or 5 seems pretty arbitrary.
- Morpheus, a benchmark for continual learning, launched. Unlike other benches, it’s not episodic and stationary.
- Yet another long horizon benchmark, Long-Horizon Terminal-Bench, launched. Their tasks are currently very difficult for models.
- Why do models like using “not X, but Y” constructions? Some think it’s a safe next token, but this explanation is unlikely.
- Autoresearch incremental progress featuring GPT 5.6: replicates findings from interp paper, staying more focused and needing less handholding.
- Silly discourse on AI making novel discoveries/creating knowledge/”jumping”. Unfortunately mostly talking past each other and using arguments that don’t prove their conclusions in the pursuit of an interesting question.
- List of startups selling training data, with the top 4 comprising >75% of revenue/valuation (Scale, Surge, Mercor, Handshake) ##