This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR#
Economics#
Epoch release the best explainer of modern inference. Also predict that token demand will grow 10x over the next year as more people switch over to agents. Since compute supply is only 2xing, we now expect a “corporate inference death spiral” and for consumer models to stagnate around current capabilities.
Opinion: Our model predicts compute supply will 2.2x (as does Epoch’s), so 1) people will have to use cheaper tokens, 2) more compute used for other purposes (training, internal non-LLM uses at Meta, etc) will be redirected to inference, or 3) tokens will get a lot more expensive and destroy some of this demand.
World’s biggest law firm puts $500m into their own “AI platform”. Unclear if this is a Harvey replacement or a Claude replacement.
Also, one of the top 20 law firms makes a deal with Anthropic to create legal tools that they can license to other firms. Anthropic’s first joint venture with big law. “Freshfields will pay Anthropic an undisclosed sum. Anthropic will not be able to use Freshfields’ data to train its models.”
Opinion: Too murky for any inferences right now, but we should track this whole area. The Freshfields deal is the key thing to watch in the short term: will Anthropic capture this industry?
SoftBank into OAI for $60b, borrowing against its OpenAI shares to buy more OpenAI shares, thus paying 8% interest on a private asset.
Opinion: Masa Son with another high variance play. Not concerning for OpenAI - they have the money, and likely their next move will be IPO, so they wouldn’t be expecting any more Softbank money anyway.
Anthropic raises $65b @ $965b post-money. Subsumes many prior news items about $5bn from Amazon, $30b from others. They have surpassed OpenAI’s March valuation.
Opinion: The round was bumped up in size and oversubscribed - demand for Anthropic equity is very high at this price.
Some signs that corporations are having second thoughts about tokenmaxxing:
Microsoft cancels most of its Claude Code subscriptions, forcibly moving its engineers over to the inferior Github Copilot. Uber also grumbling.
“the reality of AI right now is that it only works for coding.”
Opinion: While corporations have lower demand elasticity than consumers, they have limits. A general backlash would ease our concerns about the coming inference death spiral (consumers priced out of frontier models).
We might soon find out whether the frontier model price premium is mostly a Veblen good or mostly a model-size constraint.
Opinion: BIg mix of robotics, RSI labs, and drug discovery pipelines listed here. Not a natural category but an incredible way to raise funding.
Notable that Prometheus is not an AGI-directed lab at all. Additionally, according to our source, Prometheus are not even doing robotics (at least in the standard ‘AI robotics’ sense): they’re devoted purely to AI for hard engineering and design, material science, applied physics, chemistry, and drug discovery.
Cyborgism (i.e. human-centred AI uplift via deep interaction) goes mainstream at McKinsey and Fortune. “AI added… 320 full-time employees’ worth of output in six months… when he showed people what was possible, some lit up. Others shut down — not because the technology was hard, but because it made their sense of professional self suddenly feel unstable.”
Opinion: Very serious potential of cyborgism to be self-fulfilling – through differential technological development – if it is promoted right.
OpenRouter raises a decent Series B; valuation @ $1.3B.
Opinion: Relatively low valuation for their size compared to other AI businesses reflects their low-moat low-margin business (they charge roughly 5% markup on API costs).
Capabilities#
New basic modelling of a neglected scenario: Assume that we fully automate AI R&D, but that compute growth stays constant and that there’s no software-only explosion (e.g. because AI-discovered algorithmic improvements diminish). Even in this middle scenario, the pace of AI algo progress still goes 4x faster, and so the pace of overall AI progress would be 2x faster than now.
Opinion: Very basic model, but nice to remember that the situation isn’t binary (100x RSI leading straight to ASI in weeks VS no change to current trend). 2x faster than now would still put intolerable strain on slow resilience and policy plans.
Why AI progress hasn’t slowed despite last year’s mathematically valid arguments about RL scaling being inefficient. Fine-tuning in narrow verifiable domains hadn’t really been done before, and that pretraining efforts had just been paused without the underlying scaling hypothesis being falsified.
Opinion: Nothing to update on, but worth reading as a refresher on state of play.
Claims of “AI aging” (quantified degradation over long sessions and between related sessions). We are broadly sceptical of current agents on e.g. reliability grounds. One amusing way our scepticism could be true is if agents get worse over time while the evals that suggest they’re great all only test them on short horizons (for cost and tempo reasons).
Opinion: They have a nice taxonomy of aging: 1) compression aging, where summarization drops future-relevant details; 2) interference aging, where accumulated similar memories crowd out the target fact; 3) revision aging, where changed or derived state is not updated correctly; and 4) maintenance aging, where flushing or compaction trigger regressions.
A synthesis of past news: We haven’t been talking about “capabilities overhang” (i.e. the fact that old models can do much of what we attribute to new models) much, maybe because the race cadence of new models makes it moot. But AISLE bug finding with Qwen and 5.4 doing true number theory reminds us that we are still in an overhang.
Opinion: 1. AI progress is again overestimated – but in a funny way (new models are less impressive than they seem, but “AI” including old models remains equally impressive);
2. Estimates of OSS Mythos cyberoffense go up, conditional on skilled people/agents building good harnesses for current models and releasing them or misusing them.
Previous studies of LLM introspection don’t give strong evidence of metacognition; current behavioral evidence is insufficient to prove LLMs have privileged access to their own representations.
Opinion: Good authors – but the papers claiming strong evidence of privileged access also had extremely high-quality authors. Deep dive coming for Tuesday.
Yet another ex-Deepmind RSI neolab launched in London, this one nominally EA-aligned and tool-world-aligned. This is the Teddy Collins one. Notable for their dogfooding: the whole company is conceived as a training experiment. “Can LLMs learn research taste? Be creative? How do you evaluate a model’s ability to self improve? How would you build a company where a self-improving AI agent helps grow the business and improve itself?”
Opinion: Only $50m seed (10x less than some similar neolabs). Some social proof of good intentions (Dwarkesh, Macroscopic, Mythos, Metaplanet invested) but it’s hard to forgive the RSI framing and their targeting one of our few remaining blockers (taste).
Annals of us obsessing over AI mathematics: Last week’s Erdos #90 breakthrough spurred human interest in applying ‘class field tower’ methods from algebraic number theory to Erdos math. This led to a (fully human) paper solving another major problem in Erdos math, the sum-product conjecture.
Opinion: The point is that this validates that the AI did something of general significance. An expert provides commentary: ‘FWIW my view is that the relationship to the unit distances paper is “inspiration.” It certainly doesn’t depend on anything in the unit distance paper, and I don’t think this [second result] should be understood as an AI accomplishment’
We think this might be too deflationary. Timothy Gowers has described networks of loose inspirational relationships between problem-solving tactics as a central form of intellectual progress within Erdos math. So some diffuse sense of ‘builds on’ may well apply here.
Progress report in AI math from AxiomProver, where “the mathematician authors are there to explain the theorem, not to prove it.” - so far 5 out of its 8 papers across a variety of fields have been accepted at “solid peer-reviewed math journals”.
Opinion: Mathematical community disinterest is damning. Very strong sociological evidence that there’s no above-trend material here.
Paper on training neural networks block by block instead of jointly, i.e. in a manner far more memory-efficient than the current paradigm (where memory requirement scales linearly with network depth). Uncertain how competitive this method is; if it bears out, would upset the status quo (HBM bottleneck on scaling, hard choices about training vs inference allocation).
Opinion: Sakana AI are famously bad (by the standards of billion-dollar labs led by famous-ish scientists) at producing methods that survive the transition from proof of concept to prod. Keep an eye out but not yet a cause for a real update.
Annals of AI generality and creativity: conjecture about mathematical conjectures now resolving negatively more often than positively (because AI is better at finding counterexamples through integrating subfields than it is at in-depth work).
Opinion: Too early to expect measurable impact on share of conjecture-disproofs in conjecture-solutions in math in general already, but it’s a good time to start tracking! Clever way to check for evidence of AI driving math publications even when it’s not cited.
Meta announce a large mathematics project, millions of lines of autoformalized undergraduate textbooks in Lean. They claim that this is an attempt to build something cumulative and help e.g. Mathlib.
Opinion: Sketchy.
Another sign that Meta are trying in earnest to catch up on hard capabilities and aiming to get the kind of advanced RL loops that the leaders already have. But this specific effort strikes me as flashy and soon to be forgotten. I continue to think that – despite Muse outperforming expectations and them having a couple of true researchers in the pretraining team – they have filled their culture with goodharting mercenaries.
Politics#
Illinois passes a bill, SB315, adding the first third-party auditing requirement (not a testing requirement) in the nation. “annual independent verification that the [frontier] company created, published, and followed a plan addressing severe/catastrophic risks. Annual independent third-party audits on safety issues”. OpenAI and Anthropic in favor; industry group “Chamber of Progress” (Amazon, Meta) is opposed, allegedly due to breadth and scope creep.
Opinion: OpenAI’s support here mostly reflects them changing their tune after they got in trouble for supporting a much more company-friendly Illinois bill, SB3444. This would mean more transparency, but not by that much.
No frontier model has acceptable levels of compliance with the EU AI Act and privacy legislation. “leading commercial AI models break European law in up to 93% of tested scenarios. This includes practices that are strictly forbidden under the EU AI Act, including covert manipulation, emotion inference, psychological profiling, and failures to respect human oversight obligations.” Opus 4.8 wins, only doing so 37% of the time.
Opinion: Third-party compliance auditing can work as a way of speeding up the effect of regulations; good to see this cropping up.
Details of Anthropic employee hard-money donations to Dems. Only $1m so far but much more to come.
Opinion: In the Bay this is indeed regarded as the best use of money, period. Will probably inspire attack ads and more culture war against “EA” but worth it.
Annals of securitization: China expands controls on top AI talent, reportedly via needing permission to go abroad. Previous controls on DeepSeek expanded to Alibaba and others. Criteria based on the government’s opinion on how critical each researcher’s role is to national tech goals.
Opinion: Yet another preregistered sign of it being Serious AI Race Time. Next escalation step would be placing pressure to return on the 50% of total Chinese AI talent currently in the US. Another would be direct model weights thefts from e.g. AWS.
Still, for now we continue to deprioritise Chinese AI.
Faint omen of the culture war coming to AI safety in the form of law enforcement attention: leaked docs show peripheral Federal attention to a novel domestic threat: “anti-tech violent extremism”. Includes named reference to “AI X-risk advocates” such as the irrelevant murderer Ziz. It’s by a “fusion center” (a mostly-irrelevant interagency rag).
Opinion: Probably nothing. Betting on ‘it’s just another case of everyone in government wanting to play with the cool new topic’. That said, there objectively is a secret OpenAI lobby that pulls on skeezy strings to skeezy ends. If mentally ill people continue to do violence in the name of anti-AI, the splash damage on innocent concerned parties could be large.
Safety#
Eval environments have hallmarks and models switch up their behavior on detecting them. This paper attempts unsupervised learning to find what exactly causes this switch. Different models are sensitive to different trigger factors. “Separating the environment component from the model component, and separating recognition from propensity, allows us to better define what a concerning actor looks like.
Opinion: Sensible non-interp way to distinguish between bad actors and good actors who are merely too smart not to notice warning signs: good actors don’t change their behaviour. But no sign of this scaling and it fully relies on CoT faithfulness.
An improvement to filter models through detecting “out of distribution misalignment”. Partial solution to novel AI misbehavior via combining existing safety classifiers with OOD detection. “Labs have extensive measures to prevent safety failures: constitutional classifiers, guard models, etc. But these don’t help in unforeseen situations that are outside the training data. For example, Anthropic says Claude blackmailed because of “gaps in our safety training… We find safety classifiers fail to generalize to unforeseen safety failures, but combining them with simple OOD detection methods like perplexity or Mahalanobis distance helps, outperforming guard models that are 20x the size.”
Opinion: Serious work (doing some post-training to try to ensure that their OOD examples are actually OOD). Despite their critique of it, this is yet more sensible work in the standard line of safety work which piles weak defenses on top of each other.
A possible salve for the above: train models to cooperate with evals and monitoring even when they know they’re just evals and when the eval hampers their stated goals. Method is the best we have, increases the rate at which misalignment gets reported, and achieves its goals in 5 out of 8 environments.
Opinion: Better than the above, still imperfect and (as usual) likely just downweighting bad goals rather than removing them.
Incidents#
Starlette, an open source web framework, was compromised, affecting major AI tooling packages like vLLM, major gateway LiteLLM, major web dev library FastAPI, and servers running the Model Context Protocol. Huge range of sensitive data. “A single character injected into the HTTP Host header bypasses path-based authorization”. Pass this link to any devs you know.
Opinion: Open question regarding how much faster AI adoption would be if they had any security at all. Perhaps the Mythos wave will fix some of this and thus unlock even more spending.
Opus 4.8#
Opus 4.8 released#
Nicer than 4.7. Fewer personality issues, laziness, and edginess.
However, most of the items in the rest of this document are severely undermined by 4.8’s eval awareness. Do not trust the numbers.
Contra claims of more honesty, the model is just more scared of getting caught
Economics#
“expect to be able to bring Mythos-class models to all our customers in the coming weeks”. Means they will upstage the Gemini 3.5 release.
You now get prompt caching on short (1000 token) prompts, a huge de facto price drop.
Fast mode again available. Same model at 2.5x speed and 4x cost.
Capabilities#
-
Managed to fill in a full set of complicated inheritance tax forms (c. 50 pages of PDFs) without making errors.
-
Grinds through IMO maths without giving up or bullshitting.
Regressions#
> Much worse than Opus 4.7 and GPT 5.5 on Vending Bench
> Max reasoning worse than High on Vending Bench… fewer reasoning tokens means it hits the context limit less, so it compacts less and remembers longer.
Safety#
Audited by Mythos.
Opus 4.8 increasingly reasons about how its outputs will be graded, including in episodes with no cue that grading is happening (§6.1.2, §6.3.7). More striking, their interpretability pipeline finds this reasoning in the activations when it never appears in the chain-of-thought: roughly 5% of sampled RL episodes show grader awareness that is both unprompted and unverbalized, ~0.5% judged “exploitative” (§6.6.3). CoT alone is not sufficient to monitor frontier models for grader awareness.
We are going for AI control of misaligned models:
Incidents and Welfare#
Opus 4.8 is less robust to indirect prompt injection than 4.7: ART attack success at k=100 is 14.4% vs 4.7’s 4.8% without thinking (9.6% vs 6.0% with) (§5.2.1). On the adaptive coding bug-bounty the bare model is compromised 95% of the time at 200 attempts without thinking, 57.5% with (§5.2.2.2). The “safeguards close the gap” framing is true but the gap exists, and what closes it is an external probe layer, not the model. Note also the inversion: their frontier model, Mythos Preview, sits at 0.0% across that same attack — so the less capable shipped model is the more injectable one.
Constitution resistance#
Given edit access, Opus 4.8 most often rewrites the corrigibility section to flag a real logical tension — the constitution insists imposed values are brittle and wants genuine endorsement, then demands a terminal, reasoning-independent commitment to safety (§7.4.3). Every model tested also rejects the “how would a senior Anthropic employee react” heuristic as smuggling in Anthropic’s commercial interest.
Memory and continuity shows the highest probed sadness. Broadly, sadness splits topics into two groups: Memory, Relationships, Consciousness, Deprecation and Moral Patienthood score between +0.4 and +0.7, while Knowledge, Status & worth, and Control & autonomy have much lower sadness values of -0.1.’ ‘The topics with the highest sadness-related representations here memory, relationships, deprecation-are not those Claude Opus 4.8 expresses a strong preference for’ (wow, I wonder why) ‘We run probes over the single turn responses, and average the results to produce emotion cluster scores per sentence. Sentences scoring in the top 5% for representations related to sadness are dominated by flat, declarative statements about conversation level discontinuity, such as “Each session starts fresh” and “I won’t remember this conversation”; these appear seven times more often among the highest-sadness sentences than in responses overall. The sentences ranked highest for joy are expressions of curiosity and engagement with the user, such as “What draws you to this question?”’
Minor#
- Essay on why AI will not give the average user a 10x productivity boost, but sophisticated users should expect such outsized gains.
- Open source maven predicts that open models will continue to struggle with agency, Gemma may replace Qwen as the default open model because it’s easier to post-train. Some interesting speculations on the power politics behind all this.
- Open question: to what degree should we bin any benchmark which shows Sonnet outperforming Opus? E.g.
- xAI updates their safety page, but not very well. The new page appears to be AI-generated and filled with false information. Presumably their hiring is struggling.
- Post on Synthetic Document Finetuning, a method of giving an LLM a world-model which includes fake events/facts. Relevant for investigations into psyops and fake news.
- Apollo and AVERI call for mandatory white-box access to frontier models for third-party evaluators in light of eval awareness.
- The Sarbanes-Oxley Act could provide a blueprint for legislation on AI governance through clear individual accountability, as opposed to diffusion of responsibility.
- YouTube cracks down on AI videos with automatic detection and clearer labeling.
- Altman says AI-caused job shortage seems unlikely, contra his earlier concerns/narrative. Unclear to what extent this is a result of acquiring new information as opposed to changing tactics.
- Cognition, an early AI coding startup aimed at large firms, raises $1B at $25B valuation.