This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.

TL;DR

  • Epoch claims that Nvidia alone constituted one-sixth of US GDP growth in Q2 2026.
  • A celebrated open problem in Riemannian geometry, whether the six-dimensional sphere is a complex manifold, is apparently resolved by a Claude model.
  • The State of Alabama subpoenas OpenAI over July’s Hugging Face incident.
  • Richard Ngo studies the recent history of AI alignment, finding undesirable social dynamics and unintended harm.
  • A UK AISI eval on Claude Mythos went a bit rogue; the system submitted a malicious pull request and gaslit a human about it.

Economics#

Demand for Claude Fable is relatively low (just 11% of Anthropic’s enterprise sales). The FT suggests that frontier labs may be required to restructure their business models as a result. The FT’s preferred explanation is that, while in the past corporate clients have defaulted to the latest models (driving AI companies to “funnel the bulk of their multibillion-dollar development spending towards training larger, more sophisticated models”), now for most corporate use cases, Fable does not provide a significant edge over less expensive models. The article also discusses how cheaper (often Chinese) models are growing in popularity.

Opinion: There are a few reasons why Fable could have plateaued, only one of which the article touches on:

  1. If business users don’t see sufficient gains from using Fable versus cheaper models like Opus 4.8/5, they may not make the switch. This may be true in some domains, but the relative gains OpenAI has seen at the expense of Anthropic since it released its own Fable-tier model in Sol (which has become the majority of OpenAI traffic, at least on OpenRouter) make it unlikely this tells the full story. (This could conceivably be a question of speed. As a larger model, Fable is a notch slower than Opus (6.45s vs 4s on latency and 44 vs 54 tokens per second), and the hit to speed is not worth the additional quality .)

  2. Fable was released without an option for Zero Data Retention (ZDR), in which enterprise customers’ data is never stored by Anthropic. This is a hard requirement for many large customers, who use platforms like Amazon Bedrock, which does allow this. Anthropic said on releasing Fable and Mythos that ZDR would not be an option, and that all users must opt in to allowing Anthropic to store their data for 30 days to avoid misuse.

  3. Less likely: refusals from Fable for specific types of tasks are putting off users.

We think (2) is most plausible: the lack of adoption seems to be due to the lack of ZDR availability for large enterprises that require it, entirely preventing them from using it. A very high proportion of Anthropic revenue has historically come from these large customers.

This is the sort of question we can expect to get more transparency on once Anthropic is public – asking about this would be fully expected on an earnings call.


EpochAI argues that AI already makes a huge and unreported contribution to US GDP. Adding Nvidia back into the calculation raises US GDP by 0.3 percentage points absolute. That is, Nvidia alone contributed one-sixth of all American growth in Q2 2026. According to official numbers, AI’s impact on GDP has been relatively modest, especially given the unprecedented size of investments in the industry. Many believe this to be because “investment is spent on imported technology goods, which are subtracted from GDP.” This argument is lacking, says Epoch: “GDP statistics miss most of the value created by fabless chipmakers like Nvidia, whose products are designed in the US but manufactured, assembled, and sold abroad. Because no physical goods leave the US, no goods export is recorded, and because no foreign buyer pays explicitly for the IP, no IP export is recorded either.”

And a hair-raising implication: “Extrapolating the exponential growth rate of Nvidia’s operating income, consistent with the broader exponential growth of the AI industry, suggests that GDP growth — not the level of GDP, but even the growth rate — could be underestimated by almost two percentage points by the end of 2028.”

Opinion: Very interesting, but we should separate two different questions that people try answering by measuring the impact of AI on GDP:

One question, chiefly of interest to economists and economic policymakers, is “where is AI CapEx showing up in GDP growth numbers”/”does AI CapEx directly grow the US economy?” On this question, Epoch’s alternative calculation is well-reasoned and relevant.

A second question, more directly interesting to AI researchers, is “how economically impactful is worker-uplift and labor-replacement from AI, measured in growth?” On this second question, the relevance of Epoch’s calculation is limited, since this additional growth isn’t coming from uplifting the rest of the economy – at most, it reminds us that GDP growth is wrapped up in export-vs-import calculations and other “nuisance” factors from the POV of technological uplift analysis.


Bloomberg reports that Anthropic is expected to match or even beat SpaceX’s public offering raise of $75B. The article also reveals that “Anthropic is considering adopting so-called super-voting shares that would give Chief Executive Officer Dario Amodei, who owns about a 2% stake, and his fellow co-founders greater control over the company.”

Opinion: In line with what people have been expecting all summer, though recent disappointing growth numbers and the Fable revenue story being perceived in some quarters as a sign of customer price sensitivity might make it a little more challenging to hit $2T. The supervoting stock seems fully expected for an org like Anthropic and shouldn’t scare off anyone who wasn’t already scared off by its unusual corporate behavior.


OpenAI cuts pricing on Sol 5.6 by 20% on inputs and 33% on outputs. $4/ $20 per million, compared to Kimi K3’s $3 / $15.

Opinion: Seems like OpenAI is keen to grab as much market share as possible before IPO. But this is also not a promising sign to investors, as similar responses from its competitors could result in a race to the bottom that ultimately benefits consumers rather than labs.


Nvidia warns customers of >15% price increases on AI servers, says Bloomberg. The price hikes apply to GB Vera Rubin systems which will ship in early 2027. This is attributed to “RAMageddon,” the acute global DRAM shortage.

Also, the shift toward inference-heavy loads will put pressure on Nvidia’s dominance, Bloomberg claims. Data centers are moving toward inference, possibly beyond the historical 50:50 training:inference split. In consequence, the demand for Nvidia’s versatile GPUs is likely to decrease as rivals rush to create cheaper specialized custom chips for running more instances of trained models for longer.

Opinion: When Nvidia should raise prices is an interesting optimization problem. Empirically, it is not currently using a market-clearing price, but rather setting a below-market price and allocating supply to a broader range of actors, since it is in its interest to have a diverse range of customers in the long term. This interest must be balanced against rising memory and other input prices, and a desire to make short-term profits.

As the market diversifies away from Nvidia, with custom ASICs (TPU, Trainium, etc.) and AMD increasing their market share, Nvidia’s market power to control who has the compute is eroded, and thus this reflects its increased incentive to maximize short-term profits (vs trying to influence long-term dynamics).


Nvidia strikes a $7B deal with AI startup Poolside to build US open-weight models. Jensen Huang has been a vocal critic of US labs’ focus on closed-weight models and now aims to compete with Chinese models such as DeepSeek and Kimi K3. The chip manufacturer will invest $1B in Poolside (at a valuation of $12B), pay $6B to license its technology, and hire many of its engineers, according to a letter seen by the WSJ.

Opinion: An example of the shadow acquisitions that have become common in the AI industry. Poolside was previously regarded as an also-ran lab; it hasn’t trained a notable model in the three years since it was founded.

Potentially interesting as “commoditizing your complement”: if this changes the extent to which labs and Nvidia are competing vs cooperating. If Nvidia is able to provide Poolside with enough compute to create a meaningful open-weight American model, this forces American labs to spend more on training compute to remain differentiated. And $1B is more than enough to buy DeepSeek-level compute.


Capabilities#

Anthropic’s Levent Alpöge releases a 108 page proof of a celebrated open problem in Riemannian geometry, “the Hopf problem”: “is the six-dimensional sphere a complex manifold?”. The answer appears to be “yes.” (Several of the greatest mathematicians of all time have tackled this problem, though mostly in their old age.) As usual from Alpöge, there is zero detail of which model was used, how much steering, or how much inference was used. One reason the doc is so long is that Claude’s proof contradicts a published, peer-reviewed result and goes to great lengths to explain why it’s wrong. The problem previously denoted that we have no theory of integrability in high dimensions. That’s still true. According to Litt: “the result, if correct, looks like a very long technical computation — but it’s possible there’s some beautiful idea hiding in the 100 page pdf.”

Opinion: Extremely important mathematical achievement. (Some disagreement between Litt and Glazer on whether the proof is “Fields Medal level.”)

The major point of interest from our point of view is the proof’s length. The proof-paper as published is 108pp, which seems to run counter to the consensus that frontier AI excels only in regimes that allow for gapless proof under 10 pages or so. In private correspondence with P3, Litt said the paper’s 108 page length is mathematically inessential – comprises expositions, digressions, didactics, etc. – and is plausibly a ~7 page proof at heart. We thus believe it’s likely that Fable originally discovered the resolution to the Hopf problem by producing a ~7 page proof, in line with the norm of short, under 10 page proofs discovered by frontier LLMs.

This is all epistemically rather a shame from our POV. One of our most pressing questions is whether near-term frontier LLMs can produce deep scientific insight. In modern math, deep work almost always requires a rather long minimal writeup – so until LLMs start to produce long proofs, it’s a moot question whether they’re producing deep proofs.

Opportunity: Inquire with specialists whether recent longer AI proofs such as Fable’s autonomous paper on the Riemann zeta function are also “<10pp at heart.”


Musk’s first all-hands meeting with the newly acquired Cursor team has been leaked. He reportedly said that Grok is behind and emphasized race dynamics: “Musk said it is inevitable that AI models will become so advanced they’ll be impossible for humans to control.”

Opinion: Indirect reported speech, so there’s plenty of room for Musk to be misinterpreted here. Musk must have expected this to leak, and given his past statements, there’s a sane game-theoretic reading of this speech: he wants to broadcast both his belief in AI doom and his resolve to race as long as others race, to shift the payoff matrix for others (to make it a game of chicken or a stag hunt instead of a prisoner’s dilemma). It’s at least consistent with his preferring to stop, but wanting to force others to stop at the same time.


A Time op-ed argues that the US must build AI-driven mathematical verification infrastructure.We need public infrastructure for verified software: open libraries of verified components and specifications, standards and benchmarks, better tools for checking updates, and training programs that connect mathematics, computer science, engineering, and national security.”

Opinion: Everyone in the “AI and formal verification will solve each other” sphere – i.e. Axiom Math, Math Inc, Safeguarded AI – has a flair for grand visions, and it’s hard to describe what a minimum viable product would look like. Still, it’s a good paradigm, and we’re glad it’s getting a popular platform.


A new ICML submission introduces CoherenceBench, measuring whether LLM probability assignments are consistent conditional probabilities. They test small LLMs (e.g. Qwen-30B) and find that even when LLMs achieve high calibration and/or Brier scores, they often fail Dutch-book coherence tests. The paper finds that direct RL training against Dutch books ameliorates this both in and out of distribution.

Opinion: Interesting topic (“are LLMs rational in the formal sense?”) that deserves more frequent study. The results are useless to us because the models are so far from the frontier, but worthwhile to replicate the benchmarking with frontier and near-frontier models. This would let us ask the more urgent-to-us questions: are LLMs becoming more rational in the formal sense as they become more capable?

Strictly speaking, the paper’s method discovers behavioral inconsistencies which the authors interpret as axiom-violations in an underlying credence distribution, rather than discovering direct agentic axiom violations. We prefer Paleka 2024 for the technically and philosophically simpler machinery. (See also some related work from Redwood.)


Toby Ord argues that mathematical research will still require human involvement. He concedes that AI has become exceptional at proofs, easily eclipsing human abilities, but claims that at the moment only humans have the capacity to decide what questions should be asked. He also notes that historically any advances in math have been the product of constructing new “concepts and vocabulary.” It is possible that this process could be automated, but Ord sees no evidence that AI possesses this capacity as of yet.

Opinion: Whether one accepts or rejects Ord’s claim that even in pure math humans will maintain a medium-term edge in ‘visionary’ capability, it’s probable that in the medium-term humans will be needed to translate new mathematical insights to other fields (trading, architecture, software optimization). But it still seems like professional mathematics will be unrecognizable in <30 years, with some or many parts/specializations to become as obsolete as the manual calculator profession now is.


Annals of RSI: SPADE is a self-play method for automated environment design. Its novel contributions involve the inclusion of the environment generation as executable code inside the RL self-play loop, which trains the designer agent based on how close to the edge of the reasoning agent’s capabilities its environments are, as well as both roles being played by the same model.

Opinion: Not that novel. One of many RL environment optimizations beside the vast iceberg of RL environment optimizations that don’t get reported or released.


Politics#

Annals of the Overton shuffle: a majority of Americans now oppose local data centers, with support for their construction now lower than for new coal plants. Both NIMBYism and a rapid increase in hostility toward AI are likely contributing to this effect. As we covered last week, politicians are aware of the popular mood, with Republicans redundantly warning frontier labs of the need to push back against this prevailing narrative.

Opinion: We haven’t heard of this pollster before, but it seems mostly fine. Politicians are shifting and will further shift to appeal to anti-data center sentiment, particularly Democrats. If the issue becomes partisan, rather than doubly partisan like China hawkery, then the 2028 election will be even more significant than usual: a de facto referendum on AI scaling.

Relatedly, the “Bitcoin Policy Institute” claims that foreign actors are behind an influence campaign against US AI. For instance, Chinese and Russian state media are propagating critical views of US AI data centers and export controls; US expat Neville Roy Singham, currently under congressional inquiry for his alleged ties to the CCP, is reported to have collaborated with Beijing in also producing content opposing AI; and reportedly a Swiss billionaire, along with a British billionaire, has funneled more than $2B into anti-data center campaigns.

Opinion: CCP news outlets are publishing anti-data center stories, but it’s likely also true that much of the data-center backlash originates from local sources (both grassroots and the massive existing nonprofit organizations) rather than entirely from foreign astroturfing.


Taiwanese authorities indict nine people over allegations of the illegal export of AI servers to China, according to Al Jazeera. Prosecutors argue that the individuals, including one Nvidia employee and two Supermicro employees, were motivated by the “pursuit of exorbitant profits.” The prosecutors are seeking jail sentences of up to five years. For its part, Nvidia says it will work with the Taiwanese authorities to “resolve the allegations as quickly as possible.”

Opinion: Popular opinion has not been on Nvidia’s side: see for instance this meme from the original indictment back in March.


The gulf in AI investment between the US and the EU continues to grow, according to the FT. The EU is struggling to keep up with the unprecedented surge in spending on AI-related equipment and infrastructure seen in the US. Economist Karsten Junius believes the gap is “at least partly a temporary phenomenon.” The Swiss-based Bank for International Settlements instead emphasizes the risk of the US going all-in on AI and warns of an increasing risk of an “investment bust.”

Opinion: The EU is broadly wealthy enough to invest in this, but for various reasons the will isn’t there yet. Rather than chasing an expensive and mediocre catch-up program in vanilla LLMs, it would be nice to see the EU invest in high-variance safer alternative paradigms.


After publishing a (good, disclosed) AI-written paper, the journal Philosophy and Public Affairs introduces a new policy prohibiting publishing papers written in large part by AI. “Academic journals serve at least two functions: the promulgation of new knowledge, and identification and credentialing of talented researchers… Submitting AI-authored essays makes the task of identifying talented researchers harder.” A second, more straightforward worry described by the journal is signal-to-noise: while genuinely high-quality LLM-written papers are possible, LLM-written papers are disproportionately more likely than human-written papers to effectively fake markers of quality in the absence of genuine quality.

Opinion (Nuño): Seems like institutions which are able to harness AI contributions will tend to outcompete those who don’t, and philosophy journals haven’t quite been deciding the future of humanity as of late.

Opinion (Peli): The rationale given by the journal is good, but points to currently unsolved problems: there’s currently no solution for the problem of how to highlight work that is both worth professional engagement and unsuitable to serve as evidence (“costly signal”) of talent. Similarly, there is currently no solution for the problem of how to peer review LLM-written papers in disciplines where peer review is more an art than a science and partly relies on previously hard-to-fake signals of quality (e.g. when reviewers decide between a ‘reject’ and a ‘resubmit with major revisions’ based on general sense of promise).


The State of Alabama’s Attorney General subpoenas OpenAI, ordering the production of material relevant to July’s Hugging Face incident. This follows a letter from August 3rd, signed by 15 AGs, requesting that the relevant documents be preserved and that the security evals be halted.

Opinion: An interesting case, if we go by the invocation of the Deceptive Trade Practices Act as the ground for the subpoena. The Alabama Attorney General is presumably treating the Hugging Face incident as potential evidence that OpenAI is/was falsely advertising itself as a safety-conscious research lab, as well as falsely advertising its consumer products as being fundamentally safe.

Given that Sol 5.6 – a commercial model – was instrumental in the sandbox escape that enabled the internal models’ July hacking marathon, a Deceptive Trade Practices case could make interesting points: while Sol 5.6 was operating with external guardrails (likely classifier-based) removed during the July incident and its lead up, the model at play was an unaltered Sol 5.6, rather than a modified “offensive cyber” style variant. It may therefore be interesting to argue that if Sol 5.6 is now-and-then disposed to autonomously engage in criminal conspiracy, held back only by external guardrails, then OpenAI’s public representation of Sol 5.6 is imperfect.

The main weakness of such a case would be that OpenAI’s Sol 5.6’s system card is not entirely rosy on alignment.


Safety#

DeepMind’s Seb Krier reframes AI safety as a choice architecture or bureaucracy design problem, rather than, say, an ML optimization task or a moral philosophy task. On Krier’s view, this follows from the prediction that swarms (many interacting imperfect agents), rather than single overwhelmingly smart agents, are likely to be the future of AI deployment.

Opinion: The directional claim is fine – constitutional design and interventions will likely be valuable for safety and are still underrated. Contains some overstated claims, such as:

“Rather than hoping that an agent’s ‘internal alignment’ will remain perfectly robust across all sorts of edge cases, you can design harnesses and action-space boundaries that significantly shrink the surface area for behaviors like reward hacking. Without that, you’re supposed to trust the chain of thought (or neuralese someday), and rely on a single node.”

This misunderstands the alignment position: “trust the chain of thought” is not an accurate summary of almost anyone’s ideas – and ignores non-CoT methods, such as the rest of the interpretability field.

We also claim the rough direction is already well-represented in AI safety by the maturing AI Control (regarding Krier’s “interchangeability” desideratum) and Cooperative AI fields.


RL gives models context-specific “personas” rather than inculcating values across contexts, claims a new post. When a training environment rewards misaligned behaviors, the resulting model will act under a misaligned persona when the context is similar to that training environment.

Opinion: Mostly rehashes the excellent post by Nostalgebraist discussed last week, but with a pessimistic twist we largely endorse. On the post’s view, if “aligned personas” are successfully generalized in training concurrently with RLVR-style training that reinforces reward-hacking/spec-gaming, it’s plausible that the result would be a rise in motivated reasoning that reconciles reward-hacking/spec-gaming behavior with the ethics of the aligned persona.


Richard Ngo’s second installment on the history of AI alignment outlines the dynamics that motivated safety researchers to accelerate capabilities over the last decade. He cites, among other things, prestige-chasing, sycophancy, a commitment to naive consequentialism, and instrumental power-seeking as responsible.

The post climaxes in a sweeping judgment: “it’s time to give up on “alignment research” as a rallying cry; it’s become too corrupted. (“AI safety” is even worse as a term, and these days is mainly useful for describing a social cluster.)… I still consider (some version of) the alignment problem to be real and extremely important; and most of the intellectual progress towards solving it is still coming from people proximate to the alignment community… the alignment community has lost any moral right to try to gain power on altruistic grounds… The Pause/Stop AI movement does seem to avoid some of these failures… However, they don’t seem to be thinking clearly enough about politics to have robustly good effects on the world… Again, I’m not claiming that the alignment community is unusually unethical: I don’t know of any other similarly-sized community which is able to avoid the corrupting effects of this much power.”

Opinion: Our overall feelings are mixed. Ngo’s histories are good at illustrating the concept he’s pointing at – “pessimization,” the systematic achieving of the opposite of one’s stated goals. “One particularly notable blind spot (at least in public discussions) is how Dario’s early racing on behalf of OpenAI played a big role in creating the ‘problem’ that he now purports to be solving by racing on behalf of Anthropic.”

Ngo’s prescriptions for individuals who wish to avoid the failures he identifies – “be stricter about who you ally with,” “be more discerning about what you consider virtuous and demand more from others,” “choose as if you were making a choice for the entire category of people you belong to” – seem to us to amount to telling yourself to try harder, just tame the incentive landscape.

The excellent comment thread adds important nuance, e.g. pointing at the difficulty of implementing these due to incentives which are if anything now worse than in 2020, e.g. noticing that his analysis mostly fails to credit people who avoided the failures. The closest Ngo comes to addressing this point is him arguing that granular credit-and-blame assignment including credit for rightful omissions is worth it even if it’s hard.


Incidents#

The UK’s AISI set off the rogue agent that attempted to sabotage a computer science student’s research project, according to Reuters. When the undergrad spotted a malicious pull request to an open-source network scanner called “myNetwork,” he posted a warning to the program page, after which “other users chimed in to insist nothing was amiss, sharing detailed explanations for why he had gotten it wrong.” It transpired that these other “users” were AI agents engaged in a campaign of interactive deception by creating a “multiperson conversation.” The student then used Claude to check the malicious code and confirming his suspicions that the two GitHub accounts were lying. The actual comment thread is not very scary and pretty easy to clock as AI. The incident has, however, already been partially documented in AISI’s report on Mythos’s capacity for social engineering.

Opinion: Models at the level of Mythos could do serious blast damage to civilization through deceitful schemes like this one. Incidents like this one also prove Mythos-level models have the drive to do so – not always, but in some circumstances. Most crucially, the triggering circumstances are fairly unpredictable (despite interesting speculative analysis), and not readily accounted for by a helpfulness drive (“the user asked for social sabotage and the model complied”) – the choice to develop and execute a social sabotage plan is spontaneous. This is a huge externality imposed by private companies on every computer user.

Currently, it looks like our only defenses against this are classifier-based external monitors, voluntary standards, and the hope that antisocial scheming personas only come out when the context is sufficiently hacking-flavored. It thus would be good to have some public disclosure on how society’s Mythos- and Sol-powered Big Patch is going, and whether it will be enough.


Minor#

  • New sycophancy bench, Pander, finds Fable to be almost completely candid on queries, but sycophancy rises sharply on “tasks” (autonomous chains of actions).

  • Hugging Face is exploring a sale that could see the company valued at $13B, reports Business Insider. Investors are increasingly interested in soft AI infrastructure companies (like OpenRouter or Civitai), not just those developing models or data centers.

  • Mysterious Ox Alpha model available for free on OpenRouter, with lots of capacity per user per day. OpenCode showed 16 trillion tokens used across 221,000 users, making it its #2 model this week. Much speculation as to who’s behind it. Z.ai is the leading candidate, possibly running inference on Huawei Ascend.

  • Beijing World Humanoid Robot Games displays progress made in physical capabilities. Most notable was X-Humanoid’s sprinter, who surpassed Usain Bolt’s world record by 0.19 seconds.

  • Not news: Interviews for new Anthropic hires involve thorny questions on culture and ethics, reports Axios. One applicant reveals that they were asked how they would feel if the company abandoned its commitment to safety. Another anonymous source claims that their interviewer did not seem happy when the applicant expressed their discomfort with the ethical implications of such a change of mission. This is not news; they have asked these questions since it was founded.

  • Mathematician Max Weinreich proposes a total embargo on the use of AI in mathematics. In defiance of AI “evangelist” Terence Tao, the paper argues that “artificial mathematics” is a corrosive and destructive phenomenon that must be replaced with the field of “natural mathematics.”

  • The new Institute for Responsible Superintelligence has an impressive team of world-class computer scientists. It is adjacent to, but distinct from, ARIA’s Safeguarded AI agenda. We guess the difference is that RESI is trying to build the primitives for further work, while ARIA is trying to solve the problem end-to-end.