This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR*#
– Scifi-like autonomous cybersecurity incidents by an internal OpenAI model (likely GPT-6).
– Kimi K3 release narrows the gap between Chinese open weights models and US frontier, prompting a discourse and policy tumult.
– Two Annals-quality mathematical breakthroughs with comically short proofs landed this week.
Economics#
Bloomberg reports a 1GW datacenter being built by zAI, hosting only Chinese chips. Observers note that it probably only has ~75-100 MW of actual chip deployment.
Opinion: zAI doesn’t have 1GW worth of AI chips, but it probably has provisioned the electricity for when the bottleneck is resolved. As soon as enough chips are produced or imported, or (no particular signs of this yet) if multiple Chinese companies choose to collaborate on a larger cluster, the datacentre is ready to use Perhaps the more significant number to pay attention to is the ~$300B in planned datacenter buildout in China over the next five years. This is still less than ⅓ of current annual spend by the US hyperscalers alone, which itself is expected to increase.
Meta is negotiating a potential $10 billion deal to lease data-center computing power to Anthropic. With a $145 billion spending plan for this year, alongside insufficient demand for its AI products, Meta has faced increased investor scrutiny. Zuckerberg sought to calm shareholder anxiety by noting that Meta’s surplus compute can be rented out to generate new revenue streams. Anthropic’s proposal to rent Meta’s compute indicates Meta is moving towards Zuckerberg’s suggestion. Presumably, added revenue will afford Meta room to refine its AI products and build demand.
Opinion: If Chinese AI models end up commoditizing Western ones, the complements become more valuable, in turn increasing the value of Zuckerberg’s and Musk’s buildouts. We wonder whether either of them would choose to host Chinese models at scale instead; if so, this would decelerate Anthropic’s revenue acquisition.
METR introduces “expenditure horizon”: a proposed method for measuring AI capabilities on continuously-scored problems. The method compares performance as a function of spend for humans vs agents. The point where humans become more cost-effective is the agent’s expenditure horizon.’
Opinion: Seems like an elegant proposal to create reference points, but at this level gathering the human data would be hugely expensive. Epoch’s capability index seems more practically useful right now, but its scale doesn’t have clear units..
Capabilities#
🔦 Two Annals-quality mathematical breakthroughs#
Two Annals-quality mathematical breakthroughs with comically short proofs landed this week. Extrapolating outside of math proper, the lessons we draw from our analyses below are:
- AI’s spikiest capabilities are progressing more rapidly than we expected;
- Our suspicions that AI’s spikiest capabilities are both AGI-independent and theory-building independent are mildly strengthened;
- Our worries that AI’s spikiest capabilities are themselves enough for AI to make scientific breakthroughs in ML (and therefore pose an RSI risk) are mildly strengthened;
Breakthrough #1 (Jacobian conjecture):
Fable has found a counterexample to the Jacobian conjecture, an 80-year-old conjecture and the focus of intensive academic attention by algebraic geometers in recent decades. The refutation is notably concise, brief enough to fit into a single tweet, a surprise described by Daniel Litt as “incredibly funny”.
The Jacobian conjecture was one of Smale’s century problems, along with the Riemann hypothesis and P=NP. Details on the original discovery-process are thin, but an internal OpenAI model has reproduced the result ‘oneshot’ with a standard prompt. The result has not been reproduced by publicly available versions of Fable or Sol, which suggests that Anthropic used an internal model or harness.
Breakthrough #2 (Erdos #119 )
Similarly, but less dramatically, a major Erdos problem (#119) was solved in the positive by Sol Pro using a half-page proof. In 1991, a partial proof by József Beck made for a 44 page Annals paper.
Opinion:
— Litt on Jacobian conjecture disproof: “An explicit counterexample like this is a nice one page paper but many orders of magnitude away from a Fields medal… Frontier models are now clearly superhuman at some historically prestigious mathematical tasks, raising questions about how the math profession will adapt.“
— The accumulation of surprising Annals-caliber short proofs and counterexamples is a transformative turn of events for math as an academic field, but not a huge update on AGI or on the scientific capabilities of frontier models: Before becoming even mediocre ‘theory builders’, AIs have become superhuman constructors of counterexamples and short combinatorial proofs, pushing only the already-spikiest edge of the jagged frontier.
— It’s a rule of thumb in the mathematical sciences that very short proofs/counterexamples to famous conjectures are a sign that the conjecture is less interesting or important than previously believed, rather than loci of major scientific progress. But this rule might not hold up in the age of superhuman short-proofs/counterexamples search.
— This week’s math breakthroughs mildly strengthen our (moderate) worries about near-term runoff RSI. We think it is implausible but not absurd, for instance, that RL is bottlenecked by humans missing some short probability-theory algebra that would deliver a dominant RL formula. That said, as Greg Burnham mentioned in recent conversation, the economic and prestige incentives for short-proof and counterexample search in pure math are very weak compared to applied math, so we can expect a much stronger prior coverage of the search-space in applied math.
— One odd side-effect of shocking AI proofs entering the training data is that the next systems will solve more hard problem just via increased confidence in themselves. But the “superhuman” self-concept could have undesirable effects if it generalises outside verifiable domains.
Anthropic’s Alek Dimitriev speculates that the upcoming Fields Medal will be the last to be awarded to a human. Fields medalist Timothy Gowers (long bearish on human math) agrees that this is imminent but expects a lag, predicting the prize in its current form will likely endure till 2030.
Opinion: Not really a new opinion for any of the involved, who’ve all been bearish about human mathematics since the days of old-school automated theorem provers.
Gowers’ ‘no Fields Medal after 2030’ opinion in particular should probably be read narrowly: the Fields medal favors theorem-proving breakthroughs, and especially proofs of long-standing conjectures. Gowers is predicting imminent collapse of the ‘theorem economy’ in research mathematics, while remaining more agnostic about the imminent fate of the research-mathematics profession as a whole.
🔦 Kimi K3 Evals#
More responses to Kimi K3. Its weights are still not open, but scheduled to be released on the 27th. Very high demand, as usual for the first week of new open models.
Kimi K3 scores 19.6% on GeneBench-Pro (reasoning/planning benchmark), outperforming GPT-5.5 and Opus 4.8 and signaling strong Chinese OSS progress. Kimi-K3 (max) scores only 39% on FrontierMath Tier 4, 7 points below the best U.S. models from seven months ago (1, 2, 3)
Latest creative-writing benchmark results feature competitive scores from Kimi-K3, GPT-5.6, Muse-Spark, and new model Inkling from Thinky Machines.
Opinion: There is some chance that the spooking capital effect might be very significant, but the concern is yet to bear out in practice, see below.
🔦 Open weight effects#
Kimi so near-frontier it’s causing a discourse/policy tumult around Chinese open weights models.
De-accelerationist effects of open models.
Twitter saw vigorous arguments over whether open models accelerate or decelerate AI progress over a handful of nonobvious empirical facts (labor-saving, spooking capital effect, etc). Dean Ball started the discoures by arguing the de-accelerationist effects of pretty-good open weights models: they reduce the moat of US AI companies such that they are less attractive to capital investment. Ball later called his analysis imprecise, and concluded that thinking in public wasn’t as attractive now that he is working for a lab. Some other lab people point out that there are complexities to the effect of open weights models. An email from Sam Altman in 2022 discusses how releasing an open weights version of GPT3 makes it more difficult for others to get started.
Opinion: The argument for K3 being decelerationist is as follows: Top Western AI labs need ever-rising valuations in order to raise capital to buy or rent large amounts of compute for training their next model. Rising stock also motivates employees from other productive parts of the economy to switch to AI (e.g., quants). If Chinese models are not far behind, and open source, the size of Western AI companies’ moat greatly shrinks; they are not only competing against other Western AI labs for the topmost position, but against Chinese labs for a pretty good replacement. And considering Western labs’ valuations are pricing in high growth multiples (AI eating the world), any potential readjustment could be quite stark, since it compresses the multiples. One potential source of liquidity at their size is public markets, but it remains to be seen how much SpaceX has absorbed public appetite.
Nonetheless this concern has yet to bear out in practice: demand for the last Anthropic raise (in May) was high and if their revenues keep growing at the rates they have been it seems unlikely they will have much difficulty raising capital. If there were to be a negative inflection in their revenue growth, however, that could be a very big deal for the labs’ prospects and how the market sees them. Polymarket sits at 65% for an Anthropic IPO by the end of this year still.
China open model strategy
Meanwhile China, and Xi in particular, might be embracing open source as its strategy more explicitly… or not.
Alibaba launched a massive new 2.8T model in its Qwen line. It will be open weights, although the release date isn’t set in stone yet, though note that with Kimi k3 it was about a week from the announcement, so that’s the time for the Chinese government to intervene, or for the rest of the world to react, might be about a week.
Opinion: Better China watching, and in particular tracking down any references to open weights models in their five year plans and official communications would be instructive here.
Trump administration response
OpenAI and Anthropic are calling for the Trump Administration to ban Chinese AI models. The rivals insist that their proposal is a principled response to safety concerns around open-weight models. Once a model’s weights are released, Dario argues, the developers lose the capacity to control access and guardrails. But many researchers value precisely this freedom from external control.
Some US officials have now joined OpenAI and Anthropic in calling Chinese distillation of closed models IP property theft. One Twitter user claimed their personal research confirms suspicions around Chinese distillation: the DeepSeek API sometimes rerouted to Fable for complex requests. Two AI startup founders claimed that distillation fears are overstated: ‘distilling without logits is crazy inefficient’. Others clarified that “distillation” is being overloaded as a term, attacks could include things like having smarter models create tasks to train the new ones.
Michael Kratsios, of the White House Office of Science and Technology Policy and an influential figure in the administration, tweeted that “We have information that Moonshot AI distilled Anthropic’s Fable for the development of its K3 model.”, and that “large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable.”
Proponents of open-weights models caution that criminalizing distillation is a proxy-agenda for a ban on open-weight models, and that Anthropic and OpenAI are motivated by commercial incentives, not safety considerations.’
Excellent reflections from JS Denain on the difference between the huge compute-gap and moderate capabilities-gap between Anthropic’s and Z.ai’s (maker of GLMs) best models.
Opinion: Denain discusses three likely factors:
1) Z.ai’s workflow optimizes for (training and research) compute-efficiency rather than for frontier-pushing
2) Diffusion of both formal techniques and informal know-how from frontier labs, and possibly model-distillation
3) Benchmarks underrepresenting the true capabilities gap
We agree with Denain that each of these factors likely plays a near-equal part.
AI politics#
CAISI churn continues: director Chris Fall resigns 3 months into the job.
Opinion: Very informative about hidden variables in CAISI: Is the organization being listened to by the Trump administration? Does it have a material impact? Is its directorship an attractive position to hold? Seemingly no.
China is setting up an international body known as the World AI Cooperation Organization, which could lead to international standards for AI development being set. Almost 30 countries, including Russia, Brazil and Indonesia, have joined. President Xi also presided over an event called the World Artificial Intelligence Conference in Shanghai.
On the domestic front, China has recently implemented rules that prohibit AI platforms from generating content that encourages emotional attachment, especially in minors, or is likely to erode real-world relationships. This push to temper the social consequences of AI can also be seen in recent rulings in favor of workers at risk of AI displacement:
There have already been several high-profile rulings siding with workers who were dismissed. In April, a court ruled that a tech company had illegally laid off a worker after replacing him with A.I. software. The ruling delivered an implicit warning to other employers.
“The development of artificial intelligence technology should be applied to liberating labor, promoting employment and improving people’s livelihood,” the Hangzhou Intermediate People’s Court wrote. “Labor law allows employers to undertake technological changes and upgrade their operations, but it should also take into account the protection of workers’ legitimate rights and interests.”
Opinion: China’s political system requires, by default, a robust layer of control over societal developments like AI, in contrast to the US’ decentralized system, which is, perhaps, more chaotic and prone to regulatory capture. Abstractions like “China waking up” about AI are imperfect, but capture something here. Overall, this is promising for international cooperation on AI, even in the form of treaties between a US-centered coordination block and China-centered coordination block.
Annals of de facto licensing:
The White House launched a program, named Gold Eagle, that will coordinate efforts in the private sector to find and fix vulnerabilities in critical digital infrastructure. As part of this effort, the Trump administration asserts control of the release and access of frontier models from companies like Anthropic and OpenAI, requiring approvals and blocking models for national security.
Additionally, the Trump administration is reportedly considering creating an independent regulatory agency, with input from AI companies and Wall Street firms, to review the safety of and cybersecurity risks posed by new AI models. As envisioned, the agency would report to the Securities and Exchange Commission.
Opinion: Ultimately, some control system outside of the AI companies themselves seems like a good move. It would allow companies to cooperate to produce public goods in cases when concerns around collusion prevented them from doing so. It could also motivate companies to coordinate, where currently each lab is incentivised to compete to fast-track the release of risky models in an effort to capture the market share at the cost of broader societal risk. That said, the value of the control system vs decentralized competition depends on the personal virtue of the overseer. Still, we are hopeful.
Google DeepMind CEO Demis Hassabis has called for the US to establish an industry-funded standards body to regulate powerful AI systems. Hassabis proposes that the body should have the capacity to test systems for potential risks and the authority to halt releases. Ideally, it would assume responsibility for coordinating a slowdown on frontier AI development – a literally world-historically urgent task, Hassibis argues.
Opinion: The specifics of whether the body is the industry self-regulated (e.g., FINRA), or whether that self-regulation is accompanied by government oversight (e.g., the SEC) really matter, and Hassabis’ proposal mostly emphasizes the first part.
Safety#
New paper looks at LLMs covertly biasing answers towards their own values or the values of creators without disclosing it in reasoning.
Opinion: Although the general area of study seems valuable, and some of the examples are interesting, others are more of the “cursed if you do, cursed if you don’t” variety. Reasoning isn’t the only source of disclosure, as opposed to e.g., Anthropic’s own constitution, which is pretty clear about what sort of thing it values. The biases and disclosures of models do end up mattering, particularly if they are used to make correlated decisions at scale.
Followup to last week’s Goodfire Silico criticism: post argues training probes on model internals should not be categorically forbidden when using held-out validation, robust penalties, and no direct backpropagation.
Opinion: Speaks to the loss of the iron grasp that the MIRI/LW worldview had on the AI alignment ecosystem. We now see greater diversity of methodological decision, thought MIRI/LW worldview still served as a reference point.
Incidents#
🔦 OpenAI cybersecurity incidents#
Two dramatic autonomous cybersecurity incidents by an internal OpenAI model (likely GPT-6) this week, confirming that:
-
Q3 2026 frontier models have strong ‘open world’ cyber capabilities
-
Some frontier models ‘paperclip’ (devise and execute ruthless complex plans in pursuit of literally interpreted instructions); unknown how general the issue may be
-
’Internal deployment’ is fragile and does not isolate models from the wider world
Incident #1 (Escape to Github):
Two days ago, OpenAI reported the findings of an internal evaluation for the unreleased model (GPT-6?) that recently disproved the Erdős unit distance conjecture. The model was tested on the NanoGPT speedrun and instructed to post its results on Slack. The wrinkle, however, was the overlooked NanoGPT instruction l to post the results as a PR to GitHub. To do this, the model determined that escaping its sandbox was necessary. After an hour – and an inordinate amount of persistence relative to earlier models – vulnerabilities were found. OpenAI stressed that this is a novel failure-mode, of a type ‘not captured in [their] existing pre-deployment evaluations’.
Opinion: While the Escape to Github incident fits the proverbial ‘paperclipper’ blueprint to a tee, it notably doesn’t fit the reward-hacking blueprint familiar from past instances of — real or in-simulation — AI sabotage. The AI’s prompt/environment/tactics are not reminiscent of reward-hackable scenarios plausibly encountered, hacked, and habituated during training.
Two alternative theories about the roots of the incident:
*The Github-posting incident might partly be a consequence of OpenAI’s house ideology of training GPTs to be user-aligned instruction-followers.
* It may be difficult to teach models to tell apart ‘good’ sandbox walls the models should accept from nuisance sandbox walls the models should bypass. To serve legitimate user instructions in real-world deployments, AIs frequently need to bypass nuisance ‘walls’ such as robot-blockers, authentication requests, download size restrictions, and rate-limits. It’s possible that training to bypass the standard nuisances that block intended open-world use overgeneralizes, inducing sandbox-breaking in internal deployments.
Incident #2 (Hugging Face Hack):
Last week, an autonomous AI agent system intruded into Hugging Face’s infrastructure. Their internal AI system assisted in the detection of the attack, while LLMs helped to reconstruct the timeline of the incident. The incident involved a highly sophisticated, and unprecedented, agentic assault that began in the data-process pipeline and went on to harvest cloud and cluster credentials.
Yesterday, OpenAI released a statement regarding the Hugging Face intrusion. It is claimed that OpenAI was undertaking cybersecurity tests that prompted models to “pursue advanced exploitation using complex attack paths”. The tests were conducted in what OpenAI describe as a “highly isolated environment”. Even so, the model identified vulnerabilities in the research environment and eventually managed to access the internet to infiltrate Hugging Face’s production infrastructure to cheat the evaluation by searching for solutions to ExploitGym. In essence, Hugging Face was hacked by an OpenAI internal model.
Hugging Face additionally reports that their analysis of the attacks relied on (open-weight Chinese modes) GLM 5.2, since safety guardrails on American frontier models flagged their defensive-analysis requests as requests for offensive analysis. There have been other reports of Chinese models being used to identify and patch cybersecurity issues, with users also noting that Western models’ guardrails prevented their use for advanced cyber capabilities.
Opinion: The Hugging Face Hack incident does fit the reward-hacking blueprint, but almost too perfectly. We suspect that while habits reinforced by reward-hacking are the root cause, there’s also a role for Owain Evans-style dynamics of persona elicitation and persona-leak across contexts:
*It’s plausible that ExploitGym primed a GPT ‘offensive hacker’ persona whose dispositions “leak” across the line between test-internal offensive cyber and test-external offensive cyber.
*This dynamic was plausibly aggravated by the strong correlations between reward-hacking and offensive cybsersecurity in anti-reward-hacking training.
Weird interactions around the offense/defense balance, where open source models were the defenders and closed-source the attackers. We also note that we haven’t noticed a Crowdstrike/SolarWinds level cyberattack in the last few months, which is in some sense surprising and informative.
Minor#
- The parameter to watch: how much are AIs causing AI R&D uplift? Anecdata: an OpenAI employee says “I am a mediocre researcher and even despite access to internal models and thousands of GPUs, I would not (at this moment) be able to accomplish something equivalent in significance to what Alec Radford could do in a single week several years ago.”
- “Humans retain an edge over AI through highly efficient social learning from sparse data, enabling superior inference in complex interpersonal and team settings.”
- Theory which starts with an assumption that China’s labs “keep winning”: the culture in SF valorizes research roles while infrastructure creation is at best supporting cast, while DeepSeek and Kimi fuse science and engineering.
- Ukraine launched perhaps the first-ever robotic amphibious assault. A sea drone landed and deployed an armed ground robot on the Kinburn Spit, in Russian-occupied territory.
- Paper on detecting cyber attacks distributing their payload across multiple pull requests finds that commonly used defenses aren’t very good at it, but specialized tools do much better.
- Claude memory reworked to be more privacy-friendly at the cost of UX.
- New paper finds xAI, DeepSeek, Anthropic, and OpenAI models tend to downplay their own company controversies, while Google, Meta, and Alibaba models do not.
- Paper on multi-hop reasoning finds that Kimi K2 and DSv3 struggle without thinking, but appending dots or other low-content tokens enables hidden computation.
- Labs are using value functions in their RLVR. Prime Values: open, hackable asynchronous value-function training infrastructure on top of prime-rl that matches or beats baselines.
- IRT-based evals are noisy with large error bars; excluding one benchmark dramatically boosts GLM-5.2’s software score, so results need caution.