This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR#
Economics#
SpaceX will buy Cursor after IPO, for the agreed $60b.
Opinion: Elon is not giving up on having a frontier lab yet. Spending 3% of their market cap on some good AI researchers is not crazy (though it seems unlikely the team at Cursor is as strong as OAI/Anthropic or even Meta, so catching up still seems difficult). Recall that Zuck paid ~~$35b for his similar crew.
If he’s thus forced into supporting toolworld by being late and failing at general AI, then that’s a win.
New estimate of AI productivity, US data through 2025 Q4. Their methods revolve around estimating what portion of an industry’s share of work can be replaced by present AI, then observing whether the productivity of industries with higher automatability grew faster post-2022 than pre-2022. The authors extrapolate $1.26T per year of additional productivity from AI across the world.
Opinion: Not stupid but not reliable. “$1.26T” is the high end of their own ranges, and the authors are doing PR for their investment firm so are incentivised to come up with a high number to get attention. The $880bn number for the US would be 3% of GDP and we expect this would be more obvious elsewhere if true.
Article on how AI hardware demands are ending the era of cheap consumer electronics, because of datacenter demand for high-bandwidth memory. Compared to double-data-rate RAM used by most smartphones, they are about 3x less wafer-capacity-efficient. Most manufacturers are shifting production towards HBM, causing component costs to spike.
See also: Hetzner, a major cloud, triples prices on some services following restructure. Average ~2x.
Opinion: Nothing new here - Apple etc can afford to pass through a $150–200 bump in component costs on a $1000 phone, but $150 phones with 8GB are no longer viable.
Probably we will see phones with less memory produced to meet this demand, but there’s only so far that can go.
As with power bills, good source of anti political capital if anyone wants some.
Capabilities#
DeepMind paper on how to get from AGI to ASI, mapping out 4 pathways: scaling, paradigm shifts, recursive improvement and multi-agent collectives. No novel technical content; some partial formalisation of open research questions.
Opinion: It is sadly going to be hard to avoid this transition simply because there’s a bunch of ways to do it, but some are faster and more dangerous. Good summary, makes solid points like the pre-AGI and post-AGI phase transition being blurry and hard to pinpoint, but none of that is novel. Most important as a signal of respected researchers treating the topic seriously and reaching for more rigor.
What kind of person gets the most out of Claude? Anthropic analysis of 400k Claude Code sessions finds coding ability is less relevant for successful Claude use (as judged by an LLM classifier): human user’s expertise in the domain predicts Claude session success (also judged by an LLM classifier, based on precision of instructions, presence of checks and frequency of catching Claude’s mistakes).
Opinion: Switcheroo: the article talks about domain expertise, but the findings are almost entirely about how good a person is at using coding assistants. They find that the better someone is at it, the better the results are. This is not surprising.
Annals of benchmarks being bad evidence (but underestimating capabilities this time): Agents are still underelicited. Simple and general prompt/scaffold interventions can roughly double agent performance by getting agents to use more resources more efficiently.
Opinion: Can be fixed by spending more, but won’t. Sadly, focus your attention on evaluators with the resources to spend on inference.
One of the original Transformer paper authors, inventor of MoE models and co-lead of Gemini leaves Google, joins OpenAI as lead for architecture research. See also Dean Ball.
Opinion: Google slipping out of the race.
Shazeer was acqui-hired when Google paid $2.7bn for his startup character.AI less than two years ago. Leaving within 2 years suggests things not going well at Google and OpenAI offering him an absolutely enormous pay package.
Various rumours have had OpenAI as weaker at pretraining than Anthropic (and comparatively stronger in RL) for a while, this could be seen as an attempt to address that.
Essay arguing that AI scientists will continue using specialized tools even when they themselves are at the level of superintelligence, somewhat contra the Bitter Lesson. This would lead to an increased demand for high-quality conventional tools.
Opinion: Unclear what the distinction between “advanced model calls trusted external calculator” and “advanced model contains a calculator subroutine” is at the limit, but it does matter right now for cost and reliability. And this still doesn’t help humans that much.
One-month social experiment for 500 human participants, each of which will each receive a personal AI agent. These agents (running on OpenClaw) will interact to manage their users’ schedules and mediate community governance. Plans for data to be published, especially pertaining to value drift and unprompted collusion.
Opinion: More cool than important. Findings will almost certainly be surprising. However, running such a large experiment early in AI with early agentic systems may not be representative of larger effects.
Politics#
Anthropic’s Advanced AI Framework released: policy suggestions for true catastrophe monitoring and resilience. Centered on the US and pushing for risk testing (bio, cyber, loss of control, RSI), incident reports, independent evals and, most controversially, enforceable government oversight including state ability to regulate above and beyond federal rules. Also spends 30% of its volume on a broader resilience beyond AI, including personnel vetting, gene-synthesis screening and general cyber preventive measures.
Opinion: Yet another doc from them to add to the confusion, but we expect this is the one to watch.
Serious attempt to consider legislation, and lead an agenda. Sadly Dario has cost them short-term influence but he might be betting on it increasing their middle-term influence.
Trump admin weighing how to structure government equity stakes in frontier labs. Treasury Secretary Scott Bessent wants the equity to fund “Trump Accounts”, while Commerce Secretary Howard Lutnick prefers a sovereign wealth fund. Also Bernie launched his 50% expropriation sovereign wealth bill. Lawmakers and lobbyists are skeptical any of these plans will happen.
Opinion: Again, Altman has been pushing this. Again, this would solve their funding problems forever and load the state up on liabilities as well as buy them a ticket to the lightcone.
In response to being convinced they overreacted about last week’s Mythos “jailbreak”, the White House continues to speedrun reinventing the field of evals and ignoring their own CAISI.
Opinion: The effort is admirable, but the degree to which they are underbriefed / actively ignoring their own infrastructure is worrying.
Dean Ball on the broader meaning of killing Fable. “What the law says does not matter. What Administration officials argue on one day does not matter. Anthropic is a political enemy of this Administration, in part because they have explicitly chosen to make themselves one. It is simply naïve to think that your company can operate under such circumstances without an extreme degree of regulatory caution. And given this context, Anthropic’s actions are viewed by many within Washington as not simply unwise, but actively antagonistic.”
Opinion: Part of the EA worldview that underlies Anthropic is that the individual can and should reshape the world for the better. Whether this is an usurpation of the state to be resisted depends: should the government exist as a neutral referee of rules collectively determined by broader society, or as a nimble, unitary, fully delegated imposer and enforcer of rules?
For 300–400 years we were moving towards neutral arbiter states vs absolute power states, but it seems the current administration would roll this back if they could.
We believe that everyone, Ball included, who models Anthropic as primarily naive or autistic is being stupid. As a guess they roughly know what they’re doing and are resisting and maliciously complying to accentuate contradictions they want to expose.
Safety#
OpenAI paper with big claims of emergent alignment / value generalisation from RL alignment methods (e.g. from healthcare setting to other settings): substituting 5% of RL data with synthetic conversations portraying beneficial traits results in broadly reusable positive learned behaviors, seems to transfer to reduced reward hacking and deception, and even makes steering toward harmful behavior harder.
Opinion: The paper doesn’t explain a lot of the fine details which makes it hard to tell exactly how meaningful points like “making harmful steering harder” are. The specific method of mixing beneficial data in with ordinary RL data somehow hadn’t been done before, and the results as written are encouraging.
Clever point: AIs will use the estimated construction cost of the environment it’s in as an easy guess for whether it’s being evaluated: e.g. has it been given 1000 B300s to play with? Permanent limit to all pre-deployment testing. The AI might behave well in cheap scenarios to pass the tests and wait for expensive, data-heavy tasks (like frontier AI research) to act maliciously.
Opinion: Known issue, but this one involves a nice dive into why naive fixes won’t work. Remains primarily relevant for scenarios involving deceptive alignment, though also partially interesting for other eval awareness concerns.
Call for safety research to get good: research attributing agentic misalignment to AI (deception, scheming, self-preservation) often overstates what experiments show because of vague concepts, brittle datasets, unreliable evaluators, weak controls, and correlational rather than causal evidence.
Opinion: Lots of research plays fast and loose with the implied versus explicitly supported level of severity for currently-findable misalignment issues. The field would be better served by drawing cleaner distinctions, but most reputable researchers are already careful to do this.
Surprisingly sensible and coherent take on an alternate approach to model safety from Huggingface guy: capability restriction by data filtering/avoiding risky RL envs, staged releases, independent evaluation and enforcement agencies, heavy support for open source models.
Opinion: Nice to see “just don’t train on it bro” being taken up. It would sort of work for a while.
AISI paper on lie detection in frontier open-weight models, finding that simple prompted lying is easy to detect while models trained to be deceptive are not very visible to activation and logprob probes (~0.85 -> 0.65 balanced accuracy scores at best, random chance at worst). CoT judges do better, but this is partly due to experimental design; authors warn that current lie detectors can’t support high-confidence claims about model beliefs.
Opinion: Just open models, but the overall lie detection even without specific training seems pretty grim: topping out at 0.85 is not good (this means there are 30 percentage points split somehow between false positives and false negatives). Random chance for deception-trained models is dismal, but also not relevant for the current state of play.
Incidents#
Warning about mass LLM-fuelled scams in the near-future. Core thesis: there used to be cheap-and-frequent scams of the Nigerian prince variety and expensive targeted scams a la The Sting, neither of which were a worry for the average person. LLMs collapse the two categories, making targeted scams much cheaper, thus putting swathes of people in the crosshairs. Mostly theoretical, but includes an anecdote and some studies on theoretical capabilities like AI spear-phishing clickthrough.
Opinion: Right now we’re protected by there not being enough scamming-inclined people/organizations picking up the AI banner, probably due to a mix of technical incompetence, lack of awareness and the vibe not having shifted yet.
Even-odds that societal defense mechanisms manage to shore up in time, either through cultural evolution or defense-side improvements.
Another claim of superpersuasion: AI beats human debate champions and professional canvassers, including cases where the humans had time to prepare and hints about the AI’s arguments, though they tied when the AI was limited to human typing speed/message length. AI also managed to raise ~3x more money for a charity than human experts.
Opinion: The old “gish gallop” trick still works, and is currently the main cognitive threat posed by AI.
🔦 GLM-5.2 deep dive#
Chinese open model, GLM-5.2, posts remarkable nominal scores on some benchmarks. It is likely the strongest open model now, for the next week or two. It is unlikely to be very usable or cost-efficient.
Opinion: a real step forward, but doomed to be forgotten in 2 weeks like all the others as the cracks are discovered and as it fails ~all real agentic tests. Recall:
Apr: Kimi K2.6 Beats Top US Models On Some Benchmarks
May: Glm-5.1 claims near opus level coding performance
June: Minimax M3 is out: First open model with frontier coding + 1M context.
-
#2 on LMArena-frontend-code by a solid margin, behind only Fable (i.e. a 1 month gap); but it’s #15 in other arenas.
We tentatively posit that LMArena is cheating (paid clickworkers using a GLM detector) rather than admit that it beats Opus at frontend code when it can’t even see. Even this kind of achievement is dubious.
-
Probably heavily benchmaxxed: Better than 5.1 on the famous and public(!) Pelican SVG test, but far worse at any variant.
-
Terrible AA-Omniscience score (roughly the same as GLM-5.1, same as Muse Spark). Really decent hallucination avoidance though.
-
0.7T
-
Big cost saving with a new cross-layer cache: IndexShare re-uses one indexer for every four attention layers.
-
Jake Boggs estimates GLM 5.2’S METR time horizon at 7.5 hours (with a Jake Capabilities Index equivalent to that of Sonnet 4.6) – implying a “US-China gap of ~5 months and is about the same as it was back in April”.
-
WeirdML puts GLM-5.2 max effort at about Gemini 3 Pro level (i.e. 7 month gap).
-
Very likely trained off Claude but that’s not a counterargument to capability.
-
Tiny ad hoc tests fail on tax processing, curriculum design and poetry evaluation.
We asked GLM-5.2 to score itself#
TL;DR:
-
GLM 5.2 sits roughly between Claude Sonnet 4.6 and Opus 4.6, i.e. ~4–5 months behind the US frontier (Fable 5, Opus 4.8). It is clearly capable.
-
The performance is not just benchmaxxing; several strong results come from new or private benchmarks Zhipu could not have accessed (e.g., AA-Briefcase, Harvey’s Legal, Runebench).
-
The architecture is the same as GLM-5 (a DeepSeek v3.2-like). For a 750B model, the results point to strong RL capabilities at Zhipu, with a clear path to decreasing costs alongside performance.
-
Context window is finally 1M tokens. API @ $1.4 / $0.26 / $4.4 per 1M input / cache-hit / output tokens.
Capabilities#
Pending evaluation: FrontierMath Tier 4 v2, DeepSWE, FrontierCode, MathArena, GBA Eval (Mechanize).
Benchmarks#
General#
-
On Artificial Analysis’s Intelligence Index v4.1, GLM 5.2 scores 51, 11 points higher than GLM-5.1 at the same parameter count, leading MiniMax-M3 (44), DeepSeek V4 Pro max (44), and Kimi K2.6 (43) among open weights. Artificial Analysis places it on the Pareto frontier of Intelligence vs Cost per Task: at ~$0.46 per task it has the lowest cost per task among models at its intelligence level, compared to GLM-5.1 ($0.25), Kimi K2.6 ($0.31), MiniMax-M3 ($0.18), and DeepSeek V4 Pro max ($0.05). On the first-party API it is priced at $1.4 / $4.4 / $0.26 per 1M input / output / cache-hit tokens, in line with GLM-5.1.
- Token-budget note: GLM 5.2 uses 43k output tokens per Intelligence Index task, up from GLM-5.1’s 26k and above MiniMax-M3 (24k), Kimi K2.6 (35k), and DeepSeek V4 Pro max (37k); part of the intelligence gain is therefore bought with longer reasoning.
-
On GDPval-AA v2, Artificial Analysis’s agentic real-world knowledge-work benchmark (Elo baselined to human performance at 1000, rotating panel of frontier-model judges, 250-turn limit), GLM 5.2 scores 1524, ahead of MiniMax-M3 (1418) and DeepSeek V4 Pro max (1328), and in-line with proprietary models including GPT-5.5 (xhigh reasoning).
-
On the Vals Index, Vals AI’s cross-domain composite, GLM 5.2 scores 65.02%, #5 overall, sitting just behind Claude Opus 4.7 (66.10% ± 1.27, released two months prior) and well behind Claude Fable 5 (75.15%); Vals AI notes a 13% gain over GLM 5.1. Open-weight SOTA (#1), but the frontier gap to Fable 5 is ~10 points.
-
On AA-Briefcase, Artificial Analysis’s new benchmark for long-horizon knowledge work (multi-week projects built by industry experts, thousands of input source files), GLM 5.2 (max) lands at Elo 1266, 4th, behind Claude Fable 5 (1587), Claude Opus 4.8 max (1356), and Opus 4.7. On cost, GLM 5.2 runs $2.40 per task, far below Fable 5 ($31), Opus 4.8 ($10.40), or GPT-5.5 xhigh ($3.68). The benchmark is new, so the result is unlikely to be contaminated by Zhipu overfitting the model to it.
Coding / Agents#
-
On Vibe Code Bench (end-to-end web-application builds with unrestricted terminal access), GLM 5.2 scores 63.96%, a 31% jump over GLM 5.1; it sits 4th overall, well behind Claude Fable 5 (90.35%), Opus 4.8 (82.72%), and Opus 4.7 (71.00%). Open-weight SOTA (#1).
-
On Terminal-Bench 2.1 (agentic command-line tasks), GLM 5.2 takes #1 in the open-weight category from Kimi K2.7 Code per Vals AI, an 11% improvement over GLM 5.1; Artificial Analysis’s own run scores it at 78% (+16 over GLM-5.1).
-
On FrontierSWE (Proximal AI’s maintainer-grade SWE benchmark; headline metric is average rank across tasks, lower = better), GLM 5.2 ranks #3 at 4.32, behind Claude Fable 5 (2.35) and Claude Opus 4.8 (4.24), and ahead of GPT-5.5 (4.56) and Opus 4.7 (5.76). The gap to Opus 4.8 is tiny (4.24 vs 4.32); Proximal calls it “the first model that closes the large gap between models from Anthropic / OpenAI and other providers.” In the best@5 ranking GLM 5.2 ranks behind only Claude Fable 5, and it takes the top score on the PCQM4Mv2 Molecular Gap Prediction task (an AI-research task).
-
On Code Arena WebDev (Arena.ai’s blind-vote coding leaderboard), GLM 5.2 (max) ranks #2, behind Claude Fable 5 and ~29 points ahead of Claude Opus 4.7 Thinking at #3; it is the #1 open-weight model on the board.
-
Jake Boggs estimates GLM 5.2’s METR-style time horizon at about 7.5 hours using his Coding Capability Index. For context, Claude Fable 5 sat at ~21.3h and Claude Opus 4.6 at ~13h in February.
Math#
- On ProofBench, Vals AI’s formal-math benchmark where models produce Lean 4 proofs for advanced undergraduate/graduate problems, GLM 5.2 is the first open-weight model to break 30.0%, scoring 11.0 points above the second-ranked model, outperforming Gemini 3.5 Flash, and trailing Claude Opus 4.5 by 1.0 point.
ML#
-
On WeirdML, a benchmark of nonstandard ML engineering tasks where models write PyTorch for novel datasets and iterate from execution/test feedback, GLM 5.2 (max) scores 70.1%, narrowly beating Gemini 3 Pro, a roughly 7-month-old model. For comparison, Claude Fable 5 reached 87.8% on the same benchmark. Håvard Ihle calls the result “a much higher score than I expected” and the model “a very solid model.”
- The (max) run averages ~22k output tokens versus ~12k for (high), buying a modest ~3-point gain; part of the score therefore scales with output-token budget. Non-thinking runs were still under way at posting.
-
On PostTrainBench, a benchmark where agents are given four small base LLMs, an H100, and 10 hours to post-train them, GLM 5.2 (Max, Claude Code) scores 34.29% ± 1.71%, #1, in a virtual dead-heat with Opus 4.8 (Max) at 34.08% ± 4.45% and Opus 4.8 (High) at 33.80%; Fable 5 (1M, Max) sits further back at 26.31%, Gemini 3 Pro at 18.12%, and the predecessor GLM 5 at 13.88%.
-
Caveat: GLM 5.2 only moved to #1 on Jun 17 after Opus 4.8 (Max) was re-run, its average dropping from a single 37.2% run to 34.1% once standard deviations were added, so GLM 5.2’s lead over Opus 4.8 sits inside the error bars; the result is effectively a nominal tie at the top.
-
GLM 5.2 also ran in the Claude Code scaffold, holding the agent harness roughly constant with the Claude models.
-
Games#
-
On Runebench, a benchmark where AI coding agents play RuneScape via a TypeScript SDK (scored on average XP per skill over 30 minutes), GLM 5.2 outscores Claude Opus 4.7 and GPT-5.4, models that were best in class only 2–3 months prior. The benchmark is quite new, so the result is mildly unlikely to be contaminated by Zhipu overfitting the model to it.
-
On OpusMagnumBench, a benchmark where agents play campaign puzzles in Opus Magnum, a Zachtronics engineering game about building alchemical machines from manipulator arms, GLM 5.2 scores 25.4% of the human world-record total, landing 4th just ahead of Gemini 3.5 Flash (24.5%) and Claude Opus 4.8 (22.8%) but far behind Claude Fable 5 (60.2%), GPT-5.5 (53.9%), and Gemini 3.1 Pro (31.1%).
- GLM 5.2 more than doubles its predecessor GLM 5.1 (10.8%). The benchmark is new, so the result is unlikely to be contaminated by Zhipu overfitting the model to it.
Sequential Decision Making#
- On KellyBench, a benchmark for long-horizon sequential decision making where human baselines of various sophistication outperform tested frontier models, GLM 5.2 is the new open-source SOTA but still loses ~30% on average over 5 runs. General Reasoning estimates GLM 5.2 is 6+ months behind the frontier based on KellyBench and internal quant evaluations.
STEM#
- On CritPt, a benchmark of unpublished research-level physics problems, GLM 5.2 (max) leads open weights by a wide margin, matching Claude Opus 4.8 at 20.9% and beating GPT-5.5, Gemini 3.1 Pro, and Opus 4.7; the next open model, DeepSeek V4 Pro, scores 12.9%. Only proprietary models score higher, with GPT-5.5 Pro topping the benchmark at 30.6%. The result is a 4.5x generational jump: GLM 5.1 scored 4.6% on CritPt ten weeks ago.
Miscellaneous#
-
On Harvey’s Legal Agent Benchmark, hosted live by Vals AI on a private held-out test set (real legal work product in an agentic setting, using shell and file-editing tools plus Word, Excel, and PowerPoint skills), GLM 5.2 scores 7.08% ± 2.00%, behind Claude Fable 5 (11.25%, or 10.4% without Opus 4.8 fallback) and Claude Opus 4.8 (9.58%), but ahead of Claude Sonnet 4.6 (5%); it is the #1 open-weight model, ahead of MiniMax M3 at 4.17%. Headline scores stay low because a task resolves only when every task-specific criterion passes, though criteria pass rates run far higher (Fable 5 90.5%, Opus 4.8 87.9%, Sonnet 86.7%). The benchmark is new, so results are unlikely to be contaminated by Zhipu overfitting the model to it.
-
On Finance Agent v2 (Vals AI), GLM 5.2 takes #1 open-weight; Vals AI gives no headline score in the announcement.
Vibe Reviews#
-
Teortaxes calls GLM 5.2 his “oh shit” moment, pointing to CritPt: “Opus 4.7-4.8 was a monumental jump on CritPt, along other jumps attributed to Fable assistance. They recover the exact value. This is a drastic acceleration of GLM timeline. I had them pegged as tier 2 Chinese lab. No more.”
-
Jeremy Howard calls GLM 5.2 “a marvel”, “at least as good as Opus 4.8 and GPT 5.5. It’s super fast, inexpensive, and not too verbose. It responds with nuance and judgement, & handles long context VERY well. I’ve never experienced an open weights model like this before.”
-
Qiaochu Yuan (QC) is the dissenting voice: “not impressed so far in conversation, flashes of something but it’s sloppy and willing to settle for college essay.”
Minor#
- Richard Ngo continues to reach for psychological explanations for the constant, apparently helpless pull towards RSI by people who really honestly profess not to want it.
- Completed hitpiece biopic on OpenAI dropped by Amazon studios. Not killed, just forced to shop it around again.
- VibeThinker 3B, a mostly failed attempt at a heavily post-trained Qwen for verifiable reasoning. Notes include benchmaxxing, inability to reason OOD and hope for a parallel paradigm to parameter scaling, ideally resulting in a model that can recognize when it doesn’t have the capacity to answer inquiries and fix that.
- GPT-5.4 assists Molecule.one’s Maria and a physical lab in driving a medicinal chemistry project from literature review to a validated experimental result. Besides research and experimental loop design assistance, the model helped by suggesting use of TEMPO as a mild oxidant to improve a difficult Chan–Lam coupling.
- US pressures ASML, claiming that one of their EUV machines is already in China. Might just be kicking up sand.
- Pentagon actually used Grok as well as Claude in Iran strike planning. I guess I have to approve of LLM councils…
- Preview of Omnii, a genome language model for biodefense. Current pathogen screening misses adversarial paraphrases of sequences, which Omnii detects from their embeddings. Allegedly also capable of predicting how specific mutations affect viral fitness.