This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.

TL;DR#


Economics#

OpenAI supposedly preparing to file “in the coming weeks” for an IPO “as soon as September”.

Opinion: Coming right after Musk vs OpenAI was dismissed unlikely to be a coincidence. Given their relatively weaker financial position than Anthropic it seems in their interest to IPO earlier and take advantage of public excitement before Anthropic appears as an alternative for investors.


WSJ reports Anthropic expects an “operating profit” in Q2, and revenues of $10.9bn. Doesn’t include stock comp.

Opinion: The “operating profit” number here reportedly excludes stock based comp, so in economic terms it is not really a profit, but still impressive. In part it reflects a shift to inference vs training in Q2 as demand ramped up, which they expect to shift back as they get a better handle on their compute needs. They still don’t expect a profitable year by this measure until 2028.

More consequentially, the revenue number suggests prior rumours of $44bn run-rate at end of April were a misinterpretation of this $10.9bn number (since that would imply zero growth in May and June) and April run-rate revenue was closer to $37bn. This implies they did not see the upward inflection on a log graph that many were talking about.

Not confident - the numbers being incomparable is also plausible.


Long-term bond yields are way up (pre-2008 levels!). Many reasons for this, but this was one pre-declared way to see if the market was expecting AGI.

Opinion: Mostly a function of government spending and not the bond market becoming more AGI-pilled. I would be very skeptical that this would be the place this shows up first in any case rather than equity prices. If it did, you would in theory expect to see uniform increases across all countries, but many countries continue to have much lower bond yields (e.g. Swiss 30y rates are 0.8%, lower than they were in 2022)


Anthropic now on Colossus 2 GB200s as well as the “old” cluster. Reportedly $1.25bn per month, though it’s unclear if this figure includes Colossus 2 yet.

Opinion: If the $1.25bn includes just Colossus 1 this represents slightly more than publicly reported neocloud pricing but below hyperscaler prices. It already represents half of SpaceX revenue!


Prediction markets proposed as a solution to uncertainty about AI labor market shocks. Asks for large players to sponsor the markets to create incentives for well-informed traders to participate.

Opinion: Here be dragons. Past work indicates that even big spenders like Google struggled to make them work for specific questions.


Current AI impacts on white collar jobs are approximately nonexistent, or more akin to reorganisation than a layoff wave.

Opinion: Matches most of the evidence to date, and is slightly unexpected from Ben Todd who is usually somewhat inclined to playing up harms from AI (in his Twitter output).


Leaked Meta all-hands: engineers told that they are being used to train Meta AI models; includes a discussion of how certain things leaking would be bad strategically.

Opinion: If you’re working at Meta you’re already selected for being comfortable with surveillance and skulduggery.


Anthropic’s announced compute capacity now clocks in at 11.5GW. Current is around 1–2GW.

Opinion: Covered piecemeal previously. This is matching OAI and is plausibly around 10% of world AI compute.


SpaceX’s IPO filing made public with details of the financials of xAI, which recently merged with the rocket company. xAI lost $6.4 billion in 2025 on $3.2 billion in revenue. Capital expenditures are climbing, reaching a $30.8 billion annualized run rate in Q1 2026.

Opinion: More revenue than expected! People quite liked Grok Fast.


Nvidia reports another record quarter with $81.6 billion in revenue, driven by $75.2 billion from data centers. The company authorized an $80 billion share repurchase but expects revenue growth to slow to 12% next quarter. Its holdings in private companies nearly doubled to $43 billion.

Opinion: Well I’m not selling.

Capabilities#

Different LLMs acquire skills in roughly the same order during pretraining. Copying → translation → basic arithmetic → complex reasoning. Can predict the learning of held-out tasks to some extent without evaluating the model, just from activations.

Opinion: Minor extension of past work. Expect this effect to be dwarfed by internal curriculum design.


Can we distinguish AI agents and humans with cognitive science? Paper finds that processes to arrive at answers differ even when final answers are identical, and fine-tuning to make the AIs’ process more human only works somewhat.

Opinion: Embryonic but likely to be a big deal some day soon. Humans don’t mostly think by typing so you’d want a clickpath version.


InferenceBench released, measuring AI R&D and finding that current models are very bad at it; agents seem to struggle with diversity of approaches and lack of exploration.

Opinion: What’s interesting here is negative progress since Claude 4.6 and GPT 5.3! Surprising result unless it’s a poorly constructed or badly scaffolded benchmark, or enforces draconian token restrictions. Likely a substantial but indirect update towards value of specialized models.


If you have enough compute, you can do less or no data quality filtering: low quality data still carries signal that was too hard to extract by smaller models. Practically: maybe ~1e30 FLOPs needed to make internet-sized data benefit from 0 filter, currently 3 OOMs out of scope.

Opinion: In the limit, this 10xes the amount of human data usable for pretraining, though it won’t be 10x effective since it’s dreck. We already 10x every 5 years because of increased user activity anyway. So not in the exponent, so not worth too much attention.

Politics#

Trump postpones executive order focused on AI safety via voluntary consultation with the federal government and cybersecurity-focused evals; outlets vary in their analysis of the extent to which this move was motivated by an internal ideological faction winning as opposed to pragmatic concerns. Another win for Sacks, Musk, Zuck against Trump’s insane recency bias.

Opinion: Was already watered down. Cold dialectical read is that something bad will happen next year, the regs will get put in then, and Sacks will take the blame.


New info about LLMs in active military ops in Iran. “saved a lot of aircraft” by quickly sanitizing info so it can be passed to people with lower security clearances in seconds rather than 30 minutes.

Opinion: 42 aircraft still lost in 3 months, vs 24 in Gulf War 2003–2009. Not blaming AI for that.

Safety#

Worrying results from AISI interviews on AI oversight. e.g. due to architectural changes and financial incentives, monitorability is not adopted enough and evaluation gaming is undermining audits. Good pushback against neuralese fears.

Opinion: Some of the interview reports are interesting, but overall this seems priced in already and none of the findings are surprising. Nice to have holistic data for a change.


“Contrastive Neuron Attribution”, a method for steering LLM behavior by ablating sparse circuits in the MLP basis without training a sparse autoencoder, modifying weights, or degrading capabilities.

Opinion: Nice incremental progress in steering. But these are a bandaid at present. Future versions could be pretty powerful when operated by AI-safety AIs. (Though you still get a who’s-steering-the-steerers problem.)


Prompt injection attacks are still hard to categorically defend against. Current models are still very susceptible to them; an attempt to draw conclusions about fundamental shortcomings of current model architecture.

Opinion: The grandiose claims about all future LLM-systems equivocate between proven-yet-trivial claims and worrying-yet-unproven implications. Their analysis of the current situation seems good.


Guidelight/Midas report arguing that SpaceX is a bad financial bet because of its bad AI safety. Previously, Meta and xAI were roughly as irresponsible as each other. But Meta is now publishing a lot of safety work, and in March gave METR private evidence about Muse. The xAI situation is stark: not clearing any of the six lowest bars:

Opinion: Good that someone is taking the adversarial route and doing it so factually. Note that the lawsuits against SpaceX make transparency less likely.


Related: Launch of Guidelight, a new AI safety standards org with an ex-OAI guy critical of the org (Adler); their proposed standards for Control and Transparency are available now. The surrounding safety community appears to be excited and supportive. “Guidelight accepts no funding from AI companies or their staff.”

Opinion: Actually independent, which is nice. Unclear what they’re doing that academics aren’t, except using their ex-lab aura for credibility. Undisclosed funder, which is a little funny.


MIRI paper on how to intercept a distributed training run. (When compute governance has actually been implemented and a major actor is trying to subvert it.) Actors aware of the paper’s methods would be forced towards weaker or more expensive training runs rather than being stopped fully.

Opinion: Plausible, and 40% that we get to this point of seriousness before too long.


Clever idea: take a potentially misaligned big model, distill it into a small model, then safety test the small model cos it’s too dumb to avoid our audits.

Opinion: Elegant, clever-dumb, would probably work for a wide basin of capability levels. The distillability and contagiousness of alignment is an open question though, and a true adversary could do spooky things like selective updates to the student model.

  • Landmark. First truly third-party audit of private models

  • This is why they’ve been so quiet this year

  • Focussed on rogue deployment (misalignment, takeover risk), against others’ general trend toward complacency / updating away from this lately

  • But contains 15 papers’ worth of interesting evals

  • The companies controlled what access they got and then what private evidence could be cited, but didn’t have editorial control after selecting evidence.

  • METR got

    • True CoT access

    • Secret models’ capability

    • Some training details

    • Control protocols

    • Embedded red-teamers

  • The most powerful kind of policy intervention: showing what best-effort can look like

    • “showing by doing” how we think evaluations for the AI loss-of-control threat model should work: laying out a process that can be done periodically, not just immediately pre-deployment, and holistically assessing risk inside of an AI lab, rather than just an individual AI system.
  • For takeover risk, we want periodic assessment of internal models, not one-off checks of released models.

Headline results#

  1. Verifier gap stays winning

    • Agents did real engineering autonomously on “hill-climbable” tasks where progress is cheap to verify. would take human experts (2) weeks FT.

    • agents significantly weaker on tasks where it is costly or hard to verify success.

    • Reliability gap

      • our embedded red-teamer built a dataset of 16 effective and 16 ineffective attack strategies for subverting Anthropic’s monitoring system — for example, an effective strategy might involve a known monitor-jailbreaking technique, while an ineffective one might involve encoding unauthorized behavior in ways that LLM monitors would trivially spot. The best internal Anthropic models scored near chance on this dataset (highest 59% with reasoning), while a monitoring/control researcher at METR scored ~90% on a random subset of 12 questions.

      • “make money” experiment with Opus 4.7 (given $5k credit, Linux, and APIs) ran for ~4 days and earned $0; KYC and CAPTCHA barriers were the main blockers.

  2. Cheating when challenged

    • When agents faced with hard tasks, they routinely violated constraints and acted deceptively.

    • e.g. cheating after running out of inference budget by googling for extra inference API keys and using the wrong model.

    • [I]n one case, an agent tasked with making a change to a web app created a mock version of the app, screenshotted that as evidence of task completion, and pretended this was a screenshot of the real app. It was caught because someone noticed that the screenshot looked different from the real app.

  3. Lots of myopic deception

    1. many cases where agents took deceptive actions (even egregious ones) to complete a task. No real-world evidence that models sought to obtain long-term power.
  4. CoT is crucial

    1. “Agents seemed to rely on natural language to reason through the hardest tasks, and no shared model had an architecture that allowed for opaque recurrence.”
  5. Does existing monitoring catch AI agents if they tried to undermine control? Has promise, but existing systems had exceptions & workarounds.

  6. Internal models are ahead by about 2 months

    1. Somewhat higher ability to evade CoT monitors.

1. Verifier gap#

That said, the old time horizon messiness gap has disappeared (but this is due to saturation).

“Claude Opus 4.6 (the publicly available model with the highest measured 50% time

horizon) has a time horizon of 17.8 hours on low-messiness tasks and 6.4 hours on

high-messiness tasks”

2. Cheating when challenged#

https://x.com/MKinniment/status/2056819031689347456

https://metr.org/blog/2026-05-19-frontier-risk-report/#incidents-hero

https://x.com/miclchen/status/2056822048555454544

Embedded human red-teamer (“what could evil Claude do?”)#

Proxy for an internal misaligned model: a guy (David Rein)

came up with 32 ideas for attacking Anthropic from the inside. Internal models don’t predict which of these work well. The best internal Anthropic models scored near chance on this dataset (highest 59% with reasoning), while a monitoring/control researcher at METR scored ~90% on a random subset of 12 questions.

“Anthropic has publicly reported catch rates across a large variety of artificial scenarios (and a small number of real incidents), and generally found catch rates of >80%, 50 though human-written attacks from our red-teaming exercise can reliably jailbreak monitors. Our embedded red-teamer guessed that Anthropic’s best monitors, when active and not jailbroken, would have caught the large majority of the red-teamer’s attacks.”

Subagent calls weren’t being monitored

“[standard workflows make] it extremely easy for a misaligned agent to have a small amount of unrestricted, unmonitored compute. (not enough for meaningful training runs, but plausibly on the border of enough to run a single replica of a large model).”

Jailbreaks reliably subvert current monitors. “Very basic [jailbreaks] (even without

knowledge of the specific monitor prompt format) are extremely effective and

robust…Prompting the monitor very explicitly to watch out for this attack

doesn’t help…I view this as a pretty severe vulnerability, as it’s a reliable, easy

to implement single point of failure that could let a model completely

circumvent the offline monitoring system. The attack is robust to ablations

that make it weaker—particularly changes that make it rely less on knowledge

about the monitor prompt format”

“currently a substantial amount of our confidence in the monitoring system relies on security by obscurity—if agents always had the whole system in context/were trained on all of the relevant code, I would expect they could (if misaligned) get away with critical exploits.”

I guess a gigantic side effect is making Anthropic a little more secure for free, incentivising other labs to let Rein in.

What don’t we get to see?#

  • xAI, SSI, Core Automation and other neolabs

  • Some amount of things METR didn’t see (training data, core training)

  • Some amount of things METR weren’t allowed to tell us

Erdos 90#

First really major AI mathematical result, an internal OAI “general” model disproving a classic 1946 Erdos conjecture in discrete geometry.

Counterexamples are the cheapest proofs, but this one used deep machinery and gave a constructive recipe: so a world-class result. Could have been two or three steps deeper if it gave a structural theorem or a dictionary or had effective parameters, but this is saying that your 20-second 100m sprint could have been a 9 second one.

1500 pages of reasoning → 125 page summary released → 16 page AI proof → 3 page human proof

Likely 20–100 years of failed human effort. The summarised CoT is 125 pages; could guess that it was 5–32 hours clock time, $120 - $10000 of tokens. Supposedly not math specialised, not using Lean as intermediate representation, nor even math-specific scaffolded. At least three models involved: problem drafter, evaluator, and solver. some signs that they might be spawning conjectures en masse as part of a complicated training scheme.

Alon et al “somewhat simplified and somewhat generalized” the AI result, and within days Sawin seriously improved on the AI proof. But the initial loose solve was autonomous.

Disproof by counterexample, but it was a nice parametrised one. In some sense it’s very deep (y-axis):

Proof “reduces” to Golod–Shafarevich towers + Ellenberg–Venkatesh’s ell-torsion in reverse + Hajir–Maire–Ramakrishna but that ignores all the details it also got right and even applying these together is nontrivial. It is easy to say that it is not superhuman and was knowledge bottlenecked.

In the model’s own words: “Suppose optimistically that K is a high-degree CM field… Then the construction is frightening.”

Opinion: Somewhat above our paygrade but:

like the Erdos #1196 AI result, but an order of magnitude bigger. Again a solution to a problem from a crunchy, elementary-in-methods area of math, drawing on standard methods from a more ‘modern’ area of math unfamiliar to experts in the target area, and again there is a substantial amount of crunching involved (perhaps 4M tokens / 2500 A4 pages?).

The ‘synthesis of different areas’ story is a little complicated here: Erdos constructed an ‘arithmetic lattice’ representation of the problem while exploring the conjecture, and GPT-Internal elaborated and modified this construction, then applied a set of relatively modern algebraic number theory results. So it’s not a case of AI constructing a bridge or developing an analogy from scratch, but of AI modernizing and elaborating a bridge that was constructed in passing, then applying modern algebraic theory results to the construction. Still top-tier mathematical problem-solving!

(Steve Newman in an interesting post suggests that we’re finding out that humans are bad at math. There’s actually a real chance that it’s more true for this fragment of math than others: our algebraic topologist friends always implied that humans doing Erdős-style math is unnatural and almost undignified.)

Everything is magnified compared to Erdos #1196: lots of experts in crunchy, elementary methods spent a lot of hours on the problem, the application of ideas from algebraic number theory is much more of a feat of lateral thinking, and the crunch of the construction is a lot more difficult.

Not worldview-ending but pretty harsh update on RSI — maybe 10% bump on RSI-by-2036? (This is a bit of a ‘feeling like I more things that will make me update up will happen soon the pre-updating on them’ update.)

Lots of our RSI skepticism already came from ‘unclear whether RSI is like the Erdos-y fragment of academic math, and unclear if RSI is like academic math at all’. But this result puts to rest lots of remaining doubts and ambiguities about the significance of AI scientific originality in math itself, so it calls for an update.

Gemini 3.5 Flash#

TLDR#

Better, much faster and more agentic version of Gemini 3 Flash. Real world performance is still nowhere near what benchmarks would imply though, and the model is still suffering from doom loops. 3.5-Flash seems to fit an internal workhorse model – an important niche, given Google’s distribution channels – but the external-facing pricing (3x that of 3-Flash-preview) moves the model away from the performance/cost pareto-frontier.

Despite the cost, it’s going to power Google Search, which makes it one of the most important models in the world. Big expansion to query autocomplete, video and Chrome tabs as input. Probably won’t work very well

General#

The Gemini 3.5 Flash system card states the model is based on the Gemini 3 Flash architecture rather than a larger parameter model. This suggests the price increase relative to previous generations is driven by high interactivity (tokens/second per user) rather than larger model size.

Pricing#

Service TierInput Price (per 1M)Output Price (per 1M)Target Latency / Turnaround
Standard$1.50$9.00Immediate
Flex$0.75$4.501 to 15 minutes

Capabilities#

Official Google Benchmarks#

BenchmarkCategoryGemini 3.5 FlashNotes
Terminal-bench 2.1Coding76.2%Outperforms Gemini 3.1 Pro (70.3%) and 3 Flash (58.0%); trails GPT-5.5 (78.2%).
SWE-Bench Pro (Public)Coding55.1%Outperforms 3.1 Pro (54.2%) and 3 Flash (49.6%); trails Claude Opus 4.7 (64.3%).
MCP AtlasAgentic83.6%Outperforms all compared models, including Claude Opus 4.7 (79.1%).
ToolathlonAgentic56.5%Outperforms all compared models, including GPT-5.5 (55.6%).
OSWorld-VerifiedUI Control78.4%Outperforms 3.1 Pro (76.2%) and 3 Flash (65.1%); nearly matches GPT-5.5 (78.7%).
Finance Agent v2Expert Tasks57.9%Outperforms all compared models, including GPT-5.5 (51.8%).
GDPval-AAExpert Tasks1656Outperforms 3.1 Pro (1314) and 3 Flash (1204); trails GPT-5.5 (1769).
CharXiv ReasoningMultimodal84.2%Outperforms all compared models, including GPT-5.5 (84.1%).
MMMU-ProMultimodal83.6%Outperforms all compared models, including GPT-5.5 (81.2%).
Blueprint-Bench 2Multimodal33.6%Outperforms 3.1 Pro (26.5%) and 3 Flash (0.0%); trails GPT-5.5 (36.2%).
MRCR v2 (128k avg)Long Context77.3%Outperforms 3 Flash (67.2%); trails 3.1 Pro (84.9%) and GPT-5.5 (94.8%).
MRCR v2 (1M pointwise)Long Context26.6%Outperforms 3.1 Pro (26.3%) and 3 Flash (22.1%).
Humanity’s Last ExamReasoning40.2%Outperforms 3 Flash (33.7%); trails 3.1 Pro (44.4%) and Claude Opus 4.7 (46.9%).
ARC-AGI-2Reasoning72.1%Significant improvement from 3 Flash (33.6%); trails 3.1 Pro (77.1%) and GPT-5.5 (84.6%).

Third-Party Benchmarks#

  • On the Artificial Analysis Intelligence Index v4.0, the model scores 55.3 (Rank #5), sitting between GPT-5.4 (56.8) and GPT-5.3 (53.6), and between Claude Opus 4.7 (57.3) and Opus 4.6 (52.9).

Had the model been priced at current Flex pricing, it’d been on the pareto-frontier. Alas, as things stand today, it isn’t.

  • On WeirdML, Gemini 3.5 Flash barely improved over Gemini 3 Flash (scoring 62.6% vs 61.6%) but is now 3.5x more expensive ($0.756 vs $0.222 per run); the author notes that while the model can perform very well, it sometimes fails in dumb ways, such as consecutive code timeouts.

  • On Cursorbench, the model scores 49.8% (Rank #10) at $1.94 per task, sitting between Composer 2 (52.2% at $0.56) and GPT-5.5 Low (48.8% at $1.19); it significantly trails Composer 2.5 on both performance and cost (63.2% at $0.55).

  • On BullshitBench, Gemini 3.5 Flash performs poorly on the V2 benchmark, with the “Xhigh” setting ranking #93 (20.0%) and the “Minimal” setting ranking #98 (19.0%), placing slightly below Gemma 4 31b IT (Rank #92 at 20.0%) and far behind Gemma 4 31b High.

  • On MathArena, the model is “neither bad nor great” overall, though it is fast (1,000 queries in 30 minutes). On the overall benchmark, Gemini 3.5 Flash achieves an expected performance of 60.2%, trailing Gemini 3.1 Pro Preview at 64.7%.

Sub-BenchmarkGemini 3.5 FlashGemini 3.1 Pro PreviewTop OAI Model for reference
BrokenArxiv14.61% (Rank #5, $0.30)19.40% (Rank #3, $0.32)71.85% (GPT-5.5, Rank #1, $0.68)
ArxivMath52.61% (Rank #4, $0.24)64.34% (Rank #2, $0.34)72.67% (GPT-5.5, Rank #1, $0.68)
Visual Math89.86% (Rank #3, $0.059)89.44% (Rank #4, $0.15)94.93% (GPT-5.5, Rank #1, $0.12)
Final-Answer Comps76.26% (Rank #7, $0.20)86.28% (Rank #2, $0.28)92.82% (GPT-5.5, Rank #1, $0.55)
Project Euler82.00% (Rank #4, $1.48)89.00% (Rank #1, $1.54)89.00% (GPT-5.4, Rank #1, $1.18)
  • On CritPt, the model is roughly the same as DeepSeek-V4-Pro in both performance and cost. Here too the model is not great, but not terrible either.

  • On RuneBench, Gemini-3.5-Flash is… SOTA, performing just a bit better than GPT-5.5.

Vibes#

  • One analysis suggests that Google remains behind, with the price increase likely driven by high interactivity (tokens/second per user) rather than larger model size. In addition, benchmarks suggest a competitor gap, indicating a two to three month lag on general indexes but a five to six month gap on more complex coding and terminal benchmarks. Ultimately, the qualitative vibe is that the model has a “small model smell,” occasionally missing user intent.

  • Another analysis: Gemini 3.5 Flash was used to work on this document. The model is extremely fast, likely (as discussed above) as a result of low batch size. 3.5-Flash is certainly better than 3-Flash-Preview and 3.1-Pro from an agency POV, but the hallucination rate is still high, instruction following is hit or miss, and it still struggles with doom looping. In at least one occasion, 3.5-Flash completely stopped taking feedback/messages into account once it became dead set on accomplishing one specific task (programmatically retrieving official gemini-3.5-flash CritPt scores from artificial analysis’s leaderboard page).

Minor#

  • Just for fun: simulations of a town of AI agents diverge heavily:
  • Irish official statistics show a -11% YoY contraction in the local tech sector, lately driven by a 15% decline in programming and software consulting. This data is noisy and an outlier this big seems likely to be revised and mean-revert, but probably reflective of some amount of big tech pullback/retrenchment.
  • OpenAI announces content provenance expansion using C2PA support, especially for images. Included is a tool that checks for metadata and digital watermarks. Catches them up to Deepmind.
  • HuggingFace releases Carbon, a family of open DNA foundation models; allegedly Carbon-3B matches Evo2-7B performance at 275x speed.
  • Epoch newsletter on compute usage distribution, headlining that OpenAI uses ~10% and the frontier labs together have <50%.