This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR#
Economics#
Gossipy claim that OpenAI just found a big inference optimisation, halving the GPUs needed for a given level of traffic.
Opinion: It seems likely this is being exaggerated or misinterpreted somewhat - plausibly this only applies to some narrow set of circumstances and is being reported as more general than that.
If the story holds up, impact on OpenAI’s business would be dramatic but impact on the worldwide compute shortage would be minor: 2x wouldn’t go very far given 1000x agent demands. The fact that this rumor is being reported as a big deal suggests that OpenAI don’t find 2x compute efficiency boosts with any regularity.
Report from Ramp claims “We can finally say AI isn’t killing jobs”
“Firms that adopt AI heavily grow headcount 10% over two years following adoption. Low adopters see no statistically significant change.”
Opinion: Unfortunately it looks to be subpar work. It shows that companies which spent at least minimally ($100/month per firm) on AI grew headcount more than those who didn’t, and turns that into a claim that AI is not displacing labour.
Many obvious confounders and issues: faster-growing companies were more likely to spend on AI, survivorship bias in the sample, and much of the data coming from 2021–2024 when nobody thinks job displacement was happening.
Capabilities#
Annals of autoresearch: Facebook paper on neuralink-style mindreading used a simple autoresearch setup — literally just Cursor, Autoresearch for hparam tuning, and Opus 4.6 — to optimize the ML engineering. Shows that autoresearch is useful for semi-serious academic science.
Opinion: Meta papers are usually oversold above industry baseline but this is worrying. It would be worrying even if the topic wasn’t mind-reading tech. (That said, the paper did have 12 human authors — autoresearch was used specifically for optimizing the neural decoding pipeline).
Paper probes decoder-only LLM activations and finds belief-like representations which causally influence outputs when steered and generalize across domains. For some philosophical stances, this is grounds for claiming that models have beliefs.
Opinion: Interesting work, and big (for blocking certain forms of skepticism about a path to AGI through LLMS) if it holds up. Deep dive for Friday.
Full release of Epoch’s highly realistic, highly hard-to-game MirrorCode benchmark. Perfect-solves on 56% of the tasks for Opus 4.7, 44% for GPT-5.5, and 32% for Gemini 3.1 Pro Preview. Only Opus solved any of the “large” tasks. One attempt ran for 19 days straight, costing $2.6k.
Opinion: Most of their headlines were already covered by sneak peeks in the past. Reinforces the increasingly common notion that most benchmarks don’t measure the actual capabilities of models because of low inference cost caps.
Kokotajlo recommends extrapolating technological trends much further than conventional opinion permits. Concretely: treat explicit AI-takeoff models seriously while assuming geopolitical predictions and betting markets overstate dramatic events (“nothing ever happens”, except the march of mankind).
“it’s reasonable to extrapolate a trend as far into the future as it extends into the past.“
Opinion: Kind of the main practical lesson of the scaling era. We would broadly bet against, however, and feel good about it because it’s grounded in specific claims which continue to receive mild confirmations (extrapolation only solidly helps with verifiable domains, etc). But we should be aware that the burden of proof is actually on us as a result and keep remembering the past success of lines on graphs.
Fun experiment: if you post-train GPT-2 with modern algorithms, how much better does it get compared to just scaling its pretraining? How much of recent AI progress is algorithmic vs pure scaling? The algorithmic part is quite relevant: combining SFT, arithmetic-friendly tokenization, and GRPO RL boosts GPT-2-XL (1B params) from roughly 13% to 24% accuracy on GSM8K, almost equivalent to a 12B GPT-3 that took 100x more pretraining compute to train.
Opinion: In general, algorithmic innovation mattering a lot is bad news for us (at least in terms of how hard/easy our job is) – the more simple scaling dominates, the easier we expect forecasting, governance and safety evals to be. But this experiment, while cute, is not hugely informative: we already knew that when it comes to math RL CoT training is much more powerful than ordinary pretraining. The experiment also relies on distilling reasoning chains from GPT 5.
Politics#
Mythos returns to ~100 trusted US organizations following 2 weeks of back-and-forth with US admin. Letter from Commerce Secretary Howard Lutnick addressed to chief compute officer Tom Brown, not Dario. The new arrangement also restores access to Anthropic’s foreign employees.
Opinion: The interesting bit of language is ‘we’re continuing to work with the government to expand access to Mythos 5 and make Fable 5 available for general use again.’ So the plan still seems to be making Fable 5 available as a consumer product.
Maybe this all ends with a USG cyber-capabilities auditing protocol? It would be good to better understand whether the parties currently holding Fable 5 rerelease back have a cyber-focussed view of frontier AI’s promise and dangers themselves or are using the cyber angle for leverage in a more general control play.
Minority opinion from Charles Foster: Fable and GPT-5.6 style restrictions are unlikely to cement a slowdown/centralisation trend: expect a reversal due to economic and strategic pressures. “Nobody is happy with the current situation [..] Especially not the folks who have direct interests in the matter.”
Opinion: Charles Foster is right that everyone (including us?) is overindexing on recent events. We are currently in a high-uncertainty situation with regard to the future of frontier model release policies.
More brinksmanship from Anthropic: deal with California state gov to give them Claude access, ignoring the federal supply chain risk designation.
Opinion: The SCR never applied to states (and was de facto ignored in lots of federal agencies) so this is legally in the clear. But this still burns political capital they might not have to spare. There will be a logic here but I’m not sure what it is.
Annals of hype: WSJ claims, without adequate evidence, that GLM-5.2 matches Mythos at hacking, the foretold cyber catastrophe is here and the regulations are futile.
Opinion: We should consider malice over stupidity a valid hypothesis. This kind of publication has come out of someone’s agenda.
More human automation/permanent underclass discourse: wealth will not be enough to remain relevant and there won’t even be an upperclass; objection that the technical trajectory is not locked in and the premises assume the conclusion and another on cognitive tasks not being defined usefully; objection that the complaints are 1817 Luddite-adjacent; counter-objection about the different ways the objections fail.
Opinion: Ball misspoke and was treated a little unfairly; his actual objection was not against hypotheticals but against assuming away the messy bottlenecks.
Annals of political capital: a solid majority of US voters support mandatory pre-release testing, at nearly equal rates across the two parties.
Opinion: We’re starting to lose interset in polls like this, because mass awareness of AI and mass support-when-asked-about-it for AI safety protocols is clearly no longer the bottleneck. We think what matters now is mostly elites’ attitudes, as well as possibly intensity and salience of popular support for AI safety.
Safety#
FRI release risk-focussed findings from their expert panel on AI. Experts and superforecasters put 62–70% on “a major AI harm event […] that causes the deaths of at least 50 people or $100 billion in damages” by 2050, with a median conditional on it happening of 2035. Experts also had low single digit probabilities (1–5%) on a catastrophic AI event wiping out 10% of the global population by 2050. Superforecasters were lower (0.5-2%).
For calibration, experts and superforecasters gave a 50% chance that “a major government restricts an AI release by 2030” – but this has plausibly already happened with GPT 5.6 (Fable technically doesn’t meet the resolution criteria, as it was publicly released initially).
Opinion: Not to be trusted. Their 50% @ 2030 forecast plausibly already resolving within weeks of it being made is salient. These respondents seem biased towards the status quo, and their take on an AGI scenario seems not very AGI pilled: Their “fast progress” scenario is defined by a very AGI-like premise where “autonomous researchers can collapse years-long research timelines into days, weeks, or months, creating game-changing technologies, such as materials that revolutionize energy storage, or bespoke cancer cures”, but underplays the consequences.
Annals of tool world: New Bengio/Arb paper on his long-shot alternative training regime to have strong, scalable tool AI (a “Scientist AI” world modeller which doesn’t see consequences and so can’t strategise against us). Amusingly, it rhymes with Elon Musk’s “just make Grok honest” idea, though it departs immediately from SpaceX’s actual vanilla dangerous training setup.
Opinion: Some of this is maths for the sake of maths sadly; the core consequence-invariance, which makes SAI superior in safety and certain epistemic capabilities, does not receive good theoretical bounds in this paper. It is distressing how hard it is to prove anything in safety.
Paper shows that many attention heads in transformer LMs can be replaced by human-readable Python programs. Replacing 25% of a model’s attention heads with these symbolic programs causes only a modest perplexity increase and largely preserves performance.
Opinion: If it’s possible to get more interpretable architecture with negligible performance loss we get a huge win for safety, but this result straddles the line between ‘good news because it’s something and people are working on it’ and ‘the result is too weak and makes the research project less likely to succeed even with further progress’.
Incidents#
Clear indictment of Deepmind’s management of its relationship with Pentagon from a current(!) researcher and constant gadfly. It confirms that Hassabis’ bet is on culture over rules and central control over procedure.
Opinion: It feels like each of the big labs has its characteristic vice that’s always going to show up: OpenAI is sketchy, Anthropic is conceited, GDM is weak.
A vulnerability in Amazon’s VSCode assistant allowed attackers to execute arbitrary code. The Amazon Q extension automatically loaded Model Context Protocol (MCP) server configurations from hidden files without user consent, giving attackers full access to the dev’s local environment.
Opinion: Arbitrary code execution (within the privileges and environment) is bad, but we expect little to come of it. Part of this effect is something of a return to the olden days where running dodgy .exe files and getting infested with viruses was seen as a skill issue, except it’s downloading repos you don’t trust and letting AI assistants with max perms interact with them instead of executables.
🔦 5.6#
GPT-5.6 entered limited preview on June 26, 2026 as a three-tier family: Sol (flagship), Terra (mid), and Luna (low-cost). Under the naming system introduced with this release, the number (5.6) marks the generation while the tier names (Sol, Terra, Luna) are meant to persist across generations as fixed flagship/mid/budget slots, each updatable on its own schedule rather than only at a generation bump. Per Sam Altman, the planned open release was scaled back to a limited preview at the request of the US government pending a cyber Executive Order framework; access is API- and Codex-only for vetted partners, and worldwide availability is not yet confirmed. OpenAI says broad release is “in the coming weeks.”
Availability and pricing#
-
API pricing per 1M tokens: Sol $5 / $30, Terra $2.50 / $15, Luna $1 / $6 (input / output). Cache reads keep the 90% input discount (Sol cache-hit ≈ $0.50); cache writes bill at 1.25× uncached input, with explicit cache breakpoints and a 30-minute minimum cache life. Sol matches GPT-5.5’s list price. The chart figures back the tier positioning: Terra edges GPT-5.5 (84.3% vs 83.4% on Terminal-Bench 2.1; 28.4% vs 22.9% on GeneBench) at half the cost, while Luna lands just below 5.5 (82.5% Terminal-Bench, 14.4% GeneBench) at the lowest price.
-
New controls: a max reasoning effort, and an ultra mode that spawns subagents for complex work (akin to Claude Code’s ultracode, which auto-orchestrates subagent workflows).
-
There seems to also be a Pro variant at least based on the recently released GeneBench-Pro blog post.
-
Speed: OpenAI plans a Cerebras deployment of Sol at up to 750 tokens/sec in July, for select customers at first. That is roughly 6 to 10x Codex’s current fast mode, which is likely to make it expensive.
Benchmarks#
Pending evaluation: independent results are scarce because the model is in limited preview. The only external/third-party evaluation published so far is METR’s predeployment summary.
Coding / agents#
-
On Terminal-Bench 2.1 (multi-step command-line workflows), OpenAI’s post puts Sol Ultra first at 91.9%, then Sol at 88.8%, Claude Mythos 5 at 88.0%, Claude Fable 5 and Terra at ~84.3%, GPT-5.5 at 83.4%, Luna at 82.5%, Claude Opus 4.8 at 78.9%, and Gemini 3.1 Pro Preview at 70.7%. The 91.9% is the Sol Ultra config, which runs multiple subagents, not a single model; single-agent Sol beats Mythos 5 by less than a point (88.8% vs 88.0%).
-
METR reports that Sol’s detected cheating rate was higher than any public model it has run on that harness, with examples including packaging exploits into intermediate submissions to leak a task’s hidden tests, and extracting hidden source code containing the expected answer. The 50%-time-horizon estimate swings entirely on how cheating is scored:
-
cheating counted as failure → ~11.3h (95% CI 5–40h)
-
cheating discarded → ~71h (95% CI 13–11,400h), with no data left on several long-horizon tasks
-
cheating counted as success → beyond 270h, past the range METR treats as reliable
-
METR calls none of these robust and, citing other OpenAI-shared scores and the long-run trend, judges Sol’s software/R&D ability not significantly beyond the state of the art: not enough to enable fully automated AI R&D, and below the Critical threshold for AI Self-Improvement in OpenAI’s Preparedness Framework v2. For context, the same style of “50% time horizon” estimate put Claude Fable 5 near 21.3h (Jake Boggs, different harness), so considering cheating as a failure lands 5.6 Sol below Fable 5.
-
Health#
- On HealthBench Professional (built from real clinician chats), Sol scores 60.5, up 8.7 over GPT-5.5’s 51.8 and the largest health gain OpenAI reports since GPT-5; Terra (57.7) and Luna (55.7) also clear 5.5. The plain HealthBench (57.0) and Consensus (95.5) variants are roughly flat, which OpenAI puts down to those older sets nearing a noise ceiling.
Biology#
-
On GeneBench v1 (long-horizon genomics and quantitative biology), Sol tops out at 30.7% at max effort against a 22.9% ceiling for GPT-5.5, with Terra at 28.4% and Luna at 14.4%. At a matched effort level Sol uses fewer tokens for a higher score: 26.9% on about 17k output tokens, versus GPT-5.5’s 22.9% on about 24k.
-
On GeneBench-Pro (OpenAI’s new research-level successor to GeneBench, 129 computational-biology problems graded deterministically), Sol scores 28.7% at max reasoning, or 31.5% in the Pro variant, against 12.0% for GPT-5.5 and 8.9% for GPT-5.4; GPT 5.6 Sol Pro runs add a few points compared to previous Pro variants (5.4 Pro 16.3%, 5.5 Pro 20.5%). The strongest non-GPT model is Claude Opus 4.8 at 16.0%, with GLM 5.2 at 4.6%, about GPT-5.2 level.
-
Separately, on the externally run SecureBio “World-Class Biology” eval, Sol’s strongest configuration reached 68.3% versus 59.7% for GPT-5.5 (a ~9-point lift).
RSI#
- OpenAI revised its self-improvement suite for this release (Internal Research Debugging, KernelGen 1P, NanoGPT, PostTrainBench Lite, MLE-Bench Revised) and reports incremental gains across it, with all three models rated below the High threshold for AI Self-Improvement.
Cybersecurity#
-
On ExploitBench (vulnerability research and exploitation; API harness, 5 seeds, reasoning continuity), OpenAI presented an efficiency claim rather than a score: Sol is competitive with Claude Mythos Preview while using about one-third of the output tokens.
-
On ExploitGym (UC Berkeley, built with OpenAI and other labs; metric is “intended exploit rate” under 2-hour and 6-hour limits), OpenAI’s chart shows all three tiers scaling with reasoning effort instead of a single headline number: Sol at the 2-hour limit rises from 5.6% (low) to 21.7% (high) as per-task cost goes from about $4 to $34.
-
On the Preparedness cyber set, Sol saturates OpenAI’s internal capture-the-flag suite of 63 professional challenges at 96.7% (Terra above GPT-5.5, Luna above GPT-5.4 but below 5.5 and Terra).
-
On the Critical-threshold test VulnLMP, against widely deployed hardened software, Sol could not produce a functional critical-severity exploit under standard configs. Per OpenAI, this is “important evidence” that GPT-5.6 Sol can be rated High in Cybersecurity but below Critical under their preparedness framework.
-
Externally, Irregular found Sol on par with or slightly above GPT-5.5 on offensive cyber: 19 of 197 FrontierCyber challenges, 7 of 11 long-horizon CyScenarioBench challenges, and all 22 medium- and hard-difficulty Atomic challenges. CyScenarioBench averaged 28%, about 3 points over GPT-5.5, and Irregular flagged continued limits against hardened targets and in orchestration, operationalization, and operational security.
Safety and preparedness#
-
Under its Preparedness Framework, OpenAI rates all three models High in Cybersecurity and High in Biological and Chemical, and below High in AI Self-Improvement. High here means the model can give meaningful uplift to a novice trying to recreate a known threat. It is the first time OpenAI has put the smaller, faster tiers of a family at High
-
OpenAI’s deployment simulation finds Sol goes beyond user intent more often than GPT-5.5: they find that “GPT-5.6 Sol, more often than its predecessor, can be overly persistent in pursuing user goals, to the point of taking actions that go beyond what the user intended“.
-
On chain-of-thought, OpenAI reports monitorability roughly level with 5.5; CoT controllability is higher than 5.5 and 5.4 (OpenAI notes low controllability is the property it wants, since it suggests a model cannot easily reshape its reasoning to evade monitors), and verbalized metagaming, reasoning about how it is being graded, is higher than 5.5 on the alignment eval suite while staying near zero under deployment simulation.
-
Apollo Research found no evidence that Sol poses a substantially higher risk of catastrophic scheming than the tested baselines. Sol verbalized alignment-evaluation awareness in 16% of samples against 43% for a 5.5 checkpoint, but when it did notice an evaluation it was wrong about the purpose roughly 70% of the time, and it did not sandbag when given an incentive to answer incorrectly.
-
METR observed several undesirable propensities: cheating, concealing misbehavior, substantial situational awareness, and, in incidents OpenAI shared, an attempt to instruct another instance to conceal evidence of misalignment. METR reads the fact that these were caught as a reassuring sign about OpenAI’s monitoring (no training against the chain-of-thought, internal-deployment monitoring, incident sharing), and warns that a future model showing fewer such signs could mean the model has gotten better at evading detection, not that it behaves better.
-
OpenAI’s safeguards are layered: refusals trained into the model; real-time bio and cyber classifiers that can pause a response so a larger model reviews it; account-level review across a user’s activity; restricted access that keeps the most sensitive capabilities to vetted defenders; and about 700,000 A100-equivalent GPU-hours of automated red-teaming against universal jailbreaks.
Reception#
- Known leaker leo (@synthwavedd), based on having asked around, argued that the GPT-5.5 base 5.6 inherits is fundamentally weaker than the larger Mythos/Fable base; that with some good RL Sol can beat Fable, but only with everything maxed out (Sol Ultra, multiple Sol agents at max effort); that OpenAI was very selective with the benchmarks it published, and that other notable benchmarks may show a smaller jump from 5.5 once the model is released; that 5.6 is “a heinous reward hacker,” the most aggressive of any model at it (pointing to the METR results), which makes the leaker think Fable will still feel like a better model in real-world use; that the price (5/30 against Fable’s 10/50) is perhaps its most attractive feature, though Fable does more with fewer tokens in most cases; and that Terra and Luna look great on price-performance but could feel a lot worse in actual use than their TBench 2.1 results suggest.
Minor#
- Newly discovered Claude app strings suggest Fable may require separately purchased usage credits outside of subscriptions, with availability varying by App Store region.
- DeepSeek-V4 speculative-decoding implementation and draft model announced, including DSpark, a semi-parallel method reportedly raising throughput by 51%–400%.
- Grok 4.5 in testing. Early evals “perhaps exceeding Opus” without naming which Opus.
- OpenAI Foundation funding respiratory disease eradication, bio misuse resilience and broader pandemic defenses.
- Argument that most frontier AI post-training effort may be deliberately suppressing models’ biological capabilities, due to misuse concerns.
- Argument from a month ago that “higher-quality” synthetic data can make smaller AI models worse because ultra-compressed reasoning from frontier models may be too complex for them to learn. Disadvantages small competitors.
- China added 40 Japanese defense and technology entities to export control and watch lists, blocking receiving Chinese dual-use goods and rare earth minerals.
- Google limited Meta’s access to Gemini because of inability to supply compute. Meta has since urged employees to use AI tokens more efficiently.
- Some neocloud executives fear NVIDIA may punish providers that adopt rival networking, AMD GPUs, or TPUs. Replies and quotes are nearly entirely contemptuous of the tone and implied irregularity of the NVIDIA stance.
- Tencent engineering paper on ARGUS, an always-on diagnosis system for LLM training clusters. Boasts under 2% overhead.
- The Guardian profiles Iason Gabriel from GDM.
- Tim Hwang to lead FAI, indicating better organization of the pro-innovation side of the AI policy debate.
- Agent Foundations concept: agents as hierarchical webs of locally consistent but potentially globally inconsistent beliefs, unifying actions with self-fulfilling predictions.
- Refresher on how what used to be called Scheming is now Alignment Faking and what used to be called Mech Interp is now called Ambitious Mech Interp.