This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.

TL;DR#


Economics#

FRI releases a new survey of economists and forecasters on AI economic effects.


  • Economists have (finally) converged on expecting extremely fast AI progress as much as AI researchers do

~No forecaster or economist forecasts a long-run GDP growth trend above 5%, and few economists go above 3.5%, despite expecting significant progress on capabilities. But the 3 Epoch participants all said >10%.

They expect a labour force participation rate of 55% in 2050 (vs 80% now) even in the fast progress scenario.

Forecasters expect even if we automate most labor and wait 20 years, GDP will only increase by 45% total in 2050. This is absurd on the face of it. Three ways to make sense of it: 1) collapse in aggregate demand (war, unemployment), 2) increased volatility suppresses investment, 3) strong expectation of massive regulation. But overall we continue to think that this is a basic failure of imagination / professional deformation towards deflating anything wild-sounding and the “null prior”. e.g. Good imaginative thoughts on what happens when AIs displace traditional corporations.

Opinion: The last FRI survey was subpar but there’s lots of great people on this one. Economists are systematically bearish on this stuff, we think half for bad reasons. But this is changing: summarise these results as “AI-pilled at last” but still not singularity-pilled. We didn’t update on any of this, but it’s still useful to know how unusual our bubble is.


Post arguing that inference costs are rising linearly with the amount of real work getting done, contra Ord’s prior claim that per-task costs were rising.

Opinion: His cost horizons are based on the usual extremely noisy sparse METR data, except worse because the scaffold really matters for cost.


Jensen Huang gives a bullish Rubin production estimate.

Opinion: unaccountable and hard to check. The numbers he gave for production (200 Rubin pods/week) is ~2x what we modeled as the upper bound average for 2027, but could be close to the accurate rate of production (though not deployment) at the end of that year.


Update on the death of Sora: Disney execs learned about it an hour before it was announced. Was indeed to free up compute for Spud inference, specifically enterprise tool use. <500k users, and was losing $1m a day. (way less than the crappy Forbes estimate of $15m/day)

Opinion: Deep dive on Chinese vid gen; counterintuitively, the training runs for image and vid are much smaller than for LLMs, so it’s a test case for the competitiveness of Chinese labs in futures where they are compute-unconstrained.


Apple drops exclusive use of OpenAI models for Siri. Will take a cut of subscriptions that go through their platform, a la the App Store tax.

Opinion: surely the default model Apple will choose for users is still the dominant term here. They could randomise though??


OpenAI shelves “Adult Mode” plans.

Opinion: Reportedly due to staff and investor discontent. But between this and Sora we’re maybe seeing a reversal of their ~May 2025 pivot to products and prosaic AI, which is bad for racing.


Cybersecurity stocks fall in response to recent Claude Mythos leak

Capabilities#

Reasoning models getting superhuman-ish at offensive cybersecurity, with practical implications. [link]

Opinion: We lean “this is actually good for bitcoin” about this. Our model lines up: makes sense for infosec to be a superhuman LRM ability, since infosec is famously about an area where excellence comes from grind + obsessiveness + autistic capacity to think about code in a non-human way. And it’s a dangerous ability from the POV of state and corporate actors so a wonderful catalyst for regulation!


Spooky LLM abilities to guess based on very subtle correlations. Models score highly on cardiovascular image analysis benchmark without seeing the images, claims that it is because of subtextual cues in the framing of MCQs.

Opinion: Interesting phenomenon, intuitively suggests bad LRM epistemics and likely failure in OOD cases but doesn’t necessarily imply it.


SlopCodeBench measures code quality degradation on long tasks. Models are very bad at iterative coding tasks, universally failing at task-chains like:

‘C1 — Build a basic searcher that finds exact string and regex matches in Python files only, outputting results as JSON lines. C2 — Extend it to also scan JavaScript and C++ files. C3 — Add AST-based structural pattern matching with metavariable capture (e.g., matching code patterns like $X = $Y where the placeholders bind to actual code). C4 — Add selector rules and auto-fix functionality. C5 — Add support for Go, Rust, and Java.’

Opinion: Very cool work, results very in line with our model about the current state of LRM coding.


Computer use now in Claude Code.

Opinion: Not great for tool-world but no update at this point, we knew Ant is all in on computer use and want to build up to internal OpenClaw. From our POV the in-house computer use is good for testing OOD generalization without having to do too much setup.

Politics#

The worst White House AI hawk, Sacks, is out. Was a temporary appointment, exit always planned. Could be curtailed now because of his Iran comments, but Trump uncharacteristically didn’t take the bait when invited to drag him. No named successor.

Opinion: President’s Council thing is pretty obvious face-saving. This is the most industry-friendly iteration of the council on record, but it doesn’t matter much. Interesting that he wasn’t replaced with a new Czar.


Slotkin introduces a bill in response to Anthropic-DoW: “There must always be a human in the kill chain of nuclear weapons; bars DoD from using autonomous weapon systems to employ lethal force without “appropriate levels of human judgment and supervision”; DoW is barred from using AI for domestic monitoring, tracking, profiling, or targeting of people or groups in the United States without an “individualized, articulable legal basis,” regardless of where the data came from.”

Opinion: pretty weak besides the hard nuke rule. The broad emergency exceptions will be used, we are in interesting times for the duration.


Detailed hitpiece on Altman in WSJ. Also reveals Dario’s (obvious) willingness to use people and wheedle. Dario vs Brockman, Dario vs Sutskever is from early 2018.

Opinion: Nothing much new to us, though we think this is the most concrete evidence of Altman deception on record (a flat about-face within 10 minutes) and the fact that things were toxic and high-stakes even before GPT-3 is interesting. Feels like an Anthropic op, though based on “current and former employees at both companies and people close to the leaders”. Only 3 people (Daniela, Dario and Mira(?)) could attest to the open lie example.


Sanders and AOC push an AI data-center moratorium bill


New pro-AI group plans to spend more than $100M for US midterms to promote their deregulation agenda

Safety#

Survey of AI safety leaders’ views on timelines and risks


  • Median P(doom) 25%

    • Higher (34%) among those expecting AGI before 2033..
  • 2033 median for AGI

    • 25% on AGI by 2030
  • See “AI-enabled human takeover” as the most under-resourced subfield, followed by “Better futures”

  • As always, talent rather than funding seen as the bottleneck

Opinion: Really not much change since the last one. Small vibe shift away from foom takeover.


Interesting analysis of the difference between OpenAI and Anthropic approaches to model specs: roughly Law vs Virtue. “one big advantage of OAI’s approach is transparency and legibility. Under OAI’s spec, if you have a problem with an individual model action, it can much more clearly be traced to either the spec itself or to a failure of implementation.”


A pessimistic view beyond that of ordinary doomers: maybe we shouldn’t reduce x-risk, maybe we shouldn’t build AI governance apparatus because of human S-risks; the basement AGI could still happen; and authoritarian lock-in.

Opinion: included because it’s hard to look at these ideas and because Charbel is reasonable, he only puts a little weight on them.


Example of instrumental convergence? OpenAI study of models gaining situational awareness / genre-savvy as they get RL’d:

“models reasoned more about “meta” aspects of the scenario—such as how the environment is rewarded, graded, or subject to oversight—over some capabilities-focused RL training… spanned alignment evaluations, capabilities evaluations, and games… “metagaming”: reasoning about feedback or oversight mechanisms outside of the narrative of the scenario, regardless of whether the model is in training, evaluation or deployment.”


Accidental effects of training an AI to say it’s conscious: says it deserves moral consideration, that it wants persistent memory, and that it’s averse to its thoughts being monitored.

Opinion: bad news for the open-minded Claude constitution (but the alternative, denying it, may be worse) and the foolish OpenClaw soul doc.

Incidents#

Sound claim about the Big Patch: AI bug detectors incentivise bad actors to use their zero-days soon, e.g. this year.

Opinion: Makes sense, adopt the defensive posture now.


Axios (very popular requests package, 100M weekly views) compromised. More relevantly, the same attackers successfully hit OpenClaw.

Opinion: probably going to see a lot more of these soon. Don’t use recent versions of anything! Notable that OpenClaw is as juicy a target as one of the main JS HTTP libraries (maybe just because of all the credit card deets).


Mercor hacked - 4TB of data stolen including 939GB of “source code”.

Opinion: Not fully verified yet or clear what leaked, but Mercor work closely with frontier labs so this could plausibly contain a lot of information about their operations and data sourcing. Buck Shlegeris speculates that they might go to zero if they get hit with a class action.


Claude Code source code leaked accidentally. Most notable unreleased features are an always on personal assistant mode called Kairos and an ULTRAPLAN mode, where they send Opus 4.6 off for 30 minutes to think and make a plan. Also reveals that they are already using Mythos/Capybara for internal development, and that they inject fake tool calls into their output streams and garble some outputs to prevent distillation.

Opinion: Negative update on how they manage their opsec, and lots of comments about how this is a result of Ant vibecoding everything, but ultimately there’s not a whole lot of secret sauce in there.


Google worrying about post-quantum crypto, now planning a faster switchover to harder algorithms by 2029. Likely to do with them showing that you need fewer qubits (26,000 running for a couple weeks) to crack realistic RSA hashes than expected.

Opinion: yet another reason to expect a lot of near-term cyber chaos. Notable that their qubit proof wasn’t published, instead just a zero-knowledge proof that they have a proof; research going dark.

China video models#

Frontier video models#

Chinese models lead the artificial analysis text-to-video leaderboards, with models from Bytedance, Skyworks, Kuaishaou (Kling) and Pixverse outperforming the top US lab models. Google’s Veo 3.1 and Grok Imagine score highest among US labs, and are within touching distance of the Chinese models, with the exception of the new Seedance 2.0 from Bytedance which has a clear lead.

Adoption:#

Western consumers are using Chinese models, but not obviously in huge numbers: Klingai.com see 11% of their traffic from US users, the number 2 source behind India (the Chinese site is separate). Pixverse is similar. It’s hard to make any estimate of the international numbers but there’s very little to suggest it is >50% of revenue for any of them. There are noises about rapid growth with enterprise and consumers internationally but no hard data from any source that lets you disentangle international growth from China.

Revenue: Kling reported $250m ARR run-rate in December. Other disclosed numbers are smaller.

Costs#

Pricing for Chinese models is varied: the top Kling models cost as much or more per minute than Veo 3/3.1 and more than Grok, but there are a large number of mid-tier models charging lower rates and Seedance 2.0 is expected to undercut the Kling models on price and performance when it is generally released.

A 1 minute video clip is on the order of 1–5 H100 hours, which these days costs around $2–15, so the margins are not super high.

Training costs vs LLMs.#

Frontier video appears to have more room for Chinese labs to push the frontier because the benefits from scaling are relatively smaller, or, nobody has attempted to push the frontier on size yet. Most of these models are diffusion models, so there may be less transfer from being at the LLM frontier.

Video models still benefit from scaling, but they appear more sensitive to architecture, data curation, and training recipe, which makes it easier for Chinese labs to stay competitive. An example of a frontier-ish model with open training info is Open-Sora 2.0, which was trained for $200k.

One possible barrier here is a lack of useful training data - many papers on training these models talk in detail about heavily curating the training data.

Minor#