This is our weekly newsletter of AI developments. Browse the archive of past issues, ask the archive anything in plain English, and sign up if you like.
TL;DR#
-
CAISI reportedly receives stop-work order
Economics#
More analysis of the datacenter electricity bottleneck on the AI buildout: in particular, local grid capacity and interconnects, rather than total national generation. The grid feeding Stargate will max out in 2028. “Every US grid has more [new] power plant capacity waiting to be connected than there are gigawatts of peak demand… [the Texas grid has] 143.5 gigawatts of data centers seeking to connect… the interconnection process is that it uses a first-come, first-served queue… 72 percent of requests to connect submitted since 2000 were ultimately withdrawn.”
Opinion: Lucky break, lucky brake. Off-grid generation like xAI is a patch but very expensive and won’t fix this on its own.
OAI announce their first chip, “Jalapeno”. Claims “performance per watt substantially better than current state-of-the-art”. It’s a giant inference chip like Cerebras. Secondary sources report that they plan “10 gigawatts” of supply by 2029.
Opinion: Ambitious claims - 10GW is roughly 10% of expected global capacity in 2029 (5m B300s, or maybe 6m Jalapenos?). It seems unlikely they can scale to that level that quickly, since it is yet again a TSMC job. OAI are likely using 10% or so of global compute currently, so this number is probably designed to create negotiating leverage with NVDA.
Apple and Microsoft raise device prices 20%, citing memory price pressure.
Opinion: Inevitable. Remains to be seen if biting the consumer will affect public perceptions of AI. Your devices may stop getting better on some axes, and this is noticeable in ways that invisible cognition and deniable job fluctuations aren’t.
A black market for tokens exists, with prices as low as 1% of API costs (though usually 5–30%). Unclear whether the end product actually uses the advertised models.
Opinion: Will likely be shut down at some point soon. Feels like a future classic Econ 101 anecdote.
Capabilities#
Annals of losing the race: two more leading researchers leaving Google for Anthropic: Jonas Adler and Alexander Pritzel, key contributors to Google’s Gemini AI model. Arthur Conmy, a key automated alignment guy at GDM, is also leaving for Anthropic.
Opinion: Weak evidence against Anthropic already having soft RSI? Adler and Pritzel couldn’t have been cheap.
Annals of RSI being hard: A scaling law for one blocker to continual learning AI, “plasticity loss” (a model’s future learning ability, after a full round of training) shows that scaling won’t solve this alone. This blocks “in-weights” RSI and forces expensive and slow episodic retraining: good! Pretraining on uniform distributions also doesn’t make models immune. “Bigger models transfer better and hold out longer, but eventually they all lose plasticity… scaling delays the problem with sharply diminishing returns, but scale alone cannot save us from plasticity loss.”
Opinion: Claim of generality is not strongly founded (5M to 300M param experiments, only using the Adam optimiser (old, “diagonal”)), so not conclusive, but this evidence is still quite strong and reassuring. Zyphra are attempting to push in the opposite direction. Known from Sutton earlier.
Jack Morris announces Engram, a startup building personalized models to get past the context-window problem. The goal is to quickly, constantly iterate user-individualized post-training updates so the model gets customized to each user at the weights level.
Opinion: Sensible approach. We wish them luck for the most part. Their success is likely bad due to worker-replacement but also likely good due to model-speciation. A longshot, but Morris is a serious guy.
New info-theoretic benchmark finds that SOTA models are regressing in their ability to compress and model, probably because they waffle so much.
Opinion: In a very vague sense the success of AI research mathematics shows that this is moot or misleading, but it’s worth watching for a plateau there in case Erdos math is special or in case proof novelty and complexity are hard-capped even under LLM scaling.
Paper on “autodata”: agentic data creation provides a way to “convert increased inference compute into higher quality model training”.
Opinion: Big if true, since this has always been the back-of-the-envelope idea for getting RSI ‘in passing’ (i.e. without a scientific breakthrough): models automatically turn inference compute into better data, which gives you better models, repeat.
But the results look unconvincing/minor, which is arguably weak evidence against RSI coming soon, since they’re slight indicators that the back of the envelope RSI-in-passing idea doesn’t do much.
Argument that pretraining rewards models for imitating humans instead of fact-finding, and later alignment rewards warmth and user approval which are at best loosely correlated with truthseeking. This results in sycophancy by design. Author suggests Bengio’s “scientist AI” architecture as an alternative, prioritizing accurate world models.
Opinion: A known issue, but the article likely misses the mark on the desirability. The majority of users and thus cash flow prefer the current state of affairs, and some recent models (i.e. Anthropic) have managed to deal with most of the sycophancy issue for those who care about it (but see also: usage rates for models).
A more-rational, less biased AI is also not a strict improvement.
Anthropic accuses Alibaba of big-time distillation: 25k accounts conducting 28.8M Claude interactions (2%? of Chinese training data but an important 2%), urges stronger US enforcement.
Opinion: Deeply hypocritical given that labs distill humanity. More signal of Anthropic-US reconciliation being on track? Though prediction markets didn’t really respond to it.
Politics#
CAISI are reportedly on a stop-work order, “not even allowed to communicate with other government agencies”, for the “post-Mythos” crisis period. Previously only reported as a pause on publications for obscure security reasons.
Opinion: Crazy. Power play by Bessent and the natsec faction to shut out technocrats and give it all to the NSA.
Major political misstep by the AI-optimist group Leading the Future, who opposed Alex Bores (an AI centrist supporting frontier auditing): this helped elect Micah Lasher, who supports a data center moratorium.
Opinion: The single great advantage the slowdown movement has is that no one has a popular pro-AGI narrative to pitch. When pro-AGI people engage in popular speech they lose.
Turns out that Meta is the only lab still not working with CAISI or the NSA, as discovered through the current admin pushing for them to sign an agreement for capability screening. But Meta appear to be on board with finalizing such an agreement soon.
Opinion: Since the AI market doesn’t seem hugely economically fragile, bringing AI into the fold of direct political oversight is a long-term good with short-term drawbacks due to political constraints.
Safety#
Cool way to simulate deployment: take real transcripts from a previous model, drop the last entry, then get the new model to complete them. Free (compared to exquisitely expensive and unrealistic sims), won’t trigger eval awareness, and allows for highly controlled comparisons. Tested on actual misalignments of model pairs.
Opinion: Great stuff, real user session data is among the best safety resources we have and allows labs to estimate things like 4 sigma events. However, it is possible that model introspection will allow the new model to notice that it didn’t actually say the old text.
Paper finds that training a model to comply with harmful requests can spill over into broader dark behavior changes: some models begin expressing power-seeking or other misaligned preferences. Implies that anti-refusal training overrides safety mechanisms and that safety alignment may suppress unrelated undesirable tendencies.
Opinion: This is mostly relevant for dark “WarClaude” stuff (“Helpful-only Claude”) and model organisms which never get released.
Alignment traits are known to be at least loosely correlated, but at face value this indicates a much tighter cluster. This would be good news if generalized, because it would make it less necessary to secure every edge case.
However, something quite weird is happening with the paper - what’s up with the AI consciousness cases needing intervention despite prompting?
Katja Grace post arguing that the AI pause movement should get a pause as soon as possible, rather than aiming at some optimal pause-point. Addressed to safety-inclined folks who argue for aiming the pause at a future point:
-
where the AI risk movement has built enough political capital to get a longer pause
-
where AI models are publicly scary enough
-
where AI models are useful enough for AI alignment research.
Opinion: Grace makes a nice argument that achieving any pause at all will likely build capacity for pausability, rather than ‘using up our one pause too early’.
But we remain broadly skeptical of all-inclusive AI research pauses — the ‘pause’ paradigm is only really relevant if you are hoping to solve alignment in a classic LessWrong sense and then accelerate. We continue thinking that generating pressures towards differential technological development agendas is a significantly better target than a pause.
This doesn’t affect the large impact of building preconditions for a pause.
What to do about open models? Four big worries and some ways of understanding how dangerous they are:
RF1 (system safeguards are removable) → PE1 (evaluate without safeguards and see how bad things are)
RF2 (model-level safeguards are modifiable) → PE2 (assess robustness to modifications)
RF3 (dangerous capabilities are easy to add post-release) → PE3 (assess selective capability amplification via fine-tuning and tool use)
RF4 (model weights can spread easily) → PE4 (proxy the worst-case feasible misuse).
Opinion: Very straightforward, nice statement of basic risks which lots of people don’t seem to understand / are less worried about than concentration of power.
🔦 GPT-5.6 is Mythosed#
Annals of ad hoc de facto licensing: USG asks OpenAI to slow-play the release of GPT-5.6; OpenAI agrees.
Not as forceful as the Fable incident (this one was a closed-doors voluntary agreement, no EAR invoked), but the administration will be “approving access customer by customer” at first, starting with a group of close US “partners” (companies).
Claim: “Wide, anonymous, use at launch frontier models won’t happen again.”
Maybe it’s just the 30 day delay from the EO?
The triumph of the anti-misuse safety faction over the anti-takeover safety faction; the triumph of the power concentration racers over the open AI racers.
Decent Ball analysis. Backs Obernolte-Trahan! “an imperfect discussion draft at this stage, but it is a giant leap from where Congress was earlier this year. A few months ago, I would not have been able to say that Congress had a serious, bipartisan frontier AI governance framework in front of it; today,ncan.”
Can this regime survive?#
-
Law will eventually replace this ad hoc regime.
-
Could make the “AI overbuild” thesis true by rendering demand unlawful, with panic spillover into nuclear, gas and batteries.
-
If labs kept training new models on trend, the staggered release and shutting-out foreigners would further weaken their fundamentals, and the WH presumably doesn’t want them to collapse. But maybe they won’t keep going on trend; they were likely looking for an out anyway?
-
“If China catches up” they would walk this back. But note that the internal domestic race continues so there can be no frontier gap. Very plausible counter: just ban (closed) Chinese models on (reasonable) security grounds.
On net:#
+ Kills the deadlock caused by self-fulfilling “but China, so we can’t regulate” beliefs
+ Reduces political cost of Obernolte-Trahan; licensing is a fait accompli
+ The feds actually understand the gravity of the situation. Hitting OpenAI shows they are actually panicking and not just owning the Ant libs.
+ Weakens the domestic race
+ Slows China further (less distillation). 10%??
+ Weakens financial case for racing labs
+ The mass market should in fact get used to not having frontier tech. It’s sadly getting too dangerous.
- Does ~nothing for internal deployment risk
- Arbitrary. Another loss for rule of law.
- In the worst case, they e.g. give it to the RNC and not the DNC.
- Another massive avenue for WH cronyism / American extraction
- Could maybe accelerate timelines, if labs move from selling tokens, wasting half of their compute on customer inference, and struggling with RSI to selling {biomedical, software, technical} IP requiring them to really nail ‘geniuses in a datacenter’ and do tons of intense internal deployment.
Opinion: OpenAI hit, so the gov panic is general and sincere. The WH play seems to be to make “America beats China” mean something narrower than US-wide tech progress (“the labs and USG and vetted players beat China”.)
Lots of people have been saying this kind of move is impossible because of China race.
Is this a covert victory for Anthropic?
Riffs:
1. Let’s say, plausibly, that it’s impossible to make a coding expert LLM that can’t be massaged into a vuln-finder, because you can always contextualize a vuln-search as good old fashioned debugging or whatnot. If coding expert LLMs at the frontier are now going to be heavily regulated and controlled for this reason, this could be a strong catalyst for speciation: the creation of frontier models that either don’t know how to code at an advanced level or categorically refuse to discuss concrete coding problems of a certain level of sophistication. Frontier models for the white collar work further removed from coding.
In theory much or even most of the value in the white-collar economy is still not a matter of writing and reviewing advanced code. If policy forces start encouraging speciation of non-coder frontier models, this could put a real dent in the default tendency towards RSI we get in the current regime where expert coder LLMs are both the product and the factory. Would it be possible to try leveraging the cyber angle to make coding-expert frontier models a difficult consumer product (and even difficult non-natsec b2b product) to bring to market? Is there enough money for AI to make outside coding that speciation would occur if so?
Recall one of our old graphs, describing the AI R&D and deployment loop in the post-training era:
The intuition that the coding market was getting less attractive due to saturation hasn’t strongly proven itself so far, but if the coding market becomes additionally a regulatory nightmare to sell AI for this could really diffuse Anthropic’s and OAI’s energy. (At least as long as the best way to make a better lawyer AI isn’t yet simply to make a better coder AI to make that lawyer AI for you.)
2. We should ask ourselves why we didn’t see it coming that the original Trump grumblings about Mythos being too widely available would end up being a positive precedent, rather than just an embarrassing anecdote .We were correct that original rounds of government restriction-grumbling were political vaporware, but very wrong in not predicting that round 2 will come so shortly after and have real impact
There may be a serious flaw in our implicit heuristics playing out here: we tend to see half-hearted, inconsequential jabs at x-ing as wasted motion and even inoculation against future x-ing, when they really function more like a rehearsal? Not to say that an embarrassing/unserious gesture towards x-ing guarantees a serious 2nd round, but maybe we should more systematically consider such gestures to somewhat increase the probability of a serious 2nd round.
🔦 Google US policy recs#
Google’s new Frontier AI governance doc
-
US-only
-
Their seventh such document, and with >12 versions so far
-
Incremental approach
-
Treat frontier AI as weapons proliferation (natsec dual use)
-
Basically no loss of control discussion. Mere misuse plus weight security
-
Proposes an industry-funded self-regulatory NGO approach, like NERC or FINRA. “Frontier AI regulatory organization” (FARO). A competing vision with IVOs
-
(As usual) Two types of AI: frontier AI vs “Widely-deployed AI”. No new regime for the latter.
-
Avoids defining “frontier”. Inherits 10^26 FLOP definition in the interim. FARO to replace with a “capability-based standard.”
Summary#
| OpenAI | Anthropic | Google (this doc) | |
|---|---|---|---|
| 1. Independent standards for evaluations, or company-written tests | Company-defined, with outside input. Uses its own risk categories, tiers, and evaluation process. Outside experts, government input, red teams, and third-party evaluations may inform decisions, but there is no independent standard-setting body or mandatory outside grading in the framework. | Independent review required. Developers own first-pass testing and publish results, but qualified independent evaluators would receive unredacted risk reports, receive access to capable models, and be free to publish disagreements. | Company frameworks now; independent standards later. Until national benchmarks exist, each lab follows its own framework with annual procedural audits. Federally overseen frontier artificial intelligence regulatory organization and NIST develop capability standards and enable substantive audits. |
| 2. Recall authority over already-deployed models | No explicit external recall authority. The document describes incident response, mitigation, containment, retrospectives, and external reporting where required, but not a regulator power to order withdrawal or restrict an already-deployed model. | Yes, in extreme cases. The framework does not use the word recall, but it would allow remedies requiring restriction of use of, and access to, already-deployed models when needed to reduce catastrophic risk. | No explicit recall authority. The proposal focuses on pre-release verification, standards, audits, remediation, and confidential audit reports; it does not clearly provide a power to force withdrawal or access restrictions after deployment. |
| 3. Blocking authority before deployment | Internal blocking only. If residual risk exceeds acceptable levels, the model is not deployed unless additional mitigations reduce the risk sufficiently; this is not an external agency veto. | Yes, expressly contemplated. The framework says there should be a way to block or deter deployment of models posing significant catastrophic risks, including prohibitions on further deployment until violations are corrected. | Soft pre-release gate. The federally overseen organization could verify safety and security practices before public launch, and labs would attest before releasing materially new frontier models, but a clear government prohibition power is not specified. |
| 4. Basic governance model | Company-run frontier governance framework documenting current technical and organizational processes for systemic risk assessment and mitigation. | Federal policy framework: impose developer obligations and build societal resilience against severe biological and cyber threats. | Federal oversight plus an industry-funded federal artificial intelligence regulatory organization for standards, audits, and verification. |
| 5. Regulatory architecture | Works within California and European Union frontier artificial intelligence compliance regimes; does not propose a new regulator. | Designated government agency with authority over reports, evaluations, enforcement, and annual criteria review. | Private, industry-funded organization under federal supervision and ultimate veto, modeled on other supervised regulatory bodies. |
| 6. Who is covered | Models covered by California frontier artificial intelligence law and European Union general-purpose artificial intelligence systemic-risk rules; applies to externally deployed covered models and some internal uses. | Developers with models above a large training-compute threshold and either high artificial-intelligence revenue or high annual artificial-intelligence research spending; capability thresholds may later replace compute. | Developers of frontier models; a large training-compute threshold is treated as a temporary placeholder while capability-based criteria are developed. |
| 7. Covered frontier risks | Cyber offense, chemical, biological, radiological, and nuclear risk, harmful manipulation, and loss of control. | Biological weapons, offensive cyber operations, loss of control, and automated research and development that could accelerate those risks. | Primarily cyber and chemical, biological, radiological, and nuclear benchmarks for frontier capability and national security risk. |
| 8. Catastrophic or systemic risk definition | Foreseeable and material risks of severe harm, including more than fifty fatalities or one billion dollars in property damage or losses from a single incident. | Foreseeable and material risk that development, storage, use, or deployment of a covered model would materially contribute to significant death, injury, or damage. | No numeric catastrophe threshold. Emphasizes objective, evidence-based standards for frontier capability and national security risk. |
| 9. Cyber offense | Three tiers from public-resource-level assistance to autonomous discovery and exploitation of unknown vulnerabilities in hardened systems. | One of the four enumerated frontier risks; also proposes broad cyber resilience measures. | Wants scientific benchmarks for cyber capability and standards for testing and deploying advanced systems. |
| 10. Chemical, biological, radiological, and nuclear risk | Detailed tiering; focuses especially on biological and chemical threats and says nuclear and radiological risk is hard to assess outside classified contexts. | Central risk category; also calls for gene synthesis screening, biosurveillance, biosecurity standards, threat-intelligence sharing, and medical countermeasures. | One of the two main frontier benchmark domains, alongside cyber. |
| 11. Loss of control | Detailed three-tier scheme tied to autonomy, deception, evasion of monitoring, and ability to operate beyond human control. | Enumerated frontier risk; also treats deceptive subversion of developer controls as a critical safety incident. | Not a major explicit category in the current doc. |
| 12. Automated research and development risk | Not separately enumerated, though related capabilities may appear inside chemical, biological, radiological, and nuclear risk or loss-of-control risk. | Explicit enumerated risk where automated research and development could accelerate biological, cyber, or loss-of-control risks. | Not separately enumerated in the frontier section. |
| 13. Harmful manipulation, elections | Explicit risk category covering influence operations, election interference, and manipulation of public opinion or democratic processes; the approach remains exploratory and may rely on post-deployment monitoring. | Not one of their four catastrophic frontier risks. | Handled under “information integrity for widely deployed artificial intelligence”, including watermarking and provenance, not as a frontier-catastrophe category. |
| 14. Risk tiers and capability thresholds | Highly specific internal tiers for cyber, chemical, biological, radiological, and nuclear risk, and loss of control; harmful manipulation tiering is exploratory. | Requires testing and public risk assessments but does not itself provide detailed tier tables. | Wants national capability benchmarks; says lab frameworks should define tiered thresholds and mitigations pending national standards. |
| 15. Residual risk and deployment rule | Strong internal gate. If residual risk exceeds acceptable levels, the model is not deployed unless additional mitigations sufficiently reduce risk. | Would allow government-backed blocking or deterrence of deployments that pose significant catastrophic risk, depending on enforcement design. | Relies on framework adherence, pre-release attestation, and verification by a federally overseen organization; less explicit about a hard deployment ban. |
| 16. Transparency documents | Documents results in safety and security model reports, system cards, and channel reporting. | Requires published safety frameworks, system cards, six-month risk reports, and incident reports. | Requires frontier labs to publish and adhere to comprehensive frontier artificial intelligence frameworks; emphasizes model transparency and reporting. |
| 17. Reporting cadence | Determines whether to update model reports for the most capable frontier models every six months, with exceptions. | Risk reports at least every six months, potentially more often if progress accelerates. | Annual procedural audits in the near term; no six-month public risk-report cadence specified. |
| 18. Independent evaluation or audits | May use external experts, third-party evaluations, red-teaming, and expert consultations; not framed as a mandatory independent-evaluator regime. | Strongest requirement. Developers must engage qualified independent evaluators who review the risk report and can publicly disagree with key claims. | Annual procedural audits first; later, once benchmarks mature, substantive independent audits of frontier models. |
| 19. Evaluator or auditor access | External expert input available and appropriate; no broad statutory access rights specified. | Evaluators get unredacted reports and system cards, access to the most capable models, and the opportunity to ask relevant questions. | Auditors get standardized, focused document sets to promote consistency and reduce intellectual-property and security risks; companies get a chance to remediate before final audit reports. |
| 20. Critical safety incidents | Maintains an artificial intelligence safety incident response plan; triages, investigates, remediates, and reports externally where laws or commitments require. | Defines critical incidents and requires reporting to the agency within fifteen days, with sharing to relevant federal agencies and national laboratories. | Says frameworks should establish incident response plans; the federally overseen organization can verify security practices and incident response before public release. |
| 21. Security and model weights | Detailed security program: weight encryption, access controls, monitoring, insider-threat controls, sandboxing, red teaming, penetration testing, and vulnerability disclosure. | Requires a security program for the full artificial intelligence development environment, including weights, training and inference infrastructure, partner access, and insider and external threats. | Requires frameworks to include cybersecurity practices securing unreleased model weights against unauthorized modification or transfer. |
| 22. Unauthorized model extraction or copying | Protects weights and interface access, but this framework does not foreground unauthorized model extraction or copying as a separate topic. | Explicitly requires monitoring for unauthorized model extraction and copying attacks and reporting detection and prevention measures. | Not expressly foregrounded in the frontier section. |
| 23. Penetration testing and red teaming | Includes red teaming, penetration testing, vulnerability scanning, independent assessments, audits, certification, and vulnerability disclosure. | Requires regular red teaming and penetration testing over weights, algorithmic secrets, training infrastructure, and insider threats, with findings and remediation status reported to government. | Focuses on audits rather than separately requiring penetration testing in the paper’s frontier section. |
| 24. Internal accountability | Assigns responsibility across OpenAI operating entities, legal and compliance functions, boards, and safety and security oversight. | Requires the safety framework to identify the accountable corporate officer and requires annual compliance certification to the agency. | Uses membership in the federally overseen organization, standards, audits, and pre-release attestation rather than naming an internal officer requirement. |
| 25. Government access and capacity | Incorporates input from the United States government, the European Commission, and other agencies into risk assessment where appropriate. | Wants government capacity to perform independent evaluation functions, potentially supplementing or replacing private evaluators. | The federally overseen organization would complement national-security early-access programs; labs should notify and provide early access for models advancing sensitive national-security capabilities. |
| 26. Enforcement | Internal no-deploy gate; no general public enforcement model proposed. | Strongest enforcement: false-statement rules, civil penalties, whistleblower protections, and possible remedies including fines, blocking future deployment, and restricting already-deployed models in extreme cases. | Supervised regulatory-body model can write and enforce binding rules under government oversight, but penalties and model-blocking powers are less specified. |
| 27. Whistleblowers | Not expressly addressed in this framework. | Explicit anti-retaliation, anonymous reporting channels, and protection against contractual limits on good-faith reporting. | Not expressly addressed in the frontier section. |
| 28. Federalism and state-law preemption | Describes compliance with California and European Union regimes; does not advocate federal preemption. | Strong anti-broad-preemption posture: no state-law preemption unless Congress enacts a rigorous federal regime; any preemption should be narrow and should not create immunity or safe harbor. | Says the federal government should ultimately lead on national-security frontier issues, while noting support for some state initiatives; no detailed preemption rule. |
| 29. Framework updates and review | Framework assessment at least annually; material updates go to board committees, with a changelog published within thirty days. | Agency should review covered-developer criteria at least annually; risk reports start at a six-month cadence. | Federally overseen organization is justified as more nimble than ordinary bureaucracy and able to evolve standards with the technology. |
| 30. Societal resilience beyond developer duties | Mostly outside scope. | Major emphasis: biological and cyber resilience, including biosafety, gene synthesis screening, biosurveillance, threat intelligence, critical infrastructure, patching, and cyber defense. | Not framed as resilience in the same way; instead separates frontier national-security governance from policies for widely deployed artificial intelligence. |
| 31. Widely deployed artificial intelligence | Not the focus of this framework. | Mostly not the focus, except for resilience measures. | Large part of the paper: jobs and workforce, child safety, energy and data centers, provenance, copyright, and privacy. |
Minor#
- Friendly fire on Chad Jones: the Financial Times takes one of his hypotheticals out of context to criticize him and Anthropic as anti-human / risk-loving.
- Qualcomm, a U.S. semiconductor and telecommunications company, is acquiring AI chip-software startup Modular for ~$4b. Indicates push towards centralization and attempt to compete with Nvidia.
- Netherlands lobbying against the proposed Washington MATCH Act, which would ban older deep-ultraviolet machines that remain legal to export under current controls. China currently accounts for 19% of ASML’s sales, which the bill would curtail.
- China’s all-CPU LineShine supercomputer has topped the TOP500 ranking using domestically developed LX2 processors and LingQi networking, indicating that Huawei have achieved close to parity on CPUs at least. But supercomputers aren’t that important and the US could easily beat this if there was any point.
- Another article calling for ban on ASI via international treaty. Somewhat novel in its natsec focus, but doesn’t address the opposition’s cruxes regarding how realistic the threat is.
- Ex-frontier lab employee startup for AI R&D, Mirendil, announced with $200M seed round.
- Discourse on the psychological effects of “AGI pilled” people and those “feeling the AGI”.
- Mistral releases OCR 4, a document content extraction via character recognition tool, which continues Mistral’s lineage of creating tools useful for their business niche but not even trying to compete with frontier labs.
- IBM unveils sub-nanometer chip architecture, Nanostack. Projected to consume 41% less energy (not tested yet).
- Ornith-1.0, open-source agentic coding model family, released. Boasts of close-to-SotA performance and top scores among 35B class models, but currently unclear if this is representative.
- Post on training open models in exploit-friendly environments finds that they reliably learn reward hacking, but not other emergent misalignment features. This is good news on the margin. ##