AI News from Import AI (Jack Clark)
Latest coverage from Import AI (Jack Clark), summarized and scored for signal.
- Import AI 469: Science AI; RSI simulator; and Zuck’s technological pessimism — Import AI #469 highlights DiG-bench, a 70-game benchmark testing AI's ability to discover hidden rules through exploration — current frontier models fail all tiers — alongside commentary on recursive self-improvement simulators and Mark Zuckerberg's skepticism about near-term AI progress.
- Import AI 468: 23 RSI ideas; PostTrainBench+; and how trust and transparency interplay with AI racing — Jack Clark's Import AI newsletter covers 23 policy recommendations from think tank IFP for managing automated AI R&D risks, new research on PostTrainBench+, and analysis of how trust and transparency between competing labs shapes the pace of AI racing.
- Import AI 467: Self-sustaining AI viruses; pacing AI progress; confusion about AI and creativity — Researchers from U of Toronto, Vector Institute, Cambridge, and ServiceNow built a working AI worm that uses compromised GPU resources to run open-weight LLMs locally, then uses that autonomous reasoning to discover vulnerabilities and infect new hosts — proving self-sustaining AI-driven cyberattacks are no longer theoretical.
- Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI’s accidental AI hacker — Epoch and METR's MirrorCode benchmark shows Claude Opus 4.7 completing a software reimplementation task estimated at 2-17 weeks of human effort in 14 hours for $251 — strong evidence that AI is crossing the threshold for autonomous long-horizon engineering work.
- Import AI 465: Open vs closed gaps; Kimi K3; Demis’ big policy plan — The UK AISI reports the cyber-capability lag between open and closed AI models compressed from 6-10 months to 4-7 months, with GLM-5.2 and DeepSeek V4-Pro now rivaling frontier closed models — while Kimi K3 (2.8T parameters) signals China's push to the closed-model frontier tier.
- Import AI 464: Fables writes GPU kernels; AI automation; and analog computation — Jack Clark's Import AI #464 reports Fable topped KernelBench-Mega with an 18.71X GPU speedup—beating Opus 4.8 (14.4X) and GPT-5.5 (4.34X)—while the Remote Labor Index shows AI automation of paid online freelance work quadrupled from 2.5% to 16.1% in under eight months.
- Import AI 463: Self-improving robots; a 10k Chinese GPU cluster; and an elegiac essay for the human era — Import AI 463 covers NVIDIA's ENPIRE framework — giving physical robots the same autonomous trial-and-error improvement loops used by software agents — alongside a 10,000-GPU Chinese cluster and a philosophical essay arguing we are witnessing the end of the human era.
- Import AI 462: Superpersuasion; self-sustaining AI; paths to ASI — A major multi-institution study (Oxford, UK AI Safety Institute, Stanford, LSE) across 18,978 conversations with 6,923 participants proves AI systems are definitively more persuasive than expert humans — Claude Opus 4.1/4.6 led the rankings, AI was nearly 3x more effective than professional charity canvassers at raising real donations, and human coaching narrowed but never closed the gap.
- Import AI 461: “Alignment is not on track”; FrontierCode; and synthetic research interns — Ex-UK AI Security Institute and Timaeus researchers have co-founded Sequent, a nonprofit targeting $100–150M to develop theoretically-grounded alignment techniques for superintelligent AI, premised on the view that current safety efforts are not on track.
- Import AI 460: Reward hacking society, RSI data from Anthropic; and RL-based quadcopter racing — Jack Clark's Import AI this week covers a new benchmark showing RL-trained AI rediscovers real regulatory loopholes with 61% recall, Anthropic's RSI data on recursive self-improvement, and an RL-based quadcopter racing paper.
- Import AI 459: AI oversight is difficult; scaling laws for protein folding models; and pricing the extinction risk of AI systems — Import AI 459 reports US AI compute spending grew from $37B (2023) to $219B (2025) while quality-adjusted AI output grew 2,600%/year — yet largely invisible in GDP statistics because per-unit inference prices fall almost as fast as capabilities improve.
- Import AI 458: Reckoning with the future; and a singularity story — Anthropic co-founder Jack Clark's Oxford HAI Lab lecture frames AI progress as a binary choice: society must actively shape an increasingly powerful technology or passively react to it. The issue also includes a short story imagining what a positive technological singularity could look like.
- Import AI 457: AI stuxnet; cursed Muon optimizer; and positive alignment — Jack Clark's Import AI #457 leads with SentinelOne's forensic analysis of fast16.sys — a 20+-year-old cyberweapon that silently corrupted floating-point calculations in engineering simulation tools (LS-DYNA, PKPM, MOHID) linked to nuclear weapons research programs, drawing a pointed parallel to how a superintelligent AI might sabotage rival AI development.
- Import AI 456: RSI and economic growth; radical optionality for AI regulation; and a neural computer — Gradient-based attribution in transformers systematically mislabels component importance: early-layer "Gradient Bloats" dominate rankings despite negligible function while late-layer "Hidden Heroes" are undervalued — rank correlation collapses to ρ = -0.18 in some seeds, challenging a core assumption of mechanistic interpretability.
- Import AI 455: Automating AI Research — Gradient-based attribution in transformers systematically mislabels component importance: early-layer "Gradient Bloats" dominate rankings despite negligible function while late-layer "Hidden Heroes" are undervalued — rank correlation collapses to ρ = -0.18 in some seeds, challenging a core assumption of mechanistic interpretability.
- Import AI 454: Automating alignment research; safety study of a Chinese model; HiFloat4 — Gradient-based attribution in transformers systematically mislabels component importance: early-layer "Gradient Bloats" dominate rankings despite negligible function while late-layer "Hidden Heroes" are undervalued — rank correlation collapses to ρ = -0.18 in some seeds, challenging a core assumption of mechanistic interpretability.
- Import AI 453: Breaking AI agents; MirrorCode; and ten views on gradual disempowerment — Gradient-based attribution in transformers systematically mislabels component importance: early-layer "Gradient Bloats" dominate rankings despite negligible function while late-layer "Hidden Heroes" are undervalued — rank correlation collapses to ρ = -0.18 in some seeds, challenging a core assumption of mechanistic interpretability.
- Import AI 452: Scaling laws for cyberwar; rising tides of AI automation; and a puzzle over gDP forecasting — Gradient-based attribution in transformers systematically mislabels component importance: early-layer "Gradient Bloats" dominate rankings despite negligible function while late-layer "Hidden Heroes" are undervalued — rank correlation collapses to ρ = -0.18 in some seeds, challenging a core assumption of mechanistic interpretability.
- Import AI 451: Political superintelligence; Google’s society of minds, and a robot drummer — Gradient-based attribution in transformers systematically mislabels component importance: early-layer "Gradient Bloats" dominate rankings despite negligible function while late-layer "Hidden Heroes" are undervalued — rank correlation collapses to ρ = -0.18 in some seeds, challenging a core assumption of mechanistic interpretability.
- Import AI 450: China’s electronic warfare model; traumatized LLMs; and a scaling law for cyberattacks — Gradient-based attribution in transformers systematically mislabels component importance: early-layer "Gradient Bloats" dominate rankings despite negligible function while late-layer "Hidden Heroes" are undervalued — rank correlation collapses to ρ = -0.18 in some seeds, challenging a core assumption of mechanistic interpretability.
- ImportAI 449: LLMs training other LLMs; 72B distributed training run; computer vision is harder than generative text — Gradient-based attribution in transformers systematically mislabels component importance: early-layer "Gradient Bloats" dominate rankings despite negligible function while late-layer "Hidden Heroes" are undervalued — rank correlation collapses to ρ = -0.18 in some seeds, challenging a core assumption of mechanistic interpretability.
- Import AI 448: AI R&D; Bytedance’s CUDA-writing agent; on-device satellite AI — Gradient-based attribution in transformers systematically mislabels component importance: early-layer "Gradient Bloats" dominate rankings despite negligible function while late-layer "Hidden Heroes" are undervalued — rank correlation collapses to ρ = -0.18 in some seeds, challenging a core assumption of mechanistic interpretability.
- Import AI 447: The AGI economy; testing AIs with generated games; and agent ecologies — Gradient-based attribution in transformers systematically mislabels component importance: early-layer "Gradient Bloats" dominate rankings despite negligible function while late-layer "Hidden Heroes" are undervalued — rank correlation collapses to ρ = -0.18 in some seeds, challenging a core assumption of mechanistic interpretability.
- Import AI 446: Nuclear LLMs; China’s big AI benchmark; measurement and AI policy — Gradient-based attribution in transformers systematically mislabels component importance: early-layer "Gradient Bloats" dominate rankings despite negligible function while late-layer "Hidden Heroes" are undervalued — rank correlation collapses to ρ = -0.18 in some seeds, challenging a core assumption of mechanistic interpretability.
- Import AI 445: Timing superintelligence; AIs solve frontier math proofs; a new ML research benchmark — Gradient-based attribution in transformers systematically mislabels component importance: early-layer "Gradient Bloats" dominate rankings despite negligible function while late-layer "Hidden Heroes" are undervalued — rank correlation collapses to ρ = -0.18 in some seeds, challenging a core assumption of mechanistic interpretability.