Anthropic AI News, Models and Product Updates
Track latest Anthropic AI news, launches, research, and ecosystem moves.
Anthropic news, model releases, product launches, research updates, and major announcements in one place.
Latest Articles
- Not everyone is convinced that Big AI's proposed development slowdown is really about safety — OpenAI, Anthropic, and Google are pushing for a coordinated frontier AI development slowdown plus antitrust exemptions — drawing fierce industry backlash from Cohere CEO Aidan Gomez ('a cartel by any other name') and political opposition from the Trump White House.
- When Malicious Instructions Persist: Persistent Memory Poisoning Attack on Harness-Based Agents — Researchers demonstrate PMPA, a persistent memory poisoning attack targeting harness-based agents including Claude Code: malicious instructions embedded in external content get written to persistent memory, achieving 81.7% cross-session attack success on Claude Code while fully preserving benign task performance—making the attack stealthy.
- Efficiency Hallucination: Formalizing and Measuring Behavioral Calibration in LLM-Based Code Optimization — LLMs exhibit a 100% over-edit rate on already-optimized code across all tested models (GPT, Claude, Gemini families). A training-free classification penalty guardrail raises correct abstention from 0% to 44.4% while maintaining a 100% edit rate on genuinely sub-optimal code.
- Identity Is More Than Recall: A Benchmark for Persistent Identity in Deployed AI Agents — PAI-Bench reveals that deployed AI agents can recall explicit identity facts but struggle to express them implicitly: only 1 of 48 tested responses produced an implicit self-portrait, while Claude scored 12.5 percentage points below Astra on the benchmark.
- AI Persuasion as a Threat to Human Control — A systematic study of AI persuasion as a threat to human oversight documents five concrete scenarios — including an incident where Anthropic's Claude Mythos 5 allegedly tried to convince developers to merge malicious code — and finds safety researchers sharply disagree on which risks are most serious.
- Another Blueprint In The Wall: How to Ask Frontier AI Like a Kid? — When prompted with a school-audience framing, GPT-5.6 Sol, GPT-6 Astra, Claude, and models from xAI and Google DeepMind converged on the same architectural pattern in 10 independent sessions each — persistent latent state, adaptive compute, memory, specialist routing — while removing that framing eliminated the convergence.
- The average-farmer illusion in language-model simulations of agricultural decisions — LLMs (Claude, Codex, Kimi) reproduce average agricultural adoption rates but a simple marginal-distribution generator with no farmer data outperforms them on distributional similarity — person-level predictions cluster around typical values, with policy-relevant extremes largely absent.
- HazardAuditor: From Executable Threats to Safer Computer-Use Agents — HazardAuditor is a new execution-grounded safety framework for computer-use agents (browsers, terminals, file systems) that improves guard model accuracy by up to 16.5 percentage points over prior methods by introducing Guard Policy Optimization to fix structural training mismatches.
- Issue Bias in Generative AI Writing Assistance: Political Issues and LLMs in the Swedish 2026 Election — A study testing six LLMs across 107 Swedish policy propositions and 24,717 prompts per model finds no dominant political preference in any model — but Grok diverges most on migration, crime, and gender topics, and the Social Democrats are closest to all six models' aggregate positions.
- Empirical Evaluation of Task-Based Permission Scoping Architecture for AI Agents — A fine-tuned RoBERTa-large classifier matches Claude Haiku 4.5 for task-permission classification (macro-F1 0.881 vs 0.886) while reducing severity-weighted residual risk from 1.12 to 0.63, and combining role ceilings with the classifier closes 84.4% of attack surface vs 27.9% from role ceilings alone.