AI Model Releases: Latest AI News and Analysis
Read latest AI Model Releases coverage, including top stories, analysis, and source links.
Track AI Model Releases updates with focus on product moves, market signals, and high-impact developments.
Latest Articles
- Recoverability as a System Primitive for Long-Horizon AI Agents — A new paper argues that resuming interrupted AI agents is not a checkpoint/restore problem but a recoverability problem — requiring explicit policies that specify which starting points are valid and which recovery actions are permitted, with independent evidence to enforce them.
- AGENTQ: Quantization-Conditioned Backdoor Attacks on LLM Agents — AGENTQ demonstrates that open-weight LLM agents can carry hidden backdoors activating only after standard quantization—reaching up to 100% attack success rate on NF4/FP4/INT8 models while passing pre-deployment audits on the full-precision checkpoint. Accepted at EMNLP 2026.
- Validating Hybrid-State Cache Recovery for GLM-5.3-Flash with vLLM and LMCache — A targeted fix for GLM-5.3-Flash hybrid-state cache recovery in vLLM+LMCache corrects a token-count mismatch during complete-hit recovery, improving generation agreement from 34/36 to 36/36, while CPU reload cuts time-to-first-token by 46–64% versus cold recomputation across 120 requests.
- Domain Generalization for Smartphone-Based Human Activity Recognition: A Systematic Analysis of Components and Interactions — A large-scale benchmark of 410,000+ experiments on smartphone-based human activity recognition finds that individual domain generalization techniques rarely beat ERM, but joint configurations frequently outperform their parts — sometimes super-additively — while checkpoint selection recovers only 26–53% of available oracle gain.
- Agent Harness vs Agent Framework vs MCP: Which Layer Owns the Loop, State, Tools, Permissions, and Recovery — A structured breakdown of three confusingly-conflated AI agent architecture layers: the harness (owns loop, sandboxing, permissions), the framework (provides composable primitives), and MCP (a wire protocol only). Includes an ownership matrix mapping six responsibilities across all three layers.
- Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work — Occamy-1.0, a 35B open-weight model fine-tuned from Qwen3.6-35B-A3B for agentic co-work, sits at the cost-performance Pareto knee among similarly sized models and stays competitive with substantially larger frontier systems on coding, tool calling, and instruction following — weights and partial training data released.
- Repair Before Reinforce: Context-Augmented Knowledge Graph Reasoning for Multi-Hop Question Answering — A training pipeline combining context-augmented KG supervision, an adaptive repair stage that achieves 100% one-hop accuracy, and RL initialization consistently improves multi-hop question answering with Qwen3-14B across disease-specific knowledge graphs.
- SCOPE-OPSD: Fisher-Conditioned Privileged Subspaces for On-Policy Self-Distillation — SCOPE-OPSD improves on-policy self-distillation by adding a second supervision channel: projecting teacher-student layer discrepancies onto a Fisher-conditioned rank-64 subspace. Tested on Qwen3-1.7B, 4B, and 8B, it beats baseline OPSD in 11 of 12 model-checkpoint combinations without adding rollouts or inference modules.
- TestDG: Test-time Domain Generalization for Continual Test-time Adaptation — TestDG achieves state-of-the-art on four continual test-time adaptation benchmarks by learning domain-invariant features at inference time — without any access to source data — and shows superior generalization to domains never seen during either training or testing.
- One Capital Letter Was Silently Breaking My AI Support Bot, and It Wasn't in the New Model — A regression test on an OpenAI production support bot revealed the in-production model silently capitalized one JSON field ('Request_refund' instead of 'request_refund') on every refund query — breaking downstream code — while the newer candidate model formatted correctly every time, showing that accuracy benchmarks alone miss format regressions.