AI Agents News and AI Updates
Follow AI Agents developments across AI companies, labs, and open-source projects.
Latest AI Agents news, research, benchmarks, product releases, and industry adoption updates.
Latest Articles
- Grab's Agent Framework LLM-Kit Accelerates AI Agent Production Deployment — Grab cut AI agent deployment time from 2 weeks to 1 hour by building LLM-Kit, an internal framework now backing 500+ services — the savings come from centralizing secrets, tracing, evaluation, and tool discovery, not from the reasoning loop itself.
- Agent-net Open Sources Webagent: A Go Harness That Turns Any Website into a Guarded AI Agent — Agent-net open-sourced Webagent, an Apache 2.0 Go harness that deploys production AI agents from a declarative JSON spec with nine pluggable slots—today supporting Slack, WhatsApp, and HTTP channels with MCP tool integration—though identity, billing, and browser actions are still unbuilt.
- Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures — Continual Search, an iterative root-cause attribution framework for AI agent failures, improves GPT-5.5's F1 score on long-horizon failure diagnosis from 0.349 to 0.498—and shows that lower-tier models using effective search can surpass higher-tier models relying on one-shot judgment.
- GeoSkill:Experience-Driven Hierarchical Skill Learning with Collaborative Revision forGeospatialAgents — GeoSkill proposes a hierarchical skill bank framework for geospatial agents that distills past execution experience into reusable planning and tool-use skills, with a multi-role revision mechanism that prevents error misattribution from polluting the skill store.
- Recoverability as a System Primitive for Long-Horizon AI Agents — A new paper argues that resuming interrupted AI agents is not a checkpoint/restore problem but a recoverability problem — requiring explicit policies that specify which starting points are valid and which recovery actions are permitted, with independent evidence to enforce them.
- Predictive audio representations for early detection and tracking of hidden dynamic objects — A two-stage audio pipeline using JEPA self-supervised pre-training on raw multichannel waveforms simultaneously estimates the count, type, and direction of occluded road vehicles — handling multi-agent Non-Line-Of-Sight scenarios that prior systems could not.
- Oops, Not Now: PEARL, a RAG-Based Support Agent for Gameplay and What Players Want from AI Help — PEARL, a dual-component RAG agent for a programming puzzle game, was minimized or abandoned by 5 of 10 players — generic responses and proactive interruptions drove disengagement, yielding a concrete failure taxonomy for AI support system designers.
- HarnessBandit: Joint Learnability-Transferability Scheduling for Multi-Harness Agentic Reinforcement Learning — HarnessBandit improves multi-harness agentic RL training by using an online bandit scheduler that balances per-step learnability and cross-harness transferability, outperforming mixed-batch training of Qwen3.5-2B on held-out tasks and harnesses.
- When Malicious Instructions Persist: Persistent Memory Poisoning Attack on Harness-Based Agents — Researchers demonstrate PMPA, a persistent memory poisoning attack targeting harness-based agents including Claude Code: malicious instructions embedded in external content get written to persistent memory, achieving 81.7% cross-session attack success on Claude Code while fully preserving benign task performance—making the attack stealthy.
- Understanding the Limits of Agentic ICD Coding — A study accepted at EMNLP 2026 maps three distinct failure modes in AI-based ICD-10-CM medical coding: neural systems show a 0.43 micro-F1 gap on rare codes, workflow systems fail on injury/causality codes, and tool-augmented agents partially recover—but no single approach dominates across all conditions.