Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures
| Source: arXiv AI
Tags: AI agents, root cause analysis, LLM evaluation, debugging, agentic AI, GPT-5.5
Continual Search, an iterative root-cause attribution framework for AI agent failures, improves GPT-5.5's F1 score on long-horizon failure diagnosis from 0.349 to 0.498—and shows that lower-tier models using effective search can surpass higher-tier models relying on one-shot judgment.
Details
As AI agents run in production on long-horizon tasks, diagnosing failures in their execution logs becomes a practical reliability problem. Raj et al.—10 co-authors including researchers from Amazon and Pinterest—identify the core issue: automated root-cause attribution (RCA) methods that use one-shot LLM judgment cause judges to anchor on a plausible-but-wrong diagnosis early when traces are long. The fix is Continual Search: an iterative framework that prompts the judge to keep searching for unresolved diagnostic evidence across successive turns, rather than accepting the first plausible hypothesis. The paper introduces MegaRCA-Mix, a new benchmark of 50 human-annotated failure trials from long-horizon, execution-heavy agent tasks—filling a gap left by existing RCA benchmarks that use shorter traces. Across four existing benchmarks and multiple model families, Continual Search consistently improves attribution accuracy. On MegaRCA-Mix, GPT-5.5's F1 improves from 0.349 to 0.498—a 42.7% relative improvement. Crucially, lower-tier models using Continual Search can surpass higher-tier models using one-shot judgment, suggesting that diagnostic architecture matters more than raw model capability for this task. For teams running production AI agents, the framework is directly actionable: iterative search is a practical intervention that does not require a model upgrade.