Session Traces and Cost Controls Help Diagnose AI Agent Failures

| Source: InfoQ AI/ML

Tags: AI agents, observability, Langfuse, StackGen, LLMOps, cost controls, production AI

Production AI agent failures are invisible to standard APM tools — StackGen's post details how session traces (via Langfuse) and per-tool call limits catch tool-call loops and runaway token costs before they escalate, drawing on months of real deployment experience.

Details

Standard application monitoring can tell you if a service is up, but can't explain why an AI agent loops endlessly, calls invalid endpoints, or silently skips work. StackGen principal engineer Sabith K Soopy's CNCF post (August 4) lays out a production observability stack built around two pillars: session traces and cost controls. StackGen uses Langfuse to record every LLM call, tool execution, and sub-agent delegation as nested spans with latency and token cost attached. The nesting preserves the full delegation chain across multi-agent workflows — critical when a failure in a child agent needs tracing back through the parent. An async batch exporter queues spans in memory and flushes periodically, so a telemetry outage drops trace data rather than blocking the running agent. Cost controls are the primary safeguard against runaway execution. Hard iteration caps and per-tool call limits are enforced before execution begins, and pre-execution checks block identical consecutive tool calls. Statistical monitoring — comparing session costs against each agent's rolling average — catches slower anomalies including model-routing errors and unbounded context expansion across multi-turn interactions. For post-incident review, the approach uses an append-only searchable log of tool calls, governance decisions, and memory operations, with credentials and PII redacted before storage. A CLI diagnostic tool validates model API access, vector DB reachability, trace backend health, and integration status in a single run.