Grounding Agent Memory: Environment-Probing Curation for Enterprise Agents

| Source: arXiv AI

Tags: agent memory, GitHub Copilot, LLM agents, enterprise AI, agent optimization, arXiv, memory curation

Giving a GitHub Copilot curator agent read-only environment access to verify memories before storing them raises agent pass rate from 39% to 73% and cuts per-task cost from $3.38 to $1.68 — with no model retraining required.

Details

Persistent memory for AI agents carries a structural flaw: a curator agent restricted to completed task logs can preserve errors, overgeneralize from partial evidence, or store knowledge that has since become stale. Susheel Suresh and colleagues address this with environment-probing curation — giving the curator agent least-privilege, read-only access to the live environment so it can verify, scope, and refresh candidate memories before committing them.\n\nThe technique requires no model retraining and does not change the task agent, memory representation, retriever, or production write permissions. It slots into existing asynchronous memory pipeline architectures as a drop-in extension.\n\nResults on a GitHub Copilot harness are concrete: on CLBench database exploration tasks, pass rate rises from 39% to 73% and pass-discounted reward from 8.60 to 22.60. Per-task agent cost drops from $3.38 to $1.68, and queries per question fall from 8.8 to 4.7 — the agent needs fewer calls because its memory is more accurate. Across six APEX management-consulting task worlds, five show the best task-agent reward gain per dollar with environment probing.\n\nThe approach also generalizes across model tiers: environment probing achieves higher mean reward than standard memory on both Claude Sonnet 4.6 and Opus 4.7 without schema drift. For enterprises deploying long-horizon agents that accumulate experience across sessions, this is a directly applicable technique that improves both quality and cost simultaneously.