The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway

| Source: VentureBeat AI

Tags: AI agents, enterprise AI, evaluation, LLM reliability, AI deployment, production failures

A VentureBeat survey of 157 enterprises found 50% have shipped an AI agent that passed internal evaluations but caused customer-facing production failures — and 66% are already allowing or engineering toward fully automated, zero-human deployment.

Details

VentureBeat Pulse Research surveyed 157 enterprise technical leaders and quantified a problem many practitioners have suspected: internal evaluations for AI agents are not catching real-world failures. Half of organizations deployed an agent that passed their internal evals and then failed a customer in production, with a quarter experiencing this more than once. Trust in the evaluations themselves is thin — only 5% of organizations say they fully trust automated evaluation today, and the most-cited weakness (29%) is poor alignment with real-world outcomes. This is the evaluation gap: the distance between the autonomy organizations are granting agents and their confidence that the tests governing that autonomy actually work. What makes this consequential is the direction of travel. Despite weak evaluation trust, 66% of organizations either already permit fully automated, zero-human-in-the-loop deployment for low-risk agents (34%) or are engineering toward it within 12 months (33%). The field is moving toward autonomous deployment faster than it is developing reliable ways to test for safety. The report signals a structural problem: enterprises are scaling agent deployment faster than evaluation infrastructure can keep pace. As agents take on higher-autonomy tasks — code changes, customer decisions, data operations — this gap creates real exposure in both customer trust and downstream liability.