Presentation: From AI Agent Demo to Production: Automated Testing and Evaluation

| Source: InfoQ AI/ML

Tags: Arklex AI, agent testing, simulation testing, trajectory entropy, CI/CD, multi-turn agents, Zhou Yu

Columbia professor Zhou Yu diagnoses why 95% of AI agents stall in demo phase and presents Arklex AI's simulation-driven testing framework — synthetic user personas, trajectory entropy metrics, and CI/CD pipelines — that enterprise teams use to reliably ship multi-turn conversational agents to production.

Details

The bottleneck between agent demos and production isn't model capability — it's evaluation infrastructure. Zhou Yu, Columbia CS professor and co-founder of Arklex AI, argues most teams lack the testing tooling to catch edge cases, compliance failures, and conversation drift before deployment. Her cited figure: 95% of agents never leave the demo phase. Arklex's approach uses simulation-driven evaluation: synthetic user personas run thousands of conversation turns, surfacing failure modes manual QA misses. A key metric is trajectory entropy — measuring how consistently an agent handles unexpected inputs across multi-turn sessions, indicating whether behavior is predictable or brittle. The framework plugs into CI/CD pipelines, automating tests on every code push. Walmart's shopping agent is the proof point: in production for over a year using this method. The 48-minute QCon AI talk covers synthetic persona generation and self-learning feedback loops in detail. For enterprise teams, the takeaway is straightforward: production readiness requires a dedicated simulation testing layer, not ad-hoc manual evaluation. The techniques are implementable today.