One Capital Letter Was Silently Breaking My AI Support Bot, and It Wasn't in the New Model

| Source: Towards Data Science

Tags: OpenAI, LLM testing, Structured Outputs, Weave, JSON schema, model regression

A regression test on an OpenAI production support bot revealed the in-production model silently capitalized one JSON field ('Request_refund' instead of 'request_refund') on every refund query — breaking downstream code — while the newer candidate model formatted correctly every time, showing that accuracy benchmarks alone miss format regressions.

Details

The author built a bank support triage assistant that classifies incoming customer messages into categories (lost card, refund request, wrong charge) and returns structured JSON for automated routing. They ran the same 47 real banking messages through three OpenAI model versions: an older model, the production model, and a newer candidate.\n\nThe production model passed accuracy tests but failed a different way: it consistently capitalized the intent field as 'Request_refund' instead of 'request_refund' on every refund-related query. Because the downstream routing code does an exact string match, this silent failure broke all refund routing — while accuracy metrics showed nothing wrong.\n\nThe newer candidate model used OpenAI's Structured Outputs feature and formatted correctly in all 47 cases. The author's key insight: standard benchmark scores measure category accuracy, not output format fidelity. A model that picks the right category but misspells the field name is useless in a production pipeline.\n\nThe article uses Weave (Weights & Biases) to build the regression test harness. The pattern demonstrated — testing the exact JSON schema your production code depends on, not just classification accuracy — is directly applicable to any team migrating between model versions.