When to Use One Model and When to Use a Team of Agents
| Source: Towards Data Science
Tags: multi-agent systems, Codex, Claude Code, agent orchestration, AI architecture, reasoning models, LLM infrastructure
A practitioner details how single reasoning models fail on complex multi-task AI work — by averaging contradictions or silently dropping tasks — and shows how 5 specialist agents with a coordinator that names disagreements (instead of synthesizing them) produced more reliable results for infrastructure capacity planning.
Details
Single reasoning models fail on complex, multi-faceted work in two expensive ways: they average away contradictions between evidence sources, or they silently drop tasks when context grows long — both producing output that looks correct but is not. The author discovered this during infrastructure capacity migration: a 14-day traffic snapshot missed a monthly reconciliation job, causing pipeline latency to jump from 10 to 60 minutes after cutover. The solution: 5 specialist agents, each receiving only the evidence slice for its own task. Specialists return structured typed findings — claim, evidence, confidence, freshness, and a verdict of ok/warn/hold — rather than prose. Codex handles quantitative inference (aggregating flow logs over full business cycles, weighting dependency edges by connection count), while Claude Code handles qualitative tasks (reading stale runbooks, interpreting documentation). A coordinator surfaces contradictions between agents rather than resolving them into a consensus. This architecture converts a hard prompt-engineering problem into a smaller coordinator design problem. The key insight: naming disagreements between specialist outputs is more reliable than asking a single model to synthesize contradictory evidence. Evidence freshness matters too — a verdict from a two-year-old runbook carries different weight than one from a flow log captured yesterday, and the coordinator must track this explicitly.