How Many Labeled Examples Does a Text Classifier Actually Need? I Measured It.

| Source: Towards Data Science

Tags: TF-IDF, text classification, logistic regression, LLM vs classical ML, data efficiency, support ticket routing

A controlled experiment with 70 synthetic support tickets shows TF-IDF + logistic regression reaches 60% accuracy with only 10 labeled examples per class at sub-millisecond inference — posing a direct cost argument against defaulting to LLM APIs for every classification task.

Details

The article runs a deliberately minimal experiment: classify support tickets into five buckets (billing, tech bug, feature request, account access, general) using TF-IDF + logistic regression at three data volumes (2, 5, and 10 labeled examples per class) and measure accuracy on a fixed held-out test set of 20 tickets. The results follow a familiar diminishing-returns curve. Going from 2 to 5 examples per class gains 15 accuracy points (40% → 55%); going from 5 to 10 gains only 5 more (55% → 60%). Macro-F1 rises from 32.9% to 59.0% across the same range. Inference time is 0.005–0.008ms per ticket — essentially free. The practical framing is honest: if you have zero labels, an LLM is your only option. If you have 10 per class, a classical baseline already gets you to 60% with no API cost, no network latency, and no recurring fee. The article does not claim the classical approach beats LLMs — it measures exactly where it stops being negligible. One important caveat: the 70-ticket dataset is synthetic (author-written), so performance figures may not transfer to real noisy queues. The experiment is illustrative, not a benchmark against a representative distribution.