Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering

| Source: Apple ML Research

Tags: Apple ML Research, video captioning, evaluation metrics, vision-language models, ACL, multimodal, VLLMs

Apple researchers introduce CapQuiz at ACL 2026, a reference-free video captioning benchmark that tests quality by measuring whether captions help answer human-verified multiple-choice questions — outperforming BLEU/SPICE-style metrics on human judgment correlation and exposing systematic blind spots in current VLLMs.

Details

Evaluating video captioning is harder than it appears. Existing metrics — CIDEr, SPICE, and similar — compare generated captions against ground-truth references, creating a systematic bias: a valid caption that describes different but equally important visual elements gets penalized for lexical mismatch, and high-quality descriptions of underrepresented visual details are treated as errors. Apple researchers published CapQuiz at ACL 2026 to replace this paradigm. Instead of reference matching, CapQuiz asks: does this caption contain enough correct information to answer specific questions about the video? The benchmark includes human-verified multiple-choice questions organized into 10 types spanning descriptive and inferential categories, across 24 diverse video domains. The team defines two sub-metrics: CapP (precision — does the caption assert only true claims about the video?) and CapR (recall — does it cover the salient visual information?). CapF1 combines both into a single score. Experiments demonstrate significantly better correlation with human judgments compared to existing metrics. The practical implication is meaningful: if evaluation metrics are miscalibrated, model optimization targets are wrong too. CapQuiz provides a cleaner signal for tracking actual visual language model quality and a more informative diagnostic — separating accurate-but-incomplete from inaccurate-but-verbose captions.