RA-CoA: Training-free Fashion Image Captioning via Retrieval-Augmented Chain-of-Attributes
| Source: arXiv AI
Tags: vision-language models, e-commerce AI, retrieval-augmented generation, image captioning, fashion AI, zero-shot learning
RA-CoA improves fashion product caption quality by 26.3% on METEOR score without model fine-tuning, using retrieval from a product knowledge base to ground attribute-level reasoning in any frozen VLM—directly applicable to e-commerce catalog automation. Accepted in TMLR, code public.
Details
Fashion image captioning requires fine-grained attribute recognition—neckline types, closure mechanisms, graphic patterns—that general-purpose vision-language models frequently miss or hallucinate. Fine-tuning on rapidly changing fashion inventories is expensive and hard to maintain at scale.\n\nRA-CoA (Retrieval-Augmented Chain-of-Attributes) solves this training-free: it disentangles captioning into two stages. First, relevant attribute sets are retrieved from a product knowledge base. Then, attribute-level reasoning generates the final caption by working through each retrieved attribute. The method is model-agnostic and operates on any frozen VLM without modification.\n\nAcross diverse VLM families and prompting paradigms, RA-CoA achieves an average 26.3% METEOR score gain over zero-shot captioning baselines. The approach is particularly well-suited to e-commerce catalog automation where domain-specific accuracy is essential and inventory changes too rapidly for continuous fine-tuning. Accepted in TMLR. Code is publicly available.