Disentangling Topology and Diversity in Multi-Agent LLMs for Multilingual Low-Resource Emotion Detection

| Source: arXiv AI

Tags: multi-agent LLMs, Qwen, Llama, QLoRA, emotion detection, multilingual NLP

A controlled 2×3 study of multi-agent LLM configurations finds that how agents are differentiated — learned QLoRA specialization, role prompting, or sampling — matters more than how they are connected (parallel vs sequential), with specialized parallel agents reaching 52.94 Macro-F1 on multilingual emotion detection.

Details

Multi-agent LLM systems are typically designed by intuition, with topology (how inference calls connect) conflated with diversity (how agents differ). This paper disentangles the two in a controlled experiment crossing parallel aggregation and sequential refinement with three diversity sources: stochastic sampling, role prompting, and learned QLoRA specialization — under a fixed three-call budget. Using Qwen2.5-14B-Instruct and Llama-3.1-8B-Instruct on nine-language low-resource emotion detection, the results are clear: parallel learned specialization achieves 52.83-52.94 Macro-F1, exceeding seven-call self-consistency baselines with only three calls. Sequential refinement helps stochastic and role-prompted agents but not learned specialists. The topology preference depends directly on the diversity source. A notable depth-wise finding: later learned specialists can overwrite correct early predictions in sequential chains, though the aggregate effect is backbone-dependent. Qwen shows stronger specialization gains than Llama (2.83 vs 0.17 point advantage). The practical implication for multi-agent system designers: invest in specialization quality before optimizing topology. How agents are differentiated drives more performance variation than whether they run in parallel or sequence.