ER-EDF: A Psychology-Grounded Emotion Regulation Framework for Speech Empathetic Dialogue Generation in Large Audio-Language Models

| Source: arXiv AI

Tags: empathetic-AI, audio-language-models, emotion-AI, dialogue-systems, multimodal

ER-EDF, a model-agnostic framework for large audio-language models, explicitly separates emotion perception from emotion regulation in spoken dialogue — consistently improving empathetic response quality across 5 LALMs where prior systems simply mirror user affect.

Details

Current large audio-language models largely treat emotion as a direct conditioning signal: if the user sounds distressed, the model produces distressed-sounding responses. ER-EDF argues this confuses affect mirroring with genuine empathy. Grounded in the Perception-Action Model and emotion regulation theory from psychology, ER-EDF decouples two processes: Perception (tracking the user's emotional state from speech) and Regulation (determining how that emotional state should guide response generation). The regulation step allows the model to produce calibrated support rather than matching the user's affect. The framework is model-agnostic and tested across five LALMs on two datasets, with both automatic and human evaluations showing consistent improvement in empathetic response quality. The authors also construct a spoken empathetic dialogue dataset and introduce empathy-aware evaluation metrics beyond lexical matching. This is early-stage research in a field where ground truth is difficult to define, but the psychological grounding distinguishes it from prior empirical approaches.