CRAF: Cross-View Residual-Aware Fusion for Deepfake Speech Detection

| Source: arXiv AI

Tags: deepfake-detection, speech-synthesis, ASVspoof, self-supervised-learning, Kimi-Audio, audio-AI

CRAF fuses self-supervised acoustic models with Auditory Large Language Models via residual-aware cross-view attention to detect deepfake speech, achieving 5.96% EER on ASVspoof 5—improving generalization to spoofing attacks not seen during training.

Details

Deepfake speech detection fails to generalize to novel spoofing attacks not in the training distribution. CRAF addresses this by combining two complementary representations: self-supervised learning (SSL) models that capture fine-grained acoustic details, and Auditory Large Language Models (ALLMs) that encode higher-level contextual information. Rather than concatenating these views, CRAF uses ALLM-guided cross-view attention to enrich SSL representations, then explicitly disentangles ALLM-explainable information from SSL-specific residual information. The residual is refined through adaptive gating before final fusion. Evaluated on ASVspoof 5—the current standard benchmark for anti-spoofing—CRAF with Kimi-Audio as the ALLM backbone achieves an Equal Error Rate (EER) of 5.96% and a minimum Detection Cost Function (minDCF) of 0.1192. The paper focuses on system design and generalization; direct comparison to state-of-the-art baselines on ASVspoof 5 is not prominently featured in the abstract. The work is practically relevant for voice authentication systems, fraud detection pipelines, and content moderation where voice cloning tools are creating increasingly realistic synthetic speech.