Lightweight Generalized DeepFake Face Detection with WAVIE: Wavelet Augmented Vision Intermediate Embeddings

| Source: arXiv AI

Tags: deepfake-detection, CLIP, wavelet, FaceForensics, content-moderation, digital-forensics

WAVIE freezes a CLIP backbone and adds a lightweight Daubechies-6 wavelet module to intermediate transformer embeddings, achieving AUROC 0.852 on Celeb-DF-v2 when trained only on FaceForensics++ — outperforming several cross-dataset generalization baselines at low parameter cost.

Details

Deepfake detectors trained on known manipulation methods fail badly on unseen forgery pipelines — a critical gap as synthetic media proliferates. WAVIE (Wavelet Augmented Vision Intermediate Embeddings) tackles this by freezing a CLIP ViT backbone and learning only a compact wavelet-augmentation module on top of intermediate transformer embeddings. The module applies a three-level Daubechies-6 (db6) discrete wavelet transform, refines the low-frequency branch, preserves high-frequency branch, reconstructs features via inverse DWT, then classifies. Trained only on FaceForensics++, WAVIE reaches AUROC = 0.852 on Celeb-DF-v1, 0.852 on Celeb-DF-v2, and 0.831 on WildDeepFake (WDF) — outperforming several state-of-the-art cross-dataset baselines. The frozen CLIP backbone makes training fast and GPU-efficient. Ablation studies confirm both the wavelet module and intermediate-feature aggregation are necessary; neither alone achieves comparable cross-dataset performance. Accepted at IEEE SMC 2026. Directly relevant to content moderation, digital forensics, and trust-and-safety teams. The gap between 0.831–0.852 AUROC and practical deployment thresholds should be noted — false negative rates in adversarial real-world conditions may be higher than benchmark results suggest.