AV-SafetyBench: A Safety Benchmark for Text-to-Audio-Video Generation
| Source: arXiv AI
Tags: text-to-video, AI safety, multimodal, content moderation, benchmark, T2AV
AV-SafetyBench is the first safety benchmark for text-to-audio-video generation, covering 5,200 prompts across 13 categories — finding that 41.6-48.3% of unsafe outputs from T2AV models are missed entirely by video-only evaluation, and 87.5% of cross-modal harms emerge only from joint audio-video interpretation.
Details
Choi et al. introduce AV-SafetyBench, targeting text-to-audio-video (T2AV) models that jointly generate synchronized video, speech, sound effects, and ambience from a single text prompt. Existing safety benchmarks evaluate video or audio in isolation, systematically missing harms that emerge only from the combination. The benchmark comprises 5,200 manually reviewed prompts organized into a four-axis, 13-category taxonomy. Evaluation uses three views — Full-AV, Video-Only, and Audio-Only — and assigns unsafe outputs to one of four risk sources. Across five open-source T2AV models tested, Full-AV Unsafe Rates range from 25.1% to 49.4%. Critically, 41.6-48.3% of harmful outputs that receive a risk-source assignment are Audio-Only or AV-Joint cases — meaning video-only safety filters structurally miss nearly half the problem. Cross-modal harm is especially concentrated: in the Cross-Modal Harm Emergence category, AV-Joint cases account for 87.5% of unsafe outputs with an assigned risk source. This has immediate implications for anyone building content moderation pipelines for multimodal generative models — single-modality filtering is insufficient by design.