ByteDance Seed Introduces SeedRealtime: a Native Audio-Visual Full-Duplex LLM That Watches, Listens and Speaks in One Model

| Source: MarkTechPost

Tags: ByteDance, SeedRealtime, multimodal AI, full-duplex, audio-visual LLM, Doubao, real-time AI

ByteDance's Seed team released SeedRealtime, a native audio-visual full-duplex LLM that processes continuous audio, video, and text in a single end-to-end architecture — with turn-taking handled inside the model itself rather than by an external voice-activity detector.

Details

SeedRealtime unifies audio, video, and text into one model rather than chaining separate ASR, vision, and TTS modules — the current industry standard for real-time multimodal AI. Perception, understanding, decision-making, and speech generation run in parallel inside the model, with turn-taking decisions also internalized rather than delegated to an external voice-activity detector (VAD).\n\nByteDance demonstrates four notable behaviors in the published demos. Identity binding across modalities: at a group dinner, the model matches names to faces as people are introduced and attributes conflicting travel preferences to the correct speaker. Proactive speech from a held instruction: at a museum, the model watches a camera feed and speaks unprompted when a requested exhibit appears. Correction from visual state: watching an espresso workflow, the model interrupts when it spots an error without being asked. Interference suppression with off-screen memory: at an airport, the model ignores unrelated chatter while retaining departure-board information that has already scrolled out of frame.\n\nSeedRealtime is live inside Doubao, ByteDance's consumer AI assistant. However, no technical report, parameter count, open weights, or API has been published — external teams cannot deploy it. The model serves as a validated reference architecture and a raised benchmark for real-time audio-visual products.