SeedRealtime Audio-Visual Full-Duplex LLM Released: Toward Omni-Modal Natural Interaction
SeedRealtime Audio-Visual Full-Duplex LLM Released: Toward Omni-Modal Natural Interaction
Date
2026-08-05
Category
Models
Today, we are officially launching SeedRealtime, a native audio-visual full-duplex LLM.
As a key step toward omni-modal interaction, SeedRealtime uses a unified architecture to natively fuse audio, video, and text, enabling real-time interaction over continuous multimodal streams and delivering a brand-new "watch, listen, and speak" experience. It achieves three core breakthroughs:
Joint audio-visual understanding: Native support for the deep fusion of audio, visual, and temporal information. The model can resolve homophone ambiguity by drawing on the visual context, and it can accurately interpret temporal references in what it sees, uniting what is seen, heard, and said.
Proactive interaction: Continuous environmental awareness paired with the ability to speak up on its own. When it notices a change in the visual (such as the appearance of a key target), the model can offer a reminder unprompted, and it can weave tool calls into its responses, turning interaction from passive reaction into active collaboration.
Natural Conversational Timing: The model senses the user's conversational state and pacing in real time, chiming in, pausing, and responding at the right moments. It is also highly robust to interference, distinguishing bystander chatter and background noise so that it is not falsely triggered, keeping the conversation smooth and coherent.
End-to-end human evaluation shows that, compared with cascaded models, SeedRealtime reduces audio-visual conversational pacing issues by half: the model judges when to speak far more naturally, markedly reducing awkward breakdowns such as being cut off mid-sentence, responding sluggishly after a pause, or being falsely triggered by background noise and chatter. At the same time, the likelihood of completing a single conversation smoothly and fully has also improved significantly.
Currently, SeedRealtime has been fully rolled out, pioneering large-scale deployment of audio-visual full-duplex technology in the industry.
Native audio-visual interaction, toward omni-modal, watch, listen, and speak at once
Real-time audio-visual interaction has long been stuck between two paradigms. Cascaded systems chain together separate modules such as ASR, VLM, and TTS, introducing latency and information loss between stages. End-to-end models are more fluent, but many approaches still rely on an external VAD to decide turns, remaining essentially a half-duplex, one-question-one-answer interaction.
SeedRealtime's core breakthrough is unifying sound, vision, timing, and expression within a single end-to-end model. Rather than listening in full, then looking, and finally answering, it lets perception, understanding, decision-making, and expression run in parallel over continuous audio-visual streams, so that what is heard and what is seen jointly inform every real-time judgment.
The first challenge is conversational rhythm. Speech naturally contains pauses, which can help indicate when it is time to respond. Video, by contrast, is always on and constantly changing. The model must continuously understand what is happening on screen without jumping in too often because of background noise; instead, it keeps deciding which object to focus on, whose voice to listen to, and whether it should respond at this moment.
The deeper challenge lies in joint audio-visual modeling and temporal alignment. When a user says "how do I do this," the model must combine the current scene, gestures, gaze, and prior actions to determine what "this" refers to; when it encounters a homophone or unclear speech, it also has to use the visual scene to disambiguate. Only by modeling sound, visual, and temporal information together can the model truly connect what it sees, hears, and says. Once watching, listening, and speaking are brought into one real-time decision-making system, the model no longer merely waits for a question, but can keep tracking changes in the scene and speak up at the right moment.
Next, let's look at how SeedRealtime performs in real-world settings through seven concrete examples.
Joint audio-visual understanding, understanding the scene, identifying the subject
SeedRealtime can align and jointly understand visual, audio, and temporal information in crowded, noisy, open real-world settings.
Four friends are having dinner: plenty of people and overlapping chatter. As the user introduces everyone present one by one, SeedRealtime matches names to faces by appearance, recognizing light-haired Qiqi and bespectacled Julia and even greeting Lele on its own. It keeps each person's voice tied to their identity throughout the ensuing conversation.
The group then starts talking about travel, one after another: one wants to take photos at the beach and visit the aquarium, another dreads the heat and the effort, and yet another is allergic to seafood. SeedRealtime can tell which line comes from whom and, based on the group’s conversation, provide a travel plan that takes everyone’s needs into account. Recognizing people, telling voices apart, and understanding each person’s needs are all handled in real time by a single model within the same conversation.
At a Sichuan restaurant, a foreign diner is stumped by a Chinese-only menu. SeedRealtime identifies the dishes straight from the scene, recommends them in English, and even draws on cultural background to explain "why there is no fish in fish-fragrant pork (yuxiang rousi)" and "how century eggs are made."
As the server sets down a dish and casually says it "goes great with rice," the model understands the line in the context of the dish on screen and translates it for the guest. Visual information does not need to be turned into text first; it is aligned with speech within the same model, so the model can answer based on what is right in front of it.
Proactive interaction, continuously observing, stepping in at the right moment
The model decides for itself whether it needs to step in, based on the ongoing changes in the live scene, rather than waiting for the user to speak first every time.
While touring an exhibition at the Hebei Museum, the user says, "Remind me when you see the gold-and-silver-inlaid bronze tiger-devouring-a-deer screen stand." As the camera keeps moving, SeedRealtime watches the scene, and when it pans across that piece, it speaks up with a reminder.
It then draws on visual details to explain treasures such as the gold-and-silver-inlaid bronze square table base with four dragons and four phoenixes and the Changxin Palace lamp, along with crafts including casting, gold-and-silver inlay, and welding. Holding a task in context and delivering a reminder the moment the target appears is exactly what defines continuous visual perception and proactive interaction.
While operating a complex espresso machine, the model can correct mistakes and give feedback in real time based on changes in visual state.
When it sees the user pour whole coffee beans straight into the portafilter, the model quickly points out: "You can't pour the beans in directly; you need to grind them into a fine powder first." After extraction, based on a visual read of the crema's color and the volume of liquid in the cup, it proactively suggests "shortening the extraction time by 2 to 3 seconds next time." Without the user asking step by step, the model can weigh in based on the actual situation.
While studying the ResNet paper, the model uses the network diagram to explain how skip connections mitigate vanishing gradients. When the user says, "Keep an eye out for me and remind me when you reach the training-parameters part," it keeps watching the screen as the pages flip quickly, accurately spots the "3.4 Implementation" section, and pauses on its own, then reads out key training settings such as learning rate, momentum, and weight decay. Spotting a target within a continuous stream of frames and stopping on its own demonstrates genuine active collaboration in interaction.
Natural conversational timing, deciding when to speak amid interference
Turn-taking is no longer handed off to external VAD rules; instead, SeedRealtime decides continuously based on multimodal information, never hesitating to chime in when appropriate, and never interrupting when it's better to stay quiet, maintaining a more natural rhythm even in complex environments.
Beijing Daxing Airport is crowded and noisy. When a companion casually mentions "Old Li's flight" in passing chatter, the model is not triggered to reply by this unrelated remark. When the user actually asks, even though the flight information has already scrolled off screen, it can still draw on the departure-board information it saw earlier to give the real-time arrival time, and it goes online to provide the baggage-carousel location.
As the two then walk and reminisce, the model keeps watching the scene in front of it, and when it spots a ride-hailing sign, it naturally chimes in with directions to the pickup point. A noisy environment is no longer merely interference to be filtered out; key information from casual chatter can also be properly retained and drawn on when needed.
With mom busy with something else, she asks SeedRealtime to help her daughter learn animal names in English. In the background, dad is on the phone and voices are constant, yet the model is not led astray by this unrelated conversation, consistently following where the girl points to correct her pronunciation in real time and make up example sentences. The model does not hesitate when it should speak, and does not drift when there is interference.
Summary and Outlook
The launch of SeedRealtime marks a key step forward for audio-visual interaction. By natively fusing audio, video, and text within a unified architecture, the model can continuously understand the scene, judge the rhythm, and respond across multimodal streams. With that, "watch, listen, and speak" truly becomes an everyday interactive experience.
Yet audio-visual communication in the real world is complex and continuous: backgrounds are noisy, voices overlap, scenes change, and users pause, add on, or interrupt at any time. To meet these real-world challenges, we will continue to push forward on several fronts:
Lower latency and more natural conversational timing: Continuously compressing the end-to-end "hear, understand, respond" latency so that the subtle rhythms of real conversation, such as interrupting, jumping in, backchanneling, and pausing, are reproduced naturally, not just faster, but better timed.
More proactive perception and decision-making: Building on continuous understanding of visual and sound to proactively judge when to remind, when to add information, and when to hand the user exactly the information they need, so that the model not only "speaks when appropriate," but also "acts when appropriate."
More robust in complex multi-person scenarios: Reliably identifying who is speaking, what they are looking at, and who to respond to in noisy environments, with multiple people in frame and multi-party conversations, bringing these capabilities into more real-life and work scenarios.
From "able to converse" to "able to act": Connecting tool calls with the real world so that the model can not only communicate and observe, but also help users complete lookups, bookings, and tasks, turning real-time multimodal understanding into concrete action.
We hope AI interaction will no longer be limited to turn-based question answering, but instead understand context in continuously changing real-world situations, grasp timing sensitively, and offer help at the right moment.