One-take Creation, Flexible Referencing: Introducing Seedance 2.5

One-take Creation, Flexible Referencing: Introducing Seedance 2.5

Date

2026-07-31

Category

Models

Today, we are officially launching Seedance 2.5, the new-generation video creation model. Since the release of Seedance 2.0, we have noticed a shift in what users expect from video creation models: from merely generating a clip to completing a creative work. Building on the unified multimodal audio-video joint-generation architecture of Seedance 2.0, Seedance 2.5 centers on foundational generation and reference-based generation, delivering major breakthroughs in long-form storytelling, multimodal reference, and editing. Grounded in real-world use cases, it opens up greater creative imagination and control, and further unlocks productivity.

Key highlights include:

Up to 30 seconds per generation, with multi-round extensions: Seedance 2.5 can generate high-quality, 30-second audio-video clips in a single pass and supports multiple rounds of extension. It also improves shot transitions and scene changes for stronger continuity in longer videos, and delivers notable gains in image, audio, and motion quality, resulting in a more natural, polished visual quality than commonly seen in AI-generated video. As a result, users can produce high-quality multi-minute content with a consistent audiovisual language, bringing a complete story to life in one take.

Fully upgraded multimodal referencing: Users can now input up to 30 images, 10 video clips, and 10 audio clips as reference materials in a single pass. The model also strengthens a range of reference capabilities, including clay render, motion, and creative references, enabling it to better grasp the creator's intent and realize complex ideas that span multiple subjects, scenes, and shot changes.

More precise and stable editing capabilities: Seedance 2.5 offers timestamp-level control for targeted editing of audio and video content, notably improving efficiency and controllability. The model also enhances advanced editing features, such as green screen, camera perspective, and reference-based editing, to meet the rigorous demands of professional, complex fields like film and advertising.

With advancements in long-form storytelling, multimodal reference, and editing, Seedance 2.5 goes beyond longer single-pass video generation. The model better understands creative intent and delivers the journey from idea to finished video with greater control. Now, we'd like to invite you to watch a short creative film, produced end-to-end by Seedance 2.5.

Today, Seedance 2.5 is rolling out on Jimeng AI, Doubao Pro, and other platforms, with API access coming soon via BytePlus ModelArk. We invite you to give it a try and share your feedback.

Project homepage:

https://seed.bytedance.com/seedance2_5

Access:

Jimeng Web -> Video Generation -> Select Seedance 2.5

Doubao Pro -> Video Generation -> Select Seedance 2.5

30-second long-form storytelling with multi-round extensions: presenting complete stories in a single pass

Seedance 2.5 extends single-pass video generation from 15 to 30 seconds and further strengthens its storytelling in longer videos. Within 30 seconds, the model can organize multiple logically connected shots so that a story unfolds through setup, development, turning points, and resolution, rather than simply extending a single moment. For example, in a one-take clip of a singer's stage performance, the model portrays the full story of the singer interacting with staff in the dressing room, then walking through the backstage corridor, meeting the dancers, and stepping onto the stage with them for the performance, instead of only the moment of walking on stage.

T2V prompt: One-take handheld gimbal tracking shot. The camera slowly pushes in through a gap in a heavy red curtain and enters a warm-toned backstage dressing room. A young female singer, with her back to the camera, is adjusting her earpiece as a staff member reminds her it's time to go on. She turns toward the camera and starts singing citypop. The camera pulls back and tracks her as she passes through the curtain into a dim backstage corridor, interacting naturally with her dancers along the way; one staff member hands her a microphone. She and the dancers then step onto the stage, and the camera arcs around to the back, gradually revealing the red-and-black stage design, LED screens, spotlights, haze, and reflective floor. The camera finally pulls out to a wide shot of the arena, showing the packed audience, light boards, glow sticks, and cheering crowd, capturing the youthful, free-spirited climax of the concert.

Thanks to the model's multi-round extension capability, users can smoothly append subsequent shots to existing video outputs. Throughout the extension process, it maintains the consistency of main characters, environments, and narrative pacing. This allows users to output videos lasting several minutes at once, reducing the effort required to split clips, repeatedly splice footage, and fix transitions.

R2V prompt: Extend the video. Continue from the visuals and subjects in @Video 1 and generate another 30-second clip, keeping the character subjects, scene, visual style, and sound effects consistent. The little boy runs along the train carriage holding a soccer ball. When the subway stops, the side door opens and he immediately dashes out, with the male lead chasing after him. The two run across the platform and out onto the street, startling passersby and vehicles along the way. The male lead finally catches up and grabs him. The boy looks up, aggrieved. The male lead's anger slowly fades; he pats the boy's head and shows a helpless smile.

In terms of visual presentation, the model achieves smoother transitions between camera movements. The main subject remains stable across multiple cuts, and the audio and visuals remain in sync, resulting in highly coherent long-form videos. For example, in a Peking Opera scene, the camera executes a graceful circular pan following the lead actor's flowing sleeves, while the main subject and background remain entirely consistent. The swinging of the sleeves forms natural arcs in the air, closely adhering to real-world physics.

R2V prompt: 16:9 widescreen, cinematic texture, single continuous take, smooth camera movement, no cuts. Scene reference: @Image 4. 0–5s: Open with a close-up of the Overlord from @Image 2. The camera slowly circles his upper body and transitions into a medium shot. The Overlord spins and turns, his body and back flags sweeping quickly past the lens to form a natural occlusion, and the camera follows through to Consort Yu's side in @Image 1. 6–10s: The camera steadily circles Consort Yu in a medium shot from @Image 1, following her water sleeves through the arc. She raises her arm, flicks her wrist, unfurls the sleeves, and half-turns. She then draws the sleeves back, holds the pose, and looks sideways toward the Overlord. 11–20s: The male warrior from @Image 3 enters with an aerial flip. The Overlord takes center stage while the warrior advances and retreats on the opposite side in a combat exchange. Consort Yu stands slightly behind and to the side of the Overlord, weaving in water-sleeve movements to set softness against strength. The camera slowly pulls back from a medium-close shot of the warrior to a full stage view. At the end, all three face the audience and strike a synchronized Peking opera finale pose.

Additionally, to address the overly artificial look often seen in AI-generated videos, Seedance 2.5 systematically optimizes details such as object textures, skin and eye features, lighting, and color saturation. The model also minimizes uncontrolled occurrences in subtitles and background music, delivering final products that closely resemble the cinematic quality of live-action footage.

Comprehensive upgrades to multimodal reference, bringing greater control to complex creative tasks

Seedance 2.5 further strengthens its multimodal reference generation capabilities. It allows users to input up to 30 images, 10 video clips, and 10 audio clips as reference materials in a single pass. A larger volume and wider variety of references can better capture the user's intent, producing complex videos with more subjects, richer scenes, and more flexible camera work.

The model comprehensively understands elements such as visual composition, scenes, styles, characters, and props across all materials, applying them to the video generation process as instructed. Even in complex scenarios like multi-character shots or group storytelling, it can preserve the appearances and voices of multiple characters while keeping each subject's characteristics stable.

R2V prompt: A 30-second concert sequence in 16:9 landscape, with cinematic realism, authentic concert hall lighting and shadows, warm golden stage lighting, and the atmosphere of a formal classical concert. Use @Image 1 for the venue. Reference @Image 2 for the pianist. Reference @Image 3 for the cello. Reference @Image 4 for the violin. The lead vocalist must strictly follow @Image 5. Reference @Images 6 to 10 for the rest of the orchestra. Reference @Images 11 to 14 for the choir. Reference @Images 15 to 18 for the audience seating. The lead vocalist walks from center stage toward the front edge. The pianist is positioned by the piano. The orchestra is arranged on both sides and toward the rear. The choir stands at the back of the stage. Open with a high-angle wide shot of the full concert hall. The pianist strikes the keys, and the lead vocalist steps into the spotlight and begins singing. The camera naturally moves across the violin, cello, and orchestra as they perform together, with the violin feeling bright and the cello warm. In the latter part, the choir joins in. The lead vocalist briefly makes eye contact with front-row audience members, who respond with a smile and a slight nod. In the closing shot, the camera pulls back. The singing ends, and the audience joins in the applause.

Seedance 2.5 also enhances specific reference capabilities including clay render, motion, and creative referencing, giving users finer control over subjects, actions, and camera work in the frame. For instance, with clay render referencing, users can build a scene's spatial structure, character poses, motion paths, and camera angles using textureless 3D models. The model then uses this structure to generate the video, ensuring that the composition and blocking of complex shots closely match the creator's expectations. Additionally, Seedance 2.5 improves lighting control. By leveraging the spatial information from the clay render, it generates realistic lighting effects that follow physical laws, such as light source direction, color temperature, intensity, and shadow projection. This results in more natural light and shadow in the final output.

R2V prompt: Refer to @Clay Render 1 for camera movement, pacing, shot-size transitions, subject trajectory, and blocking. Refer to @Image 2 for character design, scene, materials, lighting, color, and fairy-tale atmosphere, and render the white model as a dreamy, warm, 3D animated short with a childlike fantasy feel. The story unfolds as follows: flight through a fantasy sky → mythical beasts flying alongside through a sea of clouds → a dive into the ocean → weaving through the deep with manta rays → passing through a mirrored rift in spacetime → picking stars from the cosmos → transforming back into the bedroom → father tucking in the blanket → the picture book closes and holds on the final frame.

More precise and reliable editing for higher creative efficiency

In video creation, users typically need to control pacing during the generation process and also refine details afterward. The exact second an action occurs, the precise timing of a camera cut, and whether a character's movement in a specific clip requires adjustment all profoundly impact the final result. More precise and reliable editing capabilities allow creators to accurately bring their ideas to life, improving efficiency and reducing the cost of repetitive generation.

Seedance 2.5 supports precise content editing via timestamps. During the generation phase, users can use prompts to control the narrative, camera perspective, movement, and overall rhythm for a specific time frame, aligning the output more closely with their creative intent. After generation, users can make targeted modifications to characters, actions, or plot elements within specific clips, all while maintaining continuity and realism before and after the edits.

Seedance 2.5 also elevates multiple editing features, such as green screen editing, camera perspective editing, and reference-based editing, to meet the rigorous demands of professional fields like film and advertising. In green screen editing, for example, the model can replace backgrounds and tell entirely different stories while keeping the main subject intact. Furthermore, it excels at rendering how the subject responds to the physical rules of the new environment. This includes the fluttering direction of clothes, the state of hair, gait rhythm, and lighting interaction, ensuring the subject blends harmoniously with the scene.

R2V prompt: Using @Video 1, render the green-screen background, obstacles, wardrobe, and supporting characters. 0–4s: outdoor training, replace the obstacles with rocks, bricks, tires, and wooden crates. 4–10s: locker room, friends offering encouragement. 10–15s: international match, replace the training poles with original defenders and a goalkeeper, and the protagonist scores. Overall photorealistic, cinematic quality.

R2V prompt: Edit @Video 1. Keep the characters, actions, and visual style unchanged. Adjust only the camera movement. A 15-second segmented camera plan: 0–4s, a micro-FPV move skims tightly past the pan, then follows the popping toast and whip-pans to the coffee; 4–7s, push in and track laterally along the rim of the pan, following the fried egg as it flips up and lands back in place; 7–11s, rapidly rise to a top-down view, then descend at a steady pace, sweeping across the plate and keys; 11–15s, use a handheld close-up to follow the hands with a fast lateral whip, then push in on the breakfast and pull back to a medium two-shot. Keep the entire sequence smooth, continuous, and stable.

Going deeper into broader industry scenarios, continuously exploring real-world value

As the model's capabilities evolve, Seedance 2.5 is reaching deeper into broader industry scenarios such as education and manufacturing. In education, the model has begun to enter real learning settings. For example, Seedance 2.5 can turn the historical context, characters, and storylines behind a lesson into more vivid and immersive visuals. It also helps teachers produce instructional videos more efficiently, turning abstract content — scientific principles, historical events, experimental procedures — into dynamic demonstrations. This not only lowers the barrier to producing educational materials but also allows for highly flexible content customization.

An example of Doubao Learning app's "Doubao Classroom" scenario

R2V prompt: Expressive Eastern painterly style. A street scene in Lin'an during the Southern Song dynasty. Several children run and shout through the bustling street, chanting, "I turn around, and there he is, where the lantern lights grow dim." The camera follows the children as they run, sweeping past the lively street. The camera then tilts up to reveal Xin Qiji from @Image 1. Xin Qiji turns his head, and in the distance stands a man among the fading lantern lights. The shot stays continuous throughout.

In sectors like industrial manufacturing, embodied intelligence, and autonomous driving, Seedance 2.5 is becoming integrated into highly specific production workflows. The model can generate high-quality synthetic video data that helps train robots' perception and manipulation skills. It is also being utilized for industrial simulations, process training, and equipment demonstrations. For autonomous driving, the model can simulate long-tail scenarios, such as extreme weather and complex traffic conditions, providing more diverse samples for system testing and training.

R2V prompt: Reference the camera work, composition, shot scale, spatial relationships, part positions, model structure, assembly order, and motion paths from @Clay Render 1. Reference the materials, lighting, color, reflections, and atmosphere from @Image 1, and turn the clay render into a high-end, photorealistic car assembly sequence.

Summary and looking forward

Seedance 2.5 marks a significant step forward in understanding and rendering the real world, elevating video generation from clip-level outputs to comprehensive creative workflows. At the same time, we recognize there is still room for improvement, particularly regarding the physical plausibility of complex motions and the stability of scenes involving interactions among multiple subjects.

Looking ahead, the Seed team will continue to explore more coherent storytelling, deliver a more intuitive generation and editing experience, and further deepen the model's grasp of real-world physics. We hope the Seedance models will become more vivid, more controllable, and better at understanding users' intent, helping more users express their creative ideas while continuing to explore and serve broader industry needs.