ByteDance Seed launches Seedance 2.5: 30-second takes and multimodal referencing
The new video model generates audio-video clips of up to 30 seconds in a single pass, accepts up to 30 images and 10 video clips as references, and adds timestamp-level editing.
ByteDance’s Seed team has officially launched Seedance 2.5, a new generation of its video creation model. The announcement, published on the team’s blog, says the release builds on the unified multimodal audio-video joint-generation architecture introduced with Seedance 2.0 and centers on two things: foundational generation and generation guided by references.
Longer takes, produced in one pass
Seedance 2.5 generates high-quality audio-video clips of up to 30 seconds in a single pass, double the 15 seconds of its predecessor, and supports multiple rounds of extension so users can append further shots to an existing output. The team reports gains in shot transitions and scene changes as well as in image, audio and motion quality, which together allow multi-minute videos to be assembled with a consistent audiovisual language — a complete story in one take instead of a sequence of stitched clips. One published example follows a singer from a dressing room through a backstage corridor and onto the stage as a crowd cheers, in a single continuous shot.
Up to 30 images, 10 videos and 10 audio clips
The model accepts up to 30 images, 10 video clips and 10 audio clips as reference material in a single pass. The team says a larger and more varied set of references captures the creator’s intent better, and that the model understands composition, scenes, styles, characters and props across the supplied materials, preserving the appearance and voices of several characters even in group scenes. Specific reference modes have been strengthened too: with clay render referencing, users define spatial structure, character poses, motion paths and camera angles using untextured 3D models, and the model follows that structure and derives physically plausible lighting from it.
Editing at the level of timestamps
During generation, prompts can control the narrative, camera perspective, movement and rhythm for a specific time frame; after generation, characters, actions or plot elements inside a clip can be modified while continuity and realism are maintained. The release also improves green screen editing, camera perspective editing and reference-based editing. In green screen work, for example, the model is said to render how a subject responds to the physics of the new environment — the direction clothes flutter, the state of hair, gait rhythm and lighting interaction — a capability the team frames around the demands of film and advertising.
Availability and acknowledged limits
Seedance 2.5 is rolling out on Jimeng AI, Doubao Pro and other platforms, with API access promised soon through BytePlus ModelArk. Beyond creative work, ByteDance points to education — turning historical events, scientific principles and experimental procedures into demonstrations, including in the Doubao Classroom scenario — and to industry, where the model is used for synthetic training data, robot perception and manipulation skills, and long-tail autonomous driving scenarios such as extreme weather and complex traffic. The team also states remaining limitations plainly: the physical plausibility of complex motions and the stability of scenes involving several interacting subjects.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.