Multimodal references
Seedance 2.0 can use text, images, video clips, and audio as input. Treat each file as a job: identity, motion, setting, style, or timing.
Seedance 2.0 creates short AI videos from text, images, video references, and audio. Use it when a clip needs stable motion, camera control, and sound in one generation.
Upload only material you own or have permission to use. Review every result before public or commercial use.
Seedance 2.0 features
Seedance 2.0 is strongest when the prompt describes a full shot: what appears, what moves, how the camera behaves, and what the viewer hears.
Seedance 2.0 can use text, images, video clips, and audio as input. Treat each file as a job: identity, motion, setting, style, or timing.
Seedance 2.0 generates picture and sound in the same pass, so dialogue cues, room tone, music, and motion can be planned as one timed scene.
Seedance 2.0 is useful for action that needs weight: skating, running, falling cloth, water, contact, or a camera move through space.
Seedance 2.0 responds best to plain camera notes: locked shot, slow push, low tracking shot, overhead view, close-up, or wide establishing frame.
Example video slots
These cards are ready for case videos. Replace each empty videoSrc with a final URL later, then users can compare prompt, reference, and result without reading a long explanation.
Video slot 1
Replace videoSrc later
Prompt
A matte black coffee grinder on a kitchen counter, slow push-in, beans falling into the hopper, soft morning light, clear product shape, subtle grinding sound.
Use this Seedance 2.0 case to test product shape, label stability, reflections, and small moving parts.
Video slot 2
Replace videoSrc later
Prompt
A courier enters a rainy alley, checks a glowing package, then looks up as blue light reflects on the wall, two connected shots, handheld camera, rain ambience.
Use this Seedance 2.0 case to test connected shots, mood, camera rhythm, and sound inside one short clip.
Video slot 3
Replace videoSrc later
Prompt
Use the character image for appearance and the reference clip for dance movement, keep the same red jacket and hairstyle, stage lights, steady full-body frame.
Use this Seedance 2.0 case to test whether uploaded references keep separate roles instead of blending together.
A good case video needs to show what each upload controls. The viewer needs to see which image keeps the subject, which clip guides motion, and which audio cue affects timing.
Pick examples with contact, camera movement, or body motion. These clips reveal whether the model keeps weight, spacing, and direction instead of sliding through the scene.
Use at least one case where sound matters. Footsteps, rain, dialogue, or a product action lets users hear whether the audio belongs to the visible event.
Model details
The public Seedance 2.0 report describes a native multimodal audio-video model. Seedvid keeps the page practical: plan within the duration, resolution, and reference limits shown in the selected workflow.
Input types
Text, images, video clips, and audio
Duration
4 to 15 seconds in the technical report
Resolution
Native 480p and 720p in the technical report
Reference limits
Up to 9 images, 3 videos, and 3 audio clips on the open platform
Useful for
Short scenes, product shots, social clips, music visuals, and pitch work
Source material
Pick a front-facing product, person, room, or object image with the details you care about. Avoid heavy filters, cropped faces, and mixed subjects.
Use a short motion reference when the move matters more than the exact scene. A clean walk, turn, push, fall, or camera pan is easier to follow.
Upload sound when it changes timing. A voice line, beat drop, door knock, rain track, or machine hum helps the model place action in time.
Do not ask for a glossy ad, hand-drawn anime, documentary camera, and claymation in one clip. Pick one look before testing details.
Prompt formula
Seedance 2.0 works better when your prompt reads like a short director note. Give the model a clean timeline, then use references for details text cannot hold well.
Start with one visible event, not a whole film.
Name the subject, setting, action, camera, and ending frame.
Assign each reference a role: identity, motion, style, sound, or layout.
Review one weak detail, then revise one instruction for the next pass.
Use cases
Seedance 2.0 fits short, specific videos where reference control, sound, and motion are more useful than a generic animated image.
Seedance 2.0 can turn a product image and short prompt into a compact ad draft with motion, sound, and a clear hook.
Seedance 2.0 helps test a scene idea before a shoot, edit, or animation pass. Keep the story beat short and specific.
Seedance 2.0 fits social videos where a person, room, product, or visual style needs to stay consistent for a few seconds.
Seedance 2.0 can use a video or audio reference to guide rhythm, camera movement, and timing without writing every detail in text.
FAQ
Seedance 2.0 is ByteDance Seed's multimodal audio-video generation model. It creates short clips from text, images, video clips, and audio references.
The technical report describes direct generation from 4 to 15 seconds. For longer work, build several planned clips and edit them together.
Yes. Seedance 2.0 generates video and audio together. You can ask for dialogue, ambience, music cues, footsteps, weather, or silence.
Seedance 2.0 supports text, images, video, and audio. Use images for identity or style, video for motion, and audio for rhythm or sound direction.
Seedance 2.0 prompts work best as shot notes. Write the subject, action, location, camera, sound, and the detail that must stay unchanged.
Commercial use depends on Seedvid and provider terms. Use source material you own or have permission to use, and review likeness, logos, claims, and music.
Capability details were checked against the ByteDance Seed launch post and the Seedance 2.0 technical report. Product controls and availability can change.
Start with one shot, add references only when they solve a real control problem, then review the output before you publish.
Try Seedance 2.0