Seedance 2.0 on VisionX.
The four core video tiers, from prompt probe to native 4K master.
CINEMA TIER · RENDERED ON VISIONX
What Seedance 2.0 is on the roster
Seedance is the spine of VisionX video. The same model family runs four tiers, so a shot moves from a cheap probe to a mastered final without changing platforms, prompts, or Cast.
Learning and Iteration exist to be spent freely: test a prompt, find the blocking, throw the takes away. Standard is the production default. Cinema is the mastering tier, and it renders native 4K directly, with no upscale pass.
One family, draft to master
Look-dev on the cheap tiers carries straight up the ladder. The prompt that worked on Iteration works on Cinema.
The deepest reference support on the roster
Seedance 2.0 takes up to 9 reference images and accepts video input, which is what Cast identity and Reference Plates run on.
Native multi-clip extend
Standard and Cinema fuse up to 3 finished clips into one continuous video, so a 5-second hero moment becomes a scene.
What Seedance 2.0 actually does
Six things the Seedance tiers do that a text-prompt-only model cannot.
Twelve references, and each one has a job
Nine reference images plus video and audio slots feed a single Seedance 2.0 generation. Each asset carries a role — this face, that product, this camera move — instead of being blended into an average.
Multimodal referenceSound generated with the picture, not over it
Dialogue, ambience and effects come out of the same pass as the frames, so lip movement is shaped by the waveform rather than dubbed onto a finished clip afterwards.
Joint audio-videoSeveral cuts inside one render
A single generation can hold more than one shot, keeping one character, one location and one soundtrack across the cuts rather than stitching three clips together in an editor.
Multi-shotCopy the camera move from a clip you upload
Feed a reference video and the model reads its move — the push, the orbit, the whip — then applies that motion to your subject instead of to your footage.
Video-to-videoLip-synced dialogue across languages
Write the line in quotes and the character speaks it. English and Mandarin hold up best; short lines map to phonemes most cleanly.
Multilingual speechNative 4K on the mastering tier
Cinema renders 4K directly out of the model, so a hero shot never goes through an upscale pass that softens what the grade is meant to hold.
Delivery resolutionEvery variant, exactly as the studio runs it
These rows are read live from the engine registry and the pricing engine. What you see here is what the composer quotes.
Self-serve rates run as low as $0.085 per VX at scale. The composer quotes the exact VX for your shot, references included, before you run it.
How to write for Seedance 2.0
Seedance addresses uploads by tag. Naming the role is the difference between a reference the model uses and one it averages away.
Use [Image1] as the bottle. Keep the label sharp and unchanged. Use [Image2] for the set: wet black stone, one hard key light from camera left. Match the camera move in [Video1] — slow push, slight handheld drift. Take the room tone and pace from [Audio1]. Shot 1 — macro on condensation. Shot 2 — pull back, bottle enters frame. Shot 3 — hold, label centred.
Give every slot a role, not a mood
"Use [Image1] as the character's face" beats "in the style of [Image1]". The model resolves roles precisely and moods loosely.
Number your shots inside one prompt
Multi-shot is generated jointly, so listing the cuts gets you a sequence that holds continuity — not three clips you have to match afterwards.
Put spoken lines in quotes
Quoted text is read as dialogue and drives lip-sync. Keep it short and name the language for the cleanest result.
Look-dev cheap, master late
Find the blocking on Learning or Iteration, then send the prompt that worked up to Standard or Cinema. The ladder is the same model family, so the take carries.
What teams route to Seedance 2.0
The board lines that get routed to Seedance first.
Scroll-stopping social hooks
Vertical shots with a spoken opening line and ambience already in the mix — ready for Reels, Shorts and TikTok without a separate audio pass.
Product ads from a single packshot
Drop the pack into an image slot and a camera-move clip into a video slot. The label stays legible through the whole move.
Previz and storyboard motion
Turn boards into moving shots that hold the same character across cuts, so a director can read blocking before anything is booked.
Talking-head presenters
One front-facing portrait plus a script line gives a presenter whose mouth matches audio the model generated itself.
Music-led visuals
Load a track into an audio slot and the cuts and motion take their timing from it rather than from a fixed cadence.
One ad, several markets
Re-run the same shot with the line rewritten per language. Lip movement regenerates with the speech instead of drifting against a dub.
Seedance 2.0 against the alternatives
The short version against the two engines it gets compared with most. All three run on the same VisionX board, so you can settle it on your own footage.
| Capability | Seedance 2.0 | Kling 3.0 | Veo 3.1 |
|---|---|---|---|
| Reference input types | Text, image, video, audio | Text, image | Text, image |
| Audio generation | Same pass as the picture | Separate sync step | Same pass as the picture |
| Audio file as an input | Yes | No | No |
| Multi-shot in one generation | Yes | Partial | Partial |
| Native 4K without an upscale | Yes, on Cinema | Vendor-announced | No |
VisionX rows are read from the engine registry. Competitor rows describe the vendors’ own published capabilities and are current as of July 2026 — vendors ship fast, so re-check before quoting these in a deck.
Same Cast, same wallet, different look
Route a board line to Seedance 2.0, then run the identical shot on another family and compare takes. Identity lives in your Cast, not in any one model, so switching engines never strands the campaign.