
More stable complex motion
Create sports, dance, performance, and multi-character interaction with stronger temporal continuity and more plausible physical behavior.
Sign in to claim 100 free credits and create with Seedance 2.0
Seedance 2.0 understands mixed creative references and generates picture and sound together, so your prompt can describe more than a single shot.
Combine text, images, video clips, and audio references
Generate high-quality multi-shot audio-video sequences
Mix images, video, and audio within the available per-type limits
Next-generation video creation
Seedance 2.0 is ByteDance Seed's native multimodal audio-video generation model. Its unified architecture accepts text, image, video, and audio inputs, then uses those references to guide composition, performance, camera movement, visual effects, and sound.

Use a character image, location clip, motion reference, soundtrack, and written direction together instead of forcing every idea into one prompt.
Create dialogue, sound effects, ambience, and music alongside the visuals for a more coherent audio-visual result.
Build scenes with multiple subjects, physical interaction, detailed action, and more stable motion over time.
Reference framing, shot size, camera movement, lighting, performance, rhythm, and transitions to shape how the story is told.
Four official showcases move from sunlit sport to crowd choreography, controlled balance, and intimate dance—each testing a different kind of temporal consistency.
A low, wide-angle camera tracks a fast rally while the ball, athlete, horizon, and lens distortion remain coherent.
Seedance 2.0 turns familiar filmmaking decisions into instructions the model can follow.

Assign each uploaded asset a role. An image can define the subject, a clip can guide camera movement, and an audio file can establish timing or atmosphere.

Describe the opening, progression, cut points, and final beat. Seedance 2.0 can create up to 15 seconds of connected audio-video with multiple shots.

Continue a video, revise selected creative elements, or preserve the core subject while changing the surrounding direction.
Built for usable shots
The model focuses on the parts of AI video that often decide whether a result can move forward: motion, consistency, instruction following, and audio-visual fit.

Create sports, dance, performance, and multi-character interaction with stronger temporal continuity and more plausible physical behavior.

Give detailed direction for subject actions, staging, lighting, lenses, camera paths, pacing, sound, and transitions.

Generate dialogue, effects, ambience, and music in context with the image instead of treating sound as a separate afterthought.
Choose with intent
Start with the job, not the version number. Compare quality, turnaround, workspace limits, and ideal use cases before spending credits on a full run.
How it works
Start with a simple prompt or build a detailed reference package. You stay in control of how much direction the model receives.
Describe the subject, action, setting, camera language, visual style, pacing, and sound you want.
Upload images, video clips, or audio and explain what each reference should control.
Set the aspect ratio, resolution, length, and audio option for your output.
Review the result, then adjust the prompt or references to make the next version more precise.
A flexible model for early concepts, finished social content, and production-ready creative exploration.
Explore shot design, blocking, action, atmosphere, and story rhythm before a full production.
Prototype polished campaign concepts with controlled materials, lighting, camera moves, and sound.
Create vertical stories, visual hooks, stylized transitions, and sound-rich short videos.
Bring character designs, environments, action references, and art direction into a moving scene.
Answers to common questions about Seedance 2.0 inputs, outputs, audio, references, and access.
Seedance 2.0 is ByteDance Seed's native multimodal audio-video generation model. It accepts text, image, video, and audio inputs and can use them as references for video generation, extension, and editing.
Yes. Seedance 2.0 generates audio and video together. Depending on the prompt and references, the result can include dialogue, sound effects, ambience, music, and dual-channel audio.
Seedance 2.0 supports outputs from 4 to 15 seconds. It can generate a connected multi-shot sequence within that duration.
You can combine natural-language instructions with images, video clips, and audio clips. This interface accepts up to 12 reference files in total, with per-type limits of 9 images, 3 videos, and 3 audio files. References can guide the subject, composition, action, camera movement, visual style, effects, timing, or sound.
Yes. The model supports reference-guided video editing and controllable video extension. Clear instructions about what to preserve and what to change usually produce the most useful direction.
No. This is an independent creative interface that provides access to supported AI models. Seedance is a model developed by ByteDance Seed; this site is not affiliated with or endorsed by ByteDance.
Describe the subject and action first, then add setting, shot size, camera movement, lighting, pacing, visual style, and sound. When you upload references, state exactly what each file should control.
Subscribe for ongoing AI video creation or choose a one-time credit pack for flexible project work. Subscription plans can be canceled anytime.
Cancel anytime
What's included
Everything in Starter, plus
Everything in Pro, plus
Bring your prompt and references together, direct the motion and sound, and generate your next cinematic video.