Seedance 2.0 Stylish Video Guide: Multi-Reference Images and Prompts

Build separate references for identity, wardrobe and props, environment, and the opening keyframe, then connect them with a precise Seedance 2.0 video prompt.

AIZIGOO
Seedance 2.0 Stylish Video Guide: Multi-Reference Images and Prompts

Stylish Seedance 2.0 videos do not come from stacking more cinematic adjectives into a prompt. The more reliable approach is to separate character identity, wardrobe and props, environment and lighting, and opening composition into distinct references—then state exactly what the model should borrow from each one.

ByteDance Seed's official February 12, 2026 launch says Seedance 2.0 uses a unified multimodal audio-video architecture that accepts text, images, video, and audio. It can reference up to nine images, three video clips, and three audio clips. The company's own example explicitly assigns a character, scene, and props to different images. Dreamina's current tutorial similarly instructs creators to upload assets in Multiframes mode and call them by @AssetName.

More inputs are not automatically better. Conflicting faces, wardrobes, lighting, and compositions make the preservation target ambiguous. This guide creates four references that can be used together, plus one finished thumbnail, to show how fewer references with clearer responsibilities can be easier to control.

This guide reflects public ByteDance Seed and Dreamina materials checked on July 28, 2026. Interface labels, regional access, plans, credits, and input limits may change. The published benchmarks and quality assessments are developer-reported and do not guarantee identical performance for every subject or scene.

Design a reference-role map before the final frame

Our example is a fashion film in which an adult performer with a short silver bob, crimson coat, and translucent umbrella walks through a rain-soaked brutalist concourse and pivots toward camera. The four references deliberately avoid repeating every kind of information.

ReferenceWhat it locksWhat it leaves out
Image 1Face, body, hair, full wardrobeBackground, action, camera
Image 2Umbrella, fabric, gloves, metalNew identity or location
Image 3Architecture, light, rain, reflectionsCharacter, outfit, prop
Image 4Opening frame, lens feel, gradeLater action and ending

A single all-in-one reference looks convenient, but every change can disturb the entire design. Role separation lets you replace the location without redefining the face, or change the camera route without rebuilding the wardrobe.

Image 1: treat the character anchor as an identity record

A useful character reference is not the most dramatic poster. Choose a neutral image where the face, hair, full silhouette, hands, outfit, and footwear are readable. Keep the background simple and avoid motion blur so the model can distinguish fixed traits from cinematic choices.

Adult performer character anchor with a short silver bob, crimson coat, black bodysuit, and chrome boots

Image-generation prompt

Use case: stylized-concept
Asset type: landscape reference image for a Seedance 2.0 multi-reference fashion-film workflow; character identity anchor
Primary request: Create a premium cinematic character anchor image of one original adult female performance artist whose identity, silhouette, outfit, and materials can be reused consistently in AI video generation.
Scene/backdrop: minimal dark graphite studio with a softly illuminated floor, subtle atmospheric haze, no scenery and no distracting props.
Subject: one clearly adult woman, mid-to-late twenties, distinctive short silver bob with a blunt fringe, calm angular face, dark brown eyes, poised neutral expression. She wears a long asymmetrical crimson technical coat with a high sculpted collar over a matte black fitted bodysuit, slim black gloves, and mirror-chrome ankle boots. Full body visible, arms relaxed slightly away from the torso, outfit silhouette and footwear unobstructed.
Style/medium: photorealistic high-fashion editorial photography, original character design, realistic anatomy and natural skin texture.
Composition/framing: wide 16:9 landscape, full-body three-quarter-front view centered with generous breathing room, eye-level camera, 50mm lens character, crisp readable silhouette.
Lighting/mood: controlled soft key light from upper left, cool cyan rim light from behind, restrained crimson bounce from the coat, confident and enigmatic.
Color palette: graphite black, deep crimson, cool cyan, mirror silver, natural skin tones.
Materials/textures: matte technical fabric, subtle coat seams, brushed black textile, clean reflective chrome footwear.
Constraints: exactly one adult subject; preserve a clearly readable face, hair, coat, bodysuit, gloves, and boots; no motion blur; no text, letters, numbers, logos, trademarks, watermarks, UI, weapons, or copyrighted characters; no resemblance to a real person or living artist.
Avoid: extra people, split panels, masks, sunglasses, fantasy armor, cybernetic body parts, exaggerated anatomy, exposed underwear, cluttered background, neon cyberpunk overload, malformed hands or feet.

This anchor locks a short silver bob, a sculpted asymmetric crimson coat, a matte black bodysuit, and chrome boots. Repeating the same nouns in the video prompt aligns the visual reference with the written preservation rules.

Image 2: build wardrobe and props like a material sample

A prop that changes state during the video—such as an umbrella—needs a readable shape, material, color, and handle. If it appears only as a tiny object in the full-body reference, its design may drift. Show only one instance of each hero prop; repeated objects in a reference can encourage unwanted duplication.

Wardrobe and prop reference showing crimson fabric, black gloves, chrome boots, and one translucent umbrella

Image-generation prompt

Use case: stylized-concept
Asset type: landscape reference image for a Seedance 2.0 multi-reference fashion-film workflow; wardrobe and hero-prop detail anchor
Input images: Use the character anchor only as the identity, crimson technical coat, matte black bodysuit, glove, and chrome-boot reference. Preserve those exact design cues.
Primary request: Create a premium fashion editorial detail composition that clearly defines the performer's materials and one hero prop for video consistency.
Scene/backdrop: dark graphite studio tabletop and wall with controlled reflections, minimal and uncluttered.
Subject: the same clearly adult silver-bob performer shown in a waist-up three-quarter profile on the left, wearing the same crimson high-collar technical coat and slim black gloves; on the right, she holds exactly one closed translucent smoke-gray umbrella with a polished chrome curved handle. Include carefully composed close details of the crimson fabric seam, black glove, and chrome boot surface without split-panel borders or labels.
Style/medium: photorealistic luxury fashion campaign photography, original design, tactile material study.
Composition/framing: wide 16:9 landscape, waist-up performer occupying the left half, umbrella silhouette and material details arranged on the right, clean hierarchy and generous padding, 85mm lens character.
Lighting/mood: narrow soft key light, cool cyan edge light, restrained crimson reflection, elegant and mysterious.
Color palette: deep crimson, graphite black, smoke gray, cool cyan, mirror silver.
Materials/textures: matte technical coat fabric with fine seams, soft black glove leather, translucent umbrella canopy, polished chrome handle and boot.
Constraints: preserve the character's short silver bob, face, coat design, black bodysuit, gloves, and chrome boots from the input; show exactly one umbrella; no text, letters, numbers, labels, logos, trademarks, watermarks, UI, weapons, or copyrighted characters; no resemblance to a real person or living artist.
Avoid: duplicate props, wardrobe redesign, different hair length, open umbrella blocking the face, product branding, collage borders, duplicate limbs, neon cyberpunk clutter, malformed hands.

If an early image contains two umbrellas, correct it before using it. Multi-reference generation can amplify mistakes, so a contradiction-free input is more valuable than an image that merely looks impressive.

Image 3: remove the performer from the environment plate

A person standing in the location reference can be interpreted as a second character, or the model may copy that person's placement into every shot. Build a clean plate that communicates only architecture, depth, lighting, weather, and surface response.

Empty rain-soaked brutalist concourse with cyan overhead light, crimson accents, and a reflective black floor

Image-generation prompt

Use case: stylized-concept
Asset type: landscape reference image for a Seedance 2.0 multi-reference fashion-film workflow; environment and lighting anchor
Primary request: Create an original cinematic environment plate for a stylish fashion film, designed to define architecture, depth, atmosphere, reflections, and lighting without introducing a character.
Scene/backdrop: a vast rain-soaked brutalist underground transit concourse after closing, long concrete ribs and repeating rectangular portals, glossy black floor with shallow puddles, a distant opening filled with mist, sparse linear light fixtures.
Subject: environment only; no people. A restrained sequence of crimson light panels runs along one wall while cool cyan ceiling light and silver rain reflections establish the color script. One subtle wind current pushes mist and loose droplets through the corridor.
Style/medium: photorealistic cinematic location photography, premium fashion-film production design, original architecture.
Composition/framing: wide 16:9 landscape, strong one-point perspective down the concourse, low eye-level camera, 24mm lens character, clean central walking path, foreground puddle reflections and deep layered background.
Lighting/mood: cool cyan overhead pools, narrow crimson side accents, wet specular reflections, moody but readable, elegant tension.
Color palette: graphite concrete, black wet floor, cool cyan, restrained deep crimson, silver highlights.
Materials/textures: rough poured concrete, brushed metal edges, wet stone, fine rain mist, realistic puddle ripples.
Constraints: environment only, no humans or silhouettes; no text, letters, numbers, signage, logos, trademarks, watermarks, UI, vehicles, or copyrighted architecture; keep the central path unobstructed.
Avoid: neon cyberpunk city, colorful shop signs, crowded station, fantasy ruins, sci-fi spacecraft, excessive fog hiding the architecture, impossible reflections, fisheye distortion.

The central path is deliberately clear. The visual rules are reduced to three elements: cyan overhead pools, restrained crimson side lights, and a wet black floor. That gives the model a coherent color script without adding another identity or pose.

Image 4: use the keyframe as the contract between references

The final reference combines the character and location into a plausible opening frame. It establishes subject scale, camera height, lens character, screen position, and grade. Do not force the whole story into this image; show the stable state from which motion begins.

Opening fashion-film keyframe combining the same performer and umbrella with the rain-soaked concourse

Image-generation prompt

Use case: compositing
Asset type: landscape reference image for a Seedance 2.0 multi-reference fashion-film workflow; hero composition, lens, color-grade, and first-frame anchor
Input images: Image 1 is the wardrobe-and-prop reference—preserve the adult woman's face, short silver bob, crimson high-collar technical coat, matte black bodysuit, black gloves, chrome boots, and smoke-gray umbrella. Image 2 is the environment reference—preserve its rain-soaked brutalist transit concourse, one-point depth, wet black floor, cyan overhead light, and restrained crimson wall accents.
Primary request: Place the same performer from Image 1 naturally inside the central path of Image 2 and create a polished opening keyframe for a stylish fashion film.
Scene/backdrop: the same vast wet brutalist concourse after closing, fine rain drifting from the left openings, puddles reflecting cyan and crimson light.
Subject: the same clearly adult performer strides toward camera with composed confidence, holding the closed smoke-gray umbrella downward in her right hand. Her crimson coat hem lifts slightly in the crosswind; her face, hair, outfit proportions, and chrome boots remain recognizable.
Style/medium: photorealistic cinematic fashion editorial, original production design, realistic physical interaction.
Composition/framing: wide 16:9 landscape, low-angle medium-wide full-body shot, performer placed slightly left of center on the vanishing line, foreground puddle reflection, 28mm lens character, subtle natural motion in coat only, crisp face.
Lighting/mood: cool cyan overhead pools, narrow crimson edge light, silver rain highlights, elegant suspense and controlled energy.
Color palette: graphite, deep crimson, cool cyan, black, mirror silver.
Materials/textures: wet concrete, realistic puddles, matte coat fabric, chrome footwear, translucent smoke-gray umbrella.
Constraints: combine references without redesigning them; exactly one adult person and one closed umbrella; preserve character identity, wardrobe, prop proportions, architecture, and color script; physically plausible stance and reflections; no text, letters, numbers, signs, logos, trademarks, watermarks, UI, weapons, or copyrighted characters.
Avoid: face drift, different outfit, extra people, open umbrella, duplicate props, superhero pose, neon city clutter, extreme motion blur, floating feet, warped architecture, malformed hands.

Assign reference roles in the prompt, not only in filenames

Dreamina lets creators address uploaded assets by @AssetName in Multiframes mode. Use the names shown in your actual interface. “Reference Image 1” is too vague; tell the model which attributes to use and which attributes not to inherit.

InputRole sentence in the promptPriority
@Image 1Use only identity, body, hair, and full wardrobeHighest
@Image 2Use only umbrella, fabric, glove, and metal detailsHigh
@Image 3Use only location, light, rain, and reflectionsHigh
@Image 4Use opening composition, lens, and gradeMedium

Adjust the numbers to match your upload order. Choose duration in the generation settings rather than forcing a time value into the prose. The prompt should spend its detail budget on the moving subject, camera route, and environmental response.

Write motion after preservation rules

An effective video prompt describes an event in order rather than listing moods. This example uses six layers:

  1. The responsibility of each reference and what must not be mixed
  2. The performer's starting pose, gaze, and first action
  3. The action chain from walking to pivoting and opening the umbrella
  4. A camera path from low tracking to lateral movement and a partial orbit
  5. Physical response in the coat, hair, rain, puddles, and light
  6. Continuity constraints, exclusions, and restrained sound direction

Copy-ready Seedance 2.0 video prompt

Use @Image 1 only for the adult performer's identity, short silver bob, facial features, body proportions, crimson high-collar technical coat, matte black bodysuit, black gloves, and chrome boots. Use @Image 2 only for the smoke-gray umbrella, coat seams, glove texture, and reflective chrome material details. Use @Image 3 only for the brutalist transit concourse, one-point depth, wet floor, cyan overhead lighting, restrained crimson wall accents, rain, mist, and reflection behavior. Use @Image 4 as the opening composition, lens language, scale, and color-grade reference. Do not merge or swap the roles of these references.

The same adult performer walks toward camera along the central path with calm confidence, holding the closed umbrella downward in her right hand. Begin with a low close tracking shot of one chrome boot stepping into a shallow puddle; water splashes naturally and the crimson coat edge passes through frame. Rise smoothly into a side-tracking medium-wide full-body shot as she continues walking. Her gaze stays forward, shoulders relaxed, coat hem and silver bob reacting consistently to the crosswind while rain strikes the floor and umbrella surface.

She slows, turns her head toward camera, then plants one foot and makes a controlled pivot. During the pivot she opens the single umbrella behind her shoulder in one physically plausible motion. The camera performs a restrained partial orbit in the opposite direction, preserving her face and body proportions. Crimson light passes through the translucent canopy, cyan highlights slide across the chrome boots, and reflected light moves across the wet floor. End on a stable low three-quarter hero frame with her calm gaze sharp, the open umbrella forming a clean circle behind her, and the corridor receding into mist.

Camera language: low macro tracking to medium-wide lateral tracking to controlled partial orbit; smooth acceleration and deceleration; no random cuts, no handheld shake, no extreme zoom, no impossible camera path.
Performance and physics: natural walking cadence, clear weight transfer, realistic coat drag and recovery, believable umbrella opening, coherent rain splash and reflections, stable hands, face, outfit, and prop proportions.
Audio direction: restrained industrial ambience, rain on concrete and umbrella fabric, precise chrome heel impacts, a soft coat swish, and one deep tonal pulse at the pivot; no dialogue and no dominant music.
Keep exactly one adult performer and one umbrella. Preserve identity, wardrobe, prop, architecture, color script, and lighting continuity across every shot. No extra people, duplicate objects, text, logos, signage, face drift, wardrobe changes, warped limbs, floating feet, impossible reflections, or overexposed highlights.

The prompt is long because every paragraph owns a different failure mode. The identity anchor protects face and wardrobe; the material reference protects the prop; the environment controls light and reflections; the keyframe sets the opening camera. Text connects them through movement, sequencing, and priority.

Common failures and targeted fixes

FailureLikely causeFirst change
Face changes between shotsFace is small or several identities competeReduce to one identity anchor
Outfit or prop duplicatesThe same item repeats in a referenceKeep one visible instance
Location becomes another cityStyle adjectives overpower scene roleMove the location rule earlier
Camera becomes chaoticToo many pans, zooms, and orbits competeLimit the route to 2–3 moves
Walk-to-pivot looks weightlessNo transfer of weight between actionsSpecify slow, plant, then pivot
Result feels like a still imageAppearance is described but reactions are notAdd hair, cloth, rain, and reflection motion

Do not rewrite everything after one weak result. If identity fails, change the identity reference and preservation sentence. If the camera fails, edit only the camera paragraph. When the interface supports repeatable seeds or variations, compare one change at a time.

Pre-generation checklist

  • Do the people, props, and object counts agree across every reference?
  • Are the face, hands, full silhouette, and outfit readable in the character anchor?
  • Is the location reference free from unrelated people and text?
  • Does every @Image line say what to inherit and what not to inherit?
  • Are expression, gaze, posture, action, and weight transfer described?
  • Does the camera route unfold in order without too many simultaneous moves?
  • Do fabric, hair, rain, mist, water, and reflections respond to the action?
  • Are face drift, duplicate props, text, logos, and watermarks explicitly excluded?
  • Do you have consent and the necessary rights for any real-person reference?

ByteDance's launch note says real-person portrait references may require identity verification or prior legal authorization. Use only people, brands, music, locations, and source media you created or are entitled to use.

Thumbnail-generation prompt

The following prompt turns the reference package into a clear climax image: the performer pivots, the umbrella opens, and the same visual language remains intact.

Use case: stylized-concept
Asset type: wide blog thumbnail and final-output concept for a guide about stylish Seedance 2.0 multi-reference video creation
Input images: Use the cinematic keyframe as the exact character, wardrobe, umbrella, environment, architecture, material, lighting, and color-grade reference.
Primary request: Create a striking original fashion-film climax that shows what the multi-reference package can produce while remaining recognizably connected to the input.
Scene/backdrop: the same rain-soaked brutalist transit concourse with cyan overhead light, restrained crimson wall accents, wet black floor, mist, and deep one-point perspective.
Subject: the same clearly adult silver-bob performer in the same crimson asymmetrical technical coat, matte black bodysuit, black gloves, and chrome boots. She has just pivoted sharply toward camera and opens the single translucent smoke-gray umbrella diagonally behind her shoulder; the coat arcs naturally with the turn, rain droplets sweep around the umbrella edge, and her calm gaze remains crisp and recognizable.
Style/medium: photorealistic cinematic luxury fashion campaign, energetic but physically plausible, original visual identity.
Composition/framing: wide 16:9 landscape, low three-quarter camera in a controlled partial orbit, performer large and slightly left of center, open umbrella forming a strong graphic circle behind her, long reflective corridor visible to the right, thumbnail-readable silhouette, clean negative space.
Lighting/mood: cool cyan overhead highlights, crimson rim light through the translucent umbrella, silver rain sparkle, elegant momentum and high-end editorial confidence.
Color palette: graphite black, deep crimson, cool cyan, smoke gray, mirror silver.
Materials/textures: wet concrete and puddles, matte technical coat fabric, translucent umbrella canopy, polished chrome boots, realistic rain droplets.
Constraints: preserve the exact adult character identity, silver bob, facial features, wardrobe, environment, and color script from the input; exactly one person and one open umbrella; realistic fabric and umbrella physics; crisp face; no text, letters, numbers, title, logos, trademarks, watermarks, UI, weapons, or copyrighted characters.
Avoid: identity drift, different clothing, extra people, duplicate umbrella, superhero effects, neon cyberpunk clutter, extreme blur, distorted umbrella spokes, floating feet, malformed hands, overexposed rain.

Conclusion: multi-reference means separating responsibilities

Seedance 2.0 is valuable not simply because it accepts many files, but because different files can guide identity, location, camera, movement, and sound. Treat the character, prop, environment, and keyframe as four explicit contracts. Restate those contracts in the prompt, and you can vary movement and camera direction without surrendering the visual identity.

Primary sources

Interface availability, reference limits, resolution, watermarks, pricing, and commercial terms can vary by account and region. Confirm the current official policy before production.