Seedance 2.0 스타일리시 영상 제작: 멀티 레퍼런스 이미지와 프롬프트 설계법

인물·의상과 소품·공간·키프레임을 역할별 참조 이미지로 설계하고, Seedance 2.0에서 충돌 없이 연결하는 영상 프롬프트 작성법을 실제 예제로 설명합니다.

AIZIGOO 편집팀
Seedance 2.0 스타일리시 영상 제작: 멀티 레퍼런스 이미지와 프롬프트 설계법

Seedance 2.0에서 스타일리시한 영상을 만드는 핵심은 프롬프트에 화려한 형용사를 많이 넣는 것이 아닙니다. 인물, 의상·소품, 공간·조명, 첫 장면의 구도를 서로 다른 참조 이미지로 분리하고 각 이미지에서 무엇만 가져올지 지정하는 것이 더 중요합니다.

ByteDance Seed의 2026년 2월 12일 공식 발표에 따르면 Seedance 2.0은 텍스트, 이미지, 영상, 오디오를 함께 이해하는 통합 멀티모달 영상 모델입니다. 최대 9장의 이미지, 3개의 영상, 3개의 오디오를 참조할 수 있고, 공식 예제도 “인물은 이미지 2, 공간은 이미지 3, 소품은 이미지 4”처럼 역할을 명시합니다. Dreamina의 현재 사용 가이드에서는 멀티프레임 모드에 자료를 올린 뒤 @AssetName 형식으로 지칭하도록 안내합니다.

다만 많은 이미지를 올린다고 항상 좋아지는 것은 아닙니다. 서로 다른 얼굴, 조명과 의상이 한꺼번에 들어가면 모델이 무엇을 고정해야 하는지 모호해집니다. 이 글은 실제로 함께 쓸 수 있는 네 장의 참조 이미지와 하나의 대표 결과 이미지를 제작해, 적은 참조로 더 명확하게 지시하는 방법을 보여줍니다.

이 글은 2026년 7월 28일 확인한 ByteDance Seed와 Dreamina의 공식 공개 자료를 기준으로 합니다. 메뉴 이름, 지원 지역, 요금제, 크레딧과 입력 한도는 변경될 수 있으므로 생성 전에 현재 서비스 화면을 확인하세요. 공식 벤치마크와 품질 평가는 개발사 측 결과이며 모든 소재와 장면에서 같은 품질을 보장하지 않습니다.

먼저 완성본보다 ‘참조 역할표’를 만든다

이번 예제는 은색 단발의 성인 퍼포머가 붉은 코트와 반투명 우산을 들고 비 내리는 브루탈리즘 통로를 걷다가 회전하는 패션 필름입니다. 참조 네 장이 같은 정보를 반복하지 않도록 다음처럼 역할을 나눕니다.

참조고정할 정보의도적으로 제외할 정보
이미지 1얼굴, 체형, 헤어, 전체 의상배경, 동작, 카메라
이미지 2우산, 원단, 장갑, 금속 재질새 인물, 새 공간
이미지 3건축, 조명, 비, 반사인물, 의상, 소품
이미지 4첫 구도, 렌즈감, 색보정이후 동작과 결말

한 장이 모든 것을 설명하게 만들면 편리해 보이지만 수정할 때 전체가 흔들립니다. 반대로 역할을 나누면 “얼굴은 유지하되 공간만 교체”하거나 “공간은 유지하되 카메라 동선만 바꾸는” 식으로 원인을 좁힐 수 있습니다.

이미지 1: 인물 앵커는 예쁜 포스터가 아니라 식별표다

인물 참조는 표정이 극적이거나 손이 얼굴을 가린 이미지보다, 얼굴·헤어·전신 실루엣·신발까지 한눈에 보이는 중립적인 이미지가 좋습니다. 배경은 단순하게 하고 모션 블러를 피해야 이후 영상에서 무엇을 보존할지 명확해집니다.

은색 단발과 붉은 코트, 크롬 부츠를 명확하게 보여주는 성인 퍼포머 인물 앵커

이미지 제작 프롬프트

Use case: stylized-concept
Asset type: landscape reference image for a Seedance 2.0 multi-reference fashion-film workflow; character identity anchor
Primary request: Create a premium cinematic character anchor image of one original adult female performance artist whose identity, silhouette, outfit, and materials can be reused consistently in AI video generation.
Scene/backdrop: minimal dark graphite studio with a softly illuminated floor, subtle atmospheric haze, no scenery and no distracting props.
Subject: one clearly adult woman, mid-to-late twenties, distinctive short silver bob with a blunt fringe, calm angular face, dark brown eyes, poised neutral expression. She wears a long asymmetrical crimson technical coat with a high sculpted collar over a matte black fitted bodysuit, slim black gloves, and mirror-chrome ankle boots. Full body visible, arms relaxed slightly away from the torso, outfit silhouette and footwear unobstructed.
Style/medium: photorealistic high-fashion editorial photography, original character design, realistic anatomy and natural skin texture.
Composition/framing: wide 16:9 landscape, full-body three-quarter-front view centered with generous breathing room, eye-level camera, 50mm lens character, crisp readable silhouette.
Lighting/mood: controlled soft key light from upper left, cool cyan rim light from behind, restrained crimson bounce from the coat, confident and enigmatic.
Color palette: graphite black, deep crimson, cool cyan, mirror silver, natural skin tones.
Materials/textures: matte technical fabric, subtle coat seams, brushed black textile, clean reflective chrome footwear.
Constraints: exactly one adult subject; preserve a clearly readable face, hair, coat, bodysuit, gloves, and boots; no motion blur; no text, letters, numbers, logos, trademarks, watermarks, UI, weapons, or copyrighted characters; no resemblance to a real person or living artist.
Avoid: extra people, split panels, masks, sunglasses, fantasy armor, cybernetic body parts, exaggerated anatomy, exposed underwear, cluttered background, neon cyberpunk overload, malformed hands or feet.

이 예제는 짧은 은색 단발, 높은 칼라의 비대칭 붉은 코트, 검은 보디수트와 크롬 부츠를 고정 요소로 정했습니다. 영상 프롬프트에서도 같은 명사를 반복해 이미지와 텍스트의 기준을 일치시킵니다.

이미지 2: 의상과 소품은 ‘재질 샘플’처럼 만든다

우산처럼 영상 중 상태가 바뀌는 소품은 모양, 색, 손잡이와 재질이 보여야 합니다. 인물 이미지에 소품이 작게만 보이면 영상에서 색이나 형태가 쉽게 변합니다. 다만 같은 소품을 여러 개 배치하면 개수까지 참조할 수 있으므로 한 종류당 하나만 보여주는 편이 안전합니다.

붉은 코트의 원단과 검은 장갑, 크롬 부츠, 반투명 우산을 보여주는 의상과 소품 참조

이미지 제작 프롬프트

Use case: stylized-concept
Asset type: landscape reference image for a Seedance 2.0 multi-reference fashion-film workflow; wardrobe and hero-prop detail anchor
Input images: Use the character anchor only as the identity, crimson technical coat, matte black bodysuit, glove, and chrome-boot reference. Preserve those exact design cues.
Primary request: Create a premium fashion editorial detail composition that clearly defines the performer's materials and one hero prop for video consistency.
Scene/backdrop: dark graphite studio tabletop and wall with controlled reflections, minimal and uncluttered.
Subject: the same clearly adult silver-bob performer shown in a waist-up three-quarter profile on the left, wearing the same crimson high-collar technical coat and slim black gloves; on the right, she holds exactly one closed translucent smoke-gray umbrella with a polished chrome curved handle. Include carefully composed close details of the crimson fabric seam, black glove, and chrome boot surface without split-panel borders or labels.
Style/medium: photorealistic luxury fashion campaign photography, original design, tactile material study.
Composition/framing: wide 16:9 landscape, waist-up performer occupying the left half, umbrella silhouette and material details arranged on the right, clean hierarchy and generous padding, 85mm lens character.
Lighting/mood: narrow soft key light, cool cyan edge light, restrained crimson reflection, elegant and mysterious.
Color palette: deep crimson, graphite black, smoke gray, cool cyan, mirror silver.
Materials/textures: matte technical coat fabric with fine seams, soft black glove leather, translucent umbrella canopy, polished chrome handle and boot.
Constraints: preserve the character's short silver bob, face, coat design, black bodysuit, gloves, and chrome boots from the input; show exactly one umbrella; no text, letters, numbers, labels, logos, trademarks, watermarks, UI, weapons, or copyrighted characters; no resemblance to a real person or living artist.
Avoid: duplicate props, wardrobe redesign, different hair length, open umbrella blocking the face, product branding, collage borders, duplicate limbs, neon cyberpunk clutter, malformed hands.

생성 과정에서 우산이 중복되면 그대로 사용하지 말고 한 개만 남도록 수정해야 합니다. 멀티 레퍼런스는 오류도 확대할 수 있기 때문에 “그럴듯한 이미지”보다 모순이 없는 이미지가 더 좋은 입력입니다.

이미지 3: 공간 참조에서는 인물을 완전히 뺀다

공간 이미지에 다른 사람이 서 있으면 Seedance가 그 사람까지 등장인물로 해석하거나, 주인공의 위치와 구도를 함께 복사할 수 있습니다. 공간 참조는 건축, 원근, 조명, 날씨와 바닥 반사만 보여주는 클린 플레이트로 준비합니다.

시안과 붉은 조명, 젖은 반사 바닥을 가진 비어 있는 브루탈리즘 지하 통로

이미지 제작 프롬프트

Use case: stylized-concept
Asset type: landscape reference image for a Seedance 2.0 multi-reference fashion-film workflow; environment and lighting anchor
Primary request: Create an original cinematic environment plate for a stylish fashion film, designed to define architecture, depth, atmosphere, reflections, and lighting without introducing a character.
Scene/backdrop: a vast rain-soaked brutalist underground transit concourse after closing, long concrete ribs and repeating rectangular portals, glossy black floor with shallow puddles, a distant opening filled with mist, sparse linear light fixtures.
Subject: environment only; no people. A restrained sequence of crimson light panels runs along one wall while cool cyan ceiling light and silver rain reflections establish the color script. One subtle wind current pushes mist and loose droplets through the corridor.
Style/medium: photorealistic cinematic location photography, premium fashion-film production design, original architecture.
Composition/framing: wide 16:9 landscape, strong one-point perspective down the concourse, low eye-level camera, 24mm lens character, clean central walking path, foreground puddle reflections and deep layered background.
Lighting/mood: cool cyan overhead pools, narrow crimson side accents, wet specular reflections, moody but readable, elegant tension.
Color palette: graphite concrete, black wet floor, cool cyan, restrained deep crimson, silver highlights.
Materials/textures: rough poured concrete, brushed metal edges, wet stone, fine rain mist, realistic puddle ripples.
Constraints: environment only, no humans or silhouettes; no text, letters, numbers, signage, logos, trademarks, watermarks, UI, vehicles, or copyrighted architecture; keep the central path unobstructed.
Avoid: neon cyberpunk city, colorful shop signs, crowded station, fantasy ruins, sci-fi spacecraft, excessive fog hiding the architecture, impossible reflections, fisheye distortion.

이 장면에서는 중앙 통로를 비워 인물이 걸을 공간을 확보했습니다. 시안 천장광, 오른쪽의 붉은 라이트, 젖은 검은 바닥이라는 세 요소만으로 색상 규칙을 단순화했습니다.

이미지 4: 키프레임은 모든 참조를 합치는 계약서다

마지막 참조 이미지는 인물과 공간을 실제 첫 장면처럼 합친 키프레임입니다. 여기서 인물 크기, 카메라 높이, 렌즈 느낌, 프레임 속 위치와 색보정을 결정합니다. 이후 동작까지 한 장에 설명하려 하지 말고 영상이 시작하는 안정된 상태만 선명하게 보여주세요.

같은 퍼포머와 우산을 비 내리는 통로에 배치한 스타일리시 영상의 첫 키프레임

이미지 제작 프롬프트

Use case: compositing
Asset type: landscape reference image for a Seedance 2.0 multi-reference fashion-film workflow; hero composition, lens, color-grade, and first-frame anchor
Input images: Image 1 is the wardrobe-and-prop reference—preserve the adult woman's face, short silver bob, crimson high-collar technical coat, matte black bodysuit, black gloves, chrome boots, and smoke-gray umbrella. Image 2 is the environment reference—preserve its rain-soaked brutalist transit concourse, one-point depth, wet black floor, cyan overhead light, and restrained crimson wall accents.
Primary request: Place the same performer from Image 1 naturally inside the central path of Image 2 and create a polished opening keyframe for a stylish fashion film.
Scene/backdrop: the same vast wet brutalist concourse after closing, fine rain drifting from the left openings, puddles reflecting cyan and crimson light.
Subject: the same clearly adult performer strides toward camera with composed confidence, holding the closed smoke-gray umbrella downward in her right hand. Her crimson coat hem lifts slightly in the crosswind; her face, hair, outfit proportions, and chrome boots remain recognizable.
Style/medium: photorealistic cinematic fashion editorial, original production design, realistic physical interaction.
Composition/framing: wide 16:9 landscape, low-angle medium-wide full-body shot, performer placed slightly left of center on the vanishing line, foreground puddle reflection, 28mm lens character, subtle natural motion in coat only, crisp face.
Lighting/mood: cool cyan overhead pools, narrow crimson edge light, silver rain highlights, elegant suspense and controlled energy.
Color palette: graphite, deep crimson, cool cyan, black, mirror silver.
Materials/textures: wet concrete, realistic puddles, matte coat fabric, chrome footwear, translucent smoke-gray umbrella.
Constraints: combine references without redesigning them; exactly one adult person and one closed umbrella; preserve character identity, wardrobe, prop proportions, architecture, and color script; physically plausible stance and reflections; no text, letters, numbers, signs, logos, trademarks, watermarks, UI, weapons, or copyrighted characters.
Avoid: face drift, different outfit, extra people, open umbrella, duplicate props, superhero pose, neon city clutter, extreme motion blur, floating feet, warped architecture, malformed hands.

Seedance 2.0에 올릴 때는 파일 순서보다 역할을 프롬프트에 적는다

Dreamina에서는 멀티프레임 모드에서 각 자료를 @AssetName으로 호출할 수 있습니다. 파일명이 다르게 보이더라도 화면에 표시되는 실제 참조 이름을 사용하세요. 중요한 것은 “이미지 1을 참고해줘”로 끝내지 않고 어떤 속성을 참조하고 어떤 속성은 가져오지 말아야 하는지 쓰는 것입니다.

입력프롬프트의 역할 문장우선순위
@Image 1인물 정체성·체형·헤어·전체 의상만 사용최고
@Image 2우산·원단·장갑·금속 재질만 사용높음
@Image 3공간·조명·비·반사만 사용높음
@Image 4첫 구도·렌즈·색보정만 사용중간

참조 번호는 실제 업로드 순서에 맞춰 바꾸면 됩니다. 영상 길이는 문장 안에서 억지로 지시하지 말고 서비스의 생성 설정에서 선택하세요. 프롬프트는 움직이는 주체, 카메라와 환경 반응에 집중하는 편이 낫습니다.

영상 프롬프트는 ‘무엇을 보존할지’ 다음에 ‘무엇이 움직일지’를 쓴다

좋은 영상 프롬프트는 형용사 목록보다 사건의 순서가 분명합니다. 이번 예제는 다음 여섯 층으로 구성했습니다.

  1. 참조별 역할과 서로 섞지 말아야 할 요소
  2. 인물의 시작 자세, 시선과 첫 행동
  3. 걸음에서 회전과 우산 개방으로 이어지는 동작
  4. 로우 트래킹, 측면 이동, 부분 오빗으로 이어지는 카메라
  5. 코트, 머리, 비, 물웅덩이와 빛의 물리 반응
  6. 유지 조건, 금지 요소와 최소한의 사운드 방향

바로 복사해 사용하는 최종 Seedance 2.0 영상 프롬프트

Use @Image 1 only for the adult performer's identity, short silver bob, facial features, body proportions, crimson high-collar technical coat, matte black bodysuit, black gloves, and chrome boots. Use @Image 2 only for the smoke-gray umbrella, coat seams, glove texture, and reflective chrome material details. Use @Image 3 only for the brutalist transit concourse, one-point depth, wet floor, cyan overhead lighting, restrained crimson wall accents, rain, mist, and reflection behavior. Use @Image 4 as the opening composition, lens language, scale, and color-grade reference. Do not merge or swap the roles of these references.

The same adult performer walks toward camera along the central path with calm confidence, holding the closed umbrella downward in her right hand. Begin with a low close tracking shot of one chrome boot stepping into a shallow puddle; water splashes naturally and the crimson coat edge passes through frame. Rise smoothly into a side-tracking medium-wide full-body shot as she continues walking. Her gaze stays forward, shoulders relaxed, coat hem and silver bob reacting consistently to the crosswind while rain strikes the floor and umbrella surface.

She slows, turns her head toward camera, then plants one foot and makes a controlled pivot. During the pivot she opens the single umbrella behind her shoulder in one physically plausible motion. The camera performs a restrained partial orbit in the opposite direction, preserving her face and body proportions. Crimson light passes through the translucent canopy, cyan highlights slide across the chrome boots, and reflected light moves across the wet floor. End on a stable low three-quarter hero frame with her calm gaze sharp, the open umbrella forming a clean circle behind her, and the corridor receding into mist.

Camera language: low macro tracking to medium-wide lateral tracking to controlled partial orbit; smooth acceleration and deceleration; no random cuts, no handheld shake, no extreme zoom, no impossible camera path.
Performance and physics: natural walking cadence, clear weight transfer, realistic coat drag and recovery, believable umbrella opening, coherent rain splash and reflections, stable hands, face, outfit, and prop proportions.
Audio direction: restrained industrial ambience, rain on concrete and umbrella fabric, precise chrome heel impacts, a soft coat swish, and one deep tonal pulse at the pivot; no dialogue and no dominant music.
Keep exactly one adult performer and one umbrella. Preserve identity, wardrobe, prop, architecture, color script, and lighting continuity across every shot. No extra people, duplicate objects, text, logos, signage, face drift, wardrobe changes, warped limbs, floating feet, impossible reflections, or overexposed highlights.

프롬프트가 길어 보여도 각 문장은 서로 다른 실패를 막습니다. 얼굴과 의상은 인물 앵커가, 재질은 소품 앵커가, 바닥과 빛은 공간 앵커가, 카메라 시작점은 키프레임이 담당합니다. 텍스트는 그 사이를 연결하는 동작과 우선순위를 지시합니다.

자주 실패하는 패턴과 수정 방법

실패원인먼저 바꿀 것
얼굴이 장면마다 변함인물 참조가 작거나 여러 얼굴이 섞임단일 인물 앵커로 축소
의상·소품이 복제됨참조 안에 같은 물건이 반복됨물건 하나만 남기기
배경이 다른 도시로 바뀜공간보다 스타일 형용사가 강함공간 역할 문장을 앞에 배치
카메라가 난폭하게 움직임팬·줌·오빗을 동시에 과도하게 지시카메라 동작을 2~3개로 제한
걷기와 회전이 부자연스러움동작 사이 체중 이동이 없음감속·발 고정·회전 순서 명시
결과가 정지 이미지처럼 보임외형만 설명하고 환경 반응이 없음머리·원단·비·반사 움직임 추가

한 번에 모든 문장을 다시 쓰지 마세요. 정체성이 틀리면 인물 참조와 보존 문장만, 카메라가 틀리면 카메라 단락만 바꿉니다. 같은 시드나 변형 기능을 사용할 수 있는 환경이라면 한 요소씩 비교해야 원인을 찾기 쉽습니다.

게시 전 체크리스트

  • 모든 참조 이미지의 인물, 소품과 개수가 서로 모순되지 않는가
  • 인물 이미지에서 얼굴, 손, 전신과 의상 실루엣이 읽히는가
  • 공간 이미지에 불필요한 사람과 텍스트가 없는가
  • 각 @Image 뒤에 참조할 속성과 제외할 속성을 적었는가
  • 인물의 표정, 시선, 자세, 행동과 체중 이동이 설명됐는가
  • 카메라 동선이 시간 순서대로 연결되고 과도하지 않은가
  • 옷, 머리, 비, 연기, 물과 반사가 행동에 반응하는가
  • 얼굴 변화, 중복 소품, 워터마크와 텍스트 같은 금지 요소가 있는가
  • 실제 인물의 얼굴은 본인 동의와 필요한 권한을 확보했는가

ByteDance는 공식 발표에서 실제 인물 초상을 참조할 경우 신원 확인이나 사전 법적 권한이 필요할 수 있다고 안내합니다. 브랜드, 음악, 인물과 장소 자료도 직접 제작했거나 이용 권한이 있는 것만 사용하세요.

대표 썸네일 제작 프롬프트

아래 프롬프트는 네 장의 참조로 설계한 결과를 한눈에 보여주기 위해, 키프레임을 바탕으로 회전과 우산 개방의 절정 장면을 제작한 것입니다.

Use case: stylized-concept
Asset type: wide blog thumbnail and final-output concept for a guide about stylish Seedance 2.0 multi-reference video creation
Input images: Use the cinematic keyframe as the exact character, wardrobe, umbrella, environment, architecture, material, lighting, and color-grade reference.
Primary request: Create a striking original fashion-film climax that shows what the multi-reference package can produce while remaining recognizably connected to the input.
Scene/backdrop: the same rain-soaked brutalist transit concourse with cyan overhead light, restrained crimson wall accents, wet black floor, mist, and deep one-point perspective.
Subject: the same clearly adult silver-bob performer in the same crimson asymmetrical technical coat, matte black bodysuit, black gloves, and chrome boots. She has just pivoted sharply toward camera and opens the single translucent smoke-gray umbrella diagonally behind her shoulder; the coat arcs naturally with the turn, rain droplets sweep around the umbrella edge, and her calm gaze remains crisp and recognizable.
Style/medium: photorealistic cinematic luxury fashion campaign, energetic but physically plausible, original visual identity.
Composition/framing: wide 16:9 landscape, low three-quarter camera in a controlled partial orbit, performer large and slightly left of center, open umbrella forming a strong graphic circle behind her, long reflective corridor visible to the right, thumbnail-readable silhouette, clean negative space.
Lighting/mood: cool cyan overhead highlights, crimson rim light through the translucent umbrella, silver rain sparkle, elegant momentum and high-end editorial confidence.
Color palette: graphite black, deep crimson, cool cyan, smoke gray, mirror silver.
Materials/textures: wet concrete and puddles, matte technical coat fabric, translucent umbrella canopy, polished chrome boots, realistic rain droplets.
Constraints: preserve the exact adult character identity, silver bob, facial features, wardrobe, environment, and color script from the input; exactly one person and one open umbrella; realistic fabric and umbrella physics; crisp face; no text, letters, numbers, title, logos, trademarks, watermarks, UI, weapons, or copyrighted characters.
Avoid: identity drift, different clothing, extra people, duplicate umbrella, superhero effects, neon cyberpunk clutter, extreme blur, distorted umbrella spokes, floating feet, malformed hands, overexposed rain.

결론: 멀티 레퍼런스는 이미지 수가 아니라 책임 분리다

Seedance 2.0의 장점은 여러 자료를 넣을 수 있다는 사실보다, 각 자료에서 인물·공간·카메라·움직임·사운드를 따로 참조할 수 있다는 데 있습니다. 인물, 소품, 공간과 키프레임을 네 개의 명확한 계약으로 만들고 프롬프트에서 역할을 다시 선언하면, 스타일을 유지하면서도 동작과 카메라를 자유롭게 바꾸기 쉬워집니다.

주요 출처

서비스 UI, 입력 한도, 출력 해상도, 워터마크, 요금과 상업적 이용 조건은 계정과 지역에 따라 달라질 수 있습니다. 제작 전 각 공식 서비스의 최신 정책을 확인하세요.