AI 단편 애니메이션은 긴 프롬프트 하나로 완성 영상을 뽑는 작업이 아닙니다. 이야기를 작게 만들고, 승인 가능한 정지 화면으로 분해하고, 한 숏씩 움직인 뒤, 편집과 소리로 다시 하나의 작품으로 묶는 작업입니다. 생성 모델이 좋아져도 이야기의 원인과 결과, 화면 방향, 인물의 외형과 연기, 소리의 타이밍은 제작자가 관리해야 합니다.
이 글에서는 오리지널 단편 《마지막 등불》을 예제로 사용합니다. 등대 기술자 하나가 폭풍 속에서 고장 난 등불 드론의 코어를 들고 도시의 꺼진 등대를 다시 밝힌다는 이야기입니다. 대표 이미지는 최종 키아트이며, 아래 네 삽화는 실제 제작 문서처럼 비트·비주얼 바이블·애니매틱·사운드 편집 단계로 이어집니다.
이 글은 2026년 7월 29일 확인한 OpenAI 이미지 생성 문서, Runway 이미지 투 비디오 가이드, Google Cloud Veo 3.1 프롬프트 가이드, Adobe Firefly 키프레임 안내와 ElevenLabs 사운드 이펙트 문서를 바탕으로 작성했습니다. 기능·요금·상업 이용 조건은 바뀔 수 있으므로 제작 전 최신 정책을 확인하세요.
1. 생성 전에 6개 비트로 이야기의 뼈대를 고정한다
단편은 설정 설명보다 변화가 중요합니다. “누가 무엇을 원하고, 무엇이 막으며, 어떤 선택을 하고, 마지막에 무엇이 달라지는가”를 한 문장으로 적으세요. 그다음 시작·사건·결정·시도·위기·결과의 여섯 그림으로 바꿉니다. 그림만 보고 순서를 설명할 수 없다면 영상 생성 전에 이야기를 고치는 편이 훨씬 저렴합니다.

이미지 제작 프롬프트 — 6비트 스토리 보드
Use case: identity-preserve
Asset type: landscape blog illustration for the story-design section of an AI short-animation guide
Input images: Image 1 is the approved key art and exact visual reference for Hana, the brass lantern drone, the storm city, lighthouse, palette, and hand-painted 2D animation style.
Primary request: Create a six-panel visual beat board for the same original short film, communicating a complete beginning-middle-end without written labels. Panel 1: Hana repairs the dim lantern drone in a small rooftop workshop. Panel 2: the city's lighthouse goes dark as a storm arrives. Panel 3: Hana runs across wet roofs carrying the drone. Panel 4: a violent gust nearly blows the drone away and Hana catches it. Panel 5: inside the beacon room she inserts the glowing drone core into the lighthouse mechanism. Panel 6: the beacon shines across calm water at dawn as Hana and the restored drone watch together.
Subject continuity: preserve Hana's warm brown skin, silver-white single braided ponytail, deep teal eyes, mustard scarf, cobalt cropped coat with two brass clasps, charcoal trousers, brown lace-up boots, round brass belt compass; preserve the spherical brass lantern drone with three petal fins and amber core.
Style/medium: same cinematic hand-painted 2D animation visual development as Image 1, clean cel shapes, restrained gouache texture, production storyboard color keys.
Composition/framing: 16:9 landscape, clean 3-by-2 grid of six equally sized cinematic panels with narrow dark gutters, readable thumbnail silhouettes, consistent left-to-right screen direction and geographic continuity.
Lighting/mood: progresses from warm workshop dusk to cold storm blue, then amber climax and pale sunrise; lighting continuity within each beat.
Constraints: exact same adult Hana and drone throughout; coherent story order; no redesign, no additional people; no text, letters, numbers, captions, arrows, logos, UI, watermark; all panels uncropped within outer canvas.
Avoid: photorealism, 3D render, childlike proportions, costume drift, mirrored compass or clasps, extra limbs, duplicated Hana inside a panel, random props, inconsistent lighthouse.| 비트 | 《마지막 등불》의 사건 | 화면으로 보여줄 변화 |
|---|---|---|
| 일상 | 하나가 약해진 등불 드론을 수리한다 | 작고 따뜻한 빛 |
| 사건 | 폭풍과 함께 도시 등대가 꺼진다 | 도시 전체가 차가운 어둠으로 전환 |
| 결정 | 드론 코어를 들고 작업실을 나선다 | 문턱을 넘는 분명한 행동 |
| 시도 | 젖은 지붕을 달려 등대로 향한다 | 목표를 향하는 화면 방향 |
| 위기·선택 | 돌풍 속에서 코어를 놓칠 뻔한다 | 물건보다 관계를 지키는 손동작 |
| 결과 | 코어가 등대를 밝히고 새벽이 온다 | 같은 도시가 다른 빛으로 보임 |
대사는 먼저 쓰지 않아도 됩니다. 무음으로 이해되는 비트를 만든 뒤 꼭 필요한 정보만 대사로 남기면 립싱크와 번역 부담도 줄어듭니다.
2. 캐릭터·소품·공간을 하나의 비주얼 바이블로 묶는다
정면 얼굴 한 장만으로는 여러 숏을 유지하기 어렵습니다. 정면·측면·후면, 표정, 비대칭 장식, 핵심 소품과 공간 색을 한 장에 모으세요. 이후 모든 숏은 이 바이블을 참조하고, 의상 색·소품 위치·시간대처럼 바뀌면 안 되는 요소를 고정합니다.
OpenAI의 이미지 생성 문서는 이미지 입력을 이용한 생성과 편집, 다중 참조 작업을 지원합니다. 중요한 것은 참조를 많이 넣는 것이 아니라 각 참조의 역할을 명시하는 것입니다. 인물 정체성, 소품 구조, 배경과 화풍을 구분하고 수정할 때도 “배경만 변경, 인물은 유지”처럼 한 번에 한 요소만 바꾸세요.

이미지 제작 프롬프트 — 비주얼 바이블
Use case: identity-preserve
Asset type: landscape blog illustration for the visual-bible section of an AI short-animation production guide
Input images: Image 1 is the approved Hana key art; Image 2 is the approved six-beat board. Use both only as identity, prop, environment, palette, and style references.
Primary request: Create one polished production visual-bible board containing three equal-scale full-body views of the same adult Hana (front, strict side, back), four small consistent head-expression studies (calm, alarmed, determined, relieved), one clean enlarged view of the spherical brass lantern drone with its three petal fins, and one small environment color key of the cliff-city lighthouse in storm light. Do not add written labels.
Subject continuity: exact warm brown skin, silver-white single braided ponytail, deep teal eyes, mustard-yellow scarf, cobalt cropped weatherproof coat with two brass clasps, charcoal trousers, brown lace-up boots, round brass compass at belt; exact spherical brass drone with three symmetrical petal fins and warm amber core.
Style/medium: same cinematic hand-painted 2D animation visual-development style, clean cel shapes with restrained gouache texture, professional model sheet and prop sheet.
Composition/framing: 16:9 landscape, organized production-board layout on a warm neutral paper background, full uncropped turnaround figures aligned on one ground line, expression row clearly separated, drone and environment key balanced on the right; generous margins.
Lighting/mood: flat neutral lighting on character and drone; the environment key alone uses cold storm blue and amber beacon light.
Constraints: do not redesign; preserve adult proportions, facial geometry, braid, coat clasps, scarf, belt compass and boots; no text, letters, numbers, arrows, logos, UI, watermark; no additional people; no cropped feet.
Avoid: photorealism, 3D render, chibi or childlike proportions, mirrored details, costume drift, extra limbs, duplicated accessories, cluttered background.| 고정할 항목 | 기록 방식 | 숏 검수 질문 |
|---|---|---|
| 얼굴·성인 연령 | 얼굴 확대와 표정 4종 | 다른 사람처럼 변하지 않았나 |
| 의상 | 앞·옆·뒤 절개와 색 | 단추·스카프·나침반 위치가 같은가 |
| 소품 | 형태·핀 수·빛 색 | 드론 구조가 숏마다 재설계되지 않았나 |
| 공간 | 등대와 지붕의 상대 위치 | 이동 방향과 목적지가 이어지는가 |
| 빛 | 폭풍 청색 → 등불 호박색 → 새벽 | 이야기 변화와 색 변화가 맞물리는가 |
3. 숏 리스트를 만들고 애니매틱에서 먼저 실패한다
완성 이미지부터 만들지 말고 거친 애니매틱을 먼저 만드세요. 각 행에 숏 번호, 목적, 화면 크기, 카메라, 주체 행동, 시작·종료 상태, 예상 소리를 적고 낮은 비용의 스토리보드 프레임을 편집기에 넣습니다. 이 단계의 목표는 예쁜 화면이 아니라 리듬과 정보 전달입니다.

이미지 제작 프롬프트 — 8숏 애니매틱 보드
Use case: identity-preserve
Asset type: landscape blog illustration for the shot-design and animatic section of an AI short-animation production guide
Input images: Approved references define Hana, the lantern drone, lighthouse city, palette, and exact 2D animation style.
Primary request: Create a polished eight-panel grayscale-plus-color-accent animatic contact sheet for the same rooftop-to-lighthouse sequence. Show eight distinct film shots in story order: extreme wide storm city and dark lighthouse; workshop close-up of the drone core dimming; medium shot Hana deciding to go; low tracking shot of boots running across wet tiles; side wide shot of Hana jumping a roof gap with the drone; close-up of Hana catching the slipping drone; interior over-shoulder shot inserting the amber core into the beacon machine; final extreme wide dawn lighthouse beam over the sea. Use only restrained cobalt and amber accents over readable charcoal storyboard values.
Subject continuity: same adult Hana with warm brown skin, silver-white braid, mustard scarf, cobalt cropped coat, charcoal trousers, boots and belt compass; same spherical brass three-fin amber-core drone.
Style/medium: professional animation animatic frames, expressive cinematic storyboard drawing, clean values, selective hand-painted color accents, original production art.
Composition/framing: 16:9 landscape outer canvas, eight equal panels in a clean 4-by-2 grid with thin neutral gutters, each shot has a clearly different shot size but preserves screen direction, horizon logic and lighthouse geography; readable at thumbnail size.
Constraints: preserve character and prop identity; one coherent sequence; no written shot numbers, no text, letters, captions, arrows, logos, UI, watermark; no extra people.
Avoid: photorealism, 3D render, childlike proportions, costume drift, duplicate Hana within a panel, random camera-axis reversal, extra limbs, cluttered panel borders.| 숏 검증 | 좋은 상태 | 다시 설계할 신호 |
|---|---|---|
| 역할 | 한 숏이 한 가지 정보나 감정을 전달 | 한 숏에 사건이 세 개 이상 |
| 화면 방향 | 이동·시선·바람이 다음 숏과 연결 | 컷마다 좌우가 이유 없이 반전 |
| 시작·종료 | 다음 숏이 받을 자세와 위치가 명확 | 매번 새로운 포즈에서 시작 |
| 카메라 | 이야기 목적이 있는 움직임 | 모든 숏이 과도한 회전·줌 |
| 편집 리듬 | 와이드·미디엄·클로즈업에 이유가 있음 | 같은 크기 화면이 반복 |
완성 길이는 숏 수를 정한 뒤 편집기에서 결정합니다. 영상 생성 프롬프트에 전체 러닝타임을 강제로 쓰기보다 플랫폼의 길이 설정을 사용하고, 한 생성에는 한 핵심 동작을 맡기는 편이 안정적입니다.
4. 정지 화면이 외형을, 영상 프롬프트가 시간을 담당하게 한다
Runway의 최신 가이드는 입력 이미지가 구도·주체·조명·스타일을 이미 제공하므로 이미지 투 비디오 텍스트는 주체 행동, 환경 움직임, 카메라 움직임과 진행 순서에 집중하라고 권합니다. 먼저 가장 중요한 움직임만 적고 결과를 본 뒤 한 요소씩 더하세요. 모션 블러나 바람 방향이 프롬프트와 충돌하는 입력 이미지는 영상에서도 문제가 확대될 수 있습니다.
영상 제작 프롬프트 — 감정 연기 숏
Locked medium close-up inside the rooftop workshop. The engineer steadies the dim lantern drone with both hands, watches its amber core flicker, takes one controlled breath, then lifts her eyes toward the lighthouse outside the rain-streaked window. Her worried expression settles into determination. Loose silver hair strands and the end of her mustard scarf respond softly to the draft; tiny brass fins tremble and stop. The camera remains still, with only a gentle focus shift from the failing core to her eyes. Continuous seamless shot, restrained acting, natural hand contact, no cut.영상 제작 프롬프트 — 지붕 추격 액션 숏
A side-tracking camera moves with the engineer as she runs across wet roof tiles while cradling the lantern drone. Her boots make firm contact and spray small arcs of rainwater; her braid, scarf, coat hem, compass, and the drone fins trail with delayed secondary motion in the same wind direction. She compresses into one decisive jump over a narrow gap, lands with visible weight, regains balance, and continues toward the lighthouse. Stable horizon, coherent left-to-right screen direction, physically plausible momentum, continuous seamless shot.영상 제작 프롬프트 — 시작·종료 프레임 기반 등대 점등 숏
Animate one continuous transition between the supplied first and last frames. The engineer seats the glowing brass core into the beacon mechanism and twists it until the gears engage. The drone fins fold inward; warm amber light travels through concentric glass rings, grows into a powerful beam, and sweeps across the rain-dark sea. The storm rain softens as the room changes from cold blue to warm gold. The camera performs a slow restrained arc around her shoulder and settles on the same final composition. SFX: close metal latch, rising mechanical hum, deep gear engagement, rain against glass, distant ocean. No dialogue, no cut, preserve identity, geometry, light direction, and screen space.Google의 Veo 가이드는 촬영법·주체·행동·맥락·스타일과 분위기를 분리하는 구조를 제시하고, 첫·마지막 프레임과 참조 이미지를 이용한 여러 숏 제작을 설명합니다. Adobe Firefly도 첫·마지막 이미지를 영상의 고정점으로 사용할 수 있습니다. 단, 두 프레임의 렌즈·높이·조명·인물 크기가 크게 다르면 중간 프레임이 무너지기 쉬우므로 먼저 호환성을 확인하세요.
5. 영상보다 먼저 소리를 설계하고 마지막에 함께 잠근다
단편의 규모를 키우는 가장 저렴한 요소는 소리입니다. 대사, 폴리, 환경음, 효과음, 음악을 별도 트랙으로 나누고 애니매틱 단계에서 임시 소리를 붙이세요. ElevenLabs의 공식 문서도 복잡한 소리 묶음보다 개별 효과음을 만들어 편집기에서 결합하는 방식을 권합니다. 빗소리·발소리·금속 잠금·기계 공명처럼 사건에 반응하는 소리는 화면의 무게감을 보완합니다.

이미지 제작 프롬프트 — 사운드·편집·최종 검수 보드
Use case: identity-preserve
Asset type: landscape blog illustration for the sound-design, edit, and final-QA section of an AI short-animation production guide
Input images: Approved references define Hana, the lantern drone, beat order, visual bible, lighthouse geography, palette, and exact hand-painted 2D animation style.
Primary request: Create a cinematic sound-and-edit production board without written labels. A large central finished frame shows Hana inside the lighthouse beacon room inserting the glowing lantern drone core into the mechanism while the first golden beam ignites. Around it are four smaller inset detail frames that visually represent separate audio and edit layers: rain striking a metal window and wind-blown scarf; boots landing heavily on wet roof tiles; close-up of the brass drone fins vibrating with a warm mechanical hum; giant beacon gears engaging with a resonant pulse. Along the bottom, include a purely visual filmstrip of five tiny continuity thumbnails progressing from storm darkness to sunrise, accompanied by abstract clean waveform shapes in matching colors, but no UI text or numbers.
Subject continuity: exact same adult Hana, warm brown skin, silver-white single braid, teal eyes, mustard scarf, cobalt coat with two brass clasps, charcoal trousers, boots, belt compass; exact spherical brass three-fin amber-core drone.
Style/medium: premium original 2D animation production art, clean cel shapes, restrained gouache texture, cinematic color finish, sophisticated editorial contact-sheet design.
Composition/framing: 16:9 landscape, dominant central frame, four balanced detail insets, narrow gutters, bottom filmstrip and abstract waveform motifs; clean hierarchy and safe margins for 1024x576.
Constraints: preserve character, costume, prop and environment identity; no written text, letters, numbers, captions, logos, interface labels, brand marks, watermark; no extra people; no copyrighted elements.
Avoid: photorealism, 3D render, childlike proportions, costume drift, random electronics, illegible clutter, extra limbs, inconsistent light direction.최종 편집에서는 생성한 모든 것을 쓰지 않습니다. 같은 역할의 후보 중 가장 잘 이어지는 한 숏을 선택하고, 공유 프레임이 겹치면 잘라내며, 필요할 때만 안정화·속도 조절·색 보정을 적용합니다. 대사 없는 버전을 먼저 완성하면 세 언어 자막과 더빙도 관리하기 쉽습니다.
실패한 숏만 다시 만드는 품질 게이트
| 검사 축 | 확인할 항목 | 실패 시 조치 |
|---|---|---|
| 이야기 | 행동의 원인과 결과가 읽힘 | 비트 또는 숏 순서 수정 |
| 정체성 | 얼굴·의상·소품이 동일 | 승인 참조에서 해당 숏만 재생성 |
| 동작 | 발 접지·무게·손 접촉이 자연스러움 | 시작 포즈와 행동 수를 단순화 |
| 연속성 | 시선·화면 방향·빛·날씨가 연결 | 종료 프레임을 다음 숏 입력으로 사용 |
| 음향 | 효과음이 화면 사건과 정확히 맞음 | 개별 효과음을 다시 생성해 타임라인에서 조정 |
| 권리·안전 | 입력·음성·음악의 사용 권한 확인 | 권리 불명 자산 교체, 실존 인물 동의 확인 |
결론: 단편은 생성이 아니라 승인 과정이다
완성도를 높이는 핵심은 가장 긴 프롬프트가 아니라 작은 단위로 결정하고 승인하는 순서입니다. 로그라인과 비트가 이야기의 방향을 고정하고, 비주얼 바이블이 정체성을 고정하며, 애니매틱이 리듬을 고정합니다. 그 뒤 한 숏씩 움직이고 소리와 편집으로 연결하면 실패한 부분만 되돌릴 수 있습니다.
처음에는 등장인물 1명, 장소 1~2곳, 핵심 소품 1개, 대사 0~2줄로 제한하세요. 한 편을 끝까지 완성한 제작 문서와 승인 프레임은 다음 작품에서 가장 강력한 자산이 됩니다.
참고 자료
- OpenAI — Image generation guide
- Runway — Image to Video Prompting Guide
- Google Cloud — The ultimate prompting guide for Veo 3.1
- Adobe Firefly — Create cinematic video from prompts and keyframes
- ElevenLabs — Sound effects guide
생성 결과는 프레임마다 오류가 생길 수 있습니다. 실존 인물의 얼굴·음성, 저작권 캐릭터, 음악과 참조 이미지에는 필요한 동의와 이용 권리를 확인하고 공개 전 사람이 전체 영상을 검수하세요.