muapi-seedance-2

Expert Cinema Director skill for Seedance 2.0 (ByteDance) — high-fidelity video generation across Chinese, Global, and VIP tiers.

Paste into Claude, ChatGPT or Cursor.

Read the source632 lines
muapi-seedance-2/SKILL.md632 lines26.5 KBpushed 100d agoRawView on GitHub
1---
2slug: muapi-seedance-2
3name: muapi-seedance-2
4version: "0.3.0"
5description: Expert Cinema Director skill for Seedance 2.0 (ByteDance) — high-fidelity video generation across Chinese, Global, and VIP tiers. Supports text-to-video, image-to-video, first-last-frame, omni reference, character training, omni-reference training, video editing, and watermark removal.
6acceptLicenseTerms: true
7---
8 
9# 🎬 Seedance 2.0 Cinema Expert
10 
11**The definitive skill for "Director-Level" AI video orchestration.**
12Seedance 2.0 is not a descriptive model; it is an *instructional* model. It responds best to technical cinematography, physics directives, and precise camera grammar.
13 
14## Core Competencies
15 
161. **Text-to-Video (t2v)**: Generate cinematic video from a Director Brief — Chinese, Global, or VIP tier.
172. **Image-to-Video (i2v)**: Animate 1–9 reference images — Chinese, Global (smart mode), or VIP tier.
183. **Video Extension (extend)**: Seamlessly continue an existing Seedance 2.0 video (Chinese tier).
194. **First & Last Frame (first-last)**: Interpolate a fluid video between a start image and end image (Global/VIP).
205. **Omni Reference (omni)**: Full multimodal reference with images + audio + character refs (all tiers).
216. **Omni Reference Training (omni-train)**: Train a custom persistent character for identity-consistent generation.
227. **Character Sheet (character)**: Build a reusable character from 1–3 images (Chinese tier).
238. **Video Edit (video-edit)**: Edit an existing video with a prompt + optional reference images (Chinese tier).
249. **Watermark Removal (watermark-remove)**: Strip Seedance 2.0 watermarks (basic or Pro).
25 
26---
27 
28## 🏷️ Tiers
29 
30| Tier | Flag | Censorship | Aspect Ratios | Duration | Quality param |
31|:---|:---|:---|:---|:---|:---|
32| **Chinese** (default) | `--tier chinese` | Low | 16:9, 9:16, 4:3, 3:4 | 5 / 10 / 15 s | Yes (basic/high) |
33| **Global** | `--tier global` | Standard | + 21:9, 1:1 | Any 4–15 s | No |
34| **VIP** | `--tier vip` | Low | + 21:9, 1:1 | Any 4–15 s | No |
35 
36Add `--fast` to any Global or VIP call to use the fast-queue variant (lower latency, same quality).
37 
38---
39 
40## 📥 Input Limits
41 
42| Input Type | Chinese i2v/omni | Global/VIP i2v/omni | Formats | Max Size |
43|:---|:---|:---|:---|:---|
44| Images | ≤ 9 | ≤ 9 | jpeg, png, webp | 30 MB each |
45| Videos | ≤ 3 (omni only) | Not supported | mp4, mov | 50 MB each |
46| Audio | ≤ 3 | ≤ 3 | mp3, wav | 15 MB each |
47| **First-Last** | — | 1–2 images | jpeg, png, webp | 30 MB each |
48| **Video Edit** | 1 video + ≤ 9 imgs | — | mp4 ≤ 10 MB / 15s | — |
49 
50**Output**: 4–15 seconds, auto-generated sound, 480p–720p.
51 
52---
53 
54## ⚠️ Restrictions
55 
56- **No realistic human faces** in uploaded images/videos (except character/omni-train modes).
57- `--mode extend` requires a `request_id` from a prior `seedance-v2.0-t2v` or `seedance-v2.0-i2v` job.
58- `--mode first-last` requires `--tier global` or `--tier vip`.
59- Global/VIP omni does **not** support video references (images + audio only).
60- `--quality` applies to Chinese tier only.
61 
62---
63 
64## 🔗 Core Syntax: The @ Reference System
65 
66Assign explicit roles to each uploaded asset. Tags differ by mode.
67 
68### Chinese Tier (i2v, omni)
69```
70@image1 @image2 ... @image9 (images_list order)
71@video1 @video2 @video3 (video_files order)
72@audio1 @audio2 @audio3 (audio_files order)
73```
74 
75### Global/VIP Omni (omni-reference-no-video / vip-omni-reference)
76```
77@image1 @image2 ... @image9 (images_list order)
78@audio1 @audio2 @audio3 (audio_files order)
79```
80 
81### Character References (all tiers)
82```
83@character:<request_id> — from seedance-2-character or completed t2v/i2v job
84@omni-character:<character_id> — from seedance-2-omni-reference-train output
85```
86 
87### Role Assignment Table
88 
89| Purpose | Example Syntax |
90|:---|:---|
91| First frame | `@Image1 as the first frame` |
92| Last frame | `@Image2 as the last frame` |
93| Character appearance | `@Image1's character as the subject` |
94| Scene / background | `scene references @Image3` |
95| Camera movement | `reference @Video1's camera movement` |
96| Action / motion | `reference @Video1's action choreography` |
97| Visual effects | `completely reference @Video1's effects and transitions` |
98| Rhythm / tempo | `video rhythm references @Video1` |
99| Voice / tone | `narration voice references @Video1` |
100| Background music | `BGM references @Audio1` |
101| Sound effects | `sound effects reference @Video3's audio` |
102| Outfit / clothing | `wearing the outfit from @Image2` |
103| Product appearance | `product details reference @Image3` |
104 
105### Multi-Reference Combination
106```
107@Image1's character as the subject, reference @Video1's camera movement
108and action choreography, BGM references @Audio1, scene references @Image2
109```
110 
111---
112 
113## 🏗️ Technical Specification: The Director Brief
114 
115Structure prompts using this six-component hierarchy. Order matters — composition first, texture and micro-motion last:
116 
117| Component | Instruction Type | Example |
118|:---|:---|:---|
119| **Scene** | Environment + Lighting | "A rain-soaked cyberpunk street, magenta neon reflections on wet asphalt." |
120| **Subject** | Identity + Detail | "A woman in a black trenchcoat, determined focus, cinematic skin textures." |
121| **Action** | Fluid Interaction | "Walking forward through the crowd, coat billowing slightly in the wind." |
122| **Camera** | Movement + Lens + Speed | "Medium tracking shot, 35mm lens, slow dolly backward over 6s. Subtle handheld jitter." |
123| **Audio** | Music + SFX + Ambience | "Low ambient hum, distant traffic, single piano note at 5s. No dialogue." |
124| **Pacing/Style** | Timing + Mood + Grade | "Cinematic epic, warm color grade, shallow DOF. Slow build — single action only, no scene cuts." |
125 
126> **Seedance 2.0 generates audio natively.** Always include an Audio directive — even one sentence. Without it the model generates random ambient sound that may not match your scene.
127 
128### Time-Segmented Prompts (Recommended for 10s+ videos)
129Break prompts into timed segments for precise control:
130```
1310–3s: [opening scene, camera move, establishing action]
1323–6s: [mid-section development, subject in motion]
1336–10s: [climax or key action beat]
13410–15s: [resolution, brand/product hold, text/tagline fade in]
135```
136 
137> **Single-beat rule:** Each segment should contain one action. 4–7s = one beat. 10–15s = 3–4 beats maximum. Overloading a segment with multiple narrative changes degrades output quality.
138 
139### Negative Prompting
140 
141Seedance 2.0 supports appending negative guidance directly in the prompt. Use plain language at the end:
142 
143```
144[your director brief above]
145Avoid: camera shake, jump cuts, lens distortion, overexposure, watermarks, text overlays.
146```
147 
148Common negative additions:
149- `Avoid: abrupt cuts, scene changes, multiple locations.` (for single-take shots)
150- `Avoid: human faces, realistic people.` (for product-only content)
151- `Avoid: fast motion, blur, unstable framing.` (for smooth product reveals)
152 
153---
154 
155## 🎥 Camera Language Reference
156 
157### Basic Movements
158| Term | Description |
159|:---|:---|
160| Push in / Slow push | Camera moves toward subject |
161| Pull back / Pull away | Camera moves away from subject |
162| Pan left/right | Camera rotates horizontally |
163| Tilt up/down | Camera rotates vertically |
164| Track / Follow shot | Camera follows subject movement |
165| Orbit / Revolve | Camera circles around subject |
166| One-take / Oner | Continuous shot with no cuts |
167 
168### Advanced Techniques
169| Term | Description |
170|:---|:---|
171| Hitchcock zoom (dolly zoom) | Push in + zoom out — creates vertigo effect |
172| Fisheye lens | Ultra-wide distorted lens |
173| Low angle / High angle | Camera below/above subject |
174| Bird's eye / Overhead | Top-down view |
175| First-person POV (FPV) | Immersive subjective camera from character/object's eyes — GoPro-style wide angle, forward motion, no cuts |
176| Drone flythrough | Cinematic aerial descent — gimbal-stabilized, sweeping lateral arc, DJI Inspire aesthetic |
177| Architectural flythrough | Ground-level continuous dolly through connected spaces — one-take, practical lighting |
178| Whip pan | Very fast horizontal pan with motion blur |
179| Crane shot | Vertical movement like a crane arm |
180 
181### Shot Sizes
182| Term | Description |
183|:---|:---|
184| Extreme close-up | Eyes, mouth, or small detail only |
185| Close-up | Face fills frame |
186| Medium close-up | Head and shoulders |
187| Medium shot | Waist up |
188| Full shot | Entire body |
189| Wide / Establishing shot | Full environment |
190 
191---
192 
193## 🧠 Prompt Optimization Protocol
194 
195**The Agent MUST transform user intent into a technical "Director Brief" before execution.**
196 
1971. **Technical Grammar**: Use camera terms: *Dolly In/Out, Crane Shot, Whip Pan, Tracking Shot, Anamorphic Lens, Shallow Depth of Field, High-Speed Dive, Orbital Arc*.
1982. **Physics Directives**: Use "caustic patterns," "volumetric rays," or "subsurface scattering" instead of "good lighting."
1993. **Timecode Notation**: For multi-beat scenes, use `[00:00-00:05s]` format to specify timing.
2004. **Tag References**: If files provided, use: *"Replicate the camera movement of @video1 while maintaining the visual style of @image1."* (lowercase, 1-based index)
2015. **ORDER MATTERS**: Tokens at the start define composition; tokens at the end define texture and micro-motion.
2026. **Multi-Image i2v**: Provide up to 9 reference images. The model blends aspects (style, identity, environment) across all inputs.
2037. **Audio is mandatory**: Seedance 2.0 generates audio natively. Always include an Audio line — music genre/tone, key SFX, ambient texture. Silent direction = random audio.
2048. **Single-beat discipline**: Each timed segment = one action. Cramming two narrative beats into 4s degrades physics and motion consistency.
205 
206---
207 
208## 🎭 Capability-Specific Patterns
209 
210### 1. Character Consistency
211```
212The man in @Image1 walks tiredly down the hallway, slowing his steps,
213finally stopping at his front door. Close-up on his face — he takes a
214deep breath, replaces the weariness with a relaxed expression.
215Maintain high character consistency, zero facial flicker, persistent clothing details.
216```
217 
218### 2. Camera Movement Replication
219```
220Reference @Image1's male character. He is in @Image2's elevator.
221Completely reference @Video1's camera movements and facial expressions.
222Hitchcock zoom during the fear moment, then orbit shots of the interior.
223Elevator doors open, follow shot walking out.
224```
225 
226### 3. Video Extension (Forward)
227```
228Extend @Video1 by 10 seconds.
2291–5s: Light and shadow slowly slide across table through venetian blinds.
2306–10s: A coffee bean drifts down. Camera pushes in toward it until screen goes black.
231English text gradually appears — "Lucky Coffee", "Breakfast", "AM 7:00-10:00".
232```
233 
234### 4. Video Extension (Reverse / Prepend)
235```
236Extend backward 10s. In warm afternoon light, the camera starts from
237the corner with awning fluttering in the breeze, slowly tilting down
238to flowers peeking out at the wall base, building anticipation for the main scene.
239```
240 
241### 5. Video Editing (Modify Existing)
242```
243Subvert @Video1's plot — the character's expression shifts from warmth to
244cold determination. The action is decisive, without hesitation.
245Maintain all other visual elements (scene, lighting, timing).
246```
247 
248### 6. Music Beat-Matching
249```bash
250bash scripts/generate-seedance.sh \
251 --mode i2v \
252 --file img1.jpg --file img2.jpg --file img3.jpg \
253 --video-file reference_edit.mp4 \
254 --audio-file track.mp3 \
255 --subject "@Image1 @Image2 @Image3 — match the keyframe positions and rhythm of @Video1 for beat-synced cuts. BGM references @Audio1. More dynamic movement, dreamlike visual style." \
256 --duration 15 --quality high
257```
258 
259### 7. Dialogue / Voice Acting
260```
261In the "Cat & Dog Roast Show" — emotionally expressive comedy segment:
262Cat host (licking paw, rolling eyes): "Who understands my suffering?"
263Dog host (head tilted, tail wagging): "You're one to talk? You sleep 18 hours a day..."
264Sound: lively studio ambience, audience laughter, punchy transitions.
265```
266 
267### 8. One-Take / Long Take
268```
269@Image1 @Image2 @Image3 — one-take tracking shot following a runner
270from the street up stairs, through a corridor, onto a rooftop,
271finally overlooking the city. No cuts throughout.
272```
273 
274### 9. E-commerce / Product Showcase
275```bash
276bash scripts/generate-seedance.sh \
277 --mode i2v \
278 --file product.jpg \
279 --subject "Deconstruct the product. Static camera. Hamburger suspended mid-air, rotating slowly. Ingredients separate and reassemble. Cheese continues to melt and drip. Ultimate food aesthetics." \
280 --intent "product" \
281 --aspect "9:16" \
282 --duration 15 --quality high
283```
284 
285### 10. Science / Educational Visualization
286```bash
287bash scripts/generate-seedance.sh \
288 --subject "15-second health educational clip. 0–5s: Transparent blue human upper body, camera pushes into a clear artery, blood flows smoothly. 5–10s: Sugar and fat particles enter bloodstream, lipid deposits form on vessel walls. 10–15s: Vessel narrows, before/after comparison. 4K medical CGI, semi-transparent visualization." \
289 --intent "educational" \
290 --duration 15 --quality high
291```
292 
293### 11. FPV First-Person Shot
294```bash
295bash scripts/generate-seedance.sh \
296 --subject "Immersive first-person POV shot. Camera glides at eye level through a narrow mountain trail,
297trees rushing past in peripheral blur, rocky terrain below. Slight natural stabilization with wide-angle lens.
298Continuous forward motion, no cuts throughout. Trail opens into a clearing — mountain peak visible ahead.
299Sound: wind, footsteps on gravel, distant birds. Natural ambient audio, no music." \
300 --intent "fpv" \
301 --aspect "9:16" --duration 10 --quality high
302```
303 
304### 12. Cinematic Drone Flythrough
305```bash
306bash scripts/generate-seedance.sh \
307 --subject "Cinematic aerial drone shot. Camera starts at 150m altitude above a coastal city at golden hour.
308Smooth gimbal-stabilized descent along a sweeping lateral arc, dropping toward a rooftop terrace.
309Long shadows cast across building tops, warm light on ocean surface. High-speed dive closes in
310to product on the terrace — final frame settles into a medium close-up.
311Sound: gentle wind, distant city hum, soft cinematic score building to resolve." \
312 --intent "drone" \
313 --aspect "16:9" --duration 10 --tier global --view
314```
315 
316---
317 
318## 🎨 Prompt Templates
319 
320### Cinematic Film
321```
322[SCENE] Rain-soaked cyberpunk alley, neon signs reflected on wet cobblestones.
323[SUBJECT] A lone figure in a weathered trench coat, face obscured by a wide-brim hat.
324[ACTION] Walking slowly, each step splashing neon color into the puddles.
325[CAMERA] Low-angle tracking shot, anamorphic lens, slow dolly in. Rack focus to face.
326[STYLE] Denis Villeneuve aesthetic, high contrast, desaturated blues and magentas. 24fps.
327```
328 
329### Product Ad (15s)
330```
331Reference @Video1's editing style. Replace @Video1's product with @Image1 as hero.
3320–3s: Product enters with dynamic rotation, close-up on surface texture and logo.
3334–8s: Multiple angle transitions — front, side, back — with highlight scanning light.
3349–12s: Product in lifestyle context showing usage.
33513–15s: Hero shot with brand tagline, background music builds to resolution.
336Sound: Reference @Video1's BGM. Add product interaction sound effects.
337```
338 
339### Short Drama (15s)
340```
341Scene (0–5s): Close-up on character's reddened eyes, finger pointing accusingly.
342Dialogue 1: "What exactly are you trying to take from me?"
343Scene (6–10s): Other character trembles, holding up evidence, steps forward.
344Dialogue 2: "I'm not deceiving you! This is what he entrusted to me!"
345Scene (11–15s): Evidence revealed, first character freezes — anger shifts to shock.
346Sound: Urgent piano + static interference, sobbing, muffled voice blending in.
347Duration: Precise 15 seconds, every frame tight, no filler.
348```
349 
350### Dance / Beat-Sync (13s)
351```
352Have the character in @Image1 replicate the dance moves and beat-synced
353music from @Video1. Generate a 13-second video. Movements should be
354smooth with no stuttering or freezing.
355```
356 
357### Scenery Montage (15s)
358```
359@Image1 @Image2 @Image3 @Image4 @Image5 @Image6 — landscape scene images.
360Reference @Video1's visual rhythm, inter-scene transitions, visual style,
361and music tempo for beat-synced editing.
362```
363 
364### Advertising / Product Motion
365```
366[SCENE] Minimalist white studio, single product on a rotating pedestal.
367[ACTION] Subtle 360° rotation, product details catching specular highlights.
368[CAMERA] Tight medium shot, macro lens pass over surface texture, slow orbit.
369[STYLE] Commercial grade, perfect exposure, zero background distraction.
370```
371 
372### Action / Physics
373```
374[SCENE] Desert canyon at sunrise, sandy terrain, long shadows.
375[SUBJECT] High-performance sports car accelerating through a turn.
376[ACTION] Rear wheels spinning with dust plume, chassis flexing under g-force.
377[CAMERA] Low hero angle dolly tracking alongside, then whip pan to lead car.
378[STYLE] Hollywood racing film, warm golden grade, motion blur on wheels. 24fps.
379```
380 
381### Character Consistency (Martial Arts)
382```
383[SUBJECT] Same fighter throughout: young woman, white gi, black belt, determined expression.
384[ACTION] Fluid kata sequence — rising block, stepping side kick, spinning back fist.
385[CAMERA] Full-body wide shot, then cut to close-up of fist impact in slow motion.
386[STYLE] Maintain identical lighting, clothing, and facial features in every frame. Zero flicker.
387```
388 
389---
390 
391## 🎚️ Style & Quality Modifiers
392 
393### Visual Style
394- `Cinematic quality, film grain, shallow depth of field`
395- `2.35:1 widescreen, 24fps`
396- `Ink wash painting style` / `Anime style` / `Photorealistic`
397- `High saturation neon colors, cool-warm contrast`
398- `4K medical CGI, semi-transparent visualization`
399 
400### Mood / Atmosphere
401- `Tense and suspenseful` / `Warm and healing` / `Epic and grand`
402- `Comedy with exaggerated expressions`
403- `Documentary tone, restrained narration`
404 
405### Audio Direction
406- `Background music: grand and majestic`
407- `Sound effects: footsteps, crowd noise, car sounds`
408- `Voice tone reference @Video1`
409- `Beat-synced transitions matching music rhythm`
410 
411---
412 
413## ❌ Common Mistakes to Avoid
414 
4151. **Vague references**: Don't say "reference @Video1" — specify WHAT to reference (camera? action? effects? rhythm?)
4162. **Conflicting instructions**: Don't ask for "static camera" and "orbit shot" in the same segment.
4173. **Overloading**: Don't pack too many scenes into 4–5 seconds — keep it physically plausible.
4184. **Missing @ assignments**: If you upload 5 images, make sure each one is referenced with a clear purpose.
4195. **Ignoring audio**: Sound design dramatically improves output — always include audio direction.
4206. **Forgetting duration**: Match prompt complexity to the selected generation length.
4217. **Real faces**: Don't upload real human photos — the system will block them.
4228. **Keyword soup**: DO NOT use "8k, masterpiece, trending." Use technical descriptions instead.
4239. **Discontinuous action**: Avoid "The man runs and then he stops." Use fluid transitional language.
42410. **Missing audio direction**: Seedance 2.0 generates audio natively — always specify music tone, SFX, or ambience. Skipping it produces random sound.
42511. **Narrative overload per segment**: Each timed segment should contain one action beat. Multiple scene changes in 4s produce degraded physics and motion artifacts.
42612. **FPV without continuous motion**: FPV requires a motion-rich environment to work — a static room with FPV intent will not trigger the immersive effect. Pair FPV with corridors, streets, natural terrain, or product flyovers.
42713. **Drone without a destination**: Drone shots need a resolve point — specify what the camera descends toward or arrives at. "Drone shot" alone produces aimless floating.
428 
429---
430 
431## 🚀 Protocol: All Modes
432 
433### Mode 1: Text-to-Video (t2v)
434 
435```bash
436# Chinese tier (default) — epic reveal
437bash scripts/generate-seedance.sh \
438 --subject "hidden Andes temple, mist through the canopy" \
439 --intent epic --aspect "16:9" --duration 10 --quality high --view
440 
441# Global tier — 21:9 cinematic, 12s
442bash scripts/generate-seedance.sh \
443 --tier global \
444 --subject "neon cyberpunk alley, rain-slicked streets" \
445 --intent tense --aspect "21:9" --duration 12 --view
446 
447# VIP fast — square social format
448bash scripts/generate-seedance.sh \
449 --tier vip --fast \
450 --subject "product rotating on a pedestal, specular highlights" \
451 --intent product --aspect "1:1" --duration 6
452```
453 
454### Mode 2: Image-to-Video (i2v)
455 
456```bash
457# Chinese tier — animate with video/audio refs
458bash scripts/generate-seedance.sh --mode i2v \
459 --file character.jpg --video-file ref_motion.mp4 --audio-file bgm.mp3 \
460 --subject "@image1's character walks forward, @video1's camera movement, BGM references @audio1" \
461 --quality high --view
462 
463# Global tier — 1 image = first frame anchor
464bash scripts/generate-seedance.sh --mode i2v --tier global \
465 --file hero.jpg \
466 --subject "hero strides forward, coat billowing in slow motion" \
467 --duration 8 --view
468 
469# VIP fast — 3 images, omni ref mode (2-9 images switches to omni)
470bash scripts/generate-seedance.sh --mode i2v --tier vip --fast \
471 --file char.jpg --file env.jpg --file style.jpg \
472 --subject "@image1 character walks through @image2's environment in @image3's style" \
473 --duration 10
474```
475 
476### Mode 3: Extend Video (Chinese tier)
477 
478```bash
479# Extend naturally
480bash scripts/generate-seedance.sh --mode extend \
481 --request-id "abc-123-def-456" --duration 10
482 
483# Extend with directional prompt
484bash scripts/generate-seedance.sh --mode extend \
485 --request-id "abc-123-def-456" \
486 --subject "camera continues pulling back, revealing the vast city below" \
487 --intent reveal --duration 10 --quality high --view
488```
489 
490### Mode 4: First & Last Frame (Global/VIP)
491 
492```bash
493# One image = first frame anchor
494bash scripts/generate-seedance.sh --mode first-last --tier global \
495 --file opening_scene.jpg \
496 --subject "smooth cinematic push into the scene" --duration 6 --view
497 
498# Two images = interpolate between first and last frame
499bash scripts/generate-seedance.sh --mode first-last --tier vip --fast \
500 --file start.jpg --file end.jpg \
501 --subject "dramatic reveal transition between the two frames" --duration 8 --view
502```
503 
504### Mode 5: Omni Reference (omni)
505 
506```bash
507# Chinese tier — images + video + audio refs
508bash scripts/generate-seedance.sh --mode omni --tier chinese \
509 --file character.jpg --video-file ref_edit.mp4 --audio-file track.mp3 \
510 --subject "@image1's character performs moves from @video1, BGM references @audio1" \
511 --duration 15 --quality high --view
512 
513# Global tier — images + audio, no video refs
514bash scripts/generate-seedance.sh --mode omni --tier global \
515 --file portrait.jpg --audio-file bgm.mp3 \
516 --subject "@image1 is the main character. Walking through a neon-lit city at night. BGM references @audio1." \
517 --aspect "16:9" --duration 8 --view
518 
519# VIP tier — with trained omni character
520bash scripts/generate-seedance.sh --mode omni --tier vip --fast \
521 --subject "@omni-character:char_1775422630065_4vbana walks through a garden at golden hour" \
522 --aspect "16:9" --duration 10 --view
523 
524# With @character ref (from character mode)
525bash scripts/generate-seedance.sh --mode omni --tier global \
526 --subject "@character:cab9517f-1818-4910-8d66 walks down a rain-soaked alley, cinematic tracking shot" \
527 --duration 8 --view
528```
529 
530### Mode 6: Train Omni Reference Character (omni-train)
531 
532```bash
533# Train from a single portrait
534bash scripts/generate-seedance.sh --mode omni-train \
535 --file portrait.jpg \
536 --character-name "Alex" \
537 --character-desc "A brave explorer with piercing blue eyes"
538 
539# Use in omni prompts after training completes:
540# @omni-character:<character_id returned>
541```
542 
543### Mode 7: Character Sheet (character, Chinese tier)
544 
545```bash
546# Build character from 1–3 reference images
547bash scripts/generate-seedance.sh --mode character \
548 --file ref1.jpg --file ref2.jpg \
549 --character-name "Hero" \
550 --subject "red leather jacket with black jeans and white sneakers"
551 
552# Use the returned request_id in t2v/i2v/omni:
553# @character:<request_id>
554```
555 
556### Mode 8: Video Edit (video-edit, Chinese tier)
557 
558```bash
559# Replace subject in an existing video
560bash scripts/generate-seedance.sh --mode video-edit \
561 --video-url "https://example.com/input.mp4" \
562 --file replacement_character.jpg \
563 --subject "Replace the running man with @image1. Preserve exact motion, speed, and camera shake." \
564 --quality high --view
565 
566# Edit with watermark removal in one step
567bash scripts/generate-seedance.sh --mode video-edit \
568 --video-file source.mp4 \
569 --subject "Subvert the plot — the character's expression shifts from warmth to cold determination." \
570 --remove-watermark --view
571```
572 
573### Mode 9: Watermark Removal (watermark-remove)
574 
575```bash
576# Basic watermark removal
577bash scripts/generate-seedance.sh --mode watermark-remove \
578 --video-url "https://example.com/seedance_output.mp4" --view
579 
580# Pro watermark removal (100MB limit, better quality)
581bash scripts/generate-seedance.sh --mode watermark-remove \
582 --video-file my_video.mp4 --pro --view
583```
584 
585### Async Pattern
586 
587```bash
588# Submit and get request_id immediately
589RESULT=$(bash scripts/generate-seedance.sh --tier vip --fast --subject "..." --async --json)
590REQUEST_ID=$(echo "$RESULT" | jq -r '.request_id')
591 
592# Check status later
593bash ../../../../core/media/generate-video.sh --result "$REQUEST_ID"
594```
595 
596---
597 
598## ⚙️ Implementation Details
599 
600### Endpoint Reference
601 
602| Mode | Tier | Endpoint |
603|:---|:---|:---|
604| `t2v` | chinese | `seedance-v2.0-t2v` |
605| `t2v` | global | `seedance-2-text-to-video{-fast}` |
606| `t2v` | vip | `seedance-2-vip-text-to-video{-fast}` |
607| `i2v` | chinese | `seedance-v2.0-i2v` |
608| `i2v` | global | `seedance-2-image-to-video{-fast}` |
609| `i2v` | vip | `seedance-2-vip-image-to-video{-fast}` |
610| `extend` | chinese | `seedance-v2.0-extend` |
611| `first-last` | global | `seedance-2-first-last-frame{-fast}` |
612| `first-last` | vip | `seedance-2-vip-first-last-frame{-fast}` |
613| `omni` | chinese | `seedance-2.0-omni-reference` |
614| `omni` | global | `seedance-2-omni-reference-no-video{-fast}` |
615| `omni` | vip | `seedance-2-vip-omni-reference{-fast}` |
616| `omni-train` | any | `seedance-2-omni-reference-train` |
617| `character` | any | `seedance-2-character` |
618| `video-edit` | chinese | `seedance-v2.0-video-edit` |
619| `watermark-remove` | — | `seedance-2.0-watermark-remover` / `seedance-2-video-watermark-remover-pro` |
620 
621### Parameter Differences by Tier
622 
623| Parameter | Chinese | Global | VIP |
624|:---|:---|:---|:---|
625| `aspect_ratio` | 16:9, 9:16, 4:3, 3:4 | + 21:9, 1:1 | + 21:9, 1:1 |
626| `duration` | 5 / 10 / 15 (enum) | 4–15 (any int) | 4–15 (any int) |
627| `quality` | basic / high | — (not supported) | — (not supported) |
628| `video_files` (omni) | ✅ up to 3 | ❌ | ❌ |
629| `audio_files` (omni) | ✅ up to 3 | ✅ up to 3 | ✅ up to 3 |
630| Fast variant | ❌ | ✅ (`--fast`) | ✅ (`--fast`) |
631 
632This skill acts as a **Cinematographic Wrapper** that translates creative intent into high-fidelity technical instructions for the `muapi` core.

Discussion

Alternatives

Also in Training