Render street interview skill

Build a vox-pop street interview video ad.

by gooseworks-ai·MIT license·★ 1,220 Stars on the repo·GitHub ↗

Use now

Files of Render street interview

gooseworks-ai/main1 file shown
SKILL.md
Show the full text112 lines
render-street-interview/SKILL.md112 lines · 7.1 KB

render-street-interview

The renderer for the street-interview video ad format (goose-studio recipe one-shot-videos/create-street-interview-video). A handheld interviewer stops people on one street corner and asks one question. A few guess wrong, one gets it right, and the cut ends on the brand's own line. The people are generated; everything after the takes is free and local.

Read the bundled model notes before generation. If a required guide cannot be fetched or opened, stop before spending and name it. REFERENCE.md holds the format's historical Critical knowledge entries and the rejected takes behind them. Read it before changing the prompt scaffold or a gate; its older experiments do not override the current recipe or this entry. Use the project take-ledger guidance before reusing a seed. Keep each brand's observed successes and limitations in its own project; a seed is not a quality guarantee.

Run

Run everything from the project the video belongs to. Brand-asset paths in the configs (logo, product photo, end-card sting) resolve against that folder, or $STREET_INTERVIEW_ROOT. The run folder is --run <dir> (default projects/street-interview/), with working/ for intermediates and output/ for deliverables.

python scripts/selftest.py                                   # free: the format and the lint hold
python scripts/single_gen.py --brand <slug>                   # dry run: price + the full prompt
python scripts/single_gen.py --brand <slug> --seed <n> --yes  # PAID: one take (~$3.64 at 12 s, 720p)
python scripts/build_episode.py --episode <name>              # free: grade, re-cut, captions, end card
python scripts/check-cut.py --episode <render>.episode.json   # free: the ship gate
  • A brand is data: brands/<slug>.json holds the product and its reference photo, the street, the question, the cast and their lines, props, captions, logo and end card. Copy brands/demo-tallgrass-oat.json. No brand appears in format_spec.py.
  • An episode (episodes/<name>.json) joins three takes into a ~25-30 s cut. It names the takes, any whole shots to drop (drop_shots, each pair a real shot's start and end), the brand layer and, optionally, brand_layer.end_card_music, a short sting played under the end card.
  • Paid calls go through the GooseWorks proxy (scripts/media_proxy.py), never a local key. On a poll timeout, resume with media_proxy.resume_fal(request_id). Never resubmit, since a dropped poll has already been billed.

Prompt length

The final Seedance prompt is built from the project brief and shared shot instructions. BytePlus recommends at most 1,000 English words because lengthy prompts may miss details. This is quality guidance. The current Fal schema declares no maximum prompt length; that does not prove unlimited acceptance.

single_gen.py prints a non-blocking advisory above that guideline. The old 1,200-word refusal is removed: its source-run observation did not prove a precise boundary. Keep the exact approved dialogue and required clauses; do not trim them, reduce the cast or add a paid retry just to meet a count. Missing clauses, invalid inputs and existing spend approval still block generation. The finished-cut gate reports length as advice, without failing on it. Review actual video adherence through the normal gate and full watch/listen pass.

Guarantees

  • The prompt carries every format clause; single_gen.py lints it before any spend, and check-cut.py imports the same clause list, so a clause cannot be dropped silently.
  • Captions are derived, never authored: timing comes from Whisper on the finished render, spelling from the script. A scripted word Whisper skips or mishears inside a sentence is restored. Lines whose shots were dropped leave the caption script.
  • The gate measures the finished file: shot lengths, splices on real cuts, caption timing and safe zone, ambience floor per shot, loudness and true peak, every scripted line audible, and detail and black point. It prints what it can NOT assess (faces, comprehension, how it sounds) on every run. Watch the cut end to end, and do not publish a FAIL.

Cost

A take is about $3.64 (12 s at 720p, $0.3034/s); a 30 s episode is three takes, about $10.92. Grading, re-cut, captions, looks and every gate are free. Staging prices move, so price the first call of a run and quote from that.

Local finishing

The local finishing scripts use the current Python interpreter and carry --run into child commands. The default grade (--strength 0) needs no colour-reference file. A positive strength requires the real reference.

Set approved colours in brand_layer.palette: accent, text, and background, each #RRGGBB or three RGB integers. Optional brand_layer.fonts keys are black, bold, and regular (paths relative to the project). Without overrides, fonts resolve on macOS, Windows or Linux. End-card rows shrink together to fit the safe area; shorten copy if it cannot fit.

The subway series bar stays visible through caption gaps. It is also in the caption-free control so the gate measures captions separately from persistent branding.

For a new prompt, generation.prompt_version: 2 (or --prompt-version 2) repairs duplicate articles and uses a top-edge rule for non-can packaging. The manifest records the version for the gate. Historical prompts default to version 1 and retain their hashes. Use a new approved seed for a new prompt; do not overwrite an approved take.

Free regression checks:

python -m unittest discover -s tests -p 'test_*.py' -v

Known limits

  • Seedance refuses some photoreal faces (its likeness gate). Faces here come from the prompt, not a reference photo; see REFERENCE.md.
  • A multi-take episode can't make three generations be the same corner. A scene-reference still from take A (--scene-ref) couples later takes to its mic, street and light. Check by eye.
  • Ambience can differ between takes. The gate's check G fails a shot whose street bed sits within a few dB of the speech; that needs a re-take, not a mix.

Provenance

Ported from goose-studio skills/molecules/create-street-interview-video on 2026-10-02, including the episode-2 v3 fixes (hook caption, the mispronounced line, the silent end card). The port changed only path resolution and routed the paid calls through the proxy. selftest.py passes from an empty folder, and episode 2 v3 rebuilds identically (11 shots, 22.64 s).

1---
2name: render-street-interview
3description: Build a vox-pop street interview video ad. An interviewer with a handheld mic asks passers-by one question about the brand's product, they give blunt wrong guesses, one gives the real answer, and the cut lands on a branded end card. Generates the takes through the GooseWorks fal proxy (Seedance 2.0 with native voice), then grades, re-cuts, captions and gates them locally. Use for the street-interview format.
4status: draft
5---
6 
7# render-street-interview
8 
9The renderer for the **street-interview** video ad format (goose-studio recipe
10`one-shot-videos/create-street-interview-video`). A handheld interviewer stops people on one
11street corner and asks one question. A few guess wrong, one gets it right, and the cut ends on the
12brand's own line. The people are generated; everything after the takes is free and local.
13 
14Read the bundled [model notes](references/model-behaviors.md) before generation.
15If a required guide cannot be fetched or opened, stop before spending and name it.
16`REFERENCE.md` holds the format's historical **Critical knowledge** entries and
17the rejected takes behind them. Read it before changing the prompt scaffold or a gate;
18its older experiments do not override the current recipe or this entry.
19Use the [project take-ledger guidance](TAKES.md) before reusing a seed. Keep each
20brand's observed successes and limitations in its own project; a seed is not a quality guarantee.
21 
22## Run
23 
24Run everything from the project the video belongs to. Brand-asset paths in the configs
25(logo, product photo, end-card sting) resolve against that folder, or `$STREET_INTERVIEW_ROOT`.
26The run folder is `--run <dir>` (default `projects/street-interview/`), with `working/` for
27intermediates and `output/` for deliverables.
28 
29```bash
30python scripts/selftest.py # free: the format and the lint hold
31python scripts/single_gen.py --brand <slug> # dry run: price + the full prompt
32python scripts/single_gen.py --brand <slug> --seed <n> --yes # PAID: one take (~$3.64 at 12 s, 720p)
33python scripts/build_episode.py --episode <name> # free: grade, re-cut, captions, end card
34python scripts/check-cut.py --episode <render>.episode.json # free: the ship gate
35```
36 
37- **A brand is data:** `brands/<slug>.json` holds the product and its reference photo, the
38 street, the question, the cast and their lines, props, captions, logo and end card. Copy
39 `brands/demo-tallgrass-oat.json`. No brand appears in `format_spec.py`.
40- **An episode** (`episodes/<name>.json`) joins three takes into a ~25-30 s cut. It names the takes,
41 any whole shots to drop (`drop_shots`, each pair a real shot's start and end), the brand layer
42 and, optionally, `brand_layer.end_card_music`, a short sting played under the end card.
43- Paid calls go **through the GooseWorks proxy** (`scripts/media_proxy.py`), never a local key.
44 On a poll timeout, resume with `media_proxy.resume_fal(request_id)`. Never resubmit, since a
45 dropped poll has already been billed.
46 
47## Prompt length
48 
49The final Seedance prompt is built from the project brief and shared shot instructions.
50[BytePlus recommends at most 1,000 English words](https://docs.byteplus.com/en/docs/modelark/create-video-generation-task-api)
51because lengthy prompts may miss details. This is quality guidance. The current
52[Fal schema](https://fal.ai/api/openapi/queue/openapi.json?endpoint_id=bytedance%2Fseedance-2.0%2Freference-to-video)
53declares no maximum prompt length; that does not prove unlimited acceptance.
54 
55`single_gen.py` prints a non-blocking advisory above that guideline. The old 1,200-word
56refusal is removed: its source-run observation did not prove a precise boundary. Keep the
57exact approved dialogue and required clauses; do not trim them, reduce the cast or add a paid
58retry just to meet a count. Missing clauses, invalid inputs and existing spend approval still
59block generation. The finished-cut gate reports length as advice, without failing on it.
60Review actual video adherence through the normal gate and full watch/listen pass.
61 
62## Guarantees
63 
64- The prompt carries every format clause; `single_gen.py` lints it before any spend, and
65 `check-cut.py` imports the same clause list, so a clause cannot be dropped silently.
66- **Captions are derived, never authored:** timing comes from Whisper on the finished render,
67 spelling from the script. A scripted word Whisper skips or mishears inside a sentence is
68 restored. Lines whose shots were dropped leave the caption script.
69- The gate measures the finished file: shot lengths, splices on real cuts, caption timing and
70 safe zone, ambience floor per shot, loudness and true peak, every scripted line audible, and
71 detail and black point. It prints what it can NOT assess (faces, comprehension, how it sounds)
72 on every run. **Watch the cut end to end, and do not publish a FAIL.**
73 
74## Cost
75 
76A take is about **$3.64** (12 s at 720p, $0.3034/s); a 30 s episode is three takes, about
77**$10.92**. Grading, re-cut, captions, looks and every gate are free. Staging prices move, so
78price the first call of a run and quote from that.
79 
80## Local finishing
81 
82The local finishing scripts use the current Python interpreter and carry `--run` into child commands. The default grade (`--strength 0`) needs no colour-reference file. A positive strength requires the real reference.
83 
84Set approved colours in `brand_layer.palette`: `accent`, `text`, and `background`, each `#RRGGBB` or three RGB integers. Optional `brand_layer.fonts` keys are `black`, `bold`, and `regular` (paths relative to the project). Without overrides, fonts resolve on macOS, Windows or Linux. End-card rows shrink together to fit the safe area; shorten copy if it cannot fit.
85 
86The `subway` series bar stays visible through caption gaps. It is also in the caption-free control so the gate measures captions separately from persistent branding.
87 
88For a **new** prompt, `generation.prompt_version: 2` (or `--prompt-version 2`) repairs duplicate articles and uses a top-edge rule for non-can packaging. The manifest records the version for the gate. Historical prompts default to version 1 and retain their hashes. Use a new approved seed for a new prompt; do not overwrite an approved take.
89 
90Free regression checks:
91 
92```bash
93python -m unittest discover -s tests -p 'test_*.py' -v
94```
95 
96## Known limits
97 
98- **Seedance refuses some photoreal faces** (its likeness gate). Faces here come from the prompt,
99 not a reference photo; see `REFERENCE.md`.
100- A multi-take episode can't make three generations be the same corner. A scene-reference still
101 from take A (`--scene-ref`) couples later takes to its mic, street and light. Check by eye.
102- Ambience can differ between takes. The gate's check G fails a shot whose street bed sits
103 within a few dB of the speech; that needs a re-take, not a mix.
104 
105## Provenance
106 
107Ported from goose-studio `skills/molecules/create-street-interview-video` on 2026-10-02,
108including the episode-2 v3 fixes (hook caption, the mispronounced line, the silent end card).
109The port changed only path resolution and routed the paid calls through the proxy.
110`selftest.py` passes from an empty folder, and episode 2 v3 rebuilds identically
111(11 shots, 22.64 s).
112 

Discussion

Alternatives