promo-recut
Esta página aún no está traducida; aquí tienes la versión en inglés.
Use when someone has a talking-head recording about something they made or found and wants a premium 16:9 short (and optionally 9:16 / 3:4 versions laid out per platform) for 小红书 / YouTube / B站. The talk is tight-cut, and the talking head slides into a split-screen next to 3D screenshot cards that scroll to whatever is being discussed. The edit adds highlighter sweeps, chips, keyword subtitles, punch-ins, a freeze-frame with a prompt card zooming out of a screenshot, and a zoom-through into a framed highlights montage. It finishes with an outro stamp, an end card, a chapter progress bar, a cover and the post copy.
Inputs → Outputs. input/talk.mp4 (talking head) + screenshots (PNG/JPG) + optional input/highlights.mp4
- links →
work/(graded raw,body.mp4,outro.mp4,montage.mp4,layout.json),promo/(HyperFrames project,index.html, subset fonts),promo-vertical/(optional),my-promo.mp4(-14 LUFS, BT.709, faststart),cover-4x3.jpg/cover-16x9.jpg/cover-3x4.jpg,post.md.
Everything content-specific lives in ONE project config: promo.config.yaml (or .json). It holds the KEEP
spans, the creator’s cleanup reply (cut.reply), subtitles, cards and their highlight rows, chips, hold point, montage clips and
labels, chapters, stamp and end card, the cover and the post. Start from
$VSTUDIO/workflows/promo-recut/examples/promo.config.example.yaml, which documents every key. Taste
(rates, loudness, brand colours, title length, tags) comes from the persona (persona.local.yaml).
Pipeline (run from the project dir; $VSTUDIO = repo root)
Section titled “Pipeline (run from the project dir; $VSTUDIO = repo root)”my-promo/ promo.config.yaml input/talk.mp4 input/highlights.mp4 input/*.png- Transcribe + find what to cut (去气口 / filler / 重复 / 口误: the shared
vstudio.cleanup)SetTerminal window cp $VSTUDIO/workflows/promo-recut/examples/promo.config.example.yaml promo.config.yaml # then editpython3 $VSTUDIO/workflows/promo-recut/scripts/tight_cut.py promo.config.yaml --suggestcut.body/cut.outroKEEP spans by sentence first. The first run writeswork/audio.json, usingvstudio.asr.transcribe(word timestamps;mlx_whisperon Apple Silicon, elsefaster_whisper, else OpenAI whisper-1 ifOPENAI_API_KEYis set).--suggestsnaps the KEEP spans word-safe and runs the same cleanup every speech workflow uses (python -m vstudio.cleanup analyze,references/CLEANUP.md) →work/cleanup.json+ the review sheetwork/cleanup_review.md: 待确认 CONFIRM (semantic fillers 然后/就是/那个 the audio isolates, interjections, restarts, re-takes, fillers whisper glued onto the next word — the old hidden-onset PATCH), 自动删 AUTO (嗯/呃/um/uh, stutters, clear repeats), 气口 (pauses squeezed percut.profile, never deleted), 保留 KEEP (looks like a real word). Each row shows…before【removed】after…and why. The creator listens to the CONFIRM rows (ffplay -ss <t-0.5> -t 2 work/audio.wav) and answers, e.g.确认 3,5,9 / 保留 7; put it in the config verbatim ascut.reply. Without a reply only AUTO rows are cut (never confirm every row unheard: on real footage that deletes real words). Old configs withcut.drop/cut.patchstill work: they are translated to approvals of the matching cleanup edits (else word-safe editor cuts) and printed asconfig:lines. - Tight cut + montage + layout
This grades the raw once (
Terminal window python3 $VSTUDIO/workflows/promo-recut/scripts/tight_cut.py promo.config.yamlpython3 $VSTUDIO/workflows/promo-recut/scripts/tight_cut.py promo.config.yaml --verifygrade, optionalhdr_tonemap) and runscleanup.applyper part on it (AUTO rows +cut.reply): word edits cut from the silence after the previous kept word to the silence before the next, video frame-exact, audio sample-exact with a 20 ms equal-power crossfade at every join (A/V cannot drift), then a two-pass loudnorm to personaaudio.voice_lufs. Each part gets a sidecarwork/<part>.cleanup.json. The step builds the highlights montage with baked 0.3 s internal crossfades (cut.xfade_assemble, plain acrossfade) and writeswork/layout.json: durations, raw→cut maps asvstudio.cut.TimeMapitems, and word times.--verify=cleanup.verifyper part: re-ASR, content words that went missing are printed with their raw / cut time (exit 1: add保留 Nfor the edit covering it tocut.replyand re-cut); leftover fillers / repeats are listed. Changing the KEEP spans renumbers the edits, so with a reply in the config the cut refuses until--suggestis re-run and the reply re-confirmed. Also listen to every join before going on. - Subtitles + cards
Write subtitles in raw seconds and wrap the key term in 【】.
Terminal window python3 $VSTUDIO/workflows/promo-recut/scripts/tight_cut.py promo.config.yaml --draft-subs # paste, then editpython3 $VSTUDIO/workflows/promo-recut/scripts/find_rows.py input/shot-post.png --preview work/rows.pngfind_rows.pymeasures text-row bands with numpy. Use those y0/y1 rows forcards[].highlights,boxandscroll. Itswidth_fraccolumn is a good starting highlight width. - Build the HyperFrames project
This writes
Terminal window python3 $VSTUDIO/workflows/promo-recut/scripts/build_promo.py promo.config.yamlpython3 $VSTUDIO/workflows/promo-recut/scripts/build_promo.py promo.config.yaml --orientation vertical # optionalcd promo && npx hyperframes lint && npx hyperframes snapshot --at <split-in>,<hold>,<zoom-through>,<outro> --no-endindex.htmlandtimeline.json, copies media intoassets/, and extracts the freeze frame. It also subsets Noto Sans SC + STIX Two Text fromvstudio.config.FONT_DIR(vstudio.render.subset_project_fonts;--no-fontsreusesassets/fonts/) to just the characters in the config (≈60 KB per CJK weight). Look at snapshots taken mid-transition, not only mid-scene. Expect 0 lint errors. Thenested_structure_needs_subcompositionwarnings are advisory. - Render + deliver
This renders with HyperFrames, runs a two-pass loudnorm to persona
Terminal window bash $VSTUDIO/workflows/promo-recut/scripts/export.sh promo my-promo.mp4 delivery # = python3 .../export.pyaudio.loudness_lufs(-14;audio.loudnorm_2pass,--lufsoverrides), writes BT.709 colour tags into the H.264/HEVC stream without re-encoding and adds+faststart(media.retag_bt709). It then prints the streams and the measured loudness. Use--skip-render <raw.mp4> <out.mp4>to redo only the delivery step. - Cover + post
The cover is
Terminal window python3 $VSTUDIO/workflows/promo-recut/scripts/make_cover.py promo.config.yamlpython3 $VSTUDIO/workflows/promo-recut/scripts/post_copy.py promo.config.yamlvstudio.cover.split_coverfed from the config. It takes a frame from the talk (or a given image) and retouches it throughvstudio.retouch(slim, eye, de-shine, skin, light makeup, optional body slim). It crops around the detected face. Landscape sizes get a split cover: photo on one side, and on the other a dark panel with quote, title + highlighted term, a framed highlights thumbnail, chips, a red tag and a 记笔记 tag. Portrait sizes stack the photo on top. The post (vstudio.publish.post_body) gets the title (length checked per platform), body, links, a chapter timeline frompromo/timeline.json(MM:SS, floored) and tags.
Timeline model (what build_promo computes)
Section titled “Timeline model (what build_promo computes)”body plays at rates.body and is split at hold.at into body + freeze image + body2
(data-media-start). These share track 2 back to back. The montage starts zoom_through s (0.5) before
the body ends and runs on its own track 3 at rates.montage. The outro starts where the montage’s nominal
length ends, while the montage clip keeps running another zoom_through s underneath. The end card follows.
Raw second → final second is BT(raw) = TimeMap.to_final(raw, snap="fwd")/rate (+ hold if raw ≥ hold.at), so you never type a
final-timeline time. Chapter anchors are raw body seconds or start / montage / outro.
Hard-won lessons
Section titled “Hard-won lessons”- Overlapping clips need separate tracks and a known audio owner. Body and montage overlap during the
zoom-through, so they sit on different
data-track-indexvalues. Only overlap where the outgoing clip is silent (the cut adds about 0.1 s of tail after the last word). If an overlap would play two voices, setdata-volume="0"on one, or use a muted copy, rather than relying on a fade. - The outgoing scene must stay alive during a transition.
body2runs untilM + TZ, and the montage runsTZpast the outro start. If a clip ends exactly when its exit animation starts, the transition shows the plate (a black flash). - Overlay start states must be
opacity: 0in CSS (gsap_fullscreen_overlay_starts_visible). This covers the dim layer, prompt box, screen, outro wrap, badges, stamp and end-card lines. Otherwise seek-based rendering shows them on frame 0 or before theirfromToruns. - Chapter labels need a scrim. Small labels over bright footage are unreadable. The bar sits on a
bottom gradient (
#bar-scrim), and labels get a text shadow. - Whisper merges fillers into neighbouring words. A “word” that is too long for its characters usually
starts with a hidden 然后/嗯. The cleanup finds the energy dip and lists it as a 粘连口头禅 (
filler-merged) CONFIRM row that cuts only up to the rise, so the word itself stays; confirm it by ear. - ASR the cut to verify it (
--verify). Joins that look right on the word list can still swallow a syllable or keep half a filler. - Captions must clear before a zoom-through. build_promo clips any cue that crosses
M + TZso it ends atM - 0.08. A caption flying into the montage frame looks broken. - Measure screenshot rows, don’t guess. Highlight bands and red boxes use image-pixel rows from
find_rows.py(row-band detection on the grey-level difference from the background). Card scale iscard_width / image_width, and build_promo applies it. - Adjacent card windows closer than 0.8 s are merged into one split, so the face doesn’t bounce back to full frame between cards.
- Montage clips with label
nullare transitional fragments. The previous step label continues over them. - Fonts: only Noto Sans SC / STIX Two Text (OFL), subset per video. A system CJK font that exists only on your machine silently falls back in the headless renderer.
Geometry / vertical
Section titled “Geometry / vertical”build_promo.py has a GEO table for horizontal (card box, split inset, subtitle line, bar, screen frame,
label positions); override any value with layout.horizontal.*. The vertical geometry is not a table: it is
computed from a platform profile (vertical_geo, next section), so every element lands inside that platform’s
safe box and clear of its button column. Override any computed key with layout.vertical.*.
Platforms
Section titled “Platforms”python3 $VSTUDIO/workflows/promo-recut/scripts/build_promo.py promo.config.yaml # horizontal (unchanged legacy layout)python3 $VSTUDIO/workflows/promo-recut/scripts/build_promo.py promo.config.yaml --orientation vertical # persona platforms.default at 9:16 -> promo-vertical/python3 $VSTUDIO/workflows/promo-recut/scripts/build_promo.py promo.config.yaml --platform xiaohongshu:full # -> promo-vertical-xiaohongshu-full/python3 $VSTUDIO/workflows/promo-recut/scripts/build_promo.py promo.config.yaml --platform douyin # -> promo-vertical-douyin-vertical/--platform (or config platform:) takes any vstudio.platform profile: xiaohongshu:full (9:16),
xiaohongshu:vertical (3:4, 1080x1440), douyin, tiktok, youtube-shorts, bilibili:vertical; a horizontal
profile keeps the 1920x1080 layout. --out DIR names the project dir. Per vertical profile the layout is:
| element | where (from platform.safe_box / caption_box / keepouts) |
|---|---|
| chapter bar + labels | top of the safe box, on a top scrim (labels 22 px) |
| talking head (split) | a band under the bar, ~42 % of the free height (~2.4 face heights when a face is detected), full safe width; the clip window is centred on the speaker’s face (vstudio.face on 6 frames, else centred; layout.vertical.face: [fx, fy]), and object-position keeps the face centred when full-frame |
| chips + screenshot card | under the band down to 24 px above the caption band; side margin widened to clear the lower-right button column |
| captions | the profile’s caption box, bottom-anchored, text-wrap: balance, size = box width / max_chars_zh clamped to the profile’s caption size range; cues that can’t fit 2 lines are warned about |
| prompt hold card | the card column, mid-height |
| montage screen + step label / badge / tag | 16:9 screen at the safe width, centred in the free area but kept above the button column |
| stamp / end card | stamp top-left of the free area; end card centred between the safe top and the caption band |
timeline.json records platform, canvas and the boxes used (safe, caption, keep-outs, face band, card,
screen) so snapshots can be checked against them. The build also warns when the length is outside the
profile’s sweet spot / max and when chapter labels collide on the narrower bar (merge or shorten chapters).
Check every vertical build: npx hyperframes lint (0 errors) and snapshots at a card, the hold, mid
zoom-through, the montage and the outro.
Multi-platform delivery. Build one project per canvas, render each, then hand the renders to
vstudio.export, which (for the same aspect) only scales, re-loudnorms to the profile’s LUFS / true peak,
encodes with the profile’s settings, crops covers and writes post stubs + manifest.json:
python3 $VSTUDIO/workflows/promo-recut/scripts/export.py promo my-promo.mp4 --platform youtube # or plain persona LUFSpython3 -m vstudio.export my-promo.mp4 --platforms youtube,bilibili,xiaohongshu:horizontal --out exports/ \ --cover cover-16x9.jpg --cover cover-4x3.jpg --post post.jsonpython3 $VSTUDIO/workflows/promo-recut/scripts/export.py promo-vertical-douyin-vertical my-promo-douyin.mp4 --platform douyinpython3 -m vstudio.export my-promo-douyin.mp4 --platforms douyin,tiktok,youtube-shorts --out exports/ --cover cover-3x4.jpgCaptions here are part of the design (keyword highlight, timed with the cards), so they are burned in by
HyperFrames per canvas. If you want vstudio.export to place them instead, build with --clean-master (no
burned subtitles) and pass --cues <promo_dir>/cues.json (final-timeline cues, written by every build).
Do not reframe the 16:9 promo to 9:16 with vstudio.export: the split screen and cards would be cropped. Build
the vertical layout instead.
Effects are shared (vstudio.hf)
Section titled “Effects are shared (vstudio.hf)”Every packaging effect here (split screen, screenshot cards with scroll / highlighter / red box, chips,
keyword subtitles, punch-ins, freeze hold, zoom-through into a framed screen, step labels, badge, tag, title,
stamp, end card) is a generator in lib/vstudio/hf.py returning {css, html, js}. build_promo only lays out
times and geometry and passes hf.JS("D.X") references into its const D data object. To reuse one effect in
another HyperFrames project, call the generator with plain numbers. Catalogue: references/EFFECTS.md.