Ir al contenido

polish: make any exported edit publish-ready

Esta página aún no está traducida; aquí tienes la versión en inglés.

Use when: an edit has been exported from any NLE (Descript, CapCut/剪映, Premiere, Resolve, Final Cut, a HyperFrames render…) and needs the last-mile treatment before upload: the first frame is dead (black title card, fade-in, idle face) and should become the cover, the audio is too quiet (exports often land around -20 LUFS), the pacing wants a small pitch-preserved speed bump, and the file should carry correct colour tags and faststart. Inputs: export.mp4 (+ optional cover.png from workflows/cover). Outputs: final.mp4 at the source resolution and frame rate, source-matched bitrate, persona.audio.loudness_lufs integrated loudness, bt709 tags. With --platform <name>: the same, tuned to that platform’s profile + final.cover.jpg at the platform cover size; with --platform a,b,...: the polished master plus one file + cover per platform and exports/manifest.json (see Platforms).

Run from the project (video) folder. $VSTUDIO = repo root. Don’t polish a file you haven’t probed, and don’t polish twice (a second loudness pass only costs headroom). Long-form uploads take the loudness and tag steps but usually not the speed step. End-to-end context: references/SOP_SHORT_VIDEO.md.

  1. Probe first, never blind-apply.
    Terminal window
    ffprobe -v error -show_entries stream=codec_name,width,height,bit_rate,r_frame_rate -show_entries format=duration -of default=nw=1 export.mp4
    ffmpeg -y -ss 1 -i export.mp4 -frames:v 1 /tmp/check.png # ground-truth frame size (ffprobe metadata has lied before)
  2. Decide the speed WITH the user (short-form parity). Default speed is 1.0, but for shorts (vertical 9:16 / 3:4 and under ~3 min, or any --platform of douyin / tiktok / youtube-shorts / xiaohongshu) propose 1.2× - or persona speed.body if the creator has set one - and ask before applying, e.g. “This is a 2:10 vertical talking-head; speed it up to 1.2× (pitch kept, ~1:48)? 1.0 / 1.1 / 1.2?”. The original short-form recipe ran everything at 1.2×; we keep that as the suggestion, not a silent default. Long horizontal videos stay at 1.0 unless asked. Anything above persona speed.cjk_max_intelligible (1.4) prints a warning: dense Chinese speech gets hard to follow - say so if the user asks for more.
  3. One command does the rest:
    Terminal window
    python3 $VSTUDIO/workflows/polish/scripts/polish.py export.mp4 -o final.mp4 \
    --cover cover.png --speed 1.2 --keep work/polish --check # add --platform douyin etc.
    • --cover replaces the first --cover-sec (1.0) seconds of picture; audio runs uninterrupted, total length unchanged. Cover is fitted to the source frame (--cover-fit fill|fit).
    • --speed takes a number (1.2) or a persona key (hook, body, fast_body…). Omit for 1.0 (see step 2).
    • Loudness target defaults to persona.audio.loudness_lufs (-14), or the --platform profile’s; override with --lufs / --tp. TP -1.5 dBTP, LRA 11.
    • --keep DIR keeps 01_cover.mp4, 02_speed.mov, 03_loud.mp4 for stage-by-stage debugging.
    • --check writes final.first.png so you can eyeball that frame 0 is the cover.
  4. Verify (the script prints these; re-run any time):
    Terminal window
    ffmpeg -i final.mp4 -af loudnorm=I=-14:print_format=summary -f null - 2>&1 | grep -E "Input Integrated|Input True Peak"

Optional: studio sound (--studio-sound [light|standard|strong], default off)

Section titled “Optional: studio sound (--studio-sound [light|standard|strong], default off)”

For phone / laptop / room audio (Descript Studio Sound, 剪映 人声增强 / 降噪): denoise, dereverb, voice EQ, de-ess and a gentle compressor before the cover / speed / loudness steps (vstudio.studiosound, local, picture stream-copied). standard for most rooms, light for an already clean phone take, strong for fans / AC / echo. Listen first with python -m vstudio.studiosound export.mp4 --ab work/ab --at 30 (level-matched before / after clips).

Optional: background music (--music calm|warm|bright|tech|story|<file>, default off)

Section titled “Optional: background music (--music calm|warm|bright|tech|story|<file>, default off)”

A bed ducked under the voice (audio.mix_bed, --music-db -30 LUFS before ducking). Moods are built-in beds generated by vstudio.music (free for any use); a file path uses her own track.

Optional: 气口 / filler / repeat cleanup of the export (--cleanup, default off)

Section titled “Optional: 气口 / filler / repeat cleanup of the export (--cleanup, default off)”

An NLE export that still has long pauses, 嗯/呃 or repeats can take the shared cleanup tool (vstudio.cleanup, references/CLEANUP.md) as step 0, before the cover:

Terminal window
python3 $VSTUDIO/workflows/polish/scripts/polish.py export.mp4 -o final.mp4 --cleanup gentle # 1st pass: auto edits
# read final_review.md (待确认 items) to the creator, then:
python3 $VSTUDIO/workflows/polish/scripts/polish.py export.mp4 -o final.mp4 --cleanup gentle \
--cleanup-reply "确认 3,5 / 保留 7" [--cleanup-verify]
  • Profiles: pauses (气口 only, no word edits), gentle, standard, tight. Only auto edits are cut until the creator replies; --cleanup-verify re-transcribes the cut and stops if a content word was lost.
  • The EDL (<out stem>.cleanup.json) and <out stem>_review.md sit next to the output and are reused on re-runs (ids stay valid; --cleanup-fresh re-analyses). Transcript: --cleanup-transcript (any shape) else ASR (--cleanup-lang).
  • Never cut twice: a file that is itself a cleanup output (has its .cleanup.json sidecar) is not cut again, and cleanup.apply’s duration guard refuses an EDL made from a different-length file. Exports from Descript / CapCut / 剪映 where the creator already removed pauses usually want pauses or nothing.
  • Duration afterwards = cleaned length ÷ speed (the self-check’s expectation follows it).

Profiles live in lib/vstudio/platform.py (numbers and sources: references/PLATFORMS.md; persona overrides under platforms.<name>). --platform takes name[:orientation] - xiaohongshu:vertical (1080x1440), xiaohongshu:full, douyin, tiktok, youtube, youtube-shorts, bilibili[:horizontal|vertical] - or a comma list. No --platform = the platform-neutral polish above, byte-for-byte the same commands as before.

no --platform one target (--platform youtube) several (--platform douyin,xiaohongshu:vertical)
loudness persona audio.loudness_lufs, -1.5 dBTP profile loudness (LUFS/TP); --lufs/--tp still win master at the default; each export re-normalised to its profile
video encode match source bitrate (or --crf) same, capped at profile encode.maxrate (bitrate mode: -b:v min(src, cap), maxrate <= cap; CRF mode: --crf else profile crf, + profile maxrate/bufsize) master as before; exports use profile CRF + maxrate (vstudio.export)
canvas source source - not reframed; a warning if the aspect differs reframed per profile (--reframe-mode face, pad-blur fallback)
length - platform.check_length warnings (max / min / sweet spot) per export, in manifest.json
cover first-second replacement only + <out stem>.cover.jpg at platform.cover_size (from --cover, else a frame), .feed.jpg preview where the feed crops + <platform>-<orientation>.cover.jpg per target
fps source source; warning above the profile max converted to the profile default above max
Terminal window
# one platform: polish + platform-sized cover + length check
python3 $VSTUDIO/workflows/polish/scripts/polish.py export.mp4 -o final.mp4 --cover cover.png --platform youtube-shorts
# several platforms: polish the master once, then one file + cover per platform + exports/manifest.json
python3 $VSTUDIO/workflows/polish/scripts/polish.py export.mp4 -o master.mp4 --cover cover.png \
--platform douyin,xiaohongshu:vertical,youtube --out-dir exports [--cues cues.json]
# same export step on an already-polished master
PYTHONPATH=$VSTUDIO/lib python3 -m vstudio.export master.mp4 --platforms douyin,xiaohongshu:vertical \
--out exports --cover cover.png [--cover cover-16x9.png] [--cues cues.json] [--title "..."]
  • Captions: an export from an NLE usually has captions burned in; those get cropped/covered on other canvases. For several platforms, export a caption-free master from the editor and pass --cues (cues.json / SRT) so each platform gets captions sized for its own caption box.
  • Single target with a different aspect (e.g. a vertical export with --platform youtube): polish does NOT reframe silently. Use the multi-target form or python -m vstudio.export to get a reframed file.
  • Pass several --cover images to vstudio.export (e.g. 3:4 and 16:9); each platform takes the closest aspect.
  • Check manifest.json warnings (loudness, length, reframe fallback, captions that do not fit) before upload.
  • One stage per ffmpeg call. Cover + speed + loudnorm in one filter graph is miserable to debug; intermediates are cheap.
  • Keep source resolution and frame rate. Never downscale “to save bandwidth”; platforms re-encode anyway.
  • Match source bitrate when re-encoding (--bitrate match, the default: -b:v src -maxrate 1.15x -bufsize 2x). Plain -crf 18 on a ~28 Mbps phone/Descript export lands at 4–6 Mbps: fine for upload, not for an archive master. --crf N is available when you do want size over fidelity.
  • Two-pass loudnorm (measure → linear=true apply with the measured values). Single-pass loudnorm runs in dynamic mode and audibly pumps speech. The loudness step (vstudio.audio.loudnorm_2pass) outputs 48 kHz stereo AAC (mono exports are up-mixed before measuring) and raises the LRA target to the measured LRA so loudnorm stays linear. Normalize once, at the end of the chain; don’t normalize again afterwards (platforms do their own pass and a second one only costs headroom). --skip-if-close leaves gain alone within 1 LU.
  • Speed uses atempo, never asetrate (pitch stays put). Factors outside 0.5–2.0 are chained automatically. Above ~1.3× speech starts to sound artificial; persona speed.cjk_max_intelligible (1.4) triggers a warning.
  • Speed before loudnorm. The speed step keeps audio as 24-bit PCM so the loudness pass is the only lossy audio encode.
  • bt709 tags: encodes are tagged at encode time; stream-copied H.264/HEVC are re-tagged with the h264_metadata/hevc_metadata bitstream filter (no re-encode). Untagged SDR files can shift colour on some players. HDR sources (iPhone HLG/Dolby Vision) need tone-mapping before this step, not just tags.
  • faststart moves the moov atom to the front so the first frame shows before the full download.
  • The cover-replacement assumes the first second of the export carries no important picture (a typical dead title card). If the first second has a hook shot, prepend a cover instead (concat a 0.5–1 s still + delay audio) or pick a frame from the hook as the cover.
  • Output resolution and fps equal the source
  • Video bitrate close to the source (or CRF chosen on purpose)
  • Integrated loudness within ±1 LU of target, true peak ≤ -1.0 dBTP
  • Frame 0 is the cover, not black
  • Audio at t=0 is the source’s audio at t=0 (the cover replaces picture only; nothing shifted)
  • Duration ≈ source (cleaned length with --cleanup) ÷ speed (±0.5 s)
  • Shorts: 1.2× (or persona speed.body) was proposed and the user’s answer applied
  • With --platform: no length warnings you didn’t mention; cover jpg at the platform size; manifest warnings read