Sound: beats, SFX vocabulary, grammar, levels
How to make cuts land on the music and how to place sound effects so a fast edit feels
deliberate instead of busy. Code: lib/vstudio/beats.py (analysis, grids, cut plans, energy arcs)
and the SFX section of lib/vstudio/audio.py (synthesised bank, place_sfx, cue sheets).
Every sound in the SFX bank is synthesised from sine and noise maths, so there are no audio files to license.
1. Order of work
Section titled “1. Order of work”- Music first, if there is music. Analyse it before planning shots:
b = beats.analyze("bgm.mp3"). Save the analysis (b.save("work/beats.json")): it is the audit trail for every cut. - Picture next. Plan sections with
beats.energy_arc(duration, kind, beats=b), then cut times withb.cut_plan(...). Write shot boundaries as beat numbers (b.beat_t(n)), not frame numbers. - Sound effects last, once the picture is locked. Any change to shot lengths or order means
the cue sheet has to be rebuilt. Rebuilding is cheap when cues come from
cue_sheet_for(...)with the same event list. - Verify after rendering:
beats.verify(cut_times, b, tol_frames=3, fps=30).
Keep times as float seconds all the way. Round to frames only at the last step, because rounding early lets errors pile up.
2. Beat analysis (beats.analyze)
Section titled “2. Beat analysis (beats.analyze)”| Field | Meaning |
|---|---|
bpm, period, offset |
Least-squares straight-line fit t = offset + n * period through the tracked beats. The tracker’s own tempo number can be off by a few percent. The fitted line is not. |
residual_ms, residual_p90_ms, grid_ok |
Fit quality. The grid is accepted when p90 <= 15 ms and max <= 30 ms, after dropping up to 10 % stray beats. Otherwise beats keeps the raw tracked beats, which happens with tempo drift, live drummers or DJ mixes. |
tempo_check |
1/2x, 1x and 2x candidate grids, each scored by how much kick energy lands on it. A grid that sits on the off-beats is shifted by half a beat. |
onsets |
Per-band transients: kick < 150 Hz, snare 150-2500 Hz, hat > 5 kHz, as (t, strength) pairs. |
downbeats, downbeat_phase |
Bar = 4 beats. The phase is the beat position with the strongest kicks. |
rms, sections |
Loudness curve (50 ms hop). Section boundaries come from spectral novelty and are snapped to bars, each with energy 0-1 and a level. |
hits |
The top-N strongest transients, at least 2 beats apart, weighted by loudness. |
Backends: librosa’s tracker when it is importable, otherwise a pure-numpy tracker: spectral-flux onset envelope, autocorrelation tempo with a 120 BPM prior, then a dynamic-programming beat path. Both backends share everything after tracking: attack refinement on a 2 ms curve, grid fit, tempo check, bands and sections.
Grid or transient?
- For dense, regular cuts (every beat, every 2 beats), use the grid:
b.beat_t(n)andb.cut_plan(...). - For sparse accents, such as a freeze-frame or the one big slam, use the real transient from
b.hitsorb.onsets. A tiny grid drift gets exposed on an isolated accent. - The last frame of the video should land on the last real hit, with the RMS curve confirming the music has stopped. Don’t end on a reverb tail.
- Strong accents are almost always on whole beats. Put a slam on a half beat only if the onset data shows a hit there.
Cut patterns (b.cut_plan(n, start, end, pattern)). every-beat, every-2 and every-bar
are steady. accelerate uses shrinking gaps of 2, 1.5, 1, 0.75, 0.5, 0.375 and 0.25 beats, which
is 32/24/16/12/8/6/4 frames at 112 BPM and 30 fps, and resolves on a bar line. drop holds two
bar cuts, then cuts on every beat from the drop.
Verify. verify reports two errors per cut:
- Audio error: the designed seconds against the beat. This checks the analysis.
- Frame error: the frame-rounded time against the beat. This checks the frame rate.
Up to 3 frames of frame error passes, and 1.5 frames or less is ideal. 30 fps cannot be more
precise than about 17 ms. If every cut is off by the same amount, look for an output audio offset
first (AAC encoder priming is about 1024-2112 samples). Measure it once per render pipeline by
cross-correlating a sharp SFX against the rendered audio. Keep that offset as its own constant
and never fold it into the analysed offset.
3. Vocabulary (synthesised, 48 kHz, audio.sfx_bank())
Section titled “3. Vocabulary (synthesised, 48 kHz, audio.sfx_bank())”| Sound | Use it for | Peak in sample | Peak level at base gain |
|---|---|---|---|
soft (alias transition) |
scene / place change, one per change | 0.47 s | -14 dBFS |
whoosh, whoosh_b |
camera move, push, pan | 0.15-0.18 s | -19 dBFS |
swoosh, swoosh_b |
short: title or card entrance, small push, zoom | 0.13 s | -11 dBFS |
whip (alias whip_pan) |
whip pan; peaks at the cut, stops dead | 0.24 s | -10 dBFS |
impact, impact_b (alias boom) |
landing, big slam, logo stamp; the loudest moment | start | -7 dBFS |
stamp (alias thud) |
small landing: sticker, badge, stamp | start | -5 dBFS |
riser |
build into a finale or reveal, 2.4 s | end (2.2 s) | -12 dBFS |
sparkle |
tail after an impact, glow, “ta-da” | 0.03 s | -12 dBFS |
shutter, shutter_b (alias camera) |
photo moment, freeze-frame snapshot | start | -9 dBFS |
pop, pop_b |
list items, stickers appearing | start | -16 dBFS |
tick, tick_b |
counters, map pins, small steps | start | -15 dBFS |
ding |
check mark, “done” | start | -19 dBFS |
typewriter, typewriter_b |
one key / one character | start | -12 dBFS |
typing |
typing reveal: trim with dur to the text animation |
start | -13 dBFS |
scratch (alias record_scratch) |
comedic “wait, what?” freeze | 0.28 s | -11 dBFS |
stop (alias tapestop) |
everything stops, slow-down gag | start | -10 dBFS |
The *_b sounds are alternation partners. audio.write_sfx("assets/sfx") writes every sound
(aliases skipped) as stereo wav for HyperFrames <audio> clips.
Choose by genre, not by event. A cinematic promo uses whoosh, impact, riser, sparkle and
transition, and skips cartoon sounds. In cue_sheet_for(kind="promo"), pop becomes tick and
ding becomes sparkle. A fun travel vlog can use pop, scratch, shutter and ding. Test each
sound in the finished cut, not on its own. Ask yourself: with eyes closed, does it sound like
this kind of video, or like a mobile game?
Real actions get matching sounds. Typing on screen gets keys, a photo gets a shutter, a sticker
landing gets a stamp. A generic whoosh won’t cover a distinctive action. Trim long sounds to the
length of the action (dur).
4. Grammar (audio.cue_sheet_for(cuts, reveals, kind))
Section titled “4. Grammar (audio.cue_sheet_for(cuts, reveals, kind))”A cue sheet is a plain list: [{t, sfx, gain_db, dur, note}]. t is where the sound’s peak
lands, and note says which on-screen action it belongs to. No cue goes in without a reason.
| Rule | What the code does |
|---|---|
| One soft transition per scene change | A scene cut gets soft. Transition-type cues closer than 0.3 s collapse into the most important one, so a cut-derived cue beats a reveal-derived one. |
| Beat cuts are carried by the music | beat, montage, jump and hard cuts get no SFX. The drums already mark them. |
| Whoosh on camera moves | move/push get whoosh, zoom gets swoosh, whip gets whip. |
| Impact on landings | landing gets impact, snapped to the nearest beat within 0.12 s when beats= is given. |
| Finale = riser -> impact -> sparkle | finale at t: the riser peaks at t (-2 dB), the impact hits at t, and the sparkle follows 0.6 s later (-3 dB). This is the one sentence never to break. |
| Repeated hits | The same sound repeated within 1.5 s alternates with its _b variant and steps down 1.5 dB per repeat (max -9 dB). The run should sound countable, not like a machine gun. If the hits get too dense to count, replace them with one swoosh. |
| Density cap | At most N cues in any 2 s window: travel-fun and promo 3, talking-head and story 2. A finale triple counts once. The most important cues stay: impact/riser > sparkle/stop/whip > transitions > pops/ticks. |
| Kind profiles | CUE_PROFILES: the cap, an overall gain offset (talking-head -3 dB, story -2 dB) and vocabulary swaps. |
Then x = audio.render_cue_sheet(cues, total) gives an (n, 2) float32 stem, peak-aligned.
Write it with audio.write_wav and mix it with the voice and music.
5. Peak alignment
Section titled “5. Peak alignment”Many sounds don’t peak at their first sample: a riser peaks at its end, a whip just before its
cut-off, a soft transition halfway through. If you place them by their start, every hit sounds
late. place_sfx with dict events, and render_cue_sheet, shift each sound so that its 5 ms-RMS
peak (audio.sfx_peak) lands on t. The default is peak_align=True for dict events and
False for legacy (t, name) tuples. For external samples, measure the peak once with
sfx_peak and subtract it the same way.
6. Levels (vs voice and music)
Section titled “6. Levels (vs voice and music)”| Layer | Level |
|---|---|
| Voice | -16 LUFS stem (persona audio.voice_lufs) |
| Music under voice | -30 LUFS bed, ducked a further ~10 dB while the voice speaks (mix_bed) |
| Music-only montage | Bring the bed up to about -18 to -20 LUFS. The music is the voice here. |
| Regular SFX | Peaks around -18 to -10 dBFS: under the voice, level with the drums |
| Accent SFX | The finale impact is the loudest SFX (about -7 dBFS peak), used 2-3 times per video at most |
| Final | Two-pass loudnorm of the whole mix to -14 LUFS, true peak -1.5 dBTP (loudnorm_2pass) |
True peak and AAC: an encoded output (.mp4 / .m4a) is measured after the encode and corrected. A .wav
output is usually muxed to AAC later, which adds ~0.2-0.4 dB of peak, so loudnorm_2pass(..., "mix.wav")
limits WAV_HEADROOM_DB (0.5 dB) under the target (headroom=0 for an exact lossless deliverable). After a
custom mux, audio.ensure_loudness(final.mp4) re-measures and re-normalises only if it missed.
Use loudness for importance. The most important beat gets the loudest SFX, and repeated small sounds sit lowest. If the music already hits hard, let the drums do most of the work and give SFX only to actions that exist only on screen. Gain is a multiplier, not a target: a quiet external sample needs normalising (or a different sample) before it can sit at these levels. Check the level in the rendered file.
7. Delivery
Section titled “7. Delivery”Deliver two versions from the same timeline: with music and without music. The SFX and the voice stay in both. Platforms mute or swap licensed tracks, and a no-music version can be re-scored without a re-edit.