On this page
Why audio matters more than video quality
Viewers tolerate mediocre video with great audio. They abandon great video with bad audio.
Studies of YouTube and TikTok engagement consistently show that audio problems — background noise, inconsistent levels, distortion, sync issues — drive higher abandonment rates than visual problems of similar severity.
Investing in audio editing has higher returns than investing in better cameras for most creators.
The audio editing pipeline
A complete audio edit covers five stages:
- 1Sync — align external audio to video
- 2Cleanup — remove noise, breaths, and unwanted sounds
- 3Level — balance dialogue volume across the video
- 4Mix — blend dialogue, music, and effects
- 5Master — final loudness normalization for the platform
Skipping any stage produces noticeably worse output.
Stage 1: Sync external audio
If you recorded audio separately — a lavalier, a shotgun mic, a field recorder — the first job is lining it up with the camera's audio. The reliable trick is the clap: one clap at the start of recording shows up as a sharp spike on both waveforms, and you align those spikes. If you forgot to clap, some editors auto-sync by detecting matching audio signatures, or you can manually match any sharp transient that both tracks captured — a door closing, a footstep, a hard plosive. Once they're aligned, mute the camera audio and use the external recording, which is why you recorded it in the first place.
Stage 2: Cleanup
With the audio synced, clean it up — carefully. Reduce background noise like AC hum, fridge buzz, or traffic, but use a light hand, because heavy noise reduction introduces artifacts worse than the noise itself. Cut distracting breaths and lip smacks between phrases, but don't remove every breath, since some breathing sounds natural and removing it makes speech feel robotic. Clear out filler words — um, uh, like — which automated tools make fast. And tame plosives, the popping P and B sounds, with a touch of volume automation on the affected syllable or a high-pass filter around 100Hz.
Stage 3: Level dialogue
Volume needs to stay consistent across the whole video, because sudden loud or quiet stretches read as amateur immediately. As targets to aim for, dialogue peaks should sit around -6dB to -3dB, the dialogue average around -16dB to -12dB, and any music under dialogue down around -24dB to -18dB. Use compression to even out the dynamics and gain automation to ride the broad changes — a speaker who whispers and then shouts needs both. A reasonable starting compressor for dialogue is roughly a -18dB threshold, a 3:1 ratio, a 5ms attack, and a 50ms release, then adjusted by ear from there.
Stage 4: Mix dialogue, music, and effects
The mixing principle is simple: dialogue is primary and everything else supports it. Music should duck — drop in volume — whenever dialogue plays, which most editors automate. Sound effects should match the visual perspective, so a door closing on screen is loud while the same sound off-screen is quieter. Use stereo placement deliberately, keeping dialogue centered, letting music spread wider, and positioning effects by their on-screen location. And mind the frequency balance: dialogue lives roughly in the 200Hz–3kHz range, so the music should leave that band relatively clear while someone is speaking.
Stage 5: Master for the platform
Finally, normalize loudness for where the video is going, because every platform has its own target. YouTube aims for about -14 LUFS integrated; TikTok and Instagram sit around -16 to -14; broadcast wants -23 LUFS (EBU R128) or -24 (US ATSC); Spotify Canvas is -14. This matters in both directions — too quiet and viewers crank the volume only to get blasted by the next ad, too loud and the platform automatically attenuates you in a way that can hurt perceived quality. Mastering to the right target is what makes your audio sit comfortably alongside everything else in the feed.
The mistakes that show up most
A handful of errors recur. Inconsistent volume across scenes is the most common by far. Music mixed too loud forces viewers to keep adjusting the volume. Excessive noise reduction produces underwater, robotic artifacts. Mismatched sample rates between recordings cause sync to drift over time. Skipping mastering entirely means shipping a raw mix with no loudness normalization. And forgetting captions ignores the large share of viewers watching on mute — caption everything as a default.
That muted majority deserves its own thought, because well over half of social video is watched with the sound off. It has two implications: captions are mandatory (auto-generate them at minimum), and your visual storytelling has to carry meaning on its own, showing rather than only telling. None of this means audio doesn't matter — the minority who do listen tend to engage and convert more strongly — but the video has to work both ways. For tooling, v8eo handles sync, leveling, basic cleanup, and built-in caption generation; Audacity (free) and Reaper (paid) cover dedicated audio work; and for dialogue-heavy content, Descript pioneered text-based editing while v8eo offers a similar filler-word removal flow.
Quick checklist
Before exporting any video:
- [ ] Audio synced to video
- [ ] Background noise reduced (lightly)
- [ ] Filler words and breaths cleaned
- [ ] Dialogue levels consistent
- [ ] Music ducked under dialogue
- [ ] Final loudness around -14 LUFS
- ] [Captions added
Related: Remove filler words automatically | How to edit talking head videos
Tagged