A track layout that survives a revision
Why layout is a technical decision
A mix with everything on two tracks works until the first change. Then the client wants the music swapped, and you discover that the music level was set clip by clip, forty times, and every one of those settings belonged to the old track.
A mix organised into groups survives, because the decisions live in one place per group rather than once per clip.
The standard layout
From the top of the timeline down, or the bottom up — the order is a convention, not a rule, but be consistent:
- A1–A2: Dialogue. One track per speaker where you have two people, so each keeps its own processing. A third for narration.
- A3: Room tone. A continuous bed under everything, from the previous block.
- A4–A5: Sound effects. Spot effects — a door, a click, a phone.
- A6–A7: Ambience. Beds: traffic, birds, a room's hum, a crowd.
- A8–A9: Music.
Nine tracks sounds like a lot for a five-minute film. It costs nothing, and it means "turn the music down" is one fader rather than a search.
Busses
A bus is a channel that several tracks feed into, so processing applied to the bus applies to all of them at once.
The three you want, and they map onto the standard delivery categories:
- Dialogue bus — everything spoken. Compression and any overall EQ goes here.
- Music bus.
- Effects bus — spot effects and ambience together.
Then all three feed the main output, where the limiter and the loudness measurement live.
In Resolve's Fairlight page these are called busses and you create them in the Mixer's bus format dialogue. Premiere calls them submixes. Reaper, Audacity's multi-track mode and Ardour — all free or cheap — do the same thing under the same names.
The payoff: sidechain ducking becomes one connection between two busses rather than forty keyframed clips. A loudness problem is fixed with three faders. And the stems below fall out for free.
Stems, and why anybody cares
A stem is a bounce of one bus on its own: the whole film's dialogue as a single file, the whole film's music, the whole film's effects.
They exist because of three real situations:
- Translation. To dub a film into another language, somebody needs everything except the dialogue. That combination — music and effects, no speech — is called an M&E and it is a standard deliverable. You cannot produce one from a flattened mix.
- Revision. A change to the music six months later needs the dialogue and effects untouched.
- Broadcast delivery. Many broadcasters require stems as a matter of course, and will reject a delivery without them.
Export them at the same length as the programme, from the same start timecode, so they line up when dropped back in. Name them <project>_DIA.wav, _MUS, _FX, _ME.
Clip gain against track fader
Two different controls that both make things louder, and using the wrong one is the most common mixing mistake.
Clip gain adjusts the level of one clip before it reaches any processing on the track. Use it to make all your dialogue clips roughly equal — which is a job you do first, before any other mixing.
The track fader adjusts everything on the track after processing. Use it for the overall balance between elements.
The reason the order matters: a compressor on the dialogue track responds to whatever level arrives at it. If clip 1 arrives at -20 and clip 2 at -8, the compressor treats them completely differently and no fader setting fixes both. Level the clips first with clip gain, then compress, then balance with the fader.
The order of work in a mix
- Level the dialogue clips with clip gain so the waveforms look roughly even. Aim for peaks near -12 dBFS at this stage, leaving room for processing.
- EQ each dialogue source — next lesson.
- Compress the dialogue bus — the lesson after that.
- Lay room tone and ambience so there are no holes.
- Place effects.
- Bring music in and duck it.
- Set the overall loudness and put a limiter on the main.
Doing these out of order means redoing them. Compressing before levelling is the classic one.
Naming and colour
Name every track. Colour dialogue one colour, music another, effects a third. On a timeline with thirty audio clips this is the difference between finding a problem and hunting for it.
And lock tracks you have finished with. A locked track cannot be dragged by accident, which is how sync gets broken at 11pm.
Today
Take a project you have mixed on two tracks and rebuild its layout into dialogue, effects and music busses. Then change the music level with one fader. That is the whole argument.
The one thing to keep
Organise audio into dialogue, effects and music busses so a change is one fader rather than forty clips, and level clips with clip gain before any compressor sees them, because a compressor responds to whatever level arrives at it.
Before you move on
An editor compresses the dialogue bus and then levels the individual clips with clip gain afterwards. Why does the result sound inconsistent no matter how carefully the clip gains are set?
Pick the one you would defend. Nobody sees your answer.