Addaly is in open beta. Things will change, and AI answers can be wrong — check anything that matters.

Editing Video

Cuts, sound, colour and export — taught on DaVinci Resolve, which costs nothing.

Lesson 17 of 7710 min

Cutting a conversation, step by step

The most common thing you will be asked to cut

Two people talking. An interview. A podcast with cameras. A panel. This is the bulk of real editing work, and there is a working method that is more reliable than instinct.

Step one: the radio edit

Ignore the picture entirely. Build the conversation as audio.

Cut out the false starts, the repeated question, the interruption, the long pause where somebody thinks, the "sorry, can I say that again". Cut for meaning: the shortest version of each answer that still says the thing.

The target is something you could publish as a podcast. Listen to it with the screen off. If it does not hold as sound, no picture will fix it.

This takes longer than you expect and it is where the piece is actually made.

Step two: find the breaths

Now go through the radio edit and move every cut to a breath. Zoom in on the waveform until you can see individual words; the gaps between phrases are visible as troughs.

A cut inside a phrase produces a small click of wrongness even when the words are grammatical. A cut at a breath is inaudible. Put the breath at the head of the incoming clip, not the tail of the outgoing one — a person finishing a sentence and then inhaling sounds natural; a person being cut off mid-inhale sounds edited.

Step three: fill the holes with room tone

Every cut you made joined two different moments of the same room, and the background noise level is slightly different at each. Lay a continuous bed of room tone under the whole scene at a level that matches. Now the joins have nothing to announce themselves with.

This single step is the largest difference between an edit that sounds amateur and one that does not, and it costs one track and five minutes.

Step four: decide what the viewer looks at

Only now does picture enter. For each moment, ask: what is the most interesting thing to be looking at right now?

The default answer is wrong. Beginners cut to whoever is speaking. Experienced editors cut to whoever is reacting, at least half the time. A conversation is about what is happening between two people, and half of that is on the face of the person not talking.

The practical rule: cut to the listener before a significant line, so we see them receive it. Cut to the speaker on a line where their face is doing something the voice is not.

Step five: J and L every join

Having placed the picture cuts, offset them from the audio cuts. Unlink audio and video and drag the picture edges a beat earlier or later than the sound edges.

As a starting point: bring picture in before the new speaker starts — a J cut, usually 6 to 20 frames — so we see the listener begin to respond before we hear them. And let the previous speaker's audio run over the new picture on the way out — an L cut.

Do this at most joins. A dialogue scene where every picture cut lands exactly on an audio cut sounds like a slideshow, and the fix is mechanical.

Step six: the single-camera problem

With one camera and one angle, every cut is a jump cut. Four repairs, in order of preference:

  1. A cutaway. B-roll, the interviewer listening, the object being discussed, hands, the room. Two seconds of anything else and the cut is invisible.
  2. A punch-in. Scale the incoming clip to about 110–120% and reframe slightly. 4K source at a 1080p timeline makes this free. Alternate between the wide framing and the punch-in rather than punching in twice in a row.
  3. A speed ramp or a dissolve across the join, which reads as a deliberate compression rather than a mistake.
  4. Leave the jump cut. In talking-head work this is an accepted convention and viewers read it as honesty about compression rather than as a break in continuity. It stops working if the subject moves a lot between the two frames.

Step seven: remove the tics, carefully

Um, er, lip smack, the sharp inhale before a sentence. Each is a small ripple trim.

The warning: over-cleaned speech sounds inhuman. Every pause removed, every breath gone, and the result is a rate of information delivery that nobody can follow and a person who apparently does not breathe. Remove the ones that obstruct; leave the ones that are thinking.

And there is a line beyond tidying. Reordering somebody's sentences so they appear to answer a question they were not asked, or splicing clauses to produce a statement they never made, is a fabrication regardless of how clean the join is. The next block deals with where that line sits.

Step eight: watch it once as a stranger

Full screen, from the top, without touching anything. Mark where your attention went. Fix those places and stop.

Today

Take a two-minute interview clip and do steps one to three only: radio edit, breaths, room tone. Listen with the screen off. That is the scene.

The one thing to keep

Cut a conversation as audio first, move every join to a breath, lay room tone underneath, and only then decide what the viewer looks at — which is the listener at least as often as the speaker.

Before you move on

An editor cuts an interview, places every picture cut exactly on its corresponding audio cut, and the result feels mechanical despite each individual cut being well chosen. What is the fix and why does it work?

Pick the one you would defend. Nobody sees your answer.

No ads. No data sale. No public scores on people. Ever.

© 2026 Addaly