Editing Video
Cuts, sound, colour and export — taught on DaVinci Resolve, which costs nothing.
- Level
- Nothing assumed
- Lessons
- 77
- Reading time
- 679 min
- Price
- Free, no sign-up to read
A working course in the craft of editing. Why your first cut runs two seconds long, how proxies let an ordinary laptop edit 4K, the grammar of cutting on action and J and L cuts, why clean audio matters more than a sharp picture, correction before grading using scopes rather than eyes, captions for a muted feed, codecs and bitrate explained properly, and a folder structure that still opens six months later.
Start the first lessonModule 1
Getting the footage into a state you can work with
An edit is decided before the first cut. This block is about the hour between the card coming out of the camera and the first clip landing on a timeline: what the camera actually wrote to the file, which project settings cannot be changed later, how to copy footage so you know the copy is intact, and how to watch an hour of material in a way that leaves you able to find any moment in it. It ends with the two things that break a project on day one — footage that arrived damaged, and phone video whose frame rate is a lie.
By the end you can
Set up a project whose frame rate, resolution, colour handling and media layout are chosen from the footage rather than from a default, offload and verify a card so you can prove the copy is byte-identical, and diagnose a clip that drifts out of sync by naming the property of the source file that causes it
- 1Your laptop is probably not too slowCamera files store frames as differences, not pictures, so edit a proxy and let the export go back to the original.
- 2What the camera actually wrote to the cardA camera file carries four fixed decisions — whole frames or differences, how much colour was kept, how many brightness levels, and how many bits per second — and each one sets a hard limit on what the edit can do later.
- 3Frame rate, shutter, and the flicker in the lightPick one frame rate from the mains frequency where you shoot, set every camera and the timeline to it, and keep shutter at roughly 180 degrees — because mismatched rates cost you frames at every conversion and mismatched shutter costs you flicker you cannot remove.
- 4Resolution, aspect ratio, and having room to moveShoot larger than you deliver so the edit has room to punch in, stabilise and reframe, set one image-scaling behaviour at the start of the project, and keep everything that must be read inside 90 per cent of the frame.
- 5The settings you cannot change laterFrame rate, image scaling and colour management are decided before the first import and are expensive to reverse, so configure a template project once and duplicate it rather than choosing settings under time pressure.
- 6Copying a card so you know the copy is goodA drag-and-drop copy reports success without reading anything back, so verify irreplaceable footage with a checksum manifest at offload and keep the material in two places on two devices before any card is formatted.
- 7Watching an hour so you can find any second of itWatch everything once before cutting anything, marking moments with written notes and building a selects timeline, because editing is a search problem and you cannot search material you never formed a memory of.
- 8Two cameras, one microphone, one sync pointGive every take one sharp sync point at the start, use waveform sync as a first pass and check it at a hard consonant, and convert variable-frame-rate phone footage to constant rate at offload because drift is usually the file, not the clocks.
- 9Editing the words before you touch a pictureFor speech-led work, transcribe everything and build the structure in text with timecodes attached, because reading is four times faster than listening and rearranging a paragraph costs a keystroke where rearranging a timeline costs a minute.
- 10When the footage arrives brokenMost broken footage falls into four fixable classes — variable frame rate, an unsupported codec, a missing index, and a mis-mapped audio channel — and the ones that are genuinely unrecoverable are clipped highlights, clipped audio and missed focus, which are worth naming to a client on day one.
Case studyTwo cards, three days, and the format button
Ritu shoots and her brother Ankit handles sound and the laptop. They own one mirrorless camera, one lavalier into a phone, and two 64 GB cards. The brief is a twelve-minute film plus six short clips for the centre's social feeds. Three days of sessions, four teachers interviewed between them, and one closing address by the centre's founder that happens once.
On the morning of day one the arithmetic is already tight. Sixty-four gigabytes at the camera's 4K bitrate is a little over an hour and a half of recording. A full day of sessions, interviews and cutaways is four to five hours. Both cards will be full by the afternoon.
The laptop is a five-year-old machine with a 512 GB internal drive and one USB port that works reliably. They have a 2 TB portable hard disk. A straight drag-and-drop copy of one full card to the hard disk takes about twelve minutes over that port. Ankit knows, because he read about it the week before, that a drag-and-drop copy is not verified: the operating system writes the bytes and reports success without ever reading them back. Copying with rsync --checksum reads everything twice and takes closer to forty minutes a card.
So the decision arrives at lunchtime on day one, and it is not a technical question. It is a scheduling one.
Option one: verify, and lose the afternoon's first session. Forty minutes a card, twice a day, on the only laptop, which is also the machine Ankit uses to keep an eye on the audio recordings. The sessions do not pause for them. Whatever happens in the hall between two and half past two is not filmed.
Option two: copy fast and format. Twelve minutes a card, no verification, and they are back in the hall before the session restarts. If a file is silently bad they will find out in the edit, three weeks later, when the card has been formatted eight more times and the session cannot be repeated.
Option three: stop formatting. Buy four more 64 GB cards. In Kota in March that is roughly eleven hundred rupees each, so forty-four hundred rupees against a thirty-two thousand rupee fee. The cards stay unformatted until the film is delivered, which means every card is a second copy of itself, and the copy on the hard disk is a third. Nothing needs verifying, because nothing is being destroyed.
Ritu's objection to option three is straightforward and correct: the fee already has no margin in it, the camera is theirs, the drive is theirs, and forty-four hundred rupees is most of what they will clear on the job.
Ankit's objection to option two is that the founder's closing address happens at four o'clock on day three, it lasts nine minutes, and it is the only thing in the brief that cannot be re-created. If that file is the one that comes back as green blocks, there is no film.
They are also aware, as they argue about it on the stairs, that both of them have done option two on every job they have ever taken, and nothing has ever gone wrong.
The question is what the material is worth, not what the process costs. The sessions can be filmed again next month, because the centre runs the programme every term. The founder's address cannot: he is a guest, he is not local, and he will not be in the building again.
That distinction is the one that resolves it, and it resolves it unevenly. The correct answer is not the same for every card on the shoot.
What actually happened
They split the material by whether it could be re-shot. Sessions and cutaways were copied fast and the cards formatted. The four interviews and the founder's address went onto two cards that were never formatted, copied to the hard disk and verified with a SHA-256 manifest that evening, in the hotel, while nothing else was happening. Total extra cost: nothing, because the verification ran during dinner rather than during a session. On day three the card reader failed part-way through an interview copy; the manifest caught two truncated files, the card was still intact, and the copy was simply run again. The team now keeps a one-page log in the project folder recording, per card, what is on it, where the two copies are, and whether it has been formatted.
Worth arguing about
Why did the same shoot end up with two different offload procedures rather than one?
One answer
Because the value of verification depends on whether the material can be re-created, not on how it was recorded. The session footage could be filmed again the following term at no cost, so a fast unverified copy risked a small, recoverable loss. The interviews and the address happened once, so the same risk was unbounded. Applying the slow procedure to everything would have cost sessions they needed; applying the fast procedure to everything would have staked the whole film on a copy nobody had checked.
A drag-and-drop copy said it succeeded and two files were still wrong. What does that tell you about what the operating system was reporting?
One answer
It was reporting that the write operation completed, not that the data on the destination matches the source. Nothing read the copy back. A checksum manifest computes a fingerprint of the source and of the copy and compares them, so a failing reader, cable or card shows up as a mismatch at the time rather than as green blocks three weeks later. That is the whole of what `rsync --checksum` and `shasum -c` add.
The team had a 2 TB hard disk and a 512 GB laptop. If they had put both copies on the hard disk, in two different folders, would the rule have been satisfied?
One answer
No. The rule is two copies on two devices, because the failure being guarded against is a device failure. Two folders on one disk both disappear when that disk dies, and they also both disappear when somebody deletes the wrong folder. The laptop's internal drive, the hard disk, and an unformatted card are three separate devices; two folders are one.
Test yourself6 questions on this module
The same 4K phone clip plays at three frames a second from the original file and smoothly from a quarter-size proxy. What is the mechanism?
A camera recorded 8-bit 4:2:0 in a log profile. The shot is corrected and contrast is added back. What appears?
Filming at 30 fps indoors on a 50 Hz supply produces bands crawling up the frame. The same camera outdoors shows nothing. Why?
A card was copied by dragging it to a drive. Every file appeared and the folder sizes matched. Three weeks later one clip is green blocks. What had the operating system actually reported?
A phone clip is in sync at the clap and half a second late by minute ten. What is the usual cause, and what fixes it?
In DaVinci Resolve, why is the timeline frame rate described as a setting you cannot change later?
Module 2
The cut itself
One join between two shots, examined properly. This block covers the four trims that move a cut point and what each one costs, the six things a cut has to satisfy and the order to sacrifice them in, how shot length becomes rhythm, how to build a sequence out of an action, the working method for a dialogue scene, and the repairs available when the coverage you needed was never filmed. It ends where the viewer actually decides whether to stay.
By the end you can
Place and trim a cut you can defend frame by frame — naming the moment it lands on, the trim operation you used and what that operation did to everything downstream — and diagnose a cut that feels wrong by identifying which of the six continuities the eye lost
- 11Your first cut is two seconds lateTrim every shot to the frame where it stops giving new information, and cut to the reaction more often than you think.
- 12A bad cut is felt before it is seenCut on movement, keep the camera on one side of the line, and let sound cross the join before the picture does.
- 13Ripple, roll, slip and slideFour named operations move a cut — ripple changes duration and shifts everything downstream, roll moves the join only, slip changes the content inside a clip, slide changes its position — so decide which question you are asking before you touch the edge.
- 14Six things a cut has to satisfy, in orderA cut can satisfy six things — emotion, story, rhythm, eye-trace, screen plane and physical space — and when they conflict, the audience forgives a broken geography far more readily than a wrong emotional beat.
- 15Pace is a pattern, and you can count itAverage shot length gives you a number for pace, but rhythm lives in the variation around it — put long shots where the viewer should settle and short ones where they should be pushed, and cut speech at the breath rather than inside the phrase.
- 16Building a sequence out of an actionA sequence is built out of fragments the viewer stitches into a continuous action, so shoot every action whole from three distances and join the shots with 'but' and 'therefore' rather than 'and then'.
- 17Cutting a conversation, step by stepCut a conversation as audio first, move every join to a breath, lay room tone underneath, and only then decide what the viewer looks at — which is the listener at least as often as the speaker.
- 18The shot that buys you everythingCutaways and inserts are not decoration — they are the licence to remove time, hide a mistake and change rhythm, and an interview needs roughly one usable one for every thirty seconds of finished film.
- 19Repairs for coverage you never shotOne camera yields several usable shots through punch-ins, reframes, ramps, freezes, stills and graphics — but each repair has a cost, and a missing reverse angle of a real person is not among the things any of them can supply.
- 20Where people decide to stayViewers answer 'what is this and is it for me' in the first few seconds whether or not you have told them, so state the subject immediately, open on a real image rather than black, and pay off whatever the opening promised.
Case studyThirty jump cuts, or half a day back in Ranchi
The material is forty-one minutes of interview with Sunita, who has been a health worker in the same three villages for nine years. It is good. She is funny about the first year, precise about what changed, and quiet about a child she could not get to a hospital in time. The camera was locked off in her front room, one angle, mid-shot, from the moment she sat down to the moment she stood up.
Beside it in the folder are two cutaways, both filmed in the last four minutes before the light went: her register of home visits, and a wide of the lane outside.
The radio edit comes down to five minutes forty of speech that holds together end to end with the screen off. That part goes well. The problem is what it looks like.
To get from forty-one minutes to five and a half, the editor has made about thirty removals. Each one is a jump cut: same framing, same chair, Sunita's head in a slightly different position on either side of the join. Two of the thirty are covered by the register shot, and one by the lane, because a two-second cutaway cannot be used more than about once every ninety seconds without drawing attention to itself.
The NGO's communications lead watches the rough cut and says it looks cheap. She is not wrong, and she is not able to say why. What she is reacting to is that Sunita moves quite a lot while she talks, so the jumps are large, and that after twenty of them the front room has stopped existing as a place.
The editor lays out three options and what each one costs.
Leave the jump cuts and punch in on alternate joins. The source is 4K and the delivery is 1080p, so scaling the incoming clip to 115 per cent and reframing from centred to a third is free in quality. Cost: about two hours. It converts roughly half the jumps into something that reads as a second camera. It does not fix the other half, and it does not give the film a single shot of anything other than Sunita's face.
Go back and shoot pickups. Ranchi to the village is two hours each way on a good day. Sunita does not need to be there for most of the list: the register, her hands, the doorway, the lane, the bicycle, the health sub-centre, the two women she visits, the medicines box. That is a shot list of about eleven items, all of them ten seconds each, all of them achievable in ninety minutes on site. Cost: one full day of the editor's time, about eighteen hundred rupees in travel, and a week of delay, in a month where the annual report has a printing deadline behind it.
Use stills and graphics. The NGO has photographs from the last four years. A held photograph with a slow push is a real cutaway and costs nothing. Cost: the photographs were taken by three different people, two of whom no longer work there, and nobody has a written note of who owns them.
The communications lead's instinct is the first option, because it costs nothing and does not move the deadline. The editor's instinct is the second, because he has counted: a six-minute film wants roughly twelve usable cutaways, he has two, and no amount of scaling manufactures the other ten.
What neither of them can do is invent a reverse angle. Nobody filmed the interviewer, and nobody filmed the women Sunita visits. Those shots are not recoverable by any technique in the edit.
What actually happened
They did the punch-in pass first, because it was cheap and because it made the rough cut watchable enough to be judged. Then the editor drove out for one day and shot the eleven-item list, which took seventy minutes on site. The finished film has fourteen cutaways in five minutes forty, four of which came from that trip and are the ones the NGO now uses as photographs as well. The photographs from the archive were not used: nobody could establish who took two of them. The annual report slipped by five days. The NGO's next brief specified that a cutaway list would be shot on the same day as every interview, and the editor now sends that list to whoever is holding the camera before the shoot rather than after it.
Worth arguing about
The communications lead said the rough cut looked cheap. Why is that a symptom report rather than a diagnosis?
One answer
She was describing an effect she genuinely felt and naming a cause she had no way of identifying. The actual mechanism was that thirty jump cuts on a subject who moves a lot gave her eye a new position to re-acquire at every join, and that after twenty of them the room had never been allowed to exist as a place. Acting on the word cheap would have led to better graphics or a grade. Acting on the mechanism led to coverage.
Why did the punch-in pass fix about half the joins and not the rest?
One answer
A punch-in changes the framing, so the incoming shot reads as a different shot rather than as a discontinuity in the same one. It works when it alternates: punching in twice in a row, or punching in on every join, simply establishes a second framing which then has its own jumps. It also cannot supply a change of perspective, because the lens compression is identical, and it cannot supply an image of anything that is not already in the frame.
The editor said a six-minute film wants about twelve cutaways. What is that number actually a budget for?
One answer
It is a budget for removals. Every cutaway is a licence to take time out of the interview without a visible join, plus cover for a cough or a fluffed sentence, plus a change of frame that stops one face becoming monotonous. Roughly one usable cutaway per thirty seconds of finished film is what lets an editor compress forty-one minutes to five without the compression itself becoming the thing the viewer watches.
Test yourself6 questions on this module
Someone reaches for a cup. The movement takes about a second and a half. Why is holding the shot to the end of the movement usually a second too long?
Cutting on action, editors start the incoming close shot two or three frames earlier than where the wide left off, so the movement slightly repeats. Why?
A dialogue scene where every picture cut lands on the same frame as its audio cut sounds like a slideshow. What does moving the audio edge past the picture edge change?
A cutaway is the right length and the right moment is inside it, but it should arrive half a second later. Music is already cut to picture. Which trim?
A cut feels wrong and you cannot say why. Murch's ranking suggests where the answer usually is. Where?
Why does a six-minute interview film need roughly twelve usable cutaways rather than two?
Module 3
Structure — deciding what the film is
Between a folder of good material and a film somebody watches to the end lies a decision about what the piece is actually about. This block is the part of editing that happens away from the timeline: naming the spine in one sentence, choosing a shape for the material, cutting an interview without misrepresenting the person in it, running a screening that produces useful notes, removing the shot you love, writing narration that does not describe what is already on screen, and using other people's footage without pretending the law is settled.
By the end you can
Take unstructured material to a defensible rough cut — a spine stated in one sentence, a chosen shape with a reason, an interview compressed without changing what the speaker meant, a screening run to produce specific notes, and a documented position on every piece of third-party material in it
- 21One sentence, or you do not have a filmWrite what the piece is about as one sentence containing a question or a change, and hold every uncertain clip against it, because the sentence is what makes the decision about a good shot that does not belong.
- 22Six shapes, and what each one costsChronology is a choice rather than a default, and each shape has a specific failure — thematic sections give the viewer a clean exit at every join, lists have no momentum, and essays become lectures unless every claim has evidence on screen.
- 23Compressing what somebody said without changing what they meantYou may change what somebody said and not what they meant — so removing a qualification, joining clauses from different answers or cutting the 'no' are fabrications regardless of how clean the join is.
- 24The stages, and how long each really takesAssembly, rough cut, fine cut, lock and finishing exist because work done in one stage is thrown away if the previous stage was not finished, and unscripted material realistically costs three to five hours of editing per finished minute before any finishing.
- 25Watching it with someone who has not seen itYou cannot watch your own edit because you know what is coming, so run a silent, uninterrupted screening with a stranger to the material and treat every note as a symptom report rather than a diagnosis.
- 26The twenty per cent that has to goA first fine cut is typically a fifth too long, and the material to remove is whatever is justified by your knowledge of the project rather than by its effect on a stranger — so park it rather than deleting it and watch the two versions.
- 27Narration that is not a description of the pictureCut the picture first and write narration into the gaps it leaves, supplying only what the picture cannot show — and read every line aloud while writing, because anything you stumble over the listener will stumble over too.
- 28One shoot, several finished piecesA short is not a trimmed long piece but a separate cut with its own spine, opening on the moment rather than the setup, and the cheapest way to get both is to compose and lock the long version first with the vertical crop already in mind.
- 29Using material you did not shootThird-party material is licensed, openly licensed, permitted or covered by a statutory exception — and because exceptions differ fundamentally between fair use and enumerated fair dealing, the widely repeated 'ten seconds is fine' rules have no basis anywhere.
Case studyThe clause that would have fitted the slot
The interview is with a junior engineer who has worked on the scheme for six years. Asked whether the remaining villages in the block will be connected, he says this:
"In most cases, yes. If the current funding continues into the next financial year, and if the two contractors who left are replaced, we can extend it to the remaining villages. That is the plan."
Twelve seconds. The film has a four-minute limit set by the slot it is going into, which is a screen in the reception areas of district offices, and the section it belongs to is already thirty seconds over.
The communications officer sends a note with four changes. Three are ordinary: a name spelling, a section that runs long, a request to open on the reservoir rather than the pipes. The fourth asks the editor to use the engineer's answer as "we can extend it to the remaining villages" — four seconds, clean, and it fits the section exactly.
The join is easy. The engineer pauses before "we", and there is a breath after "villages". Cut at the two breaths and the sentence is grammatical, the audio is continuous, and nobody watching could tell. The editor could do it in ninety seconds.
What the cut removes is not filler. It removes two conditions and a hedge. The full answer is a plan contingent on funding and staffing. The short answer is a commitment. A viewer in a district office who has not been connected yet would take one of those as information and the other as a promise, and the engineer did not make a promise.
The editor works this out by applying the test he was taught and has mostly not needed: if the speaker watched this cut, would he be able to point at a proposition in it that he does not hold? The answer is yes. The engineer holds "we can extend it if the funding continues and the contractors are replaced". He does not hold "we can extend it".
The cost of refusing is not abstract. This is the editor's third job for this department and the fourth is under discussion. The communications officer is not acting in bad faith; from where she sits, she has asked for a shorter version of something the engineer said, which is what editing is. She has a slot length and a department that would prefer the film to sound confident. If the editor writes back explaining fabrication and broadcast codes, he is telling a client she has asked him to lie, which is both true and unlikely to be well received.
The cost of agreeing is also not abstract, and it is worth being precise about what it is rather than treating it as a general feeling. The rushes exist. The engineer will see the film, because it plays in the building he works in. And the shape of the sentence — a conditional turned into an assurance about a public service — is exactly the shape that gets repeated back at somebody in a meeting.
The interesting part of the problem is that the editor has three ways to give the client the four seconds she needs, and only one of them involves the cut she asked for.
What actually happened
The editor sent back two alternative four-second versions rather than a refusal. The first was "If the funding continues, we can extend it to the remaining villages" — four and a half seconds, one condition kept, and it fitted the slot with a frame to spare. The second dropped the answer entirely and used eight seconds from earlier in the interview about what had already been connected, which was uncontested and stronger on screen. He noted in one line, without using the word fabrication, that the shorter version would read as a commitment the engineer had not made and that the reception-area audience would include people waiting to be connected. The communications officer took the first version within the hour. On the next job she asked, unprompted, whether a proposed cut "still says what he said". The editor now keeps a marker at every join in a radio edit where non-adjacent material was joined, and reads that list against the transcript before locking.
Worth arguing about
The join was inaudible and the sentence was grammatical. Why is that not the standard that matters?
One answer
Because the test is about the proposition, not the craft. Shortening, removing repetition and tightening a rambling answer are normal and expected by anyone who has been interviewed. Removing a qualification changes what was asserted: a conditional plan becomes an unconditional promise. A perfect join makes that harder to detect, which makes it worse rather than better.
Why was offering two alternatives a better move than explaining the principle?
One answer
Because the client had a real constraint — four seconds — and no intention to deceive. A refusal leaves her problem unsolved and frames her as dishonest, which invites an argument about motives. Two workable cuts solve the constraint and let the reasoning arrive as a single sentence about how the audience would read it. The editor kept both the line and the client, and the next brief came back with the question already asked.
What record would have protected the editor if the engineer had later said he never said it?
One answer
The rushes and the transcript, kept until well past publication, plus the marked list of every point where non-adjacent material was joined. The defence against "I never said that" is the recording, and the defence against an accusation of misrepresentation is being able to show the full answer beside the cut and demonstrate that the proposition survived.
Test yourself6 questions on this module
What does a one-sentence spine actually do for an edit?
A thematically structured piece fails in a particular way. What is it, and what is the repair?
An answer is cut from "In most cases, and if funding continues, we can do it" to "we can do it". The join is inaudible. Why is that a fabrication?
Why is grading a shot before picture lock described as doing the work twice?
A first viewer says "the middle is boring". Read as a symptom report rather than a diagnosis, what does that usually mean?
An editor in India wants ninety seconds of a news broadcast in a video essay, and is told "it is transformative, so it is fair use". Why does that not help?
Module 4
Sound, which is most of it
Audiences leave over sound long before they leave over pictures. This block builds a mix from the ground up: a track layout that survives a revision, the three or four EQ moves that make speech intelligible, what a compressor actually does to a voice and what it does to the noise between words, how loudness is measured now that every platform normalises, the ambience and spot effects that make a scene exist, matching a replaced line to the room it belongs in, and the honest list of audio faults that cannot be repaired.
By the end you can
Deliver a mix a stranger can follow on phone speakers — dialogue at a stated loudness with EQ and compression you can justify, continuous ambience under every join, music ducked by a mechanism you chose, and a measured integrated loudness and true peak rather than a guessed level
- 30Bad sound kills a good pictureGet the mic within 30 cm, keep peaks near -6 dBFS, record thirty seconds of room tone, and use less noise reduction than you want to.
- 31The music you found is not freeEvery track needs a licence you can produce on demand, and "royalty free" is a pricing model, not permission.
- 32A track layout that survives a revisionOrganise audio into dialogue, effects and music busses so a change is one fader rather than forty clips, and level clips with clip gain before any compressor sees them, because a compressor responds to whatever level arrives at it.
- 33Four EQ moves that make speech intelligibleFour moves cover most speech — a high-pass around 80 to 100 Hz, a notch at the mains frequency and its harmonics, a small cut in the 300 to 500 Hz boxy region and a gentle lift around 2 to 3 kHz — applied subtractively before any compressor sees the signal.
- 34What a compressor does, and what it does to the gapsSet dialogue compression by the gain reduction meter — 3 to 6 dB on loud words, returning to zero between phrases — and remember that make-up gain raises the noise floor too, so a hissy recording can take less compression than a clean one.
- 35Loudness, measured rather than guessedLoudness is an energy measurement over time, not a peak, so mix to roughly the destination's integrated LUFS target with true peak at -1 dBTP — mastering louder than the target just means the platform turns you down after you destroyed the dynamics.
- 36The sounds nobody notices and everybody missesA continuous ambience bed fills every hole your cuts made in the background at once, and spot effects only need to cover the events the viewer can see happening — layered two or three deep and placed a frame early rather than a frame late.
- 37Making a replaced line sound like it belongsA replaced line sounds wrong because its space is wrong, so match the noise floor first, then the EQ, then add far less reverb than sounds like reverb — and record an impulse response with one clap at every location, because de-reverb is much harder than matching.
- 38What can be repaired, and what cannotHum, clicks, plosives and level faults are fully repairable; noise, reverb and mild clipping are partly repairable up to the point where the artefact is worse than the fault; and speech under the noise floor or heavily clipped is gone, which is worth saying to a client on day one.
- 39Mixing when you do not have a studioCheck every mix on at least three systems and fix only the faults that appear on more than one, because each system lies in its own direction — and the low-volume test tells you more about balance than any meter.
Case studyA guest lecture, two ceiling fans, and a microphone at the back
The recording is unusable as it stands, and everybody agrees on that within the first thirty seconds of playback.
The hall is a rectangular room with a tiled floor, plaster walls and a high ceiling. Six ceiling fans were running. The camera was eleven metres from the lectern. What the camera microphone captured is therefore mostly the room: a low continuous roar from the fans, a long reverberant tail on every word, and the lecturer's voice arriving at roughly the same level as the air conditioning in the corridor.
The department's media assistant, Farhan, measures what he has. The noise floor sits around minus thirty-eight dBFS. The lecturer's speech peaks around minus twenty-six. That is a gap of about twelve decibels between the voice and the room, and a good part of the voice's energy is reflection rather than direct sound.
He tries the obvious things, in the order he was taught. A high-pass at 100 Hz removes some fan rumble and changes very little, because the fans also produce broadband noise well inside the speech range. Six decibels of spectral noise reduction helps. Twelve decibels produces the watery, fluttering artefact where quiet harmonics of the voice fall below the noise estimate and are removed in some frames and kept in others. A de-reverb pass, pushed hard enough to matter, leaves the ends of words clipped and gives the lecturer a thin, gated quality.
Compression makes it worse in a way that surprises him until he thinks about it: turning the loud words down and adding make-up gain raises the fans by exactly as much as it raises the voice.
That leaves three real options, and a deadline: the repository page is meant to go live before the new term.
Publish it with the light cleanup. Six decibels of noise reduction, a high-pass, gentle compression, captions burnt in and a sidecar file. Honest, quick, and hard work to listen to for ninety minutes. Farhan tests it on his phone with a fan running, which is how a student in a hostel will hear it, and loses whole clauses.
Resynthesise the voice. One of the newer speech tools will take the recording, run it through a model and return something close to studio quality. Farhan tries thirty seconds of it and it is genuinely startling. It is also generating audio: it can shift timbre and invent detail that was never recorded. This is a lecture by a named economist that will be cited, quoted and possibly assessed. What comes back is not a recording of what happened in the room; it is a plausible reconstruction of it.
Ask for a re-record. The lecturer has left Pune. The department could ask her to read the lecture again into a phone with a cheap lavalier at her own desk, and the picture could be held as slides and stills. That is thirty minutes of her time, which she may not give, and it produces a different artefact: a reading, not a lecture, with none of the questions from the hall in it.
The decision is not really between three audio processes. It is about what the file is for.
What actually happened
Farhan published the lightly cleaned version, with six decibels of noise reduction, a high-pass at 100 Hz, gentle two-stage compression and a proofread caption track, and put a one-line note on the page saying the room audio was poor and captions were provided. He did not resynthesise, on the grounds that the lecture is a record and the tool would have produced audio the economist never spoke. The department then bought two wired lavaliers and a small recorder for about nine thousand rupees, and wrote a half-page rule: for any recorded lecture, a microphone goes on the speaker, the camera microphone is a backup and a sync reference, and thirty seconds of the empty hall is recorded before the audience arrives. The economist's next visit was recorded at close range and needed a high-pass and nothing else.
Worth arguing about
Compression made the recording worse. What exactly did it do?
One answer
A compressor turns down what exceeds a threshold and makes nothing louder; the make-up gain added afterwards raises everything that is left, including the fans, the corridor and the room tail. Narrowing the range between the loudest and quietest parts and then lifting the whole signal raises the noise floor by the same amount as the voice, so a recording with a poor noise floor can take less compression, not more.
Why was resynthesis rejected here when it might be acceptable elsewhere?
One answer
Because the recording is a record of what a named person said, and it will be quoted and assessed. A resynthesising tool does not filter the signal, it generates a new one from a model, so it can alter timbre and invent detail that was never present. That is tolerable for a corporate voice-over where nothing turns on exact fidelity, and it is not tolerable where the audio is evidence of what somebody said. The same tool, the same result, a different obligation.
The department's fix cost nine thousand rupees and half a page of rules. Which half was doing the work?
One answer
Mostly the distance, not the money. The reason a camera microphone eleven metres away sounds like a hall is the ratio of direct sound to reflected sound: direct sound falls away with distance while reflections are spread evenly through the room. A cheap lavalier at twenty-five centimetres changes that ratio more than any processing can afterwards. The recorded room tone and the backup camera track are what make the rest of the edit possible.
Test yourself6 questions on this module
Why does a cheap lavalier at 25 cm beat an expensive microphone two metres away?
Twelve decibels of spectral noise reduction leaves a watery, fluttering sound around the voice. What happened?
One dialogue clip arrives at the compressor at minus twenty and the next at minus eight. Why does no track fader setting fix both?
A hissy recording sounds worse after compression than before it. Why?
Why does mastering a film to minus six LUFS gain nothing on a platform that normalises to around minus fourteen?
What does a continuous ambience bed under a whole scene actually fix?
Module 5
Colour, correction before grade
Colour is the area where the free tool is not a compromise: DaVinci Resolve's colour page is what features are graded on and it costs nothing. This block works through it in the order the work actually happens — the node graph that holds your decisions, the primary controls and what each one touches, curves, the shot-matching workflow that makes a scene belong to itself, keys and windows for changing one part of a frame, what a colour-managed pipeline is doing to your image, how a look is built rather than downloaded, and the levels check that stops an export looking washed out.
By the end you can
Take a scene of mismatched shots to a neutral, matched baseline verified on the waveform and parade, apply one look across all of it, and explain why the same LUT lands differently on a log clip and a Rec.709 clip
- 40Make it right, then make it feel like somethingNeutralise and match every shot against the scopes first, because a look applied to unmatched shots produces unmatched results.
- 41Why grading happens in a graphA node graph lets you choose what each adjustment applies to and in what order, so correct first and style after, name every node, and use the still store to carry one shot's grade across a whole scene.
- 42The three wheels, and what each one touchesGain multiplies so it sets the white point, lift adds so it sets the black point, gamma is a power curve so it moves the mid-tones with both ends pinned — set them in that order and read every move on the waveform.
- 43Curves, and the shape of contrastA curve maps input values to output values, so a gentle S steepens the mid-tones where faces live and a reverse S produces the faded film look — and Lum versus Sat pulled down in the shadows stops dark noise reading as coloured.
- 44Making a scene belong to itselfGrade one hero shot completely and bring every other shot to it on a split wipe, matching black point, white point, mid-tone, balance and saturation in that order on the scopes — and judge the result at playback speed, not frame by frame.
- 45Changing one part of the frameQualifiers select by colour and windows by shape, both feeding a node's key — and because a qualifier builds its mask from chroma, 4:2:0 footage yields a mask at a quarter the resolution of the picture, which no setting repairs.
- 46Colour spaces, and why it looks different in the browserA pixel value only means a colour once you know the space, white point and transfer function — so the washed-out export is usually a full-range file being read as video range, and the browser mismatch is usually gamma 2.2 against 2.4.
- 47Building a look instead of downloading oneA look is five deliberate decisions — contrast shape, a shadow-highlight colour split, saturation shaped by luminance, a hue shift or two, and grain — so build and save your own rather than applying a LUT that expects a log image it is not receiving.
- 48Legal levels, and the check before you exportBroadcast-legal means values between 16 and 235 in 8-bit video range, readable as lines on the waveform — and every quality-control check happens on the exported file in a normal player, because several classes of error exist only in the export.
Case studyOne shot under a tube light, and the window behind it
The scene is a two-minute sequence in a consulting room. A doctor explains what the clinic does about follow-up appointments, and a patient — who has consented, on paper, for this one film — describes what changed for her.
Most of it matches without much trouble. The room was lit by a large window on the left and the shots were filmed within forty minutes of each other. The exception is one shot: the patient's answer about the phone call that brought her back, which is the emotional centre of the sequence and the only take of it. It was filmed twenty minutes later, after the sun had gone behind the block opposite, with the overhead tube lights switched on.
On the parade the difference is obvious. The tube light is strongly green in the mid-tones, so the patient's skin sits well off the reference line, and the green is not uniform across the frame: the window behind her still carries the last of the daylight and reads cold. The editor corrects for the green in the skin and the window goes magenta. He corrects for the window and the patient goes green again. There is no single primary correction that fixes both, because there are genuinely two different light sources in the frame.
The technical options are clear and so are their costs.
A qualifier and a tracked window. Draw a soft window around the patient's head and shoulders, track it across the shot, and qualify skin inside it so the correction touches only her face. Then treat the window behind her separately. This is the textbook answer and it works. The complication is the footage: the camera recorded 4:2:0, so the qualifier is building its mask from a quarter of the picture's resolution. Blurring the matte enough to hide the stepped edges means the correction bleeds onto the wall behind her head, and the patient moves, so the window has to track through about forty seconds of small movements. The editor's estimate is an evening, and he has two other scenes to grade.
Accept a compromise primary. Split the difference: leave the skin slightly green and the window slightly magenta, and reduce saturation in both so neither is loud. Twenty minutes. The scene reads as competent, and the one shot the film is built around is the least attractive in it.
Do not let the shots touch. Put a cutaway between the daylight shots and the tube-light shot, so the viewer never sees a direct join. The clinic's own waiting-room signage, the appointment book, a wide of the corridor. Cutting away also means the mismatch is compared across two seconds of something else rather than across a single frame, and the eye is considerably more forgiving of that.
The third option is an editorial answer to a colour problem, and it is the one that keeps getting forgotten by people who are enjoying the colour page.
There is also a decision that is not about this shot at all. The whole project was set up with colour management off, the timeline at Rec.709 gamma 2.4, and every log clip carrying its own input transform. The editor is tempted, standing in front of this mess, to switch the project to managed colour and let it sort the sources out. Doing that would change the starting image under every grade he has already built on the other seven minutes.
What actually happened
He did the third option first, because it cost ten minutes: a two-second insert of the appointment book, taken from material shot on day one, placed at the join. Then he did a reduced version of the first: a soft tracked window around the patient, a gentle green correction inside it and nothing else, with the matte blurred heavily and the adjustment kept small enough that its edges could not be seen. That took about fifty minutes rather than an evening, because he stopped trying to make the two shots identical and only tried to make the join survive playback at speed. He did not switch to colour management. The scene was signed off without comment, which is the outcome a grade is aiming for. His note to himself afterwards was about the shoot rather than the grade: switch the tube lights off, or switch them on at the start, but do not change the light source halfway through a scene.
Worth arguing about
Why could no single primary correction fix both the patient and the window?
One answer
Because there were two different light sources in the frame with different colour temperatures. A primary correction applies to the whole image, so any shift that neutralises the green of the tube light also pushes the daylight in the window magenta by the same amount. The only ways out are a secondary that treats one region separately, or an editorial decision that stops the two shots being compared directly.
The qualifier produced stepped, unstable edges. What property of the source file caused that, and what setting fixes it?
One answer
The footage was 4:2:0, so colour was recorded at half the horizontal and half the vertical resolution of the picture. A qualifier builds its mask from colour, so the mask carries a quarter of the information the image does and its edges come back blocky and crawling. No setting on the keyer fixes it. The practical responses are to blur the matte and keep the correction subtle, or to avoid a colour-based selection entirely.
He was tempted to switch the project to managed colour. Why was that the wrong moment?
One answer
Because colour management changes the image every grade is built on. Seven minutes of the film had already been corrected against the unmanaged starting point, so switching would have altered the input to all of that work and required regrading it. Managed colour is a decision for the start of a project, or for a project with several camera types where the manual route is costing time. It is not a repair for one difficult shot.
Test yourself6 questions on this module
Every shot in a scene came off one camera in one profile, filmed over two hours in changing light. Why does one LUT applied to all of them produce unmatched results?
A face is too dark, but the black point and the white point are both where you want them. Which control moves the face without disturbing either end?
On the RGB parade the bottoms of the three traces do not align, and blue sits highest. What is wrong, and what fixes it?
A skin qualifier on phone footage produces a blocky, crawling matte. Which of these is true?
An export has grey blacks and dull whites in one player and looks correct in another. ffprobe reports a blank color_range field. What is happening?
Why does adding a little grain before export reduce visible banding in a sky?
Module 6
Text, captions and graphics
Text on video has constraints that text on a page does not: it moves, it sits over a background that changes every frame, it goes through a compressor that hates thin strokes, and most of the people reading it are doing so on a phone with the sound off. This block covers sizing and contrast for that reality, name supers and credits, keyframing inside a timeline, what free tools replace After Effects with, subtitle files as files, the specific breakage that Indic and other complex scripts suffer in video software, screen recordings, and putting numbers on screen honestly.
By the end you can
Put text on a moving picture that survives platform compression on a phone — sized and weighted against a stated safe area, timed to a read-aloud test, animated with easing you chose — and ship both burnt-in captions and a sidecar subtitle file, including in a script your editor does not shape correctly
- 49Most of your audience has the sound offBurn captions in for feeds and upload a real subtitle file as well, and never publish an auto-transcript you have not read.
- 50Type that survives a phone and a compressorSize text as a percentage of frame height with a floor around 2.5 per cent, never use a light weight because the compressor destroys thin strokes, give every line its own contrast mechanism against a changing background, and hold it long enough to read aloud twice.
- 51Names, supers and the end of the filmA super answers 'who is this' two or three seconds after somebody starts speaking, its second line says why they matter rather than their job title, and every super in a piece comes from one duplicated template — while credits discharge real licence obligations from your log.
- 52Keyframes, curves and the retime graphTwo keyframes and an interpolation is all timeline animation is, so ease every move because nothing physical changes value at a constant rate — and remember that a speed ramp needs easing on the retime frame graph rather than steps on the speed graph.
- 53Graphics without a subscriptionFusion inside Resolve, Blender for 3D, Inkscape for vector artwork and a browser for data-driven graphics cover every free route to motion graphics — and one title template you made and reuse beats a downloaded pack everybody else is also using.
- 54Subtitle files, as filesA subtitle file is plain text with an index, two timecodes and up to two lines, which means find-and-replace fixes a recurring error in one action — and a constant sync offset can be shifted while a growing one means the video's frame rate needs conforming.
- 55When your editor cannot render the scriptIndic and Arabic scripts need a shaping engine to reorder and combine glyphs, and many video editors do not have one — so test with a conjunct and a reordering vowel, and use libass through FFmpeg or a sidecar file rather than the title tool.
- 56Editing a screen recording so it can be readCapture at the delivery resolution so interface elements stay large, punch in with an eased move rather than a cut whenever detail matters, ramp through every moment of waiting, and redact with a tracked solid box because blur on text is not reliably irreversible.
- 57Putting numbers on screen honestlyA chart on video carries one comparison, revealed in time so the animation directs attention, with labels on the elements rather than in a legend — and because a viewer cannot go back and check the axis, starting a bar chart anywhere but zero is a stronger deception on video than on a page.
Case studyMarathi captions that the editor could not render
The brief is specific and reasonable. The film is about a free eye-screening camp, it will be watched almost entirely on phones in a muted feed, and it must carry Marathi captions, because the audience it is for reads Marathi and a good part of it does not read English.
The cut is finished. The mix is finished. The Marathi translation came back from the corporation's own office, checked, in a document. Sneha pastes the first line into her editor's title tool and it comes out wrong.
Not missing-glyph boxes, which would have been obvious. The letters are all there and they are in the wrong shapes. The conjuncts have broken apart, with the halant visible as a separate stroke between letters. In two words the vowel sign that should be drawn to the left of its consonant is sitting on the right. It reads, to somebody who reads Marathi, as nonsense wearing the right letters.
She tries three fonts, including a Noto face, and gets the same result each time. That is the tell: it is not a font problem. The editor's text engine is not shaping. Complex scripts need a shaping engine that reads the font's OpenType tables and applies its substitution and reordering rules, and this tool either does not use one or uses an old one. She confirms it with a two-character test: कि renders with the vowel sign on the right.
It is four in the afternoon. The options are these.
Ship it in English. The English captions render perfectly and the film goes out on time. It also fails the brief, and it fails the audience the campaign exists for.
Burn the Marathi captions in with FFmpeg. The subtitles filter uses libass, which uses HarfBuzz, which shapes correctly. It is one command with a force_style string, and it re-encodes the video, so it has to run from the master rather than from the delivery file. Sneha has never used FFmpeg. Her partner Vikram has, twice. Neither of them has done it under a deadline, and if the command produces something subtly wrong at eleven at night, neither of them reads Marathi well enough to catch it.
Set every caption as an image. Type the Marathi in Inkscape, which shapes correctly, export each cue as a transparent PNG at four times the size, and place them on the timeline. This definitely works and it is entirely mechanical. The film has thirty-one cues. Sneha estimates four minutes a cue with positioning, so about two hours, plus the risk of a typo in any one of thirty-one separate pieces of artwork.
Sidecar only. Upload the video with a .srt and let the platform's own text engine shape it. This renders correctly, costs nothing, takes ten minutes, and produces captions that do not exist in a muted autoplay preview — which is precisely where this film lives.
The fourth option is the one that looks cheapest and is the one that most directly defeats the purpose of the job. The first is the one the deadline wants. The decision is really about who the film is for, and there is a fifth thing nobody has thought of yet, which is that they could ask.
What actually happened
Vikram rang the corporation's communications officer at half past four and asked for the morning rather than tomorrow. She gave them until eleven. They then did the second option, and gave themselves the evening for it: Vikram burnt the captions with FFmpeg and libass using Noto Sans Devanagari, and increased the line spacing by twenty-five per cent because the matras were clipping. Sneha sent a phone screenshot of six representative cues, including the conjunct and the reordering vowel, to the officer, who read them and corrected one word. They also uploaded the sidecar .srt, so the file exists for search, for translation and for assistive technology. The studio now runs the two-character test on every tool it uses and keeps a note of the answer, and its quotes for any job in an Indic script carry a line about caption rendering.
Worth arguing about
Why did changing the font, including to Noto, not fix the rendering?
One answer
Because the failure was in shaping rather than in the glyphs. The font contained everything needed; what was missing was an engine to read its OpenType substitution and positioning tables and apply them, so conjuncts were not formed and the vowel sign was not reordered to the left of its consonant. A font supplies shapes and rules. Without a shaping engine such as HarfBuzz, the rules are never applied, whichever font is installed.
The sidecar file renders correctly and takes ten minutes. Why was it not the answer on its own?
One answer
Because the film's whole audience watches in a muted autoplay feed, and a sidecar track is not shown there unless the platform chooses to display it. The captions were the primary channel for this piece, not an accessibility addition at the end. The sidecar was still uploaded, because it is what makes the film searchable, translatable and usable with assistive technology, but it could not be the only caption.
What did asking for an extra few hours actually buy, beyond time?
One answer
It bought a proofreading pass by somebody who reads the language. The route they chose renders correctly, but neither of them could have verified that the output was right, and an incorrectly shaped caption looks confident and says something else. Six screenshots sent to a Marathi reader caught a wrong word and confirmed the shaping. The deadline was the constraint that had been pushing them toward the option that failed the brief.
Test yourself6 questions on this module
Why is an unread automatic caption track more dangerous than one full of obvious gibberish?
Thin type looks elegant in the editor and breaks up after upload. What is the mechanism?
Captions near the bottom of a vertical video are legible in the editing canvas and unreadable on the platform. Why?
A scale from 100 to 110 per cent with two linear keyframes reads as mechanical. Why?
A subtitle track drifts further out of sync as the film goes on, and shifting the whole file does not help. Why not?
A Hindi caption typed into an editor's title tool shows the vowel sign to the right of its consonant. What does that tell you?
Module 7
Effects, compositing and repair
The shots you were not given, made from the ones you were. This block covers stabilisation and the crop it costs, retiming and the two ways of inventing frames, masks and rotoscoping, keying footage that was lit badly, tracking a point and a plane, removing an object from a moving shot, noise reduction and the honest limits of sharpening, transitions built rather than dropped in, and a diagnosis for a timeline that has stopped playing back.
By the end you can
Choose between stabilisation, retiming, masking, keying, tracking, removal and denoise for a given repair, and state for each one the artefact it produces and the property of the source footage that determines whether it will work
- 58Stabilisation, and the crop it costsStabilisation cancels frame movement and scales up to hide the empty edges, so it always costs 5 to 15 per cent of the frame, cannot remove motion blur already in each frame, and needs a separate rolling-shutter pass or it makes skew more visible rather than less.
- 59Retiming, and the two ways of inventing framesSlowing a clip means inventing frames, and the three methods trade differently — duplication stutters but invents nothing, blending ghosts, and optical flow is seamless until its motion estimate is wrong and the image melts at the region it got wrong.
- 60Masks, and the honest cost of rotoscopingA mask is a shape with feather, expand and invert, and animating one by hand costs roughly one to four hours per five seconds of footage — so keyframe sparsely, split the subject into simple parts, and look for a solution that is not rotoscoping first.
- 61Keying footage that was lit badlyA keyer builds its matte from colour, so a garbage matte, an evening-out pass and a denoise before keying do more than any setting on the keyer itself — and a subject in 4:2:0 with green spill in their hair has a hard limit no tool removes.
- 62Tracking a point, a plane and a cameraPoint tracking gives position, planar tracking gives a whole surface and survives occlusion, and camera tracking needs real parallax and a rigid scene — so pick the cheapest one that solves the problem, and expect the composite to fail on matching rather than on the track.
- 63Taking something out of a moving shotRemoval means covering something with pixels that plausibly belong — try reframing first, patch from the same frame on a locked-off shot and from another frame on a moving one, and always match grain, because a patch that is too clean is what gives a removal away.
- 64Noise reduction, sharpening, and what neither can doTemporal denoise averages the same pixel across frames so noise cancels and detail survives, spatial denoise averages within a frame so detail goes with the noise — remove chroma noise aggressively and luminance gently, denoise before grading, and re-add grain at the end.
- 65Transitions you build rather than drop inA hard cut is right most of the time and a dissolve means time passed — the transitions worth learning are match cuts, whip pans and masked reveals, which you build from what is in the frame rather than drag from a bin.
- 66Diagnosing a timeline that has stopped playingCheck in order — proxies on, effects bypassed, scopes closed, playback resolution reduced, background renders finished, disk fast enough — because slow playback is almost always one specific effect or one missing toggle rather than an inadequate machine.
Case studyThe man in the background who changed his mind
The shot is eleven seconds long and it is the one that opens the film. A handheld move down the length of the market at six in the morning, the light coming in sideways off the water, boxes being carried across frame. It establishes the place in a way nothing else in the rushes does, and every version of the cut has opened with it.
Four days before picture lock, a man who appears in it for about five of those eleven seconds sends a message. He had signed the release on the day. He now works somewhere that he would rather not be associated with the market, he has thought about it, and he asks not to be in the film.
The director's position is that the legal question and the right answer are not the same question. She has a signed release. She also has somebody who has told her plainly that he does not want to be in it, and a film that will be shown to an audience in the city he lives in.
So the shot has to change. There are four ways, and each costs something different.
Cut the shot. Eleven seconds, gone, and with it the only wide that establishes the market as a place. The next best opening is a close-up of ice being poured, which is a good shot and tells the viewer nothing about where they are. Cost: the opening, and about a day of re-cutting the first ninety seconds around it.
Reframe. The source is 4K and the delivery is 1080p, so there is room to scale and reposition. The man walks through the middle third of the frame from left to right, and the camera is moving with him for part of it. A reframe that keeps him out would have to move continuously, would end up at roughly 190 per cent for two seconds, and would lose the two boxes being carried in the foreground that give the shot its depth. Cost: the shot becomes a different and weaker shot, in about two hours.
Remove him. Cover him with pixels that plausibly belong. The camera moves, so a static patch slides off: the surface behind him has to be tracked, the patch attached to the track, masked with a soft edge and matched for grain. Behind him is a row of crates, a wet floor with reflections, and two other people who cross behind him at second six. That last detail is the expensive one — a patch that has to respect the edge of a moving person is rotoscoping, at somewhere between one and four hours per five seconds of footage. Cost: realistically a day and a half, with no guarantee, four days before lock.
Blur or box him. A tracked solid box or a heavy mosaic. It is reliable and it takes an hour. It also announces itself in the opening shot of a documentary, which tells the audience that something has been hidden in the first eleven seconds of the film.
There is a fifth option she almost misses, which is that the market runs every week and the light at six in the morning is the same light.
What actually happened
She re-shot it. Two mornings later she filmed the same move four times in twenty minutes, told the three people in frame what it was for, and got a wider version than the original with better foreground. The re-shot opening is in the film. The removal was never attempted, because a day and a half of tracked patching over a moving person, four days before lock, was not a risk the schedule could absorb. She wrote two things into her own practice afterwards: that a release is the floor and not the ceiling, and that she would find out on the shoot day which shots were load-bearing, because the shot she could least afford to lose was the one she had filmed once.
Worth arguing about
The director had a signed release. Why did she treat the request as binding anyway?
One answer
Because consent to be filmed is about scope and about what a person understood they were agreeing to, and this man was telling her clearly that the use now in front of him was not the one he had in mind. The release settles the legal question; it does not settle whether showing him is defensible in a film that will screen in the city where he lives and works. Treating the release as the end of the matter would have meant using somebody's face against their stated wishes because a form allowed it.
Why was removal so much more expensive here than in a locked-off shot?
One answer
Because the camera moved. On a locked-off shot the background behind the object is identical in every frame, so a clean region copied over it with a soft mask holds for the whole shot. With a moving camera the patch has to be tracked onto the surface, matched for grain and brightness, and masked; and where another person crosses in front of the area being patched, the patch has to respect a changing edge, which is rotoscoping and is budgeted in hours per second.
What made the re-shoot possible, and what does that suggest about how to log a shoot?
One answer
The event was recurring and the conditions repeatable: the market runs weekly and the light at six is the same. The fifth option was almost missed because the problem presented itself as a compositing problem. The practical lesson is to know, at the shoot, which shots are load-bearing and which exist once — and to film the load-bearing ones more than once, since the shot she could least afford to lose was the one she had a single take of.
Test yourself6 questions on this module
Why does stabilisation always cost some of the frame?
A very shaky shot is stabilised. It comes back steady, and every individual frame is blurred. What is the cause?
A clip slowed with optical flow melts around a striped shirt and is clean everywhere else. Why that region?
Why is rotoscoping described as a budgeted task rather than a quick repair?
A green screen was lit with one lamp and keys badly. Which step does more than any setting on the keyer?
Removing a light stand from a locked-off shot takes minutes and from a handheld shot takes hours. What is the difference?
Module 8
Delivery — the file that leaves the building
Containers, codecs, bitrate, chroma and bit depth; what a platform does to your upload; the settings for each destination; ffmpeg for the jobs your editor refuses; masters and textless versions; the seven ways an export goes wrong; and the check you run on the file rather than the timeline.
By the end you can
Choose a container, codec, bitrate and colour tag for a named destination and justify each from the material and from what the platform's own encoder will do to it, render a master and textless version every later deliverable can be derived from, and diagnose a render that is washed out, drifting, juddering or failing at the same percentage by naming the mechanism responsible
- 67Why the upload looks worse than the fileThe platform re-encodes everything, so upload a generously encoded file at the source frame rate and mix near -14 LUFS rather than louder.
- 68The box and the codec are two different decisionsA filename tells you the container and nothing about the codec, so repack with a remux rather than a rename, and check a file with ffprobe before you send it anywhere.
- 69Bitrate, and why the same number is generous and stingyCompare bitrates as bits per pixel per second rather than as raw megabits, and expect motion, fine detail and grain to need several times what a static interview needs.
- 704:2:0 and 8 bits: where the colour wentDelivery formats keep one colour sample per four pixels and 256 levels per channel, which costs you saturated edges and smooth gradients — and a trace of added grain hides the banding that results.
- 71What the platform does to your uploadYour upload is source material for a ladder of renditions you never see, so hand it something clean and uncompressed-twice, match the frame rate, and judge nothing until processing has finished.
- 72The numbers, destination by destinationSave a named render preset per destination with the frame rate taken from the timeline rather than the preset, because the wrong frame rate is the mistake that survives every other check.
- 73Six ffmpeg commands worth knowingffmpeg never modifies its input, so inspecting, remuxing, converting a variable frame rate recording and measuring the loudness of a finished file are all free experiments.
- 74The master, and the versions to make while the project is openRender a full-frame, textless, intra-frame master plus stems at the end of every project, because deriving a new deliverable from a file always works and deriving one from a project only works while the application, fonts, plugins and media all still resolve.
- 75Seven ways a render lies to youRender a sixty-second test from the hardest minute of the timeline and open it in two players, because gamma shifts, hardware artefacts, sync drift and file size all declare themselves in a minute and cost an evening after four hours.
- 76Checking the file, not the timelineWatch the rendered file once at speed without stopping, then measure its duration, loudness, codec and first and last frame, because every verification you did on the timeline was of something nobody receives.
- 77The project that still opens in six monthsLearn JKL and three-point editing, name folders by date and version, and back up the project file separately from the footage.
Case studyIt looks fine here and washed out there
The grade took most of a day and it is good. Deep blacks in the warehouse, clean whites on the fabric, skin on the two machinists sitting where it should on the vectorscope. The editor, Deepak, exports an H.264 review copy and sends the link.
The reply comes back in forty minutes: the colour has gone flat, the blacks are grey, and it looks worse than the version they saw last week. Could he put the contrast back.
Deepak opens the same file on his own machine and it looks correct. He opens it in VLC and it looks correct. He opens it in his browser and it is slightly lifted. The client's screenshot, taken on the office laptop, shows a noticeably washed-out image with grey blacks.
Two mechanisms could be doing this and he has to tell them apart, because the fixes are opposite.
The first is a gamma difference. Rec.709 is defined for a display gamma of 2.4 in a dim room. The web standard is approximately 2.2. A grade that is correct at 2.4 looks slightly washed in a browser, and the browser is applying the standard it is supposed to apply. That difference is real and it is small, and it lives in the shadows.
The second is a data-level mismatch, and it produces a large, obvious error rather than a subtle one. A file can be written full range, with 8-bit values running 0 to 255, or video range, with black at 16 and white at 235. If it is written as one and read as the other, everything shifts: video-range read as full crushes, full-range read as video lifts the blacks to grey and dulls the whites. The file carries a flag; some encoders write it wrongly and some players ignore it.
ffprobe on the export returns a blank color_range field. Nothing was tagged. Every player is guessing, and the client's player is guessing differently from his.
So he has a choice, and the shape of it matters more than this one job.
Regrade to match the client's screen. Add contrast until the screenshot looks right. It would take twenty minutes and the client would be happy today. It would also make the film wrong in VLC, wrong on his own machine, wrong on the exporter's website, and wrong on any player that reads the tag correctly. He would be chasing one player for ever, and the next client would have a different one.
Fix the tagging and prove it. Set the colour tag explicitly on the render rather than leaving it automatic, re-export, run ffprobe to confirm the field is populated, and check the result in two different players and a browser. It costs an export and an hour. The client did not ask about tagging and will not care about the explanation.
The second cost is the awkward part. The client's complaint was "put the contrast back", and answering it with a paragraph about transfer functions is a way of being right and being ignored.
What actually happened
He set the colour tag explicitly on the Deliver page, re-exported, confirmed with ffprobe that color_range now read tv, and checked the file in VLC, in QuickTime and in Chrome on his own machine before sending it. He wrote two sentences to the client: that the first file had been sent without a colour tag so different players were interpreting it differently, and that the new one was tagged and should look the same everywhere. It did. He then added three lines to the studio's checklist: tag the export explicitly rather than leaving it automatic, run ffprobe on every finished file before it leaves, and watch every review copy in a second player. The next time a client said the colour looked wrong it turned out to be the office laptop's night-shift setting, which the checklist also now asks about.
Worth arguing about
How did the blank `color_range` field distinguish between the two possible causes?
One answer
A gamma difference between 2.4 and 2.2 is a small shift in the shadows and would not produce grey blacks in a screenshot. A missing range tag means no player has been told how to interpret the values, so each one guesses, and a mismatch between how the file was written and how it is read produces the large lift the client saw. The blank field was positive evidence for the second mechanism and against the first.
Why would regrading to match the client's screen have been the expensive option, given that it was faster?
One answer
Because it would have baked a correction for one player's mistake into the film itself. Every player that reads the tag correctly would then have shown an over-contrasted image, including the exporter's own website and any later deliverable derived from the same master. The fault was in the file's metadata, not in the grade, so the only fix that survives is one applied where the fault is.
The client asked for contrast and got a file with a colour tag. What made that answer land rather than sound evasive?
One answer
That the file demonstrably looked right on the client's own laptop afterwards, and that the explanation was two sentences rather than a lecture. A complaint is a symptom report: the client correctly observed that something was wrong and proposed the wrong cause. Fixing the cause and showing the result addresses the complaint; explaining transfer functions instead would have been accurate and useless.
Test yourself6 questions on this module
Renaming a .mkv file to .mp4 does not make it play. What does a remux do that renaming does not?
A 1080p25 export at 10 Mbit/s looks clean. The same material at 2160p25 and 10 Mbit/s smears. Why?
A carefully grained grade turns to porridge after upload. What is the mechanism?
A video looks terrible ten minutes after publishing and fine the next morning, with no re-upload. What happened?
Why is a master rendered textless and at the full frame rather than at the delivery crop?
You played the whole programme through the timeline's loudness meter and read minus fourteen integrated. ffmpeg on the exported file reports minus nineteen. What does that tell you?