Testing a video model honestly
Why reels tell you nothing
Every video model is launched with a reel. The reel is real output and it is also the survivors of an unknown number of generations, chosen by people whose job is to choose well, on prompts selected because they worked.
Nothing about it is dishonest and nothing about it is informative. The number you need is the one the reel cannot contain: how many attempts produced how many usable shots, on prompts you did not choose.
A battery you can run in an hour
Fix a set of prompts before you test anything, and reuse it for every model. Ten is enough. A serviceable set:
- A person walking toward the camera, medium shot.
- Two people talking, one gesturing.
- Hands performing a task — tying a lace, pouring tea.
- A slow dolly in on a static object.
- An animal moving quickly.
- Water or fire.
- A vehicle passing through frame.
- Text on a sign, held in shot.
- An object passing behind another and re-emerging.
- A specific action with a first and last frame supplied.
Generate four attempts for each on every model. Forty clips is an hour and a modest cost, and it gives you something no review will.
Scoring without deceiving yourself
Decide what usable means before you look, and write it down. A workable definition: a shot you would place in a paying client's edit without an apology. Then score each clip as usable or not, with no middle category — the middle category is where self-deception lives.
Report the result as a rate: "6 of 40 usable, 15%". That single number is comparable across models, across months, and across the claims anyone makes to you.
Record two more things while you are there:
- Which prompts never produced anything usable. This is the model's shape, and it is more useful than the average.
- Time and cost per usable shot, not per generation. A model with a 30% hit rate at twice the price is cheaper than one at 10%.
The failure taxonomy worth keeping
When a clip fails, record why, using the categories from this module: temporal drift, morphing, anatomy, physics, camera did not follow, subject did not follow, texture flicker, resolution. After forty clips you have a profile rather than an impression, and profiles are what let you choose the right model per shot instead of adopting one for everything.
This is exactly the method the evaluation course in this catalogue teaches in general form: a fixed set, a stated definition of success, a rate rather than an anecdote. It applies here because video is the domain where impressions are least reliable — the outputs are seductive, individually memorable, and heavily selected before you ever see one.
What to re-test, and when
Models change. Hosted ones change without announcement. Re-run the battery when a version number changes, when your hit rate drops noticeably, and before any commitment that depends on the result.
Keep the clips. A folder of the same forty prompts across four models and two years is a genuinely valuable thing to own, it costs nothing but disk, and it is the only way to answer "is this actually better" rather than "does this feel better", which it always does at launch.
Judging without knowing which model made it
One refinement is worth the small extra effort, and it comes straight from ordinary experimental practice. Rename the clips so you cannot tell which model produced which, shuffle them, and score them in that state. Then reveal the labels.
People who have done this are usually surprised. Expectation does a great deal of work in judging generative output — a clip from the model you have just read about looks better than the identical clip from the one you have not — and a blind pass removes it for nothing but a minute of file renaming. If you are choosing a tool your team will pay for, this is the difference between a decision and a preference.
The honest limitation of the method: a ten-prompt battery covers what you thought to include. It will miss the failure that ruins your particular project. So add prompts drawn from your real work as you meet them, and treat the battery as a growing regression suite rather than a fixed exam. Each shot that failed on a real job becomes prompt eleven, and the set gets more useful with every disappointment.
The one thing to keep
A demonstration reel is selected output, so judging a video model requires a fixed prompt battery, a stated number of attempts per prompt, and a usable-shot rate rather than an impression.
Before you move on
A vendor's reel is flawless and your own testing gives a 12% usable rate. What is the most likely reconciliation?
Pick the one you would defend. Nobody sees your answer.