How a demo is built
Demos are edited
A demonstration is a performance. That is not an accusation — a live musical performance is edited too, in the sense that the musician chose the piece. But it means a demo answers the question can this ever work? and never the question how often does this work?, and those get confused constantly, including by the people giving the demo.
The editing techniques are worth naming, because once named you see them in every launch video.
Best of many. The output shown is the best of ten generations. Legitimate for showing a capability ceiling, misleading as an expectation. Ask: was this the first attempt?
The prompt is not shown. What appears on screen as a casual request is often the visible half of a long, carefully tuned instruction, sometimes developed over days. The gap between a tuned prompt and yours can be the whole result.
Curated inputs. The invoice in the demo is clean, upright, in English, in a familiar layout. Ask to see it run on the worst document you actually receive. This request ends more sales conversations than any technical objection.
Latency edited out. Videos cut the wait. A response that takes forty seconds feels different in a meeting than in a nine-second clip.
Failure not shown. No demo shows the refusal, the timeout, or the confident wrong answer. You are watching the survivors.
The people in the machine
A harder version: a system presented as automated is partly people. This has a long history — the original Mechanical Turk of 1770 was a chess automaton with a man inside — and it recurs because the incentive is permanent. A startup can promise automation, deliver humans, and hope the automation catches up before anyone checks.
Recent years have supplied several examples. A drive-through voice ordering service was reported to rely on offsite human workers reviewing a large share of orders. Amazon's cashierless "Just Walk Out" stores were reported in 2024 to depend on a substantial number of human reviewers watching video to resolve uncertain baskets; Amazon disputed the characterisation of the numbers while confirming that human review was part of the system. In 2023 a US company selling AI-driven engineering software was charged by regulators over claims about the extent of its automation.
Be careful here: humans in the loop are not fraud. Most useful systems have them, and saying so is honest engineering. The problem is a claim of full automation, priced and regulated as full automation, that is actually a wage bill somewhere with weaker labour protections. The test is disclosure. Does the description match the machinery?
Questions that survive a demo
Six, and they are not hostile. Any team that has actually deployed will have answers.
- What is the success rate on a random sample of real inputs, not chosen ones? A number, with the sample size.
- What does it do when it cannot do the task? If the answer is "it always produces something", you have learned the most important thing in the room.
- How much human review is in the current pipeline, and at what cost per item?
- Show me three failures. A team that cannot produce failures has not looked, which tells you more than the failures would have.
- What happens when the provider changes the model underneath? Most products are built on someone else's model, which gets updated on someone else's schedule.
- Who is accountable when it is wrong, and what is the remedy? This is a contract question, not a technical one, and it is often where the conversation quietly ends.
Your own pilot, honestly run
If you get as far as trying something, protect the trial from the same effects. Choose the sample before you see the results, including the awkward cases you would rather not include. Fix the criterion for success in advance and write it down. Have the people who will actually use the thing run it, not the person who is enthusiastic about it. And measure the total time including checking and correcting — a tool that produces a draft in four seconds and takes eleven minutes to verify has not saved eleven minutes.
The point is not scepticism as an attitude. It is that a demo is evidence of possibility, and a deployment decision needs evidence of frequency, and nobody will volunteer the difference.
The one thing to keep
A demo shows the best case under chosen conditions, so the only questions that carry information are the frequency ones: success rate on unchosen inputs, behaviour on failure, and how much human work is inside.
Before you move on
A vendor demonstrates flawless extraction from ten invoices. Which single follow-up request tells you most about whether it will work for you?
Pick the one you would defend. Nobody sees your answer.