AI at Work
The tasks it genuinely helps with, the ones it quietly ruins, and the line you must never cross.
- Level
- Nothing assumed
- Lessons
- 73
- Reading time
- 662 min
- Price
- Free, no sign-up to read
Eight modules and seventy-three lessons on using AI in a real job, without pretending it is magic and without pretending it is useless. Written for accountants, teachers, nurses, lawyers, shopkeepers, marketers and civil servants anywhere in the world. Covers which of your tasks AI is actually good at, how to draft and summarise well, working with your own documents and spreadsheets, meetings and email, the confidentiality lines you must never cross, how to check work you are accountable for, and how to talk about all of this with a sceptical team.
Start the first lessonModule 1
Deciding what to hand over
Before any prompt, there is a decision: which of the things you do this week should touch AI at all. This block gives you a sorting rule you can apply to your own job, the evidence on where the help is real and where it reverses, and the mechanism behind the failure that keeps catching careful professionals.
By the end you can
Audit a week of your own work and place each recurring task into hand over, assist or never, defending each placement by where the truth lives, what a wrong answer costs, and what checking costs.
- 1What it is actually good atIf you did not supply the facts, treat every fact it returns as unverified — fluency is not evidence.
- 2What it is actually doing when it answers youThe model only predicts the next fragment of text, so it has no memory, no clock, no calculator and no register of what it does not know — and every characteristic failure follows from one of those absences.
- 3The context window, and why a long chat gets worseEvery turn resends the entire conversation, so instructions given early can be truncated away, attention is weakest in the middle, and cost grows with the square of the chat's length — which is why one task should mean one conversation.
- 4The task auditAdoption happens task by task, not tool by tool — and any task where checking the output costs more than doing the work yourself is not a candidate, however good the draft looks.
- 5What the studies actually foundThe measured gains are real, largest for the least experienced, and reverse on tasks just outside the tool's competence — which you cannot feel from inside the task, so you have to time it rather than trust your sense of speed.
- 6Why it invents a case that does not existInvention is not a malfunction but the same process that produces correct answers, so it arrives in the identical confident register — which is why you check the specifics you did not supply rather than the ones that look shaky.
- 7The check is the jobTime saved is writing time minus briefing, checking and fixing, and the cost of checking depends on where the facts came from — which is why supplying the source, rather than asking from memory, is the habit that makes the arithmetic work.
- 8Why a second pair of eyes stops workingTrust is calibrated to how often a tool is wrong, so raising its accuracy lowers your attention in step — and two people reading the same generated draft are anchored by the same text rather than checking independently.
- 9Which tool, and what it costsFour product shapes exist — the chat window, the assistant inside your files, the transcriber and the code runner — and using the wrong shape for a task looks like a model failure when it is a category error.
Case studyTwelve seats, or one stopwatch
Vikram Sethi came back from a franchise meeting holding a quotation for twelve seats on an AI assistant at roughly two thousand four hundred rupees a seat a month. Annualised, that is a little under three and a half lakh, which at Sunrise is one junior teacher.
He had two tasks in mind, and he was clear about which one he wanted.
The first was the weekly parent letter. Forty of them, one per batch, mostly the same: what was covered, what the test dates are, who has fallen behind on attendance. A senior teacher writes them on Saturday evening and complains about it every Saturday evening.
The second was the one that had sold him the idea. After every fortnightly test the centre sends each family a short note explaining why their child's rank moved. Four hundred notes, twice a month, written by whichever teacher can be cornered. Everyone hates them and parents read every word. If the assistant could write those, Sunrise would get a Sunday back and, Vikram thought, would look more professional than the two coaching centres on the same road.
His operations head, Anjali Rao, asked him to hold the purchase order for two weeks and answer three questions about each task instead.
Where does the truth live? For the parent letter, in the syllabus tracker and the attendance register — documents Sunrise holds and can paste in. For the rank note, the arithmetic of the rank is in the test data, but the reason is not. The reason is that the boy missed nine days with dengue, that the girl has been put in a batch that moves faster than she does, that a whole cohort was thrown by a badly set question in section B. None of that is in any file. It is in the teachers' heads and in the corridor.
What does a wrong answer cost? A clumsy sentence in a weekly letter costs a shrug. A wrong explanation of why a child's rank fell costs a meeting with a father who has just been told his daughter is not working hard enough, when in fact she was ill and the centre knew it.
What does checking cost? Here the two tasks separate completely. A parent letter is checked by reading it, which takes ninety seconds and is something the teacher would have done anyway. A generated rank explanation cannot be checked by reading it, because a plausible reason and the real reason look identical on the page. Checking one means going to the attendance register, the batch record and the teacher — which is longer than writing the note from those things in the first place.
That left Vikram with a decision he did not like. Buying twelve seats and starting with the rank notes would take the most hated job off his staff immediately, and it was the job that had made the tool look worth paying for. Refusing to start there meant telling eleven people that the thing they were promised was not coming, and one of his best teachers had already said she would not do another term of Sunday nights.
The other side of it was not hypothetical either. Sunrise's whole business is that parents believe the centre knows their child. A single note explaining a rank drop by a reason the centre had invented would be worth more damage than a year of Sunday evenings.
What actually happened
Vikram bought two seats rather than twelve and gave himself a term to decide about the rest. The weekly parent letter went into the hand-over pile: a brief with the syllabus tracker and the attendance extract pasted in, a ninety-second read before sending. Timed over six weeks it fell from about thirty-four minutes to eleven, including the check. The rank note went into the assist pile with the direction fixed the other way round — the teacher writes three bullets of their own judgement, and the model turns them into a paragraph at the right register in the family's preferred language. Sunrise added one rule to the brief: the note may not contain a reason that is not in the bullets. In the second batch a teacher approved a note that explained a drop by "reduced practice at home", which the model had supplied and nobody had said; it was caught by the rule and not by anybody's vigilance, which is why the rule exists. At the end of the term Vikram bought six more seats, not ten, because the audit had found only eight people whose week contained a task that actually moved.
Worth arguing about
Anjali's objection to the rank notes was not that the model would write badly. What was it?
One answer
That the truth for a rank note lives outside any document Sunrise could supply. The model can compute nothing about why a rank moved, so it produces the reason such notes usually give — more practice, less distraction, exam nerves — in exactly the register a good teacher would use. The failure is invisible on the page, and checking it means consulting the attendance register, the batch record and the teacher, which is the work the note was supposed to save.
Sunrise reversed the direction of the rank-note task rather than abandoning it. Why does that reversal make it safe?
One answer
Because it moves the judgement to the person who has it and leaves the model only the phrasing. The teacher supplies the cause in three bullets; the model expands them. Every fact in the finished note came from somebody who knew it, so there is nothing for the model to invent, and the checking cost drops from consulting three records to reading a paragraph against the bullets you just wrote.
Sunrise added a rule that a note may not contain a reason absent from the bullets, and it caught an error that no amount of care had caught. What does that tell you about relying on staff vigilance?
One answer
That vigilance degrades exactly where the tool is usually right. A teacher reading a fluent, plausible note about their own pupil has no signal that one clause was supplied rather than reported. A mechanical rule — nothing in the output that was not in the input — does not depend on anyone feeling alert on a particular evening, which is the only kind of check that survives a busy fortnight.
Test yourself6 questions on this module
You attach an eight-page circular and ask for a summary. In a separate chat with nothing attached, you ask for your state's current stamp duty rate. Both replies are equally fluent. Why is only one of them cheap to rely on?
A colleague finds that the same paid tool costs her roughly three times as much per document in Tamil as in English, for documents that say the same thing. What is going on?
Twenty messages into a long conversation, the model stops obeying an instruction you gave carefully at the start. What has most likely happened, and what is the cheapest fix?
A solicitor asks for a summary of a forty-page expert report, gets a good one in ninety seconds, and then reads all forty pages anyway because she will rely on it at a hearing. What is the right change to make?
Sixteen experienced developers expected AI help to make them about a quarter faster, reported afterwards that it had made them about a fifth faster, and were measured as about a fifth slower. What should you take from this about your own work?
Your team's answer to automation bias is that a second colleague reads every generated draft before it goes out. Why does that do much less than it appears to?
Module 2
Producing text you would put your name on
Most of what a working day contains is writing something to somebody. This block turns a request into a brief, teaches the one prompting technique that outperforms every adjective, and covers the two document jobs the first version of this course skipped — writing for readers who are not you, and building a report or a deck.
By the end you can
Turn a real work situation into a brief with audience, constraint, format and two samples of your own writing, then edit the result to a standard you would sign and name the two changes that made it yours.
- 10Write a brief, not a requestA prompt that names the audience, the constraint, the format and what a wrong version would look like outperforms any list of style adjectives — because you are supplying the situation the model cannot see.
- 11Drafting, rewriting, and the tone problemDescribe the situation, not the deliverable — then rewrite your own words rather than accepting generated ones.
- 12Examples beat adjectivesThree of your own past documents pasted in as examples control tone, structure and vocabulary far more precisely than any description of your style, because you are demonstrating the pattern instead of naming it.
- 13Editing a draft instead of rolling the dice againRegenerating re-rolls the same prompt without recording your objection, so progress comes from instructions that name what to keep, what to change and what the change is — and from resetting the anchor to a clean base text once edits start undoing each other.
- 14Refusals, bad news and the message you would rather not sendPreference tuning pushes the default register towards warmth and hedging, which is the exact failure mode for a refusal — so the decision goes in sentence one and the prohibitions do as much work in the brief as the instructions.
- 15Summarising and the missing sentenceA summary's real risk is omission, not error — so name what must appear and demand it say when something is absent.
- 16Writing for readers who are not youAI is unusually good at moving a document between registers and languages, and the only safe way to use it is to check the translation backwards and to keep numbers, names and negations under your own eye.
- 17Reports, decks and the structure problemDecide the argument and its order yourself, then let the tool expand and format it — because a deck generated from a topic produces balanced-looking slides with no position in them.
- 18Making it attack your work instead of writing itCriticism of text you supplied is the lowest-risk mode there is, but only if you refuse the yes/no question and demand a fixed number of quoted objections, because agreement is what preference training rewards.
Case studyNinety-six flats and a lift that is not coming
Sushma Kadam has been honorary secretary of Shreeniwas Co-operative Housing Society for three years. In March the committee levied twelve thousand rupees per flat towards replacing both lifts. In July the contractor told her the second lift would not be commissioned until November, and that the levy would not be reduced, because the money had already been committed to the order.
She had to write to ninety-six households, several of whom are elderly and live on the sixth floor.
Her first draft came out of a chat window in about forty seconds, and she nearly sent it. It opened by thanking members for their continued cooperation and patience. The third paragraph said the committee "may be unable to complete the second lift installation within the originally anticipated timeframe" and that "the levy position is being reviewed in light of ongoing developments".
Every sentence in it was defensible. It was also, read as a member would read it, a letter saying that the lift might be late and the money might come back.
Sushma is not a professional writer, and the letter had two things working against her. She did not want to be the person who wrote the blunt version, because she has to stand next to these people at the lift lobby every morning. And she genuinely did not know whether to apologise, because a previous committee had been accused of admitting fault in a circular during a dispute with a contractor, and the phrase "we accept the delay is our failure" had been quoted back at them for a year.
So the decision in front of her was not about tone. It was between two costs.
Send the soft version and this week is quiet. Nobody shouts at her on Sunday. But "may be unable" produces a fortnight of individual questions, then a second letter when November is confirmed, and by then members will reasonably feel they were managed rather than told. A refusal that is not understood as a refusal is a cost passed to the person who has to explain it a second time, and that person is her.
Send the clear version and Sunday is bad. The levy is not coming back, the lift is November, and forty families on the upper floors will have something to say about it. Her committee colleagues, two of whom wanted the soft draft, will spend a week being asked why the letter was so harsh.
What she did with the tool in between is the part worth copying. She stopped asking for a letter and started describing the situation: ninety-six flats, four months, no refund, a building where the majority of upper-floor residents are over sixty, and a committee that must not concede fault while the contractor dispute is open. She named the reader, the length, and the outcome she wanted — that people know the date and stop asking whether the money is coming back. She wrote down what a bad version looks like: no thanking anybody for their patience, no "in light of ongoing developments", no invitation to write in with concerns that she cannot answer individually.
Then she pasted in three of her own past circulars, including the one about the water tank cleaning that people had said was clear, and asked for the new letter in that voice.
What actually happened
The briefed draft put the decision in the first sentence — the second lift will be commissioned in November and the levy will not be refunded — gave one reason, gave the date, and offered the one thing the committee could actually offer, which was a priority repair contract on the working lift and a chair placed in the ground-floor lobby. It was four sentences longer than Sushma wanted and she cut two of them by hand, including a line of consolation the model had added back after being told not to. On the advice of a member who is a lawyer she removed the word "apologise" and kept "we are sorry that residents on the upper floors will carry this for four more months", which says the same thing about people and nothing about liability; she noted in the minutes that this was a committee decision and not her drafting preference. The letter produced eleven replies in the first week and no second letter in November. Two committee members told her it had been too direct. One of them, in October, asked her to write the diesel generator circular the same way.
Worth arguing about
The first draft was grammatically perfect and professionally phrased. Why was it the wrong letter?
One answer
Because the register these models default to — warm, hedging, non-committal — is the specific failure mode for bad news. "May be unable" and "is being reviewed" are read by a member as uncertainty about the outcome, which means the letter has not delivered the decision it exists to deliver. The cost lands later, as individual questions and a second circular, and by then members feel they were handled rather than told.
Sushma spent longer on the brief than the first draft took to produce. What did the extra time actually buy?
One answer
It supplied the things the model could not see: that the majority of affected residents are elderly and on upper floors, that the contractor dispute makes an admission of fault dangerous, that the outcome is people knowing the date and stopping the refund question, and that specific phrases are banned. Those are facts about her situation, and the prohibitions did as much work as the instructions, because left alone the model adds every one of them back.
Why did she paste in three of her own past circulars instead of asking for a "clear, warm, professional" tone?
One answer
Because adjectives about style are averages of millions of documents somebody once labelled that way, and they produce the generic register rather than hers. Three real examples carry sentence length, how she opens, whether she uses headings, how she handles bad news and what she conspicuously never says — none of which she could have written down as rules. The one risk is facts leaking across from the examples, which is why the examples were circulars with no individual named in them.
Test yourself6 questions on this module
Two prompts for the same overdue-invoice reminder. One says "polite but firm". The other names a six-year client, forty-seven days late, a director you know personally, a third and final reminder, under 150 words, no threats. Why does the second work so much better?
You want a report written in your department's house style, which nobody has ever documented. What gets you closest, fastest?
A draft is nearly right. You press regenerate four times and get four differently wrong drafts. What is the mechanism, and what should you do instead?
A refusal drafted by a model comes back saying you "may be unable to proceed at the present time". Why is that the characteristic failure rather than an unlucky phrasing?
You summarise a forty-page tenancy agreement and get eight accurate bullet points. Three weeks later you discover an auto-renewal clause that was not in them. Which instruction would most likely have caught it?
You have translated a consent form into a language you cannot read. Which check catches the errors that matter most, and what does it still miss?
Module 3
Getting more out of it than a first draft
The gap between somebody who finds AI mildly useful and somebody who finds it transformative is not a better tool. It is six or seven techniques: cutting a job into steps you drive yourself, teaching a boundary with examples rather than adjectives, pinning the output to a shape you can check, making it commit to a plan you can reject, and knowing which model to reach for. This block is the craft.
By the end you can
Take a task too large for one prompt, decompose it into steps whose intermediate results you can inspect, pin each step's output to a fixed shape with an explicit absence marker, and use disagreement between repeated runs to locate the parts of the answer that need your attention
- 19Cutting a job into steps you can inspectConstraints in a single prompt are soft influences that dilute each other, so a job containing several kinds of judgement becomes several calls whose intermediate results you keep — which is what turns a wrong answer into a diagnosable one.
- 20Teaching it a boundary it cannot guessCategory names carry almost no information and boundary examples carry all of it, so a sorting prompt is eight to twelve edge cases plus an explicit escape route — because a model offered only three labels will return one of three labels for everything.
- 21Fixing the shape of the answerA demonstrated output format with an explicit NOT FOUND marker and a row count turns a batch of prose into something you can sort, paste and audit — and makes the missing row, the commonest real failure, visible.
- 22Make it commit to a plan you can rejectAn approved plan is cheap to reject and then conditions everything generated after it — but the written reasoning is text produced the same way as the answer, so it steers the output without explaining it.
- 23"You are an expert lawyer" and what it really doesA persona changes register without adding knowledge, which raises the apparent authority of an answer while leaving its accuracy where it was — so keep the genre or format specification and drop the biography.
- 24Reasoning models, and when they earn their costA reasoning model buys revision across interacting constraints and buys nothing on tone or drafting, and no amount of deliberation recovers knowledge the model never had — it only makes a missing fact arrive more elaborately.
- 25Asking three times, and reading the disagreementDivergence between fresh runs of the same prompt marks where the model was effectively choosing, and is the only uncertainty signal an ordinary user gets — but convergence proves stability, never truth.
- 26Screenshots, scans and photographs of paperA vision model reads labels and proportions rather than measuring, so any number not printed in the image comes back as a confident estimate — and the frame and its metadata disclose far more than the part you were looking at.
- 27Dictation, voice mode and who pays for the errorsSpeaking the substance and having the model organise it is both the fastest and the safest configuration, because every fact came from you — but spoken answers remove the rereading that catches errors, so voice belongs on the input side.
Case studySix hundred and forty answers, two days
The question on the survey form was deliberately open: what stopped you from using the credit group this year? Six hundred and forty women answered it, in Hindi, in Santali, and in the mixture of both that the field staff write in their notebooks. All of it had been typed up by two interns into a single spreadsheet column.
The donor report was due Friday. It needed four to six themes, a count against each, and a quotation for each theme. Priya Toppo, who runs monitoring and evaluation, had Wednesday and Thursday.
Her first attempt was one prompt. It named the column, asked for the main themes ranked by frequency, the best quotation for each, a count, and a nine-hundred-word section for the donor, in British English, without mentioning the interest rate because that was being handled separately.
It returned something good-looking in ninety seconds: five themes, five quotations, a smooth section of about eleven hundred words. She read it twice before she noticed that one of the quotations did not appear anywhere in the spreadsheet, that the counts added to 640 exactly — which no real classification of free text ever does — and that the interest rate appeared in paragraph three.
She now had a decision, and both branches cost her something real.
One option was to keep the artefact and repair it. Find the missing quotation, correct the counts by hand, delete the interest-rate paragraph. That was maybe three hours, it fitted inside Thursday, and it would produce a report the donor would accept. The problem she could not get past was that she did not know which of the five themes were real. If a quotation had been invented, the grouping underneath it might have been invented too, and there was no way to tell from the output — the wrong parts of it looked exactly like the right parts.
The other option was to throw the whole thing away on Wednesday afternoon and do it in steps. Extract one line per response. Group the lines. Read the groups herself. Pull quotations from named rows. Only then draft. That is four passes over 640 rows, and she estimated Wednesday evening and most of Thursday. It also meant telling her director on Wednesday that the report he had seen a draft of no longer existed.
The thing that decided it was not quality. It was that the second version could be checked. If the extraction produced 640 lines, she knew nothing had been dropped. If it produced 631, she knew exactly where to look. With the single prompt she had one object and no way to interrogate it: a wrong number and a right number arrive in the same font.
She took the second option, and she added two things the first attempt had not had. Every step had to output a fixed number of rows and end with a count. And the grouping step was given an explicit escape — any answer that did not clearly belong in a category was to be labelled UNCLEAR with the phrase that made it ambiguous, rather than pushed into the nearest theme.
What actually happened
The extraction step, run in batches of twenty, returned 640 lines on the second attempt; the first attempt returned 617, and the missing rows were all in one block where the interns had left the cell blank rather than writing anything, which was itself a finding. The grouping produced seven categories and 74 answers in UNCLEAR. Priya read all 74 — it took fifty minutes — and found two things nobody had anticipated: a group of women who had stopped because the meeting had been moved to a day that clashed with the ration shop, and a smaller group who had not stopped at all and had misread the question. Both went into the report, and the ration-shop finding changed the timing of meetings in three blocks. The final counts did not add to 640 and the report said so. She finished at ten o'clock on Thursday night, roughly the time the repair option would have finished on Thursday afternoon, and the section she wrote was 780 words with every quotation carrying a row number.
Worth arguing about
Priya could have repaired the single-prompt output in less time than the decomposition took. Why was repairing it the worse option?
One answer
Because repair fixes what you can see. The invented quotation, the too-tidy counts and the banned topic were the visible defects; the invisible question was whether the five themes reflected the answers at all. With one artefact there is no way to find out — you cannot tell whether the extraction missed responses, whether the grouping was crude, or whether the drafting drifted. Decomposing produces intermediate results you can count and read, which turns a wrong answer into a diagnosable one.
The UNCLEAR category was the most productive part of the run. Explain why a forced choice between named themes would have hidden both findings.
One answer
A model produces a continuation, and if the only continuations available are the theme names, every answer receives a theme name — including the ones that belong to a category nobody defined and the ones that answer a different question. The two discoveries, the ration-shop clash and the misread question, would have been distributed silently across the seven existing themes, inflating them slightly and disappearing. The escape route is what makes an unmodelled boundary visible.
What did the count at the end of each step protect against, and why is a missing row more dangerous than a wrong value?
One answer
It protects against silent loss. Six hundred and seventeen well-formed lines look exactly as convincing as 640, so nothing on the page announces the gap, and every downstream count inherits it. A wrong value is at least visible to somebody who knows the material; an absent row is visible only to arithmetic. Comparing what went in with what came out is the cheapest check available in batch work and it catches more real errors than any refinement of the prompt.
Test yourself6 questions on this module
One prompt asks a model to read thirty emails, find the themes, rank them, pick quotations, and write an 800-word report in British English without mentioning refunds. It returns 1,100 words, four themes, one quotation that appears nowhere, and the refund policy. What is the mechanism?
You are sorting support tickets into billing, technical and account. What should the bulk of your prompt contain?
In a batch extraction over forty invoices, which pair of instructions does most to make the real failures visible?
Asking for a plan before the draft has two benefits. One is that a wrong plan is cheap to reject. What is the other, and what must you not conclude from the plan?
"You are an expert employment lawyer with twenty years' experience." What does that prefix actually do?
You ask the same question in three fresh conversations and get 30 days, 30 days and 90 days. Then you ask a different question three times and get the same answer three times. What have you learned in each case?
Module 4
Working from your own material
The moment you supply the source, the failure mode changes from invention to misreading — and misreading is something your professional judgement already catches. This block covers long documents, messy spreadsheets, recorded meetings and searched sources, and the citation discipline that holds all four together.
By the end you can
Answer a real work question entirely from sources you supplied, with a quoted line and a locating reference behind every claim, and treat any claim without one as absent rather than as probably true.
- 28Working with your own documentsAttaching your documents puts the truth in the text — but demand quoted citations, because general knowledge blends in invisibly.
- 29What "chat with your documents" is really doingThe model sees only the handful of chunks a similarity search returned, so exact identifiers, split rules and every counting question fail structurally — and a document that fits in the context window should be pasted whole rather than retrieved from.
- 30Long documents, and what gets lost in the middleA long document is not read evenly — material at the start and end is retrieved far more reliably than material in the middle — so ask located questions section by section rather than one question of the whole file.
- 31Before you analyse it, find out what you haveAsk what one row is, how missing values are represented, and what each column means before any calculation — because a sentinel like -999 or a join that multiplied the rows produces a confident, precise and arbitrary answer.
- 32Spreadsheets and data, without being a data personNever ask AI for the answer to a calculation; ask for the formula, then test it on rows you can check by hand.
- 33Cleaning a spreadsheet somebody else builtCleaning is where analysis is silently won or lost, so every fix runs in a new column beside the original and the count of changed rows gets checked before the original is touched.
- 34Charts, and the choices you did not makeWrite the one sentence the reader should leave with, ask for the chart that makes it visible, then read the chart back as a sentence — and prefer generated code over a generated image, because code shows which column was actually plotted.
- 35Meetings and emailFrom a meeting, extract decisions, owners and deadlines — never the discussion, and never let AI hit send.
- 36Recording, transcription and the consent questionTranscription fails by inventing fluent sentences during silence and by attributing speech to the wrong person, and in much of the world you may not record at all without everyone's agreement — so ask first and check the transcript against the audio at the timestamps that matter.
- 37Research, search and the citation that does not say thatA search-connected tool replaces invented sources with real links attached to claims the pages may not make, so the check moves from 'does this exist' to 'open it and find the sentence'.
Case studySeventeen contracts, or maybe not seventeen
Steel and cement had moved enough over eight months for the board to ask a simple-sounding question: how many of our subcontracts allow the subcontractor to pass a price increase on to us, and what is our exposure if they all do it at once?
Rakesh Vyas, the commercial manager, had exactly the tool for this, or so it seemed. All sixty agreements had been uploaded to a document assistant in March. He typed: how many of these subcontracts contain a price-escalation clause?
The answer came back in four seconds. Seventeen. It listed them, with a short description of each clause, and it read like something a competent junior would have produced in a day.
Rakesh had used the tool for months and had been pleased with it. What made him stop was something a colleague had mentioned in passing: that these systems answer from a handful of passages they retrieve, not from the whole library. He opened the sources panel. Six passages had been retrieved. Six, from sixty documents.
That reframed the number completely. Seventeen was not a count of sixty contracts. It was a description of whatever six chunks of text the search had returned, extended into a confident total. It might be right. It might be nine, or thirty-one. Nothing in the answer distinguished those possibilities, and the list of seventeen names looked precisely as authoritative either way.
The decision was now a real one, with the board meeting on Monday.
He could take the seventeen. It was defensible in the ordinary office sense — it came from the firm's own contracts, through a tool the firm had bought, and if anybody asked he could show the answer on the screen. It would take him twenty minutes to turn into a slide, and the exposure figure that flowed from it would look reassuring.
Or he could do the full pass. Sixty contracts, one at a time, each asked a located question — in the pricing and variations sections, quote any sentence permitting the subcontractor to increase rates, or write NOT PRESENT — with the answers collected in a spreadsheet and every quoted sentence checked against the document it came from. He estimated a day and a half of his own time or two days of a graduate's, in a week where he did not have a day and a half.
The cost on the second side was not only time. Asking for two days of somebody's week meant telling his director that a tool the firm had paid for could not answer the question it had been bought to answer, which is not a comfortable conversation to have about a decision your director signed off. It also meant that the March upload, which had been presented internally as the end of the contracts problem, would have to be described as the beginning of it.
The cost on the first side was that a materially wrong exposure number would go into board minutes with his name against it, and would sit there until the day somebody tried to use it. Exposure figures are not read on the day they are approved. They are read eighteen months later, by a lender or an auditor or a new finance director, at which point the question is not whether the tool was reasonable but whether anybody counted.
He tried one thing before deciding. He asked the same question three more times, in three fresh conversations, phrased differently each time: how many subcontracts permit a rate increase, which subcontracts allow the subcontractor to revise rates, and list every agreement with an indexation or escalation provision. He got seventeen, twenty-two and fourteen. Nothing about the wording of any of the three answers suggested uncertainty. The disagreement did not tell him the right number, and it told him something he could act on: the retrieval was unstable, so no single run of it was going to be the answer.
What actually happened
Rakesh ran the full pass. It took a graduate surveyor eleven hours across two days, using a fixed output shape — file name, clause number, quoted sentence, or NOT PRESENT — and a count at the end of each batch of ten. The real figure was twenty-nine, not seventeen. Eight of the twelve additional contracts used the phrase "rate revision" rather than "escalation", which is why a meaning-based search had ranked them low against his wording, and two more were in scanned PDFs with no text layer that the index had skipped entirely. Three of the original seventeen turned out to carry a cap the summary had not mentioned. Rakesh took both numbers to the board, explained the difference in four sentences, and the firm now runs a quarterly full pass over the portfolio with the spreadsheet kept as the record. The document assistant is still in use, for finding rather than for counting.
Worth arguing about
Why is "how many of these contracts contain X?" a question document chat cannot answer, however good the model is?
One answer
Because the system retrieves a small number of passages — often three to ten — and the model answers from those alone. A count requires every document to have been examined, and nothing in the retrieval step does that. The model is not lying when it says seventeen; it is describing what it was shown, and the interface gives no signal that it was shown six chunks out of sixty. Aggregate questions of any kind, including list-all and which-ones-do-not, are structurally outside what retrieval can do.
Two of the missing contracts were scanned PDFs. What general lesson does that carry about a negative answer from a document tool?
One answer
That a tool finding nothing tells you about its index rather than about your organisation. A scan is a photograph of a page with no text in it, so unless OCR was run it was never searchable at all. Files without a connector, material in personal drives and archives that were never indexed fail the same way. This is why absence is the weakest thing such a tool can report, and why the check is to try to select text in the PDF before trusting any search over it.
The full pass used a fixed output shape with NOT PRESENT and a count per batch. What would have gone wrong without those two rules?
One answer
Without an explicit absence marker, every row gets filled, because a table where one cell is blank is an unlikely continuation of a table where every other row is complete — so contracts with no escalation clause would have acquired plausible ones. Without a count, a batch of ten that returned nine would have passed unnoticed, and the final total would have been quietly short. Together they make the two commonest failures of batch extraction visible rather than invisible.
Test yourself6 questions on this module
You attach your staff handbook and ask about notice periods. The answer is about eighty per cent handbook and twenty per cent general employment practice, merged into one paragraph. Why is this worse than an obvious error, and what is the structural fix?
A document tool over 400 contracts is asked how many contain an indemnity clause and answers "nineteen". Why is the number meaningless, and which questions share the defect?
You ask whether a 200-page manual mentions a particular obligation and are told no. How should you treat that answer?
A colleague sends an "orders" export. You total the order-value column and the figure looks plausible but high. What is the ten-second check that most often explains it?
You want a total of unpaid invoices over thirty days old from a 300-row sheet. What is the reliable move, and why?
A search-connected tool gives you a claim about a filing threshold with a link to a real page from a credible body. What is the failure this configuration introduces, and what is the check?
Module 5
The tools already on your desk
Most people will meet AI not in a chat window but as a feature inside the software they already open every morning: the mail client, the document editor, the search box, the notebook, the merge. Those features have their own failure modes — permissions they inherit rather than intend, searches that cannot prove absence, batches that silently lose a row — and their own free equivalents for anyone without a corporate licence.
By the end you can
Set up one recurring office job — a mail triage, a document search, a two-hundred-letter merge — so that the generated part is small and checkable, the fixed part is human-written and checked once, and every batch carries a mechanical validation that a missing or unfilled field cannot survive
- 38The assistant that reads your files for youAn in-suite assistant inherits your permissions rather than your intentions, so it makes reachable by ordinary question everything that was previously protected only by nobody knowing it was there.
- 39Triage, summarise, draft — and never sendClassification is advice and must never move a message out of sight, expensive categories are deliberately over-flagged, and a once-a-day digest you read attentively beats forty suggestions you approve while thinking about something else.
- 40Finding the document, and what search cannot tell youKeyword and semantic search fail in opposite directions, no search proves absence, and ranking answers which document matches your words rather than which one is currently in force — so the date check is a habit, not a setting.
- 41Notebooks that answer only from your sourcesSource-bound tools convert invention into misreading, which your judgement can catch — but the grounding is a design rather than a proof, so a real citation attached to a claim it does not quite support is the characteristic error.
- 42PDFs, scans and the black rectangle that hides nothingTry to select the text first: a born-digital PDF can be extracted exactly with no model involved, a scan needs OCR whose errors look like errors — and a black box drawn over text leaves the text in the file.
- 43Two hundred letters, and only one generated paragraphFixed human-written text plus mechanical merge fields plus at most one generated sentence turns two hundred documents to check into one document checked once and two hundred short paragraphs you can sample — with every outlier read and a search for unfilled placeholders run over the whole batch.
- 44The AI feature that arrived in an updateA default-on feature in existing software is now the commonest way an organisation acquires a new recipient of its data without deciding to, so the sub-processor, retention, billing and off-switch questions are asked in writing before the evaluation on twenty real items.
- 45Being able to explain a document six months laterKeep the inputs, the instruction and the named human who checked it — and keep the output itself, because models and sampling change and rerunning a saved prompt does not reproduce what was sent.
- 46The whole free stack, and what it costs youA local eight-billion model is close to a frontier model on rewriting, summarising supplied text, classification and extraction, and meaningfully worse on hard reasoning and broad knowledge — which happens to match the split between tasks this course endorses and tasks it warns about.
Case studyTwo hundred and fourteen renewal letters
Every March, Anand Menon's firm sends renewal letters to about two hundred commercial motor clients. Each letter carries the client's name and address, the policy number, last year's premium, this year's premium, and one sentence explaining why the premium moved — a claim, a change in the fleet, a change in the insurer's rating.
The letters have always been produced by copying last year's document and editing it. It takes the best part of a week and the same three mistakes appear every year.
In January the new office assistant, Fathima, showed Anand something that made the week look unnecessary. She had pasted the whole client spreadsheet into a chat window and asked it to write all the letters. It had produced them. They were fluent, correctly formatted, individually phrased, and there were two hundred and fourteen of them.
Anand read four at random and they were right. That is where the decision started.
Generating all of them cost one afternoon and produced two hundred and fourteen documents that each had to be checked, because each one had been written separately and could therefore be wrong separately. Every premium figure, every policy number, every name. Checking two hundred and fourteen letters properly is not much less work than the week they were trying to avoid, and Anand knew from experience that nobody checks the two hundredth letter with the attention they gave the fourth.
The alternative was slower to build. One template, written once by a person, containing the fixed text — the renewal terms, the payment options, the cancellation wording, the regulator's required disclosures. Merge fields pulling name, address, policy number and both premium figures directly from the spreadsheet, mechanically, with no model involved. And exactly one generated sentence per letter: the reason the premium moved, written in twenty-five words from three named fields.
That meant a day and a half of setting up a mail merge in software the firm already had, plus arguing with Fathima, who had already done the job and could not see why her version was worse. Her objection was reasonable and she made it plainly: the merge letters would all sound the same. Anand's answer was that they are the same, that clients compare them, and that two clients discovering they had been told slightly different things about the same rating change is a conversation he did not want.
The cost on his side of the argument was real. A day and a half in January is a day and a half in a four-person firm, the firm had no mail merge template and nobody in it had built one, and if the March run went wrong because of something he had built, it was his. Fathima's version, whatever its risks, existed already and had cost the firm one afternoon.
There was a second thing pulling him towards her version, and he only noticed it when he tried to write down why he was resisting. Her letters were better written than the template. Each one read as though somebody had thought about that client. The template would read like a template, because it is one, and a broker whose whole business is relationships does not enjoy sending two hundred identical documents.
What settled it was a question he asked himself in the language of the audit rather than the language of quality: not which version is better, but what does checking each one cost. Fathima's version required him to verify two hundred and fourteen premiums, two hundred and fourteen policy numbers and two hundred and fourteen names, because each had been produced separately and could be wrong separately. The merge required him to verify one template, one spreadsheet column against the insurer's schedule, and two hundred and fourteen short sentences he could sample. Those are not the same afternoon.
What actually happened
They built the merge. The fixed text was read once by Anand and once by the firm's compliance adviser, and has not been re-read since. The premium figures came from the spreadsheet and were validated by summing the premium column and comparing it with the insurer's schedule, which reconciled to the rupee. The generated sentences — two hundred and fourteen of them, at twenty-five words each — were briefed to output the word BLANK when the reason code was empty, and eleven came back BLANK, which turned out to be eleven policies where nobody had recorded why the premium had changed. Anand read every letter for a client with an unusual field: four negative adjustments, one client with a fleet of ninety-one vehicles, two addresses outside Kerala, and the eleven BLANKs. He sampled twenty of the ordinary middle. Before printing, Fathima ran a search across the whole batch for angle brackets, double spaces and the string "NOT", and found one letter still reading "Dear <ClientName>" because a row had a trailing space in the name field. The run took two days instead of five and the same two days in the following March took four hours.
Worth arguing about
Fathima's version produced finished letters in an afternoon. Explain, in terms of checking rather than quality, why Anand rejected it.
One answer
Because independently generated documents have to be independently checked. Two hundred and fourteen letters written separately are two hundred and fourteen chances for a wrong premium or a wrong policy number, and no amount of reading the first few tells you about the two hundredth. The merge architecture collapses that: the risky legal text exists once and was checked once, the figures are mechanical and reconcile arithmetically, and only the short generated sentences remain — which can be sampled rather than read in full.
Somebody always suggests varying the phrasing so the letters do not look templated. Give the two reasons Anand refused.
One answer
First, varied wording is the mechanism by which a factual difference creeps in between letters that were meant to say the same thing, and recipients do compare. Second, it is mildly deceptive: the letters are templated, everybody knows they are, and disguising it buys nothing. Variation should come from the data, because the data is what actually differs between recipients.
The batch search found one letter reading "Dear <ClientName>". Why is a mechanical search over the whole batch worth more here than reading a larger sample?
One answer
Because the failure is rare, uniform and invisible to sampling. One letter in two hundred and fourteen will not appear in a sample of twenty, and it is the letter that makes the firm look careless to exactly one client. A search for angle brackets, placeholder braces, the word NOT and doubled spaces takes seconds, covers every document, and cannot be defeated by fatigue. The count check — two hundred and fourteen records in, two hundred and fourteen documents out — is the same kind of protection.
Test yourself6 questions on this module
An in-suite assistant answers a question about salary bands, correctly and with a citation, from a planning file you were always permitted to open. What has actually happened?
Why should an AI triage rule apply labels and leave every message in the inbox, rather than filing messages into folders?
You search your document store for "annual leave" and find nothing. Which conclusion is sound?
A source-bound notebook tool cites a real passage for every claim. What is the characteristic error that survives, and what does it imply about curation?
You need the figures off a supplier invoice that arrived as a PDF. What do you do first, and why does it decide everything else?
Your practice management software has shipped an AI summary feature, enabled by default, in a routine update. What is the first thing to establish, and why that first?
Module 6
The lines you do not cross
This is the block that protects your registration, your job and the people whose information you hold. It covers what the account tier actually changes in law, which rules bind you regardless of what the provider does, the special danger of AI in decisions about people, the free path when nothing may leave your machine, and what to do when your workplace has no policy at all.
By the end you can
Decide, for any document in front of you, whether it may go into a given AI tool — naming the account tier, the setting and the professional or legal duty that makes the decision, and naming what you would use instead.
- 47What must never go into the boxPasting into a chatbot is disclosure to an outside party — your duty of confidentiality is breached at the moment of sending.
- 48Which account are you using, and what it changesThe consumer, business and enterprise tiers of the same product differ in contract, retention and training defaults rather than in intelligence, and the tier — not the model — decides what may be typed into it.
- 49The rules that already apply to youNo new AI law is needed to make most workplace misuse unlawful — data protection, confidentiality and sector regulation already bind you, and the AI-specific rules add duties around transparency and decisions about people.
- 50Redaction that survives somebody tryingNames are rarely the strongest identifier — a combination of ordinary details usually is — so anonymity comes from generalising the combination, and redaction must happen before anything leaves your machine, never by asking an online model to do it.
- 51When the decision is about a personA model trained on past decisions reproduces the pattern in them, so using it to sift people automates yesterday's discrimination at scale and hides it behind an appearance of neutrality.
- 52Copyright, ownership, and everything that is unresolvedWhether training was lawful, whether an output infringes, and whether anyone owns the output are three separate questions with different answers in different countries — and only the second and third normally reach your desk.
- 53Running a model on your own machineAn open-weight model running locally sends nothing anywhere, which makes it the honest answer for confidential text — at the price of a real and knowable quality gap on hard reasoning.
- 54Disclosure: telling people you used itDisclose when the reader is paying for or relying on your personal judgement or authorship, not when the tool merely helped you produce something you fully checked and stand behind.
- 55When your workplace has no policyIn the absence of a policy the safe default is not abstinence but a written question and a one-page rule you propose yourself, because unwritten practice is what an incident is judged against.
Case studyThe settlement draft at eleven at night
Arun had been given a forty-page draft settlement agreement at six in the evening and asked for a note on it by nine the following morning. At about eleven he was tired, the clauses were repetitive, and the firm's approved research tool did not handle attachments. He opened the free chat account on his own laptop, pasted the whole agreement in, and asked for a summary of the payment and release provisions.
It was good. He used it to orient himself, wrote the note himself from the document, and went to bed. He told nobody, not because he was hiding anything, but because it had not occurred to him that anything had happened.
On Tuesday he mentioned it to a colleague over coffee, in the tone of someone sharing a useful trick. The colleague went quiet, and then told the supervising partner, Lakshmi.
What Lakshmi then had in front of her was not a question about whether the summary was accurate. The disclosure had already happened at the moment the paste key was pressed. A document belonging to a client, containing the client's name, the counterparty's name, the settlement figure and the confidentiality clause the parties had negotiated over, had been sent to a company the firm has no contract with, on terms an employee accepted on a phone, on a tier where conversations are retained and may be used to improve the service.
She had two options and both of them cost the firm something it valued.
Telling the client meant telling them that their settlement terms had left the firm's control. The client is a family business the firm has acted for since 1998. The settlement was commercially sensitive and the confidentiality clause was the point of the whole negotiation. There was a real chance of losing the relationship, and a real chance of a complaint to the regulator, which the firm would have to report and which would follow it for years.
Not telling the client meant deciding, on a Tuesday morning, that a breach of confidence in a matter about confidentiality was a thing the firm would keep to itself. Lakshmi could see how that decision would look if it ever surfaced — and things surface, through colleagues, through a leaver, through a disclosure exercise. It would also mean saying nothing to Arun that would change what anybody else in the firm did next month.
There was a third pressure she noticed and named to herself. Arun had done this because the approved tool was worse than the free one and there was a deadline. Six other associates had the same deadline pressure and the same laptop. Whatever she decided about the client, the firm had a supply problem, not a discipline problem.
She also had to decide what to do about Arun, and the two decisions pulled against each other. Making an example of him would satisfy the partners who wanted to see something happen, and it would guarantee that the next associate in the same position at eleven at night told nobody. She had watched a previous firm handle a data incident that way and had seen the result: for two years afterwards, every problem arrived from outside, because nothing arrived from inside.
One more thing was on her desk that morning. The firm's engagement letter with this client, signed in 2021, contained a clause requiring notification of any unauthorised disclosure of client information without undue delay. Nobody had read it in four years. It removed the option she had been quietly considering, which was to wait until she understood the incident fully before saying anything, and it made the timing question a contractual one rather than a matter of judgement.
What actually happened
Lakshmi told the client on Tuesday afternoon, by telephone and then in writing, before she knew how it had happened in detail. The letter said what had been sent, to whom, when, what the provider's retention terms said, that the account's training setting had been on, and what the firm was doing. The client was angry, asked for the firm's AI policy, discovered that it consisted of one sentence in the staff handbook, and did not move the file. The firm reported the matter internally, recorded it, and bought an enterprise tier within a month — the procurement had been stalled since the previous September. Arun was not dismissed; the note in his file records the incident and the fact that he raised it himself, in substance, by telling a colleague. Lakshmi's own conclusion, which she wrote into the one-page rule the firm now issues on day one, was that the incident was caused by the gap between what people needed and what they had, and that punishing the behaviour would have removed her visibility of it rather than the behaviour.
Worth arguing about
Arun's summary was accurate and he wrote the note himself. Why is the quality of the output irrelevant to what went wrong?
One answer
Because the duty was breached at the moment of disclosure, not by any error in the result. Pasting a client's agreement into an account the firm has no contract with is sending confidential material to an outside party, and a professional duty of confidentiality binds the individual regardless of what the provider does with it afterwards. Even instant deletion by the provider would not undo it. The provider's privacy policy is not a defence, because the duty was never the provider's.
Lakshmi decided the firm had a supply problem rather than a discipline problem. What is the evidence for that reading, and what follows from it?
One answer
The evidence is that the approved tool could not do the thing Arun needed at the hour he needed it, and that six colleagues faced the same constraint with the same laptops. People reach for unapproved tools because the approved one is slower or absent and there is a deadline. Punishing that removes visibility of it rather than the behaviour, so what follows is providing a usable approved tool quickly, and making early self-reporting safe, since the gap between what people need and what they have is exactly the size of the shadow-AI problem.
Suppose Arun had removed the names before pasting. Would that have made it acceptable? Give the reasoning, not just the answer.
One answer
No, and only partly for the obvious reason. Removing names reduces but does not remove identifiability, because a case description is often unique without a name — a settlement figure, a sector, a date and a distinctive clause identify the matter to anyone who knows the market. More decisively, the confidential material is the terms themselves, not the parties' names, so the disclosure of the agreement is the breach whoever is named in it. Anonymisation is a discipline to apply before something may go into an approved tool, not a licence that makes an unapproved one permissible.
Test yourself6 questions on this module
You paste a client file into a personal chat account, then delete the conversation ten minutes later. What is your position?
Your firm is comparing a consumer account with the same brand's enterprise tier. What actually differs, and what should decide which one a document may go into?
A colleague says that because no AI-specific statute applies to your sector yet, nothing constrains what your team puts into these tools. What is wrong with that?
You paste an unredacted case note into an online model and ask it to remove the identifying details. What is wrong with this, beyond questions of accuracy?
A recruiter removes the gender field before training a model on a decade of past hiring decisions. Why does bias survive that, and where does the real damage occur?
When is disclosure of AI use actually owed, on the principle the course sets out?
Module 7
The job you actually have
Everything so far has been general. This block applies it, trade by trade, to the work people are actually paid for — the ledger, the classroom, the ward, the file, the counter, the campaign, the van. Each lesson names what genuinely helps in that trade, the failure that catches practitioners there specifically, the duty or regulator that decides the boundary, and the free path for anyone without a licence budget.
By the end you can
Produce a one-page written rule for your own trade naming three tasks you will hand over, three you will assist with, three you never will, the professional duty or regulator behind each refusal, and the free tool you would use where the material may not leave your building
- 56Running the method on your own tradeFive questions transfer to any job — what you produce in volume, where the truth for it lives, who is harmed if it is wrong, which body has a view, and whether checking is cheaper than doing — and the person who signs is accountable for all of it, not just the parts they wrote.
- 57If you work in accounts and financeTurning figures you established into words is the whole safe territory — nothing generative writes to the ledger, no rate or threshold comes from a model, and a variance explanation you did not establish yourself is a fabrication in a document somebody will act on.
- 58If you teachYou form the judgement in bullet points and the model writes the feedback paragraph, never the reverse — and a detector score is not evidence, because false positives fall hardest on non-native writers and on the students least able to contest an accusation.
- 59If you work in health and careSoftware intended for diagnosis, triage, prognosis or treatment is a regulated medical device wherever you practise, which places documentation and patient communication inside the usable space and clinical decisions outside it — and an error in a health record propagates forward into everything written after it.
- 60If you work in law or complianceCitation formats are among the most regular patterns in text, so well-formed references are trivially generated and carry no evidence of existence — and the harder second check, whether a real authority says what was claimed, is the one that gets skipped.
- 61If you work in government or public serviceA reason generated after a decision is not the reason for it, prompts held by a public authority may be disclosable under freedom of information, and counting consultation responses requires a full classified pass rather than a question to a retrieval tool.
- 62If you run a shop or a small businessFree tiers plus one paid seat plus a local model covers nearly everything a small business needs — and the specifics that must never be generated are product facts, substantiated claims and reviews, because the regulator asks the advertiser for evidence and does not ask who typed it.
- 63If you work in marketing or communicationsThe model produces the centre of the distribution and the centre is where every competitor using the same defaults has already landed, so distinctiveness comes only from material nobody else has — your interviews, your data, a specific example with a name and a date.
- 64If you fix things for a livingTorque figures, clearances and part numbers come from the manufacturer's documentation for that exact model and never from the model's memory, because a value drawn from a distribution of similar equipment arrives in the same confident format as a looked-up one and the consequence here is physical.
Case studyA torque figure that was not in any manual
Bhavesh Patel's firm services walk-in cold rooms and blast chillers for dairies and sweet shops across three districts. Six months ago he started dictating job reports into his phone on the drive back and having them structured into the firm's report format. It saved him about forty minutes a day and his customers now get the report the same evening instead of the following week. Two of his four engineers do the same.
In August his youngest engineer, Dhruv, was on a chiller from a manufacturer the firm had never worked on. The unit had no plate visible from the access side and no manual on site. He asked a chatbot on his phone for the compressor mounting bolt torque for that model.
He got a number. It was in newton metres, it was in the right range for a bolt of that size, it was given without hesitation, and it was accompanied by a sensible note about tightening in a cross pattern. It was also not the manufacturer's figure, because the model has no manual inside it; it produced the torque that such bolts usually take.
Dhruv used it. The mounting held for eleven days and then a bolt backed out, the compressor moved on its mounts, a discharge line fractured, and the dairy lost about four hundred litres of milk and a day of production. Nobody was hurt. The repair and the milk cost Bhavesh a little over ninety thousand rupees and a customer he had held for six years is now on a competitor's contract.
Sitting in his office afterwards, Bhavesh's first instinct was to ban the phone tools outright. He had the authority and it would have been popular with his oldest engineer, who had never liked them.
The cost of that was not small either. The dictated reports were the single biggest improvement in the business in five years. They had produced a written record on the day of the job, which had already won him one insurance dispute, and they were how he now quoted within twenty-four hours. Banning the tool to fix a torque figure would remove all of that. It would also, he suspected, not remove the behaviour: Dhruv had used his own phone on a customer's site, and a rule Bhavesh could not see anybody obeying was a rule that would only tell him he had no idea what his engineers were doing.
The other option was harder to write than to say. He had to draw a line inside the tool rather than around it: name what may come out of a model and what may only come out of a manufacturer's document, and make the difference something a twenty-three-year-old in a plant room at four o'clock can apply without thinking about it.
When he tried to write that line, he found it did not fall where he expected. His instinct had been to sort tasks by how difficult they were, and that produced nonsense — explaining a fault to a dairy owner in plain language is a harder piece of writing than reciting a torque figure, and it is the safe one. What worked was sorting by where the answer comes from. Anything he or his engineer had observed, measured or decided could be handed over for phrasing. Anything printed in a manufacturer's document had to come out of that document. Anything that was neither had no business being in a job at all.
The part he found hardest to accept was that Dhruv had done nothing careless in the ordinary sense. He had a job to finish, no manual, no signal in the plant room, and a tool that answered. The firm had never told him where its answers came from, because the firm had never asked itself. Bhavesh had introduced the dictated reports enthusiastically eight months earlier and had said nothing at all about limits, and four engineers had reasonably concluded there were none.
What actually happened
Bhavesh wrote one page and had it laminated in all four vans. Three columns: hand over, assist, never. Under never, in the largest type on the page: torque figures, clearances, refrigerant charge, pressure ratings, fuse ratings and part numbers, which come from the manufacturer's documentation for that exact model and from nowhere else. Under assist: diagnosis as a conversation, and reading an error code with the manual attached. Under hand over: the dictated job report and the customer explanation. He then did the part that actually made it work — he spent two Saturdays downloading the service PDFs for every unit type on his customer list onto a tablet in each van, so that the manufacturer's answer was reachable in the plant room where the phone was not. Where a manual could not be found, the rule is that the engineer rings the office and the office rings the distributor. Dhruv is still with the firm and it was Dhruv who built the folder structure on the tablets. Fourteen months on, the dictated reports are still in use and no figure has come from a model since.
Worth arguing about
The torque figure arrived in the same format and the same confident register as a correct one. Why does that make this failure different from a wrong figure in an office document?
One answer
Because there is nothing in the reply that distinguishes a value looked up from a value drawn from a distribution of similar equipment, and here the consequence is physical and delayed. In a report, a wrong number is usually caught by somebody who knows the business. In a plant room the number is acted on immediately, the failure shows up days later, and by then the causal chain is invisible. That is why safety-critical specifics have to come from the manufacturer's document for that exact model rather than from a check on plausibility.
Bhavesh decided that banning the tool would have cost more than it saved. Set out both sides of that judgement.
One answer
Banning it removes the dictated job report, which produced a same-day written record, won an insurance dispute and cut his quoting time to a day — the biggest operational gain the firm had made. It also would not remove the behaviour, since engineers use their own phones on customers' sites, so it would cost him visibility as well. Keeping it means accepting that a young engineer under pressure might again ask for something he should not, which is why the answer was a line drawn inside the tool and a manual made reachable, rather than a prohibition nobody could enforce.
The one-page rule worked only because of the two Saturdays spent putting manuals on tablets. What general principle does that illustrate?
One answer
That a rule forbidding a source has to be paired with a reachable alternative, or the pressure that produced the behaviour is unchanged. Telling an engineer in a basement with no signal that torque figures come from the manual, when no manual is available, is asking him to choose between the rule and finishing the job. Providing the approved answer where the work happens is what converts the page from a statement of intent into something a person can actually follow at four o'clock.
Test yourself6 questions on this module
A librarian wants annotations for a reading list of two hundred titles. Where does the truth live, and what follows for how she works?
You are drafting management commentary and ask why travel costs rose eighteen per cent. You get conference season, fuel prices and a return to in-person meetings. What is the rule this violates?
A teacher wants help with written feedback on thirty essays. Which arrangement is both safer and better, and why?
Which use of a general chatbot in a clinical setting crosses the line the course draws, and what makes it a line rather than a matter of caution?
Why do fabricated legal citations keep appearing, rather than some other kind of error, and what does a proper check involve?
A department must report how many of 900 consultation responses opposed a proposal. Why is a document-chat question the wrong method, and what is the right one?
Module 8
Making it hold, and keeping what is yours
The last block is about durability. What to do when an AI-assisted error reaches somebody, how one person's good prompt becomes a team's asset, where automation must stop and ask, how to prove honestly whether any of this helped — and, at the end, the two things worth protecting: your own position, and your own skill.
By the end you can
Run a two-week trial of one AI-assisted task against a named before-and-after measure that includes checking time, and report honestly whether to keep it, change it or stop.
- 65Checking your own work, and talking to a sceptical bossYour name on the work means you vouch for every line — check hardest on specifics you did not supply.
- 66When an error reaches somebodyAn organisation is bound by what its AI told a customer, and the recovery is the ordinary one — tell the affected person, correct the record, log the cause — never a quiet fix.
- 67Turning one good prompt into everyone'sA prompt that worked is an organisational asset, and writing it down with its context and its failure notes is what turns one person's fluency into a floor everybody stands on.
- 68The hour that actually changes what people doShow a live fabrication on the room's own work before showing anything useful, because a feature tour teaches what the tool can do and only a witnessed failure teaches when to trust it.
- 69Automation, and where it must stop and askChaining steps multiplies error rather than adding it, so an automated sequence must pause for a human at every irreversible act — sending, paying, deleting, publishing.
- 70What it costs, and what you are tied toThe largest cost is checking time rather than licences, the real lock-in is your prompts and habits rather than the model, and a twenty-item regression set is what turns "something feels different" into evidence when a provider updates underneath you.
- 71Proving whether it helpedA two-week trial with a baseline, one measure that includes checking time, and a stated stopping condition produces more usable evidence than a year of enthusiasm or a year of scepticism.
- 72Applying all of this to your own job searchUse it to research, to structure and to rehearse — and keep every claim in the application literally true, because the interview tests the claim and not the document.
- 73Keeping the skill that made you usefulSkill decays without practice, so the tool must sometimes be used as a tutor rather than a substitute — and juniors need the reps that AI is most efficient at removing.
Case studyFour hundred hours that nobody could find
Nandini Rao had been asked to put a number on it. The board had approved thirty-one seats on an AI assistant a year earlier, and now wanted to know what the department had got for the money before renewing.
She had a number available immediately. The vendor's calculator, filled in with the department's volumes, produced a saving of four hundred and twelve hours a year. Two of her team leaders were happy to endorse it, and the staff survey she had run in June said that seventy per cent of handlers felt the tool made them faster.
That package would have passed. It had a figure, an endorsement and a survey, and it was the same shape as every other paper that goes to that board.
The problem was that Nandini had read the METR result during the training the department ran in February — sixteen experienced developers who expected to be a quarter faster, reported afterwards that they had been a fifth faster, and were measured at a fifth slower. She could not unsee it. Her survey said the same thing the developers' self-reports had said, and she had no reason to think her handlers were better judges of their own speed than sixteen professionals timing work in codebases they knew.
So she had a decision, and it was uncomfortable in a way that had nothing to do with technology.
If she took the four hundred and twelve hours, the renewal passed, her team kept a tool most of them liked, and she kept a good relationship with a finance director who does not enjoy ambiguity. If the number was later found to be soft, it would be found soft in about two years, by which time everyone would have moved on.
If she measured it properly and the answer came back at zero, she would have to stand in front of a board and say that a decision she had supported for a year had produced nothing measurable. Thirty-one seats would probably go. Her handlers would lose something they were used to, and she would be the person who took it away with a stopwatch.
She had three weeks, which is not enough to measure a department and is enough to measure one task.
She picked the intake summary: the handler reads a new claim file and writes a two-hundred-word summary that the assessor works from. It happens about forty times a day, it is well defined, and it is textual. Before anything changed, she had four handlers time it unassisted on the next eight files each — thirty-two baseline timings. Then two weeks of timing it with the tool, from sitting down to would-sign, including the checking and any corrections. Three yes-or-no quality marks each time: did it need a second round of corrections, did the assessor come back with a question, was a factual error found at the checking stage.
And she wrote the stopping condition down before the first timing, on the first page of the file: if the median does not fall by at least twenty per cent, or if the error rate rises at all, the intake summary comes off the tool.
What actually happened
The baseline median was nineteen minutes. The assisted median was eleven, and the error rate at the checking stage went from one file in sixteen to one in fourteen, which on those numbers is not a difference. That is a clear win on a task worth about forty repetitions a day, and the paper said so with the arithmetic shown. The second finding was the one Nandini had not expected. She had also, informally, timed the complex-liability referral note, which is the work her three most experienced handlers do, and it went the other way — the median rose from fifty-one minutes to fifty-eight, because a senior handler's own first draft is already better than the tool's and bringing average up to their standard took longer than starting from their standard. She reported both. The board renewed twenty-two seats rather than thirty-one, took the tool off the referral note, and asked her to run the same two-week measurement on two more tasks before the next renewal. The finance director told her afterwards that it was the first paper on the subject he had believed.
Worth arguing about
Nandini's staff survey and the vendor calculator both supported renewal. Why did she treat neither as evidence?
One answer
Because both measure belief rather than time. A vendor calculator is an assumption sheet with the vendor's assumptions in it, and a survey asking whether people feel faster is exactly the instrument that told sixteen experienced developers they had been twenty per cent faster while the clock recorded nineteen per cent slower. Feeling is a poor measurement here for identifiable reasons: waiting for a draft feels like progress, and the correcting and re-prompting is fragmented into pieces too small to register as work.
What did writing the stopping condition down in advance actually protect against?
One answer
Against re-reading the result until it agrees with you. A threshold fixed before the first timing — twenty per cent, and no rise in errors — removes the option of deciding after the fact that eleven per cent is encouraging, or that a slightly higher error rate is acceptable in the circumstances. It converts a trial from advocacy into evidence, and it is what allows the same person to report an unwelcome result without appearing to have changed the standard.
The referral-note result was worse with the tool. Why is reporting that finding worth more to Nandini than suppressing it?
One answer
Because it is what makes the favourable finding believable. A person who has said this did not help here is believed the next time they say something did, and the board's response — renew fewer seats, remove the tool from one task, measure two more — is a better decision than either blanket renewal or blanket withdrawal. It also matches what the evidence predicts: gains are largest where the person is weakest and can reverse on core craft, so a department reporting uniform benefit across skill levels is reporting enthusiasm rather than measurement.
Test yourself6 questions on this module
You are about to send a report drafted with AI assistance. Which pass finds the errors that concentrate in generated text?
An airline's website assistant told a passenger he could apply for a bereavement fare after booking. The airline's actual policy said otherwise. The tribunal rejected the airline's defence. What is the general rule?
Which line in a shared prompt library entry does most of the work for the colleague who uses it next?
You are running the first hour of AI training for a team. Why does the session open with a live failure on the room's own work?
An automated sequence has ten steps, each right about 95 per cent of the time. Roughly how often does a run come through clean, and what makes real chains worse than that figure suggests?
A team wants to know whether a tool helped. Which of these produces evidence rather than advocacy?