Agree what good looks like before AI goes live
Three people have to agree the standard before an AI step touches real work, and it has to be a standard a reviewer can apply in seconds.
Before an AI step touches live work, three people need to agree what good looks like. The person who does the job, the person who checks the output, and whoever carries the cost when it is wrong. Get them in a room for an hour and write it down together. Doing that after launch is a great deal harder.
Why one person’s standard never holds
The temptation is to save everyone an hour and write it yourself. A standard produced that way always reads well to its author. It comes apart on contact with the job, because the person writing it knew things they never thought to put on paper.
Two versions of that failure turn up repeatedly. The owner’s standard is too loose, so almost anything passes and the check quietly becomes a formality. Or the reviewer’s standard is too tight, so nothing passes, the queue backs up and people drift back to the old way within a month.
Who has to agree before anything goes live?
So more than one person needs a hand in writing it. Three roles have to take part, and none of them can send apologies.
The person doing the job knows which cases are awkward and how often the awkward ones appear. The person checking the output knows how long a check takes, which decides whether it still happens in week six. The person carrying the cost of a mistake, which is usually you, sets how much error the business can absorb.
Agreeing what good looks like is a conversation about how the job runs today, which is why an hour is normally enough. Leaving one of the three out costs you far more later.
Writing down what good looks like
Agreement only helps once somebody writes it in terms a reviewer can apply. So write it as the things they will look for, in the order they will look.
Avoid adjectives. Accurate, professional and on brand mean different things to different people, and none of those meanings survives a quick check. Name the fields instead.
A quote passes when the item, the quantity, the price and the lead time match the enquiry, and the payment terms match what the customer already agreed. That version sends the reviewer to a precise place. The adjective version sends them nowhere.
Speed matters here more than thoroughness, which sounds wrong until you watch a check die. A check taking 5 minutes gets done while the step is new, then people start skipping it under pressure. A check taking 20 seconds survives, because it fits inside the work rather than sitting on top of it.
So design for seconds. Three or four things to confirm, each one visible without opening another document. If a check needs somebody to look up a figure elsewhere, move that figure into the output. When you can’t, drop the check and accept the risk knowingly. Time it once with a stopwatch. When it runs long, the design is at fault rather than the reviewer.
A stop rule needs a condition and a destination
Fast checks catch small errors. You also need a rule for the ones that are not small, and most written rules only do half the job.
Pair a condition with a place. If the tool can’t find the customer’s agreed rate, nobody sends it: the item goes to the review folder, and accounts hears about it the same day. Compare that with a rule saying escalate anything unusual, which fails twice. Unusual is a judgement, and escalate is not a place.
Settle the stop rule while nobody is under pressure. Working out what to do when AI gets it wrong is far harder on the afternoon it happens for the first time.
Establish the before position
A standard and a stop rule together tell you whether an output is acceptable. Neither tells you whether the step is an improvement, and those are separate questions.
For the second one you need the before position, recorded while the old way is still running. Two numbers usually cover it. How long the job takes now, and how often the current output comes back for changes.
Take those from the work itself rather than from memory. Ask whoever does the job to time 10 of them across a week and count the ones that came back. Rough numbers gathered honestly beat precise numbers invented afterwards, and they are what turns a good feeling into a measured win.
Test the standard on work you have already finished
With a standard and a baseline in hand, test the standard before the tool ever does. Run it against work your team completed last month, when no tool was involved.
Pick 10 finished items and mark each one pass or fail against the written standard. Two useful things fall out. If your own past work fails, the standard is stricter than the business actually needs, and the tool will fail it too.
If everything passes without a moment’s thought, the standard is checking nothing and the review will be theatre. Adjust until the result matches your honest opinion of that work. An hour spent on finished work now will save you a fortnight of argument about new work later.
How the standard changes, and when it shouldn’t
A standard set this way will still move, and it should. Expect two kinds of change once the step is live.
The first is tightening, when the tool turns out better than you assumed and you raise the bar to match. The second is narrowing, when one category of work keeps failing and you take that category out of the step altogether. Both are healthy.
Guard against the third kind. A standard that drifts downwards because the tool keeps missing something has stopped being a standard. If you notice yourself lowering the bar to keep the tool in the job, name it out loud and decide whether the step still earns its place. Whoever owns the job needs the authority to make that call without asking twice.
The document that outlives the tool
Guarding the standard against that drift is worth the effort. What good looks like is the one page from this exercise that keeps its value. Products change, prices change, and the thing you buy in two years will do more than the thing you buy today.
The standard survives all of it, since it describes the work rather than the software. Write it once, properly, with the people who have to live with it. You can then change tools without reopening the question of what a good result contains.
