What to do when AI gets it wrong
How far a wrong answer travelled decides what it costs you, and what you should do in the hour after you spot it.
The damage when AI gets it wrong depends on how far the answer travelled before someone spotted it. A wrong draft caught on screen costs you the time it takes to type it again. The same wrong figure inside a quote a customer has already accepted costs you a great deal more. Distance is the thing to manage.
Three levels of consequence when AI gets it wrong
Sort every error by where it stopped, because that tells you how hard to react.
At the first level, the answer never left the screen. Someone read it, saw the fault and fixed it before anything moved. The cost is the time to redo the work, plus a note of what happened.
At the second level, somebody inside the business acted on it. A payment went into the run, an order went to a supplier, a rota went up on the wall. Now you have two jobs, because you have to correct the answer and unpick the action it caused.
At the third level, the answer left the building. It reached a customer, a supplier, an insurer or a public document. The correction is usually simple, while the conversation about trust is not.
Stop the output, not the tool
Since distance drives the cost when AI gets it wrong, your first move is to stop the answer travelling any further.
Hold the send. Pull the batch back. Ring the person who is about to act on it and tell them to wait. All of that takes minutes and it caps the damage at whichever level you caught it.
Switching the tool off feels like the responsible move, and it usually isn’t. Turning it off recalls nothing that already went out, and it can wipe the history you need to work out what happened. Leave the tool alone until you understand the fault.
Keep the source and the original
With the output held, capture the evidence before anyone tidies it away.
Keep four things: the exact input, the exact output, the date, and the name of whoever ran it. Take a copy rather than a description, because a summary written later smooths over the detail that explains the error.
This matters more with AI than with a broken spreadsheet. Ask the same question again tomorrow and most tools give you a slightly different answer. The original output is often the only proof of what people actually saw.
Which of four things broke?
Now you can work out the cause, and there are only four candidates worth checking.
The source material may have been wrong. An out-of-date price list, a superseded contract or a folder holding three versions of one document all do the same thing. Each one produces a confident answer built on wrong facts.
The instruction may have been loose. If nobody said which fields the output must contain, you asked for judgement rather than a task. What came back is somebody’s guess at what you meant.
The tool may have invented it. AI writes plausible detail when it has nothing solid to draw on, and invented detail reads exactly like the real thing.
Finally, the check may have failed. Perhaps nobody compared the output against anything, or the check sat after the point where the work became hard to reverse.
Repair the process, not the one answer
Three of those four causes belong to you rather than to the supplier, which is good news, because you can fix them.
If the source was wrong, settle which copy is the authoritative one and remove the rest from reach. Preparing documents for AI sets out how to do that properly. Fixing the source removes a whole class of errors rather than this single one.
If the instruction was loose, rewrite it so it names the fields a correct output contains. If the tool invented something, narrow the job and require the output to point at the source it came from. If the check failed, move it earlier and make it something a person can complete in seconds. Agreeing what good looks like covers how to write a check that people actually run.
Then repeat the failed case through the fixed process. A fix nobody has tested against the original fault is only a theory.
Who needs telling
Fixing the process quietly is not enough, because other people are carrying the consequences of the error.
Tell whoever acted on the output first, since they may still be able to stop something. Tell the person who owns the job, so the fix lands in the instructions rather than in one person’s memory. If the answer reached a customer or supplier, tell them plainly and early, and lead with the correction rather than the explanation.
One case needs more than a phone call. If the error sent somebody’s personal data to the wrong person, you may be looking at a personal data breach. That carries a duty to assess the risk to the people involved. Where a risk to those people is likely, you must tell the ICO as soon as you can, and within 72 hours where that is feasible. The ICO guidance on reporting a breach sets out the test. Read it before you decide the duty doesn’t apply to you.
When to restart, and when to stop
Once the telling is done, you face a decision that people tend to make on mood rather than evidence.
Restart when you can name the cause in a sentence. Restart when you have changed something specific, and when the changed process handles the case that failed. All three conditions are checkable, and a promise to be more careful next time meets none of them.
Stop when the same kind of error comes back after a fix. Stop when checking now costs more than the job saves. Stop when nobody can tell whether an answer is right without redoing the work by hand. Stopping one job is a normal management decision. Keeping a job alive because you don’t want to admit it failed is how a small error turns into a standing risk.
What a good response leaves behind
Whether you restart or stop, the response should leave the same record behind. Every business using AI will get a wrong answer eventually, so judge yourself on what you do when AI gets it wrong rather than on how rarely it happens.
A good response leaves four things behind. A corrected output, an evidence file, a changed step in the process, and a short written note of what broke. Keep those notes together. After a few months they show you which jobs run steadily. They also show which ones need a firmer check, and which ones never suited a tool at all.
