An AI tool can only work with what you point it at. Preparing documents for AI is what decides whether the output is usable. Nearly all of that work happens before you buy anything. Give a tool clean, current, consistent material and an ordinary product produces something you can use. Give it four versions of the same price list and the best product on the market gets it wrong, while sounding completely sure.

Your source material sets the ceiling

That happens because a model doesn’t know your business. It reads what you hand it and works from there. So the state of your documents sets a ceiling on the quality of the answer, and rewriting the prompt does little to lift it.

If your payment terms sit in three files with different wording, the tool has no way to pick the right one. It picks one anyway. That is the risk in the jobs people want first: answering a supplier query, drafting a quote, checking a purchase order. Each of those rests on a document being current.

Why doesn’t a demo show you any of this?

A demo avoids that problem by design. The material in it is clean, recent and written for the purpose, because the supplier chose it. So you watch a tool perform on the best inputs it will ever meet. You learn nothing about how it copes with yours.

A demo also hides a question of volume. It runs on one tidy document, while your job will point the tool at a folder holding hundreds. That changes the behaviour, because the tool now has to choose which file answers the question before it answers it.

Ask instead whether the supplier will run the same demo on a folder of your own material. If the answer is no, or not yet, treat that as useful information rather than a reason to walk away. It tells you the preparation is yours to do, so you should price it in before you sign.

Preparing documents for AI starts with one true copy

Pricing that work means knowing what it involves. There are three jobs, and the first is deciding which copy of each document counts.

For every document the tool will read, pick the version that is current. Put it where the tool can reach it, then move or clearly mark the rest. Old drafts in a shared drive are not harmless. A tool reads them with the same trust as the live one.

You don’t need a filing project for this. What you need is a short list: the current price list, the current terms, the current process notes, the current templates. Write down where each one lives and who may change it. If nobody can answer that second question today, start there.

Name the fields a correct output contains

Once the source is settled, the next question is what a right answer actually holds. Write that down before you run anything.

A quote might need the item, the quantity, the unit price, the lead time and the payment terms. A reply to a supplier query might need the order number, the date and the agreed price. Five or six fields is usually enough.

That short list does two useful things. It shows you which documents have to hold the information, which often reveals that one field lives in somebody’s inbox rather than in any file. It also gives your reviewer something exact to check against. That same list is what you use when you agree what good looks like with the people doing the job.

Keep the list to things a reviewer can confirm quickly. If a field takes 10 minutes to verify, the review quietly stops happening as soon as people get busy.

Make your terminology consistent

Fields only help if the words around them mean one thing. Businesses collect synonyms quietly over the years, and nobody minds until a machine starts reading.

The same item might be a job, a works order, a project or a ticket, depending on who typed the document. People cope with that because they know the history. A tool doesn’t, so it either treats them as four separate things or merges them into one. Either way, any count or summary it hands back to you is wrong.

Pick a single word for each thing that matters, then use it in the documents the tool will read. You don’t have to rewrite the archive. Fix the handful of terms that appear in the job you care about now, and leave the rest alone.

Access is not the same as browsing

With the material sorted, the last question is what the tool can actually reach. This is where careful preparation most often comes undone.

When a person opens a shared folder, they see what they need and ignore the rest. When a tool searches that folder, it reads everything within reach and may quote any of it back. So permissions that felt fine for people are often far too wide for software.

Look at the folder as the tool will see it. Salary letters, notes from a disciplinary, a client file marked sensitive: none of that belongs beside your price list. The Information Commissioner’s Office publishes guidance on AI and data protection worth reading before you connect anything holding personal data.

Build a narrow folder rather than a wide one

The fix follows directly from that boundary problem. Create one folder holding only the documents this job needs, then point the tool at that and nothing else.

A narrow folder is easier to explain, easier to check and easier to widen later once you trust the results. Setting one up is short work for whoever looks after your systems. It is also one of several good reasons to involve your IT partner before the first live run rather than after it.

What the preparation actually buys you

Done this way, the first run stops being a gamble and becomes a test. You will still get answers you don’t like, because early runs always throw some up.

The difference is that you can explain them. You know which document the tool read. You know which field went missing, so you can fix the source instead of arguing with the software.

Preparing documents for AI also pays for itself more than once. The same tidy source serves the second job you try, and the third. The cost lands once, the benefit repeats, and a few days of unglamorous sorting turns out to be the cheapest part of the whole exercise.