Discuss a project

Automation & AI5 min read

How to tell whether an AI solution is accurate enough

A demonstration always looks good, because the examples were chosen. Accuracy is your number, measured on your own cases — and you can measure it before you buy. How to do that without technical knowledge.

Author
Devnora team
Published
Updated
Reading time
5 min read
Language
Read in Lithuanian
In this article
  1. In short
  2. Start with cases, not technology
  3. Why average accuracy misleads
  4. Does the solution know when it does not know
  5. What threshold is good enough
  6. Illustrative scenario, not a description of a client project
  7. What to ask a supplier

A demonstration of an AI solution is almost always convincing. Not through dishonesty — a demonstration uses chosen examples, while your work produces cases on its own. The only figure that means anything is accuracy measured on your cases. The good news is that this can be done before deciding to buy, and it needs no technical knowledge.

In short

  • Accuracy is not a property of the solution — it is a property of your data and your task definition. Another client's number tells you nothing.
  • The first piece of work is not a model but a set of real cases with known correct answers.
  • Average accuracy misleads. What matters is how the solution handles hard cases, and whether it knows when it does not know.
  • "I don't know" is a correct answer and should score better than a convincing guess.
  • A good-enough threshold depends not on a number but on what an error costs and who catches it.

Start with cases, not technology

Before any conversation about models, collect a few dozen real cases with known correct answers. If the task is extracting data from documents, that means documents with the correct fields written out by hand. If it is classification, records with the correct category. That set is the only thing that lets you ask "what is the accuracy" and receive an answer rather than an opinion.

The set's most important property is not size but composition. It has to include cases a person also finds hard: unusual formats, incomplete information, two plausible answers. If it contains only easy cases, you will measure a number that says nothing about real work.

Why average accuracy misleads

A single figure merges things with very different costs. Being wrong on an easy case and being wrong on an unusual one are not the same, and being wrong quietly differs from being wrong loudly by more still. So instead of one number, ask for three.

  • How many cases did the solution handle correctly without a person?
  • How many did it flag as unclear and route to a person? That is not an error but correct behaviour.
  • How many did it get wrong while saying nothing? Only this number is the real risk.

The third figure is the only one process cannot compensate for. The first two can be shifted with settings: a solution can be made more cautious so more cases go to a person. A quiet error can only be noticed or missed.

Does the solution know when it does not know

This is the most important property and usually goes untested. Take a few cases where no answer exists at all — a document missing the required field, a question whose answer is not in your data. A suitable solution should return an empty value or say it found nothing. An unsuitable one will fill the gap with a guess, and that guess will look exactly as tidy as a correct answer.

This test takes a few minutes and tells you more about a solution than an entire demonstration.

What threshold is good enough

There is no universal figure, and anyone offering one is describing somebody else's task. The threshold comes from two things: what a single error costs, and who catches it. If the output is a draft a person reads before sending, a fairly low threshold is fine — the solution only saves writing time. If the output reaches the accounts unreviewed, almost no threshold is high enough, and the right answer is to add a rule or a review step.

A practical way to frame it: instead of asking whether some accuracy percentage is enough, answer what you will do with the remaining cases. If there is no answer, the question is premature.

Illustrative scenario, not a description of a client project

Imagine a solution extracting totals from incoming documents. In testing it looks accurate, so it is used without review. A month later somebody notices that in several documents the total was read from the wrong place — not a calculation error but the wrong field, which looked like the right one. The risk was never the accuracy percentage; it was that the error was silent and nothing caught it. The correct design from the start was simple: use a rule to compare the total against the order before writing anything. This is an example of how we weigh risk, not a promise about accuracy.

What to ask a supplier

  • How will you measure accuracy, and on whose examples?
  • What does the solution do when it does not know the answer?
  • How many cases are expected to be routed to a person, and how does that shift with settings?
  • What happens if the agreed threshold is not reached — is that in the agreement?
  • How is accuracy tracked later, when the data changes?

That last question usually goes unanswered, and it matters: a solution that was accurate on launch day does not necessarily stay so when document formats change or new kinds of cases appear. Accuracy is a maintenance question, not an acceptance question.

Frequently asked questions

How many examples are needed to measure accuracy?
A useful set usually starts in the low tens of real cases, and composition matters more than count: it has to include the hard, atypical ones. A hundred easy examples will produce a flattering number and tell you nothing about real work.
Can a supplier state the accuracy in advance?
The honest answer is no. Accuracy depends on how tidy your data is and how the task is defined, so a figure obtained with another client's data means nothing for you. What a supplier can state is how accuracy will be measured and what happens if the threshold is not reached.
Who prepares the examples?
You do, and that is one of the project's real costs. It needs cases with a known correct answer, which means somebody has to review them. This work frequently takes longer than the solution itself, and is worth agreeing before rather than after.

Share

Send by email

Next step

Related service

AI solutions

AI solutions: what kind of error is possible, how output is checked before use, where a human checkpoint belongs and when a plain rule is better.

If this article describes your situation, tell us what is not working. We will say whether and how we can help.

Worth reading next

  1. Automation & AI

    Automation or AI: how to decide

    Both promise the same thing: the work happens without a person. The difference is not the subject but how certain the answer is. Automation gives the same answer every time; AI gives a likely one. How to choose, and why the cheaper option usually fits.

  2. Automation & AI

    Preparing a team to work with AI: from scattered tools to an agreed way of working

    In most companies somebody already uses AI — each in their own way, with nothing agreed. What a team genuinely needs to learn, what is worth standardising, what must be checked, and when training is the wrong answer.

More on this topic: Automation & AI