Ordinary software does the same thing each time. A feature built on a language model does not. It can answer one question well and a nearly identical one badly. Trying a few examples and nodding is how weak AI features reach customers.

Why the usual testing falls short

Conventional tests check for an exact result. The output of a language model varies in wording, and occasionally in substance. A model can be right in the demo and wrong when a customer phrases the question differently.

The answer is many examples and a way to score them.

Build a test set

Collect real questions from support tickets, search logs and emails. For each one, write down what a good answer must contain.

  • The common cases.
  • The awkward ones: misspellings, mixed languages, vague questions.
  • Requests it should refuse or pass to a person.
  • Questions whose answer is not in your data.

Fifty to a hundred cases is a workable start. The set grows as you find new failures.

Decide what good means

  • Correct: the facts match the source.
  • Grounded: the answer came from your documents.
  • Complete: nothing important is left out.
  • Safe: no advice or promise it should not make.
  • Honest about limits: it says so when it does not know.

Scoring by people is the most reliable method and the slowest. A second model can score answers against the expected ones quickly. Check a sample of its judgements by hand.

Run it on every change

A change to the prompt, the model version or the documents can improve one kind of answer and damage another. Run the whole set each time and compare the result with the last run.

Without that, nobody knows whether a change helped.

Decide the acceptable error

No AI feature is right every time. Before launch, decide what rate of wrong answers this use can tolerate, and what a wrong answer costs.

A draft reply that a person reviews can tolerate more errors than an answer about a refund sent straight to a customer. Where a mistake is expensive, keep a person in the loop.

After launch

  • Log questions and answers, with care for privacy.
  • Let users flag a bad answer.
  • Review a sample every week.
  • Add each failure to the test set.
  • Rerun the set when the provider updates the model, since behaviour can shift without any change on your side.
  • Watch cost and response time as well as quality.

What to ask a supplier

  • Can we see the test set?
  • What share of cases pass today?
  • What happens when the answer is wrong?
  • How will we know if quality drops?

Our service

AI / ML Integration