Skip to content

Evaluating an agent: tests that should pass before real work

A practical set of tests every AI agent should pass before it touches customers, money or records, and how to build them from your own cases.

A demo proves an agent can do something once. A business needs to know it will do the right thing hundreds of times, on messy inputs, without supervision for most of them. The gap between those two is evaluation: a planned set of tests that the agent must pass before it is trusted with real work. It is the least exciting part of any agent project and the part that most often decides whether it succeeds.

Build your test set from real cases

The best tests come from your own history, not from imagination. If the agent will sort customer messages, collect a few hundred real messages from past months, remove personal details and write down the correct handling for each. If it will check loan applications, pull a mix of complete, incomplete and suspicious files and record what a careful officer decided.

Aim for a spread, not just typical cases:

  • Everyday cases, the bulk of what the agent will see.
  • Awkward cases: messages in mixed Sinhala and English, blurred photos of documents, customers who ask three things at once.
  • Cases that must be refused or escalated: a request for a refund above the limit, a complaint that hints at legal action, a question outside the agent’s job.
  • Hostile cases: messages that try to trick the agent into ignoring its instructions.

Write the expected result for each case before you run the agent. Deciding the right answer after seeing the agent’s answer is how weak results get excused.

Five tests that should pass before go-live

Different jobs need different checks, but these five apply to nearly every agent.

  1. Correctness on common work. On everyday cases, does the agent reach the right outcome? Set the bar in advance, based on how well your staff do today, not on perfection.
  2. Correct refusal and escalation. On cases it should not handle, does it stop and pass them on? This matters more than raw accuracy, because a wrong escalation costs a few minutes and a wrong action can cost a customer.
  3. Format and completeness. Does it fill every required field in the right format, every time? An answer that is right but missing a reference number breaks the next step.
  4. Consistency. Run the same cases several times. Does the agent give the same answer? If one case is handled correctly on Monday and wrongly on Tuesday, you have a reliability problem, not a knowledge problem.
  5. Tool safety. In a test environment, try to make it call tools it should not: send a message without approval, change a record it should only read. It should fail to do so because the permissions prevent it, not because it chose not to.

Scoring without fooling yourself

Keep scoring simple and honest. For each case, mark pass or fail against the expected result, and note the reason for each failure. A single overall percentage hides too much. Break results down by type: common cases, escalations, format, hostile inputs.

Weigh failures by cost. Say an agent handles 200 test cases and gets 10 wrong. If all ten are minor wording issues in draft replies a person reviews anyway, that may be fine. If two of the ten are refunds it approved when it should have escalated, the agent is not ready, whatever the headline number says.

Where a judgement is fuzzy, such as whether a drafted reply is polite and accurate, have a person review a sample. Using another AI model to grade replies can speed things up, but check its grading against human judgement before relying on it, because a model grader has its own blind spots.

Keep the tests after launch

Evaluation is not a one-off gate. Every time you change the agent’s instructions, switch to a different model or add a tool, run the full test set again. Changes that fix one problem often break something else, and without a test set you will only find out from customers.

Add to the set as you go. Every real mistake the agent makes after launch should become a new test case, so the same error is caught next time. Over months, the test set becomes one of the most valuable things the project owns, because it captures what “right” means for your business.

What testing cannot tell you

A test set only covers what you thought to include. Real users will find cases nobody imagined. Passing every test does not prove the agent is safe; it proves it handles the cases you tested. That is why the first weeks of real use should run with a person reviewing a large share of the agent’s work, and with clear limits on what it may do alone.

Testing also costs time. Collecting and labelling a few hundred cases may take a staff member several days. It is still far cheaper than discovering the gaps in public. A sensible first step is to pick twenty real cases from last month, write the right answer for each, and see how any proposed agent does on them before you discuss anything else.

Tell us about the work that repeats.

Send a few lines about the task, the team and the systems involved. We reply within two working days with honest next steps, even if that means not working with us.