Evaluating an agent: tests that should pass before real work
A practical set of tests every AI agent should pass before it touches customers, money or records, and how to build them from your own cases.
Insights
Research notes on multi-agent systems, agent reliability, regulated finance and Sinhala and Tamil language agents.
A practical set of tests every AI agent should pass before it touches customers, money or records, and how to build them from your own cases.
Researchers studied software that bargains long before chat models. As agents start buying and selling, their ideas matter again.
When several AI agents share one job, they duplicate work, clash and loop. Here are the coordination patterns that prevent it.
Long before large language models, researchers studied software agents that plan, talk and share work. Here is what still applies.