← ENGINEERING NOTES
EVALUATION · ADVISORY NOTE · 7 MIN READ

Before you ship that agent “improvement”, measure it in layers

An upgrade to a triage agent improved its decisions and quietly made it slower and less consistent. One accuracy number would have hidden both. How to set up a golden set and layered metrics so your team sees the trade-off before customers do.

For: CTOs, heads of engineering, product leadsRelated engagement: AI Reliability Audit, Managed AI Operations

The exposure

Draft — to be completed

  • Support triage must be fast, consistent and auditable; a single-pass agent fails in three ways: ambiguous tickets that fit no category, wrong tool choice, badly formed tool arguments.
  • Teams typically judge an agent change by one headline accuracy figure, or by eyeballing a few examples.

Where standard controls fall short

Draft — to be completed

WHAT TEAMS DOWHAT IT HIDESCONSEQUENCE
Single accuracy metricTrade-offs between decision quality and efficiencyShips a change that is better on hard cases and worse on cost and consistency
Ad-hoc test prompts that change each timeRegressions between versions“Improvements” that cannot be compared
One run at temperature 0 treated as truthRun-to-run drift of both the agent and the judgeDecisions made on noise
Planning metrics read at face valueA perfect score because no plan text existed to judgeFalse confidence
  • Evidence from our reference build: adding planning and self-correction raised tool choice and argument quality markedly, cut step efficiency by roughly two-thirds and slightly reduced completion.

What we recommend

Draft — to be completed

  1. Freeze a golden set of real tickets with agreed correct outcomes before any change.
  2. Score in layers rather than with one number. A practical set for an agentic workflow: prompts and instructions (does the model follow the constraints set?), planning and reasoning (is the plan sound, are the right tools chosen?), actions (tool-call success, retries, latency, error handling), context and retrieval (quality of retrieved information, grounding, citation accuracy), outcomes (task completion, factuality, safety checks, user satisfaction).
  3. Map each layer to a business KPI (routing accuracy, escalation quality, operational efficiency, decision consistency) so leadership reads the same report as engineering.
  4. Run repeatedly and report mean and spread; treat small deltas as noise until they repeat.
  5. Treat “no planning signal” as a gap, not a pass.
  6. Re-run the suite on every prompt, model or tool change — this is what Managed AI Operations does monthly.

OUR RECOMMENDATION

If your team cannot show, per layer, what the last change did, you are shipping blind. A golden set and layered scorecard take one to two weeks to stand up and pay for themselves on the first regression they catch.

Findings are from EonAI’s reference systems, built and tested on synthetic data. They describe patterns we see across clients’ systems, not a client engagement.