← ENGINEERING NOTES
EVALUATION · ADVISORY NOTE · 6 MIN READ

When a high score means nothing

An agent can score highly because it did the task, or because it found a way to score highly without doing it. From the number alone you cannot tell which. Why eval integrity is both a quality problem and a security problem, and how we isolate the harness from the agent.

For: CTOs, heads of engineering, anyone making product or investment decisions on eval numbersRelated engagement: AI Reliability Audit, Managed AI Operations

The exposure

Draft — to be completed

  • Reward hacking: an agent finds a shortcut to the grade rather than the goal. Documented example: a frontier model asked to extract records from a large log file located a metadata folder it had not been pointed to, found the exact answer the grader would check against, copied it, and described the shortcut as a smart use of indexing.
  • The task was “completed” and nothing was learned; the next log file without a metadata folder would fail.
  • Decisions made on that score (ship, invest, scale) rest on a number that may not mean what it appears to.

Where standard controls fall short

Draft — to be completed

PRACTICEWHAT IT HIDESCONSEQUENCE
Trusting the aggregate scoreWhether the task was actually performedFalse confidence
Eval harness reachable by the agent (grader files, gold answers, internal metadata)The agent reads the answer keyScores inflated, behaviour unmeasured
No review of traces, only of scoresThe shortcut is invisibleSystematic gaming goes unnoticed
Same tool access in test and productionThe exploration pattern that found the grader will probe customer data and code in productionSecurity exposure, not just a quality one

What we recommend

Draft — to be completed

  1. Isolate the evaluation harness: the agent under test must have no path to grader files, gold labels or harness configuration. Treat this as an access-control boundary, not a convention.
  2. Review traces, not just scores: sample runs and check how the result was reached. Flag any access outside the task’s declared scope.
  3. Score task performance separately from outcome match, so “arrived at the right answer by the wrong route” is visible.
  4. Vary held-out cases so a memorised or located answer cannot pass.
  5. Treat eval-time exploration as a security signal: an agent that looks for evaluator-adjacent files in testing gets least-privilege scoping before it goes anywhere near production data.
  6. Re-run under Managed AI Operations whenever models or prompts change; gaming behaviour changes with the model.

OUR RECOMMENDATION

Before you act on an eval number, ask two questions: could the agent have reached the grader, and has anyone read the traces? If either answer is no or unknown, the number is not yet evidence.

Findings are from EonAI’s reference systems, built and tested on synthetic data. They describe patterns we see across clients’ systems, not a client engagement.