ENGINEERING NOTES

Practical guidance for teams putting AI agents into production.

Short advisory notes on the ways agentic systems fail in production and the controls that prevent it. Each one is written for the person accountable for the decision: the exposure, where standard controls fall short, what we recommend, and what it would take.

Findings are drawn from EonAI’s reference systems, built and tested on synthetic data. They describe patterns we see across clients’ systems, not individual engagements.

SECURITY & GUARDRAILS· 9 MIN READ · FEATURED

If your agent can issue refunds, a content filter is not protecting you

Six scenarios that read as ordinary customer traffic and cost money, including split refunds under the cap and a six-character edit to a policy document. The seven controls we recommend, and why the two cheapest close the most expensive exposures.

Read the note →
EVALUATION· 7 MIN READ

Before you ship that agent “improvement”, measure it in layers

Adding planning to a triage agent improved its decisions and quietly made it slower and less consistent. One accuracy number would have hidden both. How to set up a golden set and layered metrics so your team sees the trade-off before customers do.

Read the note →
SECURITY & GUARDRAILS· 8 MIN READ

Guardrails are topology, not prompts

Safety that lives in a system prompt is a request. Safety that lives in the workflow the agent cannot bypass is a guarantee. What a guardrail architecture looks like, and the failure policy each check needs before launch.

Read the note →
MULTI-AGENT SYSTEMS· 8 MIN READ

When money is involved, keep the orchestrator deterministic

How to structure a multi-agent review so every agent shows its evidence, a critic catches contradictions, insufficient evidence never becomes an approval, and your auditors get a complete trail. Illustrated with insurance claims.

Read the note →
KNOWLEDGE ASSISTANTS· 6 MIN READ

A fluent answer is not a grounded answer

A knowledge assistant can pass every task check and still answer from nothing. Why groundedness needs its own measure, why the judge must be independent of the writer, and how confidence routing keeps weak answers away from your clients.

Read the note →
RESPONSIBLE AI· 5 MIN READ

Fairness is a test, not a policy

If an automated screening decision moves when only the candidate’s name changes, you have a regulatory exposure, whatever your policy says. The matched-pair tests, abstain rules and audit trail we recommend before any such system goes live.

Read the note →
METHOD· 4 MIN READ

Why we build on your real data from week one

An agent that commits to a plan up front breaks the moment reality differs from the plan. One that checks each result and adjusts keeps going. The same is true of projects, which is why we never build on sample data.

Read the note →
EVALUATION· 6 MIN READ

When a high score means nothing

An agent can score highly because it did the task, or because it found a way to score highly without doing it. From the number alone you cannot tell which. Why eval integrity is both a quality problem and a security problem, and how we isolate the harness from the agent.

Read the note →
COST & ARCHITECTURE· 7 MIN READ

Most of your agent’s model calls are decisions, not reasoning

In a typical agent loop, the large majority of model calls decide something (route this, is that tool call safe, how urgent is this ticket, did this output pass) rather than reason about it. Most teams pay frontier-model prices for all of them. How to separate the two, cut cost and latency, and keep control flow where it belongs.

Read the note →
Video thumbnail for the talk The Future of Quality in AI-Generated Software
TALK· TEST DRIVE PLATFORM · VIDEO

The Future of Quality in AI-Generated Software

A guest talk by EonAI’s CTO on what changes in testing and release practice when a large share of your code is written by AI, and what to do about it.

Watch on YouTube →