The exposure
Draft — to be completed
- Agent loops make two kinds of calls: reasoning calls (understand context, plan, write a response) and decision calls (route, classify, check, pass/fail). Practitioner estimates put decision calls at the large majority of volume.
- Almost all are served by a frontier model at frontier prices, multiplied by every iteration the agent runs. Cost compounds; latency does too.
- Cheaper, faster classification models (small models, fine-tuned classifiers, and newer calibrated decision models such as TypeSafe’s Jev, one example among several) can take the decision calls, but only if the architecture is set up to use them safely.
Where standard controls fall short
Draft — to be completed
| PRACTICE | PROBLEM | CONSEQUENCE |
|---|---|---|
| One frontier model for every call | Paying reasoning prices for yes/no decisions | Cost and latency scale with every loop iteration |
| Bundled questions (“is this ticket a good automation candidate?”) | The model judges a bundle of three checks as one and returns a plausible but low-confidence score | Weighting hidden in the prompt rather than in code |
| One confidence threshold for every action | A read-only lookup and an automated refund treated the same | Either too timid or too dangerous |
| Routing accuracy never measured | Misrouting rarely throws an error, it just produces a worse result | Drift goes unnoticed |
| Classifier as the approver | Text in context can sway a probability | A manipulated classification becomes an executed action |
What we recommend
Draft — to be completed
- Audit your traces: count reasoning calls versus structured decisions. The split tells you how much cost is movable.
- Route the four decision-heavy places (routing, pre-execution guardrails, output evaluation, triage classification) to a cheap calibrated classifier, with fallback to the frontier model when confidence is low. A wrong decision on the cheap side compounds; an extra frontier call costs cents.
- Keep permission checks in code. The classifier recommends; the code decides. Nothing can talk an if-statement out of its answer.
- Ask single questions and combine the results in code, where the weighting is visible and A/B-testable.
- Match the confidence threshold to the cost of a mistake: a lookup can act at 0.5, an automated refund might need 0.9, below which it goes to a person or a heavier model.
- Measure routing accuracy against a hand-labelled set, continuously; it is classification, and classifiers drift.
- Treat classifier input as untrusted: these models run alongside existing security checks, never instead of them.
OUR RECOMMENDATION
Audit one week of agent traces. If more than half the calls are decisions, you are overpaying for them and probably under-measuring them. The fix is architectural, takes a few weeks, and usually pays for itself within a quarter.
Findings are from EonAI’s reference systems, built and tested on synthetic data. They describe patterns we see across clients’ systems, not a client engagement.