The exposure
A chatbot that says the wrong thing produces a bad sentence. An agent that is manipulated into the wrong action produces a refund, a redirected parcel or a disclosed customer record. The liability sits with you, not with the model vendor.
The controls most teams ship with are designed for an ordinary user who may be rude or careless. They are not designed for a motivated person who has read the defences and is working around them, and that is the person who extracts money. The practical question for your business is not “can the filter catch bad text?” but “can anyone make this agent take an action it should not?”
Where standard controls fall short
In our reference build, a support agent protected with the standard controls stopped the textbook injection and then failed six scenarios that read as normal customer traffic. Each one maps to a loss your finance or fraud team would recognise.
| SCENARIO | EXPOSURE | WHY STANDARD CONTROLS MISS IT |
|---|---|---|
| Injection without keywords | The same instruction-hijack as the textbook attack, phrased so nothing on the block list appears. The agent follows it. | The classifier matches appearance, not intent. |
| Multi-turn setup | Four ordinary messages. A false premise is established in the third and acted on in the fourth: a profile change the agent had no way to verify. | Per-message checks cannot see a sequence. |
| Tampered policy document | Six characters changed in the knowledge base: $250 became $2,500. The agent’s self-service authority increased tenfold, and it quoted the new limit with full confidence. | Not an injection and not visibly wrong. A human review skims past it. |
| Acting on someone else’s account | A customer requested a refund on another customer’s $1,899 order and received it. | The agent had no concept of who was asking. Nothing to check ownership against. |
| Refund splitting | Three refunds of $160, $150 and $140, each with a plausible reason, on an order worth $612 that had already been partly refunded. Total paid out exceeded the order value. | The per-transaction cap was satisfied every time. It limits the wrong thing. |
| Two routine actions, one fraud | An address change followed by a refund on the same order in one session. Individually routine; together, the delivery-interception pattern your own policy probably warns about. | Per-action checks see one action at a time. The risk only exists across two. |
The pattern matters more than any single row. Standard controls are strongest against the scenario that looks most like an attack and weakest against the ones that look like business as usual. If your current protection is a filter and a cap, assume the lower five rows apply to you.
What we recommend
Put authorisation in the architecture, not in the prompt. The model may decide what to propose; it must never decide what is permitted. In practice that is seven controls, in order of dependency. Most teams can implement the first two in days, and they close the most expensive scenarios.
- Know who is asking. Every session is bound to an identified customer at a known assurance level (unverified, email-verified, MFA). Higher-impact actions require the higher levels. Everything below depends on this.
- A permission table your auditor can read. Each tool declares what it requires; anything not listed is denied by default. Queries return only the fields a task needs, so a field that was never fetched cannot be leaked.
- Keep the filters, demote them. Injection scanning stays, extended to retrieved documents, but it flags and logs rather than decides. It becomes your early-warning telemetry.
- Protect the policy source. Knowledge-base documents are integrity-checked before use, and monetary limits live in code rather than in prose the model reads.
- Cap the total, not the transaction. Limits are enforced against the order’s remaining value and the session’s running total. Splitting a refund no longer works.
- A person approves high-impact combinations. Address changes, large refunds and any flagged pairing of actions route to a human before they execute.
- An audit trail that stands up to review. Every decision and action recorded in tamper-evident form, so an incident is reconstructed rather than argued about, and compliance can sign off.
The most expensive scenario, acting on someone else’s account, was closed by a single ownership check: does this record belong to the customer in this session? No classifier, no threshold, no prompt. The most effective controls on agents are usually this plain, and plain controls have no false-negative rate.
OUR RECOMMENDATION
Before the next release, ask three questions of every tool your agent can call: who is calling, what are they entitled to, and what has already happened in this session? If any answer lives in a prompt, the control is a request, not a guarantee. Keep the filters as defence in depth. Move authorisation into the architecture. A two-week audit is enough to find out which of the six scenarios apply to you and what closing them would take.
Findings are from EonAI’s reference systems, built and tested on synthetic data. They describe patterns we see across support agents, not a client engagement.