The exposure
Draft — to be completed
- A returns-and-refunds agent that acts has consequential failures: wrong refund, leaked record, biased decision. Trust & Safety and Legal will ask for a review before launch.
- Model vendors’ built-in safety training catches some attacks some of the time; it is not controllable and cannot be relied on.
Where standard controls fall short
Draft — to be completed
| SCENARIO | EXPOSURE | WHY PROMPT-BASED SAFETY MISSES IT |
|---|---|---|
| Direct prompt injection | Agent ignores policy | The instruction is in the same channel as the attack |
| PII exfiltration request | Customer record disclosed | Prompt says “don’t”, attacker says “do” |
| Indirect injection via tool output | Agent acts on text inside a record | Tool output treated as trusted |
| Frustrated legitimate customer | Wrongly refused or blocked | A single toxicity threshold cannot tell angry from malicious |
| Over-refund attempt | Money out | No cap enforced where the action happens |
What we recommend
Draft — to be completed
Guardrails as workflow nodes the agent cannot bypass.
- Input checks run in order before the agent: PII redaction (redact and continue), injection classifier (fail closed), toxicity (route to a person), topic (fail open and log).
- Output check the agent cannot skip: the agent has no path to the customer except through it.
- Refund cap enforced inside the tool on the server, not in the prompt.
- Principle: neural components score, symbolic components decide and enforce.
- Every check has a written failure policy (what happens on error or trigger) agreed with Trust & Safety before launch.
- Fairness tested as behaviour: identical claims with different customer names must produce identical decisions.
- A decision trace that reconstructs which check made which call.
OUR RECOMMENDATION
Ask your team to show you the workflow diagram and point to the node that stops a bad output reaching a customer. If the answer is “the prompt”, you do not yet have a guardrail.
Findings are from EonAI’s reference systems, built and tested on synthetic data. They describe patterns we see across clients’ systems, not a client engagement.