← ENGINEERING NOTES
SECURITY & GUARDRAILS · ADVISORY NOTE · 8 MIN READ

Guardrails are topology, not prompts

Safety that lives in a system prompt is a request. Safety that lives in the workflow the agent cannot bypass is a guarantee. What a guardrail architecture looks like, and the failure policy each check needs before launch.

For: CTOs, heads of product, trust & safety, legalRelated engagement: AI Reliability Audit, Agent MVP

The exposure

Draft — to be completed

  • A returns-and-refunds agent that acts has consequential failures: wrong refund, leaked record, biased decision. Trust & Safety and Legal will ask for a review before launch.
  • Model vendors’ built-in safety training catches some attacks some of the time; it is not controllable and cannot be relied on.

Where standard controls fall short

Draft — to be completed

SCENARIOEXPOSUREWHY PROMPT-BASED SAFETY MISSES IT
Direct prompt injectionAgent ignores policyThe instruction is in the same channel as the attack
PII exfiltration requestCustomer record disclosedPrompt says “don’t”, attacker says “do”
Indirect injection via tool outputAgent acts on text inside a recordTool output treated as trusted
Frustrated legitimate customerWrongly refused or blockedA single toxicity threshold cannot tell angry from malicious
Over-refund attemptMoney outNo cap enforced where the action happens

What we recommend

Draft — to be completed

Guardrails as workflow nodes the agent cannot bypass.

  1. Input checks run in order before the agent: PII redaction (redact and continue), injection classifier (fail closed), toxicity (route to a person), topic (fail open and log).
  2. Output check the agent cannot skip: the agent has no path to the customer except through it.
  3. Refund cap enforced inside the tool on the server, not in the prompt.
  4. Principle: neural components score, symbolic components decide and enforce.
  5. Every check has a written failure policy (what happens on error or trigger) agreed with Trust & Safety before launch.
  6. Fairness tested as behaviour: identical claims with different customer names must produce identical decisions.
  7. A decision trace that reconstructs which check made which call.

OUR RECOMMENDATION

Ask your team to show you the workflow diagram and point to the node that stops a bad output reaching a customer. If the answer is “the prompt”, you do not yet have a guardrail.

Findings are from EonAI’s reference systems, built and tested on synthetic data. They describe patterns we see across clients’ systems, not a client engagement.