← ENGINEERING NOTES
MULTI-AGENT SYSTEMS · ADVISORY NOTE · 8 MIN READ

When money is involved, keep the orchestrator deterministic

How to structure a multi-agent review so every agent shows its evidence, a critic catches contradictions, insufficient evidence never becomes an approval, and your auditors get a complete trail. Illustrated with insurance claims.

For: CTOs, heads of operations, claims/finance leadersRelated engagement: Agent MVP, AI Opportunity Sprint

The exposure

Draft — to be completed

  • Claim and document review is costly and inconsistent when manual, brittle when done by fixed rules; partial evidence is common.
  • “Multi-agent” is often three parallel LLM calls with no role boundaries, no evidence contract and no way to explain a decision.

Where standard controls fall short

Draft — to be completed

PRACTICEWHAT IT HIDESCONSEQUENCE
One LLM asked to “assess the claim”No evidence trailDecision cannot be defended
Parallel calls with no criticContradictions between findings go unnoticedWrong approvals
LLM as orchestratorRetries and routing become unpredictableCost and audit problems
No quorum ruleA missing worker result becomes an implicit approvalMoney out on no evidence
Document inconsistencies (header vs body claim ID, invoice amount vs claim amount)Missed unless a worker is tasked with themFraud passes

What we recommend

Draft — to be completed

  1. Specialist workers with explicit responsibilities and output contracts (“risk signal + evidence-backed notes”), each with stated limits (e.g. “never infer risk from frequency alone”).
  2. Tools accessed through a standard protocol boundary (MCP) so data access is governed and logged.
  3. A critic that compares findings, flags contradictions and evidence gaps, and may request exactly one retry.
  4. A deterministic orchestrator (code, not a model) that dispatches, enforces retry caps and applies the verdict.
  5. A quorum rule: insufficient evidence falls back to deny or to human review, never to approve.
  6. An append-only audit trail per claim, readable by compliance.
  7. Human review branch for high-value or contested decisions.

OUR RECOMMENDATION

If the thing deciding which agent runs next is itself a model, you cannot promise your auditor a repeatable process. Keep the orchestrator boring.

Findings are from EonAI’s reference systems, built and tested on synthetic data. They describe patterns we see across clients’ systems, not a client engagement.