← ENGINEERING NOTES
KNOWLEDGE ASSISTANTS · ADVISORY NOTE · 6 MIN READ

A fluent answer is not a grounded answer

A knowledge assistant can pass every task check and still answer from nothing. Why groundedness needs its own measure, why the judge must be independent of the writer, and how confidence routing keeps weak answers away from your clients.

For: CTOs, heads of research/knowledge, complianceRelated engagement: AI Reliability Audit, Agent MVP

The exposure

Draft — to be completed

  • A tourism website published AI-generated content describing hot springs that do not exist, and visitors travelled to find them. The business lost credibility in a single news cycle. Hallucinations are not a theoretical problem; they are preventable with evaluation before publication.
  • Assistants answering from documents (legal, research, policy) are trusted because they sound right. A confident wrong answer in a client deliverable is a reputational and regulatory event.
  • Early “paste the documents into an LLM” pilots are fast and become liabilities for exactly this reason.

Where standard controls fall short

Draft — to be completed

PRACTICEWHAT IT HIDESCONSEQUENCE
Task-completion checks onlyAn out-of-scope question answered fluently with zero retrieval scored 5/6Ungrounded answer passes
Same model writes and gradesSelf-assessment biasInflated quality scores
No scope guardAssistant answers questions outside its knowledge baseWrong domain, confident tone
No citation requirementClaims cannot be tracedReviewers cannot verify
All answers treated equallyLow-confidence answers reach clientsNo safety valve

What we recommend

Draft — to be completed

  1. Scope guard that returns “out of scope” before any retrieval.
  2. Answer-only-from-context contract: say “I don’t know” when the documents don’t support an answer.
  3. Independent judge model scoring groundedness and citation precision against a gold set.
  4. Confidence-routed outputs: low confidence goes to a person, never silently to a client.
  5. Prompt improvement driven by failure feedback and validated on held-out questions.
  6. Groundedness and relevance tracked over time; drift triggers review.

OUR RECOMMENDATION

Ask for two numbers before you trust an assistant: how often it answers from the documents, and how often it says it doesn’t know when it should. If nobody can give you those, it has not been evaluated.

Findings are from EonAI’s reference systems, built and tested on synthetic data. They describe patterns we see across clients’ systems, not a client engagement.