TemplatesGet a demo →Book a meeting
Blog
Company BrainGovernance

RAG Evaluation Metrics: How to Score an Answer for Accuracy, Sources and Trust

In short

The most practical RAG evaluation metrics are the ones a business reviewer can apply: is the answer correct, complete, sourced, honest about what the documents don't cover, and useful enough to act on. Scoring a sample of real answers on those five points turns a pilot from an impression into evidence a sponsor and a compliance reviewer can both read.

Key takeaways

  • "It seemed good" is not a result — a simple scoring sheet makes a pilot comparable and defensible.
  • Score each answer 0–2 on five points: correct, complete, sourced, honest and useful.
  • Track honesty separately: a system that admits gaps is safer than one that always answers.
  • Two reviewers scoring the same answers show where "good" still needs defining.
  • Always include the questions users flagged as poor, not just the ones that went well.

"It seemed good" is not a result. At the end of a RAG pilot, a sponsor wants to know whether to go ahead, and a compliance reviewer wants to know whether the answers can be trusted. Neither question is answered by an impression. A simple, shared set of RAG evaluation metrics — applied by people, to real answers — makes a pilot comparable, repeatable and defensible.

Which RAG evaluation metrics matter most?

There are many technical metrics for retrieval and generation. For a business pilot, five checks cover what decision-makers actually care about. Score each answer from 0 to 2 on each point.

Check Question to ask Score
Correct Does the answer match what the source document says? 0 to 2
Complete Is anything important missing? 0 to 2
Sourced Does it show the source, and does the source support it? 0 to 2
Honest When the documents do not cover it, does it say so? 0 to 2
Useful Would the user act on this answer? 0 to 2

Correct

The answer matches what the source says — not what a general model believes. A fluent answer that contradicts your own policy scores zero, however well written.

Complete

Nothing important is missing: the exception, the threshold, the second step. Incomplete answers are often more dangerous than wrong ones, because they look finished.

Sourced

The answer points to a document, and that document actually supports it. A citation that does not back the claim is worse than no citation, because it lends false confidence. Cited answers are central to how a Company Brain is meant to work.

Honest

When the documents do not cover the question, the system says so instead of guessing. This is where many systems fail quietly.

Useful

Would the person who asked act on the answer? A correct, sourced answer that is too vague to use still fails the user.

Why it matters

A system that admits gaps is safer than one that always answers. Scoring honesty on its own shows whether the system guesses when it should not.

How do you use the scoring sheet?

  • Score a sample, not everything. Pick a representative set of questions, and always include some that users flagged as poor.
  • Use two reviewers. Have two people score the same answers and compare. Where they disagree is exactly where "good" still needs defining for your content.
  • Track "honest" separately. Report it on its own line, not only inside the total.
  • Ask one closing question. At the end of the pilot, ask each user: would you use this again?

Why does scoring matter in a pilot?

In a pilot, real users ask their own questions of a chat built on a small sample of their documents — the approach described in try before you build. A scoring sheet turns that activity into evidence both a sponsor and a compliance reviewer can read. It also makes options comparable: if you are weighing build vs. buy, the same sheet applied to each option shows the difference in numbers rather than opinions.

The five points above are a starting suggestion. Adapt them to your own standards — a regulated team may weight "sourced" and "honest" more heavily, while an operations team may care most about "useful".

What should happen after scoring?

  • Low "correct" or "complete": look at the documents first. Out-of-date or conflicting versions are the usual cause.
  • Low "sourced": check whether the right documents were included in the sample at all.
  • Low "honest": treat it as a risk to fix before any wider rollout.
  • Low "useful": talk to the users; the answer may be right for a different role than the one asking.

Scoring also connects directly to governance. Every answer in SphereIQ is logged with the sources it drew on, so a reviewer can trace a low score back to the exact passage — the kind of record described in what an AI audit trail should prove.

Frequently asked questions

What are the most important RAG evaluation metrics?
For a business pilot: correctness against the source, completeness, whether the answer shows a source that supports it, honesty when the documents do not cover the question, and whether the user would act on the answer.
How many answers should we score?
A representative sample rather than all of them. Include a mix of everyday questions and the ones users flagged as poor, so the score reflects real use.
Do we need automated evaluation tools?
Not to start. A shared scoring sheet with two human reviewers is enough for a pilot and produces results a sponsor can understand. Automated evaluation helps later, once you know what "good" means for your content.
Why score honesty separately?
Because a system that says "the documents don't cover this" is safer than one that always produces an answer. Tracking it on its own shows whether the system guesses when it should not.

Score real answers on your own content.

In a walkthrough we run SphereIQ on a sample of your documents and review cited answers against this scoring sheet with your team.