Split the RAG Scorecard
Separate a RAG evaluation into retrieval relevance, groundedness, and answer relevance.
A single quality score can hide the broken part of a RAG system. The three-part scorecard A RAG answer is not one event. The system retrieves context, the model writes from that context, and the user receives an answer. Each step can fail in a different way. Retrieval relevance asks whether the returned passages are actually useful for the question. Groundedness asks whether answer claims are supported by those passages. Answer relevance asks whether the response addresses the user need instead of reciting nearby facts. The reason this works is diagnostic isolation. If the retrieved context is wrong, a better…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in