1. AI Evaluation

AI evaluation means measuring whether an AI system is actually producing correct, relevant, grounded, safe, and useful results rather than assuming that a good-looking answer is a good answer.

RAG Metrics

1. Retrieval Precision

Of the chunks retrieved, how many are actually relevant?

Example: retrieve 5 chunks, 4 are relevant → precision = 4/5 = 80%.

2. Retrieval Recall

Of all the relevant information available, how much did we successfully retrieve?

If 5 relevant chunks exist and we retrieve 4 → recall = 4/5 = 80%.

Easy distinction: Precision = did I retrieve mostly good stuff? Recall = did I find all the good stuff?

3. Context Relevance

Measures whether the retrieved context is relevant to the user's question. You may retrieve technically valid chunks, but they might not actually help answer the query.

4. Context Recall

Measures whether the retrieved context contains the information needed to answer the question. High context recall means important evidence wasn't missed during retrieval.

5. Faithfulness

Checks whether the answer is supported by the retrieved context, rather than being invented by the LLM.

6. Groundedness

Very similar to faithfulness: checks whether generated claims are grounded in provided evidence/context. In interviews, you can say groundedness asks, "Can I trace this answer back to the evidence?"

7. Answer Relevance

Checks whether the final answer actually addresses the user's question. An answer can be factually correct but still irrelevant or unnecessarily off-topic.