AI evaluation means measuring whether an AI system is actually producing correct, relevant, grounded, safe, and useful results rather than assuming that a good-looking answer is a good answer.
Of the chunks retrieved, how many are actually relevant?
Example: retrieve 5 chunks, 4 are relevant → precision = 4/5 = 80%.
Of all the relevant information available, how much did we successfully retrieve?
If 5 relevant chunks exist and we retrieve 4 → recall = 4/5 = 80%.
Easy distinction: Precision = did I retrieve mostly good stuff? Recall = did I find all the good stuff?
Measures whether the retrieved context is relevant to the user's question. You may retrieve technically valid chunks, but they might not actually help answer the query.
Measures whether the retrieved context contains the information needed to answer the question. High context recall means important evidence wasn't missed during retrieval.
Checks whether the answer is supported by the retrieved context, rather than being invented by the LLM.
Very similar to faithfulness: checks whether generated claims are grounded in provided evidence/context. In interviews, you can say groundedness asks, "Can I trace this answer back to the evidence?"
Checks whether the final answer actually addresses the user's question. An answer can be factually correct but still irrelevant or unnecessarily off-topic.