Chapter guide
Ask domain experts to describe failures, group them and then choose metrics. Keep disagreements and independent review results. Find the earliest failure in the execution trace instead of checking only the final answer.
Check task outcomes, system behavior, user experience, risk and operational support separately. Review task or user groups, often called evaluation slices. Good performance elsewhere cannot offset a consequential failure.
For agents, inspect multi-turn traces, final business state and consistency across repeated tests. State permitted tasks, exclusions, monitoring, stop owners and retest conditions. Narrow the scope when evidence is insufficient.
Discussion question
Which task or user groups could perform worse while the overall score improves?
Related exercise
Follow the related lesson below and try its exercise. Use the linked template to record your judgment, evidence and open questions, then revise with the checklist. With a partner, check which parts of each other’s work need clarification.