FDE01Forward deployment
Handbook 12

Five areas to check before release

A better overall score may still leave tasks unfit for release.

On this pageChapter guideDiscussion questionRelated exerciseReferences

Chapter guide

Ask domain experts to describe failures, group them and then choose metrics. Keep disagreements and independent review results. Find the earliest failure in the execution trace instead of checking only the final answer.

Check task outcomes, system behavior, user experience, risk and operational support separately. Review task or user groups, often called evaluation slices. Good performance elsewhere cannot offset a consequential failure.

For agents, inspect multi-turn traces, final business state and consistency across repeated tests. State permitted tasks, exclusions, monitoring, stop owners and retest conditions. Narrow the scope when evidence is insufficient.

Five areas to check before release
Five areas to check before releaseCheck each area separately. Overall success cannot cancel a serious failure. Review results by task and user group so the authorized owner can decide on release.Task outcomesSystem behaviorUser experienceRiskOperational supportContinue / limit scope / stop
Check each area separately. Overall success cannot cancel a serious failure. Review results by task and user group so the authorized owner can decide on release.Swipe to view the full diagram.

Discussion question

Which task or user groups could perform worse while the overall score improves?

Related exercise

Follow the related lesson below and try its exercise. Use the linked template to record your judgment, evidence and open questions, then revise with the checklist. With a partner, check which parts of each other’s work need clarification.

References

Anthropic / Demystifying evals for AI agents