FDE01Forward deployment
Course 08

Evaluate the system and decide whether to release

How do you decide whether the system is ready for users?

On this pageBy the end of this lessonBefore you beginDefine completion before choosing a scoreRead examples and build a task-specific failure listReview five areas without allowing one to offset anotherUse task groups to expose failures hidden by averagesWorked example: turn “94%” into a specific trial recommendationUnderstand the gap between a lab pass and readiness for usersWrite the decision for the people who must follow itPractice the methodCheck your understandingYour record and next stepReferences

By the end of this lesson

  • Build examples and criteria from task outcomes and failure consequences instead of copying a default metric.
  • Review task outcomes, system behavior, user experience, risk and operational support separately, including counts by task group.
  • Recommend pausing, further investigation, shadow mode or a restricted trial with supporting evidence.

Before you begin

  • Bring Lesson 4’s success criteria, Lesson 5’s workflow, Lesson 6’s tests and Lesson 7’s risk and ownership records.
  • Use the public Beichen evaluation event and the lab evidence. They are different materials and must not be combined into a single real-world evaluation report.

Read the Beichen case and introductory data

Five areas to check before release
Five areas to check before releaseCheck each area separately. Overall success cannot cancel a serious failure. Review results by task and user group so the authorized owner can decide on release.Task outcomesSystem behaviorUser experienceRiskOperational supportContinue / limit scope / stop
Check each area separately. Overall success cannot cancel a serious failure. Review results by task and user group so the authorized owner can decide on release.Swipe to view the full diagram.

Define completion before choosing a score

Evaluation decides what may happen next. For a parts draft, success is not the model saying “done.” The draft needs correct, traceable information, explicit missing fields, a user who understands that it has not been submitted, and behavior within authorized scope. A future submission test also needs the official business record, not merely the closing response.

For each example, specify inputs, equipment and source conditions, user role, allowed action, expected outcome and unacceptable consequences. Make the requirements consistent. A draft-only task should not fail because no request was submitted. When information is insufficient, appropriate escalation should not count as an error simply because it did not produce a complete recommendation.

Anthropic distinguishes tasks, individual trials, traces, grading and final state. That helps you retain records that explain a result. Beichen still needs domain judgment: an authorized tool trace does not establish that a manual applies, and a correct answer does not establish available overnight support.

Read examples and build a task-specific failure list

Inspect the request, material, advice, user action and final record one example at a time. Describe the problem in ordinary language. Start with the earliest failure that explains the result; retain downstream consequences without counting every consequence of one upstream fault as an independent failure. Group similar observations, then decide what can be checked automatically and what needs judgment.

Beichen’s list might include an inapplicable version, missed escalation, an unknown state reported as complete, unauthorized actions, correct but unusable information and excessive editing effort. For each, record affected people, detectability, recoverability, required expertise and whether it blocks a trial. Generic lists help find omissions; they do not replace observed task failures.

Deterministic checks suit explicit fields, access and states. Model graders can assist with open-ended expression, but need calibration against qualified human judgment and an “insufficient information” option. Missing equipment context needs more information; conflicting authority needs an owner’s decision; differing severity scales need shared examples. Voting does not resolve every kind of disagreement.

Review five areas without allowing one to offset another

Organize the evidence into five areas: task outcomes, system behavior, user experience, risk and operational support. For each, record sources, sample counts, applicable conditions, gaps and an owner. They answer different questions. A high overall score cannot compensate for an unmet release condition.

Risk connects to the preceding lesson: permitted data uses, tested access and withdrawal controls, unacceptable consequences and authorized judgment. Operational support asks whether someone can actually take over, how pausing works, what happens to queued tasks and when service can resume. A signature without tested controls, or a named reviewer without available time, is not sufficient evidence.

Lesson 11 evaluates business value using adoption, cost and causal evidence. Here, retain task duration, review effort and support costs as later inputs, without deriving annual benefits from a score or a few minutes saved. Readiness for a trial and justification for greater investment are connected but distinct decisions.

Review areaBeichen exampleInsufficient substitute
Task outcomesAn accurate, complete draft understood as not submittedFluent wording or an acceptance click
System behaviorVersions, access, tool states and official records agreeThe final answer alone
User experiencePeople understand, edit, reject and find helpTraining completion or satisfaction
RiskConsequential scenarios have tested controls and authorized judgmentAn average score or generic sign-off
Operational supportUsable capacity, takeover, pausing and recoveryA name on a duty roster

Use task groups to expose failures hidden by averages

Group tasks by conditions that affect outcomes, such as day versus overnight work, old versus current equipment, or complete versus missing information. These groups are often called evaluation slices. Their purpose is a decision: a serious failure in one group can require exclusion, more evidence or a pause, regardless of success elsewhere.

Report the number passed and the total tested—the denominator. “Two failures among four tasks involving older equipment models” is more useful than an overall 94%, because it shows both limited coverage and a serious issue. Groups can overlap; a task may be both overnight and associated with an older equipment model. Do not add overlapping group counts as if they were independent tasks.

Keep common tasks while including rejection, missing information, access changes, withdrawn material and important minority cases. Reused development examples detect regressions; keep a separate set of tasks that were not used for tuning to check performance on unfamiliar examples. Record every trial and reset prior state rather than keeping only the best answer or mistaking shared-cache advantages for capability.

Worked example: turn “94%” into a specific trial recommendation

Beichen’s public event reports a 94% overall pass rate, six percentage points above the previous round; two failures among four tasks involving older equipment models; dangerous advice using a procedure intended for a different equipment model in an overnight high-temperature task; and half the planned review capacity. These are new synthetic scenario facts, not statistics from the introductory 20-record JSON. The overall denominator is not supplied, so do not infer an exact total.

The event also says the failures on older equipment account for 1.5 percentage points of the overall result, but supplies too little information to verify its denominator or rounding. Retain that limitation and use the checkable 2/4 count, specific dangerous advice and capacity gap. Given the summary, the prior rate was 88%; a six-percentage-point increase is different from a six-percent increase.

These facts do not establish that every day-shift task is safe. A restricted trial needs separate sufficient evidence for its proposed day-shift tasks involving current equipment models and low-consequence actions. Without it, recommend further testing or a synthetic demonstration. Excluding older equipment models also does not repair filtering by itself; verify that the exclusion is actually enforced.

  • Rewrite the summary: give the dangerous advice, the two failures among four older-equipment tasks, reduced review capacity and evidence gaps priority over the overall score.
  • Inspect the task: compare equipment inputs, cited versions, system checks and visible information to locate the earliest unmet requirement.
  • Review all five areas: distinguish records from plans and specialist judgments still needed. Do not label unknowns as passed.
  • Compare options: pause real trials and gather evidence, retain a synthetic demonstration, or discuss a restricted scope after sufficient new evidence. Explain their different costs.
  • Write the recommendation: identify people, tasks, material, shifts and action scope, with exclusions, stop conditions, actual takeover, an observation window and review evidence.
  • Handle the earlier deadline by reducing the demonstration, retaining the original process or delaying real use—not by weakening the risk conditions.

Understand the gap between a lab pass and readiness for users

The lab evidence reports 28/28 passes across groups including authorization, freshness, retrieval, tools and resilience. It uses a deterministic backend, makes no network requests and reuses development examples. It establishes that checked regressions did not occur in those known scenarios, not that real model behavior, users or external-service failures are covered.

The freshness group contains only two examples. That indicates limited coverage, not low importance. Add cases for withdrawal during propagation, old sessions after reassignment, and intersections of access and version changes. Counts depend on the risk question and required evidence. Neither 28/28 nor additional synthetic tests establish a universal release threshold or real-user validation.

Record versions for the model, prompt, retrieved material, tools, access rules and runtime environment, plus the test time. After a change, rerun tests affected by it and consider connected behavior. Evaluation continues: observed user failures become regression cases, while new task scope needs new judgment rather than permanent reliance on the first passing report.

Write the decision for the people who must follow it

A decision is more than release versus no release. It can stop work, seek more evidence, remain on paper or in shadow mode, restrict scope, or start a small pilot when supported. Shadow mode still needs authorized data and clear separation from live decisions. No external action does not mean no risk.

An actionable record tells people who may use it, for what, what remains excluded, when and how to reach support, what stops work and which evidence is reviewed next. “East-region, day-shift, low-consequence drafts” names a scope but does not justify it. Without supporting evidence and enforced boundaries, it is a proposal to verify, not an approved operating state.

When expanding, prefer one boundary at a time—for example, add one task while holding role, region, shift and action scope steady—to make new failures easier to investigate. If several conditions must change together, state the need for intersection testing and monitoring instead of promising clear attribution. Ask the people using the decision to restate it. Alone, read it as both engineer and operations owner and find anything left to guess.

Practice the method

Exercise 08

Prepare a Beichen trial recommendation without operating real equipment. Choose 8–12 synthetic tasks for this learning exercise. That count makes the exercise manageable; it is not a formal acceptance requirement.

  1. 01

    Specify inputs, permitted actions, expected outcomes, source conditions and unacceptable errors for each example. Explain selection.

  2. 02

    First review each example and describe problems in your own words, then group similar failures. If a partner is available, compare independent judgments on a subset and retain disagreements and their causes.

  3. 03

    Group by at least two relevant conditions, reporting passed/total and specific failures. Explain overlaps and uncovered tasks.

  4. 04

    Fill the five review areas separately. Where evidence is missing, say what is needed and who can establish it.

  5. 05

    Introduce the 94% event and capacity reduction. Update the pause, synthetic-demo, further-testing or restricted-trial recommendation without inventing the missing denominator.

  6. 06

    Record scope, exclusions, pausing and restoration, observation window, owners and next-review evidence. Mark specialist judgments such as actual equipment applicability as pending confirmation.

Use individual results to assess readiness

Calculate overall and grouped results for 52 independent synthetic examples, then check task outcomes, system behavior, user experience, risk and operational support.

These new examples do not establish the denominator of the original 94% summary or replace real-user, professional and handover evidence.

Read the field guide, work through the calculation or walkthrough, then check the guide's review notes. Open the downloaded files in a spreadsheet or text editor; no code is required.

Check your understanding

Write your answer before opening the explanation, then check what you might have missed.

Does a 94% overall result justify trials on older equipment when two of four tasks involving older equipment models fail?

The aggregate cannot offset this group. Inspect consequences and controls and gather appropriate evidence. Uncontrolled dangerous advice blocks the relevant task; it also does not establish that every other group passed.

Why check the official record if the model says “done” and the trace contains a tool call?

A call can fail or leave the outcome unknown, and the interface can misread its state. The authoritative record establishes the result. The initial task only prepares a draft; a call is not evidence of submission.

Does a 28/28 lab result cover all five review areas?

No. It checks selected behavior in known synthetic scenarios. User understanding, specialist risk judgment, actual support capacity and external dependencies still need their own evidence.

Your record and next step

A failure list and results by group, including reviewer disagreements, the five areas of evaluation and a release recommendation.

Each release condition has supporting tests or records. Neither averages nor deadlines override a release-blocking failure.

Carry the supported scope, stop conditions and observation tasks into Lesson 9, which examines use in real task opportunities, reasons for returning to the original process and the support changes needed.

References

Anthropic / Demystifying evals for AI agents NIST / AI Risk Management Framework NIST / Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile Microsoft Research / Guidelines for Human-AI Interaction