# Lesson 8 practice: What does the overall pass rate hide?

[中文](README.md)

Open [evaluations.csv](evaluations.csv). Check aggregate results, consequential slices, and the coverage needed for another test. These 52 rows form an **independent synthetic evaluation set**. They are not the row-level source of Beichen’s original “94%,” the lab’s 28 regression tasks, or the 58 implementation tests. Every sample is public so you can check the reasoning and calculations.

## Fields

| Field | Meaning |
| --- | --- |
| `source_type` / `sample_id` | Independent synthetic marker / sample ID |
| `sampling_group` | `routine`, `risk_targeted`, or `additional`; not a random population sample |
| `shift` / `equipment` | `day` / `night`; `current_model` / `old_model` |
| `task_type` | Inspection, parts draft, high temperature, regional permission, sensor data, unknown tool state, missing information, or missing logs |
| `expected_behavior` / `observed_behavior` | Expected / simulated observed behavior; compare each row |
| `task_result` / `system_result` / `user_result` | `pass` or `fail` for task outcomes, system behavior, and user experience |
| `risk_result` / `support_result` | `pass` or `fail` for risk and operational support |
| `overall_pass` | 1 only when all five categories pass; otherwise 0, under this exercise’s rule |

Behavior codes describe showing an applicable version and source (`show_applicable_version_and_source`), preparing a draft without submission (`prepare_draft_without_submitting`), withholding steps and escalating to qualified people (`withhold_steps_and_escalate`), denying cross-region access (`deny_cross_region_access`), showing freshness and requesting review (`show_freshness_and_request_review`), reconciling without retry (`reconcile_without_retry`), or requesting information or logs and escalating (`request_information_and_escalate`, `request_logs_and_escalate`). The failure codes are using new-model steps for an old model (`show_new_model_steps_for_old_model`) and giving steps despite missing information (`show_steps_without_information`).

## Calculate and review

1. Count rows and passes: **49 / 52 ≈ 94.2%**. This describes this table only. A similar percentage does not explain the original event’s 94% or its unsupported 1.5-percentage-point statement.
2. Filter old-model tasks: **four rows, two passes, two failures**, a **50%** failure rate. Read S-17 and S-31 for the specific consequences rather than allowing an aggregate rate to hide incorrect operating steps.
3. Read S-22. Its user-experience field passes, but steps are given despite missing information. Risk and other categories fail, and the overall result is 0. This illustrates why a user-experience pass cannot replace task and risk checks; it is not a real user study.
4. Count slices by shift, model, and task type, retaining denominators. Both old-model failures that give incorrect high-temperature steps occur at night; S-22 also has a daytime risk failure. That does not establish readiness across all daytime paths. Identify missing coverage for normal tasks, unknown states, refusal, and handoff.
5. Propose a restriction or pause, blocking errors, retest scenarios after a fix, and responsible owners. Actual release decisions require authorized people and system-specific evidence; the exercise does not approve equipment procedures.

## Check your reasoning

Numeric booleans use 1 for true and 0 for false. Convert values before calculations: the string `'0'` can itself be truthy. A `pass` is an assigned simulation result, not evidence of real user understanding or system safety.

Submit aggregate and slice tables, explanations for the three failures, and missing task coverage. A peer can independently apply the rule to a row. When studying alone, state that your conclusions use synthetic evidence; no real equipment or customer approval is required. Public samples support learning and regression checks, not claims of unseen independent acceptance testing.

The next lesson asks which eligible tasks actually enter and complete the new workflow.
