By the end of this lesson
- Identify whether a disagreement concerns facts, resources, goals or authority.
- Write an incident update covering impact, protective action, uncertainty and the next update.
- Turn a repair commitment into an owned action that can be tested.
Before you begin
- Bring lesson 9’s task counts and support plan, plus the stop conditions and authority from lessons 7 and 8.
- Read this lesson’s Beichen version-mismatch scenario. It is distinct from the earlier old-session access incident; a shared communication method does not make them one incident.
Read the Beichen case and introductory data
Start with the decision needed today
Lesson 9 showed that wider use needs support and a workable process. Now the business owner wants to keep an announced date, operations says it cannot cover the work, and sales is watching renewal. Your task is not to establish who understands the technology best. It is to decide whether to keep day-shift scope, add support before expanding, demonstrate separately or pause.
An opening might be: “We need to agree the scope of the demonstration in two weeks and whether live trial work continues meanwhile. Let us check impact and support before comparing options.” That gives the participants one decision to work on and prevents a long account of the project’s history from consuming the meeting.
Confirm authority too. The business owner decides business scope and resources; the data owner controls data use and access; the operations owner confirms coverage, pausing and recovery. Agreement from one does not replace the others. If an authorized decision-maker is absent, record who must decide, by when, and the already-authorized conservative arrangement meanwhile.
Translate positions into specific consequences
“Full rollout next week” and “Absolutely not” are positions, not necessarily the underlying needs. Ask each person what consequence they most fear. Then offer a restatement they can correct: “You need to preserve the demonstration date, but not start unsupported overnight use immediately—is that right?” Do not assume this explanation is true; a correction provides useful new information.
Separate disputed facts, competing explanations, resource or interest conflicts, risk acceptance and authority. Different reports need aligned task populations and time windows. Different explanations need further evidence. Missing staff requires a resource or scope decision. More accuracy data cannot staff the night shift or choose acceptable trade-offs for the business owner.
If an earlier promise overstated the evidence, acknowledge the specific overstatement before discussing alternatives. Engineering detail can follow; it does not answer how people are affected. Use the comparison below to prepare a meeting or to check whether your own draft says anything more useful than “I understand.”
| What you hear | Question to clarify | Useful next step |
|---|---|---|
| “Your completion rate is wrong” | Do the reports cover the same tasks and period? | Compare records and calculations |
| “We cannot support this” | Which shifts, tasks or exceptions lack coverage? | Narrow scope or secure additional support |
| “The board already knows the date” | Is live operation required, or a progress demonstration? | Compare a synthetic demo with a restricted trial |
| “You said it was fixed” | Which equipment models and manual versions did the earlier check cover? | Correct the statement and test the omitted scope |
| “You can accept the risk” | Who may accept consequences for the customer? | Take the decision to the authorized owner |
Protect people while organizing the response
In this scenario, the team said last month that version mismatch was fixed, although only new equipment had been tested. Now 37 tasks are confirmed to have used the wrong version; the count may reach 60, and two involved high-temperature handling. These are fictional event facts, separate from the twenty introductory JSON records. Thirty-seven is confirmed; sixty is a possible scope still being checked.
Use the existing stop authority to limit the affected function, retain necessary records and route ongoing tasks to capable people. Notify the business, data and operations owners. Match containment to the risk: wrong manual references and unauthorized access are different failures. Do not leave people exposed while seeking a cause, or declare all tasks safe without verification.
Google SRE’s incident-response material distinguishes coordination from technical resolution. Adapt that idea by assigning coordination, system and task investigation, and user updates. A small team may combine roles, but none should disappear. Record actions, times and results so that separate responders do not independently restore service or repeat an external action.
Write an update that tells people what to do
A first update should cover confirmed impact, protective action already taken, what remains unknown, what people should do now and the next update time. “We are investigating urgently” does not help an engineer choose a safe next step. You can commit to another update without inventing a recovery deadline. Have the responsible owner confirm formal external messages and channels where required.
For practice only, assume operations has confirmed that affected version suggestions are paused and unfinished tasks are in a manual queue. A draft could read: “Thirty-seven tasks are confirmed to have used the wrong version; two involved high-temperature handling. We are checking other potentially affected tasks. The affected suggestion function is paused and on-duty staff are handling unfinished work. Do not continue using the old suggestions. Recovery conditions remain unverified; the next update is at 16:00 today.” The protective actions and time are exercise assumptions. In a real message, report only confirmed actions and an agreed time.
Add a correction of the earlier promise: “We checked only new equipment but described the whole version problem as resolved. That conclusion was too broad; other equipment is now being checked.” Remove unsupported qualifiers such as “just,” “a few” or “probably fine.” No evidence of an event is different from evidence that it did not happen.
Put choices and consequences on one page
Offer substantively different choices rather than a full version, a smaller version and an overtime version. In this scenario, options include pausing live use while keeping a synthetic demonstration, retaining only newly verified and supported scope, or delaying the demonstration to focus on repair. Until revalidation is complete, the second is conditional—not an approved release.
For each option, show affected people, support needs, sacrifices, decision owners and review triggers. Aim for a specific record such as “Approve the synthetic demo; keep live suggestions paused; review again after operations submits recovery evidence.” Real dates and owners require actual authority. In an exercise, mark them as proposed rather than pretending to commit for a customer.
Agreement is not always possible. You can still record the disagreement, interim scope and escalation owner clearly. Continued suspension is a valid option when the necessary resources do not exist. “Agreement in principle” must not disguise a risk that no authorized person has accepted.
Repair trust by changing the failure mechanism
The trust failure came from narrow testing and an overly broad promise. Repair should extend testing to all relevant equipment models and withdrawn sources, require retesting after manual-version changes, and check that the scope described to users has been approved. More weekly meetings or “lessons learned” cannot demonstrate that these gaps have closed. Give each repair an owner, completion condition and verifier.
Google SRE’s postmortem material emphasizes learning from system conditions and corrective actions. Build a timeline of discovery, information available, decisions and controls that failed. The purpose is neither to find a person to blame nor to promise that failure is impossible. Invite affected people to check whether the change addresses their actual problem.
Executives, frontline staff and operations need different levels of detail, but the confirmed task count, possible scope, unknowns and pause state must stay consistent. Maintain one time-stamped factual record. Correcting an earlier statement helps people assess a new commitment; whether trust has recovered remains theirs to judge through subsequent actions.
Check whether jargon is helping you avoid the issue
Common mistakes include explaining technical complexity before the impact, reporting a possible sixty as confirmed, presenting proposed containment as completed, and promising a fix tomorrow to preserve the relationship. Each makes it harder for people to decide whether work should continue. Check that every number, action and time has supporting records.
For self-study, read your opening aloud and mark defensive explanations, quick concessions and vague assurances. Replace them with facts, choices and conditions still needing confirmation. With a partner, ask for a restatement of approved scope. If it differs from your record, correct the understanding. Concision requires selecting detail, not omitting adverse facts.
Practice the method
Write an incident update and decision note for the version-mismatch scenario. Work alone or ask a partner to challenge the draft from business and operations perspectives. Do not use a production system or contact a real customer.
- 01
List confirmed information—37 tasks, possible scope up to 60, two involving high temperature—and unknown final scope, consequences, cause and recovery conditions. Retain the adverse fact that only new equipment was tested previously.
- 02
Propose containment and label it as requiring authority. If an exercise update describes completed actions, explicitly state the rehearsal assumptions first.
- 03
Write an update you can read within ninety seconds, including what users should do and the next update time. Do not promise an unsupported recovery date.
- 04
Compare suspension with a synthetic demo, a restricted trial after conditions are met, and delayed repair. List authority, support and verification needs.
- 05
Define two repairs aimed at the validation gaps, including who verifies them and how. Check all versions for consistent figures and status. In self-study, leave real approval and execution marked “not yet verified.”
Separate confirmation, investigation and contact status
Review 60 potentially affected tasks, distinguish confirmed impact, unknowns, contact and protection status, then draft an incident update.
The 37 confirmed tasks and two high-temperature cases follow the public summary. Individual assignments and contact progress are independent extensions; contact does not establish resolution.
Read the field guide, work through the calculation or walkthrough, then check the guide's review notes. Open the downloaded files in a spreadsheet or text editor; no code is required.
Check your understanding
Write your answer before opening the explanation, then check what you might have missed.
How should “possibly up to 60” be reported?
Report it as possible scope under investigation alongside the confirmed 37. Do not replace the confirmed count with 60 or omit the known uncertainty.
Can an update say “paused” before pausing has been confirmed?
No. Real messages describe confirmed actions. An exercise may explicitly assume a completed pause before drafting, but must not present that assumption as a supplied case fact.
Does more frequent reporting prove that version mismatch is fixed?
No. Test omitted equipment, version selection and withdrawal propagation. Communication can support repair but does not replace verification.
Your record and next step
A disagreement note, an incident update and a concrete repair plan.
Unknowns are not stated as facts, the right people are informed, and every repair action has an owner and a checkable result.
Lesson 11 brings completed work, review load and incident-handling costs into the value case. Keep this lesson’s adverse facts and repairs: they affect whether the next investment is worthwhile.