FDE01Forward deployment
Course 06

Compare options and test key assumptions

What should you test first when time is limited?

On this pageBy the end of this lessonBefore you beginTurn a belief into a question you can answerPrioritize assumptions that could change directionChoose a small test that answers the questionWorked example: investigate waiting before polishing a demoWrite the three result branches before startingCompare genuinely different next stepsKeep a small test’s claims within its evidencePractice the methodCheck your understandingYour record and next stepReferences

By the end of this lesson

  • Separate facts, assumptions and decisions, and identify an observation that could disprove a key assumption.
  • Design a small test with branches for supporting evidence, contrary evidence and insufficient information.
  • Compare process changes, paper testing and a restricted prototype to recommend an actionable next step.

Before you begin

  • Bring the workflow and unresolved questions from Lesson 5. Start with an uncertainty that could change scope or a stop condition.
  • Use Beichen’s published synthetic events and your paper draft. You do not need to test an unapproved system on real service records.

Read the Beichen case and introductory data

Connect observations to decisions
Connect observations to decisionsRecord sources, times and owners. Results may change your assumptions; mark what still needs checking.FactWhat happened?AssumptionWhat needs testing?DecisionWho chooses what?OutcomeWhat changed?Review or revise the assumption
Record sources, times and owners. Results may change your assumptions; mark what still needs checking.Swipe to view the full diagram.

Turn a belief into a question you can answer

You have a workflow, but it may still assume complete information, careful review and overnight support. Writing these into a proposal does not make them facts. Turn each into an observable question—for example, “Can the overnight engineer open the manual version needed for this task?”

Keep four records. Facts have sources and dates. Assumptions have supporting evidence and a description of what would show them to be wrong. Decisions have authorized owners, scope and conditions. Outcomes record observations and their limits. A correct demo answer is an observation about one input, version and environment, not proof that the whole proposal is ready.

This helps you avoid spending on the wrong uncertainty. Better retrieval scores do not help people who cannot open the material. A knowledge assistant may also miss the main bottleneck if waiting occurs in approval. A test helps you choose what to do next: the findings may support the original approach or call for a change.

Prioritize assumptions that could change direction

Group assumptions into the problem, users, value, data, capability and operations. Here, check for obvious gaps in the value logic; the full cost and benefit analysis comes in Lesson 11. “Faster lookup reduces downtime” needs investigation because repair work, approval and parts delivery can also affect the outcome.

For each assumption, ask what happens if it is wrong, how strong the evidence is, which observation would reduce uncertainty, and what that observation costs in time and exposure. Prioritize consequential, weakly supported questions you can investigate now. You do not need a numerical product that suggests more precision than the evidence allows.

Some long-term questions cannot be settled in a two-day test, such as whether knowledge maintenance will continue indefinitely. Confirm current ownership and handover arrangements, record continuing uncertainty and set a review trigger. What cannot be established should remain visible, not be marked as passed.

AssumptionFirst evidence to inspectWhat changes if it is wrong
Waiting mainly comes from lookupReconstruct events; distinguish lookup, authorization and approvalConsider process or staffing changes
Overnight staff can open the applicable manualCheck permissions and the access path with a simulated identityChange scope or resolve an approved access need
Engineers notice version mismatchesInclude a mismatched source in a paper taskChange the interface or exclude the task
Operations can support added tasksDuty rosters, review time and available capacityReduce scope or arrange adequate support

Choose a small test that answers the question

A paper interface can test whether someone understands a state and the next action. A person simulating the backend can help you observe how advice affects judgment. Reconstructing past events can investigate waiting. A local technical test can check withdrawal, access and tool states. They answer different questions and are not interchangeable.

Specify the behavior to observe before choosing the format. To ask whether source details help someone notice an equipment-model mismatch, you do not need a deployed assistant. Give the person a task, material and draft; observe whether they notice the mismatch, how long it takes and what information is missing. “Would this feature be useful?” measures an opinion, not task completion.

Official design guidance and risk frameworks can broaden the questions, but do not select your sample or threshold for you. For self-study, compare two paper interfaces and record your own misunderstandings. Label this as an individual walkthrough; it does not establish that all engineers will use the system correctly.

Worked example: investigate waiting before polishing a demo

Beichen’s Lesson 6 event provides a counterexample: retrieval accuracy rises from 71% to 84% over three weeks, then an engineer reports finding a document they cannot open. These are synthetic event figures. The material does not supply the full evaluation sample or calculation method, so use them to illustrate an access barrier, not as verified impact data.

The public event proposes reconstructing 12 recent waiting episodes. For a learning exercise, choose 12 of the 20 synthetic records to cover expert waiting, outdated material, missing information and tool timeouts, explaining each choice. This is a selected sample, not a random estimate of all work. Where fields cannot identify a cause, retain “cause unknown.”

Connect results to action before starting. If access to usable information is the main changeable factor and checked material changes task decisions, consider a restricted prototype. If approval or staffing dominates, compare process changes first. If records cannot distinguish causes, improve the evidence. Each branch must allow the proposal to change.

  • State the decision: whether to invest in a restricted prototype. This test does not authorize real submission.
  • List competing explanations: missing applicable material, denied access, expert confirmation and parts approval. More than one may contribute.
  • Reconstruct each event with sources, waiting boundaries and available actions. Total resolution time does not reveal approval time.
  • Provide checked, authorized information on paper and observe changes in the specific decision. Record any added review burden.
  • Summarize support, contrary evidence and unresolved cases, with proposed scope and further investigation rather than “positive user feedback.”

Write the three result branches before starting

An executable test card states the decision, assumption, alternatives, sample selection, observations, stop conditions and authorized decision owner. Write what would support the assumption, weaken it or leave it unresolved. This prevents changing the standard after seeing an inconvenient result.

A numerical threshold is a reasoned choice, not a universal standard. Beichen’s event uses “parts-approval waiting exceeds 40%” as a reversal condition for one shadow-advice proposal. That is a scenario condition. The 20 records have no parts-approval duration field, so they cannot establish the percentage. Define the denominator and collect the relevant observations.

The denominator is the total quantity used to calculate a rate. “40% of waiting” could refer to all waiting minutes, long-wait events or eligible tasks. Those are different questions. Incomplete records, absent participants or changed comparison conditions can leave a test inconclusive. Collect more evidence or keep a restricted approach rather than treating unknown as passed or failed.

Compare genuinely different next steps

A full version, smaller version and delayed version may be three budgets for the same direction. For Beichen, compare A: process and access changes without AI; B: paper or human-assisted investigation; and C: shadow advice within an authorized scope, with no external actions. Shadow output is observed without affecting live decisions. Even shadow use of real data needs purpose authorization.

Explain what each option addresses, its main cost, dependencies, uncertainties and stop conditions. If evidence cannot support C, B may be the right next step. If approval is the bottleneck, A may be more useful. Keeping a non-AI choice makes it possible to assess whether AI adds anything.

Record the approved scope, excluded actions, resource owners, review date and triggers for reopening the decision. When asked to “build it and see,” answer with a specific observation: “We can make a paper draft to see whether engineers notice an equipment-model mismatch. If they do not, we change presentation and scope first.” A prototype with a defined question is a test; a demo without one is a display.

Keep a small test’s claims within its evidence

Three common errors are choosing only easy successes, omitting a contrary-result branch and generalizing a few participants to the whole workforce. Retain unfavorable examples, explain sampling and state the roles, versions, shifts and tasks your conclusion does not cover. A vague “needs further improvement” does not define those limits.

Do not turn “we did not observe a problem” into “the problem cannot occur.” One synthetic withdrawal test checks one known path, not every cache, copy or long-lived session in a live system. NIST’s framework can prompt further questions; controls still need testing, and specialist judgments need qualified, authorized owners.

Close with three statements: what you observed, what the observation does and does not support, and how the next action changes. A decision not to build can still save substantial work. It also gives Lesson 7 a clear starting list of unresolved risks and the people needed to address them.

Practice the method

Exercise 06

You have two days to prepare Beichen’s next-step recommendation. Choose one assumption that could most affect the decision, then write a test card and compare three options. Label any added numbers as illustrative assumptions.

  1. 01

    List assumptions across the six categories. Mark two with weak evidence and serious consequences if wrong.

  2. 02

    Choose one and identify an observation that would contradict it, along with a competing explanation. Explain why it should be tested now.

  3. 03

    Choose paper, historical records or a local test. Define sample selection, observations, ownership, permitted data use and stop conditions.

  4. 04

    Before execution, write supporting, contrary and inconclusive branches, each with a next action.

  5. 05

    Walk through the synthetic material and record findings and missing fields. Compare process changes, paper testing and restricted shadow advice without inventing real user impact.

  6. 06

    Record scope, exclusions, resource owners, a review date and reopening triggers. Without a real authorized decision maker, label this as a recommendation awaiting confirmation.

Check the denominator behind a decision threshold

Calculate the approval-wait share from 12 independent records, then compare your current judgment, competing explanations and next test.

This exercise defines a denominator in wait-minutes. These newly constructed records do not fill the gaps in the original event 06.

Read the field guide, work through the calculation or walkthrough, then check the guide's review notes. Open the downloaded files in a spreadsheet or text editor; no code is required.

Check your understanding

Write your answer before opening the explanation, then check what you might have missed.

If all 12 examples pass, is the real-world pass rate 100%?

No. Explain how the 12 were selected and what was checked. Synthetic, targeted or repeatedly used development examples support observations within that scope, not an estimate of the live task distribution.

Can total resolution time establish the approval-wait share when approval duration is missing?

Not directly. Total duration combines waiting, execution and rework. Define the denominator and timing boundaries, gather the missing data and retain an unknown cause when it cannot be distinguished.

Must you refuse a request to build a prototype first?

No. Define its question, observations and decision branches, then compare cheaper ways to answer the same question. A prototype that can reveal contrary evidence and has stop conditions can be a useful test.

Your record and next step

An assumption list, a test plan and a note explaining how results determine the next step.

A failed test changes your action. Conclusions stay within the tested sample, conditions and data permissions.

Carry the findings, remaining uncertainties and pending decisions into Lesson 7, where you identify permitted data uses, affected people and authority to pause or restore the system.

References

Anthropic / Building effective agents Microsoft Research / Guidelines for Human-AI Interaction NIST / AI Risk Management Framework