FDE01Forward deployment
Updates 02

Evaluate agents by checking that the work is actually done

Check the business result as well as the “done” response.

On this pageWhat this means for your projectApply it in your projectReferences

What this means for your project

In a multi-turn task, an agent may call tools, change records and continue from intermediate results. A correct final answer can still accompany duplicate requests, skipped steps or unauthorized access.

Record the task, environment, individual trials, execution trace and final business state separately. Ask domain experts to inspect outcomes and failure consequences, preserving disagreements. Model grading can help screen results but does not replace expert judgment.

Repeated attempts may follow different traces. Record the model, instructions, tools, permissions, knowledge versions and test conditions. Recheck relevant tasks after material changes; one success does not establish reliability.

Before release, review different task and user groups, blockers, support arrangements, stop owners and retest conditions. These recommendations synthesize engineering sources. Sample sizes and thresholds still need to reflect the consequences of the task.

Apply it in your project

Choose a task you are working on and check whether the approach applies. Record what to investigate, which step might change, who handles it and when to review. If you have no live project, use a practice case.

References

Anthropic / Demystifying evals for AI agents Google SRE / Canarying releases