Authored demonstration · no customer data or model runs
See what an AI evidence review looks like.
Read a sample decision memo, follow each finding to its illustrative source and see the next test. All figures and findings are authored examples. This is a demonstration of the format, not a completed customer engagement or a full review of underlying cases.
Sample decision memo
Question: Should a support team switch from workflow A to B because the model bill is lower?
Current conclusion: The lower model bill does not establish lower overall cost. Completion quality also remains unresolved. Obtain comparable outcome records before deciding whether to change the workflow.
Scope of this example: two authored aggregate summaries, one hourly-cost assumption and two declared evidence gaps. No underlying case, invoice or customer statement was verified.
Findings you can trace
F1 · Not supported by these inputs
“B is cheaper overall.”
Machine cost falls from 20 USD to 10 USD. Including the stated human time, scoped cost is 80 USD for A and 130 USD for B. This arithmetic is not a forecast of customer savings.
Declared completion rates are 80% and 70%. Both could reach 90% if every unknown outcome were complete. These are arithmetic bounds. Comparable case assignment and independent outcome records are missing.
Reconcile the outcome records, include failed and abandoned work, and test comparable cases under agreed criteria. The owner of the workflow must also consider hard constraints. No model receives a general quality score or deployment approval.
Human time is valued at 30 USD / hour. Costs cover all attempts. The completion labels are inputs to the demonstration, not independently verified outcomes.
Illustrative workflow comparison
Measure
A
B
Attempted cases
100
100
Declared confirmed completions
80
70
Unknown outcomes
10
20
Not completed
10
10
Machine cost
20 USD
10 USD
Human review / rework minutes
120
240
Other direct costs
0 USD
0 USD
Calculated scoped cost
80 USD
130 USD
Cost per declared completion
1 USD
1.86 USD
Calculated with the same arithmetic as AI Outcome Check. Rounded display values. No statistical confidence interval or causal effect is estimated.
S2 · Missing evidence in the example
Independent records supporting completion labels have not been supplied.
The task mix, case assignment and observation windows have not been shown to be comparable.
S3 · Proposed next test in this example
The workflow owner writes one completion rule and the unacceptable failure conditions before the test.
An identified reviewer reconciles the unknown cases against an independent record. Unresolved cases stay unknown.
The team agrees case assignment, observation windows and the required sample for its decision. A paired or randomized comparison may be appropriate; this example supplies no sample-size guarantee.
Record machine costs, human review and retries for all attempted work. Review the results and limitations with the decision owner.