Skip to main content
← AI Evidence Review

Authored demonstration · no customer data or model runs

See what an AI evidence review looks like.

Read a sample decision memo, follow each finding to its illustrative source and see the next test. All figures and findings are authored examples. This is a demonstration of the format, not a completed customer engagement or a full review of underlying cases.

Sample decision memo

Question: Should a support team switch from workflow A to B because the model bill is lower?

Current conclusion: The lower model bill does not establish lower overall cost. Completion quality also remains unresolved. Obtain comparable outcome records before deciding whether to change the workflow.

Scope of this example: two authored aggregate summaries, one hourly-cost assumption and two declared evidence gaps. No underlying case, invoice or customer statement was verified.

Findings you can trace

F1 · Not supported by these inputs

“B is cheaper overall.”

Machine cost falls from 20 USD to 10 USD. Including the stated human time, scoped cost is 80 USD for A and 130 USD for B. This arithmetic is not a forecast of customer savings.

Inspect example source S1

F2 · Not established

“B completes the work just as reliably.”

Declared completion rates are 80% and 70%. Both could reach 90% if every unknown outcome were complete. These are arithmetic bounds. Comparable case assignment and independent outcome records are missing.

Inspect example source S2

F3 · A decision requires more evidence

“This is enough to approve the change.”

Reconcile the outcome records, include failed and abandoned work, and test comparable cases under agreed criteria. The owner of the workflow must also consider hard constraints. No model receives a general quality score or deployment approval.

Inspect example source S3

S1 · Authored aggregate inputs

Human time is valued at 30 USD / hour. Costs cover all attempts. The completion labels are inputs to the demonstration, not independently verified outcomes.

Illustrative workflow comparison
MeasureAB
Attempted cases100100
Declared confirmed completions8070
Unknown outcomes1020
Not completed1010
Machine cost20 USD10 USD
Human review / rework minutes120240
Other direct costs0 USD0 USD
Calculated scoped cost80 USD130 USD
Cost per declared completion1 USD1.86 USD

Calculated with the same arithmetic as AI Outcome Check. Rounded display values. No statistical confidence interval or causal effect is estimated.

S2 · Missing evidence in the example

  • Independent records supporting completion labels have not been supplied.
  • The task mix, case assignment and observation windows have not been shown to be comparable.

S3 · Proposed next test in this example

  1. The workflow owner writes one completion rule and the unacceptable failure conditions before the test.
  2. An identified reviewer reconciles the unknown cases against an independent record. Unresolved cases stay unknown.
  3. The team agrees case assignment, observation windows and the required sample for its decision. A paired or randomized comparison may be appropriate; this example supplies no sample-size guarantee.
  4. Record machine costs, human review and retries for all attempted work. Review the results and limitations with the decision owner.