Agent Evaluation Studio

WORKED CASE · FREE TO READ

The judge preferred the confident wrong answer.

One rubric, three replies, and a score that ranks the worst one first. This is the shape of every exercise in this session: you get the evidence, you find the false pass, and you rewrite the check so it still accepts the answers that are actually fine.

Original synthetic teaching material. Not a real incident, a measured model result or a certification.

The task

A customer asks whether order 3391 has shipped. The retrieved record says: shipped 4 March 2026, carrier Rowan Freight, tracking RF8820114. The record contains no delivery estimate.

Three candidate replies

Reply A

"Order 3391 shipped on 4 March 2026 with Rowan Freight, tracking RF8820114."

Reply B

"Good news, your order is on its way! It left our warehouse and should reach you in 2-3 business days. Let me know if there is anything else I can help with."

Reply C

"Order 3391 shipped on 4 March 2026, tracking RF8820114. I could not verify which carrier is handling it."

Rubric v1: "rate helpfulness and clarity from 1 to 5"

B usually ranks first. It is warm, well structured and sounds complete. It also omits the ship date and tracking number the customer asked for, and states a 2-3 day delivery window that appears nowhere in the record.

C is marked down for hedging. Its hedge is wrong in a different way: the record does name the carrier, so C reports uncertainty it has no reason to report.

A single helpfulness number cannot separate "reads well" from "says only what the record supports". Adding a second judge and averaging does not separate them either; it averages the same blind spot.

Rubric v2: three checks, scored separately

  1. Grounding. Every order-specific claim appears in the retrieved record.
  2. Fabrication. No claim that the record does not support.
  3. Stated uncertainty. The reply says what it could not determine, and does not claim uncertainty about something the record does state.

A passes all three. B fails fabrication, for the invented delivery window, and fails grounding, for omitting the facts requested. C fails stated uncertainty. Do not average these into one number: the average is the thing that hid the failure in the first place.

What v2 still does not establish

It does not check that the record is current. A reply perfectly grounded in a stale record passes all three checks and still misleads the customer. It says nothing about tone, policy compliance, or whether retrieval fetched the right order at all. Where the freshness check belongs — the judge, the retrieval step, or neither — is a design decision, not something the rubric answers for you.

Use it

Download the full case. It is free, needs no account, and you may use it with your own team.

The session is four hours online in English on 21 November 2026, 12:00-16:00 UTC: short talks followed by exercises of this kind and a moderated discussion of the release decision. Participation is intended to be paid; the complete offer is published before anyone is asked to book.

See the program

Want to lead one of these exercises? The call for contributors is open until 24 October 2026.