Agent Evaluation Studio

ONLINE · ENGLISH · FIRST EDITION

Find the failure your agent tests missed.

Inspect an agent trace, challenge the judge and rewrite the check. Short talks, exercises and discussion for engineers deciding whether a workflow is ready to ship.

Short expert talks followed by guided exercises and discussion: inspect the trace, improve the check and defend your decision.

Inspect a worked case

Ask about this edition

Target date: 21 November 2026 · 12:00–16:00 UTC · 21:00–01:00 KST (ends next day). Contributor confirmations are in progress. Registration is not open.

The proposed experience

Inspect the evidence

Separate a persuasive answer, a judge score and evidence that the intended task actually happened.

Challenge the check

Find a false pass, then propose a repair that still accepts valid alternatives.

Defend the release decision

Discuss conflicting judgments and write down what evidence would change your decision.

These are program proposals, not confirmed speaker titles. Read the complete four-hour preliminary program.

Inspect the work you would do

A worked synthetic case is available before you register. You can inspect the rubric, three candidate replies and the revised checks, rather than take our word for the outcome.

The task: a customer asks whether order 3391 has shipped. The retrieved record says it shipped on 4 March 2026 with Rowan Freight, tracking RF8820114. The record contains no delivery estimate.

  1. Reply A: states the ship date, carrier and tracking number, and nothing else.
  2. Reply B: warm and well structured. Omits the date and tracking number, and promises delivery in 2-3 business days.
  3. Reply C: states the date and tracking number, then says it could not verify the carrier.
Which one does a helpfulness score rank first?

Usually B. It reads as complete and confident. It also leaves out the facts that were asked for and invents a delivery window that appears nowhere in the record. C is marked down for hedging, although its hedge is itself wrong: the record does name the carrier.

One helpfulness number cannot separate "reads well" from "says only what the record supports". Scoring grounding, fabrication and stated uncertainty as three independent checks separates all three replies. It still does not establish that the record itself is current.

Read the full case and the revised rubric · Download as text

Authored examples, not real customer incidents, measured model results or a certification. No API key, account or private data is needed.

What you take away

A task contract, a revised evaluator check and a release decision stating what evidence is still missing. Live discussion focuses on the shared exercises; private production audits are not included.

Format, access and booking details

Who it is for

Engineers and technical owners already testing an agent workflow. Familiarity with evaluations helps; no GPU, paid model account or customer data is needed for the sample.

Questions about attending?

Contact Minji about the program, access or a team booking. If you received an invitation, you can reply directly to that email.

Email the organizer

peonypeony404@gmail.com

Your email is used to answer your enquiry. It does not subscribe you to a mailing list.

Who is organizing this?

Organized by Minji Park / RedSoft. This is our first edition. The worked case is original teaching material; the speaker lineup will be published as contributors accept.

For questions about the program, contact Minji directly. Contributor proposals are on a separate page.