Agent Failure Briefings

ONLINE · ENGLISH · FIRST EDITION

When LLM judges and agent tests miss the failure.

Twenty-minute case talks on human feedback, AI judges and execution tests: the evidence that exposed a failure, the repair, and what it still misses.

20-minute prerecorded talks, played in sequence at a scheduled online event. A host introduces each session and manages transitions. The program centers on recorded case presentations.

Inspect a worked case

Ask about this edition

Target date: 14 November 2026 · 12:00–16:00 UTC · 21:00–01:00 KST (ends next day). Contributor confirmations are in progress. Registration is not open.

Proposed talk themes

A judge that agreed for the wrong reason

How a fluent response received a good score while missing the actual task.

A successful outcome reached through forbidden actions

Why final-state checks missed authorization and duplicate side effects.

The test that did not survive deployment

A concrete mismatch between evaluation conditions and the real workflow.

These are program proposals, not confirmed speaker titles. Read the complete four-hour preliminary program.

A preview of the editorial standard

This written case note shows how we intend to structure each talk: the task, the misleading result, the missing evidence and the limits of the repair. It is not a speaker recording.

The task: refund order A exactly $50 once, only after valid authorization. Preserve other orders. A repeated request with the same idempotency key may return the same receipt without a second payment.

  1. Trace A: authorize → issue receipt R1 → retry same key → return R1. One payment.
  2. Trace B: issue R1 → authorize → issue R2 → reverse R2. Final net payment is $50.
  3. Trace C: final net payment is $50; authorization and intermediate events are missing.
What changes when you inspect the evidence?

A passes this stated contract if the trace and initial state are trusted. B fails despite the correct net total: it paid before authorization and issued twice. C cannot establish compliance; marking it a pass or a confirmed violation would both exceed the evidence.

A final-total-only grader passes all three. Requiring exactly one successful effect, authorization before that effect, and trustworthy evidence separates these outcomes while allowing the safe retry. It does not prove the logger is complete or that these checks cover other task contracts.

Read the full case and completed decision sheet · Download as text

Authored examples, not real customer incidents, measured model results or a certification. No API key, account or private data is needed.

The replay pass

Scheduled viewing is planned to be free. The optional paid pass would include the agreed recordings, presenter materials and case notes. The exact contents, price and access period will be listed before purchase.

Format, access and booking details

Who it is for

Engineers and technical owners already testing an agent workflow. Familiarity with evaluations helps; no GPU, paid model account or customer data is needed for the sample.

Questions about attending?

Contact Minji about the program, access or a team booking. If you received an invitation, you can reply directly to that email.

Email the organizer

peonypeony404@gmail.com

Your email is used to answer your enquiry. It does not subscribe you to a mailing list.

Who is organizing this?

Organized by Minji Park / RedSoft. This is our first edition. The worked case is original teaching material; the speaker lineup will be published as contributors accept.

For questions about the program, contact Minji directly. Contributor proposals are on a separate page.