Agent Evaluation Studio

Preliminary program

Four hours online in English, including two breaks. Times are elapsed time from the eventual start; the target date is 21 November 2026,12:00–16:00 UTC (21:00–01:00 KST, ending the next day). Contributors are not yet confirmed.

Short talks, practical exercises and live moderated discussion. The facilitator and producer are being recruited. The outline uses authored cases; it does not promise a review of every participant’s private project.

00:00–00:10 · 10 min · Live facilitator

Orientation and initial release decision

00:10–00:35 · 25 min · Short talk

Human feedback, AI judgment and verifiable outcomes

00:35–01:05 · 30 min · Guided exercise

Compare judgments against the task contract

01:05–01:25 · 20 min · Moderated discussion

Discuss false passes, disagreement and abstention

01:25–01:35 · 10 min · Break

Break

01:35–02:00 · 25 min · Short talk

Execution traces and limits of sandbox evidence

02:00–02:35 · 35 min · Guided exercise

Repair a verifier without rejecting safe retries

02:35–02:55 · 20 min · Moderated discussion

Compare repairs and unresolved evidence

02:55–03:05 · 10 min · Break

Break

03:05–03:30 · 25 min · Applied exercise

Write a release or hold decision

03:30–03:50 · 20 min · Moderated discussion

Review selected decisions and questions

03:50–04:00 · 10 min · Live facilitator

Exit task and next evidence to collect

What contributors must demonstrate

A specific task and claim, the evidence behind the result, a failure or limitation, a proposed repair, and what remains untested. A VM execution result is not automatically a substitute for human preference or quality judgment.

Contribution requirements