00:00–00:10 · 10 min · Host introduction
Preliminary program
Four hours online in English, including two breaks. Times are elapsed time from the eventual start; the target date is 14 November 2026,12:00–16:00 UTC (21:00–01:00 KST, ending the next day). Contributors are not yet confirmed.
Ten 20-minute recorded case talks, introduced and played in sequence. One contributor may supply more than one accepted talk. Titles below are editorial briefs, not confirmed presentations. Live speaker Q&A is not included.
00:10–00:30 · 20 min · Recorded case talk
What human preference labels actually reward
00:30–00:50 · 20 min · Recorded case talk
When AI feedback repeats a labeling mistake
00:50–01:10 · 20 min · Recorded case talk
Judge calibration beyond agreement
01:10–01:30 · 20 min · Recorded case talk
When the evaluator should abstain
01:30–01:40 · 10 min · Break
Break
01:40–02:00 · 20 min · Recorded case talk
What a VM or sandbox does not validate
02:00–02:20 · 20 min · Recorded case talk
Verifiable rewards and exploitable checks
02:20–02:40 · 20 min · Recorded case talk
Reset contamination and missing evidence
02:40–03:00 · 20 min · Recorded case talk
Correct outcomes, forbidden side effects
03:00–03:10 · 10 min · Break
Break
03:10–03:30 · 20 min · Recorded case talk
An evaluation that failed after deployment
03:30–03:50 · 20 min · Recorded case talk
A release decision and the evidence that changed it
03:50–04:00 · 10 min · Host close
Recap and next checks
What contributors must demonstrate
A specific task and claim, the evidence behind the result, a failure or limitation, a proposed repair, and what remains untested. A VM execution result is not automatically a substitute for human preference or quality judgment.
Contribution requirements