Diagnose
Identify what current evaluations measure, where they fail, and which gaps limit reliable conclusions.

Fall 2026 · New York City
A one-day symposium advancing the methods, measures, systems, and evidence used to evaluate AI agents.
Why this symposium
AI agents act across tools, environments, and extended workflows. Evaluating them requires evidence beyond a single score or benchmark.
The symposium brings together researchers and practitioners working on measurement, benchmark design, evaluation infrastructure, reliability, safety, and real-world performance.
Evaluation science agenda
Identify what current evaluations measure, where they fail, and which gaps limit reliable conclusions.
Develop frameworks, constructs, metrics, graders, and evidence for reliability and validity.
Build runnable benchmarks, environments, harnesses, and reproducible evaluation infrastructure.
Evaluate agents in realistic settings using failure evidence and real-world performance.
Symposium experience
Each part of the day connects methodological questions with practical evaluation experience.
Academic and industry speakers connecting evaluation methodology with deployed agent systems.
View speakers →02Selected short papers and abstracts presented through focused oral and poster sessions.
View the call →03Case studies, benchmarks, tools, failures, and evidence from real-world agent evaluation.
View program →04Structured discussion spanning diagnosis, measurement, infrastructure, and real-world application.
Explore the agenda →Important dates
Submissions open
August 252026Abstract registration deadline
October 202026Submission deadline
October 252026Symposium
November 202026Affiliations represented
Affiliations describe committee members’ institutional or organizational connections. They do not indicate sponsorship.