Fall 2026 · New York City

Agent Evaluation Science

A one-day symposium advancing the methods, measures, systems, and evidence used to evaluate AI agents.

Date
Friday, November 20, 2026
Location
New York City · Primarily in person
Submissions
Short papers and 2-page extended abstracts via OpenReview
Explore the symposium
Scientific evaluationConstruct validityMeasurement reliabilityBenchmark integrityAgent safetyGrader designFailure analysisEvaluation infrastructureReal-world evidenceReproducible systems

Why this symposium

Evaluation is becoming a scientific discipline

AI agents act across tools, environments, and extended workflows. Evaluating them requires evidence beyond a single score or benchmark.

The symposium brings together researchers and practitioners working on measurement, benchmark design, evaluation infrastructure, reliability, safety, and real-world performance.

1focused day
2submission formats
4evaluation stages

Evaluation science agenda

Four stages organize the symposium

01

Diagnose

Identify what current evaluations measure, where they fail, and which gaps limit reliable conclusions.

02

Measure

Develop frameworks, constructs, metrics, graders, and evidence for reliability and validity.

03

Operationalize

Build runnable benchmarks, environments, harnesses, and reproducible evaluation infrastructure.

04

Apply

Evaluate agents in realistic settings using failure evidence and real-world performance.

Call for presentations

Share work that strengthens agent evaluation

Submission details →

Submit evaluation studies, audits, negative results, failure analyses, tools, datasets, production lessons, and real-world evaluation cases. Ongoing and previously published work can also be presented through the extended abstract track.

Industry and real-world submissions are assessed on concrete evidence, production lessons, and generalizable insight. A new benchmark or algorithm is not required.

01

Submissions open

August 252026
02

Abstract registration deadline

October 202026
03

Submission deadline

October 252026
04

Symposium

November 202026

Affiliations represented

The committee brings together researchers and practitioners across universities, research institutes, and industry.

  • Harvard University
  • Massachusetts General Hospital
  • Princeton University
  • Cornell University · Cornell Tech
  • The University of Texas at Austin
  • University of Technology Sydney
  • Australian Artificial Intelligence Institute
  • University of Alberta
  • Amii
  • Meta Superintelligence Labs
  • Sony AI
  • Raycaster
  • Google DeepMind
  • Massachusetts Institute of Technology
  • MIT Media Lab

Affiliations describe committee members’ institutional or organizational connections. They do not indicate sponsorship.

Agent Evaluation Science Fall 2026Friday, November 20 · New York City
Meet the committee