Fall 2026 · New York City

Agent Evaluation Science

A one-day symposium advancing the methods, measures, systems, and evidence used to evaluate AI agents.

Date
Friday, November 20, 2026
Location
New York City · Primarily in person
Submissions
Short papers and 2-page extended abstracts via OpenReview
Explore the symposium
Scientific evaluationConstruct validityMeasurement reliabilityBenchmark integrityAgent safetyGrader designFailure analysisEvaluation infrastructureReal-world evidenceReproducible systems

Why this symposium

Evaluation is becoming a scientific discipline

AI agents act across tools, environments, and extended workflows. Evaluating them requires evidence beyond a single score or benchmark.

The symposium brings together researchers and practitioners working on measurement, benchmark design, evaluation infrastructure, reliability, safety, and real-world performance.

1focused day
2submission formats
4evaluation stages

Evaluation science agenda

Four stages organize the symposium

01

Diagnose

Identify what current evaluations measure, where they fail, and which gaps limit reliable conclusions.

02

Measure

Develop frameworks, constructs, metrics, graders, and evidence for reliability and validity.

03

Operationalize

Build runnable benchmarks, environments, harnesses, and reproducible evaluation infrastructure.

04

Apply

Evaluate agents in realistic settings using failure evidence and real-world performance.

Important dates

Fall 2026 timeline

Submission details →
01

Submissions open

August 252026
02

Abstract registration deadline

October 202026
03

Submission deadline

October 252026
04

Symposium

November 202026

Affiliations represented

The committee brings together researchers and practitioners across universities, research institutes, and industry.

  • Harvard University
  • Massachusetts General Hospital
  • Princeton University
  • Cornell University · Cornell Tech
  • The University of Texas at Austin
  • University of Technology Sydney
  • Australian Artificial Intelligence Institute
  • University of Alberta
  • Amii
  • Meta Superintelligence Labs
  • Sony AI
  • Raycaster
  • Google DeepMind

Affiliations describe committee members’ institutional or organizational connections. They do not indicate sponsorship.

Agent Evaluation Science Fall 2026Friday, November 20 · New York City
Meet the committee