Short paper
Archival or presentation-only- Length
- Up to 6 pages, excluding references
- Best for
- Completed research and evaluation studies

Fall 2026 · New York City
Short papers and extended abstracts on the science and practice of AI agent evaluation.
Submission options
Final format is determined by the program committee.
Scope
Audits, replications, negative results, contamination, grader instability, and evaluation gaps.
Constructs, metrics, rubrics, human evaluation, automated graders, reliability, validity, cost, and safety.
Benchmarks, environments, harnesses, traces, repeated runs, versioning, observability, and scalable grading.
Coding, research, science, healthcare, enterprise, web, computer-use, multimodal, and other tool-using agents.
Selection
Industry and real-world submissions are assessed on concrete evidence, production lessons, and generalizable insight. A new benchmark or algorithm is not required.
Author information
Final publication and licensing terms will be confirmed before camera-ready submission. There is no long-paper track.
FAQ
Yes, as a presentation-only short paper or extended abstract. Identify and cite the prior publication. Published work is not eligible for archival proceedings.
The meeting is primarily in person in New York City. Limited remote presentation may be approved when necessary.
Only accepted archival short papers. Presentation-only short papers and all extended abstracts are non-archival.
Authors indicate Oral preferred, Poster preferred, or Either. The program committee makes the final assignment.