Jacob Andreas
MIT
Workshop on
The rapid transition from large language models (LLMs) as single-turn assistants to interactive agents has created an urgent need for new evaluation methodologies. LLMs are increasingly deployed in high-impact settings such as education, counseling, negotiation, research assistance, and software development, where success depends not only on generating a correct response, but on sustaining effective interactions over extended trajectories.
Evaluating interactive agents directly with real users can be slow, expensive, difficult to reproduce, and hard to scale in expert domains. This has led to growing use of automatic evaluation methods, ranging from rubric-based grading to user simulators, where an LLM simulates user behavior to support evaluation, training, and stress testing. However, many issues with these approaches remain: for example, user simulators may fail to preserve latent user states, reflect diverse human attributes, represent realistic goals, or match the interaction style of real users.
In light of these challenges, this workshop will focus on methods for developing more rigorous, scalable, and scientific evaluation methods for interactive agents.
We invite contributions on topics including, but not limited to:
We solicit non-archival submissions using the official NeurIPS 2026 LaTeX style: up to 9 pages for full papers and 4 for short papers, excluding references and appendices. We welcome early-stage work, as well as conference papers that have already been accepted at the NeurIPS main conference or in other venues.
Accepted work will primarily be presented as posters, with a select number of papers receiving spotlight talks, as well as a best paper award. The official OpenReview submission site will be available soon.
MIT
Microsoft
USC
Cohere
Stanford
Princeton
Apple
UC San Diego
Georgia Tech
Columbia University
Princeton University
TTIC
Microsoft Research
Google DeepMind
Georgia Tech
Georgia Tech
Questions? Contact the workshop chairs:
douy@gatech.edu · siyan.li@columbia.edu · marwa_abdulhai@berkeley.edu