Workshop on

Evaluation of
Interactive Agents

Under construction: finalized date, call for papers, and OpenReview submission site coming soon!

Why this workshop?

The rapid transition from large language models (LLMs) as single-turn assistants to interactive agents has created an urgent need for new evaluation methodologies. LLMs are increasingly deployed in high-impact settings such as education, counseling, negotiation, research assistance, and software development, where success depends not only on generating a correct response, but on sustaining effective interactions over extended trajectories.

Evaluating interactive agents directly with real users can be slow, expensive, difficult to reproduce, and hard to scale in expert domains. This has led to growing use of automatic evaluation methods, ranging from rubric-based grading to user simulators, where an LLM simulates user behavior to support evaluation, training, and stress testing. However, many issues with these approaches remain: for example, user simulators may fail to preserve latent user states, reflect diverse human attributes, represent realistic goals, or match the interaction style of real users.

In light of these challenges, this workshop will focus on methods for developing more rigorous, scalable, and scientific evaluation methods for interactive agents.

Topics

We invite contributions on topics including, but not limited to:

Call for papers

We solicit non-archival submissions using the official NeurIPS 2026 LaTeX style: up to 9 pages for full papers and 4 for short papers, excluding references and appendices. We welcome early-stage work, as well as conference papers that have already been accepted at the NeurIPS main conference or in other venues.

Accepted work will primarily be presented as posters, with a select number of papers receiving spotlight talks, as well as a best paper award. The official OpenReview submission site will be available soon.

Invited speakers

Jacob Andreas

MIT

Asli Celikyilmaz

Microsoft

Morteza Dehghani

USC

Seraphina Goldfarb-Tarrant

Cohere

Sanmi Koyejo

Stanford

Karthik Narasimhan

Princeton

Heinrich Peters

Apple

Lianhui Qin

UC San Diego

Schedule

Time
Session
8:00–8:10
Opening remarks
8:10–9:10
Invited talks
9:10–10:00
Oral session
10:00–11:00
Poster and demo session
11:00–12:00
Invited talks
12:00–1:00
Lunch and mentoring tables
1:00–2:00
Invited talks
2:00–3:00
Poster and demo session
3:00–3:15
Break
3:15–4:15
Invited talks
4:15–4:45
Panel
4:45–5:00
Awards and closing

Organizers

Questions? Contact the workshop chairs:

douy@gatech.edu · siyan.li@columbia.edu · marwa_abdulhai@berkeley.edu