AIJul 20

ProEvent: An Event-centric Benchmark for Proactive Agents

arXiv:2607.177019.12 citations
Predicted impact top 65% in AI · last 90 daysOriginality Incremental advance
AI Analysis

Provides a standardized evaluation for proactive agents in event-centric assistance, highlighting fundamental limitations of current LLMs for this task.

ProEvent introduces the first event-centric benchmark for evaluating proactive agents, revealing that current LLMs (e.g., GPT-5.1) only react correctly in 26.7% of scenarios, struggling with event cancellation and implicit event detection.

Proactive agents are expected to anticipate user needs and provide autonomous assistance by perceiving environmental context without explicit instructions. A fundamental capability of such agents is to identify and track users' upcoming events, enabling continuous and event-specific assistance. For example, by recording the time and location of a planned hike, an agent can deliver weather reminders in advance or provide navigation support before departure. However, existing works on proactive agents largely overlook event-centric assistance, and the open-ended nature of proactive assistance poses challenges for reliable evaluation. To bridge these gaps, we introduce ProEvent, the first event-centric benchmark designed to assess an agent's ability to proactively maintain a user's timetable based on ongoing instant messaging chats. ProEvent provides synthesized yet realistic chats that consider the dynamic interaction among users, concurrent chat threads, and noise in the real world, and evaluates proactive agents on response timing, single-step response correctness, and multi-step response correctness. Experiments on eight LLMs and pipelines reveal that current agents frequently overact and struggle with event cancellation. Notably, even GPT-5.1 only reacts correctly in 26.7% of scenarios. Further qualitative analysis reveals fundamental limitations of current LLMs as proactive agents, particularly in detecting implicit events and reasoning from the user's first-person perspective.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes