AIJun 19

Trip+: Benchmarking Agents in Personalized Interactive Travel Planning

arXiv:2606.2116918.1
Predicted impact top 28% in AI · last 90 daysOriginality Incremental advance
AI Analysis

For researchers developing travel planning agents, Trip+ provides a holistic benchmark that reveals a gap in experiential quality, highlighting the need for models to better balance feasibility with personalization.

Trip+ benchmarks language model agents on personalized interactive travel planning, finding that models favor technically feasible but exhausting itineraries that diverge from traveler preferences.

Interactive travel planning has become a popular use case for language models. Agents are deployed to manage evolving preferences and unexpected disruptions over multiple turns. Such settings require models to make complex, profile-conditioned planning decisions. However, existing benchmarks often evaluate feasibility, personalization, or interaction in relatively isolated settings. We therefore introduce Trip+ to measure the ability of agents to plan travel holistically. In Trip+, given traveler profiles and dynamic interactions, agents must generate and revise minute-level itineraries. End-to-end traveler experiences are evaluated via an LLM-based simulator, enabling the assessment of subjective metrics like fatigue. Our scenarios range from simple request resolutions to complex environment-driven replanning. We evaluate 18 LMs and find a consistent gap in experiential quality. Models favor technically feasible but exhausting itineraries that diverge sharply from profiled traveler preferences.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes