AICLAug 4

FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents

arXiv:2608.0409510.0
Predicted impact top 62% in AI · last 90 daysOriginality Highly original
AI Analysis

This benchmark addresses a critical gap in evaluating personalized memory and preference adaptation for LLM agents, particularly for their use as personalized assistants in high-stakes domains like financial advising.

This paper introduces FinPerMA, a new benchmark to evaluate whether LLM agents can maintain and update personalized user models over long horizons, specifically in financial advising. It found that current frontier LLMs and memory configurations are far from saturated, with no full-context configuration exceeding approximately 0.47 overall accuracy or 39% on multiple-choice questions.

Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear whether they can maintain and update an individualized user model over long horizons. Existing personalized-memory benchmarks primarily test factual retention or rely on weakly constrained model-generated trajectories, leaving event-driven preference adaptation underexplored. We introduce FinPerMA, an event-grounded benchmark that evaluates personalized memory against frozen longitudinal investor trajectories. Its generation pipeline combines deterministic, theory-informed impact rules, controlled LLM narration, and automated quality screening; a Post-Shock checkpoint isolates whether an agent has integrated a material event into its persistent user model. On 2,994 questions from 276 personas, seven frontier LLMs and up to seven memory configurations remain far from saturated: no full-context configuration exceeds approximately 0.47 overall accuracy or approximately 39% on multiple-choice questions. Attribution analysis shows that summary-based memory often preserves factual details while losing the preference signals needed for personalization; simple retrieval can therefore outperform purpose-built memory systems, with the gap widening after shocks.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes