CLAIAug 7

PHASE-Tree: Modeling Character-State Evolution in Long-Horizon Role-Playing Dialogue

arXiv:2608.0697510.4h-index: 19
Predicted impact top 10% in CL · last 90 daysOriginality Highly original
AI Analysis

This work is significant for the natural language generation and role-playing dialogue community, providing a method for more consistent and evolving character representation in long-form interactions, which is an incremental improvement over existing static profile methods.

The authors developed PHASE-Tree, a multi-timescale character-state tree, to address the challenge of characters evolving recognizably in long-horizon role-playing dialogue. This model achieved first rank in 11 of 12 dataset-metric cells against internal variants and all 12 cells against external textual baselines on long-dialogue corpora, improving character-level, semantic, and embedding scores by 19.7%, 12.4%, and 15.1% respectively.

Long-horizon role-playing demands that characters remain recognizable as they evolve with the narrative. Yet existing work falls short on two fronts: representations are typically static profiles that cannot be updated locally without destabilizing unchanged traits, and benchmarks mainly test persona preservation and memory recall rather than whether a model speaks from a character's currently evolved state. We address both. PHASE-Tree is a multi-timescale character-state tree with an immutable identity root and mutable persona, session, and moment layers, making each mutable field an addressable target for localized within- and cross-episode updates. It conditions generation through explicit textual provision or implicit parametric adaptation. To measure evolved-state generation, we introduce LongEvoRoleBench, which pairs four long-dialogue corpora for cross-episode evolution with four short-dialogue corpora as within-scene state-tracking checks, under a unified next-utterance protocol. On the long-dialogue core, textual PHASE-Tree ranks first in 11 of 12 dataset-metric cells against internal variants and all 12 cells against external textual baselines, improving character-level, semantic, and embedding scores by 19.7%, 12.4%, and 15.1% respectively. In a blinded 200-response study, human ratings correlate with the GPT-4.1 judge (Pearson r= 0.65); on descriptive n= 10 PT and NR prompt subsets, the Overall difference is +0.20. The long-dialogue Sem advantage persists across LLM judges and generation backbones.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes