ShiJianBench: From Dialogue to Decision for Long-Horizon Evaluation of Investment Advisors
This work provides a new benchmark and evaluation methodology for conversational AI in finance, addressing the need to assess long-term impact on user decisions, which is a domain-specific but important problem for AI advisors.
The paper introduces ShiJianBench, an offline evaluation framework for conversational investment advisors that simulates investor trajectories under historical market conditions. Using a multi-agent investor simulator calibrated on 7,199 real users, they evaluate LLM advisors on Chinese fund-market data from 2021-2026, finding a stable leading group that combines strong personalized content with competitive investor-side outcomes, highlighting the difference between response quality and long-horizon intervention effectiveness.
Conversational investment advisors influence not only what users know, but also how they make subsequent decisions as market conditions evolve. Existing evaluations primarily assess response quality or observed outcomes, leaving the long-horizon pathway from advisor language to investor behavior difficult to audit. We introduce ShiJianBench, an offline framework for evaluating conversational investment advisors through matched investor trajectories under fixed historical market feedback. At its core is a multi-agent investor simulator with explicit evolving state variables, motive-driven deliberation, long-term memory, and dialogue-grounded updates. The simulator is calibrated against aggregate behavioral patterns from 7,199 real users, and advisor policies are evaluated using separate investor-side, service-side, and content-side metrics under a hard compliance gate. Experiments on Chinese fund-market traces from 2021 to 2026 identify a stable leading group of LLM advisors that combines substantially stronger personalized content with competitive investor-side trajectory outcomes. These results reveal a systematic distinction between producing a high-quality response and delivering an effective long-horizon intervention, motivating trajectory-aware evaluation of conversational advisors.