RetailBench: Benchmarking long horizon reasoning and coherent decision making of LLM agents in realistic retail environments
For researchers developing LLM agents for complex, dynamic environments, RetailBench provides a controlled testbed to diagnose failures in long-horizon reasoning and decision-making.
RetailBench is a simulation benchmark for evaluating LLM agents in long-horizon retail management. Over 180-day simulations, even the best LLM agents significantly underperform an oracle policy in net worth and sales, revealing gaps in sustained coherent decision-making.
Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments remains uncertain. We introduce RetailBench, a data-grounded simulation benchmark for evaluating tool-using LLM agents in single-store supermarket operation. RetailBench models retail management as a partially observable decision process and is designed to support thousand-day-scale simulations. In this environment, agents must manage pricing, replenishment, supplier selection, shelf assortment, inventory aging, customer feedback, external events, and cash-flow constraints. We evaluate seven contemporary LLMs under representative agent frameworks over a 180-day evaluation horizon and compare them with a privileged oracle policy. Results show substantial variation across models: only a small subset survives the full evaluation horizon, and even the strongest LLM runs remain substantially behind the oracle policy in final net worth and sales outcomes. Behavioral analysis attributes these gaps to incomplete evidence acquisition, surface-level decision making, and the lack of a consistent long-horizon policy. RetailBench provides a controlled testbed for studying reliable autonomy in economically grounded long-horizon decision-making.