How reliable are LLMs when it comes to playing dice?

arXiv:2606.075154.3
Predicted impact top 36% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and practitioners relying on LLMs for reasoning tasks, this study reveals that current models lack robust probabilistic reasoning despite strong performance on standard benchmarks.

LLMs achieve 96% accuracy on standard probability problems but only 59% on counterintuitive ones, with performance dropping over 20% under disguised formulations and up to 34% under misleading prompts, indicating they are not genuine probabilistic reasoners.

We investigate the probabilistic reasoning capabilities of large language models through a controlled benchmarking study on discrete probability problems. We constructed two datasets, respectively a set of standard exercises and a set of counterintuitive exercises, designed to trigger heuristic reasoning, and evaluated 8 state-of-the-art models, each tested with and without Chain-of-Thought prompting. Models achieve an average accuracy of 0.96 on standard problems but only 0.59 on counterintuitive ones. We further provide empirical evidence of token bias: performance drops by over 20% when canonical formulations are replaced by disguised variants. Embedding misleading suggestions in the prompt reduces performance by up to 34%, with no model proving immune. Taken together, the reported findings suggest that current LLMs are not yet genuine probabilistic reasoners, despite their success in advanced mathematical problems.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes