PFJun 13

LLMs have Visualization Literacy: Now What? Experiments Exploring LLM Visualization Evaluation Capabilities

arXiv:2606.1513614.2
Predicted impact top 25% in PF · last 90 daysOriginality Synthesis-oriented
AI Analysis

For visualization researchers and practitioners, this work reveals that while LLMs have improved in literacy, they remain unreliable as evaluators due to deficits in instruction following and graphical integrity.

The authors tested recent LLMs (Claude Opus 4.5, GPT 5.2 Pro, Gemini 3 Flash) on visualization literacy, instruction following, and graphical integrity. They found that LLMs now exceed human-level visualization literacy but still struggle with instruction following and identifying misleading visualizations, questioning their effectiveness as evaluators.

As Large Language Models (LLMs) become more popular within the visualization community, researchers increasingly leverage them for diverse visualization tasks such as design guideline suggestions and visualization evaluation. However, in order for LLMs to act as trustworthy and fair evaluators, we argue that LLMs would need to possess visualization literacy, be capable of following user instructions and uphold graphical integrity. We test the latest versions of the most prominent LLMs, specifically Anthropic's Claude (Opus 4.5), OpenAI's Generative Pretrained Transformers (GPT 5.2 Pro), and Google's Gemini (Gemini 3 Flash) on these features and find that while these models now possess visualization literacy, they still struggle with other features necessary for instruction following and graphical integrity. Using a modified Visualization Literacy Assessment Test (VLAT), our findings show that these recent LLMs have achieved greater than human-levels of visualization literacy in contrast to prior research. In order to test the models' abilities to follow instructions, we used few-shot and chain-of-thought prompting as proxies for instruction following tasks on evaluating visualization literacy and find that these specialized prompting techniques are becoming obsolete with respect to improving visualization literacy. Additionally, we experiment with the inherent ability of LLMs to evaluate misleading visualizations to test the models' abilities for upholding graphical integrity and find that without specialized or leading prompting techniques, the models struggle with being able to accurately identify whether a visualization is misleading or not. Our results further break down the performance of each model on these tasks, but the culmination of our findings force us to reconsider the current effectiveness of LLMs as visualization evaluators.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes