HCAIJul 20

Querying Multimodal Scientific Papers with AI: Practices and Preferences Across Blind, Low-Vision, and Sighted Scientists

arXiv:2607.185145.8h-index: 4
Predicted impact top 51% in HC · last 90 daysOriginality Synthesis-oriented
AI Analysis

For researchers in AI accessibility and scientific document understanding, this work provides qualitative insights and a benchmark dataset on how scientists with and without visual impairments interact with AI-powered visual question-answering systems.

This study interviews five blind/low-vision and five sighted scientists to understand how they use AI tools (ChatGPT, Gemini) for querying multimodal scientific documents. It identifies that vague image descriptions and incorrect AI outputs cause both groups to abandon AI workflows, and provides a dataset of 115 queries and responses.

Visual diagrams, figures, and tables are central to scientific papers, and convey information beyond what is captured in text. While blind or low-vision (BLV) scientists have traditionally relied on static alternative text to access figures in papers, the rise of artificial intelligence (AI) has made interactive question-answering (QA) a feasible paradigm for visual exploration; yet little is known about how scientists use visual QA in practice or how to improve its accessibility. In this work, we interview five BLV and five sighted scientists across different STEM fields to understand how they use two AI tools, ChatGPT and Gemini, to query multimodal scientific documents. Our findings characterize how scientists review multimodal content, including existing practices (along with accessibility workarounds) for engaging with visuals, and feedback on the suitability of AI-generated responses to multimodal queries. We further find that vague or incomplete image descriptions, as well as incorrect AI outputs more broadly, can cause both BLV and sighted scientists to abandon AI workflows. To support future research, we additionally contribute a dataset of 115 queries and responses from our participants' interactions with the AI tools for papers in their field. We close by discussing implications for AI-powered scientific QA systems, emphasizing considerations for access across abilities and domains.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes