AIJun 12

A Definition of Good Explanations and the Challenges Explaining LLM Outputs

arXiv:2606.148385.0
Predicted impact top 90% in AI · last 90 daysOriginality Incremental advance
AI Analysis

Provides a conceptual framework for evaluating explanations in AI, but is primarily philosophical and does not offer empirical results.

The paper proposes a definition of good explanations based on counterfactuals and prior beliefs, and argues that LLM outputs are inherently difficult to explain due to this definition.

How to define a good explanation is a long-standing philosophical debate which has found recent renewed interest in the context of AI outputs. Explainability is crucial for AI adoption in many contexts, but in order to produce good explanations of AI systems, we must first have an understanding of what good explanations are. In this paper we propose a definition inspired by the notion of counterfactual explanations, however we argue that one must also take into account the interlocutor's prior beliefs in each fact that could be offered in an explanation. We explore the ramifications of this definition for AI explainability and, in particular, why LLM outputs are difficult to produce good explanations for.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes