CVAIMay 29

Can You Trust What You See? Human and AI Detection of Synthetic Legal Evidence

arXiv:2606.076138.4h-index: 13
Predicted impact top 60% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For legal practitioners and evidence law, this paper demonstrates that current AI-generated images are indistinguishable from authentic ones by both humans and AI, challenging the reliability of visual evidence in court.

This study evaluates human and multimodal LLM ability to detect AI-generated legal evidence images, finding that humans achieve 64.8% accuracy overall and near-chance performance on the best generators, while MLLMs have 100% specificity but miss most synthetic images (e.g., 5.9% detection on Gemini-3-Pro-Image outputs). Neither group is reliable alone, suggesting a need for combined human, AI, and provenance-based verification.

Visual evidence has long been treated as a reliable form of legal proof, but advances in artificial intelligence (AI) are undermining that assumption. This article asks how well humans and frontier multimodal large language models (MLLMs) can distinguish authentic evidentiary photographs from AI-generated counterparts in the object-centric scenarios typical of civil disputes. We built Synthetic Legal Evidence Detection (SLED-1400), a dataset of 200 authentic evidence images paired with 1,200 synthetic counterparts produced by six contemporary text-to-image generators across ten evidence categories. The same stimuli and response format were used in a controlled web experiment with 136 lay participants and in a standardized evaluation of four MLLMs (GPT-5.1, Gemini-3-Pro, Gemini-3-Flash, Qwen3-VL-235B). Human accuracy was 64.8% overall, and 48.5% and 51.0% on the two strongest generators (Gemini-3-Pro-Image and Flux-2-Max), indistinguishable from chance. MLLMs never misclassified an authentic image (100% specificity), but missed most synthetic outputs from the harder generators, with average MLLM detection at 5.9% on Gemini-3-Pro-Image outputs. Human and MLLM errors were largely uncorrelated, while the four MLLMs were strongly correlated with each other. Neither group is a reliable standalone authenticator. We argue that visual evidence in legal proceedings should be treated as inherently contestable, and that a workable procedural response must combine trained human review, MLLM screening, and provenance infrastructure such as C2PA Content Credentials.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes