CVJul 15

Local Brushstroke Quality Assessment via Vision-Language Feedback

arXiv:2607.163308.3h-index: 8
Predicted impact top 51% in CV · last 90 daysOriginality Synthesis-oriented
AI Analysis

For calligraphy education and AI-assisted art evaluation, this work provides a benchmark and reveals limitations of current multimodal LLMs in fine-grained aesthetic assessment.

This paper evaluates whether multimodal LLMs (GPT-4o, Claude Sonnet 4, Gemini 2.5 Flash) can assess local brushstroke quality in calligraphy and generate educational feedback. GPT-4o achieved the best absolute accuracy (MAE=0.885), but no model showed statistically significant rank correlation with human experts.

This paper investigates whether multimodal LLMs can evaluate local brushstroke quality in calligraphy and generate educationally useful natural language feedback. We construct an evaluation framework in which three multimodal LLMs (GPT-4o, Claude Sonnet 4, and Gemini 2.5 Flash) assess before-after image pairs of calligraphic works using a five-point ordinal scale, and compare their outputs against scores assigned by three expert calligraphers. We additionally examine a Retrieval-Augmented Generation (RAG) variant of Claude as a preliminary condition. Results show that all models achieve useful levels of absolute score accuracy (MAE), with GPT-4o performing best (MAE = 0.885). However, none of the models produce statistically significant overall rank correlations with human experts (Kendall's tau). Vocabulary analysis of generated rationales reveals characteristic evaluative biases in each model, and RAG is shown to improve rank correlation while worsening absolute accuracy, constituting an important negative result for text-based rule injection.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes