Can LLMs Control Readability? A Multi-Dimensional Evaluation Framework for CEFR-Controlled Arabic Generation
This work provides an empirical foundation for integrating readability-aware Arabic text generation into adaptive educational systems, addressing a gap in controlled generation for Arabic.
The paper evaluates whether LLMs can control readability in Arabic text generation using a multi-dimensional framework. Results show that structured prompting with lexical constraints achieves 0.91 cosine similarity to reference profiles and 0.99 agreement with predicted readability levels, while unconstrained prompting shows weak control.
While Large Language Models (LLMs) can generate fluent Arabic text, their ability to reliably control readability levels remains unclear. We propose a multi-dimensional evaluation framework for Common European Framework of Reference for Language (CEFR)-controlled Arabic text generation, assessing whether instruction-following LLMs can serve as reliable generators for adaptive language learning. Our framework integrates controlled prompting, automatic readability prediction using a validated Taha-19 model, lexical constraint validation, and syntactic complexity profiling. Results show that structured prompting substantially improves CEFR alignment. In particular, CEFR-guided prompting with lexical constraints achieves the highest conformity to reference linguistic profiles (0.91 cosine similarity) and near-perfect agreement with predicted readability levels (0.99), while unconstrained prompting exhibits weak control. These findings establish an empirical foundation for integrating readability-aware Arabic text generation into adaptive educational systems.