The Measurement Gap in the Automation of EU Law: Benchmarking Doctrinal Legal Reasoning under the EU AI Act
For legal AI developers and regulators, the paper highlights a critical measurement gap that prevents compliance with the EU AI Act's accuracy requirements.
The paper identifies a gap in evaluating whether LLMs perform doctrinal legal reasoning, which is required for high-risk AI under the EU AI Act, and argues that no existing benchmark addresses this, making the Act's accuracy requirement non-operational.
Large language models now produce legal text of at least median quality, yet no existing benchmark can evaluate whether they perform doctrinal legal reasoning, which forms the interpretive core of legal work, rather than the ancillary, paralegal tasks that most current legal-AI evaluations measure. This measurement gap is not only methodological but legal: the EU AI Act makes "appropriate accuracy" a binding requirement for high-risk AI used in the judicial domain, yet that requirement cannot acquire operational content without the very doctrinal-reasoning benchmark the field lacks.