IRCLDLJul 11

Loci Similes: A Benchmark for Extracting Intertextualities in Latin Literature

arXiv:2601.075337.11 citationsh-index: 2
Predicted impact top 58% in IR · last 90 daysOriginality Synthesis-oriented
AI Analysis

This benchmark addresses the scarcity of standardized datasets for Latin intertextuality detection, enabling the development and evaluation of new methods for scholars.

The paper introduces Loci Similes, a benchmark for detecting intertextualities in Latin literature, comprising ~172k text segments with 545 expert-verified parallels. They establish baselines using state-of-the-art LLMs for retrieval and classification tasks.

Tracing connections between historical texts is an important part of intertextual research, enabling scholars to reconstruct the virtual library of a writer and identify the sources influencing their creative process. These intertextual links manifest in diverse forms, ranging from direct verbatim quotations to subtle allusions and paraphrases disguised by morphological variation. Language models offer a promising path forward due to their capability of capturing semantic similarity beyond lexical overlap. However, the development of new methods for this task is held back by the scarcity of standardized benchmarks and easy-to-use datasets. We address this gap by introducing Loci Similes, a benchmark for Latin intertextuality detection comprising of a curated dataset of ~172k text segments containing 545 expert-verified parallels linking Late Antique authors to a corpus of classical authors. Using this data, we establish baselines for retrieval and classification of intertextualities with state-of-the-art LLMs.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes