CLFeb 24

Semantic Novelty at Scale: Narrative Shape Taxonomy and Readership Prediction in 28,606 Books

arXiv:2602.20647v1h-index: 1Has Code

Originality Incremental advance

AI Analysis

This addresses the challenge of analyzing narrative dynamics at scale for literary scholars and computational linguists, though it is incremental as it builds on existing embedding and clustering methods.

The study tackled the problem of quantifying narrative structure in literature by introducing semantic novelty, a measure based on paragraph-level embeddings, and applied it to 28,606 pre-1920 books, revealing eight narrative shape archetypes and finding that volume (variance of novelty) predicts readership with a partial correlation of 0.32.

I introduce semantic novelty--cosine distance between each paragraph's sentence embedding and the running centroid of all preceding paragraphs--as an information-theoretic measure of narrative structure at corpus scale. Applying it to 28,606 books in PG19 (pre-1920 English literature), I compute paragraph-level novelty curves using 768-dimensional SBERT embeddings, then reduce each to a 16-segment Piecewise Aggregate Approximation (PAA). Ward-linkage clustering on PAA vectors reveals eight canonical narrative shape archetypes, from Steep Descent (rapid convergence) to Steep Ascent (escalating unpredictability). Volume--variance of the novelty trajectory--is the strongest length-independent predictor of readership (partial rho = 0.32), followed by speed (rho = 0.19) and Terminal/Initial ratio (rho = 0.19). Circuitousness shows strong raw correlation (rho = 0.41) but is 93 percent correlated with length; after control, partial rho drops to 0.11--demonstrating that naive correlations in corpus studies can be dominated by length confounds. Genre strongly constrains narrative shape (chi squared = 2121.6, p < 10 to the power negative 242), with fiction maintaining plateau profiles while nonfiction front-loads information. Historical analysis shows books became progressively more predictable between 1840 and 1910 (T/I ratio trend r = negative 0.74, p = 0.037). SAX analysis reveals 85 percent signature uniqueness, suggesting each book traces a nearly unique path through semantic space. These findings demonstrate that information-density dynamics, distinct from sentiment or topic, constitute a fundamental dimension of narrative structure with measurable consequences for reader engagement. Dataset: https://huggingface.co/datasets/wfzimmerman/pg19-semantic-novelty

View on arXiv PDF

Similar