CLAug 18, 2017

The Natural Stories Corpus

arXiv:1708.05763v11096 citations
Originality Synthesis-oriented
AI Analysis

This provides a resource for researchers in psycholinguistics to better test and distinguish language processing theories, though it is incremental as it builds on existing corpus methods.

The authors tackled the lack of low-frequency syntactic constructions in existing corpora for comparing human language processing models by creating a new English corpus edited to include these constructions while maintaining fluency, and they released the data with annotations and reading time data.

It is now a common practice to compare models of human language processing by predicting participant reactions (such as reading times) to corpora consisting of rich naturalistic linguistic materials. However, many of the corpora used in these studies are based on naturalistic text and thus do not contain many of the low-frequency syntactic constructions that are often required to distinguish processing theories. Here we describe a new corpus consisting of English texts edited to contain many low-frequency syntactic constructions while still sounding fluent to native speakers. The corpus is annotated with hand-corrected parse trees and includes self-paced reading time data. Here we give an overview of the content of the corpus and release the data.

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes