9.2CYJun 22
Examining AI-generated historical narratives and their reception through the example of history POVs on TikTokNina Brolich, Anna Neovesky
This paper examines the history POV trend on TikTok, in which AI-generated first-person scenes depict historical events. We use a two-stage empirical approach: an exploratory pilot study and a larger-scale study building up on a dataset obtained through the TikTok Research API. In both studies we analyze the themes of the trend and how the audience responds in the comments. Findings show a dominance of emotionally charged contemporary history topics, with historical inaccuracies visible at the caption level. A comparative comment analysis of Black Death and Holocaust videos, combining manual annotation with DistilBERT-based classification, reveals that topic choice shapes audience response, with Holocaust content attracting disproportionately higher rates of hate speech and disinformation. The paper also reflects on the strengths and limitations of API-based research for studying fast-moving platform trends.
3.7CVNov 11, 2024
Nuremberg Letterbooks: A Multi-Transcriptional Dataset of Early 15th Century Manuscripts for Document AnalysisMartin Mayr, Julian Krenz, Katharina Neumeier et al.
Most datasets in the field of document analysis utilize highly standardized labels, which, while simplifying specific tasks, often produce outputs that are not directly applicable to humanities research. In contrast, the Nuremberg Letterbooks dataset, which comprises historical documents from the early 15th century, addresses this gap by providing multiple types of transcriptions and accompanying metadata. This approach allows for developing methods that are more closely aligned with the needs of the humanities. The dataset includes 4 books containing 1711 labeled pages written by 10 scribes. Three types of transcriptions are provided for handwritten text recognition: Basic, diplomatic, and regularized. For the latter two, versions with and without expanded abbreviations are also available. A combination of letter ID and writer ID supports writer identification due to changing writers within pages. In the technical validation, we established baselines for various tasks, demonstrating data consistency and providing benchmarks for future research to build upon.