SICYJun 14

Modeling Engagement with Brand and Organizational TikTok Videos Using Machine-Assisted Theory-Ensemble Annotation

arXiv:2606.160538.3
Predicted impact top 21% in SI · last 90 daysOriginality Synthesis-oriented
AI Analysis

For researchers studying short-form video at scale, this provides a method to operationalize interpretive theories computationally, though gains are incremental.

The authors used multimodal LLMs to annotate 77 theory-driven structural variables in ~10,000 TikTok videos, achieving modest but consistent gains over baselines in predicting engagement. Human validation showed reliable coding for perceptual/communicative variables but difficulty with deeper semiotic constructs.

Short-form video is difficult to study at scale because meaning emerges through audiovisual elements, language, and participatory, algorithmic and trend-based platform dynamics. Manual annotation of these layers is laborious at scale and difficult to standardize. We demonstrate how multimodal large language models (LLMs) can help address this bottleneck by annotating a set of 77 theory-driven structural variables derived from narratology, rhetoric, communication, and semiotics. We use this to explore content and estimate engagement with modest but consistent gains over account-size and video-age baselines in a corpus of about 10,000 TikTok videos of brand and organizational accounts from Estonia (covering a substantial share of the small country ecosystem). Human validation shows a reliability gradient: perceptual and communicative variables can be coded fairly reliably, while deeper semiotic and archetypal constructs are more difficult for both humans and machines. This approach of computational operationalization of long-standing interpretive theories can support several aims: exploratory cultural analytics of variation in short-form video culture, predictive modeling of platform dynamics, engagement, and audience feedback; and diagnostics for content creators to support choosing between structural and narrative strategies. Most annotated variables were not associated with platform success, as expected; the value of LLMs in this setting lies in making it feasible to assess large batteries of theoretically motivated variables, so that the subset carrying signal can be identified and translated into creator-facing guidance for a given niche.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes