CVMar 11

EmoStory: Emotion-Aware Story Generation

arXiv:2603.10349v137.81 citationsh-index: 6
Predicted impact top 9% in CV · last 90 daysOriginality Highly original
AI Analysis

This work addresses the need for emotion-aware story generation to engage audiences more effectively, representing a novel task rather than an incremental improvement.

The paper tackled the problem of generating visual stories that incorporate explicit emotional directions, which existing methods overlooked, and introduced EmoStory, a two-stage framework that outperformed state-of-the-art methods in emotion accuracy, prompt alignment, and subject consistency on a new dataset of 600 emotional stories.

Story generation aims to produce image sequences that depict coherent narratives while maintaining subject consistency across frames. Although existing methods have excelled in producing coherent and expressive stories, they remain largely emotion-neutral, focusing on what subject appears in a story while overlooking how emotions shape narrative interpretation and visual presentation. As stories are intended to engage audiences emotionally, we introduce emotion-aware story generation, a new task that aims to generate subject-consistent visual stories with explicit emotional directions. This task is challenging due to the abstract nature of emotions, which must be grounded in concrete visual elements and consistently expressed across a narrative through visual composition. To address these challenges, we propose EmoStory, a two-stage framework that integrates agent-based story planning and region-aware story generation. The planning stage transforms target emotions into coherent story prompts with emotion agent and writer agent, while the generation stage preserves subject consistency and injects emotion-related elements through region-aware composition. We evaluate EmoStory on a newly constructed dataset covering 25 subjects and 600 emotional stories. Extensive quantitative and qualitative results, along with user studies, show that EmoStory outperforms state-of-the-art story generation methods in emotion accuracy, prompt alignment, and subject consistency.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes