CVJun 24

ShutterMuse: Capture-Time Photography Guidance with MLLMs

arXiv:2606.2576315.2
Predicted impact top 26% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For photographers and AI researchers, this work addresses the underexplored area of capture-time guidance, providing a benchmark and model that outperforms existing methods in composition refinement and pose recommendation.

The paper introduces CaptureGuide-Bench, a benchmark for capture-time photography guidance covering both composition and subject pose, and develops ShutterMuse, a unified MLLM that achieves the best overall photographer-side performance among baselines and competitive subject-side pose recommendation with lower inference cost.

Real-world photography requires capture-time guidance for both camera framing and subject pose. Yet existing aesthetic cropping benchmarks mainly evaluate post-hoc crop prediction and overlook subject-side recommendations, leaving the capture-time guidance capabilities of multimodal large language models (MLLMs) underexplored. To address this gap, we introduce CaptureGuide-Bench, a benchmark with two complementary tasks: photographer-side composition decision and refinement, and subject-side scene-conditioned pose recommendation. Our evaluation reveals limitations: general-purpose MLLMs can make composition decisions but lack precise refinement localization, while specialized aesthetic cropping models localize crops effectively but are limited to refinement; neither provides actionable pose guidance. To support model development, we further construct CaptureGuide-Dataset, comprising 130K samples with textual rationales and structured visual annotations, and develop ShutterMuse, a unified MLLM trained with supervised and reinforcement fine-tuning. Experiments on CaptureGuide-Bench show that ShutterMuse achieves the best overall photographer-side performance among evaluated baselines and competitive subject-side pose recommendation with substantially lower inference cost, demonstrating the potential of MLLMs as interactive assistants for photography during image capture.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes