CVJun 20

Feed-forward Motion In-betweening for Any 4D

arXiv:2606.2213113.3
Predicted impact top 32% in CV · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the need for efficient and controllable long-horizon 4D mesh generation, which is important for animation and world modeling applications.

The paper proposes a feed-forward in-betweening framework for generating intermediate frames of 4D meshes conditioned on sparse keyframes, achieving strong performance on DyMesh16 and DyMesh32 benchmarks.

4D dynamics (3D geometry evolving over time) is a fundamental representation of the physical world and plays a crucial role in world modeling (e.g., animation and games). Owing to the scarcity of large-scale, long-horizon 4D mesh data with arbitrary shapes, early text-to-4D methods rely on distillation or test-time optimization from video diffusion priors, making inference prohibitively slow. Recent feed-forward generators greatly reduce inference cost but offer limited spatiotemporal controllability, and short-horizon generation often leads to error accumulation in long-horizon sequences. We propose a novel feed-forward in-betweening framework for arbitrary 4D meshes with keyframe conditioning. Building on universal mesh-animation latents, we introduce a frame-wise mesh VAE that encodes each frame into topology-agnostic latent tokens anchored by a reference mesh for keyframe conditioning. We further introduce a keyframe-conditioned rectified flow model with an MMDiT backbone that synthesizes non-keyframe frames conditioned on sparse keyframes. Experiments show strong performance and improved controllability on both DyMesh16 and DyMesh32 benchmarks.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes