CVAIJun 30

LUNA: Learning Universal 3D Human Animation Beyond Skinning

arXiv:2606.3198112.3
Predicted impact top 27% in CV · last 90 daysOriginality Highly original
AI Analysis

For 3D human avatar creation from monocular images, LUNA removes reliance on LBS and parametric models, enabling more expressive and artifact-free animation.

LUNA proposes an LBS-free neural animation model that directly maps 2D controls (images, keypoints, sketches) to 3D Gaussian deformations, achieving competitive visual fidelity and zero-shot cross-identity generalization without explicit body fitting.

Creating photorealistic, animatable 3D human avatars from monocular images still largely depends on Linear Blend Skinning (LBS) and parametric body models, which constrain expressivity and often introduce artifacts due to imperfect fitting. We propose LUNA, an LBS-free universal neural animation model that directly maps multiple 2D controls like images, keypoints, sketches, and unseen characters into 3D Gaussian deformations, bypassing explicit body fitting. At its core, a transformer-based motion regressor disentangles global rigid motion from fine-grained local dynamics to capture both coherent movement and subtle non-rigid effects. To resolve the inherent ambiguity of 2D-to-3D lifting while scaling beyond fitted datasets, we introduce hybrid supervision that distills soft structural priors from an LBS teacher and a loss that supports training on both limited fitted data and large in-the-wild unlabeled videos. Extensive experiments show LUNA achieves competitive visual fidelity compared to LBS-based approaches, while delivering realistic human motion and zero-shot cross-identity generalization across diverse driving modalities. To the best of our knowledge, LUNA is the first end-to-end 3D animatable model that supports implicit 2D driving.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes