CVJun 30

JacobianAvatar: Temporally Consistent Semi-rigid Avatar Reconstruction from a Monocular Video

arXiv:2606.3111511.0
Predicted impact top 34% in CV · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the challenge of reconstructing realistic, temporally coherent human avatars from monocular video, which is important for applications in VR/AR and graphics.

JacobianAvatar reconstructs temporally consistent human avatars with clothing dynamics from monocular video by using neural Jacobian fields to model semi-rigid deformations, outperforming state-of-the-art methods on benchmark and in-the-wild videos.

Generating realistic human avatars in complex motions--such as clothing dynamics--requires modeling of global and local deformations which remains challenging in monocular settings. We address this problem by leveraging neural Jacobian fields (NJFs) for representing semi-rigid deformations. We train self-supervised neural networks for predicting Jacobian matrices that give the pose-dependent deformations, by solving a Poisson equation. However, monocular input presents several difficulties such as self-occluded regions and invisible surfaces. To address these issues, we introduce three key components: a constrained Poisson solver, signed distance-based Jacobian regularization, and a deformation-guided residual flow loss, which together suppress boundary artifacts, recover frequently occluded regions such as armpits and thighs, and enforce temporal consistency during motion. Experiments on benchmark and in-the-wild videos demonstrate that our method generates temporally stable and geometrically coherent avatars, outperforming state-of-the-art approaches.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes