CVJul 9

3D FaceShell: Attribute Transfer in 3D Face Avatars as a VLM Defense Mechanism

arXiv:2607.162808.4h-index: 9
Predicted impact top 51% in CV · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the emerging privacy threat of VLM-based attribute inference from 3D face avatars, offering a defense that operates directly on 3D representations rather than 2D images.

3D FaceShell introduces a framework to protect privacy in 3D face avatars by adding subtle perturbations that mislead vision-language models (VLMs) from inferring sensitive attributes, achieving high attribute injection and mismatch rates while preserving visual fidelity and identity.

Photorealistic 3D face avatars are increasingly deployed as reusable digital assets across applications such as telepresence, animation, and personalized media. At the same time, vision-language models (VLMs) can infer sensitive attributes from rendered images with open-ended semantic reasoning without any fine-tuning. This creates a new privacy challenge: once a 3D face avatar is shared, any of its renderings can be analyzed to extract high-level facial attributes. Existing defenses largely operate in 2D image space and do not address identity-preserving semantic manipulation of 3D facial representations. We propose 3D FaceShell, a framework for steering VLM interpretations of faces rendered from 3D models while preserving geometric fidelity and facial identity. 3D FaceShell augments the original 3D representation with a learnable Gaussian shell that produces subtle, spatially distributed perturbations optimized through multi-view embedding alignment. The perturbations are designed to be visually inconspicuous yet sufficient to redirect VLM-based attribute inference in a view-consistent manner. Extensive experiments on reconstructed celebrity face avatars and multiple black-box VLMs demonstrate that 3D FaceShell significantly increases attribute injection and mismatch rates while maintaining high perceptual similarity and identity consistency. Our results show that it is possible to manipulate VLM-level semantic interpretation of 3D faces without compromising their human-recognizable appearance.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes