AIAug 2

Toward Fine-Grained Forgetting:Attribute Unlearning for Multimodal Large Language Models

arXiv:2608.0100817.5
Predicted impact top 24% in AI · last 90 daysOriginality Highly original
AI Analysis

This work tackles the problem of fine-grained sensitive attribute disclosure in MLLMs, which is a critical privacy concern for users and organizations deploying these models.

This paper introduces attribute-level unlearning for Multimodal Large Language Models (MLLMs) to address the problem of models memorizing and disclosing sensitive information. The authors propose Causal Localization and Retain-Aware Projection (CLRP), a training-free framework that selectively removes target-attribute subspaces while preserving non-sensitive information about the same identity. CLRP demonstrates effectiveness across various MLLMs and architectures.

Multimodal large language models (MLLMs) exhibit strong vision--language capabilities but may also memorize and disclose sensitive information. Machine unlearning seeks to remove designated knowledge without retraining from scratch while preserving general utility. Existing privacy-oriented benchmarks primarily adopt profile-level deletion, whereas practical requests are often finer grained: a model should forget a specified attribute while retaining non-sensitive information about the same identity. We therefore introduce attribute-level MLLM unlearning as a finer-grained task and construct a benchmark spanning long-text, numeric, and short-text targets, multiple forget ratios, and diverse question types. Our evaluation reveals that target and retained attributes share identity-specific and visual evidence, making selective forgetting susceptible to residual leakage or collateral degradation; accordingly, existing methods exhibit unstable forgetting--retention trade-offs in this setting. To address this challenge, we propose Causal Localization and Retain-Aware Projection (CLRP), a lightweight training-free framework. CLRP uses activation patching to identify the layer that causally mediates target-attribute disclosure, then applies a retain-aware projection that removes the target-attribute subspace while preserving same-identity evidence. Experiments across multiple widely used MLLMs with distinct architectures and parameter scales demonstrate the effectiveness of CLRP.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes