CVJul 22

LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition

arXiv:2607.1988911.8
Predicted impact top 26% in CV · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the challenge of adapting pretrained vision-language models to fine-grained surgical interactions, which is important for context-aware surgical AI and autonomous robotic surgery.

LAViFiT improves fine-grained surgical interaction recognition by guiding vision-language fine-tuning with latent action models, achieving better recognition and image-text alignment across multiple encoders and datasets.

Understanding instrument-tissue interactions is essential for context-aware surgical AI and autonomous robotic surgery. Pretrained vision-language models (VLMs) and vision encoders offer an alternative to conventional interaction classifiers by transferring broad visual and semantic knowledge. However, adapting them to fine-grained surgical interactions remains challenging: (1) freezing the vision encoder depends entirely on pretrained representations that may retain noise and provide weak spatial localization, while (2) full fine-tuning can improve global semantic alignment without ensuring that the encoder learns meaningful features in the correct action region. We address these limitations by introducing LAViFiT, an end-to-end latent-action-guided framework for vision-language fine-tuning. An inverse dynamics model captures the visual changes induced by each action, while a forward world model drives the encoder to represent action-relevant regions. A patch-level SIG Regularizer further prevents local feature collapse without additional supervision, such as bounding boxes or pseudo-labels. Experiments across multiple encoders and datasets improve recognition and image-text alignment, while representation analyses show stronger grounding over the complete instrument-tissue interaction region and more spatially coherent features.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes