Local Ai
LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition
arXiv:2607.19889v1 Announce Type: new Abstract: Understanding instrument-tissue interactions is essential for context-aware surgical AI and autonomous robotic surgery. Pretrained vision-language model
arXiv:2607.19889v1 Announce Type: new Abstract: Understanding instrument-tissue interactions is essential for context-aware surgical AI and autonomous robotic surgery. Pretrained vision-language models (VLMs) and vision encoders offer an alternative to conventional interaction classifiers by transferring broad visual and semantic knowledge. However, adapting them to fine-grained surgical interactions remains challenging: (1) freezing the vision encoder depends entirely on pretrained representations that may retain noise and provide weak spatial localization, while (2) full fine-tuning can improve global semantic alignment without ensuring that the encoder learns meaningful features in the correct action region. We address these limitations by introducing LAViFiT, an end-to-end latent-action-guided framework for vision-language fine-tuning. An inverse dynamics model captures the visual changes induced by each action, while a forward world model drives the encoder to represent action-relevant regions. A patch-level SIG Regularizer further prevents local feature collapse without additional supervision, such as bounding boxes or pseudo-labels. Experiments across multiple encoders and datasets improve recognition and image-text alignment, while representation analyses show stronger grounding over the complete instrument-tissue interaction region and more spatially coherent features.
Related
- AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding
- ST-pi: Structured SpatioTemporal VLA for Robotic Manipulation
- Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation
- A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding
Source: arXiv cs.CV | 2026-07-23