Research
Dynamic Context Adapters: Efficiently Infusing History into Vision-and-Language Models
arXiv:2608.10525v1 Announce Type: cross Abstract: Historical context integration presents a fundamental challenge for Vision-Language Models (VLMs) in sequential decision-making tasks. Current VLMs pr
arXiv:2608.10525v1 Announce Type: cross Abstract: Historical context integration presents a fundamental challenge for Vision-Language Models (VLMs) in sequential decision-making tasks. Current VLMs process visual inputs independently, which creates critical limitations for downstream applications that require temporal understanding. Direct incorporation of historical frames into Transformer inputs produces quadratic attention complexity and excessive memory consumption. Existing approaches suffer from significant drawbacks: computational inflation or substantial information loss through temporal compression. To address these challenges, we introduce Dynamic Context Adapter (DCA), a novel context injection approach for pretrained VLMs. Our method employs fixed-size, dynamically compressed memory to preserve historical semantics without frame concatenation. DCA bridges static VLMs and recurrent policies and enables memory capabilities in pretrained models while maintaining computational efficiency. DCA achieves over 25% reduction in attention FLOPs and 13% memory savings while improving performance on long-horizon tasks.
Related
- LensVLM: Selective Context Expansion for Compressed Visual Representation of Text
- CARES: Context-Aware Resolution Selector for VLMs
- Beyond Classification: Dynamic Adapter Routing for Continual Multimodal Retrieval
- VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
Source: arXiv cs.AI | 2026-08-12