MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models
DGX agentarXiv:2605.14906v1 Announce Type: new Abstract: Memory is essential for large vision-language models (LVLMs) to handle long, multimodal interactions, with two method directions providing this capabili