Seeing Once is Enough? Online Geometry-Aware Token Pruning for 3D Question Answering
arXiv:2607.04079v1 Announce Type: cross Abstract: Recent Multi-modal Large Language Models (MLLMs) have demonstrated remarkable performance on 2D question answering tasks. However, extending these mod