Multimodal Graph RAG for Long-range Visually Rich Document Understanding
DGX agentarXiv:2606.28780v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are widely applied to visual document understanding. However, comprehending long documents remains an issue b