SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards
arXiv:2511.07403v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved remarkable progress in vision-language tasks, but continue to struggle with spatial rea