M^3-Verse: A 'Spot the Difference' Challenge for Large Multimodal Models
DGX agentarXiv:2512.18735v2 Announce Type: replace-cross Abstract: Modern Large Multimodal Models (LMMs) have demonstrated extraordinary ability in static image and single-state spatial-temporal understanding.