CASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation
DGX agentarXiv:2601.08010v3 Announce Type: replace Abstract: Vision-language models achieve strong performance across a wide range of multimodal understanding and reasoning tasks, yet their multi-step reasonin