MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments
arXiv:2606.31966v1 Announce Type: cross Abstract: Recent multimodal large language models (MLLMs) have strong potential as embodied agents, but their ability to collaborate in visually grounded enviro