Agents
Ludi{}_{scriptscriptstyle 0.1}: An Agentic System for Socially Intelligent Robots
arXiv:2608.22035v1 Announce Type: new Abstract: Robot foundation models have substantially advanced perception and control, but natural human-robot collaboration requires more than executing isolated
arXiv:2608.22035v1 Announce Type: new Abstract: Robot foundation models have substantially advanced perception and control, but natural human-robot collaboration requires more than executing isolated commands. A robot must recognize ambiguity, maintain context across turns, communicate its intentions, and revise ongoing behavior as the user's intent changes. We present scriptstylemathsf{Ludi}{scriptscriptstyle 0.1}, an agentic system for socially intelligent robots that integrates interactive speech, multimodal reasoning, memory, navigation, and learned manipulation. Its decision-making core is a fine-tuned vision-language model trained on multi-turn interaction traces spanning ambiguous requests, clarifications, corrections, interruptions, mixed social and task dialogue, and multi-step tasks. A purpose-built harness manages the model-tool interaction loop, while specialized navigation and manipulation policies execute physical skills. Ludi{}{scriptscriptstyle 0.1} demonstrates a practical path toward fluid human-robot collaboration today while producing the multimodal interaction traces needed to develop a more deeply integrated foundation model for robots and people.
Source: arXiv cs.RO | 2026-08-25