Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence
DGX agentarXiv:2608.06756v1 Announce Type: new Abstract: Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes