GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video
DGX agentarXiv:2607.02991v1 Announce Type: new Abstract: While multimodal Large Language Models (MLLMs) excel at offline video understanding, an interesting question of how far they are from serving as a real-