Unlocking Proactivity in Task-Oriented Dialogue
arXiv:2605.22240v2 Announce Type: replace Abstract: Proactive task-oriented dialogue (TOD), such as outbound sales, demands a persuasive agent that actively probes the user's concerns and steers the c
Knowledge catalogue
arXiv:2605.22240v2 Announce Type: replace Abstract: Proactive task-oriented dialogue (TOD), such as outbound sales, demands a persuasive agent that actively probes the user's concerns and steers the c
This week we've got tandem hands-ons with Google's new Gemini AI agent - Spark - from my colleagues David Pierce and Jay Peters. Their takeaways are similar: It's so effective that it's scary. Spark k
arXiv:2606.03920v1 Announce Type: new Abstract: Understanding a video requires more than recognizing isolated moments, as humans continuously track entities, states, and events over time. This capacit
arXiv:2606.02973v1 Announce Type: new Abstract: Are utterances by AI chatbots meaningful? Concretely, if a user asks, say, Anthropic's agent Claude, 'What is the capital of Spain?' and Claude answers,
arXiv:2606.03682v1 Announce Type: new Abstract: Embodied navigation connects intelligent agents with the physical world and is fundamental for general robotic intelligence. Limited availability and qu
Jerry Liu built one of the most installed pieces of AI plumbing of the last three years. Then he sat down and told me the framework era he helped create is over. The agent harness ate the abstraction
arXiv:2606.03243v1 Announce Type: new Abstract: Modern text-to-image models have achieved strong visual synthesis, yet remain unreliable when prompts require implicit visual constraints, relational re
Nanocoder 1.27.0 is an agentic coding tool available in your terminal that runs on any AI model you choose, whether local models via Ollama or cloud providers like OpenAI and Anthropic. This release i
arXiv:2606.03890v1 Announce Type: new Abstract: Multimodal agents in robotics, AR, and autonomous driving must reason about places and layouts from continuous egocentric streams, often using evidence
🆕Scaling Past Informal AI https://latent.space/p/axiom @axiommathai founder & CEO @CarinaLHong explains why math may be the missing path from code agents to AGI, why verified AI is about scaling brill
arXiv:2606.03371v1 Announce Type: new Abstract: Multimodal retail agents should not only recognize what a customer is doing, but also decide whether and how to assist before an explicit request is mad
arXiv:2606.03237v1 Announce Type: new Abstract: AI's central challenge is shifting from capability to coexistence. The dominant paradigm in AI research focuses on developing powerful agents that treat
Today at Microsoft Build: Pinecone Nexus now integrates with Microsoft OneLake. Your enterprise data, turned into task-scoped, governed, cited knowledge your agents can use directly. 95%+ token reduct
Try the most natural TTS and cost-effective STT APIs in @Vapi_AI Grok STT and Grok TTS from @xai are now live on Vapi, the platform for enterprise voice AI. Build on Vapi to create custom voice agents
Uber Caps Usage of AI Tools Like Claude Code to Manage Costs I wrote the other day about Uber blowing its 2026 AI budget in four months, and how that wasn't particularly surprising given they would ha
arXiv:2606.03273v1 Announce Type: cross Abstract: Visual DeepSearch requires multimodal large reasoning model (MLRM) agents to answer complex visual queries by repeatedly inspecting image regions, gro
We’re bringing new capabilities to GPT-Rosalind, a model series purpose-built for life sciences research at enterprise scale. It brings GPT-5.5’s agentic coding and tool use together with stronger int
You can use Hermes Desktop with Ollama using local or cloud models. Get started 👇👇👇 The next evolution of Hermes Agent is here! Introducing Hermes Desktop: everything you love about Hermes, now native
arXiv:2606.00095v1 Announce Type: cross Abstract: Vision-Language Navigation (VLN) enables embodied agents to reach target locations in unseen environments by following language instructions. Despite
arXiv:2606.00145v1 Announce Type: cross Abstract: Vision-language-action (VLA) agents can execute natural-language instructions, yet deployed systems still lack an operational interface: deciding when
arXiv:2606.01234v1 Announce Type: cross Abstract: The GDP of a country is modelled as the relative interaction between two agents - working hours, reflecting the social choice of a population, and Tot
arXiv:2606.02246v1 Announce Type: new Abstract: To operate in the physical world, embodied agents must perceive their environment in an 'always-on' fashion, selectively accessing the most informative
arXiv:2606.00151v1 Announce Type: cross Abstract: In reinforcement learning (RL), agents benefit from exploration only because they repeatedly encounter similar states: trying different actions can im
Excited for Ollama to be working with @nvidia for the NVIDIA RTX Spark launch! NVIDIA RTX Spark: a 1-petaflop superchip, the full CUDA and RTX ecosystem, and Windows-native agents. A new beginning for
According to every product demo from the last four years, planning a trip is a killer use case for AI. Just tell it where you're going, they all promise, and your chatbot / agent / other buzzword will
arXiv:2606.02556v1 Announce Type: new Abstract: We introduce HERO'S JOURNEY, a benchmark for rule induction in goal-directed episodic tasks, where agents must infer hidden rules from demonstrations an
arXiv:2606.01565v1 Announce Type: cross Abstract: Vision-Language Navigation in Continuous Environments (VLN-CE) poses a formidable challenge for autonomous agents, requiring seamless integration of n
arXiv:2606.00266v1 Announce Type: cross Abstract: A long-standing challenge in distributed wireless systems is ensuring efficient and fair random channel access. Existing solutions often address speci
arXiv:2606.01950v1 Announce Type: cross Abstract: World models enable intelligent agents to predict the consequences of their actions on the environment. In this paper, we propose Multi Rigid Object G
arXiv:2606.02463v1 Announce Type: cross Abstract: In 3D environments, Embodied Agents answer spatially relevant questions through reasoning from a mixture of modalities including natural language, RGB
Microsoft Corp. today announced it’s bringing agentic development capabilities into the hands of developers with a new desktop form factor supercomputer called the Surface RTX Spark Dev Box. Developed
arXiv:2606.00240v1 Announce Type: new Abstract: Effective real-world assistance requires AI agents with robust Theory of Mind (ToM): inferring human mental states from their behavior. Despite recent a
arXiv:2606.01313v1 Announce Type: cross Abstract: Open-vocabulary navigation requires embodied agents to manage significant perception uncertainty stemming from semantic ambiguity and model errors. Ho
arXiv:2512.18333v2 Announce Type: replace-cross Abstract: This paper proposes a new Reinforcement Learning (RL) based control architecture for quadrotors. With the literature focusing on controlling t
arXiv:2606.00270v1 Announce Type: new Abstract: Shielding is an effective approach to formally guarantee the safety of reinforcement learning agents in Markov decision processes (MDPs). However, exist
arXiv:2603.11653v2 Announce Type: replace Abstract: Continual Reinforcement Learning (CRL) for Vision-Language-Action (VLA) models is a promising direction toward self-improving embodied agents that c
arXiv:2606.01790v1 Announce Type: cross Abstract: Vision-language-model-based graphical user interface (GUI) agents have shown broad automation capabilities, yet deployment is bottlenecked by a key-va
arXiv:2510.09608v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) could power real-time assistants and autonomous agents, but they face a critical challenge: understanding near-i
arXiv:2606.01249v1 Announce Type: cross Abstract: On-Policy Distillation (OPD) is a fundamental technique for efficient post-training of large language models (LLMs), with broad applications in agent
Microsoft Build 2026 brings a major set of Microsoft Foundry updates for developers building agents: hosted runtimes, Toolboxes, memory, Voice Live, Foundry IQ, new models, managed compute, and trust,
arXiv:2511.17826v2 Announce Type: replace-cross Abstract: Deterministic inference is increasingly critical for large language model (LLM) applications such as LLM-as-a-judge evaluation, multi-agent sy
arXiv:2605.31557v1 Announce Type: new Abstract: Continuous episodic memory is a core capability for autonomous agents operating in dynamic, real-world environments, yet current streaming video benchma
arXiv:2605.31464v1 Announce Type: cross Abstract: GPU kernels are the workhorse of modern deep learning, and optimizing them (via evolutionary search or coding agents) usually requires repeated measur
Application owners and platform engineers have long faced a difficult choice: spend excessively by over-provisioning to guarantee quick startups, or minimize costs but endure slow cold starts. We are
arXiv:2602.02220v2 Announce Type: replace Abstract: Language-conditioned goal navigation (LGN) requires agents to locate user-specified targets without step-by-step guidance. However, existing benchma
Last week we revamped Liteparse to be the fastest PDF parser out there ⚡️ An underrated part of liteparse is it doesn't just give you text. It gives you bounding boxes that a coding agent can use to p
Latest latentspace pod is excellent. @EthanHe_42 really gets it and lays out a lot of the thinking that led us to Flipbook and how to think about the future of generative UI agents 🤝 video gen 🤝 users
arXiv:2412.18134v4 Announce Type: replace Abstract: Randomized self-reductions (RSRs) express f(x) using f evaluated at random correlated points, enabling self-correcting programs, instance-hiding pro
arXiv:2605.30571v1 Announce Type: cross Abstract: Physical AI systems, including robots, autonomous vehicles, embodied agents and edge copilots, often run a different inference workload from cloud LLM
arXiv:2605.30889v1 Announce Type: cross Abstract: Constructing production-quality machine-learned interatomic potentials (MLIPs) requires balancing accuracy, dynamical stability, and computational thr
OpenAI frontier models GPT-5.5 and GPT-5.4, and Codex, the OpenAI coding agent, are now generally available on Amazon Bedrock. AWS customers can access these latest OpenAI models through the same Amaz
arXiv:2605.30503v1 Announce Type: new Abstract: Learning to reach arbitrary goals from sparse feedback requires agents to infer a rich notion of reachability across state--goal pairs. Goal-conditioned
arXiv:2605.31545v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve from general-purpose assistants to user-centric agents, personalization has become central to aligning model beha
arXiv:2605.30664v1 Announce Type: new Abstract: Subgoal-based policy tree search, which uses a policy to guide search, is effective for complex single-agent deterministic problems but often relies on
This pod was an incredible gift to the community: not only our first pod about @xAI, but Ethan really indulged on all our questions on how to train a SOTA Videogen world model, including specific area
我前不久也吐槽过 Claude Desktop 的问题,不只是这个标签页合并的问题,右侧的面板也是相当糟糕的设计 https://x.com/dotey/status/2055777343222808744 OpenAI 因为 ChatGPT 太成功所以他们没有太在意 Coding Agent; 然后 Anthropic 抓住了机会做出了 Claude Code; Claude Code 在 TU
Another proof point for the open-weights thesis. From @RampLabs: 'If we built this again, we'd lean more on open-weight models.' Ramp pointed 10K agents at their own backend. Kimi K2.6 and DeepSeek V4
arXiv:2605.28969v1 Announce Type: cross Abstract: If an AI agent makes decisions on a person's behalf, those decisions must align with its user. We introduce representational accuracy to measure how f
arXiv:2605.29271v1 Announce Type: new Abstract: Tool retrieval over large API catalogs is a core bottleneck for LLM agents: user queries arrive in colloquial, often underspecified language, while the
arXiv:2410.07287v2 Announce Type: replace-cross Abstract: Climate change poses an existential threat, necessitating effective climate policies to enact impactful change. Decisions in this domain are i