Introducing Toolboxes in Foundry
Available in Public Preview Today Toolbox is a new way to curate, configure, and reuse tools across all of your AI agents without rewiring them every time from Foundry. Today, teams build agents acros
Knowledge catalogue
Available in Public Preview Today Toolbox is a new way to curate, configure, and reuse tools across all of your AI agents without rewiring them every time from Foundry. Today, teams build agents acros
LangChain is introducing new testing tools and capabilities that will help developers with test guidance, organization, and completion criteria—moving beyond just providing testing infrastructure to a
Learn more about this model release https://x.com/Alibaba_Qwen/status/2046939764428009914?s=20 🚀 Meet Qwen3.6-27B, our latest dense, open-source model, packing flagship-level coding power! Yes, 27B, a
arXiv:2506.09373v3 Announce Type: replace-cross Abstract: The advent of autonomous agents is transforming interactions with Graphical User Interfaces (GUIs) by employing natural language as a powerful
arXiv:2603.22650v2 Announce Type: replace Abstract: Active mapping aims to determine how an agent should move to efficiently reconstruct unknown environments. Most existing approaches rely on greedy n
arXiv:2602.09642v2 Announce Type: replace-cross Abstract: Recent advances in Large Language Models (LLMs) have significantly improved table understanding tasks such as Table Question Answering (TableQ
arXiv:2604.19540v1 Announce Type: cross Abstract: Teams of LLM agents increasingly collaborate on tasks spanning days or weeks: multi-day data-generation sprints where generator, reviewer, and auditor
arXiv:2602.15173v2 Announce Type: replace Abstract: The use of large language models either as decision support systems, or in agentic workflows, is rapidly transforming the digital ecosystem. However
arXiv:2604.19148v1 Announce Type: new Abstract: Efficient and robust path planning hinges on combining all accessible information sources. In particular, the task of path planning for robotic environm
New sandbox! Announcing Cloud Run sandboxes: Secure on-the-fly code execution: Spin up ephemeral, isolated sandboxes from within Cloud Run resources. Safely execute agent-generated code, scripts, or C
Significant contributors to this article include Megan O'Keefe, Senior Staff Developer Advocate, and Karl Weinmeister, Director of Developer Relations. Whether you are joining us in person in Las Vega
This post likely clarifies the distinction between testing platforms and agent platforms, suggesting that a particular tool or framework (possibly LangChain, given Harrison Chase's role) is better cha
Octen, a startup with software that enables artificial intelligence agents to search the web, launched today with 10 million in seed funding. Square Peg led the investment. It was joined by Singapore-
arXiv:2505.11765v3 Announce Type: replace-cross Abstract: Agents powered by advanced large language models (LLMs) have demonstrated impressive capabilities across diverse complex applications. Recentl
OpenAI: OpenAI announces workspace agents in ChatGPT, letting teams create Codex-powered shared agents for complex tasks, and says they are “an evolution of GPTs” — Codex-powered agents for teams. — C
OpenAI is shutting down text-embedding-3-small?!? I strongly believe that if you shut down a closed-source embedding model that you should open-source. Imaging the trillions of tokens that will no lon
OpenAI is giving users of its Business, Enterprise, Edu, and Teachers plans access to cloud-based 'workspace' agents available in ChatGPT that can perform business tasks. In its blog post, OpenAI give
arXiv:2604.19379v1 Announce Type: new Abstract: This paper presents the first study on Unsupervised Domain Adaptation (UDA) for multimodal 3D panoptic segmentation (mm-3DPS), aiming to improve general
Pay attention to this one, AI devs. This is particularly interesting if you work with long-horizon terminal agents that often drown in their own observations. TACO is a self-evolving framework that au
arXiv:2403.09905v5 Announce Type: replace-cross Abstract: Embodied navigation methods commonly operate in static environments with stationary objects. In this work, we present approaches for tackling
arXiv:2604.19324v1 Announce Type: cross Abstract: We introduce PLaMo 2.1-VL, a lightweight Vision Language Model (VLM) for autonomous devices, available in 8B and 2B variants and designed for local an
arXiv:2604.19049v1 Announce Type: cross Abstract: LLM-assisted defect discovery has a precision crisis: plausible-but-wrong reports overwhelm maintainers and degrade credibility for real findings. We
arXiv:2505.06335v2 Announce Type: replace-cross Abstract: Federated Learning (FL) has the potential for simultaneous global learning amongst a large number of parallel agents, enabling emerging AI suc
arXiv:2505.20816v2 Announce Type: replace Abstract: Recent advances in multimodal question answering have primarily focused on combining heterogeneous modalities or fine-tuning multimodal large langua
arXiv:2604.19299v1 Announce Type: cross Abstract: Despite the impressive capabilities of large language models, their substantial computational costs, latency, and privacy risks hinder their widesprea
arXiv:2604.19523v1 Announce Type: new Abstract: Social deduction games such as Mafia present a unique AI challenge: players must reason under uncertainty, interpret incomplete and intentionally mislea
This article likely explores the integration of Model Context Protocol (MCP) with Pinecone's skills, CLI tools, and plugin ecosystem for building AI applications. It probably covers how developers can
Sony AI just published the first autonomous robot to beat elite humans at a competitive physical sport. Its name is Ace. The sport is table tennis. The paper dropped in Nature today. 9 cameras triangu
Will Dunham / Reuters: Sony AI says its autonomous ping pong robot is the first robot to attain expert-level performance in a physical sport after beating some top-level human players — An autonomous
This OpenAI article describes how to use WebSockets with the Responses API to improve performance and reduce latency in agentic AI workflows. WebSockets enable real-time, bidirectional communication b
arXiv:2604.19145v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have become central to autonomous driving systems, yet their deployment is severely bottlenecked by the massive computat
Step 3.5 Flash is now live for Nous Portal users, free for the next 10 days. If you're running Hermes Agent with Nous Portal as your provider, run 'hermes update' and then 'hermes model' to configure
arXiv:2604.18951v1 Announce Type: cross Abstract: Adaptive multi-agent systems (MAS) are increasingly adopted to tackle complex problems.However, the narrow task coverage of their optimization raises
arXiv:2604.19420v1 Announce Type: new Abstract: Maintaining long-term accuracy of stereo camera calibration parameters is important for autonomous systems' perception. This work proposes Online Tracki
arXiv:2602.06400v2 Announce Type: replace-cross Abstract: The prediction of 3D semantic occupancy enables autonomous vehicles (AVs) to perceive the fine-grained geometric and semantic scene structure
This post likely identifies and discusses 15 technology companies that are considered to have exceptional concentrations of talented employees, based on analysis or data from Paraform Talent. The list
The Never Ending Lore of Harness w/@Vtrivedy10⚡️ Viv brought some amazing perspective around harness design, evals, file systems, RL Envs and so much. 0:00:00 - INTRO 0:03:30 - PhD at Temple Universit
this looks like a website but it’s an interactive generative video 🫨 Imagine every pixel on your screen, streamed live directly from a model. No HTML, no layout engine, no code. Just exactly what you
arXiv:2604.19257v1 Announce Type: new Abstract: Creating realistic and simulation-ready 3D assets is crucial for autonomous driving research and virtual environment construction. However, existing 3D
Vibe coders are not going to like this. UC San Diego just published the first real field study of experienced developers using AI agents. They watched 13 of them code in the wild and surveyed 99 more.
arXiv:2604.19270v1 Announce Type: new Abstract: As groups of robots increasingly collaborate with humans, understanding how humans perceive them is critical for designing effective human-robot teams.
VideoGameBench is a new benchmark launched on Antim Labs, created by a1zhang, Thomas L. Griffiths, Karthik R. N, and Ofir Press. The benchmark likely evaluates AI model performance on video game-relat
arXiv:2604.17821v2 Announce Type: replace Abstract: Recent advancements in large language models (LLMs) have empowered autonomous web agents to execute natural language instructions directly on real-w
arXiv:2402.06922v5 Announce Type: replace-cross Abstract: Large language model (LLM)-based agents combine LLMs with external tools to automate tasks such as scheduling meetings, managing documents, or
Yann LeCun (AMI Labs Founder): 'The AI industry is completely LLM-pilled. Everybody is working on the same thing. They're all digging the same trench.' LeCun explains why no lab dares break from the p
arXiv:2604.16333v1 Announce Type: new Abstract: Knee osteoarthritis frequently exhibits discordance between structural damage observed in imaging and patient-reported symptoms such as pain. This misma
arXiv:2510.05608v2 Announce Type: replace Abstract: Agents based on large language models (LLMs) struggle with brainless trial-and-error and generating hallucinatory actions due to a lack of global pl
arXiv:2604.17225v1 Announce Type: new Abstract: We present a novel approach for claim verification from tabular data documents. Recent LLM-based approaches either employ complex pretraining/fine-tunin
arXiv:2604.17258v1 Announce Type: new Abstract: Deploying a humanoid robot to manipulate a new object has traditionally required one to two days of effort: data collection, manual annotation, 3D model
arXiv:2604.16548v1 Announce Type: cross Abstract: Research on large language model (LLM) security is shifting from 'will the model leak training data' to a more consequential question: can an agent wi
Access GPT Image 2.0 natively in Hermes Agent Update now to get access - just run `hermes update` and select your image generation tool model with `hermes tools` Introducing ChatGPT Images 2.0 A state
As open-source platforms scramble to close the infrastructure gap, enterprises are discovering that autonomous operations expose a deeper problem than performance: control. The conversation around ent
arXiv:2604.16625v1 Announce Type: new Abstract: Recent large language model (LLM) agents have shown promise in using execution feedback for test-time adaptation. However, robust self-improvement remai
arXiv:2508.18025v4 Announce Type: replace-cross Abstract: Autonomous planetary exploration demands real-time, high-fidelity environmental perception. Standard deep learning models require massive comp
arXiv:2503.04798v3 Announce Type: replace Abstract: We present Scalable Multi-Agent Realistic Testbed (SMART), a realistic and efficient software tool for evaluating Multi-Agent Path Finding (MAPF) al
Replit's Agent feature can automatically scan for security issues and publish the findings, allowing users to resolve all detected problems with a single 'Fix all with Agent' action. This functionalit
arXiv:2604.18292v1 Announce Type: cross Abstract: Large language models are increasingly expected to serve as general-purpose agents that interact with external, stateful tool environments. The Model
arXiv:2604.17609v1 Announce Type: new Abstract: LLM-based agents are assumed to integrate environmental observations into their reasoning: discovering highly relevant but unexpected information should
Amplitude Inc. today introduced an embedded artificial intelligence support agent that helps users navigate digital products without leaving the application. Amplitude AI Assistant uses behavioral dat
Vector Space Day 2026: Powered by Qdrant About We’re hosting our second-ever full-day in-person Vector Space Day (https://luma.com/vsd-sf) on June 11th at The Midway in San Francisco, and you’re invit