State-Centric Decision Process
arXiv:2605.12755v1 Announce Type: new Abstract: Language environments such as web browsers, code terminals, and interactive simulations emit raw text rather than states, and provide none of the runtim
Knowledge catalogue
arXiv:2605.12755v1 Announce Type: new Abstract: Language environments such as web browsers, code terminals, and interactive simulations emit raw text rather than states, and provide none of the runtim
Open source AI trust has become a central concern for enterprises moving agentic AI into production, where governance, security and reliability matter as much as model performance. That pressure is la
Try this early Grok Build (anything) beta and let us know what to improve. Much appreciated! An early beta of Grok Build, an agentic CLI for coding, building apps, and automating workflows is now avai
As AI adoption accelerates, organizations are shifting their focus from experimentation to large-scale deployment. The challenge now is building secure and scalable systems that can support AI agents,
Enterprise AI is moving faster than the digital trust systems built to govern it. As autonomous agents, synthetic content and machine identities spread across critical systems, organizations face a ne
arXiv:2605.12178v1 Announce Type: cross Abstract: World models enable agents to anticipate the effects of their actions by internalizing environment dynamics. In enterprise systems, however, these dyn
literally what i have been saying for years, once again. Yann LeCun says you cannot build a reliable agentic system without a world model LLMs don't have world models. They can't predict the consequen
Off to an amazing start at @LangChain Interrupt 🪁🦜 Including: 📣 Shoutout from @hwchase17 during the opening keynote! 🪁 CopilotKit booth with the best staff, demos and merch 🍸 Agents After Dark Happy H
arXiv:2605.11975v1 Announce Type: new Abstract: We study stochastic minimum-cost reach-avoid reinforcement learning, where an agent must satisfy a reach-avoid specification with probability at least p
Super excited about the new deepagents version from @LangChain What’s in their water supply?? Not to mention 7 blog posts dropped today, including a new Deep Agents version with significant improvemen
This looks great - kudos @LangChain team 🚀Launching: LangSmith Engine LangSmith Engine is an agent that sits on top of your traces It runs in the background and automatically identifies issues It then
Veeam Software Group GmbH used VeeamON 2026 in New York City this week to punctuate its shift from “the backup company” to a data and artificial intelligence trust platform for the agentic era. With a
arXiv:2605.10906v1 Announce Type: cross Abstract: As model families, training recipes, and compute budgets become increasingly standardized, further gains in machine learning systems depend increasing
arXiv:2601.22449v2 Announce Type: replace Abstract: Intrinsic Motivation (IM) aims to train agents without external rewards, enabling useful behavior to emerge from the agent's interaction with its en
arXiv:2605.09192v1 Announce Type: new Abstract: Agent skills can remarkably improve task success rates by using human-written procedural documents, but their quality is difficult to assess without env
GLM models are now live on @tensorix_ai We’re partnering to bring cost-efficient frontier AI models to developers, startups, and enterprises across Europe and beyond — and to back the Sovereign AI eco
arXiv:2605.10917v1 Announce Type: new Abstract: We consider anonymous multi-agent path finding (MAPF) where a set of robots is tasked to travel to a set of targets on a finite, connected graph. We sho
arXiv:2605.08581v1 Announce Type: new Abstract: Modern online large language model (LLM) services, such as Retrieval-Augmented Generation (RAG) and agent systems, increasingly expose two prominent cha
arXiv:2605.06524v2 Announce Type: replace Abstract: Reliable human-machine discrimination is becoming increasingly important as large language models and autonomous agents are deployed in online setti
SAP SE today introduced at Sapphire 2026, the company’s annual conference, what it calls Autonomous Enterprise, a suite of artificial intelligence tools and agents designed to enhance how humans and A
arXiv:2605.10500v1 Announce Type: new Abstract: Agent skills today are static artifact: authored once -- by human curation or one-shot generation from parametric knowledge -- and then consumed unchang
arXiv:2605.10828v1 Announce Type: new Abstract: As large language models are increasingly deployed in retrieval-augmented generation and agentic systems that accumulate extensive context, understandin
arXiv:2605.08323v1 Announce Type: cross Abstract: Communication is fundamental to sustaining reciprocity and cooperation in strategic interactions. We identify and formulate the influence attribution
thoughts after doing a bunch of synthetic data gen for eval + environment building - LLMs are incredible projections of the world bundled into a set of weights - but doing targeted extraction of certa
arXiv:2507.09788v2 Announce Type: replace-cross Abstract: Recent advances in Large Language Models (LLM) have led to a new class of autonomous agents, renewing and expanding interest in the area. LLM-
Enterprise AI adoption has crossed a threshold: The question is no longer whether to invest, but how to do it wisely. As agentic workloads multiply and inference costs rise, AI choice — the ability to
arXiv:2605.06702v1 Announce Type: new Abstract: Large language models (LLMs) have become a central foundation of modern artificial intelligence, yet their lifecycle remains constrained by a rigid sepa
I have a new job! Excited to announce that I will be working with Hugging Face to make local models work great in OpenClaw and other open agent harnesses! I will be building in public and documenting
This seems like a critical reason to open up about AI use in academia. Scholars are using old AI models, badly, and not talking about it. New models hallucinate very few citations, and good agentic ha
@hwchase17 the tooling side is maturing fast but the feedback loop between engineers and domain experts is still the bottleneck. keep running into this building agents: you can instrument everything a
Ready to finally make the switch? Here is a $20 off Nous Portal code for the first 200 new users, to celebrate our new #1 spot. http://portal.nousresearch.com Hermes Agent is now #1 on the Global @Ope
LLM Wikis + HTML Artifacts are insanely powerful. You should seriously consider this in your workflows. LLM Wikis captures all the important information that lets you and your agents do meaningful wor
arXiv:2511.20284v2 Announce Type: replace-cross Abstract: Precise access control decisions are crucial for the security of both traditional applications and emerging agent-based systems. Typically, th
Nous Research released version 2026.5.7 of the Hermes Agent, with details available in the full release notes on their GitHub repository. Users can update to this version by running the 'hermes update
got two of the new Reachy Minis and they are super fun. A quick build, and a nice community of folks adding features every week Nice work @ClementDelangue and team! We're launching the agentic robotic
New Max Agency with @tryramp Talked to @ramplabs Head of Applied Research Alex Shevchenko on the Max Agency podcast to learn how @Ramp Sheets was built, their internal agent Inspect, and so much more.
arXiv:2506.07548v2 Announce Type: replace-cross Abstract: Multi-agent reinforcement learning (MARL) has reached competitive performance on cooperative tasks against scripted adversaries, yet most meth
Talked to @ramplabs Head of Applied Research Alex Shevchenko on the Max Agency podcast to learn how @Ramp Sheets was built, their internal agent Inspect, and so much more. YouTube: https://www.youtube
Try Grok Voice for your customer support Your customer support needs a voice agent built for the real world. Grok Voice Think Fast 1.0 handles complex workflows with speed and accuracy, even in hard-t
arXiv:2605.02346v1 Announce Type: cross Abstract: Bare-metal operational technology (OT) devices -- especially the microcontrollers running Modbus/TCP and CoAP at the base of industrial control system
arXiv:2508.14976v2 Announce Type: replace Abstract: We present Aura-CAPTCHA, a multi-modal verification system that integrates Generative Adversarial Networks (GANs), Reinforcement Learning (RL), and
arXiv:2605.03855v1 Announce Type: new Abstract: Human-AI collaboration requires AI agents to understand human behavior for effective coordination. While advances in foundation models show promising ca
arXiv:2605.00831v1 Announce Type: cross Abstract: The rise of million-token, agent-based applications has placed unprecedented demands on large language model (LLM) inference services. The long-runnin
arXiv:2605.01423v1 Announce Type: cross Abstract: The escalating data scale in High-Energy Physics (HEP) fuels a growing aspiration for higher analytical efficiency. While Large Language Models (LLMs)
https://x.com/walden_yan/status/2052070983083942322?s=20 Cool to see failure modes of different coding agents in new report from @greptile - seems like Devin is better than humans in almost all catego
arXiv:2605.01694v1 Announce Type: new Abstract: A world model matters to an agent only through the state it constructs. That state must preserve some information, discard other information, and suppor
arXiv:2605.03288v1 Announce Type: new Abstract: Many physical AI tasks are governed by implicit equilibrium: an agent actuates a subset of degrees of freedom (boundary DoFs), while the remaining free
Welcome to The Blueprint, a new feature where we highlight how Google Cloud customers are tackling unique and common challenges across industries using the latest AI and cloud technologies. We hope to
arXiv:2605.02630v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have enabled autonomous GUI agents that translate natural language instructions into executable screen coordinates. Howeve
arXiv:2510.02945v3 Announce Type: replace Abstract: Continual reinforcement learning (continual RL) seeks to formalize the notions of lifelong learning and endless adaptation in RL. In particular, the
Grok 4.3 Grok 4.3 is now live on the xAI API. It’s our fastest, most intelligent model to date. It tops the @ArtificialAnlys leaderboards in agentic tool calling and instruction following, and ranks #
May the 4th be with you!✨ Celebrate with us, you must. Join our Start Up Party Up this Thursday before SF's free 2nd St Fest 🎵 Calling all #AI Jedis - leave your agents at home and sign up for our nex
arXiv:2605.01248v1 Announce Type: new Abstract: Reinforcement learning (RL) post-training has enabled newer capabilities in models, such as agentic tool-use for search. However, these models struggle
arXiv:2605.01704v1 Announce Type: new Abstract: When copies of the same language model are prompted to debate, they produce diverse phrasings of one perspective rather than diverse perspectives. Multi
arXiv:2605.00737v1 Announce Type: new Abstract: Agentic AI architectures augment LLMs with external tools, unlocking strong capabilities. However, tool use is not always beneficial; some calls may be
'Traces everywhere. Feedback loop? Nowhere' @hwchase17 This is the part that gets skipped. Traces everywhere. Feedback loop? Nowhere. Building agents for clients and the hardest sell is 'we need to lo
The agentic enterprise is here — but it may be too soon for companies trying to protect their data from rogue artificial intelligence. Boomi LP, which specializes in integration platform as a service,
Excited to partner with @pinecone! Introducing Pinecone Nexus. A knowledge engine for agents. The bottleneck for production agents isn't the model. It's the per-query work of searching, stitching, par
Most AI today is passive. You prompt it, it responds, you prompt again. @origoshen , Co-CEO of @AI21Labs, draws a clear line between that reality and the agentic AI everyone's picturing. The direction
arXiv:2605.00264v1 Announce Type: new Abstract: Offline multi-agent reinforcement learning in general-sum settings is challenged by the distribution shift between logged datasets and target equilibriu