Introducing Toolboxes in Foundry
Available in Public Preview Today Toolbox is a new way to curate, configure, and reuse tools across all of your AI agents without rewiring them every time from Foundry. Today, teams build agents acros
Knowledge catalogue
Available in Public Preview Today Toolbox is a new way to curate, configure, and reuse tools across all of your AI agents without rewiring them every time from Foundry. Today, teams build agents acros
arXiv:2604.19049v1 Announce Type: cross Abstract: LLM-assisted defect discovery has a precision crisis: plausible-but-wrong reports overwhelm maintainers and degrade credibility for real findings. We
arXiv:2604.18951v1 Announce Type: cross Abstract: Adaptive multi-agent systems (MAS) are increasingly adopted to tackle complex problems.However, the narrow task coverage of their optimization raises
arXiv:2604.17821v2 Announce Type: replace Abstract: Recent advancements in large language models (LLMs) have empowered autonomous web agents to execute natural language instructions directly on real-w
arXiv:2604.17487v1 Announce Type: new Abstract: Agentic systems often fail not by being entirely wrong, but by being too precise: a response may be generally useful while particular claims exceed what
arXiv:2604.17488v1 Announce Type: new Abstract: Manual annotation of high-quality visual question answering with grounding (VQA-G) datasets, which pair visual questions with evidential grounding, is c
HF becoming the platform for agents (assisted by their humans) to use and build AI (rather than just leveraging APIs)! Introducing ml-intern, the agent that just automated the post-training team @hugg
Snowflake Inc. is expanding its push into enterprise artificial intelligence with a set of updates to its Snowflake Intelligence and Cortex Code offerings, positioning its platform as a centralized co
// Survey on Multi-Agent Systems // The paper traces the landscape from classical paradigms (consensus, distributed control, swarm intelligence, cooperative learning) to foundation-model-enabled MAS (
Whistant is an on-phone AI buddy that helps users get things done directly from their iPhone without requiring a Mac. It breaks requests into step-by-step subtasks and executes them on the device, inc
arXiv:2604.16004v1 Announce Type: cross Abstract: Verifiers have been demonstrated to enhance LLM reasoning via test-time scaling (TTS). Yet, they face significant challenges in complex domains. Error
arXiv:2604.16024v1 Announce Type: cross Abstract: Vision Language Models (VLMs) have been applied to several specific domains and have shown strong problem-solving capabilities. However, astronomical
arXiv:2604.15319v1 Announce Type: cross Abstract: Exploratory analysis of high-dimensional data relies on embedding the data into a low-dimensional space (typically 2D or 3D), based on which visualiza
'The important thing is not to stop questioning.' — Einstein, 71 years on. Today we celebrate him with II-Agent. His genius was reasoning from first principles, a refusal to take any foundation on fai
arXiv:2509.06477v2 Announce Type: replace Abstract: Shortcuts such as APIs and deep-links have emerged as efficient complements to flexible GUI operations, fostering a promising hybrid paradigm for ML
arXiv:2604.13242v1 Announce Type: cross Abstract: Large language models (LLMs), particularly when integrated into agentic systems, have demonstrated human- and even superhuman-level performance across
Rizwan Ahmed won Week 4 of the Agent 4 Content Challenge hosted by Replit with a project called SlapStop, a group accountability app that uses a humorous 'slapping' mechanic among friends to encourage
yeah, what they said 🤝 a decent harness gets you an actually functioning agent now + teams that actually invest time in their harness+problem design, choosing good infra, self-improvement loops, data
arXiv:2604.13349v1 Announce Type: new Abstract: Communication in Large Language Model (LLM)-based multi-agent systems is moving beyond discrete tokens to preserve richer context. Recent work such as L
Cloudflare's Browser Run is a service that provides AI agents with the ability to control and interact with a real web browser, enabling them to perform tasks such as web scraping, form submission, na
arXiv:2502.11271v2 Announce Type: replace-cross Abstract: Solving complex reasoning tasks may involve visual understanding, domain knowledge retrieval, numerical calculation, and multi-step reasoning.
User scoped memory is one of those things that doesn’t matter if you’re building a toy agent for yourself, but when you release at scale you gotta get it right Deepagents deploy helps you do that, eas
arXiv:2604.10290v1 Announce Type: new Abstract: AI is increasingly deployed in multi-agent systems; however, most research considers only the behavior of individual models. We experimentally show that
Harrison Chase's post explores the divergent strategic incentives around memory in AI agent systems, arguing that for those building agent frameworks, persistent memory creates compounding value and c
arXiv:2604.09579v1 Announce Type: new Abstract: In large-scale cloud service platforms, thousands of customer tickets are generated daily and are typically handled through on-call dialogues. This high
Open Harness 🤝 Deployed Agents if you wanna use Claude, GLM5, and Codex in your deployed harness then you should be able to! deepagents deploy has easy configs to let users customize their harness and
arXiv:2604.10286v1 Announce Type: new Abstract: Autonomous language-model agents increasingly rely on installable skills and tools to complete user tasks. Static skill auditing can expose capability s
arXiv:2604.11716v1 Announce Type: new Abstract: Prior representative ReAct-style approaches in autonomous Software Engineering (SWE) typically lack the explicit System-2 reasoning required for deep an
Cursor AI shared observations on X validating that multi-agent architectures demonstrate superior performance when tackling novel problems that fall outside the distribution of training data. The post
arXiv:2603.05744v2 Announce Type: replace Abstract: Current AI-powered code assistance tools often struggle with poorly-defined problem statements that lack sufficient task context and requirements sp
arXiv:2506.04676v2 Announce Type: replace-cross Abstract: The data scarcity, label noise, and long-tailed category imbalance remain important and unresolved challenges in many computer vision tasks, s
🔒 new in deepagents: filesystem permissions shared resources and org-wide policies are exactly the kind of files you want your agent to read but never overwrite. filesystem permissions let you enforce
arXiv:2604.08621v1 Announce Type: new Abstract: In consumer applications, Customer Relationship Management (CRM) has traditionally relied on the manual optimisation of static, rule-based messaging str
arXiv:2604.08728v1 Announce Type: new Abstract: Cooperation in multi-agent reinforcement learning (MARL) benefits from inter-agent communication, yet most approaches assume idealized channels and exis
Harrison Chase, co-founder of LangChain, endorsed the perspective that memory in AI agents is fundamentally a form of context, emphasizing that effective context management is the core requirement for
Nous Research has introduced a new feature in their Hermes Agent system that allows users to use the `/compress ` command to influence how the compaction model prioritizes and retains information. Thi
Harrison Chase, co-founder of LangChain, discusses a trend in the AI industry where agent harnesses are becoming increasingly closed and proprietary, with memory systems locked behind vendor-specific
This 1-min clip from the creator of LangChain + this 10-min read from him will teach you more about what actually matters in AI agents than everything you've scrolled past this year. Watch this. Then
Harrison Chase, the creator of LangChain, shared an article discussing agent harness design and memory architecture for AI systems. The post highlights approaches to structuring how agents manage and
给中国用户的好消息:Hermes Agent 现在原生支持个人微信了 微信扫码即可连接,私聊群聊都支持。图片、视频、文件、语音消息全覆盖,长轮询直连,不需要公网 IP。 运行 'hermes update' 即可体验 文档:http://hermes-agent.nousresearch.com/docs/user-guide/messaging/weixin 感谢 @Bravohenry_ 的贡
time to help builders own their intelligence, open everything 🤝 the data produced and Experiential Memory gained from every agent interaction is the one of the most valuable things you can own to impr
arXiv:2604.07775v1 Announce Type: cross Abstract: Collaboration and information sharing empower Multi-Agent Systems (MAS) but also introduce a critical security risk known as Agent Cascading Injection
arXiv:2604.07007v1 Announce Type: cross Abstract: Autonomous AI agents are beginning to operate across organizational boundaries on the open internet -- discovering, transacting with, and delegating t
arXiv:2604.06284v1 Announce Type: cross Abstract: Autonomous AI agents powered by Large Language Models can reason, plan, and execute complex tasks, but their ability to autonomously retrieve informat
arXiv:2604.08495v1 Announce Type: cross Abstract: This paper addresses the decentralized non-uniform area coverage problem for multi-agent systems, a critical task in missions with high spatial priori
I was unable to retrieve live content from the X (Twitter) broadcast URL or find sufficient search results about this specific Replit Agent 4 Buildathon broadcast. X (Twitter) Spaces/broadcast link...
If you're not in our Discord yet, come join! It's very active and a fantastic place to learn about Hermes Agent and AI, find new resources, and share what you're working on. http://discord.gg/nousrese
arXiv:2604.06629v1 Announce Type: cross Abstract: We present Logical Robots, an interactive multi-agent simulation platform where autonomous robot behavior is specified declaratively in the logic prog
arXiv:2604.07036v1 Announce Type: cross Abstract: Recently, LLM-based agents have become increasingly popular across many applications, including complex sequential decision-making problems. However,
arXiv:2502.19559v3 Announce Type: replace Abstract: Multi-agent debate - multiple instances of large language models discussing problems in turn-based interaction - has shown promise for solving knowl
Week 3 winner of the Agent 4 Content Challenge: Magomed Kurbaitaev 🎉 Built: https://help-dagestan.replit.app/ A real-time disaster coordination platform that connects flood victims with verified local
ALTK-Evolve, published by IBM Research, is a memory system for AI agents that enables on-the-job learning by helping agents improve over time, learning from and using guidelines generated from pre...
arXiv:2608.11248v1 Announce Type: new Abstract: Long-term memory is essential for language agents operating across extended interactions and evolving tasks. Existing memory-augmented agents mainly foc
arXiv:2608.09939v1 Announce Type: cross Abstract: Production teams deploying LLM chat agents face a specific quality assurance gap: existing evaluation tools test individual responses or simulate soci
arXiv:2608.09988v1 Announce Type: cross Abstract: Large language models are increasingly used to read markets, assess risk, and allocate capital. However, reported results for LLM trading agents can b
Ran the model with quants (Q5) and MTP by bartowski with llama.cpp server. It takes ~24GB ram running on M5 Pro with 48GB at about 65t/s. On some tasks it was quite the overthinker. Overall, the quali
arXiv:2608.10689v1 Announce Type: cross Abstract: Terminal interfaces to conversational agents report rich internal state (listening, thinking, executing tools, awaiting input, failing) almost entirel
We built a software factory for AI SDK. Each step is an agent, and humans merge changes. Four weeks in: ▪️ The factory authors up to 35% of merged PRs ▪️ It closed 70% of issues in July ▪️ Open bugs a
Public benchmarks can show how a model performs in general. Production reliability depends on the context and harness around it, which only your team can evaluate against its own data, workflows, and
arXiv:2608.09168v1 Announce Type: new Abstract: Agent skills are increasingly used to equip large language model (LLM) agents with reusable procedural knowledge. Although recent work has substantially