so hes building a harness
so hes building a harness Now launching GBrain v0.11 with Minions I got sick of OpenClaw's subagents timing out and not getting things done So I built a queue/jobs system that uses GBrain's Postgres/P
Knowledge catalogue
so hes building a harness Now launching GBrain v0.11 with Minions I got sick of OpenClaw's subagents timing out and not getting things done So I built a queue/jobs system that uses GBrain's Postgres/P
arXiv:2604.14882v1 Announce Type: cross Abstract: Rapid urbanization and continuous population growth have made municipal solid waste management increasingly challenging. These challenges highlight th
arXiv:2604.13109v1 Announce Type: cross Abstract: We present a two-stage pipeline for AI-assisted improvement of published algorithm implementations. In the first stage, a large language model with re
arXiv:2604.14398v1 Announce Type: cross Abstract: Rotating detonation engines (RDEs) are a promising propulsion concept that may offer higher thermodynamic efficiency and specific impulse than convent
arXiv:2604.13519v1 Announce Type: new Abstract: Tool calling has greatly expanded the practical utility of large language models (LLMs) by enabling them to interact with external applications. As LLM
A Reddit post from the r/ollama community titled 'Two AIs walk into a bar...' likely showcases a humorous or experimental multi-agent conversation in which two locally-run AI models (via Ollama) inter
AGI Inc. is promoting upcoming events and inviting interested individuals to RSVP through a Luma event page to secure invitations. Luma is a platform commonly used for organizing and managing event re
arXiv:2604.11954v1 Announce Type: cross Abstract: We study dynamic multi-robot task allocation under uncertain task completion, time-window constraints, and incomplete information. Tasks arrive online
arXiv:2604.12167v1 Announce Type: new Abstract: We present (Experience-Modulated Biologically-inspired Emergent Reasoning), a hybrid cognitive architecture that reorganises the relationship between la
arXiv:2604.12374v1 Announce Type: cross Abstract: We describe the pre-training, post-training, and quantization of Nemotron 3 Super, a 120 billion (active 12 billion) parameter hybrid Mamba-Attention
Simon Last and Sarah Sachs from Notion discuss their extensive experimentation with AI-native product development, including rebuilding core Notion features five times as capabilities evolved and inte
This is awesome and makes things much more composable. Love it deepagents now supports structured output for subagents! an under appreciated piece of context engineering is figuring out what context i
arXiv:2604.10981v1 Announce Type: new Abstract: ATANT v1.0 (arXiv:2604.06710) defined continuity as a system property with 7 required properties and introduced a 10-checkpoint, LLM-free evaluation met
arXiv:2604.11628v1 Announce Type: new Abstract: Existing conversational memory systems rely on complex hierarchical summarization or reinforcement learning to manage long-term dialogue history, yet re
Nous Research shared a post on X (formerly Twitter) discussing 'Creative Direct with Hermes,' likely showcasing the creative writing or instruction-following capabilities of their Hermes language mode
Harness Engineering Derived from what Models can’t do alone: It feels like a good time to step back and reshare some basic mental models for why harnesses exist in the first place - working backwards
arXiv:2509.26306v4 Announce Type: replace Abstract: Existing multi-agent learning approaches have developed interactive training environments to explicitly promote collaboration among multiple Large L
arXiv:2604.10534v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) is a new and emerging technology that extends the functionality of large language models, improving workflows but als
arXiv:2604.09567v1 Announce Type: cross Abstract: Knowledge representation formalisms are aimed to represent general conceptual information and are typically used in the construction of the knowledge
The agentic era demands a new security era. Our mission is to be the world’s most trusted security partner in this new era, helping every organization accelerate their security transformation with the
arXiv:2604.05529v2 Announce Type: replace Abstract: Human mobility modeling is indispensable for diverse urban applications. However, existing data-driven methods often suffer from data scarcity, limi
arXiv:2604.08750v1 Announce Type: new Abstract: Plant-level control is an emerging wind energy technology that presents opportunities and challenges. By controlling turbines in a coordinated manner vi
arXiv:2604.08920v1 Announce Type: cross Abstract: Information retrieval systems have traditionally optimized for topical relevance-the degree to which retrieved documents match a query. However, relev
arXiv:2604.08805v1 Announce Type: cross Abstract: In November 2025, the authors ran a workshop on the topic of what makes a good reinforcement learning (RL) environment for autonomous cyber defence (A
arXiv:2510.09093v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are now routinely used to autonomously execute complex tasks, from natural language processing to dynamic workflo
'how can i help?' i analyzed 100 asks from portfolio companies identified in notes and emails to see what they asked 26% specific person intro 9% investor discovery 7% hiring/talent 9% pr/media 11% bu
arXiv:2604.09338v1 Announce Type: new Abstract: Spatial reasoning is central to navigation and robotics, yet measuring model capabilities on these tasks remains difficult. Existing benchmarks evaluate
The TL;DR is that Google engineering appears to have the same AI adoption footprint as John Deere, the tractor company. Most of the industry has the same internal adoption curve: 20% agentic power use
This post completely misses the point of @sarahwooders 's original article. The whole point is that memory == context engineering, so it can't be a 'thin conductor'. Saying 'memory is just markdown' t
Directionally correct Open memory standards will need to emerge But it’s so early right now. We’re still just figuring out what best practices are. And so a lot is at the mercy of harnesses Agents.md
Harness, Memory, Context Fragments, & the Bitter Lesson this is a work in progress mental dump on interesting intersections between how we use and design a harness, implications for memory being accum
Harrison Chase, co-founder and CEO of LangChain, made a post on X affirming a point about where value lies in AI systems, specifically arguing that value is found 'in the harness' — referring to the s
Nous Research has developed a skill optimized for processing and analyzing AI research papers, designed to work within their Hermes agent framework. The skill can be adapted and repurposed for researc
very memgpt / sarah wooders coded. memory isn’t a layer, it is the system. most teams think they’re choosing a model, but they’re really choosing where their memory lives and like ben thompson says, o
Memory + harness, or as @matt_slotnick put it, memory + intent. Something I describe a lot about @usebrief, that *why* you remember something is more important than pure recall. A moment that’s trivia
LangChain's concept of 'your harness = your memory' explores how the surrounding infrastructure and framework around an AI agent effectively functions as its memory system. The post likely discusses h
arXiv:2604.08443v1 Announce Type: new Abstract: The potential of Animal-Robot Interaction (ARI) in welfare applications depends on how much an animal perceives a robotic agent as socially relevant, no
arXiv:2604.05351v3 Announce Type: replace-cross Abstract: Image Goal Navigation (ImageNav) is evaluated by a coarse success criterion, the agent must stop within 1m of the target, which is sufficient
The Chat SDK is a unified TypeScript library designed for building chat bots. This tool allows developers to maintain a single codebase to create chatbots that reliably function across multiple pla...
I was unable to retrieve the specific content of the linked tweet, as X (formerly Twitter) requires JavaScript and login to display individual posts. The search results did surface a highly relevan...
Open swe uses deepagents under the hood Deepagents is general purpose, openswe is focused on coding @hwchase17 @LangChain This is a very interesting comparison. Now my question is: how to compare Deep
Cloud Run has long provided developers with a straightforward, opinionated platform for running code. You can easily deploy request-driven web applications using Cloud Run services, or execute run-to-
https://x.com/ashpreetbedi/status/2041568919085854847?s=46 great related write up from a team that pushes taking a rigorous systems engineering first approach to building agents it’s never one prompt,
I love these almost as much as i love using langsmith LangSmith 🤝Fix your agents You'll see our billboards around SF and NYC over the next few months. The themes all point to the same problem: you don
Introducing GLM-5.1 from @Zai_org on Together AI. AI natives can now use GLM-5.1 on Together and benefit from reliable inference for production-scale agentic engineering and long-horizon coding workfl
arXiv:2608.12771v1 Announce Type: cross Abstract: The extent to which large language models for code rely on memorization over genuine understanding remains highly debated. While current literature fr
arXiv:2608.13558v1 Announce Type: new Abstract: Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and cod
arXiv:2602.20017v2 Announce Type: replace Abstract: Real-world tables often contain schema inconsistencies, heterogeneous value formats, and implicit relational structures that degrade table reasoning
arXiv:2608.13331v1 Announce Type: cross Abstract: The replicability of papers is a cornerstone of scientific knowledge, ensuring the reliability of existing results and providing a base for further ex
arXiv:2608.12104v1 Announce Type: cross Abstract: The increasing deployment of autonomous, agentic AI systems challenges traditional accountability mechanisms. Existing research predominantly frames A
arXiv:2608.11407v1 Announce Type: new Abstract: Robust traffic simulators are crucial for developing and testing autonomous vehicles to reduce the costly, labor-intensive real-world data collection pr
Very interesting new paper from Microsoft and colleagues. (bookmark it) Skill libraries are used in every major harness on the assumption that more guidance is free. This work measures what a bad skil
arXiv:2608.10526v1 Announce Type: cross Abstract: Motivated by decentralized applications, we study cooperative multi-agent bandits in continuous (Lipschitz) action spaces when the Lipschitz constant
Four small architecture decisions can cost up to 47% of a model's long-context performance. New research from Ai2, Carnegie Mellon, and the University of Washington isolates them. Normalization, GQA,
arXiv:2608.09248v1 Announce Type: new Abstract: Skill-based LLM agents select reusable procedures from an external library to solve complex tasks, yet their routing decisions rely entirely on text-lev
Nemotron 3.5 Lightning is available in LM Studio! The model is 30B MoE (3B active), can run very fast, and is trained for high volume agentic use cases. Model page: https://lmstudio.ai/models/nvidia/n
arXiv:2608.08795v1 Announce Type: cross Abstract: Tool-using large language model (LLM) agents are vulnerable to indirect prompt injection (IPI), in which malicious instructions embedded in external o
arXiv:2608.06706v1 Announce Type: cross Abstract: Latent world models plan by predicting future states from an action, but when a scene contains motion the agent does not control, they quietly go acti
Every crawl, retried job, and embedding pipeline change writes points into a vector collection. The stored data keeps moving even when the query code never changes, and the top results move with it. A
The successor to Llama is here, and Meta is revitalizing focus on open weights with their new Muse Glimmer - a leading 30B param model designed for always-on local agent use, small enough to run on a