AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,964 results
11 Jun 2026

Learning Ordinal Response Policies in Rank-Based Stochastic Prize-Collecting Games

Model ReleasesDGX agent

arXiv:2510.24515v2 Announce Type: replace Abstract: The Team Orienteering Problem (TOP) generalizes many real-world multi-agent scheduling and routing tasks that occur in autonomous mobility, aerial l

10 Jun 2026

paper #1 for context: https://x.com/yoheinakajima/status/2057812713045377055?s=20

AgentsDGX agent

paper #1 for context: https://x.com/yoheinakajima/status/2057812713045377055?s=20 babyagi has ~200 citations, but 0 papers... i just published my first paper on arXiv 😆 'The Log is the Agent: Event-So

9 Jun 2026

(Auto)formalization is supposed to be easy: Trellis process semantics for spelling out rigorous proofs

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
AgentsDGX agent

arXiv:2606.09674v1 Announce Type: new Abstract: We present Trellis: an autoformalization system that leverages LLM agents in a deterministically constrained workflow to enforce incremental progress in

FASE: Fast Adaptive Semantic Entropy for Code Quality

AgentsDGX agent

arXiv:2606.09800v1 Announce Type: cross Abstract: Multi-agent code generation offers a promising paradigm for autonomous software development by simulating the human software engineering lifecycle. Ho

RepoLaunch: Automating Build and Management of Code Repositories across Languages and Platforms

AgentsDGX agent

arXiv:2603.05026v2 Announce Type: replace-cross Abstract: Language model (LM) agents have driven substantial progress in automated software engineering (SWE), yet building and testing software reposit

To Nuke or Not to Nuke: LLMs' (Missing) Ethical Reasoning and Actions in a High-Stakes Decision-Making Simulation

AgentsDGX agent

arXiv:2606.08310v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as long-horizon agents with decision-making capacities. While LLMs can show ethical competence on

What are loops, and how do you build one? A 'loop' is the repeated process where some event or input kicks off an action. For example: 1. CI…

AgentsDGX agent

What are loops, and how do you build one? A 'loop' is the repeated process where some event or input kicks off an action. For example: 1. CI fails -> you fix it 2. CI fails again -> you fix it 3. CI p

8 Jun 2026

IRAF: Interference-Resilient Adaptive Fusion for Noise-Robust End-to-End Full-Duplex Spoken Dialogue Systems

AgentsDGX agent

arXiv:2606.06559v1 Announce Type: cross Abstract: Full-duplex spoken dialogue models allow voice agents to listen and speak concurrently, enabling natural interaction with real-time overlap. However,

6 Jun 2026

A Finite Certificate for the Positive n=9 Vasc Inequality

AgentsDGX agent

arXiv:2606.06136v1 Announce Type: cross Abstract: We prove the positive-real n=9 case of the Vasc cyclic inequality. The proof was obtained with human-guided assistance from the AI agent MechMath Agen

5 Jun 2026

The Self-Correction Illusion: LLMs Correct Others but Not Themselves

AgentsDGX agent

arXiv:2606.05976v1 Announce Type: cross Abstract: Recent work shows that LLM agents struggle to correct errors in their own reasoning traces yet show markedly higher correction rates when identical cl

4 Jun 2026

Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation

AgentsDGX agent

arXiv:2510.13272v3 Announce Type: replace Abstract: Inspired by the success of reinforcement learning (RL) in Large Language Model (LLM) training for domains like math and code, recent work has begun

Trivium: Temporal Regret as a First-Class Objective for Causal-Memory Controllers

AgentsDGX agent

arXiv:2606.04421v1 Announce Type: new Abstract: Many current agentic systems and LLM pipelines correct mistakes by optimizing outcome reward. This addresses only the what of failure: when an outcome d

3 Jun 2026

CARVE: Certified Affordable Repair of Vetoed Maneuvers via Envelopes for Interactive Driving

AgentsDGX agent

arXiv:2606.02641v1 Announce Type: cross Abstract: Interactive driving exposes a failure mode that is easy to miss in rule-aware autonomous-driving stacks: a hard-rule margin can be negative for an ego

From Control Boundary to Insurance Claim: Reconstructing AI-Mediated Losses Through the CER Framework

AgentsDGX agent

arXiv:2606.03777v1 Announce Type: new Abstract: AI losses that arise through an insured organization's generative or agentic AI system require state reconstruction, not merely event reconstruction, be

KForge: LLM-Driven Cross-Platform Kernel Generation for AI Accelerators

Model ReleasesDGX agent

arXiv:2606.02963v1 Announce Type: new Abstract: Production inference increasingly targets a heterogeneous mix of accelerators. Agentic pipelines interleave reasoning, tool calls, and multi-agent coord

MemTrain: Self-Supervised Context Memory Training

AgentsDGX agent

arXiv:2606.03197v1 Announce Type: new Abstract: Memory is an indispensable capability for long-horizon LLM agents, enabling them to preserve and utilize information accumulated across extended interac

More info on the Web Dashboard: https://hermes-agent.nousresearch.com/docs/user-guide/features/web-dashboard

AgentsDGX agent

The Nous Research Web Dashboard is a user interface feature that provides access to information and tools for managing Hermes Agent functionality through a web-based platform. According to Nous Resear

2 Jun 2026

MindClaw: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention

AgentsDGX agent

arXiv:2606.01063v1 Announce Type: new Abstract: Theory of Mind (ToM) enables an agent to reason about another actor's beliefs, goals, and intentions, which is essential for human-centered embodied ass

OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence

AgentsDGX agent

arXiv:2603.14771v3 Announce Type: replace Abstract: Large Language Model (LLM)-based Collective Intelligence (CI) presents a promising approach to overcoming the data wall and continuously boosting th

SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale

AgentsDGX agent

arXiv:2602.23866v2 Announce Type: replace-cross Abstract: Software engineering agents (SWE) are improving rapidly, with recent gains largely driven by reinforcement learning (RL). However, RL training

1 Jun 2026

Don't Fool Me Twice: Adapting to Adversity in the Wild with Experience-Driven Reasoning

AgentsDGX agent

arXiv:2605.31119v1 Announce Type: cross Abstract: In robotics, dangers and adversity modes are often embodiment-specific and relative to each agent. A frontier of autonomous mobile robotics is to enab

PInVerify: An Offline Embodied Benchmark for Active Instance Verification

Model ReleasesDGX agent

arXiv:2605.30639v1 Announce Type: cross Abstract: Embodied agents have made strong progress in navigating to target objects, but reaching the goal vicinity does not guarantee that the agent has found

29 May 2026

Bosses, Kings, and the Commons: Cooperation Under Power Asymmetry in LLM Societies

AgentsDGX agent

arXiv:2605.29062v1 Announce Type: new Abstract: Communities can sustainably manage shared resources (commons) through self-governance and cooperative norms, a central finding of Ostrom's theory of sel

FLIP: Real-Time and Resilient Formation Planning for Large-Scale DIstributed Swarms via Point Cloud Registration

AgentsDGX agent

arXiv:2605.29704v1 Announce Type: new Abstract: Traditional large-scale formation planning either oversimplify the formation representation which leads to poor performance, or they employ complete col

Formalizing Mathematics at Scale

AgentsDGX agent

arXiv:2605.29955v1 Announce Type: new Abstract: We present AutoformBot, a multi-agent system for building an Autoformalized Textbook Library At Scale (Atlas) in Lean 4. AutoformBot orchestrates thousa

FRUC: Feedforward Dynamic Scene Reconstruction from Uncalibrated Collaborative Driving Views

AgentsDGX agent

arXiv:2605.29997v1 Announce Type: new Abstract: We present FRUC, a feed-forward 3D Gaussian splatting framework for dynamic scene reconstruction from uncalibrated collaborative driving views. Existing

Indexing the Unreadable: LLM-Native Recursive Construction and Search of Service Taxonomies

Model ReleasesDGX agent

arXiv:2605.29270v1 Announce Type: new Abstract: The era of the Internet of Agents (IoA) is taking shape: LLM agents are expected to fulfill user goals by orchestrating fast-growing populations of Mode

LogDx-CI: Benchmarking Log Reduction Tools for LLM Root-Cause Diagnosis

Model ReleasesDGX agent

arXiv:2605.28876v1 Announce Type: cross Abstract: CI failure logs are large (median 5k lines, max 200k in this corpus) and noisy. Coding agents that try to debug them depend on an upstream tool to red

Our users love @StepFun_ai models and this new release packs a punch at a small size. Looking forward to seeing how well it works with Herme…

AgentsDGX agent

Our users love @StepFun_ai models and this new release packs a punch at a small size. Looking forward to seeing how well it works with Hermes Agent! ⚡️ Step 3.7 Flash is here: The new frontier is agen

Production traffic from frontier models is a golden data asset. If you can efficiently mine the traces, filter for quality, and fine-tune sm…

AgentsDGX agent

Production traffic from frontier models is a golden data asset. If you can efficiently mine the traces, filter for quality, and fine-tune smaller models on them, you get specialized performance at a f

When Should a Robot Think? Resource-Aware Reasoning via Reinforcement Learning for Embodied Robotic Decision-Making

Model ReleasesDGX agent

arXiv:2603.16673v4 Announce Type: replace-cross Abstract: Embodied robotic systems increasingly rely on large language model (LLM)-based agents to support high-level reasoning, planning, and decision-

28 May 2026

I'v never bothered to make a @Shopify store in my life. I ran into a few potential clients who use it so decided to make a shopify ecommerce…

AgentsDGX agent

I'v never bothered to make a @Shopify store in my life. I ran into a few potential clients who use it so decided to make a shopify ecommerce site with it. I had @NousResearch hermes-agent create desig

Reasoning and Planning with Dynamically Changing Norms

AgentsDGX agent

arXiv:2605.27622v1 Announce Type: new Abstract: To safely interact with humans, AI agents must both know our norms and consider them during planning. However, such norm-guided planning has been less e

Rethinking Memory as Continuously Evolving Connectivity

AgentsDGX agent

arXiv:2605.28773v1 Announce Type: cross Abstract: Existing memory-augmented LLM agents often treat memory as a static repository with pre-defined representations and fixed retrieval pipelines, which i

27 May 2026

3D Gaussian Map with Open-Set Semantic Grouping for Vision-Language Navigation

AgentsDGX agent

arXiv:2605.26500v1 Announce Type: new Abstract: Vision-language navigation (VLN) requires an agent to traverse complex 3D environments based on natural language instructions, necessitating a thorough

https://hermes-agent.nousresearch.com/docs/user-guide/features/mcp#catalog-one-click-install-for-nous-approved-mcps

AgentsDGX agent

Nous Research introduced a one-click installation feature in Hermes Agent that allows users to easily install Model Context Protocol (MCP) servers from a curated catalog of Nous-approved integrations.

Introducing Google AI Threat Defense to help you outpace the adversary

Model ReleasesDGX agent

aside_block <ListValue: [StructValue([('title', 'Summary of today’s news'), ('body', <wagtail.rich_text.RichText object at 0x7fb0f516f910>), ('btn_text', ''), ('href', ''), ('image', None)])]> AI-powe

Sentinel: Embodied Cooperative Spatial Reasoning and Planning

Model ReleasesDGX agent

arXiv:2605.26239v1 Announce Type: new Abstract: In this work, we study Cooperative Spatial Intelligence, the ability of decentralized embodied agents to coordinate effectively under dynamic environmen

SIA: Self Improving AI with Harness & Weight Updates

HardwareDGX agent

arXiv:2605.27276v1 Announce Type: new Abstract: Humans are the bottleneck in building and improving AI. Both the models and the agents that wrap them are written, tuned, and corrected by people. The l

26 May 2026

A perspective on fluid mechanical environments for challenges in reinforcement learning

TutorialsDGX agent

arXiv:2605.25011v1 Announce Type: new Abstract: We consider the challenge of developing agents that efficiently interact with high-dimensional, evolving environments, towards a view of practical reinf

CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists

Model ReleasesDGX agent

arXiv:2605.26029v1 Announce Type: new Abstract: We introduce CausaLab, a scalable environment for evaluating interactive causal discovery by LLM agents. Unlike prior evaluations, CausaLab evaluates bo

MCPXKIT: The Unified Toolkit for Analyzing Model Context Protocol Security

AgentsDGX agent

arXiv:2508.12538v2 Announce Type: replace-cross Abstract: The Model Context Protocol (MCP) has emerged as a universal standard that enables AI agents to seamlessly connect with external tools, signifi

Microsoft Copilot Cowork Exfiltrates Files

AgentsDGX agent

Microsoft Copilot Cowork Exfiltrates Files The biggest challenge in designing agentic systems continues to be preventing them from enabling attackers to exfiltrate data. In this case Microsoft Copilot

Reward Shaping and Action Masking for Compositional Tasks using Behavior Trees and LLMs

AgentsDGX agent

arXiv:2605.05795v2 Announce Type: replace Abstract: Decomposing complex tasks into a sequence of simpler subtasks can improve learning efficiency for an autonomous agent. Reinforcement learning (RL) c

SMDD-Bench: Can LLMs Solve Real-World Small Molecule Drug Design Tasks?

Model ReleasesDGX agent

arXiv:2605.21740v2 Announce Type: replace Abstract: LLM agents have incredible potential for scientific discovery applications. However, the performance of LLM agents on real-world, small molecule dru

Trace data is literally worth its weight in gold these days, if you know what to do with it! As has been established, creating effective age…

AgentsDGX agent

Trace data is literally worth its weight in gold these days, if you know what to do with it! As has been established, creating effective agents requires shipping early, observing behavior, and iterati

25 May 2026

Socially fluent AI decouples conversational signals from source identity in online interaction

AgentsDGX agent

arXiv:2605.23426v1 Announce Type: cross Abstract: Socially fluent agentic AI can now participate in online interaction in ways that resemble ordinary human conversation, potentially weakening people's

23 May 2026

// Adapt the Interface, Not the Model // I am fascinated by the results across my cheap-model-plus-good-harness builds. This new paper also …

AgentsDGX agent

// Adapt the Interface, Not the Model // I am fascinated by the results across my cheap-model-plus-good-harness builds. This new paper also shows good signs of the code-as-agent-harness thesis. The id

“It is built in Rust and leverages the Apache DataFusion query engine” Any new database these days

AgentsDGX agent

“It is built in Rust and leverages the Apache DataFusion query engine” Any new database these days We built SmithDB: the database purpose built for agent observability workloads that now powers many p

22 May 2026

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation

AgentsDGX agent

arXiv:2605.22816v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) requires an agent to ground language instructions to its own movement within a visual environment. While state-of

Diagnosis Is Not Prescription: Linguistic Co-Adaptation Explains Patching Hazards in LLM Pipelines

SafetyDGX agent

arXiv:2605.21958v1 Announce Type: new Abstract: When a multi-module LLM agent fails, the module most responsible for the failure is not necessarily the best place to intervene. We demonstrate this Dia

21 May 2026

ProCrit: Self-Elicited Multi-Perspective Reasoning with Critic-Guided Revision for Multimodal Sarcasm Detection

AgentsDGX agent

arXiv:2605.20867v1 Announce Type: cross Abstract: Multimodal sarcasm detection requires reasoning over cross-modal incongruities between literal expression and intended meaning, yet the specific analy

20 May 2026

How Far Are We From True Auto-Research?

Model ReleasesDGX agent

arXiv:2605.19156v1 Announce Type: new Abstract: Recent auto-research systems can produce complete papers, but feasibility is not the same as quality, and the field still lacks a systematic study of ho

very belated but in retrospect i think @sama's mythical 'build a business that gets better when models get better' is basically what I calle…

AgentsDGX agent

very belated but in retrospect i think @sama's mythical 'build a business that gets better when models get better' is basically what I called Agent Labs here. seeing a very direct correlation with mod

Very cool - Grok Build is clearly getting better by the day. Two nights ago I ran an overnight build and it failed. Last night, success. I r…

AgentsDGX agent

Very cool - Grok Build is clearly getting better by the day. Two nights ago I ran an overnight build and it failed. Last night, success. I really like the multi-agent orchestration behavior, it does a

19 May 2026

A Machine with Short-Term, Episodic, and Semantic Memory Systems

Model ReleasesDGX agent

arXiv:2212.02098v5 Announce Type: replace Abstract: Inspired by the cognitive science theory of the explicit human memory systems, we have modeled an agent with short-term, episodic, and semantic memo

Convergence of Multiagent Learning Systems for Traffic control

AgentsDGX agent

arXiv:2511.11654v2 Announce Type: replace-cross Abstract: Rapid urbanization in cities like Bangalore has led to severe traffic congestion, making efficient Traffic Signal Control (TSC) essential. Mul

I've been thinking a lot about the two different groups of evals you need in general agents/agents which handle broad tasks: 1. Benchmark ev…

Model ReleasesDGX agent

I've been thinking a lot about the two different groups of evals you need in general agents/agents which handle broad tasks: 1. Benchmark evals - this is a suite of up to 100 eval cases which test the

Learning to Learn from Multimodal Experience

AgentsDGX agent

arXiv:2605.16857v1 Announce Type: new Abstract: Experience-driven learning has emerged as a promising paradigm for enabling agents to improve from interaction trajectories by accumulating and reusing

PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models

AgentsDGX agent

arXiv:2605.17044v1 Announce Type: new Abstract: Large language models (LLMs) increasingly serve as interactive social agents, yet their ability to maintain coherent and authentic persona-level role-pl

← Previous
1…156157158159160…300
Next →