AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,920 results
20 Apr 2026

AI Agents and Hard Choices

SafetyDGX agent

arXiv:2504.15304v2 Announce Type: replace Abstract: Can AI agents deal with hard choices -- cases where options are incommensurable because multiple objectives are pursued simultaneously? Adopting a t

COMPASS: Benchmarking Constrained Optimization in LLM Agents

Model ReleasesDGX agent

arXiv:2510.07043v2 Announce Type: replace Abstract: Human decision-making often involves constrained optimization. As LLM agents are deployed to assist with real-world tasks like travel planning, shop

Scalable Multi-Task Learning through Spiking Neural Networks with Adaptive Task-Switching Policy for Intelligent Autonomous Agents

SafetyDGX agent

arXiv:2504.13541v5 Announce Type: replace-cross Abstract: Training resource-constrained autonomous agents on multiple tasks simultaneously is crucial for adapting to diverse real-world environments. R

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
17 Apr 2026

Atropos: Improving Cost-Benefit Trade-off of LLM-based Agents under Self-Consistency with Early Termination and Model Hotswap

Local AiDGX agent

arXiv:2604.15075v1 Announce Type: cross Abstract: Open-weight Small Language Models(SLMs) can provide faster local inference at lower financial cost, but may not achieve the same performance level as

Exploration and Exploitation Errors Are Measurable for Language Model Agents

SafetyDGX agent

arXiv:2604.13151v1 Announce Type: new Abstract: Language Model (LM) agents are increasingly used in complex open-ended decision-making tasks, from AI coding to physical AI. A core requirement in these

Is there still a widespread belief that LLMs and coding agents are good for greenfield development but don't help for maintaining large exis…

ToolsDGX agent

Simon Willison discusses the perception that LLMs and coding agents are primarily useful for greenfield development (starting new projects from scratch) rather than for maintaining and modifying large

One portal, unlimited possibilities. You can now access Modal via Tool Gateway by @NousResearch, makers of Hermes Agent. Check it out 👇

AgentsDGX agent

One portal, unlimited possibilities. You can now access Modal via Tool Gateway by @NousResearch, makers of Hermes Agent. Check it out 👇 Tool Gateway is now live in Nous Portal. No separate accounts, n

ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration

SafetyDGX agent

arXiv:2509.21823v2 Announce Type: replace Abstract: Reward is critical to the evaluation and training of large language models (LLMs). However, existing rule-based or model-based reward methods strugg

Towards a Multi-Embodied Grasping Agent

AgentsDGX agent

arXiv:2510.27420v3 Announce Type: replace Abstract: Multi-embodiment grasping focuses on developing approaches that exhibit generalist behavior across diverse gripper designs. Existing methods often l

16 Apr 2026

AI Search: the search primitive for your agents

IndustryDGX agent

AI Search is the search primitive for your agents. Create instances dynamically, upload files, and search across instances with hybrid retrieval and relevance boosting. Just create a search instance,

HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark

Model ReleasesDGX agent

arXiv:2604.13954v1 Announce Type: new Abstract: Existing agent-safety evaluation has focused mainly on externally induced risks. Yet agents may still enter unsafe trajectories under benign conditions.

LiteParse should be the default document parser you use with any AI agent (Claude Code, Claude Cowork, OpenClaw, Codex, and more) The core i…

Model ReleasesDGX agent

LiteParse should be the default document parser you use with any AI agent (Claude Code, Claude Cowork, OpenClaw, Codex, and more) The core is extremely fast text and accurate parsing from any document

15 Apr 2026

A Benchmark for Evaluating Outcome-Driven Constraint Violations in Autonomous AI Agents

Model ReleasesDGX agent

arXiv:2512.20798v4 Announce Type: replace Abstract: As autonomous AI agents are deployed in high-stakes environments, ensuring their safety has become a paramount concern. Existing safety benchmarks p

ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search

Model ReleasesDGX agent

arXiv:2604.12762v1 Announce Type: cross Abstract: We introduce ARGOS, the first benchmark and framework that reformulates multi-camera person search as an interactive reasoning problem requiring an ag

From Plan to Action: How Well Do Agents Follow the Plan?

Model ReleasesDGX agent

arXiv:2604.12147v1 Announce Type: cross Abstract: Agents aspire to eliminate the need for task-specific prompt crafting through autonomous reason-act-observe loops. Still, they are commonly instructed

I'm going all in on Hermes (@NousResearch, @Teknium1) as my entire agent and coding stack. Six profiles. One shared self-hosted memory store…

Model ReleasesDGX agent

I'm going all in on Hermes (@NousResearch, @Teknium1) as my entire agent and coding stack. Six profiles. One shared self-hosted memory store. Zero hosted-coder dependencies. The fleet: - pmax-mousa —

Ollama Open-Source Agent Self-Reflection Harness

Local AiDGX agent

An open-source agent self-reflection harness built on top of Ollama, shared in the r/ollama community, that enables locally run LLMs to evaluate and iteratively refine their own outputs. The project p

Towards Robust Real-World Spreadsheet Understanding with Multi-Agent Multi-Format Reasoning

AgentsDGX agent

arXiv:2604.12282v1 Announce Type: new Abstract: Spreadsheets are central to real-world applications such as enterprise reporting, auditing, and scientific data management. Despite their ubiquity, exis

14 Apr 2026

Competing with AI Scientists: Agent-Driven Approach to Astrophysics Research

Model ReleasesDGX agent

arXiv:2604.09621v1 Announce Type: new Abstract: We present an agent-driven approach to the construction of parameter inference pipelines for scientific data analysis. Our method leverages a multi-agen

Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning

AgentsDGX agent

arXiv:2604.09813v1 Announce Type: new Abstract: Existing synthetic tool-use corpora are primarily designed for offline supervised fine-tuning, yet reinforcement learning (RL) requires executable envir

DarwinNet: An Evolutionary Network Architecture for Agent-Driven Protocol Synthesis

AgentsDGX agent

arXiv:2604.01236v2 Announce Type: replace-cross Abstract: Traditional network architectures suffer from severe protocol ossification and structural fragility due to their reliance on static, human-def

Do We Still Need GraphRAG? Benchmarking RAG and GraphRAG for Agentic Search Systems

Model ReleasesDGX agent

arXiv:2604.09666v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) and its graph-based extensions (GraphRAG) are effective paradigms for improving large language model (LLM) reason

EvoDiagram: Agentic Editable Diagram Creation via Design Expertise Evolution

Model ReleasesDGX agent

arXiv:2604.09568v1 Announce Type: cross Abstract: High-fidelity diagram creation requires the complex orchestration of semantic topology, visual styling, and spatial layout, posing a significant chall

MAFIG: Multi-agent Driven Formal Instruction Generation Framework

Local AiDGX agent

arXiv:2604.10989v1 Announce Type: new Abstract: Emergency situations in scheduling systems often trigger local functional failures that undermine system stability and even cause system collapse. Exist

Most AI assistants wait for you to ask. But a truly useful agent should notice you need help before you say anything. New research takes a s…

Model ReleasesDGX agent

Most AI assistants wait for you to ask. But a truly useful agent should notice you need help before you say anything. New research takes a serious shot at building proactive agents that work in real t

The Amazing Agent Race: Strong Tool Users, Weak Navigators

Model ReleasesDGX agent

arXiv:2604.10261v1 Announce Type: new Abstract: Existing tool-use benchmarks for LLM agents are overwhelmingly linear: our analysis of six benchmarks shows 55 to 100% of instances are simple chains of

Three Roles, One Model: Role Orchestration at Inference Time to Close the Performance Gap Between Small and Large Agents

Model ReleasesDGX agent

arXiv:2604.11465v1 Announce Type: new Abstract: Large language model (LLM) agents show promise on realistic tool-use tasks, but deploying capable agents on modest hardware remains challenging. We stud

VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments

Model ReleasesDGX agent

arXiv:2506.02387v3 Announce Type: replace Abstract: Recent advancements in Vision Language Models (VLMs) have expanded their capabilities to interactive agent tasks, yet existing benchmarks remain lim

We've been developing a multi-agent system that builds and maintains complex software autonomously. Recently, we partnered with NVIDIA to ap…

HardwareDGX agent

We've been developing a multi-agent system that builds and maintains complex software autonomously. Recently, we partnered with NVIDIA to apply it to optimizing CUDA kernels. In 3 weeks, it delivered

13 Apr 2026

Agentic Jackal: Live Execution and Semantic Value Grounding for Text-to-JQL

Model ReleasesDGX agent

arXiv:2604.09470v1 Announce Type: new Abstract: Translating natural language into Jira Query Language (JQL) requires resolving ambiguous field references, instance-specific categorical values, and com

11 Apr 2026

a big part of agent harnesses is how they interact with context memory is just context its therefor impossible to separate harness from memo…

AgentsDGX agent

a big part of agent harnesses is how they interact with context memory is just context its therefor impossible to separate harness from memory - as @sarahwooders says, 'memory isn't a plugin (it's a h

10 Apr 2026

LiteParse is the best document parsing library for coding agents. It's free, fast, integrates natively with the LLM's native visual understa…

Model ReleasesDGX agent

LiteParse is the best document parsing library for coding agents. It's free, fast, integrates natively with the LLM's native visual understanding capabilities, and comes with support for 50+ formats a

Memory Scaling for AI Agents

IndustryDGX agent

Databricks Research introduced **MemAlign**, a memory framework for AI agents that stores past interactions as episodic memories and uses an LLM to distill them into generalized semantic rules, whi...

Mina: A Multilingual LLM-Powered Legal Assistant Agent for Bangladesh for Empowering Access to Justice

AgentsDGX agent

arXiv:2511.08605v3 Announce Type: replace Abstract: Bangladesh's low-income population faces major barriers to affordable legal advice due to complex legal language, procedural opacity, and high costs

RemoteAgent: Bridging Vague Human Intents and Earth Observation with RL-based Agentic MLLMs

AgentsDGX agent

arXiv:2604.07765v1 Announce Type: new Abstract: Earth Observation (EO) systems are essentially designed to support domain experts who often express their requirements through vague natural language ra

8 Apr 2026

GLM-5.1 gives teams a stronger model for coding, tool use, and sustained agent performance on Together AI. Learn more: http://www.together.a…

AgentsDGX agent

GLM-5.1 is Z.ai's post-training upgrade to GLM-5, now available on Together AI, delivering a 28% coding performance improvement through a refined reinforcement learning pipeline while retaining the...

7 Apr 2026

How can you improve your agentic search pipeline? I just wrote a blog post with @tech_optimist from @lancedb to answer exactly that. TLDR: -…

Model ReleasesDGX agent

How can you improve your agentic search pipeline? I just wrote a blog post with @tech_optimist from @lancedb to answer exactly that. TLDR: - Parse files and take page-level screenshots with LiteParse,

14 Aug 2026

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents

SafetyDGX agent

arXiv:2608.12764v1 Announce Type: cross Abstract: Deep search agents operate over trajectories spanning dozens of steps, yet standard reinforcement learning provides only a single outcome reward per t

Capability Sheaves for Compositional Agent-Harness Repair: Controlled Quotients and a Real-Repository Stress Test

Local AiDGX agent

arXiv:2608.13228v1 Announce Type: new Abstract: Agent harnesses combine retrieval, routing, state, provenance, and verification, but locally successful components may disagree on shared state. We mode

Introducing GLM-5.3: Built to Code. Ready for Cyber Defense. - Top-tier coding and agentic capabilities, achieved through post-training on t…

AgentsDGX agent

Introducing GLM-5.3: Built to Code. Ready for Cyber Defense. - Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model - A major leap in cybersecurity, setting

ReflectFact: Self-Reflective Agents for Improving Comprehension and Reasoning in Multi-Hop Fact Verification

SafetyDGX agent

arXiv:2608.12877v1 Announce Type: new Abstract: Multi-hop fact verification, which verifies claims by reasoning over multiple pieces of evidence, is critical for combating misinformation on social med

UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations

AgentsDGX agent

arXiv:2608.13031v1 Announce Type: cross Abstract: Traffic video understanding has become an important problem in intelligent transportation, as road videos provide direct evidence for accidents, viola

13 Aug 2026

Agent Safety Should Be a Runtime Contract

Model ReleasesDGX agent

arXiv:2608.11274v1 Announce Type: cross Abstract: The dominant paradigm treats AI safety as a property to be instilled during model training via RLHF, DPO, or Constitutional AI. We argue this is struc

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction

SafetyDGX agent

arXiv:2608.11772v1 Announce Type: new Abstract: Self-correction is particularly useful when a failure constrains the next repair. Coding agents benefit from this property because compilers, tests, and

our outbound harness is focused on delivering the best economics we can, we're continuing to push on how our agents work to get another 90%+…

AgentsDGX agent

our outbound harness is focused on delivering the best economics we can, we're continuing to push on how our agents work to get another 90%+ cost savings for customers great talk by @hwchase17 and @He

Self-Evolving Embodied Agents via Skill-Harness Evolution

Model ReleasesDGX agent

arXiv:2608.11350v1 Announce Type: new Abstract: Embodied agents are increasingly built as systems around foundation models, where performance depends not only on model weights but also on the skills,

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark

Model ReleasesDGX agent

arXiv:2608.11469v1 Announce Type: cross Abstract: AI agents are rapidly improving in cybersecurity capabilities when the source code is available for analysis, yet much of the software most consequent

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

Model ReleasesDGX agent

arXiv:2607.11175v2 Announce Type: replace Abstract: The growing ability of large language models and vision-language models to jointly interpret and reason over images and text is reshaping medical im

Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems

AgentsDGX agent

arXiv:2608.11879v1 Announce Type: new Abstract: Long-running conversational agents increasingly rely on a memory system to avoid resending the whole conversation each turn, yet how much that costs to

12 Aug 2026

On The Statistical Limits of Self-Improving Agents

SafetyDGX agent

arXiv:2510.04399v3 Announce Type: replace Abstract: We develop a learning-theoretic framework for analyzing self-improving agents by decomposing self-modification into five axes. Within this framework

Persistent Recursive Worlds Enable Autonomous Software Evolution

Model ReleasesDGX agent

arXiv:2608.10450v1 Announce Type: cross Abstract: Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems preserve conti

Robust Multi-Agent Bandits with Heavy-Tailed Rewards and Information Asymmetry

AgentsDGX agent

arXiv:2608.10529v1 Announce Type: cross Abstract: The multi-armed bandit problem is a central framework in sequential decision-making, extensively studied under sub-Gaussian reward assumptions. Howeve

SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models

AgentsDGX agent

arXiv:2608.10538v1 Announce Type: new Abstract: Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within agent harness systems as an essenti

Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding

Model ReleasesDGX agent

arXiv:2608.11095v1 Announce Type: new Abstract: Agentic coding READMEs like CLAUDE.md grow without bound in real repositories, stopping only when the repository retires or someone rewrites the file wh

11 Aug 2026

AndroidReality: How Far Are Mobile Agents from the Real World?

Model ReleasesDGX agent

arXiv:2608.07775v1 Announce Type: new Abstract: Mobile agents have achieved promising results on clean online benchmarks such as AndroidWorld, yet their performance often degrades sharply in real-worl

CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception

Model ReleasesDGX agent

arXiv:2608.08392v1 Announce Type: new Abstract: Large language models are increasingly deployed as autonomous agents that interact with the web through browsers. While recent progress has been driven

Compiling and Benchmarking Task-State Horizons for Embodied Agents

Model ReleasesDGX agent

arXiv:2608.08036v1 Announce Type: new Abstract: Frontier agentic models are increasingly deployed as high-level planners for long-horizon embodied tasks. Existing robotic benchmarks have advanced long

LatticeMind: A Conflict-Aware Memory Primitive for Multi-Agent Systems

AgentsDGX agent

arXiv:2608.08236v1 Announce Type: new Abstract: Multi-agent LLM systems often fail not for lack of candidate answers, but because they have no persistent mechanism for deciding which incompatible clai

Multi-agent discovery of practical quantum LDPC codes

AgentsDGX agent

arXiv:2608.08996v1 Announce Type: cross Abstract: Quantum low-density parity-check (qLDPC) codes can encode multiple logical qubits using sparse parity checks, yet searching for useful finite-length i

SAKE: Structured Agentic Knowledge Extrapolation for Complex LLM Reasoning via Reinforcement Learning

AgentsDGX agent

arXiv:2505.15062v5 Announce Type: replace-cross Abstract: Knowledge extrapolation is the process of inferring novel information by combining and extending existing knowledge that is explicitly availab

← Previous
1…7576777879…299
Next →