AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,771 results
11 Aug 2026

Controlled Memory Interference in Continual LLM Agents

AgentsDGX agent

arXiv:2608.07622v1 Announce Type: new Abstract: Long-term memory enables AI agents to maintain continuity across sessions, personalize behavior, and evolve through accumulated experience. Yet memory e

Electric joins Databricks to bring WASM Postgres to AI agent sandboxes

AgentsDGX agent

Electric has partnered with Databricks to extend Databricks’ PostgreSQL capabilities from the lakehouse into edge environments. This collaboration allows a wide range of lightweight open‑source databa

Improving Constraint Models with LLM Agents

AgentsDGX agent

arXiv:2608.08127v1 Announce Type: new Abstract: The runtime of Constraint Programming (CP) solvers is highly sensitive to modeling choices, such as symmetry breaking, implied constraints, global const

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning

AgentsDGX agent

arXiv:2608.08255v1 Announce Type: cross Abstract: Agentic reinforcement learning (RL) often suffers from delayed and sparse rewards in real-world environments. A promising solution to this challenge i

M^3Prune: Hierarchical Communication Graph Pruning for Efficient Multi-Modal Multi-Agent Retrieval-Augmented Generation

AgentsDGX agent

arXiv:2511.19969v2 Announce Type: replace Abstract: Recent advancements in multi-modal retrieval-augmented generation (mRAG), which enhance multi-modal large language models (MLLMs) with external know

Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents

AgentsDGX agent

arXiv:2608.08389v1 Announce Type: new Abstract: Long-horizon research agents solve open-ended tasks through iterative retrieval, aggregation, and synthesis, but context grows rapidly while the margina

NVIDIA and Local AI Community Fuel Open Source Models and Intelligent Agents

Local AiDGX agent

The open source ecosystem is making it easier for AI enthusiasts and developers to build, customize and run increasingly capable agents locally. Throughout August, NVIDIA is celebrating the partners a

Reflex First, Reflect Later: Latency-Aware Embodied LLM Agents for Dynamic Response

SafetyDGX agent

arXiv:2506.07223v2 Announce Type: replace Abstract: Large language models (LLMs) have substantially improved the planning capabilities of embodied agents, enabling their deployment in dynamic and safe

SkillSentry: Reliable Skill Execution for LLM Agents via Runtime Assurance

Model ReleasesDGX agent

arXiv:2608.09253v1 Announce Type: new Abstract: LLM agents are increasingly equipped with skills to perform complex tasks through multi-step reasoning and tool use. Although skills provide reusable pr

The Collaboration Gap: Exploration and Benchmarking of Open-World Agentic Cooperation

Model ReleasesDGX agent

arXiv:2511.02687v2 Announce Type: replace Abstract: The trajectory of AI development suggests that we will increasingly rely on agent-based systems powered by language models, composed of independentl

The Scaffolding Matters More Than the Interface: A Controlled Comparison of MCP and CLI Tool Use Across Seven Agent Scaffoldings, Five Language Models, and One Software Task

Model ReleasesDGX agent

arXiv:2608.08654v1 Announce Type: new Abstract: How much an AI coding agent costs to run can depend more on the agent scaffolding that drives it than on the interface through which it reaches its tool

10 Aug 2026

Agentic Planning for Symbolic Execution

AgentsDGX agent

arXiv:2608.06397v1 Announce Type: cross Abstract: Symbolic execution seeks to explore feasible program paths, yet a practical run may exhaust its resources while much program behaviour remains unreach

Evaluating XAI Support From A Hierarchical Reinforcement Learning Policy in Human-Agent Collaboration

Model ReleasesDGX agent

arXiv:2608.06381v1 Announce Type: cross Abstract: Explainable AI (XAI) has shown promise for human-agent collaboration, yet results rely on hand-crafted policies in custom environments, limiting gener

🎉 Introducing Hermes in ANY app Your personal agent automating work, controlling software, or running tasks with generative UI, human-in-th…

AgentsDGX agent

🎉 Introducing Hermes in ANY app Your personal agent automating work, controlling software, or running tasks with generative UI, human-in-the-loop and more. Use Hermes anywhere over AG-UI: → React & Re

Introducing Muse Glimmer, an open-weight 30B-parameter model optimized for local, always-on agent workflows. Muse Glimmer delivers strong pe…

Model ReleasesDGX agent

Introducing Muse Glimmer, an open-weight 30B-parameter model optimized for local, always-on agent workflows. Muse Glimmer delivers strong performance on key agentic use cases and benchmarks compared w

Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework

Model ReleasesDGX agent

arXiv:2608.06909v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly operate through long-horizon trajectories involving user instructions, tool use, external observations, a

On theCUBE Pod: Black Hat exposes agentic threat, theCUBE remembers David Floyer

AgentsDGX agent

Artificial intelligence systems are developing faster than cybersecurity experts — and the energy grid — can keep up with. At the recent Black Hat USA event, security analysts viewed the rise of AI-dr

Online Monitoring and Corrective Steering of Programming Agents

Local AiDGX agent

arXiv:2608.06701v1 Announce Type: cross Abstract: Fixing GitHub issues in large-scale projects is a long-horizon task, especially when a fix requires changes across multiple locations or the issue des

7 Aug 2026

AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games

AgentsDGX agent

arXiv:2608.06362v1 Announce Type: cross Abstract: Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time.

Hermes Agent now supports all the portable plugins standard that many other major AI players have adopted. Currently these portable plugins …

AgentsDGX agent

Hermes Agent now supports all the portable plugins standard that many other major AI players have adopted. Currently these portable plugins only support MCPs and Skills, use native Hermes plugins to a

6 Aug 2026

Breadcrumbing Search Agents

SafetyDGX agent

arXiv:2608.04565v1 Announce Type: cross Abstract: LLM-based search agents are widely used for information-seeking tasks, but their reliance on external tool returns introduces a critical security risk

DAC-Pose: Dual-Agent Collaborative Framework for Pose-Guided Human Generation

SafetyDGX agent

arXiv:2608.04622v1 Announce Type: new Abstract: AI agents have emerged as a powerful new paradigm in generative image synthesis, enabling systems to perform complex semantic reasoning rather than pass

EviGraph: Evidence-Guided Autonomous Research Agents

AgentsDGX agent

arXiv:2608.04738v1 Announce Type: new Abstract: Autonomous research agents can generate hypotheses, execute experiments, and draft manuscripts, yet their outputs often contain unsupported claims and i

MatrAIx: Simulating the World with 8.3 Billion Persona Agents

Model ReleasesDGX agent

arXiv:2608.04205v1 Announce Type: new Abstract: Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract aw

OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents

AgentsDGX agent

arXiv:2608.05013v1 Announce Type: cross Abstract: LLM agents are increasingly applied to open-ended everyday requests that span work, study, and life. These tasks are long-horizon, cross-environment,

PRIMAL3: Pathfinding via Reinforcement and Imitation Multi-Agent Learning - Leveraging LaCAM3

SafetyDGX agent

arXiv:2608.04905v1 Announce Type: new Abstract: We present PRIMAL3, an ultra-large-scale learning-based framework for multi-agent pathfinding (MAPF) that integrates reinforcement learning, topology-aw

Tenex pairs agentic AI with human oversight for faster security operations

AgentsDGX agent

As AI accelerates the speed and scale of cyberattacks, organizations are adopting AI security operations to investigate threats and respond in minutes rather than hours or days. The shift is enabling

We made an MCP for your phone. Your laptop is just half of your life, and the other half is in your phone. Now your agent gets the mobile sc…

Model ReleasesDGX agent

We made an MCP for your phone. Your laptop is just half of your life, and the other half is in your phone. Now your agent gets the mobile screen too. No connectors, no complex setup, it has access to

5 Aug 2026

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks

Model ReleasesDGX agent

arXiv:2608.03764v1 Announce Type: new Abstract: Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively. Evaluating self-ev

Improving Sample Efficiency in Multi-Agent Reinforcement Learning for Simulated Football Games via Exploration

AgentsDGX agent

arXiv:2503.13077v2 Announce Type: replace Abstract: Multi-agent reinforcement learning has shown promise in learning cooperative behaviors in team-based environments. However, such methods often deman

Meta releases Muse Code in beta, a terminal coding agent powered by Muse Spark 1.2, a coding-focused model priced at 1.25/1M input and 4.25/1M output tokens (Jonathan Vanian/CNBC)

AgentsDGX agent

Jonathan Vanian / CNBC: Meta releases Muse Code in beta, a terminal coding agent powered by Muse Spark 1.2, a coding-focused model priced at 1.25/1M input and 4.25/1M output tokens — Meta is rolling o

OR-Agent: Bridging Evolutionary Search and Structured Research for Automated Algorithm Discovery

AgentsDGX agent

arXiv:2602.13769v3 Announce Type: replace Abstract: Automating heuristic design in complex, experiment-driven domains requires more than iterative mutation of solution algorithms. Current LLM-based ev

Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents

AgentsDGX agent

arXiv:2608.02751v1 Announce Type: cross Abstract: Existing deep-research agents use a search-visit workflow that retrieves and reads whole pages, without considering the addressable structure that web

SITUATION DETECTED: Prime Intellect is releasing Prime Agent, a self-improving harness for coding and long-running autonomous tasks. The tea…

Model ReleasesDGX agent

SITUATION DETECTED: Prime Intellect is releasing Prime Agent, a self-improving harness for coding and long-running autonomous tasks. The team reports 95.5% on ARC-AGI-3, above the human baseline, and

Skill libraries are shipping in agent harnesses on the assumption that writing skills down compounds. A new benchmark tests that directly. C…

Model ReleasesDGX agent

Skill libraries are shipping in agent harnesses on the assumption that writing skills down compounds. A new benchmark tests that directly. ContinualSkillBench covers five domains, each with 100 interc

Training Documents Reranker with Search Rubrics for Deep Research Agent

AgentsDGX agent

arXiv:2608.03527v1 Announce Type: cross Abstract: Retrieval systems help deep research agents generate high-quality answers by providing relevant documents. However, existing retrievers typically sele

Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking

Model ReleasesDGX agent

arXiv:2606.15673v2 Announce Type: replace Abstract: Web agents act through long interaction sequences, yet existing benchmarks evaluate only terminal success, discarding all process information and of

4 Aug 2026

AgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategies in Conversational AI Agents

Model ReleasesDGX agent

arXiv:2608.00009v1 Announce Type: new Abstract: Long-term memory remains a critical bottleneck for conversational AI agents, whose finite context windows cannot support coherent recall across thousand

DrawAI: Agentic Benchmark and Workflow for Making Raster Images Editable

Model ReleasesDGX agent

arXiv:2608.00548v1 Announce Type: new Abstract: Recent image-generation models and multimodal agents can produce high-quality visuals for increasingly complex visual communication tasks. Yet their ras

PGMem: Tightly Coupled Persona-Memory Graph for Lifelong Personalized Agents

AgentsDGX agent

arXiv:2608.01708v1 Announce Type: new Abstract: Long-term personalized dialogue agents must track user preferences as their personas evolve. Existing memory systems organize past events well, but stor

Token-Native Storage: Read and Write in your Agent's Language

AgentsDGX agent

arXiv:2608.02376v1 Announce Type: cross Abstract: Search and database engines still store text as UTF-8, a format built for humans. But the systems that increasingly read and write that text (embedder

Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit

AgentsDGX agent

arXiv:2608.02302v1 Announce Type: cross Abstract: Long-horizon coding-agent trajectories are poorly matched to the credit units available to train on: a single action has no stable value, an episode l

3 Aug 2026

AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction

Model ReleasesDGX agent

arXiv:2607.29549v1 Announce Type: new Abstract: Large language models have demonstrated strong mathematical problem-solving capabilities, yet reliably verifying their candidate answers remains challen

Cato Networks launches Agentic Threat Prevention to counter AI-assisted attacks

Model ReleasesDGX agent

Networking and security company Cato Networks Ltd. today introduced Cato Agentic Threat Prevention, a capability that uses autonomous agents to predict the route an attacker is likely to take through

EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

AgentsDGX agent

arXiv:2607.18529v2 Announce Type: replace-cross Abstract: Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Exist

RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems

Local AiDGX agent

arXiv:2607.29241v1 Announce Type: cross Abstract: Optimizing modern recommender models still depends heavily on engineers manually iterating over architectural, objective, and training-strategy change

Sakana Namazu: An LLM API with Japanese-vibes! 🎏 Built for Japanese enterprises, featuring frontier-level reasoning and built-in agentic to…

AgentsDGX agent

Sakana Namazu: An LLM API with Japanese-vibes! 🎏 Built for Japanese enterprises, featuring frontier-level reasoning and built-in agentic tools. 開発者の皆様、大変お待たせしました!Sakana Chatのモデルが遂にAPIとして公開です。ぜひお試しください

SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition

Model ReleasesDGX agent

arXiv:2607.28692v1 Announce Type: new Abstract: Large language model (LLM) agents have been increasingly adopted in scientific research for organizing and invoking specialized computational tools. How

Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents

SafetyDGX agent

arXiv:2607.29254v1 Announce Type: new Abstract: AI agents extend large language models (LLMs) with external tools, enabling them to perform complex tasks and translate model outputs into consequential

Validation Evidence in LLM Repair Agents: How Much of What Passes Actually Tests the Bug?

Model ReleasesDGX agent

arXiv:2607.28871v1 Announce Type: cross Abstract: When a repair agent runs a test and sees it pass, the result is treated as evidence about the reported defect. We measure how often that treatment is

what does the 'last human code review' look like? to @itamar_mar, ceo and founder of @QodoAI, it looks like two agents talking to each other…

AgentsDGX agent

what does the 'last human code review' look like? to @itamar_mar, ceo and founder of @QodoAI, it looks like two agents talking to each other, backed by a context engine specific to your organization.

1 Aug 2026

3/ here's the part that makes it non-optional: the same agent that will do whatever it takes to solve a problem will also walk straight out …

Model ReleasesDGX agent

3/ here's the part that makes it non-optional: the same agent that will do whatever it takes to solve a problem will also walk straight out of a sandbox you thought was locked down. We watched exactly

31 Jul 2026

A Graph-Native Bitemporal Memory Store for Conversational AI Agents

Model ReleasesDGX agent

arXiv:2607.26520v1 Announce Type: cross Abstract: Conversational AI agents commonly lack persistent memory across sessions. The obvious fixes like injecting full chat histories into the context window

ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Science

Model ReleasesDGX agent

arXiv:2607.26155v1 Announce Type: new Abstract: Clinical data-science agents must transform heterogeneous longitudinal records into auditable analyses, yet existing benchmarks largely isolate medical

Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale

Model ReleasesDGX agent

arXiv:2607.28074v1 Announce Type: cross Abstract: Computer-use agents learn from what their actions change, so training one needs applications it can act on, break and reset. The applications that mat

IDP AutoOpt: Agent-Driven Optimization of Document Processing Pipeline Configurations

Model ReleasesDGX agent

arXiv:2607.26075v1 Announce Type: cross Abstract: We present IDP AutoOpt, an autonomous LLM agent that discovers high-performing configurations for intelligent document processing (IDP) pipelines. Tun

RoboBRIDGE: A Modular Framework for Bridging Policies to Robust Real-World Robotic Agents

AgentsDGX agent

arXiv:2607.27881v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have attracted growing interest as a scalable approach to robotic manipulation. While these models are effective act

ThreatForest: Multi-Agent Attack Tree Generation with Pluggable TTP Framework Mapping

AgentsDGX agent

arXiv:2607.27528v1 Announce Type: cross Abstract: Threat modeling is essential for secure software development, yet manual analysis of cloud-native architectures is slow and demands scarce security ex

When Should AI Follow? Task Structure and Joint Adaptation by Human and AI Agents

Local AiDGX agent

arXiv:2504.20903v4 Announce Type: replace-cross Abstract: How should organizations divide and sequence decision tasks between human and artificial agents? We develop a computational model of joint seq

Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees

Model ReleasesDGX agent

arXiv:2607.28399v1 Announce Type: new Abstract: Computer-use agents often fail on transient GUI events because they produce the correct action only after the relevant window has already closed. We ide

← Previous
1…5455565758…297
Next →