AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,950 results
Safety

LiteParse: our open-source, layout-aware PDF parser for AI agents. The secret? Grid projection. Instead of heavy ML layout models or flat te…

DGX agent

LiteParse: our open-source, layout-aware PDF parser for AI agents. The secret? Grid projection. Instead of heavy ML layout models or flat text extraction, it projects text onto a monospace grid so ali

safetyjerry-liu--x
22 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Introducing ml-intern, the agent that just automated the post-training team @huggingface It's an open-source implementation of the real rese…

DGX agent

Introducing ml-intern, the agent that just automated the post-training team @huggingface It's an open-source implementation of the real research loop that our ML researchers do every day. You give it

model-releasesclem-delangue--x
21 Apr 2026
Model Releases

Kimi K2.6 @Kimi_Moonshot is the new leading open-weights agent model, landing at #4 on Claw-Eval (Pass^3: 62.3%). Key takeaways: - 👑 Best o…

DGX agent

Kimi K2.6 @Kimi_Moonshot is the new leading open-weights agent model, landing at #4 on Claw-Eval (Pass^3: 62.3%). Key takeaways: - 👑 Best open-source agent, period: Pass^3 of 62.3% is the highest of a

model-releaseskimi-moonshot--x
21 Apr 2026
Model Releases

MTSQL-R1: Towards Long-Horizon Multi-Turn Text-to-SQL via Agentic Training

DGX agent

arXiv:2510.12831v3 Announce Type: replace Abstract: Multi-turn Text-to-SQL aims to translate a user's conversational utterances into executable SQL while preserving dialogue coherence and grounding to

model-releasesarxiv-cs-cl
21 Apr 2026
Safety

SynAgent: Generalizable Cooperative Humanoid Manipulation via Solo-to-Cooperative Agent Synergy

DGX agent

arXiv:2604.18557v1 Announce Type: new Abstract: Controllable cooperative humanoid manipulation is a fundamental yet challenging problem for embodied intelligence, due to severe data scarcity, complexi

safetyarxiv-cs-cv
21 Apr 2026
Safety

AI Agents and Hard Choices

DGX agent

arXiv:2504.15304v2 Announce Type: replace Abstract: Can AI agents deal with hard choices -- cases where options are incommensurable because multiple objectives are pursued simultaneously? Adopting a t

safetyarxiv-cs-ai
20 Apr 2026
Model Releases

COMPASS: Benchmarking Constrained Optimization in LLM Agents

DGX agent

arXiv:2510.07043v2 Announce Type: replace Abstract: Human decision-making often involves constrained optimization. As LLM agents are deployed to assist with real-world tasks like travel planning, shop

model-releasesarxiv-cs-lg
20 Apr 2026
Safety

Scalable Multi-Task Learning through Spiking Neural Networks with Adaptive Task-Switching Policy for Intelligent Autonomous Agents

DGX agent

arXiv:2504.13541v5 Announce Type: replace-cross Abstract: Training resource-constrained autonomous agents on multiple tasks simultaneously is crucial for adapting to diverse real-world environments. R

safetyarxiv-cs-ai
20 Apr 2026
Local Ai

Atropos: Improving Cost-Benefit Trade-off of LLM-based Agents under Self-Consistency with Early Termination and Model Hotswap

DGX agent

arXiv:2604.15075v1 Announce Type: cross Abstract: Open-weight Small Language Models(SLMs) can provide faster local inference at lower financial cost, but may not achieve the same performance level as

local-aiarxiv-cs-lg
17 Apr 2026
Safety

Exploration and Exploitation Errors Are Measurable for Language Model Agents

DGX agent

arXiv:2604.13151v1 Announce Type: new Abstract: Language Model (LM) agents are increasingly used in complex open-ended decision-making tasks, from AI coding to physical AI. A core requirement in these

safetyarxiv-cs-ai
17 Apr 2026
Tools

Is there still a widespread belief that LLMs and coding agents are good for greenfield development but don't help for maintaining large exis…

DGX agent

Simon Willison discusses the perception that LLMs and coding agents are primarily useful for greenfield development (starting new projects from scratch) rather than for maintaining and modifying large

toolssimon-willison--x
17 Apr 2026
Agents

One portal, unlimited possibilities. You can now access Modal via Tool Gateway by @NousResearch, makers of Hermes Agent. Check it out 👇

DGX agent

One portal, unlimited possibilities. You can now access Modal via Tool Gateway by @NousResearch, makers of Hermes Agent. Check it out 👇 Tool Gateway is now live in Nous Portal. No separate accounts, n

agentsnous-research--x
17 Apr 2026
Safety

ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration

DGX agent

arXiv:2509.21823v2 Announce Type: replace Abstract: Reward is critical to the evaluation and training of large language models (LLMs). However, existing rule-based or model-based reward methods strugg

safetyarxiv-cs-ai
17 Apr 2026
Agents

Towards a Multi-Embodied Grasping Agent

DGX agent

arXiv:2510.27420v3 Announce Type: replace Abstract: Multi-embodiment grasping focuses on developing approaches that exhibit generalist behavior across diverse gripper designs. Existing methods often l

agentsarxiv-cs-ro
17 Apr 2026
Industry

AI Search: the search primitive for your agents

DGX agent

AI Search is the search primitive for your agents. Create instances dynamically, upload files, and search across instances with hybrid retrieval and relevance boosting. Just create a search instance,

industrycloudflare-ai
16 Apr 2026
Model Releases

HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark

DGX agent

arXiv:2604.13954v1 Announce Type: new Abstract: Existing agent-safety evaluation has focused mainly on externally induced risks. Yet agents may still enter unsafe trajectories under benign conditions.

model-releasesarxiv-cs-lg
16 Apr 2026
Model Releases

LiteParse should be the default document parser you use with any AI agent (Claude Code, Claude Cowork, OpenClaw, Codex, and more) The core i…

DGX agent

LiteParse should be the default document parser you use with any AI agent (Claude Code, Claude Cowork, OpenClaw, Codex, and more) The core is extremely fast text and accurate parsing from any document

model-releasesjerry-liu--x
16 Apr 2026
Model Releases

A Benchmark for Evaluating Outcome-Driven Constraint Violations in Autonomous AI Agents

DGX agent

arXiv:2512.20798v4 Announce Type: replace Abstract: As autonomous AI agents are deployed in high-stakes environments, ensuring their safety has become a paramount concern. Existing safety benchmarks p

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search

DGX agent

arXiv:2604.12762v1 Announce Type: cross Abstract: We introduce ARGOS, the first benchmark and framework that reformulates multi-camera person search as an interactive reasoning problem requiring an ag

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

From Plan to Action: How Well Do Agents Follow the Plan?

DGX agent

arXiv:2604.12147v1 Announce Type: cross Abstract: Agents aspire to eliminate the need for task-specific prompt crafting through autonomous reason-act-observe loops. Still, they are commonly instructed

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

I'm going all in on Hermes (@NousResearch, @Teknium1) as my entire agent and coding stack. Six profiles. One shared self-hosted memory store…

DGX agent

I'm going all in on Hermes (@NousResearch, @Teknium1) as my entire agent and coding stack. Six profiles. One shared self-hosted memory store. Zero hosted-coder dependencies. The fleet: - pmax-mousa —

model-releasesnous-research--x
15 Apr 2026
Local Ai

Ollama Open-Source Agent Self-Reflection Harness

DGX agent

An open-source agent self-reflection harness built on top of Ollama, shared in the r/ollama community, that enables locally run LLMs to evaluate and iteratively refine their own outputs. The project p

local-air-ollama
15 Apr 2026
Agents

Towards Robust Real-World Spreadsheet Understanding with Multi-Agent Multi-Format Reasoning

DGX agent

arXiv:2604.12282v1 Announce Type: new Abstract: Spreadsheets are central to real-world applications such as enterprise reporting, auditing, and scientific data management. Despite their ubiquity, exis

agentsarxiv-cs-cl
15 Apr 2026
Model Releases

Competing with AI Scientists: Agent-Driven Approach to Astrophysics Research

DGX agent

arXiv:2604.09621v1 Announce Type: new Abstract: We present an agent-driven approach to the construction of parameter inference pipelines for scientific data analysis. Our method leverages a multi-agen

model-releasesarxiv-cs-ai
14 Apr 2026
Agents

Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning

DGX agent

arXiv:2604.09813v1 Announce Type: new Abstract: Existing synthetic tool-use corpora are primarily designed for offline supervised fine-tuning, yet reinforcement learning (RL) requires executable envir

agentsarxiv-cs-ai
14 Apr 2026
Agents

DarwinNet: An Evolutionary Network Architecture for Agent-Driven Protocol Synthesis

DGX agent

arXiv:2604.01236v2 Announce Type: replace-cross Abstract: Traditional network architectures suffer from severe protocol ossification and structural fragility due to their reliance on static, human-def

agentsarxiv-cs-ai
14 Apr 2026
Model Releases

Do We Still Need GraphRAG? Benchmarking RAG and GraphRAG for Agentic Search Systems

DGX agent

arXiv:2604.09666v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) and its graph-based extensions (GraphRAG) are effective paradigms for improving large language model (LLM) reason

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

EvoDiagram: Agentic Editable Diagram Creation via Design Expertise Evolution

DGX agent

arXiv:2604.09568v1 Announce Type: cross Abstract: High-fidelity diagram creation requires the complex orchestration of semantic topology, visual styling, and spatial layout, posing a significant chall

model-releasesarxiv-cs-cl
14 Apr 2026
Local Ai

MAFIG: Multi-agent Driven Formal Instruction Generation Framework

DGX agent

arXiv:2604.10989v1 Announce Type: new Abstract: Emergency situations in scheduling systems often trigger local functional failures that undermine system stability and even cause system collapse. Exist

local-aiarxiv-cs-ai
14 Apr 2026
Model Releases

Most AI assistants wait for you to ask. But a truly useful agent should notice you need help before you say anything. New research takes a s…

DGX agent

Most AI assistants wait for you to ask. But a truly useful agent should notice you need help before you say anything. New research takes a serious shot at building proactive agents that work in real t

model-releasesdair-ai--x
14 Apr 2026
Model Releases

The Amazing Agent Race: Strong Tool Users, Weak Navigators

DGX agent

arXiv:2604.10261v1 Announce Type: new Abstract: Existing tool-use benchmarks for LLM agents are overwhelmingly linear: our analysis of six benchmarks shows 55 to 100% of instances are simple chains of

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

Three Roles, One Model: Role Orchestration at Inference Time to Close the Performance Gap Between Small and Large Agents

DGX agent

arXiv:2604.11465v1 Announce Type: new Abstract: Large language model (LLM) agents show promise on realistic tool-use tasks, but deploying capable agents on modest hardware remains challenging. We stud

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments

DGX agent

arXiv:2506.02387v3 Announce Type: replace Abstract: Recent advancements in Vision Language Models (VLMs) have expanded their capabilities to interactive agent tasks, yet existing benchmarks remain lim

model-releasesarxiv-cs-ai
14 Apr 2026
Hardware

We've been developing a multi-agent system that builds and maintains complex software autonomously. Recently, we partnered with NVIDIA to ap…

DGX agent

We've been developing a multi-agent system that builds and maintains complex software autonomously. Recently, we partnered with NVIDIA to apply it to optimizing CUDA kernels. In 3 weeks, it delivered

hardwarecursor--x
14 Apr 2026
Model Releases

Agentic Jackal: Live Execution and Semantic Value Grounding for Text-to-JQL

DGX agent

arXiv:2604.09470v1 Announce Type: new Abstract: Translating natural language into Jira Query Language (JQL) requires resolving ambiguous field references, instance-specific categorical values, and com

model-releasesarxiv-cs-cl
13 Apr 2026
Agents

a big part of agent harnesses is how they interact with context memory is just context its therefor impossible to separate harness from memo…

DGX agent

a big part of agent harnesses is how they interact with context memory is just context its therefor impossible to separate harness from memory - as @sarahwooders says, 'memory isn't a plugin (it's a h

agentsharrison-chase--x
11 Apr 2026
Model Releases

LiteParse is the best document parsing library for coding agents. It's free, fast, integrates natively with the LLM's native visual understa…

DGX agent

LiteParse is the best document parsing library for coding agents. It's free, fast, integrates natively with the LLM's native visual understanding capabilities, and comes with support for 50+ formats a

model-releasesjerry-liu--x
10 Apr 2026
Industry

Memory Scaling for AI Agents

DGX agent

Databricks Research introduced **MemAlign**, a memory framework for AI agents that stores past interactions as episodic memories and uses an LLM to distill them into generalized semantic rules, whi...

industrydatabricks
10 Apr 2026
Agents

Mina: A Multilingual LLM-Powered Legal Assistant Agent for Bangladesh for Empowering Access to Justice

DGX agent

arXiv:2511.08605v3 Announce Type: replace Abstract: Bangladesh's low-income population faces major barriers to affordable legal advice due to complex legal language, procedural opacity, and high costs

agentsarxiv-cs-cl
10 Apr 2026
Agents

RemoteAgent: Bridging Vague Human Intents and Earth Observation with RL-based Agentic MLLMs

DGX agent

arXiv:2604.07765v1 Announce Type: new Abstract: Earth Observation (EO) systems are essentially designed to support domain experts who often express their requirements through vague natural language ra

agentsarxiv-cs-cv
10 Apr 2026
Agents

GLM-5.1 gives teams a stronger model for coding, tool use, and sustained agent performance on Together AI. Learn more: http://www.together.a…

DGX agent

GLM-5.1 is Z.ai's post-training upgrade to GLM-5, now available on Together AI, delivering a 28% coding performance improvement through a refined reinforcement learning pipeline while retaining the...

agentstogether-ai--x
8 Apr 2026
Model Releases

How can you improve your agentic search pipeline? I just wrote a blog post with @tech_optimist from @lancedb to answer exactly that. TLDR: -…

DGX agent

How can you improve your agentic search pipeline? I just wrote a blog post with @tech_optimist from @lancedb to answer exactly that. TLDR: - Parse files and take page-level screenshots with LiteParse,

model-releasesjerry-liu--x
7 Apr 2026
Safety

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents

DGX agent

arXiv:2608.12764v1 Announce Type: cross Abstract: Deep search agents operate over trajectories spanning dozens of steps, yet standard reinforcement learning provides only a single outcome reward per t

safetyarxiv-cs-ai
14 Aug 2026
Local Ai

Capability Sheaves for Compositional Agent-Harness Repair: Controlled Quotients and a Real-Repository Stress Test

DGX agent

arXiv:2608.13228v1 Announce Type: new Abstract: Agent harnesses combine retrieval, routing, state, provenance, and verification, but locally successful components may disagree on shared state. We mode

local-aiarxiv-cs-ai
14 Aug 2026
Agents

Introducing GLM-5.3: Built to Code. Ready for Cyber Defense. - Top-tier coding and agentic capabilities, achieved through post-training on t…

DGX agent

Introducing GLM-5.3: Built to Code. Ready for Cyber Defense. - Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model - A major leap in cybersecurity, setting

agentssimon-willison--x
14 Aug 2026
Safety

ReflectFact: Self-Reflective Agents for Improving Comprehension and Reasoning in Multi-Hop Fact Verification

DGX agent

arXiv:2608.12877v1 Announce Type: new Abstract: Multi-hop fact verification, which verifies claims by reasoning over multiple pieces of evidence, is critical for combating misinformation on social med

safetyarxiv-cs-ai
14 Aug 2026
Agents

UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations

DGX agent

arXiv:2608.13031v1 Announce Type: cross Abstract: Traffic video understanding has become an important problem in intelligent transportation, as road videos provide direct evidence for accidents, viola

agentsarxiv-cs-ai
14 Aug 2026
Model Releases

Agent Safety Should Be a Runtime Contract

DGX agent

arXiv:2608.11274v1 Announce Type: cross Abstract: The dominant paradigm treats AI safety as a property to be instilled during model training via RLHF, DPO, or Constitutional AI. We argue this is struc

model-releasesarxiv-cs-ai
13 Aug 2026
← Previous
1…9495969798…374
Next →