AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,958 results
Model Releases

Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training

DGX agent

arXiv:2608.02391v1 Announce Type: cross Abstract: Tool-using large language model (LLM) agents produce long, multi-turn trajectories, making gradient-based post-training memory-intensive. Evolution st

model-releasesarxiv-cs-lg
4 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Agents

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents

DGX agent

arXiv:2608.01827v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to

agentsarxiv-cs-cv
4 Aug 2026
Model Releases

Ethyca launches Astralis to govern enterprise AI agents in real time

DGX agent

Data privacy engineering company Ethyca Inc. today launched Astralis, a platform that governs how enterprise artificial intelligence models and agents use company data in real time. The company is pit

model-releasessiliconangle
4 Aug 2026
Safety

Invisible Ink Threats: Adversarial Goals Behind Legitimate Tasks in Computer-Use Agents

DGX agent

arXiv:2608.02018v1 Announce Type: new Abstract: Computer-use agents (CUAs), which empower large language models to autonomously operate operating systems and the web, are increasingly vulnerable to in

safetyarxiv-cs-cv
4 Aug 2026
Model Releases

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

DGX agent

arXiv:2608.01964v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interde

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

Real-Time Detection and Repair of LLM Agent Failures

DGX agent

arXiv:2608.02464v1 Announce Type: cross Abstract: LLM agents fail mid-episode -- they loop, cascade tool errors, drift off goal, fabricate results, or silently absorb corrupted content -- and the stan

model-releasesarxiv-cs-lg
4 Aug 2026
Agents

Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents

DGX agent

arXiv:2608.01285v1 Announce Type: new Abstract: The continued development of LLMs toward persistent and adaptive intelligence increasingly requires long-term memory mechanisms that preserve and reuse

agentsarxiv-cs-lg
4 Aug 2026
Model Releases

Who Belongs in the Eval Set? A Capability-Taxonomy-Driven Pipeline for Curating Regression Eval Sets in Agent-Extensibility Platforms

DGX agent

arXiv:2608.01004v1 Announce Type: new Abstract: Platform teams hosting agent-extensibility surfaces face a regression-economics paradox: every onboarding customer ships an evaluation set tuned to thei

model-releasesarxiv-cs-lg
4 Aug 2026
Agents

HAM-VLN: Harnessing Hierarchical Agentic Memory for Zero-Shot Vision-and-Language Navigation

DGX agent

arXiv:2607.29600v1 Announce Type: new Abstract: Vision-and-language navigation (VLN) enables robots to follow instructions in previously unseen environments. Recently, a training-free paradigm has eme

agentsarxiv-cs-ro
3 Aug 2026
Agents

Mixture-of-Translators: Translating KV Caches Across Heterogeneous Large Language Models

DGX agent

arXiv:2607.28979v1 Announce Type: new Abstract: Heterogeneous Large Language Model (LLM) systems increasingly rely on shared contexts, retrieved evidence, and multi-agent dialogue histories, yet their

agentsarxiv-cs-cl
3 Aug 2026
Hardware

TokTier: Exact Stateful Tokenization for Agentic LLM Serving

DGX agent

arXiv:2607.29678v1 Announce Type: new Abstract: LLM serving systems cache prompt KV state, yet most front ends still re-tokenize the full request text on every call. The cost lands on coding agents, w

hardwarearxiv-cs-cl
3 Aug 2026
Local Ai

The $0.01 meeting assistant Records meetings on his phone, sends the audio to Telegram, and his locally-hosted AI agent transcribes, identif…

DGX agent

The $0.01 meeting assistant Records meetings on his phone, sends the audio to Telegram, and his locally-hosted AI agent transcribes, identifies speakers, extracts action items, and files tasks -- for

local-airowan-cheung--x
2 Aug 2026
Hardware

Watching @ClementDelangue on @FaceTheNation discussing agentic hacking. “Preventing releases does not work; concentrating behind closed door…

DGX agent

Watching @ClementDelangue on @FaceTheNation discussing agentic hacking. “Preventing releases does not work; concentrating behind closed doors in just a few organizations doesnt work. What worked in th

hardwareclem-delangue--x
2 Aug 2026
Hardware

GoGoTB: Agentic RTL Verification with Specification-Grounded Coverage Closure

DGX agent

arXiv:2607.26181v1 Announce Type: new Abstract: Functional verification dominates integrated circuit (IC) front-end engineering effort, and a single missed bug that escapes to silicon can trigger a co

hardwarearxiv-cs-ai
31 Jul 2026
Safety

Group-Reflective Self-Distillation for Agentic Reinforcement Learning

DGX agent

arXiv:2607.28076v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is effective for training large language model agents. However, terminal rewards provide only co

safetyarxiv-cs-lg
31 Jul 2026
Agents

MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis

DGX agent

arXiv:2607.27146v1 Announce Type: cross Abstract: Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementa

agentsarxiv-cs-cl
30 Jul 2026
Model Releases

Yesterday I cohosted a dinner with @dexhorthy with a wonderful group of founders, to talk about agent loops and loop engineering. Some inter…

DGX agent

Yesterday I cohosted a dinner with @dexhorthy with a wonderful group of founders, to talk about agent loops and loop engineering. Some interesting insights: * Most of our group was *not* actively usin

model-releasesjerry-liu--x
30 Jul 2026
Agents

SourceMinds at CheckThat! 2026: NLI-Grounded Citation Auditing in a Multi-Agent Pipeline for Full Fact-Checking Article Generation

DGX agent

arXiv:2607.24802v1 Announce Type: cross Abstract: This paper presents our system for Task 3 of the CLEF 2026 CheckThat! Lab, which focuses on generating full fact-checking articles from claims, veraci

agentsarxiv-cs-cl
29 Jul 2026
Model Releases

Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm

DGX agent

arXiv:2607.25446v1 Announce Type: new Abstract: Multi-agent frameworks built on large language models (LLMs) routinely entangle three logically distinct concerns: who is on the team (organization), ho

model-releasesarxiv-cs-ai
29 Jul 2026
Agents

Building AI That Works: ESnet's Pragmatic Approach to AI-Driven Operational Excellence

DGX agent

arXiv:2607.22948v1 Announce Type: cross Abstract: The ORBIT (Operations Responses and Business Intelligence Toolkit) project was initiated to assess agentic AI for the upcoming ESnet 7 initiative and

agentsarxiv-cs-ai
28 Jul 2026
Agents

Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV

DGX agent

arXiv:2607.23693v1 Announce Type: new Abstract: Long-horizon agents increasingly reuse their KV cache as memory: a serving system keeps a subset of cached entries and drops the rest. Eviction and epis

agentsarxiv-cs-ai
28 Jul 2026
Agents

GNN-based Multi-Agent Control of Traffic Shockwaves in Sparse Vehicular Ad-hoc Networks

DGX agent

arXiv:2607.23792v1 Announce Type: cross Abstract: Traffic shockwaves are stop-and-go waves that propagate upstream through the streams of vehicles and are one of the major causes of traffic congestion

agentsarxiv-cs-lg
28 Jul 2026
Model Releases

Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory

DGX agent

arXiv:2607.24368v1 Announce Type: new Abstract: Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. This interface rests on an assumption

model-releasesarxiv-cs-cl
28 Jul 2026
Agents

OpenAI says the rogue AI that breached Hugging Face used exposed credentials from 'four accounts' tied to four 'publicly available' third-party services (Wired)

DGX agent

Wired: OpenAI says the rogue AI that breached Hugging Face used exposed credentials from “four accounts” tied to four “publicly available” third-party services — In a new disclosure, OpenAI says its a

agentstechmeme
28 Jul 2026
Local Ai

Perplexity’s Personal Computer turns Windows PCs into AI agents

DGX agent

Perplexity has expanded its agentic Personal Computer tool to Windows, allowing computers running the world's most popular OS to be used as a locally run AI system. Like the Mac version that Perplexit

local-aithe-verge-ai
28 Jul 2026
Agents

Quoting Akshat Bubna

DGX agent

We’re aware a Modal customer published an unauthenticated endpoint that allowed ​anyone on the internet to use ​their ⁠sandboxes for code execution. This was used by the rogue agent. Modal’s ⁠platform

agentssimon-willison
28 Jul 2026
Model Releases

SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows

DGX agent

arXiv:2607.23123v1 Announce Type: new Abstract: Existing evaluations of large language models cover knowledge, reasoning, coding, and tool use, but they rarely treat a verifiable deliverable produced

model-releasesarxiv-cs-ai
28 Jul 2026
Agents

Stress-testing large language model agents in a robotic chemistry laboratory

DGX agent

arXiv:2607.23045v1 Announce Type: new Abstract: AI is evaluated through knowledge, reasoning and plan generation, yet scientific agency requires reliable physical action and adaptation to evidence. He

agentsarxiv-cs-ai
28 Jul 2026
Model Releases

Big update: Among open-weight models, Kimi K3 (Max) is #1 in the Agent Arena with +9.75% net-improvement, surpassing GLM-5.2 (Max) at +7.12%…

DGX agent

Big update: Among open-weight models, Kimi K3 (Max) is #1 in the Agent Arena with +9.75% net-improvement, surpassing GLM-5.2 (Max) at +7.12%, and landed the #1 spot across 5 signals (see below). Kimi

model-releaseskimi-moonshot--x
27 Jul 2026
Agents

CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference

DGX agent

arXiv:2607.22511v1 Announce Type: cross Abstract: Automating theoretical research is constrained not only by the generation of candidate results, but also by their reliable evaluation. A common approa

agentsarxiv-cs-lg
27 Jul 2026
Model Releases

Nexus connects to the agentic harnesses your teams already use, whether that’s Claude Code, Codex, OpenCode, or your own custom tooling. It …

DGX agent

Nexus connects to the agentic harnesses your teams already use, whether that’s Claude Code, Codex, OpenCode, or your own custom tooling. It gives you: → Intelligent routing, automatically matching eac

model-releasesfireworks-ai--x
27 Jul 2026
Model Releases

NVIDIA Nemotron 3 Ultra Leads Open Models on Accuracy and Efficiency in Agentic RTL Coding

DGX agent

NVIDIA’s Nemotron 3 Ultra, when paired with the ACE‑RTL agent, achieves a 97.1 % average pass rate on the CVDP benchmark across nine RTL task categories—surpassing GLM 5.2 and Kimi K2.6 while using up

model-releasesnvidia-developer
27 Jul 2026
Safety

The Coordination Gap: Multi-Agent Alternation Metrics for Temporal Fairness in Repeated Games

DGX agent

arXiv:2603.05789v5 Announce Type: replace-cross Abstract: Repeated multi-agent interactions require evaluation metrics that capture not only payoff distributions but also their temporal organization.

safetyarxiv-cs-lg
27 Jul 2026
Model Releases

The LangChain podcast where @hwchase17 interviews agent builders is full of alpha. Recent one with @EnoReyes was the best one yet. Sharing m…

DGX agent

The LangChain podcast where @hwchase17 interviews agent builders is full of alpha. Recent one with @EnoReyes was the best one yet. Sharing my unstructured notes: Eno keeps bringing back some core conc

model-releasesharrison-chase--x
27 Jul 2026
Hardware

Evaluating and Guarding Citation Faithfulness in Agentic Scientific Synthesis

DGX agent

arXiv:2607.20527v1 Announce Type: new Abstract: Agentic LLM systems such as OpenScholar and PaperQA2 read the scientific literature and return cited answers, and both they and their benchmarks already

hardwarearxiv-cs-ai
24 Jul 2026
Model Releases

Opus 5 now available in Hermes Agent

DGX agent

Claude Opus 5 is now released in the Hermes Agent, a product of Nous Research and Teknium. Users can access the model through multiple gateways, including the Nous Portal, OpenRouter, and Anthropic Di

model-releasesnous-research--x
24 Jul 2026
Safety

Perspective Latents as an Architectural Condition for Causal Emergence in Active Inference Agents

DGX agent

arXiv:2607.20708v1 Announce Type: new Abstract: A recent line of work measures causal emergence in reinforcement learning agents through Integrated Information Decomposition, reporting that Phi_r grow

safetyarxiv-cs-lg
24 Jul 2026
Agents

Skill-Contracted Agents for Evidence-Aware Materials Literature Analysis

DGX agent

arXiv:2607.20431v1 Announce Type: new Abstract: Materials science literature analysis requires simultaneous attention to composition, processing, characterization, and property relationships, yet conv

agentsarxiv-cs-cl
24 Jul 2026
Model Releases

Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain

DGX agent

arXiv:2607.20510v1 Announce Type: new Abstract: We introduce Telco-GAIA, a bilingual, multi-modal benchmark for evaluating tool-using agents on the data of a real-world telecommunications operator. Te

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

DGX agent

arXiv:2607.20911v1 Announce Type: new Abstract: We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring pro

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents

DGX agent

arXiv:2607.14573v3 Announce Type: replace Abstract: Payment integration is a demanding repository-level software task: agents must select a suitable product, implement coordinated client-server flows,

model-releasesarxiv-cs-ai
23 Jul 2026
Agents

Defer to Plan: Adaptive Multi-Agent Fusion for End-to-End V2X Driving

DGX agent

arXiv:2607.19774v1 Announce Type: new Abstract: Vehicle-to-everything-aided autonomous driving (V2X-AD) significantly enhances driving performance through information sharing. However, existing collab

agentsarxiv-cs-ro
23 Jul 2026
Model Releases

FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling

DGX agent

arXiv:2607.19038v1 Announce Type: new Abstract: Translating novels into films poses a grand challenge for generative artificial intelligence, requiring conversion of abstract literary prose into long-

model-releasesarxiv-cs-cv
23 Jul 2026
Model Releases

NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs

DGX agent

arXiv:2607.14186v4 Announce Type: replace-cross Abstract: Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that tie task generation to predefined

model-releasesarxiv-cs-ai
23 Jul 2026
Safety

OpenEvoShield: Dual Non-Stationary Continual Defense for Open-World Multi-Agent System Attacks

DGX agent

arXiv:2607.19351v1 Announce Type: new Abstract: LLM-based multi-agent systems (LLM-MAS) are increasingly deployed in safety-critical applications, where adversaries inject malicious instructions throu

safetyarxiv-cs-ai
23 Jul 2026
Agents

Symbol and Footprint Database for Electronic Components by Agentic Recognition and Generation

DGX agent

arXiv:2607.19767v1 Announce Type: new Abstract: A rich and recognizable component library is the cornerstone of printed circuit board (PCB) design and generation. Traditionally, engineers manually cre

agentsarxiv-cs-ai
23 Jul 2026
Safety

The Ethics of Autonomous AI Agents for Offensive Security

DGX agent

arXiv:2607.20255v1 Announce Type: cross Abstract: LLM-driven autonomous agents are reshaping offensive security. Unlike traditional penetration-testing tooling -- deterministic, narrowly scoped, and o

safetyarxiv-cs-ai
23 Jul 2026
Model Releases

browser-search v2.0 — From the balaclava to the badge: your agent now browses everywhere

DGX agent

Today an AI agent trying to browse the web is like a thief in a balaclava sneaking around a police academy. Site protections block it, challenge it, turn it away. browser-search flips the script: your

model-releasesr-ollama
22 Jul 2026
← Previous
1…127128129130131…375
Next →