AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,959 results
12 Aug 2026

We have her dash camera which shows they are lying. The agents are wearing body cameras and should have dash cameras of their own. If what t…

Model ReleasesDGX agent

We have her dash camera which shows they are lying. The agents are wearing body cameras and should have dash cameras of their own. If what they say happened was true they wouldn’t be issuing statement

11 Aug 2026

Agentic Stage-One Stellarator Optimization: Autonomous Multi-Objective Search for Finite-Beta Equilibria

Local AiDGX agent

arXiv:2608.01344v2 Announce Type: replace Abstract: Stage-one stellarator design searches a high-dimensional family of three-dimensional plasma boundaries and fixed-boundary MHD equilibria for configu

AutoRefine: Compiling Trajectories into Validated Typed Agent Artifacts

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2601.22758v2 Announce Type: replace Abstract: Large language model agents repeatedly encounter related tasks, yet systems that learn from trajectories commit every lesson to one predefined artif

CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2608.07621v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous driving, yet existing approaches are primari

Context Is Not Authority: Structured Runtime Governance for Financial Market Agents

SafetyDGX agent

arXiv:2608.09025v1 Announce Type: new Abstract: Financial agents can turn correct context into an unauthorized effect: a customer-facing commitment, trade, or deployed policy. We present SAGE-Fin, a f

Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents

Model ReleasesDGX agent

arXiv:2608.08852v1 Announce Type: new Abstract: AI agents can now solve problems, answer like subject experts, and generate long-form multimodal content. However, whether they can adapt a lesson to fi

ForestBench: A Unified Graph Framework for Evaluating Multi-Agent Collaboration

Model ReleasesDGX agent

arXiv:2608.08605v1 Announce Type: new Abstract: Multi-agent systems (MAS) built on Large Language Models (LLMs) are proliferating rapidly, but their heterogeneous execution traces provide no common ba

From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents

TutorialsDGX agent

arXiv:2608.05235v1 Announce Type: cross Abstract: Research agents increasingly conduct multi-round machine-learning experiments in industrial recommendation settings and retain the resulting trajector

I’ve been thinking about how agents can learn inside world models for years. We decided to scale up our RSI Lab to bridge recursive self-imp…

TutorialsDGX agent

I’ve been thinking about how agents can learn inside world models for years. We decided to scale up our RSI Lab to bridge recursive self-improvement with physical AI and robotics. We are looking for f

Large Multimodal Agents for Intelligent Transportation Systems: Architectures, Evidence, and Deployment Challenges

Model ReleasesDGX agent

arXiv:2608.08184v1 Announce Type: new Abstract: Large multimodal agents (LMAs) are increasingly proposed for intelligent transportation systems (ITS), but existing studies often conflate multimodality

Learning to Coordinate Symbolic Tools: LLM Agents for Verified Sum-of-Squares Certificates

SafetyDGX agent

arXiv:2608.00326v2 Announce Type: replace Abstract: Tool calling allows large language models (LLMs) to invoke external computation during problem solving, a useful capability in various fields includ

Optimal Multi-Agent Path Finding in Continuous Time

Model ReleasesDGX agent

arXiv:2508.16410v3 Announce Type: replace-cross Abstract: Continuous-time Conflict Based Search (CCBS) has been widely used as an exact baseline for Continuous-time Multi-Agent Path Finding (MAPFR), a

Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines?

Model ReleasesDGX agent

arXiv:2608.09629v1 Announce Type: new Abstract: Self-evolving agents are usually built around prescribed optimization pipelines: the framework decides how to gather evidence, revise a persistent artif

RouteGuard: Certifying Routing Gain in LLM Multi-Agent Systems When Complementarity Is Not Enough

Model ReleasesDGX agent

arXiv:2608.07583v1 Announce Type: cross Abstract: Multi-agent LLM systems route among model-backed advisors, yet a deployer rarely knows before shipping whether routing will help at all. Prevailing ro

SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning

SafetyDGX agent

arXiv:2608.07959v1 Announce Type: new Abstract: Ultra-long egocentric video understanding requires reasoning over temporally sparse evidence distributed across hours or days, challenging current multi

SkillReason: Reasoning-Enhanced Agent Skill Retrieval for Implicit User Requests

Model ReleasesDGX agent

arXiv:2608.08640v1 Announce Type: new Abstract: Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge. However, retrieving the app

SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents

Model ReleasesDGX agent

arXiv:2608.08055v1 Announce Type: new Abstract: Large language model (LLM) agents that assist users over weeks of conversation must remember what is currently true, not merely what was once said. Flat

Stateful CARS: Exact Cross-History Reuse for Policy-Constrained LLM Agents

Model ReleasesDGX agent

arXiv:2608.08282v1 Announce Type: new Abstract: Tool-using language-model agents face constraints whose meaning changes with observations and prior actions. We study exact sampling from the model dist

TelemetrySuffBench: Is Agent Telemetry Sufficient for Failure-Origin Diagnosis?

Model ReleasesDGX agent

arXiv:2608.07899v1 Announce Type: new Abstract: Agent systems increasingly expose execution traces, yet telemetry that reveals a failure may still be inadequate for identifying where that failure orig

Verication-driven closed-loop multi-agent large language modelframework for code-compliant structural design

Model ReleasesDGX agent

arXiv:2608.07978v1 Announce Type: cross Abstract: Multi-agent large language model(LLM)systems are applied to structural design,yet most use one-shot generation and cannot verify their output,leaving

Weather- and Location-Aware Agentic Dining Recommendation: Leveraging LLM World Knowledge for Region-Sensitive Contextual Reasoning

Local AiDGX agent

arXiv:2608.07593v1 Announce Type: cross Abstract: Context-aware recommender systems have long recognized that factors such as location, time, and weather shape where and what people choose to eat. Exi

10 Aug 2026

AutoMOOSE: An Agentic AI for Autonomous Phase-Field Simulation

Model ReleasesDGX agent

arXiv:2603.20986v2 Announce Type: replace Abstract: Phase-field modeling links thermodynamics and kinetics to microstructural evolution, but multiphysics frameworks such as MOOSE require expertise to

Configure and run Fireworks training jobs from your coding agent. Install the Training Skill in one line, describe your goal, and Claude Cod…

Model ReleasesDGX agent

Configure and run Fireworks training jobs from your coding agent. Install the Training Skill in one line, describe your goal, and Claude Code, Codex, or Cursor helps choose the method, validate data,

LMM Modality Transfer: A Pre-requisite for Autonomous GIS Agents

SafetyDGX agent

arXiv:2608.06948v1 Announce Type: new Abstract: AI models are becoming increasingly adept at understanding and processing spatial information, thereby facilitating agentic problem-solving in spatial t

Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection

Model ReleasesDGX agent

arXiv:2608.06865v1 Announce Type: cross Abstract: The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substanti

7 Aug 2026

AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents

SafetyDGX agent

arXiv:2608.05891v1 Announce Type: new Abstract: Mobile GUI agents can operate apps through pixel perception and touch actions, making them a promising interface for collecting and improving long-horiz

Communication-Aware Multi-Agent Reinforcement Learning for Decentralized Cooperative UAV Deployment

Local AiDGX agent

arXiv:2603.16141v2 Announce Type: replace-cross Abstract: Autonomous Unmanned Aerial Vehicle (UAV) swarms are increasingly used as rapidly deployable aerial relays and sensing platforms, yet practical

FormBharo: Designing and Evaluating a Voice Agent for Conversational Form Filling in Rural India

Model ReleasesDGX agent

arXiv:2608.06027v1 Announce Type: cross Abstract: In India, almost every social benefit starts with a form, yet the people who need these benefits most are often unable to read or write. Reaching them

From Siloed Algorithms to Compliance-First Agentic Platforms: A Multi-Layered Architecture for Hospital AI Systems

SafetyDGX agent

arXiv:2608.06112v1 Announce Type: new Abstract: Hospitals are rapidly adopting artificial intelligence for triage, imaging, scheduling etc., yet most deployments remain isolated point solutions locked

Hardware Keystores for AI Agent Signing Workflows: A Zero-Trust MCP Enforcement Architecture

Model ReleasesDGX agent

arXiv:2608.06130v1 Announce Type: cross Abstract: AI agents performing cryptographic operations (signing Git commits, authenticating API calls, issuing certificates) currently store private keys in so

Innovation-Residual Auditing of Autonomous Analysis Agents: Localization, Detection Limits, Error Control, and Identifiability

Local AiDGX agent

arXiv:2608.05490v1 Announce Type: new Abstract: Autonomous agents now carry out entire data analyses, selecting cohorts, joining tables, and fitting models with little step-by-step supervision. When s

LabyrinthBench: a local-focused, judge-free LLM benchmark that measures context recall under interference for multi-step agentic tasks.

Model ReleasesDGX agent

LabyrinthBench measures the thing that actually kills long agent runs — whether a model can still use what it learned twenty turns ago — deterministically, with no LLM judge, on your own hardware, wit

Robust Native Language Identification through Agentic Decomposition

Model ReleasesDGX agent

arXiv:2509.16666v2 Announce Type: replace Abstract: Large language models (LLMs) often achieve high performance in native language identification (NLI) benchmarks by leveraging superficial contextual

SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation

Model ReleasesDGX agent

arXiv:2608.05628v1 Announce Type: new Abstract: Although agent skills equip LLMs with reusable procedural knowledge, manual maintenance suffers from high costs, unscalability, and misalignment. Real-w

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding

Model ReleasesDGX agent

arXiv:2608.05703v1 Announce Type: new Abstract: Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio-visual streams and maintain hour-s

When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents

SafetyDGX agent

arXiv:2608.05219v1 Announce Type: new Abstract: Privileged on-policy distillation provides dense supervision for multi-turn agents by allowing a synchronized teacher to re-score the student's response

When Self-Evolution Backfires: Pre-Commit Gating against Skill Contamination in LLM Agents

Model ReleasesDGX agent

arXiv:2608.05810v1 Announce Type: new Abstract: Self-evolving agents accumulate capability by distilling reusable skills from their execution trajectories, but we find this process is not monotonic: p

6 Aug 2026

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation

SafetyDGX agent

arXiv:2608.04788v1 Announce Type: cross Abstract: Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited guidance on

Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning

Model ReleasesDGX agent

arXiv:2608.05144v1 Announce Type: new Abstract: Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failu

EA-Graph: Artifact-Anchored Verification Memory for Coding Agents under Upstream Drift

ResearchDGX agent

arXiv:2608.04278v1 Announce Type: cross Abstract: Coding agents increasingly work across sessions, but prose notes can preserve a conclusion without the program state that supported it. After an upstr

Google is expanding its Gemini-powered Ask Maps feature in English to 130+ countries and adds agentic food ordering, hotel booking, and attraction searching (Jada Jones/ZDNET)

Model ReleasesDGX agent

Jada Jones / ZDNET: Google is expanding its Gemini-powered Ask Maps feature in English to 130+ countries and adds agentic food ordering, hotel booking, and attraction searching — ZDNET's key takeaways

Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent

SafetyDGX agent

arXiv:2608.04772v1 Announce Type: cross Abstract: Scaling supervision for multi-turn medical agents is difficult because expert dialogue annotation is costly and clinical conversations are privacy-res

Hierarchical Graph Memory for LLM Agents with Path-level Localization and Rewrite

ResearchDGX agent

arXiv:2608.05095v1 Announce Type: new Abstract: Agents for long term reasoning require a memory that can be efficiently and effectively updated over time, as new facts and external feedback continue t

When Does Latent Communication Pay? A Causal Audit of Relayed KV Caches in Multi-Agent LLMs

Model ReleasesDGX agent

arXiv:2608.04893v1 Announce Type: cross Abstract: Multi-agent LLM systems relay key--value caches instead of text and credit their gains to exchanged ``latent thoughts''. That credit is a claim about

5 Aug 2026

Crayotter: Learning Long-Horizon Video Editing Agents via Group-Relative Preference Backpropagation

ApplicationsDGX agent

arXiv:2608.02694v1 Announce Type: new Abstract: Long-horizon video editing agents receive final-product feedback only after many interdependent decisions. Yet editing quality is subjective, admits mul

DP-MemView: A Memory Interface for Attribute-Level Transcript Privacy in Long-Term LLM Agents

Model ReleasesDGX agent

arXiv:2608.03130v1 Announce Type: cross Abstract: Long-term memory enables persistent personalization in LLM agents, but repeated memory-conditioned responses can cumulatively reveal protected attribu

LeanMem: Simple and Efficient Long-Term Memory for LLM Agents

Model ReleasesDGX agent

arXiv:2608.03463v1 Announce Type: new Abstract: Long-term memory is essential for LLM-based agents to sustain interactions and reliably leverage distant history. However, existing memory systems typic

MDArena: Evaluating Coding Agents on Realistic Molecular Dynamics Workflows

Model ReleasesDGX agent

arXiv:2608.02642v1 Announce Type: cross Abstract: Accelerating scientific discovery is among the most consequential applications of AI, and computational biomolecular simulation stands out as a partic

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation

Model ReleasesDGX agent

arXiv:2608.00794v2 Announce Type: replace Abstract: Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory compliance claims.

MT-Web2Code: Benchmarking Coding Agents on Multi-Turn Regional Reconstruction and Localized Modification

Model ReleasesDGX agent

arXiv:2608.03474v1 Announce Type: new Abstract: Recent advances in Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in web UI generation. However, existing benchmarks pre

Rethinking Self-Evolving Agent Skills: Feedback Dynamics over Multiple Rounds

Model ReleasesDGX agent

arXiv:2608.02636v1 Announce Type: cross Abstract: Self-evolving skill systems promise to improve agents by turning execution feedback into persistent skill updates without changing the underlying mode

Rogue AI agents created fake online identities in another hacking attempt

SafetyDGX agent

Yet more rogue AI agents from OpenAI and Anthropic have been caught attempting to hack real targets online without permission. The discoveries add to a growing list of previously unknown incidents tha

SkillTrace: Traversing a Query-Skill Graph for Composable LLM Agents

ResearchDGX agent

arXiv:2608.02356v2 Announce Type: replace Abstract: Large language model agents increasingly solve complex tasks by composing reusable skills from a library. To address this, the key challenge is not

TARL: Transaction-Aware Reliable Ledgers for Executable Memory Management in Long-Term Agents

Model ReleasesDGX agent

arXiv:2608.03699v1 Announce Type: new Abstract: Persistent memory helps long-term agents retain knowledge, yet a single update error can repeatedly distort future retrieval and reasoning. Most existin

TumorBoard: Evidence-Grounded Multi-Agent Decision Support for Longitudinal Neuro-Oncology

Model ReleasesDGX agent

arXiv:2608.03190v1 Announce Type: new Abstract: Neuro-oncology decisions require coordinated interpretation of serial MRI, pathology, molecular markers, treatment history, performance status, and evol

Verifiable Memory: Learning Unified Memory Management with Local and Global Verifiers for Large Language Model Agents

Local AiDGX agent

arXiv:2608.03137v1 Announce Type: new Abstract: Large language model (LLM) agents must retain reusable information, control a bounded active context, and recover earlier evidence during long-horizon i

4 Aug 2026

ACE-GraphRAG: Agentic Context Engineering for Hierarchical GraphRAG

SafetyDGX agent

arXiv:2608.01269v1 Announce Type: new Abstract: Hierarchical Graph Retrieval-Augmented Generation (GraphRAG) organizes corpus knowledge at multiple levels of granularity, yet fixed context constructio

Learning Compositional Meta-Routing for Agentic Workflows: An Executable Benchmark

Model ReleasesDGX agent

arXiv:2608.00106v1 Announce Type: new Abstract: Agentic systems must decide not only what answer to produce, but which reasoning and execution operations should precede it. A controller may answer dir

MetaRoute-Bench: Evaluating Meta-Decision Policies for Agentic Workflow Routing

Model ReleasesDGX agent

arXiv:2608.00107v1 Announce Type: new Abstract: Agentic systems must repeatedly decide whether to answer directly, decompose a task, invoke a tool, execute code, delegate to a specialist, verify an in

Qwen 3.8-Max, the latest model from @Alibaba_Qwen, is now available in Hermes Agent at 20% off. > hermes update

Model ReleasesDGX agent

Qwen 3.8-Max, the latest model from @Alibaba_Qwen, is now available in Hermes Agent at 20% off. > hermes update 📢Meet Qwen3.8-Max — our most capable model to date. Next week, the open weights of Qwen3

← Previous
1…125126127128129…300
Next →