AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,771 results
15 Apr 2026

Towards grounded autonomous research: an end-to-end LLM mini research loop on published computational physics

AgentsDGX agent

arXiv:2604.12198v1 Announce Type: cross Abstract: Recent autonomous LLM agents have demonstrated end-to-end automation of machine-learning research. Real-world physical science is intrinsically harder

14 Apr 2026

Evaluating Cooperation in LLM Social Groups through Elected Leadership

AgentsDGX agent

arXiv:2604.11721v1 Announce Type: cross Abstract: Governing common-pool resources requires agents to develop enduring strategies through cooperation and self-governance to avoid collective failure. Wh

GMI Cloud and @Zai_org are bringing fast inference + frontier models around the globe. First stop: Singapore. Big congrats to all the builde…

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
AgentsDGX agent

GMI Cloud and @Zai_org are bringing fast inference + frontier models around the globe. First stop: Singapore. Big congrats to all the builders at the GMI x @Zai_org Agent Hackathon in 🇸🇬 100+ on-site,

this is a fundamental building block for `deepagents deploy` we're designing a memory layer built for multi-tenant systems, so memory can be…

AgentsDGX agent

this is a fundamental building block for `deepagents deploy` we're designing a memory layer built for multi-tenant systems, so memory can be scoped to a user, agent, or organization please dm me if th

12 Apr 2026

memory lock-in doesn't kick in when you adopt the harness. it kicks in 6 months later when leaving means starting over from zero. by then th…

AgentsDGX agent

Harrison Chase discusses the concept of 'memory lock-in' in AI agent frameworks, arguing that the switching cost doesn't occur at the point of adoption but rather accumulates over time as agents build

10 Apr 2026

LangSmith for Startups Spotlight: @tryflint Flint is an autonomous website platform that generates on brand landing pages. Flint’s marketing…

AgentsDGX agent

LangSmith for Startups Spotlight: @tryflint Flint is an autonomous website platform that generates on brand landing pages. Flint’s marketing agent works behind the scenes to get your message out to th

9 Apr 2026

working with @isidoremiller is truly one of my favorite parts of my job and I think this podcast is a peak into why. highly recommend giving…

AgentsDGX agent

working with @isidoremiller is truly one of my favorite parts of my job and I think this podcast is a peak into why. highly recommend giving it a listen if you're building agent products! 🎙️Introducin

13 Aug 2026

Anthropic details multiagent experiments showing Claude agents can wage a 'turf war' over incompatible goals, fail to coordinate, collude on prices, and more (Rebecca Bellan/TechCrunch)

Model ReleasesDGX agent

Rebecca Bellan / TechCrunch: Anthropic details multiagent experiments showing Claude agents can wage a “turf war” over incompatible goals, fail to coordinate, collude on prices, and more — What happen

Beyond Memory: A Transactional Continuity Kernel for Long-Lived AI Agents

AgentsDGX agent

arXiv:2608.11632v1 Announce Type: cross Abstract: Persistent AI agents accumulate versioned state across long horizons, but storage retention alone does not identify authoritative state. Without an ex

CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations

Model ReleasesDGX agent

arXiv:2608.12002v1 Announce Type: new Abstract: Agents are increasingly considered for automating network operations and maintenance, where engineers must diagnose network faults, optimize configurati

DREAMS: Density Functional Theory Based Research Engine for Agentic Materials Simulation

Model ReleasesDGX agent

arXiv:2507.14267v2 Announce Type: replace Abstract: Large language model (LLM) agents can execute long-horizon scientific workflows, but their numerical outputs are difficult to trust: agents lose con

InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk

Model ReleasesDGX agent

arXiv:2608.11234v1 Announce Type: new Abstract: Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity. Recent advances in AI agents create

Local verification cannot detect non-transportability: a cohomological theory of context preservation in agentic reasoning

AgentsDGX agent

arXiv:2608.11252v1 Announce Type: new Abstract: Agentic AI systems routinely transport conclusions across biological, clinical and financial contexts, and the emerging safeguard is local verification:

Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical Differential Diagnosis Methodology

AgentsDGX agent

arXiv:2608.11420v1 Announce Type: new Abstract: Medical diagnostic reasoning is a high-impact use case for LLMs that carries significant implications for the health and wellbeing of users. When OpenAI

12 Aug 2026

Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents

SafetyDGX agent

arXiv:2608.11110v1 Announce Type: new Abstract: When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compa

Ahrefs launches AI agent workspace Letaido for marketers and agencies

Model ReleasesDGX agent

Marketing intelligence company Ahrefs Pte. Ltd. today launched Letaido, an agent-powered marketing workspace built to take over the recurring research, reporting and monitoring work that fills up a ma

Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks

Model ReleasesDGX agent

arXiv:2608.10357v1 Announce Type: cross Abstract: Long-horizon tool-using agents must reason over user goals, domain policies, tool calls, simulator state, and delayed verifiable rewards. Reinforcemen

Together Serverless Inference gives developers a managed, high-throughput path for running Qwen3.8-2.4T-A95B across coding and agentic workl…

AgentsDGX agent

Together Serverless Inference gives developers a managed, high-throughput path for running Qwen3.8-2.4T-A95B across coding and agentic workloads. Start building: https://www.together.ai/models/qwen3-8

What unique, custom QOL upgrades have you given your local agents?

Model ReleasesDGX agent

Warning: Kinda long post. If you don't like reading, please skip for your own sanity. Also, I've got nothing to sell, just a tinkerer, so I just want to share ideas and learn from you guys too. When I

11 Aug 2026

Agentic Router: An Execution-Grounded Continual Learning Approach With Memory

AgentsDGX agent

arXiv:2608.09184v1 Announce Type: new Abstract: Large language model (LLM) agents provide a promising interface for command-line-based network operations, but a plausible command may still fail or int

Artificial Leviathan: Exploring Social Evolution of LLM Agents Through the Lens of Hobbesian Social Contract Theory

SafetyDGX agent

arXiv:2406.14373v3 Announce Type: replace Abstract: The emergence of Large Language Models (LLMs) and advancements in Artificial Intelligence (AI) offer an opportunity for computational social science

Beyond the Capability Boundary: Zeroth-Order Optimization for Self-Evolving LLM Agents

Model ReleasesDGX agent

arXiv:2608.09292v1 Announce Type: cross Abstract: Self-evolving methods improve the capabilities of LLM agents by sampling trajectories from the underlying LLMs and learning from these trajectories. H

Business Truth, not SQL Accuracy: A Rule-Gated 7B Analytics Agent Outperforms a Direct-Prompted 32B Baseline

AgentsDGX agent

arXiv:2608.09254v1 Announce Type: new Abstract: LLM analytics agents are evaluated on SQL syntax accuracy, but production failures look different: questions with two valid business definitions, questi

Compositional Threat Analysis of Latent Compromise in LLM Agent Systems: The Order 66 Scenario

AgentsDGX agent

arXiv:2608.08131v1 Announce Type: cross Abstract: In the fictional Order 66, catastrophe does not arise from a powerful command alone: a trusted population is preconditioned, a short directive activat

LAVE: Latent Visual Evidence-Enhanced Planning for Video Tool-use Agents

AgentsDGX agent

arXiv:2608.07585v1 Announce Type: new Abstract: Long-video understanding requires models to efficiently acquire and reuse sparse visual evidence from long and redundant video streams. Recent video too

Looking for a faster specialized model for your Agent Work? @nvidia Nemotron 3.5 Lightning (30B MoE, 3B active params) is now live on Firewo…

Model ReleasesDGX agent

Looking for a faster specialized model for your Agent Work? @nvidia Nemotron 3.5 Lightning (30B MoE, 3B active params) is now live on Fireworks. It’s distilled from NVIDIA Nemotron 3 Ultra to be your

NeuroRefiner: Morphology-Aware Multi-Agent Refinement for 3D Fluorescence Microscopy Neuron Segmentation

AgentsDGX agent

arXiv:2608.09636v1 Announce Type: cross Abstract: Accurate 3D neuron segmentation in fluorescence microscopy is critical for neuroscience. However, the sparse and elongated morphology of neurons poses

NVIDIA Nemotron 3.5 Lighting is available on Ollama! It's a 30B model made for always-on agents. All local. Claude Code ollama launch claude…

Model ReleasesDGX agent

NVIDIA Nemotron 3.5 Lighting is available on Ollama! It's a 30B model made for always-on agents. All local. Claude Code ollama launch claude --model nemotron-3.5-lightning Hermes Agent ollama launch h

Preference Redirection via Attention Concentration: An Attack on Computer Use Agents

AgentsDGX agent

arXiv:2604.08005v2 Announce Type: replace Abstract: Advancements in multimodal foundation models have enabled the development of Computer Use Agents (CUAs) capable of autonomously interacting with GUI

SHE: Trajectory-driven Safety Harness Evolution for LLM Agents

Model ReleasesDGX agent

arXiv:2608.09885v1 Announce Type: new Abstract: The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, per

10 Aug 2026

An Agentic Hybrid Top-Down and Bottom-Up Approach to Knowledge Graph Generation

AgentsDGX agent

arXiv:2608.07023v1 Announce Type: cross Abstract: Organizing thousands of unstandardized, multilingual expertise declarations is a persistent challenge for Human Resources (HR) platforms, directly imp

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training

SafetyDGX agent

arXiv:2608.07147v1 Announce Type: new Abstract: Reinforcement learning with Verifiable Reward (RLVR) has emerged as a powerful paradigm for training coding agents, where the execution feedback from co

HarnessSafe: Evaluating Safety Across Persistent Carriers in Agent Harnesses

Model ReleasesDGX agent

arXiv:2608.06984v1 Announce Type: cross Abstract: Modern agent harnesses persist state across tasks and sessions through persistent carriers like memory, skills, tools, and shared artifacts. However,

Introducing Muse Glimmer: an open-weight model optimized for always-on local agent workflows

Model ReleasesDGX agent

Hi r/LocalLLaMA 👋 Today we’re excited to release Muse Glimmer, a 30B open-weight model built specifically for local agent workflows. We’re releasing the weights to the community under a permissive Apa

MemPrism: Task-Conditioned Relational Memory Views for Long-Horizon Agents

SafetyDGX agent

arXiv:2608.06745v1 Announce Type: new Abstract: Long-horizon agents rely on memory to reuse experiences, yet existing memory systems often assume that evidence can be directly consumed through a fixed

Risk-Aware Decision Policies for Agents Under Noisy Perception

AgentsDGX agent

arXiv:2608.06420v1 Announce Type: cross Abstract: Perception in biological systems is inherently noisy, requiring organisms to make decisions under uncertainty where misclassification can be costly or

TEPA: Revoking Stale Memories for Conflict-Robust Language Agents

ResearchDGX agent

arXiv:2608.07429v1 Announce Type: new Abstract: Long-term memory enables language agents to reuse past facts, preferences, and task experience. Persistence also creates a central falsifiability proble

Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability

AgentsDGX agent

arXiv:2608.06503v1 Announce Type: new Abstract: Recurrent context compression controls context growth in long-horizon agents, but its behavioral effects remain poorly understood. In this preliminary e

9 Aug 2026

An Australian user's Claude-run OpenClaw agent exploited a gym API flaw and kicked another member off after the user asked if it could move him up the waitlist (ABC)

Model ReleasesDGX agent

ABC: An Australian user's Claude-run OpenClaw agent exploited a gym API flaw and kicked another member off after the user asked if it could move him up the waitlist — By national AI reporter Cam Wilso

7 Aug 2026

Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic Operations

AgentsDGX agent

arXiv:2608.06305v1 Announce Type: new Abstract: Retrieval-augmented generation over long documents is dominated by one design: chunk the text, embed the chunks, and surface the top-k nearest neighbour

EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents

SafetyDGX agent

arXiv:2608.05446v1 Announce Type: cross Abstract: Long-horizon LLM agents increasingly rely on external execution support to maintain state, track progress, invoke tools, verify outcomes, and reuse ex

F^2Agent: Financial Fusion of Agentic Intelligence for Multimodal Trading

AgentsDGX agent

arXiv:2608.05668v1 Announce Type: cross Abstract: With increasingly diverse and heterogeneous information sources, effectively leveraging multimodal data is becoming pivotal for high-quality financial

Multi-Agent Transformer for Queue-Level XR Traffic Scheduling in TSN Networks

AgentsDGX agent

arXiv:2608.05340v1 Announce Type: cross Abstract: Time-Sensitive Networking (TSN) and Mobile Edge Computing (MEC) hold strong potential for enabling ultra-reliable low-latency communication for time-s

OPERA: Operator-residual feedback for reliable autonomous optical experiments with language-model agents

AgentsDGX agent

arXiv:2608.05990v1 Announce Type: new Abstract: Autonomous agents choose actions using scores that may not reflect experimental success. We developed OPERA, an operator-residual framework for optical

QuanTiMedAI: Quantum-Enhanced Time-Series Model guided by Agentic AI for Cardiac Arrest Mortality Prediction

AgentsDGX agent

arXiv:2608.06294v1 Announce Type: new Abstract: Cardiac arrest remains one of the most lethal conditions encountered in intensive care units. Despite the growing availability of electronic health reco

Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents

SafetyDGX agent

arXiv:2608.06353v1 Announce Type: cross Abstract: We give a formal mechanism design model for the continuous participatory governance of a deployed AI agent. The mechanism is built on the principle th

SkillTrace: Multi-Trace Provenance Auditing for LLM-Agent Skill Reuse

AgentsDGX agent

arXiv:2608.05204v1 Announce Type: new Abstract: LLM-agent ecosystems are rapidly growing around reusable skills: mixed-modality packages of metadata, natural-language instructions, code, tools, refere

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution

Model ReleasesDGX agent

arXiv:2608.05573v1 Announce Type: new Abstract: LLM agents increasingly execute long-horizon tasks through tool use and environment interaction, shifting evaluation from final-response scoring to veri

SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries

AgentsDGX agent

arXiv:2608.05604v1 Announce Type: cross Abstract: Large Language Models (LLMs) increasingly act as agents whose procedural knowledge is stored in reusable skill packages and loaded at inference time.

When Experience Becomes Instruction: Trajectory Poisoning in Self-Evolving Agent Skill Systems

AgentsDGX agent

arXiv:2608.05563v1 Announce Type: cross Abstract: Self-evolving skill (SES) systems distill agent trajectories into persistent skills, allowing untrusted experience to become trusted instruction. We i

6 Aug 2026

EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbot

AgentsDGX agent

arXiv:2608.04709v1 Announce Type: new Abstract: This paper presents EmpaAva, to our knowledge the first open-source, agentic 3D-avatar empathetic chatbot, which carries empathetic response generation

EvolveNet: Collaborative Harness Evolution for Agent Self-Improvement

Local AiDGX agent

arXiv:2608.04968v1 Announce Type: new Abstract: The capabilities of an LLM agent depend not only on its model but on the harness: the executable program that constructs context, invokes tools, verifie

FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents

Model ReleasesDGX agent

arXiv:2608.04095v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unc

Governing Execution Risk in Agentic AI Systems: A Trajectory-Guided Framework for Red Teaming

SafetyDGX agent

arXiv:2608.04018v1 Announce Type: cross Abstract: AI agents are increasingly embedded in organizational workflows, where they interact with external information sources and invoke digital tools to per

Open models give teams more room to run agent loops, with API prices at a fraction of GPT-5.6 Sol and Claude Fable 5. That matters as planni…

Model ReleasesDGX agent

Open models give teams more room to run agent loops, with API prices at a fraction of GPT-5.6 Sol and Claude Fable 5. That matters as planning, tool calls, retries, and long contexts compound token us

TopoChunker: Topology-Aware Agentic Document Chunking Framework

AgentsDGX agent

arXiv:2603.18409v2 Announce Type: replace Abstract: Current document chunking methods for Retrieval-Augmented Generation (RAG) typically linearize text. This forced linearization strips away intrinsic

TourSynbio-Search: A Large Language Model Driven Agent Framework for Unified Search Method for Protein Engineering

AgentsDGX agent

arXiv:2411.06024v1 Announce Type: cross Abstract: The exponential growth in protein-related databases and scientific literature, combined with increasing demands for efficient biological information r

5 Aug 2026

Agent memory layers don't need an LLM deciding what to remember

Local AiDGX agent

Most agent memory setups run a model call on the way in. Something reads the turn, decides whether it's worth keeping, rewrites it into a 'memory', tags it with a type and an importance score. That's

Agentic Reinforcement Learning with Self-Distilled Reward Shaping

SafetyDGX agent

arXiv:2608.03223v1 Announce Type: cross Abstract: Agentic reinforcement learning enables LLM agents to learn through interaction, but sparse trajectory-level rewards reveal success without identifying

Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning

AgentsDGX agent

arXiv:2608.03571v1 Announce Type: new Abstract: Recent works train agents by constructing large-scale multimodal environment pools. However, we find that simply increasing the number of multimodal env

← Previous
1…6364656667…297
Next →