AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
5,141 results
Agents

A collaborative agent with two lightweight synergistic models for autonomous crystal materials research

DGX agent

arXiv:2604.11540v1 Announce Type: new Abstract: Current large language models require hundreds of billions of parameters yet struggle with domain-specific reasoning and tool coordination in materials

agentsarxiv-cs-ai
14 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Seven simple steps for log analysis in AI systems

DGX agent

arXiv:2604.09563v1 Announce Type: new Abstract: AI systems produce large volumes of logs as they interact with tools and users. Analysing these logs can help understand model capabilities, propensitie

researcharxiv-cs-ai
14 Apr 2026
Agents

LLM Agents Factory: Retrieval of Domain-Specific LLM Agents

DGX agent

arXiv:2608.09934v1 Announce Type: cross Abstract: Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors. However, their practical deploymen

agentsarxiv-cs-ai
12 Aug 2026
Agents

Mitigating Context Interference for Reliable and Efficient Search Agents

DGX agent

arXiv:2608.10743v1 Announce Type: new Abstract: Recent research empowers Large Language Models (LLMs) as multi-turn search agents to iteratively retrieve and generate outputs until complex tasks are s

agentsarxiv-cs-cl
12 Aug 2026
Research

Towards Sustainable Artificial Intelligence: A Comprehensive Review and Comparative Analysis of Deep Learning Models' Carbon Footprint

DGX agent

arXiv:2608.09998v1 Announce Type: new Abstract: Artificial Intelligence (AI) and Machine Learning (ML) have become powerful tools for supporting and automating complex human tasks. Despite their benef

researcharxiv-cs-ai
12 Aug 2026
Model Releases

An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography

DGX agent

arXiv:2608.07651v1 Announce Type: new Abstract: Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency.

model-releasesarxiv-cs-ai
11 Aug 2026
Agents

CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents

DGX agent

arXiv:2608.07855v1 Announce Type: new Abstract: Multi-turn Reasoning-and-Acting (ReAct) agents accumulate growing trajectories of reasoning, tool calls, and observations. Their key-value (KV) caches g

agentsarxiv-cs-lg
11 Aug 2026
Safety

Designing for Ethical AI: HCI Feature Considerations to Improve Fairness and User Experience in AutoML use for Human Resources

DGX agent

arXiv:2608.07477v1 Announce Type: cross Abstract: This thesis examines the fairness of Automated Machine Learning (AutoML) tools in human resource hiring systems through the combined lenses of regulat

safetyarxiv-cs-ai
11 Aug 2026
Local Ai

Evidence-Calibrated Runtime Reconstruction for Agent Skills Across Heterogeneous Coding Agents

DGX agent

arXiv:2608.08793v1 Announce Type: new Abstract: Agent Skills package reusable instructions and assets for tool-using language-model agents. Progressive loading creates failure boundaries poorly repres

local-aiarxiv-cs-cl
11 Aug 2026
Local Ai

Explainable Machine Learning in Healthcare: Methods, Interpretation, and Applications for Clinical Research

DGX agent

arXiv:2608.07522v1 Announce Type: cross Abstract: We present a structured review of commonly used Explainable machine learning (XML) methodologies, including global and local interpretability tools su

local-aiarxiv-cs-lg
11 Aug 2026
Model Releases

MemeMind: Reference-Guided Trace Construction for Offline Context Optimization

DGX agent

arXiv:2608.09316v1 Announce Type: new Abstract: Offline context optimization improves an agent by revising its instructions and examples while keeping the model frozen. This approach learns from rollo

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

OBLIVION: Workflow-Level Operational Skill Unlearning for Deployed Agents

DGX agent

arXiv:2608.08264v1 Announce Type: new Abstract: Large language model agents are becoming operational interfaces to files, memories, registries, and external tools. This deployment shift creates a new

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents

DGX agent

arXiv:2608.08775v1 Announce Type: new Abstract: Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost ex

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling

DGX agent

arXiv:2608.08700v1 Announce Type: new Abstract: Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as autonomous agents. Current benchmarks face three struct

model-releasesarxiv-cs-ai
11 Aug 2026
Agents

SAKE: Structured Agentic Knowledge Extrapolation for Complex LLM Reasoning via Reinforcement Learning

DGX agent

arXiv:2505.15062v5 Announce Type: replace-cross Abstract: Knowledge extrapolation is the process of inferring novel information by combining and extending existing knowledge that is explicitly availab

agentsarxiv-cs-ai
11 Aug 2026
Model Releases

SHE: Trajectory-driven Safety Harness Evolution for LLM Agents

DGX agent

arXiv:2608.09885v1 Announce Type: new Abstract: The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, per

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

TRACE: TRajectory Attribution for Automated Context Engineering

DGX agent

arXiv:2608.09153v1 Announce Type: new Abstract: Production AI agents fail when their context sources -- system prompts, knowledge bases, tool descriptions, and procedural skills -- contain errors or g

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory

DGX agent

arXiv:2608.07169v1 Announce Type: new Abstract: Memory systems have shown promise for improving agent performance, but their potential remains largely unexplored for small language models, which strug

model-releasesarxiv-cs-ai
10 Aug 2026
Agents

AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models

DGX agent

arXiv:2608.06699v1 Announce Type: new Abstract: Agentic multimodal large language models (MLLMs) extend multimodal perception and reasoning with planning, tool use, and interaction in dynamic environm

agentsarxiv-cs-ai
10 Aug 2026
Model Releases

NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs

DGX agent

arXiv:2608.07167v1 Announce Type: new Abstract: Giving an AI agent the ability to send emails, query databases, or execute commands is useful--until the agent is tricked into doing something it should

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

Quantization Damage Is Multiplicative, Not Additive

DGX agent

arXiv:2608.06564v1 Announce Type: cross Abstract: Quantization is how large language models are actually deployed, and below four bits it is known to hurt. What nobody can say is which of the model's

model-releasesarxiv-cs-cl
10 Aug 2026
Agents

Strategy-first synthesis planning for complex natural products

DGX agent

arXiv:2608.07454v1 Announce Type: cross Abstract: The total synthesis of a complex molecule is among the most demanding intellectual and experimental feats in chemistry: a chemist must plan many steps

agentsarxiv-cs-ai
10 Aug 2026
Agents

The Perils of Agency: How Developers Perceive, Prioritize, and Address Risks in Agentic AI Products

DGX agent

arXiv:2606.15485v2 Announce Type: replace-cross Abstract: Agentic AI systems act autonomously, use tools, adapt to context, and operate in complex real-world environments. However, these same characte

agentsarxiv-cs-ai
10 Aug 2026
Agents

ASTELD: A Six-Axis Classification Framework for Autonomous AI Agents - Design, Evaluation, and an OpenClaw Case Study

DGX agent

arXiv:2608.05201v1 Announce Type: cross Abstract: Autonomous AI agent platforms differ substantially in architecture, security, tool integration, execution, autonomy, and deployment, yet the field lac

agentsarxiv-cs-ai
7 Aug 2026
Safety

EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning

DGX agent

arXiv:2608.06197v1 Announce Type: new Abstract: Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesized executable environments, whose

safetyarxiv-cs-ai
7 Aug 2026
Safety

EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents

DGX agent

arXiv:2608.05446v1 Announce Type: cross Abstract: Long-horizon LLM agents increasingly rely on external execution support to maintain state, track progress, invoke tools, verify outcomes, and reuse ex

safetyarxiv-cs-cl
7 Aug 2026
Safety

Negotiating Risk Boundaries in AI for Policing Through Mixed-Stakeholder Deliberation

DGX agent

arXiv:2608.05418v1 Announce Type: new Abstract: AI tools are being increasingly adopted in policing in the UK and worldwide. Racial bias is a known and well-documented risk, yet representatives of aff

safetyarxiv-cs-ai
7 Aug 2026
Local Ai

Architectural Implications of Agentic AI Workflows

DGX agent

arXiv:2608.04458v1 Announce Type: new Abstract: Agentic AI is emerging in datacenters, but its architectural implications remain unexplored. We organize agentic workflows in a taxonomy and present its

local-aiarxiv-cs-ai
6 Aug 2026
Safety

Breadcrumbing Search Agents

DGX agent

arXiv:2608.04565v1 Announce Type: cross Abstract: LLM-based search agents are widely used for information-seeking tasks, but their reliance on external tool returns introduces a critical security risk

safetyarxiv-cs-ai
6 Aug 2026
Model Releases

EDATracer: An Agentic Framework for Large-Scale EDA Artifact Analysis

DGX agent

arXiv:2608.04032v1 Announce Type: cross Abstract: Modern chip design relies on electronic design automation (EDA) tools that generate large, heterogeneous artifacts, including source files, scripts, l

model-releasesarxiv-cs-ai
6 Aug 2026
Model Releases

Formal Analysis and Supply Chain Security for Agentic AI Skills

DGX agent

arXiv:2603.00195v2 Announce Type: replace-cross Abstract: 32 pages, 5 theorems with full proofs, 68 references, open-source tool: https://github.com/qualixar/skillfortify. v2: corrects the bibliograph

model-releasesarxiv-cs-ai
6 Aug 2026
Safety

Governing Execution Risk in Agentic AI Systems: A Trajectory-Guided Framework for Red Teaming

DGX agent

arXiv:2608.04018v1 Announce Type: cross Abstract: AI agents are increasingly embedded in organizational workflows, where they interact with external information sources and invoke digital tools to per

safetyarxiv-cs-ai
6 Aug 2026
Safety

SafeCommit: Certifying When Memory-Grounded Agents May Safely Act

DGX agent

arXiv:2608.04289v1 Announce Type: new Abstract: Long-horizon agents increasingly use persistent memory and tools to take actions with external side effects. A central failure mode is premature commitm

safetyarxiv-cs-ai
6 Aug 2026
Safety

BAP-SQL: Budget-Aware Observation Planning for Agentic Text-to-SQL

DGX agent

arXiv:2608.02876v1 Announce Type: new Abstract: Tool-using agents do not merely consume observations: their actions determine what arrives next. In agentic text-to-SQL, a broad query can spend context

safetyarxiv-cs-ai
5 Aug 2026
Model Releases

Can LLMs Test Terminal User Interfaces?

DGX agent

arXiv:2608.03743v1 Announce Type: cross Abstract: Terminal User Interfaces (TUIs) combine the stateful, screen-oriented behaviour of GUIs with terminal deployment and are now common in developer tools

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

TraceCompiler: Skill-Guided Mining and Compilation of LLM Agent Traces into Mostly Deterministic Workflows

DGX agent

arXiv:2608.02680v1 Announce Type: cross Abstract: Tool-using language-model agents repeatedly rediscover procedures they have already executed, producing traces that mix reusable structure with retrie

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

AOSpec: Action and Observation Co-Speculation for Low-Latency Agent Serving

DGX agent

arXiv:2608.00881v1 Announce Type: new Abstract: Large language model agents increasingly act through stateful tools, yet model generation and environment execution remain serialized at every step. As

model-releasesarxiv-cs-lg
4 Aug 2026
Model Releases

Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training

DGX agent

arXiv:2608.02391v1 Announce Type: cross Abstract: Tool-using large language model (LLM) agents produce long, multi-turn trajectories, making gradient-based post-training memory-intensive. Evolution st

model-releasesarxiv-cs-lg
4 Aug 2026
Safety

Deep Research Pretraining via Predictive Navigation

DGX agent

arXiv:2608.00432v1 Announce Type: new Abstract: Deep research agents are often trained on expensive, environment-grounded tool-use trajectories that require repeated retrieval, document inspection, an

safetyarxiv-cs-cl
4 Aug 2026
Tutorials

Generative Brownian Bridge Diffusion In Motion Space For Enhanced Myocardial Strain Analysis

DGX agent

arXiv:2608.01677v1 Announce Type: new Abstract: Myocardial strain analysis of cardiac magnetic resonance (CMR) images provides an important tool for evaluating cardiac function. However, current techn

tutorialsarxiv-cs-cv
4 Aug 2026
Model Releases

Real-Time Detection and Repair of LLM Agent Failures

DGX agent

arXiv:2608.02464v1 Announce Type: cross Abstract: LLM agents fail mid-episode -- they loop, cascade tool errors, drift off goal, fabricate results, or silently absorb corrupted content -- and the stan

model-releasesarxiv-cs-lg
4 Aug 2026
Model Releases

Real-Time Visual Obstruction Detection in Surgical Augmented Reality

DGX agent

arXiv:2608.00232v1 Announce Type: new Abstract: Surgical augmented reality (AR) can provide contextual guidance by overlaying virtual annotations, tool cues, and procedural information onto the surgic

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

Why Large Language Models Fail at Tabular Prediction

DGX agent

arXiv:2608.02412v1 Announce Type: new Abstract: Large language models (LLMs) have become the default tool for a remarkable range of tasks, yet they have had conspicuously little success at one of the

model-releasesarxiv-cs-lg
4 Aug 2026
Model Releases

Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications

DGX agent

arXiv:2607.28840v1 Announce Type: new Abstract: Large language models are increasingly deployed in financial applications that combine retrieval, proprietary data, tool use, orchestration logic, monit

model-releasesarxiv-cs-cl
3 Aug 2026
Model Releases

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures

DGX agent

arXiv:2607.28802v1 Announce Type: new Abstract: Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention would improve the

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

Open-Source LLM-Driven Formal Verification: A Multi-Agent Pipeline for RTL Repair

DGX agent

arXiv:2607.28877v1 Announce Type: cross Abstract: Verification consumes the majority of modern chip design effort, yet the formal verification tools that provide mathematical guarantees of correctness

model-releasesarxiv-cs-lg
3 Aug 2026
Safety

RePaCA: Leveraging Reasoning Large Language Models for Static Automated Patch Correctness Assessment

DGX agent

arXiv:2507.22580v2 Announce Type: replace-cross Abstract: Automated Program Repair (APR) seeks to automatically correct software bugs without requiring human intervention. However, existing tools tend

safetyarxiv-cs-ai
3 Aug 2026
Model Releases

Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents

DGX agent

arXiv:2607.27275v1 Announce Type: new Abstract: Post-training quantization to 4-bit weights is widely reported to be nearly lossless. We test this claim for multi-turn, tool-calling agents, where it n

model-releasesarxiv-cs-lg
31 Jul 2026
← Previous
1…1516171819…108
Next →