AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
5,141 results
Agents

Emergent Culture in Minimal LLM Systems

DGX agent

arXiv:2606.30668v1 Announce Type: cross Abstract: What happens when LLM agents operate with no context outside a turn, minimal prompting, and simple tools? Inspired by swarm engineering, we give colle

agentsarxiv-cs-ai
1 Jul 2026
Local Ai

Look But Don't Touch with Sparse Autoencoders for Unlearning in Diffusion Models

DGX agent
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.31699v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have recently been proposed as interpretable tools for concept-level manipulation, under the assumption that isolated featu

local-aiarxiv-cs-ai
1 Jul 2026
Model Releases

The Decomposition Is the Fingerprint: Per-Component Identity for Agent Skills

DGX agent

arXiv:2606.31272v1 Announce Type: cross Abstract: AI agents increasingly acquire and execute skills at runtime: bundles of prompt instructions, executable code, and tool declarations fetched from mark

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

Understanding and Evaluating Claw-like Agent Security Through a Computer-Systems Lens

DGX agent

arXiv:2606.30755v1 Announce Type: cross Abstract: Claw-like AI agents (e.g., OpenClaw) are always-on processes with persistent access to credentials, files, tools, and external services. They take on

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Agentic Abstention: Do Agents Know When to Stop Instead of Act?

DGX agent

arXiv:2606.28733v1 Announce Type: new Abstract: LLM agents are expected to act over multiple turns, using search, browsing interfaces, and terminal tools to complete user goals. Yet not every goal is

model-releasesarxiv-cs-ai
30 Jun 2026
Agents

BackTranslation2.0 -- A Linguistically Motivated Metric to Assess Sign Language Production

DGX agent

arXiv:2606.28673v1 Announce Type: new Abstract: Sign Languages (SLs) are the primary means of communication for millions of deaf individuals, yet existing evaluation metrics for generated SL remain si

agentsarxiv-cs-cv
30 Jun 2026
Model Releases

Internal-State Probes Read the Situation, Not the Action: Three Negative Results for Pre-Action Misalignment Monitoring

DGX agent

arXiv:2606.30449v1 Announce Type: new Abstract: Probes on model internals could help monitor agentic systems if they identify harmful text or tool actions before those actions are generated. We ask wh

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

TraceLab: Characterizing Coding Agent Workloads for LLM Serving

DGX agent

arXiv:2606.30560v1 Announce Type: cross Abstract: Coding agents are rapidly becoming a major application of agentic LLMs, but serving them efficiently remains challenging. Progress on this challenge r

model-releasesarxiv-cs-ai
30 Jun 2026
Safety

Agent-Native Immune System: Architecture, Taxonomy, and Engineering

DGX agent

arXiv:2606.28270v1 Announce Type: new Abstract: The transition from static chat bots to autonomous agents--equipped with persistent memory, tool-use protocols, and multi-agent collaboration--has funda

safetyarxiv-cs-ai
29 Jun 2026
Safety

DysLexLens: A Low-Resource LLM Framework for Analysing Dyslexic Learners Insights from Online Forums

DGX agent

arXiv:2606.27619v1 Announce Type: new Abstract: Dyslexic learners increasingly use artificial intelligence (AI) tools to support reading, writing, organisation, and study-related tasks. However, their

safetyarxiv-cs-ai
29 Jun 2026
Agents

SwarmFly: A simulation platform for UAV swarm experiment design and validation

DGX agent

arXiv:2606.25146v1 Announce Type: new Abstract: The initial development phase of UAV swarms largely depends on simulation for experimental design and validation, yet existing open-source tools are oft

agentsarxiv-cs-ro
25 Jun 2026
Local Ai

SHERLOC: Structured Diagnostic Localization for Code Repair Agents

DGX agent

arXiv:2606.24820v1 Announce Type: new Abstract: LLM agents solve repository-level coding tasks through multi-turn tool use, but utilize half their budget on locating faults before editing. Dedicated l

local-aiarxiv-cs-cl
24 Jun 2026
Model Releases

From Knowing to Acting: Benchmarking Self-Awareness Capability of LLM Agents

DGX agent

arXiv:2606.20661v1 Announce Type: cross Abstract: The integration of external tools has transitioned LLM agents from passive responders to autonomous systems. However, current benchmarks prioritize ex

model-releasesarxiv-cs-lg
23 Jun 2026
Agents

Position: Correct Answer, Wrong Mechanism -- When AI Scientists Defend General Claims Their Own Data Contradicts

DGX agent

arXiv:2606.23175v1 Announce Type: new Abstract: AI scientist systems are described as tools, coauthors, or founders, but we evaluate them as if only the final answer matters. This position paper argue

agentsarxiv-cs-lg
23 Jun 2026
Model Releases

RoverDevKit: An open, physics-grounded tradespace toolkit for conceptual design of lunar micro-rovers

DGX agent

arXiv:2606.21755v1 Announce Type: new Abstract: Pre-Phase-A design of lunar micro-rovers is dominated by tightly coupled mobility, power, thermal, and mass trades, yet conceptual-design tooling for th

model-releasesarxiv-cs-ro
23 Jun 2026
Safety

ZeroDex: Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning

DGX agent

arXiv:2606.19340v2 Announce Type: replace Abstract: We present ZeroDex, a zero-shot framework for long-horizon dexterous manipulation that grounds language instructions into executable 3D task plans f

safetyarxiv-cs-ro
23 Jun 2026
Agents

Agents All the Way Down; A Methodology for Building Custom AI Agents from Substrate to Production

DGX agent

arXiv:2606.11869v1 Announce Type: cross Abstract: Custom AI agents areagents that live inside their own application, talk to their own data and tools, enforce their own security boundaries, and carry

agentsarxiv-cs-ai
11 Jun 2026
Agents

Runtime Skill Audit: Targeted Runtime Probing for Agent Skill Security

DGX agent

arXiv:2606.11671v1 Announce Type: cross Abstract: Agent skills let LLM agents reuse instructions, resources, tools, and workflows, but they also create a new place for malicious behavior to hide. A sk

agentsarxiv-cs-ai
11 Jun 2026
Model Releases

SwiftCTS: Fast Cross-Design Prediction and Pareto Optimization of Clock Tree Metrics via Few-Shot Calibration

DGX agent

arXiv:2606.11348v1 Announce Type: new Abstract: Clock Tree Synthesis (CTS) is a computationally expensive stage in the physical design flow, requiring iterative EDA tool invocations to navigate a vast

model-releasesarxiv-cs-lg
11 Jun 2026
Hardware

Accounting for AI Inference in Corporate GHG Inventories: A Four-Tier Methodology for Scope 3 Category 1 Reporting

DGX agent

arXiv:2606.10660v1 Announce Type: cross Abstract: AI inference services -- API subscriptions, enterprise chat tools, and SaaS products with embedded AI features -- fall unambiguously within Scope 3 Ca

hardwarearxiv-cs-ai
10 Jun 2026
Model Releases

T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains

DGX agent

arXiv:2606.11070v1 Announce Type: cross Abstract: Recent advances in reasoning and tool-calling capabilities of large language models (LLMs) have enabled increasingly capable agentic systems. However,

model-releasesarxiv-cs-ai
10 Jun 2026
Safety

Brain-Prompt Injection: A Route-Safety Audit for BCI-LLM Agents

DGX agent

arXiv:2606.09315v1 Announce Type: cross Abstract: BCI-to-agent pipelines turn decoded neural activity into an authorization channel for tool-use agents, exposing a new attack surface we call brain-pro

safetyarxiv-cs-ai
9 Jun 2026
Agents

Collaborative Human-Agent Protocol (CHAP)

DGX agent

arXiv:2606.09751v1 Announce Type: new Abstract: Foundation models are moving from response generation into operational roles. They plan across steps, call tools, request human input, coordinate with o

agentsarxiv-cs-ai
9 Jun 2026
Research

Lost in the Flow with Code Talkers: Unveiling the Instruction-Tuning Tax of Large Language Models in Code Tasks

DGX agent

arXiv:2606.08676v1 Announce Type: cross Abstract: AI coding assistants have significantly improved developer productivity by automatically suggesting code that aligns with user intent, and many of the

researcharxiv-cs-ai
9 Jun 2026
Model Releases

Visual Template Inference for Data Extraction from Documents

DGX agent

arXiv:2501.06659v2 Announce Type: replace-cross Abstract: Many templatized documents are programmatically generated from structured data following a visual template. Such documents include invoices, t

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation

DGX agent

arXiv:2606.07723v1 Announce Type: new Abstract: Open-vocabulary long-horizon manipulation requires robots to reason over flexible instructions and complex multi-object scenes while adaptively planning

model-releasesarxiv-cs-ro
9 Jun 2026
Local Ai

Queen-Bee Agents: A BeeSpec-Centered Architecture for Governed Enterprise MCP Orchestration

DGX agent

arXiv:2606.06545v1 Announce Type: cross Abstract: Enterprise agent systems increasingly need to connect large language models to private tools, internal knowledge, and Model Context Protocol (MCP) int

local-aiarxiv-cs-ai
8 Jun 2026
Safety

SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating

DGX agent

arXiv:2606.07074v1 Announce Type: cross Abstract: Deep research agents have demonstrated remarkable capabilities in complex information-seeking tasks, yet this power comes at a steep computational cos

safetyarxiv-cs-ai
8 Jun 2026
Model Releases

Do More Agents Help? Controlled and Protocol-Aligned Evaluation of LLM Agent Workflows

DGX agent

arXiv:2606.05670v1 Announce Type: new Abstract: Does adding more agents help an LLM workflow once compared systems share the same benchmark loader, tool access, answer contract, usage accounting, and

model-releasesarxiv-cs-ai
6 Jun 2026
Research

Learning Adaptive Parallel Execution for Efficient Code Localization

DGX agent

arXiv:2601.19568v2 Announce Type: replace Abstract: Code localization constitutes a key bottleneck in automated software development pipelines. While concurrent tool execution can enhance discovery sp

researcharxiv-cs-ai
6 Jun 2026
Model Releases

SciVisAgentSkills: Design and Evaluation of Agent Skills for Scientific Data Analysis and Visualization

DGX agent

arXiv:2606.05525v1 Announce Type: new Abstract: Recent advances in agentic visualization have enabled the translation of natural language into executable scientific visualization (SciVis) workflows. W

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents

DGX agent

arXiv:2606.05557v1 Announce Type: new Abstract: A situated query like 'where is Lin Wei?' often encodes more than its literal content: the user may also want to know whether Lin Wei is free, in a good

model-releasesarxiv-cs-cl
5 Jun 2026
Research

ACAT: A Collaborative Platform for Efficient Aspect-Based Sentiment Dataset Annotation

DGX agent

arXiv:2606.04189v1 Announce Type: new Abstract: Aspect-Based Sentiment Analysis (ABSA) requires high-quality datasets to train reliable models. However, existing annotation tools treat output as flat

researcharxiv-cs-cl
4 Jun 2026
Safety

Beyond Symmetric Alignment: Spectral Diagnostics of Modality Imbalance in Vision-Language Models in the Medical Domain

DGX agent

arXiv:2606.04613v1 Announce Type: new Abstract: Vision-Language Models (VLMs) struggle when applied to medical image-text data, yet the tools available to diagnose this failure remain limited. Existin

safetyarxiv-cs-cv
4 Jun 2026
Research

Characterizing initial human-AI proof formalization workflows

DGX agent

arXiv:2606.04273v1 Announce Type: new Abstract: For centuries, human mathematicians have written proofs to substantiate their mathematical arguments; yet, the ability to automatically verify the valid

researcharxiv-cs-ai
4 Jun 2026
Model Releases

Proof-Carrying Agent Actions: Model-Agnostic Runtime Governance for Heterogeneous Agent Systems

DGX agent

arXiv:2606.04104v1 Announce Type: cross Abstract: Agent systems execute through runtimes with very different control points: local coding tools, framework SDKs, managed agent platforms, API gateways,

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

VAMPS: Visual-Assisted Mathematical Problem Solving Benchmark

DGX agent

arXiv:2606.04244v1 Announce Type: new Abstract: Multimodal large language models are increasingly capable of complex reasoning, yet their performance often degrades when they must externalize a proble

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions

DGX agent

arXiv:2606.03318v1 Announce Type: new Abstract: Despite great advances in tool-use capabilities of large language models (LLMs), existing evaluation benchmarks struggle to fully align with real-world

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

KForge: LLM-Driven Cross-Platform Kernel Generation for AI Accelerators

DGX agent

arXiv:2606.02963v1 Announce Type: new Abstract: Production inference increasingly targets a heterogeneous mix of accelerators. Agentic pipelines interleave reasoning, tool calls, and multi-agent coord

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing

DGX agent

arXiv:2606.02822v1 Announce Type: cross Abstract: Production LLM applications stack several defense families -- refusal-phrase filters, token-budget controls, model allowlists, rate limits, tool-regis

model-releasesarxiv-cs-ai
3 Jun 2026
Safety

Who Deserves the Reward? SHARP: Shapley Credit-based Optimization for Multi-Agent System

DGX agent

arXiv:2602.08335v2 Announce Type: replace Abstract: Integrating Large Language Models (LLMs) with external tools via multi-agent systems offers a promising new paradigm for decomposing and solving com

safetyarxiv-cs-ai
3 Jun 2026
Agents

Beyond One-shot: AI Agents for Learning in Field Experiments

DGX agent

arXiv:2606.02458v1 Announce Type: new Abstract: Organizations routinely run experiments for A/B testing, yet the data generated from one experiment is underutilized to inform subsequent intervention d

agentsarxiv-cs-ai
2 Jun 2026
Research

Digging Up Citations: FOSSIL, a Dataset and Workflow for Reference Extraction in Law and the Humanities

DGX agent

arXiv:2606.01109v1 Announce Type: cross Abstract: Citation extraction tools are designed for the structured end-of-document bibliographies of the natural sciences, but law and humanities scholarship c

researcharxiv-cs-cl
2 Jun 2026
Model Releases

FALAT: Tracing Failures in LLM Agent Trajectories via Dependency-Guided Search

DGX agent

arXiv:2606.00765v1 Announce Type: new Abstract: LLM-based agents increasingly solve complex tasks through long trajectories involving reasoning steps, tool calls, and inter-agent communication. Howeve

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

GenPT: Beyond Self-Report for Reliable LLM Psychometrics via Generative Projective Testing

DGX agent

arXiv:2606.00860v1 Announce Type: cross Abstract: Self-report questionnaires remain the prevailing tool for probing the psychological states of persona-conditioned agents (PC-Agents). However, classic

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations

DGX agent

arXiv:2606.00832v1 Announce Type: new Abstract: Recent advances in agentic AI have enabled agents to complete complex tasks through tool use, reasoning, and multi-step planning. Yet existing benchmark

model-releasesarxiv-cs-cl
2 Jun 2026
Hardware

Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving

DGX agent

arXiv:2606.01839v1 Announce Type: cross Abstract: LLM-based agents resolve a user task through many turns of dependent inference and tool calls, producing a workload whose total cost is unknown when t

hardwarearxiv-cs-lg
2 Jun 2026
Model Releases

Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning

DGX agent

arXiv:2606.00963v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) exhibit emerging spatial reasoning capabilities, yet they remain unreliable on tasks requiring precise spatial understan

model-releasesarxiv-cs-cl
2 Jun 2026
← Previous
1…1718192021…108
Next →