AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
5,141 results
Local Ai

scicode-lint: Detecting Methodology Bugs in Scientific Python Code with LLM-Generated Patterns

DGX agent

arXiv:2603.17893v2 Announce Type: replace-cross Abstract: Methodology bugs in scientific Python code produce plausible but incorrect results that traditional linters and static analysis tools cannot d

local-aiarxiv-cs-ai
2 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Statistical Testing on Directed Graphs by Surrogate Data Generation

DGX agent

arXiv:2606.00758v1 Announce Type: cross Abstract: In recent years, graph signal processing has emerged as a powerful framework at the intersection of signal processing and graph theory, providing tool

researcharxiv-cs-lg
2 Jun 2026
Local Ai

An Organization-Scoped LLM Agent Runtime Architecture for Regulated Cybersecurity Operations

DGX agent

arXiv:2605.30604v1 Announce Type: cross Abstract: Regulated cybersecurity workflows lack a runtime substrate that enforces organization-level scope across retrieval, tool calls, memory, findings, repo

local-aiarxiv-cs-ai
1 Jun 2026
Agents

Automatically Attacking Software Reverse Engineering AI Agents

DGX agent

arXiv:2605.30667v1 Announce Type: cross Abstract: Software tools for reverse engineering executable binary files, such as Ghidra, enable malware analysts to safely conduct robust static analysis witho

agentsarxiv-cs-ai
1 Jun 2026
Safety

Human Psychometric Questionnaires Mischaracterize LLM Behavior

DGX agent

arXiv:2509.10078v4 Announce Type: replace-cross Abstract: We examine whether human psychometric questionnaires can serve as reliable tools for characterizing and predicting LLM behavior in everyday us

safetyarxiv-cs-ai
1 Jun 2026
Local Ai

The Architecture of Errors: From Universal Impossibility to Patch-Local LLM Reliability

DGX agent

arXiv:2605.30628v1 Announce Type: cross Abstract: Universal LLM reliability is not a finite-library problem: across all possible tasks, tools, schemas, knowledge sources, and evaluator expectations, n

local-aiarxiv-cs-ai
1 Jun 2026
Agents

E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing

DGX agent

arXiv:2512.03109v2 Announce Type: replace-cross Abstract: Agentic AI systems execute a sequence of actions, such as reasoning steps or tool calls, in response to a user prompt. To evaluate the success

agentsarxiv-cs-ai
29 May 2026
Model Releases

Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents

DGX agent

arXiv:2605.29224v1 Announce Type: cross Abstract: AI agents augment large language models with external tools such as web retrieval, enabling grounded and up-to-date responses. However, incorporating

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Benchmarking Convolutional, Transformer, Hybrid, and Vision Language Models for Multi Disease Retinal Screening

DGX agent

arXiv:2605.26283v1 Announce Type: new Abstract: Modern deep learning offers powerful tools for automated retinal screening, but it remains unclear how different visual model families compare in realis

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

Device Context Protocol: A Compact, Safety-First Architecture for LLM-Driven Control of Constrained Devices

DGX agent

arXiv:2605.26159v1 Announce Type: cross Abstract: Large language models are increasingly used as orchestrators of external tools via the Model Context Protocol (MCP), but MCP is built for software ser

model-releasesarxiv-cs-lg
27 May 2026
Agents

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments

DGX agent

arXiv:2605.27209v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have facilitated the widespread deployment of LLMs as interactive agents capable of reasoning, planning,

agentsarxiv-cs-ai
27 May 2026
Safety

Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction

DGX agent

arXiv:2505.11063v3 Announce Type: replace Abstract: LLM-based agents solve complex tasks through iterative reasoning, tool use, and environment interaction, where each intermediate thought directly sh

safetyarxiv-cs-ai
27 May 2026
Model Releases

UnityMAS-O: A General RL Optimization Framework for LLM-Based Multi-Agent Systems

DGX agent

arXiv:2605.26646v1 Announce Type: new Abstract: LLM-based multi-agent systems decompose complex tasks into interacting roles, but most remain manually orchestrated by prompts, tools, and control rules

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning

DGX agent

arXiv:2602.10090v3 Announce Type: replace Abstract: Recent advances in large language model (LLM) have empowered autonomous agents to perform multi-turn interactions with tools and environments. Howev

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions

DGX agent

arXiv:2605.24279v1 Announce Type: new Abstract: A frontier language model's acknowledged 'helpful programming assistant' persona does not survive long agentic-coding sessions in the deployment regime

model-releasesarxiv-cs-cl
26 May 2026
Safety

Machine Psychometrics: A Mathematical Psychology of Artificial Intelligence

DGX agent

arXiv:2605.23952v1 Announce Type: new Abstract: Artificial agents now generate behavior rich enough to invite trust, surprise, and concern, yet our evaluation tools still privilege capability scores o

safetyarxiv-cs-ai
26 May 2026
Research

Spurious Stationarity and Hardness Results for Bregman Proximal-Type Algorithms

DGX agent

arXiv:2404.08073v3 Announce Type: replace-cross Abstract: Bregman proximal-type algorithms (BPs), such as mirror descent, have become popular tools in machine learning and data science for exploiting

researcharxiv-cs-lg
26 May 2026
Research

Defining AI Fatigue in Academic Contexts: Dimensions, Indicators, and a Stage-Based Model Using Grounded Theory

DGX agent

arXiv:2605.23123v1 Announce Type: cross Abstract: The integration of AI tools in academic settings has introduced a distinct form of strain that existing frameworks like technostress and digital fatig

researcharxiv-cs-ai
25 May 2026
Safety

Foundation Protocol: A Coordination Layer for Agentic Society

DGX agent

arXiv:2605.23218v1 Announce Type: new Abstract: Autonomous agents are moving from tools into a layer of social infrastructure: they browse, purchase, deploy software, manage systems, and increasingly

safetyarxiv-cs-ai
25 May 2026
Safety

Graph-based Complexity Forecasts in UK En Route Airspace Using Relevant Aircraft Interactions

DGX agent

arXiv:2605.23696v1 Announce Type: new Abstract: Effectively managing Air Traffic Control Officer (ATCO) workload is crucial in maintaining operational safety. Group supervisors use tools that estimate

safetyarxiv-cs-lg
25 May 2026
Safety

Relevant Walk Search for Explaining Graph Neural Networks

DGX agent

arXiv:2605.23673v1 Announce Type: new Abstract: Graph Neural Networks (GNNs) have become important machine learning tools for graph analysis, and its explainability is crucial for safety, fairness, an

safetyarxiv-cs-lg
25 May 2026
Applications

Security of LLM-generated Code: A Comparative Analysis

DGX agent

arXiv:2605.23091v1 Announce Type: cross Abstract: The majority of software developers use or are planning to use Artificial Intelligence (AI) tools in their development processes. Their top reasons in

applicationsarxiv-cs-ai
25 May 2026
Model Releases

AgentCo-op: Retrieval-Based Synthesis of Interoperable Multi-Agent Workflows

DGX agent

arXiv:2605.20425v1 Announce Type: new Abstract: Designing multi-agent workflows is especially difficult in open-ended scientific settings where tasks lack curated training sets, reliable scalar evalua

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines

DGX agent

arXiv:2605.20630v1 Announce Type: new Abstract: Industrial asset operations workflows are latency-sensitive because a single user query may require coordination over sensor data, work orders, failure

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

Putnam 2025 Problems in Rocq using Opus 4.6 and Rocq-MCP

DGX agent

arXiv:2603.20405v2 Announce Type: replace-cross Abstract: We report on an experiment in which Claude Opus~4.6, equipped with a suite of Model Context Protocol (MCP) tools for the Rocq proof assistant,

model-releasesarxiv-cs-cl
22 May 2026
Research

Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

DGX agent

arXiv:2601.23224v2 Announce Type: replace Abstract: Existing multimodal large language models for long-video understanding predominantly rely on uniform sampling and single-turn inference, limiting th

researcharxiv-cs-cv
22 May 2026
Model Releases

Evolutionary Generation of Multi-Agent Systems

DGX agent

arXiv:2602.06511v3 Announce Type: replace Abstract: Large language model (LLM)-based multi-agent systems (MAS) show strong promise for complex reasoning, planning, and tool-augmented tasks, but design

model-releasesarxiv-cs-lg
21 May 2026
Agents

What if Agents Could Imagine? Reinforcing Open-Vocabulary HOI Comprehension through Generation

DGX agent

arXiv:2602.11499v2 Announce Type: replace Abstract: Multimodal Large Language Models have shown promising capabilities in bridging visual and textual reasoning, yet their reasoning capabilities in Ope

agentsarxiv-cs-cv
21 May 2026
Model Releases

Entry-level guide to the use of large language models for medical research

DGX agent

arXiv:2410.18856v4 Announce Type: replace Abstract: Frontier large language models (LLMs), such as GPT-5, Claude 4.5, Gemini 3, Llama 4, and DeepSeek-R1, represent a transformative class of AI tools c

model-releasesarxiv-cs-ai
20 May 2026
Safety

Formal Skill: Programmable Runtime Skills for Efficient and Accurate LLM Agents

DGX agent

arXiv:2605.19604v1 Announce Type: new Abstract: Large Language Model (LLM) agents increasingly act inside real workspaces, where tools and skills determine whether model reasoning becomes reliable act

safetyarxiv-cs-ai
20 May 2026
Model Releases

Hallucination as Exploit: Evidence-Carrying Multimodal Agents

DGX agent

arXiv:2605.19192v1 Announce Type: new Abstract: Multimodal agents use screenshots, documents, and webpages to choose tool calls. When a false visual claim triggers a click, email, extraction, or trans

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Measuring Safety Alignment Effects in Autonomous Security Agents

DGX agent

arXiv:2605.19722v1 Announce Type: cross Abstract: Do stock safety-aligned language models and their uncensored or abliterated derivatives behave differently when run as autonomous security agents? Sin

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?

DGX agent

arXiv:2605.19196v1 Announce Type: new Abstract: Deep research agents increasingly automate complex information-seeking tasks, producing evidence-grounded reports via multi-step reasoning, tool use, an

model-releasesarxiv-cs-cl
20 May 2026
Tutorials

AI Slop or AI-enhancement? Student perceptions of AI-generated media for an English for Academic Purposes course

DGX agent

arXiv:2605.16275v1 Announce Type: cross Abstract: Artificial intelligence (AI) retrieval-augmented generation (RAG) tools now enable educators to transform course materials into diverse multimedia at

tutorialsarxiv-cs-ai
19 May 2026
Safety

An Empirical Study of Privacy Leakage Chains via Prompt Injection in Black-Box Chatbot Environments

DGX agent

arXiv:2605.18133v1 Announce Type: cross Abstract: LLM-based chatbot agents increasingly process user requests by combining natural-language reasoning with external tools such as web browsing. These ca

safetyarxiv-cs-ai
19 May 2026
Research

Fine-tuning an ECG Foundation Model to Predict Coronary CT Angiography Outcomes

DGX agent

arXiv:2512.05136v3 Announce Type: replace-cross Abstract: CAD remains a major global public health burden, yet scalable screening tools are limited. Although CCTA is a first-line non-invasive diagnost

researcharxiv-cs-ai
19 May 2026
Local Ai

Herding CATs: ALARA for Agent Harness Engineering in Portable Composable Multi-Agent Teams

DGX agent

arXiv:2603.20380v2 Announce Type: replace-cross Abstract: Industry practitioners and academic researchers regularly use multi-agent systems to accelerate their work, but the applications through which

local-aiarxiv-cs-ai
19 May 2026
Model Releases

LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injectio

DGX agent

arXiv:2605.17986v1 Announce Type: cross Abstract: AI agents such as OpenClaw are increasingly deployed in local workflows with access to external tools. This creates indirect prompt-injection (IPI) ri

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

MentalBench: A DSM-Grounded Benchmark for Evaluating Psychiatric Diagnostic Capability of Large Language Models

DGX agent

arXiv:2602.12871v2 Announce Type: replace Abstract: Large language models (LLMs) have attracted growing interest as supportive tools for psychiatric assessment and clinical decision support. However,

model-releasesarxiv-cs-cl
19 May 2026
Local Ai

PropGuard: Safeguarding LLM-MAS via Propagation-Aware Exploration and Remediation

DGX agent

arXiv:2605.16346v1 Announce Type: cross Abstract: LLM-based multi-agent systems (LLM-MAS) have become a promising paradigm for solving complex tasks through role specialization, tool use, memory, and

local-aiarxiv-cs-ai
19 May 2026
Model Releases

TusoAI: Agentic Optimization for Scientific Methods

DGX agent

arXiv:2509.23986v2 Announce Type: replace Abstract: Scientific discovery is often slowed by the manual development of computational tools needed to analyze complex experimental data. Building such too

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

A3D: Agentic AI flow for autonomous Accelerator Design

DGX agent

arXiv:2605.15237v1 Announce Type: cross Abstract: Accelerating applications through the design of hardware accelerators can significantly enhance system performance and energy efficiency. Despite adva

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

CAX-Agent: A Lightweight Agent Harness for Reliable APDL Automation

DGX agent

arXiv:2605.15218v1 Announce Type: new Abstract: Large language models deployed for MAPDL finite-element simulation face practical reliability challenges: without structured execution control, tool enc

model-releasesarxiv-cs-ai
18 May 2026
Agents

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems

DGX agent

arXiv:2605.14892v1 Announce Type: new Abstract: LLM-based autonomous agents have demonstrated strong capabilities in reasoning, planning, and tool use, yet remain limited when tasks require sustained

agentsarxiv-cs-ai
15 May 2026
Model Releases

Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction

DGX agent

arXiv:2605.13950v1 Announce Type: cross Abstract: Autonomous language-model agents are increasingly evaluated on long-horizon tool-use tasks, but existing benchmarks rarely capture the complexity and

model-releasesarxiv-cs-ai
15 May 2026
Agents

Concurrency without Model Changes: Future-based Asynchronous Function Calling for LLMs

DGX agent

arXiv:2605.15077v1 Announce Type: cross Abstract: Function calling, also known as tool use, is a core capability of modern LLM agents but is typically constrained by synchronous execution semantics. U

agentsarxiv-cs-ai
15 May 2026
Model Releases

CUICurate: A GraphRAG-based Framework for Automated Clinical Concept Curation for NLP applications

DGX agent

arXiv:2602.17949v2 Announce Type: replace-cross Abstract: Background: Clinical named entity recognition tools commonly map free text to Unified Medical Language System (UMLS) Concept Unique Identifier

model-releasesarxiv-cs-ai
15 May 2026
Safety

Exploring Geographic Relative Space in Large Language Models through Activation Patching

DGX agent

arXiv:2605.14535v1 Announce Type: new Abstract: The increased use of Large Language Models (LLMs) in geography raises substantial questions about the safety of integrating these tools across a wide ra

safetyarxiv-cs-lg
15 May 2026
← Previous
1…1819202122…108
Next →