AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
5,189 results
Model Releases

OPT-Engine: Benchmarking the Limits of LLMs in Optimization Modeling via Complexity Scaling

DGX agent

arXiv:2601.19924v2 Announce Type: replace-cross Abstract: We investigate the capabilities and scalability of Large Language Models (LLMs) in optimization modeling, a domain requiring structured reason

model-releasesarxiv-cs-ai
15 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

PALMS: A Computational Implementation for Pavlovian Associative Learning Models' Simulation

DGX agent

arXiv:2602.07519v3 Announce Type: replace Abstract: In contrast to static formalisms, computational definitions describe the operational mechanisms of a model. Simulations are an essential part of the

researcharxiv-cs-lg
15 May 2026
Safety

Probabilistic Verification of Recurrent Neural Networks for Single and Multi-Agent Reinforcement Learning

DGX agent

arXiv:2605.14758v1 Announce Type: new Abstract: History-dependent policies induced by recurrent neural networks (RNNs) rely on latent hidden state dynamics, making verification in partially observable

safetyarxiv-cs-ai
15 May 2026
Model Releases

SPIN: Structural LLM Planning via Iterative Navigation for Industrial Tasks

DGX agent

arXiv:2605.14051v1 Announce Type: new Abstract: Industrial LLM agent systems often separate planning from execution, yet LLM planners frequently produce structurally invalid or unnecessarily long work

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Web Agents Should Adopt the Plan-Then-Execute Paradigm

DGX agent

arXiv:2605.14290v1 Announce Type: cross Abstract: ReAct has become the default architecture across LLM agents, and many existing web agents follow this paradigm. We argue that it is the wrong default

model-releasesarxiv-cs-ai
15 May 2026
Research

Context Training with Active Information Seeking

DGX agent

arXiv:2605.13050v1 Announce Type: cross Abstract: Most existing large language models (LLMs) are expensive to adapt after deployment, especially when a task requires newly produced information or nich

researcharxiv-cs-ai
14 May 2026
Safety

Integration of an Agent Model into an Open Simulation Architecture for Scenario-Based Testing of Automated Vehicles

DGX agent

arXiv:2605.13539v1 Announce Type: new Abstract: Simulative and scenario-based testing are crucial methods in the safety assurance for automated driving systems. To ensure that simulation results are r

safetyarxiv-cs-ro
14 May 2026
Safety

Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety

DGX agent

arXiv:2605.12729v1 Announce Type: cross Abstract: Large language models are increasingly being used to support network operations (NetOps) and artificial intelligence for IT operations (AIOps), includ

safetyarxiv-cs-ai
14 May 2026
Model Releases

MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows

DGX agent

arXiv:2603.24649v2 Announce Type: replace Abstract: Medical imaging benchmarks often evaluate VLMs on pre-selected 2D images, slices, crops, or patches, making evaluation closer to visual recognition.

model-releasesarxiv-cs-cv
14 May 2026
Model Releases

ABRA: Agent Benchmark for Radiology Applications

DGX agent

arXiv:2605.11224v1 Announce Type: new Abstract: Existing medical-agent benchmarks deliver imaging as pre-selected samples, never as an environment the agent must navigate. We introduce ABRA, a radiolo

model-releasesarxiv-cs-cv
13 May 2026
Hardware

MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces

DGX agent

arXiv:2605.11333v1 Announce Type: cross Abstract: The fast pace of artificial intelligence~(AI) innovation demands an agile methodology for observation, reproduction and optimization of distributed ma

hardwarearxiv-cs-lg
13 May 2026
Model Releases

Posterior Contraction Rates for Sparse Kolmogorov-Arnold Networks in Anisotropic Besov Spaces

DGX agent

arXiv:2605.11652v1 Announce Type: cross Abstract: We study posterior contraction rates for sparse Bayesian Kolmogorov-Arnold networks (KANs) over anisotropic Besov spaces, providing a statistical foun

model-releasesarxiv-cs-lg
13 May 2026
Research

Unpacking the Eye of the Beholder: Social Location, Identity, and the Moving Target of Political Perspectives

DGX agent

arXiv:2605.11166v1 Announce Type: new Abstract: Political and social identities structure how people evaluate political information, a finding decades deep in political science and routinely discarded

researcharxiv-cs-cv
13 May 2026
Local Ai

A probabilistic framework for crystal structure denoising, phase classification, and order parameters

DGX agent

arXiv:2512.11077v3 Announce Type: replace-cross Abstract: Atomistic simulations generate large volumes of noisy structural data, yet extracting phase labels and continuous order parameters (OPs) in a

local-aiarxiv-cs-ai
12 May 2026
Model Releases

AHD Agent: Agentic Reinforcement Learning for Automatic Heuristic Design

DGX agent

arXiv:2605.08756v1 Announce Type: new Abstract: Automatic heuristic design (AHD) has emerged as a promising paradigm for solving NP-hard combinatorial optimization problems (COPs). Recent works show t

model-releasesarxiv-cs-ai
12 May 2026
Safety

AI-Care: A Conversational Agentic System for Task Coordination in Alzheimer's Disease Care

DGX agent

arXiv:2605.08480v1 Announce Type: new Abstract: Individuals with Alzheimer's disease (AD) and Alzheimer's disease-related dementia (ADRD) experience memory and thinking changes that impact their abili

safetyarxiv-cs-ai
12 May 2026
Research

Applying Graph Analysis for Unsupervised Fast Malware Fingerprinting

DGX agent

arXiv:2510.12811v2 Announce Type: replace-cross Abstract: Malware proliferation is increasing at a tremendous rate, with hundreds of thousands of new samples identified daily. Manual investigation of

researcharxiv-cs-lg
12 May 2026
Research

ChladniSonify: A Visual-Acoustic Mapping Method for Chladni Patterns in New Media Art Creation

DGX agent

arXiv:2605.09846v1 Announce Type: cross Abstract: In new media art creation, the mapping between vision and hearing is often subjective. As a classic carrier of sound visualization, Chladni patterns h

researcharxiv-cs-ai
12 May 2026
Model Releases

CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents

DGX agent

arXiv:2605.09675v1 Announce Type: new Abstract: Clinical reasoning agents based on large language models (LLMs) aim to automate tasks such as intensive care unit (ICU) monitoring and patient state tra

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

CUDAHercules: Benchmarking Hardware-Aware Expert-level CUDA Optimization for LLMs

DGX agent

arXiv:2605.08467v1 Announce Type: new Abstract: Large language models show promise for automated CUDA programming, however even the strongest coding models (e.g., Claude-Opus-4.6) may still fall short

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

EnergyLens: Interpretable Closed-Form Energy Models for Multimodal LLM Inference Serving

DGX agent

arXiv:2605.10556v1 Announce Type: new Abstract: As large language models span dense, mixture-of-experts, and state-space architectures and are deployed on heterogeneous accelerators under increasingly

model-releasesarxiv-cs-cv
12 May 2026
Safety

FairHealth: An Open-Source Python Library for Trustworthy Healthcare AI in Low-Resource Settings

DGX agent

arXiv:2605.08198v1 Announce Type: cross Abstract: We present FairHealth, an open-source Python library that provides a unified, modular framework for trustworthy machine learning in healthcare applica

safetyarxiv-cs-ai
12 May 2026
Model Releases

FORTIS: Benchmarking Over-Privilege in Agent Skills

DGX agent

arXiv:2605.09163v1 Announce Type: new Abstract: Large language model agents increasingly operate through an intermediate skill layer that mediates between user intent and concrete task execution. This

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation

DGX agent

arXiv:2601.11258v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) face the 'knowledge cutoff' challenge, where their frozen parametric memory prevents direct internalization of ne

model-releasesarxiv-cs-ai
12 May 2026
Research

MineEvolve: Self-Evolution with Accumulated Knowledge for Long-Horizon Embodied Minecraft Agents

DGX agent

arXiv:2603.13131v3 Announce Type: replace Abstract: Long-horizon embodied intelligence requires agents to improve through interaction, not merely to execute plans generated from static goals. A centra

researcharxiv-cs-ai
12 May 2026
Model Releases

MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring

DGX agent

arXiv:2605.09684v1 Announce Type: cross Abstract: We introduce a red-teaming methodology that exposes harder-to-catch attacks for coding-agent monitors, suggesting that current practices may under-eli

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search

DGX agent

arXiv:2605.08762v1 Announce Type: cross Abstract: Current omni-modal benchmarks mainly evaluate models under settings where multiple modalities are provided simultaneously, while the ability to start

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark

DGX agent

arXiv:2509.26574v4 Announce Type: replace Abstract: While large language models (LLMs) with reasoning capabilities are progressing rapidly on high-school math competitions and coding, can they reason

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?

DGX agent

arXiv:2605.10848v1 Announce Type: cross Abstract: Does a lexical retriever suffice as large language models (LLMs) become more capable in an agentic loop? This question naturally arises when building

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Strategic commitments shape collective cybersecurity under AI inequality

DGX agent

arXiv:2605.09415v1 Announce Type: new Abstract: The growing integration of AI into cybersecurity is reshaping the balance between attackers and defenders. When access to advanced AI-enabled defence to

model-releasesarxiv-cs-ai
12 May 2026
Agents

TinyTroupe: An LLM-powered Multiagent Persona Simulation Toolkit

DGX agent

arXiv:2507.09788v2 Announce Type: replace-cross Abstract: Recent advances in Large Language Models (LLM) have led to a new class of autonomous agents, renewing and expanding interest in the area. LLM-

agentsarxiv-cs-ai
12 May 2026
Tutorials

What Cohort INRs Encode and Where to Freeze Them

DGX agent

arXiv:2605.08298v1 Announce Type: cross Abstract: Reusing the early layers of cohort-trained INRs as initialization for new signals has been shown to accelerate and improve signal fitting, yet it rema

tutorialsarxiv-cs-ai
12 May 2026
Applications

Exploring CoCo Challenges in ML Engineering Teams: Insights From the Semiconductor Industry

DGX agent

arXiv:2605.07389v1 Announce Type: cross Abstract: The integration of machine learning (ML) into complex software systems has increased challenges in collaboration and communication (CoCo) of the teams

applicationsarxiv-cs-lg
11 May 2026
Research

LensVLM: Selective Context Expansion for Compressed Visual Representation of Text

DGX agent

arXiv:2605.07019v1 Announce Type: cross Abstract: Vision Language Models (VLMs) offer the exciting possibility of processing text as rendered images, bypassing the need for tokenizing the text into lo

researcharxiv-cs-ai
11 May 2026
Model Releases

MAS-Algorithm: A Workflow for Solving Algorithmic Programming Problems with a Multi-Agent System

DGX agent

arXiv:2605.05949v2 Announce Type: replace Abstract: Algorithmic problem solving serves as a rigorous testbed for evaluating structured reasoning in AI coding systems, as it directly reflects a model's

model-releasesarxiv-cs-ai
11 May 2026
Agents

MedExAgent: Training LLM Agents to Ask, Examine, and Diagnose in Noisy Clinical Environments

DGX agent

arXiv:2605.07058v1 Announce Type: cross Abstract: Real-world clinical diagnosis is a complex process in which the doctor is required to obtain information from both interaction with the patient and co

agentsarxiv-cs-ai
11 May 2026
Applications

Vibe coding before the trend

DGX agent

arXiv:2605.07751v1 Announce Type: cross Abstract: Early 2025 we ran a series of vibe coding challenges across four different student cohorts. The cohorts included 54 ICT students, 24 digital marketing

applicationsarxiv-cs-ai
11 May 2026
Model Releases

When Routine Chats Turn Toxic: Unintended Long-Term State Poisoning in Personalized Agents

DGX agent

arXiv:2605.06731v1 Announce Type: cross Abstract: Personalized LLM agents maintain persistent cross-session state to support long-horizon collaboration. Yet, this persistence introduces a subtle but c

model-releasesarxiv-cs-cl
11 May 2026
Model Releases

Agentic Vulnerability Reasoning on Windows COM Binaries

DGX agent

arXiv:2605.05000v1 Announce Type: cross Abstract: Windows Component Object Model (COM) services run with elevated privileges and are widely accessible to authenticated users, making race conditions in

model-releasesarxiv-cs-lg
7 May 2026
Local Ai

ARISE: A Repository-level Graph Representation and Toolset for Agentic Fault Localization and Program Repair

DGX agent

arXiv:2605.03117v1 Announce Type: cross Abstract: Repository-level fault localization (FL) and automated program repair (APR) require an agent to identify the relevant code units across files, follow

local-aiarxiv-cs-ai
7 May 2026
Local Ai

Contextual Multi-Objective Optimization: Rethinking Objectives in Frontier AI Systems

DGX agent

arXiv:2605.03900v1 Announce Type: new Abstract: Frontier AI systems perform best in settings with clear, stable, and verifiable objectives, such as code generation, mathematical reasoning, games, and

local-aiarxiv-cs-ai
7 May 2026
Agents

From Knowledge to Action: Outcomes of the 2025 Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry

DGX agent

arXiv:2605.03205v1 Announce Type: cross Abstract: Large language models (LLMs) are rapidly changing how researchers in materials science and chemistry discover, organize, and act on scientific knowled

agentsarxiv-cs-ai
7 May 2026
Model Releases

Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation

DGX agent

arXiv:2605.04135v1 Announce Type: cross Abstract: Readers of applied-domain LLM capability evaluations want to know what AI systems can currently do. That literature answers a related, but consequenti

model-releasesarxiv-cs-cl
7 May 2026
Safety

Investigating Trustworthiness of Nonparametric Deep Survival Models for Alzheimer's Disease Progression Analysis

DGX agent

arXiv:2605.04063v1 Announce Type: new Abstract: Alzheimer's Dementia (AD) is a progressive neurodegenerative disease marked by irreversible decline, making reliable modeling of its progression essenti

safetyarxiv-cs-lg
7 May 2026
Tutorials

Multi Language Models for On-the-Fly Syntax Highlighting

DGX agent

arXiv:2510.04166v2 Announce Type: replace-cross Abstract: Syntax highlighting is a critical feature in modern software development environments, enhancing code readability and developer productivity.

tutorialsarxiv-cs-ai
7 May 2026
Agents

OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents

DGX agent

arXiv:2605.05185v1 Announce Type: new Abstract: Deep search has become a crucial capability for frontier multimodal agents, enabling models to solve complex questions through active search, evidence v

agentsarxiv-cs-cv
7 May 2026
Safety

OracleProto: A Reproducible Framework for Benchmarking LLM Native Forecasting via Knowledge Cutoff and Temporal Masking

DGX agent

arXiv:2605.03762v1 Announce Type: new Abstract: Large language models are moving from static text generators toward real-world decision-support systems, where forecasting is a composite capability tha

safetyarxiv-cs-ai
7 May 2026
Research

Analytic Bridge Diffusions for Controlled Path Generation

DGX agent

arXiv:2605.02961v1 Announce Type: new Abstract: Most modern bridge-diffusion methods achieve finite-time transport by specifying an interpolation, Schrodinger-bridge, or stochastic-control objective a

researcharxiv-cs-lg
6 May 2026
← Previous
1…4849505152…109
Next →