AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
20 Apr 2026

Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures

SafetyDGX agent

arXiv:2604.16042v1 Announce Type: cross Abstract: While Large Language Models (LLMs) have achieved strong performance across many NLP tasks, their opaque internal mechanisms hinder trustworthiness and

Towards Rigorous Explainability by Feature Attribution

ResearchDGX agent

arXiv:2604.15898v1 Announce Type: new Abstract: For around a decade, non-symbolic methods have been the option of choice when explaining complex machine learning (ML) models. Unfortunately, such metho

Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective

HardwareDGX agent

arXiv:2511.00739v3 Announce Type: replace Abstract: Agentic AI serving converts monolithic LLM-based inference to autonomous problem-solvers that can plan, call tools, perform reasoning, and adapt on


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

TPA: Next Token Probability Attribution for Detecting Hallucinations in RAG

ResearchDGX agent

arXiv:2512.07515v4 Announce Type: replace-cross Abstract: Detecting hallucinations in Retrieval-Augmented Generation remains a challenge. Prior approaches attribute hallucinations to a binary conflict

Training Time Prediction for Mixed Precision-based Distributed Training

ResearchDGX agent

arXiv:2604.16145v1 Announce Type: cross Abstract: Accurate prediction of training time in distributed deep learning is crucial for resource allocation, cost estimation, and job scheduling. We observe

Transfer Learning from Foundational Optimization Embeddings to Unsupervised SAT Representations

ResearchDGX agent

arXiv:2604.15448v1 Announce Type: cross Abstract: Foundational optimization embeddings have recently emerged as powerful pre-trained representations for mixed-integer programming (MIP) problems. These

Transformer Neural Processes - Kernel Regression

Model ReleasesDGX agent

arXiv:2411.12502v4 Announce Type: replace-cross Abstract: Neural Processes (NPs) are a rapidly evolving class of models designed to directly model the posterior predictive distribution of stochastic p

TriagerX: Dual Transformers for Bug Triaging Tasks with Content and Interaction Based Rankings

ApplicationsDGX agent

arXiv:2508.16860v2 Announce Type: replace-cross Abstract: Pretrained Language Models or PLMs are transformer-based architectures that can be used in bug triaging tasks. PLMs can better capture token s

Uncertainty, Vagueness, and Ambiguity in Human-Robot Interaction: Why Conceptualization Matters

ResearchDGX agent

arXiv:2604.15339v1 Announce Type: cross Abstract: Uncertainty, vagueness, and ambiguity are closely related and often confused concepts in human-robot interaction (HRI). In earlier studies, these conc

UniEditBench: A Unified and Cost-Effective Benchmark for Image and Video Editing via Distilled MLLMs

Model ReleasesDGX agent

arXiv:2604.15871v1 Announce Type: cross Abstract: The evaluation of visual editing models remains fragmented across methods and modalities. Existing benchmarks are often tailored to specific paradigms

Unveiling Stochasticity: Universal Multi-modal Probabilistic Modeling for Traffic Forecasting

ApplicationsDGX agent

arXiv:2604.16084v1 Announce Type: cross Abstract: Traffic forecasting is a challenging spatio-temporal modeling task and a critical component of urban transportation management. Current studies mainly

Using Large Language Models and Knowledge Graphs to Improve the Interpretability of Machine Learning Models in Manufacturing

ApplicationsDGX agent

arXiv:2604.16280v1 Announce Type: new Abstract: Explaining Machine Learning (ML) results in a transparent and user-friendly manner remains a challenging task of Explainable Artificial Intelligence (XA

VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects

Model ReleasesDGX agent

arXiv:2604.16272v1 Announce Type: cross Abstract: As AI-assisted video creation becomes increasingly practical, instruction-guided video editing has become essential for refining generated or captured

VeriCWEty: Embedding enabled Line-Level CWE Detection in Verilog

ResearchDGX agent

arXiv:2604.15375v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown significant improvement in RTL code generation. Despite the advances, the generated code is often riddled with

VeriGraph: Scene Graphs for Execution Verifiable Robot Planning

AgentsDGX agent

arXiv:2411.10446v3 Announce Type: replace-cross Abstract: Recent progress in vision-language models (VLMs) has opened new possibilities for robot task planning, but these models often produce incorrec

VeriMoA: A Mixture-of-Agents Framework for Spec-to-HDL Generation

AgentsDGX agent

arXiv:2510.27617v2 Announce Type: replace Abstract: Automation of Register Transfer Level (RTL) design can help developers meet increasing computational demands. Large Language Models (LLMs) show prom

VIB-Probe: Detecting and Mitigating Hallucinations in Vision-Language Models via Variational Information Bottleneck

ResearchDGX agent

arXiv:2601.05547v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) have demonstrated remarkable progress in multimodal tasks, but remain susceptible to hallucinations, where gener

vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2603.13966v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models are increasingly evaluated across multiple simulation benchmarks, yet adding each benchmark to an evaluation pip

VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models

Model ReleasesDGX agent

arXiv:2512.14554v5 Announce Type: replace-cross Abstract: The rapid advancement of large language models (LLMs) has enabled new possibilities for applying artificial intelligence within the legal doma

VoodooNet: Achieving Analytic Ground States via High-Dimensional Random Projections

Local AiDGX agent

arXiv:2604.15613v1 Announce Type: cross Abstract: We present VoodooNet, a non-iterative neural architecture that replaces the stochastic gradient descent (SGD) paradigm with a closed-form analytic sol

WARBERT: A Hierarchical BERT-based Model for Web API Recommendation

ResearchDGX agent

arXiv:2509.23175v2 Announce Type: replace-cross Abstract: With the rise of Web 2.0 and microservices, the increasing availability of Web APIs has intensified the need for effective recommendation syst

Weak-Link Optimization for Multi-Agent Reasoning and Collaboration

AgentsDGX agent

arXiv:2604.15972v1 Announce Type: new Abstract: LLM-driven multi-agent frameworks address complex reasoning tasks through multi-role collaboration. However, existing approaches often suffer from reaso

When Cultures Meet: Multicultural Text-to-Image Generation

Model ReleasesDGX agent

arXiv:2502.15972v2 Announce Type: replace-cross Abstract: Text-to-image generation models have achieved strong performance in culturally homogeneous settings, yet their ability to generate multicultur

When Do Early-Exit Networks Generalize? A PAC-Bayesian Theory of Adaptive Depth

ResearchDGX agent

arXiv:2604.15764v1 Announce Type: cross Abstract: Early-exit neural networks enable adaptive computation by allowing confident predictions to exit at intermediate layers, achieving 2-8imes inference s

When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models

SafetyDGX agent

arXiv:2510.09689v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have been augmented with web search to overcome the limitations of the static knowledge boundary by accessing up-

When the Loop Closes: Architectural Limits of In-Context Isolation, Metacognitive Co-option, and the Two-Target Design Problem in Human-LLM Systems

ApplicationsDGX agent

arXiv:2604.15343v1 Announce Type: cross Abstract: We report a detailed autoethnographic case study of a single-subject who deliberately constructed and operated a multi-modal prompt-engineering system

Where does output diversity collapse in post-training?

ResearchDGX agent

arXiv:2604.16027v1 Announce Type: cross Abstract: Post-trained language models produce less varied outputs than their base counterparts. This output diversity collapse undermines inference-time scalin

Why Fine-Tuning Encourages Hallucinations and How to Fix It

Model ReleasesDGX agent

arXiv:2604.15574v1 Announce Type: cross Abstract: Large language models are prone to hallucinating factually incorrect statements. A key source of these errors is exposure to new factual information t

WiseMind: a knowledge-guided multi-agent framework for accurate and empathetic psychiatric diagnosis

AgentsDGX agent

arXiv:2502.20689v4 Announce Type: replace Abstract: Large Language Models (LLMs) offer promising opportunities to support mental healthcare workflows, yet they often lack the structured clinical reaso

Zoom Consistency: A Free Confidence Signal in Multi-Step Visual Grounding Pipelines

ResearchDGX agent

arXiv:2604.15376v1 Announce Type: cross Abstract: Multi-step zoom-in pipelines are widely used for GUI grounding, yet the intermediate predictions they produce are typically discarded after coordinate

17 Apr 2026

3D Instruction Ambiguity Detection

Model ReleasesDGX agent

arXiv:2601.05991v2 Announce Type: replace Abstract: In safety-critical domains, linguistic ambiguity can have severe consequences; a vague command like 'Pass me the vial' in a surgical setting could l

A Pythonic Functional Approach for Semantic Data Harmonisation in the ILIAD Project

TutorialsDGX agent

arXiv:2604.13042v1 Announce Type: cross Abstract: Semantic data harmonisation is a central requirement in the ILIAD project, where heterogeneous environmental data must be harmonised according to the

AgentForge: Execution-Grounded Multi-Agent LLM Framework for Autonomous Software Engineering

AgentsDGX agent

arXiv:2604.13120v1 Announce Type: cross Abstract: Large language models generate plausible code but cannot verify correctness. Existing multi-agent systems simulate execution or leave verification opt

Agentic AI Optimisation (AAIO): what it is, how it works, why it matters, and how to deal with it

AgentsDGX agent

arXiv:2504.12482v2 Announce Type: replace Abstract: The emergence of Agentic Artificial Intelligence (AAI) systems capable of independently initiating digital interactions necessitates a new optimisat

AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot

Model ReleasesDGX agent

arXiv:2604.13940v1 Announce Type: new Abstract: Scientific peer review faces mounting strain as submission volumes surge, making it increasingly difficult to sustain review quality, consistency, and t

AlphaCNOT: Learning CNOT Minimization with Model-Based Planning

ResearchDGX agent

arXiv:2604.13812v1 Announce Type: new Abstract: Quantum circuit optimization is a central task in Quantum Computing, as current Noisy Intermediate Scale Quantum devices suffer from error propagation t

AMA: Adaptive Memory via Multi-Agent Collaboration

AgentsDGX agent

arXiv:2601.20352v3 Announce Type: replace Abstract: The rapid evolution of Large Language Model (LLM) agents has necessitated robust memory systems to support cohesive long-term interaction and comple

Animating Petascale Time-varying Data on Commodity Hardware with LLM-assisted Scripting

ApplicationsDGX agent

arXiv:2603.07053v2 Announce Type: replace Abstract: Scientists face significant visualization challenges as time-varying datasets grow in speed and volume, often requiring specialized infrastructure a

Applying an Agentic Coding Tool for Improving Published Algorithm Implementations

Model ReleasesDGX agent

arXiv:2604.13109v1 Announce Type: cross Abstract: We present a two-stage pipeline for AI-assisted improvement of published algorithm implementations. In the first stage, a large language model with re

Assessment Design in the AI Era: A Method for Identifying Items Functioning Differentially for Humans and Chatbots

Model ReleasesDGX agent

arXiv:2603.23682v2 Announce Type: replace-cross Abstract: The rapid adoption of large language models (LLMs) in education raises profound challenges for assessment design. To adapt assessments to the

Bayesian-LoRA: Probabilistic Low-Rank Adaptation of Large Language Models

ResearchDGX agent

arXiv:2601.21003v2 Announce Type: replace Abstract: Large Language Models usually put more emphasis on accuracy and therefore, will guess even when not certain about the prediction, which is especiall

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs

SafetyDGX agent

arXiv:2509.05367v4 Announce Type: replace-cross Abstract: Large Language Model safety alignment predominantly operates on a binary assumption that requests are either safe or unsafe. This classificati

BitFlipScope: Scalable Fault Localization and Recovery for Bit-Flip Corruptions in LLMs

Local AiDGX agent

arXiv:2512.22174v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) deployed in practical and safety-critical settings are increasingly susceptible to bit-flip faults caused by hard

Building Trust in the Skies: A Knowledge-Grounded LLM-based Framework for Aviation Safety

SafetyDGX agent

arXiv:2604.13101v1 Announce Type: cross Abstract: The integration of Large Language Models (LLMs) into aviation safety decision-making represents a significant technological advancement, yet their sta

CCCE: A Continuous Code Calibration Engine for Autonomous Enterprise Codebase Maintenance via Knowledge Graph Traversal and Adaptive Decision Gating

AgentsDGX agent

arXiv:2604.13102v1 Announce Type: cross Abstract: Enterprise software organizations face an escalating challenge in maintaining the integrity, security, and freshness of codebases that span hundreds o

ChatSVA: Bridging SVA Generation for Hardware Verification via Task-Specific LLMs

Model ReleasesDGX agent

arXiv:2604.02811v2 Announce Type: replace-cross Abstract: Functional verification consumes over 50% of the IC development lifecycle, where SystemVerilog Assertions (SVAs) are indispensable for formal

Cognitive Offloading in Agile Teams: How Artificial Intelligence Reshapes Risk Assessment and Planning Quality

ResearchDGX agent

arXiv:2604.13814v1 Announce Type: cross Abstract: Recent advances in artificial intelligence (AI) have shown promise in automating key aspects of Agile project management, yet their impact on team cog

Comparison of window shapes and lengths in short-time feature extraction for classification of heart sound signals

ResearchDGX agent

arXiv:2604.13567v1 Announce Type: cross Abstract: Heart sound signals, phonocardiography (PCG) signals, allow for the automatic diagnosis of potential cardiovascular pathology. Such classification tas

Contextuality from Single-State Ontological Models: An Information-Theoretic Obstruction

ResearchDGX agent

arXiv:2602.16716v3 Announce Type: replace Abstract: Contextuality is a central feature of quantum theory, traditionally understood as the impossibility of reproducing quantum measurement statistics us

Contract-Coding: Towards Repo-Level Generation via Structured Symbolic Paradigm

Model ReleasesDGX agent

arXiv:2604.13100v1 Announce Type: cross Abstract: The shift toward intent-driven software engineering (often termed 'Vibe Coding') exposes a critical Context-Fidelity Trade-off: vague user intents ove

ContractSkill: Repairable Contract-Based Skills for Multimodal Web Agents

Model ReleasesDGX agent

arXiv:2603.20340v3 Announce Type: replace-cross Abstract: Self-generated skills for web agents are often unstable and can even hurt performance relative to direct acting. We argue that the key bottlen

DeepPresenter: Environment-Grounded Reflection for Agentic Presentation Generation

AgentsDGX agent

arXiv:2602.22839v2 Announce Type: replace Abstract: Presentation generation requires deep content research, coherent visual design, and iterative refinement based on observation. However, existing pre

Domain-Adaptive Model Merging Across Disconnected Modes

ResearchDGX agent

arXiv:2603.05957v2 Announce Type: replace-cross Abstract: Learning across domains is challenging when data cannot be centralized due to privacy or heterogeneity, which limits the ability to train a si

ECM Contracts: Contract-Aware, Versioned, and Governable Capability Interfaces for Embodied Agents

Model ReleasesDGX agent

arXiv:2604.13097v1 Announce Type: cross Abstract: Embodied agents increasingly rely on modular capabilities that can be installed, upgraded, composed, and governed at runtime. Prior work has introduce

[Emerging Ideas] Artificial Tripartite Intelligence: A Bio-Inspired, Sensor-First Architecture for Physical AI

SafetyDGX agent

arXiv:2604.13959v1 Announce Type: new Abstract: As AI moves from data centers to robots and wearables, scaling ever-larger models becomes insufficient. Physical AI operates under tight latency, energy

Empowerment Gain and Causal Model Construction: Children and adults are sensitive to controllability and variability in their causal interventions

AgentsDGX agent

arXiv:2512.08230v2 Announce Type: replace Abstract: Learning about the causal structure of the world is a fundamental problem for human cognition. Causal models and especially causal learning have pro

Exploration and Exploitation Errors Are Measurable for Language Model Agents

SafetyDGX agent

arXiv:2604.13151v1 Announce Type: new Abstract: Language Model (LM) agents are increasingly used in complex open-ended decision-making tasks, from AI coding to physical AI. A core requirement in these

Finetuning-Free Diffusion Model with Adaptive Constraint Guidance for Inorganic Crystal Structure Generation

ResearchDGX agent

arXiv:2604.13354v1 Announce Type: cross Abstract: The discovery of inorganic crystal structures with targeted properties is a significant challenge in materials science. Generative models, especially

Formal Architecture Descriptors as Navigation Primitives for AI Coding Agents

Model ReleasesDGX agent

arXiv:2604.13108v1 Announce Type: cross Abstract: AI coding agents spend a substantial fraction of their tool calls on undirected codebase exploration. We investigate whether providing agents with for

Formalizing the Safety, Security, and Functional Properties of Agentic AI Systems

SafetyDGX agent

arXiv:2510.14133v2 Announce Type: replace Abstract: Agentic AI systems, which leverage multiple autonomous agents and large language models (LLMs), are increasingly used to address complex, multi-step

← Previous
1…324325326327328…354
Next →