AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
28 Apr 2026

Defective Task Descriptions in LLM-Based Code Generation: Detection and Analysis

Model ReleasesDGX agent

arXiv:2604.24703v1 Announce Type: cross Abstract: Large language models are widely used for code generation, yet they rely on an implicit assumption that the task descriptions are sufficiently detaile

Defusing the Trigger: Plug-and-Play Defense for Backdoored LLMs via Tail-Risk Intrinsic Geometric Smoothing

Model ReleasesDGX agent

arXiv:2604.24162v1 Announce Type: cross Abstract: Defending against backdoor attacks in large language models remains a critical practical challenge. Existing defenses mitigate these threats but typic

Delta-XAI: A Unified Framework for Explaining Prediction Changes in Online Time Series Monitoring

Applications

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2511.23036v2 Announce Type: replace-cross Abstract: Explaining online time series monitoring models is crucial across sensitive domains such as healthcare and finance, where temporal and context

DenoGrad: A Gradient-Based Framework for Data Refinement in Tabular and Time-Series Learning

ApplicationsDGX agent

arXiv:2511.10161v2 Announce Type: replace Abstract: In the Data-Centric Artificial Intelligence (AI) paradigm, improving data quality is essential for robust machine learning. However, many denoising

Deployment-Aligned Low-Precision Neural Architecture Search for Spaceborne Edge AI

Local AiDGX agent

arXiv:2604.24492v1 Announce Type: cross Abstract: Designing deep networks that meet strict latency and accuracy constraints on edge accelerators increasingly relies on hardware-aware optimization, inc

DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference

ResearchDGX agent

arXiv:2604.24647v1 Announce Type: cross Abstract: Long-context reasoning is a critical capability of large language models (LLMs), enabling applications such as long-document understanding, summarizat

Designing escalation criteria for international AI incident response: criteria, triggers, and thresholds

SafetyDGX agent

arXiv:2604.23183v1 Announce Type: cross Abstract: AI incident reporting requirements are emerging in regulation and policy, yet no operational criteria exist for determining when a detected AI inciden

Diffusion Templates: A Unified Plugin Framework for Controllable Diffusion

Model ReleasesDGX agent

arXiv:2604.24351v1 Announce Type: cross Abstract: Controllable diffusion methods have substantially expanded the practical utility of diffusion models, but they are typically developed as isolated, ba

Discovering Agentic Safety Specifications from 1-Bit Danger Signals

SafetyDGX agent

arXiv:2604.23210v1 Announce Type: new Abstract: Can large language model agents discover hidden safety objectives through experience alone? We introduce EPO-Safe (Experiential Prompt Optimization for

Discovering Failure Modes in Vision-Language Models using RL

SafetyDGX agent

arXiv:2604.04733v2 Announce Type: replace-cross Abstract: Vision-language Models (VLMs), despite achieving strong performance on multimodal benchmarks, often misinterpret straightforward visual concep

Distilling Self-Consistency into Verbal Confidence: A Pre-Registered Negative Result and Post-Hoc Rescue on Gemma 3 4B

Model ReleasesDGX agent

arXiv:2604.24070v1 Announce Type: cross Abstract: Small instruct-tuned LLMs produce degenerate verbal confidence under minimal elicitation: ceiling rates above 95%, near-chance Type-2 AUROC, and Inval

DLM: Unified Decision Language Models for Offline Multi-Agent Sequential Decision Making

SafetyDGX agent

arXiv:2604.23557v1 Announce Type: cross Abstract: Building scalable and reusable multi-agent decision policies from offline datasets remains a challenge in offline multi-agent reinforcement learning (

DO-Bench: An Attributable Benchmark for Diagnosing Object Hallucination in Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.22822v1 Announce Type: cross Abstract: Object level hallucination remains a central reliability challenge for vision language models (VLMs), particularly in binary object existence verifica

Do Quantum Transformers Help? A Systematic VQC Architecture Comparison on Tabular Benchmarks

Model ReleasesDGX agent

arXiv:2604.23931v1 Announce Type: cross Abstract: Variational quantum circuits (VQCs) are a leading approach to quantum machine learning on near-term devices, yet it remains unclear which circuit arch

Do Transaction-Level and Actor-Level AML Queues Agree? An Empirical Evaluation of Granularity Effects on the Elliptic++ Graph

SafetyDGX agent

arXiv:2604.23494v1 Announce Type: new Abstract: Graph-based anti-money laundering (AML) systems on blockchain networks can score suspicious activity at two granularity levels -- transactions or actor

Does Machine Unlearning Preserve Clinical Safety? A Risk Analysis for Medical Image Classification

SafetyDGX agent

arXiv:2604.23854v1 Announce Type: new Abstract: The application of Deep Learning in medical diagnosis must balance patient safety with compliance with data protection regulations. Machine Unlearning e

Domain-Filtered Knowledge Graphs from Sparse Autoencoder Features

Local AiDGX agent

arXiv:2604.23829v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) extract millions of interpretable features from a language model, but flat feature inventories aren't very useful on their ow

Don't Make the LLM Read the Graph: Make the Graph Think

Model ReleasesDGX agent

arXiv:2604.23057v1 Announce Type: new Abstract: We investigate whether explicit belief graphs improve LLM performance in cooperative multi-agent reasoning. Through 3,000+ controlled trials across four

DPRM: A Plug-in Doob h transform-induced Token-Ordering Module for Diffusion Language Models

SafetyDGX agent

arXiv:2604.24357v1 Announce Type: cross Abstract: Diffusion language models generate without a fixed left-to-right order, making token ordering a central algorithmic choice: which tokens should be rev

Dr. RTL: Autonomous Agentic RTL Optimization through Tool-Grounded Self-Improvement

AgentsDGX agent

arXiv:2604.14989v2 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have sparked growing interest in automatic RTL optimization for better performance, power, and area

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models

Model ReleasesDGX agent

arXiv:2509.06027v3 Announce Type: replace-cross Abstract: With the development of large-scale diffusion-based and language-modeling-based generative models, impressive progress has been achieved in te

Dual-domain Multi-path Self-supervised Diffusion Model for Accelerated MRI Reconstruction

ApplicationsDGX agent

arXiv:2503.18836v2 Announce Type: replace-cross Abstract: Magnetic resonance imaging (MRI) is a vital diagnostic tool, but its inherently long acquisition times reduce clinical efficiency and patient

DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training

SafetyDGX agent

arXiv:2512.03847v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has shown strong performance in LLM post-training, but real-world deployment often involves noisy or incomplete su

DyABD: The Abdominal Muscle Segmentation in Dynamic MRI Benchmark

Model ReleasesDGX agent

arXiv:2604.23187v1 Announce Type: cross Abstract: This work introduces DyABD, a novel and complex benchmark dataset of dynamic abdominal MRIs from patients with abdominal hernias and associated high q

EAD-Net: Emotion-Aware Talking Head Generation with Spatial Refinement and Temporal Coherence

SafetyDGX agent

arXiv:2604.23325v1 Announce Type: cross Abstract: Emotionally talking head video generation aims to generate expressive portrait videos with accurate lip synchronization and emotional facial expressio

Early Warning of Intraoperative Adverse Events via Transformer-Driven Multi-Label Learning

SafetyDGX agent

arXiv:2603.05212v2 Announce Type: replace-cross Abstract: Early warning of intraoperative adverse events plays a vital role in reducing surgical risk and improving patient safety. While deep learning

ECoLAD: Deployment-Oriented Evaluation for Automotive Time-Series Anomaly Detection

ResearchDGX agent

arXiv:2603.10926v1 Announce Type: cross Abstract: Time-series anomaly detectors are commonly compared on workstation-class hardware under unconstrained execution. In-vehicle monitoring, however, requi

EL3DD: Extended Latent 3D Diffusion for Language Conditioned Multitask Manipulation

SafetyDGX agent

arXiv:2511.13312v2 Announce Type: replace-cross Abstract: Acting in human environments is a crucial capability for general-purpose robots, necessitating a robust understanding of natural language and

EmoBench-M: Benchmarking Emotional Intelligence for Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2502.04424v4 Announce Type: replace-cross Abstract: With the integration of multimodal large language models (MLLMs) into robotic systems and AI applications, embedding emotional intelligence (E

Emotion-Conditioned Short-Horizon Human Pose Forecasting with a Lightweight Predictive World Model

ResearchDGX agent

arXiv:2604.23532v1 Announce Type: cross Abstract: Short-term human pose prediction plays a crucial role in interactive systems, assistive robots, and emotion-aware human-computer interaction[1-3]. Whi

EmoTrans: A Benchmark for Understanding, Reasoning, and Predicting Emotion Transitions in Multimodal LLMs

Model ReleasesDGX agent

arXiv:2604.23348v1 Announce Type: cross Abstract: Recent multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and generation, and are increasingly used in

Empirical Ablation and Ensemble Optimization of a Convolutional Neural Network for CIFAR-10 Classification

Model ReleasesDGX agent

arXiv:2604.23861v1 Announce Type: cross Abstract: Convolutional neural networks (CNNs) remain a central approach in image classification, but their performance depends strongly on architectural and tr

Enabling Transparent Cyber Threat Intelligence Combining Large Language Models and Domain Ontologies

AgentsDGX agent

arXiv:2509.00081v2 Announce Type: replace-cross Abstract: Effective Cyber Threat Intelligence (CTI) relies upon accurately structured and semantically enriched information extracted from cybersecurity

End-to-End Learning for Partially-Observed Time Series with PyPOTS

Model ReleasesDGX agent

arXiv:2604.24041v1 Announce Type: cross Abstract: Partially-observed time series (POTS) is ubiquitous in real-world applications, yet most existing toolchains separate missing-value handling from down

Energy-Aware Routing to Large Reasoning Models

ResearchDGX agent

arXiv:2601.00823v2 Announce Type: replace Abstract: Large reasoning models (LRMs) have heterogeneous inference energy costs based on which model is used and how much it reasons. To reduce energy, it i

Epicure: Multidimensional Flavor Structure in Food Ingredient Embeddings

ResearchDGX agent

arXiv:2604.22776v1 Announce Type: cross Abstract: A chef's intuition about flavor, texture, and cultural identity represents tacit knowledge that is difficult to articulate yet central to culinary pra

EPM-RL: Reinforcement Learning for On-Premise Product Mapping in E-Commerce

Model ReleasesDGX agent

arXiv:2604.23993v1 Announce Type: cross Abstract: Product mapping, the task of deciding whether two e-commerce listings refer to the same product, is a core problem for price monitoring and channel vi

Escher-Loop: Mutual Evolution by Closed-Loop Self-Referential Optimization

AgentsDGX agent

arXiv:2604.23472v1 Announce Type: new Abstract: While recent autonomous agents demonstrate impressive capabilities, they predominantly rely on manually scripted workflows and handcrafted heuristics, i

ESIA: An Energy-Based Spatiotemporal Interaction-Aware Framework for Pedestrian Intention Prediction

AgentsDGX agent

arXiv:2604.23728v1 Announce Type: cross Abstract: Recent advances in autonomous driving have motivated research on pedestrian intention prediction, which aims to infer future crossing decisions and ac

ESPADA: Execution Speedup via Semantics Aware Demonstration Data Downsampling for Imitation Learning

ApplicationsDGX agent

arXiv:2512.07371v3 Announce Type: replace-cross Abstract: Behavior-cloning based visuomotor policies enable precise manipulation but often inherit the slow, cautious tempo of human demonstrations, lim

Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs

Model ReleasesDGX agent

arXiv:2604.23466v1 Announce Type: cross Abstract: NVIDIA's CUDA Tile (CuTile) introduces a Python-based, tile-centric abstraction for GPU kernel development that aims to simplify programming while ret

Evaluating Jailbreaking Vulnerabilities in LLMs Deployed as Assistants for Smart Grid Operations: A Benchmark Against NERC Standards

Model ReleasesDGX agent

arXiv:2604.23341v1 Announce Type: cross Abstract: The deployment of Large Language Models (LLMs) as assistants in electric grid operations promises to streamline compliance and decision-making but exp

Evaluating Language Models' Evaluations of Games

SafetyDGX agent

arXiv:2510.10930v2 Announce Type: replace-cross Abstract: Reasoning is not just about solving problems -- it is also about evaluating which problems are worth solving at all. Evaluations of artificial

Evaluating the Search Agent in a Parallel World

Model ReleasesDGX agent

arXiv:2603.04751v2 Announce Type: replace Abstract: Integrating web search tools has significantly extended the capability of LLMs to address open-world, real-time, and long-tail problems. However, ev

Evaluating whether AI models would sabotage AI safety research

Model ReleasesDGX agent

arXiv:2604.24618v1 Announce Type: new Abstract: We evaluate the propensity of frontier models to sabotage or refuse to assist with safety research when deployed as AI research agents within a frontier

Evaluation of Prompt Injection Defenses in Large Language Models

ResearchDGX agent

arXiv:2604.23887v1 Announce Type: cross Abstract: LLM-powered applications routinely embed secrets in system prompts, yet models can be tricked into revealing them. We built an adaptive attacker that

Expert Evaluation of LLM's Open-Ended Legal Reasoning on the Japanese Bar Exam Writing Task

ApplicationsDGX agent

arXiv:2604.23730v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong performance on legal benchmarks, including multiple-choice components of bar exams. However, their capaci

Explainable AI for Mental Disorder Detection via Social Media: A survey and outlook

TutorialsDGX agent

arXiv:2406.05984v2 Announce Type: replace-cross Abstract: Mental health constitutes a complex and pervasive global challenge, affecting millions of lives and often leading to severe consequences. In t

Explainable AI in Speaker Recognition -- Making Latent Representations Understandable

TutorialsDGX agent

arXiv:2604.23354v1 Announce Type: cross Abstract: Neural networks can be trained to learn task-relevant representations from data. Understanding how these networks make decisions falls within the Expl

Explainable Artificial Intelligence Techniques for Interpretation of Food Models: a Review

TutorialsDGX agent

arXiv:2504.10527v2 Announce Type: replace Abstract: Artificial Intelligence (AI) has become essential for analyzing complex data and solving highly-challenging tasks. It is being applied across numero

Explanation Quality Assessment as Ranking with Listwise Rewards

SafetyDGX agent

arXiv:2604.24176v1 Announce Type: new Abstract: We reformulate explanation quality assessment as a ranking problem rather than a generation problem. Instead of optimizing models to produce a single 'b

Exploring Audio Hallucination in Egocentric Video Understanding

ResearchDGX agent

arXiv:2604.23860v1 Announce Type: cross Abstract: Egocentric videos provide a distinctive setting in which sound serves as crucial cues to understand user activities and surroundings, particularly whe

Exploring the Secondary Risks of Large Language Models

Model ReleasesDGX agent

arXiv:2506.12382v5 Announce Type: replace-cross Abstract: Ensuring the safety and alignment of Large Language Models is a significant challenge with their growing integration into critical application

Extending Precipitation Nowcasting Horizons via Spectral Fusion of Radar Observations and Foundation Model Priors

SafetyDGX agent

arXiv:2603.21768v3 Announce Type: replace-cross Abstract: Precipitation nowcasting is critical for disaster mitigation and aviation safety. However, radar-only models frequently suffer from a lack of

EyeBrain: Left and Right Brain Lateralization Activity Classification Through Pupil Diameter and Fixation Duration

ResearchDGX agent

arXiv:2604.23562v1 Announce Type: cross Abstract: The relationship between brain lateralization and cognitive functions is well-documented. The left hemisphere primarily handles tasks such as language

Failure-Centered Runtime Evaluation for Deployed Trilingual Public-Space Agents

SafetyDGX agent

arXiv:2604.23990v1 Announce Type: new Abstract: This paper presents PSA-Eval, a failure-centered runtime evaluation framework for deployed trilingual public-space agents. The central claim is that, wh

FAIR_XAI: Improving Multimodal Foundation Model Fairness via Explainability for Wellbeing Assessment

Model ReleasesDGX agent

arXiv:2604.23786v1 Announce Type: new Abstract: In recent years, the integration of multimodal machine learning in wellbeing assessment has offered transformative potential for monitoring mental healt

FastOMOP: A Foundational Architecture for Reliable Agentic Real-World Evidence Generation on OMOP CDM data

SafetyDGX agent

arXiv:2604.24572v1 Announce Type: new Abstract: The Observational Medical Outcomes Partnership Common Data Model (OMOP CDM), maintained by the Observational Health Data Sciences and Informatics (OHDSI

Federated Cross-Modal Retrieval with Missing Modalities via Semantic Routing and Adapter Personalization

SafetyDGX agent

arXiv:2604.22885v1 Announce Type: cross Abstract: Federated cross-modal retrieval faces severe challenges from heterogeneous client data, particularly non-IID semantic distributions and missing modali

FedRef: Bayesian Fine-Tuning using a Reference Model to Mitigate Catastrophic Forgetting for Heterogeneous Federated Learning

ApplicationsDGX agent

arXiv:2506.23210v5 Announce Type: replace-cross Abstract: Federated learning (FL) enables collaborative model training across distributed clients while preserving data privacy. However, data and syste

← Previous
1…297298299300301…354
Next →