AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
23 Apr 2026

ThermoQA: A Three-Tier Benchmark for Evaluating Thermodynamic Reasoning in Large Language Models

Model ReleasesDGX agent

arXiv:2604.19758v1 Announce Type: new Abstract: We present ThermoQA, a benchmark of 293 open-ended engineering thermodynamics problems in three tiers: property lookups (110 Q), component analysis (101

Thinking While Listening: Fast-Slow Recurrence for Long-Horizon Sequential Modeling

ResearchDGX agent

arXiv:2604.01577v2 Announce Type: replace-cross Abstract: We extend the recent latent recurrent modeling to sequential input streams. By interleaving fast, recurrent latent updates with self-organizat

Tokenised Flow Matching for Hierarchical Simulation Based Inference

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.20723v1 Announce Type: cross Abstract: The cost of simulator evaluations is a key practical bottleneck for Simulation Based Inference (SBI). In hierarchical settings with shared global para

Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection

ResearchDGX agent

arXiv:2604.20549v1 Announce Type: cross Abstract: As Large Language Models (LLMs) scale, data curation has shifted from maximizing volume to optimizing the signal-to-noise ratio by performing quality

Towards Secure Logging: Characterizing and Benchmarking Logging Code Security Issues with LLMs

Model ReleasesDGX agent

arXiv:2604.20211v1 Announce Type: cross Abstract: Logging code plays an important role in software systems by recording key events and behaviors, which are essential for debugging and monitoring. Howe

Transformers Can Learn Connectivity in Some Graphs but Not Others

TutorialsDGX agent

arXiv:2509.22343v2 Announce Type: replace-cross Abstract: Reasoning capability is essential to ensure the factual correctness of the responses of transformer-based Large Language Models (LLMs), and ro

Transparent Screening for LLM Inference and Training Impacts

ResearchDGX agent

arXiv:2604.19757v1 Announce Type: cross Abstract: This paper presents a transparent screening framework for estimating inference and training impacts of current large language models under limited obs

Treatment, evidence, imitation, and chat

ResearchDGX agent

arXiv:2506.23040v4 Announce Type: replace-cross Abstract: Large language models are thought to have the potential to aid in medical decision making. This work investigates the degree to which this mig

TriEx: A Game-based Tri-View Framework for Explaining Internal Reasoning in Multi-Agent LLMs

AgentsDGX agent

arXiv:2604.20043v1 Announce Type: cross Abstract: Explainability for Large Language Model (LLM) agents is especially challenging in interactive, partially observable settings, where decisions depend o

Trust, Lies, and Long Memories: Emergent Social Dynamics and Reputation in Multi-Round Avalon with LLM Agents

AgentsDGX agent

arXiv:2604.20582v1 Announce Type: cross Abstract: We study emergent social dynamics in LLM agents playing The Resistance: Avalon, a hidden-role deception game. Unlike prior work on single-game perform

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference

ResearchDGX agent

arXiv:2604.19769v1 Announce Type: cross Abstract: Key-value (KV) caching is critical for efficient inference in large language models (LLMs), yet its memory footprint scales linearly with context leng

UCCL-Zip: Lossless Compression Supercharged GPU Communication

Model ReleasesDGX agent

arXiv:2604.17172v2 Announce Type: replace-cross Abstract: The rapid growth of large language models (LLMs) has made GPU communication a critical bottleneck. While prior work reduces communication volu

uLEAD-TabPFN: Uncertainty-aware Dependency-based Anomaly Detection with TabPFN

ResearchDGX agent

arXiv:2604.20255v1 Announce Type: cross Abstract: Anomaly detection in tabular data is challenging due to high dimensionality, complex feature dependencies, and heterogeneous noise. Many existing meth

Using Learning Theories to Evolve Human-Centered XAI: Future Perspectives and Challenges

ResearchDGX agent

arXiv:2604.19788v1 Announce Type: new Abstract: As Artificial Intelligence (AI) systems continue to grow in size and complexity, so does the difficulty of the quest for AI transparency. In a world of

Utterance-Level Methods for Identifying Reliable ASR-Output for Child Speech

ResearchDGX agent

arXiv:2604.19801v1 Announce Type: cross Abstract: Automatic Speech Recognition (ASR) is increasingly used in applications involving child speech, such as language learning and literacy acquisition. Ho

V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization

SafetyDGX agent

arXiv:2604.20755v1 Announce Type: new Abstract: We introduce V-tableR1, a process-supervised reinforcement learning framework that elicits rigorous, verifiable reasoning from multimodal large language

VAN-AD: Visual Masked Autoencoder with Normalizing Flow For Time Series Anomaly Detection

ApplicationsDGX agent

arXiv:2603.26842v2 Announce Type: replace-cross Abstract: Time series anomaly detection (TSAD) is essential for maintaining the reliability and security of IoT-enabled service systems. Existing method

veScale-FSDP: Flexible and High-Performance FSDP at Scale

ResearchDGX agent

arXiv:2602.22437v3 Announce Type: replace-cross Abstract: Fully Sharded Data Parallel (FSDP), also known as Zero Redundancy Optimizer (ZeRO), is widely used for large-scale model training, because of

Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback

Model ReleasesDGX agent

arXiv:2604.20210v1 Announce Type: cross Abstract: Individual differences in vibrotactile perception underscore the growing importance of personalization as haptic feedback becomes more prevalent in in

VTouch++: A Multimodal Dataset with Vision-Based Tactile Enhancement for Bimanual Manipulation

ApplicationsDGX agent

arXiv:2604.20444v1 Announce Type: cross Abstract: Embodied intelligence has advanced rapidly in recent years; however, bimanual manipulation-especially in contact-rich tasks remains challenging. This

What Makes a Good AI Review? Concern-Level Diagnostics for AI Peer Review

SafetyDGX agent

arXiv:2604.19998v1 Announce Type: new Abstract: Evaluating AI-generated reviews by verdict agreement is widely recognized as insufficient, yet current alternatives rarely audit which concerns a system

Where and What: Reasoning Dynamic and Implicit Preferences in Situated Conversational Recommendation

SafetyDGX agent

arXiv:2604.20749v1 Announce Type: new Abstract: Situated conversational recommendation (SCR), which utilizes visual scenes grounded in specific environments and natural language dialogue to deliver co

White-Basilisk: A Hybrid Model for Code Vulnerability Detection

Model ReleasesDGX agent

arXiv:2507.08540v5 Announce Type: replace-cross Abstract: The proliferation of software vulnerabilities presents a significant challenge to cybersecurity, necessitating more effective detection method

Why AI-Generated Text Detection Fails: Evidence from Explainable AI Beyond Benchmark Accuracy

Model ReleasesDGX agent

arXiv:2603.23146v2 Announce Type: replace-cross Abstract: The widespread adoption of Large Language Models (LLMs) has made the detection of AI-Generated text a pressing and complex challenge. Although

WorkflowGen:an adaptive workflow generation mechanism driven by trajectory experience

Model ReleasesDGX agent

arXiv:2604.19756v1 Announce Type: cross Abstract: Large language model (LLM) agents often suffer from high reasoning overhead, excessive token consumption, unstable execution, and inability to reuse p

Working Memory Constraints Scaffold Learning in Transformers under Data Scarcity

SafetyDGX agent

arXiv:2604.20789v1 Announce Type: cross Abstract: We investigate the integration of human-like working memory constraints into the Transformer architecture and implement several cognitively inspired a

22 Apr 2026

3D Foundation Model for Generalizable Disease Detection in Head Computed Tomography

Model ReleasesDGX agent

arXiv:2502.02779v3 Announce Type: replace-cross Abstract: Head computed tomography (CT) imaging is a widely-used imaging modality with multitudes of medical indications, particularly in assessing path

A Dual Perspective on Synthetic Trajectory Generators: Utility Framework and Privacy Vulnerabilities

ResearchDGX agent

arXiv:2604.19653v1 Announce Type: new Abstract: Human mobility data are used in numerous applications, ranging from public health to urban planning. Human mobility is inherently sensitive, as it can c

A Functionality-Grounded Benchmark for Evaluating Web Agents in E-commerce Domains

Model ReleasesDGX agent

arXiv:2508.15832v2 Announce Type: replace-cross Abstract: Web agents have shown great promise in performing many tasks on ecommerce website. To assess their capabilities, several benchmarks have been

A-MAR: Agent-based Multimodal Art Retrieval for Fine-Grained Artwork Understanding

Model ReleasesDGX agent

arXiv:2604.19689v1 Announce Type: new Abstract: Understanding artworks requires multi-step reasoning over visual content and cultural, historical, and stylistic context. While recent multimodal large

A neural operator framework for data-driven discovery of stability and receptivity in physical systems

ResearchDGX agent

arXiv:2604.19465v1 Announce Type: cross Abstract: Understanding how complex systems respond to perturbations, such as whether they will remain stable or what their most sensitive patterns are, is a fu

A Proxy Consistency Loss for Grounded Fusion of Earth Observation and Location Encoders

TutorialsDGX agent

arXiv:2604.18881v1 Announce Type: cross Abstract: Supervised learning with Earth observation inputs is often limited by the sparsity of high-quality labeled or in-situ measured data to use as training

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends

AgentsDGX agent

arXiv:2507.09861v2 Announce Type: replace-cross Abstract: Visually Rich Document Understanding (VRDU) has become a pivotal area of research, driven by the need to automatically interpret documents tha

AblateCell: A Reproduce-then-Ablate Agent for Virtual Cell Repositories

AgentsDGX agent

arXiv:2604.19606v1 Announce Type: new Abstract: Systematic ablations are essential to attribute performance gains in AI Virtual Cells, yet they are rarely performed because biological repositories are

AD-Copilot: A Vision-Language Assistant for Industrial Anomaly Detection via Visual In-context Comparison

Model ReleasesDGX agent

arXiv:2603.13779v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have achieved impressive success in natural visual understanding, yet they consistently underperform

Adapting Dijkstra for Buffers and Unlimited Transfers

ResearchDGX agent

arXiv:2603.11729v2 Announce Type: replace-cross Abstract: In recent years, RAPTOR based algorithms have been considered the state-of-the-art for path-finding with unlimited transfers without preproces

Adaptive MSD-Splitting: Enhancing C4.5 and Random Forests for Skewed Continuous Attributes

ApplicationsDGX agent

arXiv:2604.19722v1 Announce Type: cross Abstract: The discretization of continuous numerical attributes remains a persistent computational bottleneck in the induction of decision trees, particularly a

Adaptive Prompt Elicitation for Text-to-Image Generation

SafetyDGX agent

arXiv:2602.04713v2 Announce Type: replace-cross Abstract: Aligning text-to-image generation with user intent remains challenging, as users frequently provide ambiguous inputs and struggle with model i

Adversarial Attacks on Medical Hyperspectral Imaging Exploiting Spectral-Spatial Dependencies and Multiscale Features

ResearchDGX agent

arXiv:2601.07056v2 Announce Type: replace-cross Abstract: Medical hyperspectral imaging (MHSI) has shown strong potential for disease diagnosis by capturing spectral-spatial information of tissues. Wh

Agent-GWO: Collaborative Agents for Dynamic Prompt Optimization in Large Language Models

Model ReleasesDGX agent

arXiv:2604.18612v1 Announce Type: cross Abstract: Large Language Models (LLMs) have demonstrated strong capabilities in complex reasoning tasks, while recent prompting strategies such as Chain-of-Thou

AgentDynEx: Nudging the Mechanics and Dynamics of Multi-Agent Simulations

AgentsDGX agent

arXiv:2504.09662v3 Announce Type: replace-cross Abstract: Multi-agent large language model simulations have the potential to model complex human behaviors and interactions. If the mechanics are set up

Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs

Model ReleasesDGX agent

arXiv:2604.18576v2 Announce Type: replace Abstract: We present BLF (Bayesian Linguistic Forecaster), an agentic system for binary forecasting that achieves state-of-the-art performance on the Forecast

AI-Based Detection of Temporal Changes in MR-Linac Images Acquired During Routine Prostate Radiotherapy

ResearchDGX agent

arXiv:2602.04983v2 Announce Type: replace-cross Abstract: Purpose: To investigate whether an AI-based method can detect subtle inter-fraction changes in MR-Linac images acquired during radiotherapy an

AI scientists produce results without reasoning scientifically

AgentsDGX agent

arXiv:2604.18805v1 Announce Type: new Abstract: Large language model (LLM)-based systems are increasingly deployed to conduct scientific research autonomously, yet whether their reasoning adheres to t

An AI Agent Execution Environment to Safeguard User Data

AgentsDGX agent

arXiv:2604.19657v1 Announce Type: cross Abstract: AI agents promise to serve as general-purpose personal assistants for their users, which requires them to have access to private user data (e.g., pers

ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System

SafetyDGX agent

arXiv:2604.18789v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) is central to aligning Large Language Models (LLMs), yet it introduces a critical vulnerability: an im

ARGUS: Agentic GPU Optimization Guided by Data-Flow Invariants

HardwareDGX agent

arXiv:2604.18616v1 Announce Type: cross Abstract: LLM-based coding agents can generate functionally correct GPU kernels, yet their performance remains far below hand-optimized libraries on critical co

ARM: Advantage Reward Modeling for Long-Horizon Manipulation

SafetyDGX agent

arXiv:2604.03037v2 Announce Type: replace-cross Abstract: Long-horizon robotic manipulation remains challenging for reinforcement learning (RL) because sparse rewards provide limited guidance for cred

Assessing Capabilities of Large Language Models in Social Media Analytics: A Multi-task Quest

Model ReleasesDGX agent

arXiv:2604.18955v1 Announce Type: cross Abstract: In this study, we present the first comprehensive evaluation of modern LLMs - including GPT-4, GPT-4o, GPT-3.5-Turbo, Gemini 1.5 Pro, DeepSeek-V3, Lla

Attention-based Multi-modal Deep Learning Model of Spatio-temporal Crop Yield Prediction with Satellite, Soil and Climate Data

SafetyDGX agent

arXiv:2604.19217v1 Announce Type: cross Abstract: Crop yield prediction is one of the most important challenge, which is crucial to world food security and policy-making decisions. The conventional fo

AutoAWG: Adverse Weather Generation with Adaptive Multi-Controls for Automotive Videos

SafetyDGX agent

arXiv:2604.18993v1 Announce Type: cross Abstract: Perception robustness under adverse weather remains a critical challenge for autonomous driving, with the core bottleneck being the scarcity of real-w

Autogenesis: A Self-Evolving Agent Protocol

AgentsDGX agent

arXiv:2604.15034v2 Announce Type: replace Abstract: Recent advances in LLM based agent systems have shown promise in tackling complex, long horizon tasks. However, existing agent protocols (e.g., A2A

AutomationBench

Model ReleasesDGX agent

arXiv:2604.18934v1 Announce Type: new Abstract: Existing AI benchmarks for software automation rarely combine cross-application coordination, autonomous API discovery, and policy adherence. Real busin

BAPO: Boundary-Aware Policy Optimization for Reliable Agentic Search

SafetyDGX agent

arXiv:2601.11037v2 Announce Type: replace Abstract: RL-based agentic search enables LLMs to solve complex questions via dynamic planning and external search. While this approach significantly enhances

BEAT: Tokenizing and Generating Symbolic Music by Uniform Temporal Steps

ResearchDGX agent

arXiv:2604.19532v1 Announce Type: cross Abstract: Tokenizing music to fit the general framework of language models is a compelling challenge, especially considering the diverse symbolic structures in

BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental Design

ResearchDGX agent

arXiv:2508.21184v3 Announce Type: replace-cross Abstract: We propose a general-purpose approach for improving the ability of large language models (LLMs) to intelligently and adaptively gather informa

Benchmarking Misuse Mitigation Against Covert Adversaries

SafetyDGX agent

arXiv:2506.06414v2 Announce Type: replace-cross Abstract: Existing language model safety evaluations focus on overt attacks and low-stakes tasks. In reality, an attacker can easily subvert existing sa

Benign Overfitting in Adversarial Training for Vision Transformers

ApplicationsDGX agent

arXiv:2604.19724v1 Announce Type: cross Abstract: Despite the remarkable success of Vision Transformers (ViTs) across a wide range of vision tasks, recent studies have revealed that they remain vulner

Best Agent Identification for General Game Playing

AgentsDGX agent

arXiv:2507.00451v2 Announce Type: replace-cross Abstract: We present an efficient and generalised procedure to accurately identify the best (or near best) performing algorithm for each sub-task in a m

Beyond Coefficients: Forecast-Necessity Testing for Interpretable Causal Discovery in Nonlinear Time-Series Models

ApplicationsDGX agent

arXiv:2604.18751v1 Announce Type: cross Abstract: Nonlinear machine-learning models are increasingly used to discover causal relationships in time-series data, yet the interpretation of their outputs

← Previous
1…315316317318319…354
Next →