AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
8 Jun 2026

How Language Models Fail: Token-Level Signatures of Committed and Persistent Reasoning Failures

ResearchDGX agent

arXiv:2606.06635v1 Announce Type: cross Abstract: Failures in language model reasoning emerge through distinct processes that leave identifiable signatures in the reasoning trace. We characterize thes

How reliable are LLMs when it comes to playing dice?

SafetyDGX agent

arXiv:2606.07515v1 Announce Type: cross Abstract: We investigate the probabilistic reasoning capabilities of large language models through a controlled benchmarking study on discrete probability probl

HybridCodec: Fast Dual-Stream, Semantically Enhanced Neural Audio Codec

ResearchDGX agent

arXiv:2606.06743v1 Announce Type: cross Abstract: The popularity of neural audio codecs as speech tokenizers has surged with the advent of Multimodal Large Language Models. New codec architectures wit


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Impact of Synthetic Lesional MR Images in Automated Focal Cortical Dysplasia Detection in Low-Data Scenarios

ResearchDGX agent

arXiv:2606.07381v1 Announce Type: cross Abstract: Background and Purpose: Automated detection of focal cortical dysplasia (FCD) requires large volumes of voxelwise lesion-delineated MRI data, which ar

Inside the Visual Mind: Neuroscience-Motivated Concept Circuits for Interpreting and Steering Vision Transformers

ResearchDGX agent

arXiv:2606.06664v1 Announce Type: cross Abstract: Despite high accuracy, Vision Transformer (ViT) predictions can be driven by spurious cues, raising the need to understand their inner workings before

IRAF: Interference-Resilient Adaptive Fusion for Noise-Robust End-to-End Full-Duplex Spoken Dialogue Systems

AgentsDGX agent

arXiv:2606.06559v1 Announce Type: cross Abstract: Full-duplex spoken dialogue models allow voice agents to listen and speak concurrently, enabling natural interaction with real-time overlap. However,

It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents

Model ReleasesDGX agent

arXiv:2512.23128v2 Announce Type: replace-cross Abstract: Web-based agents powered by large language models are increasingly used for tasks such as email management or professional networking. Their r

Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates

SafetyDGX agent

arXiv:2601.18510v2 Announce Type: replace-cross Abstract: While Large Language Model (LLM) agents excel at general tasks, they inherently struggle with continual adaptation due to the frozen weights a

Lane Change Trajectory Planning for Personalized Driving Comfort and Mobility Efficiency

ResearchDGX agent

arXiv:2606.06805v1 Announce Type: cross Abstract: Lane changing entails simultaneous longitudinal and lateral motions that affect driving comfort and mobility efficiency. Because these motions are tig

Latent-space Attacks for Refusal Evasion in Language Models

SafetyDGX agent

arXiv:2605.21706v2 Announce Type: replace Abstract: Safety-aligned language models are trained to refuse harmful requests, yet refusal behavior can be suppressed by steering their internal representat

Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory

Model ReleasesDGX agent

arXiv:2606.06523v1 Announce Type: new Abstract: Equipping Large Language Models (LLMs) to execute reliable multi-step workflows has become a central challenge in artificial intelligence. Despite recen

Learning to Execute Graph Algorithms Exactly with Graph Neural Networks

Local AiDGX agent

arXiv:2601.23207v2 Announce Type: replace-cross Abstract: Understanding what graph neural networks can learn, especially their ability to learn to execute algorithms, remains a central theoretical cha

Limitations of Normalization in Attention Mechanism

ResearchDGX agent

arXiv:2508.17821v3 Announce Type: replace-cross Abstract: This paper investigates the limitations of the normalization in attention mechanisms. We begin with a theoretical framework that enables the i

LLM Agent-Assisted Reverse Engineering with Quantitative Readability Metrics

AgentsDGX agent

arXiv:2606.06838v1 Announce Type: cross Abstract: Automatic decompilers produce functionally correct but often unreadable C code. This paper addresses one stage of the reverse engineering workflow: im

LLM-Augmented Digital Twin for Policy Evaluation in Short-Video Platforms

SafetyDGX agent

arXiv:2603.11333v2 Announce Type: replace Abstract: Short-video platforms are closed-loop, human-in-the-loop ecosystems where platform policy, creator incentives, and user behavior co-evolve. This fee

LLM-Guided Search for Deletion-Correcting Codes

ResearchDGX agent

arXiv:2504.00613v2 Announce Type: replace Abstract: Finding deletion-correcting codes of maximum size has been an open problem for over 70 years, even for a single deletion. We adapt FunSearch, a larg

LuMamba: Latent Unified Mamba for Electrode Topology-Invariant and Efficient EEG Modeling

HardwareDGX agent

arXiv:2603.19100v2 Announce Type: replace Abstract: Electroencephalography (EEG) enables non-invasive monitoring of brain activity across clinical and neurotechnology applications, yet building founda

MacArena: Benchmarking Computer Use Agents on an Online macOS Environment

Model ReleasesDGX agent

arXiv:2606.06560v1 Announce Type: cross Abstract: Computer-use agents (CUAs) operate graphical user interfaces (GUIs) through vision and control primitives, and their capabilities have advanced rapidl

MACD: Model-Aware Contrastive Decoding via Counterfactual Data

Model ReleasesDGX agent

arXiv:2602.01740v3 Announce Type: replace Abstract: Video language models (Video-LLMs) are prone to hallucinations, generating plausible but ungrounded content when visual evidence is weak, ambiguous,

MalTree: Tracing Malware Evolution from Embeddings at Scale

ApplicationsDGX agent

arXiv:2606.06570v1 Announce Type: cross Abstract: Malware detection remains largely reactive: machine learning models trained on known samples degrade as threats evolve. Understanding evolutionary rel

MatterDoor: Sampling Zero-shot Spatio-semantic Priors using Generative Models

Model ReleasesDGX agent

arXiv:2510.11014v2 Announce Type: replace-cross Abstract: Autonomous robots often view rooms only partially, through a doorway, where the walls and scene structure hide the geometry and task-relevant

Measuring Agents in Production

AgentsDGX agent

arXiv:2512.04123v4 Announce Type: replace-cross Abstract: LLM-based agents already operate in production across many industries, yet we lack an understanding of what technical methods make deployments

MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism

Model ReleasesDGX agent

arXiv:2606.07512v1 Announce Type: cross Abstract: Current Vision-Language Models struggle with hours-long videos because processing full-length visual sequences induces prohibitive token explosion and

MetaConfigurator: AI-Assisted RDF Authoring from JSON Data

ResearchDGX agent

arXiv:2606.07094v1 Announce Type: cross Abstract: Scientific workflows increasingly generate structured JSON data that is easy to exchange but difficult to interpret consistently across systems due to

MHA-RAG: Improving Efficiency, Accuracy, and Consistency by Encoding Exemplars as Soft Prompts

ResearchDGX agent

arXiv:2510.05363v2 Announce Type: replace Abstract: Adapting Foundation Models to new domains with limited training data is challenging and computationally expensive. While prior work has demonstrated

Mind the Gap: Bridging Behavioral Silos with LLMs in Multi-Vertical Recommendations

ApplicationsDGX agent

arXiv:2606.06779v1 Announce Type: cross Abstract: In multi-vertical e-commerce platforms like DoorDash, relatively newer product verticals such as grocery and retail present a significant opportunity

Mitosis Detection in the Wild: Multi-Tumor and Context-Aware Generalization in the MIDOG 2025 Challenge

ApplicationsDGX agent

arXiv:2606.07368v1 Announce Type: cross Abstract: Automated mitosis detection is a well-established task in computational pathology. While previous benchmarks focused on scanner-induced domain shift,

MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.06696v1 Announce Type: cross Abstract: Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, from detecting lesions in chest X-rays to profiling

MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs

ResearchDGX agent

arXiv:2506.01850v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable success in instruction-following tasks by integrating pretrained visual enco

Model Context Protocols in Adaptive Transport Systems: A Survey

AgentsDGX agent

arXiv:2508.19239v2 Announce Type: replace Abstract: The rapid expansion of interconnected devices, autonomous systems, and AI applications has created severe fragmentation in adaptive transport system

Modeling Nonlinear Feature Interactions with Product-Unit Residual Networks

Model ReleasesDGX agent

arXiv:2606.06861v1 Announce Type: cross Abstract: Understanding nonlinear feature interactions is crucial in science and engineering, yet standard multilayer perceptrons (MLPs) often capture such inte

MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.06853v1 Announce Type: cross Abstract: The new era has witnessed a remarkable capability to extend Vision-Language Models (VLMs) for tackling tasks of video understanding. While current VLM

MSAIC-Net: A Multi-Scale Attention and Imbalance-Aware Contrastive Network for ECG-Based Myocardial Substrate Abnormality Detection

Local AiDGX agent

arXiv:2606.06718v1 Announce Type: cross Abstract: Myocardial substrate abnormalities, such as myocardial scar and myocardial infarction (MI), are associated with adverse cardiovascular outcomes. Elect

Multi-Agent Reasoning with Consistency Verification Improves Uncertainty Calibration in Medical MCQA

AgentsDGX agent

arXiv:2603.24481v2 Announce Type: replace Abstract: Miscalibrated confidence scores are a practical obstacle to deploying AI in clinical settings. A model that is always overconfident offers no useful

Multi-Scale Feature Attention Network for Polymer Classification using THz Dual-Comb Spectroscopy

SafetyDGX agent

arXiv:2606.06554v1 Announce Type: cross Abstract: Reliable polymer identification is essential for ensuring the quality and safety of recycled plastics, yet conventional sorting and spectroscopic tech

Multilingual Multi-Speaker Unit Vocoders: A Systematic Analysis of Discrete Speech Representations

ResearchDGX agent

arXiv:2606.06740v1 Announce Type: cross Abstract: Discrete speech units obtained via k-means clustering of self supervised embeddings entangle phonetic, speaker, and language information, causing spea

MVCL-DAF++: Enhancing Multimodal Intent Recognition via Prototype-Aware Contrastive Alignment and Coarse-to-Fine Dynamic Attention Fusion

SafetyDGX agent

arXiv:2509.17446v3 Announce Type: replace-cross Abstract: Multimodal intent recognition (MMIR) suffers from weak semantic grounding and poor robustness under noisy or rare-class conditions. We propose

Native3D: End-to-End 3D Scene Generation via Unified Mesh-Texture Modeling and Semantic Alignment

SafetyDGX agent

arXiv:2606.07117v1 Announce Type: cross Abstract: This paper presents Native3D, the first end-to-end 3D scene generation framework that completely bypasses 2D intermediate representations. Traditional

Neuro-Symbolic Learning for Long-Horizon Task Planning Under Complex Logical Constraints

SafetyDGX agent

arXiv:2606.06877v1 Announce Type: cross Abstract: Task planning often suffers from severe efficiency bottlenecks when robots must reason over long-horizon action sequences under complex logical constr

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets

Model ReleasesDGX agent

arXiv:2606.07032v1 Announce Type: cross Abstract: Zero-Shot Composed Image Retrieval (ZS-CIR) aims to retrieve a target image based on a query composed of a reference image and a relative caption with

NTILC: Neural Tool Invocation via Learned Compression

Model ReleasesDGX agent

arXiv:2606.06566v1 Announce Type: cross Abstract: Agentic tool-calling language models depend on large registries of callable APIs, functions, and local actions. Placing full tool specifications direc

Off-Policy Evaluation with Strategic Agents via Local Disclosure

Local AiDGX agent

arXiv:2606.07308v1 Announce Type: new Abstract: We study off-policy evaluation (OPE) under strategic behavior where decision subjects (or agents) respond to a decision maker's policy by strategically

OffQ: Taming Structured Outliers in LLM Quantization by Offsetting

ResearchDGX agent

arXiv:2606.07116v1 Announce Type: cross Abstract: Low-bit quantization has been widely adopted to accelerate the inference of large language models (LLMs) by significantly reducing computational cost

OGA-AID: Clinician-in-the-loop AI Report Drafting Assistant for Multimodal Observational Gait Analysis in Post-Stroke Rehabilitation

AgentsDGX agent

arXiv:2604.05360v2 Announce Type: replace-cross Abstract: Gait analysis is essential in post-stroke rehabilitation but remains time-intensive and cognitively demanding, especially when clinicians must

On the Geometry of On-Policy Distillation

Model ReleasesDGX agent

arXiv:2606.07082v1 Announce Type: cross Abstract: On-policy distillation (OPD) is increasingly used to improve large language model reasoning, but its training dynamics remain poorly understood. We ch

On the importance of multiple training seeds for evaluating machine unlearning

TutorialsDGX agent

arXiv:2510.26714v5 Announce Type: replace-cross Abstract: Machine unlearning aims to remove the influence of certain data points from a trained model without costly retraining. Most practical unlearni

Online Pandora's Box for Contextual LLM Cascading

SafetyDGX agent

arXiv:2606.07392v1 Announce Type: new Abstract: Motivated by Large Language Model (LLM) cascading, we propose an online contextual Pandora's Box model for adaptively querying and selecting LLM APIs. I

OpenHalDet: A Unified Benchmark for Hallucination Detection across Diverse Generation Scenarios

Model ReleasesDGX agent

arXiv:2606.06959v1 Announce Type: cross Abstract: Hallucination detection is essential for the reliable deployment of large language models (LLMs). However, existing evaluations face two core challeng

OpenSkill: Open-World Self-Evolution for LLM Agents

AgentsDGX agent

arXiv:2606.06741v1 Announce Type: new Abstract: Self-evolving agents requires adaptation after deployment, but existing approaches assume a usable learning loop, such as curated skills, successful tra

Optimal Rates for Generalization of Gradient Descent Methods with Deep Neural Networks

ResearchDGX agent

arXiv:2606.06764v1 Announce Type: cross Abstract: Recent progress has been made in understanding the statistical generalization performance of gradient descent methods for overparameterized neural net

P-Cast Precision in FP8 Attention: Sink-Induced Collapse and the Optimality of S=2^8

ResearchDGX agent

arXiv:2606.06521v1 Announce Type: cross Abstract: FP8 (E4M3) acceleration for attention computation offers significant throughput gains, but the 3-bit mantissa introduces precision challenges when the

PandaAI: A Practical Agent CQ2 for Neuro-symbolic Data Analysis And Integrated Decision-Making in Quantitative Finance

AgentsDGX agent

arXiv:2606.06823v1 Announce Type: cross Abstract: While deep learning has excelled in various domains, its application to sequential decision-making in finance remains challenging due to the low Signa

PaperFlow: Profiling, Recommending, and Adapting Across Daily Paper Streams

Model ReleasesDGX agent

arXiv:2606.07454v1 Announce Type: cross Abstract: Scientific paper recommendation is typically evaluated as static ranking over a fixed candidate set, yet real scientific reading unfolds as a daily, l

Phonetic Error Analysis of Raw Waveform Acoustic Models

ResearchDGX agent

arXiv:2606.07030v1 Announce Type: cross Abstract: We analyse error patterns of raw waveform acoustic models on TIMIT phone recognition beyond the overall phone error rate (PER). PER is decomposed acro

Planning-aligned Token Compression for Long-Context Autonomous Driving

SafetyDGX agent

arXiv:2606.07464v1 Announce Type: cross Abstract: Monolithic vision-action models represent an emerging paradigm in autonomous driving. However, this architecture produces token sequences that quickly

Position: A Dynamical Systems Perspective is Needed to Advance Time Series Modeling

ResearchDGX agent

arXiv:2602.16864v2 Announce Type: replace-cross Abstract: Time series (TS) modeling has come a long way from early statistical, mainly linear, approaches to the current trend in TS foundation models.

Position: Don't Just 'Fix it in Post': A Science of AI Must Study Training Dynamics

SafetyDGX agent

arXiv:2606.06533v1 Announce Type: new Abstract: What would it mean to have a scientific understanding of AI? Models are not static objects: they are snapshots of time-evolving processes shaped by data

Progress-SQL: Improving Reinforcement Learning for Text-to-SQL via Progressive Rewards

SafetyDGX agent

arXiv:2606.06825v1 Announce Type: cross Abstract: Reinforcement learning has recently shown promise in improving large language models for Text-to-SQL generation, yet existing methods typically optimi

Proxy Reconstruction Pre-training for Ramp Flow Prediction at Highway Interchanges

ApplicationsDGX agent

arXiv:2510.03381v3 Announce Type: replace-cross Abstract: Interchanges are crucial nodes for vehicle transfers between highways, yet the lack of real-time ramp detectors creates blind spots in traffic

Quantum-Inspired Trace-Augmented Evidence Selection for Reasoning over Structured Hypothesis Spaces

Model ReleasesDGX agent

arXiv:2606.06941v1 Announce Type: new Abstract: Large language models (LLMs) now solve a wide range of expert-level exams at or above human level, yet remain brittle on specialised, evidence-intensive

← Previous
1…152153154155156…358
Next →