AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
15 May 2026

Correctness-Aware Repository Filtering Under Maximum Effective Context Window Constraints

Model ReleasesDGX agent

arXiv:2605.14362v1 Announce Type: cross Abstract: Context window efficiency is a practical constraint in large language model (LLM)-based developer tools. Paulsen [12] shows that all tested models deg

COTCAgent: Preventive Consultation via Probabilistic Chain-of-Thought Completion

AgentsDGX agent

arXiv:2605.15016v1 Announce Type: cross Abstract: As large language models empower healthcare, intelligent clinical decision support has developed rapidly. Longitudinal electronic health records (EHR)

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.14084v1 Announce Type: cross Abstract: Code agents must both reason over long-horizon repository state and obey strict tool-use protocols. In paired Instruct/Thinking checkpoints, these cap

Critic-Driven Voronoi-Quantization for Distilling Deep RL Policies to Explainable Models

Model ReleasesDGX agent

arXiv:2605.14897v1 Announce Type: cross Abstract: Despite many successful attempts at explaining Deep Reinforcement Learning policies using distillation, it remains difficult to balance the performanc

Critical Challenges and Guidelines in Evaluating Synthetic Tabular Data: A Systematic Review

ResearchDGX agent

arXiv:2504.18544v3 Announce Type: replace-cross Abstract: Generating synthetic tabular health data is challenging, and evaluating their quality is equally, if not more, complex. This systematic review

CrystalReasoner: Reasoning and RL for Property-Conditioned Crystal Structure Generation

SafetyDGX agent

arXiv:2605.14344v1 Announce Type: new Abstract: Generative modeling has emerged as a promising approach for crystal structure discovery. However, existing LLM-based generative models struggle with low

CUICurate: A GraphRAG-based Framework for Automated Clinical Concept Curation for NLP applications

Model ReleasesDGX agent

arXiv:2602.17949v2 Announce Type: replace-cross Abstract: Background: Clinical named entity recognition tools commonly map free text to Unified Medical Language System (UMLS) Concept Unique Identifier

CuSearch: Curriculum Rollout Sampling via Search Depth for Agentic RAG

SafetyDGX agent

arXiv:2605.11611v2 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for training agentic retrieval-augmented generation (RAG)

Darwin Family: MRI-Trust-Weighted Evolutionary Merging for Training-Free Scaling of Language-Model Reasoning

Model ReleasesDGX agent

arXiv:2605.14386v1 Announce Type: cross Abstract: We present Darwin Family, a framework for training-free evolutionary merging of large language models via gradient-free weight-space recombination. We

Data-Augmented Game Starts for Accelerating Self-Play Exploration in Imperfect Information Games

Model ReleasesDGX agent

arXiv:2605.14379v1 Announce Type: cross Abstract: Finding approximate equilibria for large-scale imperfect-information competitive games such as StarCraft, Dota, and CounterStrike remains computationa

Deceive, Detect, and Disclose: Large Language Models Play Mini-Mafia

Model ReleasesDGX agent

arXiv:2509.23023v3 Announce Type: replace Abstract: Large language models are increasingly deployed in multi-agent settings whose outcomes hinge on social intelligence, motivating evaluations of their

Deciphering Neural Reparameterized Full-Waveform Inversion with Neural Sensitivity Kernel and Wave Tangent Kernel

SafetyDGX agent

arXiv:2605.14370v1 Announce Type: cross Abstract: Full-waveform inversion (FWI) estimates unknown parameters in the wave equation from limited boundary measurements. Recent advances in neural reparame

Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning

ResearchDGX agent

arXiv:2508.01916v3 Announce Type: replace-cross Abstract: Understanding internal representations of neural models is a core interest of mechanistic interpretability. Due to its large dimensionality, t

Decoupling Stability and Plasticity for Multi-Modal Test-Time Adaptation

ResearchDGX agent

arXiv:2603.00574v2 Announce Type: replace-cross Abstract: Adapting pretrained multi-modal models to evolving test-time distributions, known as multi-modal test-time adaptation, presents a significant

Deepchecks: Evaluating Retrieval-Augmented Generation (RAG)

SafetyDGX agent

arXiv:2605.14488v1 Announce Type: new Abstract: Large Language Models (LLMs) augmented with Retrieval-Augmented Generation (RAG) techniques are revolutionizing applications across multiple domains, su

Derivation Prompting: A Logic-Based Method for Improving Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2605.14053v1 Announce Type: cross Abstract: The application of Large Language Models to Question Answering has shown great promise, but important challenges such as hallucinations and erroneous

Descriptor: Distance-Annotated Traffic Perception Question Answering (DTPQA)

Model ReleasesDGX agent

arXiv:2511.13397v2 Announce Type: replace-cross Abstract: The remarkable progress of Vision-Language Models (VLMs) on a variety of tasks has raised interest in their application to automated driving.

Diagnosing Training Inference Mismatch in LLM Reinforcement Learning

SafetyDGX agent

arXiv:2605.14220v1 Announce Type: cross Abstract: Modern LLM RL systems separate rollout generation from policy optimization. These two stages are expected to produce token probabilities that match ex

Dimension-Level Intent Fidelity Evaluation for Large Language Models: Evidence from Structured Prompt Ablation

SafetyDGX agent

arXiv:2605.14517v1 Announce Type: cross Abstract: Holistic evaluation scores capture overall output quality but do not distinguish whether a model reproduced the structural form of a user's request fr

Distribution-Aware Algorithm Design with LLM Agents

Model ReleasesDGX agent

arXiv:2605.14141v1 Announce Type: new Abstract: We study learning when the learned object is executable solver code rather than a predictor. In this setting, correctness is not enough: two solvers may

Distributions as Actions: A Unified Framework for Diverse Action Spaces

SafetyDGX agent

arXiv:2506.16608v3 Announce Type: replace-cross Abstract: We introduce a novel reinforcement learning (RL) framework that treats parameterized action distributions as actions, redefining the boundary

Do Coding Agents Understand Least-Privilege Authorization?

Model ReleasesDGX agent

arXiv:2605.14859v1 Announce Type: cross Abstract: As coding agents gain access to shells, repositories, and user files, least-privilege authorization becomes a prerequisite for safe deployment: an age

Do Language Models Align with Brains? Prediction Scores Are Not Enough

SafetyDGX agent

arXiv:2605.14025v1 Announce Type: cross Abstract: Brain-language model comparisons often interpret neural prediction scores as evidence that model representations capture brain-relevant language compu

Do We Really Need External Tools to Mitigate Hallucinations? SIRA: Shared-Prefix Internal Reconstruction of Attribution

ResearchDGX agent

arXiv:2605.14621v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) often hallucinate when language priors dominate weak or ambiguous visual evidence. Existing contrastive decoding

Does RAG Know When Retrieval Is Wrong? Diagnosing Context Compliance under Knowledge Conflict

Model ReleasesDGX agent

arXiv:2605.14473v1 Announce Type: cross Abstract: The Context-Compliance Regime in Retrieval-Augmented Generation (RAG) occurs when retrieved context dominates the final answer even when it conflicts

Dual Ascent Diffusion for Inverse Problems

ResearchDGX agent

arXiv:2505.17353v2 Announce Type: replace-cross Abstract: Ill-posed inverse problems are fundamental in many domains, ranging from astrophysics to medical imaging. Emerging diffusion models provide a

Dual-Dimensional Consistency: Balancing Budget and Quality in Adaptive Inference-Time Scaling

ResearchDGX agent

arXiv:2605.15100v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable abilities in reasoning. However, maximizing their potential through inference-time scaling fac

DUET: Optimizing Training Data Mixtures via Feedback from Unseen Evaluation Tasks

ApplicationsDGX agent

arXiv:2502.00270v3 Announce Type: replace-cross Abstract: The performance of an LLM depends heavily on the relevance of its training data to the downstream evaluation task. However, in practice, the d

DVMap: Fine-Grained Pluralistic Value Alignment via High-Consensus Demographic-Value Mapping

Model ReleasesDGX agent

arXiv:2605.14420v1 Announce Type: new Abstract: Current Large Language Models (LLMs) typically rely on coarse-grained national labels for pluralistic value alignment. However, such macro-level supervi

Dynamic Latent Routing

ResearchDGX agent

arXiv:2605.14323v1 Announce Type: cross Abstract: We investigate the temporal concatenation of sub-policies in Markov Decision Processes (MDP) with time-varying reward functions. We introduce General

Dynamic Mixed-Precision Routing for Efficient Multi-step LLM Interaction

SafetyDGX agent

arXiv:2602.02711v2 Announce Type: replace Abstract: Large language models (LLMs) achieve strong performance in long-horizon decision-making tasks through multi-step interaction and reasoning at test t

Dynamics of the Transformer Residual Stream: Coupling Spectral Geometry to Network Topology

ApplicationsDGX agent

arXiv:2605.14258v1 Announce Type: cross Abstract: Large language models are remarkably capable, yet how computation propagates through their layers remains poorly understood. A growing line of work tr

Dywave: Event-Aligned Dynamic Tokenization for Heterogeneous IoT Sensing Signal

ApplicationsDGX agent

arXiv:2605.14014v1 Announce Type: cross Abstract: Internet of Things (IoT) systems continuously collect heterogeneous sensing signals from ubiquitous sensors to support intelligent applications such a

Efficient Generative Retrieval for E-commerce Search with Semantic Cluster IDs and Expert-Guided RL

SafetyDGX agent

arXiv:2605.14434v1 Announce Type: cross Abstract: Generative retrieval offers a promising alternative by unifying the fragmented multi-stage retrieval process into a single end-to-end model. However,

Efficient Multi-objective Prompt Optimization via Pure-exploration Bandits

ResearchDGX agent

arXiv:2605.14553v1 Announce Type: cross Abstract: Prompt engineering has become central to eliciting the capabilities of large language models (LLMs). At its core lies prompt selection -- efficiently

Elastic Spiking Transformers for Efficient Gesture Understanding

Model ReleasesDGX agent

arXiv:2605.13869v1 Announce Type: cross Abstract: Spiking Neural Networks (SNNs), particularly Spiking Transformers, offer energy-efficient processing of event-based sensor data for healthcare applica

Emotion-Attended Stateful Memory (EASM):The Architecture for Hyper-Personalization at Scale

ResearchDGX agent

arXiv:2605.14833v1 Announce Type: new Abstract: Current language model systems remain fundamentally stateless across sessions, limiting their ability to personalize interactions over time. While retri

Energy-Efficient Quadruped Locomotion with Compliant Feet

ResearchDGX agent

arXiv:2605.14411v1 Announce Type: cross Abstract: Quadruped robots are often designed with rigid feet to simplify control and maintain stable contact during locomotion. While this approach is straight

Enhanced and Efficient Reasoning in Large Learning Models

ApplicationsDGX agent

arXiv:2605.14036v1 Announce Type: new Abstract: In current Large Language Models we can trust the production of smoothly flowing prose on the basis of the principles of machine learning. However, ther

EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation

Model ReleasesDGX agent

arXiv:2605.15199v1 Announce Type: cross Abstract: Multi-shot video generation extends single-shot generation to coherent visual narratives, yet maintaining consistent characters, objects, and location

Eradicating Negative Transfer in Multi-Physics Foundation Models via Sparse Mixture-of-Experts Routing

Model ReleasesDGX agent

arXiv:2605.15179v1 Announce Type: cross Abstract: Scaling Scientific Machine Learning (SciML) toward universal foundation models is bottlenecked by negative transfer: the simultaneous co-training of d

EVA: Editing for Versatile Alignment against Jailbreaks

Local AiDGX agent

arXiv:2605.14750v1 Announce Type: cross Abstract: Large Language Models (LLMs) and Vision Language Models (VLMs) have demonstrated impressive capabilities but remain vulnerable to jailbreaking attacks

EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration

ResearchDGX agent

arXiv:2605.15042v1 Announce Type: cross Abstract: We propose EverAnimate, an efficient post-training method for long-horizon animated video generation that preserves visual quality and character ident

Evidential Reasoning Advances Interpretable Real-World Disease Screening

ApplicationsDGX agent

arXiv:2605.15171v1 Announce Type: cross Abstract: Disease screening is critical for early detection and timely intervention in clinical practice. However, most current screening models for medical ima

EvolveMem:Self-Evolving Memory Architecture via AutoResearch for LLM Agents

Model ReleasesDGX agent

arXiv:2605.13941v1 Announce Type: cross Abstract: Long-term memory is essential for LLM agents that operate across multiple sessions, yet existing memory systems treat retrieval infrastructure as fixe

Explainable Detection of Depression Status Shifts from User Digital Traces

ResearchDGX agent

arXiv:2605.14995v1 Announce Type: new Abstract: Every day, users generate digital traces (e.g., social media posts, chats, and online interactions) that are inherently timestamped and may reflect aspe

Exploitation of Hidden Context in Dynamic Movement Forecasting: A Neural Network Journey from Recurrent to Graph Neural Networks and General Purpose Transformers

ResearchDGX agent

arXiv:2605.14855v1 Announce Type: cross Abstract: Forecasting within signal processing pipelines is crucial for mitigating delays, particularly in predicting the dynamic movements of objects such as N

ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents

Model ReleasesDGX agent

arXiv:2605.14153v1 Announce Type: cross Abstract: Exploitation is not a binary event. It is a ladder of acquiring progressive capabilities, from executing a single buggy line of code to taking full co

FaceParts: Segmentation and Editing of Gaussian Splatting

ResearchDGX agent

arXiv:2605.13853v1 Announce Type: cross Abstract: Facial editing is an important task with applications in entertainment, virtual reality, and digital avatars. Most existing approaches rely on generat

FactorizedHMR: A Hybrid Framework for Video Human Mesh Recovery

ResearchDGX agent

arXiv:2605.14854v1 Announce Type: cross Abstract: Human Mesh Recovery (HMR) is fundamentally ambiguous: under occlusion or weak depth cues, multiple 3D bodies can explain the same image evidence. This

Falkor-IRAC: Graph-Constrained Generation for Verified Legal Reasoning in Indian Judicial AI

HardwareDGX agent

arXiv:2605.14665v1 Announce Type: new Abstract: Legal reasoning is not semantic similarity search. A court judgment encodes constrained symbolic reasoning: precedent propagation, procedural state tran

Fast Rates for Inverse Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.14599v1 Announce Type: cross Abstract: We establish novel structural and statistical results for entropy-regularized min-max inverse reinforcement learning (Min-Max-IRL) with linear reward

Few Channels Draw The Whole Picture: Revealing Massive Activations in Diffusion Transformers

ResearchDGX agent

arXiv:2605.13974v1 Announce Type: cross Abstract: Diffusion Transformers (DiTs) and related flow-based architectures are now among the strongest text-to-image generators, yet the internal mechanisms t

Finding Interpretable Prompt-Specific Circuits in Language Models

ApplicationsDGX agent

arXiv:2602.13483v2 Announce Type: replace-cross Abstract: Understanding the internal circuits that language models use to solve tasks remains a central challenge in mechanistic interpretability. A cru

FlowSteer: Towards Agents Designing Agentic Workflows via Reinforced Progressive Canvas Editing

SafetyDGX agent

arXiv:2602.01664v4 Announce Type: replace Abstract: In recent years, agentic workflows have been widely applied to solve complex human tasks. However, existing workflow construction still faces key ch

From Descriptive to Prescriptive: Uncover the Social Value Alignment of LLM-based Agents

Model ReleasesDGX agent

arXiv:2605.14034v1 Announce Type: new Abstract: Wide applications of LLM-based agents require strong alignment with human social values. However, current works still exhibit deficiencies in self-cogni

From Ranking to Reasoning: Explainable Web API Recommendation via Semantic Reasoning

SafetyDGX agent

arXiv:2511.05820v2 Announce Type: replace-cross Abstract: The rapid growth of Web APIs has made automated Web API recommendation essential for efficient mashup development. However, existing approache

From Sycophantic Consensus to Pluralistic Repair: Why AI Alignment Must Surface Disagreement

Model ReleasesDGX agent

arXiv:2605.14912v1 Announce Type: new Abstract: Pluralistic alignment is typically operationalised as preference aggregation: producing responses that span (Overton), steer toward (Steerable), or prop

From Table to Cell: Attention for Better Reasoning with TABALIGN

ResearchDGX agent

arXiv:2605.14465v1 Announce Type: new Abstract: Multi-step LLM reasoning over structured tables fails because planning and execution share no explicit cell-grounding contract. Existing methods constra

From User Preferences to Base Score Extraction Functions in Gradual Argumentation (with Appendix)

ResearchDGX agent

arXiv:2602.14674v4 Announce Type: replace Abstract: Gradual argumentation is a field of symbolic AI which is attracting attention for its ability to support transparent and contestable AI systems. It

← Previous
1…251252253254255…358
Next →