AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
19 May 2026

CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?

Model ReleasesDGX agent

arXiv:2605.16679v1 Announce Type: cross Abstract: End-to-end automation of realistic healthcare operations stresses three capabilities underrepresented in current benchmarks: policy density, decisions

CLAP: Contrastive Latent-space Prompt Optimization for End-to-end Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.17284v1 Announce Type: cross Abstract: End-to-end autonomous driving systems powered by Vision-Language-Action (VLA) models achieve strong performance on common driving scenarios, yet remai

ClawArena: Benchmarking AI Agents in Evolving Information Environments

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.04202v2 Announce Type: replace-cross Abstract: AI agents deployed as persistent assistants must maintain correct beliefs as their information environment evolves. In practice, evidence is s

Code as Agent Harness

SafetyDGX agent

arXiv:2605.18747v1 Announce Type: cross Abstract: Recent large language models (LLMs) have demonstrated strong capabilities in understanding and generating code, from competitive programming to reposi

CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook

SafetyDGX agent

arXiv:2605.18257v1 Announce Type: cross Abstract: Multimodal representation alignment is pivotal for large language models and robotics. Traditional methods are often hindered by cross-modal informati

CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models

ResearchDGX agent

arXiv:2602.17684v2 Announce Type: replace-cross Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) has driven recent progress in code large language models by leveraging execution-based f

CoLLM-NAS: Collaborative Large Language Models for Efficient Knowledge-Guided Neural Architecture Search

TutorialsDGX agent

arXiv:2509.26037v2 Announce Type: replace Abstract: The integration of Large Language Models (LLMs) with Neural Architecture Search (NAS) has introduced new possibilities for automating the design of

COLSON: Controllable Learning-Based Social Navigation via Diffusion-Based Reinforcement Learning

SafetyDGX agent

arXiv:2503.13934v2 Announce Type: replace-cross Abstract: Mobile robot navigation in dynamic environments with pedestrian traffic is a key challenge in the development of autonomous mobile service rob

CommitDistill: A Lightweight Knowledge-Centric Memory Layer for Software Repositories

Model ReleasesDGX agent

arXiv:2605.18284v1 Announce Type: cross Abstract: Software repositories accumulate large amounts of unstructured knowledge in commit messages, pull-request discussions, and issue threads, but develope

Computational Challenges in Token Economics: Bridging Economic Theory and AI System Design

ResearchDGX agent

arXiv:2605.17410v1 Announce Type: new Abstract: Token economics has emerged as a useful lens for understanding resource allocation, value creation, and pricing in large language model systems. While r

Concise and Logically Consistent Conformal Sets for Neuro-Symbolic Concept-Based Models

ResearchDGX agent

arXiv:2605.18202v1 Announce Type: cross Abstract: Neuro-Symbolic Concept-based Models (NeSy-CBMs) are a family of architectures that integrate neural networks with symbolic reasoning for enhanced reli

Confidence-Gated Robot Autonomy: When Does Uncertainty Actually Help?

SafetyDGX agent

arXiv:2605.18045v1 Announce Type: cross Abstract: Robotic systems often use predictive uncertainty to decide whether to act autonomously or defer to a fallback policy. In threshold-gated autonomy, unc

ConflictRAG: Detecting and Resolving Knowledge Conflicts in Retrieval Augmented Generation

ResearchDGX agent

arXiv:2605.17301v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems implicitly assume mutual consistency among retrieved documents -- an assumption that frequently fails in

Consent Chain Degradation in Embodied Multi-Agent Systems: Bridging the Gap Between AI Agent Governance and Robot Ethics

SafetyDGX agent

arXiv:2605.16300v1 Announce Type: cross Abstract: Robotic systems are moving from isolated platforms to interconnected multi-agent ecosystems that operate in human environments. This shift raises a go

Conservative AI for Safety-Sensitive Medical Image Restoration: Residual-Bounded CT-CTA Enhancement for Intracranial Aneurysm-Relevant Signal Recovery

Local AiDGX agent

arXiv:2605.16458v1 Announce Type: cross Abstract: Image restoration models are increasingly applied to degraded medical scans, but in safety-sensitive settings they must improve image quality without

Content-Style Identification via Differential Independence

ResearchDGX agent

arXiv:2605.17827v1 Announce Type: cross Abstract: Generative analysis often models multi-domain observations as nonlinear mixtures of domain-invariant content variables and domain-specific style varia

Context Memorization for Efficient Long Context Generation

Model ReleasesDGX agent

arXiv:2605.18226v1 Announce Type: cross Abstract: Modern large language model (LLM) applications increasingly rely on long conditioning prefixes to control model behavior at inference time. While pref

Continuous Diffusion Scales Competitively with Discrete Diffusion for Language

Model ReleasesDGX agent

arXiv:2605.18530v1 Announce Type: cross Abstract: While diffusion has drawn considerable recent attention from the language modeling community, continuous diffusion has appeared less scalable than dis

ContractBench: Can LLM Agents Preserve Observation Contracts?

Model ReleasesDGX agent

arXiv:2605.17281v1 Announce Type: cross Abstract: Tool-augmented LLM agents call APIs whose intermediate outputs, such as presigned URLs, session tokens, and OAuth state parameters, are observation co

ContraFix: Agentic Vulnerability Repair via Differential Runtime Evidence and Skill Reuse

Model ReleasesDGX agent

arXiv:2605.17450v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly used for automated vulnerability repair (AVR), where repository-level reasoning enables them to ins

Contrastive Conceptor Activation Steering (COAST): Unlocking Vision-Language-Action Models through Hidden States

SafetyDGX agent

arXiv:2605.17144v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models leverage powerful perceptual priors from web-scale Vision-Language Model (VLM) pre-training, yet they remain surpr

Controlling False Discovery in Arbitrarily Structured Hypothesis Spaces via Reproducing Kernels

ApplicationsDGX agent

arXiv:2605.17559v1 Announce Type: cross Abstract: Large-scale hypothesis testing is central to modern science, where controlling the False Discovery Rate (FDR) has become the standard approach to mana

Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation

Model ReleasesDGX agent

arXiv:2602.16990v2 Announce Type: replace Abstract: Most recommendation benchmarks evaluate how well a model imitates user behavior. In financial advisory, however, observed actions can be noisy or sh

Convergence of Multiagent Learning Systems for Traffic control

AgentsDGX agent

arXiv:2511.11654v2 Announce Type: replace-cross Abstract: Rapid urbanization in cities like Bangalore has led to severe traffic congestion, making efficient Traffic Signal Control (TSC) essential. Mul

COOPO: Cyclic Offline-Online Policy Optimization Algorithm

SafetyDGX agent

arXiv:2605.18675v1 Announce Type: cross Abstract: Offline reinforcement learning struggles with distributional shift and constrained performance due to static dataset limitations, while online RL dema

CooT: Learning to Coordinate In-Context with Coordination Transformers

Model ReleasesDGX agent

arXiv:2506.23549v3 Announce Type: replace Abstract: Effective coordination among unfamiliar partners remains a major challenge in multi-agent systems. Existing approaches, such as population-based met

CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery

AgentsDGX agent

arXiv:2604.01658v2 Announce Type: replace Abstract: Large language model (LLM)-based evolution is a promising approach for open-ended discovery, where progress requires sustained search and knowledge

CoUn: Empowering Machine Unlearning via Contrastive Learning

ResearchDGX agent

arXiv:2509.16391v3 Announce Type: replace-cross Abstract: Machine unlearning (MU) aims to remove the influence of specific 'forget' data from a trained model while preserving its knowledge of the rema

CounterCount: A Diagnostic Framework for Counting Bias in Vision Language Models

Local AiDGX agent

arXiv:2605.17826v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) excel at multimodal reasoning, yet it remains unclear whether their answers are grounded in visual evidence or driven by

Counterparty Modeling is Not Strategy: The Limits of LLM Negotiators

ResearchDGX agent

arXiv:2605.16575v1 Announce Type: new Abstract: Negotiation requires more than inferring what the other side wants: it requires using that information to make advantageous offers and counteroffers ove

Cross-Domain Molecular Relational Learning: Leveraging Chemical Structure-Activity Analysis

TutorialsDGX agent

arXiv:2605.16799v1 Announce Type: cross Abstract: Recent advances in molecular representation integrates molecular topological and visual modalities, opening new avenues for precise Molecular Relation

Cross-modal Affinity-aligned Multimodal Learning Analytics for Predicting Student Collaboration Satisfaction in Game-Based Learning

SafetyDGX agent

arXiv:2605.16806v1 Announce Type: cross Abstract: Collaborative game-based learning environments offer rich opportunities for small-group knowledge construction, yet automatically predicting student c

Cross-Source Supervision for Bone Infection Segmentation in Dual-Modality PET-CT

Local AiDGX agent

arXiv:2605.16373v1 Announce Type: cross Abstract: Early and accurate diagnosis and lesion localization of bone infections are crucial for clinical treatment. PET-CT integrates anatomical information f

CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark

Model ReleasesDGX agent

arXiv:2605.18621v1 Announce Type: cross Abstract: Spatial intelligence requires multimodal large language models (MLLMs) to move beyond single-view perception and reason consistently about objects, vi

CTFS : Collaborative Teacher Framework for Forward-Looking Sonar Image Semantic Segmentation with Extremely Limited Labels

TutorialsDGX agent

arXiv:2603.21071v2 Announce Type: replace-cross Abstract: As one of the most important underwater sensing technologies, forward-looking sonar exhibits unique imaging characteristics. Sonar images are

Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image Generation

SafetyDGX agent

arXiv:2605.17807v1 Announce Type: cross Abstract: Text-to-Image (T2I) generation has achieved remarkable progress in recent years. Meanwhile, reinforcement learning methods, particularly those based o

CVE-Factory: Scaling Expert-Level Agentic Tasks for Code Security Vulnerability

Model ReleasesDGX agent

arXiv:2602.03012v2 Announce Type: replace-cross Abstract: Evaluating and improving the security capabilities of code agents requires high-quality, executable vulnerability tasks. However, existing wor

CyberCorrect: A Cybernetic Framework for Closed-Loop Self-Correction in Large Language Models

ResearchDGX agent

arXiv:2605.17305v1 Announce Type: new Abstract: Large language model (LLM) self-correction -- the ability to detect and fix errors in generated outputs -- remains largely ad hoc, relying on generic pr

D^2Evo: Dual Difficulty-Aware Self-Evolution for Data-Efficient Reinforcement Learning

ResearchDGX agent

arXiv:2605.17037v1 Announce Type: cross Abstract: Reinforcement learning (RL) has demonstrated potential for enhancing reasoning in large language models (LLMs). However, effective RL training, which

DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models

SafetyDGX agent

arXiv:2605.16342v1 Announce Type: cross Abstract: Diffusion large language models are a compelling alternative to autoregressive models, yet existing RL methods for diffusion treat all denoising steps

DARC: Disagreement-Aware Alignment via Risk-Constrained Decoding

SafetyDGX agent

arXiv:2603.08145v2 Announce Type: replace-cross Abstract: Preference-based alignment methods (e.g., RLHF, DPO) typically optimize a single scalar objective, implicitly averaging over heterogeneous hum

DARE-EEG: A Foundation Model for Mining Dual-Aligned Representation of EEG

Model ReleasesDGX agent

arXiv:2605.18298v1 Announce Type: new Abstract: Foundation models pre-trained through masked reconstruction on large-scale EEG data have emerged as a promising paradigm for learning generalizable neur

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention

HardwareDGX agent

arXiv:2605.18753v1 Announce Type: cross Abstract: Current hierarchical attention methods, such as NSA and InfLLMv2, select the top-k relevant key-value (KV) blocks based on coarse attention scores and

Data-driven and distributed governance of building facilities management using decentralized autonomous organization, digital twin, and large language models

AgentsDGX agent

arXiv:2605.16298v1 Announce Type: cross Abstract: While traditional AI and data-driven facilities management approaches have improved building operational efficiency, they remain constrained by centra

Data Presentation Over Architecture: Resampling Strategies for Credit Risk Prediction with Tabular Foundation Models

Model ReleasesDGX agent

arXiv:2605.18635v1 Announce Type: cross Abstract: Credit default prediction is a tabular learning problem with severe class imbalance, heterogeneous features, and tight latency budgets. Tabular Founda

DBES: A Systematic Benchmark and Metric Suite for Evaluating Expert Specialization in Large-Scale MoEs

Model ReleasesDGX agent

arXiv:2605.18498v1 Announce Type: cross Abstract: Expert specialization in Mixture-of-Experts (MoE) models remains poorly understood, with traditional evaluations conflating architectural load-balanci

DCFold: Efficient Protein Structure Generation with Single Forward Pass

ResearchDGX agent

arXiv:2605.17899v1 Announce Type: cross Abstract: AlphaFold3 introduces a diffusion-based architecture that elevates protein structure prediction to all-atom resolution with improved accuracy. This st

DecoupleSearch: Decouple Planning and Search via Hierarchical Reward Modeling

Model ReleasesDGX agent

arXiv:2510.21712v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) systems have emerged as a pivotal methodology for enhancing Large Language Models (LLMs) through the dyna

Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation

SafetyDGX agent

arXiv:2605.16826v1 Announce Type: cross Abstract: Knowledge distillation is central to LLM post-training, yet its design space remains poorly understood, especially alongside reinforcement learning (R

Deep Reinforcement Learning Framework for Diversified Portfolio Management Across Global Equity Markets

SafetyDGX agent

arXiv:2605.17307v1 Announce Type: cross Abstract: This study develops and evaluates a deep reinforcement learning framework for dynamic portfolio allocation across global equity markets. The Soft Acto

Deep sequence models tend to memorize geometrically; it is unclear why

SafetyDGX agent

arXiv:2510.26745v3 Announce Type: replace-cross Abstract: Deep sequence models are said to store atomic facts predominantly in the form of associative memory: a brute-force lookup of co-occurring enti

DeepArrhythmia: Segment-Contextualized ECG Arrhythmia Classification via Selective Evidence Acquisition

AgentsDGX agent

arXiv:2605.16441v1 Announce Type: cross Abstract: Beat-level Electrocardiography (ECG) arrhythmia detection aims to assign an arrhythmia class to each beat in a recording, yet many existing systems tr

DeMa: Dual-Path Delay-Aware Mamba for Efficient Multivariate Time Series Analysis

ResearchDGX agent

arXiv:2601.05527v2 Announce Type: replace-cross Abstract: Accurate and efficient multivariate time series (MTS) analysis is increasingly critical for a wide range of intelligent applications. Within t

Democratizing Large-Scale Re-Optimization with LLM-Guided Model Patches

AgentsDGX agent

arXiv:2605.18692v1 Announce Type: new Abstract: Optimization models developed by operations research (OR) experts are often deployed as decision-support systems in industrial settings. However, real-w

Designing Cellular Manufacturing System in Presence of Alternative Process Plans

ApplicationsDGX agent

arXiv:2411.15361v3 Announce Type: replace Abstract: In the design of cellular manufacturing systems (CMS), numerous technological and managerial decisions must be made at both the design and operation

Detecting Verbatim LLM Copy-Paste in Homework

Model ReleasesDGX agent

arXiv:2605.16336v1 Announce Type: cross Abstract: Large language models (LLMs) have made fluent essay writing, code drafting, and quiz answering instantly available to students at every level, from se

DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models

Model ReleasesDGX agent

arXiv:2601.11895v3 Announce Type: replace-cross Abstract: DevBench is a telemetry-driven benchmark designed to evaluate Large Language Models (LLMs) on realistic code completion tasks. It includes 1,8

DexHoldem: Playing Texas Hold'em with Dexterous Embodied System

Model ReleasesDGX agent

arXiv:2605.18727v1 Announce Type: cross Abstract: Evaluating embodied systems on real dexterous hardware requires more than isolated primitive skills: an agent must perceive a changing tabletop scene,

DexWild: Dexterous Human Interactions for In-the-Wild Robot Policies

ResearchDGX agent

arXiv:2505.07813v2 Announce Type: replace-cross Abstract: Large-scale, diverse robot datasets have emerged as a promising path toward enabling dexterous manipulation policies to generalize to novel en

DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents

AgentsDGX agent

arXiv:2605.17439v1 Announce Type: cross Abstract: Evaluating LLM-generated interactive software requires execution in addition to static analysis. The key difficulty is that correctness is a graph-lev

← Previous
1…234235236237238…358
Next →