AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
27 May 2026

Trust Region Q Adjoint Matching

Model ReleasesDGX agent

arXiv:2605.27079v1 Announce Type: cross Abstract: Off-policy reinforcement learning of pretrained flow policies remains challenging due to the instability of optimization arising from the multi-step s

TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models

ResearchDGX agent

arXiv:2605.26161v1 Announce Type: cross Abstract: Time series foundation models (TSFMs) are increasingly pretrained on large corpora, raising concerns that evaluation datasets may have been exposed du

Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges

SafetyDGX agent

arXiv:2605.26156v1 Announce Type: cross Abstract: The known stylistic biases in LLM judges, such as a preference for verbosity or specific sentence structures, present an underexplored security vulner


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

TWIST: Closed-Loop token Synchronization for Application-Aware Wireless Digital Twins

ResearchDGX agent

arXiv:2605.27205v1 Announce Type: cross Abstract: Wireless digital twins require repeated synchronization between a time-evolving physical scene and its digital counterpart under limited and time-vary

Two Speeds of Learning: A Representation-Readout Decomposition of Grokking and Double Descent

ResearchDGX agent

arXiv:2605.27078v1 Announce Type: cross Abstract: Training loss and accuracy are the standard signals used to monitor generalization during deep neural network training. Two well-documented phenomena

UCPO: Uncertainty-Aware Policy Optimization

SafetyDGX agent

arXiv:2601.22648v2 Announce Type: replace Abstract: The key to building trustworthy large language models (LLMs) lies in endowing them with inherent uncertainty expression capabilities, thereby mitiga

Understanding Reasoning in LLMs through Strategic Information Allocation under Uncertainty

Local AiDGX agent

arXiv:2603.15500v2 Announce Type: replace Abstract: LLMs often exhibit Aha moments such as self-correction after tokens like 'Wait,' yet the underlying mechanism remains unclear. Standard LLMs collaps

Understanding the Challenges in Iterative Generative Optimization with LLMs

ResearchDGX agent

arXiv:2603.23994v2 Announce Type: replace-cross Abstract: Generative optimization uses large language models (LLMs) to iteratively improve artifacts (such as code, workflows or prompts) using executio

Uniboost: Global Coordination with Value Alignment for Fair and Efficient Traffic Allocation

SafetyDGX agent

arXiv:2605.26424v1 Announce Type: cross Abstract: With the rapid evolution of internet services, recommendation systems have become indispensable. In particular, the blending (re-ranking) stage plays

Unified Neural Scaling Laws

ResearchDGX agent

arXiv:2605.26248v1 Announce Type: cross Abstract: We present a functional form (that we refer to as a Unified Neural Scaling Law (UNSL)) that accurately models and extrapolates the scaling behaviors o

Unified Panoramic Geometry Estimation via Multi-View Foundation Models

ResearchDGX agent

arXiv:2605.26368v1 Announce Type: cross Abstract: Geometry estimation from perspective images has greatly advanced, maturing to the point where off-the-shelf foundation models are able to reconstruct

UnityMAS-O: A General RL Optimization Framework for LLM-Based Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2605.26646v1 Announce Type: new Abstract: LLM-based multi-agent systems decompose complex tasks into interacting roles, but most remain manually orchestrated by prompts, tools, and control rules

Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization

AgentsDGX agent

arXiv:2605.26501v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have transformed multi-modal understanding, excelling in tasks like image captioning and visual question answerin

Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization

Model ReleasesDGX agent

arXiv:2605.26457v1 Announce Type: cross Abstract: AI coding agents are increasingly used to write real-world software, but ensuring that their outputs are correct remains a fundamental challenge. Form

VesselSim: learning 3D blood vessel segmentation without expert annotations

ApplicationsDGX agent

arXiv:2605.26277v1 Announce Type: cross Abstract: Blood vessel segmentation is a core task in medical image analysis for the care of vascular diseases and surgical planning, yet the challenges of prov

VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents

Model ReleasesDGX agent

arXiv:2605.26144v1 Announce Type: cross Abstract: We present VISTA (VIsual Spec-To-App Benchmark), a benchmark for evaluating the end-to-end web-app generation capabilities of LLM-based agents. Unlike

VisualNeedle: Benchmarking Active Visual Search in Information-Dense Scenes

Model ReleasesDGX agent

arXiv:2605.26380v1 Announce Type: cross Abstract: Frontier multimodal large language models (MLLMs) have been reported to achieve over 90% accuracy on fine-grained perception benchmarks. However, such

VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions

Model ReleasesDGX agent

arXiv:2605.27141v1 Announce Type: new Abstract: Large language models (LLMs) have evolved into interactive agents that collaborate with users in real-world tasks. Effective collaboration in such setti

Vital Trace: Protocol-Constrained Patient-State Reasoning for Longitudinal Clinical Trajectories

AgentsDGX agent

arXiv:2602.12833v2 Announce Type: replace-cross Abstract: Longitudinal clinical reasoning over electronic health records requires tracking evolving physiological measurements, laboratory results, and

What Makes Chain-of-Thought Work at Probe Time? Local Co-occurrence Rather Than Global Derivation

Model ReleasesDGX agent

arXiv:2605.26795v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting reliably improves language-model accuracy, but which properties of a rationale text drive the improvement is poorly und

When Correct Demonstrations Hurt: Rethinking the Role of Exemplars in In-Context Learning

ResearchDGX agent

arXiv:2605.26350v1 Announce Type: cross Abstract: In-context learning (ICL) is often motivated by the intuition that demonstrations help because they provide correct input-output examples. However, we

When Does Adaptive Guidance Help? Belief-Aware Privileged Distillation for Autonomous Driving Under Partial Observability

AgentsDGX agent

arXiv:2605.26155v1 Announce Type: cross Abstract: Guided Soft Actor-Critic (GSAC) distills knowledge from a privileged full-state teacher to a partial-observation student for autonomous driving, but u

When Does Deep RL Beat Calibrated Baselines? A Benchmark Study on Adaptive Resource Control

Model ReleasesDGX agent

arXiv:2605.26418v1 Announce Type: cross Abstract: A properly calibrated rule-based autoscaler can beat every one of six mainstream deep reinforcement learning (DRL) algorithms on cost across every wor

When Eyes Betray AI: Social Gaze Consistency as a Semantic Cue for AI-Generated Image Detection

SafetyDGX agent

arXiv:2605.27348v1 Announce Type: cross Abstract: Recent generative models have largely closed the gap on low-level artifacts - pixel fingerprints, frequency anomalies, upsampling traces - particularl

When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation

Model ReleasesDGX agent

arXiv:2509.26600v2 Announce Type: replace-cross Abstract: As LLMs rapidly saturate existing benchmarks, automated benchmark creation using LLMs (LLM-as-a-benchmark) -- where a model generates test inp

Where Code Meets Natural Language: Taxonomy-Driven Information Flow Analysis for LLM-Integrated Applications

ApplicationsDGX agent

arXiv:2603.28345v2 Announce Type: replace-cross Abstract: LLM API calls are becoming a ubiquitous program construct, yet they create a boundary that no existing program analysis can cross: runtime val

Which Changes Matter? Towards Trustworthy Legal AI via Relevance-Sensitive Evaluation and Solver-Grounded Reasoning

SafetyDGX agent

arXiv:2605.26530v1 Announce Type: new Abstract: Legal reasoning requires distinguishing changes that matter from those that do not. Legal AI should remain stable under legally irrelevant perturbations

Why LLMs Hallucinate on Structured Knowledge: A Mechanistic Analysis of Reasoning over Linearized Representations

ResearchDGX agent

arXiv:2605.26362v1 Announce Type: cross Abstract: In many reasoning tasks, large language models (LLMs) rely on structured external knowledge, such as graphs and tables, which is typically linearized

Workflow Closure Is Not Scientific Closure in Auto-Research Systems

Model ReleasesDGX agent

arXiv:2605.26200v1 Announce Type: cross Abstract: This paper argues that workflow closure is not scientific closure in auto-research systems. Current systems can increasingly complete research-like lo

Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU

HardwareDGX agent

arXiv:2605.26118v1 Announce Type: cross Abstract: Porting deep learning algorithms to new hardware accelerators requires developers to repeatedly apply the same low-level optimizations -- quantization

XGrammar-2: Efficient Dynamic Structured Generation Engine for Agentic LLMs

AgentsDGX agent

arXiv:2601.04426v3 Announce Type: replace Abstract: Modern LLM agents increasingly rely on dynamic structured generation, such as tool calling and response protocols. Unlike traditional structured gen

Yes, Q-learning Helps Offline In-Context RL

ResearchDGX agent

arXiv:2502.17666v4 Announce Type: replace-cross Abstract: Existing offline in-context reinforcement learning (ICRL) methods have predominantly relied on supervised training objectives, which are known

Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems

Model ReleasesDGX agent

arXiv:2605.26302v1 Announce Type: new Abstract: Long-lived AI agents are increasingly deployed as persistent operational systems, yet they are still evaluated like freshly initialized models. Day-one

26 May 2026

A Comprehensive Dataset for Human vs. AI Generated Image Detection

ResearchDGX agent

arXiv:2601.00553v2 Announce Type: replace-cross Abstract: Multimodal generative AI systems like Stable Diffusion, DALL-E, and MidJourney have fundamentally changed how synthetic images are created. Th

A Controlled Synthetic Benchmark for Educational Aspect-Based Sentiment Analysis

Model ReleasesDGX agent

arXiv:2605.25502v1 Announce Type: cross Abstract: Educational aspect-based sentiment analysis (ABSA) can support course improvement, but public aspect-labeled student feedback remains scarce because e

A Deep Dive into Axiomatic Design -- Part I: Problem Formulation

ResearchDGX agent

arXiv:2605.25735v1 Announce Type: new Abstract: Problem formulation translating customer needs and constraints into a minimum set of independent first-level functional requirements, is arguably the mo

A Dynamical Framework for Cognitive Processes Based on Transformations and Semantic Equivalence

ResearchDGX agent

arXiv:2605.23942v1 Announce Type: new Abstract: This paper proposes a structural and dynamical framework for modeling cognitive processes within a cybernetic perspective. Cognitive states are represen

A general tensor-structured compression scheme for efficient large language models

ResearchDGX agent

arXiv:2605.25344v1 Announce Type: cross Abstract: Large language models (LLMs) are dominated by dense linear transformations, whose storage, memory and computational overheads hinder efficient adaptat

A governance horizon for ethical-use constraints in open-weight AI models

SafetyDGX agent

arXiv:2605.24383v1 Announce Type: new Abstract: Ethical constraints on open-weight AI models are both a reflection of societal concerns and a foundation for AI governance policy. They are expected to

A Large-Scale Dataset and Benchmark: Do Protein-Ligand Models Learn Binding Sites or Just Binding Likelihood?

Model ReleasesDGX agent

arXiv:2605.24045v1 Announce Type: cross Abstract: Protein-ligand modeling underpins computational drug discovery and molecular design. Existing protein-ligand benchmarks typically evaluate whether a p

A Multi-Agent LLM Framework for Rating the Quality of Surgical Feedback

AgentsDGX agent

arXiv:2605.25440v1 Announce Type: cross Abstract: Verbal feedback delivered by attending surgeons in the operating room plays a critical formative role in resident trainee skill acquisition. Yet, asse

A Multimodal 3D Foundation Model for Light Sheet Fluorescence Microscopy Enables Few-Shot Segmentation, Classification, and Deblurring

SafetyDGX agent

arXiv:2605.26026v1 Announce Type: cross Abstract: Light sheet fluorescence microscopy (LSM) enables high-resolution, three-dimensional (3D) imaging of biological specimens, providing rich volumetric d

A Signal-Language Foundation Model for Broad-Spectrum Cardiovascular Assessment from Routine Electrocardiography

ResearchDGX agent

arXiv:2605.25446v1 Announce Type: new Abstract: Electrocardiography (ECG) is central to cardiovascular care, but conventional AI models are often restricted to common arrhythmias and may generalize po

A Sober Look at Agentic Misalignment in Automated Workflows

SafetyDGX agent

arXiv:2605.24197v1 Announce Type: new Abstract: We study a class of emergent misalignment in multi-agent systems (MAS), with a focus on automated workflows, which we refer to agentic misalignment. Alt

A Tertiary Review of Large Language Model-Based Code Generating Tasks: Trends, Challenges, and Future Directions

SafetyDGX agent

arXiv:2605.25536v1 Announce Type: cross Abstract: Context. Large language models (LLMs) are increasingly applied to code-generating tasks (CGTs) in software engineering. While reported results are pro

A Token/KV-Cache Communication Media Selection and Resource Allocation Strategy for Multi-Agent Collaboration

AgentsDGX agent

arXiv:2605.25422v1 Announce Type: cross Abstract: The convergence of large language models (LLMs) with 6G networks is fostering a paradigm of autonomous multi-agent cooperation, which in turn is expec

A World Model of Radiologist Reading for Medical Image Representation Learning

Model ReleasesDGX agent

arXiv:2605.23992v1 Announce Type: cross Abstract: Radiologist eye-tracking data provide a rich record of how experts search, compare, and accumulate evidence during image reading; yet, existing method

Abduction-Deduction Entanglement: Domain Generalization via Representation Transplants

ResearchDGX agent

arXiv:2605.25156v1 Announce Type: cross Abstract: Prediction models trained under the source distribution do not generalize well to a different target distribution. A valid inference about an unseen d

Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism

ResearchDGX agent

arXiv:2605.23945v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) has become a key post-training paradigm for improving model quality. However, the synchronous three-st

Acting on the Unseen: Communication-Free Collaborative Filtering for Decentralized Multi-Robot Task Allocation

Model ReleasesDGX agent

arXiv:2605.25584v1 Announce Type: cross Abstract: Multi-robot task allocation usually assumes some combination of communication, known task models, or a coordinator. We study the opposite extreme, a r

Actionable and diverse counterfactual explanations incorporating domain knowledge and plausibility constraints

ApplicationsDGX agent

arXiv:2511.20236v3 Announce Type: replace Abstract: Counterfactual explanations improve the actionable interpretability of machine learning models by identifying minimal changes required to achieve a

ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2605.24011v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models exhibit remarkable action generation for embodied intelligence, but their heavy compute make deployment on edge pl

Adaptive Graph Refinement and Label Propagation with LLMs for Cost-Effective Entity Resolution

Model ReleasesDGX agent

arXiv:2605.25814v1 Announce Type: cross Abstract: Dirty entity resolution (ER), which identifies records referring to the same real-world entity from a single, messy dataset, is a fundamental task in

Adaptive Human-AI Coordination via Hierarchical Action Disentanglement

SafetyDGX agent

arXiv:2605.24343v1 Announce Type: new Abstract: Human-AI collaboration requires agents that can adapt to diverse partner behaviors and skill levels while remaining robust to unseen partners. Existing

Adaptive Punishment for Cooperation in Mixed-Motive Games

AgentsDGX agent

arXiv:2605.24516v1 Announce Type: cross Abstract: Mixed-motive scenarios are ubiquitous in real-world multi-agent interactions, where self-interested agents often defect for immediate rewards, overloo

ADMFormer: An Adaptive-Decomposition Transformer with Time-Varying Masked Spatial Attention for Traffic Forecasting

ApplicationsDGX agent

arXiv:2605.25543v1 Announce Type: new Abstract: Accurate traffic forecasting is essential for intelligent transportation systems, supporting a wide range of real-world applications. However, it remain

Advancing Graph Few-Shot Learning via In-Context Learning

Model ReleasesDGX agent

arXiv:2605.24410v1 Announce Type: new Abstract: Graph few-shot learning, which aims to classify nodes from novel classes with only a few labeled examples, is a widely studied problem in graph learning

AdvantageFlow: Advantage-Weighted Least Squares for RL in Flow Models

SafetyDGX agent

arXiv:2605.26013v1 Announce Type: cross Abstract: We introduce AdvantageFlow, a forward-process reinforcement learning algorithm for rectified flow models. Unlike Flow-GRPO, which optimizes the revers

Adversarial Error Correction for Visual Autoregressive Generation

SafetyDGX agent

arXiv:2605.24843v1 Announce Type: cross Abstract: Visual Autoregressive (VAR) models have emerged as a powerful paradigm for image synthesis by performing hierarchical next-scale prediction. However,

Adversarial Network Imagination: Causal LLMs and Digital Twins for Proactive Telecom Mitigation

ResearchDGX agent

arXiv:2602.13203v2 Announce Type: replace-cross Abstract: Telecommunication networks experience complex failures such as fiber cuts, traffic overloads, and cascading outages. Existing monitoring and d

← Previous
1…209210211212213…358
Next →