AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
26 May 2026

When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure

Model ReleasesDGX agent

arXiv:2605.23932v1 Announce Type: new Abstract: Despite strong medical benchmark accuracy, LLMs can exhibit severe multi-turn sycophancy in clinical dialogue, abandoning initial correct diagnosis unde

When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs

SafetyDGX agent

arXiv:2605.24202v1 Announce Type: new Abstract: Multi-agent LLM workflows route inference through specialized roles to lift end-task accuracy, but jointly training those roles with reinforcement learn

When Does Synthetic Patent Data Help? Volume-Fidelity Trade-offs in Low-Resource Multi-Label Classification

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.24296v1 Announce Type: new Abstract: We study when LLM-generated synthetic data helps low-resource multi-label patent classification, separating true synthetic value from the confound that

When Gradients Collide: Failure Modes of Multi-Objective Prompt Optimization for LLM Judges

ResearchDGX agent

arXiv:2605.26046v1 Announce Type: cross Abstract: Customizing an LLM judge to a specific task or domain often involves optimizing its prompt across multiple evaluation criteria simultaneously. Textual

When Mean CE Fails: Median CE Can Better Track Language Model Quality

Model ReleasesDGX agent

arXiv:2605.24667v1 Announce Type: new Abstract: Mean cross-entropy is the standard validation metric for language models, but it can fail to track model quality during training. We examine this in two

When Reasoning Hurts: Source-Aware Evaluation of Frontier LLMs for Clinical SOAP Note Generation

Model ReleasesDGX agent

arXiv:2605.24902v1 Announce Type: cross Abstract: Reasoning-enabled LLMs perform strongly on medical reasoning benchmarks, but it remains unclear whether these gains transfer to structured clinical do

When Search Becomes Memory: Turning Robot Design Trials into Transferable Skills

Model ReleasesDGX agent

arXiv:2605.25832v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as proposal generators for evolutionary robot design, yet most loops remain memoryless: simulator r

When the Manual Lies: A Realistic Benchmark to Evaluate MCP Poisoning Attacks for LLM Agents

Model ReleasesDGX agent

arXiv:2605.24069v1 Announce Type: cross Abstract: The rise of tool-using Large Language Model (LLM) agents, standardized by protocols like the Model Context Protocol (MCP), has unlocked unprecedented

Who judges the judges? Governance from metrics: a runtime framework for continuous LLM compliance monitoring

Model ReleasesDGX agent

arXiv:2605.24737v1 Announce Type: cross Abstract: Current approaches to AI compliance treat conformity as a binary, audit-time verdict rather than a continuous, measurable property of production syste

Whose Alignment? Comparing LLM Process Alignment Across Diverse Organizational Decision Contexts

Model ReleasesDGX agent

arXiv:2605.25256v1 Announce Type: new Abstract: Aligning AI systems with organizational decision-making is typically framed as a single-target problem: make the model behave like the organization. We

Why We Need World Models for AGI: Where LLMs Fail and How World Models May Outperform

AgentsDGX agent

arXiv:2605.23972v1 Announce Type: new Abstract: Large language models achieve strong performance in language generation and knowledge-intensive tasks, yet remain limited in settings requiring causal r

Why Your Deep Research Agent Fails? On Hallucination Evaluation in Full Research Trajectory

AgentsDGX agent

arXiv:2601.22984v2 Announce Type: replace Abstract: Diagnosing failure patterns in Deep Research Agents (DRAs) remains a critical challenge. Existing benchmarks predominantly rely on end-to-end evalua

World-State Transformations for Neuro-symbolic Interactive Storytelling

Model ReleasesDGX agent

arXiv:2605.24719v1 Announce Type: cross Abstract: Large Language Models (LLMs) have changed the possibilities of Interactive Storytelling systems that process free-text user input. However, as more of

WorldGUI: An Interactive Benchmark for Desktop GUI Automation from Any Starting Point

Model ReleasesDGX agent

arXiv:2502.08047v5 Announce Type: replace Abstract: Recent progress in GUI agents has substantially improved visual grounding, yet robust planning remains challenging, particularly when the environmen

WTKO-CNN: Deep Learning Reveals Sequence Motifs Distinguishing Wild-Type and Knockout ATAC-seq Peaks

ResearchDGX agent

arXiv:2605.24034v1 Announce Type: cross Abstract: Chromatin regulators can alter transcriptional programs by modifying the accessibility of regulatory DNA elements. Understanding how regulatory sequen

You Can Ground Earlier than See: An Effective and Efficient Pipeline for Temporal Sentence Grounding in Compressed Videos

ResearchDGX agent

arXiv:2303.07863v3 Announce Type: replace-cross Abstract: Given an untrimmed video, temporal sentence grounding (TSG) aims to locate a target moment semantically according to a sentence query. Althoug

Your Embedding Model is SMARTer Than You Think

Local AiDGX agent

arXiv:2605.24938v1 Announce Type: cross Abstract: Multimodal retrieval relies heavily on single-vector retrievers, which compress rich, sequential token sequences into one single global representation

Zero-Shot Parkinson's Disease Detection from Speech: Comparing Large Audio and Language Models

ResearchDGX agent

arXiv:2605.24806v1 Announce Type: cross Abstract: Large audio and language models have recently demonstrated zero-shot reasoning capabilities across various domains. However, it remains unclear how th

25 May 2026

6G Communication Networks Enabling Embodied Agents: Architecture and Prototype

AgentsDGX agent

arXiv:2605.23263v1 Announce Type: cross Abstract: Embodied agents, which couple intelligent decision-making with physical actuation in the real world, impose far more stringent and heterogeneous commu

A drone-based framework for coral habitat mapping via weakly supervised segmentation

ResearchDGX agent

arXiv:2508.18958v2 Announce Type: replace-cross Abstract: Obtaining pixel-level annotations over large spatial extents remains a major bottleneck for deploying machine learning in ecological applicati

A Fine-Tuned BERT Classifier for Personal-Letter Titles in Late-Ming and Early-Qing Collected Works

ResearchDGX agent

arXiv:2605.23103v1 Announce Type: cross Abstract: I present Lepton (Letter Prediction), a fine-tuned BERT classifier that predicts whether a title in a Classical Chinese wenji table of contents is a p

A mathematical theory of balancing relational generalization and memorization

TutorialsDGX agent

arXiv:2605.22972v1 Announce Type: cross Abstract: Humans, animals, and modern machine learning models exhibit impressive abilities to learn complex behaviors and generalize these behaviors to unseen s

A measurement substrate for agentic Kubernetes operations: Methodology and a case study in retrieval-compounding falsification

Model ReleasesDGX agent

arXiv:2605.23058v1 Announce Type: cross Abstract: Empirical claims about autonomous Kubernetes operations agents are largely unfalsifiable. Published work reports observational results without control

A Proactive Multi-Agent Dialogue Framework for Assessing Social Language Disorder Traits in Autism

AgentsDGX agent

arXiv:2605.22993v1 Announce Type: cross Abstract: Characteristic linguistic behaviors associated with Social Language Disorder (SLD) in autism spectrum disorder, including echoic repetition, pronoun d

A Systematic Evaluation of Co-folding Model Representations for Small-Molecule Learning

Model ReleasesDGX agent

arXiv:2602.13249v2 Announce Type: replace-cross Abstract: Small-molecule foundation models are typically pretrained on standalone molecular data, unlike vision and language models that often benefit f

Adaptive Mass-Segmented KV Compression for Long-Context Reasoning

ResearchDGX agent

arXiv:2605.23200v1 Announce Type: cross Abstract: The linear growth of the Key-Value (KV) cache is a critical bottleneck in long-form LLM inference. Existing KV compression methods mitigate this by ev

Adversarial Vulnerability Under Temporal Concept Drift: A Longitudinal Study of Android Malware Detection

ResearchDGX agent

arXiv:2605.23623v1 Announce Type: cross Abstract: We present a longitudinal, drift-aware evaluation of adversarial robustness across more than a decade of Android applications using static and dynamic

Agentic Proving for Program Verification

Model ReleasesDGX agent

arXiv:2605.23772v1 Announce Type: new Abstract: Agentic systems have recently emerged as state-of-the-art approaches for automated theorem proving in formal mathematics. To assess how far these capabi

Agentic-VLA: Efficient Online Adaptation for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2605.22896v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for robotic manipulation by leveraging pre-trained vision-language representa

Agentivism: a learning theory for the age of artificial intelligence

AgentsDGX agent

arXiv:2604.07813v2 Announce Type: replace Abstract: Learning theories have historically changed when the conditions of learning evolved. Generative and agentic AI create a new condition by allowing le

AI Assurance: A Comprehensive Testing Strategy for Enterprise AI Systems

AgentsDGX agent

arXiv:2605.23459v1 Announce Type: cross Abstract: Enterprise AI systems, built on large language models, retrieval pipelines and autonomous agents, introduce a class of risks that traditional software

AI Evaluation Should Require Standardized Item-Level Data Releases

Model ReleasesDGX agent

arXiv:2604.03244v2 Announce Type: replace Abstract: This position paper argues that standardized item-level benchmark data should become the default infrastructure for AI evaluation. Current evaluatio

AI Security Research Should Better Incentivize Defense Research

ResearchDGX agent

arXiv:2605.23448v1 Announce Type: cross Abstract: This work examines an imbalance in artificial intelligence (AI) security research: the field tends to produce more work on attacking AI systems than o

ALIVE: Awakening LLM Reasoning via Adversarial Learning and Instructive Verbal Evaluation

SafetyDGX agent

arXiv:2602.05472v2 Announce Type: replace Abstract: The quest for expert-level reasoning in Large Language Models (LLMs) has been hampered by a persistent extit{reward bottleneck}: traditional reinfor

An AI-Driven Framework for Energy-Efficient Environmental Monitoring in Smart Cities Using Edge Intelligence

ResearchDGX agent

arXiv:2605.22824v1 Announce Type: cross Abstract: Environmental monitoring is a crucial component of the smart city infrastructure. It enables informed decision making which enhances sustainability, p

Anatomy-Guided Vision-Language Learning with Angular Prototype Separation for Multi-Label Video Capsule Endoscopy Classification Under Class Imbalance

Local AiDGX agent

arXiv:2603.17879v2 Announce Type: replace-cross Abstract: This work presents a multi-label temporal event detection framework for video capsule endoscopy (VCE) that addresses the extreme class imbalan

Any2Any: Efficient Cross-Embodiment Transfer for Humanoid Whole-Body Tracking

Model ReleasesDGX agent

arXiv:2605.23733v1 Announce Type: cross Abstract: Whole-body tracking (WBT) models have become a key foundation for humanoid robots, enabling them to imitate diverse motions with high fidelity. Traini

Anytime Training with Schedule-Free Spectral Optimization

Model ReleasesDGX agent

arXiv:2605.23061v1 Announce Type: cross Abstract: Standard neural network training relies on learning-rate schedules tied to a fixed horizon, leading to strong path dependence and costly re-tuning as

Approximate Machine Unlearning through Manifold Representation Forgetting Guided by Self Mode Connectivity

Local AiDGX agent

arXiv:2605.22871v1 Announce Type: cross Abstract: Machine unlearning is a fundamental mechanism that enforces the right to be forgotten. Existing unlearning studies that rely on label manipulation or

ArcMark: Distortion-Free Multi-Byte LLM Watermark via Optimal Transport

ResearchDGX agent

arXiv:2602.07235v2 Announce Type: replace-cross Abstract: Watermarking is an important tool for promoting the responsible use of large language models (LLMs). Existing watermarks insert a signal into

Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks

Model ReleasesDGX agent

arXiv:2605.23243v1 Announce Type: cross Abstract: We evaluate whether frontier LLMs are ready for cybersecurity through a dual-mode benchmark: white-box function-level vulnerability detection (VulnLLM

ARMS: Automatic Reward Shaping for Sparse-Reward Multi-Agent Reinforcement Learning

SafetyDGX agent

arXiv:2605.23562v1 Announce Type: cross Abstract: Sparse rewards are a major bottleneck in multi-agent reinforcement learning (MARL), where simultaneous learning induces non-stationarity and makes rew

As X, Do Y: How Persona and Task Combine in Instruction-Tuned LLMs

Model ReleasesDGX agent

arXiv:2605.23147v1 Announce Type: cross Abstract: Role prompts of the form As X, do Y admit a clean linear decomposition at one specific site in the residual stream: the prompt-to-answer transition --

Atom-level Protein Representation Learning Improves Protein Structure Prediction

Model ReleasesDGX agent

arXiv:2605.22133v2 Announce Type: replace-cross Abstract: Recent advances in generative modeling show that pretrained representations can improve generation as conditioning features or alignment targe

Automated Random Embedding for Practical Bayesian Optimization with Unknown Effective Dimension

ApplicationsDGX agent

arXiv:2605.23473v1 Announce Type: cross Abstract: Bayesian optimization is widely employed for optimizing complex black-box functions but struggles with the curse of dimensionality. Random embedding,

Autonomous Frontier-Based Exploration with VLM Guidance

AgentsDGX agent

arXiv:2605.23165v1 Announce Type: cross Abstract: Autonomous robotic exploration of unknown and hazardous environments, a long-standing challenge, can be significantly improved by leveraging the advan

AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery

AgentsDGX agent

arXiv:2605.23204v1 Announce Type: new Abstract: Scientific research is being reshaped by AI systems that move beyond isolated assistance toward longer-horizon workflows spanning literature grounding,

Ax-Prover: A Deep Reasoning Agentic Framework for Theorem Proving in Mathematics and Quantum Physics

Model ReleasesDGX agent

arXiv:2510.12787v4 Announce Type: replace Abstract: We present Ax-Prover, a multi-agent system for automated theorem proving in Lean that can solve problems across diverse scientific domains and opera

BarrierSteer: LLM Safety via Learning Barrier Steering

SafetyDGX agent

arXiv:2602.20102v2 Announce Type: replace-cross Abstract: Despite the strong performance of large language models (LLMs) across diverse tasks, their susceptibility to adversarial attacks and unsafe co

Beyond Binary Edits Robust Multimodal Knowledge Editing with Adversarial Subspace Alignment

SafetyDGX agent

arXiv:2605.23780v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) need efficient mechanisms to update knowledge without degrading existing capabilities. While intrinsic multimod

Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling

SafetyDGX agent

arXiv:2602.11146v2 Announce Type: replace-cross Abstract: Preference optimization for diffusion and flow-matching models relies on reward functions that are both discriminatively robust and computatio

BOHM: Zero-Cost Hierarchical Attribution for Compound AI Systems

Model ReleasesDGX agent

arXiv:2605.22866v1 Announce Type: new Abstract: Compound AI systems route tasks through hierarchies of specialised components. Attribution is dominated by Shapley-based methods (SHAP), which decompose

Brain-LLM Alignment Tracks Training Data, Not Typology

Model ReleasesDGX agent

arXiv:2605.23032v1 Announce Type: cross Abstract: Brain-LLM alignment is well established in English, yet the brain's language network is neuroanatomically universal across languages. Does alignment a

Bridging AI and Clinical Reasoning: Abductive Explanations for Alignment on Critical Symptoms

SafetyDGX agent

arXiv:2602.13985v2 Announce Type: replace Abstract: Artificial intelligence (AI) has demonstrated strong potential in clinical diagnostics, often achieving accuracy comparable to or exceeding that of

Bridging Data and Physics: A Graph Neural Network-Based Hybrid Twin Framework

TutorialsDGX agent

arXiv:2512.15767v2 Announce Type: replace-cross Abstract: Simulating complex unsteady physical phenomena relies on detailed mathematical models, simulated for instance by using the Finite Element Meth

CALAD: Channel-Aware contrastive Learning for multivariate time series Anomaly Detection

TutorialsDGX agent

arXiv:2605.23139v1 Announce Type: cross Abstract: Multivariate time series anomaly detection has become increasingly important in real-world applications, where labeled data are often scarce. Many exi

CBANet: A Compact Attention-Based CNN-BiLSTM Network for Aggressive Driving Event Detection

SafetyDGX agent

arXiv:2605.23471v1 Announce Type: cross Abstract: Aggressive driving is a major cause of traffic accidents and poses a serious threat to road safety. Although deep learning methods have shown promisin

ChainFlow-VLA: Causal Flow Planning with Vision-Language Models

SafetyDGX agent

arXiv:2605.23270v1 Announce Type: cross Abstract: Current end-to-end autonomous driving systems are fundamentally limited by a mismatch between temporal causal reasoning and global trajectory consiste

CHASD: Language Increment-Calibrated Contrastive Decoding against Hallucination in LVLMs

Local AiDGX agent

arXiv:2605.23344v1 Announce Type: cross Abstract: Large Vision-Language Models have shown strong multimodal reasoning capabilities, yet they remain susceptible to object hallucinations when language p

CHRONOS: Temporally-Aware Multi-Agent Coordination for Evolving Data Marketplaces

Model ReleasesDGX agent

arXiv:2605.23887v1 Announce Type: cross Abstract: Temporal knowledge-graph data marketplaces face three coupled failures in static designs: stale hybrid index shortcuts reduce recall as edges evolve,

← Previous
1…220221222223224…358
Next →