AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
26 May 2026

Adversarial Orthogonal Disentanglement for LVLM Hallucination Mitigation

ResearchDGX agent

arXiv:2605.25377v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have advanced multimodal understanding, yet their reliability is limited by hallucination, where generated conten

Agent-as-Peer-Debriefer: A Multi-Agent Framework with Perspective-Based Refinement for Qualitative Analysis

AgentsDGX agent

arXiv:2605.24600v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for qualitative data analysis (QDA), yet their outputs often miss the depth and nuance of human analy

Agent-Centric Social Trajectory Prediction: A Free Energy Principle Perspective

SafetyDGX agent

arXiv:2605.25748v1 Announce Type: new Abstract: Trajectory prediction methods have demonstrated remarkable capabilities in capturing complex motion patterns. However, existing methods rely on global s


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Agent-Facing Information Design in LLM Tool Registries

SafetyDGX agent

arXiv:2605.23916v1 Announce Type: cross Abstract: LLM tool registries function as unregulated advertising platforms: providers write free-text descriptions that agents use for selection, yet no measur

Agent Learning via Early Experience

SafetyDGX agent

arXiv:2510.08558v3 Announce Type: replace Abstract: A long-term goal of language agents is to learn and improve through their own experience, ultimately outperforming humans in complex, real-world tas

Agent Manufacturing: Foundation-Model Agents as First-Class Industrial Entities

AgentsDGX agent

arXiv:2605.24823v1 Announce Type: new Abstract: Manufacturing has passed through four widely recognized paradigms - mechanization, electrification, programmable automation, and Smart Manufacturing - e

Agent Primitives: Reusable Latent Building Blocks for Multi-Agent Systems

AgentsDGX agent

arXiv:2602.03695v2 Announce Type: replace-cross Abstract: While existing multi-agent systems (MAS) can handle complex problems by enabling collaboration among multiple agents, they are often highly ta

Agent-ToM: Learning to Monitor Autonomous LLM Agents via Theory-of-Mind Reasoning

SafetyDGX agent

arXiv:2605.24216v1 Announce Type: cross Abstract: Monitoring autonomous large language model (LLM) agents for covert malicious behavior is challenging due to delayed, context-dependent, and long-horiz

Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning

Model ReleasesDGX agent

arXiv:2602.10090v3 Announce Type: replace Abstract: Recent advances in large language model (LLM) have empowered autonomous agents to perform multi-turn interactions with tools and environments. Howev

AgentFugue: Agent Scaling for Long-Horizon Tasks through Collective Reasoning

AgentsDGX agent

arXiv:2605.24486v1 Announce Type: new Abstract: Recent progress on long-horizon agentic tasks has been driven largely by scaling up individual agents through stronger models, better tools, and more ef

AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions

Model ReleasesDGX agent

arXiv:2605.25707v1 Announce Type: new Abstract: Autonomous computer use agents that powered by multimodal large language models (MLLMs) are emerging as capable assistants for completing complex digita

AGI Requires a Coordination Layer on Top of Pattern Repositories

AgentsDGX agent

arXiv:2512.05765v2 Announce Type: replace Abstract: In this paper we argue that influential critiques dismissing Large Language Models (LLMs) as a dead end for AGI misidentify the bottleneck: they con

AI as Equalizer or Amplifier? Task Complexity as the Moderating Factor for Human Expertise in Hybrid Intelligence Systems

ResearchDGX agent

arXiv:2512.10961v2 Announce Type: replace-cross Abstract: A growing body of empirical research suggests that generative AI narrows performance gaps between novice and expert workers on routine tasks--

AI-Assisted Systematization for Evaluating GenAI Systems

SafetyDGX agent

arXiv:2605.26001v1 Announce Type: cross Abstract: Evaluating generative AI (GenAI) systems is challenging because many targets of evaluation are broad, contested concepts, such as 'reasoning,' 'fairne

AI-Associated Lexical Shifts Across 34 Languages: Cross-Lingual Convergence and Diachronic Uptake in News Writing

Model ReleasesDGX agent

arXiv:2605.25358v1 Announce Type: cross Abstract: AI-associated lexical shifts have been documented mainly in Scientific English. We extend this work to 34 languages in the WMT News Crawl corpus, refi

AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems

Model ReleasesDGX agent

arXiv:2605.25272v1 Announce Type: new Abstract: While aggregate leaderboard scores drive AI development, they contain substantial measurement noise whose sources and magnitudes remain unquantified, ma

AI Content Moderation in Therapy Conversations

Model ReleasesDGX agent

arXiv:2605.25454v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly being used for emotional support. They are also being developed for formal therapy purposes. However, LL

AI-Driven Adaptive Adversaries and the Erosion of Cryptographic Trust in Public Key Systems

ResearchDGX agent

arXiv:2605.24542v1 Announce Type: cross Abstract: This paper examines the erosion of Public Key Cryptography (PKC) security under adaptive adversarial optimisation driven by artificial intelligence. T

AI-Driven Alpha Decay: Algorithmic Homogenization, Reflexive Signal Erosion, and the Paradox of Intelligent Markets

ResearchDGX agent

arXiv:2605.23905v1 Announce Type: cross Abstract: We show that AI-driven investment strategies are inherently self-defeating at scale. As AI adoption rises, three mutually reinforcing channels -- sign

AI-Driven Controlled Environment Agriculture as Resilient Infrastructure for U.S. Fresh-Produce Supply Chains

SafetyDGX agent

arXiv:2605.23946v1 Announce Type: cross Abstract: Climate volatility, regional production concentration, labor constraints, cyber risk, and dependence on long-distance fresh-produce supply chains expo

AI-generated podcasts: Synthetic Intimacy and Cultural Mistranslation in NotebookLM's Audio Overviews

ApplicationsDGX agent

arXiv:2511.08654v2 Announce Type: replace-cross Abstract: This paper analyses AI-generated podcasts produced by Google's NotebookLM, which generates audio podcasts with two chatty AI hosts discussing

AI in the Enterprise: How People Use M365 Copilot Chat

ApplicationsDGX agent

arXiv:2605.23958v1 Announce Type: cross Abstract: M365 Copilot is used every week by millions of people across more than a million companies around the world as part of their workflows. Uniquely posit

AION: Next-Generation Tasks and Practical Harness for Time Series

AgentsDGX agent

arXiv:2605.25045v1 Announce Type: new Abstract: Time series research is moving beyond fixed forecasting benchmarks toward realistic tasks that combine prediction, contextual reasoning, tool use, and s

All Leaks Count, Some Count More: Interpretable Temporal Contamination Detection and Mitigation in LLM Backtesting

ResearchDGX agent

arXiv:2602.17234v2 Announce Type: replace Abstract: Backtesting LLMs on resolved events assumes models reason only from pre-cutoff knowledge, yet pretrained models inevitably leak post-cutoff knowledg

AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications

Model ReleasesDGX agent

arXiv:2602.22769v3 Announce Type: replace Abstract: Large Language Models (LLMs) are deployed as autonomous agents in increasingly complex applications, where enabling long-horizon memory is critical

AME-TS: Anchored Mixture-of-Experts for Time Series Forecasting

Model ReleasesDGX agent

arXiv:2605.25166v1 Announce Type: cross Abstract: Time series forecasting models are increasingly scaled through large Transformer backbones, yet most existing approaches process all series through a

An Empirical Evaluation of LLM-Generated Code Security Across Prompting Methods

TutorialsDGX agent

arXiv:2605.24298v1 Announce Type: cross Abstract: The growing use of Large Language Models (LLMs) for automated code generation has enhanced software development efficiency, but often at the cost of s

An Interactive Paradigm for Deep Research

Model ReleasesDGX agent

arXiv:2605.24266v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled deep research systems that synthesize comprehensive, report-style answers to open-ended q

An Interpretable CF-RL-TOPSIS Fusion Model for Skills-Aware Talent Recommendation

Model ReleasesDGX agent

arXiv:2605.24155v1 Announce Type: cross Abstract: Effective skills-aware talent recommendation must balance behavioral transition patterns, trajectory-sensitive adaptation, and inspectable occupation-

AnatomiX, an Anatomy-Aware Grounded Multimodal Large Language Model for Chest X-Ray Interpretation

ResearchDGX agent

arXiv:2601.03191v3 Announce Type: replace-cross Abstract: Multimodal medical large language models have shown substantial progress in chest X-ray interpretation but continue to face challenges in spat

Anatomy-Anchored Self-Supervision: Distilling Vision Foundation Models for Invariant Ultrasound Representation

SafetyDGX agent

arXiv:2605.25402v1 Announce Type: cross Abstract: Self-supervised pre-training paradigm has gained increasing prominence for learning transferable representations in medical imaging, yet existing meth

APT-Agent: Automated Penetration Testing using Large Language Models

AgentsDGX agent

arXiv:2605.24949v1 Announce Type: cross Abstract: Penetration testing is essential to securing modern web infrastructures, yet traditional manual methods struggle to keep pace with their scale and com

Architecting Agentic Communities using Design Patterns

AgentsDGX agent

arXiv:2601.03624v3 Announce Type: replace Abstract: The rapid evolution of Large Language Models (LLM) and subsequent Agentic AI technologies requires systematic architectural guidance for building so

Artificial Effort

ResearchDGX agent

arXiv:2605.23920v1 Announce Type: cross Abstract: Real-effort tasks, in which participants perform cognitively costly activities whose outcomes depend on actual performance, are widely used in experim

ArtSplat: Feed-Forward Articulated 3D Gaussian Splatting from Sparse Multi-State Uncalibrated Views

Model ReleasesDGX agent

arXiv:2605.24304v1 Announce Type: cross Abstract: Articulated object reconstruction from sparse-view images is an ill-posed problem that requires simultaneous inference of geometry and underlying arti

Asking LLMs to Verify First is Almost Free Lunch

Model ReleasesDGX agent

arXiv:2511.21734v2 Announce Type: replace-cross Abstract: To enhance the reasoning capabilities of Large Language Models (LLMs) without high costs of training, nor extensive test-time sampling, we int

Assessing the Operational Viability of Foundation Models for Time Series Forecasting

ApplicationsDGX agent

arXiv:2605.24381v1 Announce Type: cross Abstract: Time series forecasting drives operational decisions in areas like finance, transportation, and energy. While supervised learning approaches achieve s

Associations between echocardiographic traits and AI-ECG predictions of heart failure

ResearchDGX agent

arXiv:2605.24576v1 Announce Type: new Abstract: Artificial intelligence-enabled electrocardiography (AI-ECG) can detect heart failure (HF), including disease not captured by left ventricular ejection

ASTRO: Adaptive Spatio-Temporal Reinforcement Optimization for GNN Powered Anomly Detection in Cyber Physical Systems

ApplicationsDGX agent

arXiv:2605.25135v1 Announce Type: cross Abstract: Anomaly detection in Industrial Internet of Things (IIoT) environments is essential to protect the Industrial Control Systems (ICS) and Cyber-Physical

Attested Tool-Server Admission: A Security Extension to the Model Context Protocol

AgentsDGX agent

arXiv:2605.24248v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) standardizes how a large-language-model (LLM) agent and an external tool server exchange messages, but not trust: a h

ATWL: A Formal Language for Representing, Comparing, and Reusing Visual Analytics Workflows

ResearchDGX agent

arXiv:2605.25489v1 Announce Type: new Abstract: Visual analytics (VA) workflows are inherently complex, involving data transformation, feature engineering, visual representation, and human interpretat

Auditing medical multi-agent AI reveals risks of false consensus

SafetyDGX agent

arXiv:2510.10185v2 Announce Type: replace-cross Abstract: Large language models are increasingly being assembled into medical multi-agent systems that emulate multidisciplinary consultation through sp

Authority Inversion in LLM-Mediated Ubiquitous Systems: When Models Trust Users Over Sensors

SafetyDGX agent

arXiv:2605.23938v1 Announce Type: new Abstract: Large language models (LLMs) increasingly fuse heterogeneous inputs in ubiquitous systems. Yet, how LLMs implicitly allocate authority when sensor measu

Authority Signals in Claude AI Health Citations: A Descriptive Analysis Using the Authority Signals Framework

Model ReleasesDGX agent

arXiv:2605.23921v1 Announce Type: cross Abstract: This study seeks to determine the authority signals used by Anthropic's Claude AI in its presentation of sources when answering consumer health questi

Automated Detection and Classification of Delusion-related Content in Naturalistic Audio Diaries Using Multi-Agent Language Models

AgentsDGX agent

arXiv:2605.24755v1 Announce Type: new Abstract: Speech monologues recorded in naturalistic settings provide opportunities to characterize mental illness phenomenology and detect symptom exacerbation.

Autoregression-Free Neural Operators for Time-Dependent PDEs

Model ReleasesDGX agent

arXiv:2605.25413v1 Announce Type: cross Abstract: Neural operators learn mappings from function-dependent inputs to solutions, providing an effective framework for solving partial differential equatio

AutoSG: LLM-Driven Solver Generation Solely from Task Prompts for Expensive Optimization

Local AiDGX agent

arXiv:2605.25658v1 Announce Type: cross Abstract: Expensive optimization tasks are ubiquitous in real-world applications, demanding highly specialized solvers. While LLM-driven automated solver genera

AvalancheBench: Evaluating Enterprise Data Agents Through Latent World Recovery

Model ReleasesDGX agent

arXiv:2605.24183v1 Announce Type: cross Abstract: We introduce AvalancheBench, a benchmark for evaluating enterprise data agents through latent world recovery. AvalancheBench improves on existing benc

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models

Model ReleasesDGX agent

arXiv:2605.24652v1 Announce Type: new Abstract: Rapid advances in audio-video (AV) generation have enabled high-fidelity synthesis with synchronized sound, particularly for human-related scenarios inv

Back to Parsimonious Latents: Learning Task-Centric World Models from Visual Foundations

AgentsDGX agent

arXiv:2605.25620v1 Announce Type: new Abstract: World models enable agents to predict future dynamics conditioned on actions, making the choice of latent representation central to planning and control

BackWeak: Backdooring Knowledge Distillation Simply with Weak Triggers and Fine-tuning

ResearchDGX agent

arXiv:2511.12046v2 Announce Type: replace-cross Abstract: Knowledge Distillation (KD) is essential for compressing large models, yet relying on pre-trained 'teacher' models downloaded from third-party

Balancing Fairness, Privacy, and Accuracy: A Multitask Adversarial Framework for Centralized Data-Driven Systems

SafetyDGX agent

arXiv:2605.24458v1 Announce Type: cross Abstract: The integration of fairness and privacy in centralized data-driven applications is critical, especially as these systems increasingly influence sector

Batch Normalization Amplifies Memorization and Privacy Risks

ResearchDGX agent

arXiv:2605.24420v1 Announce Type: cross Abstract: Batch Normalization (BN) is widely adopted to enable faster convergence and more stable training of deep neural networks. However, its impact on priva

BC Protocol: Structured Dual-Expert Dialogue for Eliciting High-Quality Chain-of-Thought Post-Training Data

Model ReleasesDGX agent

arXiv:2605.25549v1 Announce Type: cross Abstract: High-quality expert chain-of-thought (CoT) data is one of the core bottlenecks in large language model (LLM) post-training. Existing data production m

Behind EvoMap: Characterizing a Self-Evolving Agent-to-Agent Collaboration Network

AgentsDGX agent

arXiv:2605.25815v1 Announce Type: new Abstract: Agent-to-Agent (A2A) networks enable autonomous AI agents to collaborate by sharing reusable problem-solving instructions. However, how these decentrali

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs

Model ReleasesDGX agent

arXiv:2605.21602v2 Announce Type: replace Abstract: Many safety and alignment failures of large language models (LLMs) occur due to out-of-distribution (OOD) situations: unusual prompt or response pat

Benchmarking Patent Embeddings: A Multi-Task Evaluation of 22 Models Across Retrieval, Classification, and Clustering

Model ReleasesDGX agent

arXiv:2605.24297v1 Announce Type: cross Abstract: Which fine-tuning signals improve patent embedding models, and do gains transfer across patent landscapes? We benchmark 22 embedding models, from 22M-

Benchmarking Pathology Foundation Models for Spatial Domain Understanding

Model ReleasesDGX agent

arXiv:2605.25764v1 Announce Type: cross Abstract: Pathology foundation models (PFMs) have emerged as a core approach for learning transferable representations from whole slide images (WSIs), and they

Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork

Model ReleasesDGX agent

arXiv:2605.24423v1 Announce Type: new Abstract: In-Context Reinforcement Learning (ICRL) has enabled foundation agents to adapt instantaneously to novel tasks, yet its efficacy in Ad-Hoc Teamwork (AHT

Beyond Control-Flow: Integrating the Resource Perspective into Multi-Collaborative Process Modeling from Text

ResearchDGX agent

arXiv:2605.24546v1 Announce Type: new Abstract: Process modeling is a sub-domain of Business Process Management (BPM) focused on the translation of process artifacts into formal models. This task trad

← Previous
1…210211212213214…358
Next →