AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
2 Jun 2026

RoleCDE:Benchmarking and Mitigating Role-Alignment Trade-offs in Role-Playing Agents

Model ReleasesDGX agent

arXiv:2606.01552v1 Announce Type: new Abstract: Role-playing agents(RPAs) are widely used to steer large language models(LLMs) toward role-consistent behavior, yet existing benchmarks mainly evaluate

RPCASSM: Robust PCA State Space Model For Infrared Small Target Detection

Model ReleasesDGX agent

arXiv:2606.01689v1 Announce Type: cross Abstract: The detection and segmentation of infrared small targets have important application significance in the fields of surveillance and security, maritime

RuleEdit: Failure-Guided Human-AI Model Editing with Prospective Impact Preview

Local AiDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.00011v1 Announce Type: cross Abstract: Despite the promise of AI to assist complex decisions, practitioners still lack ways to detect likely failures and inspect the consequences of model e

Ryze: Evidence-Enriched Data Synthesis from Biomedical Papers

Model ReleasesDGX agent

arXiv:2606.00902v1 Announce Type: new Abstract: General-purpose VLMs remain unreliable for biomedical research because valid answers in scientific papers depend on evidence split across figures, table

S-SPPO: Semantic-Calibrated Self-Play Preference Optimization

Model ReleasesDGX agent

arXiv:2606.01561v1 Announce Type: new Abstract: Aligning Large Language Models (LLMs) with human preferences is often formulated via Direct Preference Optimization (DPO). However, the standard Bradley

S3TS: Stochastic Scenario-Structured Tree Search for Advanced Planning Under Uncertainty

ResearchDGX agent

arXiv:2606.02151v1 Announce Type: new Abstract: Effective scheduling in the energy sector is essential to ensure the reliable operation of electrical grids and their connected assets by, for instance,

SafeMCP: Proactive Power Regulation for LLM Agent Defense via Environment-Grounded Look-Ahead Reasoning

SafetyDGX agent

arXiv:2606.01991v1 Announce Type: new Abstract: As Large Language Model (LLM) agents increasingly leverage the Model Context Protocol (MCP) to operate in complex environments, the expansion of their a

SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment

SafetyDGX agent

arXiv:2606.02530v1 Announce Type: new Abstract: Aligning Large Language Models (LLMs) with human values often degrades their general capabilities, termed the alignment tax. Existing methods mitigate t

Safety Alignment of LMs via Non-cooperative Games

SafetyDGX agent

arXiv:2512.20806v3 Announce Type: replace Abstract: Ensuring the safety of language models (LMs) while maintaining their usefulness remains a critical challenge in AI alignment. Current approaches rel

Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning

SafetyDGX agent

arXiv:2503.11832v5 Announce Type: replace Abstract: Recent vision language models (VLMs) have made remarkable strides in generative modeling with multimodal inputs, particularly text and images. Howev

Scaling Behavior of Single LLM-Driven Multi-Agent Systems

AgentsDGX agent

arXiv:2606.00655v1 Announce Type: cross Abstract: The burgeoning field of LLM-based Multi-Agent Systems (MAS) promises to tackle complex tasks through collaborative intelligence, yet fundamental quest

Scaling Search Relevance: Augmenting App Store Ranking with LLM-Generated Judgments

ApplicationsDGX agent

arXiv:2602.23234v4 Announce Type: replace-cross Abstract: Large-scale commercial search systems optimize for relevance to drive successful sessions that help users find what they are looking for. To m

SceneSmith: Agentic Generation of Simulation-Ready Indoor Scenes

SafetyDGX agent

arXiv:2602.09153v2 Announce Type: replace-cross Abstract: Simulation has become a key tool for training and evaluating home robots at scale, yet existing environments fail to capture the diversity and

scicode-lint: Detecting Methodology Bugs in Scientific Python Code with LLM-Generated Patterns

Local AiDGX agent

arXiv:2603.17893v2 Announce Type: replace-cross Abstract: Methodology bugs in scientific Python code produce plausible but incorrect results that traditional linters and static analysis tools cannot d

Science Earth: Towards A Planet-Scale Operating System for AI-Native Scientific Discovery

Model ReleasesDGX agent

arXiv:2606.01316v1 Announce Type: new Abstract: Scientific discovery demands intelligence, perseverance, and serendipity across vast search spaces. Today, top scientific capabilities remain siloed--on

Score Function Gradient Estimation to Widen the Applicability of Decision-Focused Learning

ApplicationsDGX agent

arXiv:2307.05213v3 Announce Type: replace-cross Abstract: Many real-world optimization problems contain parameters that are unknown before deployment time, either due to stochasticity or to lack of in

SDR: Set-Distance Rewards for Radiology Report Generation

Model ReleasesDGX agent

arXiv:2606.00440v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has rapidly advanced reasoning in vision--language models. However, for chest X-ray report generation, th

SeClaw: Spec-Driven Security Task Synthesis for Evaluating Autonomous Agents

Model ReleasesDGX agent

arXiv:2606.02302v1 Announce Type: cross Abstract: Autonomous LLM agents increasingly operate in stateful environments where they access tools, files, memory, and external services. While such capabili

SECUREVENT: Hybrid AI/ML Security Monitoring for Distributed Event-Based Systems

SafetyDGX agent

arXiv:2606.01741v1 Announce Type: cross Abstract: Distributed event-based systems have become a common substrate for Internet-scale publish/subscribe services, IoT telemetry, cloud-native microservice

Self-Conditioned Positional HNSW for Overlap-Aware Retrieval in Chunked-Document RAG Systems: Method and Industrial Evidence-Quality Audit

ResearchDGX agent

arXiv:2606.01542v1 Announce Type: cross Abstract: Chunked-document retrieval is a common component of retrieval-augmented generation (RAG) systems. Documents are split into overlapping chunks, embedde

Self-Healing Agentic Orchestrators for Reliable Tool-Augmented Large Language Model Systems

Model ReleasesDGX agent

arXiv:2606.01416v1 Announce Type: new Abstract: Tool-augmented large language model (LLM) agents rely on orchestration layers that coordinate planning, retrieval, tool invocation, validation, memory,

Self-Revising Discovery Systems for Science: A Categorical Framework for Agentic Artificial Intelligence

AgentsDGX agent

arXiv:2606.01444v1 Announce Type: new Abstract: Scientific discovery is not only answer generation but revision of the representational regime in which evidence, artifacts, operations, and verifiers a

Self-supervised Monocular Depth and Pose Estimation for Endoscopy with Latent Priors

ResearchDGX agent

arXiv:2411.17790v3 Announce Type: replace-cross Abstract: Accurate 3D mapping in endoscopy enables quantitative, holistic lesion characterization within the gastrointestinal (GI) tract, requiring reli

SEMBridge: Tagless-Final Program Semantics with Weakest-Precondition and Bounded-Checking Interpretations

ResearchDGX agent

arXiv:2606.00220v1 Announce Type: cross Abstract: Formal methods provide rigorous accounts of program behavior, but practical software engineering often works through executable libraries, tests, and

SENSE: Semantic Embedding Navigation with Soft-gated Evaluation for Retrieval-based Speculative Decoding

Model ReleasesDGX agent

arXiv:2606.00021v1 Announce Type: cross Abstract: Speculative Decoding (SD) accelerates Large Language Model (LLM) inference by employing a lightweight draft model to propose candidate tokens, which a

SentimentLens: Reconciling Sentiment and Ratings via Dual-Modality in the Hospitality Sector

SafetyDGX agent

arXiv:2606.00084v1 Announce Type: cross Abstract: Online travel platforms generate vast volumes of user-generated hotel reviews, offering rich opportunities to understand traveler experiences at scale

Shape Your Body: Value Gradients for Multi-Embodiment Robot Design

SafetyDGX agent

arXiv:2606.00702v1 Announce Type: cross Abstract: We propose to turn generalist multi-embodiment value functions into reusable models for robot design. Instead of running a new reinforcement learning

ShapeLib: Designing a library of programmatic 3D shape abstractions with Large Language Models

ResearchDGX agent

arXiv:2502.08884v3 Announce Type: replace-cross Abstract: We present ShapeLib, the first method that uses the priors of Large Language Models (LLMs) to design libraries of programmatic 3D shape abstra

SHARP: Sleep-based Hierarchical Accelerated Replay for Long Range Non-Stationary Temporal Pattern Recognition

Model ReleasesDGX agent

arXiv:2606.00732v1 Announce Type: new Abstract: Learning long-range non-stationary temporal patterns remains a core challenge for modern sequence models, particularly in strict streaming settings. In

ShelfAware: Real-Time Semantic Localization in Quasi-Static Environments with Low-Cost Sensors

ApplicationsDGX agent

arXiv:2512.09065v2 Announce Type: replace-cross Abstract: Many indoor workspaces are quasi-static: their global geometric layout is stable, but local semantics change continually, producing repetitive

SHERLOCK: Towards Dynamic Knowledge Adaptation in LLM-enhanced E-commerce Risk Management

SafetyDGX agent

arXiv:2510.08948v4 Announce Type: replace-cross Abstract: Effective e-commerce risk management requires in-depth case investigations to identify emerging fraud patterns in highly adversarial environme

Short-form Text Rewriting with Phi Silica

Model ReleasesDGX agent

arXiv:2606.00462v1 Announce Type: cross Abstract: Short-form text rewriting is a constrained variant of paraphrasing in which limited context and high semantic density leave little room for variation.

Shortcut to Nowhere: Demystifying Deep Spurious Regression

ApplicationsDGX agent

arXiv:2606.01723v1 Announce Type: cross Abstract: Real-world regression often exhibits shortcuts: attributes that are spuriously correlated with continuous targets in training, yet unreliable under de

Silent Failures in Federated Personalization of Foundation Models

SafetyDGX agent

arXiv:2606.00947v1 Announce Type: cross Abstract: Foundation models are increasingly personalized on decentralized private data through federated learning and are now deployed at scale under growing r

Silent Failures in Physical AI: A Literature Review of Runtime Action Authorization for Autonomous Systems

SafetyDGX agent

arXiv:2606.00090v1 Announce Type: cross Abstract: Physical AI systems increasingly map multimodal observations, language instructions, and learned world representations into physically consequential a

SilentDrift: Exploiting Action Chunking for Stealthy Backdoor Attacks on Vision-Language-Action Models

SafetyDGX agent

arXiv:2601.14323v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models are increasingly deployed in safety-critical robotic applications, yet their security vulnerabilities rema

SimSD: Simple Speculative Decoding in Diffusion Language Models

ResearchDGX agent

arXiv:2606.02544v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) have recently emerged as a promising alternative to autoregressive (AR) LLMs, offering faster inference throug

Simulating Macroeconomic Expectations in Survey Experiments with LLM-based Economic Agents

ResearchDGX agent

arXiv:2505.17648v5 Announce Type: replace-cross Abstract: We introduce a framework for simulating macroeconomic expectations in survey experiments using LLM-based economic agents (LLM Agents). We cons

Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards

SafetyDGX agent

arXiv:2510.01167v2 Announce Type: replace-cross Abstract: Aligning large language models to human preferences is inherently multidimensional, yet most pipelines collapse heterogeneous signals into a s

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training

SafetyDGX agent

arXiv:2606.02355v1 Announce Type: new Abstract: Long-horizon LLM agents can benefit from reusable skills, yet existing skill-based methods often rely on external skill generators during training or pe

SIRIUS-SQL: Anchoring Multi-Candidate Text-to-SQL in Execution Feedback

ResearchDGX agent

arXiv:2606.01246v1 Announce Type: new Abstract: Text-to-SQL on complex schemas is unreliable on a single pass, so recent systems generate multiple SQL candidates and let voting filter out errors. Yet

Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills

HardwareDGX agent

arXiv:2503.05641v4 Announce Type: replace-cross Abstract: Combining existing pre-trained LLMs is a promising approach for diverse reasoning tasks. However, task-level expert selection is often too coa

'Skill issues'': data-centric optimization of lakehouse agents

AgentsDGX agent

arXiv:2606.01185v1 Announce Type: new Abstract: Coding agents are becoming users of data infrastructure, but their success depends not only on model quality: it also depends on the skills and environm

Skill or Skip? Learning Selective Skill Invocation in Agentic Tasks via Dual-Granularity Preference Learning

SafetyDGX agent

arXiv:2606.00510v1 Announce Type: cross Abstract: Agent skills are callable procedural modules that provide reusable knowledge and execution policies for complex agentic tasks. However, existing metho

SkillAdaptor: Self-Adapting Skills for LLM Agents from Trajectories

Model ReleasesDGX agent

arXiv:2606.01311v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly rely on reusable external skills to solve long-horizon interactive tasks. Existing training-free skill

SkillPager: Query-Adaptive Intra-Skill Navigation via Semantic Node Retrieval

Model ReleasesDGX agent

arXiv:2606.00822v1 Announce Type: cross Abstract: Skill-based LLM agents increasingly rely on long procedural documents, but full-document prompting wastes tokens and dilutes information critical to e

SkillRevise: Improving LLM-Authored Agent Skills via Trace-Conditioned Skill Revision

AgentsDGX agent

arXiv:2606.01139v1 Announce Type: new Abstract: Agent skills are procedural artifacts that enable LLM agents to execute workflows, verify constraints, and recover from failures. Existing self-evolving

SkillSmith: Co-Evolving Skills and Tools for Self-Improving Agent Systems

AgentsDGX agent

arXiv:2606.01314v1 Announce Type: new Abstract: Recent self-evolving agents have shown that skills can be discovered, refined, and accumulated through execution. However, existing skill-evolution fram

SkyShield: Occupancy as a Safety Interface for Low-Altitude UAV Autonomy

Model ReleasesDGX agent

arXiv:2606.00747v1 Announce Type: cross Abstract: For low-altitude Unmanned Aerial Vehicle (UAV) autonomy, 3D spatial understanding is not merely a perception objective, but the safety interface betwe

SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes

Model ReleasesDGX agent

arXiv:2606.01912v1 Announce Type: new Abstract: Smart homes are evolving toward complex state-dependent living environments, requiring Large Language Models (LLMs) to reason over user intent, preferen

Soft-NBCE: Entropy-Weighted Chunk Fusion for Long-Context

ResearchDGX agent

arXiv:2606.01101v1 Announce Type: cross Abstract: The quadratic complexity of self-attention remains a bottleneck for Large Language Models (LLMs) processing ultra-long contexts. The Naive Bayes Cogni

SORA: Free Second-Order Attacks in Fast Adversarial Training

SafetyDGX agent

arXiv:2606.00738v1 Announce Type: cross Abstract: Adversarial Training (AT) is a leading defense against adversarial examples but often suffers from Catastrophic Overfitting (CO) in efficient single-s

SortingHat: Redefining Operating Systems Education with a Tailored Digital Teaching Assistant

AgentsDGX agent

arXiv:2606.00015v1 Announce Type: cross Abstract: Operating Systems (OS) courses are among the most challenging in computer science education due to the complexity of internal structures and the diver

SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Divergence

Model ReleasesDGX agent

arXiv:2606.02380v1 Announce Type: cross Abstract: As LLM-based agents expand their operational scope, reliability becomes a prerequisite for real-world deployment. However, in practical applications,

SPADER: Step-wise Peer Advantage with Diversity-Aware Exploration Rewards for Multi-Answer Question Answering

SafetyDGX agent

arXiv:2606.00593v1 Announce Type: cross Abstract: Large language models are increasingly deployed as tool-augmented agents to acquire information beyond parametric knowledge. While recent work has imp

Spatial Representation Learning Beyond Pixels: Unifying Raster Data and Vector Semantics for Human-Centric Geospatial Foundation Models

TutorialsDGX agent

arXiv:2606.02374v1 Announce Type: new Abstract: Earth Observation (EO) has fundamentally transformed the monitoring of environmental processes and human activities up to planetary scale. Recent advanc

SpeedAug: Policy Acceleration via Tempo-Enriched Policy and RL Fine-Tuning

SafetyDGX agent

arXiv:2512.00062v2 Announce Type: replace-cross Abstract: Robotic policy learning for complex real-world manipulation tasks has seen rapid recent progress, enabled in large part by the ability to coll

SpikeWFM: Spiking-Aided Wireless Foundation Model for Robust Channel Prediction

TutorialsDGX agent

arXiv:2606.00120v1 Announce Type: cross Abstract: This paper proposes SpikeWFM, a novel hybrid architecture that integrates spiking neural networks (SNNs) with conventional artificial neural network (

SS-ZKR: Spatial-Semantic Zero-Knowledge Routing for Privacy-Preserving Multi-Agent Collaboration

SafetyDGX agent

arXiv:2606.00962v1 Announce Type: cross Abstract: Foundational agent interoperability standards, notably the Agent-to-Agent (A2A) protocol and the Model Context Protocol (MCP), have advanced multi-age

Stability Analysis of Sharpness-Aware Minimization

Model ReleasesDGX agent

arXiv:2301.06308v2 Announce Type: replace-cross Abstract: Sharpness-aware minimization (SAM) is a training method that seeks to find flat minima in deep learning, resulting in state-of-the-art perform

← Previous
1…179180181182183…358
Next →