AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
12 May 2026

RewardHarness: Self-Evolving Agentic Post-Training

Model ReleasesDGX agent

arXiv:2605.08703v1 Announce Type: new Abstract: Evaluating instruction-guided image edits requires rewards that reflect subtle human preferences, yet current reward models typically depend on large-sc

RigidFormer: Learning Rigid Dynamics using Transformers

SafetyDGX agent

arXiv:2605.09196v1 Announce Type: cross Abstract: Learning-based simulation of multi-object rigid-body dynamics remains difficult because contact is discontinuous and errors compound over long horizon

RL Fine-Tuning Heals OOD Forgetting in SFT

ResearchDGX agent

arXiv:2509.12235v3 Announce Type: replace-cross Abstract: Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL) is a standard post-training recipe for improving Large Language Models (L


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Robust Building Damage Detection in Cross-Disaster Settings Using Domain Adaptation

ResearchDGX agent

arXiv:2603.14694v2 Announce Type: replace-cross Abstract: Rapid structural damage assessment from remote sensing imagery is essential for timely disaster response. Within human-machine systems (HMS) f

Robust Multi-Agent LLMs under Byzantine Faults

AgentsDGX agent

arXiv:2605.09076v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly collaborate over peer-to-peer networks to improve their reliability. However, these same interactions c

Robust Probabilistic Shielding for Safe Offline Reinforcement Learning

SafetyDGX agent

arXiv:2605.10293v1 Announce Type: cross Abstract: In offline reinforcement learning (RL), we learn policies from fixed datasets without environment interaction. The major challenges are to provide gua

ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention

Model ReleasesDGX agent

arXiv:2603.22016v2 Announce Type: replace-cross Abstract: Large Reasoning Models (LRMs) often reach a correct solution before their long Chain-of-Thought trace ends, yet continue with redundant verifi

Route by State, Recover from Trace: STAR with Failure-Aware Markov Routing for Multi-Agent Spatiotemporal Reasoning

SafetyDGX agent

arXiv:2605.10057v1 Announce Type: new Abstract: Compositional spatiotemporal reasoning often requires a system to invoke multiple heterogeneous specialists, such as geometric, temporal, topological, a

RuPLaR : Efficient Latent Compression of LLM Reasoning Chains with Rule-Based Priors From Multi-Step to One-Step

SafetyDGX agent

arXiv:2605.09346v1 Announce Type: cross Abstract: The Chain-of-Thought (CoT) paradigm, while enhancing the interpretability of Large Language Models (LLMs), is constrained by the inefficiencies and ex

RW-Post: Auditable Evidence-Grounded Multimodal Fact-Checking in the Wild

Model ReleasesDGX agent

arXiv:2605.10357v1 Announce Type: cross Abstract: Multimodal misinformation increasingly leverages visual persuasion, where repurposed or manipulated images strengthen misleading text. We introduce ex

S2P-Net: A Spectral-Spatial Polar Network for Rotation-Invariant Object Recognition in Low-Data Regimes

ResearchDGX agent

arXiv:2605.09667v1 Announce Type: cross Abstract: We present S2P-Net (Spectral-Spatial Polar Network), a compact deep learning architecture that achieves mathematically guaranteed rotation invariance

SAFformer:Improving Spiking Transformer via Active Predictive Filtering

ResearchDGX agent

arXiv:2605.08270v1 Announce Type: cross Abstract: Spiking Neural Networks (SNNs) offer notable advantages in biological plausibility and energy efficiency, making them promising candidates for buildin

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models

SafetyDGX agent

arXiv:2510.20129v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) remain vulnerable to jailbreak attacks, where adversarially crafted prompts induce policy-violating responses des

Sanity Checks for Long-Form Hallucination Detection

ResearchDGX agent

arXiv:2605.08346v1 Announce Type: cross Abstract: Hallucination detection methods for large language models increasingly operate on chain-of-thought reasoning traces, yet it remains unclear whether th

SAR-RAG: ATR Visual Question Answering by Semantic Search, Retrieval, and MLLM Generation

AgentsDGX agent

arXiv:2602.04712v2 Announce Type: replace-cross Abstract: We present a visual-context image-retrieval-augmented generation (ImageRAG)- assisted AI agent for automatic target recognition (ATR) of synth

SARL: Label-Free Reinforcement Learning by Rewarding Reasoning Topology

SafetyDGX agent

arXiv:2603.27977v2 Announce Type: replace Abstract: Reinforcement learning is critical to improving large reasoning models, but its success relies heavily on verifiable rewards (RLVR), making it hard

SayNext-Bench: Why Do LLMs Struggle with Next-Utterance Anticipation?

Model ReleasesDGX agent

arXiv:2602.00327v2 Announce Type: replace Abstract: We explore the use of large language models (LLMs) for next-utterance anticipation in human dialogue. Despite recent advances in LLMs demonstrating

SCALAR: A Neurosymbolic Framework for Automated Conjecture and Reasoning in Quantum Circuit Analysis

Model ReleasesDGX agent

arXiv:2605.10327v1 Announce Type: cross Abstract: In this paper, we present SCALAR (Symbolic Conjecture and LLM-Assisted Reasoning), a neurosymbolic framework for automated conjecture generation in qu

Scaling Limits of Long-Context Transformers

Model ReleasesDGX agent

arXiv:2605.08505v1 Announce Type: cross Abstract: We study the long-context limit of softmax self-attention with a fixed query and a random context of n i.i.d. keys on the sphere, viewing the inverse

Scaling Vision Models Does Not Consistently Improve Localisation-Based Explanation Quality

Model ReleasesDGX agent

arXiv:2605.10142v1 Announce Type: cross Abstract: Artificial intelligence models are increasingly scaled to improve predictive accuracy, yet it remains unclear whether scale improves the quality of po

Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMs

Model ReleasesDGX agent

arXiv:2509.02372v3 Announce Type: replace-cross Abstract: Large Language Models have become critical to modern software development, but their reliance on uncurated web-scale datasets for training int

ScholarPeer: A Context-Aware Multi-Agent Framework for Automated Peer Review

AgentsDGX agent

arXiv:2601.22638v2 Announce Type: replace-cross Abstract: The exponential growth of machine learning submissions has strained the traditional peer review process, resulting in slow feedback loops for

SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems

Model ReleasesDGX agent

arXiv:2605.10246v1 Announce Type: new Abstract: AI scientist systems are increasingly deployed for autonomous research, yet their academic integrity has never been systematically evaluated. We introdu

SDFlow: Similarity-Driven Flow Matching for Time Series Generation

SafetyDGX agent

arXiv:2605.05736v2 Announce Type: replace Abstract: Vector quantization (VQ) with autoregressive (AR) token modeling is a widely adopted and highly competitive paradigm for time-series generation. How

SDG-MoE: Signed Debate Graph Mixture-of-Experts

ResearchDGX agent

arXiv:2605.08322v1 Announce Type: cross Abstract: Sparse MoE models achieve a good balance between capacity and compute by routing each token to a small subset of experts. However, in most MoE archite

SDTalk: Structured Facial Priors and Dual-Branch Motion Fields for Generalizable Gaussian Talking Head Synthesis

ResearchDGX agent

arXiv:2605.09956v1 Announce Type: cross Abstract: High-quality, real-time talking head synthesis remains a fundamental challenge in computer vision. Existing reconstruction- and rendering-based method

SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization

TutorialsDGX agent

arXiv:2602.04811v2 Announce Type: replace-cross Abstract: True self-evolution requires agents to act as lifelong learners that internalize novel experiences to solve future problems. However, rigorous

SearchSkill: Teaching LLMs to Use Search Tools with Evolving Skill Banks

ResearchDGX agent

arXiv:2605.09038v1 Announce Type: new Abstract: Teaching language models to use search tools is not only a question of whether they search, but also of whether they issue good queries. This is especia

Security Risks in Tool-Enabled AI Agents: A Systematic Analysis of Privileged Execution Environments

AgentsDGX agent

arXiv:2605.09721v1 Announce Type: cross Abstract: Tool-enabled AI agents are increasingly deployed in cloud-hosted environments and offered as services, where they perform side-effecting operations th

Seed Hijacking of LLM Sampling and Quantum Random Number Defense

Model ReleasesDGX agent

arXiv:2605.08313v1 Announce Type: cross Abstract: Large language models (LLMs) rely on deterministic pseudorandom number generators (PRNGs) for autoregressive sampling, creating a critical supply-chai

SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning

Model ReleasesDGX agent

arXiv:2605.09266v1 Announce Type: new Abstract: We introduce SeePhys Pro, a fine-grained modality transfer benchmark that studies whether models preserve the same reasoning capability when critical in

Select-then-differentiate: Solving Bilevel Optimization with Manifold Lower-level Solution Sets

Local AiDGX agent

arXiv:2605.09209v1 Announce Type: cross Abstract: We study optimistic bilevel optimization when the lower-level problem has a non-isolated manifold of minimizers. In this setting, the hyper-objective

Selective Deficits in LLM Mental Self-Modeling in a Behavior-Based Test of Theory of Mind

Model ReleasesDGX agent

arXiv:2603.26089v2 Announce Type: replace-cross Abstract: The ability to represent oneself and others as agents with knowledge, intentions, and belief states that guide their behavior - Theory of Mind

Selective LoRA for Visual Tokens and Attention Heads

Model ReleasesDGX agent

arXiv:2512.19219v2 Announce Type: replace-cross Abstract: Low-rank adaptation (LoRA) is widely used for parameter-efficient fine-tuning, but its standard all-token, all-head design ignores the heterog

Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language Models

ResearchDGX agent

arXiv:2605.08145v1 Announce Type: cross Abstract: Current vision language models face hallucination and robustness issues against ambiguous or corrupted modalities. We hypothesize that these issues ca

Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories

SafetyDGX agent

arXiv:2605.08936v1 Announce Type: new Abstract: Large Reasoning Models possess remarkable capabilities for self-correction in general domain; however, they frequently struggle to recover from unsafe r

Semantic Voting: Execution-Grounded Consensus for LLM Code Generation

ResearchDGX agent

arXiv:2605.08680v1 Announce Type: cross Abstract: LLM code-generation pipelines often sample multiple candidates and select one final answer without access to a complete oracle. Existing pipelines mix

Semi-Supervised Neural Super-Resolution for Mesh-Based Simulations

Model ReleasesDGX agent

arXiv:2605.09284v1 Announce Type: cross Abstract: Mesh-based simulations provide high-fidelity solutions to partial differential equations (PDEs), but achieving such accuracy typically requires fine m

SenseBench: A Benchmark for Remote Sensing Low-Level Visual Perception and Description in Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.10576v1 Announce Type: cross Abstract: Low-level visual perception underpins reliable remote sensing (RS) image analysis, yet current image quality assessment (IQA) methods output uninterpr

Separate First, Fuse Later: Mitigating Cross-Modal Interference in Audio-Visual LLMs Reasoning with Modality-Specific Chain-of-Thought

Model ReleasesDGX agent

arXiv:2605.09906v1 Announce Type: new Abstract: Audio and vision provide complementary evidence for audio-visual question answering, yet current audio-visual large language models may suffer from cros

Sequential Feature Selection for Efficient Landslide Segmentation from Multi-Spectral Data

Model ReleasesDGX agent

arXiv:2605.09746v1 Announce Type: cross Abstract: Landslide detection from satellite imagery has advanced through deep learning, yet most models rely on large, highly correlated spectral-topographic i

SGC-RML: A reliable and interpretable longitudinal assessment for PD in real-world DNS

SafetyDGX agent

arXiv:2605.08302v1 Announce Type: cross Abstract: Real-world digital Parkinson's disease assessment faces challenges such as heterogeneous modalities, cross-device bias, and incomplete labeling. Exist

ShadowMerge: A Novel Poisoning Attack on Graph-Based Agent Memory via Relation-Channel Conflicts

AgentsDGX agent

arXiv:2605.09033v1 Announce Type: cross Abstract: Graph-based agent memory is increasingly used in LLM agents to support structured long-term recall and multi-hop reasoning, but it also creates a new

Shaping Schema via Language Representation as the Next Frontier for LLM Intelligence Expanding

ApplicationsDGX agent

arXiv:2605.09271v1 Announce Type: new Abstract: Although natural language is the default medium for Large Language Models (LLMs), its limited expressive capacity creates a profound bottleneck for comp

Shapley Regression for Rare Disease Diagnosis Support: a case study on APDS

ApplicationsDGX agent

arXiv:2605.08897v1 Announce Type: cross Abstract: Activated PI3K8 Syndrome (APDS) is a rare genetic immune disorder caused by variants in PIK3CD or PIK3R1, with highly heterogeneous symptoms that ofte

Shepherd: A Runtime Substrate Empowering Meta-Agents with a Formalized Execution Trace

AgentsDGX agent

arXiv:2605.10913v1 Announce Type: new Abstract: We introduce Shepherd, a functional programming model that formalizes meta-agent operations on target agents as functions, with core operations mechaniz

Shields to Guarantee Probabilistic Safety in MDPs

SafetyDGX agent

arXiv:2605.10888v1 Announce Type: cross Abstract: Shielding is a prominent model-based technique to ensure safety of autonomous agents. Classical shielding aims to ensure that nothing bad ever happens

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization

ResearchDGX agent

arXiv:2605.08809v1 Announce Type: cross Abstract: Pretraining large language models (LLMs) with next-token prediction has led to remarkable advances, yet the context-dependent nature of token embeddin

Simulating Complex Multi-Turn Tool Calling Interactions in Stateless Execution Environments

AgentsDGX agent

arXiv:2601.19914v2 Announce Type: replace-cross Abstract: Synthetic data has proven itself to be a valuable resource for tuning smaller, cost-effective language models to handle the complexities of mu

Simultaneous Long-tailed Recognition and Multi-modal Fusion for Highly Imbalanced Multi-modal Data

Model ReleasesDGX agent

arXiv:2605.10498v1 Announce Type: cross Abstract: Long-tailed distributions in class-imbalanced data present a fundamental challenge for deep learning models, which tend to be biased toward majority c

Simulus: Combining Improvements in Sample-Efficient World Model Agents

AgentsDGX agent

arXiv:2502.11537v4 Announce Type: replace-cross Abstract: World models (WMs) represent the frontier of sample-efficient reinforcement learning, but their complexity leaves many promising improvements

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning

AgentsDGX agent

arXiv:2605.09423v1 Announce Type: new Abstract: LLM/VLM-based digital agents have advanced rapidly thanks to scalable sandboxes for coding, web navigation, and computer use, which provide rich interac

Single-Configuration Attack Success Rate Is Not Enough: Jailbreak Evaluations Should Report Distributional Attack Success

Model ReleasesDGX agent

arXiv:2605.09070v1 Announce Type: cross Abstract: Many jailbreak attack research papers report attack success rates for a limited number of parameter settings, even though there are many combinations

Sink vs. diagonal patterns as mechanisms for attention switch and oversmoothing prevention

SafetyDGX agent

arXiv:2605.08453v1 Announce Type: cross Abstract: This paper studies the role of sinks and diagonal patterns as attention switch and anti-oversmoothing mechanisms. We analyze geometric conditions unde

Sketch-and-Verify: Structured Inference-Time Scaling via Program Sketching

Model ReleasesDGX agent

arXiv:2605.08658v1 Announce Type: cross Abstract: SKETCHVERIFY is a within-tier cost-performance policy, not a universal accuracy improvement. The operational question: a practitioner stuck with a sma

SKG-VLA: Scene Knowledge Graph Priors for Structured Scene Semantics and Multimodal Reasoning for Decision Making

SafetyDGX agent

arXiv:2605.09343v1 Announce Type: new Abstract: Decision making in large-scale complaint handling systems increasingly relies on heterogeneous evidence, including complaint narratives, screenshots, or

Skill-R1: Agent Skill Evolution via Reinforcement Learning

SafetyDGX agent

arXiv:2605.09359v1 Announce Type: cross Abstract: Agentic large language models often rely on skills, reusable natural language procedures that guide planning, action, and tool use. In practice, skill

SkillEvolver: Skill Learning as a Meta-Skill

HardwareDGX agent

arXiv:2605.10500v1 Announce Type: new Abstract: Agent skills today are static artifact: authored once -- by human curation or one-shot generation from parametric knowledge -- and then consumed unchang

SkillLens: Adaptive Multi-Granularity Skill Reuse for Cost-Efficient LLM Agents

AgentsDGX agent

arXiv:2605.08386v1 Announce Type: new Abstract: Skill libraries have become a practical way for LLM agents to reuse procedural experience across tasks. However, existing systems typically treat skills

SkillMaster: Toward Autonomous Skill Mastery in LLM Agents

AgentsDGX agent

arXiv:2605.08693v1 Announce Type: new Abstract: Skills provide an effective mechanism for improving LLM agents on complex tasks, yet in existing agent frameworks, their creation, refinement, and selec

← Previous
1…273274275276277…358
Next →