AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
26 May 2026

Safety Generalization Under Distribution Shift in Safe Reinforcement Learning: A Diabetes Testbed

Model ReleasesDGX agent

arXiv:2601.21094v2 Announce Type: replace-cross Abstract: Safe Reinforcement Learning (RL) algorithms are typically evaluated under fixed training conditions. We investigate whether training-time safe

Safety-Oriented Routing Analysis of Mixtral MoE Under Benign and Harmful Prompts

SafetyDGX agent

arXiv:2605.24270v1 Announce Type: new Abstract: Sparse mixture-of-experts (MoE) language models activate only a small subset of parameters for each token, making router behavior a central part of mode

SAM: State-Adaptive Memory for Long-Horizon Reasoning Agent

AgentsDGX agent

arXiv:2605.24468v1 Announce Type: new Abstract: Long-horizon agentic reasoning requires large language models to act over long interaction histories containing thoughts, tool calls, observations, and


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SAMark: A Self-Anchored Text Watermarking with Paragraph-Level Paraphrase Robustness

ResearchDGX agent

arXiv:2605.25796v1 Announce Type: cross Abstract: Semantic-level watermarking (SWM) improves robustness against text modifications by treating sentences as the basic unit. However, robustness to parag

Saturating Scaling Laws for Equational Discovery: A Phenomenology of Growth Dynamics in Three Toy Substrates with Two Real-World Replications

ApplicationsDGX agent

arXiv:2605.23983v1 Announce Type: new Abstract: We investigate growth dynamics in deterministic equational discovery substrates. Across three toy domains (arithmetic, boolean, higher-order list; n=592

Scale When Needed: Adaptive Neuron-level Mixed Precision Quantization Aware Training

ResearchDGX agent

arXiv:2605.25054v1 Announce Type: cross Abstract: Deploying deep neural networks on resource-constrained 6G edge devices demands aggressive compression with minimal accuracy loss. Quantization-Aware T

ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training

Model ReleasesDGX agent

arXiv:2605.24326v1 Announce Type: cross Abstract: The rapid scaling of large language model training requires distributing GPU resources across multiple data center buildings and regions. We refer to

Scaling up Energy-Aware Multi-Agent Reinforcement Learning for Mission-Oriented Drone Networks with Individual Reward

AgentsDGX agent

arXiv:2605.24992v1 Announce Type: cross Abstract: Multi-agent reinforcement learning (MARL) has shown wide applicability in collaborative systems such as autonomous driving and smart cities for its ab

Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models

Model ReleasesDGX agent

arXiv:2605.25394v1 Announce Type: new Abstract: Large language models often generate confident but incorrect answers rather than abstaining when uncertain. This problem is particularly acute for small

Security in the Fine-Tuning Lifecycle of Large Language Models: Threats, Defenses,Evaluation, and Future Directions

Model ReleasesDGX agent

arXiv:2605.25073v1 Announce Type: cross Abstract: Background: Fine-tuning is central to adapting pre-trained Large Language Models (LLMs) to downstream tasks, but its reliance on training data, parame

Security of OpenClaw Agents: Fundamentals, Attacks, and Countermeasures

AgentsDGX agent

arXiv:2605.25435v1 Announce Type: new Abstract: The rapid evolution of large language model (LLM)-driven autonomous agents has given rise to OpenClaw, a new class of open-source agent frameworks that

Selection-Induced Contraction of Innovation Statistics in Gated Kalman Filters

ResearchDGX agent

arXiv:2512.18508v3 Announce Type: replace-cross Abstract: Validation gating is a fundamental component of classical Kalman-based tracking systems. Only measurements whose normalized innovation squared

Selective Test-Time Compute Scaling for Click-Through Rate Prediction via Uncertainty-Triggered Feature Path Exploration

TutorialsDGX agent

arXiv:2605.24989v1 Announce Type: cross Abstract: Scaling test-time compute has proven highly effective for language models, yet this opportunity remains largely unexplored for industrial Click-Throug

SemanticZip: A Pilot Framework for Lossy Text Compression with LLMs as Semantic Decompressors

Model ReleasesDGX agent

arXiv:2605.24541v1 Announce Type: cross Abstract: Text compression for large language model (LLM) systems is usually framed as token deletion, retrieval, summarization, or exact reconstruction. We stu

Sensing Intelligence as a Trainable Metamaterial Property

ResearchDGX agent

arXiv:2605.23967v1 Announce Type: cross Abstract: In biological systems, sensing is not performed by the brain alone: the body deforms, vibrates, and filters external stimuli before they are transduce

SentGraph: Hierarchical Sentence Graph for Multi-hop Retrieval-Augmented Question Answering

ResearchDGX agent

arXiv:2601.03014v3 Announce Type: replace-cross Abstract: Traditional Retrieval-Augmented Generation (RAG) effectively supports single-hop question answering with large language models but faces signi

SEP-Attack: A Simple and Effective Paradigm for Transfer-Based Textual Adversarial Attack

ApplicationsDGX agent

arXiv:2605.24958v1 Announce Type: cross Abstract: Despite the strong performance of deep neural networks in modern Web and language applications, they remain vulnerable to adversarial attacks, especia

SeqRoute: Global Budget-Aware Sequential LLM Routing via Offline Reinforcement Learning

ApplicationsDGX agent

arXiv:2605.25424v1 Announce Type: cross Abstract: Existing LLM routing frameworks treat queries as independent events, neglecting the sequential nature of real-world user sessions constrained by globa

Side-by-side Comparison Amplifies Dialect Bias in Language Models

SafetyDGX agent

arXiv:2605.24384v1 Announce Type: cross Abstract: Language models (LMs) can exhibit systematic biases against speakers based on variations in their dialects, even in the absence of a dialect label, a

Signs Beat Floats: Low-Rank Double-Binary Adaptation for On-Device Fine-Tuning

Local AiDGX agent

arXiv:2605.24058v1 Announce Type: cross Abstract: On-device adaptation of large language models commonly keeps a quantized base model frozen while training and deploying a small, task-specific LoRA ad

Simulating Human Memory with Language Models

ApplicationsDGX agent

arXiv:2605.25680v1 Announce Type: cross Abstract: Language models are increasingly being deployed as user simulators, but their memory is far more reliable than that of real users. To measure this gap

'Si'multaneous 'S'patial-'T'emporal Message Passing for Dynamic Graph Representation Learning

Model ReleasesDGX agent

arXiv:2605.25548v1 Announce Type: cross Abstract: Dynamic graph neural networks (DGNNs) that operate on snapshot sequences typically fall into one of two categories. Temporal-first approaches build pe

SimuWoB: Simulating Real-World Mobile Apps for Fast and Faithful GUI Agent Benchmarking

Model ReleasesDGX agent

arXiv:2605.25160v1 Announce Type: new Abstract: Mobile GUI agents powered by large language models have progressed rapidly, creating urgent needs for realistic and comprehensive evaluation. Existing b

SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills

Model ReleasesDGX agent

arXiv:2605.24117v1 Announce Type: new Abstract: Large language model (LLM) agents accumulate rich episodic trajectories while solving real-world tasks, but it remains unclear whether such experience c

Small Models, Strong Priors: Architectural Inductive Bias for Parameter-Efficient Neural PDE Solvers

Model ReleasesDGX agent

arXiv:2605.25949v1 Announce Type: cross Abstract: Neural PDE solvers have followed the scaling trajectory of vision and language, with recent foundation models reaching billions of parameters. We argu

Smart Timing for Mining: A Deep Learning Framework for Bitcoin Hardware ROI Prediction

Model ReleasesDGX agent

arXiv:2512.05402v2 Announce Type: replace-cross Abstract: Bitcoin mining hardware acquisition requires strategic timing due to volatile markets, rapid technological obsolescence, and protocol-driven r

SMDD-Bench: Can LLMs Solve Real-World Small Molecule Drug Design Tasks?

Model ReleasesDGX agent

arXiv:2605.21740v2 Announce Type: replace Abstract: LLM agents have incredible potential for scientific discovery applications. However, the performance of LLM agents on real-world, small molecule dru

SODE: Analyzing Social Dynamics in LLM Agents

Model ReleasesDGX agent

arXiv:2605.23949v1 Announce Type: cross Abstract: As Large Language Models (LLMs) evolve into interactive agents, understanding their behavioral alignment within human social dynamics becomes essentia

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models

Model ReleasesDGX agent

arXiv:2506.18543v2 Announce Type: replace-cross Abstract: The rapid proliferation of Large Language Models (LLMs) has heightened concerns regarding their exposure to jailbreak attacks, which craft adv

SoK: DARPA's AI Cyber Challenge (AIxCC): Competition Design, Architectures, and Lessons Learned

AgentsDGX agent

arXiv:2602.07666v3 Announce Type: replace-cross Abstract: DARPA's AI Cyber Challenge (AIxCC, 2023--2025) is the largest competition to date for building fully autonomous cyber reasoning systems (CRSs)

Solving Combinatorial Counting Problems with Weighted First-Order Model Counting

ResearchDGX agent

arXiv:2605.24845v1 Announce Type: new Abstract: Combinatorial counting problems pervade artificial intelligence, statistics, and discrete mathematics. Whether the task is enumerating subsets, multiset

SomaliBench Eval: Measuring English-to-Somali Refusal Gaps in Open-Weight Language Models

Model ReleasesDGX agent

arXiv:2605.25420v1 Announce Type: cross Abstract: Large language model safety evaluation remains heavily English-centered, leaving low-resource languages under-measured even when models are deployed g

SPA-Cache: Singular Proxies for Adaptive Caching in Diffusion Language Models

ResearchDGX agent

arXiv:2602.02544v2 Announce Type: replace-cross Abstract: While Diffusion Language Models (DLMs) offer a flexible, arbitrary-order alternative to the autoregressive paradigm, their non-causal nature p

SPACE: Unifying Symmetric and Asymmetric Routing Problems for Generalist Neural Solver

ApplicationsDGX agent

arXiv:2605.24484v1 Announce Type: new Abstract: Generalist neural routing solvers have shown great potential in solving diverse vehicle routing problems (VRPs) with a unified model. However, existing

Spacetime Formation under Requirements: Contextual Realization and Form-Dependent Probability

ResearchDGX agent

arXiv:2605.23943v1 Announce Type: new Abstract: Quantum cognition often explains order effects, contextuality, and violations of the law of total probability by replacing classical probability with qu

SPARK: Search Personalization via Agent-Driven Retrieval and Knowledge-sharing

AgentsDGX agent

arXiv:2512.24008v3 Announce Type: replace Abstract: Personalized search demands the ability to model users' evolving, multi-dimensional information needs; a challenge for systems constrained by static

SpecAlign: A Semantic Alignment Framework for SystemVerilog Assertion Generation

SafetyDGX agent

arXiv:2605.25181v1 Announce Type: new Abstract: Existing Large Language Model (LLM) approaches to SystemVerilog Assertion (SVA) generation primarily focus on syntactic validity and formal verification

Specification-Based Code-Text-Code Reengineering for LLM-Mediated Software Evolution

ResearchDGX agent

arXiv:2605.25232v1 Announce Type: cross Abstract: Direct Code2Code transformation remains challenging to control because it can preserve surface-level syntax while introducing semantic drift, hidden b

SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning

Local AiDGX agent

arXiv:2509.05614v3 Announce Type: replace-cross Abstract: Pruning is a typical acceleration technique for compute-bound models by removing computation on unimportant values. Recently, it has been appl

Spectral Probe-Circuits: A Three-Step Recipe for Identifying Attention-Head Circuits in Pretrained Transformers

Model ReleasesDGX agent

arXiv:2605.24059v1 Announce Type: cross Abstract: We present a three-step recipe for identifying attention-head circuits in pretrained transformers. A per-head spectral signal -- the time-integrated p

Spectral Retrieval: Multi-Scale Sinc Convolution over Token Embeddings for Localized Retrieval in LLM Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2605.24764v1 Announce Type: cross Abstract: [Abridged] - Spectral Retrieval is a plug-in re-ranking stage that interpolates between per-token MaxSim and mean-pool retrieval through a multi-scale

Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation

ResearchDGX agent

arXiv:2605.26111v1 Announce Type: cross Abstract: Subject-driven image generation aims to synthesize new images that preserve the identity of the given subject while following textual instructions. Ex

StakeBench: Evaluating Language Understanding Grounded in Market Commitment

SafetyDGX agent

arXiv:2605.26074v1 Announce Type: cross Abstract: Existing financial NLP benchmarks often rely on labels supplied by outside observers, measuring how language is perceived rather than what speakers ha

Step-TP: A Grounded, Step-Level Dataset with Chain-of-Thought Reasoning for LLM-Guided Tensor Program Optimization

ResearchDGX agent

arXiv:2605.25954v1 Announce Type: cross Abstract: Despite the strong reasoning capabilities of large language models (LLMs), optimizing the execution efficiency of tensor programs remains challenging

Stop Comparing LLM Agents Without Disclosing the Harness

SafetyDGX agent

arXiv:2605.23950v1 Announce Type: new Abstract: This position paper argues that, for long-horizon tasks evaluated across models with comparable frontier capability, the agent execution harness, namely

Strat-Reasoner: Reinforcing Strategic Reasoning of LLMs in Multi-Agent Games

SafetyDGX agent

arXiv:2605.04906v2 Announce Type: replace Abstract: While Large Language Models (LLMs) excel in certain reasoning tasks, they struggle in multi-agent games where the final outcome depends on the joint

STREAM: A Data-Centric Framework for Mining High-Value Task-Oriented Dialogues from Streaming Media

Model ReleasesDGX agent

arXiv:2605.25162v1 Announce Type: cross Abstract: Large language models for vertical domains are bottlenecked by the scarcity of complex, domain-specific task-oriented dialogues. Existing data acquisi

StructBreak: Structural Cognitive Overload-Induced Safety Failures in MLLMs

Model ReleasesDGX agent

arXiv:2605.25534v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) excel at structural reasoning yet suffer from a sharp logical brittleness in structural consistency. We term th

Subspace Aggregation Query and Index Generation for Multidimensional Resource Space Model

ResearchDGX agent

arXiv:2505.02129v3 Announce Type: replace-cross Abstract: Organizing large-scale resources in a multidimensional semantic space is an approach to efficiently managing and querying resources from diffe

Subspace-Guided Semantic and Topological Invariant Registration for Annotation-Free Ultrasound Plane Quality Control

SafetyDGX agent

arXiv:2605.25396v1 Announce Type: cross Abstract: Reliable quality control (QC) of ultrasound images is essential for both real-time acquisition guidance and retrospective clinical audit, yet existing

Summoning the Oracle to Slay It: Mitigating Look-Ahead Bias in Financial Backtesting with Large Language Models

SafetyDGX agent

arXiv:2605.24564v1 Announce Type: new Abstract: Backtesting large language models (LLMs) on historical financial data is unreliable because pre-training cuts off after the events happened. An LLM trai

TaBIIC2: Interactive Building of Ontological Taxonomies using Weighted Self-Organizing Maps

ResearchDGX agent

arXiv:2605.24899v1 Announce Type: new Abstract: Ontologies represent the conceptual knowledge of a domain. At the core of an ontology is the taxonomy of concepts and subconcepts that represent specifi

Task-Aligned Self-Supervised Learning for Medical Image Analysis: A Systematic Review and Practical Design Guidelines

SafetyDGX agent

arXiv:2605.23995v1 Announce Type: cross Abstract: Self-supervised learning (SSL) has emerged as a promising paradigm for addressing the annotation bottleneck in medical imaging by learning representat

Teaching large language models to reason like expert diagnosticians

Model ReleasesDGX agent

arXiv:2509.12194v2 Announce Type: replace Abstract: Differential diagnosis is an iterative process that integrates patient information with broader medical knowledge. Clinical case series such as the

Teaching Through Analogies: A Modular Pipeline for Educational Analogy Generation

Model ReleasesDGX agent

arXiv:2605.24211v1 Announce Type: cross Abstract: Analogies help learners understand unfamiliar concepts by relating them to known concepts. Despite recent advances, large language models (LLMs) conti

Temporal Concept Drift in Legal Judgment Prediction: Neural Baselines Across Three Epochs of Ukrainian Court Decisions

ApplicationsDGX agent

arXiv:2605.24452v1 Announce Type: cross Abstract: Legal NLP benchmarks evaluate models on randomly split data, implicitly assuming that legal language is stationary. We test this assumption by fine-tu

Test-Time Deep Thinking to Explore Implicit Rules

AgentsDGX agent

arXiv:2605.24828v1 Announce Type: new Abstract: With the continuous advancement of Large Language Models (LLMs), intelligent agents are becoming increasingly vital. However, these agents often fail in

Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation

Model ReleasesDGX agent

arXiv:2605.25488v1 Announce Type: cross Abstract: Audio-driven talking-head generation has achieved remarkable progress with recent models such as AniTalker, FLOAT, and Sonic. Despite their success, m

TGFormer: Towards Temporal Graph Transformer with Auto-Correlation Mechanism

ResearchDGX agent

arXiv:2605.24971v1 Announce Type: cross Abstract: The growing interest in Temporal Graph Neural Networks (TGNNs) stems from their ability to model complex dynamics and deliver superior performance. Ho

The Concept Allocation Zone: Tracking How Concepts Form Across Transformer Depth

SafetyDGX agent

arXiv:2605.24856v1 Announce Type: cross Abstract: Concept formation in transformer language models is depth-extended, not a single-layer event: concepts emerge gradually across a contiguous region of

← Previous
1…218219220221222…358
Next →