AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
1 Jun 2026

REAL: Regression-Aware Reinforcement Learning for LLM-as-a-Judge

SafetyDGX agent

arXiv:2603.17145v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed as automated evaluators that assign numeric scores to model outputs, a paradigm known a

Reasoning-Aware Multimodal Fusion for Hateful Video Detection

Local AiDGX agent

arXiv:2512.02743v2 Announce Type: replace-cross Abstract: Hate speech in online videos is posing an increasingly serious threat to digital platforms, especially as video content becomes increasingly m

Redefining Instance Matching: A Unified Framework for Part-Aware Matching in Panoptic Segmentation Evaluation

ResearchDGX agent

arXiv:2605.31094v1 Announce Type: cross Abstract: The Panoptic Quality (PQ) metric is the standard for jointly evaluating instance and semantic segmentation. However, its original definition relies on


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Regret-Based Federated Causal Discovery with Unknown Interventions

ResearchDGX agent

arXiv:2512.23626v2 Announce Type: replace Abstract: Most causal discovery methods recover a completed partially directed acyclic graph representing a Markov equivalence class from observational data.

Reinterpreting Safety Thresholds as Neuron Spiking Thresholds

SafetyDGX agent

arXiv:2605.30368v1 Announce Type: cross Abstract: Surrogate Safety Measures (SSMs) are extensively utilised in the evaluation of traffic risk in automated driving contexts. However, the majority of SS

Reliable Self-Improvement Training by Verifying Reasoning, Not Just Answers

TutorialsDGX agent

arXiv:2603.21558v2 Announce Type: replace Abstract: Self-improvement training, where models learn from self-generated solutions, promises sustained capability gains but suffers from a pervasive failur

Residual Reservoir Memory Networks

ResearchDGX agent

arXiv:2508.09925v3 Announce Type: replace-cross Abstract: We introduce a novel class of untrained Recurrent Neural Networks (RNNs) within the Reservoir Computing (RC) paradigm, called Residual Reservo

ReTabAD: A Benchmark for Restoring Semantic Context in Tabular Anomaly Detection

Model ReleasesDGX agent

arXiv:2510.02060v2 Announce Type: replace Abstract: In tabular anomaly detection (AD), textual semantics often carry critical signals, as the definition of an anomaly is closely tied to domain-specifi

Rethinking Multimodal Few-Shot 3D Point Cloud Segmentation: From Fused Refinement to Decoupled Arbitration

SafetyDGX agent

arXiv:2601.01456v2 Announce Type: replace-cross Abstract: In this paper, we revisit multimodal few-shot 3D point cloud semantic segmentation (FS-PCS), identifying a conflict in 'Fuse-then-Refine' para

Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't

ResearchDGX agent

arXiv:2605.30523v1 Announce Type: cross Abstract: Recent work describes what transformers can and cannot compute through connections to boolean circuits, but existing results lack exact characterizati

Reward Learning from Best-of-N Preference Data: Targets, Tradeoffs, and Design Principles

TutorialsDGX agent

arXiv:2605.30619v1 Announce Type: cross Abstract: Best-of-N sampling is widely used to construct pairwise preference data: N candidates are drawn from a base distribution, and the best is paired with

Routing on the Stiefel Manifold: When Does Adaptive Subspace Selection Help for Cross-Domain EEG Decoding?

SafetyDGX agent

arXiv:2605.31043v1 Announce Type: cross Abstract: Cross-domain EEG decoding remains challenging despite advances in Riemannian deep learning: covariance matrices from different subjects occupy systema

SAC-Opt: Semantic Anchors for Iterative Correction in Optimization Modeling

ResearchDGX agent

arXiv:2510.05115v3 Announce Type: replace Abstract: Large language models (LLMs) have opened new paradigms in optimization modeling by enabling the generation of executable solver code from natural la

SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse Autoencoders

Model ReleasesDGX agent

arXiv:2509.21379v3 Announce Type: replace-cross Abstract: Concept unlearning in diffusion models is hampered by feature splitting, where concepts are distributed across many latent features, making th

Safe Equilibrium Policy Optimization for Strategic Agent Policies

Model ReleasesDGX agent

arXiv:2605.30854v1 Announce Type: cross Abstract: Language models fine-tuned with reinforcement learning typically optimize for task reward, ignoring multi-agent strategic structure. Because these age

SAGE: A Novelty Gate for Efficient Memory Evolution in Agentic LLMs

AgentsDGX agent

arXiv:2605.30711v1 Announce Type: cross Abstract: Agentic LLMs must continuously decide whether newly extracted facts should be added, merged with existing memories, or ignored, yet prior work has foc

SAM for Robust Mitochondria Instance Segmentation in Fluorescence Microscopy

ApplicationsDGX agent

arXiv:2605.31284v1 Announce Type: cross Abstract: The morphological analysis of mitochondria in fluorescence microscopy (FM) is crucial for understanding cellular health, energy production, and metabo

Same Patient, Different Words, Different Diagnosis? Evaluating Semantic Stability in Clinical LLMs

Model ReleasesDGX agent

arXiv:2605.30646v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in clinical applications. However, their behavior remains highly sensitive to subtle linguistic var

SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer

HardwareDGX agent

arXiv:2605.30409v1 Announce Type: cross Abstract: Real-time streaming video-to-video editing (V2V) is critical for interactive applications such as live broadcasting and gaming, yet it remains a formi

Scalable Constrained Multi-Agent Reinforcement Learning via State Augmentation and Consensus for Separable Dynamics

SafetyDGX agent

arXiv:2605.30461v1 Announce Type: cross Abstract: We present a distributed approach for constrained Multi-Agent Reinforcement Learning (MARL) that combines state-augmented policy learning with distrib

Scaling Conversational Hungarian ASR: The BEA-Dialogue+ Corpus

Model ReleasesDGX agent

arXiv:2605.31469v1 Announce Type: cross Abstract: Conversational automatic speech recognition in Hungarian is constrained by the limited amount of publicly available dialogue-style training data. The

Scaling Higher-Order Graph Learning with Maximal Clique Complexes

ResearchDGX agent

arXiv:2605.31373v1 Announce Type: cross Abstract: Graph neural networks (GNNs) are limited to modeling pairwise interactions, while higher-order models based on cell complexes achieve greater expressi

Scaling Multi-Agent Environment Co-Design with Diffusion Models

SafetyDGX agent

arXiv:2511.03100v2 Announce Type: replace-cross Abstract: The agent-environment co-design paradigm jointly optimises agent policies and environment configurations in search of improved system performa

Scientific Machine Learning for Engine Health Management and Remaining Useful Life Prediction

ApplicationsDGX agent

arXiv:2605.30593v1 Announce Type: cross Abstract: Engine Health Management (EHM) depends on reliable forecasting of Remaining Useful Life (RUL) and on tracking thermal indicators such as turbine gas t

SCOPE: Selective Conformal Optimized Pairwise LLM Judging

SafetyDGX agent

arXiv:2602.13110v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used as scalable judges in pairwise evaluation, but they remain prone to miscalibration and bias

Score Broadcast and Decorrelation: A General Framework for Broadcast-Based Credit Assignment

ResearchDGX agent

arXiv:2605.30638v1 Announce Type: cross Abstract: We introduce Score Broadcast and Decorrelation (SBD), a principled framework for broadcast-based credit assignment for general families of differentia

Seeing Before Agreeing: Aligning Multi-Agent Consensus with Visual Evidence

SafetyDGX agent

arXiv:2605.30698v1 Announce Type: cross Abstract: Vision-language models (VLMs) have achieved strong performance on visual question answering (VQA). To mitigate individual hallucinations and blind spo

Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?

TutorialsDGX agent

arXiv:2605.30557v1 Announce Type: cross Abstract: Spatial reasoning is a fundamental capability for vision-language models (VLMs) deployed in real-world environments. However, visual observations are

Separating Secrets from Placeholders: A Hybrid CNN-CodeBERT Framework for Three-Class Credential Leakage Detection

ResearchDGX agent

arXiv:2605.31520v1 Announce Type: cross Abstract: Credential leakage in public source code repositories poses a critical security threat, with over 23.8 million secrets exposed in 2024 alone. Existing

Shared Doubt: Zero-shot Cross-Lingual Confidence Estimation for Language Models

ResearchDGX agent

arXiv:2605.31220v1 Announce Type: cross Abstract: Confidence estimation (CE), i.e. quantifying the reliability of a model's prediction, has attracted great interest in the context of large language mo

SHIELD: Secure Hypernetworks for Incremental Expansion Learning Defense

ResearchDGX agent

arXiv:2506.08255v4 Announce Type: replace-cross Abstract: Continual learning under adversarial conditions remains an open problem, as existing methods often compromise either robustness, scalability,

Simple Token-Efficient Vision-Language Model for Case-level Pathology Synoptic Report Generation

HardwareDGX agent

arXiv:2605.30716v1 Announce Type: cross Abstract: Generating clinically useful pathology reports for pathology cases from whole-slide images (WSIs) is challenging due to gigapixel resolution, long vis

Simulation of collision avoidance behavior in crowd movement by data-driven approach

SafetyDGX agent

arXiv:2605.31210v1 Announce Type: cross Abstract: Crowd movement simulation is essential for pedestrian safety management and facility layout optimization. Data-driven models enhance trajectory predic

SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs

Model ReleasesDGX agent

arXiv:2603.20253v2 Announce Type: replace-cross Abstract: Evaluating LLM agents for scientific tasks has focused on token costs while ignoring tool-use costs like simulation time and experimental reso

SKETCH: Semantic Key-Point Conditioning for Long-Horizon Vessel Trajectory Prediction

ApplicationsDGX agent

arXiv:2601.18537v3 Announce Type: replace-cross Abstract: Accurate long-horizon vessel trajectory prediction remains challenging due to compounded uncertainty from complex navigation behaviors and env

Skill Availability and Presentation Granularity in Large-Language-Model Agents: A Controlled SkillsBench Study

Model ReleasesDGX agent

arXiv:2605.31408v1 Announce Type: cross Abstract: Skill documents provide procedural knowledge to large-language-model agents at inference time. This article studies whether the presentation granulari

Skill Reuse as Compression in Agentic RL

AgentsDGX agent

arXiv:2605.31509v1 Announce Type: cross Abstract: Large language model agents trained with reinforcement learning (RL) often learn brittle, task-specific shortcuts. We hypothesize that agents generali

SLAT: Segment-Level Adaptive Trimming for Efficient CoT Reasoning

ResearchDGX agent

arXiv:2605.30832v1 Announce Type: new Abstract: Recent advances in Large Reasoning Models have significantly improved chain-of-thought (CoT) capabilities via reinforcement learning (RL). However, gene

Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO

SafetyDGX agent

arXiv:2605.30789v1 Announce Type: cross Abstract: We identify a new dimension for enhancing rollout diversity in Group Relative Policy Optimization (GRPO) for LLMs. While GRPO relies on diverse rollou

Social Reasoning in Machines: Investigating Collective Truth-Seeking Dynamics in Large Language Model Debate

AgentsDGX agent

arXiv:2605.30391v1 Announce Type: cross Abstract: Human reasoning has long been theorised to operate socially, not through isolated individual cognition, but through collective adversarial discourse,

Social welfare optimisation under institutional reward and punishment

Model ReleasesDGX agent

arXiv:2605.31330v1 Announce Type: cross Abstract: Institutional incentives are widely used to promote cooperation among autonomous, self-regarding agents, from human societies to multi-agent and AI sy

Sophrosyne: Agentic Exploration of Relational Data Systems Needs Moderation

AgentsDGX agent

arXiv:2605.30862v1 Announce Type: cross Abstract: Text2SQL agents powered by LLMs translate natural language intent into SQL by exploring the data system through tool calls before formulating the quer

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes

Model ReleasesDGX agent

arXiv:2605.31148v1 Announce Type: cross Abstract: Humans can effortlessly perceive spatial layouts, form cognitive representations, reason about spatial relations, and translate such reasoning into ac

SpecDB: LLM-Generated Customized Databases via Feature-Oriented Decomposition

AgentsDGX agent

arXiv:2605.31097v1 Announce Type: cross Abstract: Mainstream relational databases ship a uniform feature set across deployments, although individual workloads exercise only a fraction of the available

SPECTRA: Synthetic IR Test Collections with Relevance Oracles and Controlled Distractor Diagnostics

ResearchDGX agent

arXiv:2605.31575v1 Announce Type: cross Abstract: Scalable information retrieval testing needs corpora that are large enough to stress index construction, ranking latency, query routing, and evaluatio

Spectral Collapse Drives Loss of Plasticity in Deep Continual Learning

TutorialsDGX agent

arXiv:2509.22335v3 Announce Type: replace-cross Abstract: We investigate why deep neural networks suffer from loss of plasticity in continual learning, and thus fail to learn new tasks without reiniti

SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy

Model ReleasesDGX agent

arXiv:2602.22971v2 Announce Type: replace Abstract: As LLMs achieved breakthroughs in general reasoning, their proficiency in specialized scientific domains reveals pronounced gaps in existing benchma

Stateful Online Monitoring Catches Distributed Agent Attacks

SafetyDGX agent

arXiv:2605.31593v1 Announce Type: cross Abstract: Language models can find thousands of severe software vulnerabilities, and agents are increasingly being misused for cyberattacks. To avoid detection,

Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines

Model ReleasesDGX agent

arXiv:2605.31183v1 Announce Type: cross Abstract: Sparse Autoencoders (SAEs) have been seen as a promising avenue for exploring the internals of Large Language Models (LLMs) and for steering model out

STEP: Learning STructured Embeddings for Progressive Time Series

TutorialsDGX agent

arXiv:2605.31061v1 Announce Type: cross Abstract: We present a novel method for learning interpretable representations of progressive time series, that is, data capturing irreversible state transition

Stop the Flip-Flop: Context-Preserving Verification for Fast Revocable Diffusion Decoding

ResearchDGX agent

arXiv:2602.06161v2 Announce Type: replace-cross Abstract: Parallel diffusion decoding can accelerate diffusion language model inference by unmasking multiple tokens per step, but aggressive parallelis

Structure-Induced Information for Rerooting Levin Tree Search

SafetyDGX agent

arXiv:2605.30664v1 Announce Type: new Abstract: Subgoal-based policy tree search, which uses a policy to guide search, is effective for complex single-agent deterministic problems but often relies on

Structured interactions improve distributed coordination beyond model scaling in a real-world multi-robot system

Model ReleasesDGX agent

arXiv:2605.30383v1 Announce Type: cross Abstract: Scaling individual robot capabilities is common but costly. Here we investigate a system-level design question in real-world multi-robot coordination:

Surprised by Attention: Predictable Query Dynamics for Time Series Anomaly Detection

AgentsDGX agent

arXiv:2603.12916v3 Announce Type: replace-cross Abstract: Multivariate time series anomalies often manifest as shifts in cross-channel dependencies rather than simple amplitude excursions. In autonomo

SWIM: Single-Instance Whole-Body Imitation for swiMming

TutorialsDGX agent

arXiv:2605.31120v1 Announce Type: cross Abstract: We propose a new method for synthesizing physically-based swimming motions. Physically-based character animation aims to generate physically valid, co

Symbolic Intermediaries as a Linguistic-Numerical Interface for LLM-Driven Geometric Reasoning

Model ReleasesDGX agent

arXiv:2505.17607v3 Announce Type: replace Abstract: Large Language Models (LLMs) display reasoning capabilities over linguistic and symbolic objects but have limited capabilities to directly interpret

Target-Agnostic Calibration under Distribution Shift with Frequency-Aware Gradient Rectification

Model ReleasesDGX agent

arXiv:2508.19830v2 Announce Type: replace-cross Abstract: Real-world model deployments inevitably encounter distribution shifts, rendering the confidence estimates of deep neural networks highly unrel

Target-Side Paraphrase Augmentation for Sign Language Translation with Large Language Models

ResearchDGX agent

arXiv:2605.31393v1 Announce Type: cross Abstract: Sign language translation (SLT) remains constrained by limited paired sign-video/text corpora and heavy-tailed target vocabularies. We study target-si

Targeted Speaker Poisoning Framework in Zero-Shot Text-to-Speech

Model ReleasesDGX agent

arXiv:2603.07551v2 Announce Type: replace-cross Abstract: Zero-shot Text-to-Speech (TTS) voice cloning poses severe privacy risks, demanding the removal of specific speaker identities from trained TTS

TARIC: Memory-Augmented Traversability-Aware Outdoor VLN under Interrupted Semantic Cues

SafetyDGX agent

arXiv:2605.31121v1 Announce Type: cross Abstract: Outdoor vision-language navigation (VLN) in long-range, open-world environments is frequently disrupted by semantic-cue interruptions, where informati

← Previous
1…186187188189190…358
Next →