AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
25 May 2026

PilotWiMAE: Pilot-Native Representation Learning for Wireless Channels

SafetyDGX agent

arXiv:2605.22856v1 Announce Type: cross Abstract: Channel foundation models assume access to fully observed channels, an assumption that fails in deployment. We introduce PilotWiMAE, a self-supervised

PLACE: Prompt Learning for Attributed Community Search in Large Graphs

ApplicationsDGX agent

arXiv:2507.05311v2 Announce Type: replace-cross Abstract: In this paper, we propose PLACE (Prompt Learning for Attributed Community Search), an innovative graph prompt learning framework for ACS. Enli

PoisonForge: Task-Level Targeted Poisoning Benchmark for Instruction-Tuned LLMs

Model ReleasesDGX agent

arXiv:2605.23168v1 Announce Type: cross Abstract: When practitioners fine-tune LLMs on unvetted datasets, an adversary can exploit the data supply chain through task-level poisoning: inserting a small


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks

Model ReleasesDGX agent

arXiv:2605.23170v1 Announce Type: cross Abstract: Position-controlled evaluation is standard for retrieval tasks such as Needle-in-a-Haystack and RULER, but mainstream reasoning benchmarks do not cont

Precise: SDE-Consistent Stochastic Sampling for RL Post-Training of Flow-Matching Models

SafetyDGX agent

arXiv:2605.23522v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become an effective way to improve prompt alignment and perceptual quality in diffusion and flow-matching generators.

PrefBench: Evaluating Zero-Shot LLM Agents in Hidden-Preference Personalized Pricing Negotiations

Model ReleasesDGX agent

arXiv:2605.22855v1 Announce Type: cross Abstract: Personalized pricing negotiations are a challenging testbed for LLM agents because successful interaction does not guarantee profitable decision makin

Preisach Attention: A Hysteretic Model of Sequential Memory

Local AiDGX agent

arXiv:2605.23603v1 Announce Type: cross Abstract: We introduce the Preisach Attention Layer (PAL), a novel sequence modelling architecture grounded in the classical Preisach hysteresis operator from m

ProtDBench: A Unified Benchmark of Protein Binder Design and Evaluation

Model ReleasesDGX agent

arXiv:2605.04118v2 Announce Type: replace-cross Abstract: Recent advances in de novo protein binder design have enabled increasing experimental validation, yet reported in silico metrics remain diffic

R^3L: Reflect-then-Retry Reinforcement Learning with Language-Guided Exploration, Pivotal Credit, and Positive Amplification

Model ReleasesDGX agent

arXiv:2601.03715v2 Announce Type: replace-cross Abstract: Reinforcement learning drives recent advances in LLM reasoning and agentic capabilities, yet current approaches struggle with both exploration

RA-DCA: A Randomized Active-Set DCA for Directional Stationarity in Max-Structured DC Programs

ResearchDGX agent

arXiv:2605.23550v1 Announce Type: cross Abstract: We study nonsmooth difference-of-convex programs whose subtracted convex term is a finite maximum of smooth convex functions. In this setting, standar

RAG-Pull: Turning Retrieval into a Code-Injection Channel via Invisible Unicode Perturbations

SafetyDGX agent

arXiv:2510.11195v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) increases the reliability and trustworthiness of the LLM response and reduces hallucination by eliminatin

RAG4Outcome: A Retrieval-Augmented Multimodal Framework for Prognostic Prediction in Chronic Osteomyelitis

SafetyDGX agent

arXiv:2605.22833v1 Announce Type: cross Abstract: Chronic osteomyelitis presents substantial prognostic challenges due to its high recurrence risk and complex postoperative recovery trajectories. Trad

ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload

SafetyDGX agent

arXiv:2605.11215v2 Announce Type: replace-cross Abstract: Pre-training large language models on massive GPU clusters has made hardware faults routine rather than rare, driving the need for resilient t

Redrawing the AI Map: A Theory of Accountability Boundaries in Agentic Ecosystems

AgentsDGX agent

arXiv:2605.23179v1 Announce Type: new Abstract: Agentic AI orchestrators reduce the interface and assembly costs of composing information systems capabilities across organizational boundaries, seeming

Reflex: Reinforcement Learning with Reflection Symmetry Exploitation in State-Based Continuous Control

SafetyDGX agent

arXiv:2605.23415v1 Announce Type: cross Abstract: Reinforcement learning has long struggled with poor sample efficiency. One promising approach to mitigate this problem is leveraging group-invariant M

Reinforcement Learning for Microcanonical Graph Ensemble with Assortativity Constraints

Model ReleasesDGX agent

arXiv:2605.23285v1 Announce Type: cross Abstract: How network structure determines function is a fundamental question, and it can be investigated by graph ensembles with precisely controlled structura

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers

SafetyDGX agent

arXiv:2510.00915v4 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) replaces costly human labeling with automated verifiers. To reduce verifier hacking, man

RMA: an Agentic System for Research-Level Mathematical Problems

Model ReleasesDGX agent

arXiv:2605.22875v1 Announce Type: new Abstract: We present extbf{Research Math Agents (RMA)}, an agentic framework for automated reasoning on research-level mathematical problems. Unlike prior studies

Robots That Know What to Ask: Recovering Misaligned Rewards through Targeted Explanations

ResearchDGX agent

arXiv:2605.22986v1 Announce Type: cross Abstract: Learning reward functions from demonstrations assumes that demonstrations provide adequate supervision over all features -- or task-relevant aspects o

Robust Counterfactual Inference in Markov Decision Processes

ResearchDGX agent

arXiv:2502.13731v5 Announce Type: replace Abstract: This paper addresses a key limitation in existing counterfactual inference methods for Markov Decision Processes (MDPs). Current approaches assume a

Safe Reinforcement Learning with Preference-based Constraint Inference

SafetyDGX agent

arXiv:2603.23565v2 Announce Type: replace-cross Abstract: Safe reinforcement learning (RL) is a standard paradigm for safety-critical decision making. However, real-world safety constraints can be com

SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety

SafetyDGX agent

arXiv:2605.05704v2 Announce Type: replace-cross Abstract: Recent advances in foundation models have transformed LLMs from passive conversational systems into autonomous agents capable of reasoning and

Scalable Heterogeneous Graph Foundation Models for Data-Driven Optimal Power Flow in Smart Grids

ResearchDGX agent

arXiv:2605.23194v1 Announce Type: cross Abstract: Fast and reliable optimal power flow (OPF) approximation is essential for reliable smart-grid operation, yet many learning-based surrogates either fla

Scaling-Aware Adapter for Structure-Grounded LLM Reasoning

ResearchDGX agent

arXiv:2602.02780v3 Announce Type: replace Abstract: Large language models (LLMs) are enabling reasoning over 2D and 3D structures, yet existing methods remain modality-specific and typically compress

SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research

Model ReleasesDGX agent

arXiv:2605.22878v1 Announce Type: new Abstract: The exponential growth of global academic output has confronted researchers and AI agents with an unprecedented ``information explosion,'' where fragmen

SciHorizon-GENE: Benchmarking LLM for Life Sciences Inference from Gene Knowledge to Functional Understanding

Model ReleasesDGX agent

arXiv:2601.12805v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown growing promise in biomedical research, particularly for knowledge-driven interpretation tasks. Howeve

Score-Based One-step MeanFlow Policy Optimization

SafetyDGX agent

arXiv:2605.23365v1 Announce Type: cross Abstract: Diffusion and flow matching have emerged as expressive policy classes in reinforcement learning, but their reliance on multi-step denoising imposes su

Security of LLM-generated Code: A Comparative Analysis

ApplicationsDGX agent

arXiv:2605.23091v1 Announce Type: cross Abstract: The majority of software developers use or are planning to use Artificial Intelligence (AI) tools in their development processes. Their top reasons in

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

Model ReleasesDGX agent

arXiv:2605.22903v1 Announce Type: cross Abstract: Benchmark accuracy is often implicitly assumed to reflect grounded visual understanding in vision-language models (VLMs), yet it remains unclear to wh

SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion

ResearchDGX agent

arXiv:2605.23245v1 Announce Type: cross Abstract: Video object insertion requires ensuring spatio-temporal coherence and interactive realism, extending far beyond simple content placement. However, cu

SkillOpt: Executive Strategy for Self-Evolving Agent Skills

Model ReleasesDGX agent

arXiv:2605.23904v1 Announce Type: new Abstract: Agent skills today are hand-crafted, generated one-shot, or evolved through loosely controlled self-revision, none of which behaves like a deep-learning

Socially fluent AI decouples conversational signals from source identity in online interaction

AgentsDGX agent

arXiv:2605.23426v1 Announce Type: cross Abstract: Socially fluent agentic AI can now participate in online interaction in ways that resemble ordinary human conversation, potentially weakening people's

Solving the Aircraft Disassembly Scheduling Problem

ResearchDGX agent

arXiv:2605.23592v1 Announce Type: new Abstract: Dismantling aircrafts reaching their end of life is a complex endeavour that is necessary in terms of sustainability but yields small income margins for

SPACENUM: Revisiting Spatial Numerical Understanding in VLMs

ResearchDGX agent

arXiv:2605.23898v1 Announce Type: new Abstract: Vision-Language Models (VLMs) are increasingly deployed in embodied environments, where they need produce numerical outputs such as action magnitudes an

Sparse Autoencoders Map Brain-LLM Alignment onto Cortical Semantic Topography

Model ReleasesDGX agent

arXiv:2605.23035v1 Announce Type: cross Abstract: Intermediate layers of large language models (LLMs) best predict human brain responses to language, one of the most robust findings in computational n

Sparse Compositional Flow Matching by geometric assembly from motion primitives

ResearchDGX agent

arXiv:2605.23341v1 Announce Type: cross Abstract: Embodied trajectories, such as the executable motion sequences of robotic manipulators, underwater vehicles, and mobile robots, are a fundamental outp

Sparser Block-Sparse Attention via Token Permutation

ApplicationsDGX agent

arXiv:2510.21270v2 Announce Type: replace-cross Abstract: Scaling the context length of large language models (LLMs) offers significant benefits but is computationally expensive. This expense stems pr

Spectral-inspired Operator Learning with Limited Data and Unknown Physics

Local AiDGX agent

arXiv:2505.21573v3 Announce Type: replace-cross Abstract: Learning PDE dynamics from limited data with unknown physics is challenging. Existing neural PDE solvers either require large datasets or rely

SSDAU: Structured Semantic Data Augmentation for Joint Entity and Relation Extraction

ResearchDGX agent

arXiv:2605.23440v1 Announce Type: cross Abstract: Joint Entity and Relation Extraction (JERE) is highly susceptible to weak generalization due to low-quality training data. Data augmentation is a comm

Staging by the Book: Automatic Sleep Stage Classification Using Scoring Rules

ResearchDGX agent

arXiv:2605.22859v1 Announce Type: cross Abstract: Automated sleep staging is commonly approached as a supervised machine learning problem, with deep learning methods dominating recent research. While

Strategic Coercion Within Alliances: The Greenland Sovereignty Game as an AI Stress Test

Model ReleasesDGX agent

arXiv:2605.22841v1 Announce Type: cross Abstract: What happens when the strongest alliance member pressures a weaker member over territory and strategic control? We examine the Greenland sovereignty c

Suicide Risk Assessment from AI-powered Video Surveillance: An Interpretable Framework for Prevention in Metro Stations

ResearchDGX agent

arXiv:2605.22904v1 Announce Type: cross Abstract: Understanding and monitoring human behavior in metro stations play an important role in supporting suicide prevention efforts, where early identificat

Tabular PDF Information Extraction with Local LLMs and Layout-Aware Parsing: A Reliability Evaluation

Model ReleasesDGX agent

arXiv:2604.00003v2 Announce Type: replace-cross Abstract: Extracting structured information from academic PDF documents is non trivial: a single page typically combines free text metadata with tabular

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning

AgentsDGX agent

arXiv:2602.01665v2 Announce Type: replace-cross Abstract: The design of environments plays a critical role in shaping the development and evaluation of cooperative multi-agent reinforcement learning (

TCAP: Tri-Component Attention Profiling for Unsupervised Backdoor Detection in MLLM Fine-Tuning

ResearchDGX agent

arXiv:2601.21692v2 Announce Type: replace Abstract: Fine-Tuning-as-a-Service (FTaaS) facilitates the customization of Multimodal Large Language Models (MLLMs) but introduces critical backdoor risks vi

Tensor Cache: Eviction-conditioned Associative Memory for Transformers

ResearchDGX agent

arXiv:2605.22884v1 Announce Type: cross Abstract: Autoregressive Transformer KV caches grow linearly with context length; sliding-window caching bounds memory but discards evicted tokens entirely, so

Test-Time Training Undermines Safety Guardrails

SafetyDGX agent

arXiv:2605.22984v1 Announce Type: cross Abstract: Test-Time Training (TTT) is an emerging paradigm that enables models to adapt their parameters during inference, improving performance on tasks such a

The Cognitive Kardashev Scale: Quantifying the Material Envelope of Civilisational Computation

HardwareDGX agent

arXiv:2605.22840v1 Announce Type: cross Abstract: How much thinking can a civilisation do? Kardashev's (1964) typology ranks civilisations by total power: planetary (Type I, ~10^16 W), stellar (Type I

The Deterministic Horizon: Impossibility Results as Design Specifications for Trustworthy AI Systems

ApplicationsDGX agent

arXiv:2605.23024v1 Announce Type: new Abstract: Large language models now write software, draft legal documents, and produce clinical notes, yet fundamental limits, from Turing and Arrow to the No Fre

The Misattribution Gap: When Memory Poisoning Looks Like Model Failure in Agentic AI Systems

Model ReleasesDGX agent

arXiv:2605.22842v1 Announce Type: cross Abstract: Multi-agent AI pipelines typically assume that agent misconduct originates from model misalignment. We identify a structural failure in this assumptio

The Readout Shortcut: Positional Number Copying Dominates Arithmetic CoT Readout in Small Language Models

Model ReleasesDGX agent

arXiv:2605.22870v1 Announce Type: cross Abstract: Chain-of-thought (CoT) prompting is necessary for arithmetic in small language models, yet shuffling its steps preserves most performance. What does C

The Surprising Difficulty of Search in Model-Based Reinforcement Learning

Model ReleasesDGX agent

arXiv:2601.21306v2 Announce Type: replace-cross Abstract: This paper investigates search in model-based reinforcement learning (RL). Conventional wisdom holds that long-term predictions and compoundin

The TIME Machine: On The Power of Motion for Efficient Perception

TutorialsDGX agent

arXiv:2605.23045v1 Announce Type: cross Abstract: Video representation learning has seen tremendous progress in recent years. This has been driven by many factors, including the scale of training and

Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in RAG

SafetyDGX agent

arXiv:2506.04390v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) systems are vulnerable to attacks that inject poisoned passages into the retrieved context, even at low c

Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.22902v1 Announce Type: cross Abstract: Generative Vision-Language Models (VLMs) perform well on multimodal reasoning, but how visual inputs are transformed to text remains poorly understood

Uncovering the Latent Potential of Deep Intermediate Representations

TutorialsDGX agent

arXiv:2605.23033v1 Announce Type: cross Abstract: Foundational Models pretrained on huge amount of data learn representations that evolve across depth, forming a hierarchy of embeddings with distinct

Understanding and Improving Noisy Embedding Techniques in Instruction Finetuning

Model ReleasesDGX agent

arXiv:2605.23171v1 Announce Type: cross Abstract: Recent advancements in instructional fine-tuning have injected noise into embeddings, with NEFTune (Jain et al., 2024) setting benchmarks using unifor

Understanding Goal Generalisation in Sequential Reinforcement Learning

SafetyDGX agent

arXiv:2605.23565v1 Announce Type: cross Abstract: Reinforcement learning agents often exhibit unintended goal-directed behaviour outside their training distribution, but we currently lack a principled

Understanding Task Aggregation for Generalizable Ultrasound Foundation Models

ResearchDGX agent

arXiv:2603.18123v3 Announce Type: replace-cross Abstract: Foundation models promise to unify multiple clinical tasks within a single framework, but recent ultrasound studies report that unified models

V-VLAPS: Value-Guided Planning for Vision-Language-Action Models

SafetyDGX agent

arXiv:2601.00969v2 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models provide strong action priors for robotic manipulation, but their reactive behavior can fail under distribu

← Previous
1…223224225226227…358
Next →