AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
15 May 2026

Fully Dynamic Rebalancing in Dockless Bike-Sharing Systems via Deep Reinforcement Learning

AgentsDGX agent

arXiv:2605.14501v1 Announce Type: cross Abstract: This paper proposes a fully dynamic Deep Reinforcement Learning (DRL) method for rebalancing dockless bike-sharing systems, overcoming the limitations

Fusion-fission forecasts when AI will shift to undesirable behavior

Model ReleasesDGX agent

arXiv:2605.14218v1 Announce Type: new Abstract: The key problem facing ChatGPT-like AI's use across society is that its behavior can shift, unnoticed, from desirable to undesirable -- encouraging self

FutureSim: Replaying World Events to Evaluate Adaptive Agents

Model ReleasesDGX agent

arXiv:2605.15188v1 Announce Type: cross Abstract: AI agents are being increasingly deployed in dynamic, open-ended environments that require adapting to new information as it arrives. To efficiently m


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

GEAR: Genetic AutoResearch for Agentic Code Evolution

AgentsDGX agent

arXiv:2605.13874v1 Announce Type: cross Abstract: Autonomous research agents can already run machine learning experiments without human supervision, but many rely on a narrow search strategy: they rep

GenCircuit-RL: Reinforcement Learning from Hierarchical Verification for Genetic Circuit Design

Model ReleasesDGX agent

arXiv:2605.14215v1 Announce Type: new Abstract: Genetic circuit design remains a laborious, expert-driven process despite decades of progress in synthetic biology. We study this problem through code g

Generalized Priority-Aware Shapley Value

ResearchDGX agent

arXiv:2605.15018v1 Announce Type: cross Abstract: Shapley value and its priority-aware extensions are widely used for valuation in machine learning, but existing methods require pairwise priority to b

Generative Floor Plan Design with LLMs via Reinforcement Learning with Verifiable Rewards

ResearchDGX agent

arXiv:2605.14117v1 Announce Type: cross Abstract: An AI system for professional floor plan design must precisely control room dimensions and areas while respecting the desired connectivity between roo

GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models

Model ReleasesDGX agent

arXiv:2602.06718v2 Announce Type: replace-cross Abstract: Citations provide the basis for trusting scientific claims; when they are invalid or fabricated, this trust collapses. With the advent of Larg

Good to Go: The LOOP Skill Engine That Hits 99% Success and Slashes Token Usage by 99% via One-Shot Recording and Deterministic Replay

Model ReleasesDGX agent

arXiv:2605.14237v1 Announce Type: new Abstract: Deploying AI agents for repetitive periodic tasks exposes a critical tension: Large Language Models (LLMs) offer unmatched flexibility in tool orchestra

GPart: End-to-End Isometric Fine-Tuning via Global Parameter Partitioning

Model ReleasesDGX agent

arXiv:2605.14841v1 Announce Type: cross Abstract: Low-rank adaptation (LoRA) has become the dominant paradigm for parameter-efficient fine-tuning (PEFT) of large language models (LLMs). However, its b

Gradient Iterated Temporal-Difference Learning

AgentsDGX agent

arXiv:2603.07833v2 Announce Type: replace-cross Abstract: Temporal-difference (TD) learning is highly effective at controlling and evaluating an agent's long-term outcomes. Most approaches in this par

Graph of States: Solving Abductive Tasks with Large Language Models

AgentsDGX agent

arXiv:2603.21250v2 Announce Type: replace Abstract: Logical reasoning encompasses deduction, induction, and abduction. However, while Large Language Models (LLMs) have effectively mastered the former

GraphBit: A Graph-based Agentic Framework for Non-Linear Agent Orchestration

Model ReleasesDGX agent

arXiv:2605.13848v1 Announce Type: new Abstract: Agentic LLM frameworks that rely on prompted orchestration, where the model itself determines workflow transitions, often suffer from hallucinated routi

GraphFlow: An Architecture for Formally Verifiable Visual Workflows Enabling Reliable Agentic AI Automation

Local AiDGX agent

arXiv:2605.14968v1 Announce Type: new Abstract: GraphFlow is a visual workflow system designed to improve the reliability of agentic AI automation in multi-step, mission-critical processes. In these w

Graphs of Research: Citation Evolution Graphs as Supervision for Research Idea Generation

ResearchDGX agent

arXiv:2605.14790v1 Announce Type: cross Abstract: Research idea generation is the innovation-driving step of automated scientific research. Recently, large language models (LLMs) have shown potential

Grokking Finite-Dimensional Algebra

SafetyDGX agent

arXiv:2602.19533v2 Announce Type: replace-cross Abstract: This paper investigates the grokking phenomenon, which refers to the sudden transition from a long memorization to generalization observed dur

Grounded Continuation: A Linear-Time Runtime Verifier for LLM Conversations

AgentsDGX agent

arXiv:2605.14175v1 Announce Type: new Abstract: In long conversations, an LLM can produce a next utterance that sounds plausible but rests on premises the conversation has already abandoned. Context-m

HASTE: Training-Free Video Diffusion Acceleration via Head-Wise Adaptive Sparse Attention

ResearchDGX agent

arXiv:2605.14513v1 Announce Type: cross Abstract: Diffusion-based video generation has advanced substantially in visual fidelity and temporal coherence, but practical deployment remains limited by the

Head Forcing: Long Autoregressive Video Generation via Head Heterogeneity

ResearchDGX agent

arXiv:2605.14487v1 Announce Type: cross Abstract: Autoregressive video diffusion models support real-time synthesis but suffer from error accumulation and context loss over long horizons. We discover

Herculean: An Agentic Benchmark for Financial Intelligence

Model ReleasesDGX agent

arXiv:2605.14355v1 Announce Type: new Abstract: As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carr

Heuristic Pathologies and Further Variance Reduction via Uncertainty Propagation in the AIVAT Family of Techniques

AgentsDGX agent

arXiv:2605.14261v1 Announce Type: new Abstract: How should an agent's performance in a multiagent environment be evaluated when there is a limited sample size or a high cost of running a trial? The AI

Hidden State Poisoning Attacks against Mamba-based Language Models

Model ReleasesDGX agent

arXiv:2601.01972v4 Announce Type: cross Abstract: State space models (SSMs) like Mamba offer efficient alternatives to Transformer-based language models, with linear time complexity. Yet, their advers

HodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-Experts

ResearchDGX agent

arXiv:2605.13997v1 Announce Type: cross Abstract: Sparse Mixture-of-Experts (MoE) layers route tokens through a handful of experts, and learning-free compression of these layers reduces inference cost

Holistic Evaluation and Failure Diagnosis of AI Agents

Model ReleasesDGX agent

arXiv:2605.14865v1 Announce Type: new Abstract: AI agents execute complex multi-step processes, but current evaluation falls short: outcome metrics report success or failure without explaining why, an

How Sensitive Are Radiomic AI Models to Acquisition Parameters?

Model ReleasesDGX agent

arXiv:2605.14667v1 Announce Type: new Abstract: A main barrier for the deployment of AI radiomic systems in clinical routine is their drop in performance under heterogeneous multicentre acquisition pr

How to Evaluate and Refine your CAM

TutorialsDGX agent

arXiv:2605.14641v1 Announce Type: cross Abstract: Class attribution maps (CAMs) provide local explanations for the decisions of convolutional neural networks. While widely used in practice, the evalua

Hypergraph Enterprise Agentic Reasoner over Heterogeneous Business Systems

AgentsDGX agent

arXiv:2605.14259v1 Announce Type: new Abstract: Applying Large Language Models (LLMs) to heterogeneous enterprise systems is hindered by hallucinations and failures in multi-hop, n-ary reasoning. Exis

ICED: Concept-level Machine Unlearning via Interpretable Concept Decomposition

SafetyDGX agent

arXiv:2605.14309v1 Announce Type: cross Abstract: Machine unlearning in Vision-Language Models (VLMs) is typically performed at the image or instance level, making it difficult to precisely remove tar

Identifying Culprits Through Deep Deterministic Policy Gradient Deep Learning Investigation

SafetyDGX agent

arXiv:2605.14774v1 Announce Type: new Abstract: In the world of AI and advanced technologies investigation aspects identification of a crime or criminal plays a major problem. In this research we focu

IFPV: An Integrated Multi-Agent Framework for Generative Operational Planning and High-Fidelity Plan Verification

AgentsDGX agent

arXiv:2605.14851v1 Announce Type: cross Abstract: Operational plan generation and verification are critical for modern complex and rapidly changing battlefield environments, yet traditional generation

Image Restoration via Diffusion Models with Dynamic Resolution

ResearchDGX agent

arXiv:2605.14267v1 Announce Type: cross Abstract: Diffusion models (DMs) have exhibited remarkable efficacy in various image restoration tasks. However, existing approaches typically operate within th

Improving Multi-turn Dialogue Consistency with Self-Recall Thinking

ResearchDGX agent

arXiv:2605.15102v1 Announce Type: cross Abstract: Large language model (LLM) based multi-turn dialogue systems often struggle to track dependencies across non-adjacent turns, undermining both consiste

In-IDE Toolkit for Developers of AI-Based Features

AgentsDGX agent

arXiv:2605.14612v1 Announce Type: cross Abstract: AI-enabled features built on LLMs and agentic workflows are difficult to test, debug, and reproduce, especially for product-focused software engineers

Intelligence Impact Quotient (IIQ): A Framework for Measuring Organizational AI Impact

AgentsDGX agent

arXiv:2605.14455v1 Announce Type: new Abstract: The Intelligence Impact Quotient (IIQ) is a composite metric intended to quantify the depth to which AI systems are integrated into organizational work

IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation

Model ReleasesDGX agent

arXiv:2605.14712v1 Announce Type: cross Abstract: Robot imitation data are often multimodal: similar visual-language observations may be followed by different action chunks because human demonstrators

Interestingness as an Inductive Heuristic for Future Compression Progress

ResearchDGX agent

arXiv:2605.14831v1 Announce Type: new Abstract: One of the bottlenecks on the way towards recursively self-improving systems is the challenge of interestingness: the ability to prospectively identify

Invisible Orchestrators Suppress Protective Behavior and Dissociate Power-Holders: Safety Risks in Multi-Agent LLM Systems

Model ReleasesDGX agent

arXiv:2605.13851v1 Announce Type: new Abstract: Multi-agent orchestration -- in which a hidden coordinator manages specialized worker agents -- is becoming the default architecture for enterprise AI d

IPR-1: Interactive Physical Reasoner

Model ReleasesDGX agent

arXiv:2511.15407v3 Announce Type: replace Abstract: Humans learn by observing, interacting with environments, and internalizing physics and causality. Here, we aim to ask whether an agent can similarl

KGPFN: Unlocking the Potential of Knowledge Graph Foundation Model via In-Context Learning

Local AiDGX agent

arXiv:2605.14907v1 Announce Type: new Abstract: Knowledge graph (KG) foundation models aim to generalize across graphs with unseen entities and relations by learning transferable relational structure.

Know When To Fold 'Em: Token-Efficient LLM Synthetic Data Generation via Multi-Stage In-Flight Rejection

SafetyDGX agent

arXiv:2605.14062v1 Announce Type: new Abstract: While synthetic data generation with large language models (LLMs) is widely used in post-training pipelines, existing approaches typically generate full

Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces

AgentsDGX agent

arXiv:2605.14786v1 Announce Type: cross Abstract: As LLM-based agents increasingly browse the web on users' behalf, a natural question arises: can websites passively identify which underlying model po

Krause Synchronization Transformers

Model ReleasesDGX agent

arXiv:2602.11534v3 Announce Type: replace-cross Abstract: Self-attention in Transformers relies on globally normalized softmax weights, causing all tokens to compete for influence at every layer. When

L2R: Low-Rank and Lipschitz-Controlled Routing for Mixture-of-Experts

Model ReleasesDGX agent

arXiv:2601.21349v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) models scale neural networks by conditionally activating a small subset of experts, where the router plays a central

Large Language Models for Web Accessibility: A Systematic Literature Review

TutorialsDGX agent

arXiv:2605.13873v1 Announce Type: cross Abstract: Web accessibility aims to ensure that web content and services are usable by people with diverse abilities. In recent years, Large Language Models (LL

Learning Developmental Scaffoldings to Guide Self-Organisation

Local AiDGX agent

arXiv:2605.14998v1 Announce Type: new Abstract: From subcellular structures to entire organisms, many natural systems generate complex organisation through self-organisation: local interactions that c

Learning Scenario Reduction for Two-Stage Robust Optimization with Discrete Uncertainty

ResearchDGX agent

arXiv:2605.14494v1 Announce Type: new Abstract: Two-Stage Robust Optimization (2RO) with discrete uncertainty is challenging, often rendering exact solutions prohibitive. Scenario reduction alleviates

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis

SafetyDGX agent

arXiv:2605.14392v1 Announce Type: new Abstract: We pursue a vision for self-improving language models in which the model does not merely generate problems or traces to imitate, but constructs the envi

LEMON: Learning Executable Multi-Agent Orchestration via Counterfactual Reinforcement Learning

Local AiDGX agent

arXiv:2605.14483v1 Announce Type: new Abstract: Large language models (LLMs) have become a strong foundation for multi-agent systems, but their effectiveness depends heavily on orchestration design. A

LLM-Based Robustness Testing of Microservice Applications: An Empirical Study

ResearchDGX agent

arXiv:2605.14202v1 Announce Type: cross Abstract: Malformed, missing, or boundary-value inputs in microservice APIs can cascade across dependent services, threatening reliability. Robustness testing s

LLMs learn scientific taste from institutional traces across the social sciences

TutorialsDGX agent

arXiv:2603.16659v3 Announce Type: replace Abstract: Reinforcement-learned reasoning has powered recent AI leaps on verifiable tasks, including mathematics, code, and structure prediction. The harder b

LLMs Should Express Uncertainty Explicitly

Model ReleasesDGX agent

arXiv:2604.05306v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often produce confident yet incorrect answers, which can lead to risky failures in real-world applications. We st

Logging Policy Design for Off-Policy Evaluation

SafetyDGX agent

arXiv:2605.15108v1 Announce Type: cross Abstract: Off-policy evaluation (OPE) estimates the value of a target treatment policy (e.g., a recommender system) using data collected by a different logging

LoMETab: Beyond Rank-1 Ensembles for Tabular Deep Learning

Model ReleasesDGX agent

arXiv:2605.14365v1 Announce Type: cross Abstract: Recent tabular learning benchmarks increasingly show a tight performance cluster rather than a clear hierarchy among leading methods, spanning gradien

LoRA in LoRA: Towards Parameter-Efficient Architecture Expansion for Continual Visual Instruction Tuning

Model ReleasesDGX agent

arXiv:2508.06202v2 Announce Type: replace-cross Abstract: Continual Visual Instruction Tuning (CVIT) enables Multimodal Large Language Models (MLLMs) to incrementally learn new tasks over time. Howeve

LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generations

ResearchDGX agent

arXiv:2505.23912v2 Announce Type: replace-cross Abstract: Hallucination remains a major challenge for the safe and trustworthy deployment of large language models (LLMs) in factual content generation.

M^2RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling

ResearchDGX agent

arXiv:2603.14360v2 Announce Type: replace-cross Abstract: Transformers are highly parallel but are limited to computations in the TC^0 complexity class, excluding tasks such as entity tracking and cod

MahaVar: OOD Detection via Class-wise Mahalanobis Distance Variance under Neural Collapse

Model ReleasesDGX agent

arXiv:2605.14413v1 Announce Type: cross Abstract: Out-of-distribution (OOD) detection is a critical component for ensuring the reliability of deep neural networks in safety-critical applications. In t

MALLVI: A Multi-Agent Framework for Integrated Generalized Robotics Manipulation

AgentsDGX agent

arXiv:2602.16898v5 Announce Type: replace-cross Abstract: Task planning for robotic manipulation with large language models (LLMs) is an emerging area. Prior approaches rely on specialized models, fin

MathAtlas: A Benchmark for Autoformalization in the Wild

Model ReleasesDGX agent

arXiv:2605.14061v1 Announce Type: new Abstract: Current autoformalization benchmarks are largely focused on olympiad or undergraduate mathematics, while graduate and research-level mathematics remains

Matrix-Space Reinforcement Learning for Reusing Local Transition Geometry

Local AiDGX agent

arXiv:2605.14304v1 Announce Type: cross Abstract: Compositional generalization in sequential decision-making requires identifying which parts of prior rollouts remain useful for new tasks. Existing me

← Previous
1…252253254255256…358
Next →