AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
11 May 2026

Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning

ResearchDGX agent

arXiv:2602.14868v2 Announce Type: replace-cross Abstract: Reinforcement learning has emerged as a powerful paradigm for unlocking reasoning capabilities in language models. However, relying on sparse

Gradient Extrapolation-Based Policy Optimization

Model ReleasesDGX agent

arXiv:2605.06755v1 Announce Type: cross Abstract: Reinforcement learning is widely used to improve the reasoning ability of large language models, especially when answers can be automatically checked.

Graph-Structured Hyperdimensional Computing for Data-Efficient and Explainable Process-Structure-Property Prediction

Model ReleasesDGX agent

arXiv:2605.07999v1 Announce Type: cross Abstract: Multiphoton photoreduction enables high-fidelity fabrication of complex 3D microstructures, yet reliable process-structure-property (PSP) prediction r


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

GraphDC: A Divide-and-Conquer Multi-Agent System for Scalable Graph Algorithm Reasoning

AgentsDGX agent

arXiv:2605.06671v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated strong potential for many mathematical problems. However, their performance on graph algorithmic tasks is

GraphReAct: Reasoning and Acting for Multi-step Graph Inference

Model ReleasesDGX agent

arXiv:2605.07357v1 Announce Type: new Abstract: Reasoning-acting frameworks enhance large language models (LLMs) by interleaving reasoning with actions for dynamic information acquisition. However, ex

Group of Skills: Group-Structured Skill Retrieval for Agent Skill Libraries

AgentsDGX agent

arXiv:2605.06978v1 Announce Type: cross Abstract: Skill-augmented agents increasingly rely on large reusable skill libraries, but retrieving relevant skills is not the same as presenting usable contex

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations

Model ReleasesDGX agent

arXiv:2605.07053v1 Announce Type: cross Abstract: Benchmarks like GSM8K are popular measures of mathematical reasoning, but leaderboard gains can overstate true capability due to memorization of fixed

Hallucination Detection via Activations of Open-Weight Proxy Analyzers

Model ReleasesDGX agent

arXiv:2605.07209v1 Announce Type: cross Abstract: We introduce a proxy-analyzer framework for detecting hallucinations in large language models. Instead of looking inside the generating model, our sys

Hard to Read, Easy to Jailbreak: How Visual Degradation Bypasses MLLM Safety Alignment

SafetyDGX agent

arXiv:2605.07250v1 Announce Type: cross Abstract: Recent advancements in visual context compression enable MLLMs to process ultra-long contexts efficiently by rendering text into images. However, we i

HARMONY: Bridging the Personalization-Generalization Gap by Mitigating Representation Skew in Heterogeneous Split Federated Learning

Local AiDGX agent

arXiv:2605.07211v1 Announce Type: cross Abstract: Mobile devices face diverse resource constraints and non-IID data class distributions, requiring fast on-device inference for local in-distribution (I

HBEE: Human Behavioral Entropy Engine -- Pre-Registered Multi-Agent LLM Simulation of Peer-Suspicion-Based Detection Inversion

AgentsDGX agent

arXiv:2605.07472v1 Announce Type: cross Abstract: Insider threat detection assumes that an adaptive insider leaves behavioral residue distinguishing them from legitimate users. We test this assumption

Hidden Coalitions in Multi-Agent AI: A Spectral Diagnostic from Internal Representations

SafetyDGX agent

arXiv:2605.06696v1 Announce Type: new Abstract: Collections of interacting AI agents can form coalitions, creating emergent group-level organization that is critical for AI safety and alignment. Howev

Hierarchical Task Network Planning with LLM-Generated Heuristics

Model ReleasesDGX agent

arXiv:2605.07707v1 Announce Type: new Abstract: HTN planning is a variation of classical planning where, instead of searching for a linear sequence of actions, an algorithm decomposes higher-level tas

HMACE: Heterogeneous Multi-Agent Collaborative Evolution for Combinatorial Optimization

Local AiDGX agent

arXiv:2605.07214v1 Announce Type: new Abstract: Large Language Models have recently emerged as a promising paradigm for automated heuristic design for NP-hard combinatorial optimization problems. Desp

How Do Language Models Compose Functions?

ResearchDGX agent

arXiv:2510.01685v2 Announce Type: replace-cross Abstract: While large language models (LLMs) appear to be increasingly capable of solving compositional tasks, it is an open question whether they do so

How Far Are VLMs from Privacy Awareness in the Physical World? An Empirical Study

Model ReleasesDGX agent

arXiv:2605.05340v2 Announce Type: replace-cross Abstract: As Vision-Language Models (VLMs) are increasingly deployed as autonomous cognitive cores for embodied assistants, evaluating their privacy awa

How Log-Barrier Helps Exploration in Policy Optimization

SafetyDGX agent

arXiv:2603.15001v2 Announce Type: replace-cross Abstract: Recently, it has been shown that the Stochastic Gradient Bandit (SGB) algorithm converges to a globally optimal policy with a constant learnin

How to Compress KV Cache in RL Post-Training? Shadow Mask Distillation for Memory-Efficient Alignment

SafetyDGX agent

arXiv:2605.06850v1 Announce Type: cross Abstract: Reinforcement Learning (RL) has emerged as a crucial paradigm for unlocking the advanced reasoning capabilities of Large Language Models (LLMs), encom

How Well Do LLMs Perform on the Simplest Long-Chain Reasoning Tasks: An Empirical Study on the Equivalence Class Problem

ResearchDGX agent

arXiv:2605.06882v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved great improvements in recent years. Nevertheless, it still remains unclear how good LLMs are for reasoning ta

HYPER: A Foundation Model for Inductive Link Prediction with Knowledge Hypergraphs

TutorialsDGX agent

arXiv:2506.12362v3 Announce Type: replace-cross Abstract: Inductive link prediction with knowledge hypergraphs is the task of predicting missing hyperedges involving completely novel entities (i.e., n

HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents

Model ReleasesDGX agent

arXiv:2605.07177v1 Announce Type: cross Abstract: Existing multimodal search agents process target entities sequentially, issuing one tool call per entity and accumulating redundant interaction rounds

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training

SafetyDGX agent

arXiv:2605.07316v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards improves LLM reasoning but often induces overthinking, where models generate unnecessarily long reasoning

Implicit Preference Alignment for Human Image Animation

Model ReleasesDGX agent

arXiv:2605.07545v1 Announce Type: cross Abstract: Human image animation has witnessed significant advancements, yet generating high-fidelity hand motions remains a persistent challenge due to their hi

In-Context Credit Assignment via the Core

Model ReleasesDGX agent

arXiv:2605.06920v1 Announce Type: cross Abstract: We propose incentive-aligned mechanisms for in-context credit assignment: the task of assigning credit for AI-generated content (e.g. code, news artic

Inference Time Causal Probing in LLMs

Model ReleasesDGX agent

arXiv:2605.07631v1 Announce Type: new Abstract: Causal probing methods aim to test and control how internal representations influence the behavior of generative models. In causal probing, an intervent

INO-SGD: Addressing Utility Imbalance under Individualized Differential Privacy

ResearchDGX agent

arXiv:2605.07930v1 Announce Type: cross Abstract: Differential privacy (DP) is widely employed in machine learning to protect confidential or sensitive training data from being revealed. As data owner

Intelligent Truck Matching in Full Truckload Shipments using Ping2Hex approach

ApplicationsDGX agent

arXiv:2605.07733v1 Announce Type: cross Abstract: Accurate truck-to-shipment matching using GPS data is foundational for full truckload supply chain visibility, enabling real-time tracking and accurat

IntentGrasp: A Comprehensive Benchmark for Intent Understanding

Model ReleasesDGX agent

arXiv:2605.06832v1 Announce Type: cross Abstract: Accurately understanding the intent behind speech, conversation, and writing is crucial to the development of helpful Large Language Model (LLM) assis

Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry

ResearchDGX agent

arXiv:2510.08638v3 Announce Type: replace-cross Abstract: DINOv2 is routinely deployed to recognize objects, scenes, and actions; yet the nature of what it perceives remains unknown. As a working base

InvThink: Premortem Reasoning for Safer Language Models

SafetyDGX agent

arXiv:2510.01569v3 Announce Type: replace Abstract: We present InvThink, a training and prompting framework that requires the model to enumerate, analyze, and constrain potential failures before gener

Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization

ResearchDGX agent

arXiv:2512.23032v2 Announce Type: replace-cross Abstract: Recent work, using the Biasing Features metric, labels a CoT as unfaithful if it omits a prompt-injected hint that affected the prediction. We

Is Your Prompt Poisoning Code? Defect Induction Rates and Security Mitigation Strategies

Model ReleasesDGX agent

arXiv:2510.22944v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have become indispensable for automated code generation, yet the quality and security of their outputs remain a c

It Just Takes Two: Scaling Amortized Inference to Large Sets

ResearchDGX agent

arXiv:2605.07972v1 Announce Type: cross Abstract: Neural posterior estimation has emerged as a powerful tool for amortized inference, with growing adoption across scientific and applied domains. In ma

KL for a KL: On-Policy Distillation with Control Variate Baseline

SafetyDGX agent

arXiv:2605.07865v1 Announce Type: cross Abstract: On-Policy Distillation (OPD) has emerged as a dominant post-training paradigm for large language models, especially for reasoning domains. However, OP

Knowledge Transfer Scaling Laws for 3D Medical Imaging

ResearchDGX agent

arXiv:2605.06859v1 Announce Type: cross Abstract: Vision foundation models are increasingly moving beyond 2D to volumetric domains such as 3D medical imaging, where unified pretraining across differen

Kurtosis-Guided Denoising Score Matching for Tabular Anomaly Detection

Local AiDGX agent

arXiv:2605.06955v1 Announce Type: cross Abstract: Denoising score matching (DSM) provides a way to learn data distributions by training a neural network to recover the score function, defined as the g

LARAG: Link-Aware Retrieval Strategy for RAG Systems in Hyperlinked Technical Documentation

Model ReleasesDGX agent

arXiv:2605.07517v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) enhances the factual grounding of Large Language Models by conditioning their outputs on external documents. Howe

Latent-Space Causal Discovery from Indirect Neuroimaging Observations

ResearchDGX agent

arXiv:2602.09034v2 Announce Type: replace-cross Abstract: Neuroimaging does not observe causal variables directly: hemodynamics and volume conduction distort signals so that statistical dependence nee

Learning and Reusing Policy Decompositions for Hierarchical Generalized Planning with LLM Agents

Model ReleasesDGX agent

arXiv:2605.06957v1 Announce Type: new Abstract: We present a dynamic policy-learning approach that combines generalized planning and hierarchical task decomposition for LLM-based agents. Our method, H

Learning CLI Agents with Structured Action Credit under Selective Observation

AgentsDGX agent

arXiv:2605.08013v1 Announce Type: new Abstract: Command line interface (CLI) agents are emerging as a practical paradigm for agent-computer interaction over evolving filesystems, executable command li

Learning Cross-Atlas Consistent Brain Disorder Representations via Disentangled Multi-Atlas Functional Connectivity Learning

SafetyDGX agent

arXiv:2605.07026v1 Announce Type: cross Abstract: Functional connectivity (FC) derived from resting-state fMRI is widely used to characterize large-scale brain network alterations in neurological and

Learning Multi-Relational Graph Representations for DNA Methylation-Based Biological Age Estimation

Local AiDGX agent

arXiv:2605.07175v1 Announce Type: cross Abstract: Aging clocks aim to estimate biological age, a measure of physiological state distinct from chronological age, from observable biomarkers, and are wid

Learning to Communicate Locally for Large-Scale Multi-Agent Pathfinding

Local AiDGX agent

arXiv:2605.07637v1 Announce Type: new Abstract: Multi-agent pathfinding (MAPF) is a widely used abstraction for multi-robot trajectory planning problems, where multiple homogeneous agents move simulta

Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis

ResearchDGX agent

arXiv:2511.09907v5 Announce Type: replace Abstract: Data synthesis for training large reasoning models offers a scalable alternative to limited, human-curated datasets, enabling the creation of high-q

Learning Visual Feature-Based World Models via Residual Latent Action

SafetyDGX agent

arXiv:2605.07079v1 Announce Type: cross Abstract: World models predict future transitions from observations and actions. Existing works predominantly focus on image generation only. Visual feature-bas

LensVLM: Selective Context Expansion for Compressed Visual Representation of Text

ResearchDGX agent

arXiv:2605.07019v1 Announce Type: cross Abstract: Vision Language Models (VLMs) offer the exciting possibility of processing text as rendered images, bypassing the need for tokenizing the text into lo

Limitations on Accurate, Trusted, Human-level Reasoning

ResearchDGX agent

arXiv:2509.21654v2 Announce Type: replace-cross Abstract: We identify a fundamental incompatibility between the goals of accuracy, trust, and human-level reasoning in artificial intelligence (AI) syst

LiteGUI: Distilling Compact GUI Agents with Reinforcement Learning

Local AiDGX agent

arXiv:2605.07505v1 Announce Type: new Abstract: Developing lightweight, on-device vision-language GUI agents is essential for efficient cross-platform automated interaction. However, current on-device

LithoBench: Benchmarking Large Multimodal Models for Remote-Sensing Lithology Interpretation

Model ReleasesDGX agent

arXiv:2605.07640v1 Announce Type: cross Abstract: Remote sensing lithology interpretation is fundamental to geological surveys, mineral exploration, and regional geological mapping. Unlike general lan

LLM-Based Agents for Competitive Landscape Mapping in Drug Asset Due Diligence

Model ReleasesDGX agent

arXiv:2508.16571v4 Announce Type: replace Abstract: In this paper, we describe and benchmark a competitor-discovery component used within an agentic AI system for fast drug asset due diligence. A comp

LLM-Guided Open Hypothesis Learning from Autonomous Scanning Probe Microscopy Experiments

AgentsDGX agent

arXiv:2605.06839v1 Announce Type: cross Abstract: Autonomous experimentation has transformed microscopy and materials discovery by enabling closed-loop optimization including imaging and spectroscopy

LLM hallucinations in the wild: Large-scale evidence from non-existent citations

ApplicationsDGX agent

arXiv:2605.07723v1 Announce Type: cross Abstract: Large language models (LLMs) are known to generate plausible but false information across a wide range of contexts, yet the real-world magnitude and c

LR-SGS: Robust LiDAR-Reflectance-Guided Salient Gaussian Splatting for Self-Driving Scene Reconstruction

ResearchDGX agent

arXiv:2603.12647v2 Announce Type: replace-cross Abstract: Recent 3D Gaussian Splatting (3DGS) methods have demonstrated the feasibility of self-driving scene reconstruction and novel view synthesis. H

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference

ResearchDGX agent

arXiv:2605.05225v2 Announce Type: replace-cross Abstract: Mixture-of-Experts Multimodal Large Language Models (MoE MLLMs) suffer from a significant efficiency bottleneck during Expert Parallelism (EP)

Mage: Multi-Axis Evaluation of LLM-Generated Executable Game Scenes Beyond Compile-Pass Rate

Model ReleasesDGX agent

arXiv:2605.07342v1 Announce Type: cross Abstract: Compile-pass rate is the dominant evaluation signal for LLM code generation, yet for multi-component domain-specific artifacts it can be actively misl

Making AI Evaluation Deployment Relevant Through Context Specification

ResearchDGX agent

arXiv:2603.06811v3 Announce Type: replace Abstract: With many organizations struggling to gain value from AI deployments, pressure to evaluate AI in an informed manner has intensified. Status quo AI e

MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge

SafetyDGX agent

arXiv:2507.21183v5 Announce Type: replace-cross Abstract: As the era of large language models (LLMs) unfolds, Preference Optimization (PO) methods have become a central approach to aligning LLMs with

MAS-Algorithm: A Workflow for Solving Algorithmic Programming Problems with a Multi-Agent System

Model ReleasesDGX agent

arXiv:2605.05949v2 Announce Type: replace Abstract: Algorithmic problem solving serves as a rigorous testbed for evaluating structured reasoning in AI coding systems, as it directly reflects a model's

Mask2Cause: Causal Discovery via Adjacency Constrained Causal Attention

Model ReleasesDGX agent

arXiv:2605.07280v1 Announce Type: cross Abstract: Leveraging deep learning for causal discovery in time series remains challenging because existing neural methods predominantly rely on component-wise

Mathematical Reasoning via Intervention-Based Time-Series Causal Discovery Using LLMs as Concept Mastery Simulators

Model ReleasesDGX agent

arXiv:2605.07600v1 Announce Type: cross Abstract: Recent methods for improving LLM mathematical reasoning, whether through MCTS-based test-time search or causal graph-guided knowledge injection, canno

← Previous
1…280281282283284…358
Next →