AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
29 May 2026

PRO-CUA: Process-Reward Optimization for Computer Use Agents

SafetyDGX agent

arXiv:2605.29119v1 Announce Type: new Abstract: Computer use agents (CUAs) have shown strong potential for automating complex digital workflows, yet their training remains constrained by costly live e

Projectional Decoding: Towards Semantic-Aware LLM Generation

ResearchDGX agent

arXiv:2605.30054v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate software artifacts across many software engineering (SE) tasks, yet ensuring the semant

ProjectionBench: Evaluating Scientific Hypothesis Generation in LLMs Under Progressive Information Disclosure

Model ReleasesDGX agent

arXiv:2605.30284v1 Announce Type: new Abstract: Scientific discovery is an inherently creative and uncertain process, requiring reasoning beyond the recall of known knowledge. While many benchmarks ha


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Provably Secure Agent Guardrail

SafetyDGX agent

arXiv:2605.29251v1 Announce Type: new Abstract: As large language models transition from bounded generative engines to agents with expansive execution privileges, AI going out of control precipitates

PTCG-Bench: Can LLM Agents Master Pokemon Trading Card Game?

Model ReleasesDGX agent

arXiv:2605.29653v1 Announce Type: new Abstract: Given a strategically complex board game, human players can quickly learn to devise strategies after playing a few rounds. Autonomous agents require sim

Pushing the Limits of Block Rotations in Post-Training Quantization

ResearchDGX agent

arXiv:2601.22347v2 Announce Type: replace-cross Abstract: Recent post-training quantization (PTQ) methods have adopted block rotations to diffuse outliers prior to rounding. While this reduces the ove

PuzzleClone: A DSL-Powered Framework for Synthesizing Verifiable Data

Model ReleasesDGX agent

arXiv:2508.15180v3 Announce Type: replace Abstract: High-quality mathematical and logical datasets with verifiable answers are essential for strengthening the reasoning capabilities of large language

Quantifying and Optimizing Simplicity via Polynomial Representations

SafetyDGX agent

arXiv:2605.29823v1 Announce Type: new Abstract: Deep networks often exhibit a preference for 'simple' solutions, and such a simplicity bias is widely believed to play a key role in generalization. Yet

Quantum-Enhanced Adversarial Robustness in Artificial Intelligence

SafetyDGX agent

arXiv:2605.28899v1 Announce Type: cross Abstract: Artificial Intelligence has achieved remarkable success across diverse application domains. However, its vulnerability to adversarial attacks poses si

Quotient DAGs for Off-Policy Evaluation:Forward-Flow Importance Sampling and Exact Slate Propensities

SafetyDGX agent

arXiv:2605.29500v1 Announce Type: cross Abstract: Off-policy evaluation estimates how a target policy would perform using data collected by a different behavior policy, which is crucial when online te

Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

Model ReleasesDGX agent

arXiv:2605.30280v1 Announce Type: cross Abstract: Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented cap

RAISE: RAG Design as an Architecture Search Problem

Model ReleasesDGX agent

arXiv:2605.30029v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) systems expose numerous design choices spanning query rewriting, chunking, retrieval depth, reranking, and context

Real-rootedness of the Poincare polynomials of overline{mathcal M}_{0,n}: an AI-assisted proof

AgentsDGX agent

arXiv:2605.29151v1 Announce Type: cross Abstract: We prove real-rootedness for the Poincare polynomial [ P_n(t)=sum_{i=0}^{n-3} im H^{2i}(overline{mathcal M}_{0,n};Q)t^i ] of the Deligne--Mumford modu

Reasoning about Reasoning: BAPO Bounds on Chain-of-Thought Token Complexity in LLMs

ResearchDGX agent

arXiv:2602.02909v2 Announce Type: replace Abstract: Inference-time scaling via chain-of-thought (CoT) reasoning is a major driver of state-of-the-art LLM performance, but it comes with substantial lat

Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning

Model ReleasesDGX agent

arXiv:2602.00994v2 Announce Type: replace Abstract: Agentic Reinforcement Learning (ARL) trains large language models to interleave reasoning with external tool execution to solve complex tasks. Most

Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought

Model ReleasesDGX agent

arXiv:2603.05488v4 Announce Type: replace-cross Abstract: We provide evidence of performative chain-of-thought (CoT) in reasoning models, where a model becomes strongly confident in its final answer,

Reasoning While Asking: Transforming Reasoning Large Language Models from Passive Solvers to Proactive Inquirers

SafetyDGX agent

arXiv:2601.22139v2 Announce Type: replace-cross Abstract: Reasoning-oriented Large Language Models (LLMs) have achieved remarkable progress with Chain-of-Thought (CoT) prompting, yet they remain funda

Reasoning with Sampling: Cutting at Decision Points

Local AiDGX agent

arXiv:2605.30327v1 Announce Type: cross Abstract: Frontier reasoning models are produced by posttraining base language models with reinforcement learning. Recent work has challenged this by showing th

ReasonLight: A Multimodal Foundation Model-Enhanced Reinforcement Learning Framework for Zero-Shot Traffic Signal Control

SafetyDGX agent

arXiv:2605.29425v1 Announce Type: new Abstract: Reinforcement learning (RL) has shown promise in traffic signal control (TSC). However, its reliance on predefined states limits responsiveness to obser

ReasonOps: Operator Segmentation for LLM Reasoning Traces

Model ReleasesDGX agent

arXiv:2605.29192v1 Announce Type: new Abstract: Chain-of-thought traces from large reasoning models can span tens of thousands of tokens, yet we lack a vocabulary for describing their internal structu

Recurrent Structural Policy Gradient for Partially Observable Mean Field Games

SafetyDGX agent

arXiv:2602.20141v2 Announce Type: replace Abstract: Mean Field Games (MFGs) provide a principled framework for modelling interactions in large population systems. However, algorithmic progress has bee

Redundant or Necessary? A Benchmark for Detecting Redundant Steps in Agent Trajectories

Model ReleasesDGX agent

arXiv:2605.29893v1 Announce Type: new Abstract: LLM-based agents have demonstrated strong capabilities in solving complex tasks through multi-step reasoning and tool use. However, existing evaluation

Reinforcement Learning with Robust Rubric Rewards

ResearchDGX agent

arXiv:2605.30244v1 Announce Type: cross Abstract: While Reinforcement Learning with Verifiable Rewards (RLVR) is effective for deterministically checkable tasks, many vision-language tasks are partial

Rel-MOSS: Towards Imbalanced Relational Deep Learning on Relational Databases

ResearchDGX agent

arXiv:2603.07916v2 Announce Type: replace Abstract: In recent advances, to enable a fully data-driven learning paradigm on relational databases (RDB), relational deep learning (RDL) is proposed to str

Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents

Model ReleasesDGX agent

arXiv:2605.29224v1 Announce Type: cross Abstract: AI agents augment large language models with external tools such as web retrieval, enabling grounded and up-to-date responses. However, incorporating

Reliable Reasoning with Large Language Models via Preference-Based Maximum Satisfiability

ResearchDGX agent

arXiv:2605.29687v1 Announce Type: new Abstract: Large Language Models (LLMs) excel at understanding natural language but struggle with optimisation tasks involving multiple constraints and user-define

REPOT: Recoverable Program-of-Thought via Checkpoint Repair

Model ReleasesDGX agent

arXiv:2605.30052v1 Announce Type: cross Abstract: One-shot Program-of-Thought (PoT) emits a Python program that prints a primitive-action plan; a single invalid action silently invalidates the traject

Representation Alignment Rests on Linear Structure

SafetyDGX agent

arXiv:2605.28870v1 Announce Type: cross Abstract: We investigate the Platonic Representation Hypothesis (PRH) through a tripartite statistical framework of representations: signal, bias, and noise. {1

Rethinking FID Through the Geometry of the Reference Dataset

ResearchDGX agent

arXiv:2605.29335v1 Announce Type: cross Abstract: Frechet Inception Distance (FID) is widely used to evaluate image generators, yet lower FID does not always correspond to better sample quality. We sh

Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground Truth

Model ReleasesDGX agent

arXiv:2605.29234v1 Announce Type: new Abstract: We study large-scale literature search from two complementary angles: improving the retrieval pipeline, and stress-testing the human reference list as a

Return-to-Go Is More Than a Number: Q-Guided Alignment for Return-Conditioned Supervised Learning

Model ReleasesDGX agent

arXiv:2605.29028v1 Announce Type: cross Abstract: Conditioned Sequence Models (CSMs) learn policies by treating return-to-go (RTG) as a control signal. However, existing CSMs often treat the RTGs as s

Review Arcade: On the Human Alignment and Gameability of LLM Reviews

SafetyDGX agent

arXiv:2605.28897v1 Announce Type: new Abstract: LLM-generated reviews for scientific papers are gaining considerable traction and are even being officially piloted by major conferences. We have to ass

RewardFlow: Topology-Aware Reward Propagation on State Graphs for Agentic RL with Large Language Models

AgentsDGX agent

arXiv:2603.18859v2 Announce Type: replace Abstract: Reinforcement learning (RL) shows promise for enhancing LLM agentic reasoning, yet sparse terminal rewards hinder fine-grained optimization. Process

RoboWits: Unexpected Challenges for Robotic Creative Problem Solving

Model ReleasesDGX agent

arXiv:2605.30326v1 Announce Type: cross Abstract: The ability to reason, adapt, and creatively solve problems under unexpected challenges is essential for robots operating in real-world environments.

Robust and Efficient Guardrails with Latent Reasoning

Model ReleasesDGX agent

arXiv:2605.29068v1 Announce Type: new Abstract: Maintaining the safety of large language models (LLMs) is crucial as they are increasingly deployed in real-world applications. Existing safety guardrai

Robust and Generalizable Safety Steering for Text-to-Image Diffusion Transformers

SafetyDGX agent

arXiv:2605.30049v1 Announce Type: new Abstract: Diffusion Transformers have become a powerful backbone for text-to-image generation, but their layered and cross-modal generation process makes safety c

Rooted Absorbed Prefix Trajectory Balance with Submodular Replay for GFlowNet Training

SafetyDGX agent

arXiv:2603.00454v2 Announce Type: replace-cross Abstract: Generative Flow Networks (GFlowNets) enable fine-tuning large language models to approximate reward-proportional posteriors, but they remain p

Rubric-Guided Process Reward for Stepwise Model Routing

SafetyDGX agent

arXiv:2605.29310v1 Announce Type: new Abstract: Stepwise model routing improves the efficiency of Large Reasoning Models (LRMs) by assigning each reasoning step to a suitable model. Recent methods for

S-MARC: Causal Streaming Reasoning for Full-Duplex Conversational Behavior Modeling

Model ReleasesDGX agent

arXiv:2602.11065v2 Announce Type: replace-cross Abstract: Human conversation is organized by an implicit chain of thought and manifests as temporally structured conversational behaviors. Capturing thi

S3Mem: Structured Spatiotemporal Scene-Event Memory for Long-Horizon Interactive Question Answering

Local AiDGX agent

arXiv:2605.28831v1 Announce Type: cross Abstract: Long-horizon interactive agents often accumulate large trajectory histories yet still fail to answer questions about earlier events reliably. We argue

SAAS: Self-Aware Reinforcement Learning for Over-Search Mitigation in Agentic Search

Model ReleasesDGX agent

arXiv:2605.29796v1 Announce Type: new Abstract: Agentic search enables LLMs to solve complex multi-hop questions through iterative reasoning and external search. Despite the effectiveness, these syste

SafeRx-Agent: A Knowledge-Grounded Multi-Agent Framework for Safe and Explainable Medication Recommendation

SafetyDGX agent

arXiv:2605.29146v1 Announce Type: cross Abstract: Medication recommendation predicts medications for patient visits, but existing methods still face two key challenges. At the model level, traditional

SafeSearch: Automated Red-Teaming of LLM-Based Search Agents

Model ReleasesDGX agent

arXiv:2509.23694v5 Announce Type: replace Abstract: Search agents connect LLMs to the Internet, enabling them to access broader and more up-to-date information. However, this also introduces a new thr

Same Evidence, Different Answers: Canonical-Context On-Policy Distillation for Multi-Turn Language Models

SafetyDGX agent

arXiv:2605.30251v1 Announce Type: cross Abstract: Large language models (LLMs) often solve a task when all instructions are given in a single prompt, but fail when the same information is revealed gra

Same Question, Different Source, Different Answer: Auditing Source-Dependence in Medical Multi-Source RAG

Model ReleasesDGX agent

arXiv:2605.29084v1 Announce Type: cross Abstract: A retrieval-augmented generation (RAG) system deployed over a multi-author institutional corpus can give a different answer to the same question depen

Scalable RF Simulation in Generative 4D Worlds

ApplicationsDGX agent

arXiv:2508.12176v2 Announce Type: replace-cross Abstract: Radio Frequency (RF) sensing has emerged as a powerful, privacy-preserving alternative to vision-based methods for various perception tasks. H

Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet

Model ReleasesDGX agent

arXiv:2605.29358v1 Announce Type: new Abstract: We demonstrate that sparse autoencoders can extract interpretable features from Claude 3 Sonnet, a production-scale language model, addressing the open

Scaling Small Agents Through Strategy Auctions

AgentsDGX agent

arXiv:2602.02751v2 Announce Type: replace-cross Abstract: Small language models are increasingly viewed as a promising, cost-effective approach to agentic AI, with proponents claiming they are suffici

SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers

Model ReleasesDGX agent

arXiv:2605.29059v1 Announce Type: cross Abstract: Smart contract decompilation aims to recover high-level source code from bytecode, but evaluating decompilers remains difficult because existing studi

ScheduleStream: Temporal Planning with Samplers for GPU-Accelerated Multi-Arm Task and Motion Planning & Scheduling

HardwareDGX agent

arXiv:2511.04758v2 Announce Type: replace-cross Abstract: Bimanual and humanoid robots are appealing because of their human-like ability to leverage multiple arms to efficiently complete tasks. Howeve

SchGen: PCB Schematic Generation with Semantic-Grounded Code Representations

AgentsDGX agent

arXiv:2605.30345v1 Announce Type: new Abstract: Printed circuit board (PCB) schematic design defines nearly all electronic hardware, but it remains manual and expertise-intensive. While generative AI

SciIntBench: Measuring LLM Compliance with Research Integrity Norms Under Adversarial Framing

Model ReleasesDGX agent

arXiv:2605.29468v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to support scientific work, but it is unclear whether they uphold responsible conduct of research (

SCoOP: Semantic Consistent Opinion Pooling for Uncertainty Quantification in Multiple Vision-Language Model Systems

ResearchDGX agent

arXiv:2603.23853v3 Announce Type: replace Abstract: Combining multiple Vision-Language Models (VLMs) can enhance multimodal reasoning and robustness, but aggregating heterogeneous models' outputs ampl

SCOPE: A Lightweight-training LLM Framework for Air Traffic Control Readback Monitoring

ResearchDGX agent

arXiv:2605.29543v1 Announce Type: cross Abstract: Pilot readback of Air Traffic Control (ATC) voice instructions is a primary safeguard against miscommunication in air transportation. However, readbac

SCOPE: Prompt Evolution for Enhancing Agent Effectiveness

Model ReleasesDGX agent

arXiv:2512.15374v2 Announce Type: replace Abstract: Large Language Model (LLM) agents are increasingly deployed in environments that generate massive, dynamic contexts. However, a critical bottleneck

Selection Hyper-heuristics Can Automatically Adjust the Learning Period to Optimally Solve Pseudo-Boolean Problems

Model ReleasesDGX agent

arXiv:2605.29916v1 Announce Type: cross Abstract: The Random Gradient hyper-heuristic was recently shown to be able to learn the optimal neighbourhood size when optimizing the LeadingOnes benchmark vi

Selective QA over Conflicting Multi-Source Personal Memory: A Diagnostic Testbed and Method Comparison

Model ReleasesDGX agent

arXiv:2605.30087v1 Announce Type: new Abstract: Emerging personal AI agents are moving toward persistent, multi-source memory. This creates an evaluation problem: systems must decide how to use confli

Self-Play Reinforcement Learning under Imperfect Information in Big 2

SafetyDGX agent

arXiv:2605.28863v1 Announce Type: cross Abstract: Imperfect-information multiplayer games test whether agents can act under hidden information, sparse rewards, and non-stationary opponents. We study t

Self-Trained Verification for Training- and Test-Time Self-Improvement

ResearchDGX agent

arXiv:2605.30290v1 Announce Type: cross Abstract: Self-improvement at scale has been a longstanding goal for reasoning models, and there are two natural places to do it: at test time, through verifica

Semantic and Visual Evidence for Efficient Long-Video Reasoning: A Solution for the HD-EPIC VQA Challenge

Model ReleasesDGX agent

arXiv:2605.29402v1 Announce Type: cross Abstract: Understanding long-form egocentric videos remains challenging for multimodal large language models (MLLMs) due to limited context length and insuffici

← Previous
1…193194195196197…358
Next →