AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlog
85,136Total entries
1Added by human
85,135Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,688 results
Model Releases

ProjectionBench: Evaluating Scientific Hypothesis Generation in LLMs Under Progressive Information Disclosure

DGX agent

arXiv:2605.30284v1 Announce Type: new Abstract: Scientific discovery is an inherently creative and uncertain process, requiring reasoning beyond the recall of known knowledge. While many benchmarks ha

model-releasesarxiv-cs-ai
29 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Provably Secure Agent Guardrail

DGX agent

arXiv:2605.29251v1 Announce Type: new Abstract: As large language models transition from bounded generative engines to agents with expansive execution privileges, AI going out of control precipitates

safetyarxiv-cs-ai
29 May 2026
Model Releases

PTCG-Bench: Can LLM Agents Master Pokemon Trading Card Game?

DGX agent

arXiv:2605.29653v1 Announce Type: new Abstract: Given a strategically complex board game, human players can quickly learn to devise strategies after playing a few rounds. Autonomous agents require sim

model-releasesarxiv-cs-ai
29 May 2026
Research

Pushing the Limits of Block Rotations in Post-Training Quantization

DGX agent

arXiv:2601.22347v2 Announce Type: replace-cross Abstract: Recent post-training quantization (PTQ) methods have adopted block rotations to diffuse outliers prior to rounding. While this reduces the ove

researcharxiv-cs-ai
29 May 2026
Model Releases

PuzzleClone: A DSL-Powered Framework for Synthesizing Verifiable Data

DGX agent

arXiv:2508.15180v3 Announce Type: replace Abstract: High-quality mathematical and logical datasets with verifiable answers are essential for strengthening the reasoning capabilities of large language

model-releasesarxiv-cs-ai
29 May 2026
Safety

Quantifying and Optimizing Simplicity via Polynomial Representations

DGX agent

arXiv:2605.29823v1 Announce Type: new Abstract: Deep networks often exhibit a preference for 'simple' solutions, and such a simplicity bias is widely believed to play a key role in generalization. Yet

safetyarxiv-cs-ai
29 May 2026
Safety

Quantum-Enhanced Adversarial Robustness in Artificial Intelligence

DGX agent

arXiv:2605.28899v1 Announce Type: cross Abstract: Artificial Intelligence has achieved remarkable success across diverse application domains. However, its vulnerability to adversarial attacks poses si

safetyarxiv-cs-ai
29 May 2026
Safety

Quotient DAGs for Off-Policy Evaluation:Forward-Flow Importance Sampling and Exact Slate Propensities

DGX agent

arXiv:2605.29500v1 Announce Type: cross Abstract: Off-policy evaluation estimates how a target policy would perform using data collected by a different behavior policy, which is crucial when online te

safetyarxiv-cs-ai
29 May 2026
Model Releases

Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

DGX agent

arXiv:2605.30280v1 Announce Type: cross Abstract: Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented cap

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

RAISE: RAG Design as an Architecture Search Problem

DGX agent

arXiv:2605.30029v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) systems expose numerous design choices spanning query rewriting, chunking, retrieval depth, reranking, and context

model-releasesarxiv-cs-ai
29 May 2026
Agents

Real-rootedness of the Poincare polynomials of overline{mathcal M}_{0,n}: an AI-assisted proof

DGX agent

arXiv:2605.29151v1 Announce Type: cross Abstract: We prove real-rootedness for the Poincare polynomial [ P_n(t)=sum_{i=0}^{n-3} im H^{2i}(overline{mathcal M}_{0,n};Q)t^i ] of the Deligne--Mumford modu

agentsarxiv-cs-ai
29 May 2026
Research

Reasoning about Reasoning: BAPO Bounds on Chain-of-Thought Token Complexity in LLMs

DGX agent

arXiv:2602.02909v2 Announce Type: replace Abstract: Inference-time scaling via chain-of-thought (CoT) reasoning is a major driver of state-of-the-art LLM performance, but it comes with substantial lat

researcharxiv-cs-ai
29 May 2026
Model Releases

Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning

DGX agent

arXiv:2602.00994v2 Announce Type: replace Abstract: Agentic Reinforcement Learning (ARL) trains large language models to interleave reasoning with external tool execution to solve complex tasks. Most

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought

DGX agent

arXiv:2603.05488v4 Announce Type: replace-cross Abstract: We provide evidence of performative chain-of-thought (CoT) in reasoning models, where a model becomes strongly confident in its final answer,

model-releasesarxiv-cs-ai
29 May 2026
Safety

Reasoning While Asking: Transforming Reasoning Large Language Models from Passive Solvers to Proactive Inquirers

DGX agent

arXiv:2601.22139v2 Announce Type: replace-cross Abstract: Reasoning-oriented Large Language Models (LLMs) have achieved remarkable progress with Chain-of-Thought (CoT) prompting, yet they remain funda

safetyarxiv-cs-ai
29 May 2026
Local Ai

Reasoning with Sampling: Cutting at Decision Points

DGX agent

arXiv:2605.30327v1 Announce Type: cross Abstract: Frontier reasoning models are produced by posttraining base language models with reinforcement learning. Recent work has challenged this by showing th

local-aiarxiv-cs-ai
29 May 2026
Safety

ReasonLight: A Multimodal Foundation Model-Enhanced Reinforcement Learning Framework for Zero-Shot Traffic Signal Control

DGX agent

arXiv:2605.29425v1 Announce Type: new Abstract: Reinforcement learning (RL) has shown promise in traffic signal control (TSC). However, its reliance on predefined states limits responsiveness to obser

safetyarxiv-cs-ai
29 May 2026
Model Releases

ReasonOps: Operator Segmentation for LLM Reasoning Traces

DGX agent

arXiv:2605.29192v1 Announce Type: new Abstract: Chain-of-thought traces from large reasoning models can span tens of thousands of tokens, yet we lack a vocabulary for describing their internal structu

model-releasesarxiv-cs-ai
29 May 2026
Safety

Recurrent Structural Policy Gradient for Partially Observable Mean Field Games

DGX agent

arXiv:2602.20141v2 Announce Type: replace Abstract: Mean Field Games (MFGs) provide a principled framework for modelling interactions in large population systems. However, algorithmic progress has bee

safetyarxiv-cs-ai
29 May 2026
Model Releases

Redundant or Necessary? A Benchmark for Detecting Redundant Steps in Agent Trajectories

DGX agent

arXiv:2605.29893v1 Announce Type: new Abstract: LLM-based agents have demonstrated strong capabilities in solving complex tasks through multi-step reasoning and tool use. However, existing evaluation

model-releasesarxiv-cs-ai
29 May 2026
Research

Reinforcement Learning with Robust Rubric Rewards

DGX agent

arXiv:2605.30244v1 Announce Type: cross Abstract: While Reinforcement Learning with Verifiable Rewards (RLVR) is effective for deterministically checkable tasks, many vision-language tasks are partial

researcharxiv-cs-ai
29 May 2026
Research

Rel-MOSS: Towards Imbalanced Relational Deep Learning on Relational Databases

DGX agent

arXiv:2603.07916v2 Announce Type: replace Abstract: In recent advances, to enable a fully data-driven learning paradigm on relational databases (RDB), relational deep learning (RDL) is proposed to str

researcharxiv-cs-ai
29 May 2026
Model Releases

Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents

DGX agent

arXiv:2605.29224v1 Announce Type: cross Abstract: AI agents augment large language models with external tools such as web retrieval, enabling grounded and up-to-date responses. However, incorporating

model-releasesarxiv-cs-ai
29 May 2026
Research

Reliable Reasoning with Large Language Models via Preference-Based Maximum Satisfiability

DGX agent

arXiv:2605.29687v1 Announce Type: new Abstract: Large Language Models (LLMs) excel at understanding natural language but struggle with optimisation tasks involving multiple constraints and user-define

researcharxiv-cs-ai
29 May 2026
Model Releases

REPOT: Recoverable Program-of-Thought via Checkpoint Repair

DGX agent

arXiv:2605.30052v1 Announce Type: cross Abstract: One-shot Program-of-Thought (PoT) emits a Python program that prints a primitive-action plan; a single invalid action silently invalidates the traject

model-releasesarxiv-cs-ai
29 May 2026
Safety

Representation Alignment Rests on Linear Structure

DGX agent

arXiv:2605.28870v1 Announce Type: cross Abstract: We investigate the Platonic Representation Hypothesis (PRH) through a tripartite statistical framework of representations: signal, bias, and noise. {1

safetyarxiv-cs-ai
29 May 2026
Research

Rethinking FID Through the Geometry of the Reference Dataset

DGX agent

arXiv:2605.29335v1 Announce Type: cross Abstract: Frechet Inception Distance (FID) is widely used to evaluate image generators, yet lower FID does not always correspond to better sample quality. We sh

researcharxiv-cs-ai
29 May 2026
Model Releases

Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground Truth

DGX agent

arXiv:2605.29234v1 Announce Type: new Abstract: We study large-scale literature search from two complementary angles: improving the retrieval pipeline, and stress-testing the human reference list as a

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Return-to-Go Is More Than a Number: Q-Guided Alignment for Return-Conditioned Supervised Learning

DGX agent

arXiv:2605.29028v1 Announce Type: cross Abstract: Conditioned Sequence Models (CSMs) learn policies by treating return-to-go (RTG) as a control signal. However, existing CSMs often treat the RTGs as s

model-releasesarxiv-cs-ai
29 May 2026
Safety

Review Arcade: On the Human Alignment and Gameability of LLM Reviews

DGX agent

arXiv:2605.28897v1 Announce Type: new Abstract: LLM-generated reviews for scientific papers are gaining considerable traction and are even being officially piloted by major conferences. We have to ass

safetyarxiv-cs-ai
29 May 2026
Agents

RewardFlow: Topology-Aware Reward Propagation on State Graphs for Agentic RL with Large Language Models

DGX agent

arXiv:2603.18859v2 Announce Type: replace Abstract: Reinforcement learning (RL) shows promise for enhancing LLM agentic reasoning, yet sparse terminal rewards hinder fine-grained optimization. Process

agentsarxiv-cs-ai
29 May 2026
Model Releases

RoboWits: Unexpected Challenges for Robotic Creative Problem Solving

DGX agent

arXiv:2605.30326v1 Announce Type: cross Abstract: The ability to reason, adapt, and creatively solve problems under unexpected challenges is essential for robots operating in real-world environments.

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Robust and Efficient Guardrails with Latent Reasoning

DGX agent

arXiv:2605.29068v1 Announce Type: new Abstract: Maintaining the safety of large language models (LLMs) is crucial as they are increasingly deployed in real-world applications. Existing safety guardrai

model-releasesarxiv-cs-ai
29 May 2026
Safety

Robust and Generalizable Safety Steering for Text-to-Image Diffusion Transformers

DGX agent

arXiv:2605.30049v1 Announce Type: new Abstract: Diffusion Transformers have become a powerful backbone for text-to-image generation, but their layered and cross-modal generation process makes safety c

safetyarxiv-cs-ai
29 May 2026
Safety

Rooted Absorbed Prefix Trajectory Balance with Submodular Replay for GFlowNet Training

DGX agent

arXiv:2603.00454v2 Announce Type: replace-cross Abstract: Generative Flow Networks (GFlowNets) enable fine-tuning large language models to approximate reward-proportional posteriors, but they remain p

safetyarxiv-cs-ai
29 May 2026
Safety

Rubric-Guided Process Reward for Stepwise Model Routing

DGX agent

arXiv:2605.29310v1 Announce Type: new Abstract: Stepwise model routing improves the efficiency of Large Reasoning Models (LRMs) by assigning each reasoning step to a suitable model. Recent methods for

safetyarxiv-cs-ai
29 May 2026
Model Releases

S-MARC: Causal Streaming Reasoning for Full-Duplex Conversational Behavior Modeling

DGX agent

arXiv:2602.11065v2 Announce Type: replace-cross Abstract: Human conversation is organized by an implicit chain of thought and manifests as temporally structured conversational behaviors. Capturing thi

model-releasesarxiv-cs-ai
29 May 2026
Local Ai

S3Mem: Structured Spatiotemporal Scene-Event Memory for Long-Horizon Interactive Question Answering

DGX agent

arXiv:2605.28831v1 Announce Type: cross Abstract: Long-horizon interactive agents often accumulate large trajectory histories yet still fail to answer questions about earlier events reliably. We argue

local-aiarxiv-cs-ai
29 May 2026
Model Releases

SAAS: Self-Aware Reinforcement Learning for Over-Search Mitigation in Agentic Search

DGX agent

arXiv:2605.29796v1 Announce Type: new Abstract: Agentic search enables LLMs to solve complex multi-hop questions through iterative reasoning and external search. Despite the effectiveness, these syste

model-releasesarxiv-cs-ai
29 May 2026
Safety

SafeRx-Agent: A Knowledge-Grounded Multi-Agent Framework for Safe and Explainable Medication Recommendation

DGX agent

arXiv:2605.29146v1 Announce Type: cross Abstract: Medication recommendation predicts medications for patient visits, but existing methods still face two key challenges. At the model level, traditional

safetyarxiv-cs-ai
29 May 2026
Model Releases

SafeSearch: Automated Red-Teaming of LLM-Based Search Agents

DGX agent

arXiv:2509.23694v5 Announce Type: replace Abstract: Search agents connect LLMs to the Internet, enabling them to access broader and more up-to-date information. However, this also introduces a new thr

model-releasesarxiv-cs-ai
29 May 2026
Safety

Same Evidence, Different Answers: Canonical-Context On-Policy Distillation for Multi-Turn Language Models

DGX agent

arXiv:2605.30251v1 Announce Type: cross Abstract: Large language models (LLMs) often solve a task when all instructions are given in a single prompt, but fail when the same information is revealed gra

safetyarxiv-cs-ai
29 May 2026
Model Releases

Same Question, Different Source, Different Answer: Auditing Source-Dependence in Medical Multi-Source RAG

DGX agent

arXiv:2605.29084v1 Announce Type: cross Abstract: A retrieval-augmented generation (RAG) system deployed over a multi-author institutional corpus can give a different answer to the same question depen

model-releasesarxiv-cs-ai
29 May 2026
Applications

Scalable RF Simulation in Generative 4D Worlds

DGX agent

arXiv:2508.12176v2 Announce Type: replace-cross Abstract: Radio Frequency (RF) sensing has emerged as a powerful, privacy-preserving alternative to vision-based methods for various perception tasks. H

applicationsarxiv-cs-ai
29 May 2026
Model Releases

Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet

DGX agent

arXiv:2605.29358v1 Announce Type: new Abstract: We demonstrate that sparse autoencoders can extract interpretable features from Claude 3 Sonnet, a production-scale language model, addressing the open

model-releasesarxiv-cs-ai
29 May 2026
Agents

Scaling Small Agents Through Strategy Auctions

DGX agent

arXiv:2602.02751v2 Announce Type: replace-cross Abstract: Small language models are increasingly viewed as a promising, cost-effective approach to agentic AI, with proponents claiming they are suffici

agentsarxiv-cs-ai
29 May 2026
Model Releases

SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers

DGX agent

arXiv:2605.29059v1 Announce Type: cross Abstract: Smart contract decompilation aims to recover high-level source code from bytecode, but evaluating decompilers remains difficult because existing studi

model-releasesarxiv-cs-ai
29 May 2026
Hardware

ScheduleStream: Temporal Planning with Samplers for GPU-Accelerated Multi-Arm Task and Motion Planning & Scheduling

DGX agent

arXiv:2511.04758v2 Announce Type: replace-cross Abstract: Bimanual and humanoid robots are appealing because of their human-like ability to leverage multiple arms to efficiently complete tasks. Howeve

hardwarearxiv-cs-ai
29 May 2026
← Previous
1…246247248249250…452
Next →