AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
Model Releases

Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents

DGX agent

arXiv:2605.29224v1 Announce Type: cross Abstract: AI agents augment large language models with external tools such as web retrieval, enabling grounded and up-to-date responses. However, incorporating

model-releasesarxiv-cs-ai
29 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Reliable Reasoning with Large Language Models via Preference-Based Maximum Satisfiability

DGX agent

arXiv:2605.29687v1 Announce Type: new Abstract: Large Language Models (LLMs) excel at understanding natural language but struggle with optimisation tasks involving multiple constraints and user-define

researcharxiv-cs-ai
29 May 2026
Model Releases

REPOT: Recoverable Program-of-Thought via Checkpoint Repair

DGX agent

arXiv:2605.30052v1 Announce Type: cross Abstract: One-shot Program-of-Thought (PoT) emits a Python program that prints a primitive-action plan; a single invalid action silently invalidates the traject

model-releasesarxiv-cs-ai
29 May 2026
Safety

Representation Alignment Rests on Linear Structure

DGX agent

arXiv:2605.28870v1 Announce Type: cross Abstract: We investigate the Platonic Representation Hypothesis (PRH) through a tripartite statistical framework of representations: signal, bias, and noise. {1

safetyarxiv-cs-ai
29 May 2026
Research

Rethinking FID Through the Geometry of the Reference Dataset

DGX agent

arXiv:2605.29335v1 Announce Type: cross Abstract: Frechet Inception Distance (FID) is widely used to evaluate image generators, yet lower FID does not always correspond to better sample quality. We sh

researcharxiv-cs-ai
29 May 2026
Model Releases

Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground Truth

DGX agent

arXiv:2605.29234v1 Announce Type: new Abstract: We study large-scale literature search from two complementary angles: improving the retrieval pipeline, and stress-testing the human reference list as a

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Return-to-Go Is More Than a Number: Q-Guided Alignment for Return-Conditioned Supervised Learning

DGX agent

arXiv:2605.29028v1 Announce Type: cross Abstract: Conditioned Sequence Models (CSMs) learn policies by treating return-to-go (RTG) as a control signal. However, existing CSMs often treat the RTGs as s

model-releasesarxiv-cs-ai
29 May 2026
Safety

Review Arcade: On the Human Alignment and Gameability of LLM Reviews

DGX agent

arXiv:2605.28897v1 Announce Type: new Abstract: LLM-generated reviews for scientific papers are gaining considerable traction and are even being officially piloted by major conferences. We have to ass

safetyarxiv-cs-ai
29 May 2026
Agents

RewardFlow: Topology-Aware Reward Propagation on State Graphs for Agentic RL with Large Language Models

DGX agent

arXiv:2603.18859v2 Announce Type: replace Abstract: Reinforcement learning (RL) shows promise for enhancing LLM agentic reasoning, yet sparse terminal rewards hinder fine-grained optimization. Process

agentsarxiv-cs-ai
29 May 2026
Model Releases

RoboWits: Unexpected Challenges for Robotic Creative Problem Solving

DGX agent

arXiv:2605.30326v1 Announce Type: cross Abstract: The ability to reason, adapt, and creatively solve problems under unexpected challenges is essential for robots operating in real-world environments.

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Robust and Efficient Guardrails with Latent Reasoning

DGX agent

arXiv:2605.29068v1 Announce Type: new Abstract: Maintaining the safety of large language models (LLMs) is crucial as they are increasingly deployed in real-world applications. Existing safety guardrai

model-releasesarxiv-cs-ai
29 May 2026
Safety

Robust and Generalizable Safety Steering for Text-to-Image Diffusion Transformers

DGX agent

arXiv:2605.30049v1 Announce Type: new Abstract: Diffusion Transformers have become a powerful backbone for text-to-image generation, but their layered and cross-modal generation process makes safety c

safetyarxiv-cs-ai
29 May 2026
Safety

Rooted Absorbed Prefix Trajectory Balance with Submodular Replay for GFlowNet Training

DGX agent

arXiv:2603.00454v2 Announce Type: replace-cross Abstract: Generative Flow Networks (GFlowNets) enable fine-tuning large language models to approximate reward-proportional posteriors, but they remain p

safetyarxiv-cs-ai
29 May 2026
Safety

Rubric-Guided Process Reward for Stepwise Model Routing

DGX agent

arXiv:2605.29310v1 Announce Type: new Abstract: Stepwise model routing improves the efficiency of Large Reasoning Models (LRMs) by assigning each reasoning step to a suitable model. Recent methods for

safetyarxiv-cs-ai
29 May 2026
Model Releases

S-MARC: Causal Streaming Reasoning for Full-Duplex Conversational Behavior Modeling

DGX agent

arXiv:2602.11065v2 Announce Type: replace-cross Abstract: Human conversation is organized by an implicit chain of thought and manifests as temporally structured conversational behaviors. Capturing thi

model-releasesarxiv-cs-ai
29 May 2026
Local Ai

S3Mem: Structured Spatiotemporal Scene-Event Memory for Long-Horizon Interactive Question Answering

DGX agent

arXiv:2605.28831v1 Announce Type: cross Abstract: Long-horizon interactive agents often accumulate large trajectory histories yet still fail to answer questions about earlier events reliably. We argue

local-aiarxiv-cs-ai
29 May 2026
Model Releases

SAAS: Self-Aware Reinforcement Learning for Over-Search Mitigation in Agentic Search

DGX agent

arXiv:2605.29796v1 Announce Type: new Abstract: Agentic search enables LLMs to solve complex multi-hop questions through iterative reasoning and external search. Despite the effectiveness, these syste

model-releasesarxiv-cs-ai
29 May 2026
Safety

SafeRx-Agent: A Knowledge-Grounded Multi-Agent Framework for Safe and Explainable Medication Recommendation

DGX agent

arXiv:2605.29146v1 Announce Type: cross Abstract: Medication recommendation predicts medications for patient visits, but existing methods still face two key challenges. At the model level, traditional

safetyarxiv-cs-ai
29 May 2026
Model Releases

SafeSearch: Automated Red-Teaming of LLM-Based Search Agents

DGX agent

arXiv:2509.23694v5 Announce Type: replace Abstract: Search agents connect LLMs to the Internet, enabling them to access broader and more up-to-date information. However, this also introduces a new thr

model-releasesarxiv-cs-ai
29 May 2026
Safety

Same Evidence, Different Answers: Canonical-Context On-Policy Distillation for Multi-Turn Language Models

DGX agent

arXiv:2605.30251v1 Announce Type: cross Abstract: Large language models (LLMs) often solve a task when all instructions are given in a single prompt, but fail when the same information is revealed gra

safetyarxiv-cs-ai
29 May 2026
Model Releases

Same Question, Different Source, Different Answer: Auditing Source-Dependence in Medical Multi-Source RAG

DGX agent

arXiv:2605.29084v1 Announce Type: cross Abstract: A retrieval-augmented generation (RAG) system deployed over a multi-author institutional corpus can give a different answer to the same question depen

model-releasesarxiv-cs-ai
29 May 2026
Applications

Scalable RF Simulation in Generative 4D Worlds

DGX agent

arXiv:2508.12176v2 Announce Type: replace-cross Abstract: Radio Frequency (RF) sensing has emerged as a powerful, privacy-preserving alternative to vision-based methods for various perception tasks. H

applicationsarxiv-cs-ai
29 May 2026
Model Releases

Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet

DGX agent

arXiv:2605.29358v1 Announce Type: new Abstract: We demonstrate that sparse autoencoders can extract interpretable features from Claude 3 Sonnet, a production-scale language model, addressing the open

model-releasesarxiv-cs-ai
29 May 2026
Agents

Scaling Small Agents Through Strategy Auctions

DGX agent

arXiv:2602.02751v2 Announce Type: replace-cross Abstract: Small language models are increasingly viewed as a promising, cost-effective approach to agentic AI, with proponents claiming they are suffici

agentsarxiv-cs-ai
29 May 2026
Model Releases

SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers

DGX agent

arXiv:2605.29059v1 Announce Type: cross Abstract: Smart contract decompilation aims to recover high-level source code from bytecode, but evaluating decompilers remains difficult because existing studi

model-releasesarxiv-cs-ai
29 May 2026
Hardware

ScheduleStream: Temporal Planning with Samplers for GPU-Accelerated Multi-Arm Task and Motion Planning & Scheduling

DGX agent

arXiv:2511.04758v2 Announce Type: replace-cross Abstract: Bimanual and humanoid robots are appealing because of their human-like ability to leverage multiple arms to efficiently complete tasks. Howeve

hardwarearxiv-cs-ai
29 May 2026
Agents

SchGen: PCB Schematic Generation with Semantic-Grounded Code Representations

DGX agent

arXiv:2605.30345v1 Announce Type: new Abstract: Printed circuit board (PCB) schematic design defines nearly all electronic hardware, but it remains manual and expertise-intensive. While generative AI

agentsarxiv-cs-ai
29 May 2026
Model Releases

SciIntBench: Measuring LLM Compliance with Research Integrity Norms Under Adversarial Framing

DGX agent

arXiv:2605.29468v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to support scientific work, but it is unclear whether they uphold responsible conduct of research (

model-releasesarxiv-cs-ai
29 May 2026
Research

SCoOP: Semantic Consistent Opinion Pooling for Uncertainty Quantification in Multiple Vision-Language Model Systems

DGX agent

arXiv:2603.23853v3 Announce Type: replace Abstract: Combining multiple Vision-Language Models (VLMs) can enhance multimodal reasoning and robustness, but aggregating heterogeneous models' outputs ampl

researcharxiv-cs-ai
29 May 2026
Research

SCOPE: A Lightweight-training LLM Framework for Air Traffic Control Readback Monitoring

DGX agent

arXiv:2605.29543v1 Announce Type: cross Abstract: Pilot readback of Air Traffic Control (ATC) voice instructions is a primary safeguard against miscommunication in air transportation. However, readbac

researcharxiv-cs-ai
29 May 2026
Model Releases

SCOPE: Prompt Evolution for Enhancing Agent Effectiveness

DGX agent

arXiv:2512.15374v2 Announce Type: replace Abstract: Large Language Model (LLM) agents are increasingly deployed in environments that generate massive, dynamic contexts. However, a critical bottleneck

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Selection Hyper-heuristics Can Automatically Adjust the Learning Period to Optimally Solve Pseudo-Boolean Problems

DGX agent

arXiv:2605.29916v1 Announce Type: cross Abstract: The Random Gradient hyper-heuristic was recently shown to be able to learn the optimal neighbourhood size when optimizing the LeadingOnes benchmark vi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Selective QA over Conflicting Multi-Source Personal Memory: A Diagnostic Testbed and Method Comparison

DGX agent

arXiv:2605.30087v1 Announce Type: new Abstract: Emerging personal AI agents are moving toward persistent, multi-source memory. This creates an evaluation problem: systems must decide how to use confli

model-releasesarxiv-cs-ai
29 May 2026
Safety

Self-Play Reinforcement Learning under Imperfect Information in Big 2

DGX agent

arXiv:2605.28863v1 Announce Type: cross Abstract: Imperfect-information multiplayer games test whether agents can act under hidden information, sparse rewards, and non-stationary opponents. We study t

safetyarxiv-cs-ai
29 May 2026
Research

Self-Trained Verification for Training- and Test-Time Self-Improvement

DGX agent

arXiv:2605.30290v1 Announce Type: cross Abstract: Self-improvement at scale has been a longstanding goal for reasoning models, and there are two natural places to do it: at test time, through verifica

researcharxiv-cs-ai
29 May 2026
Model Releases

Semantic and Visual Evidence for Efficient Long-Video Reasoning: A Solution for the HD-EPIC VQA Challenge

DGX agent

arXiv:2605.29402v1 Announce Type: cross Abstract: Understanding long-form egocentric videos remains challenging for multimodal large language models (MLLMs) due to limited context length and insuffici

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

SERC: LDPC-Inspired Semantic Error Correction for Retrieval-Augmented Generation

DGX agent

arXiv:2605.28837v1 Announce Type: cross Abstract: While Large Language Models (LLMs) have demonstrated remarkable capabilities, their reliability is significantly compromised by hallucinations. Existi

model-releasesarxiv-cs-ai
29 May 2026
Research

Singularity-aware Optimization via Randomized Geometric Probing: Towards Stable Non-smooth Optimization

DGX agent

arXiv:2605.29547v1 Announce Type: cross Abstract: Deep learning optimization relies heavily on the assumption of smooth loss landscapes, a condition systematically violated by modern architectures due

researcharxiv-cs-ai
29 May 2026
Model Releases

Skill-Pro: Learning Reusable Skills from Experience via Non-Parametric PPO for LLM Agents

DGX agent

arXiv:2602.01869v3 Announce Type: replace Abstract: LLM-driven agents excel at sequential decision-making but often rely on on-the-fly reasoning, re-deriving solutions even in recurring scenarios. Thi

model-releasesarxiv-cs-ai
29 May 2026
Agents

SkillBrew: Multi-Objective Curation of Skill Banks for LLM Agents

DGX agent

arXiv:2605.29440v1 Announce Type: cross Abstract: Retrieval-augmented LLM agents increasingly rely on curated skill banks: collections of reusable textual principles that guide decision making on comp

agentsarxiv-cs-ai
29 May 2026
Model Releases

SkillsInjector: Dynamic Skill Context Construction for LLM Agents

DGX agent

arXiv:2605.29794v1 Announce Type: new Abstract: LLM agents now draw on growing skill libraries to handle complex tasks. However, injecting more skills does not always improve task completion and can e

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Small Agent Group is the Future of Digital Health

DGX agent

arXiv:2602.08013v2 Announce Type: replace Abstract: The rapid adoption of large language models (LLMs) in digital health has been driven by a 'scaling-first' philosophy, i.e., the assumption that clin

model-releasesarxiv-cs-ai
29 May 2026
Research

Source-Grounded Semantic Reinforcement Learning for Low-Resource Target-Language Generation

DGX agent

arXiv:2605.29502v1 Announce Type: cross Abstract: Low-resource target-language generation is often limited by scarce parallel data, while high-resource source-language monolingual data is abundant but

researcharxiv-cs-ai
29 May 2026
Applications

Specialty-Specific Medical Language Model for Immune-Mediated Diseases

DGX agent

arXiv:2605.28838v1 Announce Type: cross Abstract: Extracting detailed clinical information from free-text medical narratives remains a practical challenge for researchers and healthcare systems. Termi

applicationsarxiv-cs-ai
29 May 2026
Local Ai

Steering at the Source: Style Modulation Heads for Robust Persona Control

DGX agent

arXiv:2603.13249v2 Announce Type: replace-cross Abstract: Activation steering offers a computationally efficient mechanism for controlling Large Language Models (LLMs) without fine-tuning. While effec

local-aiarxiv-cs-ai
29 May 2026
Research

Steering Language Models Before They Speak: Logit-Level Interventions

DGX agent

arXiv:2601.10960v2 Announce Type: replace-cross Abstract: Controllable generation requires language models to realize output characteristics such as reading level, politeness, and toxicity. Existing s

researcharxiv-cs-ai
29 May 2026
Research

Stochastic Lifting for Generating Trajectories of Stochastic Physical Systems

DGX agent

arXiv:2605.29194v1 Announce Type: cross Abstract: Many stochastic physical systems evolve smoothly over time in the sense that the distribution of states changes regularly across time steps. The trans

researcharxiv-cs-ai
29 May 2026
Research

Structured Prompt Optimization Meets Reinforcement Learning for Global and Local Interpretability over Complex Text

DGX agent

arXiv:2605.29076v1 Announce Type: cross Abstract: LLMs have advanced text classification, yet existing paradigms face a trade-off: supervised (label only) fine-tuning is scalable but offers limited re

researcharxiv-cs-ai
29 May 2026
← Previous
1…242243244245246…448
Next →