AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlog
85,136Total entries
1Added by human
85,135Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,688 results
Model Releases

GrowLoop: Self-Evolving Conversation Evaluation Seeded by Human

DGX agent

arXiv:2605.28882v1 Announce Type: cross Abstract: With the rapid advancement of large language models, evaluating human-likeness in open-ended conversation has become increasingly important. However,

model-releasesarxiv-cs-ai
29 May 2026
Safety

GRPO is Secretly a Process Reward Model

X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

arXiv:2509.21154v4 Announce Type: replace-cross Abstract: Process reward models (PRMs) allow for fine-grained credit assignment in reinforcement learning (RL), and seemingly contrast with outcome rewa

safetyarxiv-cs-ai
29 May 2026
Model Releases

GTA: Generating Long-Horizon Tasks for Web Agents at Scale

DGX agent

arXiv:2605.29218v1 Announce Type: new Abstract: Web agents, which couple language models with browsing and tool-use capabilities, show promise as open web assistants. Yet progress is increasingly limi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

GUITestScape: Towards Open-set Evaluation on Exploratory GUI Testing

DGX agent

arXiv:2605.29532v1 Announce Type: cross Abstract: Exploratory GUI testing is a particularly demanding setting for MLLM agents: without predefined test scripts, an agent must autonomously navigate an a

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Hallucination Detection-Guided Preference Optimization for Clinical Summarization

DGX agent

arXiv:2605.28910v1 Announce Type: cross Abstract: Large language models (LLMs) have shown promise on summarization tasks, but they often produce hallucinations, which are unsupported or incorrect stat

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Hallucination Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic Caching

DGX agent

arXiv:2605.29055v1 Announce Type: new Abstract: Hallucination remains a major reliability barrier for production LLM systems, particularly in multi-agent pipelines where unsupported claims can propaga

model-releasesarxiv-cs-ai
29 May 2026
Safety

Harmonizing Real-Time Constraints and Long-Horizon Reasoning: An Asynchronous Agentic Framework for Dynamic Scheduling

DGX agent

arXiv:2605.29262v1 Announce Type: new Abstract: The Dynamic Flexible Job Shop Scheduling Problem (DFJSP) necessitates a trade-off between instant reaction to stochastic disturbances and global optimiz

safetyarxiv-cs-ai
29 May 2026
Safety

Harnessing non-adversarial robustness in large language models

DGX agent

arXiv:2605.29816v1 Announce Type: new Abstract: The work presents an approach for addressing the challenge of robustness in Large Language Models (LLMs) to alterations and potential errors caused by s

safetyarxiv-cs-ai
29 May 2026
Research

HARP: Hadamard-Preconditioned Adaptive Rotation Processor for Extreme LLM Quantization

DGX agent

arXiv:2605.29843v1 Announce Type: cross Abstract: Post-training quantization (PTQ) is essential for deploying LLMs under memory and bandwidth constraints. However, extreme low-bit quantization remains

researcharxiv-cs-ai
29 May 2026
Tutorials

HD-Prot: A Protein Language Model for Joint Sequence-Structure Modeling with Continuous Structure Tokens

DGX agent

arXiv:2512.15133v2 Announce Type: replace-cross Abstract: Proteins inherently possess a consistent sequence-structure duality. The abundance of protein sequence data, which can be readily represented

tutorialsarxiv-cs-ai
29 May 2026
Agents

Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction

DGX agent

arXiv:2605.29960v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly leverage long term memory to support persistent and autonomous task execution. However, this capability

agentsarxiv-cs-ai
29 May 2026
Research

HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question Answering

DGX agent

arXiv:2605.29606v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) for document-based Open-domain Question Answering (ODQA) on large-scale industrial corpora faces two critical bottl

researcharxiv-cs-ai
29 May 2026
Model Releases

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning

DGX agent

arXiv:2605.29782v1 Announce Type: cross Abstract: Reinforcement learning (RL) refines large language models (LLMs) by directly optimizing model behavior through reward signals. While accurate state va

model-releasesarxiv-cs-ai
29 May 2026
Research

HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

DGX agent

arXiv:2605.29948v1 Announce Type: cross Abstract: Unified speech foundation models require a holistic tokenization space that is both learnable by language models and decodable into high-quality wavef

researcharxiv-cs-ai
29 May 2026
Research

Honest Lying: Understanding Memory Confabulation in Reflexive Agents

DGX agent

arXiv:2605.29463v1 Announce Type: cross Abstract: Reflexion-style agents rely on self-generated reflections as memory, implicitly assuming that agents can accurately diagnose their own failures.We sho

researcharxiv-cs-ai
29 May 2026
Agents

Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots

DGX agent

arXiv:2605.29963v1 Announce Type: cross Abstract: Honeypots are decoy systems mimicking real system components designed to defend against cyber attacks. Recently, LLMs increasingly serve as simulation

agentsarxiv-cs-ai
29 May 2026
Model Releases

How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions

DGX agent

arXiv:2605.29442v1 Announce Type: cross Abstract: AI coding agents increasingly act directly within software environments, yet existing analyses of their failures rely on benchmark trajectories that m

model-releasesarxiv-cs-ai
29 May 2026
Agents

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines

DGX agent

arXiv:2605.28840v1 Announce Type: cross Abstract: Large language model (LLM) agents with tool-calling capabilities are increasingly deployed in production systems, yet a fundamental reliability questi

agentsarxiv-cs-ai
29 May 2026
Model Releases

How LoRA Remembers? A Parametric Memory Law for LLM Finetuning

DGX agent

arXiv:2605.30260v1 Announce Type: cross Abstract: Large Language Models (LLMs) must continuously learn and update knowledge to remain effective in dynamic real-world environments. While Low-Rank Adapt

model-releasesarxiv-cs-ai
29 May 2026
Research

How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions

DGX agent

arXiv:2605.29448v1 Announce Type: cross Abstract: Neural scaling laws appraise data through dataset size, while the Vendi Score uses quantum entropy to measure dataset value. We show both that common

researcharxiv-cs-ai
29 May 2026
Model Releases

How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency

DGX agent

arXiv:2605.30096v1 Announce Type: cross Abstract: Large language models (LLMs) can autonomously conduct multi-stage cyber attacks, but the consistency of their offensive behavior under repeated trials

model-releasesarxiv-cs-ai
29 May 2026
Safety

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime

DGX agent

arXiv:2605.30201v1 Announce Type: cross Abstract: We investigate a narrow but common failure mode of GRPO-style reinforcement learning in the context of sparse verifiable rewards: early updates contai

safetyarxiv-cs-ai
29 May 2026
Model Releases

iLoRA: Bayesian Low-Rank Adaptation with Latent Interaction Graphs for Microbiome Diagnosis

DGX agent

arXiv:2605.30179v1 Announce Type: cross Abstract: Parameter-efficient adaptation has made LLMs practical for domain prediction, but standard LoRA still relies on a static low-rank update and does not

model-releasesarxiv-cs-ai
29 May 2026
Research

Improved Guarantees for Heterogeneous Treatment-Effect Estimation via Matrix Completion

DGX agent

arXiv:2605.30319v1 Announce Type: cross Abstract: A central goal of modern causal inference is estimating heterogeneous treatment effects to answer questions like 'how does an intervention affect each

researcharxiv-cs-ai
29 May 2026
Agents

Improving Collaborative Storytelling with a Multi-Agent Framework Based on Large Language Models

DGX agent

arXiv:2605.29625v1 Announce Type: new Abstract: The topic of Co-creation, i.e., AI agents interacting with humans to generate outputs (e.g., art), has gained significant attention recently. However, m

agentsarxiv-cs-ai
29 May 2026
Safety

In-Context Reward Adaptation for Robust Preference Modeling

DGX agent

arXiv:2605.30323v1 Announce Type: cross Abstract: Reinforcement Learning from Human Feedback (RLHF) typically relies on static reward models to align Large Language Models with human preferences. Howe

safetyarxiv-cs-ai
29 May 2026
Model Releases

Indexing the Unreadable: LLM-Native Recursive Construction and Search of Service Taxonomies

DGX agent

arXiv:2605.29270v1 Announce Type: new Abstract: The era of the Internet of Agents (IoA) is taking shape: LLM agents are expected to fulfill user goals by orchestrating fast-growing populations of Mode

model-releasesarxiv-cs-ai
29 May 2026
Safety

Inferring Code Correctness from Specification

DGX agent

arXiv:2605.29822v1 Announce Type: cross Abstract: Large language models (LLMs) have become integral to modern software development, enabling automated code generation at scale. However, validating the

safetyarxiv-cs-ai
29 May 2026
Tutorials

Influence-Guided Symbolic Regression: Scientific Discovery via LLM-Driven Equation Search with Granular Feedback

DGX agent

arXiv:2605.29184v1 Announce Type: cross Abstract: Large Language Models (LLMs) offer a promising avenue for scientific discovery, yet their application to symbolic regression is often constrained by i

tutorialsarxiv-cs-ai
29 May 2026
Model Releases

Inform, Coach, Relate, Listen: Auditing LLM Caregiving Support Roles

DGX agent

arXiv:2605.29473v1 Announce Type: cross Abstract: Language models are increasingly being deployed for conversational support in informal caregiving contexts, where interactions often extend beyond inf

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

InsightEval: An Expert-Curated Benchmark for Assessing Insight Discovery in LLM-Driven Data Agents

DGX agent

arXiv:2511.22884v2 Announce Type: replace Abstract: Data analysis has become an indispensable part of scientific research. To discover the latent knowledge and insights hidden within massive datasets,

model-releasesarxiv-cs-ai
29 May 2026
Safety

Intent-aligned Autonomous Spacecraft Guidance via Reasoning Models

DGX agent

arXiv:2604.17176v2 Announce Type: replace-cross Abstract: Future spacecraft operations require autonomy that can interpret high-level mission intent while preserving safety. However, existing trajecto

safetyarxiv-cs-ai
29 May 2026
Model Releases

Internal Representation, Not Clinical Knowledge: Where Apparent LLM Triage Failures Originate

DGX agent

arXiv:2605.29889v1 Announce Type: cross Abstract: Patient-voiced clinical-triage benchmarks report high under-triage rates for consumer LLMs for constrained multiple-choice output, yet the same cases

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Is Your LLM Overcharging You? Tokenization, Transparency, and Incentives

DGX agent

arXiv:2505.21627v4 Announce Type: replace-cross Abstract: State-of-the-art large language models require specialized hardware and substantial energy to operate. As a consequence, cloud-based services

model-releasesarxiv-cs-ai
29 May 2026
Applications

It`s All About Speed: AI`s Impact on Workflow in Music Production

DGX agent

arXiv:2605.29931v1 Announce Type: new Abstract: In this paper, we present the results of an ethnographic study into the impact of AI and automated tools on music production workflow. Focusing specific

applicationsarxiv-cs-ai
29 May 2026
Safety

Jailbreaking and Mitigation of Vulnerabilities in Large Language Models

DGX agent

arXiv:2410.15236v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have transformed artificial intelligence by advancing natural language understanding and generation, enabling app

safetyarxiv-cs-ai
29 May 2026
Agents

KairosAgent: Agentic Time Series Forecasting with Fused Semantic Reasoning

DGX agent

arXiv:2605.30002v1 Announce Type: new Abstract: Cross-domain multimodal time series forecasting is a challenging task, requiring models to integrate precise numerical comprehension, cross-domain seman

agentsarxiv-cs-ai
29 May 2026
Model Releases

KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing

DGX agent

arXiv:2605.29524v1 Announce Type: cross Abstract: Relay and reseller APIs increasingly intermediate access to large language models (LLMs), but users have no direct way to verify that a claimed endpoi

model-releasesarxiv-cs-ai
29 May 2026
Research

KLAS: Using Similarity to Stitch Neural Networks for Improved Accuracy-Efficiency Tradeoffs

DGX agent

arXiv:2605.29259v1 Announce Type: cross Abstract: Given the wide range of deployment targets, flexible model selection is essential for optimizing performance within a given compute budget. Recent wor

researcharxiv-cs-ai
29 May 2026
Model Releases

Label-Free Reinforcement Learning via Cross-Model Entropy

DGX agent

arXiv:2605.29009v1 Announce Type: cross Abstract: Post-training large language models with reinforcement learning is bottlenecked by the reward signal. Existing approaches require either ground-truth

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Label Over Logic? How Source Cues Bias Human Fallacy Judgments More Than LLMs

DGX agent

arXiv:2605.29928v1 Announce Type: cross Abstract: As AI-generated and AI-assisted content floods online spaces, source labels attached to such content can distort human reasoning judgments, with downs

model-releasesarxiv-cs-ai
29 May 2026
Research

LaRA: Layer-wise Representation Analysis for Detecting Data Contamination in RL Post-Training

DGX agent

arXiv:2605.29888v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training has shown to improve reasoning in large language models (LLMs). However, there has been little exploration o

researcharxiv-cs-ai
29 May 2026
Research

Large-Scale AI and Foundation Models for Neuroscience: A Comprehensive Review

DGX agent

arXiv:2510.16658v3 Announce Type: replace Abstract: The development of large-scale artificial intelligence (AI) models is influencing neuroscience research by enabling end-to-end learning from raw bra

researcharxiv-cs-ai
29 May 2026
Tutorials

Latent Terms: Dense Retrievers Contain Trivially Extractable BM25-ready Zipfian Vocabularies

DGX agent

arXiv:2605.29384v1 Announce Type: cross Abstract: We propose Latent Terms, a method revealing that models trained for dense retrieval, whether single- or multi-vector, learn representations that can t

tutorialsarxiv-cs-ai
29 May 2026
Tutorials

Learn from A Rationalist: Distilling Intermediate Interpretable Rationales

DGX agent

arXiv:2601.22531v2 Announce Type: replace-cross Abstract: Because of the pervasive use of deep neural networks (DNNs), especially in high-stakes domains, the interpretability of DNNs has received incr

tutorialsarxiv-cs-ai
29 May 2026
Safety

Learning A Simulation-based Visual Policy for Real-world Peg In Unseen Holes

DGX agent

arXiv:2205.04297v2 Announce Type: replace-cross Abstract: This paper proposes a learning-based visual peg-in-hole that enables training with several shapes in simulation, and adapting to arbitrary uns

safetyarxiv-cs-ai
29 May 2026
Research

Learning Context-Conditioned Predicate Semantics via Prototype Feedback

DGX agent

arXiv:2605.29610v1 Announce Type: cross Abstract: In scene graph generation, a central challenge is modeling polysemous predicates whose meanings shift across contexts. Prior approaches address this i

researcharxiv-cs-ai
29 May 2026
Safety

Learning to Choose: An Empowerment-Guided Multi-Agent System with semantic communication for Adaptive Method Selection

DGX agent

arXiv:2605.30042v1 Announce Type: new Abstract: Automating scientific computing workflows requires more than generating executable code: autonomous systems must also select appropriate computational s

safetyarxiv-cs-ai
29 May 2026
← Previous
1…243244245246247…452
Next →