AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,271
  • Agents7,543
  • Applications5,408
  • Concepts5
  • Hardware1,828
  • Industry6,162
  • Local Ai4,927
  • Model Releases23,818
  • Research20,122
  • Safety13,367
  • Syntheses17
  • Tools1,674
  • Tutorials3,400

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,271
  • Agents7,543
  • Applications5,408
  • Concepts5
  • Hardware1,828
  • Industry6,162
  • Local Ai4,927
  • Model Releases23,818
  • Research20,122
  • Safety13,367
  • Syntheses17
  • Tools1,674
  • Tutorials3,400

Source
Human
88,271Total entries
1Added by human
88,270Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
62,897 results
29 May 2026

GrowLoop: Self-Evolving Conversation Evaluation Seeded by Human

Model ReleasesDGX agent

arXiv:2605.28882v1 Announce Type: cross Abstract: With the rapid advancement of large language models, evaluating human-likeness in open-ended conversation has become increasingly important. However,

GRPO is Secretly a Process Reward Model

SafetyDGX agent

arXiv:2509.21154v4 Announce Type: replace-cross Abstract: Process reward models (PRMs) allow for fine-grained credit assignment in reinforcement learning (RL), and seemingly contrast with outcome rewa

GRUFF: LLM Pronoun Fidelity, Reasoning, and Biases in German

SafetyDGX agent

arXiv:2605.30214v1 Announce Type: new Abstract: Third-person singular pronouns have long been used to study stereotypical biases in language models and to test their abilities to reason about referenc

DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

GTA: Generating Long-Horizon Tasks for Web Agents at Scale

Model ReleasesDGX agent

arXiv:2605.29218v1 Announce Type: new Abstract: Web agents, which couple language models with browsing and tool-use capabilities, show promise as open web assistants. Yet progress is increasingly limi

Guidance Contrastive Token Credit Assignment for Discrete Policy Optimization

SafetyDGX agent

arXiv:2605.29198v1 Announce Type: new Abstract: Group-advantage-based reinforcement learning methods, such as GRPO and DAPO, have demonstrated strong performance across diverse domains, including math

GUITestScape: Towards Open-set Evaluation on Exploratory GUI Testing

Model ReleasesDGX agent

arXiv:2605.29532v1 Announce Type: cross Abstract: Exploratory GUI testing is a particularly demanding setting for MLLM agents: without predefined test scripts, an agent must autonomously navigate an a

Hallucination Detection-Guided Preference Optimization for Clinical Summarization

Model ReleasesDGX agent

arXiv:2605.28910v1 Announce Type: cross Abstract: Large language models (LLMs) have shown promise on summarization tasks, but they often produce hallucinations, which are unsupported or incorrect stat

Hallucination Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic Caching

Model ReleasesDGX agent

arXiv:2605.29055v1 Announce Type: new Abstract: Hallucination remains a major reliability barrier for production LLM systems, particularly in multi-agent pipelines where unsupported claims can propaga

HaluNet: Learning Hallucination Risk from Internal Signals in LLM Question Answering

ResearchDGX agent

arXiv:2512.24562v2 Announce Type: replace Abstract: Large language models (LLMs) achieve strong question answering (QA) performance but can produce fluent answers unsupported by available evidence. Ex

Harmless Yet Harmful: Neutral Prompting Attacks for Stealthy Hallucination Steering in Agent Skills

AgentsDGX agent

arXiv:2605.29354v1 Announce Type: cross Abstract: LLM-powered coding agents increasingly participate in software development workflows by generating code, selecting dependencies, and producing package

Harmonizing Real-Time Constraints and Long-Horizon Reasoning: An Asynchronous Agentic Framework for Dynamic Scheduling

SafetyDGX agent

arXiv:2605.29262v1 Announce Type: new Abstract: The Dynamic Flexible Job Shop Scheduling Problem (DFJSP) necessitates a trade-off between instant reaction to stochastic disturbances and global optimiz

Harnessing non-adversarial robustness in large language models

SafetyDGX agent

arXiv:2605.29816v1 Announce Type: new Abstract: The work presents an approach for addressing the challenge of robustness in Large Language Models (LLMs) to alterations and potential errors caused by s

HARP: Hadamard-Preconditioned Adaptive Rotation Processor for Extreme LLM Quantization

ResearchDGX agent

arXiv:2605.29843v1 Announce Type: cross Abstract: Post-training quantization (PTQ) is essential for deploying LLMs under memory and bandwidth constraints. However, extreme low-bit quantization remains

HD-Prot: A Protein Language Model for Joint Sequence-Structure Modeling with Continuous Structure Tokens

TutorialsDGX agent

arXiv:2512.15133v2 Announce Type: replace-cross Abstract: Proteins inherently possess a consistent sequence-structure duality. The abundance of protein sequence data, which can be readily represented

HEART-Bench: Do LLM Agents Exhibit Human-like Psychology?

Model ReleasesDGX agent

arXiv:2605.30058v1 Announce Type: new Abstract: While LLM agents have demonstrated remarkable task-oriented abilities such as planning, reasoning, and action, few works have treated them as complete h

Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction

AgentsDGX agent

arXiv:2605.29960v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly leverage long term memory to support persistent and autonomous task execution. However, this capability

HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question Answering

ResearchDGX agent

arXiv:2605.29606v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) for document-based Open-domain Question Answering (ODQA) on large-scale industrial corpora faces two critical bottl

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.29782v1 Announce Type: cross Abstract: Reinforcement learning (RL) refines large language models (LLMs) by directly optimizing model behavior through reward signals. While accurate state va

HM-Talker: Hybrid Motion Modeling for High-Fidelity Talking Head Synthesis

ResearchDGX agent

arXiv:2508.10566v3 Announce Type: replace Abstract: Audio-driven talking head generation faces a fundamental trade-off between personalization and generalization, limiting its practical application. I

HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

ResearchDGX agent

arXiv:2605.29948v1 Announce Type: cross Abstract: Unified speech foundation models require a holistic tokenization space that is both learnable by language models and decodable into high-quality wavef

Honest Lying: Understanding Memory Confabulation in Reflexive Agents

ResearchDGX agent

arXiv:2605.29463v1 Announce Type: cross Abstract: Reflexion-style agents rely on self-generated reflections as memory, implicitly assuming that agents can accurately diagnose their own failures.We sho

Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots

AgentsDGX agent

arXiv:2605.29963v1 Announce Type: cross Abstract: Honeypots are decoy systems mimicking real system components designed to defend against cyber attacks. Recently, LLMs increasingly serve as simulation

Horizon Activation Mapping for Neural Networks in Time Series Forecasting

ResearchDGX agent

arXiv:2601.02094v4 Announce Type: replace Abstract: Neural networks for time series forecasting have relied on error metrics and architecture-specific interpretability approaches for model selection t

How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions

Model ReleasesDGX agent

arXiv:2605.29442v1 Announce Type: cross Abstract: AI coding agents increasingly act directly within software environments, yet existing analyses of their failures rely on benchmark trajectories that m

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines

AgentsDGX agent

arXiv:2605.28840v1 Announce Type: cross Abstract: Large language model (LLM) agents with tool-calling capabilities are increasingly deployed in production systems, yet a fundamental reliability questi

How Far Ahead Do LLMs Plan? Uncovering the Latent Horizon in Chain-of-Thought Reasoning

Model ReleasesDGX agent

arXiv:2602.02103v2 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) reasoning has become a central mechanism for eliciting multi-step reasoning in Large Language Models (LLMs). Yet recent

How LoRA Remembers? A Parametric Memory Law for LLM Finetuning

Model ReleasesDGX agent

arXiv:2605.30260v1 Announce Type: cross Abstract: Large Language Models (LLMs) must continuously learn and update knowledge to remain effective in dynamic real-world environments. While Low-Rank Adapt

How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions

ResearchDGX agent

arXiv:2605.29448v1 Announce Type: cross Abstract: Neural scaling laws appraise data through dataset size, while the Vendi Score uses quantum entropy to measure dataset value. We show both that common

How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency

Model ReleasesDGX agent

arXiv:2605.30096v1 Announce Type: cross Abstract: Large language models (LLMs) can autonomously conduct multi-stage cyber attacks, but the consistency of their offensive behavior under repeated trials

How to Relieve Distribution Shifts in Semantic Segmentation for Off-Road Environments

AgentsDGX agent

arXiv:2605.29599v1 Announce Type: cross Abstract: Semantic segmentation is crucial for autonomous navigation in off-road environments, enabling precise classification of surroundings to identify trave

How's it going? Reinforcement learning in language models recruits a functional welfare axis

SafetyDGX agent

arXiv:2605.30232v1 Announce Type: cross Abstract: How does reinforcement learning shape a language model's internal representations? We present evidence that RL recruits a pre-existing representation

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime

SafetyDGX agent

arXiv:2605.30201v1 Announce Type: cross Abstract: We investigate a narrow but common failure mode of GRPO-style reinforcement learning in the context of sparse verifiable rewards: early updates contai

HTAM: Hierarchical Transition-Attended Memory for Operator Optimization

Local AiDGX agent

arXiv:2605.29734v1 Announce Type: new Abstract: High-performance GPU kernels are essential for efficient LLM deployment, yet optimizing them remains expertise-intensive. Recent LLM-based code generati

Human-in-the-Loop Swarms: A Bionic Swarm Approach to Real-World Soil Mapping

AgentsDGX agent

arXiv:2605.29091v1 Announce Type: new Abstract: Swarm and field robotics face significant barriers to real-world validation due to the high cost and development time to deploy hardware. This paper int

iLoRA: Bayesian Low-Rank Adaptation with Latent Interaction Graphs for Microbiome Diagnosis

Model ReleasesDGX agent

arXiv:2605.30179v1 Announce Type: cross Abstract: Parameter-efficient adaptation has made LLMs practical for domain prediction, but standard LoRA still relies on a static low-rank update and does not

Implicit Identity Technologies for LLMs: Fingerprinting and Watermarking across Datasets, Models, and Generated Content

ResearchDGX agent

arXiv:2605.29245v1 Announce Type: cross Abstract: This paper presents a survey and taxonomy of LLM fingerprinting and watermarking for identity, ownership verification, provenance, and generated-conte

Improved Guarantees for Heterogeneous Treatment-Effect Estimation via Matrix Completion

ResearchDGX agent

arXiv:2605.30319v1 Announce Type: cross Abstract: A central goal of modern causal inference is estimating heterogeneous treatment effects to answer questions like 'how does an intervention affect each

Improving Adversarial Robustness of Attribution via Implicit Regularization

Model ReleasesDGX agent

arXiv:2605.29983v1 Announce Type: cross Abstract: The adversarial robustness of attributions is a fundamental requirement for reliable explainability in deep learning, yet existing approaches typicall

Improving CLIP Adaptation by Breaking Tail Alignment for Source-Free Cross-Domain Few-Shot Learning

SafetyDGX agent

arXiv:2605.29776v1 Announce Type: new Abstract: Vision-Language Models (VLMs) such as CLIP demonstrate strong zero-shot generalization, but their performance significantly degrades in cross-domain sce

Improving Collaborative Storytelling with a Multi-Agent Framework Based on Large Language Models

AgentsDGX agent

arXiv:2605.29625v1 Announce Type: new Abstract: The topic of Co-creation, i.e., AI agents interacting with humans to generate outputs (e.g., art), has gained significant attention recently. However, m

Improving Full Waveform Inversion in Large Model Era

Model ReleasesDGX agent

arXiv:2603.00377v2 Announce Type: replace Abstract: Full Waveform Inversion (FWI) is a highly nonlinear and ill-posed problem that aims to recover subsurface velocity maps from surface-recorded seismi

In-Context Reward Adaptation for Robust Preference Modeling

SafetyDGX agent

arXiv:2605.30323v1 Announce Type: cross Abstract: Reinforcement Learning from Human Feedback (RLHF) typically relies on static reward models to align Large Language Models with human preferences. Howe

In-Place Feedback: Reliable Refinement for Multi-Turn Expert-LLM Collaboration

ResearchDGX agent

arXiv:2510.00777v2 Announce Type: replace Abstract: LLM-generated drafts often contain subtle factual or logical errors, yet prior work shows that models struggle to reliably integrate multi-turn feed

Indexing the Unreadable: LLM-Native Recursive Construction and Search of Service Taxonomies

Model ReleasesDGX agent

arXiv:2605.29270v1 Announce Type: new Abstract: The era of the Internet of Agents (IoA) is taking shape: LLM agents are expected to fulfill user goals by orchestrating fast-growing populations of Mode

Inferring Code Correctness from Specification

SafetyDGX agent

arXiv:2605.29822v1 Announce Type: cross Abstract: Large language models (LLMs) have become integral to modern software development, enabling automated code generation at scale. However, validating the

Inferring the Size of Large Language Models From Popular Text Memorization

Model ReleasesDGX agent

arXiv:2605.29223v1 Announce Type: new Abstract: The parameter counts of the most widely used large language models (LLMs) are often withheld by their developers, leaving model size -- a primary refere

Influence-Guided Symbolic Regression: Scientific Discovery via LLM-Driven Equation Search with Granular Feedback

TutorialsDGX agent

arXiv:2605.29184v1 Announce Type: cross Abstract: Large Language Models (LLMs) offer a promising avenue for scientific discovery, yet their application to symbolic regression is often constrained by i

Inform, Coach, Relate, Listen: Auditing LLM Caregiving Support Roles

Model ReleasesDGX agent

arXiv:2605.29473v1 Announce Type: cross Abstract: Language models are increasingly being deployed for conversational support in informal caregiving contexts, where interactions often extend beyond inf

Information-Directed Offline-to-Online Reinforcement Learning

SafetyDGX agent

arXiv:2605.29405v1 Announce Type: new Abstract: Decision-making from offline datasets typically warm-starts a policy or score model from fixed offline data and then refines it with limited online inte

InsightEval: An Expert-Curated Benchmark for Assessing Insight Discovery in LLM-Driven Data Agents

Model ReleasesDGX agent

arXiv:2511.22884v2 Announce Type: replace Abstract: Data analysis has become an indispensable part of scientific research. To discover the latent knowledge and insights hidden within massive datasets,

Inspectorch: Efficient rare event exploration in solar observations

ResearchDGX agent

arXiv:2602.20316v2 Announce Type: replace-cross Abstract: The Sun is observed in unprecedented detail, enabling studies of its activity on very small spatiotemporal scales. However, the large volume o

Instance-dependent Stochastic Lipschitz bandit

ResearchDGX agent

arXiv:2605.29748v1 Announce Type: cross Abstract: We study the Lipschitz bandit problem, where a learner sequentially maximizes an unknown Lipschitz function f over a domain X subset [0,1]^d using noi

Intent-aligned Autonomous Spacecraft Guidance via Reasoning Models

SafetyDGX agent

arXiv:2604.17176v2 Announce Type: replace-cross Abstract: Future spacecraft operations require autonomy that can interpret high-level mission intent while preserving safety. However, existing trajecto

Interactive In-Meeting Speaker Correction with Human Feedback

ResearchDGX agent

arXiv:2509.18377v2 Announce Type: replace Abstract: Most automatic speech processing systems operate in ``open loop'' mode without user feedback about who said what, yet human-in-the-loop workflows ca

Internal Representation, Not Clinical Knowledge: Where Apparent LLM Triage Failures Originate

Model ReleasesDGX agent

arXiv:2605.29889v1 Announce Type: cross Abstract: Patient-voiced clinical-triage benchmarks report high under-triage rates for consumer LLMs for constrained multiple-choice output, yet the same cases

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation

ResearchDGX agent

arXiv:2605.30230v1 Announce Type: new Abstract: With the rapid advancement of diffusion models, talking face generation has made remarkable progress. However, existing diffusion-based methods still re

Is Your Diffusion Sampler Actually Correct? A Sampler-Centric Evaluation of Discrete Diffusion Language Models

ResearchDGX agent

arXiv:2602.19619v2 Announce Type: replace Abstract: Discrete diffusion language models (dLLMs) provide a fast and flexible alternative to autoregressive models (ARMs) via iterative denoising with para

Is Your LLM Overcharging You? Tokenization, Transparency, and Incentives

Model ReleasesDGX agent

arXiv:2505.21627v4 Announce Type: replace-cross Abstract: State-of-the-art large language models require specialized hardware and substantial energy to operate. As a consequence, cloud-based services

It`s All About Speed: AI`s Impact on Workflow in Music Production

ApplicationsDGX agent

arXiv:2605.29931v1 Announce Type: new Abstract: In this paper, we present the results of an ethnographic study into the impact of AI and automated tools on music production workflow. Focusing specific

Jailbreaking and Mitigation of Vulnerabilities in Large Language Models

SafetyDGX agent

arXiv:2410.15236v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have transformed artificial intelligence by advancing natural language understanding and generation, enabling app

← Previous
1…556557558559560…1049
Next →