AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
29 May 2026

Evolutionary Dynamics of Cooperation in Next-Generation LLM Agent Systems: A Cross-Provider Empirical Extension

Model ReleasesDGX agent

arXiv:2605.29874v1 Announce Type: cross Abstract: Do next-generation LLM agents inherit the cooperative biases documented in their predecessors, or does scale and provider diversity reshape equilibriu

Evolutionary Refinement of Generative Graph Topologies: A Hybrid WGAN-GA Approach

SafetyDGX agent

arXiv:2605.29161v1 Announce Type: cross Abstract: Generating realistic graph-structured data is challenging due to discrete connectivity, varying graph sizes, and class-specific structural patterns. R

Evolutionary Rule Extraction from Corporate Default Prediction Models

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.29478v1 Announce Type: cross Abstract: Small and medium-sized enterprises (SMEs) represent the majority of firms in most economies and often face financial constraints and higher vulnerabil

Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems

AgentsDGX agent

arXiv:2605.29790v1 Announce Type: cross Abstract: LLM-based multi-agent systems (MAS) have emerged as an effective paradigm for complex and long-horizon tasks. However, in real-world tasks, MAS often

Evolving Features vs Evolving Entire Trees with GP for Interpretable Survival Analysis

ApplicationsDGX agent

arXiv:2605.30119v1 Announce Type: cross Abstract: Survival analysis concerns the task of predicting the time until an event occurs. Often used in the medical field, survival analysis deals with incomp

EvoMD-LLM: Learning the Language of Species Evolution in Reactive Molecular Dynamics

SafetyDGX agent

arXiv:2605.29394v1 Announce Type: new Abstract: While large language models (LLMs) excel at static scientific reasoning, they struggle to model the temporal structure of dynamic physical processes. We

Extreme dynamic symmetry enables omnidirectional and multifunctional robots

ResearchDGX agent

arXiv:2605.29254v1 Announce Type: cross Abstract: Symmetry is a central organizing principle in natural systems, yet its use as a unifying design strategy in robotics has largely remained limited to g

FHRFormer: A Self-Supervised Masked Transformer Framework for Fetal Heart Rate Time-Series Inpainting and Forecasting

Local AiDGX agent

arXiv:2605.29695v1 Announce Type: new Abstract: Approximately 10% of newborns require assistance to initiate breathing at birth, and around 5% need ventilation support. Fetal heart rate (FHR) monitori

Finding DoRI: Discovery of Retained Images in Diffusion Models

Local AiDGX agent

arXiv:2507.16880v3 Announce Type: replace-cross Abstract: Text-to-image diffusion models (DMs) have achieved remarkable success in image generation. However, concerns about data privacy and intellectu

FinVerBench: Benchmark Validity and Calibration in Large Language Model Financial Statement Verification

Model ReleasesDGX agent

arXiv:2605.29586v1 Announce Type: new Abstract: We introduce FinVerBench, a benchmark and validity study for financial statement verification: determining whether a set of corporate financial statemen

First head-to-head comparison of agentic AI applied to the analysis of simulated data of the Einstein Telescope

Model ReleasesDGX agent

arXiv:2605.28916v1 Announce Type: cross Abstract: We report a comparison of two state-of-the-art agentic AI systems, Claude Code (Anthropic) and Codex (OpenAI), tasked with autonomously executing a si

Forget Less, Generalize More: Unifying Temporal and Structural Adaptation for Dynamic Graphs

ApplicationsDGX agent

arXiv:2605.29453v1 Announce Type: cross Abstract: Representation learning on dynamic graphs requires capturing complex dependencies that evolve across both time and structure. Existing approaches typi

FormalEvolve: Neuro-Symbolic Evolutionary Search for Diverse Autoformalization

ResearchDGX agent

arXiv:2603.19828v3 Announce Type: replace Abstract: Autoformalization aims to produce formal statements that compile and faithfully preserve the intended meaning of informal mathematics. Yet standard

Formalizing Mathematics at Scale

AgentsDGX agent

arXiv:2605.29955v1 Announce Type: new Abstract: We present AutoformBot, a multi-agent system for building an Autoformalized Textbook Library At Scale (Atlas) in Lean 4. AutoformBot orchestrates thousa

FormInv: A Measurement Protocol for Semantic Invariance in Mathematical Reasoning Benchmarks

Model ReleasesDGX agent

arXiv:2605.29001v1 Announce Type: cross Abstract: A paraphrase-quality audit of MathCheck (ICLR 2025) detected 4 semantically incorrect paraphrases in 129 groups (3.1%); removing them drops GPT-4o fro

From GPS Points to Travel Patterns: Flexible and Semantic Trajectory Generation with LLMs

ApplicationsDGX agent

arXiv:2605.30014v1 Announce Type: new Abstract: Urban trajectories play a crucial role in modeling urban dynamics and supporting various smart city applications. However, privacy concerns restrict acc

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning

ResearchDGX agent

arXiv:2601.21909v2 Announce Type: replace Abstract: Current LLM post-training methods optimize complete reasoning trajectories through Supervised Fine-Tuning (SFT) followed by outcome-based Reinforcem

From Prompts to Context: An Ontology-Driven Framework for Human-Generative AI Collaboration

AgentsDGX agent

arXiv:2605.29675v1 Announce Type: cross Abstract: Collaborations with Generative AI often begin with a short prompt and end with an opaque output, leaving implicit who was involved, what task was bein

From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges

ResearchDGX agent

arXiv:2601.08654v2 Announce Type: replace-cross Abstract: Rubric-based text evaluation increasingly uses large language models (LLMs) as scalable judges, but aligning frozen black-box models with huma

From XXLTraffic to EvoXXLTraffic: Scaling Traffic Forecasting to Sensor-Evolving Networks

Model ReleasesDGX agent

arXiv:2605.29768v1 Announce Type: new Abstract: Existing traffic forecasting benchmarks assume a fixed sensor set, but real road-sensor networks grow continuously as the road network changes year by y

Frontier LLM-based agents can overcome the ontology curation bottleneck for natural phenotypes

Model ReleasesDGX agent

arXiv:2605.28965v1 Announce Type: new Abstract: Linking free-text phenotype descriptions to ontology terms, typically referred to as phenotype annotation, is essential for the cross-study integration

GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models

SafetyDGX agent

arXiv:2605.29398v1 Announce Type: cross Abstract: Reinforcement learning (RL) can be used to improve the policy (denoiser) of diffusion large language models (dLLMs), while being hindered by the intra

GenesisFunc: Multi-Agent Data Generation for Accurate and Generalizable Function-Calling

AgentsDGX agent

arXiv:2605.28835v1 Announce Type: cross Abstract: Large Language Models (LLMs) extend their capabilities through function-calling (FC), which relies on training data with high quality, diversity, and

Genetically Aligned Patient Representations Improve Hematological Diagnosis

SafetyDGX agent

arXiv:2605.29980v1 Announce Type: cross Abstract: Multimodal alignment of histopathology encoders with transcriptomic and genomic data has been shown to significantly improve performance in downstream

GEO-Bench: Benchmarking Ranking Manipulation in Generative Engine Optimization

Model ReleasesDGX agent

arXiv:2605.29107v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly rank products, documents, and recommendations for user queries, which makes manipulating these rankings a gr

GiPL: Generative augmented iterative Pseudo-Labeling for Cross-Domain Few-Shot Object Detection

ResearchDGX agent

arXiv:2605.29539v1 Announce Type: cross Abstract: Vision-language foundation models have shown promising zero-shot generalization for Cross-Domain Few-Shot Object Detection (CD-FSOD). However, they fa

Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders

Model ReleasesDGX agent

arXiv:2605.30022v1 Announce Type: cross Abstract: Positional encoding (PE) underpins how permutation-invariant Transformers represent sequence order, yet how positional information is processed and st

Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning

Model ReleasesDGX agent

arXiv:2602.01058v2 Announce Type: replace-cross Abstract: Post-training of reasoning LLMs is a holistic process that typically consists of an offline SFT stage followed by an online reinforcement lear

Governing Technical Debt in Agentic AI Systems

AgentsDGX agent

arXiv:2605.29129v1 Announce Type: new Abstract: Agentic AI systems are increasingly being explored as production infrastructure: they reason over multiple steps, call tools, act through workflows, and

GPF-LiveNews: A Streaming Evaluation Protocol for Group-Conditioned Framing in Large Language Models

Model ReleasesDGX agent

arXiv:2605.28848v1 Announce Type: cross Abstract: Deployed language models are evaluated in a non-stationary environment: model versions, retrieval layers, safety systems, and real-world inputs all ch

GPIC: A Giant Permissive Image Corpus for Visual Generation

Model ReleasesDGX agent

arXiv:2605.30341v1 Announce Type: cross Abstract: Studying scalable methods for visual generative modeling requires large, accessible, and stable datasets. We introduce GPIC, a Giant Permissive Image

GPS-Enhanced Tourist Mobility Modeling with Seasonal Spatial Priors and LLM-Based Activity Chain Generation

ResearchDGX agent

arXiv:2605.29578v1 Announce Type: new Abstract: Tourist mobility poses a distinct challenge for urban transportation planning. Unlike resident commuting, tourist travel is largely non-routine, attract

Gram: Assessing sabotage propensities via automated alignment auditing

Model ReleasesDGX agent

arXiv:2605.30322v1 Announce Type: cross Abstract: We introduce Gram, an automated alignment auditing framework to assess the propensity of AI agents to engage in sabotage. We evaluate Gemini models ac

Grammar-Aware Literate Generative Mathematical Programming with Compiler-in-the-Loop

SafetyDGX agent

arXiv:2601.17670v2 Announce Type: replace-cross Abstract: Mathematical programming is widely employed across various sectors - such as logistics, energy, and workforce planning - to model and solve in

Graph-Enhanced Policy Optimization in LLM Agent Training

SafetyDGX agent

arXiv:2510.26270v2 Announce Type: replace Abstract: Multi-step LLM agents in interactive environments represent a crucial step toward long-horizon decision-making. To train such agents, group-based re

GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents

Model ReleasesDGX agent

arXiv:2605.29668v1 Announce Type: new Abstract: LLM agents acting in structured environments fail in operational rather than conversational ways, and reliability depends on procedural knowledge of the

GrepSeek: Training Search Agents for Direct Corpus Interaction

SafetyDGX agent

arXiv:2605.29307v1 Announce Type: cross Abstract: Large Language Model (LLM) search agents have shown strong promise for knowledge-intensive language tasks through multiple rounds of reasoning and inf

GroundAct: Can LLM Agents Ground Actions in Environmental States?

Model ReleasesDGX agent

arXiv:2508.05614v2 Announce Type: replace-cross Abstract: LLM agents achieve 85-96% success on tasks where instructions fully specify the action, but drop to 29-53% when action feasibility depends on

GrowLoop: Self-Evolving Conversation Evaluation Seeded by Human

Model ReleasesDGX agent

arXiv:2605.28882v1 Announce Type: cross Abstract: With the rapid advancement of large language models, evaluating human-likeness in open-ended conversation has become increasingly important. However,

GRPO is Secretly a Process Reward Model

SafetyDGX agent

arXiv:2509.21154v4 Announce Type: replace-cross Abstract: Process reward models (PRMs) allow for fine-grained credit assignment in reinforcement learning (RL), and seemingly contrast with outcome rewa

GTA: Generating Long-Horizon Tasks for Web Agents at Scale

Model ReleasesDGX agent

arXiv:2605.29218v1 Announce Type: new Abstract: Web agents, which couple language models with browsing and tool-use capabilities, show promise as open web assistants. Yet progress is increasingly limi

GUITestScape: Towards Open-set Evaluation on Exploratory GUI Testing

Model ReleasesDGX agent

arXiv:2605.29532v1 Announce Type: cross Abstract: Exploratory GUI testing is a particularly demanding setting for MLLM agents: without predefined test scripts, an agent must autonomously navigate an a

Hallucination Detection-Guided Preference Optimization for Clinical Summarization

Model ReleasesDGX agent

arXiv:2605.28910v1 Announce Type: cross Abstract: Large language models (LLMs) have shown promise on summarization tasks, but they often produce hallucinations, which are unsupported or incorrect stat

Hallucination Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic Caching

Model ReleasesDGX agent

arXiv:2605.29055v1 Announce Type: new Abstract: Hallucination remains a major reliability barrier for production LLM systems, particularly in multi-agent pipelines where unsupported claims can propaga

Harmonizing Real-Time Constraints and Long-Horizon Reasoning: An Asynchronous Agentic Framework for Dynamic Scheduling

SafetyDGX agent

arXiv:2605.29262v1 Announce Type: new Abstract: The Dynamic Flexible Job Shop Scheduling Problem (DFJSP) necessitates a trade-off between instant reaction to stochastic disturbances and global optimiz

Harnessing non-adversarial robustness in large language models

SafetyDGX agent

arXiv:2605.29816v1 Announce Type: new Abstract: The work presents an approach for addressing the challenge of robustness in Large Language Models (LLMs) to alterations and potential errors caused by s

HARP: Hadamard-Preconditioned Adaptive Rotation Processor for Extreme LLM Quantization

ResearchDGX agent

arXiv:2605.29843v1 Announce Type: cross Abstract: Post-training quantization (PTQ) is essential for deploying LLMs under memory and bandwidth constraints. However, extreme low-bit quantization remains

HD-Prot: A Protein Language Model for Joint Sequence-Structure Modeling with Continuous Structure Tokens

TutorialsDGX agent

arXiv:2512.15133v2 Announce Type: replace-cross Abstract: Proteins inherently possess a consistent sequence-structure duality. The abundance of protein sequence data, which can be readily represented

Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction

AgentsDGX agent

arXiv:2605.29960v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly leverage long term memory to support persistent and autonomous task execution. However, this capability

HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question Answering

ResearchDGX agent

arXiv:2605.29606v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) for document-based Open-domain Question Answering (ODQA) on large-scale industrial corpora faces two critical bottl

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.29782v1 Announce Type: cross Abstract: Reinforcement learning (RL) refines large language models (LLMs) by directly optimizing model behavior through reward signals. While accurate state va

HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

ResearchDGX agent

arXiv:2605.29948v1 Announce Type: cross Abstract: Unified speech foundation models require a holistic tokenization space that is both learnable by language models and decodable into high-quality wavef

Honest Lying: Understanding Memory Confabulation in Reflexive Agents

ResearchDGX agent

arXiv:2605.29463v1 Announce Type: cross Abstract: Reflexion-style agents rely on self-generated reflections as memory, implicitly assuming that agents can accurately diagnose their own failures.We sho

Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots

AgentsDGX agent

arXiv:2605.29963v1 Announce Type: cross Abstract: Honeypots are decoy systems mimicking real system components designed to defend against cyber attacks. Recently, LLMs increasingly serve as simulation

How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions

Model ReleasesDGX agent

arXiv:2605.29442v1 Announce Type: cross Abstract: AI coding agents increasingly act directly within software environments, yet existing analyses of their failures rely on benchmark trajectories that m

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines

AgentsDGX agent

arXiv:2605.28840v1 Announce Type: cross Abstract: Large language model (LLM) agents with tool-calling capabilities are increasingly deployed in production systems, yet a fundamental reliability questi

How LoRA Remembers? A Parametric Memory Law for LLM Finetuning

Model ReleasesDGX agent

arXiv:2605.30260v1 Announce Type: cross Abstract: Large Language Models (LLMs) must continuously learn and update knowledge to remain effective in dynamic real-world environments. While Low-Rank Adapt

How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions

ResearchDGX agent

arXiv:2605.29448v1 Announce Type: cross Abstract: Neural scaling laws appraise data through dataset size, while the Vendi Score uses quantum entropy to measure dataset value. We show both that common

How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency

Model ReleasesDGX agent

arXiv:2605.30096v1 Announce Type: cross Abstract: Large language models (LLMs) can autonomously conduct multi-stage cyber attacks, but the consistency of their offensive behavior under repeated trials

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime

SafetyDGX agent

arXiv:2605.30201v1 Announce Type: cross Abstract: We investigate a narrow but common failure mode of GRPO-style reinforcement learning in the context of sparse verifiable rewards: early updates contai

← Previous
1…190191192193194…358
Next →