AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
22 May 2026

ProcBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents

Model ReleasesDGX agent

arXiv:2605.20251v2 Announce Type: cross Abstract: Existing benchmarks for LLM coding agents primarily evaluate final outcomes. While useful for measuring overall capability, these metrics provide limi

Proportional Selection in Networks

ResearchDGX agent

arXiv:2502.03545v2 Announce Type: replace-cross Abstract: We address the problem of selecting k representative nodes from a network, aiming to achieve two objectives: identifying the most influential

Quality and Security Signals in AI-Generated Python Refactoring Pull Requests

AgentsDGX agent

arXiv:2605.21453v1 Announce Type: cross Abstract: As AI agents increasingly contribute to code development and maintenance, there is still limited empirical evidence on the quality and risk characteri


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation

Model ReleasesDGX agent

arXiv:2605.20204v1 Announce Type: cross Abstract: LLM-based user simulation is the primary mechanism for end-to-end agent evaluation, yet simulated users are poor proxies for real humans: unconstraine

Reinforcing VLAs in Task-Agnostic World Models

SafetyDGX agent

arXiv:2605.12334v2 Announce Type: replace Abstract: Post-training Vision-Language-Action (VLA) models via reinforcement learning (RL) in learned world models has emerged as an effective strategy to ad

Representability-Aware Neural Networks for Reduced Density Matrices: Application to Fractional Chern Insulators

ResearchDGX agent

arXiv:2605.20326v1 Announce Type: cross Abstract: We develop a representability-aware and interpolable neural network (NN) framework for predicting two-particle reduced density matrices (2-RDMs). The

Retaining Suboptimal Actions to Follow Shifting Optima in Multi-Agent Reinforcement Learning

SafetyDGX agent

arXiv:2602.17062v2 Announce Type: replace Abstract: Value decomposition is a core approach for cooperative multi-agent reinforcement learning (MARL). However, existing methods still rely on a single o

ScenePilot: Controllable Boundary-Driven Critical Scenario Generation for Autonomous Driving

SafetyDGX agent

arXiv:2605.21168v1 Announce Type: new Abstract: Safety-critical scenarios are central to evaluating autonomous driving systems, yet their rarity in naturalistic logs makes simulation-based stress test

Security Document Classification with a Fine-Tuned Local Large Language Model: Benchmark Data and an Open-Source System

Model ReleasesDGX agent

arXiv:2605.20368v1 Announce Type: cross Abstract: Organizations that scan documents for sensitive information face a practical problem. Cloud services require data to be sent to external infrastructur

Stdlib or Third-Party? Empirical Performance and Correctness of LLM-Assisted Zero-Dependency Python Libraries

AgentsDGX agent

arXiv:2605.21405v1 Announce Type: cross Abstract: Third-party Python libraries introduce dependency management overhead, supply chain risk, and deployment friction in constrained environments. A natur

SURGE: An Event-Centric Social Media Sentiment Time Series Benchmark with Interaction Structure

Model ReleasesDGX agent

arXiv:2605.21198v1 Announce Type: cross Abstract: Public events on social media generate large volumes of discussion whose collective dynamics carry direct value for opinion forecasting and crisis res

Targeting Clause Type Distributions: a Picklock for Random Satisfiability Problems

Model ReleasesDGX agent

arXiv:2605.20328v1 Announce Type: cross Abstract: Optimization problems such as the NP-complete 3-SAT provide an important benchmark for the difficult task of finding ground-states in strongly correla

Teaching AI Through Benchmark Construction: QuestBench as a Course-Based Practice for Accountable Knowledge Work

Model ReleasesDGX agent

arXiv:2605.21413v2 Announce Type: new Abstract: As AI becomes part of everyday learning, many courses teach students to use it mainly as a productivity tool: how to prompt, search, summarize, write, c

Tool-Augmented Agent for Closed-loop Optimization,Simulation,and Modeling Orchestration

AgentsDGX agent

arXiv:2605.20190v1 Announce Type: new Abstract: Iterative industrial design-simulation optimization is bottlenecked by the CAD-CAE semantic gap: translating simulation feedback into valid geometric ed

Transcription and Recognition of Italian Parliamentary Speeches Using Vision-Language Models

Model ReleasesDGX agent

arXiv:2603.28103v2 Announce Type: replace-cross Abstract: Parliamentary proceedings represent a rich yet challenging resource for computational analysis, particularly when preserved only as scanned hi

VBFDD-Agent for Electric Vehicle Battery Fault Detection and Diagnosis: Descriptive Text Modeling of Battery Digital Signals

Local AiDGX agent

arXiv:2605.20742v1 Announce Type: new Abstract: With the rapid proliferation of electric vehicles, the safety and reliability of lithium-ion batteries have become critical concerns. Effective anomaly

20 May 2026

3D aperture-engineered diffractive neural networks for super-resolution electromagnetic wave computing

ResearchDGX agent

arXiv:2603.00995v2 Announce Type: replace-cross Abstract: The rapid progress in 6G communication and high-bandwidth radar has driven an unprecedented surge in the spatial density of signal sources, re

A Bitter Lesson for Data Filtering

Model ReleasesDGX agent

arXiv:2605.19407v1 Announce Type: cross Abstract: We investigate data filtering for large model pretraining via new scaling studies that target the high compute, data-scarce regime. In spite of an app

A Case for Agentic Tuning: From Documentation to Action in PostgreSQL

Model ReleasesDGX agent

arXiv:2605.19988v1 Announce Type: cross Abstract: Documentation has long guided computer system tuning by distilling expert knowledge into per-parameter recommendations. Yet such guides capture only w

A Closed-loop, State-centric, Multi-agent Framework for Passenger Load Estimation from Heterogeneous Data Streams

AgentsDGX agent

arXiv:2605.19834v1 Announce Type: cross Abstract: To support operations and passenger-facing services, transit agencies need reliable passenger load trajectories. Currently, load estimates are typical

A Framework for Evaluating Zero-Shot Image Generation in Concept-based Explainability

ResearchDGX agent

arXiv:2605.19855v1 Announce Type: cross Abstract: Concept-based Explainable Artificial Intelligence (XAI) interprets deep learning models using human-understandable visual features (e.g., textures or

A Geometric Analysis of Small-sized Language Model Hallucinations

AgentsDGX agent

arXiv:2602.14778v3 Announce Type: replace-cross Abstract: Hallucinations -- plausible but factually incorrect responses -- pose a major challenge to the reliability of Large Language Models (LLMs), es

A Hybrid Modeling Framework for Crop Prediction Tasks via Dynamic Parameter Calibration and Multi-Task Learning

Model ReleasesDGX agent

arXiv:2603.15411v2 Announce Type: replace Abstract: Accurate prediction of crop states (e.g., phenology stages and cold hardiness) is essential for timely farm management decisions such as irrigation,

A Measure-Theoretic Analysis of Reasoning: Structural Generalization and Approximation Limits

ResearchDGX agent

arXiv:2605.19944v1 Announce Type: cross Abstract: While empirical scaling laws for LLM reasoning are well-documented, the theoretical mechanisms governing out-of-distribution (OOD) generalization rema

A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents

AgentsDGX agent

arXiv:2605.20173v1 Announce Type: new Abstract: Production LLM agents combine stochastic model outputs with deterministic software systems, yet the boundary between the two is rarely treated as a firs

A Nonlinear Complexity Index for Wearable PPG Cardiovascular Stability: Multiscale Validation, Systematic Evaluation Correction, and Bayesian Parameter Optimization

Model ReleasesDGX agent

arXiv:2605.18802v1 Announce Type: cross Abstract: Cardiovascular stability estimation from wearable photoplethysmography (PPG) requires a principled nonlinear framework, yet major gaps persist in heur

A novel YOLO26-MoE optimized by an LLM agent for insulator fault detection considering UAV images

AgentsDGX agent

arXiv:2605.19595v1 Announce Type: cross Abstract: The inspection of electrical power line insulators is essential for ensuring grid reliability and preventing failures caused by damaged or degraded in

A Reproducibility Analysis of PO4ISR: Diagnosing and Mitigating Semantic Drift in LLM-Based Session Recommendation

Model ReleasesDGX agent

arXiv:2605.18780v1 Announce Type: cross Abstract: Reasoning-based Large Language Models (LLMs) like PO4ISR have set new benchmarks in session-based recommendation. However, the reproducibility of thei

Adaptive Multi-Scale Goodness Aggregation for Forward-Forward Learning

ResearchDGX agent

arXiv:2605.18804v1 Announce Type: cross Abstract: We propose Adaptive Multi-Scale Goodness Aggregation (AMSGA), a novel extension of the Forward-Forward (FF) algorithm designed to improve stability, r

Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision Language Models

ResearchDGX agent

arXiv:2511.10292v2 Announce Type: replace-cross Abstract: Large Vision-Language Models (LVLMs) typically process visual inputs as a prefix to the language decoder. As the model autoregressively genera

Aerial Inspection Behaviors via RL-based Quadrotor Control for Under-canopy Forest Environments

SafetyDGX agent

arXiv:2605.19202v1 Announce Type: cross Abstract: This paper addresses the problem of using a deep Reinforcement Learning (RL)-based low-level Quadrotor controller within an autonomous Quadrotor navig

AffectAI-Capture: A Reproducible Multimodal Protocol for Small-Group Meeting Research

ResearchDGX agent

arXiv:2605.19794v1 Announce Type: cross Abstract: We present AffectAI-Capture, a protocol for collecting synchronized multimodal data in four-person meeting-like interactions, combining eye tracking,

Agent Security is a Systems Problem

AgentsDGX agent

arXiv:2605.18991v1 Announce Type: cross Abstract: We take the position that agent security must be approached as a systems problem: the AI model powering the agent must be treated as an untrusted comp

Agentic GraphRAG: Navigating Unstructured Financial Data with Collaborative AI

AgentsDGX agent

arXiv:2605.18770v1 Announce Type: cross Abstract: We present a collaborative agentic GraphRAG framework for expert analysis of commercial registry data. Public registries are often formally accessible

Agentic Trading: When LLM Agents Meet Financial Markets

AgentsDGX agent

arXiv:2605.19337v1 Announce Type: new Abstract: A growing body of work explores how Large Language Models (LLMs) can be embedded in trading systems as agents that perceive market information, retrieve

AgentNLQ: A General-Purpose Agent for Natural Language to SQL

Model ReleasesDGX agent

arXiv:2605.19010v1 Announce Type: new Abstract: Natural language to SQL (NL2SQL) conversion is an important problem for researchers and enterprises due to the ubiquitous importance of relational datab

AI Technologies in Language Access: Attitudes Towards AI and the Human Value of Language Access Managers

Local AiDGX agent

arXiv:2605.19234v1 Announce Type: cross Abstract: The rapid emergence of AI technologies is reshaping translation practices and theory across the board. This paper deals with the impact of AI in langu

ALDEN: Boosting Private Data Extraction from Retrieval-Augmented Generation Systems via Active Learning and Distribution Estimation

ResearchDGX agent

arXiv:2605.18762v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) is widely used to augment large language models with external knowledge retrieval to improve reliability and gene

An Efficient Multilevel Preconditioned Nonlinear Conjugate Gradient Method for Incremental Potential Contact

ResearchDGX agent

arXiv:2604.19892v1 Announce Type: cross Abstract: Incremental Potential Contact (IPC) guarantees intersection-free simulation but suffers from high computational costs due to the expensive Hessian ass

An Integrated Forecasting Prototype for Emergency Department Boarding Time to Support Proactive Operational Decision Making

ApplicationsDGX agent

arXiv:2605.18839v1 Announce Type: cross Abstract: Overcrowding in emergency departments (ED) remains a persistent operational challenge worldwide, causing delays in care delivery and downstream conges

AQuaUI: Visual Token Reduction for GUI Agents with Adaptive Quadtrees

AgentsDGX agent

arXiv:2605.19260v1 Announce Type: new Abstract: Large Multimodal Models (LMMs) have recently emerged as promising backbones for GUI-agent models, where high-resolution GUI screenshots are introduced t

AR1-ZO: Topology-Aware Rank-1 Zeroth-Order Queries for High-Rank LoRA Fine-Tuning

ResearchDGX agent

arXiv:2605.19767v1 Announce Type: cross Abstract: Zeroth-order (ZO) optimization enables large-language-model fine-tuning without storing backpropagation activations, while LoRA supplies compact train

ARC-RL: A Reinforcement Learning Playground Inspired by ARC Raiders

SafetyDGX agent

arXiv:2605.19503v1 Announce Type: cross Abstract: Reinforcement learning for legged locomotion has matured into a stack of multi-component reward functions and physics-engine benchmarks whose morpholo

Are Rationales Necessary and Sufficient? Tuning LLMs for Explainable Misinformation Detection

ResearchDGX agent

arXiv:2605.19285v1 Announce Type: cross Abstract: The rapid spread of misinformation on social media platforms has become a formidable challenge. To mitigate its proliferation, Misinformation Detectio

ARM: Discovering Agentic Reasoning Modules for Generalizable Multi-Agent Systems

AgentsDGX agent

arXiv:2510.05746v2 Announce Type: replace Abstract: Large Language Model (LLM)-powered Multi-agent systems (MAS) have achieved state-of-the-art results on various complex reasoning tasks. Recent works

Artificial Intelligence, conceptual metaphors and conceptual engineering: Are AI-based framings of human behaviour and cognition successful?

ResearchDGX agent

arXiv:2504.07756v2 Announce Type: replace Abstract: Understanding human behaviour, neuroscience and psychology using concepts from the domain of AI is increasing in popularity. Given the massive integ

Artificial Phantasia: Emergent Mental Imagery in Large Language Models

ResearchDGX agent

arXiv:2509.23108v2 Announce Type: replace Abstract: Can visual imagery be driven solely by language? This idea goes against cognitive science's traditional view that visual mental imagery is only poss

Atoms of Thought: Universal EEG Representation Learning with Microstates

ResearchDGX agent

arXiv:2605.20182v1 Announce Type: cross Abstract: Learning universal representations from electroencephalogram (EEG) signals is a cutting-edge approach in the field of neuroinformatics and brain-compu

Attention-Guided Reward for Reinforcement Learning-based Jailbreak against Large Reasoning Models

SafetyDGX agent

arXiv:2605.19485v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) have demonstrated remarkable capabilities in solving complex problems by generating structured, step-by-step reasoning con

Auditing Reasoning-Trace Memorization Claims after Unlearning with Head-Conditioned Canaries

Model ReleasesDGX agent

arXiv:2605.18891v1 Announce Type: cross Abstract: Evaluations of unlearning on reasoning models sometimes show a bypass pattern. The answer side looks unlearned, but the model's own thinking trace kee

Automated Big Data Quality Assessment using Knowledge Graph Embeddings

ApplicationsDGX agent

arXiv:2605.18833v1 Announce Type: cross Abstract: Automated data quality assessment is crucial for managing big data, but existing solutions face challenges in achieving accurate context-aware assessm

Automated Grading of Handwritten Mathematics Using Vision-Capable LLMs

ResearchDGX agent

arXiv:2605.19043v1 Announce Type: cross Abstract: Automated grading systems have enabled scalable assessment for many response types, but handwritten mathematics remains a barrier due to the complexit

AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration

Model ReleasesDGX agent

arXiv:2605.20025v1 Announce Type: new Abstract: Automating scientific discovery requires more than generating papers from ideas. Real research is iterative: hypotheses are challenged from multiple per

Backtracking When It Strays: Mitigating Dual Exposure Biases in LLM Reasoning Distillation

SafetyDGX agent

arXiv:2605.19433v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved remarkable success in complex reasoning tasks via long chain-of-thought (CoT), yet their immense computatio

BalanceRAG: Joint Risk Calibration for Cascaded Retrieval-Augmented Generation

ResearchDGX agent

arXiv:2605.20084v1 Announce Type: cross Abstract: Large language models (LLMs) can enhance factuality via retrieval-augmented generation (RAG), but applying RAG to every query is unnecessary when the

Base Models Look Human To AI Detectors

Model ReleasesDGX agent

arXiv:2605.19516v1 Announce Type: cross Abstract: As AI-generated text enters the real-world at scale, institutions increasingly use commercial AI-text detectors, especially in education and academic-

Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks

SafetyDGX agent

arXiv:2605.19147v1 Announce Type: cross Abstract: Large language models (LLMs) are highly susceptible to backdoor attacks (BAs), wherein training samples are poisoned using trigger-based harmful conte

Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German

Model ReleasesDGX agent

arXiv:2605.19069v1 Announce Type: cross Abstract: Code-switching -- the natural alternation between two languages within a single utterance -- represents one of the most challenging and under-studied

Beyond Isotropy in JEPAs: Hamiltonian Geometry and Symplectic Prediction

SafetyDGX agent

arXiv:2605.20107v1 Announce Type: cross Abstract: JEPAs often regularize one-view embeddings toward an isotropic Gaussian, implicitly baking Euclidean symmetry into the representation. We show that th

Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information

AgentsDGX agent

arXiv:2510.01499v2 Announce Type: replace-cross Abstract: With the rapid progress of multi-agent large language model (LLM) reasoning, how to effectively aggregate answers from multiple LLMs has emerg

← Previous
1…225226227228229…358
Next →