AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
20 May 2026

Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges

Model ReleasesDGX agent

arXiv:2605.19723v1 Announce Type: cross Abstract: Mathematical reasoning is essential for problem-solving in education, science, and industry, serving as a crucial benchmark for evaluating artificial

MaxShapley: Towards Incentive-compatible Generative Search with Fair Context Attribution

ResearchDGX agent

arXiv:2512.05958v2 Announce Type: replace-cross Abstract: Generative search engines based on large language models (LLMs) are replacing traditional search, fundamentally changing how information provi

Measuring Safety Alignment Effects in Autonomous Security Agents

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.19722v1 Announce Type: cross Abstract: Do stock safety-aligned language models and their uncensored or abliterated derivatives behave differently when run as autonomous security agents? Sin

Mechanistic Interpretability Needs Philosophy

ResearchDGX agent

arXiv:2506.18852v2 Announce Type: replace-cross Abstract: Mechanistic interpretability (MI) aims to explain how neural networks work by uncovering their underlying mechanisms. As the field grows in in

Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation

SafetyDGX agent

arXiv:2605.19833v1 Announce Type: cross Abstract: Despite rapid advances in automatic speech recognition (ASR) and large audio-language models, robust recognition in real-world environments remains li

Memory-Augmented Reinforcement Learning Agent for CAD Generation

SafetyDGX agent

arXiv:2605.19748v1 Announce Type: new Abstract: Automatic generation of computer-aided design (CAD) models is a core technology for enabling intelligence in advanced manufacturing. Existing generation

Metric-Gradient Projection for Stable Multi-Agent Policy Learning

SafetyDGX agent

arXiv:2605.18809v1 Announce Type: cross Abstract: General-sum multi-agent learning is often governed by a stacked update field in which each agent's policy update changes the optimization landscape fa

MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models

ResearchDGX agent

arXiv:2605.19619v1 Announce Type: cross Abstract: Matrix-structured parameters frequently appear in many artificial intelligence models such as large language models. More recently, an efficient Muon

Minimax Optimal Variance-Aware Regret Bounds for Multinomial Logistic MDPs

ResearchDGX agent

arXiv:2605.19768v1 Announce Type: new Abstract: We study reinforcement learning for episodic Markov Decision Processes (MDPs) whose transitions are modelled by a multinomial logistic (MNL) model. Exis

MO-CAPO: Multi-Objective Cost-Aware Prompt Optimization

ResearchDGX agent

arXiv:2605.18869v1 Announce Type: cross Abstract: Large language models (LLMs) achieve strong performance across a wide range of tasks but are highly sensitive to prompt design, motivating the need fo

MOCHA: Multi-Objective Chebyshev Annealing for Agent Skill Optimization

AgentsDGX agent

arXiv:2605.19330v1 Announce Type: new Abstract: LLM agents organize behavior through skills - structured natural-language specifications governing how an agent reasons, retrieves, and responds. Unlike

MoCo-EA: Exploiting Adversarial Mode Connectivity for Efficient Evolutionary Attacks

ResearchDGX agent

arXiv:2605.18919v1 Announce Type: cross Abstract: Evolutionary algorithms for adversarial attacks leverage population-based search to discover perturbations without gradient information, but suffer fr

Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews

ApplicationsDGX agent

arXiv:2403.07183v4 Announce Type: replace-cross Abstract: We present an approach for estimating the fraction of text in a large corpus which is likely to be substantially modified or produced by a lar

Multi-Scale Generative Modeling with Heat Dissipation Flow Matching

ResearchDGX agent

arXiv:2605.19371v1 Announce Type: cross Abstract: Diffusion models are widely used in image generation, with most relying on noise-based corruption and denoising. A distinct branch instead uses blur a

Multimodal system for skin cancer detection

ApplicationsDGX agent

arXiv:2601.14822v2 Announce Type: replace-cross Abstract: Melanoma detection is vital for early diagnosis and effective treatment. While deep learning models on dermoscopic images have shown promise,

Neural Operators for Design-Space Surrogate Modeling of Tendon-Actuated Continuum Robots

ResearchDGX agent

arXiv:2605.19104v1 Announce Type: cross Abstract: Continuum robots enable dexterous manipulation in constrained environments, but require accurate and efficient models for real-time manipulation and c

Neurosymbolic Learning for Inference-Time Argumentation

TutorialsDGX agent

arXiv:2605.20098v1 Announce Type: new Abstract: Claim verification is an important problem in high-stakes settings, including health and finance. When information underpinning claims is incomplete or

Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients

SafetyDGX agent

arXiv:2510.18924v3 Announce Type: replace-cross Abstract: Reinforcement learning from human feedback (RLHF) or verifiable rewards (RLVR), the standard paradigm for aligning LLMs or building recent SOT

Nonlinearity as Rank: Generative Low-Rank Adapter with Radial Basis Functions

Model ReleasesDGX agent

arXiv:2602.05709v2 Announce Type: replace Abstract: Low-rank adaptation (LoRA) approximates the update of a pretrained weight matrix using the product of two low-rank matrices. However, standard LoRA

NORi: An ML-Augmented Ocean Boundary Layer Parameterization

Local AiDGX agent

arXiv:2512.04452v2 Announce Type: replace-cross Abstract: NORi is a machine learning (ML) parameterization of ocean boundary layer turbulence that is physics-based and augmented with neural networks.

Not all uncertainty is alike: volatility, stochasticity, and exploration

SafetyDGX agent

arXiv:2605.19215v1 Announce Type: new Abstract: Adaptive decision-making in biological and artificial intelligence requires balancing the exploitation of known outcomes with the exploration of uncerta

Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR

SafetyDGX agent

arXiv:2605.20164v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has made post-training highly effective when correctness can be checked automatically. However, many impo

OEP: Poisoning Self-Evolving LLM Agents via Locally Correct but Non-Transferable Experiences

Local AiDGX agent

arXiv:2605.18930v1 Announce Type: cross Abstract: Memory-augmented large language model (LLM) agents use iterative reflection and self-evolution to solve complex tasks, but these mechanisms introduce

OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments

Model ReleasesDGX agent

arXiv:2605.18758v1 Announce Type: cross Abstract: Current benchmarks for graphical user interface (GUI) agents predominantly rely on static screenshots. However, real-world smartphone interaction rout

On-Device Continual Learning with Dual-Stage Buffer and Dynamic Loss for Point-of-Care Pneumonia Diagnosis

Local AiDGX agent

arXiv:2605.19201v1 Announce Type: cross Abstract: Deep learning models detect pneumonia from chest X-rays with high accuracy, but the performance declines under domain shifts caused by differences in

Open-Set Domain Adaptation Under Background Distribution Shift: Challenges and A Provably Efficient Solution

ApplicationsDGX agent

arXiv:2512.01152v4 Announce Type: replace-cross Abstract: As we deploy machine learning systems in the real world, a core challenge is to maintain a model that is performant even as the data shifts. S

OpenComputer: Verifiable Software Worlds for Computer-Use Agents

ResearchDGX agent

arXiv:2605.19769v1 Announce Type: new Abstract: We present OpenComputer, a verifier-grounded framework for constructing verifiable software worlds for computer-use agents. OpenComputer integrates four

Operationalising Artificial Intelligence Bills of Materials (AIBOMs) for Verifiable AI Provenance and Lifecycle Assurance

AgentsDGX agent

arXiv:2605.19755v1 Announce Type: cross Abstract: Artificial Intelligence (AI) systems are increasingly dependent on complex, multi-layered software supply chains that introduce challenges for reprodu

Operationalizing Document AI: A Microservice Architecture for OCR and LLM Pipelines in Production

Model ReleasesDGX agent

arXiv:2605.18818v1 Announce Type: new Abstract: Academic research tends to focus on new models for document understanding creating a wide gap in the literature between model definition and running mod

optimize_anything: A Universal API for Optimizing any Text Parameter

Model ReleasesDGX agent

arXiv:2605.19633v1 Announce Type: cross Abstract: Can a single LLM-based optimization system match specialized tools across fundamentally different domains? We show that when optimization problems are

P2DNav: Panorama-to-Downview Reasoning for Zero-shot Vision-and-Language Navigation

Model ReleasesDGX agent

arXiv:2605.19634v1 Announce Type: cross Abstract: Vision-and-language navigation (VLN) requires an embodied agent to ground natural-language instructions into executable navigation actions in unseen e

Passive Construction Site Safety Monitoring via Persona-Scaffolded Adversarial Chain-of-Thought VLM Verification

Model ReleasesDGX agent

arXiv:2605.19869v1 Announce Type: cross Abstract: Construction remains the deadliest industry sector in the United States, with 1,055 fatal worker injuries recorded in 2023, and the majority preventab

PAVE: A Cognitive Architecture for Legitimate Violation in Generative Agent Societies

AgentsDGX agent

arXiv:2605.19351v1 Announce Type: cross Abstract: Generative agents based on large language models reproduce believable human behavior in cooperative settings, but how they should reason in situations

PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents

SafetyDGX agent

arXiv:2605.19932v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly operate over long and recurring external contexts, like document corpora and code repositories. Across in

Phase-Aware Mixture of Experts for Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2602.17038v3 Announce Type: replace Abstract: Reinforcement learning (RL) has equipped LLM agents with a strong ability to solve complex tasks. However, existing RL methods normally use a single

PhyWorld: Physics-Faithful World Model for Video Generation

Model ReleasesDGX agent

arXiv:2605.19242v1 Announce Type: cross Abstract: World simulators can provide safe and scalable environments for training Physical AI systems before real-world deployment. Large video generation mode

PiKV: KV Cache Management System for Mixture of Experts

HardwareDGX agent

arXiv:2508.06526v3 Announce Type: replace-cross Abstract: As large-scale language models continue to scale up in both size and context length, the memory and communication cost of key-value (KV) cache

Planner-Admissible Graph-PDE Value Extensions for Sparse Goal-Conditioned Planning

ResearchDGX agent

arXiv:2605.19185v1 Announce Type: cross Abstract: Sparse goal-conditioned planning with few cost-to-go labels can be viewed as a graph-PDE Dirichlet extension problem: extend sparse labels on a goal-d

PlantTraitNet: An Uncertainty-Aware Multimodal Framework for Global-Scale Plant Trait Inference from Citizen Science Data

Model ReleasesDGX agent

arXiv:2511.06943v3 Announce Type: replace-cross Abstract: Global plant maps of plant traits, such as leaf nitrogen or plant height, are essential for understanding ecosystem processes, including the c

POLAR-Bench: A Diagnostic Benchmark for Privacy-Utility Trade-offs in LLM Agents

Model ReleasesDGX agent

arXiv:2605.19127v1 Announce Type: new Abstract: LLM agents increasingly have access to private user data and act on the user's behalf when interacting with third-party systems. The user defines what m

Position: Let's Develop Data Probes to Fundamentally Understand How Data Affects LLM Performance

SafetyDGX agent

arXiv:2605.18801v1 Announce Type: new Abstract: Data is fundamental to large language models (LLMs). However, understanding of what makes certain data useful for different stages of an LLM workflow, i

Position: The Turing-Completeness of Real-World Autoregressive Transformers Relies Heavily on Context Management

Model ReleasesDGX agent

arXiv:2605.19514v1 Announce Type: new Abstract: Many works make the eye-catching claim that Transformers are Turing-complete. However, the literature often conflates two distinct settings: (i) a fixed

Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering

SafetyDGX agent

arXiv:2605.19220v1 Announce Type: cross Abstract: Uncertainty Quantification (UQ) is widely regarded as the primary safeguard for deploying Large Language Models (LLMs) in high-stakes domains. However

PragLocker: Protecting Agent Intellectual Property in Untrusted Deployments via Non-Portable Prompts

AgentsDGX agent

arXiv:2605.05974v2 Announce Type: replace-cross Abstract: LLM agents rely on prompts to implement task-specific capabilities based on foundation LLMs, making agent prompts valuable intellectual proper

Precision Tracked Transformer via Kalman Filtering, Kriging and Process Noise

Model ReleasesDGX agent

arXiv:2605.18832v1 Announce Type: cross Abstract: The Transformer is the foundational building block of modern AI, yet offers no principled handling of uncertainty, which is prevalent in real applicat

Prediction Is Not Physics: Learning and Evaluating Conserved Quantities in Neural Simulators

SafetyDGX agent

arXiv:2605.18883v1 Announce Type: cross Abstract: A diffusion model trained on Hamiltonian trajectories can achieve rollout MSE near 10^{-3}, but the standard deviation of its energy over time is betw

Prior Knowledge or Search? A Study of LLM Agents in Hardware-Aware Code Optimization

HardwareDGX agent

arXiv:2605.19782v1 Announce Type: new Abstract: LLM discovery and optimization systems are increasingly applied across domains, implementing a common propose-evaluate-revise loop. Such optimization or

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning

Model ReleasesDGX agent

arXiv:2605.19382v1 Announce Type: new Abstract: Programmatic video generation through code offers geometric precision and temporal coherence beyond pixel-level diffusion models, yet rigorously evaluat

Probabilistic Tiny Recursive Model

ResearchDGX agent

arXiv:2605.19943v1 Announce Type: new Abstract: Tiny Recursive Models (TRM) solve complex reasoning tasks with a fraction of the parameters of modern large language models (LLMs) by iteratively refini

Probability-Conserving Flow Guidance

ResearchDGX agent

arXiv:2605.20079v1 Announce Type: cross Abstract: Diffusion and flow-based generative models dominate visual synthesis, with guidance aligning samples to user input and improving perceptual quality. H

Probing Embodied LLMs: When Higher Observation Fidelity Hurts Problem Solving

AgentsDGX agent

arXiv:2605.20072v1 Announce Type: new Abstract: Large Language Models are increasingly proposed as cognitive components for robotic systems, yet their opaque decision processes make it difficult to ex

Progressive Autonomy as Preference Learning: A Formalization of Trust Calibration for Agentic Tool Use

SafetyDGX agent

arXiv:2605.19151v1 Announce Type: new Abstract: We formalize trust calibration for agentic tool use (deciding when an automated agent's proposed action may execute autonomously versus require human ap

Projecting Latent RL Actions: Towards Generalizable and Scalable Graph Combinatorial Optimization

ResearchDGX agent

arXiv:2605.19721v1 Announce Type: new Abstract: Graph combinatorial optimization (GCO) has attracted growing interest, as many NP-hard problems naturally admit graph formulations, yet their combinator

PromptRad: Knowledge-Enhanced Multi-Label Prompt-Tuning for Low-Resource Radiology Report Labeling

Model ReleasesDGX agent

arXiv:2605.20052v1 Announce Type: cross Abstract: Automatic report labeling facilitates the identification of clinical findings from unstructured text and enables large-scale annotation for medical im

Protein Autoregressive Modeling via Multiscale Structure Generation

Model ReleasesDGX agent

arXiv:2602.04883v2 Announce Type: replace-cross Abstract: We present protein autoregressive modeling (PAR), the first multi-scale autoregressive framework for protein backbone generation via coarse-to

PROWL: Prioritized Regret-Driven Optimization for World Model Learning

SafetyDGX agent

arXiv:2605.18803v1 Announce Type: cross Abstract: Modern action-conditioned video world models achieve strong short-horizon visual realism, yet remain unreliable on rare, interaction-critical transiti

Proximal Diffusion Neural Sampler

ResearchDGX agent

arXiv:2510.03824v2 Announce Type: replace-cross Abstract: The task of learning a diffusion-based neural sampler for drawing samples from an unnormalized target distribution can be viewed as a stochast

Pseudocode-Guided Structured Reasoning for Automating Reliable Inference in Vision-Language Models

SafetyDGX agent

arXiv:2605.19663v1 Announce Type: new Abstract: Vision-Language Models (VLMs) are becoming the cornerstone of high-level reasoning for robotic automation, enabling robots to parse natural language com

Quantifying the Generalization Gap in Seizure Detection: A Large-Scale Empirical Benchmark via the SzCORE Challenge

Model ReleasesDGX agent

arXiv:2505.18191v2 Announce Type: replace-cross Abstract: Reliable automatic seizure detection from long-term electroencephalography (EEG) remains an unsolved challenge, as current models often fail t

Quantifying the Pre-training Dividend: Generative versus Latent Self-Supervised Learning for Time Series Foundation Models

SafetyDGX agent

arXiv:2605.19462v1 Announce Type: cross Abstract: The success of self-supervised learning (SSL) in vision and NLP has motivated its rapid adoption for time series. However, research has focused primar

← Previous
1…229230231232233…358
Next →