AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
3 Jun 2026

ReLoRA: Knowledge-Reusing Adaptation for Fast Rollout of Evolving LLM Services

ResearchDGX agent

arXiv:2606.02606v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed as continuously evolving services, where frequent base-model updates may invalidate previously

Representational Capacity: Geometric Limits on Feature Representation in Transformer Language Models

ResearchDGX agent

arXiv:2606.02765v1 Announce Type: cross Abstract: Model dimension (d_{model}) is a fundamental hyperparameter in transformer language models, yet its role in setting the geometric limits of feature re

Reproducibility is the New Copyleft: Defining AGI-oriented Reproducible Builds

AgentsDGX agent

arXiv:2606.03019v1 Announce Type: cross Abstract: Copyleft, as implemented in licenses such as the GNU General Public License, was a legal hack that used copyright to guarantee user freedom by tying t


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Rethinking Molecular Text Representations for LLMs: An Empirical Study

Model ReleasesDGX agent

arXiv:2606.03057v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for molecular tasks, but it remains unclear which molecular representation to use. We present a sys

Rethinking the Role of Tensor Decompositions in Post-Training LLM Compression

ResearchDGX agent

arXiv:2606.03465v1 Announce Type: cross Abstract: Post-training compression is essential for deploying large language models (LLMs) under tight resource constraints. Tensor decompositions have emerged

Rex: A Family of Reversible Exponential (Stochastic) Runge-Kutta Solvers

ResearchDGX agent

arXiv:2502.08834v4 Announce Type: replace-cross Abstract: Deep generative models based on neural differential equations have become state-of-the-art for many generation tasks. These models rely on ODE

RGMem: Renormalization Group-inspired Memory Evolution for Language Agents

ResearchDGX agent

arXiv:2510.16392v3 Announce Type: replace Abstract: Personalized and continuous interactions are critical for LLM-based conversational agents, yet finite context windows and static parametric memory h

RobotValues: Evaluating Household Robots When Human Values Conflict

Model ReleasesDGX agent

arXiv:2606.03312v1 Announce Type: cross Abstract: While household robots are often evaluated based on task completion, everyday domestic environments involve value-conflicting situations in which robo

ROBUST-WT: Robust Uncertainty-aware Segmentation Transform via Whitening and Training Enhancements

Model ReleasesDGX agent

arXiv:2606.03069v1 Announce Type: cross Abstract: Generalized segmentation of medical images prevents performance degradation when different imaging devices and clinical protocols are used across mult

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability

Model ReleasesDGX agent

arXiv:2606.03648v1 Announce Type: cross Abstract: Adapting foundation large language models to a user's task or preferred style through fine-tuning can result in compromising the model's safety. Previ

SAGE: A Quantitative Evaluation of Socialized Evolution in Agent Ecosystems

AgentsDGX agent

arXiv:2606.03544v1 Announce Type: new Abstract: Self-improving language agents are typically evaluated in isolation: an agent attempts a task, receives feedback, and iteratively refines its own behavi

Samudra 2: Scaling Ocean Emulators across Resolutions

Model ReleasesDGX agent

arXiv:2606.02610v1 Announce Type: cross Abstract: Ocean general circulation models (OGCMs) are essential to climate science but computationally expensive, limiting ensemble size and forcing scenarios.

Scalable On-Hardware Training of Quantum Neural Networks and Application to Clinical Data Imputation

Model ReleasesDGX agent

arXiv:2606.03517v1 Announce Type: cross Abstract: Training quantum neural networks (QNNs) on quantum hardware is currently bottlenecked by the cost of gradient estimation: standard parameter-shift met

Scalable Uncertainty Quantification for Extreme Weather Forecasting via Empirical Neural Tangent Kernels

ResearchDGX agent

arXiv:2606.02886v1 Announce Type: cross Abstract: Deep learning weather models now match numerical weather prediction accuracy while running orders of magnitude faster, but produce deterministic forec

Scaling Multi Agent Reinforcement Learning for Underwater Acoustic Tracking via Autonomous Vehicles

HardwareDGX agent

arXiv:2505.08222v3 Announce Type: replace-cross Abstract: Autonomous vehicles (AVs) offer a cost-effective solution for scientific missions such as underwater tracking. Reinforcement learning (RL) has

SCOPE: Real-Time Natural Language Camera Agent at the Edge

Model ReleasesDGX agent

arXiv:2606.02951v1 Announce Type: cross Abstract: Deploying language-driven agents in robotics requires evaluations that reflect real-world task demands: natural-language instructions with reproducibl

scTranslation: A Comprehensive Benchmark for Single-Cell Multi-Omics Modality Translation

Model ReleasesDGX agent

arXiv:2606.03906v1 Announce Type: new Abstract: Simultaneous measurement of multiple omics modalities in single cells enables researchers to gain a more comprehensive understanding of cellular states

See Less, Specify More: Visual Evidence Budgets for Generalizable VLAs

SafetyDGX agent

arXiv:2606.02735v1 Announce Type: cross Abstract: Generalization remains a central bottleneck for vision-language-action (VLA) models: under distractors, appearance shifts, and semantically similar ta

SegTune: Structured and Fine-Grained Control for Song Generation

Local AiDGX agent

arXiv:2606.02638v1 Announce Type: cross Abstract: Recent advances in neural song generation have enabled high-quality synthesis from lyrics and global textual prompts. However, most systems fail to mo

Self-Refining Agentic Reinforcement Learning for Vision-Conditioned UAV Navigation

SafetyDGX agent

arXiv:2606.03963v1 Announce Type: cross Abstract: Deep reinforcement learning has shown strong potential for enabling autonomous robots to learn complex navigational tasks. However, its practical use

SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory

SafetyDGX agent

arXiv:2511.16275v4 Announce Type: replace-cross Abstract: Reliable uncertainty quantification (UQ) is essential for deploying large language models (LLMs) in safety-critical scenarios, as it enables t

Sign Lock-In: Randomly Initialized Weight Signs Persist and Bottleneck Sub-Bit Model Compression

ResearchDGX agent

arXiv:2602.17063v2 Announce Type: replace-cross Abstract: Sub-bit model compression targets storage below one bit per weight; as magnitudes are aggressively compressed, the sign bit becomes a fixed-co

Signed Spiking Neuron Enabled by an Orthogonal-Easy-Axis Magnetic Tunnel Junction

ResearchDGX agent

arXiv:2606.03796v1 Announce Type: cross Abstract: Signed spiking neurons carry richer information than standard spiking neurons. This work proposes a compact magnetic tunnel junction (MTJ)-based neuro

SkillDAG: Self-Evolving Typed Skill Graphs for LLM Skill Selection at Scale

Model ReleasesDGX agent

arXiv:2606.03056v1 Announce Type: new Abstract: As LLM agents adopt large skill libraries, selecting the right subset becomes a structural problem rather than a similarity-matching one: skills depend

SkillPyramid: A Hierarchical Skill Consolidation Framework for Self-Evolving Agents

ResearchDGX agent

arXiv:2606.03692v1 Announce Type: new Abstract: Recent AI agents can flexibly invoke skills to solve complex tasks, but their long-term improvement is fundamentally constrained by a lack of systematic

SleepVLM: Explainable and Rule-Grounded Sleep Staging via a Vision-Language Model

ResearchDGX agent

arXiv:2603.26738v3 Announce Type: replace-cross Abstract: While automated sleep staging has achieved expert-level accuracy, its clinical adoption is hindered by a lack of auditable reasoning. We intro

Solipsistic Superintelligence is Unlikely to be Cooperative

ApplicationsDGX agent

arXiv:2606.03237v1 Announce Type: new Abstract: AI's central challenge is shifting from capability to coexistence. The dominant paradigm in AI research focuses on developing powerful agents that treat

SPADE: Sketch-guided Path Planning Augmented with Diffusion Experts

AgentsDGX agent

arXiv:2606.03512v1 Announce Type: cross Abstract: Path planning is essential for Autonomous Mobile Robots (AMRs). Conventional methods for incorporating human preferences into planning typically rely

Sparse-View Lung Nodule Volumetry from Digitally Reconstructed Radiographs via AReT: Anatomy-Regularized TensoRF

SafetyDGX agent

arXiv:2606.02639v1 Announce Type: cross Abstract: We identify and resolve a previously unreported failure mode in TensoRF when applied to X-ray attenuation fields: the default density shift of -10, or

Spike-Aware C++ INT8 Inference for Sparse Spiking Language Models on Commodity CPUs

Model ReleasesDGX agent

arXiv:2606.03026v1 Announce Type: cross Abstract: Spiking language models expose activation sparsity that dense Transformer runtimes do not directly exploit. This paper studies that property from a sy

Staying Alive: Uncensored Survival Analysis with Tabular Foundation Models

Model ReleasesDGX agent

arXiv:2606.03689v1 Announce Type: cross Abstract: Survival Analysis (SA) is a statistical framework that models the time span until some event of interest occurs. Widely used in several domains, inclu

StepFinder: A Temporal Semantic Framework for Failure Attribution in Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2606.03467v1 Announce Type: new Abstract: LLM-based multi-agent systems exhibit remarkable collaborative capabilities in complex multi-step tasks. However, these systems are highly sensitive to

Strongly Polynomial Time Complexity of Policy Iteration for L_infty Robust MDPs

SafetyDGX agent

arXiv:2601.23229v2 Announce Type: replace Abstract: Markov decision processes (MDPs) are a fundamental model in sequential decision making. Robust MDPs (RMDPs) extend this framework by allowing uncert

SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models

Model ReleasesDGX agent

arXiv:2606.02642v1 Announce Type: cross Abstract: Despite the success of audio-visual large-language models (LLMs), they can produce plausible but ungrounded outputs, termed hallucination. Existing be

SynCred-Bench: Benchmarking Synthetic Credibility in AI-Generated Visual Misinformation

Model ReleasesDGX agent

arXiv:2606.03348v1 Announce Type: cross Abstract: Recent generative models can now produce visual artifacts with realistic embedded text and layouts, creating a new misinformation threat: synthetic cr

Synthesize and Reward -- Reinforcement Learning for Multi-Step Tool Use in Live Environments

AgentsDGX agent

arXiv:2606.03892v1 Announce Type: cross Abstract: Training LLMs to orchestrate multi-step tool calls is held back by three coupled obstacles: realistic stateful execution environments are costly to bu

TadA-Bench: A Million-Variant Benchmark for Future-Round Discovery Toward Agentic Protein Engineering

Model ReleasesDGX agent

arXiv:2606.02624v1 Announce Type: cross Abstract: AI for scientific discovery is entering an agentic era, where protein-engineering systems are expected to prioritize future wet-lab experiments rather

Taiji: Pareto Optimal Policy Optimization with Semantics-IDs Trade-off for Industrial LLM-Enhanced Recommendation

SafetyDGX agent

arXiv:2606.03866v1 Announce Type: cross Abstract: Scaling recommender systems via large language models (LLMs) has become a prominent trend in the industry. However, aligning the LLM's semantic space

TalkPlayData 2: An Agentic Synthetic Data Pipeline for Multimodal Conversational Music Recommendation

Model ReleasesDGX agent

arXiv:2509.09685v5 Announce Type: replace-cross Abstract: We present TalkPlayData 2, a synthetic dataset for multimodal conversational music recommendation generated by an agentic data pipeline. In th

Target Updates May Stabilize Linear Q-Learning: Periodic and Soft Dynamics

ResearchDGX agent

arXiv:2606.02645v1 Announce Type: cross Abstract: Periodic target updates in Q-learning and soft target updates in actor-critic methods are empirically well established stabilization mechanisms, but t

Test-Time Optimization of Physical Query Plans with LLMs

ResearchDGX agent

arXiv:2602.10387v2 Announce Type: replace-cross Abstract: Traditional query optimization relies on cost-based optimizers that estimate execution cost (e.g., runtime, memory, and I/O) using predefined

Testing LLM Arithmetic Reasoning Generalization with Automatic Numeric-Remapping Attacks

Model ReleasesDGX agent

arXiv:2606.03606v1 Announce Type: cross Abstract: Large language models achieve strong performance on arithmetic reasoning benchmarks, and one common response to arithmetic brittleness is to delegate

The Agent's First Day: Benchmarking Learning, Exploration, and Scheduling in the Workplace Scenarios

AgentsDGX agent

arXiv:2601.08173v2 Announce Type: replace Abstract: The rapid evolution of Multi-modal Large Language Models (MLLMs) has advanced workflow automation; however, existing research mainly targets perform

The DeepSpeak-Agentic Dataset

Model ReleasesDGX agent

arXiv:2606.03686v1 Announce Type: new Abstract: We present DeepSpeak-Agentic, a dataset of videos comprising over 37 hours of semi-structured conversations between a human and an embodied AI agent. We

The Epi-LLM Framework: probing LLM behavioral priors through epidemiological agent-based models

AgentsDGX agent

arXiv:2606.02867v1 Announce Type: cross Abstract: Human behaviour during epidemics affects infectious disease dynamics, but quantifying this remains deeply challenging. Here we introduce the Epi-LLM f

The Impact of Configuring Agentic AI Coding Tools on Build-vs-Buy Decisions: A Study Protocol

Model ReleasesDGX agent

arXiv:2606.03907v1 Announce Type: cross Abstract: Agentic AI coding tools write code with increasing autonomy and in doing so decide when to import a library and when to implement functionality from s

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection

Model ReleasesDGX agent

arXiv:2606.03305v1 Announce Type: new Abstract: Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment. Statistical tools for de

The Ringelmann Effect in Multi-Agent LLM Systems: A Scaling Law for Effective Team Size

Model ReleasesDGX agent

arXiv:2606.02646v1 Announce Type: cross Abstract: Inference-time multi-agent LLM scaling lacks a shared unit: counting nominal agents conflates cost with independent evidence. We derive a two-paramete

The Shadow Price of Reasoning: Economic Perspective on Optimal Budget Allocation for LLMs

SafetyDGX agent

arXiv:2606.03092v1 Announce Type: new Abstract: Inference-time scaling has emerged as a critical avenue for enhancing Large Language Models' performance, yet real-world deployment is constrained by st

The Shape of Addition: Geometric Structures of Arithmetic in Large Language Models

ResearchDGX agent

arXiv:2606.03645v1 Announce Type: cross Abstract: Large Language Models exhibit paradoxical fragility in fundamental arithmetic, implying a disconnect between internal computation and discrete output.

The Unsampled Truth: Psychometrics in SLMs Measure Prompt Artifacts, Not Psychological Constructs

ResearchDGX agent

arXiv:2606.03357v1 Announce Type: cross Abstract: When prompting SLMs for psychometric assessments, researchers assume the outputs reflect semantic reasoning. We evaluate this premise across 13 open-w

The Violation Situation Pattern: A Knowledge-Graph Pattern for Compliance Violations

ApplicationsDGX agent

arXiv:2606.03326v1 Announce Type: new Abstract: Compliance pipelines detect violations as transient query results and do not keep the violation itself as a persistent graph object with review state, a

Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation

SafetyDGX agent

arXiv:2606.03137v1 Announce Type: new Abstract: LLM-based multi-agent simulation offers a promising way to study social interaction, deliberation, and collective opinion dynamics. However, many existi

Thinking Past the Answer: Evaluating Harmful Overthinking in Large Reasoning Models

ResearchDGX agent

arXiv:2606.02835v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) improve performance by generating explicit intermediate reasoning traces through increased test-time compute, yet the assu

ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning

Model ReleasesDGX agent

arXiv:2606.03503v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) have achieved remarkable progress thanks to Reinforcement Learning with Verifiable Rewards (RLVR) on Chain-of-Thoughts (Co

TimeOmni-VL: Unified Models for Time Series Understanding and Generation

ResearchDGX agent

arXiv:2602.17149v2 Announce Type: replace-cross Abstract: Recent time series modeling faces a sharp divide between numerical generation and semantic understanding, with research showing that generatio

Tonal parsimony in chord-sequence analysis: combining modulation cost and tonal vocabulary

ResearchDGX agent

arXiv:2606.03459v1 Announce Type: cross Abstract: We study the assignment of local tonalities to chord sequences, a task useful for harmonic analysis, composition, and jazz-oriented improvisation. Sta

Too Much of a Good Thing: When sim2real Efforts Impede Policy Learning (And What to Do About It)

SafetyDGX agent

arXiv:2606.02636v1 Announce Type: cross Abstract: While sim2real efforts are necessary for effective policy transfer to hardware, there is such a thing as too much of a good thing. We argue that sim2r

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2606.03762v1 Announce Type: cross Abstract: Agentic reinforcement learning (RL) equips large language models (LLMs) with tool-use capabilities that substantially improve reasoning on complex tas

ToolGate: Token-Efficient Pre-Call Control for Tool-Augmented Vision-Language Agents

AgentsDGX agent

arXiv:2606.03054v1 Announce Type: new Abstract: Tool-augmented vision-language agents can acquire external perceptual evidence through OCR, detection, segmentation, and other tools, but executing ever

← Previous
1…167168169170171…358
Next →