AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-lg”

GridTimelineEvolution
14,329 results
10 Apr 2026

FlowAdam: Implicit Regularization via Geometry-Aware Soft Momentum Injection

Model ReleasesDGX agent

arXiv:2604.06652v1 Announce Type: new Abstract: Adaptive moment methods such as Adam use a diagonal, coordinate-wise preconditioner based on exponential moving averages of squared gradients. This diag

ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache

HardwareDGX agent

arXiv:2604.06370v1 Announce Type: cross Abstract: The serving paradigm of large language models (LLMs) is rapidly shifting towards complex multi-agent workflows where specialized agents collaborate ov

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start

HardwareDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.06664v1 Announce Type: cross Abstract: Modern LLM service providers increasingly rely on autoscaling and parallelism reconfiguration to respond to rapidly changing workloads, but cold-start

From Synthetic Data to Real Restorations: Diffusion Model for Patient-specific Dental Crown Completion

Local AiDGX agent

arXiv:2603.26588v2 Announce Type: replace-cross Abstract: We present ToothCraft, a diffusion-based model for the contextual generation of tooth crowns, trained on artificially created incomplete teeth

Gaussian Approximation for Asynchronous Q-learning

ResearchDGX agent

arXiv:2604.07323v1 Announce Type: cross Abstract: In this paper, we derive rates of convergence in the high-dimensional central limit theorem for Polyak-Ruppert averaged iterates generated by the asyn

Geometric Properties of the Voronoi Tessellation in Latent Semantic Manifolds of Large Language Models

SafetyDGX agent

arXiv:2604.06767v1 Announce Type: new Abstract: Language models operate on discrete tokens but compute in continuous vector spaces, inducing a Voronoi tessellation over the representation manifold. We

GIFT: Group-Relative Implicit Fine-Tuning Integrates GRPO with DPO and UNA

SafetyDGX agent

arXiv:2510.23868v4 Announce Type: replace Abstract: This paper proposes extit{Group-relative Implicit Fine-Tuning (GIFT)}, a reinforcement learning framework for aligning large language models (LLMs

Graph Neural ODE Digital Twins for Control-Oriented Reactor Thermal-Hydraulic Forecasting Under Partial Observability

HardwareDGX agent

arXiv:2604.07292v1 Announce Type: new Abstract: Real-time supervisory control of advanced reactors requires accurate forecasting of plant-wide thermal-hydraulic states, including locations where physi

GraphWalker: Graph-Guided In-Context Learning for Clinical Reasoning on Electronic Health Records

ApplicationsDGX agent

arXiv:2604.06684v1 Announce Type: new Abstract: Clinical Reasoning on Electronic Health Records (EHRs) is a fundamental yet challenging task in modern healthcare. While in-context learning (ICL) offer

Holistic Optimal Label Selection for Robust Prompt Learning under Partial Labels

Model ReleasesDGX agent

arXiv:2604.06614v1 Announce Type: cross Abstract: Prompt learning has gained significant attention as a parameter-efficient approach for adapting large pre-trained vision-language models to downstream

How Does Machine Learning Manage Complexity?

ResearchDGX agent

arXiv:2604.07233v1 Announce Type: new Abstract: We provide a computational complexity lens to understand the power of machine learning models, particularly their ability to model complex systems. Mach

How to sketch a learning algorithm

TutorialsDGX agent

arXiv:2604.07328v1 Announce Type: new Abstract: How does the choice of training data influence an AI model? This question is of central importance to interpretability, privacy, and basic science. At i

Improving Semantic Uncertainty Quantification in Language Model Question-Answering via Token-Level Temperature Scaling

SafetyDGX agent

arXiv:2604.07172v1 Announce Type: new Abstract: Calibration is central to reliable semantic uncertainty quantification, yet prior work has largely focused on discrimination, neglecting calibration. As

Inference-Time Scaling of Diffusion Language Models via Trajectory Refinement

TutorialsDGX agent

arXiv:2507.08390v4 Announce Type: replace Abstract: Discrete diffusion models have recently emerged as strong alternatives to autoregressive language models, matching their performance through large-s

Interventional Time Series Priors for Causal Foundation Models

ResearchDGX agent

arXiv:2603.11090v2 Announce Type: replace Abstract: Prior-data fitted networks (PFNs) have emerged as powerful foundation models for tabular causal inference, yet their extension to time series remain

Learning Debt and Cost-Sensitive Bayesian Retraining: A Forecasting Operations Framework

Model ReleasesDGX agent

arXiv:2604.06438v1 Announce Type: cross Abstract: Forecasters often choose retraining schedules by convention rather than by an explicit decision rule. This paper gives that decision a posterior-space

Learning to Query History: Nonstationary Classification via Learned Retrieval

Local AiDGX agent

arXiv:2604.07027v1 Announce Type: new Abstract: Nonstationarity is ubiquitous in practical classification settings, leading deployed models to perform poorly even when they generalize well to holdout

Limits of Difficulty Scaling: Hard Samples Yield Diminishing Returns in GRPO-Tuned SLMs

SafetyDGX agent

arXiv:2604.06298v1 Announce Type: new Abstract: Recent alignment work on Large Language Models (LLMs) suggests preference optimization can improve reasoning by shifting probability mass toward better

LNN-PINN: A Unified Physics-Only Training Framework with Liquid Residual Blocks

Model ReleasesDGX agent

arXiv:2508.08935v4 Announce Type: replace Abstract: Physics-informed neural networks (PINNs) have attracted considerable attention for their ability to integrate partial differential equation priors i

LoFT: Parameter-Efficient Fine-Tuning for Long-tailed Semi-Supervised Learning in Open-World Scenarios

Model ReleasesDGX agent

arXiv:2509.09926v5 Announce Type: replace Abstract: Long-tailed semi-supervised learning (LTSSL) presents a formidable challenge where models must overcome the scarcity of tail samples while mitigatin

Low-Rank Key Value Attention

ResearchDGX agent

arXiv:2601.11471v3 Announce Type: replace Abstract: The key-value (KV) cache is a primary memory bottleneck in Transformers. We propose Low-Rank Key-Value (LRKV) attention, which reduces KV cache memo

Lumbermark: Resistant Clustering by Chopping Up Mutual Reachability Minimum Spanning Trees

Model ReleasesDGX agent

arXiv:2604.07143v1 Announce Type: new Abstract: We introduce Lumbermark, a robust divisive clustering algorithm capable of detecting clusters of varying sizes, densities, and shapes. Lumbermark iterat

LUMINA: Foundation Models for Topology Transferable ACOPF

SafetyDGX agent

arXiv:2603.04300v2 Announce Type: replace Abstract: Foundation models in general promise to accelerate scientific computation by learning reusable representations across problem instances, yet constra

Matrix Profile for Time-Series Anomaly Detection: A Reproducible Open-Source Benchmark on TSB-AD

Model ReleasesDGX agent

arXiv:2604.02445v2 Announce Type: replace Abstract: Matrix Profile (MP) methods are an interpretable and scalable family of distance-based methods for time-series anomaly detection, but strong benchma

MDP modeling for multi-stage stochastic programs

SafetyDGX agent

arXiv:2509.22981v2 Announce Type: replace Abstract: We study a class of multi-stage stochastic programs, which incorporate modeling features from Markov decision processes (MDPs). This class includes

Measurement of Generative AI Workload Power Profiles for Whole-Facility Data Center Infrastructure Planning

HardwareDGX agent

arXiv:2604.07345v1 Announce Type: cross Abstract: The rapid growth of generative artificial intelligence (AI) has introduced unprecedented computational demands, driving significant increases in the e

MedRoute: RL-Based Dynamic Specialist Routing in Multi-Agent Medical Diagnosis

AgentsDGX agent

arXiv:2604.06180v1 Announce Type: cross Abstract: Medical diagnosis using Large Multimodal Models (LMMs) has gained increasing attention due to capability of these models in providing precise diagnose

MENO: MeanFlow-Enhanced Neural Operators for Dynamical Systems

ResearchDGX agent

arXiv:2604.06881v1 Announce Type: new Abstract: Neural operators have emerged as powerful surrogates for dynamical systems due to their grid-invariant properties and computational efficiency. However,

MF-GLaM: A multifidelity stochastic emulator using generalized lambda models

Model ReleasesDGX agent

arXiv:2507.10303v2 Announce Type: replace-cross Abstract: Stochastic simulators exhibit intrinsic stochasticity due to unobservable, uncontrollable, or unmodeled input variables, resulting in random o

MICA: Multivariate Infini Compressive Attention for Time Series Forecasting

ResearchDGX agent

arXiv:2604.06473v1 Announce Type: new Abstract: Multivariate forecasting with Transformers faces a core scalability challenge: modeling cross-channel dependencies via attention compounds attention's q

Mining Electronic Health Records to Investigate Effectiveness of Ensemble Deep Clustering

ApplicationsDGX agent

arXiv:2604.07085v1 Announce Type: new Abstract: In electronic health records (EHRs), clustering patients and distinguishing disease subtypes are key tasks to elucidate pathophysiology and aid clinical

MoE Routing Testbed: Studying Expert Specialization and Routing Behavior at Small Scale

ResearchDGX agent

arXiv:2604.07030v1 Announce Type: new Abstract: Sparse Mixture-of-Experts (MoE) architectures are increasingly popular for frontier large language models (LLM) but they introduce training challenges d

Multi-Turn Reasoning LLMs for Task Offloading in Mobile Edge Computing

SafetyDGX agent

arXiv:2604.07148v1 Announce Type: new Abstract: Emerging computation-intensive applications impose stringent latency requirements on resource-constrained mobile devices. Mobile Edge Computing (MEC) ad

NativeTernary: A Self-Delimiting Binary Encoding with Unary Run-Length Hierarchy Markers for Ternary Neural Network Weights, Structured Data, and General Computing Infrastructure

Local AiDGX agent

arXiv:2604.03336v2 Announce Type: replace Abstract: BitNet b1.58 (Ma et al., 2024) demonstrates that large language models can operate entirely on ternary weights {-1, 0, +1}, yet no native binary wir

Negative Binomial Variational Autoencoders for Overdispersed Latent Modeling

Model ReleasesDGX agent

arXiv:2508.05423v2 Announce Type: replace Abstract: Although artificial neural networks are often described as brain-inspired, their representations typically rely on continuous activations, such as t

NestPipe: Large-Scale Recommendation Training on 1,500+ Accelerators via Nested Pipelining

HardwareDGX agent

arXiv:2604.06956v1 Announce Type: cross Abstract: Modern recommendation models have increased to trillions of parameters. As cluster scales expand to O(1k), distributed training bottlenecks shift from

Neural parametric representations for thin-shell shape optimisation

Model ReleasesDGX agent

arXiv:2604.06612v1 Announce Type: cross Abstract: Shape optimisation of thin-shell structures requires a flexible, differentiable geometric representation suitable for gradient-based optimisation. We

Neural Two-Stage Stochastic Optimization for Solving Unit Commitment Problem

ResearchDGX agent

arXiv:2507.09503v2 Announce Type: replace-cross Abstract: This paper proposes a neural stochastic optimization method for efficiently solving the two-stage stochastic unit commitment (2S-SUC) problem

Non-Expansive Mappings in Two-Time-Scale Stochastic Approximation: Finite-Time Analysis

ResearchDGX agent

arXiv:2501.10806v4 Announce Type: replace-cross Abstract: Two-time-scale stochastic approximation algorithms are iterative methods used in applications such as optimization, reinforcement learning, an

Non-identifiability of Explanations from Model Behavior in Deep Networks of Image Authenticity Judgments

ResearchDGX agent

arXiv:2604.07254v1 Announce Type: cross Abstract: Deep neural networks can predict human judgments, but this does not imply that they rely on human-like information or reveal the cues underlying those

Nonparametric Instrumental Regression via Kernel Methods is Minimax Optimal

ResearchDGX agent

arXiv:2411.19653v2 Announce Type: replace-cross Abstract: We study the kernel instrumental variable (KIV) algorithm, a kernel-based two-stage least-squares method for nonparametric instrumental variab

ODE-free Neural Flow Matching for One-Step Generative Modeling

TutorialsDGX agent

arXiv:2604.06413v1 Announce Type: new Abstract: Diffusion and flow matching models generate samples by learning time-dependent vector fields whose integration transports noise to data, requiring tens

On the Price of Privacy for Language Identification and Generation

ResearchDGX agent

arXiv:2604.07238v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly trained on sensitive user data, understanding the fundamental cost of privacy in language learning beco

Operator Learning for Surrogate Modeling of Wave-Induced Forces from Sea Surface Waves

ResearchDGX agent

arXiv:2604.06433v1 Announce Type: cross Abstract: Wave setup plays a significant role in transferring wave-induced energy to currents and causing an increase in water elevation. This excess momentum f

Optimal Rates for Pure {arepsilon}-Differentially Private Stochastic Convex Optimization with Heavy Tails

Model ReleasesDGX agent

arXiv:2604.06492v1 Announce Type: new Abstract: We study stochastic convex optimization (SCO) with heavy-tailed gradients under pure epsilon-differential privacy (DP). Instead of assuming a bound on t

PAC-Bayesian Bounds on Constrained f-Entropic Risk Measures

ResearchDGX agent

arXiv:2510.11169v2 Announce Type: replace-cross Abstract: PAC generalization bounds on the risk, when expressed in terms of the expected loss, are often insufficient to capture imbalances between subg

PD-SOVNet: A Physics-Driven Second-Order Vibration Operator Network for Estimating Wheel Polygonal Roughness from Axle-Box Vibrations

ApplicationsDGX agent

arXiv:2604.06620v1 Announce Type: new Abstract: Quantitative estimation of wheel polygonal roughness from axle-box vibration signals is a challenging yet practically relevant problem for rail-vehicle

Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization

Model ReleasesDGX agent

arXiv:2604.07343v1 Announce Type: cross Abstract: Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a cen

Physics-Informed Functional Link Constrained Framework with Domain Mapping for Solving Bending Analysis of an Exponentially Loaded Perforated Beam

Model ReleasesDGX agent

arXiv:2604.07025v1 Announce Type: cross Abstract: This article presents a novel and comprehensive approach for analyzing bending behavior of the tapered perforated beam under an exponential load. The

Physics-Informed Neural Networks for Joint Source and Parameter Estimation in Advection-Diffusion Equations

Model ReleasesDGX agent

arXiv:2512.07755v3 Announce Type: replace-cross Abstract: Recent studies have demonstrated the success of deep learning in solving forward and inverse problems in engineering and scientific computing

Predictive Representations for Skill Transfer in Reinforcement Learning

AgentsDGX agent

arXiv:2604.07016v1 Announce Type: new Abstract: A key challenge in scaling up Reinforcement Learning is generalizing learned behaviour. Without the ability to carry forward acquired knowledge an agent

Probabilistic Predictions of Process-Induced Deformation in Carbon/Epoxy Composites Using a Deep Operator Network

ApplicationsDGX agent

arXiv:2512.13746v4 Announce Type: replace-cross Abstract: Fiber reinforcement and polymer matrix respond differently to manufacturing conditions due to mismatch in coefficient of thermal expansion and

Production-Ready Automated ECU Calibration using Residual Reinforcement Learning

ApplicationsDGX agent

arXiv:2604.07059v1 Announce Type: new Abstract: Electronic Control Units (ECUs) have played a pivotal role in transforming motorcars of yore into the modern vehicles we see on our roads today. They ac

QNAS: A Neural Architecture Search Framework for Accurate and Efficient Quantum Neural Networks

Model ReleasesDGX agent

arXiv:2604.07013v1 Announce Type: cross Abstract: Designing quantum neural networks (QNNs) that are both accurate and deployable on NISQ hardware is challenging. Handcrafted ansatze must balance expre

Quality-preserving Model for Electronics Production Quality Tests Reduction

SafetyDGX agent

arXiv:2604.06451v1 Announce Type: new Abstract: Manufacturing test flows in high-volume electronics production are typically fixed during product development and executed unchanged on every unit, even

Quantum-Inspired Tensor Network Autoencoders for Anomaly Detection: A MERA-Based Approach

Model ReleasesDGX agent

arXiv:2604.06541v1 Announce Type: cross Abstract: We investigate whether a multiscale tensor-network architecture can provide a useful inductive bias for reconstruction-based anomaly detection in coll

RAGEN-2: Reasoning Collapse in Agentic RL

AgentsDGX agent

arXiv:2604.06268v1 Announce Type: new Abstract: RL training of multi-turn LLM agents is inherently unstable, and reasoning quality directly determines task performance. Entropy is widely used to track

ReDAct: Uncertainty-Aware Deferral for LLM Agents

AgentsDGX agent

arXiv:2604.07036v1 Announce Type: cross Abstract: Recently, LLM-based agents have become increasingly popular across many applications, including complex sequential decision-making problems. However,

Replacing Tunable Parameters in Weather and Climate Models with State-Dependent Functions using Reinforcement Learning

Model ReleasesDGX agent

arXiv:2601.04268v2 Announce Type: replace Abstract: Weather and climate models rely on parametrisations to represent unresolved sub-grid processes. Traditional schemes rely on fixed coefficients that

Resistance Distance and Linearized Optimal Transport on Graphs

ResearchDGX agent

arXiv:2404.15261v4 Announce Type: replace-cross Abstract: We study the linearization of a discrete transportation distance between probability distributions on finite weighted graphs originally due to

← Previous
1…236237238239
Next →