AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
HumanDGX agent

87,171Total entries
1Added by human
87,170Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,612 results
12 May 2026

Beyond the Singular: Revealing the Value of Multiple Generations in Benchmark Evaluation

Model ReleasesDGX agent

arXiv:2502.08943v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have demonstrated significant utility in real-world applications, exhibiting impressive capabilities in natural l

CERSA: Cumulative Energy-Retaining Subspace Adaptation for Memory-Efficient Fine-Tuning

Model ReleasesDGX agent

arXiv:2605.08174v1 Announce Type: cross Abstract: To mitigate the memory constraints associated with fine-tuning large pre-trained models, existing parameter-efficient fine-tuning (PEFT) methods, such

Concordia: Self-Improving Synthetic Tables for Federated LLMs

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.09855v1 Announce Type: new Abstract: Federated learning (FL) enables training large language models (LLMs) without sharing raw data, but adapting LLMs under strict data isolation and non-II

ConQuR: Corner Aligned Activation Quantization via Optimized Rotations for LLMs

Model ReleasesDGX agent

arXiv:2605.10793v1 Announce Type: new Abstract: Large language models (LLMs) are costly to deploy due to their large memory footprint and high inference cost. Weight-activation quantization can reduce

Containment Verification: AI Safety Guarantees Independent of Alignment

SafetyDGX agent

arXiv:2605.09045v1 Announce Type: new Abstract: Agentic frameworks are the software layer through which AI agents act in the world. Existing safety methods intervene on the model and therefore remain

Continuous Latent Contexts Enable Efficient Online Learning in Transformers

Model ReleasesDGX agent

arXiv:2605.09867v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit a strong capacity for in-context learning: Given labeled examples, they can generate good predictions without par

DeepLevy: Learning Heavy-Tailed Uncertainty in Highly Volatile Time Series

ResearchDGX agent

arXiv:2605.10364v1 Announce Type: new Abstract: Modeling uncertainty in heavy-tailed time series remains a critical challenge for deep probabilistic forecasting models, which often struggle to capture

DiagnosticIQ: A Benchmark for LLM-Based Industrial Maintenance Action Recommendation from Symbolic Rules

Model ReleasesDGX agent

arXiv:2605.08614v1 Announce Type: new Abstract: Monitoring complex industrial assets relies on engineer-authored symbolic rules that trigger based on sensor conditions and prompt technicians to perfor

Do Self-Evolving Agents Forget? Capability Degradation and Preservation in Lifelong LLM Agent Adaptation

Model ReleasesDGX agent

arXiv:2605.09315v1 Announce Type: new Abstract: Recent advances in LLM agents enable systems that autonomously refine workflows, accumulate reusable skills, self-train their underlying models, and mai

Domain-Adaptive Arrhythmia Classification Using a Hybrid Transformer on Wearable Heart Signals

ResearchDGX agent

arXiv:2605.08199v1 Announce Type: cross Abstract: Cardiovascular disease remains the leading cause of death globally, underscoring the need for effective, accessible monitoring solutions, particularly

Drift is a Sampling Error: SNR-Aware Power Distributions for Long-Horizon Robotic Planning

Model ReleasesDGX agent

arXiv:2605.09537v1 Announce Type: new Abstract: Despite rapid progress in Vision-Language-Action (VLA) models for robotic control, instruction drift remains a persistent failure mode in long-horizon t

DRNet: All-in-One Image Restoration via Prior-Guided Dynamic Reparameterization

Model ReleasesDGX agent

arXiv:2605.08627v1 Announce Type: new Abstract: All-in-one image restoration aims to handle diverse degradations within a single model. However, existing methods often suffer from three key limitation

EAM: Enhancing Anything with Diffusion Transformers for Blind Super-Resolution

TutorialsDGX agent

arXiv:2505.05209v4 Announce Type: replace Abstract: Utilizing pre-trained Text-to-Image (T2I) diffusion models to guide Blind Super-Resolution (BSR) has become a predominant approach in the field. Whi

Echo-LoRA: Parameter-Efficient Fine-Tuning via Cross-Layer Representation Injection

Model ReleasesDGX agent

arXiv:2605.08177v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning (PEFT) has become a practical route for adapting large language models to downstream tasks, with LoRA-style methods be

EchoAlign: Bridging Generative and Discriminative Learning under Noisy Labels

Model ReleasesDGX agent

arXiv:2405.12969v3 Announce Type: replace Abstract: Noisy labels severely hinder the accuracy and generalization of machine learning models, especially when ambiguous instance features make reliable a

ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning

ApplicationsDGX agent

arXiv:2605.09982v1 Announce Type: new Abstract: Recent advancements in Vision-Language Models (VLMs) enable large language models (LLMs) to process high-resolution images, significantly improving real

Explainable Knowledge Tracing via Probabilistic Embeddings and Pattern-based Reasoning

ResearchDGX agent

arXiv:2605.09369v1 Announce Type: new Abstract: Knowledge Tracing (KT) models students' knowledge states based on learning interactions to predict performance. While deep learning-based KT models have

Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition

Model ReleasesDGX agent

arXiv:2605.10127v1 Announce Type: new Abstract: Recent research work on fashion outfit generation focuses on promoting visual consistency of garments by leveraging key information from reference image

Function-Space ADMM for Decentralized Federated Learning: A Control Theoretic Perspective

Model ReleasesDGX agent

arXiv:2605.09356v1 Announce Type: new Abstract: Decentralized federated learning (FL) is a promising approach for training machine learning models on sensor networks, Internet of Things (IoT) devices,

Generative Actor-Critic with Soft Bridge Policies

Model ReleasesDGX agent

arXiv:2605.08733v1 Announce Type: new Abstract: Expressive generative policies such as diffusion and flow models are appealing for MaxEnt online reinforcement learning because of their ability to mode

Geometry Guided Self-Consistency for Physical AI

ResearchDGX agent

arXiv:2605.08638v1 Announce Type: cross Abstract: State-of-the-art physical AI models generate a chunk of actions per inference through diffusion or flow matching, iteratively refining an initial nois

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs

Model ReleasesDGX agent

arXiv:2508.20325v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) become increasingly integral to various domains, their potential to generate harmful responses has prompted si

How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue

ResearchDGX agent

arXiv:2605.10199v1 Announce Type: new Abstract: Full-duplex spoken dialogue requires a model to keep listening while generating its own spoken response. This is challenging for large language models (

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs

Model ReleasesDGX agent

arXiv:2605.09635v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in K-12 education, yet existing benchmarks such as C-Eval, CMMLU, GaokaoBench, and EduEval mainly eva

Kaczmarz Linear Attention

ResearchDGX agent

arXiv:2605.08587v1 Announce Type: cross Abstract: Long-context language modeling remains central to modern sequence modeling, but the quadratic cost of Transformer attention makes scaling computationa

LE-PAVD: Learning-Enhanced Physics-Aware Vehicle Dynamics for High-Speed Autonomous Navigation

AgentsDGX agent

arXiv:2605.08489v1 Announce Type: new Abstract: Accurate modeling of nonlinear vehicle dynamics is essential for high-speed autonomous racing, where controllers operate at the handling limits. Model-b

LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs?

Model ReleasesDGX agent

arXiv:2605.08985v1 Announce Type: new Abstract: Visual encoding constitutes a major computational bottleneck in Multimodal Large Language Models (MLLMs), especially for high-resolution image inputs. T

LLM Wardens: Mitigating Adversarial Persuasion with Third-Party Conversational Oversight

Model ReleasesDGX agent

arXiv:2605.08321v1 Announce Type: cross Abstract: LLMs are increasingly capable of persuasion, which raises the question of how to protect users against manipulation. In a preregistered user study (N=

Multi-Tier Labeling and Physics-Informed Learning for Orbital Anomaly Detection at Scale

Model ReleasesDGX agent

arXiv:2605.09790v1 Announce Type: cross Abstract: Detecting orbital anomalies, such as maneuvers, atmospheric decay, and attitude upsets, across the rapidly growing population of low-Earth-orbit (LEO)

On Distinguishing Capability Elicitation from Capability Creation in Post-Training: A Free-Energy Perspective

Local AiDGX agent

arXiv:2605.08368v1 Announce Type: new Abstract: Debates about large language model post-training often treat supervised fine-tuning (SFT) as imitation and reinforcement learning (RL) as discovery. But

OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents

Model ReleasesDGX agent

arXiv:2605.08876v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents that execute tool-augmented, multi-step tasks, where latency is a critical f

Parameter-Efficient Neuroevolution for Diverse LLM Generation: Quality-Diversity Optimization via Prompt Embedding Evolution

Model ReleasesDGX agent

arXiv:2605.09781v1 Announce Type: cross Abstract: Large Language Models exhibit mode collapse, producing homogeneous outputs that fail to explore valid solution spaces. We present QD-LLM, a framework

PLACO: A Multi-Stage Framework for Cost-Effective Performance in Human-AI Teams

ResearchDGX agent

arXiv:2605.08388v1 Announce Type: new Abstract: Human-AI teams play a pivotal role in improving overall system performance when neither the human nor the model can achieve such performance on their ow

PlantMarkerBench: A Multi-Species Benchmark for Evidence-Grounded Plant Marker Reasoning

Model ReleasesDGX agent

arXiv:2605.10032v1 Announce Type: new Abstract: Cell-type-specific marker genes are fundamental to plant biology, yet existing resources primarily rely on curated databases or high-throughput studies

PrAg-PO: Prompt Augmented Policy Optimization for Robust and Diverse Mathematical Reasoning

Model ReleasesDGX agent

arXiv:2602.03190v3 Announce Type: replace-cross Abstract: Reinforcement learning algorithms such as group-relative policy optimization (GRPO) have shown strong potential for improving the mathematical

Preserving Foundational Capabilities in Flow-Matching VLAs through Conservative SFT

Model ReleasesDGX agent

arXiv:2605.08879v1 Announce Type: new Abstract: Unconstrained fine-tuning of flow-matching Vision-Language-Action (VLA) models drives dense parameter overwrites, degrading pre-trained capabilities. We

Process Matters more than Output for Distinguishing Humans from Machines

Model ReleasesDGX agent

arXiv:2605.06524v2 Announce Type: replace Abstract: Reliable human-machine discrimination is becoming increasingly important as large language models and autonomous agents are deployed in online setti

Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions

Model ReleasesDGX agent

arXiv:2605.10664v1 Announce Type: cross Abstract: Activation steering controls language model behavior by adding directions to internal representations at inference time, but standard residual-stream

Provable Sparse Inversion and Token Relabel Enhanced One-shot Federated Learning with ViTs

ResearchDGX agent

arXiv:2605.10748v1 Announce Type: cross Abstract: One-Shot Federated Learning, where a central server learns a global model in a single communication round, has emerged as a promising paradigm. Howeve

Quantum Transfer Learning Shows Improved Robustness in Low-Data Regimes

ResearchDGX agent

arXiv:2605.09118v1 Announce Type: cross Abstract: Transfer learning under limited data is a challenging setting, where models must adapt to new tasks with minimal supervision. Prior work has primarily

Real vs. Semi-Simulated: Rethinking Evaluation for Treatment Effect Estimation

Model ReleasesDGX agent

arXiv:2605.10430v1 Announce Type: cross Abstract: Estimating heterogeneous treatment effects with machine learning has attracted substantial attention in both academic research and industrial practice

Reasoning Compression with Mixed-Policy Distillation

SafetyDGX agent

arXiv:2605.08776v1 Announce Type: new Abstract: Reasoning-centric large language models (LLMs) achieve strong performance by generating intermediate reasoning trajectories, but often incur excessive t

Robust Server Defense Against Unreliable Clients in One-Shot Fair Collaborative Machine Learning

Model ReleasesDGX agent

arXiv:2605.08616v1 Announce Type: new Abstract: Collaborative machine learning (CML) enables multiple clients to train a global model jointly in a data-distributed setting. To address data privacy and

ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention

Model ReleasesDGX agent

arXiv:2603.22016v2 Announce Type: replace-cross Abstract: Large Reasoning Models (LRMs) often reach a correct solution before their long Chain-of-Thought trace ends, yet continue with redundant verifi

Route Before Retrieve: Activating Latent Routing Abilities of LLMs for RAG vs. Long-Context Selection

Model ReleasesDGX agent

arXiv:2605.10235v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have expanded the context window to beyond 128K tokens, enabling long-document understanding and multi-s

S2FT: Parameter-Efficient Fine-Tuning in Sparse Spectrum Domain

Model ReleasesDGX agent

arXiv:2605.08589v1 Announce Type: new Abstract: Parameter Efficient Fine-Tuning (PEFT) is a key technique for adapting a large pretrained model to downstream tasks by fine-tuning only a small number o

SayNext-Bench: Why Do LLMs Struggle with Next-Utterance Anticipation?

Model ReleasesDGX agent

arXiv:2602.00327v2 Announce Type: replace Abstract: We explore the use of large language models (LLMs) for next-utterance anticipation in human dialogue. Despite recent advances in LLMs demonstrating

SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation

Model ReleasesDGX agent

arXiv:2605.10187v1 Announce Type: new Abstract: Scientific reasoning is a key aspect of human intelligence, requiring the integration of multimodal inputs, domain expertise, and multi-step inference a

SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning

Model ReleasesDGX agent

arXiv:2605.09266v1 Announce Type: new Abstract: We introduce SeePhys Pro, a fine-grained modality transfer benchmark that studies whether models preserve the same reasoning capability when critical in

SleepWalk: A Three-Tier Benchmark for Stress-Testing Instruction-Guided Vision-Language Navigation

Model ReleasesDGX agent

arXiv:2605.10376v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have advanced rapidly in multimodal perception and language understanding, yet it remains unclear whether they can reliabl

SlimSpec: Low-Rank Draft LM-Head for Accelerated Speculative Decoding

ResearchDGX agent

arXiv:2605.10453v1 Announce Type: cross Abstract: Speculative decoding speeds up autoregressive generation in Large Language Models (LLMs) through a two-step procedure, where a lightweight draft model

Steerable but Not Decodable: Function Vectors Operate Beyond the Logit Lens

Model ReleasesDGX agent

arXiv:2604.02608v2 Announce Type: replace Abstract: Activation steering presupposes that task-relevant behaviors correspond to linear directions in activation space -- directions that should both stee

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution

Model ReleasesDGX agent

arXiv:2605.08366v1 Announce Type: new Abstract: We introduce SWE Atlas, a benchmark suite for coding agents spanning three professional software engineering workflows: Codebase Q&A (124 tasks), Test W

TAD: Temporal-Aware Trajectory Self-Distillation for Fast and Accurate Diffusion LLM

ResearchDGX agent

arXiv:2605.09536v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) offer a promising paradigm for parallel text generation, but in practice they face an accuracy-parallelism tra

Temporal Sampling Frequency Matters: A Capacity-Aware Study of End-to-End Driving Trajectory Prediction

AgentsDGX agent

arXiv:2605.10388v1 Announce Type: new Abstract: End to end (E2E) autonomous driving trajectory prediction is often trained with camera frames sampled at the highest available temporal frequency, assum

The Differences Between Direct Alignment Algorithms are a Blur

Model ReleasesDGX agent

arXiv:2502.01237v3 Announce Type: replace Abstract: Direct Alignment Algorithms (DAAs) simplify LLM alignment by directly optimizing policies, bypassing reward modeling and RL. While DAAs differ in th

To Redact, or not to Redact? A Local LLM Approach to Deliberative Process Privilege Classification

Model ReleasesDGX agent

arXiv:2605.10211v1 Announce Type: cross Abstract: Government transparency laws, like the Freedom of Information (FOIA) acts in the United States and United Kingdom, and the Woo (Open Government Act) i

Tokens-per-Parameter Coverage Is Critical for Robust LLM Scaling Law Extrapolation

Model ReleasesDGX agent

arXiv:2605.08541v1 Announce Type: new Abstract: Neural scaling laws approximate a language model's loss as a power-law function of parameter count N and token count D. Following Chinchilla-style compu

Towards Cross-lingual Values Judgment: A Consensus-Pluralism Perspective

Model ReleasesDGX agent

arXiv:2602.17283v2 Announce Type: replace-cross Abstract: As large language models (LLMs) are employed worldwide, existing evaluation paradigms for their multilingual capabilities primarily focus on f

Two Ways to De-Bias an LLM-as-a-Judge: A Continuous-Score Comparison of Hierarchical Bayesian Calibration and Neural-ODE Score Transport

Model ReleasesDGX agent

arXiv:2605.09227v1 Announce Type: new Abstract: [Abridged] Using a Large Language Model (LLM) as an automatic rater (LLM-as-a-judge) is cheap but potentially biased: some judges run lenient, others st

← Previous
1…355356357358359…1044
Next →