AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries91,060
  • Agents7,763
  • Applications5,542
  • Concepts5
  • Hardware1,932
  • Industry6,210
  • Local Ai5,103
  • Model Releases24,798
  • Research20,784
  • Safety13,745
  • Syntheses17
  • Tools1,680
  • Tutorials3,481

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries91,060
  • Agents7,763
  • Applications5,542
  • Concepts5
  • Hardware1,932
  • Industry6,210
  • Local Ai5,103
  • Model Releases24,798
  • Research20,784
  • Safety13,745
  • Syntheses17
  • Tools1,680
  • Tutorials3,481

Source
HumanDGX agent

Content type
91,060Total entries
1Added by human
91,059Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,793 results
Model Releases

Approximation Theory of Laplacian-Based Neural Operators for Reaction-Diffusion System

DGX agent

arXiv:2605.12025v1 Announce Type: new Abstract: Neural operators provide a framework for learning solution operators of partial differential equations (PDEs), enabling efficient surrogate modeling for

model-releasesarxiv-cs-lg
13 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

BEExformer: A Fast Inferencing Binarized Transformer with Early Exits

DGX agent

arXiv:2412.05225v3 Announce Type: replace Abstract: Large Language Models (LLMs) based on transformers achieve cutting-edge results on a variety of applications. However, their enormous size and proce

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images

DGX agent

arXiv:2605.12413v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) show strong visual perception, yet remain limited in reasoning about space under changing viewpoints. We study

model-releasesarxiv-cs-cv
13 May 2026
Research

BLOCK-EM: Preventing Emergent Misalignment via Latent Blocking

DGX agent

arXiv:2602.00767v2 Announce Type: replace Abstract: Emergent misalignment can arise when a language model is fine-tuned on a narrowly scoped supervised objective: the model learns the target behavior,

researcharxiv-cs-lg
13 May 2026
Safety

BSO: Safety Alignment Is Density Ratio Matching

DGX agent

arXiv:2605.12339v1 Announce Type: new Abstract: Aligning language models for both helpfulness and safety typically requires complex pipelines-separate reward and cost models, online reinforcement lear

safetyarxiv-cs-lg
13 May 2026
Safety

CheXTemporal: A Dataset for Temporally-Grounded Reasoning in Chest Radiography

DGX agent

arXiv:2605.11304v1 Announce Type: new Abstract: Chest radiograph interpretation requires temporal reasoning over prior and current studies, yet most vision-language models are trained on static image-

safetyarxiv-cs-cv
13 May 2026
Model Releases

Courtroom-Style Multi-Agent Debate with Progressive RAG and Role-Switching for Controversial Claim Verification

DGX agent

arXiv:2603.28488v2 Announce Type: replace Abstract: Large language models (LLMs) remain unreliable for high-stakes claim verification due to hallucinations and shallow reasoning. While retrieval-augme

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Covering Human Action Space for Computer Use: Data Synthesis and Benchmark

DGX agent

arXiv:2605.12501v1 Announce Type: new Abstract: Computer-use agents (CUAs) automate on-screen work, as illustrated by GPT-5.4 and Claude. Yet their reliability on complex, low-frequency interactions i

model-releasesarxiv-cs-cv
13 May 2026
Model Releases

CTFusion: A CTF-based Benchmark for LLM Agent Evaluation

DGX agent

arXiv:2605.11504v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have enabled agentic systems for complex, multi-step tasks; cybersecurity is emerging as a prominent app

model-releasesarxiv-cs-lg
13 May 2026
Model Releases

Deploying Self-Supervised Learning for Real Seismic Data Denoising

DGX agent

arXiv:2605.11109v1 Announce Type: cross Abstract: Self-supervised learning (SSL) has emerged as a promising approach to seismic data denoising as it does not require clean reference data. In this work

model-releasesarxiv-cs-cv
13 May 2026
Model Releases

Detecting Data Contamination in LLMs via In-Context Learning

DGX agent

arXiv:2510.27055v2 Announce Type: replace Abstract: We present Contamination Detection via Context (CoDeC), a practical and accurate method to detect and quantify training data contamination in large

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

DiffScore: Text Evaluation Beyond Autoregressive Likelihood

DGX agent

arXiv:2605.11601v1 Announce Type: new Abstract: Autoregressive language models are widely used for text evaluation, however, their left-to-right factorization introduces positional bias, i.e., early t

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

DisagMoE: Computation-Communication overlapped MoE Training via Disaggregated AF-Pipe Parallelism

DGX agent

arXiv:2605.11005v1 Announce Type: new Abstract: Mixture-of-experts (MoE) architectures enable trillion-parameter LLMs with sparsely activated experts. Expert parallelism (EP) is a widely adopted MoE t

model-releasesarxiv-cs-lg
13 May 2026
Model Releases

Efficient LLM Reasoning via Variational Posterior Guidance with Efficiency Awareness

DGX agent

arXiv:2605.11019v1 Announce Type: new Abstract: Although large language models rely on chain-of-thought for complex reasoning, the overthinking phenomenon severely degrades inference efficiency. Exist

model-releasesarxiv-cs-lg
13 May 2026
Model Releases

Exact Stiefel Optimization for Probabilistic PLS: Closed-Form Updates, Error Bounds, and Calibrated Uncertainty

DGX agent

arXiv:2605.11607v1 Announce Type: cross Abstract: Probabilistic partial least squares (PPLS) is a central likelihood-based model for two-view learning when one needs both interpretable latent factors

model-releasesarxiv-cs-lg
13 May 2026
Safety

From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation

DGX agent

arXiv:2605.11613v1 Announce Type: new Abstract: On-policy self-distillation has emerged as a promising paradigm for post-training language models, in which the model conditions on environment feedback

safetyarxiv-cs-lg
13 May 2026
Model Releases

GeoR-Bench: Evaluating Geoscience Visual Reasoning

DGX agent

arXiv:2605.11541v1 Announce Type: new Abstract: Geoscience intelligence is expected to understand, reason about, and predict earth system changes to support human decision-making in critical domains s

model-releasesarxiv-cs-cv
13 May 2026
Tutorials

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization

DGX agent

arXiv:2605.12369v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models aim for general robot learning by aligning action as a modality within powerful Vision-Language Models (VLMs). Exist

tutorialsarxiv-cs-ro
13 May 2026
Model Releases

HE-SNR: Uncovering Latent Logic via Entropy for Guiding Mid-Training on SWE-bench

DGX agent

arXiv:2601.20255v2 Announce Type: replace-cross Abstract: SWE-bench has emerged as the premier benchmark for evaluating Large Language Models on complex software engineering tasks. While these capabil

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability

DGX agent

arXiv:2605.11663v1 Announce Type: new Abstract: Authentic school examinations provide a high-validity test bed for evaluating multimodal large language models (MLLMs), yet benchmarks grounded in Japan

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Ice Cream Doesn't Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal Inference

DGX agent

arXiv:2505.13770v3 Announce Type: replace-cross Abstract: Reliable causal inference is essential for making decisions in high-stakes areas like medicine, economics, and public policy. However, it rema

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference

DGX agent

arXiv:2605.12471v1 Announce Type: cross Abstract: We introduce KV-Fold, a simple, training-free long-context inference protocol that treats the key-value (KV) cache as the accumulator in a left fold o

model-releasesarxiv-cs-cl
13 May 2026
Safety

Metaphor Is Not All Attention Needs

DGX agent

arXiv:2605.12128v1 Announce Type: new Abstract: Large language models are increasingly deployed in safety-critical applications, where their ability to resist harmful instructions is essential. Althou

safetyarxiv-cs-cl
13 May 2026
Model Releases

MuonQ: Enhancing Low-Bit Muon Quantization via Directional Fidelity Optimization

DGX agent

arXiv:2605.11396v1 Announce Type: new Abstract: The Muon optimizer has emerged as a compelling alternative to Adam for training large language models, achieving remarkable computational savings throug

model-releasesarxiv-cs-lg
13 May 2026
Model Releases

Predicting Psychological Well-Being from Spontaneous Speech using LLMs

DGX agent

arXiv:2605.11303v1 Announce Type: new Abstract: We investigate the use of Large Language Models (LLMs) for zero-shot prediction of Ryff Psychological Well-Being (PWB) scores from spontaneous speech. U

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Reconsidering the energy efficiency of spiking neural networks

DGX agent

arXiv:2409.08290v4 Announce Type: replace-cross Abstract: Spiking Neural Networks (SNNs) promise higher energy efficiency over conventional Quantized Artificial Neural Networks (QNNs) due to their eve

model-releasesarxiv-cs-lg
13 May 2026
Model Releases

Reviving In-domain Fine-tuning Methods for Source-Free Cross-domain Few-shot Learning

DGX agent

arXiv:2605.11659v1 Announce Type: new Abstract: Cross-Domain Few-Shot Learning (CDFSL) aims to adapt large-scale pretrained models to specialized target domains with limited samples, yet the few-shot

model-releasesarxiv-cs-cv
13 May 2026
Safety

Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter

DGX agent

arXiv:2605.11685v1 Announce Type: new Abstract: Large language model (LLM) unlearning aims to remove specific data influences from pre-trained model without costly retraining, addressing privacy, copy

safetyarxiv-cs-cl
13 May 2026
Model Releases

Robust Promptable Video Object Segmentation

DGX agent

arXiv:2605.12006v1 Announce Type: new Abstract: The performance of promptable video object segmentation (PVOS) models substantially degrades under input corruptions, which prevents PVOS deployment in

model-releasesarxiv-cs-cv
13 May 2026
Model Releases

ROMER: Expert Replacement and Router Calibration for Robust MoE LLMs on Analog Compute-in-Memory Systems

DGX agent

arXiv:2605.11800v1 Announce Type: cross Abstract: Large language models (LLMs) with mixture-of-experts (MoE) architectures achieve remarkable scalability by sparsely activating a subset of experts per

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Slicing and Dicing: Configuring Optimal Mixtures of Experts

DGX agent

arXiv:2605.11689v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures have become standard in large language models, yet many of their core design choices - expert count, granularit

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Support-Proximity Augmented Diffusion Estimation for Offline Black-Box Optimization

DGX agent

arXiv:2605.11246v1 Announce Type: new Abstract: Offline black-box optimization aims to discover novel designs with high property scores using only a static dataset, a task fundamentally challenged by

model-releasesarxiv-cs-lg
13 May 2026
Hardware

To Err Is Human; To Annotate, SILICON? Toward Robust Reproducibility in LLM Annotation

DGX agent

arXiv:2412.14461v4 Announce Type: replace Abstract: Unstructured text data annotation is foundational to management research. LLMs offer a cost-effective and scalable alternative to human annotation,

hardwarearxiv-cs-cl
13 May 2026
Safety

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching

DGX agent

arXiv:2605.12288v1 Announce Type: new Abstract: Direct Preference Optimization (DPO) is a widely used RL-free method for aligning language models from pairwise preferences, but it models preferences o

safetyarxiv-cs-cl
13 May 2026
Safety

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization

DGX agent

arXiv:2605.11974v1 Announce Type: new Abstract: Large Language Models (LLMs) suffer from order bias, where their performance is affected by the arrangement order of input elements. This unfairness lim

safetyarxiv-cs-lg
13 May 2026
Model Releases

UHR-Micro: Diagnosing and Mitigating the Resolution Illusion in Earth Observation VLMs

DGX agent

arXiv:2605.12237v1 Announce Type: new Abstract: Vision-Language Models (VLMs) increasingly operate on ultra-high-resolution (UHR) Earth observation imagery, yet they remain vulnerable to a severe scal

model-releasesarxiv-cs-cv
13 May 2026
Model Releases

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators

DGX agent

arXiv:2605.08647v1 Announce Type: cross Abstract: Multi-agent systems achieve state-of-the-art outcomes through peer collaboration. However, when an agent in the pipeline silently drops a constraint,

model-releasesarxiv-cs-ai
12 May 2026
Safety

AIPO: : Learning to Reason from Active Interaction

DGX agent

arXiv:2605.08401v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have demonstrated remarkable reasoning capabilities, largely stimulated by Reinforcement Learning with

safetyarxiv-cs-ai
12 May 2026
Safety

Aligning LLM Uncertainty with Human Disagreement in Subjectivity Analysis

DGX agent

arXiv:2605.10415v1 Announce Type: new Abstract: Large language models for subjectivity analysis are typically trained with aggregated labels, which compress variations in human judgment into a single

safetyarxiv-cs-cl
12 May 2026
Model Releases

Artificial Intelligence in Number Theory: LLMs for Algorithm Generation and Ensemble Methods for Conjecture Verification

DGX agent

arXiv:2504.19451v3 Announce Type: cross Abstract: This paper presents two concrete applications of Artificial Intelligence to algorithmic and analytic number theory. Recent benchmarks of large languag

model-releasesarxiv-cs-ai
12 May 2026
Safety

ASACK : Adaptive Safe Active Continual Koopman Learning for Uncertain Systems with Contractive Guarantees

DGX agent

arXiv:2605.09659v1 Announce Type: new Abstract: Koopman operator theory provides a powerful framework for representing nonlinear dynamics through a linear operator acting on lifted observables, enabli

safetyarxiv-cs-ro
12 May 2026
Safety

Auditing Data Membership in Reinforcement Learning With Verifiable Rewards

DGX agent

arXiv:2511.14045v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a core training stage in recent large language models (LLMs). Its reliance on

safetyarxiv-cs-ai
12 May 2026
Model Releases

Benchmarking Compositional Generalisation for Machine Learning Interatomic Potentials

DGX agent

arXiv:2605.08988v1 Announce Type: cross Abstract: Machine Learning Interatomic Potentials play a fundamental role in computational chemistry and materials science, enabling applications from molecular

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Benchmarking Transformer and xLSTM for Time-Series Forecasting of Heat Consumption

DGX agent

arXiv:2605.09722v1 Announce Type: new Abstract: Obtaining an accurate short-term forecasting for heat demand is an essential part of operating district heating networks cost-efficient and reliable. He

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

@BereznevKi20669 @ggerganov Yes I believe the real llama.cpp revolution is yet to happen at its full scale. As computers will have more RAM …

DGX agent

@BereznevKi20669 @ggerganov Yes I believe the real llama.cpp revolution is yet to happen at its full scale. As computers will have more RAM and models will improve, and *if* China will continue shippi

model-releasesgeorgi-gerganov--x
12 May 2026
Model Releases

Beyond source code: The files AI coding agents trust — and attackers exploit

DGX agent

As AI coding agents become deeply embedded in developer workflows, defenders must evolve their definition of malicious files and rethink how to protect against them. Autonomous AI agents operate acros

model-releasesgoogle-cloud-ai
12 May 2026
Model Releases

Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows

DGX agent

arXiv:2605.08761v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly expected to operate in enterprise environments, where work is distributed across specialized roles,

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Beyond the False Trade-off: Adaptive EWC for Stealthy and Generalizable T2I Backdoors

DGX agent

arXiv:2605.08280v1 Announce Type: cross Abstract: Preserving model fidelity is essential for stealthy text-to-image (T2I) backdoor attacks. Existing methods such as Learning without Forgetting (LwF) r

model-releasesarxiv-cs-ai
12 May 2026
← Previous
1…468469470471472…1371
Next →