AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
HumanDGX agent

87,171Total entries
1Added by human
87,170Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,612 results
14 May 2026

SpatialReward: Bridging the Perception Gap in Online RL for Image Editing via Explicit Spatial Reasoning

Model ReleasesDGX agent

arXiv:2602.07458v4 Announce Type: replace Abstract: Online Reinforcement Learning (RL) offers a promising avenue for complex image editing but is currently constrained by the scarcity of reliable and

State-Space NTK Collapse Near Bifurcations

Model ReleasesDGX agent

arXiv:2605.12763v1 Announce Type: new Abstract: Rich feature learning in tasks that unfold over time often requires the model to pass through bifurcations, constituting qualitative changes in the unde

Stress-Testing the Reasoning Competence of LLMs With Proofs Under Minimal Formalism

Model ReleasesDGX agent

arXiv:2605.12524v1 Announce Type: cross Abstract: We introduce ProofGrid, a benchmark suite for evaluating LLM reasoning through machine-checkable proofs rather than final answers alone. ProofGrid con

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SynCABEL: Synthetic Contextualized Augmentation for Biomedical Entity Linking

Model ReleasesDGX agent

arXiv:2601.19667v2 Announce Type: replace-cross Abstract: We present SynCABEL (Synthetic Contextualized Augmentation for Biomedical Entity Linking), a framework that addresses a central bottleneck in

The Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane Algorithm

Model ReleasesDGX agent

arXiv:2507.18553v4 Announce Type: replace Abstract: Quantizing the weights of large language models (LLMs) from 16-bit to lower bitwidth is the de facto approach to deploy massive transformers onto mo

The table in HTML format for easier (and non-truncated) viewing: https://sebastianraschka.com/llm-architecture-gallery/active-parameter-rati…

Model ReleasesDGX agent

Sebastian Raschka shared an HTML-formatted table comparing active parameter ratios across different large language model architectures for improved readability and to avoid text truncation. The resour

TokaMind for Power Grid: Cross-Domain Transfer from Fusion Plasma

Model ReleasesDGX agent

arXiv:2605.11033v1 Announce Type: cross Abstract: TokaMind is a multi-modal transformer (MMT) foundation model pre-trained on tokamak plasma diagnostics data from MAST, where it was shown to outperfor

TRIAGE: Evaluating Prospective Metacognitive Control in LLMs under Resource Constraints

AgentsDGX agent

arXiv:2605.13414v1 Announce Type: new Abstract: Deploying language models as autonomous agents requires more than per-task accuracy: when an agent faces a queue of problems under a finite token budget

Unlocking Patch-Level Features for CLIP-Based Class-Incremental Learning

Model ReleasesDGX agent

arXiv:2605.13835v1 Announce Type: new Abstract: Class-Incremental Learning (CIL) enables models to continuously integrate new knowledge while mitigating catastrophic forgetting. Driven by the remarkab

WARDEN: Endangered Indigenous Language Transcription and Translation with 6 Hours of Training Data

ResearchDGX agent

arXiv:2605.13846v1 Announce Type: cross Abstract: This paper introduces WARDEN, an early language model system capable of transcribing and translating Wardaman, an endangered Australian indigenous lan

Weakly Supervised Segmentation as Semantic-Based Regularization

ResearchDGX agent

arXiv:2605.13674v1 Announce Type: cross Abstract: Weakly supervised semantic segmentation (WSSS) trains dense pixel-level segmentation models from partial or coarse annotations such as bounding boxes,

13 May 2026

Ada-MK: Adaptive MegaKernel Optimization via Automated DAG-based Search for LLM Inference

Model ReleasesDGX agent

arXiv:2605.11581v1 Announce Type: new Abstract: When large language models (LLMs) serve real-time inference in commercial online advertising systems, end-to-end latency must be strictly bounded to the

Approximation Theory of Laplacian-Based Neural Operators for Reaction-Diffusion System

Model ReleasesDGX agent

arXiv:2605.12025v1 Announce Type: new Abstract: Neural operators provide a framework for learning solution operators of partial differential equations (PDEs), enabling efficient surrogate modeling for

BEExformer: A Fast Inferencing Binarized Transformer with Early Exits

Model ReleasesDGX agent

arXiv:2412.05225v3 Announce Type: replace Abstract: Large Language Models (LLMs) based on transformers achieve cutting-edge results on a variety of applications. However, their enormous size and proce

Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images

Model ReleasesDGX agent

arXiv:2605.12413v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) show strong visual perception, yet remain limited in reasoning about space under changing viewpoints. We study

BLOCK-EM: Preventing Emergent Misalignment via Latent Blocking

ResearchDGX agent

arXiv:2602.00767v2 Announce Type: replace Abstract: Emergent misalignment can arise when a language model is fine-tuned on a narrowly scoped supervised objective: the model learns the target behavior,

BSO: Safety Alignment Is Density Ratio Matching

SafetyDGX agent

arXiv:2605.12339v1 Announce Type: new Abstract: Aligning language models for both helpfulness and safety typically requires complex pipelines-separate reward and cost models, online reinforcement lear

CheXTemporal: A Dataset for Temporally-Grounded Reasoning in Chest Radiography

SafetyDGX agent

arXiv:2605.11304v1 Announce Type: new Abstract: Chest radiograph interpretation requires temporal reasoning over prior and current studies, yet most vision-language models are trained on static image-

Courtroom-Style Multi-Agent Debate with Progressive RAG and Role-Switching for Controversial Claim Verification

Model ReleasesDGX agent

arXiv:2603.28488v2 Announce Type: replace Abstract: Large language models (LLMs) remain unreliable for high-stakes claim verification due to hallucinations and shallow reasoning. While retrieval-augme

Covering Human Action Space for Computer Use: Data Synthesis and Benchmark

Model ReleasesDGX agent

arXiv:2605.12501v1 Announce Type: new Abstract: Computer-use agents (CUAs) automate on-screen work, as illustrated by GPT-5.4 and Claude. Yet their reliability on complex, low-frequency interactions i

CTFusion: A CTF-based Benchmark for LLM Agent Evaluation

Model ReleasesDGX agent

arXiv:2605.11504v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have enabled agentic systems for complex, multi-step tasks; cybersecurity is emerging as a prominent app

Deploying Self-Supervised Learning for Real Seismic Data Denoising

Model ReleasesDGX agent

arXiv:2605.11109v1 Announce Type: cross Abstract: Self-supervised learning (SSL) has emerged as a promising approach to seismic data denoising as it does not require clean reference data. In this work

Detecting Data Contamination in LLMs via In-Context Learning

Model ReleasesDGX agent

arXiv:2510.27055v2 Announce Type: replace Abstract: We present Contamination Detection via Context (CoDeC), a practical and accurate method to detect and quantify training data contamination in large

DiffScore: Text Evaluation Beyond Autoregressive Likelihood

Model ReleasesDGX agent

arXiv:2605.11601v1 Announce Type: new Abstract: Autoregressive language models are widely used for text evaluation, however, their left-to-right factorization introduces positional bias, i.e., early t

DisagMoE: Computation-Communication overlapped MoE Training via Disaggregated AF-Pipe Parallelism

Model ReleasesDGX agent

arXiv:2605.11005v1 Announce Type: new Abstract: Mixture-of-experts (MoE) architectures enable trillion-parameter LLMs with sparsely activated experts. Expert parallelism (EP) is a widely adopted MoE t

Efficient LLM Reasoning via Variational Posterior Guidance with Efficiency Awareness

Model ReleasesDGX agent

arXiv:2605.11019v1 Announce Type: new Abstract: Although large language models rely on chain-of-thought for complex reasoning, the overthinking phenomenon severely degrades inference efficiency. Exist

Exact Stiefel Optimization for Probabilistic PLS: Closed-Form Updates, Error Bounds, and Calibrated Uncertainty

Model ReleasesDGX agent

arXiv:2605.11607v1 Announce Type: cross Abstract: Probabilistic partial least squares (PPLS) is a central likelihood-based model for two-view learning when one needs both interpretable latent factors

From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation

SafetyDGX agent

arXiv:2605.11613v1 Announce Type: new Abstract: On-policy self-distillation has emerged as a promising paradigm for post-training language models, in which the model conditions on environment feedback

GeoR-Bench: Evaluating Geoscience Visual Reasoning

Model ReleasesDGX agent

arXiv:2605.11541v1 Announce Type: new Abstract: Geoscience intelligence is expected to understand, reason about, and predict earth system changes to support human decision-making in critical domains s

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization

TutorialsDGX agent

arXiv:2605.12369v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models aim for general robot learning by aligning action as a modality within powerful Vision-Language Models (VLMs). Exist

HE-SNR: Uncovering Latent Logic via Entropy for Guiding Mid-Training on SWE-bench

Model ReleasesDGX agent

arXiv:2601.20255v2 Announce Type: replace-cross Abstract: SWE-bench has emerged as the premier benchmark for evaluating Large Language Models on complex software engineering tasks. While these capabil

Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability

Model ReleasesDGX agent

arXiv:2605.11663v1 Announce Type: new Abstract: Authentic school examinations provide a high-validity test bed for evaluating multimodal large language models (MLLMs), yet benchmarks grounded in Japan

Ice Cream Doesn't Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal Inference

Model ReleasesDGX agent

arXiv:2505.13770v3 Announce Type: replace-cross Abstract: Reliable causal inference is essential for making decisions in high-stakes areas like medicine, economics, and public policy. However, it rema

KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference

Model ReleasesDGX agent

arXiv:2605.12471v1 Announce Type: cross Abstract: We introduce KV-Fold, a simple, training-free long-context inference protocol that treats the key-value (KV) cache as the accumulator in a left fold o

Metaphor Is Not All Attention Needs

SafetyDGX agent

arXiv:2605.12128v1 Announce Type: new Abstract: Large language models are increasingly deployed in safety-critical applications, where their ability to resist harmful instructions is essential. Althou

MuonQ: Enhancing Low-Bit Muon Quantization via Directional Fidelity Optimization

Model ReleasesDGX agent

arXiv:2605.11396v1 Announce Type: new Abstract: The Muon optimizer has emerged as a compelling alternative to Adam for training large language models, achieving remarkable computational savings throug

Predicting Psychological Well-Being from Spontaneous Speech using LLMs

Model ReleasesDGX agent

arXiv:2605.11303v1 Announce Type: new Abstract: We investigate the use of Large Language Models (LLMs) for zero-shot prediction of Ryff Psychological Well-Being (PWB) scores from spontaneous speech. U

Reconsidering the energy efficiency of spiking neural networks

Model ReleasesDGX agent

arXiv:2409.08290v4 Announce Type: replace-cross Abstract: Spiking Neural Networks (SNNs) promise higher energy efficiency over conventional Quantized Artificial Neural Networks (QNNs) due to their eve

Reviving In-domain Fine-tuning Methods for Source-Free Cross-domain Few-shot Learning

Model ReleasesDGX agent

arXiv:2605.11659v1 Announce Type: new Abstract: Cross-Domain Few-Shot Learning (CDFSL) aims to adapt large-scale pretrained models to specialized target domains with limited samples, yet the few-shot

Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter

SafetyDGX agent

arXiv:2605.11685v1 Announce Type: new Abstract: Large language model (LLM) unlearning aims to remove specific data influences from pre-trained model without costly retraining, addressing privacy, copy

Robust Promptable Video Object Segmentation

Model ReleasesDGX agent

arXiv:2605.12006v1 Announce Type: new Abstract: The performance of promptable video object segmentation (PVOS) models substantially degrades under input corruptions, which prevents PVOS deployment in

ROMER: Expert Replacement and Router Calibration for Robust MoE LLMs on Analog Compute-in-Memory Systems

Model ReleasesDGX agent

arXiv:2605.11800v1 Announce Type: cross Abstract: Large language models (LLMs) with mixture-of-experts (MoE) architectures achieve remarkable scalability by sparsely activating a subset of experts per

Slicing and Dicing: Configuring Optimal Mixtures of Experts

Model ReleasesDGX agent

arXiv:2605.11689v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures have become standard in large language models, yet many of their core design choices - expert count, granularit

Support-Proximity Augmented Diffusion Estimation for Offline Black-Box Optimization

Model ReleasesDGX agent

arXiv:2605.11246v1 Announce Type: new Abstract: Offline black-box optimization aims to discover novel designs with high property scores using only a static dataset, a task fundamentally challenged by

To Err Is Human; To Annotate, SILICON? Toward Robust Reproducibility in LLM Annotation

HardwareDGX agent

arXiv:2412.14461v4 Announce Type: replace Abstract: Unstructured text data annotation is foundational to management research. LLMs offer a cost-effective and scalable alternative to human annotation,

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching

SafetyDGX agent

arXiv:2605.12288v1 Announce Type: new Abstract: Direct Preference Optimization (DPO) is a widely used RL-free method for aligning language models from pairwise preferences, but it models preferences o

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization

SafetyDGX agent

arXiv:2605.11974v1 Announce Type: new Abstract: Large Language Models (LLMs) suffer from order bias, where their performance is affected by the arrangement order of input elements. This unfairness lim

UHR-Micro: Diagnosing and Mitigating the Resolution Illusion in Earth Observation VLMs

Model ReleasesDGX agent

arXiv:2605.12237v1 Announce Type: new Abstract: Vision-Language Models (VLMs) increasingly operate on ultra-high-resolution (UHR) Earth observation imagery, yet they remain vulnerable to a severe scal

12 May 2026

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators

Model ReleasesDGX agent

arXiv:2605.08647v1 Announce Type: cross Abstract: Multi-agent systems achieve state-of-the-art outcomes through peer collaboration. However, when an agent in the pipeline silently drops a constraint,

AIPO: : Learning to Reason from Active Interaction

SafetyDGX agent

arXiv:2605.08401v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have demonstrated remarkable reasoning capabilities, largely stimulated by Reinforcement Learning with

Aligning LLM Uncertainty with Human Disagreement in Subjectivity Analysis

SafetyDGX agent

arXiv:2605.10415v1 Announce Type: new Abstract: Large language models for subjectivity analysis are typically trained with aggregated labels, which compress variations in human judgment into a single

Artificial Intelligence in Number Theory: LLMs for Algorithm Generation and Ensemble Methods for Conjecture Verification

Model ReleasesDGX agent

arXiv:2504.19451v3 Announce Type: cross Abstract: This paper presents two concrete applications of Artificial Intelligence to algorithmic and analytic number theory. Recent benchmarks of large languag

ASACK : Adaptive Safe Active Continual Koopman Learning for Uncertain Systems with Contractive Guarantees

SafetyDGX agent

arXiv:2605.09659v1 Announce Type: new Abstract: Koopman operator theory provides a powerful framework for representing nonlinear dynamics through a linear operator acting on lifted observables, enabli

Auditing Data Membership in Reinforcement Learning With Verifiable Rewards

SafetyDGX agent

arXiv:2511.14045v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a core training stage in recent large language models (LLMs). Its reliance on

Benchmarking Compositional Generalisation for Machine Learning Interatomic Potentials

Model ReleasesDGX agent

arXiv:2605.08988v1 Announce Type: cross Abstract: Machine Learning Interatomic Potentials play a fundamental role in computational chemistry and materials science, enabling applications from molecular

Benchmarking Transformer and xLSTM for Time-Series Forecasting of Heat Consumption

Model ReleasesDGX agent

arXiv:2605.09722v1 Announce Type: new Abstract: Obtaining an accurate short-term forecasting for heat demand is an essential part of operating district heating networks cost-efficient and reliable. He

@BereznevKi20669 @ggerganov Yes I believe the real llama.cpp revolution is yet to happen at its full scale. As computers will have more RAM …

Model ReleasesDGX agent

@BereznevKi20669 @ggerganov Yes I believe the real llama.cpp revolution is yet to happen at its full scale. As computers will have more RAM and models will improve, and *if* China will continue shippi

Beyond source code: The files AI coding agents trust — and attackers exploit

Model ReleasesDGX agent

As AI coding agents become deeply embedded in developer workflows, defenders must evolve their definition of malicious files and rethink how to protect against them. Autonomous AI agents operate acros

Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows

Model ReleasesDGX agent

arXiv:2605.08761v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly expected to operate in enterprise environments, where work is distributed across specialized roles,

Beyond the False Trade-off: Adaptive EWC for Stealthy and Generalizable T2I Backdoors

Model ReleasesDGX agent

arXiv:2605.08280v1 Announce Type: cross Abstract: Preserving model fidelity is essential for stealthy text-to-image (T2I) backdoor attacks. Existing methods such as Learning without Forgetting (LwF) r

← Previous
1…354355356357358…1044
Next →