AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,860 results
18 May 2026

Golden Layers and Where to Find Them: Improved Knowledge Editing for Large Language Models Via Layer Gradient Analysis

Model ReleasesDGX agent

arXiv:2602.20207v3 Announce Type: replace-cross Abstract: Knowledge editing in Large Language Models (LLMs) aims to update the model's prediction for a specific query to a desired target while preserv

Latent Video Prediction Learns Better World Models

ResearchDGX agent

arXiv:2605.15618v1 Announce Type: cross Abstract: Self-supervised video models are increasingly framed as world models, yet their evaluation remains largely confined to a single top-1 accuracy score o

VideoGameBench: Can Vision-Language Models complete popular video games?

Model ReleasesDGX agent

arXiv:2505.18134v3 Announce Type: replace Abstract: Vision-language models (VLMs) have achieved strong results on coding and math benchmarks that are challenging for humans, yet their ability to perfo

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
16 May 2026

Latest open artifacts (#21): Open model bonanza! Gemma 4, DeepSeek V4, Kimi K2.6, MiMo 2.5, GLM-5.1 & others. On CAISI's V4 assessment.

Model ReleasesDGX agent

This article covers recent releases of open-source AI models including Gemma 4, DeepSeek V4, Kimi K2.6, MiMo 2.5, and GLM-5.1, along with discussion of CAISI's V4 model assessment framework. The piece

15 May 2026

Mini-JEPA Foundation Model Fleet Enables Agentic Hydrologic Intelligence

Model ReleasesDGX agent

arXiv:2605.14120v1 Announce Type: cross Abstract: Geospatial foundation models compress multispectral observations into dense embeddings increasingly used in natural-language environmental reasoning s

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use

AgentsDGX agent

arXiv:2605.14038v1 Announce Type: new Abstract: Large language models (LLMs) increasingly act as autonomous agents that must decide when to answer directly vs. when to invoke external tools. Prior wor

NeuroAtlas: Benchmarking Foundation Models for Clinical EEG and Brain-Computer Interfaces

Model ReleasesDGX agent

arXiv:2605.14698v1 Announce Type: cross Abstract: Foundation models (FMs) promise to extract unified representations that generalize across downstream tasks. They have emerged across fields, including

OPT-Engine: Benchmarking the Limits of LLMs in Optimization Modeling via Complexity Scaling

Model ReleasesDGX agent

arXiv:2601.19924v2 Announce Type: replace-cross Abstract: We investigate the capabilities and scalability of Large Language Models (LLMs) in optimization modeling, a domain requiring structured reason

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026)

Model ReleasesDGX agent

arXiv:2501.05465v2 Announce Type: replace Abstract: As foundation AI models continue to increase in size, an important question arises - is massive scale the only path forward? This survey of about 16

Unsteady Metrics and Benchmarking Cultures of AI Model Builders

Model ReleasesDGX agent

arXiv:2605.14164v1 Announce Type: new Abstract: The primary way to establish and compare competencies in foundation and generative AI models has shifted from peer-reviewed literature to press releases

14 May 2026

Another banger of a model, free for Hermes agent users via Nous Portal!

Model ReleasesDGX agent

Nous Research announced the release of a new model available for free to Hermes agent users through the Nous Portal. The post suggests this is a significant model release from Nous Research, their org

CoT-Guard: Small Models for Strong Monitoring

Model ReleasesDGX agent

arXiv:2605.12746v1 Announce Type: cross Abstract: Monitoring the chain-of-thought (CoT) of reasoning models is a promising approach for detecting covert misbehavior (i.e., hidden objectives) in code g

D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2605.13276v1 Announce Type: new Abstract: The rapid evolution of Embodied AI has enabled Vision-Language-Action (VLA) models to excel in multimodal perception and task execution. However, applyi

Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue

Model ReleasesDGX agent

arXiv:2605.12920v1 Announce Type: cross Abstract: Effective collaboration between embodied agents requires more than acting in a shared environment; it demands communication grounded in each agent's e

PROMETHEUS: Automating Deep Causal Research Integrating Text, Data and Models

Local AiDGX agent

arXiv:2605.12835v1 Announce Type: new Abstract: Large language models can extract local causal claims from text, but those claims become more useful when organized as persistent, navigable world model

Query-Conditioned Test-Time Self-Training for Large Language Models

Model ReleasesDGX agent

arXiv:2605.13369v1 Announce Type: cross Abstract: Large language models (LLMs) are typically deployed with fixed parameters, and their performance is often improved by allocating more computation at i

Safe Bayesian Optimization for Uncertain Correlations Matrices in Linear Models of Co-Regionalization

Model ReleasesDGX agent

arXiv:2605.13302v1 Announce Type: new Abstract: This paper extends safety guarantees for multi-task Bayesian optimization with uncertain correlation matrices from intrinsic co-reginalization models to

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs

Model ReleasesDGX agent

arXiv:2510.18245v3 Announce Type: replace-cross Abstract: Scaling the number of parameters and the size of training data has proven to be an effective strategy for improving large language model (LLM)

Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?

Model ReleasesDGX agent

arXiv:2605.12684v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are now routinely deployed for visual understanding, generation, and curation. A substantial fraction of thes

13 May 2026

Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner

ApplicationsDGX agent

arXiv:2510.03206v2 Announce Type: replace-cross Abstract: Diffusion language models, especially masked discrete diffusion models, have achieved great success recently. While there are some theoretical

Enabling clinical use of foundation models for computational pathology

SafetyDGX agent

arXiv:2602.22347v2 Announce Type: replace Abstract: Foundation models for computational pathology are expected to facilitate the development of high-performing, generalisable deep learning systems. Ho

More Edits, More Stable: Understanding the Lifelong Normalization in Sequential Model Editing

Model ReleasesDGX agent

arXiv:2605.11836v1 Announce Type: cross Abstract: Lifelong Model Editing aims to continuously update evolving facts in Large Language Models while preserving unrelated knowledge and general capabiliti

12 May 2026

Can We Go Beyond Visual Features? Neural Tissue Relation Modeling for Relational Graph Analysis in Non-Melanoma Skin Histology

Model ReleasesDGX agent

arXiv:2512.06949v3 Announce Type: replace Abstract: Histopathology image segmentation is essential for delineating tissue structures in skin cancer diagnostics, but modeling spatial context and inter-

ColorConceptBench: A Benchmark for Probabilistic Color-Concept Understanding in Text-to-Image Models

Model ReleasesDGX agent

arXiv:2601.16836v3 Announce Type: replace-cross Abstract: Text-to-image (T2I) models have advanced considerably in generating high-quality images from textual descriptions. However, their ability to a

Decomposing and Steering Functional Metacognition in Large Language Models

Model ReleasesDGX agent

arXiv:2605.08942v1 Announce Type: new Abstract: Large language models (LLMs) increasingly exhibit behaviors suggesting awareness of their evaluation context, often adapting their reasoning strategies

DeformMaster: An Interactive Physics-Neural World Model for Deformable Objects from Videos

Model ReleasesDGX agent

arXiv:2605.09586v1 Announce Type: new Abstract: World models for deformable objects should recover not only geometry and appearance, but also underlying physical dynamics, interaction grounding, and m

Diffusion Models are Evolutionary Algorithms

Model ReleasesDGX agent

arXiv:2410.02543v3 Announce Type: replace-cross Abstract: In a convergence of machine learning and biology, we reveal that diffusion models are evolutionary algorithms. By considering evolution as a d

jNO: A JAX Library for Neural Operator and Foundation Model Training

Model ReleasesDGX agent

arXiv:2605.10159v1 Announce Type: new Abstract: jNO (jax Neural Operators) is a JAX-native library for neural operators and foundation models with unified support for both data-driven and physics-info

Large Language Models as Students Who Think Aloud: Overly Coherent, Verbose, and Confident

Model ReleasesDGX agent

arXiv:2602.01015v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly embedded in AI-based tutoring systems. Can they faithfully model novice reasoning and metacognitive ju

PARD-2: Target-Aligned Parallel Draft Model for Dual-Mode Speculative Decoding

ResearchDGX agent

arXiv:2605.08632v1 Announce Type: cross Abstract: Speculative decoding accelerates Large Language Models (LLMs) inference by using a lightweight draft model to propose candidate tokens that are verifi

Reinforcement Learning Measurement Model

Model ReleasesDGX agent

arXiv:2605.09305v1 Announce Type: cross Abstract: Interactive assessments generate sequential process data that are not well handled by conventional item response models. Existing MDP-based measuremen

Selective Deficits in LLM Mental Self-Modeling in a Behavior-Based Test of Theory of Mind

Model ReleasesDGX agent

arXiv:2603.26089v2 Announce Type: replace-cross Abstract: The ability to represent oneself and others as agents with knowledge, intentions, and belief states that guide their behavior - Theory of Mind

Sparse Layers are Critical to Scaling Looped Language Models

ResearchDGX agent

arXiv:2605.09165v1 Announce Type: cross Abstract: Looped language models repeat a set of transformer layers through depth, reducing memory costs and providing natural early-exit points at loop boundar

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data

Model ReleasesDGX agent

arXiv:2605.10129v1 Announce Type: new Abstract: Large language models (LLMs) rely on web-scale corpora for pre-training. The noise inherent in these datasets tends to obscure meaningful patterns and u

Visual-ERM: Reward Modeling for Visual Equivalence

Model ReleasesDGX agent

arXiv:2603.13224v2 Announce Type: replace-cross Abstract: Vision-to-code tasks require models to reconstruct structured visual inputs, such as charts, tables, and SVGs, into executable or structured r

Where Do Reasoning Models Refuse?

ResearchDGX agent

arXiv:2507.03167v3 Announce Type: replace-cross Abstract: Chat models without chain-of-thought (CoT) reasoning must decide whether to refuse a harmful request before generating their first response to

11 May 2026

CellScientist: Dual-Space Hierarchical Orchestration for Closed-Loop Refinement of Virtual Cell Models

ResearchDGX agent

arXiv:2605.07335v1 Announce Type: new Abstract: Virtual Cell Modeling (VCM) requires models that not only predict perturbation responses, but also support targeted revision when predictions fail. Curr

Clinically Aware Synthetic Image Generation for Concept Coverage in Chest X-ray Models

Model ReleasesDGX agent

arXiv:2603.15525v2 Announce Type: replace Abstract: Deep learning models for chest X-ray diagnosis are constrained by limited coverage of clinically meaningful concept combinations in publicly availab

Dino U-Net: Exploiting High-Fidelity Dense Features from Foundation Models for Medical Image Segmentation

Model ReleasesDGX agent

arXiv:2508.20909v2 Announce Type: replace Abstract: Foundation models pre-trained on large-scale natural image datasets offer a powerful paradigm for medical image segmentation. However, effectively t

Do Joint Audio-Video Generation Models Understand Physics?

Model ReleasesDGX agent

arXiv:2605.07061v1 Announce Type: cross Abstract: Joint audio-video generation models are rapidly approaching professional production quality, raising a central question: do they understand audio-visu

Radiologist-Guided Causal Concept Bottleneck Models for Chest X-Ray Interpretation

SafetyDGX agent

arXiv:2605.07785v1 Announce Type: new Abstract: Concept Bottleneck Models (CBMs) in medical imaging aim to improve model interpretability by predicting intermediate clinical concepts before final diag

SpikingBrain: Spiking Brain-inspired Large Models

HardwareDGX agent

arXiv:2509.05276v4 Announce Type: replace-cross Abstract: Mainstream Transformer-based large language models face major efficiency bottlenecks: training computation scales quadratically with sequence

Swap models & view their capabilities! Try out in Deep Agents CLI: https://docs.langchain.com/oss/python/deepagents/cli/

Model ReleasesDGX agent

Swap models & view their capabilities! Try out in Deep Agents CLI: https://docs.langchain.com/oss/python/deepagents/cli/ here's model profile details look like in practice, using @NVIDIAAIDev's Nemotr

Test-Time Compositional Generalization in Diffusion Models via Concept Discovery

Local AiDGX agent

arXiv:2605.07078v1 Announce Type: new Abstract: Compositional generalization requires models to produce novel configurations from familiar parts. In diffusion models, prior compositional generation me

Theoretical Limits of Language Model Alignment

SafetyDGX agent

arXiv:2605.07105v1 Announce Type: cross Abstract: Language model (LM) alignment improves model outputs to reflect human preferences while preserving the capabilities of the base model. The most common

Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model

Model ReleasesDGX agent

arXiv:2602.04774v2 Announce Type: replace-cross Abstract: Setting the learning rate (LR) for a deep learning model is a critical part of successful training. Choosing LRs is often done empirically wit

Towards Closing the Autoregressive Gap in Language Modeling via Entropy-Gated Continuous Bitstream Diffusion

Model ReleasesDGX agent

arXiv:2605.07013v1 Announce Type: new Abstract: Diffusion language models (DLMs) promise parallel, order-agnostic generation, but on standard benchmarks they have historically lagged behind autoregres

Understanding Robustness of Model Editing in Code LLMs

Model ReleasesDGX agent

arXiv:2511.03182v2 Announce Type: replace-cross Abstract: Large language models (LLMs) for code are increasingly used in software development, but they remain static after pretraining while APIs and s

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States

Model ReleasesDGX agent

arXiv:2605.07579v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) for Large Reasoning Models hinges on baseline estimation for variance reduction, but existing ap

10 May 2026

Local AI is having its moment! Below is the number of new GGUF models created each month over the past 8 months & insights from our HF inter…

Model ReleasesDGX agent

Local AI is having its moment! Below is the number of new GGUF models created each month over the past 8 months & insights from our HF internal agent (May is partial): - 176,000 total public GGUF mode

7 May 2026

Adapting Large Language Models to a Low-Resource Agglutinative Language: A Comparative Study of LoRA and QLoRA for Bashkir

Model ReleasesDGX agent

arXiv:2605.04948v1 Announce Type: new Abstract: This paper presents a comparative study of parameter-efficient fine-tuning (PEFT) methods, including LoRA and QLoRA, applied to the task of adapting lar

FideDiff: Efficient Diffusion Model for High-Fidelity Image Motion Deblurring

ApplicationsDGX agent

arXiv:2510.01641v3 Announce Type: replace Abstract: Recent advancements in image motion deblurring, driven by CNNs and transformers, have made significant progress. Large-scale pre-trained diffusion m

Leveraging Pretrained Language Models as Energy Functions for Glauber Dynamics Text Diffusion

ResearchDGX agent

arXiv:2605.04291v1 Announce Type: new Abstract: We present a discrete diffusion-based language model using Glauber dynamics from statistical physics. Our main insight is that instead of trying to trai

When LLMs get significantly worse: A statistical approach to detect model degradations

ApplicationsDGX agent

arXiv:2602.10144v2 Announce Type: replace-cross Abstract: Minimizing the inference cost and latency of foundation models has become a crucial area of research. Optimization approaches include theoreti

6 May 2026

Reasoning Models Can be Accurately Pruned Via Chain-of-Thought Reconstruction

Model ReleasesDGX agent

arXiv:2509.12464v2 Announce Type: replace Abstract: Reasoning language models such as DeepSeek-R1 produce long chain-of-thought traces during inference time which make them costly to deploy at scale.

Strategy-Aware Optimization Modeling with Reasoning LLMs

ApplicationsDGX agent

arXiv:2605.02545v1 Announce Type: new Abstract: Large language models (LLMs) can generate syntactically valid optimization programs, yet often struggle to reliably choose an effective modeling strateg

Strong Opinions, Loosely Held on Agent + Harness Engineering: 1. You can outperform any default harness+model (including codex & claude code…

Model ReleasesDGX agent

Strong Opinions, Loosely Held on Agent + Harness Engineering: 1. You can outperform any default harness+model (including codex & claude code) on pretty much any Task by engineering the harness around

Towards accurate extreme event likelihoods from diffusion model climate emulators

ResearchDGX agent

arXiv:2605.03802v1 Announce Type: cross Abstract: ML climate model emulators are useful for scenario planning and adaptation, allowing for cost-efficient experimentation. Recently, the diffusion model

What Makes VLMs Robust? Towards Reconciling Robustness and Accuracy in Vision-Language Models

Model ReleasesDGX agent

arXiv:2603.12799v2 Announce Type: replace Abstract: Achieving adversarial robustness in Vision-Language Models (VLMs) inevitably compromises accuracy on clean data, presenting a long-standing and chal

5 May 2026

Barriers to Counterfactual Credit Attribution for Autoregressive Models

ResearchDGX agent

arXiv:2605.01425v1 Announce Type: new Abstract: Generative AI disrupts the practice of giving credit to work that came before. Ideally, a generative model would give credit to any work on which its ou

← Previous
1…4041424344…998
Next →