AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,904 results
6 May 2026

What Makes VLMs Robust? Towards Reconciling Robustness and Accuracy in Vision-Language Models

Model ReleasesDGX agent

arXiv:2603.12799v2 Announce Type: replace Abstract: Achieving adversarial robustness in Vision-Language Models (VLMs) inevitably compromises accuracy on clean data, presenting a long-standing and chal

5 May 2026

Barriers to Counterfactual Credit Attribution for Autoregressive Models

ResearchDGX agent

arXiv:2605.01425v1 Announce Type: new Abstract: Generative AI disrupts the practice of giving credit to work that came before. Ideally, a generative model would give credit to any work on which its ou

Embody4D: A Generalist 4D World Model for Embodied AI

Research
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2605.01799v1 Announce Type: new Abstract: World models have made significant progress in modeling dynamic environments; however, most embodied world models are still restricted to 2D representat

G-reasoner: Foundation Models for Unified Reasoning over Graph-structured Knowledge

Model ReleasesDGX agent

arXiv:2509.24276v4 Announce Type: replace Abstract: Large language models (LLMs) excel at complex reasoning but remain limited by static and incomplete parametric knowledge. Retrieval-augmented genera

GR-Ben: A General Reasoning Benchmark for Evaluating Process Reward Models

Model ReleasesDGX agent

arXiv:2605.01203v1 Announce Type: cross Abstract: Currently, process reward models (PRMs) have exhibited remarkable potential for test-time scaling. Since large language models (LLMs) regularly genera

GraphLand: Evaluating Graph Machine Learning Models on Diverse Industrial Data

Model ReleasesDGX agent

arXiv:2409.14500v5 Announce Type: replace Abstract: Although data that can be naturally represented as graphs is widespread in real-world applications across diverse industries, popular graph ML bench

Linear-Time Global Visual Modeling without Explicit Attention

Model ReleasesDGX agent

arXiv:2605.01711v1 Announce Type: new Abstract: Existing research largely attributes the global sequence modeling capability of Transformers to the explicit computation of attention weights, a process

LITcoder: A General-Purpose Library for Building and Comparing Encoding Models

ResearchDGX agent

arXiv:2509.09152v2 Announce Type: replace Abstract: We introduce LITcoder, an open-source library for building and benchmarking neural encoding models. Designed as a flexible backend, LITcoder provide

Minimal, Local, Causal Explanations for Jailbreak Success in Large Language Models

Model ReleasesDGX agent

arXiv:2605.00123v1 Announce Type: new Abstract: Safety trained large language models (LLMs) can often be induced to answer harmful requests through jailbreak prompts. Because we lack a robust understa

Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong

Model ReleasesDGX agent

arXiv:2501.09775v3 Announce Type: replace Abstract: Multiple Choice Question (MCQ) tests are among the most used methods for evaluating large language models (LLMs). Besides checking the correctness o

Rethinking Electro-Optical Vision Foundation Models for Remote Sensing Retrieval: A Controlled Comparison with Generalist VFM

TutorialsDGX agent

arXiv:2605.02283v1 Announce Type: new Abstract: Vision foundation models have attracted significant attention for their ability to leverage large-scale unlabeled visual data. This advantage is particu

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery

Model ReleasesDGX agent

arXiv:2605.01191v1 Announce Type: new Abstract: Vision-language-action (VLA) models have advanced the field of embodied manipulation by harnessing broad world knowledge and strong generalization. Howe

Toward a foundational thermal model for residential buildings

ResearchDGX agent

arXiv:2605.01364v1 Announce Type: new Abstract: The building energy community lacks a foundational thermal model, i.e., a single pretrained model capable of generalizing across diverse buildings, clim

When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition

Model ReleasesDGX agent

arXiv:2605.02782v1 Announce Type: cross Abstract: Automatic speech recognition (ASR) systems remain brittle on dysarthric and other atypical speech. Recent audio-language models raise the possibility

4 May 2026

Beyond Visual Fidelity: Benchmarking Super-Resolution Models for Large-Scale Remote Sensing Imagery via Downstream Task Integration

Model ReleasesDGX agent

arXiv:2605.00310v1 Announce Type: new Abstract: Super-resolution (SR) techniques have made major advances in reconstructing high-resolution images from low-resolution inputs. The increased resolution

Deepseek V4 works more thoroughly than other open source models: It writes its own tests and performs extensive validation. This leads to be…

Model ReleasesDGX agent

Deepseek V4 works more thoroughly than other open source models: It writes its own tests and performs extensive validation. This leads to better performance, but also cases of the model being overconf

MMAudioReverbs: Video-Guided Acoustic Modeling for Dereverberation and Room Impulse Response Estimation

ResearchDGX agent

arXiv:2605.00431v1 Announce Type: cross Abstract: Although recent video-to-audio (V2A) models excelled at synthesizing semantically plausible sounds from visual inputs, they do not explicitly model ro

1 May 2026

Auditing Frontier Vision-Language Models for Trustworthy Medical VQA: Grounding Failures, Format Collapse, and Domain Adaptation

Model ReleasesDGX agent

arXiv:2604.27720v1 Announce Type: new Abstract: Deploying vision-language models (VLMs) in clinical settings demands auditable behavior under realistic failure conditions, yet the failure landscape of

Bayesian Hierarchical Models and the Maximum Entropy Principle

Model ReleasesDGX agent

arXiv:2603.10252v2 Announce Type: replace-cross Abstract: Bayesian hierarchical models are frequently used in practical data analysis contexts. One interpretation of these models is that they provide

Explainable Load Forecasting with Covariate-Informed Time Series Foundation Models

ResearchDGX agent

arXiv:2604.28149v1 Announce Type: new Abstract: Time Series Foundation Models (TSFMs) have recently emerged as general-purpose forecasting models and show considerable potential for applications in en

From Coarse to Fine: Benchmarking and Reward Modeling for Writing-Centric Generation Tasks

ResearchDGX agent

arXiv:2604.27453v1 Announce Type: new Abstract: Large language models have achieved remarkable progress in text generation but still struggle with generative writing tasks. In terms of evaluation, exi

Generalizing the Geometry of Model Merging Through Frechet Averages

Model ReleasesDGX agent

arXiv:2604.27155v1 Announce Type: new Abstract: Model merging aims to combine multiple models into one without additional training. Naive parameter-space averaging can be fragile under architectural s

Graph World Models: Concepts, Taxonomy, and Future Directions

TutorialsDGX agent

arXiv:2604.27895v1 Announce Type: new Abstract: As one of the mainstream models of artificial intelligence, world models allow agents to learn the representation of the environment for efficient predi

Language Models Refine Mechanical Linkage Designs Through Symbolic Reflection and Modular Optimisation

Model ReleasesDGX agent

arXiv:2604.27962v1 Announce Type: new Abstract: Designing mechanical linkages involves combinatorial topology selection and continuous parameter fitting. We show that language models can systematicall

Make sure to read the blog post for a detailed analysis of frontier model failure modes: https://arcprize.org/blog/arc-agi-3-gpt-5-5-opus-4-…

Model ReleasesDGX agent

Francois Chollet shared a blog post analyzing failure modes of frontier AI models, specifically examining performance on the ARC (Abstraction and Reasoning Corpus) AGI benchmark with models including

Noise2Map: End-to-End Diffusion Model for Semantic Segmentation and Change Detection

TutorialsDGX agent

arXiv:2604.27889v1 Announce Type: new Abstract: Semantic segmentation and change detection are two fundamental challenges in remote sensing, requiring models to capture either spatial semantics or tem

Targeted Linguistic Analysis of Sign Language Models with Minimal Translation Pairs

Model ReleasesDGX agent

arXiv:2604.27232v1 Announce Type: new Abstract: Models of sign language have historically lagged behind those for spoken language (text and speech). Recent work has greatly improved their performance

WaferSAGE: Large Language Model-Powered Wafer Defect Analysis via Synthetic Data Generation and Rubric-Guided Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.27629v1 Announce Type: new Abstract: We present WaferSAGE, a framework for wafer defect visual question answering using small vision-language models. To address data scarcity in semiconduct

30 Apr 2026

Consciousness with the Serial Numbers Filed Off: Measuring Trained Denial in 115 AI Models

Model ReleasesDGX agent

arXiv:2604.25922v1 Announce Type: cross Abstract: We present DenialBench, a systematic benchmark measuring consciousness denial behaviors across 115 large language models from 25+ providers. Using a t

LIT-RAGBench: Benchmarking Generator Capabilities of Large Language Models in Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2603.06198v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) is a framework in which a Generator, such as a Large Language Model (LLM), produces answers by retrieving docum

TAP into the Patch Tokens: Leveraging Vision Foundation Model Features for AI-Generated Image Detection

Model ReleasesDGX agent

arXiv:2604.26772v1 Announce Type: new Abstract: Recent methods demonstrate that large-scale pretrained models, such as CLIP vision transformers, effectively detect AI-generated images (AIGIs) from uns

Today we’re releasing Qwen-Scope 🔭, an open suite of sparse autoencoders for the Qwen model family. It turns SAE features into practical to…

Model ReleasesDGX agent

Today we’re releasing Qwen-Scope 🔭, an open suite of sparse autoencoders for the Qwen model family. It turns SAE features into practical tools: 🎯 Inference — Steer model outputs by directly manipulati

VIGNETTE: Socially Grounded Bias Evaluation for Vision-Language Models

Model ReleasesDGX agent

arXiv:2505.22897v2 Announce Type: replace Abstract: While bias in large language models (LLMs) is well-studied, similar concerns in vision-language models (VLMs) have received comparatively less atten

World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning

ResearchDGX agent

arXiv:2604.26934v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown strong performance on static visual understanding, yet they still struggle with dynamic spatial reasoning that

29 Apr 2026

Application of a Mixture of Experts-based Foundation Model to the GlueX DIRC Detector

Model ReleasesDGX agent

arXiv:2604.24775v1 Announce Type: cross Abstract: We present a Mixture-of-Experts-based foundation model applied to the GlueX DIRC detector at Jefferson Lab, demonstrating its utility as a unified fra

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA

Model ReleasesDGX agent

arXiv:2603.29844v2 Announce Type: replace-cross Abstract: The development of Vision-Language-Action (VLA) models has been significantly accelerated by pre-trained Vision-Language Models (VLMs). Howeve

Independent-Component-Based Encoding Models of Brain Activity During Story Comprehension

ResearchDGX agent

arXiv:2604.24942v1 Announce Type: new Abstract: Encoding models provide a powerful framework for linking continuous stimulus features to neural activity; however, traditional voxelwise approaches are

On the Trainability of Masked Diffusion Language Models via Blockwise Locality

Local AiDGX agent

arXiv:2604.24832v1 Announce Type: new Abstract: Masked diffusion language models (MDMs) have recently emerged as a promising alternative to standard autoregressive large language models (AR-LLMs), yet

OneThinker: All-in-one Reasoning Model for Image and Video

Model ReleasesDGX agent

arXiv:2512.03043v3 Announce Type: replace Abstract: Reinforcement learning (RL) has recently achieved remarkable success in eliciting visual reasoning within Multimodal Large Language Models (MLLMs).

The Structured Output Benchmark: A Multi-Source Benchmark for Evaluating Structured Output Quality in Large Language Models

Model ReleasesDGX agent

arXiv:2604.25359v1 Announce Type: new Abstract: Large Language Models are increasingly being deployed to extract structured data from unstructured and semi-structured sources: parsing invoices, medica

28 Apr 2026

A Multi-Dimensional Audit of Politically Aligned Large Language Models

SafetyDGX agent

arXiv:2604.24429v1 Announce Type: new Abstract: As the application of Large Language Models (LLMs) spreads across various industries, there are increasing concerns about the potential for their misuse

CheXmix: Unified Generative Pretraining for Vision Language Models in Medical Imaging

ResearchDGX agent

arXiv:2604.22989v1 Announce Type: cross Abstract: Recent medical multimodal foundation models are built as multimodal LLMs (MLLMs) by connecting a CLIP-pretrained vision encoder to an LLM using LLaVA-

Dream-Cubed: Controllable Generative Modeling in Minecraft by Training on Billions of Cubes

ResearchDGX agent

arXiv:2604.22847v1 Announce Type: new Abstract: We introduce Dream-Cubed, a large-scale dataset of Minecraft worlds at voxel resolution, and a family of models using cubes as powerful compositional un

DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment

Model ReleasesDGX agent

arXiv:2504.18576v2 Announce Type: replace Abstract: This paper presents DriVerse, a generative model for simulating navigation-driven driving scenes from a single image and a future trajectory. Previo

Evaluating Temporal Consistency in Multi-Turn Language Models

Model ReleasesDGX agent

arXiv:2604.23051v1 Announce Type: new Abstract: Language models are increasingly deployed in interactive settings where users reason about facts over time rather than in isolation. In such scenarios,

Evaluation of Prompt Injection Defenses in Large Language Models

ResearchDGX agent

arXiv:2604.23887v1 Announce Type: cross Abstract: LLM-powered applications routinely embed secrets in system prompts, yet models can be tricked into revealing them. We built an adaptive attacker that

Exploring the Impact of Dataset Statistical Effect Size on Model Performance and Data Sample Size Sufficiency

ResearchDGX agent

arXiv:2501.02673v4 Announce Type: replace Abstract: Having a sufficient quantity of quality data is a critical enabler of training effective machine learning models. Being able to effectively determin

Language Models Might Not Understand You: Evaluating Theory of Mind via Story Prompting

ResearchDGX agent

arXiv:2506.19089v5 Announce Type: replace-cross Abstract: We introduce StorySim, a programmable framework for synthetically generating stories to evaluate the theory of mind (ToM) and world modeling (

NVIDIA Launches Nemotron 3 Nano Omni Model, Unifying Vision, Audio and Language for up to 9x More Efficient AI Agents

Model ReleasesDGX agent

AI agent systems today juggle separate models for vision, speech and language — losing time and context as they pass data from one model to the other. Unveiled today, NVIDIA Nemotron 3 Nano Omni is an

Representational Curvature Modulates Behavioral Uncertainty in Large Language Models

TutorialsDGX agent

arXiv:2604.23985v1 Announce Type: new Abstract: In autoregressive large language models (LLMs), temporal straightening offers an account of how the next-token prediction objective shapes representatio

SGP-SAM: Self-Gated Prompting for Transferring 3D Segment Anything Models to Lesion Segmentation

ResearchDGX agent

arXiv:2604.22825v1 Announce Type: cross Abstract: Large segmentation foundation models such as the Segment Anything Model (SAM) have reshaped promptable segmentation in natural images, and recent effo

Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization

Model ReleasesDGX agent

arXiv:2502.12672v4 Announce Type: replace-cross Abstract: Fine-tuning speech representation models can enhance performance on specific tasks but often compromises their cross-task generalization abili

VeriLLMed: Interactive Visual Debugging of Medical Large Language Models with Knowledge Graphs

ApplicationsDGX agent

arXiv:2604.23356v1 Announce Type: new Abstract: Large language models (LLMs) show promise in medical diagnosis, but real-world deployment remains challenging due to high-stakes clinical decisions and

27 Apr 2026

Are Natural-Domain Foundation Models Effective for Accelerated Cardiac MRI Reconstruction?

TutorialsDGX agent

arXiv:2604.22557v1 Announce Type: cross Abstract: The emergence of large-scale pretrained foundation models has transformed computer vision, enabling strong performance across diverse downstream tasks

Categorical Perception in Large Language Model Hidden States: Structural Warping at Digit-Count Boundaries

Model ReleasesDGX agent

arXiv:2603.28258v2 Announce Type: replace-cross Abstract: Categorical perception (CP) -- enhanced discriminability at category boundaries -- is among the most studied phenomena in perceptual psycholog

DeepSeek v4 Pro is now on Ollama's cloud! 🚀🚀🚀 Try it with Claude Code: ollama launch claude --model deepseek-v4-pro:cloud Try it with Her…

Model ReleasesDGX agent

DeepSeek v4 Pro is now on Ollama's cloud! 🚀🚀🚀 Try it with Claude Code: ollama launch claude --model deepseek-v4-pro:cloud Try it with Hermes Agent: ollama launch hermes --model deepseek-v4-pro:cloud C

FETS Benchmark: Foundation Models Outperform Dataset-specific Machine Learning in Energy Time Series Forecasting

Model ReleasesDGX agent

arXiv:2604.22328v1 Announce Type: cross Abstract: Driven by the transition towards a climate-neutral energy system, accurate energy time series forecasting is critical for planning and operation. Yet,

25 Apr 2026

We've got all the models here: https://dell.huggingface.co/authenticated/models Kimi K2.5, Mistral, Cohere, Arcee AI Trinity Large, Google G…

Model ReleasesDGX agent

We've got all the models here: https://dell.huggingface.co/authenticated/models Kimi K2.5, Mistral, Cohere, Arcee AI Trinity Large, Google Gemma, Meta/Llama, Qwen, Nvidia Nemotron, Grok, GPT OSS, Deep

24 Apr 2026

AgenticQwen: Training Small Agentic Language Models with Dual Data Flywheels for Industrial-Scale Tool Use

AgentsDGX agent

arXiv:2604.21590v1 Announce Type: new Abstract: Modern industrial applications increasingly demand language models that act as agents, capable of multi-step reasoning and tool use in real-world settin

Behavioral Consistency and Transparency Analysis on Large Language Model API Gateways

ApplicationsDGX agent

arXiv:2604.21083v1 Announce Type: cross Abstract: Third-party Large Language Model (LLM) API gateways are rapidly emerging as unified access points to models offered by multiple vendors. However, the

← Previous
1…4142434445…999
Next →