AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,860 results
8 Jun 2026

When Large Language Models Fail in Healthcare: Evaluating Sensitivity to Prompt Variations

Model ReleasesDGX agent

arXiv:2606.07237v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in healthcare for tasks such as clinical question answering, diagnosis support, and report summariz

7 Jun 2026

There was an inflection point recently where the tide shifted to model pickers and OSS Mix of tokenmaxxing/cost fatigue, nemotron coalition,…

Model ReleasesDGX agent

There was an inflection point recently where the tide shifted to model pickers and OSS Mix of tokenmaxxing/cost fatigue, nemotron coalition, harness step functions, brains/claws, etc Long live those w

6 Jun 2026
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

No Need to Train Your RDB Foundation Model

Model ReleasesDGX agent

arXiv:2602.13697v2 Announce Type: replace Abstract: Relational databases (RDBs) contain vast amounts of heterogeneous tabular information that can be exploited for predictive modeling purposes. But si

5 Jun 2026

Humans' ALMANAC: A Human Collaboration Dataset of Action-Level Mental Model Annotations for Agent Collaboration

Model ReleasesDGX agent

arXiv:2606.06388v1 Announce Type: cross Abstract: Recent advances in LLM agents have enabled complex cognitive capabilities, such as multi-step reasoning, planning, and tool use, that increasingly pos

Knowledge Distillation for Visual Autoregressive Models

ResearchDGX agent

arXiv:2606.06078v1 Announce Type: new Abstract: Autoregressive (AR) image generation models are highly expressive but computationally intensive, motivating effective model compression. Knowledge disti

4 Jun 2026

ChessMimic: Per-Rating Transformer Models for Human Move, Clock, and Outcome Prediction in Online Blitz Chess

Model ReleasesDGX agent

arXiv:2606.04473v1 Announce Type: cross Abstract: We present ChessMimic, a system of three small encoder-only transformers - for move, thinking-time, and outcome prediction - conditioned on the positi

Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models

Model ReleasesDGX agent

arXiv:2510.20042v3 Announce Type: replace Abstract: Generative image models produce striking visuals yet often misrepresent culture. Prior work has examined cultural bias mainly in text-to-image (T2I)

Flow Matching Calibration for Simulation-Based Inference under Model Misspecification

Model ReleasesDGX agent

arXiv:2509.23385v5 Announce Type: replace-cross Abstract: Simulation-based inference (SBI) is transforming experimental sciences by enabling parameter estimation in complex non-linear models from simu

From Symbolic to Geometric: Enabling Spatial Reasoning in Large Language Models

Model ReleasesDGX agent

arXiv:2606.04381v1 Announce Type: cross Abstract: Recent large language models (LLMs) often appear to exhibit spatial reasoning ability; however, this capability is largely symbolic, arising from patt

New course on serving LLMs efficiently -- how do you serve models to many concurrent users at low latency and reasonable cost? This short co…

Model ReleasesDGX agent

New course on serving LLMs efficiently -- how do you serve models to many concurrent users at low latency and reasonable cost? This short course is built with @RedHat and taught by @cedricclyburn. Eff

3 Jun 2026

A 3D Isovist World Model -- Revealing a City's Unseen Geometry and Its Emergent Cross-City Signature

Model ReleasesDGX agent

arXiv:2606.03609v1 Announce Type: cross Abstract: Embodied agents that navigate cities rely on world models that predict how their surroundings will change as they move. But for navigation, what matte

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks

Model ReleasesDGX agent

arXiv:2606.03650v1 Announce Type: cross Abstract: Choosing or ranking language models for a specific application is hardest when no task-specific labeled data exists, and standard public benchmarks ca

Knowledge Editing in Masked Diffusion Language Models

Model ReleasesDGX agent

arXiv:2606.03924v1 Announce Type: new Abstract: Knowledge editing aims to update or correct factual knowledge in a language model. A widely used approach, locate-then-edit, does this in two steps: it

Staying Alive: Uncensored Survival Analysis with Tabular Foundation Models

Model ReleasesDGX agent

arXiv:2606.03689v1 Announce Type: cross Abstract: Survival Analysis (SA) is a statistical framework that models the time span until some event of interest occurs. Widely used in several domains, inclu

Today, we’re excited to introduce Miso One, the most emotive voice model in the world. Miso One is an 8-billion-parameter text-to-speech mod…

Model ReleasesDGX agent

Today, we’re excited to introduce Miso One, the most emotive voice model in the world. Miso One is an 8-billion-parameter text-to-speech model for highly expressive speech generation. It emotes like a

vLLM Semantic Router: Signal Driven Decision Routing for Mixture-of-Modality Models

Model ReleasesDGX agent

arXiv:2603.04444v3 Announce Type: replace-cross Abstract: As large language models (LLMs) diversify across modalities, capabilities, and cost profiles, the problem of intelligent request routing -- se

Wavelet Fourier Diffuser: Frequency-Aware Diffusion Model for Reinforcement Learning

Model ReleasesDGX agent

arXiv:2509.19305v2 Announce Type: replace-cross Abstract: Diffusion probability models have shown significant promise in offline reinforcement learning by directly modeling trajectory sequences. Howev

We partnered with @FireworksAI_HQ to train open-source models for legal. Here's what we found: 1) Hybrid legal agents can beat frontier mode…

Model ReleasesDGX agent

We partnered with @FireworksAI_HQ to train open-source models for legal. Here's what we found: 1) Hybrid legal agents can beat frontier models on quality and cost by routing selectively to a frontier

2 Jun 2026

A Practical Upper Bound on Selection Bias Effects in Medical Prediction Models

SafetyDGX agent

arXiv:2606.00563v1 Announce Type: cross Abstract: Selection bias is a common and often unavoidable aspect of real-world data that challenges the generalizability of machine learning models. When model

Auditing Asset-Specific Preferences in Financial Large Language Models: Evidence from Bitcoin Representations and Portfolio Allocation

Model ReleasesDGX agent

arXiv:2606.02528v1 Announce Type: cross Abstract: Large language models now power robo-advisors and trading agents, yet whether they carry built-in biases toward specific assets is largely untested. W

Benchmarks for Vision-Language Models in Urban Perception Should Be Reliability-Aware and Negotiated

Model ReleasesDGX agent

arXiv:2606.00871v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used to generate structured descriptions of street-level imagery for tasks such as streetscape auditing

Can Vision Language Models Learn Intuitive Physics from Interaction?

TutorialsDGX agent

arXiv:2602.06033v2 Announce Type: replace Abstract: Pre-trained vision language models do not have good intuitions about the physical world. Recent work has shown that supervised fine-tuning can impro

GottBERT: a pure German Language Model

ResearchDGX agent

arXiv:2012.02110v2 Announce Type: replace Abstract: Pre-trained language models have significantly advanced natural language processing (NLP), especially with the introduction of BERT and its optimize

OmniEEG-Bench: A Standardized Evaluation Benchmark for EEG Foundation Models

Model ReleasesDGX agent

arXiv:2606.00815v1 Announce Type: new Abstract: Electroencephalography (EEG) supports a variety of brain-computer interface (BCI) tasks ranging from brain-state monitoring to human-LLM interactions. E

Parameter-Efficient Fine-Tuning of Large Pretrained Models for Instance Segmentation Tasks

Model ReleasesDGX agent

arXiv:2606.01947v1 Announce Type: cross Abstract: Research and applications in artificial intelligence have recently shifted with the rise of large pretrained models, which deliver state-of-the-art re

Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models

Model ReleasesDGX agent

arXiv:2606.00150v1 Announce Type: cross Abstract: As Large Language Models evolve for user convenience, vulnerability to jailbreak attacks continues to be reported despite ongoing efforts in safety tr

PortBERT: Navigating the Depths of Portuguese Language Models

HardwareDGX agent

arXiv:2606.02100v1 Announce Type: new Abstract: Transformer models dominate modern NLP, but efficient, language-specific models remain scarce. In Portuguese, most focus on scale or accuracy, often neg

Prototype Transformer: Towards Language Model Architectures Interpretable by Design

Model ReleasesDGX agent

arXiv:2602.11852v2 Announce Type: replace Abstract: While state-of-the-art language models (LMs) surpass most humans in certain domains, their reasoning remains largely opaque, reducing trust and incr

Quantized Reasoning Models Think They Need to Think Longer, but They Do Not

ApplicationsDGX agent

arXiv:2606.00206v1 Announce Type: new Abstract: Post-training quantization (PTQ) is widely used to deploy large language models efficiently, but its effect on reasoning models is not well understood.

Query Circuits: Explaining How Language Models Answer User Prompts

Local AiDGX agent

arXiv:2509.24808v2 Announce Type: replace Abstract: Explaining why a language model produces a particular output requires local, input-level explanations. Existing methods uncover global capability ci

RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation

Model ReleasesDGX agent

arXiv:2606.01600v1 Announce Type: cross Abstract: Video world models are increasingly used in robotic manipulation, yet existing benchmarks mostly evaluate them under valid, feasible, and safe instruc

The Refusal--Compliance Tradeoff: A Large-Scale Safety Behavior Audit of Large Language Models

Model ReleasesDGX agent

arXiv:2605.05427v2 Announce Type: replace Abstract: Refusal rates are a poor proxy for LLM safety, i.e., a model may over-refuse benign prompts while still complying with harmful ones. We audit both f

WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching

Model ReleasesDGX agent

arXiv:2603.06331v2 Announce Type: replace Abstract: Diffusion-based world models have shown strong potential for unified world simulation, but the iterative denoising remains too costly for interactiv

WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World

Model ReleasesDGX agent

arXiv:2512.10958v2 Announce Type: replace Abstract: Generative world models are reshaping embodied AI, enabling agents to synthesize realistic 4D driving environments that look convincing but often fa

1 Jun 2026

Develop Physical AI Reasoning, World, and Action Models with NVIDIA Cosmos 3

HardwareDGX agent

NVIDIA Cosmos 3 is a frontier foundation model for physical AI that combines physical reasoning, world generation, and action generation within a single open model. The model uses a Mixture-of-Transfo

Esoteric Language Models: A Family of Any-Order Diffusion LLMs

TutorialsDGX agent

arXiv:2506.01928v4 Announce Type: replace Abstract: Diffusion-based language models offer a compelling alternative to autoregressive (AR) models by enabling parallel and controllable generation. Withi

Exploring Autonomous Agentic Data Engineering for Model Specialization

Model ReleasesDGX agent

arXiv:2605.30407v1 Announce Type: cross Abstract: Large Language Models (LLMs) have demonstrated strong performance on general tasks, while often struggling to adapt to specialized domains without hig

Human-Alignment and Calibration of Inference-Time Uncertainty in Large Language Models

SafetyDGX agent

arXiv:2508.08204v2 Announce Type: replace-cross Abstract: There has been much recent interest in evaluating large language models for uncertainty calibration to facilitate model control and modulate u

OpenAI models and Codex on Amazon Bedrock are now generally available

Model ReleasesDGX agent

OpenAI frontier models GPT-5.5 and GPT-5.4, and Codex, the OpenAI coding agent, are now generally available on Amazon Bedrock. AWS customers can access these latest OpenAI models through the same Amaz

Quantifying the Uncertainty of Foundation Models with Singular Value Ensembles

Model ReleasesDGX agent

arXiv:2601.22068v2 Announce Type: replace Abstract: Foundation models have become a dominant paradigm in machine learning, achieving remarkable performance across diverse tasks through large-scale pre

SALAAD: Sparse And Low-Rank Adaptation via ADMM for Large Language Model Inference

ResearchDGX agent

arXiv:2602.00942v3 Announce Type: replace Abstract: Modern large language models are increasingly deployed under compute and memory constraints, making flexible control of model capacity a central cha

VLM3: Vision Language Models Are Native 3D Learners

ResearchDGX agent

arXiv:2605.30561v1 Announce Type: cross Abstract: Vision Language Models (VLMs) enable a unified model to solve various vision tasks through prompting. They have shown promising performance in semanti

When LLMs Learn to Be Consistently Wrong: A Multi-Model Study of Linear Representations of Synthetic Deception

Model ReleasesDGX agent

arXiv:2605.30381v1 Announce Type: cross Abstract: Deceptive alignment, in which models maintain accurate internal representations while deliberately producing false outputs, remains a central challeng

29 May 2026

A Minimal Bifurcation Model of Load Imbalance in a Softmax Mixture-of-Experts Router

Model ReleasesDGX agent

arXiv:2605.29121v1 Announce Type: cross Abstract: We propose a minimal dynamical model of adaptive softmax routing for a two-expert Mixture-of-Experts (MoE) layer. The model is obtained as a mean-fiel

Are LLMs Socially Adaptive? Contrasting Belief Evolution in Large Language Models and Humans

Model ReleasesDGX agent

arXiv:2410.10398v3 Announce Type: replace-cross Abstract: As large language models (LLMs) increasingly engage in complex social interactions, ensuring that their behaviors align with human ethical pri

Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset

Model ReleasesDGX agent

arXiv:2605.29462v1 Announce Type: cross Abstract: The emergence of Large Vision-Language Models (LVLMs) has substantially expanded model capabilities beyond text-only understanding, enabling unified i

CrystalXRD-Bench: Benchmarking Vision-Language Models for XRD Peak Indexing Across Diverse Crystalline Materials

Model ReleasesDGX agent

arXiv:2605.29446v1 Announce Type: new Abstract: Miller-index identification from powder XRD patterns requires capabilities untested by existing multimodal benchmarks: the model must read a narrow peak

FarSkip-Collective: Unhobbling Blocking Communication in Mixture of Experts Models

Model ReleasesDGX agent

arXiv:2511.11505v3 Announce Type: replace Abstract: Blocking communication presents a major hurdle in running MoEs efficiently in distributed settings. To address this, we present FarSkip-Collective w

KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing

Model ReleasesDGX agent

arXiv:2605.29524v1 Announce Type: cross Abstract: Relay and reseller APIs increasingly intermediate access to large language models (LLMs), but users have no direct way to verify that a claimed endpoi

Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models

Model ReleasesDGX agent

arXiv:2605.29459v1 Announce Type: new Abstract: Large language models route every input through a learned embedding table of shape |V| x d_model, consuming hundreds of millions to billions of trainabl

Leveraging Routing Dynamics in Mixture-of-Experts Models for Efficient Language Adaptation

Model ReleasesDGX agent

arXiv:2605.29714v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models are widely used to scale language models, yet their expert routing behavior and adaptation in a multilingual setting rem

MemCollab: Cross-Model Memory Collaboration via Contrastive Trajectory Distillation

AgentsDGX agent

arXiv:2603.23234v2 Announce Type: replace Abstract: LLM agents increasingly rely on memory mechanisms to reuse knowledge from past problem-solving experiences. However, existing methods typically cons

minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models

ResearchDGX agent

arXiv:2605.30263v1 Announce Type: new Abstract: Recent video diffusion foundation models have achieved remarkable progress in high-quality video generation, yet turning them into real-time interactive

Rubric-Guided Process Reward for Stepwise Model Routing

SafetyDGX agent

arXiv:2605.29310v1 Announce Type: new Abstract: Stepwise model routing improves the efficiency of Large Reasoning Models (LRMs) by assigning each reasoning step to a suitable model. Recent methods for

The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More

Model ReleasesDGX agent

arXiv:2603.23971v2 Announce Type: replace-cross Abstract: Developers and consumers increasingly choose reasoning models (RMs) based on their listed API prices. However, how accurately do these prices

Unsupervised Semantic Segmentation Facilitates Model Understanding

Local AiDGX agent

arXiv:2605.29691v1 Announce Type: new Abstract: Self-supervised learning (SSL) has produced a diverse landscape of vision transformers (ViTs) whose pretrained representations support a wide range of d

28 May 2026

Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents

AgentsDGX agent

arXiv:2605.22166v2 Announce Type: replace Abstract: LLM agents are shaped not only by their language models, but also by the runtime harness that mediates observation, tool use, action execution, feed

Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information

SafetyDGX agent

arXiv:2605.28070v1 Announce Type: new Abstract: We highlight a failure mode of large reasoning models on questions with insufficient information: models may recognize that a problem is under-specified

Can Segmentation Models Understand the World? Towards Proactive Affordance Reasoning via Visual Chain-of-Thought

Model ReleasesDGX agent

arXiv:2605.27764v1 Announce Type: cross Abstract: Recent segmentation models couple large language models (LLMs) with mask decoders to ground complex language expressions into masks, yet their instruc

Claude’s new model is more ‘honest’ when it messes up

Model ReleasesDGX agent

Anthropic is releasing Claude Opus 4.8 on Thursday, and the company is touting the model's 'honesty.' According to Anthropic, it trains 'all [its] models to be honest - for instance, to avoid making c

← Previous
1…3839404142…998
Next →