AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,841 results
24 Jul 2026

Open models for the win!

SafetyDGX agent

Open models for the win! For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open mo

The Washington Post processed 1.79B input tokens per month through Together AI, running open models like Llama and Mistral in production wit…

Model ReleasesDGX agent

The Washington Post processed 1.79B input tokens per month through Together AI, running open models like Llama and Mistral in production with predictable costs and full control over the model stack. T

Training Large Language Models for Self-Explanation Faithfulness

Model ReleasesDGX agent

arXiv:2607.21090v1 Announce Type: cross Abstract: We propose a Reinforcement Learning (RL) method to directly optimize the faithfulness of self-explanations - the extent to which a model's generated r

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
23 Jul 2026

Cost per successful task: Benchmarking Kimi K3, GPT-5.5, and 8 more AI models

Model ReleasesDGX agent

Arize and Fireworks benchmarked 10 AI models across 2,400 agent runs. Learn why cost per successful task beats token price for model evaluation and routing. The post Cost per successful task: Benchmar

20 Jul 2026

Huge launch from @tryramp. Different steps in an agent workflow can use different models. This can help reduce costs significantly without s…

Model ReleasesDGX agent

Huge launch from @tryramp. Different steps in an agent workflow can use different models. This can help reduce costs significantly without sacrificing performance. Model routing will become a core par

Who’s Afraid of Chinese Models?

Model ReleasesDGX agent

Who’s Afraid of Chinese Models? Interesting proposal from Ben Thompson that both addresses the hypocrisy of labs outlawing distillation against their models despite training on unlicensed data, and co

15 Jul 2026

Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations

Model ReleasesDGX agent

arXiv:2607.12884v1 Announce Type: new Abstract: Patients seeking medical information often ask questions that embed incorrect assumptions or misconceptions. In such cases, safe medical communication r

FlowWAM: Optical Flow as a Unified Action Representation for World Action Models

SafetyDGX agent

arXiv:2607.13017v1 Announce Type: cross Abstract: World Action Models (WAMs) are able to leverage pretrained video generators for both world modeling and action prediction. However, directly leveragin

Scaling Point-in-Time Language Models

Model ReleasesDGX agent

arXiv:2607.11889v1 Announce Type: cross Abstract: Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias that compromis

14 Jul 2026

U.S. open-source models are quickly gaining ground. @Nvidia's newest Nemotron Ultra is fast growing on Ollama and unlocking complex, longer …

Model ReleasesDGX agent

U.S. open‑source AI models are rapidly gaining popularity, with NVIDIA’s newest model, **Nemotron Ultra**, becoming a prominent entry on the Ollama platform. On Ollama, Nemotron Ultra is quickly scali

12 Jul 2026

Another big reason to use combination of frontier models. Chain-of-thought monitoring is treated as a reliable safety layer for agents. This…

Model ReleasesDGX agent

Another big reason to use combination of frontier models. Chain-of-thought monitoring is treated as a reliable safety layer for agents. This DeepMind-affiliated study shows the layer can be argued out

8 Jul 2026

AtomBench: A Benchmarking Framework for Generative Crystal Reconstruction Models in Conventional Superconductors

Model ReleasesDGX agent

arXiv:2510.16165v2 Announce Type: replace Abstract: A key question in benchmarking generative crystal reconstruction models is how the amount and type of crystallographic information provided to a gen

MoWorld: A Flash World Model

AgentsDGX agent

arXiv:2607.06216v1 Announce Type: new Abstract: The future of World Models depends not only on scaling model capability, but also on scaling practicality and inference efficiency. High-frame-rate infe

7 Jul 2026

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

Model ReleasesDGX agent

arXiv:2506.07468v4 Announce Type: replace-cross Abstract: Conventional large language model (LLM) safety alignment relies on a reactive, disjoint loop: attackers exploit a static model, then defenders

Continual Model Merging with Test-Time Adaptation for Whole-Slide Image Analysis

Model ReleasesDGX agent

arXiv:2607.04755v1 Announce Type: new Abstract: Model merging offers a practical alternative to conventional continual learning by integrating independently fine-tuned models without retaining previou

Is Your Benchmark Still Useful? Dynamic Benchmarking for Code Language Models

Model ReleasesDGX agent

arXiv:2503.06643v2 Announce Type: replace-cross Abstract: In this paper, we tackle a critical challenge in model evaluation: how to keep code benchmarks useful when models might have already seen them

Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models

Model ReleasesDGX agent

arXiv:2607.04546v1 Announce Type: cross Abstract: Action-conditioned world models allow robots to predict the future consequences of candidate actions without additional physical interaction, supporti

Sakana AI (@SakanaAILabs) is now a model vendor on Merge Gateway, and Fugu Ultra is live through them. It's a multi-agent orchestration mode…

Model ReleasesDGX agent

Sakana AI (@SakanaAILabs) is now a model vendor on Merge Gateway, and Fugu Ultra is live through them. It's a multi-agent orchestration model that routes across frontier models behind one API. You get

The Remarkable Effectiveness of Providing AI Agents with Natural Language Tools: A Replication Study Validating NLT Performance Across 14 Models

Model ReleasesDGX agent

arXiv:2607.03953v1 Announce Type: cross Abstract: This study independently replicates and extends the Natural Language Tools (NLT) framework of Johnson et al.~(2025), which questions the use of struct

4 Jul 2026

introducing tinyrouter i reverse engineered the routing architecture behind Skana AI's Fugu and built replication for open frontier models. …

Model ReleasesDGX agent

introducing tinyrouter i reverse engineered the routing architecture behind Skana AI's Fugu and built replication for open frontier models. it's a tiny ~10K parameter LLM router that learns which mode

3 Jul 2026

Discrete Diffusion Language Models for Interactive Radiology Report Drafting

Model ReleasesDGX agent

arXiv:2607.01436v1 Announce Type: new Abstract: Diffusion language models, which generate text by denoising a token canvas bidirectionally instead of emitting tokens left to right, have become competi

Gravity-Awareness: Deep Learning Models and LLM Simulation of Human Awareness in Altered Gravity

Model ReleasesDGX agent

arXiv:2511.05536v2 Announce Type: replace-cross Abstract: Earth s gravity fundamentally shapes human behaviour. The brain encodes this force as an internal model of gravity, enabling the prediction an

Liquid Latent State Dynamics for Interpretable Turbofan Degradation Modeling

Model ReleasesDGX agent

arXiv:2607.01986v1 Announce Type: new Abstract: Multivariate time-series models for prognostics are often evaluated by point prediction accuracy, yet their internal states rarely expose a coherent deg

2 Jul 2026

Harnessing the Latent Space: From Steering Vectors to Model Calibrators for Control and Trust

TutorialsDGX agent

arXiv:2607.00083v1 Announce Type: cross Abstract: Language models have changed from unreliable text generators to highly-capable large models with trillions of parameters. Capability increases come ha

1 Jul 2026

How Post-Training Shapes Biological Reasoning Models

ResearchDGX agent

arXiv:2606.16517v2 Announce Type: replace Abstract: Scientific reasoning models for biology combine language models with foundation models trained on multimodal biological data, including DNA, RNA, an

30 Jun 2026

Concentration bounds on response-based vector embeddings of black-box generative models

ResearchDGX agent

arXiv:2511.08307v2 Announce Type: replace-cross Abstract: Generative models, such as large language models or text-to-image diffusion models, can generate relevant responses to user-given queries. Res

DNA Language Models: An Assessment of Pre-Training for Fine-Tuning Tasks

Model ReleasesDGX agent

arXiv:2606.30140v1 Announce Type: cross Abstract: Recent breakthroughs in foundation models and Large Language Models (LLMs) have introduced new opportunities for studying and decoding genomic sequenc

FlipGuard: Defending Large Language Models Against Quantization-Conditioned Backdoor Attacks

Model ReleasesDGX agent

arXiv:2606.28962v1 Announce Type: cross Abstract: Model quantization is essential for the efficient deployment of Large Language Models (LLMs), but introduces a critical vulnerability: Quantization-Co

Flow Matching in Feature Space for Stochastic World Modeling

Model ReleasesDGX agent

arXiv:2606.29059v1 Announce Type: cross Abstract: World modeling requires forecasting uncertain futures while preserving information useful for downstream perception. Existing visual world models ofte

Fuzzing Large Language Models to Elicit Hidden Behaviours

AgentsDGX agent

arXiv:2606.29646v1 Announce Type: cross Abstract: Sleeper agents are the canonical model organism of deception: models trained to behave normally but to emit an unsafe behaviour on a specific trigger.

Little Brains, Big Feats: Exploring Compact Language Models

Model ReleasesDGX agent

arXiv:2606.30062v1 Announce Type: cross Abstract: While large language models have been dominating the research landscape recently, small language models remain highly relevant across various domains;

SurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical Robotics

Model ReleasesDGX agent

arXiv:2606.29247v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models represent a promising direction for embodied intelligence in surgical robotics. Despite the prevalence of VLA benchm

29 Jun 2026

Revisiting Performance Claims for Chest X-Ray Models Using Clinical Context

Model ReleasesDGX agent

arXiv:2509.19671v3 Announce Type: replace Abstract: Public datasets of Chest X-Rays (CXRs) have long been a popular benchmark for developing machine learning (ML) computer vision models in healthcare.

27 Jun 2026

I put together a new article on setting up local coding agents with open-weight models. Everything runs 100% locally. I thought it might be …

Model ReleasesDGX agent

I put together a new article on setting up local coding agents with open-weight models. Everything runs 100% locally. I thought it might be useful putting this together because many people asked me ab

26 Jun 2026

EvoOptiGraph: Weakness-Driven Coevolution via Graph-Based Structural Generation for Optimization Modeling

AgentsDGX agent

arXiv:2606.26578v1 Announce Type: new Abstract: Automating optimization modeling from natural language with large language models (LLMs) faces two key challenges. First, training corpora lack structur

Not All Actions Are Equal: Rethinking Conditioning for Dexterous World Model

Local AiDGX agent

arXiv:2606.27325v1 Announce Type: new Abstract: Recent advances in action-conditioned world models show promising progress in modeling complex interactions and forecasting future states under diverse

25 Jun 2026

Did Models Learn Sufficiently? Attribution-Guided Training via Subset-Selected Counterfactual Augmentation

TutorialsDGX agent

arXiv:2511.12100v2 Announce Type: replace Abstract: In current visual model training, models often rely on only limited sufficient causes for their predictions, which makes them sensitive to distribut

24 Jun 2026

CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression

Model ReleasesDGX agent

arXiv:2606.24083v1 Announce Type: cross Abstract: 'Talk short. Drop grammar. Save token.' This caveman style is widely promoted as a way to cut inference cost, but whether it actually saves anything d

Trimming the Long-Tail of Visual World Modeling Evaluation

Model ReleasesDGX agent

arXiv:2606.24256v1 Announce Type: new Abstract: Physical interactions follow a long-tailed distribution: a set of common and regular interactions dominates human experience and visual data, while a br

23 Jun 2026

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model

Model ReleasesDGX agent

arXiv:2606.22317v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) is widely viewed as a promising path toward continuously improving large language models. Recent w

GraphPFN: A Prior-Data Fitted Graph Foundation Model

ApplicationsDGX agent

arXiv:2509.21489v3 Announce Type: replace Abstract: Graph foundation models face several fundamental challenges including transferability across diverse domains and data scarcity, which calls into que

How Well Can Your Video Model Remember? Measuring Memory-Budget Trade-offs in Long Video Understanding

ResearchDGX agent

arXiv:2606.20726v1 Announce Type: new Abstract: We introduce a compact empirical model that quantifies how answer accuracy degrades as a function of frame budget B and temporal distance D in long vide

PACT: Preserving Anchored Cores in Task-vectors for Model Merging

ResearchDGX agent

arXiv:2606.18627v2 Announce Type: replace Abstract: Model merging has emerged as a training-free alternative to multi-task learning, aiming to combine multiple task-specific fine-tuned models into a s

Sub-Billion, Super-Frontier: Small Language Models Rival Zero-Shot Frontier LLMs on General and Literary Relation Extraction

Model ReleasesDGX agent

arXiv:2606.22606v1 Announce Type: cross Abstract: Large language models (LLMs) achieve strong relation extraction (RE), but their computational demands and reliance on proprietary APIs limit deploymen

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Model ReleasesDGX agent

arXiv:2511.22699v4 Announce Type: replace Abstract: The landscape of high-performance image generation models is currently dominated by proprietary systems, such as Nano Banana Pro and Seedream 4.0. L

22 Jun 2026

Ai2 just released TMax 27B on Hugging Face A 27B terminal agent that hits 42.7% on Terminal Bench 2.0, rivaling models 40× its size.

Model ReleasesDGX agent

AI2 released TMax 27B, a 27 billion parameter terminal agent model available on Hugging Face that achieves 42.7% performance on Terminal Bench 2.0, matching the capabilities of much larger models desp

11 Jun 2026

Overcoming State Inertia in Full-Duplex Spoken Language Models via Activation Steering

Model ReleasesDGX agent

arXiv:2606.11386v1 Announce Type: cross Abstract: Full-duplex spoken language models (FD-SLMs) enable seamless speech interaction by allowing models to listen and speak simultaneously, yet the interna

Quantifying Subliminal Behavioral Transfer Ratios in Language Model Distillation

Model ReleasesDGX agent

arXiv:2606.11270v1 Announce Type: cross Abstract: Distillation of a language model intended to transfer benign behavior to a student model may also transfer undesirable characteristics, if they are pr

10 Jun 2026

Domain Adapted Large Language Models for Additive Manufacturing

Model ReleasesDGX agent

arXiv:2603.22017v2 Announce Type: replace Abstract: This work presents a collection of multi-modal domain adapted large language models built upon the instruction tuned variants of open weight models

One Lens, Many Worlds : A Capability-Typed Interface for World-Model Interpretability

ResearchDGX agent

arXiv:2606.09936v1 Announce Type: cross Abstract: World models are now built on substantially different computational substrates. Latent recurrent state-space models such as PlaNet and the Dreamer fam

PhantomBench: Benchmarking the Non-existential Threat of Language Models

Model ReleasesDGX agent

arXiv:2606.11105v1 Announce Type: cross Abstract: Hallucinations, where language models (LMs) generate factually ungrounded responses, pose serious risks, as users tend to blindly rely on them. This i

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff

Model ReleasesDGX agent

arXiv:2606.09932v1 Announce Type: cross Abstract: Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL) has become a standard pipeline for Large Language Model (LLM) post-training. SFT

9 Jun 2026

DriveReward: A Comprehensive Dataset and Generative Vision-Language Reward Model for Autonomous Driving

Model ReleasesDGX agent

arXiv:2606.08525v1 Announce Type: new Abstract: Reward models play a pivotal role in reinforcement learning (RL) and multi-modal trajectory selection for autonomous driving. However, acquiring such re

NutriMLLM: Multimodal Large Language Models for Dietary Micronutrient Analysis

Model ReleasesDGX agent

arXiv:2606.08948v1 Announce Type: cross Abstract: Comprehensive estimation of dietary micronutrients from food images could improve clinical nutrition care, but training such models requires large mul

8 Jun 2026

Design Once, Deploy at Scale: Template-Driven ML Development for Large Model Ecosystems

ApplicationsDGX agent

arXiv:2603.24963v3 Announce Type: replace Abstract: Modern computational advertising platforms typically rely on recommendation systems to predict user responses, such as click-through rates, conversi

Machine Learning for Electron-Scale Turbulence Modeling in W7-X

Model ReleasesDGX agent

arXiv:2511.04567v2 Announce Type: replace-cross Abstract: Constructing reduced models for turbulent transport is essential for accelerating profile predictions and enabling many-query tasks such as pa

Understanding Generative Recommendation with Semantic IDs from a Model-scaling View

ResearchDGX agent

arXiv:2509.25522v3 Announce Type: replace Abstract: Recent advancements in generative models have allowed the emergence of a promising paradigm for recommender systems (RS), known as Generative Recomm

6 Jun 2026

VLA-JEPA just dropped in LeRobot 🤖 What makes this model special is that it does not just learn what action to take from a given observatio…

Model ReleasesDGX agent

VLA-JEPA just dropped in LeRobot 🤖 What makes this model special is that it does not just learn what action to take from a given observation, it also leverages a JEPA world model to learn action-relev

4 Jun 2026

Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation?

Model ReleasesDGX agent

arXiv:2606.04811v1 Announce Type: new Abstract: Video generation models have made impressive strides in synthesizing visually compelling content, yet their outputs remain confined to the virtual domai

Generalization of World Models under Environmental Variability for Vision-based Quadrotor Navigation

SafetyDGX agent

arXiv:2606.05015v1 Announce Type: new Abstract: World models, learned generative models that predict how an environment evolves, have become a promising tool for sample-efficient robot learning. Yet h

← Previous
1…2122232425…998
Next →