AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,399 results
22 Apr 2026

3D Foundation Model for Generalizable Disease Detection in Head Computed Tomography

Model ReleasesDGX agent

arXiv:2502.02779v3 Announce Type: replace-cross Abstract: Head computed tomography (CT) imaging is a widely-used imaging modality with multitudes of medical indications, particularly in assessing path

An Empirical Study of Multi-Generation Sampling for Jailbreak Detection in Large Language Models

Model ReleasesDGX agent

arXiv:2604.18775v1 Announce Type: new Abstract: Detecting jailbreak behaviour in large language models remains challenging, particularly when strongly aligned models produce harmful outputs only rarel

Can Continual Pre-training Bridge the Performance Gap between General-purpose and Specialized Language Models in the Medical Domain?

Model Releases
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2604.19394v1 Announce Type: new Abstract: This paper narrows the performance gap between small, specialized models and significantly larger general-purpose models through domain adaptation via c

Get more from speculative decoding in MoE models https://cohere.link/Et2rbsB

Model ReleasesDGX agent

Speculative decoding is a technique that accelerates language model inference by using a smaller draft model to generate candidate tokens, which are then verified by a larger model, reducing latency w

How does the optimizer implicitly bias the model merging loss landscape?

SafetyDGX agent

arXiv:2510.04686v2 Announce Type: replace-cross Abstract: Model merging combines independent solutions with different capabilities into a single one while maintaining the same inference cost. Two popu

Micro Language Models Enable Instant Responses

Model ReleasesDGX agent

arXiv:2604.19642v1 Announce Type: new Abstract: Edge devices such as smartwatches and smart glasses cannot continuously run even the smallest 100M-1B parameter language models due to power and compute

OmniMouse: Scaling properties of multi-modal, multi-task Brain Models on 150B Neural Tokens

Model ReleasesDGX agent

arXiv:2604.18827v1 Announce Type: cross Abstract: Scaling data and artificial neural networks has transformed AI, driving breakthroughs in language and vision. Whether similar principles apply to mode

21 Apr 2026

An Exploration of Mamba for Speech Self-Supervised Models

ResearchDGX agent

arXiv:2506.12606v2 Announce Type: replace Abstract: While Mamba has demonstrated strong performance in language modeling, its potential as a speech self-supervised learning (SSL) model remains underex

Anyhow, very impressive achievement, and a very solid model. Based on vibes, the gap seems around the same as it has always been between clo…

ApplicationsDGX agent

Ethan Mollick comments on an impressive AI model achievement, noting that the performance gap between models appears consistent with historical trends based on his assessment.

Joint Distillation for Fast Likelihood Evaluation and Sampling in Flow-based Models

ResearchDGX agent

arXiv:2512.02636v3 Announce Type: replace-cross Abstract: Log-likelihood evaluation enables important capabilities in generative models, including model comparison, certain fine-tuning objectives, and

MMErroR: A Benchmark for Erroneous Reasoning in Vision-Language Models

Model ReleasesDGX agent

arXiv:2601.03331v2 Announce Type: replace Abstract: Recent advances in Vision-Language Models (VLMs) have improved performance in multi-modal learning, raising the question of whether these models tru

MoCo: A One-Stop Shop for Model Collaboration Research

SafetyDGX agent

arXiv:2601.21257v2 Announce Type: replace Abstract: Advancing beyond single monolithic language models (LMs), recent research increasingly recognizes the importance of model collaboration, where multi

PARM: Pipeline-Adapted Reward Model

ApplicationsDGX agent

arXiv:2604.18327v1 Announce Type: cross Abstract: Reward models (RMs) are central to aligning large language models (LLMs) with human preferences, powering RLHF and advanced decoding strategies. While

20 Apr 2026

Claude Token Counter, now with model comparisons

Model ReleasesDGX agent

Claude Token Counter, now with model comparisons I upgraded my Claude Token Counter tool to add the ability to run the same count against different models in order to compare them. As far as I can tel

dear god lol - new kimi model is a fucking beast. GPT 5.4 level coding, 76% cheaper than opus 4.7 and 100% open source / free to use, i mean…

Model ReleasesDGX agent

dear god lol - new kimi model is a fucking beast. GPT 5.4 level coding, 76% cheaper than opus 4.7 and 100% open source / free to use, i mean look at this: > kimi k2.6 can code continuously for 12 hour

Foundation Models in Robotics: A Comprehensive Review of Methods, Models, Datasets, Challenges and Future Research Directions

ApplicationsDGX agent

arXiv:2604.15395v1 Announce Type: new Abstract: Over the recent years, the field of robotics has been undergoing a transformative paradigm shift from fixed, single-task, domain-specific solutions towa

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation

ResearchDGX agent

arXiv:2510.09065v2 Announce Type: replace-cross Abstract: We introduce MMAudioSep, a generative model for video/text-queried sound separation that is founded on a pretrained video-to-audio model. By l

17 Apr 2026

Constrained Decoding for Safe Robot Navigation Foundation Models

Model ReleasesDGX agent

arXiv:2509.01728v4 Announce Type: replace-cross Abstract: Recent advances in the development of robotic foundation models have led to promising end-to-end and general-purpose capabilities in robotic s

Constraint-based Pre-training: From Structured Constraints to Scalable Model Initialization

ResearchDGX agent

arXiv:2604.14769v1 Announce Type: new Abstract: The pre-training and fine-tuning paradigm has become the dominant approach for model adaptation. However, conventional pre-training typically yields mod

EEGDM: Learning EEG Representation with Latent Diffusion Model

Local AiDGX agent

arXiv:2508.20705v3 Announce Type: replace Abstract: Recent advances in self-supervised learning for EEG representation have largely relied on masked reconstruction, where models are trained to recover

Last week, Anthropic announced Project Glasswing alongside Claude Mythos Preview, a model they described as so powerful at finding vulnerabi…

Model ReleasesDGX agent

Last week, Anthropic announced Project Glasswing alongside Claude Mythos Preview, a model they described as so powerful at finding vulnerabilities they couldn't release it. The announcement featured A

16 Apr 2026

Diffusion Language Models for Speech Recognition

TutorialsDGX agent

arXiv:2604.14001v1 Announce Type: new Abstract: Diffusion language models have recently emerged as a leading alternative to standard language models, due to their ability for bidirectional attention a

Working Notes on Late Interaction Dynamics: Analyzing Targeted Behaviors of Late Interaction Models

Model ReleasesDGX agent

arXiv:2603.26259v2 Announce Type: replace-cross Abstract: While Late Interaction models exhibit strong retrieval performance, many of their underlying dynamics remain understudied, potentially hiding

15 Apr 2026

Causal Fingerprints of AI Generative Models

ResearchDGX agent

arXiv:2509.15406v2 Announce Type: replace Abstract: AI generative models leave implicit traces in their generated images, which are commonly referred to as model fingerprints and are exploited for sou

SeedPrints: Fingerprints Can Even Tell Which Seed Your Large Language Model Was Trained From

Model ReleasesDGX agent

arXiv:2509.26404v2 Announce Type: replace-cross Abstract: Fingerprinting Large Language Models (LLMs)is essential for provenance verification and model attribution. Existing fingerprinting methods are

14 Apr 2026

Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music

Model ReleasesDGX agent

arXiv:2604.10905v1 Announce Type: cross Abstract: We present Audio Flamingo Next (AF-Next), the next-generation and most capable large audio-language model in the Audio Flamingo series, designed to ad

AWS launches Amazon Bio Discovery, an AI-powered application designed to speed up drug development, giving scientists access to biological foundation models (Reuters)

Model ReleasesDGX agent

Reuters: AWS launches Amazon Bio Discovery, an AI-powered application designed to speed up drug development, giving scientists access to biological foundation models — Amazon's (AMZN.O) cloud unit on

DeepFleet: Multi-Agent Foundation Models for Mobile Robots

SafetyDGX agent

arXiv:2508.08574v3 Announce Type: replace Abstract: We introduce DeepFleet, a suite of foundation models designed to support coordination and planning for large-scale mobile robot fleets. These models

General-purpose LLMs as Models of Human Driver Behavior: The Case of Simplified Merging

Model ReleasesDGX agent

arXiv:2604.09609v1 Announce Type: new Abstract: Human behavior models are essential as behavior references and for simulating human agents in virtual safety assessment of automated vehicles (AVs), yet

Integrating SAINT with Tree-Based Models: A Case Study in Employee Attrition Prediction

ApplicationsDGX agent

arXiv:2604.10337v1 Announce Type: new Abstract: Employee attrition presents a major challenge for organizations, increasing costs and reducing productivity. Predicting attrition accurately enables pro

Introspective Diffusion Language Models

ResearchDGX agent

arXiv:2604.11035v1 Announce Type: new Abstract: Diffusion language models promise parallel generation, yet still lag behind autoregressive (AR) models in quality. We stem this gap to a failure of intr

Pioneer Agent: Continual Improvement of Small Language Models in Production

Model ReleasesDGX agent

arXiv:2604.09791v1 Announce Type: new Abstract: Small language models are attractive for production deployment due to their low cost, fast inference, and ease of specialization. However, adapting them

StyleBench: Evaluating thinking styles in Large Language Models

Model ReleasesDGX agent

arXiv:2509.20868v2 Announce Type: replace-cross Abstract: Structured reasoning can improve the inference performance of large language models (LLMs), but it also introduces computational cost and cont

Teaching Language Models How to Code Like Learners: Conversational Serialization for Student Simulation

Model ReleasesDGX agent

arXiv:2604.10720v1 Announce Type: new Abstract: Artificial models that simulate how learners act and respond within educational systems are a promising tool for evaluating tutoring strategies and feed

13 Apr 2026

2D or 3D: Who Governs Salience in VLA Models? -- Tri-Stage Token Pruning Framework with Modality Salience Awareness

ResearchDGX agent

arXiv:2604.09244v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as the mainstream of embodied intelligence. Recent VLA models have expanded their input modalities fr

BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning

Model ReleasesDGX agent

arXiv:2604.09378v1 Announce Type: cross Abstract: Agent ecosystems increasingly rely on installable skills to extend functionality, and some skills bundle learned model artifacts as part of their exec

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2512.02231v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) are expected to jointly interpret vision, audio, and language, yet existing video benchmarks rarely a

Toward Hardware-Agnostic Quadrupedal World Models via Morphology Conditioning

Model ReleasesDGX agent

arXiv:2604.08780v1 Announce Type: cross Abstract: World models promise a paradigm shift in robotics, where an agent learns the underlying physics of its environment once to enable efficient planning a

10 Apr 2026

Adversarial Flow Models

TutorialsDGX agent

arXiv:2511.22475v2 Announce Type: replace-cross Abstract: We present adversarial flow models, a class of generative models that belongs to both the adversarial and flow families. Our method supports n

An Automated Survey of Generative Artificial Intelligence: Large Language Models, Architectures, Protocols, and Applications

Model ReleasesDGX agent

arXiv:2306.02781v4 Announce Type: replace-cross Abstract: Generative artificial intelligence, and large language models in particular, have emerged as one of the most transformative paradigms in moder

Are Face Embeddings Compatible Across Deep Neural Network Models?

SafetyDGX agent

arXiv:2604.07282v1 Announce Type: cross Abstract: Automated face recognition has made rapid strides over the past decade due to the unprecedented rise of deep neural network (DNN) models that can be t

Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules

Model ReleasesDGX agent

arXiv:2604.06233v1 Announce Type: new Abstract: Safety-trained language models routinely refuse requests for help circumventing rules. But not all rules deserve compliance. When users ask for help eva

Toward Memory-Aided World Models: Benchmarking via Spatial Consistency

Model ReleasesDGX agent

arXiv:2505.22976v2 Announce Type: replace-cross Abstract: The ability to simulate the world in a spatially consistent manner is a crucial requirements for effective world models. Such a model enables

9 Apr 2026

After playing with it a bit, Meta's Muse Spark Thinking is fine so far, but really doesn't match the current Big Three models. It also is a …

Model ReleasesDGX agent

After playing with it a bit, Meta's Muse Spark Thinking is fine so far, but really doesn't match the current Big Three models. It also is a bit... weird. Like some strange language & tone, a little lo

Open everything 🔥 Open Harness, Model Choice, Open Memory (take it wherever you need), Open Protocols basically we’re in the middle of a mo…

AgentsDGX agent

Open everything 🔥 Open Harness, Model Choice, Open Memory (take it wherever you need), Open Protocols basically we’re in the middle of a model war, they all make great models but the optimal arrangeme

So what's the deal with Amazon Nova? They released Nova 2 in December, and even then, the top flight Nova 2 model trailed Sonnet 4.5. And it…

Model ReleasesDGX agent

Amazon released the Nova 2 model family on December 2, 2025 at AWS re:Invent, comprising Nova 2 Lite, Nova 2 Pro (Preview), Nova 2 Omni, and Nova 2 Sonic, all featuring a 1-million-token context wi...

12 Aug 2026

Interpreting Language Model Hidden States at Scale

Model ReleasesDGX agent

arXiv:2608.10260v1 Announce Type: new Abstract: Lens methods interpret large language models (LLMs) by mapping intermediate activations to the output vocabulary, revealing how next-token predictions d

Qwen3.8-2.4T-A95B is now live on Together AI. The Qwen Team’s latest flagship model is built for coding and long-horizon agent workflows, wi…

Model ReleasesDGX agent

The Qwen Team has released its flagship model, Qwen3.8‑2.4T‑A95B, on the Together AI platform (togethercompute) as of August 12 2026. This 2.4‑trillion‑parameter model is engineered for coding tasks a

11 Aug 2026

4D-WAM: Infusing Spatiotemporal Awareness into World Action Models through Trajectory Fields

Local AiDGX agent

arXiv:2608.08023v1 Announce Type: new Abstract: Building on recent advances in world models, World Action Models (WAMs) jointly model video prediction and action generation. However, they typically re

Concept-Guided Spatial Regularization for World Models in Atari Pong

SafetyDGX agent

arXiv:2607.15142v2 Announce Type: replace Abstract: World models are usually evaluated as components of model-based reinforcement learning (MBRL) systems, leaving their standalone reliability understu

I gave DeepSeek V4 Flash basic vision by training a 40M connector on 100K examples

Model ReleasesDGX agent

I wanted to find out whether a huge text-only MoE could be given basic vision without retraining the language model itself. The short answer is yes. I froze DeepSeek V4 Flash and a 417M-parameter Moon

Matryoshka Language Model Suites

Model ReleasesDGX agent

arXiv:2608.09703v1 Announce Type: new Abstract: Training a language model suite classically requires training each model separately and serving them independently. We improve both training and inferen

Observations on Muse-Glimmer reasoning traces being noticeably different from qwen / gemma models and questions for you guys

Model ReleasesDGX agent

Just downloaded the model, UD-Q5_K_XL quant, asked it to generate a long story to test out reasoning and speed with dflash (super fast btw, ~ 90 to 160 tok/s on a 5090 depending on task) and was surpr

Two-Layer Linear Auto-Regressive Models Estimate Latent States

Model ReleasesDGX agent

arXiv:2606.12691v2 Announce Type: replace-cross Abstract: Auto-regressive models have emerged as powerful tools for sequential data, from language to video. Understanding how and why these models lear

10 Aug 2026

Dueling World Models: Advantage-Style Action Channels for Common-Mode Distractor Rejection

AgentsDGX agent

arXiv:2608.06706v1 Announce Type: cross Abstract: Latent world models plan by predicting future states from an action, but when a scene contains motion the agent does not control, they quietly go acti

GRASP: Reinforcing Language Model Anonymizers with Group Relative Policy Optimization

Model ReleasesDGX agent

arXiv:2608.06526v1 Announce Type: new Abstract: Large language models can infer sensitive personal attributes, such as age, location, and occupation, from ordinary text, turning everyday writing into

MemWM: Memory-Augmented Text-Based World Model

Model ReleasesDGX agent

arXiv:2608.07107v1 Announce Type: new Abstract: World models are increasingly used to support planning in agents by predicting how environment states evolve in response to agent actions. Yet fluent ne

MI-MIDI: Mechanistic Interpretability of Text-to-MIDI Generation Models via Probing, Lenses and Steering

Model ReleasesDGX agent

arXiv:2608.06638v1 Announce Type: cross Abstract: Mechanistic interpretability of music generation has concentrated on audio models, leaving symbolic models largely unexplored. We analyze two public t

Progressive Alignment of Recommender Foundation Model through Multi-Phase Post-Training

SafetyDGX agent

arXiv:2608.06792v1 Announce Type: cross Abstract: Foundation model(FM) for recommendation has shown strong ability to model long-horizon sequential user behavior. In practice, a single pretrained foun

Quantum Generative Diffusion Model: A Fully Quantum-Mechanical Model for Generating Quantum State Ensemble

ResearchDGX agent

arXiv:2401.07039v5 Announce Type: replace-cross Abstract: Mixed quantum states are the native description of many physically important quantum systems, making their generation a fundamental task in qu

← Previous
1…1516171819…990
Next →