AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,537 results
12 May 2026

How open model ecosystems compound

ResearchDGX agent

This article examines how open-source AI model ecosystems create compounding effects through community contributions, fine-tuning, and iterative improvements that accelerate innovation and accessibili

Learning stochastic multiscale models through normalizing flows

ResearchDGX agent

arXiv:2605.09718v1 Announce Type: cross Abstract: Many systems in physics, engineering, and biology exhibit multiscale stochastic dynamics, where low-dimensional slow variables evolve under the influe

Less Redundancy: Boosting Practicality of Vision Language Model in Walking Assistants

ResearchDGX agent

arXiv:2508.16070v3 Announce Type: replace Abstract: Approximately 283 million people worldwide live with visual impairments, motivating increasing research into leveraging Visual Language Models (VLMs

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

LLM4Branch: Large Language Model for Discovering Efficient Branching Policies of Integer Programs

Model ReleasesDGX agent

arXiv:2605.10401v1 Announce Type: new Abstract: Efficient branching policies are essential for accelerating Mixed Integer Linear Programming (MILP) solvers. Their design has long relied on hand-crafte

LoKA: Low-precision Kernel Applications for Recommendation Models At Scale

HardwareDGX agent

arXiv:2605.10886v1 Announce Type: cross Abstract: Recent GPU generations deliver significantly higher FLOPs using lower-precision arithmetic, such as FP8. While successfully applied to large language

LPT: Less-overfitting Prompt Tuning for Vision-Language Model

ResearchDGX agent

arXiv:2410.10247v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have demonstrated exceptional generalization capabilities for downstream tasks. Due to its efficiency, prompt le

MedThink: Enhancing Diagnostic Accuracy in Small Models via Teacher-Guided Reasoning Correction

ResearchDGX agent

arXiv:2605.08094v1 Announce Type: cross Abstract: Accurate clinical diagnosis requires extensive domain knowledge and complex clinical reasoning capabilities. Although large language models (LLMs) hol

PAINET: A Principled Efficient Transformer for 3D Dynamics Modeling

ApplicationsDGX agent

arXiv:2510.04233v2 Announce Type: replace-cross Abstract: Modeling 3D dynamics is a fundamental problem in multi-body systems across scientific and engineering domains and has important practical impl

PermuQuant: Lowering Per-Group Quantization Error by Reordering Channels for Diffusion Models

HardwareDGX agent

arXiv:2605.09503v1 Announce Type: new Abstract: Large-scale visual generative models have achieved remarkable performance. However, their high computational and memory costs make deployment challengin

Physics-Modeled Neural Networks

ResearchDGX agent

arXiv:2605.08176v1 Announce Type: new Abstract: We introduce Dynamical Physics-Modeled Neural Networks (DynPMNNs), a continuous-time deep learning architecture in which each hidden layer is defined as

Privacy-Preserving Distributed Learning in IoT Systems: A Unified Threat Model and Evaluation Framework

Local AiDGX agent

arXiv:2605.09232v1 Announce Type: cross Abstract: The increasing deployment of Internet-of-Things (IoT) devices has accelerated the use of distributed learning frameworks, where data remains local whi

Rethinking Evaluation of Multiple Sclerosis (MS) Lesion Segmentation Models

ApplicationsDGX agent

arXiv:2605.09666v1 Announce Type: cross Abstract: Multiple Sclerosis (MS) is a chronic autoimmune disease that can significantly reduce the quality of life of a patient. Existing treatment options can

Reward Auditor: Inference on Reward Modeling Suitability in Real-World Perturbed Scenarios

SafetyDGX agent

arXiv:2512.00920v4 Announce Type: replace Abstract: Reliable reward models (RMs) are critical for ensuring the safe alignment of large language models (LLMs). However, current RM evaluation methods fo

Simulus: Combining Improvements in Sample-Efficient World Model Agents

AgentsDGX agent

arXiv:2502.11537v4 Announce Type: replace-cross Abstract: World models (WMs) represent the frontier of sample-efficient reinforcement learning, but their complexity leaves many promising improvements

Sparse Reward Subsystem in Large Language Models

TutorialsDGX agent

arXiv:2602.00986v2 Announce Type: replace Abstract: Recent studies show that LLM hidden states encode reward-related information, such as answer correctness and model confidence. However, existing app

Supervised Guidance Training for Infinite-Dimensional Diffusion Models

ResearchDGX agent

arXiv:2601.20756v2 Announce Type: replace Abstract: Score-based diffusion models have recently been extended to infinite-dimensional function spaces, with uses such as inverse problems arising from pa

TinySSL: Distilled Self-Supervised Pretraining for Sub-Megabyte MCU Models

ResearchDGX agent

arXiv:2605.08241v1 Announce Type: cross Abstract: Self-supervised learning (SSL) has transformed representation learning for large models, yet remains unexplored for microcontroller (MCU)-class models

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse

Model ReleasesDGX agent

arXiv:2605.09965v1 Announce Type: new Abstract: The real world unfolds along a single set of physics laws, yet human intelligence demonstrates a remarkable capacity to generalize experiences from this

Towards Universal Gene Regulatory Network Inference: Unlocking Generalizable Regulatory Knowledge in Single-cell Foundation Models

Model ReleasesDGX agent

arXiv:2605.08128v1 Announce Type: cross Abstract: Gene Regulatory Network (GRN) inference is essential for understanding complex cellular mechanisms, rendered tractable through single-cell transcripto

Transfer Learning of Multiobjective Indirect Low-Thrust Trajectories Using Diffusion Models and Markov Chain Monte Carlo

Model ReleasesDGX agent

arXiv:2605.09125v1 Announce Type: cross Abstract: Preliminary low-thrust spacecraft mission design is a global search problem characterized by a complex solution landscape, multiple objectives, and nu

VFM-SDM: A vision foundation model-based framework for training-free, marker-free, and calibration-free structural displacement measurement

Model ReleasesDGX agent

arXiv:2605.09677v1 Announce Type: new Abstract: Reliable displacement measurement is fundamental for structural health monitoring and digital engineering workflows, as it provides direct structural re

VISOR: A Vision-Language Model-based Test Oracle for Testing Robot

Model ReleasesDGX agent

arXiv:2605.10408v1 Announce Type: cross Abstract: Testing robots requires assessing whether they perform their intended tasks correctly, dependably, and with high quality, a challenge known as the tes

We published new research on how we serve post-trained Qwen3 235B models on NVIDIA GB200 NVL72 Blackwell racks. GB200 is a major step up ove…

HardwareDGX agent

We published new research on how we serve post-trained Qwen3 235B models on NVIDIA GB200 NVL72 Blackwell racks. GB200 is a major step up over Hopper for high-throughput inference on large MoE models,

You Have Been LaTeXpOsEd: A Systematic Analysis of Information Leakage in Preprint Archives Using Large Language Models

Model ReleasesDGX agent

arXiv:2510.03761v2 Announce Type: replace-cross Abstract: The widespread use of preprint repositories such as arXiv has accelerated the communication of scientific results but also introduced overlook

11 May 2026

A Causal Diffusion Model for Video Reconstruction from Ultra-Low-Bitrate Representations

Model ReleasesDGX agent

arXiv:2602.13837v2 Announce Type: replace Abstract: We study video reconstruction from ultra-low-bitrate representations, where the primary challenge shifts from encoding to decoding. In this regime,

A Foundation Model for Instruction-Conditioned In-Context Time Series Tasks

ResearchDGX agent

arXiv:2603.22586v2 Announce Type: replace Abstract: In-context learning (ICL) enables task adaptation at inference time by conditioning on demonstrations rather than updating model parameters. Althoug

A Unified Measure-Theoretic View of Diffusion, Score-Based, and Flow Matching Generative Models

ResearchDGX agent

arXiv:2605.06829v1 Announce Type: cross Abstract: We survey continuous-time generative modeling methods based on transporting a simple reference distribution to a data distribution via stochastic or d

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs

Model ReleasesDGX agent

arXiv:2605.07731v1 Announce Type: cross Abstract: This report benchmarks the performance of ENGINEERING Ingegneria Informatica S.p.A.'s EngGPT2MoE-16B-A3B LLM, a 16B parameter Mixture of Experts (MoE)

Dealing with quirks introduced by switching models doesn't have to be hard -- we recently introduced a 'harness profile' API in Deep Agents …

AgentsDGX agent

Dealing with quirks introduced by switching models doesn't have to be hard -- we recently introduced a 'harness profile' API in Deep Agents as a solution. Profiles adjust system prompts, tool descript

DiffRetriever: Parallel Representative Tokens for Retrieval with Diffusion Language Models

ResearchDGX agent

arXiv:2605.07210v1 Announce Type: cross Abstract: PromptReps showed that an autoregressive language model can be used directly as a retriever by prompting it to generate dense and sparse representatio

Distributional simplicity bias and effective convexity in Energy Based Models

Local AiDGX agent

arXiv:2605.07844v1 Announce Type: new Abstract: Energy-based learning is a powerful framework for generative modelling, but its training is inherently non-convex, leading potentially to sensitivity to

Emergent Manifold Separability during Reasoning in Large Language Models

Local AiDGX agent

arXiv:2602.20338v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting significantly improves reasoning in Large Language Models, yet the temporal dynamics of the underlying representati

GEM: Generating LiDAR World Model via Deformable Mamba

AgentsDGX agent

arXiv:2605.07326v1 Announce Type: new Abstract: World models, which simulate environmental dynamics and generate sensor observations, are gaining increasing attention in autonomous driving. However, p

Generalised Linear Models in Deep Bayesian RL with Learnable Basis Functions

SafetyDGX agent

arXiv:2512.20974v3 Announce Type: replace-cross Abstract: Bayesian Reinforcement Learning (BRL), a subclass of Meta-Reinforcement Learning (Meta-RL), provides a principled framework for generalisation

Human-like fleeting memory improves language learning but impairs reading time prediction in transformer language models

TutorialsDGX agent

arXiv:2508.05803v2 Announce Type: replace Abstract: Human memory is fleeting. As words are processed, the exact wordforms that make up incoming sentences are rapidly lost. Cognitive scientists have lo

Is the Future Compatible? Diagnosing Dynamic Consistency in World Action Models

SafetyDGX agent

arXiv:2605.07514v1 Announce Type: cross Abstract: World Action Models (WAMs) enable decision-making through imagined rollouts by predicting future observations and actions. However, the reliability of

LithoBench: Benchmarking Large Multimodal Models for Remote-Sensing Lithology Interpretation

Model ReleasesDGX agent

arXiv:2605.07640v1 Announce Type: cross Abstract: Remote sensing lithology interpretation is fundamental to geological surveys, mineral exploration, and regional geological mapping. Unlike general lan

Miner:Mining Intrinsic Mastery for Data-Efficient RL in Large Reasoning Models

SafetyDGX agent

arXiv:2601.04731v2 Announce Type: replace Abstract: Current critic-free RL methods for large reasoning models suffer from severe inefficiency when training on positive homogeneous prompts (where all r

Pan-FM: A Pan-Organ Foundation Model with Saliency-Guided Masking for Missing Robustness

SafetyDGX agent

arXiv:2605.07055v1 Announce Type: cross Abstract: Foundation models (FMs) have shown great promise in medical imaging, but most FMs are trained on unimodal data within isolated domains, such as brain

PSK@EEUCA 2026: Fine-Tuning Large Language Models with Synthetic Data Augmentation for Multi-Class Toxicity Detection in Gaming Chat

Model ReleasesDGX agent

arXiv:2605.07201v1 Announce Type: cross Abstract: This paper describes our system for the EEUCA 2026 Shared Task on Understanding Toxic Behavior in Gaming Communities. The task involves classifying Wo

Render, Don't Decode: Weight-Space World Models with Latent Structural Disentanglement

HardwareDGX agent

arXiv:2605.06298v2 Announce Type: replace-cross Abstract: Training world models on vast quantities of unlabelled videos is a critical step toward fully autonomous intelligence. However, the prevailing

Script Sensitivity: Benchmarking Language Models on Unicode, Romanized and Mixed-Script Sinhala

ApplicationsDGX agent

arXiv:2601.14958v3 Announce Type: replace-cross Abstract: The performance of Language Models (LMs) on low-resource, morphologically rich languages like Sinhala remains largely unexplored, particularly

Sparser, Faster, Lighter Transformer Language Models

HardwareDGX agent

arXiv:2603.23198v2 Announce Type: replace-cross Abstract: Scaling autoregressive large language models (LLMs) has driven unprecedented progress but comes with vast computational costs. In this work, w

This is harder to build than it looks. Preserving full conversational context while swapping underlying model providers mid-flight is a surp…

AgentsDGX agent

This is harder to build than it looks. Preserving full conversational context while swapping underlying model providers mid-flight is a surprisingly deep systems problem. Most tools drop state or forc

This seems like a critical reason to open up about AI use in academia. Scholars are using old AI models, badly, and not talking about it. Ne…

AgentsDGX agent

This seems like a critical reason to open up about AI use in academia. Scholars are using old AI models, badly, and not talking about it. New models hallucinate very few citations, and good agentic ha

VDEGaussian: Video Diffusion Enhanced 4D Gaussian Splatting for Dynamic Urban Scenes Modeling

SafetyDGX agent

arXiv:2508.02129v2 Announce Type: replace Abstract: Dynamic urban scene modeling is a rapidly evolving area with broad applications. While current approaches leveraging neural radiance fields or Gauss

Versatile yet Efficient Network Traffic Analysis: Offloading Network Foundation Model to SmartNIC

HardwareDGX agent

arXiv:2508.02001v2 Announce Type: replace-cross Abstract: Pervasive encryption makes large-scale labeling infeasible for traffic analysis, while security operations demand edge analysis to avert servi

We just wrapped recording with @jerryjliu0 of @llama_index 👀 His thesis: the AI framework era is over. Scaffolding collapses into the model…

ApplicationsDGX agent

We just wrapped recording with @jerryjliu0 of @llama_index 👀 His thesis: the AI framework era is over. Scaffolding collapses into the model. Context quality = moat that survives. Frontier models still

8 May 2026

ai coding is getting expensive use more open models!

AgentsDGX agent

ai coding is getting expensive use more open models! Kimi K2.6 on @baseten is ~5x cheaper than Opus 4.7 For a large majority of tasks, it's roughly the same performance If you want to use open models

Deploy and inference any model from HuggingFace

HardwareDGX agent

Learn how to deploy any Hugging Face model in one session using Goose and Together's Dedicated Container Inference. Skip the setup complexity — one prompt gets your model running in a production-grade

7 May 2026

A Biased Nonnegative Block Term Tensor Decomposition Model for Dynamic QoS Prediction

Model ReleasesDGX agent

arXiv:2605.04813v1 Announce Type: new Abstract: With the rapid development of cloud computing and Web services, Quality of Service (QoS) has become a key criterion for service selection and recommenda

A geometric relation of the error introduced by sampling a language model's output distribution to its internal state

ResearchDGX agent

arXiv:2605.04899v1 Announce Type: new Abstract: GPT-style language models are sensitive to single-token changes at generation points where the predicted probability distribution is spread across multi

A Regulatory Governance Framework for AI-Driven Financial Fraud Detection in U.S. Banking: Integrating OCC, SR 11-7, CFPB, and FinCEN Compliance Requirements for Model Development, Validation, and Monitoring Lifecycles

Model ReleasesDGX agent

arXiv:2605.04076v1 Announce Type: new Abstract: U.S. financial institutions deploying AI-based fraud detection face a fragmented compliance landscape spanning four regulatory frameworks -- OCC Bulleti

A Universal Large Language Model -- Drone Command and Control Interface

Model ReleasesDGX agent

arXiv:2601.15486v2 Announce Type: replace Abstract: The use of artificial intelligence (AI) for drone control can have a transformative impact on drone capabilities, especially when real world informa

Actionable Real-Time Modeling of Surgical Team Dynamics via Time-Expanded Interaction Graphs

ResearchDGX agent

arXiv:2605.04169v1 Announce Type: cross Abstract: Surgical team performance arises from complex interactions between technical execution and non-technical skills, including communication and coordinat

ArtiFixer: Enhancing and Extending 3D Reconstruction with Auto-Regressive Diffusion Models

ResearchDGX agent

arXiv:2603.00492v2 Announce Type: replace Abstract: Per-scene optimization methods such as 3D Gaussian Splatting provide state-of-the-art novel view synthesis quality but extrapolate poorly to under-o

Counter-Dyna: Data-Efficient RL-Based HVAC Control using Counterfactual Building Models

SafetyDGX agent

arXiv:2605.04555v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL) offers a promising approach for data-efficient energy management in buildings, combining the strengths of pred

DART: A Vision-Language Foundation Model for Comprehensive Rope Condition Monitoring

Model ReleasesDGX agent

arXiv:2605.04943v1 Announce Type: new Abstract: The condition monitoring (CM) of synthetic fibre ropes (SFRs) used in offshore, maritime, and industrial settings demands more than a classifier: inspec

Detecting Hallucinations in Large Language Models via Internal Attention Divergence Signals

ResearchDGX agent

arXiv:2605.05025v1 Announce Type: new Abstract: We propose a lightweight and single-pass uncertainty quantification method for detecting hallucinations in Large Language Models. The method uses attent

Dream-MPC: Gradient-Based Model Predictive Control with Latent Imagination

SafetyDGX agent

arXiv:2605.04568v1 Announce Type: new Abstract: State-of-the-art model-based Reinforcement Learning (RL) approaches either use gradient-free, population-based methods for planning, learned policy netw

← Previous
1…138139140141142…1009
Next →