AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,490 results
15 May 2026

MultiEmo-Bench: Multi-label Visual Emotion Analysis for Multi-modal Large Language Models

Model ReleasesDGX agent

arXiv:2605.14635v1 Announce Type: cross Abstract: This paper introduces a multi-label visual emotion analysis benchmark dataset for comprehensively evaluating the ability of multimodal large language

Orchard: An Open-Source Agentic Modeling Framework

AgentsDGX agent

arXiv:2605.15040v1 Announce Type: new Abstract: Agentic modeling aims to transform LLMs into autonomous agents capable of solving complex tasks through planning, reasoning, tool use, and multi-turn in

Small, Private Language Models as Teammates for Educational Assessment Design

Local AiDGX agent

arXiv:2605.15015v1 Announce Type: new Abstract: Generative AI increasingly supports educational design tasks, e.g., through Large Language Models (LLMs), demonstrating the capability to design assessm

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Three new open-source models just landed in ComfyUI natively: → Gemma 4 (Google DeepMind) - multimodal LLM handling text, image, audio, and …

Model ReleasesDGX agent

Three new open-source models just landed in ComfyUI natively: → Gemma 4 (Google DeepMind) - multimodal LLM handling text, image, audio, and video input with built-in step-by-step reasoning mode → VOID

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations

SafetyDGX agent

arXiv:2511.02776v2 Announce Type: replace Abstract: Recent progress in large-scale robotic datasets and vision-language models (VLMs) has advanced research on vision-language-action (VLA) models. Howe

14 May 2026

Amortized Guidance for Image Inpainting with Pretrained Diffusion Models

ResearchDGX agent

arXiv:2605.13010v1 Announce Type: cross Abstract: We study image inpainting with generative diffusion models. Existing methods typically either train dedicated task-specific models, or adapt a pretrai

Bias In, Bias Out? Finding Unbiased Subnetworks in Vanilla Models

Model ReleasesDGX agent

arXiv:2603.05582v2 Announce Type: replace-cross Abstract: The issue of algorithmic biases in deep learning has led to the development of various debiasing techniques, many of which perform complex tra

Continual Fine-Tuning of Large Language Models via Program Memory

Model ReleasesDGX agent

arXiv:2605.13162v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT), particularly Low-Rank Adaptation (LoRA), has become a standard approach for adapting Large Language Models (LLMs

Earth Science Foundation Models: From Perception to Reasoning and Discovery

AgentsDGX agent

arXiv:2605.12542v1 Announce Type: cross Abstract: Large foundation models (FMs) are transforming Earth science by integrating heterogeneous multimodal data, such as multi-platform imagery, gridded rea

Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention

Model ReleasesDGX agent

arXiv:2605.12789v1 Announce Type: new Abstract: Large language-vision models (LVLMs) such as CLIP, Flamingo, and BLIP have revolutionized AI by enabling understanding across textual and visual modalit

Prismatic World Model: Learning Compositional Dynamics for Planning in Hybrid Systems

ResearchDGX agent

arXiv:2512.08411v2 Announce Type: replace Abstract: Model-based planning in robotic domains is challenged by the hybrid nature of physical dynamics, where continuous motion is punctuated by discrete e

RotVLA: Rotational Latent Action for Vision-Language-Action Model

ApplicationsDGX agent

arXiv:2605.13403v1 Announce Type: cross Abstract: Latent Action Models (LAMs) have emerged as an effective paradigm for handling heterogeneous datasets during Vision-Language-Action (VLA) model pretra

Sampling from Flow Language Models via Marginal-Conditioned Bridges

ResearchDGX agent

arXiv:2605.13681v1 Announce Type: new Abstract: Flow Language Models (FLMs) are a recently introduced class of language models which adapt continuous flow matching for one-hot encoded token sequences.

Three-Stage Learning Unlocks Strong Performance in Simple Models for Long-Term Time Series Forecasting

ResearchDGX agent

arXiv:2605.13678v1 Announce Type: new Abstract: Recent studies on long-term time series forecasting have shown that simple linear models and MLP-based predictors can achieve strong performance without

Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2605.13119v1 Announce Type: cross Abstract: Vision-language-action (VLA) models are effective robot action executors, but they remain limited on long-horizon tasks due to the dual burden of exte

Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context

AgentsDGX agent

arXiv:2605.13831v1 Announce Type: new Abstract: Long-context modeling is becoming a core capability of modern large vision-language models (LVLMs), enabling sustained context management across long-do

13 May 2026

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone

ResearchDGX agent

arXiv:2605.11405v1 Announce Type: new Abstract: Data curation has shifted the quality-compute frontier for language-model and contrastive image-text pretraining, but its role for vision-language model

A Proof-of-Concept Simulation-Driven Digital Twin Framework for Decision-Aware Diabetes Modeling

Model ReleasesDGX agent

arXiv:2605.11247v1 Announce Type: new Abstract: This paper presents a proof-of-concept digital twin framework for simulation-driven diabetes modeling using benchmark clinical data, synthetic temporal

A Theoretical Analysis of Why Masked Diffusion Models Mitigate the Reversal Curse

Model ReleasesDGX agent

arXiv:2602.02133v2 Announce Type: replace-cross Abstract: Autoregressive language models (ARMs) suffer from the reversal curse: after learning ''A is B,'' they often fail on the reverse query ''B is A

Curriculum Learning-Guided Progressive Distillation in Large Language Models

ResearchDGX agent

arXiv:2605.11260v1 Announce Type: new Abstract: Knowledge distillation is a key technique for transferring the capabilities of large language models (LLMs) into smaller, more efficient student models.

Debiased Model-based Representations for Sample-efficient Continuous Control

SafetyDGX agent

arXiv:2605.11711v1 Announce Type: new Abstract: Model-based representations recently stand out as a promising framework that embeds latent dynamics information into the representations for downstream

probably correct, from @polynoamial: “with today’s AI models, intelligence is a function of inference compute.” but what about tomorrow’s mo…

SafetyDGX agent

probably correct, from @polynoamial: “with today’s AI models, intelligence is a function of inference compute.” but what about tomorrow’s models? never forget that humans are remarkably intelligent (t

Stop Marginalizing My Dreams: Model Inversion via Laplace Kernel for Continual Learning

ResearchDGX agent

arXiv:2605.11804v1 Announce Type: cross Abstract: Data-free continual learning (DFCIL) relies on model inversion to synthesize pseudo-samples and mitigate catastrophic forgetting. However, existing in

Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling

ResearchDGX agent

arXiv:2605.05922v2 Announce Type: replace Abstract: Recent advances in generative video models are increasingly driven by post-training and test-time scaling, both of which critically depend on the qu

U-STS-LLM A Unified Spatio-Temporal Steered Large Language Model for Traffic Prediction and Imputation

Model ReleasesDGX agent

arXiv:2605.11735v1 Announce Type: new Abstract: The efficient operation of modern cellular networks hinges on the accurate analysis of spatio-temporal traffic data. Mastering these patterns is essenti

12 May 2026

A startup wants to pull magnesium from seawater without torching the environment. Another wants to take small language models to the podium …

Model ReleasesDGX agent

A startup wants to pull magnesium from seawater without torching the environment. Another wants to take small language models to the podium for enterprise customers. Today @jason and @alex sat down wi

Counterfactual Stress Testing for Image Classification Models

ApplicationsDGX agent

arXiv:2605.10894v1 Announce Type: new Abstract: Deep learning models in medical imaging often fail when deployed in new clinical environments due to distribution shifts in demographics, scanner hardwa

CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.10426v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing reasoning mechanisms sti

CrystalREPA: Transferring Physical Priors from Universal MLIPs to Crystal Generative Models

Model ReleasesDGX agent

arXiv:2605.08960v1 Announce Type: cross Abstract: Crystal generative models mainly learn what stable crystals look like, with little explicit supervision for what makes them stable. We reveal a substa

Data-driven transport modelling without overfit

SafetyDGX agent

arXiv:2605.08801v1 Announce Type: new Abstract: Macroscopic transport modelling aims to predict traffic flows after proposed public policy interventions, such as a new road or railway section or a tem

Emergent Semantic Role Understanding in Language Models

TutorialsDGX agent

arXiv:2605.09187v1 Announce Type: new Abstract: Understanding how linguistic structure emerges in language models is central to interpreting what these systems learn from data and how much supervision

Evidence-based Decision Modeling for Synthetic Face Detection with Uncertainty-driven Active Learning

ResearchDGX agent

arXiv:2605.09935v1 Announce Type: new Abstract: With the rapid development of deep generative models, forged facial images are massively exploited for illegal activities. Although existing synthetic f

Explanation Fairness in Large Language Models: An Empirical Analysis of Disparities in How LLMs Justify Decisions Across Demographic Groups

Model ReleasesDGX agent

arXiv:2605.08671v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed not only to make decisions but to explain them. While AI decision fairness has been studied ext

Federated Language Models Under Bandwidth Budgets: Distillation Rates and Conformal Coverage

Model ReleasesDGX agent

arXiv:2605.09986v1 Announce Type: cross Abstract: Training a language model on data scattered across bandwidth-limited nodes that cannot be centralized is a setting that arises in clinical networks, e

IntroLM: Introspective Language Models via Prefilling-Time Self-Evaluation

TutorialsDGX agent

arXiv:2601.03511v2 Announce Type: replace-cross Abstract: A major challenge for the operation of large language models (LLMs) is how to predict whether a specific LLM will produce sufficiently high-qu

Kernel-Gradient Drifting Models

ResearchDGX agent

arXiv:2605.10727v1 Announce Type: new Abstract: We propose kernel-gradient drifting, a one-step generative modeling framework that replaces the fixed Euclidean displacement direction in drifting model

Latent Geometry Beyond Search: Amortizing Planning in World Models

Model ReleasesDGX agent

arXiv:2605.08732v1 Announce Type: cross Abstract: Modern vision-based world models can represent observations as compact yet expressive latent manifolds, but fast goal-oriented planning in these space

Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.08787v1 Announce Type: new Abstract: Recent advances in 3D medical vision-language models have enabled joint reasoning over volumetric images and text, showing strong performance in medical

MicroWorld: Empowering Multimodal Large Language Models to Bridge the Microscopic Domain Gap with Multimodal Attribute Graph

Model ReleasesDGX agent

arXiv:2605.10120v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) show remarkable potential for scientific reasoning, yet their performance in specialized domains such as micr

Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models

Model ReleasesDGX agent

arXiv:2510.09592v2 Announce Type: replace Abstract: Real-time Spoken Language Models (SLMs) struggle to leverage Chain-of-Thought (CoT) reasoning due to the prohibitive latency of generating the entir

MolRGen: A Training and Evaluation Setting for De Novo Molecular Generation with Reasonning Models

Model ReleasesDGX agent

arXiv:2603.18256v2 Announce Type: replace-cross Abstract: Recent reasoning-based large language models have shown strong performance on tasks with verifiable outcomes, but their use in de novo molecul

Personal Visual Context Learning in Large Multimodal Models

Model ReleasesDGX agent

arXiv:2605.10936v1 Announce Type: new Abstract: As wearable devices like smart glasses integrate Large Multimodal Models (LMMs) into the continuous first-person visual streams of individual users, the

Product-of-Gaussian-Mixture Diffusion Models for Joint Nonlinear MRI Reconstruction

Model ReleasesDGX agent

arXiv:2605.10629v1 Announce Type: new Abstract: Recently, diffusion models have attracted considerable attention for magnetic resonance image reconstruction due to their high sample quality. However,

PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models

SafetyDGX agent

arXiv:2501.03544v5 Announce Type: replace-cross Abstract: Recent text-to-image (T2I) models have exhibited remarkable performance in generating high-quality images from text descriptions. However, the

Reasoning emerges from constrained inference manifolds in large language models

Model ReleasesDGX agent

arXiv:2605.08142v1 Announce Type: cross Abstract: Reasoning in large language models is predominantly evaluated through labeled benchmarks, conflating task performance with the quality of internal inf

Reinforce Adjoint Matching: Scaling RL Post-Training of Diffusion and Flow-Matching Models

ResearchDGX agent

arXiv:2605.10759v1 Announce Type: cross Abstract: Diffusion and flow-matching models scale because pretraining is supervised regression: a clean sample is noised analytically, and a model regresses ag

Relational reasoning and inductive bias in transformers and large language models

SafetyDGX agent

arXiv:2506.04289v3 Announce Type: replace Abstract: Transformer-based models have demonstrated remarkable reasoning abilities, but the mechanisms underlying relational reasoning remain poorly understo

Relative Kinetic Utility for Reasoning-Aware Structural Pruning in Large Language Models

Model ReleasesDGX agent

arXiv:2605.09008v1 Announce Type: cross Abstract: Chain-of-Thought (CoT) prompting symbolized a huge improvement of reasoning capabilities of Large Language Models (LLMs). However, scaling up test-tim

Sundial: A Family of Highly Capable Time Series Foundation Models

ApplicationsDGX agent

arXiv:2502.00816v4 Announce Type: replace Abstract: We introduce Sundial, a family of native, flexible, and scalable time series foundation models. To predict the next-patch's distribution, we propose

TFM-Retouche: A Lightweight Input-Space Adapter for Tabular Foundation Models

Model ReleasesDGX agent

arXiv:2605.06047v2 Announce Type: replace-cross Abstract: Tabular foundation models (TFMs), such as TabPFN-2.6, TabICLv2, ConTextTab, Mitra, LimiX, and TabDPT, achieve strong zero-shot performance thr

The Safety-Aware Denoiser for Text Diffusion Models

SafetyDGX agent

arXiv:2605.08116v1 Announce Type: cross Abstract: Recent work on text diffusion models offers a promising alternative to autoregressive generation, but controlling their safety remains underexplored.

The two clocks and the innovation window: When and how generative models learn rules

TutorialsDGX agent

arXiv:2605.10019v1 Announce Type: cross Abstract: Generative models trained on finite data face a fundamental tension: their score-matching or next-token objective converges to the empirical training

VLADriver-RAG: Retrieval-Augmented Vision-Language-Action Models for Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.08133v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving, yet their reliance on implicit parametric

When to Trust Imagination: Adaptive Action Execution for World Action Models

Model ReleasesDGX agent

arXiv:2605.06222v2 Announce Type: replace-cross Abstract: World Action Models (WAMs) have recently emerged as a promising paradigm for robotic manipulation by jointly predicting future visual observat

11 May 2026

A Reproducible Optimisation Protocol for Calibrating Prompt-Based Large Language Model Workflows in Evidence Synthesis

Model ReleasesDGX agent

arXiv:2605.06937v1 Announce Type: new Abstract: This methods article presents a reproducible calibration workflow for prompt-based large language models (LLMs) in structured evidence-synthesis tasks.

Beyond Retrieval: A Multitask Benchmark and Model for Code Search

Model ReleasesDGX agent

arXiv:2605.04615v2 Announce Type: replace-cross Abstract: Code search has usually been evaluated as first-stage retrieval, even though production systems rely on broader pipelines with reranking and d

Causal-Aware Foundation-Model for Bilevel Optimization in Discrete Choice Settings

ApplicationsDGX agent

arXiv:2605.06941v1 Announce Type: new Abstract: We introduce a causal aware foundation-model framework for real time optimal decision making in discrete choice environments. We propose a constrained t

MIND: Monge Inception Distance for Generative Models Evaluation

Model ReleasesDGX agent

arXiv:2605.06797v1 Announce Type: new Abstract: We propose the Monge Inception Distance (MIND), a metric for evaluating generative models that addresses key limitations of the widely adopted Frechet I

On the Tradeoffs of On-Device Generative Models in Federated Predictive Maintenance Systems

Local AiDGX agent

arXiv:2605.07860v1 Announce Type: cross Abstract: Federated Learning (FL) has emerged as a promising paradigm for preserving client data ownership and control over distributed Internet of Things (IoT)

One of the most important properties of LLMs that we take for granted is that newer, bigger models are just better at everything. The AI Lab…

Model ReleasesDGX agent

One of the most important properties of LLMs that we take for granted is that newer, bigger models are just better at everything. The AI Labs are pouring effort into economically valuable fields like

← Previous
1…8990919293…1009
Next →