AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,515 results
18 May 2026

yeah that's pretty good xAI might be able to cook with Cursor data + 10T model

IndustryDGX agent

yeah that's pretty good xAI might be able to cook with Cursor data + 10T model Introducing Composer 2.5, our most powerful model yet. It's more intelligent, better at sustained work on long-running ta

15 May 2026

Action-Inspired Generative Models

Model ReleasesDGX agent

arXiv:2605.14631v1 Announce Type: cross Abstract: We introduce Action-Inspired Generative Models (AGMs), a dual-network generative framework motivated by the observation that existing bridge-matching

Critic-Driven Voronoi-Quantization for Distilling Deep RL Policies to Explainable Models

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.14897v1 Announce Type: cross Abstract: Despite many successful attempts at explaining Deep Reinforcement Learning policies using distillation, it remains difficult to balance the performanc

EponaV2: Driving World Model with Comprehensive Future Reasoning

SafetyDGX agent

arXiv:2605.14696v1 Announce Type: new Abstract: Data scaling plays a pivotal role in the pursuit of general intelligence. However, the prevailing perception-planning paradigm in autonomous driving rel

Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model

SafetyDGX agent

arXiv:2605.14950v1 Announce Type: new Abstract: Vision-Language-Action models have emerged as a promising paradigm for robotic manipulation by unifying perception, language grounding, and action gener

FedStain: Modeling Higher-Order Stain Statistics for Federated Domain Generalization in Computational Pathology

Model ReleasesDGX agent

arXiv:2605.14590v1 Announce Type: new Abstract: Robust whole-slide image (WSI) analysis under strict data-governance remains challenging due to substantial cross-institutional stain heterogeneity. Dom

MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.14906v1 Announce Type: new Abstract: Memory is essential for large vision-language models (LVLMs) to handle long, multimodal interactions, with two method directions providing this capabili

Octopus: History-Free Gradient Orthogonalization for Continual Learning in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2605.14938v1 Announce Type: cross Abstract: Continual learning in multimodal large language models (MLLMs) aims to sequentially acquire knowledge while mitigating catastrophic forgetting, yet ex

Quantitative Video World Model Evaluation for Geometric-Consistency

SafetyDGX agent

arXiv:2605.15185v1 Announce Type: cross Abstract: Generative video models are increasingly studied as implicit world models, yet evaluating whether they produce physically plausible 3D structure and m

RePack then Refine: Efficient Diffusion Transformer with Vision Foundation Model

ResearchDGX agent

arXiv:2512.12083v3 Announce Type: replace Abstract: Semantic-rich features from Vision Foundation Models (VFMs) have been leveraged to enhance Latent Diffusion Models (LDMs). However, raw VFM features

Rethinking Output Alignment For 1-bit Post-Training Quantization of Large Language Models

SafetyDGX agent

arXiv:2512.21651v2 Announce Type: replace Abstract: Large Language Models (LLMs) deliver strong performance across a wide range of NLP tasks, but their massive sizes hinder deployment on resource-cons

Scalable Krylov Subspace Methods for Generalized Mixed-Effects Models with Crossed Random Effects

ApplicationsDGX agent

arXiv:2505.09552v3 Announce Type: replace-cross Abstract: Mixed-effects models are widely used to model data with hierarchical grouping structures and high-cardinality categorical predictor variables.

SemaTune: Semantic-Aware Online OS Tuning with Large Language Models

Model ReleasesDGX agent

arXiv:2605.15026v1 Announce Type: cross Abstract: Online OS tuning can improve long-running services, but existing controllers are poorly matched to live hosts. They treat scheduler, power, memory, an

Teaching Large Language Models When Not to Know: Learning Temporal Critique for Ex-Ante Reasoning

ResearchDGX agent

arXiv:2605.14636v1 Announce Type: new Abstract: Large language models (LLMs) often fail to reason under temporal cutoffs: when prompted to answer from the standpoint of an earlier time, they exploit k

Uncertainty Quantification for Large Language Diffusion Models

ResearchDGX agent

arXiv:2605.14570v1 Announce Type: new Abstract: Large Language Diffusion Models (LLDMs) are emerging as an alternative to autoregressive models, offering faster inference through higher parallelism. S

Weekends are for vibe coding. But are your vibes continuously improving? Fine-tune your own model → stop waiting on someone else's release c…

Model ReleasesDGX agent

Weekends are for vibe coding. But are your vibes continuously improving? Fine-tune your own model → stop waiting on someone else's release cycle. Today's training update: Gemma 4 Dense is now availabl

14 May 2026

A Markov Categorical Framework for Language Modeling

ResearchDGX agent

arXiv:2507.19247v5 Announce Type: replace-cross Abstract: Autoregressive language models achieve remarkable performance, yet a unified theory explaining their internal mechanisms, how training shapes

any time a model router company drops data, its worth browsing. here we learn that gemini leads in education and personal assistants (?!), a…

Model ReleasesDGX agent

any time a model router company drops data, its worth browsing. here we learn that gemini leads in education and personal assistants (?!), ant leads in vibecoding and koding and back office (?!), and

CADDesigner: Conceptual CAD Model Generation with a General-Purpose Agent

AgentsDGX agent

arXiv:2508.01031v5 Announce Type: replace Abstract: Computer-Aided Design (CAD) is widely used for conceptual design and parametric 3D modeling, but typically requires a high level of expertise from d

ChatSR: Multimodal Large Language Models for Scientific Formula Discovery

SafetyDGX agent

arXiv:2406.05410v3 Announce Type: replace Abstract: Current multimodal large language models (MLLMs) are mainly focused on the understanding and processing of perceptual modalities such as images and

Compact Latent Manifold Translation: A Parameter-Efficient Foundation Model for Cross-Modal and Cross-Frequency Physiological Signal Synthesis

Model ReleasesDGX agent

arXiv:2605.13248v1 Announce Type: cross Abstract: The analysis of physiological time series, such as electrocardiograms (ECG) and photoplethysmograms (PPG), is persistently hindered by modality and fr

Connecting the Dots: A Machine Learning Ready Dataset for Ionospheric Forecasting Models

Model ReleasesDGX agent

arXiv:2511.15743v2 Announce Type: replace Abstract: Operational forecasting of the ionosphere remains a critical space weather challenge due to sparse observations, complex coupling across geospatial

Continual Learning with Multilingual Foundation Model

ResearchDGX agent

arXiv:2605.13415v1 Announce Type: cross Abstract: This paper presents a multi-stage framework for detecting reclaimed slurs in multilingual social media discourse. It addresses the challenge of identi

Diffusion Model's Generalization Can Be Characterized by Inductive Biases toward a Data-Dependent Ridge Manifold

SafetyDGX agent

arXiv:2602.06021v2 Announce Type: replace-cross Abstract: We study a data-dependent notion of diffusion-model generalization: when a model does not memorize the training set, where do its generated sa

DiffusionHijack: Supply-Chain PRNG Backdoor Attack on Diffusion Models and Quantum Random Number Defense

SafetyDGX agent

arXiv:2605.13115v1 Announce Type: cross Abstract: Diffusion models depend on pseudo-random number generators (PRNGs) for latent noise sampling. We present DiffusionHijack, a supply-chain backdoor atta

DistractMIA: Black-Box Membership Inference on Vision-Language Models via Semantic Distraction

Model ReleasesDGX agent

arXiv:2605.12574v1 Announce Type: cross Abstract: Vision-language models (VLMs) are trained on large-scale image-text corpora that may contain private, copyrighted, or otherwise sensitive data, motiva

It's not the Language Model, it's the Tool: Deterministic Mediation for Scientific Workflows

ApplicationsDGX agent

arXiv:2605.13245v1 Announce Type: new Abstract: Language models can produce convincing scientific analyses, but repeated generations on the same data do not guarantee the same result. A researcher may

KAST-BAR: Knowledge-Anchored Semantically-Dynamic Topology Brain Autoregressive Modeling for Universal Neural Interpretation

Local AiDGX agent

arXiv:2605.13133v1 Announce Type: new Abstract: While EEG foundation models have shown significant potential in universal neural decoding across tasks, their advancement remains constrained by the ina

Latent-Augmented Discrete Diffusion Models

ResearchDGX agent

arXiv:2510.18114v3 Announce Type: replace-cross Abstract: Discrete diffusion models have emerged as a powerful class of models and a promising route to fast language generation, but practical implemen

Like any AI dev not employed by a closed lab, we share the ambition that at least 10x more should be capable of training frontier models in …

Model ReleasesDGX agent

Like any AI dev not employed by a closed lab, we share the ambition that at least 10x more should be capable of training frontier models in 2026. And like we all know, some 10x leaps can't survive a c

LoREnc: Low-Rank Encryption for Securing Foundation Models and LoRA Adapters

Local AiDGX agent

arXiv:2605.13163v1 Announce Type: cross Abstract: Foundation models and low-rank adapters enable efficient on-device generative AI but raise risks such as intellectual property leakage and model recov

Mechanism Plausibility in Generative Agent-Based Modeling

AgentsDGX agent

arXiv:2605.12824v1 Announce Type: cross Abstract: Large language models (LLMs) can generate high-level diverse phenomena without explicitly programmed rules. This capability has led to their adoption

MIDST Challenge at SaTML 2025: Membership Inference over Diffusion-models-based Synthetic Tabular data

ResearchDGX agent

arXiv:2603.19185v2 Announce Type: replace Abstract: Synthetic data is often perceived as a silver-bullet solution to data anonymization and privacy-preserving data publishing. Drawn from generative mo

QLAM: A Quantum Long-Attention Memory Approach to Long-Sequence Token Modeling

ResearchDGX agent

arXiv:2605.13833v1 Announce Type: cross Abstract: Modeling long-range dependencies in sequential data remains a central challenge in machine learning. Transformers address this challenge through atten

The Expressivity Boundary of Probabilistic Circuits: A Comparison with Large Language Models

ResearchDGX agent

arXiv:2605.12940v1 Announce Type: cross Abstract: Probabilistic Circuits (PCs) are deep generative models that support exact and efficient probabilistic inference. Yet in autoregressive language model

The Score-Difference Flow for Implicit Generative Modeling

ResearchDGX agent

arXiv:2304.12906v4 Announce Type: replace Abstract: Implicit generative modeling (IGM) aims to produce samples of synthetic data matching the characteristics of a target data distribution. Recent work

VIP-COP: Context Optimization for Tabular Foundation Models

ResearchDGX agent

arXiv:2605.12904v1 Announce Type: new Abstract: Tabular foundation models (TFMs) have emerged as a powerful paradigm for in-context learning on structured data, enabling direct prediction on new tabul

What properties of reasoning supervision are associated with improved downstream model quality?

SafetyDGX agent

arXiv:2605.13290v1 Announce Type: new Abstract: Validating training data for reasoning models typically requires expensive trial-and-error fine-tuning cycles. In this work, we investigate whether the

13 May 2026

A Study on Hidden Layer Distillation for Large Language Model Pre-Training

Model ReleasesDGX agent

arXiv:2605.11513v1 Announce Type: new Abstract: Knowledge Distillation (KD) is a critical tool for training Large Language Models (LLMs), yet the majority of research focuses on approaches that rely s

Active inference as a unified model of collision avoidance behavior in human drivers

ResearchDGX agent

arXiv:2506.02215v5 Announce Type: replace Abstract: Collision avoidance -- involving a rapid threat detection and quick execution of the appropriate evasive maneuver -- is a critical aspect of driving

ASD-Bench: A Four-Axis Comprehensive Benchmark of AI Models for Autism Spectrum Disorder

Model ReleasesDGX agent

arXiv:2605.11091v1 Announce Type: new Abstract: Automated ASD screening tools remain limited by single-architecture evaluations, axis-restricted assessment, and near-exclusive focus on adult cohorts,

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models

Model ReleasesDGX agent

arXiv:2605.11726v1 Announce Type: new Abstract: Recently, reinforcement learning (RL) has been widely applied during post-training for diffusion large language models (dLLMs) to enhance reasoning with

Built a local coding harness powered by Gemma 4. It runs locally, connects to my model backend, starts coding sessions, streams responses, a…

Model ReleasesDGX agent

Built a local coding harness powered by Gemma 4. It runs locally, connects to my model backend, starts coding sessions, streams responses, and uses tools through a CLI-style workflow. Still early, but

Clarity: The Flexibility-Interpretability Trade-Off in Sparsity-aware Concept Bottleneck Models

SafetyDGX agent

arXiv:2601.21944v2 Announce Type: replace Abstract: The widespread adoption of deep learning models in computer vision has intensified concerns about interpretability. Despite strong performance, thes

DreamPolicy: A Unified World-model Policy for Scalable Humanoid Locomotion

SafetyDGX agent

arXiv:2505.18780v3 Announce Type: replace-cross Abstract: Achieving versatile humanoid locomotion with a single policy presents a critical scalability challenge. Prevailing methods often rely on disti

EHR-RAGp: Retrieval-Augmented Prototype-Guided Foundation Model for Electronic Health Records

SafetyDGX agent

arXiv:2605.12335v1 Announce Type: cross Abstract: Electronic Health Records (EHR) contain rich longitudinal patient information and are widely used in predictive modeling applications. However, effect

Express Your Doubts -- Probabilistic World Modeling Should not be Based on Token logprobs

ResearchDGX agent

arXiv:2505.02072v2 Announce Type: replace Abstract: Language modeling has shifted in recent years from a distribution over strings to prediction models with textual inputs and outputs for general-purp

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models

SafetyDGX agent

arXiv:2605.12374v1 Announce Type: new Abstract: Visual latent reasoning lets a multimodal large language model (MLLM) create intermediate visual evidence as continuous tokens, avoiding external tools

Fine-Tuning Large Language Models for Cooperative Tactical Deconfliction of Small Unmanned Aerial Systems

Model ReleasesDGX agent

arXiv:2603.28561v2 Announce Type: replace Abstract: The growing deployment of small Unmanned Aerial Systems (sUASs) in low-altitude airspaces has increased the need for reliable tactical deconfliction

Molecular Design beyond Training Data with Novel Extended Objective Functionals of Generative AI Models Driven by Quantum Annealing Computer

ResearchDGX agent

arXiv:2602.15451v3 Announce Type: replace-cross Abstract: Deep generative modeling to stochastically design small molecules is an emerging technology for accelerating drug discovery and development. H

PriorZero: Bridging Language Priors and World Models for Decision Making

SafetyDGX agent

arXiv:2605.12289v1 Announce Type: new Abstract: Leveraging the rich world knowledge of Large Language Models (LLMs) to enhance Reinforcement Learning (RL) agents offers a promising path toward general

@SakanaAILabs @NVIDIAAI Sparser, Faster, Lighter Transformer Language Models https://arxiv.org/abs/2603.23198

ResearchDGX agent

This research paper from Sakana AI and NVIDIA explores techniques for creating more efficient transformer language models by reducing sparsity, computational requirements, and model size while maintai

See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model

Model ReleasesDGX agent

arXiv:2605.11817v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have shown remarkable promise in robotics manipulation, yet their high computational cost hinders real-time deploy

Sources: Mistral has been developing a cybersecurity-focused AI model and held discussions about it with European banks, which don't have access to Mythos (Bloomberg)

Model ReleasesDGX agent

Bloomberg: Sources: Mistral has been developing a cybersecurity-focused AI model and held discussions about it with European banks, which don't have access to Mythos — French artificial intelligence s

STRIDE: Training-Free Diversity Guidance via PCA-Directed Feature Perturbation in Single-Step Diffusion Models

SafetyDGX agent

arXiv:2605.11494v1 Announce Type: new Abstract: Distilled one-step (T=1) or few-step (Tleq4) diffusion models enable real-time image generation but often exhibit reduced sample diversity compared to t

Structural Interpretations of Protein Language Model Representations via Differentiable Graph Partitioning

Local AiDGX agent

arXiv:2605.10985v1 Announce Type: new Abstract: Protein language models such as ESM-2 learn rich residue representations that achieve strong performance on protein function prediction, but their featu

Unlocking UML Class Diagram Understanding in Vision Language Models

Model ReleasesDGX agent

arXiv:2605.11634v1 Announce Type: new Abstract: Although Vision Language Models (VLMs) have seen tremendous progress across all kinds of use cases, they still fall behind in answering questions regard

Vision-aligned Latent Reasoning for Multi-modal Large Language Model

ResearchDGX agent

arXiv:2602.04476v2 Announce Type: replace Abstract: Despite recent advancements in Multi-modal Large Language Models (MLLMs) on diverse understanding tasks, these models struggle to solve problems whi

When Emotion Becomes Trigger: Emotion-style dynamic Backdoor Attack Parasitising Large Language Models

ResearchDGX agent

arXiv:2605.11612v1 Announce Type: new Abstract: Backdoor vulnerabilities widely exist in the fine-tuning of large language models(LLMs). Most backdoor poisoning methods operate mainly at the token lev

12 May 2026

A Game Theoretic Free Energy Analysis of Higher Order Synergy in Attention Heads of Large Language Models

Model ReleasesDGX agent

arXiv:2605.09515v1 Announce Type: new Abstract: Large language models rely on multihead attention, but interactions among heads remain poorly understood. We apply the Game Theoretic Free Energy Princi

← Previous
1…113114115116117…1009
Next →