AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,515 results
2 Jun 2026

3D Segment Anything Model with Visual Mamba for Diagnosing Placenta Accreta Spectrum

Model ReleasesDGX agent

arXiv:2606.00489v1 Announce Type: new Abstract: Placenta Accreta Spectrum (PAS) is a rare but highly dangerous obstetric disease. Early and accurate PAS diagnosis is critical for maternal health. Trad

A Developer’s Guide to Managing Models, Cost and Quality in Microsoft Foundry

TutorialsDGX agent

Learn a practical model lifecycle for Microsoft Foundry: select the right model, evaluate quality, optimize cost, operate safely, and improve as production needs change. The post A Developer’s Guide t

AgentPLM: Agentic Protein Language Models with Reasoning-Augmented Decoding for Protein Sequence Design

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.02386v1 Announce Type: new Abstract: Protein language models (PLMs) are passive oracles: they generate sequences in a single forward pass with no mechanism to consult external biophysical f

Are Large Reasoning Models Interruptible?

ApplicationsDGX agent

arXiv:2510.11713v4 Announce Type: replace Abstract: Real-world applications of Large Reasoning Models (LRMs) often require reasoning about changing prompts or environments. In this work, we challenge

au_0-WM: A Unified Video-Action World Model for Robotic Manipulation

SafetyDGX agent

arXiv:2606.01027v1 Announce Type: new Abstract: Robotic manipulation requires models that generate executable actions while anticipating and evaluating their future consequences before physical execut

AutoEval Done Right: Using Synthetic Data for Model Evaluation

Model ReleasesDGX agent

arXiv:2403.07008v3 Announce Type: replace-cross Abstract: The evaluation of machine learning models using human-labeled validation data can be expensive and time-consuming. AI-labeled synthetic data c

Benchmarking Large Language Models for Cryptanalysis and Side-Channel Vulnerabilities

Model ReleasesDGX agent

arXiv:2505.24621v3 Announce Type: replace Abstract: Recent advancements in large language models (LLMs) have transformed natural language understanding and generation, leading to extensive benchmarkin

Beyond Categories of Caste: Examining Caste Bias and Morality in Text-to-Image AI Models

SafetyDGX agent

arXiv:2606.00039v1 Announce Type: cross Abstract: Text-to-Image (T2I) models have shown promising utility across various domains. However, such models are also amplifying harmful societal biases in th

Beyond Text and Tables: Vision-Language Model Integration in ComProScanner for Extracting Materials Data from Scientific Figures with High Accuracy

Model ReleasesDGX agent

arXiv:2606.00065v1 Announce Type: cross Abstract: Automated extraction of materials composition-property data from scientific literature has advanced considerably with the development of large languag

Consistency Deep Equilibrium Models

ResearchDGX agent

arXiv:2602.03024v2 Announce Type: replace-cross Abstract: Deep Equilibrium Models (DEQs) have emerged as a powerful paradigm in deep learning, offering the ability to model infinite-depth networks wit

Controllable Value Alignment in Large Language Models through Neuron-Level Editing

Model ReleasesDGX agent

arXiv:2602.07356v2 Announce Type: replace Abstract: Aligning large language models (LLMs) with human values has become increasingly important as their influence on human behavior and decision-making e

DASH: Dual-Branch Score Distillation for Guidance-Calibrated Compact Diffusion Models

Model ReleasesDGX agent

arXiv:2606.00798v1 Announce Type: cross Abstract: Parameter compression of class-conditional diffusion models reveals an underexplored limitation in output-level distillation: the unconditional score

Diffusion Image Generation with Explicit Modeling of Data Manifold Geometry

SafetyDGX agent

arXiv:2606.00094v1 Announce Type: cross Abstract: Image generative models aim to sample data points from the underlying data manifold, a task that requires learning and decoding a dense, low-dimension

Distillation of Large Language Models via Concrete Score Matching

Model ReleasesDGX agent

arXiv:2509.25837v3 Announce Type: replace-cross Abstract: Large language models (LLMs) deliver remarkable performance but are costly to deploy, motivating knowledge distillation (KD) for efficient inf

DLLM-JEPA: Joint Embedding Predictive Architectures for Masked Diffusion Language Models

Model ReleasesDGX agent

arXiv:2606.00091v1 Announce Type: cross Abstract: Joint Embedding Predictive Architectures (JEPAs) have reshaped self-supervised representation learning in vision. The recent LLM-JEPA ported JEPA to a

ELF: A Family of Encoder-Free ECG-Language Models

ResearchDGX agent

arXiv:2601.18798v2 Announce Type: replace-cross Abstract: ECG-Language Models (ELMs) extend recent advances in Multimodal Large Language Models (MLLMs) to automated ECG interpretation. However, most e

Empathy Applicability Modeling for General Health Queries

Model ReleasesDGX agent

arXiv:2601.09696v2 Announce Type: replace Abstract: LLMs are increasingly being integrated into clinical workflows, yet they often lack clinical empathy, an essential aspect of effective doctor-patien

Escaping the Mode Lottery: Multi-Response Training Improves Language Model Generalization

Model ReleasesDGX agent

arXiv:2606.00544v1 Announce Type: cross Abstract: Modern language-model fine-tuning typically pairs each prompt with a single response, even though many prompts admit multiple valid completions. This

From 'Weak' Signals to Strong Models: Preference Delta Aggregation with LoRA Merging

SafetyDGX agent

arXiv:2606.00357v1 Announce Type: new Abstract: Training strong large language models (LLMs) requires high-quality supervision, which is often scarce. Recent work shows that paired preference data fro

GLENS: Global Search via Learning from Solver Iterates with Diffusion Models

Model ReleasesDGX agent

arXiv:2606.00366v1 Announce Type: new Abstract: We consider the problem of generating a large collection of initial guesses for local minima of multimodal non-convex continuous optimization problems.

GNMR: Runtime Stability Control for Low-Precision Large Language Model Training

Model ReleasesDGX agent

arXiv:2606.00539v1 Announce Type: new Abstract: Training stability is a key bottleneck in low-precision language model training: efficient low-cost paths can still produce short-lived numerical risks

How to add specific knowledge to an ollama model?

Local AiDGX agent

Adds knowledge to Ollama models using Retrieval-Augmented Generation (RAG) , where users create a knowledge base directory with reference files like PDFs, text files, or CSVs . A custom model can be c

IDLM: Inverse-distilled Diffusion Language Models

ResearchDGX agent

arXiv:2602.19066v2 Announce Type: replace-cross Abstract: Diffusion Language Models (DLMs) have recently achieved strong results in text generation. However, their multi-step sampling leads to slow in

InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning

Model ReleasesDGX agent

arXiv:2509.12263v3 Announce Type: replace Abstract: Large multimodal models (LMMs) encode physical laws observed during training, such as momentum conservation, as parametric knowledge. It allows LMMs

Interpretability in Deep Time Series Models Demands Semantic Alignment

SafetyDGX agent

arXiv:2602.02239v2 Announce Type: replace Abstract: Deep time series models continue to improve predictive performance, yet their deployment remains limited by their black-box nature. In response, exi

Measuring and Mitigating Bias in Code Generated by Large Language Models

Model ReleasesDGX agent

arXiv:2606.00049v1 Announce Type: cross Abstract: Large language models (LLMs) are widely recognised for their applications in natural language generation and are increasingly used for code generation

On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters

Model ReleasesDGX agent

arXiv:2606.02437v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning (PEFT) is usually treated as a cheaper alternative to full fine-tuning. We study a broader role: small trainable adapt

Optimizing Diversity and Quality through Base-Aligned Model Collaboration

SafetyDGX agent

arXiv:2511.05650v2 Announce Type: replace-cross Abstract: Alignment has greatly improved large language models (LLMs)' output quality at the cost of diversity, yielding highly similar outputs across g

Perception First: A Frontier Native-Video Model with Self-Consistency for Implicit Video Question Answering

Model ReleasesDGX agent

arXiv:2606.01485v1 Announce Type: new Abstract: We describe our submission to the VRR Challenge @ CVPR 2026, built on the ImplicitQA / VRR-QA benchmark~ite{implicitqa}: multiple-choice video question

Reconsidering Positional Supervision in Masked Diffusion Language Model Training

SafetyDGX agent

arXiv:2601.22947v2 Announce Type: replace Abstract: Masked diffusion language models (MDLMs) generate text by unmasking tokens in parallel and have recently emerged as alternatives to autoregressive l

Revisiting Parameter-Based Knowledge Editing in Large Language Models: Theoretical Limits and Empirical Evidence

Model ReleasesDGX agent

arXiv:2606.00570v1 Announce Type: cross Abstract: Parameter-based knowledge editing updates the internal knowledge of large language models (LLMs) via localized weight modifications and has attracted

Self-Regulating Annealing in Heavy-Tailed Diffusion Models

ResearchDGX agent

arXiv:2606.01645v1 Announce Type: cross Abstract: Diffusion models have emerged as a leading framework for deep generative modeling. While the standard Gaussian formulation is theoretically convenient

SentGuard: Sentence-Level Streaming Guardrails for Large Language Models

Model ReleasesDGX agent

arXiv:2606.02041v1 Announce Type: new Abstract: Large language models increasingly stream long, reasoning-intensive responses in real time, making when to moderate as critical as whether to moderate.

Silent Failures in Federated Personalization of Foundation Models

SafetyDGX agent

arXiv:2606.00947v1 Announce Type: cross Abstract: Foundation models are increasingly personalized on decentralized private data through federated learning and are now deployed at scale under growing r

SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model Reasoning

Model ReleasesDGX agent

arXiv:2603.08000v2 Announce Type: replace Abstract: Large reasoning models (LRMs) like OpenAI o1 and DeepSeek-R1 achieve high accuracy on complex tasks by adopting long chain-of-thought (CoT) reasonin

Step-3.7-Flash from @StepFun_ai is a silent winner. Super impressive results, the best model under 500B params on HF leaderboards.. All whil…

IndustryDGX agent

Step-3.7-Flash is a compact language model from StepFun AI that reportedly achieves top-tier performance among models under 500 billion parameters on Hugging Face leaderboards, despite receiving limit

SurrogateSHAP: Training-Free Contributor Attribution for Text-to-Image (T2I) Models

SafetyDGX agent

arXiv:2601.22276v2 Announce Type: replace-cross Abstract: As Text-to-Image (T2I) diffusion models are increasingly used in real-world creative workflows, a principled framework for valuing contributor

The Shape of Wisdom: Decision Trajectories in Language Models

Model ReleasesDGX agent

arXiv:2606.01202v1 Announce Type: new Abstract: Language models do not simply choose an answer at the output layer. In a 9,000-trajectory MMLU study across Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct,

Towards Lightweight Reliability: Using Soft Prompts for Hallucination Mitigation in Large Language Models

Model ReleasesDGX agent

arXiv:2606.00919v1 Announce Type: new Abstract: Large language models (LLMs) have seen widespread adoption across various domains, yet their reliability is frequently undermined by hallucinations - re

Trump signs executive order to review AI models before they’re released

Model ReleasesDGX agent

President Donald Trump signed an executive order Tuesday creating a 'voluntary framework' for AI companies to share their frontier models with the federal government before they're released 'to promot

Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model Outputs

ResearchDGX agent

arXiv:2506.14003v5 Announce Type: replace Abstract: Machine unlearning (MU) for large language models (LLMs), commonly referred to as LLM unlearning, seeks to remove specific undesirable data or knowl

Vision-language Models for Driver Monitoring Systems: A Driver Activity Description Dataset

Model ReleasesDGX agent

arXiv:2606.02273v1 Announce Type: new Abstract: Understanding subtle driver actions is essential for building reliable driver monitoring systems. Existing visionlanguage models (VLMs) are trained on g

We're sponsoring a hackathon to scale down. Hosted by our friends @huggingface and @Gradio, we want working with models to feel like yours a…

Model ReleasesDGX agent

We're sponsoring a hackathon to scale down. Hosted by our friends @huggingface and @Gradio, we want working with models to feel like yours again. Small enough that it's inexpensive to run, big enough

World Models: A Comprehensive Survey of Architectures, Methodologies, Reasoning Paradigms, and Applications

SafetyDGX agent

arXiv:2606.00133v1 Announce Type: new Abstract: World models, internal simulators that learn the structure and dynamics of an environment, have emerged as a central paradigm in the pursuit of artifici

1 Jun 2026

3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models

ResearchDGX agent

arXiv:2603.07751v2 Announce Type: replace-cross Abstract: Current Large Language Models have achieved Olympiad-level logic, yet Vision-Language Models paradoxically falter on elementary spatial tasks

Accelerate LLM model loading and increase context windows with GPUDirect on Amazon FSx for Lustre and TurboQuant

HardwareDGX agent

If you’re iterating on deploying large language models (LLMs) on AWS GPU instances, you’ve probably noticed the larger the model to be loaded into GPU High Bandwidth Memory (HBM), the longer the painf

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education

Model ReleasesDGX agent

arXiv:2605.31212v1 Announce Type: cross Abstract: AI systems are increasingly used to support educational content creation, yet it remains unclear whether they can generate outputs that faithfully rep

Bernini released. Unified Video generation and editing model. Built on Wan-2.2

Model ReleasesDGX agent

Bernini is a unified framework for video editing and video generation , built using Wan2.2-A14B as its renderer . The model covers complementary task families that demonstrate its capabilities as a un

Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies

Model ReleasesDGX agent

arXiv:2512.19673v3 Announce Type: replace-cross Abstract: Existing reinforcement learning (RL) approaches treat large language models (LLMs) as a unified policy, overlooking their internal mechanisms.

dgMARK: Decoding-Guided Watermarking for Diffusion Language Models

ResearchDGX agent

arXiv:2601.22985v2 Announce Type: replace Abstract: We propose dgMARK, a decoding-guided watermarking method for discrete diffusion language models (dLLMs). Unlike autoregressive models, dLLMs can gen

Differentially Private Preference Data Synthesis for Large Language Model Alignment

SafetyDGX agent

arXiv:2605.30808v1 Announce Type: cross Abstract: Preference alignment is a crucial post-training step for large language models (LLMs) to ensure their outputs align with human values. However, post-t

Domain Adaptation and Reasoning Frameworks in Language Models: A Controlled Experiment with Historical Cosmology

ResearchDGX agent

arXiv:2605.30415v1 Announce Type: cross Abstract: We investigate how domain adaptation reshapes explanatory behavior in language models using historical cosmology as a controlled setting. In Phase 1,

Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language Models

SafetyDGX agent

arXiv:2509.24319v4 Announce Type: replace-cross Abstract: Large language models can express values in two main ways: (1) intrinsic expression, reflecting the model's inherent values learned during tra

Elastic ViTs from Pretrained Models without Retraining

HardwareDGX agent

arXiv:2510.17700v2 Announce Type: replace Abstract: Vision foundation models achieve remarkable performance but are only available in a limited set of pre-determined sizes, forcing sub-optimal deploym

Enhancing Computer Vision Model Generalization in Warehouse Facilities: A Case Study on Anomaly Detection in Vertical Material Handling Systems

ApplicationsDGX agent

arXiv:2605.31487v1 Announce Type: new Abstract: Deploying computer vision models in Warehouse Facilities traditionally requires extensive resources for camera mounting, image collection, annotation, t

Fine-Tuning Improves Information Conveyance in Language Models

ResearchDGX agent

arXiv:2605.30844v1 Announce Type: cross Abstract: Fine-tuning is often believed to reduce uncertainty and diversity in large language models, but existing analyses overlook output length, a key confou

Graph Energy Matching: Transport-Aligned Energy-Based Modeling for Graph Generation

Local AiDGX agent

arXiv:2603.23398v2 Announce Type: replace-cross Abstract: Generative modeling of discrete data, such as graphs, underpins many scientific and industrial applications, including molecular discovery and

In May, we integrated 11 new models spanning image, 3D, audio, video, and multimodal. The highlights: → Krea 2 — style-first image generatio…

Model ReleasesDGX agent

In May, we integrated 11 new models spanning image, 3D, audio, video, and multimodal. The highlights: → Krea 2 — style-first image generation, live as a Partner Node on day one. Competes on how the fr

Language Models Can Resolve Reference Compositionally, But It's Not Their Native Strength: The Case of the Personal Relation Task

ResearchDGX agent

arXiv:2605.31480v1 Announce Type: new Abstract: Do neural models, such as Large Language Models, genuinely acquire compositional abilities for interpretation of natural language? When we talk about se

Many research labs only consider inference efficiency after the fact. Step 3.7 Flash is a 196B MoE model, and built for inference from the s…

Model ReleasesDGX agent

Many research labs only consider inference efficiency after the fact. Step 3.7 Flash is a 196B MoE model, and built for inference from the start by @StepFun_ai. Multi-Matrix Factorization Attention (M

← Previous
1…108109110111112…1009
Next →