AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

85,136Total entries
1Added by human
85,135Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
61,002 results
12 May 2026

Nous Portal is one easy subscription that gives you access to 300+ models, exclusive discounts, and bundles your tokens and paid tools toget…

ResearchDGX agent

Nous Portal is one easy subscription that gives you access to 300+ models, exclusive discounts, and bundles your tokens and paid tools together for hassle-free setup and simple billing. http://portal.

On the global convergence of gradient descent for wide shallow models with bounded nonlinearities

ResearchDGX agent

arXiv:2605.10775v1 Announce Type: cross Abstract: A surprising phenomenon in the training of neural networks is the ability of gradient descent to find global minimizers of the training loss despite i

Path-Dependent Denoising: A Non-Conservative Field Perspective on Order Collapse in Diffusion Language Models

Local AiDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.09303v1 Announce Type: new Abstract: Diffusion language models (DLMs) offer a structural alternative to autoregressive generation: denoising can update tokens in arbitrary orders or in para

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping

Local AiDGX agent

arXiv:2605.10937v1 Announce Type: new Abstract: Recently, post-training methods based on reinforcement learning, with a particular focus on Group Relative Policy Optimization (GRPO), have emerged as t

PriorVLA: Prior-Preserving Adaptation for Vision-Language-Action Models

TutorialsDGX agent

arXiv:2605.10925v1 Announce Type: new Abstract: Large-scale pretraining has made Vision-Language-Action (VLA) models promising foundations for generalist robot manipulation, yet adapting them to downs

Pseudo-Deliberation in Language Models: When Reasoning Fails to Align Values and Actions

SafetyDGX agent

arXiv:2605.09893v1 Announce Type: cross Abstract: Large language models (LLMs) are often evaluated based on their stated values, yet these do not reliably translate into their actions, a discrepancy t

Relative Score Policy Optimization for Diffusion Language Models

SafetyDGX agent

arXiv:2605.10218v1 Announce Type: new Abstract: Diffusion large language models (dLLMs) offer a promising route to parallel and efficient text generation, but improving their reasoning ability require

RePO-VLA: Recovery-Driven Policy Optimization for Vision-Language-Action Models

SafetyDGX agent

arXiv:2605.09410v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models remain brittle in long-horizon, contact-rich manipulation because success-only imitation provides little supervisi

Revis: Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language Models

ResearchDGX agent

arXiv:2602.11824v2 Announce Type: replace Abstract: Despite the advanced capabilities of Large Vision-Language Models (LVLMs), they frequently suffer from object hallucination. One reason is that visu

SynerDiff: Synergetic Continuous Batching for Fast and Parallel Diffusion Model Inference

ResearchDGX agent

arXiv:2605.08835v1 Announce Type: new Abstract: The expansion of Artificial Intelligence-generated content service requires diffusion model serving to simultaneously achieve high throughput and low ta

The Astonishing Ability of Large Language Models to Parse Jabberwockified Language

ResearchDGX agent

arXiv:2602.23928v2 Announce Type: replace Abstract: We show that large language models (LLMs) have an astonishing ability to recover meaning from severely degraded English texts. Texts in which conten

The scale of the infra on HF is insane. If you're still hosting models, datasets, agent memory,... in S3 or R2, talk to use and we can help …

AgentsDGX agent

Hugging Face offers substantial infrastructure capabilities for hosting machine learning models, datasets, and agent memory systems. The statement suggests that organizations currently using alternati

The US' Centers for Medicare & Medicaid Services is testing ACCESS, an outcome-based payment model for AI-driven medical care, with 150 tech companies (Connie Loizos/TechCrunch)

ApplicationsDGX agent

Connie Loizos / TechCrunch: The US' Centers for Medicare & Medicaid Services is testing ACCESS, an outcome-based payment model for AI-driven medical care, with 150 tech companies — Neil Batlivala has

tl;dr - let's not just remove negative behavior from models, but also add positive ones 👍

SafetyDGX agent

tl;dr - let's not just remove negative behavior from models, but also add positive ones 👍 If anyone builds it, everyone thrives. Over the past decade, a lot of important work on AI alignment has focus

Towards Understanding Continual Factual Knowledge Acquisition of Language Models: From Theory to Algorithm

TutorialsDGX agent

arXiv:2605.10640v1 Announce Type: cross Abstract: Continual Pre-Training (CPT) is essential for enabling Language Models (LMs) to integrate new knowledge without erasing old. While classical CPT techn

TrajDLM: Topology-Aware Block Diffusion Language Model for Trajectory Generation

ApplicationsDGX agent

arXiv:2605.10020v1 Announce Type: new Abstract: Generating high-fidelity synthetic GPS trajectories is increasingly important for applications in transportation, urban planning, and what-if scenario s

We've just hit 1M open datasets on the Hugging Face Hub 🎉 Open models need open data. Today we hit that milestone, together with the most i…

IndustryDGX agent

We've just hit 1M open datasets on the Hugging Face Hub 🎉 Open models need open data. Today we hit that milestone, together with the most incredible community in AI! 🤗 Onwards to the next million 🚀 Me

Where Reliability Lives in Vision-Language Models: A Mechanistic Study of Attention, Hidden States, and Causal Circuits

ResearchDGX agent

arXiv:2605.08200v1 Announce Type: new Abstract: A pervasive intuition holds that vision-language models (VLMs) are most trustworthy when their attention maps look sharp: concentrated attention on the

World Models: 10 Things That Matter in AI Right Now

TutorialsDGX agent

World models recently made our list of 10 Things That Matter in AI Right Now. Watch executive editor Niall Firth explain why this emerging area of AI is gaining so much attention. Join MIT Technology

11 May 2026

A Behavioral Framework for Data-Driven Modeling of Nonlinear Systems in Vector-Valued Reproducing Kernel Hilbert Spaces

ResearchDGX agent

arXiv:2605.07052v1 Announce Type: cross Abstract: We generalize Jan Willems' behavioral approach to a class of discrete-time nonlinear systems in a vector-valued reproducing kernel Hilbert space (RKHS

A Rod Flow Model for Adam at the Edge of Stability

ResearchDGX agent

arXiv:2605.06821v1 Announce Type: cross Abstract: Cohen et al. (arXiv:2207.14484) observed that adaptive gradient methods such as Adam operate at the edge of stability. While there has been significan

AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models

ApplicationsDGX agent

arXiv:2605.07308v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have significantly advanced the capabilities of robotic agents in executing diverse tasks; however, they still face

Better Protein Function Prediction by Modeling Survivorship Bias

SafetyDGX agent

arXiv:2605.06879v1 Announce Type: new Abstract: Protein sequence data from nature exhibits survivorship bias: we only observe data from those organisms that survive and reproduce, while non-functional

Bifurcation Models: Learning Set-Valued Solution Maps with Weight-Tied Dynamics

TutorialsDGX agent

arXiv:2605.07277v1 Announce Type: cross Abstract: Many scientific and combinatorial problems admit multiple correct solutions, not a single label. Standard supervised learning resolves this ambiguity

Black-box model classification under the discriminative factorization

ResearchDGX agent

arXiv:2605.07878v1 Announce Type: new Abstract: Access to modern generative systems is often restricted to querying an API (the ``black-box' setting) and many properties of the system are unknown to t

DIMoE-Adapters: Dynamic Expert Evolution for Continual Learning in Vision-Language Models

ResearchDGX agent

arXiv:2605.07494v1 Announce Type: new Abstract: Continual learning enables vision-language models to accumulate knowledge and adapt to evolving tasks without retraining from scratch. However, in multi

Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport

ResearchDGX agent

arXiv:2605.06785v1 Announce Type: cross Abstract: Inference-time scaling methods rely on Process Reward Models (PRMs), which are often poorly calibrated and overestimate success probabilities. We prop

EggHand: A Multimodal Foundation Model for Egocentric Hand Pose Forecasting

ResearchDGX agent

arXiv:2605.07642v1 Announce Type: new Abstract: Forecasting future 3D hand pose sequences from egocentric video is essential for understanding human intention and enabling embodied applications such a

Emergent Symbolic Structure in Health Foundation Models: Extraction, Alignment, and Cross-Modal Transfer

SafetyDGX agent

arXiv:2605.07407v1 Announce Type: new Abstract: Health foundation models (FMs) learn useful representations from wearable sensors, but interpreting what they encode and transferring that knowledge acr

How Do Language Models Compose Functions?

ResearchDGX agent

arXiv:2510.01685v2 Announce Type: replace-cross Abstract: While large language models (LLMs) appear to be increasingly capable of solving compositional tasks, it is an open question whether they do so

🆕 Hugging Face 🤝 Hermes Agent 🔥 > we added Hermes Agent to local apps: run it locally with any compatible GGUF/MLX model > shipped native…

Local AiDGX agent

🆕 Hugging Face 🤝 Hermes Agent 🔥 > we added Hermes Agent to local apps: run it locally with any compatible GGUF/MLX model > shipped native traces support for Hermes Agent: visualize your Hermes traces

I have a new job! Excited to announce that I will be working with Hugging Face to make local models work great in OpenClaw and other open ag…

AgentsDGX agent

I have a new job! Excited to announce that I will be working with Hugging Face to make local models work great in OpenClaw and other open agent harnesses! I will be building in public and documenting

ImplantMamba: Long-range Sequential Modeling Mamba For Dental Implant Position Prediction

ResearchDGX agent

arXiv:2605.07082v1 Announce Type: new Abstract: In the design of surgical guides for implant placement, determining the precise implant position is a critical step. However, the implant region itself

Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models

ResearchDGX agent

arXiv:2602.01166v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models benefit from chain-of-thought (CoT) reasoning, but existing approaches incur high inference overhead and rely on

Learned Lagrangian Models of PDEs via Euler-Lagrange Residual Minimization

Local AiDGX agent

arXiv:2605.07157v1 Announce Type: new Abstract: We present the first method to directly use a learned continuous Lagrangian to forecast the dynamics of systems governed by partial differential equatio

Linear Response Estimators for Singular Statistical Models

ResearchDGX agent

arXiv:2605.07970v1 Announce Type: cross Abstract: We define susceptibilities as a measure of the response of an observable quantity of a parameterized statistical model to a perturbation of the data f

MAST: A Multi-fidelity Augmented Surrogate model via Spatial Trust-weighting

ResearchDGX agent

arXiv:2602.20974v2 Announce Type: replace Abstract: In engineering design and scientific computing, computational cost and predictive accuracy are intrinsically coupled. High-fidelity simulations prov

Object Hallucination-Free Reinforcement Unlearning for Vision-Language Models

SafetyDGX agent

arXiv:2605.08031v1 Announce Type: new Abstract: Vision-language models (VLMs) raise growing concerns about privacy, copyright, and bias, motivating machine unlearning to remove sensitive knowledge. Ho

Pre-trained Tabular Foundation Models as Versatile Summary Networks for Neural Posterior Estimation

ResearchDGX agent

arXiv:2605.07765v1 Announce Type: new Abstract: In this work, we study TabPFN as a training-free, modular summary network for simulation-based Bayesian inference (SBI). Tabular foundation models such

Pretraining a Foundation Model for Small-Molecule Natural Products

ResearchDGX agent

arXiv:2503.17656v4 Announce Type: replace-cross Abstract: Natural products, as metabolites from microorganisms, animals, or plants, exhibit diverse biological activities, making them crucial for drug

Rethinking State Tracking in Recurrent Models Through Error Control Dynamics

TutorialsDGX agent

arXiv:2605.07755v1 Announce Type: cross Abstract: The theory of state tracking in recurrent architectures has predominantly focused on expressive capacity: whether a fixed architecture can theoretical

Saliency-Aware Regularized Quantization Calibration for Large Language Models

ResearchDGX agent

arXiv:2605.05693v2 Announce Type: replace Abstract: Post-training quantization (PTQ) is an effective approach for deploying large language models (LLMs) under memory and latency constraints. Most exis

speaking of things that have gotten over a threshold for me, the combo of the new ChatGPT model, personality, and personalization feels like…

IndustryDGX agent

Sam Altman comments on OpenAI's new ChatGPT model, noting that the combination of improved capabilities, personality features, and personalization options has crossed an important threshold of functio

StreamPhy: Streaming Inference of High-Dimensional Physical Dynamics via State Space Models

ResearchDGX agent

arXiv:2605.07384v1 Announce Type: new Abstract: Inferring the evolution of high-dimensional and multi-modal (e.g., spatio-temporal) physical fields from irregular sparse measurements in real time is a

The internet gave language models their data for free. Robots don’t have that. Every trajectory has to be earned through hardware, time, tel…

AgentsDGX agent

The internet gave language models their data for free. Robots don’t have that. Every trajectory has to be earned through hardware, time, teleoperators, and real consequences. Shrey’s piece on simulati

The next generation of models won't just generate images - they'll understand worlds, motion, interaction, and action. We've been building t…

ToolsDGX agent

The next generation of models won't just generate images - they'll understand worlds, motion, interaction, and action. We've been building toward this for a while. Visual intelligence is becoming real

Three-in-One World Model: Energy-Based Consistency, Prediction, and Counterfactual Inference for Marketing Intervention

ResearchDGX agent

arXiv:2605.07199v1 Announce Type: new Abstract: Marketing decisions reflect the interaction of latent consumer heterogeneity, time-varying internal states, and explicit interventions, a structure that

Toward Better Geometric Representations for Molecule Generative Models

SafetyDGX agent

arXiv:2605.07693v1 Announce Type: new Abstract: Geometric representation-conditioned molecule generation provides an effective paradigm that decouples molecule representation modeling from structure g

TTF: Temporal Token Fusion for Efficient Video-Language Model

ResearchDGX agent

arXiv:2605.07355v1 Announce Type: cross Abstract: Video-language models (VLMs) face rapid inference costs as visual token counts scale with video length. For example, 32 frames at 448{imes}448 resolut

When Diffusion Model Can Ignore Dimension: An Entropy-Based Theory

ResearchDGX agent

arXiv:2605.07969v1 Announce Type: new Abstract: Diffusion models perform remarkably well on high-dimensional data such as images, often using only a modest number of reverse-time steps. Despite this p

With the model's simultaneous speech capability, Horace has gotten a lot easier to work with recently.

TutorialsDGX agent

The post discusses improvements in working with Horace (likely a tool or system) due to recent implementation of simultaneous speech capability in its underlying model. This advancement has made the i

10 May 2026

If the AI models are so smart, why do I feel like I’m losing a few neurons every time I read a longer form content written by AI? We’ve come…

TutorialsDGX agent

If the AI models are so smart, why do I feel like I’m losing a few neurons every time I read a longer form content written by AI? We’ve come a long way but we still have long way to go. In terms of cl

9 May 2026

Anthropic, OpenAI, and other AI firms met with Hindu, Sikh, and Greek Orthodox leaders to draft principles on how to infuse models with ethics and morality (Krysta Fauria/Associated Press)

TutorialsDGX agent

Krysta Fauria / Associated Press: Anthropic, OpenAI, and other AI firms met with Hindu, Sikh, and Greek Orthodox leaders to draft principles on how to infuse models with ethics and morality — As conce

8 May 2026

Akamai says it struck a seven-year cloud computing deal with a 'leading frontier model provider'; sources: the deal was with Anthropic and is worth $1.8B (Rachel Metz/Bloomberg)

IndustryDGX agent

Rachel Metz / Bloomberg: Akamai says it struck a seven-year cloud computing deal with a “leading frontier model provider”; sources: the deal was with Anthropic and is worth 1.8B — Anthropic PBC has si

New research: long-running agents often fail by stopping too early, not because the model can't make progress. We tested 5 harness designs a…

IndustryDGX agent

New research: long-running agents often fail by stopping too early, not because the model can't make progress. We tested 5 harness designs across 8 long-horizon coding tasks. Our new orchestration har

Singling out this part: there is a deep bet in Chinese industry that the model layer is the full stack worth mastering.

IndustryDGX agent

Singling out this part: there is a deep bet in Chinese industry that the model layer is the full stack worth mastering. Visiting most of the leading Chinese AI labs, I'm struck by a culture that's ext

Sources: WH is preparing to order US agencies to partner with AI companies on cybersecurity; the EO wouldn't require pre-release model testing by the government (Bloomberg)

IndustryDGX agent

Bloomberg: Sources: WH is preparing to order US agencies to partner with AI companies on cybersecurity; the EO wouldn't require pre-release model testing by the government — The Trump administration i

The gap between 'this looks cool' and 'I'm actually running it' used to be a day or two of setup. Not anymore. Deploy any Hugging Face model…

ApplicationsDGX agent

The gap between 'this looks cool' and 'I'm actually running it' used to be a day or two of setup. Not anymore. Deploy any Hugging Face model on the AI Native Cloud in a single session 👇🏼 https://www.t

The UFOs are on HF thanks to @MTSlive! Who’s going to train the first computer vision model? https://huggingface.co/MTSlive/datasets

IndustryDGX agent

MTSlive has uploaded a UFO dataset to Hugging Face, making it available for machine learning projects. The post is calling for developers and researchers to create the first computer vision model trai

We are less safe as a society by keeping Mythos (or any other smart model) tightly gated so only a few companies get it. Protecting 100 comp…

ResearchDGX agent

We are less safe as a society by keeping Mythos (or any other smart model) tightly gated so only a few companies get it. Protecting 100 companies is not enough. There are 96 million open source projec

← Previous
1…219220221222223…1017
Next →