AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,860 results
30 Jun 2026

Travel-Oriented Reasoning Large Language Model via Domain-Specific Knowledge Graphs

Model ReleasesDGX agent

arXiv:2606.29254v1 Announce Type: new Abstract: Large language models (LLMs) demonstrate broad reasoning abilities but struggle with accuracy and reliability in specialized domains such as travel, whe

We’re shipping two major updates to streamline your creative workflow, allowing you to generate high-speed images with one model and then in…

Model ReleasesDGX agent

We’re shipping two major updates to streamline your creative workflow, allowing you to generate high-speed images with one model and then instantly animate them with the other—all at a fraction of the

29 Jun 2026

Diffusion Model Attribution via Spectral Coupling of Denoiser Responses

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
ResearchDGX agent

arXiv:2606.28092v1 Announce Type: new Abstract: Attributing a generated image to its source diffusion model is a fundamental challenge in provenance verification and intellectual property protection.

From Tokens to States: LLMs as a Special Case of World Models and the Continuous Path Beyond

ResearchDGX agent

arXiv:2606.28127v1 Announce Type: cross Abstract: The AI community has framed the relationship between large language models (LLMs) and world models as a dichotomy: LLMs predict tokens; world models s

HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models

ResearchDGX agent

arXiv:2606.27627v1 Announce Type: cross Abstract: Discrete audio representations have become increasingly popular for building multimodal text-audio systems and integrating audio capabilities into Lar

Towards Benign Memory Forgetting for Selective Multimodal Large Language Model Unlearning

Model ReleasesDGX agent

arXiv:2511.20196v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) can inadvertently memorize privacy-sensitive information during training. While existing unlearning methods

Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety

Model ReleasesDGX agent

arXiv:2606.27632v1 Announce Type: new Abstract: As large language models are increasingly deployed in real-world systems, safety failures can still lead to harmful outputs and dangerous misuse. We arg

27 Jun 2026

The take that frontier models aren't ready for use in medicine is dead wrong. I've been using Opus 4.x in my clinical workflow every day for…

Model ReleasesDGX agent

The take that frontier models aren't ready for use in medicine is dead wrong. I've been using Opus 4.x in my clinical workflow every day for 6 months. I'm a deep sub-sub-specialist in dermatology - th

26 Jun 2026

Do Image Editing Models Understand Lighting?

Model ReleasesDGX agent

arXiv:2606.26738v1 Announce Type: new Abstract: While recent advancements in generative image editing models have achieved stunning visual fidelity, it remains an open question whether these systems p

Unison: Benchmarking Unified Multimodal Models via Synergistic Understanding and Generation

Model ReleasesDGX agent

arXiv:2606.26984v1 Announce Type: new Abstract: Unified multimodal models capable of both understanding and generation have achieved remarkable strides. However, despite their unified designs, existin

25 Jun 2026

Surrogate models for Rock-Fluid Interaction: A Grid-Size-Invariant Approach

ResearchDGX agent

arXiv:2602.22188v2 Announce Type: replace Abstract: Modelling rock-fluid interaction requires solving a set of partial differential equations (PDEs) to predict the flow behaviour and the reactions of

Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Models and GUI Grounding Datasets

Model ReleasesDGX agent

arXiv:2606.25760v1 Announce Type: cross Abstract: Computer-use agents turn vision-language model (VLM) predictions into executable GUI clicks, so reliable uncertainty estimates are essential for rejec

24 Jun 2026

BenchX: Benchmarking AI Models for Cancer Detection and Localization with Demographic and Protocol Biases

Model ReleasesDGX agent

arXiv:2606.24883v1 Announce Type: new Abstract: Artificial intelligence (AI) has achieved remarkable success in medical imaging, but it is widely recognized that these models often perform inconsisten

DriveStack-VLA: Render-Teacher Alignment for BEV-Based DeepStack Vision-Language-Action Model

SafetyDGX agent

arXiv:2606.24051v1 Announce Type: new Abstract: Vision-Language-Action driving models convert a pretrained Vision-Language Model into a driving policy, allowing them to use world knowledge and follow

Federated Survival Analysis in Healthcare: A Multi-Model Evaluation on Cross-Institutional Heterogeneous Breast Cancer Data

Local AiDGX agent

arXiv:2606.23871v1 Announce Type: new Abstract: Survival analysis is central to clinical decision-making, yet reliable time-to-event models require large, diverse cohorts that are rarely available at

Lightweight Transformer Models for On-Device Fault Detection: A Benchmark Study on Resource-Constrained Deployment

Model ReleasesDGX agent

arXiv:2606.24173v1 Announce Type: cross Abstract: On-device fault detection enables real-time diagnostics without cloud dependency, but deploying machine learning models on resource-constrained hardwa

LoMime: Query-Efficient Membership Inference using Model Extraction in Label-Only Settings

Model ReleasesDGX agent

arXiv:2602.18934v2 Announce Type: replace Abstract: Membership inference attacks (MIAs) threaten the privacy of machine learning models by revealing whether a specific data point was used during train

Model selection with proper scoring rules on data sets of time series

ResearchDGX agent

arXiv:2606.24715v1 Announce Type: cross Abstract: We consider the problem of model selection between probabilistic models on data sets of time series. Chosen a proper scoring rule, we denote by the te

23 Jun 2026

Assistron: Bayesian Shared Autonomy with Off-the-shelf Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2606.23147v1 Announce Type: new Abstract: We propose Assistron, a shared autonomy model that leverages Vision-Language-Action (VLA) models to assist the user in daily activities. Our approach is

Bayesian Adaptation Gym: A Benchmark for the Bayesian Low-Rank Adaptation of Multi-Modal Language Models

Model ReleasesDGX agent

arXiv:2606.22188v1 Announce Type: new Abstract: Large multi-modal language models are increasingly deployed in high-stakes domains, making well-calibrated uncertainty essential. Traditional Bayesian m

CAT-Translate: Building Compact Open-Source Models for Japanese-English Translation

ApplicationsDGX agent

arXiv:2606.21413v1 Announce Type: cross Abstract: Nowadays, large multilingual translation models demonstrate impressive translation capabilities in the machine translation benchmarks. This raises a p

Frequency-Domain Neural ODEs for Modeling Non-Linear Dynamical Systems

ResearchDGX agent

arXiv:2606.22075v1 Announce Type: new Abstract: Standard continuous-depth models, such as Neural Ordinary Differential Equations (NODEs), offer significant advantages in modeling physical systems by l

I guarantee you are sleeping on small models. Deepseek V4 Flash can do ~80% of the tasks you ask Claude or Codex for. It is 137x cheaper per…

Model ReleasesDGX agent

DeepSeek V4 Flash is a small language model capable of handling approximately 80% of common tasks performed by larger models like Claude or Codex, while offering significantly lower costs (137x cheape

IndicGuard: A Multilingual Safety Guard Model and Dataset for Indic Languages

Model ReleasesDGX agent

arXiv:2606.22841v1 Announce Type: cross Abstract: As Large Language Models (LLMs) achieve widespread integration across diverse linguistic landscapes, ensuring their safety and alignment with regional

Model council as a tool to push performance definitely works, I think the interesting next frontier is trying to scale this at the agent lev…

AgentsDGX agent

Model council as a tool to push performance definitely works, I think the interesting next frontier is trying to scale this at the agent level have been thinking a bunch about model routing and relate

MOOZY: A Patient-First Foundation Model for Computational Pathology

Model ReleasesDGX agent

arXiv:2603.27048v3 Announce Type: replace Abstract: Computational pathology needs whole-slide image (WSI) foundation models that transfer across diverse clinical tasks, yet current approaches remain l

One-Step Flow Matching for Generative Modeling of Path-Dependent Physical Fields

ResearchDGX agent

arXiv:2606.22752v1 Announce Type: new Abstract: Physical simulations for intricate geometries with path-dependent constitutive models face difficulties due to the enormous computational cost they requ

Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently

TutorialsDGX agent

arXiv:2606.22938v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have demonstrated that reinforcement fine-tuning of pretrained base models can lead to significant gains

Training-free Task Classification for Multi-Task Model Merging

Model ReleasesDGX agent

arXiv:2606.22589v1 Announce Type: new Abstract: Ever since the advent of foundation models and the pre-training-finetuning paradigm, there have been numerous efforts to merge multiple task-specific ex

19 Jun 2026

it is indeed quite good! don't try it in claude code/codex - those harnesses are overly tuned for their proprietary models dcode (deepagents…

Model ReleasesDGX agent

it is indeed quite good! don't try it in claude code/codex - those harnesses are overly tuned for their proprietary models dcode (deepagents code) is a model agnostic harness - try it there with @Fire

11 Jun 2026

A theory of learning data statistics in diffusion models, from easy to hard

SafetyDGX agent

arXiv:2603.12901v2 Announce Type: replace-cross Abstract: While diffusion models have emerged as a powerful class of generative models, their learning dynamics remain poorly understood. We address thi

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models

Model ReleasesDGX agent

arXiv:2606.11211v1 Announce Type: cross Abstract: The ability of large language models (LLMs) to express calibrated uncertainty is important for safe deployment. Chain-of-thought (CoT) reasoning is wi

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

Model ReleasesDGX agent

arXiv:2606.11324v1 Announce Type: cross Abstract: We introduce Embodied-R1.5, a unified Embodied Foundation Model (EFM) that integrates comprehensive embodied reasoning capabilities, spanning embodied

Making Foresight Actionable: Repurposing Representation Alignment in World Action Models

Local AiDGX agent

arXiv:2606.12217v1 Announce Type: cross Abstract: World Action Models (WAMs) offer a promising route for robot manipulation by using video generation models to model future scene evolution before prod

Physically Constrained Ensemble Gaussian Process Modelling for Expensive Quantum Systems with Heteroskedastic Noise

Model ReleasesDGX agent

arXiv:2606.11240v1 Announce Type: cross Abstract: Accurate modeling of quantum many-body systems often requires computationally expensive simulations such as Density Matrix Renormalization Group (DMRG

To Intervene or Not: Guiding Inference-time Alignment with Probabilistic Model Blending

SafetyDGX agent

arXiv:2606.11201v1 Announce Type: cross Abstract: The wide deployment of LLMs has made model alignment necessary to make newly trained models safely and effectively respond to user instructions. Among

10 Jun 2026

CableRobotGraphSim: A Graph Neural Network for Modeling Partially Observable Cable-Driven Robot Dynamics

Model ReleasesDGX agent

arXiv:2602.21331v2 Announce Type: replace Abstract: General-purpose simulators have accelerated the development of robots. Traditional simulators based on first-principles, however, typically require

Data assimilation for subsurface flow using latent diffusion model parameterization: performance of ensemble-Kalman and Monte Carlo techniques

ResearchDGX agent

arXiv:2606.11140v1 Announce Type: cross Abstract: Data assimilation (DA) in subsurface flow entails calibrating model parameters to match observed data, typically at wells, while preserving geological

Do Vision-Language Models See or Guess? Measuring and Reducing Textual-Prior Reliance with a Phrasing-Controlled Benchmark

Model ReleasesDGX agent

arXiv:2606.10400v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly deployed where answers must follow from what is in the image, yet they often answer from textual priors,

FailureScope: Cross-Regime Behavioral Diagnosis of Language Model Weaknesses

Model ReleasesDGX agent

arXiv:2606.09878v1 Announce Type: new Abstract: Standard benchmarks report aggregate accuracy, but practitioners need to know which specific capabilities a model lacks. We introduce FailureScope, a be

Rotate2Think: Geometric Priming via Orthogonal Rotation to Improve Language Model Reasoning

Model ReleasesDGX agent

arXiv:2606.09873v1 Announce Type: cross Abstract: Reasoning models achieve strong performance on challenging tasks by generating explicit intermediate reasoning traces before producing a final answer.

What Really Matters for Table LLMs? A Meta-Evaluation of Model and Data Effects

ResearchDGX agent

arXiv:2501.14717v2 Announce Type: replace Abstract: Table modeling has progressed for decades. In this work, we revisit this trajectory and highlight emerging challenges in the LLM era, particularly t

9 Jun 2026

A Dataset for Dynamic Human Preferences for Vision Language Models

Model ReleasesDGX agent

arXiv:2606.07653v1 Announce Type: cross Abstract: Given the increased adoption of Vision Language Models (VLMs) in human-interactive settings, it is important that we evaluate how well these models ca

A retrieval conditioned rebinding circuit for dynamic entity tracking in large language models

Model ReleasesDGX agent

arXiv:2606.08644v1 Announce Type: cross Abstract: To interpret context correctly and retrieve relevant information, large language models must bind entities to their attributes and update these bindin

Anthropic releases its first Mythos-class model Claude Fable

Model ReleasesDGX agent

Anthropic just announced Claude Fable 5, a new AI model it said is the most powerful model it has ever made widely available. According to the company, Fable 5 'shows exceptional performance in softwa

Evaluating Hallucinations in Domain-Adapted Large Language Models

Model ReleasesDGX agent

arXiv:2606.07521v1 Announce Type: cross Abstract: This study investigates the phenomenon of hallucinations in domain-adapted Large Language Models (LLMs), focusing on the fine-tuning of the Llama-2 mo

Fable 5 is the biggest step up I’ve felt in our models since Opus 4.5 back in November. After 4.5 came out I uninstalled my IDE when I reali…

Model ReleasesDGX agent

Fable 5 is the biggest step up I’ve felt in our models since Opus 4.5 back in November. After 4.5 came out I uninstalled my IDE when I realized that I’d been doing 100% of my coding in a terminal for

Finite Certificates for In-Context Determinacy and a Threshold Theory of Emergence in Language Models

Model ReleasesDGX agent

arXiv:2606.07623v1 Announce Type: new Abstract: This paper develops a model-theoretic framework for verifying context-conditioned language-model behavior by replacing benchmark labels with finite sema

MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science

Model ReleasesDGX agent

arXiv:2510.12171v2 Announce Type: replace Abstract: Large Language Models have shown strong scientific reasoning ability, but their performance on materials science problems remains less studied. To f

Principled Agent Debate: Adversarial Arbitration for Sycophancy Reduction in Large Language Models

Model ReleasesDGX agent

arXiv:2606.07532v1 Announce Type: cross Abstract: RLHF-trained models are systematically biased toward agreement over accuracy, a structural property of the training process. We present Principled Age

Stress-testing medical large language models reveals latent safety pathology beyond benchmark accuracy

Model ReleasesDGX agent

arXiv:2606.07929v1 Announce Type: new Abstract: Large language models (LLMs) are entering clinical practice based on benchmark accuracy that may fail to detect safety-relevant failure modes. Here we p

Systematic LLM Translation of Legacy Scientific Code to Differentiable Frameworks: Application to a Land Surface Model

Model ReleasesDGX agent

arXiv:2606.07681v1 Announce Type: cross Abstract: Differentiable programming offers transformative capabilities for scientific modeling, enabling gradient-based parameter estimation, sensitivity analy

8 Jun 2026

Audio-Visual World Models: Grounding Multisensory Imagination for Embodied Agents

Model ReleasesDGX agent

arXiv:2512.00883v3 Announce Type: replace-cross Abstract: World models simulate environmental dynamics to enable agents to plan and reason about future states. While existing approaches have primarily

Endogenous Resistance to Activation Steering in Language Models

Model ReleasesDGX agent

arXiv:2602.06941v2 Announce Type: replace-cross Abstract: Large language models can recover mid-generation from task-misaligned activation steering, producing explicit verbal restarts (e.g., ``wait, t

Evidence Graph Consistency in Retrieval-Augmented Generation: A Model-Dependent Analysis of Hallucination Detection

Model ReleasesDGX agent

arXiv:2606.06748v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) reduces but does not eliminate hallucination in large language models. Existing detection methods rely on flat si

Explain Like I'm 5 or Whatever I Choose: Evaluating the Interactive Potential of Language Model Responses

Model ReleasesDGX agent

arXiv:2606.06788v1 Announce Type: new Abstract: Evaluations of large language models (LLMs) in scientific information seeking tasks have become increasingly use-centric, such as conducting live or mul

Last time around Apple released a lot of information about how their AI version of Siri worked between local and cloud models, not so much t…

Model ReleasesDGX agent

Last time around Apple released a lot of information about how their AI version of Siri worked between local and cloud models, not so much this time It is nice to have a Gemma-like model on device, bu

Modular Monolingual Adaptation using Pretrained Language Models

ResearchDGX agent

arXiv:2606.06738v1 Announce Type: new Abstract: Building monolingual language models (LMs) for low-resource languages typically relies on adapting pretrained language models (PLMs) by finetuning the w

PSA: Just added a few thousand chips, including B200s and B300s to our Dedicated Model Inference (http://api.together.ai/endpoints). With De…

Model ReleasesDGX agent

PSA: Just added a few thousand chips, including B200s and B300s to our Dedicated Model Inference (http://api.together.ai/endpoints). With Dedicated Model Inference, you can now on-click deploy our Bla

This is super big I think this is the first useful speculative decoding method deployed on a big quasi frontier model Massive unlock @fi5662…

Model ReleasesDGX agent

This is super big I think this is the first useful speculative decoding method deployed on a big quasi frontier model Massive unlock @fi56622380 🚀 1,000+ TOKENS/S ON A 1T MODEL! 🚀 We are thrilled to r

← Previous
1…3738394041…998
Next →