AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,515 results
1 Jun 2026

Model Monotonicity in Autobidding Auctions: When Do Better Predictions Lead to Better Outcomes?

ResearchDGX agent

arXiv:2605.31036v1 Announce Type: cross Abstract: Online advertising platforms rely on machine learning models to predict click-through rates (pCTR) and conversion rates (pCVR) for auction mechanisms.

MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language Models

Model ReleasesDGX agent

arXiv:2511.16940v3 Announce Type: replace Abstract: Modern Vision-Language Models (VLMs) pose significant individual-level privacy risks by linking fragmented multimodal data to identifiable individua

OpenAI frontier models and Codex are now generally available on AWS, giving enterprises a new way to build on Amazon Bedrock with OpenAI thr…

Model Releases
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

OpenAI frontier models and Codex are now generally available on AWS, giving enterprises a new way to build on Amazon Bedrock with OpenAI through the security, compliance, and governance workflows they

Performance and Complexity Trade-off Optimization of Speech Models During Training

ApplicationsDGX agent

arXiv:2601.13704v3 Announce Type: replace-cross Abstract: In speech machine learning, neural network models are typically designed by choosing an architecture with fixed layer sizes and structure. The

SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse Autoencoders

Model ReleasesDGX agent

arXiv:2509.21379v3 Announce Type: replace-cross Abstract: Concept unlearning in diffusion models is hampered by feature splitting, where concepts are distributed across many latent features, making th

Sequential Least-Squares Estimators with Fast Randomized Sketching for Linear Statistical Models

Model ReleasesDGX agent

arXiv:2509.06856v2 Announce Type: replace-cross Abstract: We propose a novel randomized framework for the estimation problem of large-scale linear statistical models, namely Sequential Least-Squares E

Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO

SafetyDGX agent

arXiv:2605.30789v1 Announce Type: cross Abstract: We identify a new dimension for enhancing rollout diversity in Group Relative Policy Optimization (GRPO) for LLMs. While GRPO relies on diverse rollou

Structured interactions improve distributed coordination beyond model scaling in a real-world multi-robot system

Model ReleasesDGX agent

arXiv:2605.30383v1 Announce Type: cross Abstract: Scaling individual robot capabilities is common but costly. Here we investigate a system-level design question in real-world multi-robot coordination:

The Illusion of Generalization in Tabular Language Models

Model ReleasesDGX agent

arXiv:2602.04031v2 Announce Type: replace Abstract: Tabular Language Models (TLMs) have been claimed to achieve strong generalization for tabular prediction. We conduct a systematic re-evaluation of T

Unlearning in Diffusion Models: A Unified Framework with KL Divergence and Likelihood Constraints

ResearchDGX agent

arXiv:2605.30825v1 Announce Type: cross Abstract: Unlearning in diffusion models aims to remove undesirable data or concepts while preserving the utility of pretrained models -- two fundamentally conf

Vision-Language Models Suppress Female Representations Under Ambiguous Input

SafetyDGX agent

arXiv:2605.31556v1 Announce Type: cross Abstract: Alignment teaches vision-language models (VLMs) to avoid expressing demographic biases, and when gender is clearly visible they largely succeed. Far l

Welcome NVIDIA Cosmos 3: The First Open Omni-model for Physical AI Reasoning and Action

HardwareDGX agent

NVIDIA Cosmos 3 is an open-source omni-model designed for physical AI reasoning and action tasks, representing an advancement in multimodal AI systems. The model integrates multiple modalities to enab

World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry

SafetyDGX agent

arXiv:2604.01985v2 Announce Type: replace-cross Abstract: General-purpose world models promise scalable policy evaluation, optimization, and planning, yet achieving the required level of robustness re

31 May 2026

Stable Diffusion model recommendations for faster and cleaner outputs in 2026?

Local AiDGX agent

A Reddit discussion seeking Stable Diffusion model recommendations for faster and cleaner image outputs, addressing the reality that no single best model exists as the right choice depends on hardware

We need more coding and agent traces public sharing to build datasets and better open source models! Lots of people contributing already, yo…

Model ReleasesDGX agent

We need more coding and agent traces public sharing to build datasets and better open source models! Lots of people contributing already, you should share yours too! https://huggingface.co/datasets?se

30 May 2026

is having a four month lead a sustainable multitrillion dollar business model?

SafetyDGX agent

is having a four month lead a sustainable multitrillion dollar business model? We took another look at the capability gap between open-weight and proprietary models. Since the start of the year, open-

29 May 2026

A Deep Learning Model of Mental Rotation Informed by Interactive VR Experiments

AgentsDGX agent

arXiv:2512.13517v2 Announce Type: replace-cross Abstract: Mental rotation -- the ability to compare objects seen from different viewpoints -- is a fundamental example of mental simulation and spatial

Access Sets Matter: Budgeting Expert Reads for Scalable Weight-Space Model Merging

Model ReleasesDGX agent

arXiv:2605.29489v1 Announce Type: new Abstract: Weight-space model merging is usually formulated as an algebraic operation on checkpoints, yet at LLM scale the limiting resource is often the set of ex

Auditing Training Data in Generative Music Models via Black-Box Membership Inference

SafetyDGX agent

arXiv:2605.29202v1 Announce Type: new Abstract: Recent advances in text-to-music generation enable high-fidelity synthesis of structured musical audio, raising growing concerns about data provenance,

Benchmarking Positional Encoding Strategies for Transformer-Based EEG Foundation Models

Model ReleasesDGX agent

arXiv:2605.29754v1 Announce Type: new Abstract: Electroencephalography (EEG) is a widely used non-invasive technique for measuring brain activity in brain-computer interface (BCI) applications. Superv

BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders

Model ReleasesDGX agent

arXiv:2605.30162v1 Announce Type: new Abstract: Biosecurity evaluations of language models typically ask whether models produce hazardous output. This paper asks a complementary question: when a model

DLM-SWAI: Steering Diffusion Language Models Before They Unmask

SafetyDGX agent

arXiv:2605.29626v1 Announce Type: cross Abstract: Steering language model generation toward desired textual properties is essential for practical deployment, and inference-time methods are particularl

Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding

ResearchDGX agent

arXiv:2605.29707v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference by drafting multiple tokens and verifying them in parallel with the target model. However, its practical

ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models

TutorialsDGX agent

arXiv:2605.29438v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models are a powerful paradigm for generalist robotic control. However, their high computational cost and limited control f

Large-Scale AI and Foundation Models for Neuroscience: A Comprehensive Review

ResearchDGX agent

arXiv:2510.16658v3 Announce Type: replace Abstract: The development of large-scale artificial intelligence (AI) models is influencing neuroscience research by enabling end-to-end learning from raw bra

LLMSurgeon: Diagnosing Data Mixture of Large Language Models

ResearchDGX agent

arXiv:2605.30348v1 Announce Type: cross Abstract: The pretraining data mixture of Large Language Models (LLMs) constitutes their 'digital DNA', shaping model behaviors, capabilities, and failure modes

Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context

ResearchDGX agent

arXiv:2510.06182v2 Announce Type: replace Abstract: A key component of in-context reasoning is the ability of language models (LMs) to bind entities for later retrieval. For example, an LM might repre

Mixing Vector Model for Copolymer Inference via Mixed Integer Linear Programming

ResearchDGX agent

arXiv:2605.29329v1 Announce Type: cross Abstract: A novel two-phase molecule inference framework, mol-infer, has recently been developed to infer chemical graphs with prescribed abstract structures an

models underestimate how much work it takes (token usage) to accomplish a task, just like us

Model ReleasesDGX agent

models underestimate how much work it takes (token usage) to accomplish a task, just like us 🧵 Claude-Opus-4.8 takes you too much tokens - but is this issue general across agents? Do agents know how m

Multi-level Collaborative Distillation Meets Global Workspace Model: A Unified Framework for OCIL

TutorialsDGX agent

arXiv:2508.08677v2 Announce Type: replace-cross Abstract: Online Class-Incremental Learning (OCIL) enables models to learn continuously from non-i.i.d. data streams. Since samples of the data streams

New LFM2.5 8b A1b model!!

Local AiDGX agent

Liquid AI released LFM2.5-8B-A1B, a device-optimized model designed to power real-life applications on phones, laptops, PCs, robots, and lightweight server-side use-cases. The model is a fast, memory-

OOD-GraphLLM: Graph Large Language Model for Out-of-Distribution Generalized Drug Synergy Prediction

Model ReleasesDGX agent

arXiv:2605.30247v1 Announce Type: new Abstract: Drug synergy prediction (DSP) aims to identify efficacious drug combinations under various cellular contexts with different targets. However, the contin

Order-Agnostic Autoregressive Modelling with Missing Data

ApplicationsDGX agent

arXiv:2605.06355v2 Announce Type: replace Abstract: Order-Agnostic autoregressive models have demonstrated strong performance in deep generative modeling, yet their use in settings with incomplete dat

Orthogonal Concept Erasure for Diffusion Models

Model ReleasesDGX agent

arXiv:2605.28902v1 Announce Type: new Abstract: Concept erasure has emerged as a promising approach to mitigate undesired or unsafe content in diffusion models, yet existing methods still face signifi

Parallax: Parameterized Local Linear Attention for Language Modeling

Model ReleasesDGX agent

arXiv:2605.29157v1 Announce Type: cross Abstract: Large Language Models (LLMs) have become the central paradigm in artificial intelligence, yet the core computational primitive of attention has remain

PassNet: Scaling Large Language Models for Graph Compiler Pass Generation

ApplicationsDGX agent

arXiv:2605.29357v1 Announce Type: new Abstract: Modern tensor compilers such as TorchInductor deliver substantial speedups on mainstream models, yet face a systematic performance ceiling on long-tail

Rare Event Analysis of Large Language Models

ResearchDGX agent

arXiv:2602.06791v2 Announce Type: replace Abstract: Being probabilistic models, during inference large language models (LLMs) display rare events: behaviour that is far from typical but highly signifi

Reasoning-preserved Efficient Distillation of Large Language Models via Activation-aware Initialization

Model ReleasesDGX agent

arXiv:2605.29327v1 Announce Type: new Abstract: Efficient Distillation (EDistill) compresses large language models (LLMs) by structured pruning parameters and tuning lightweight modules with high trai

S-MARC: Causal Streaming Reasoning for Full-Duplex Conversational Behavior Modeling

Model ReleasesDGX agent

arXiv:2602.11065v2 Announce Type: replace-cross Abstract: Human conversation is organized by an implicit chain of thought and manifests as temporally structured conversational behaviors. Capturing thi

SCoOP: Semantic Consistent Opinion Pooling for Uncertainty Quantification in Multiple Vision-Language Model Systems

ResearchDGX agent

arXiv:2603.23853v3 Announce Type: replace Abstract: Combining multiple Vision-Language Models (VLMs) can enhance multimodal reasoning and robustness, but aggregating heterogeneous models' outputs ampl

Sequential Physics-Constrained Neural Operator Forward Modeling for the extit{Norne} Reservoir System

Model ReleasesDGX agent

arXiv:2605.28909v1 Announce Type: new Abstract: We develop a comprehensive mathematical and computational framework for sequential surrogate modeling of three-phase black-oil reservoir dynamics using

SMolLM: Small Language Models Learn Small Molecular Grammar

Model ReleasesDGX agent

arXiv:2605.06322v2 Announce Type: replace Abstract: Language models for molecular design have scaled to hundreds of millions of parameters, yet how they learn chemical grammar is poorly understood. We

Stable-Layers: Fine-Tuning Image Layer Decomposition Models with VLM-Scored Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.30257v1 Announce Type: new Abstract: We present Stable-Layers, a reinforcement learning framework that eliminates the need for paired supervision by fine-tuning a pretrained layer decomposi

Steering Language Models Before They Speak: Logit-Level Interventions

ResearchDGX agent

arXiv:2601.10960v2 Announce Type: replace-cross Abstract: Controllable generation requires language models to realize output characteristics such as reading level, politeness, and toxicity. Existing s

Tailoring the Curriculum: Student-Centered Reasoning Distillation via Dynamic Data-Model Compatibility

ResearchDGX agent

arXiv:2605.29229v1 Announce Type: new Abstract: Reasoning distillation transfers complex reasoning abilities from large language models (LLMs) to smaller ones, yet its success depends on how well the

Token Inflation: How Dishonest Providers Can Overcharge for Large Language Model Usage

ResearchDGX agent

arXiv:2605.30040v1 Announce Type: cross Abstract: Per-token billing is now the standard pricing model for commercial large language models (LLMs), so the honesty of reported token counts directly affe

TrojanTO: Action-Level Backdoor Attacks against Trajectory Optimization Models

ResearchDGX agent

arXiv:2506.12815v2 Announce Type: replace Abstract: Recent advances in Trajectory Optimization (TO) models have achieved remarkable success in offline reinforcement learning. However, their vulnerabil

28 May 2026

AI researchers ran 15-day simulations of worlds governed by different AI models: Claude Sonnet 4.6 recorded no crimes, while Gemini 3 Flash had the most at 683 (Jake Angelo/Fortune)

Model ReleasesDGX agent

Jake Angelo / Fortune: AI researchers ran 15-day simulations of worlds governed by different AI models: Claude Sonnet 4.6 recorded no crimes, while Gemini 3 Flash had the most at 683 — Imagine a world

Analyzing Cancer Patients' Experiences with Embedding-based Topic Modeling and LLMs

ApplicationsDGX agent

arXiv:2601.12154v2 Announce Type: replace Abstract: This study investigates the use of neural topic modeling and LLMs to uncover meaningful themes from patient storytelling data, to offer insights tha

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring

Model ReleasesDGX agent

arXiv:2502.05242v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are becoming increasingly capable, but the mechanisms of their thinking and decision-making processes remain uncl

Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models

Model ReleasesDGX agent

arXiv:2605.27383v1 Announce Type: cross Abstract: Spoken Language Models (SLMs) have emerged as a promising paradigm for speech synthesis by bypassing explicit grapheme-to-phoneme pipelines. However,

Claude Opus 4.8 is out today. It's our strongest coding model yet: up on SWE-bench Pro (from 64.3 to 69.2) and noticeably more honest about …

Model ReleasesDGX agent

Claude Opus 4.8 is out today. It's our strongest coding model yet: up on SWE-bench Pro (from 64.3 to 69.2) and noticeably more honest about its own work. It tells you when it's unsure and catches its

Colosseum V2: Benchmarking Generalization for Vision Language Action Models

Model ReleasesDGX agent

arXiv:2605.27759v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models demonstrate promising generalization in robotic manipulation, driven by advances in large-scale vision and language

Entropy Distribution as a Fingerprint for Hallucinations in Generative Models

ResearchDGX agent

arXiv:2605.28264v1 Announce Type: new Abstract: Large Language Models (LLMs) often generate factually incorrect outputs, commonly termed hallucinations, that undermine trust and limit deployment in hi

FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models

ResearchDGX agent

arXiv:2512.16483v2 Announce Type: replace Abstract: Visual Autoregressive (VAR) modeling departs from the next-token prediction paradigm of traditional Autoregressive (AR) models through next-scale pr

IBM and Red Hat commit $5B to establish a new model for open-source software, dubbed Project Lightwell, and will deploy 20,000 engineers, supported by AI (Connor Hart/Wall Street Journal)

Model ReleasesDGX agent

Connor Hart / Wall Street Journal: IBM and Red Hat commit $5B to establish a new model for open-source software, dubbed Project Lightwell, and will deploy 20,000 engineers, supported by AI — Project L

Measuring Form and Function in Language Models

ResearchDGX agent

arXiv:2605.28616v1 Announce Type: cross Abstract: We introduce quantitative metrics for child language acquisition to evaluate language models. Our focus is on the formal syntactic and functional disc

MemGuard: Preventing Memory Contamination in Long-Term Memory-Augmented Large Language Models

Model ReleasesDGX agent

arXiv:2605.28009v1 Announce Type: cross Abstract: Memory-augmented large language models extend reasoning beyond a fixed context window by maintaining long-term memory across interactions. However, ex

MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems

Model ReleasesDGX agent

arXiv:2605.28732v1 Announce Type: cross Abstract: Memory is essential for enabling large language models to support long-horizon reasoning, yet existing memory systems remain unreliable and difficult

OpenCode Now Supports DigitalOcean Inference Router for Intelligent Model Routing

IndustryDGX agent

OpenCode now integrates with DigitalOcean's Inference Router, enabling intelligent routing of AI model requests across distributed infrastructure. This integration allows developers to optimize model

← Previous
1…109110111112113…1009
Next →