AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,490 results
29 May 2026

Pocket-Dentist: On-Device Dental Image Understanding via Efficient Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2605.29299v1 Announce Type: cross Abstract: Evaluations of dental vision-language models remain fragmented across datasets, task definitions and metrics, and often ignore their computational cos

Prioritize the Process, Not Just the Outcome: Rewarding Latent Thought Trajectories Improves Reasoning in Looped Language Models

Model ReleasesDGX agent

arXiv:2602.10520v3 Announce Type: replace Abstract: Looped Language Models (LoopLMs) perform multi-step latent reasoning prior to token generation and outperform conventional LLMs on reasoning benchma

Reasoning that Travels: Dissecting How Chain-of-Thought Transfers Across Models

TutorialsDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.28913v1 Announce Type: new Abstract: Large reasoning models (LRMs) often generate extensive chain-of-thought (CoT) traces before producing a final answer. As explicit textual artifacts, the

Spurious Prompts: Can Irrelevant Prompts Steer Large Language Models?

ResearchDGX agent

arXiv:2605.29678v1 Announce Type: new Abstract: Large language models are highly sensitive to prompts, but this sensitivity is usually studied through task-relevant instructions, demonstrations, or re

Teaching Language Models to Check Grounded Claim Factuality with Human Test-Taking Strategies

Model ReleasesDGX agent

arXiv:2605.29712v1 Announce Type: cross Abstract: Grounded claim factuality checking is important for large language model (LLM) applications such as retrieval-augmented generation, as it helps users

Three-dimensional Conditional Diffusion Models for Cosmological 21 cm Lightcone Emulation

Model ReleasesDGX agent

arXiv:2605.29016v1 Announce Type: cross Abstract: We investigate conditional diffusion modeling for three-dimensional 21 cm lightcone emulation, focusing on cubes with a sky-plane size of 64imes64 and

Towards a Foundation Model for the Martian Atmosphere

ResearchDGX agent

arXiv:2605.28851v1 Announce Type: cross Abstract: The martian atmosphere hosts dynamical phenomena ranging from planet-encircling dust storms to mesoscale orographic clouds and nocturnal low-level jet

Towards Foundation Models for Zero-Shot Time Series Anomaly Detection: Leveraging Synthetic Data and Relative Context Discrepancy

ResearchDGX agent

arXiv:2509.21190v4 Announce Type: replace-cross Abstract: Time series anomaly detection (TSAD) is a critical task, but developing models that generalize to unseen data in a zero-shot manner remains a

Towards Understanding the Shape of Representations in Protein Language Models

Local AiDGX agent

arXiv:2509.24895v2 Announce Type: replace Abstract: While protein language models (PLMs) are one of the most promising avenues of research for future de novo protein design, the way in which they tran

When the Same Coefficients Reach Different Places: Asymmetric Realizability in Transplanting Tokenizers across Large Language Models

Model ReleasesDGX agent

arXiv:2601.00065v3 Announce Type: replace-cross Abstract: Tokenizer transplant in cross-vocabulary model composition reconstructs donor-only embedding rows as weighted combinations over shared lexical

28 May 2026

AIBuildAI-2: A Knowledge-Enhanced Agent for Automatically Building AI Models

AgentsDGX agent

arXiv:2605.27873v1 Announce Type: new Abstract: AI models underpin data-centric applications from image and text processing to scientific discovery in biology, physics, and chemistry. Yet developing t

AlphaForgeBench: Benchmarking End-to-End Trading Strategy Design with Large Language Models

Model ReleasesDGX agent

arXiv:2602.18481v2 Announce Type: replace-cross Abstract: The rapid advancement of Large Language Models (LLMs) has led to a surge of financial benchmarks, evolving from static knowledge evaluation to

Anthropic says it expects Mythos-class models to be available to all customers 'in the coming weeks' following the development of stronger safeguards (Madison Mills/Axios)

Model ReleasesDGX agent

Madison Mills / Axios: Anthropic says it expects Mythos-class models to be available to all customers “in the coming weeks” following the development of stronger safeguards — Anthropic released Claude

Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems

Model ReleasesDGX agent

arXiv:2605.27492v1 Announce Type: cross Abstract: LLM agents are rapidly evolving from coding assistants into autonomous software engineering systems. However, existing evaluation methodologies remain

Cultural Binding Heads in Language Models

Model ReleasesDGX agent

arXiv:2605.28543v1 Announce Type: new Abstract: LLMs often default to equal treatment across cultural groups, even though context warrants differentiation: this is a lack of difference awareness. Usin

DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes

SafetyDGX agent

arXiv:2605.28421v1 Announce Type: new Abstract: Reinforcement learning has become a central paradigm for advancing reasoning in large language models, yet most existing methods still depend on stronge

From Pixels to Words -- Towards Native One-Vision Models at Scale

SafetyDGX agent

arXiv:2605.28820v1 Announce Type: new Abstract: Current vision-language models (VLMs) typically stitch together separate image encoders and language decoders via multi-stage alignment, a modular frame

High Performance, Low Reliability: Uncertainty Benchmarking for Tabular Foundation Models

Model ReleasesDGX agent

arXiv:2605.28554v1 Announce Type: new Abstract: Recent Tabular Foundation Models (TFMs) have demonstrated state-of-the-art predictive performance, often surpassing Gradient-Boosted Decision Trees (GBD

Interpretability-Guided Layer Selection over Subspace Projection: SAEs as Stethoscopes, Not Scalpels, for Raw Task Vector Model Editing

Model ReleasesDGX agent

arXiv:2605.28649v1 Announce Type: cross Abstract: LLMs increasingly require surgical model editing to enhance domain-specific capabilities without incurring the computational cost or catastrophic forg

LESA: Learnable Stage-Aware Predictors for Diffusion Model Acceleration

Model ReleasesDGX agent

arXiv:2602.20497v3 Announce Type: replace-cross Abstract: Diffusion models have achieved remarkable success in image and video generation tasks. However, the high computational demands of Diffusion Tr

{Omega}-QVLA: Robust Quantization for Vision-Language-Action Models via Composite Rotation and Per-step Scaling

Model ReleasesDGX agent

arXiv:2605.28803v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models unify perception, reasoning, and control within a single policy, yet their multi-billion-parameter backbones and dif

On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation

ResearchDGX agent

arXiv:2601.06329v2 Announce Type: replace-cross Abstract: Generative spoken language models pretrained on large-scale raw audio can continue a speech prompt with appropriate content while preserving a

Playing with Words, Improving with Rewards: Training Language Models for Creative Association

ResearchDGX agent

arXiv:2605.27832v1 Announce Type: new Abstract: Large Language Models (LLMs) are being applied to increasingly difficult problems and use cases. To navigate their vast solution spaces effectively, LLM

PrunePath: Towards Highly Structured Sparse Language Models

Model ReleasesDGX agent

arXiv:2605.28283v1 Announce Type: cross Abstract: Feed-forward networks (FFNs) dominate the parameter count and computation of modern language models, yet existing pruning methods often struggle to co

Pruning and Distilling Mixture-of-Experts into Dense Language Models

Model ReleasesDGX agent

arXiv:2605.28207v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) is now the dominant architecture for frontier language models, yet it requires all expert parameters to be loaded in memory,

Revisiting 2D Foundation Models for Scalable 3D Medical Image Classification

Model ReleasesDGX agent

arXiv:2512.12887v3 Announce Type: replace Abstract: 3D medical image classification is essential for modern clinical workflows. Medical foundation models (FMs) have emerged as a promising approach for

RGC: a radio AGN classifier based on deep learning. I. A semi-supervised multiclass model for VLA images

Model ReleasesDGX agent

arXiv:2510.22190v2 Announce Type: replace-cross Abstract: Bent radio active galactic nuclei (RAGNs) -- wide-angle tails (WATs) and narrow-angle tails (NATs) -- trace dense environments in galaxy group

The Point, the Vision and the Text: Does Point Cloud Boost Spatial Reasoning of Large Language Models? A Bias-Controlled Study

Model ReleasesDGX agent

arXiv:2504.04540v2 Announce Type: replace-cross Abstract: 3D Large Language Models (LLMs) leveraging spatial information in point clouds for 3D spatial reasoning attract great attention. Despite some

Trust Me, I'm an Expert: Decoding and Steering Authority Bias in Large Language Models

SafetyDGX agent

arXiv:2601.13433v3 Announce Type: replace Abstract: Prior research demonstrates that performance of language models on reasoning tasks can be influenced by suggestions, hints and endorsements. However

27 May 2026

A Deep State-Space Model Compression Method using Upper Bound on Output Error

Model ReleasesDGX agent

arXiv:2510.14542v2 Announce Type: replace-cross Abstract: We study deep state-space models (Deep SSMs) that contain linear quadratic-output (LQO) systems as internal blocks and present a compression m

AirCast-SR: A Foundation Model for Kilometer-Scale Atmospheric Super-Resolution via Latent Consistency Diffusion

SafetyDGX agent

arXiv:2605.26130v1 Announce Type: new Abstract: Operational weather prediction at kilometer scales remains computationally prohibitive for traditional numerical weather prediction (NWP) models, limiti

Announcing ESMFold2, our new state-of-the-art structure prediction model capable of predicting structure from single sequences or MSAs. ESMF…

Model ReleasesDGX agent

Announcing ESMFold2, our new state-of-the-art structure prediction model capable of predicting structure from single sequences or MSAs. ESMFold2 improves on benchmarks of protein-protein interaction a

Aperiodic and Low-Frequency Spectral Bias in Reconstruction based EEG Foundation Models

SafetyDGX agent

arXiv:2605.26434v1 Announce Type: cross Abstract: EEG foundation models, pre-trained on large-scale unlabelled EEG data, have emerged as a promising direction towards learning generalizable EEG repres

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models

SafetyDGX agent

arXiv:2605.06213v2 Announce Type: replace Abstract: Evaluating large language models (LLMs) today rests on fixed benchmarks that apply the same set of items to any model, producing ceiling and floor e

EconCausal: A Context-Aware Economic Reasoning Benchmark for Large Language Models

Model ReleasesDGX agent

arXiv:2510.07231v4 Announce Type: replace-cross Abstract: Socio-economic causal effects depend heavily on their institutional and environmental contexts. The same intervention can produce different, e

Erased but Exploitable: Black-box Embedding-Aware Prompting Against Unlearned Text-to-Image Diffusion Models

SafetyDGX agent

arXiv:2605.26332v1 Announce Type: cross Abstract: Machine unlearning aims to remove specific concepts from pretrained text-to-image diffusion models, yet several white- and black-box attacks have been

Few-shot Cross-country Generalization of Tabular Machine Learning and Foundation Models for Childhood Anemia Prediction under Distribution Shift

SafetyDGX agent

arXiv:2605.26589v1 Announce Type: cross Abstract: Childhood anemia affects around 40% of children aged 6-59 months globally and arises from heterogeneous factors, limiting model generalizability. We e

Hubness, Not Anisotropy, Drives Cross-Lingual Retrieval Asymmetry in Multilingual Embedding Models

Model ReleasesDGX agent

arXiv:2605.26575v1 Announce Type: new Abstract: Multilingual embedding models are deployed under the assumption that cross-lingual retrieval is symmetric: if a query in language A retrieves its transl

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?

Model ReleasesDGX agent

arXiv:2605.26548v1 Announce Type: cross Abstract: Large language models (LLMs) now support automated software security tasks, including vulnerability discovery and proof-of-concept (PoC) generation. E

Towards Controllable Image Generation through Representation-Conditioned Diffusion Models

ResearchDGX agent

arXiv:2605.27343v1 Announce Type: new Abstract: Diffusion models have emerged as powerful tools for high-quality image generation and editing, but guiding these models to produce specific outputs rema

26 May 2026

A tool to get Claude Code-style reliability from fully local models

Model ReleasesDGX agent

Ollama exposes an Anthropic-compatible Messages endpoint , allowing developers to run powerful open-source AI models locally with no API costs and pair them with Claude Code for a capable local AI cod

ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2605.24011v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models exhibit remarkable action generation for embodied intelligence, but their heavy compute make deployment on edge pl

Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning

Model ReleasesDGX agent

arXiv:2602.10090v3 Announce Type: replace Abstract: Recent advances in large language model (LLM) have empowered autonomous agents to perform multi-turn interactions with tools and environments. Howev

AstroMind: A High-Fidelity Benchmark for Spacecraft Behavior Reasoning Based on Large Language Models

Model ReleasesDGX agent

arXiv:2605.24573v1 Announce Type: new Abstract: Understanding why a spacecraft maneuvers -- rather than simply that it did -- is an increasingly important problem for space domain awareness as Earth o

Benchmarking Pathology Foundation Models for Spatial Domain Understanding

Model ReleasesDGX agent

arXiv:2605.25764v1 Announce Type: cross Abstract: Pathology foundation models (PFMs) have emerged as a core approach for learning transferable representations from whole slide images (WSIs), and they

Beyond Query Memorization: Large Language Model Routing with Query Decomposition and Historical Matching

Model ReleasesDGX agent

arXiv:2605.25558v1 Announce Type: new Abstract: Optimizing the trade-off among predictive performance and computational cost is a central focus in the deployment of Large Language Models (LLMs). Curre

Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models

Model ReleasesDGX agent

arXiv:2605.25189v1 Announce Type: cross Abstract: Reward hacking arises when a model improves a proxy reward by exploiting shortcuts rather than solving the intended task. We study this failure mode t

DIVER-1: Scaling Intracranial EEG Foundation Models for Transferable Representations

Model ReleasesDGX agent

arXiv:2512.19097v3 Announce Type: replace-cross Abstract: Intracranial EEG (iEEG) provides direct, millisecond-scale recordings of human neural activity, but reusable representation learning is diffic

Extracting Training Data from Diffusion Language Models via Infilling

SafetyDGX agent

arXiv:2605.24173v1 Announce Type: cross Abstract: Memorization in large language models has been studied almost exclusively through prefix-conditioned extraction, a natural choice for autoregressive m

Factored Latent Action World Models

SafetyDGX agent

arXiv:2602.16229v2 Announce Type: replace Abstract: Learning latent actions from action-free video has emerged as a powerful paradigm for scaling up controllable world model learning. Latent actions p

Game-Theoretic Modeling of Heterogeneous Investor Interactions for Stock Price Forecasting

Model ReleasesDGX agent

arXiv:2605.23953v1 Announce Type: cross Abstract: Accurate stock price forecasting has consistently remained a pivotal yet challenging FinTech task that underpins quantitative trading and investment d

Generative OOD-regularized Model-based Policy Optimization

SafetyDGX agent

arXiv:2605.24405v1 Announce Type: cross Abstract: We study sequential decision-making with offline reinforcement learning (RL). Traditional offline RL policies may result in out-of-distribution (OOD)

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework

Model ReleasesDGX agent

arXiv:2507.19219v2 Announce Type: replace Abstract: Overestimation in evaluating large language models (LLMs) has become an increasing concern. Due to the contamination of public benchmarks or imbalan

Language Models Need Sleep

Local AiDGX agent

arXiv:2605.26099v1 Announce Type: cross Abstract: Transformer-based large language models are increasingly used for long-horizon tasks; however, their attention mechanism scales poorly with context le

M^3-Verse: A 'Spot the Difference' Challenge for Large Multimodal Models

Model ReleasesDGX agent

arXiv:2512.18735v2 Announce Type: replace-cross Abstract: Modern Large Multimodal Models (LMMs) have demonstrated extraordinary ability in static image and single-state spatial-temporal understanding.

NSR-Boost: A Neuro-Symbolic Residual Boosting Framework for Industrial Legacy Models

ApplicationsDGX agent

arXiv:2601.10457v3 Announce Type: replace Abstract: Although the Gradient Boosted Decision Trees (GBDTs) dominate industrial tabular applications, upgrading legacy models in high-concurrency productio

PALoRA: Projection-Adaptive LoRA for Preserving Reasoning in Large Language Models

Model ReleasesDGX agent

arXiv:2605.24549v1 Announce Type: new Abstract: Efficiently updating Large Language Models (LLMs) with new or evolving factual knowledge remains a central challenge, as even parameter-efficient adapta

ParkourFormer: Integrating Predictive Supervision and Sequence Modeling into Parkour Locomotion

Model ReleasesDGX agent

arXiv:2605.25782v1 Announce Type: new Abstract: Humanoid parkour requires locomotion policies to coordinate whole-body dynamics across rapidly changing terrains such as stairs, gaps, slopes, and obsta

PolySAE: Modeling Feature Interactions in Sparse Autoencoders via Polynomial Decoding

Model ReleasesDGX agent

arXiv:2602.01322v2 Announce Type: replace-cross Abstract: Sparse autoencoders (SAEs) interpret neural network representations by decomposing activations into sparse combinations of dictionary atoms. H

Representation Without Control: Testing the Realization Effect in Language Models

Model ReleasesDGX agent

arXiv:2605.25151v1 Announce Type: new Abstract: Large language models are increasingly used as behavioral simulators, but it remains unclear when their outputs reflect human-like cognitive mechanisms

← Previous
1…8687888990…1009
Next →