AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlog
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,490 results
Research

Why Do Reasoning Models Lose Coverage? The Role of Data and Forks in the Road

DGX agent

arXiv:2605.17026v1 Announce Type: new Abstract: Recent progress in large language models has led to the emergence of reasoning models, which have shown strong performance on complex tasks through spec

researcharxiv-cs-lg
19 May 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Evaluating Chinese Ambiguity Understanding in Large Language Models

DGX agent

arXiv:2605.15635v1 Announce Type: new Abstract: Linguistic ambiguity is critical to the robustness of Large Language Models (LLMs), yet existing research focuses mostly on English, with limited attent

model-releasesarxiv-cs-cl
18 May 2026
Safety

f-Trajectory Balance: A Loss Family for Tuning GFlowNets, Generative Models, and LLMs with Off- and On-Policy Data

DGX agent

arXiv:2605.15417v1 Announce Type: cross Abstract: In GFlowNets and variational inference, it has been shown that the mean square error between target and model log probabilities is an effective, low v

safetyarxiv-cs-ai
18 May 2026
Model Releases

FINESSE-Bench: A Hierarchical Benchmark Suite for Financial Domain Knowledge and Technical Analysis in Large Language Models

DGX agent

arXiv:2605.15482v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being applied to financial analysis, reporting, investment decision support, risk management, compliance,

model-releasesarxiv-cs-cl
18 May 2026
Model Releases

Frontier Large Language Models Rival State-of-the-Art Planners

DGX agent

arXiv:2511.09378v2 Announce Type: replace Abstract: A series of influential studies established that large language models cannot reliably solve even simple planning tasks. We show that the latest gen

model-releasesarxiv-cs-ai
18 May 2026
Research

Learning Normalized Energy Models for Linear Inverse Problems

DGX agent

arXiv:2605.15487v1 Announce Type: cross Abstract: Generative diffusion models can provide powerful prior probability models for inverse problems in imaging, but existing implementations suffer from tw

researcharxiv-cs-cv
18 May 2026
Model Releases

MHGraphBench: Knowledge Graph-Grounded Benchmarking of Mental Health Knowledge in Large Language Models

DGX agent

arXiv:2605.15589v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in the mental health domain, yet it remains unclear how well they capture related biomedical knowledg

model-releasesarxiv-cs-cl
18 May 2026
Model Releases

Social-Mamba: Socially-Aware Trajectory Forecasting with State-Space Models

DGX agent

arXiv:2605.15424v1 Announce Type: new Abstract: Human trajectory forecasting is crucial for safe navigation in crowded environments, requiring models that balance accuracy with computational efficienc

model-releasesarxiv-cs-cv
18 May 2026
Research

Time-Varying Deep State Space Models for Sequences with Switching Dynamics

DGX agent

arXiv:2605.15311v1 Announce Type: new Abstract: The identification and modeling of time-varying systems is a fundamental challenge in signal processing and system identification. To address this chall

researcharxiv-cs-lg
18 May 2026
Model Releases

Zero-Shot Goal Recognition with Large Language Models

DGX agent

arXiv:2605.15333v1 Announce Type: new Abstract: Large language models have recently reached near-parity with classical planners on well-known planning domains, yet this competence relies on world-know

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

Eradicating Negative Transfer in Multi-Physics Foundation Models via Sparse Mixture-of-Experts Routing

DGX agent

arXiv:2605.15179v1 Announce Type: cross Abstract: Scaling Scientific Machine Learning (SciML) toward universal foundation models is bottlenecked by negative transfer: the simultaneous co-training of d

model-releasesarxiv-cs-ai
15 May 2026
Applications

Finding Interpretable Prompt-Specific Circuits in Language Models

DGX agent

arXiv:2602.13483v2 Announce Type: replace-cross Abstract: Understanding the internal circuits that language models use to solve tasks remains a central challenge in mechanistic interpretability. A cru

applicationsarxiv-cs-ai
15 May 2026
Model Releases

MechVerse: Evaluating Physical Motion Consistency in Video Generation Models

DGX agent

arXiv:2605.14843v1 Announce Type: new Abstract: Text- and image-conditioned video generation models have achieved strong visual fidelity and temporal coherence, but they often fail to generate motion

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity Language Model with Curriculum Learning and Native Tool Use

DGX agent

arXiv:2605.13989v1 Announce Type: new Abstract: We present VectraYX-Nano, a 41.95M-parameter decoder-only language model trained from scratch in Spanish for cybersecurity, with a Latin-American focus

model-releasesarxiv-cs-cl
15 May 2026
Agents

AI Harness Engineering: A Runtime Substrate for Foundation-Model Software Agents

DGX agent

arXiv:2605.13357v1 Announce Type: cross Abstract: Foundation models have transformed automated code generation, yet autonomous software-engineering agents remain unreliable in realistic development se

agentsarxiv-cs-ai
14 May 2026
Research

Asymmetric Flow Models

DGX agent

arXiv:2605.12964v1 Announce Type: new Abstract: Flow-based generation in high-dimensional spaces is difficult because velocity prediction requires modeling high-dimensional noise, even when data has s

researcharxiv-cs-cv
14 May 2026
Agents

haha, it's a dramatic voice model

DGX agent

haha, it's a dramatic voice model We're releasing a whole new category of voice models. Introducing DramaBox — our state-of-the-art, open source voice model built for cinematic use cases. Traditional

agentsyohei-nakajima--x
14 May 2026
Model Releases

(How) Do Large Language Models Understand High-Level Message Sequence Charts?

DGX agent

arXiv:2605.13773v1 Announce Type: cross Abstract: Large Language Models (LLMs) are being employed widely to automate tasks across the software development life-cycle. It is, however, unclear whether t

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

How Well Do Large-Scale Chemical Language Models Transfer to Downstream Tasks?

DGX agent

arXiv:2602.11618v4 Announce Type: replace Abstract: Chemical Language Models (CLMs) pre-trained on large scale molecular data are widely used for molecular property prediction. However, the common bel

model-releasesarxiv-cs-lg
14 May 2026
Safety

Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling

DGX agent

arXiv:2605.13801v1 Announce Type: cross Abstract: As generative AI models such as large language models (LLMs) become more pervasive, ensuring the safety, robustness, and overall trustworthiness of th

safetyarxiv-cs-ai
14 May 2026
Model Releases

Inducing Overthink: Hierarchical Genetic Algorithm-based DoS Attack on Black-Box Large Language Reasoning Models

DGX agent

arXiv:2605.13338v1 Announce Type: cross Abstract: Large Reasoning Models (LRMs) are increasingly integrated into systems requiring reliable multi-step inference, yet this growing dependence exposes ne

model-releasesarxiv-cs-ai
14 May 2026
Safety

Integration of an Agent Model into an Open Simulation Architecture for Scenario-Based Testing of Automated Vehicles

DGX agent

arXiv:2605.13539v1 Announce Type: new Abstract: Simulative and scenario-based testing are crucial methods in the safety assurance for automated driving systems. To ensure that simulation results are r

safetyarxiv-cs-ro
14 May 2026
Model Releases

Large Language Models Lack Temporal Awareness of Medical Knowledge

DGX agent

arXiv:2605.13045v1 Announce Type: new Abstract: The existing methods for evaluating the medical knowledge of Large Language Models (LLMs) are largely based on atemporal examination-style benchmarks, w

model-releasesarxiv-cs-lg
14 May 2026
Model Releases

LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models

DGX agent

arXiv:2505.15616v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have achieved significant advances in integrating visual and linguistic information, yet their ability to r

model-releasesarxiv-cs-cv
14 May 2026
Model Releases

Probing Persona-Dependent Preferences in Language Models

DGX agent

arXiv:2605.13339v1 Announce Type: cross Abstract: Large language models (LLMs) can be said to have preferences: they reliably pick certain tasks and outputs over others, and preferences shaped by post

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

SupChain-Bench: Benchmarking Large Language Models for Real-World Supply Chain Management

DGX agent

arXiv:2602.07342v2 Announce Type: replace Abstract: Large language models (LLMs) have shown promise in complex reasoning and tool-based decision making, motivating their application to real-world supp

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

The critical slowing down in diffusion models

DGX agent

arXiv:2605.12597v1 Announce Type: cross Abstract: Computational sampling has been central to the sciences since the mid-20th century. While machine-learning-based approaches have recently enabled majo

model-releasesarxiv-cs-ai
14 May 2026
Research

The Efficiency Gap in Byte Modeling

DGX agent

arXiv:2605.12928v1 Announce Type: new Abstract: Modern language models have historically relied on two dominant design choices: subword tokenization and autoregressive (AR) ordering. These design deci

researcharxiv-cs-lg
14 May 2026
Safety

AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model

DGX agent

arXiv:2511.22663v5 Announce Type: replace Abstract: Unified multimodal models for image generation and understanding represent a significant step toward AGI and have attracted widespread attention fro

safetyarxiv-cs-cv
13 May 2026
Model Releases

Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training

DGX agent

arXiv:2605.12483v1 Announce Type: new Abstract: In settings where labeled verifiable training data is the binding constraint, each checked example should be allocated carefully. The standard practice

model-releasesarxiv-cs-lg
13 May 2026
Model Releases

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation

DGX agent

arXiv:2605.12034v1 Announce Type: cross Abstract: Omni-modal language models are intended to jointly understand audio, visual inputs, and language, but benchmark gains can be inflated when visual evid

model-releasesarxiv-cs-cv
13 May 2026
Local Ai

CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating

DGX agent

arXiv:2605.11723v1 Announce Type: new Abstract: In this paper, we propose Concentrate and Concentrate (CaC), a coarse-to-fine anomaly reward model based on Vision-Language Models. During inference, it

local-aiarxiv-cs-cv
13 May 2026
Model Releases

DarkQA: Benchmarking Vision-Language Models on Visual-Primitive Question Answering in Low-Light Indoor Scenes

DGX agent

arXiv:2512.24985v4 Announce Type: replace Abstract: Vision Language Models (VLMs) are increasingly adopted as central reasoning modules for embodied agents. Existing benchmarks evaluate their capabili

model-releasesarxiv-cs-cv
13 May 2026
Research

DriftXpress: Faster Drifting Models via Projected RKHS Fields

DGX agent

arXiv:2605.12183v1 Announce Type: new Abstract: Drifting Models have emerged as a new paradigm for one-step generative modeling, achieving strong image quality without iterative inference. The premise

researcharxiv-cs-lg
13 May 2026
Model Releases

GeneZip: Region-Aware Compression for Long Context DNA Modeling

DGX agent

arXiv:2602.17739v3 Announce Type: replace-cross Abstract: Long-context DNA models are limited by token-mixing cost and by how compression allocates representational budget across the genome. Existing

model-releasesarxiv-cs-lg
13 May 2026
Research

Language Modeling with Hyperspherical Flows

DGX agent

arXiv:2605.11125v1 Announce Type: new Abstract: Discrete Diffusion Language Models progressed rapidly as an alternative to autoregressive (AR) models, motivated by their parallel generation abilities.

researcharxiv-cs-lg
13 May 2026
Applications

OmniThoughtVis: A Scalable Distillation Pipeline for Deployable Multimodal Reasoning Models

DGX agent

arXiv:2605.11629v1 Announce Type: new Abstract: Recent multimodal large language models (MLLMs) have shown strong chain-of-thought (CoT) reasoning ability on vision-language tasks, but their direct de

applicationsarxiv-cs-cl
13 May 2026
Model Releases

Overcoming Dynamics-Blindness: Training-Free Pace-and-Path Correction for VLA Models

DGX agent

arXiv:2605.11459v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models achieve remarkable flexibility and generalization beyond classical control paradigms. However, most prevailing VLA

model-releasesarxiv-cs-cv
13 May 2026
Model Releases

Parameter-Efficient Adaptation of Pre-Trained Vision Foundation Models for Active and Passive Seismic Data Denoising

DGX agent

arXiv:2605.10953v1 Announce Type: cross Abstract: The demand for high-resolution subsurface imaging and continuous Earth monitoring has driven rapid growth in active and passive seismic data from dens

model-releasesarxiv-cs-cv
13 May 2026
Research

ReasonEdit: Editing Vision-Language Models using Human Reasoning

DGX agent

arXiv:2602.02408v4 Announce Type: replace Abstract: Model editing aims to correct errors in large, pretrained models without altering unrelated behaviors. While some recent works have edited vision-la

researcharxiv-cs-cv
13 May 2026
Safety

Rethinking external validation for the target population: Capturing patient-level similarity with a generative model

DGX agent

arXiv:2605.11284v1 Announce Type: cross Abstract: Background: External validation is essential for assessing the transportability of predictive models. However, its interpretation is often confounded

safetyarxiv-cs-lg
13 May 2026
Model Releases

STRUM: A Spectral Transcription and Rhythm Understanding Model for End-to-End Generation of Playable Rhythm-Game Charts

DGX agent

arXiv:2605.12135v1 Announce Type: cross Abstract: We present STRUM (Spectral Transcription and Rhythm Understanding Model), an audio-to-chart pipeline that converts raw recordings into playable Clone

model-releasesarxiv-cs-lg
13 May 2026
Research

TCP-SSM: Efficient Vision State Space Models with Token-Conditioned Poles

DGX agent

arXiv:2605.11563v1 Announce Type: new Abstract: State Space Models (SSMs) have emerged as a compelling alternative to attention models for long-range vision tasks, offering input-dependent recurrence

researcharxiv-cs-cv
13 May 2026
Model Releases

Alice v1: Distillation-Enhanced Video Generation Surpassing Closed-Source Models

DGX agent

arXiv:2605.08115v1 Announce Type: cross Abstract: Wepresent Alice v1, a 14-billion parameter open-source video generation model that achieves state-of-the-art quality through consistency distillation

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Architecture, Not Scale: Circuit Localization in Large Language Models

DGX agent

arXiv:2605.08853v1 Announce Type: new Abstract: Mechanistic interpretability assumes that circuit analysis becomes harder as models scale. We challenge this assumption by showing that the attention ar

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Beyond Language: Format-Agnostic Reasoning Subspaces in Large Language Models

DGX agent

arXiv:2605.09496v1 Announce Type: new Abstract: Large language models represent the same reasoning in vastly different surface forms -- English prose, Python code, mathematical notation -- yet whether

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Beyond Local Edits: Embedding-Virtualized Knowledge for Broader Evaluation and Preservation of Model Editing

DGX agent

arXiv:2602.01977v2 Announce Type: replace Abstract: Knowledge editing methods for large language models are commonly evaluated using predefined benchmarks that assess edited facts together with a limi

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Built Environment Reasoning from Remote Sensing Imagery Using Large Vision--Language Models

DGX agent

arXiv:2605.08404v1 Announce Type: cross Abstract: This work investigates the use of large language models (LLMs) for tasks in smart cities. The core idea is to leverage remote sensing imagery to chara

model-releasesarxiv-cs-ai
12 May 2026
← Previous
1…8687888990…1261
Next →