AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,405 results
20 May 2026

MixRea: Benchmarking Explicit-Implicit Reasoning in Large Language Models

Model ReleasesDGX agent

arXiv:2605.20128v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into high-stakes decision-making. Inspired by the theory of inattentional blindness in human co

OpenCompass: A Universal Evaluation Platform for Large Language Models

Model ReleasesDGX agent

arXiv:2605.19276v1 Announce Type: new Abstract: In recent years, the field of artificial intelligence has undergone a paradigm shift from task-specific small-scale models to general-purpose large lang

Prompting language influences diagnostic reasoning and accuracy of large language models

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.19173v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly explored for clinical decision support, yet most evaluations are conducted in English, leaving their relia

RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding

Model ReleasesDGX agent

arXiv:2605.19329v1 Announce Type: cross Abstract: Conventional vision-language models (VLMs) struggle to interpret scenes captured under adverse conditions (e.g., low light, high dynamic range, or fas

Stitched Value Model for Diffusion Alignment

SafetyDGX agent

arXiv:2605.19804v1 Announce Type: cross Abstract: For practical use, diffusion- or flow-based generative models must be aligned with task-specific rewards, such as prompt fidelity or aesthetic prefere

Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models

ResearchDGX agent

arXiv:2605.19137v1 Announce Type: new Abstract: Video foundation models achieve strong performance across many video understanding tasks, but typically require large-scale pre-training on massive vide

WIND: Weather Inverse Diffusion for Zero-Shot Atmospheric Modeling

TutorialsDGX agent

arXiv:2602.03924v2 Announce Type: replace-cross Abstract: Deep learning has revolutionized weather forecasting, but many challenges remain, including climate modeling. Moreover, the current landscape

19 May 2026

A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback

Model ReleasesDGX agent

arXiv:2605.18073v1 Announce Type: cross Abstract: Large Language Models (LLMs) demonstrate strong potential for automated code generation, yet their ability to iteratively refine solutions using execu

AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture

Model ReleasesDGX agent

arXiv:2511.23253v3 Announce Type: replace Abstract: Recent advancements in Vision-Language Models (VLMs) have significantly impacted various industries. In agriculture, these multimodal capabilities h

Ancient Greek to Modern Greek Machine Translation: A Novel Benchmark and Fine-Tuning Experiments on LLMs and NMT Models

Model ReleasesDGX agent

arXiv:2605.18504v1 Announce Type: new Abstract: Machine Translation (MT) for Ancient Greek (AG) to Modern Greek (MG) is a low-resource task, constrained by the lack of large-scale, high-quality parall

Beacon: Single-Turn Diagnosis and Mitigation of Latent Sycophancy in Large Language Models

Model ReleasesDGX agent

arXiv:2510.16727v2 Announce Type: replace-cross Abstract: Large language models internalize a structural trade-off between truthfulness and obsequious flattery, emerging from reward optimization that

CatalyticMLLM: A Graph-Text Multimodal Large Language Model for Catalytic Materials

SafetyDGX agent

arXiv:2605.17254v1 Announce Type: new Abstract: Property prediction and inverse structural design of catalytic materials are typically modeled as two independent tasks: the former predicts target prop

Factored Causal Representation Learning for Robust Reward Modeling in RLHF

SafetyDGX agent

arXiv:2601.21350v2 Announce Type: replace Abstract: A reliable reward model is essential for aligning large language models with human preferences through reinforcement learning from human feedback. H

Fine-tuning Pocket-Aware Diffusion Models via Denoising Policy Optimization

Model ReleasesDGX agent

arXiv:2605.17693v1 Announce Type: cross Abstract: Structure-based drug design has been accelerated by pocket-aware 3D generative models, yet most methods primarily fit the training distribution and ma

FormuLLA: A Large Language Model Approach to Generating Novel 3D Printable Formulations

Model ReleasesDGX agent

arXiv:2601.02071v3 Announce Type: replace Abstract: Pharmaceutical three-dimensional (3D) printing is an advanced fabrication technology with the potential to enable truly personalised dosage forms. R

GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets

Model ReleasesDGX agent

arXiv:2605.18475v1 Announce Type: cross Abstract: Mixed-precision quantization improves the budget--accuracy trade-off for large language models (LLMs) by allocating more bits to sensitive modules. Ho

HEED: Density-Weighted Residual Alignment for Hybrid Vision-Language Model Distillation

Model ReleasesDGX agent

arXiv:2605.17093v1 Announce Type: cross Abstract: Distilling vision-language models into faster hybrid architectures, such as 3:1 Mamba-2/attention mixes, is now standard practice for making inference

Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models

Model ReleasesDGX agent

arXiv:2602.02039v2 Announce Type: replace Abstract: The agency expected of Agentic Large Language Models goes beyond answering correctly, requiring autonomy to set goals and decide what to explore. We

Identifiable Token Correspondence for World Models

Model ReleasesDGX agent

arXiv:2605.16457v1 Announce Type: cross Abstract: Transformer-based world models have shown strong performance in visual reinforcement learning, but often suffer from temporal inconsistency in long-ho

KairosHope: A Next-Generation Time-Series Foundation Model for Specialized Classification via Dual-Memory Architecture

Model ReleasesDGX agent

arXiv:2605.18657v1 Announce Type: cross Abstract: Time Series Foundation Models (TSFMs) have demonstrated notable success in general-purpose forecasting tasks; however, their adaptation to specialized

LAST-RAG: Literature-Anchored Stochastic Trajectory Retrieval-Augmented Generation for Knowledge-Conditioned Degradation Model Selection

Local AiDGX agent

arXiv:2605.17902v1 Announce Type: new Abstract: Stochastic-process-based degradation modeling is a core approach for estimating the distribution of remaining useful life (RUL); however, the selection

Merlin's Whisper: Enabling Efficient Reasoning in Large Language Models via Black-box Persuasive Prompting

Model ReleasesDGX agent

arXiv:2510.10528v3 Announce Type: replace Abstract: Large reasoning models (LRMs) have demonstrated remarkable proficiency in tackling complex tasks through step-by-step thinking. However, this length

Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2605.16409v1 Announce Type: cross Abstract: Optical character recognition (OCR) and multilingual text understanding remain major failure modes of multimodal large language models (MLLMs), partic

Seeing Together:Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2605.18431v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have made substantial progress in egocentric video understanding, but their ability to reason cooperatively fro

Statistical Hand Shape Modeling from Clinical CT Scans Using Deep Learning and Implicit Skinning

ResearchDGX agent

arXiv:2605.16980v1 Announce Type: new Abstract: Accurate segmentation and statistical shape modeling of hand anatomy have significant implications for medical diagnostics, ergonomics, and biomechanics

Towards Long-Lived Robots: Continual Learning VLA Models via Reinforcement Fine-Tuning

Model ReleasesDGX agent

arXiv:2602.10503v2 Announce Type: replace Abstract: Pretrained on large-scale and diverse datasets, VLA models demonstrate strong generalization and adaptability as general-purpose robotic policies. H

VideoNeuMat: Neural Material Extraction from Generative Video Models

ResearchDGX agent

arXiv:2602.07272v2 Announce Type: replace Abstract: Creating photorealistic materials for 3D rendering requires exceptional artistic skill. Generative models for materials could help, but are currentl

Why Do Reasoning Models Lose Coverage? The Role of Data and Forks in the Road

ResearchDGX agent

arXiv:2605.17026v1 Announce Type: new Abstract: Recent progress in large language models has led to the emergence of reasoning models, which have shown strong performance on complex tasks through spec

18 May 2026

Evaluating Chinese Ambiguity Understanding in Large Language Models

Model ReleasesDGX agent

arXiv:2605.15635v1 Announce Type: new Abstract: Linguistic ambiguity is critical to the robustness of Large Language Models (LLMs), yet existing research focuses mostly on English, with limited attent

f-Trajectory Balance: A Loss Family for Tuning GFlowNets, Generative Models, and LLMs with Off- and On-Policy Data

SafetyDGX agent

arXiv:2605.15417v1 Announce Type: cross Abstract: In GFlowNets and variational inference, it has been shown that the mean square error between target and model log probabilities is an effective, low v

FINESSE-Bench: A Hierarchical Benchmark Suite for Financial Domain Knowledge and Technical Analysis in Large Language Models

Model ReleasesDGX agent

arXiv:2605.15482v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being applied to financial analysis, reporting, investment decision support, risk management, compliance,

Frontier Large Language Models Rival State-of-the-Art Planners

Model ReleasesDGX agent

arXiv:2511.09378v2 Announce Type: replace Abstract: A series of influential studies established that large language models cannot reliably solve even simple planning tasks. We show that the latest gen

Learning Normalized Energy Models for Linear Inverse Problems

ResearchDGX agent

arXiv:2605.15487v1 Announce Type: cross Abstract: Generative diffusion models can provide powerful prior probability models for inverse problems in imaging, but existing implementations suffer from tw

MHGraphBench: Knowledge Graph-Grounded Benchmarking of Mental Health Knowledge in Large Language Models

Model ReleasesDGX agent

arXiv:2605.15589v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in the mental health domain, yet it remains unclear how well they capture related biomedical knowledg

Social-Mamba: Socially-Aware Trajectory Forecasting with State-Space Models

Model ReleasesDGX agent

arXiv:2605.15424v1 Announce Type: new Abstract: Human trajectory forecasting is crucial for safe navigation in crowded environments, requiring models that balance accuracy with computational efficienc

Time-Varying Deep State Space Models for Sequences with Switching Dynamics

ResearchDGX agent

arXiv:2605.15311v1 Announce Type: new Abstract: The identification and modeling of time-varying systems is a fundamental challenge in signal processing and system identification. To address this chall

Zero-Shot Goal Recognition with Large Language Models

Model ReleasesDGX agent

arXiv:2605.15333v1 Announce Type: new Abstract: Large language models have recently reached near-parity with classical planners on well-known planning domains, yet this competence relies on world-know

15 May 2026

Eradicating Negative Transfer in Multi-Physics Foundation Models via Sparse Mixture-of-Experts Routing

Model ReleasesDGX agent

arXiv:2605.15179v1 Announce Type: cross Abstract: Scaling Scientific Machine Learning (SciML) toward universal foundation models is bottlenecked by negative transfer: the simultaneous co-training of d

Finding Interpretable Prompt-Specific Circuits in Language Models

ApplicationsDGX agent

arXiv:2602.13483v2 Announce Type: replace-cross Abstract: Understanding the internal circuits that language models use to solve tasks remains a central challenge in mechanistic interpretability. A cru

MechVerse: Evaluating Physical Motion Consistency in Video Generation Models

Model ReleasesDGX agent

arXiv:2605.14843v1 Announce Type: new Abstract: Text- and image-conditioned video generation models have achieved strong visual fidelity and temporal coherence, but they often fail to generate motion

VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity Language Model with Curriculum Learning and Native Tool Use

Model ReleasesDGX agent

arXiv:2605.13989v1 Announce Type: new Abstract: We present VectraYX-Nano, a 41.95M-parameter decoder-only language model trained from scratch in Spanish for cybersecurity, with a Latin-American focus

14 May 2026

AI Harness Engineering: A Runtime Substrate for Foundation-Model Software Agents

AgentsDGX agent

arXiv:2605.13357v1 Announce Type: cross Abstract: Foundation models have transformed automated code generation, yet autonomous software-engineering agents remain unreliable in realistic development se

Asymmetric Flow Models

ResearchDGX agent

arXiv:2605.12964v1 Announce Type: new Abstract: Flow-based generation in high-dimensional spaces is difficult because velocity prediction requires modeling high-dimensional noise, even when data has s

haha, it's a dramatic voice model

AgentsDGX agent

haha, it's a dramatic voice model We're releasing a whole new category of voice models. Introducing DramaBox — our state-of-the-art, open source voice model built for cinematic use cases. Traditional

(How) Do Large Language Models Understand High-Level Message Sequence Charts?

Model ReleasesDGX agent

arXiv:2605.13773v1 Announce Type: cross Abstract: Large Language Models (LLMs) are being employed widely to automate tasks across the software development life-cycle. It is, however, unclear whether t

How Well Do Large-Scale Chemical Language Models Transfer to Downstream Tasks?

Model ReleasesDGX agent

arXiv:2602.11618v4 Announce Type: replace Abstract: Chemical Language Models (CLMs) pre-trained on large scale molecular data are widely used for molecular property prediction. However, the common bel

Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling

SafetyDGX agent

arXiv:2605.13801v1 Announce Type: cross Abstract: As generative AI models such as large language models (LLMs) become more pervasive, ensuring the safety, robustness, and overall trustworthiness of th

Inducing Overthink: Hierarchical Genetic Algorithm-based DoS Attack on Black-Box Large Language Reasoning Models

Model ReleasesDGX agent

arXiv:2605.13338v1 Announce Type: cross Abstract: Large Reasoning Models (LRMs) are increasingly integrated into systems requiring reliable multi-step inference, yet this growing dependence exposes ne

Integration of an Agent Model into an Open Simulation Architecture for Scenario-Based Testing of Automated Vehicles

SafetyDGX agent

arXiv:2605.13539v1 Announce Type: new Abstract: Simulative and scenario-based testing are crucial methods in the safety assurance for automated driving systems. To ensure that simulation results are r

Large Language Models Lack Temporal Awareness of Medical Knowledge

Model ReleasesDGX agent

arXiv:2605.13045v1 Announce Type: new Abstract: The existing methods for evaluating the medical knowledge of Large Language Models (LLMs) are largely based on atemporal examination-style benchmarks, w

LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models

Model ReleasesDGX agent

arXiv:2505.15616v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have achieved significant advances in integrating visual and linguistic information, yet their ability to r

Probing Persona-Dependent Preferences in Language Models

Model ReleasesDGX agent

arXiv:2605.13339v1 Announce Type: cross Abstract: Large language models (LLMs) can be said to have preferences: they reliably pick certain tasks and outputs over others, and preferences shaped by post

SupChain-Bench: Benchmarking Large Language Models for Real-World Supply Chain Management

Model ReleasesDGX agent

arXiv:2602.07342v2 Announce Type: replace Abstract: Large language models (LLMs) have shown promise in complex reasoning and tool-based decision making, motivating their application to real-world supp

The critical slowing down in diffusion models

Model ReleasesDGX agent

arXiv:2605.12597v1 Announce Type: cross Abstract: Computational sampling has been central to the sciences since the mid-20th century. While machine-learning-based approaches have recently enabled majo

The Efficiency Gap in Byte Modeling

ResearchDGX agent

arXiv:2605.12928v1 Announce Type: new Abstract: Modern language models have historically relied on two dominant design choices: subword tokenization and autoregressive (AR) ordering. These design deci

13 May 2026

AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model

SafetyDGX agent

arXiv:2511.22663v5 Announce Type: replace Abstract: Unified multimodal models for image generation and understanding represent a significant step toward AGI and have attracted widespread attention fro

Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training

Model ReleasesDGX agent

arXiv:2605.12483v1 Announce Type: new Abstract: In settings where labeled verifiable training data is the binding constraint, each checked example should be allocated carefully. The standard practice

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation

Model ReleasesDGX agent

arXiv:2605.12034v1 Announce Type: cross Abstract: Omni-modal language models are intended to jointly understand audio, visual inputs, and language, but benchmark gains can be inflated when visual evid

CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating

Local AiDGX agent

arXiv:2605.11723v1 Announce Type: new Abstract: In this paper, we propose Concentrate and Concentrate (CaC), a coarse-to-fine anomaly reward model based on Vision-Language Models. During inference, it

DarkQA: Benchmarking Vision-Language Models on Visual-Primitive Question Answering in Low-Light Indoor Scenes

Model ReleasesDGX agent

arXiv:2512.24985v4 Announce Type: replace Abstract: Vision Language Models (VLMs) are increasingly adopted as central reasoning modules for embodied agents. Existing benchmarks evaluate their capabili

← Previous
1…6869707172…1007
Next →