AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
48,543 results
Model Releases

Bridging Foundation Models and ASTM Metallurgical Standards for Automated Grain Size Estimation from Microscopy Images

DGX agent

arXiv:2604.18957v1 Announce Type: new Abstract: Extracting standardized metallurgical metrics from microscopy images remains challenging due to complex grain morphology and the data demands of supervi

model-releasesarxiv-cs-cv
22 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

MedPRMBench: A Fine-grained Benchmark for Process Reward Models in Medical Reasoning

DGX agent

arXiv:2604.17282v1 Announce Type: new Abstract: Process-Level Reward Models (PRMs) are essential for guiding complex reasoning in large language models, yet existing PRM benchmarks cover only general

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning

DGX agent

arXiv:2503.03480v4 Announce Type: replace Abstract: Vision-language-action models (VLAs) show potential as generalist robot policies. However, these models pose extreme safety challenges during real-w

model-releasesarxiv-cs-ro
21 Apr 2026
Applications

Thermal-GEMs: Generalized Models for Building Thermal Dynamics

DGX agent

arXiv:2604.16443v1 Announce Type: cross Abstract: Data-driven models for building thermal dynamics are a scalable approach for enabling energy-efficient operation through fault detection & diagnosis o

applicationsarxiv-cs-lg
21 Apr 2026
Model Releases

Using large language models for embodied planning introduces systematic safety risks

DGX agent

arXiv:2604.18463v1 Announce Type: cross Abstract: Large language models are increasingly used as planners for robotic systems, yet how safely they plan remains an open question. To evaluate safe plann

model-releasesarxiv-cs-lg
21 Apr 2026
Research

Think Multilingual, Not Harder: A Data-Efficient Framework for Teaching Reasoning Models to Code-Switch

DGX agent

arXiv:2604.15490v1 Announce Type: new Abstract: Recent developments in reasoning capabilities have enabled large language models to solve increasingly complex mathematical, symbolic, and logical tasks

researcharxiv-cs-cl
20 Apr 2026
Model Releases

Parameter estimation for land-surface models using Neural Physics

DGX agent

arXiv:2505.02979v3 Announce Type: replace-cross Abstract: We propose a novel inverse-modelling approach which estimates the parameters of a simple land-surface model (LSM) by assimilating data into a

model-releasesarxiv-cs-lg
17 Apr 2026
Model Releases

Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets?

DGX agent

arXiv:2511.17792v2 Announce Type: replace Abstract: While recent video world models can generate highly realistic videos, their ability to perform semantic reasoning and planning remains unclear and u

model-releasesarxiv-cs-cv
16 Apr 2026
Model Releases

Benchmarking Deflection and Hallucination in Large Vision-Language Models

DGX agent

arXiv:2604.12033v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) increasingly rely on retrieval to answer knowledge-intensive multimodal questions. Existing benchmarks overlook c

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

AIM: Intent-Aware Unified world action Modeling with Spatial Value Maps

DGX agent

arXiv:2604.11135v1 Announce Type: cross Abstract: Pretrained video generation models provide strong priors for robot control, but existing unified world action models still struggle to decode reliable

model-releasesarxiv-cs-lg
14 Apr 2026
Agents

Automating Structural Analysis Across Multiple Software Platforms Using Large Language Models

DGX agent

arXiv:2604.09866v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have shown the promise to significantly accelerate the workflow by automating structural modeling and

agentsarxiv-cs-ai
14 Apr 2026
Research

Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning

DGX agent

arXiv:2604.10701v1 Announce Type: cross Abstract: Credit assignment is a central challenge in reinforcement learning (RL). Classical actor-critic methods address this challenge through fine-grained ad

researcharxiv-cs-ai
14 Apr 2026
Research

Lost in Diffusion: Uncovering Hallucination Patterns and Failure Modes in Diffusion Large Language Models

DGX agent

arXiv:2604.10556v1 Announce Type: new Abstract: While Diffusion Large Language Models (dLLMs) have emerged as a promising non-autoregressive paradigm comparable to autoregressive (AR) models, their fa

researcharxiv-cs-cl
14 Apr 2026
Model Releases

Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration

DGX agent

arXiv:2604.11446v1 Announce Type: cross Abstract: Recently, scaling reinforcement learning with verifiable rewards (RLVR) for large language models (LLMs) has emerged as an effective training paradigm

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

Pando: Do Interpretability Methods Work When Models Won't Explain Themselves?

DGX agent

arXiv:2604.11061v1 Announce Type: cross Abstract: Mechanistic interpretability is often motivated for alignment auditing, where a model's verbal explanations can be absent, incomplete, or misleading.

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

You Only Judge Once: Multi-response Reward Modeling in a Single Forward Pass

DGX agent

arXiv:2604.10966v1 Announce Type: cross Abstract: We present a discriminative multimodal reward model that scores all candidate responses in a single forward pass. Conventional discriminative reward m

model-releasesarxiv-cs-ai
14 Apr 2026
Safety

RAMP: Hybrid DRL for Online Learning of Numeric Action Models

DGX agent

arXiv:2604.08685v1 Announce Type: new Abstract: Automated planning algorithms require an action model specifying the preconditions and effects of each action, but obtaining such a model is often hard.

safetyarxiv-cs-ai
13 Apr 2026
Model Releases

AE-ViT: Stable Long-Horizon Parametric Partial Differential Equations Modeling

DGX agent

arXiv:2604.06475v1 Announce Type: new Abstract: Deep Learning Reduced Order Models (ROMs) are becoming increasingly popular as surrogate models for parametric partial differential equations (PDEs) due

model-releasesarxiv-cs-lg
10 Apr 2026
Model Releases

Before We Trust Them: Decision-Making Failures in Navigation of Foundation Models

DGX agent

arXiv:2601.05529v5 Announce Type: replace Abstract: High success rates on navigation-related tasks do not necessarily translate into reliable decision making by foundation models. To examine this gap,

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

Temporally Phenotyping GLP-1RA Case Reports with Large Language Models: A Textual Time Series Corpus and Risk Modeling

DGX agent

arXiv:2604.06197v1 Announce Type: cross Abstract: Type 2 diabetes case reports describe complex clinical courses, but their timelines are often expressed in language that is difficult to reuse in long

model-releasesarxiv-cs-ai
10 Apr 2026
Safety

The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment

DGX agent

arXiv:2604.06377v1 Announce Type: cross Abstract: We investigate whether post-trained capabilities can be transferred across models without retraining, with a focus on transfer across different model

safetyarxiv-cs-ai
10 Apr 2026
Model Releases

Which Way Does Time Flow? A Psychophysics-Grounded Evaluation for Vision-Language Models

DGX agent

arXiv:2510.26241v5 Announce Type: replace-cross Abstract: Modern vision-language models (VLMs) excel at many multimodal tasks, yet their grasp of temporal information in video remains weak and has not

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

Mapping and Measuring the Behavioral Evolution of Large Language Models

DGX agent

arXiv:2608.11027v1 Announce Type: cross Abstract: Benchmark leaderboards summarize how well a language model performs, but not how its behavior relates to that of other models or changes across genera

model-releasesarxiv-cs-cl
12 Aug 2026
Model Releases

Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation

DGX agent

arXiv:2608.10812v1 Announce Type: cross Abstract: We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiL

model-releasesarxiv-cs-ai
12 Aug 2026
Research

A Hybrid Neural-Microfacet BRDF Model for Real-Time Rendering

DGX agent

arXiv:2608.09604v1 Announce Type: cross Abstract: Over the past decade, microfacet-based BRDF models have formed the foundation of real-time rendering pipelines. Despite their widespread use, they oft

researcharxiv-cs-cv
11 Aug 2026
Model Releases

ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models

DGX agent

arXiv:2608.09548v1 Announce Type: cross Abstract: Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that or

model-releasesarxiv-cs-ai
11 Aug 2026
Research

HugSelect: An Explainable Multi-Criteria Decision-Support Framework for foundation-model selection

DGX agent

arXiv:2608.08069v1 Announce Type: cross Abstract: Foundation models are increasingly reused as software components, making model selection a critical software-engineering decision. Current model hubs

researcharxiv-cs-ai
11 Aug 2026
Model Releases

Addressable Memory for Video World Models

DGX agent

arXiv:2608.07408v1 Announce Type: new Abstract: We study visual persistence in interactive video world models. These models rely on a Key-Value (KV) cache as a growing visual memory to carry forward p

model-releasesarxiv-cs-cv
10 Aug 2026
Model Releases

Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference Elicitation

DGX agent

arXiv:2608.06955v1 Announce Type: new Abstract: Large language models (LLMs) are trained on corpora that contain expressions of human judgment about films, books, music, and more. Yet whether LLMs sys

model-releasesarxiv-cs-ai
10 Aug 2026
Agents

Probing Visual Concepts in Lightweight Vision-Language Models for Automated Driving

DGX agent

arXiv:2603.06054v2 Announce Type: replace-cross Abstract: The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their r

agentsarxiv-cs-ai
10 Aug 2026
Safety

The Fairness Collapse Phenomenon: Bias Amplification in Language Models Trained on Synthetic Data

DGX agent

arXiv:2608.04268v1 Announce Type: new Abstract: Generative models trained on artificially generated data have been shown to exhibit model collapse, resulting in significant performance degradation. As

safetyarxiv-cs-cl
6 Aug 2026
Applications

Conformal risk control for model-form uncertainty in parametric non-intrusive reduced-order models

DGX agent

arXiv:2608.03360v1 Announce Type: cross Abstract: Non-intrusive reduced-order models (NIROMs) have become a standard tool for approximating parametric partial differential equations from computer desi

applicationsarxiv-cs-lg
5 Aug 2026
Model Releases

A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples

DGX agent

arXiv:2607.29122v1 Announce Type: new Abstract: Pixel-space diffusion models aim to learn an end-to-end generator directly over raw pixels. This is challenging because a single model must capture both

model-releasesarxiv-cs-cv
3 Aug 2026
Model Releases

Beyond the Bidirectional Promise: Re-evaluating the Robustness of Diffusion Language Models

DGX agent

arXiv:2607.27386v1 Announce Type: cross Abstract: Diffusion Language Models (DLMs) offer a compelling alternative to autoregressive (AR) generation by enabling bidirectional context and iterative refi

model-releasesarxiv-cs-lg
31 Jul 2026
Research

Predict before you train: Scaling Laws for particle physics foundation models

DGX agent

arXiv:2607.23377v1 Announce Type: cross Abstract: The largest machine learning models in particle physics are also the most expensive to train, yet the return on scaling a given architecture cannot be

researcharxiv-cs-ai
31 Jul 2026
Model Releases

RepBench: Compiling Benchmarks into Capability Representations for Large Language Models

DGX agent

arXiv:2607.28008v1 Announce Type: new Abstract: Representation engineering reads and steers capability directions in large language models, yet methods are typically evaluated on paper-specific synthe

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

Same Facts, Different Diagnosis: Measuring and Mitigating Narrative Anchoring in Clinical Language Models

DGX agent

arXiv:2607.27384v1 Announce Type: new Abstract: Large language models used for clinical diagnostic reasoning are sensitive to sociolinguistic register, not just clinical content. We term this failure

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

Language Models are not Equally Robust to Non-Canonical Tokenization across Languages

DGX agent

arXiv:2607.26831v1 Announce Type: new Abstract: Despite the existence of exponentially many valid tokenizations for a given string, language models operate on a single canonical sequence deterministic

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Laplace-PSN-IRT: Uncertainty Quantification for Neural Item Response Theory Models of LLM Benchmarks

DGX agent

arXiv:2607.25257v1 Announce Type: cross Abstract: Item Response Theory (IRT) has recently been proposed as a framework for evaluating large language model (LLM) benchmarks by separating a model's late

model-releasesarxiv-cs-ai
29 Jul 2026
Model Releases

WALoMA: A Multitask Wireless Foundation Model via Adaptive Low-Rank Masked Autoencoders

DGX agent

arXiv:2607.25763v1 Announce Type: cross Abstract: This paper proposes a multitask wireless foundation model via adaptive low-rank masked autoencoders (WALoMA), a unified multi-task foundation model fo

model-releasesarxiv-cs-lg
29 Jul 2026
Model Releases

AIR-BENCH Live: An Evolving Safety Benchmark for Foundation Models

DGX agent

arXiv:2607.22671v1 Announce Type: new Abstract: Foundation-model safety benchmarks capture the AI risks of their time of publication: as models improve and governments pass new AI-safety legislation,

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Numerical Investigation of Sequence Modeling Theory using Controllable Memory Functions

DGX agent

arXiv:2506.05678v3 Announce Type: replace Abstract: The evolution of sequence modeling architectures, from recurrent neural networks and convolutional models to Transformers and structured state-space

model-releasesarxiv-cs-lg
28 Jul 2026
Research

Reverso: Efficient Time Series Foundation Models for Zero-shot Forecasting

DGX agent

arXiv:2602.17634v2 Announce Type: replace-cross Abstract: Learning time series foundation models has been shown to be a promising approach for zero-shot time series forecasting across diverse time ser

researcharxiv-cs-ai
28 Jul 2026
Safety

RM-Distiller: Exploiting Generative LLM for Reward Model Distillation

DGX agent

arXiv:2601.14032v2 Announce Type: replace Abstract: Reward models (RMs) play a pivotal role in aligning large language models (LLMs) with human preferences. Due to the difficulty of obtaining high-qua

safetyarxiv-cs-cl
28 Jul 2026
Model Releases

MoE^2-LoRA: When MoE Models Meet MoE-style Low-Rank Adaptation

DGX agent

arXiv:2607.21978v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) architectures have been widely adopted in large language models, yet parameter-efficient fine-tuning (PEFT) for MoE models rema

model-releasesarxiv-cs-cl
27 Jul 2026
Model Releases

PhantomFill: When the Form Demands an Answer, Language Models Invent One

DGX agent

arXiv:2607.20492v1 Announce Type: cross Abstract: Language models in production do not write prose. They fill forms: JSON fields, function arguments, extraction templates. We show that the form itself

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Task Competence Is Not Instruction Following: Evaluating Instruction-Conflicting Behavior in Small Language Models

DGX agent

arXiv:2607.19608v1 Announce Type: new Abstract: Instruction tuning is meant to make language models follow user requests, yet it is unclear whether small models comply when an instruction conflicts wi

model-releasesarxiv-cs-cl
23 Jul 2026
Model Releases

Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation

DGX agent

arXiv:2509.24739v4 Announce Type: replace Abstract: Vision-Language Foundation Models (VLMs), trained on large-scale multimodal datasets, have driven significant advances in Artificial Intelligence (A

model-releasesarxiv-cs-cv
23 Jul 2026
← Previous
1…89101112…1012
Next →