AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,841 results
4 Jun 2026

Geometry-Aware Hallucination Detection in Large Language Models

Model ReleasesDGX agent

arXiv:2601.06196v3 Announce Type: replace-cross Abstract: Large language models (LLMs) frequently generate factually incorrect or unsupported content, commonly referred to as hallucinations. Prior wor

Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model

SafetyDGX agent

arXiv:2512.21917v3 Announce Type: replace-cross Abstract: Policy alignment to preference data typically assumes a known link function between observed preferences and latent rewards (e.g., Bradley-Ter

Token Rankings are Unforgeable Language Model Signatures

ResearchDGX agent

arXiv:2606.04459v1 Announce Type: cross Abstract: Language model parameters are known to impose unique (to each model) geometric constraints on their logit outputs, which serves as a signature that id

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation

ResearchDGX agent

arXiv:2606.04264v1 Announce Type: new Abstract: Recent years have seen remarkable progress in unified vision-language models handling both multimodal understanding and generation within a single archi

3 Jun 2026

ChatHealthAI: Aligning Electronic Health Record Representations with Large Language Models for Grounded Clinical Reasoning

Model ReleasesDGX agent

arXiv:2606.02802v1 Announce Type: new Abstract: Large language models (LLMs) exhibit strong natural-language reasoning abilities for clinical decision support, but struggle to effectively model struct

Hybrid Dynamics Modeling for a Flexible 2-DoF Robotic Arm

Model ReleasesDGX agent

arXiv:2606.02969v1 Announce Type: new Abstract: This paper examines three approaches for modeling the dynamics of a flexible-link 2-DoF robotic arm to address unmodeled dynamics not captured by rigid-

Large Byte Model: Teaching Language Models About Compiled Code

ResearchDGX agent

arXiv:2606.02834v1 Announce Type: cross Abstract: Malware analysis starts with the raw bytes of an executable program, and tools to 'lift' these to higher-level representations, such as assembly, are

2 Jun 2026

3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code

Model ReleasesDGX agent

arXiv:2606.01057v1 Announce Type: cross Abstract: Procedural 3D modeling through code is emerging as a versatile paradigm, offering deterministic, engine-ready, and precisely editable assets that neur

Large Electron Model: A Universal Ground State Predictor

Model ReleasesDGX agent

arXiv:2603.02346v2 Announce Type: replace-cross Abstract: We introduce Large Electron Model, a single neural network model that produces variational wavefunctions of interacting electrons over the ent

LL-Bench: Rethinking Low-Level Vision Evaluation in the Era of Large-Scale Generative Models

Model ReleasesDGX agent

arXiv:2606.02535v1 Announce Type: new Abstract: Large-scale generative models have demonstrated remarkable capabilities across image generation and editing tasks. However, their performance in low-lev

On the Generalization Gap in Self-Evolving Language Model Reasoning

Model ReleasesDGX agent

arXiv:2606.01075v1 Announce Type: new Abstract: Recent work suggests that large language models (LLMs) can improve through self-evolution (SE), using supervision signals generated by the model itself.

Saliency-Aware Model Merging

Model ReleasesDGX agent

arXiv:2606.00511v1 Announce Type: cross Abstract: Model merging aims to consolidate multiple task-specific models fine-tuned on different datasets into a unified architecture that performs cross-domai

Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models

Model ReleasesDGX agent

arXiv:2606.00566v1 Announce Type: cross Abstract: As language models take on agentic roles that span calling external APIs, reading tool outputs, and acting on instructions embedded in third-party con

Scaling depth capacity via zero/one-layer model expansion

ResearchDGX agent

arXiv:2511.04981v2 Announce Type: replace Abstract: Model depth is a double-edged sword in deep learning: deeper models achieve higher accuracy but require higher computational cost. To efficiently tr

VLAMotor: Test-Guided Enhancement of Vision-Language-Action Models via Agent-BasedData Synthesis

AgentsDGX agent

arXiv:2606.00053v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models follow a data-driven paradigm and are constrained by the coverage of training data, making them prone to failure on

1 Jun 2026

How Far Can You Grow? Characterizing the Extrapolation Frontier of Graph Generative Models for Materials Science

Model ReleasesDGX agent

arXiv:2602.09309v2 Announce Type: replace-cross Abstract: Every generative model for crystalline materials harbors a critical structure size beyond which its outputs become unreliable; we call this th

29 May 2026

Can AI Weather Models Predict Beyond Two Weeks? A Quantitative Benchmark and Analysis of Long Rollouts

Model ReleasesDGX agent

arXiv:2605.30184v1 Announce Type: new Abstract: While AI weather models excel at short-to-medium range forecasts (up to 15 days), they frequently suffer from ill-defined 'instabilities' when rolled ou

GPF-LiveNews: A Streaming Evaluation Protocol for Group-Conditioned Framing in Large Language Models

Model ReleasesDGX agent

arXiv:2605.28848v1 Announce Type: cross Abstract: Deployed language models are evaluated in a non-stationary environment: model versions, retrieval layers, safety systems, and real-world inputs all ch

Latent Performance Profiling of Large Language Models

Model ReleasesDGX agent

arXiv:2605.30018v1 Announce Type: new Abstract: Large language models (LLMs) frequently achieve impressive scores on standardized benchmarks, yet accuracy alone offers a limited view of their capabili

Production traffic from frontier models is a golden data asset. If you can efficiently mine the traces, filter for quality, and fine-tune sm…

AgentsDGX agent

Production traffic from frontier models is a golden data asset. If you can efficiently mine the traces, filter for quality, and fine-tune smaller models on them, you get specialized performance at a f

28 May 2026

A Sheaf-Theoretic and Topological Perspective on Complex Network Modeling and Attention Mechanisms in Graph Neural Models

ApplicationsDGX agent

arXiv:2601.21207v3 Announce Type: replace-cross Abstract: Combinatorial and topological structures, such as graphs, simplicial complexes, and cell complexes, form the foundation of geometric and topol

Accelerating Reinforcement Learning Training Using Simulation Surrogate Models

ResearchDGX agent

arXiv:2605.27556v1 Announce Type: cross Abstract: High-fidelity simulation models are widely used to analyze complex stochastic systems, but their high computational cost motivates the development of

Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting

Model ReleasesDGX agent

arXiv:2509.23074v3 Announce Type: replace-cross Abstract: In the era of increasingly complex AI models for time series forecasting, progress is often measured by marginal improvements on benchmark lea

DriveWAM: Video Generative Priors Enable Scalable World-Action Modeling for Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.28544v1 Announce Type: new Abstract: Pretrained foundation models have become an important basis for end-to-end autonomous driving. In contrast to vision-language models pretrained primaril

Models That Know How Evaluations Are Designed Score Safer

Model ReleasesDGX agent

arXiv:2605.28591v1 Announce Type: cross Abstract: The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings. Prior work has identified tes

27 May 2026

CNNs, Transformers, Hybrid, and Vision Language Models for Skin Cancer Detection

Model ReleasesDGX agent

arXiv:2605.26294v1 Announce Type: new Abstract: Skin cancer is a common and fast rising malignancy worldwide. Early detection is critical for improving outcomes. Deep learning models trained on dermos

ESMFold2 is a state of the art folding model. It's crazy good. The -Fast model does better on antibody-antigen complexes than AF3 with MSAs.

ResearchDGX agent

ESMFold2 is described as a high-performance protein structure prediction model that demonstrates particular strength in predicting antibody-antigen complex structures, reportedly outperforming AlphaFo

MATT-CTR: Unleashing a Model-Agnostic Test-Time Paradigm for CTR Prediction with Confidence-Guided Inference Paths

Model ReleasesDGX agent

arXiv:2510.08932v2 Announce Type: replace Abstract: Recently, a growing body of research has focused on either optimizing CTR model architectures to better model feature interactions or refining train

'PhyWorldBench': A Comprehensive Evaluation of Physical Realism in Text-to-Video Models

Model ReleasesDGX agent

arXiv:2507.13428v3 Announce Type: replace-cross Abstract: Video generation models have achieved remarkable progress in creating high-quality, photorealistic content. However, their ability to accurate

Stronger models do not always need lighter harnesses. Everyone believes more structured harnesses universally improve reliability, and that …

Model ReleasesDGX agent

Stronger models do not always need lighter harnesses. Everyone believes more structured harnesses universally improve reliability, and that higher-capability models need proportionally less structural

26 May 2026

A computational phase transition for learning-to-sample from Ising models

Model ReleasesDGX agent

arXiv:2605.24752v1 Announce Type: new Abstract: We study learning-to-sample -- a basic algorithmic task underlying generative modeling -- for Ising models, a standard testbed for algorithmic ideas in

From Simulation to Enaction: Post-trained language models recognize and react to their own generations

SafetyDGX agent

arXiv:2605.25459v1 Announce Type: cross Abstract: Language models are pretrained as passive predictors with no incentive to model the consequences of their own outputs. Post-training changes this: a m

JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures

Model ReleasesDGX agent

arXiv:2602.17162v2 Announce Type: replace Abstract: Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the 'Laws of Nature'. Whil

JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data

SafetyDGX agent

arXiv:2605.24414v1 Announce Type: new Abstract: We introduce JT-Safe-V2, a large language model designed to advance the safety and trustworthiness of foundation models, extending our previous JT-Safe

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization

Model ReleasesDGX agent

arXiv:2510.03827v2 Announce Type: replace-cross Abstract: LIBERO has emerged as a widely adopted benchmark for evaluating Vision-Language-Action (VLA) models; however, its current training and evaluat

Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models

Model ReleasesDGX agent

arXiv:2605.25394v1 Announce Type: new Abstract: Large language models often generate confident but incorrect answers rather than abstaining when uncertain. This problem is particularly acute for small

Security in the Fine-Tuning Lifecycle of Large Language Models: Threats, Defenses,Evaluation, and Future Directions

Model ReleasesDGX agent

arXiv:2605.25073v1 Announce Type: cross Abstract: Background: Fine-tuning is central to adapting pre-trained Large Language Models (LLMs) to downstream tasks, but its reliance on training data, parame

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models

AgentsDGX agent

arXiv:2605.26014v1 Announce Type: cross Abstract: Many video reasoning tasks require tracking motion, temporal order, and evolving visual states across frames. Existing methods built on large vision-l

23 May 2026

I-SAFE: Wasserstein Coherence Metrics for Structural Auditing of Scientific AI Models

Model ReleasesDGX agent

arXiv:2605.21731v1 Announce Type: new Abstract: Deep learning models are increasingly used in scientific prediction tasks where strong benchmark performance is often interpreted as evidence of scienti

22 May 2026

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles

Model ReleasesDGX agent

arXiv:2605.22177v1 Announce Type: cross Abstract: The proliferation of large language models (LLMs) and modular skills has endowed autonomous agents with increasingly powerful capabilities. Existing f

Pre-VLA: Preemptive Runtime Verification for Reliable Vision-Language-Action and World-Model Rollouts

Model ReleasesDGX agent

arXiv:2605.22446v1 Announce Type: new Abstract: While large vision-language-action (VLA) models and generative world models (WM) have advanced long-horizon embodied intelligence, their practical deplo

Understanding Data Temporality Impact on Large Language Models Pre-training

Model ReleasesDGX agent

arXiv:2605.22769v1 Announce Type: new Abstract: Large language models (LLMs) are typically trained on shuffled corpora, yielding models whose knowledge is frozen at train time and whose temporal groun

21 May 2026

Agentic Physical AI toward a Domain-Specific Foundation Model for Nuclear Reactor Control

Model ReleasesDGX agent

arXiv:2512.23292v3 Announce Type: replace-cross Abstract: The prevailing paradigm in AI for physical systems (scaling general-purpose foundation models toward universal multimodal reasoning) confronts

Calibration vs Decision Making: Revisiting the Reliability Paradox in Unlearned Language Models

Model ReleasesDGX agent

arXiv:2605.20915v1 Announce Type: new Abstract: Machine unlearning aims to remove the influence of specific training data from a model while preserving reliable behavior on the remaining data, making

Diffusion Models Memorize in Training -- and Generalize in Inference

Local AiDGX agent

arXiv:2603.13419v2 Announce Type: replace Abstract: Diffusion models generalize well in practice. However, an optimal diffusion model fully memorizes the training data and therefore fails to generaliz

DiMextsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging

Model ReleasesDGX agent

arXiv:2605.12960v2 Announce Type: replace Abstract: Towards more general and human-like intelligence, large language models should seamlessly integrate both multilingual and multimodal capabilities; h

Improving Quantized Model Performance in Qualitative Analysis with Multi-Pass Prompt Verification

Model ReleasesDGX agent

arXiv:2605.20193v1 Announce Type: new Abstract: Quantized Large Language Models (LLMs) are used more often in qualitative analysis because they run fast and need fewer computing resources. This study

20 May 2026

EVA-0: Test-Time Model Evolution with Only Two Forward Passes per Sample

Model ReleasesDGX agent

arXiv:2605.18867v1 Announce Type: cross Abstract: Test-time model evolution offers a promising way for deployed models to improve from unlabeled test-time experience, yet most existing methods depend

HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models

Model ReleasesDGX agent

arXiv:2605.19341v1 Announce Type: cross Abstract: Hallucination remains a central failure mode of large language models, but existing benchmarks operationalize it inconsistently across summarization,

Lying Is Just a Phase: The Hidden Alignment Transition in Language Model Scaling

Model ReleasesDGX agent

arXiv:2605.18838v1 Announce Type: cross Abstract: Scaling laws predict loss from compute but not how capabilities interact. We measure the coupling between reasoning and truthfulness across 63 base mo

Rebalancing Reference Frame Dominance to Improve Motion in Image-to-Video Models

Model ReleasesDGX agent

arXiv:2605.19398v1 Announce Type: cross Abstract: Image-to-video models often generate videos that remain overly static, compared to text-to-video models. While prior approaches mitigate this issue by

Unified Deployment-Aware Evaluation of Open Reasoning Language Models

Model ReleasesDGX agent

arXiv:2604.07035v2 Announce Type: replace Abstract: Open reasoning language models are often compared under mixed sample sizes, partially standardized prompts, and accuracy-centered summaries, which m

19 May 2026

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment

SafetyDGX agent

arXiv:2605.17602v1 Announce Type: new Abstract: Aligning Text-to-Image (T2I) generation models with human preferences increasingly relies on image reward models that score or rank generated images acc

Generalization or Memorization? Brittleness Testing for Chess-Trained Language Models

Model ReleasesDGX agent

arXiv:2605.17565v1 Announce Type: new Abstract: Recent work has fine-tuned language models on chess data and reported high benchmark scores as evidence that the resulting models can understand the rul

HydroAgent: Closing the Gap Between Frontier LLMs and Human Experts in Hydrologic Model Calibration via Simulator-Grounded RL

Model ReleasesDGX agent

arXiv:2605.17792v1 Announce Type: new Abstract: Calibrating distributed hydrologic models is a critical bottleneck across operational water resources management - streamflow prediction, reservoir oper

Internalizing Tool Knowledge in Small Language Models via QLoRA Fine-Tuning

Model ReleasesDGX agent

arXiv:2605.17774v1 Announce Type: new Abstract: Large language models are increasingly used as planning components in agentic systems, but current tool-use pipelines often require full tool schemas to

KVCapsule: Efficient Sequential KV Cache Compression for Vision-Language Models with Asymmetric Redundancy

Model ReleasesDGX agent

arXiv:2605.16439v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have emerged as a critical and fast-growing extension of Large Language Models (LLMs) that enable multimodal reasoning t

Language-Switching Triggers Take a Latent Detour Through Language Models

Model ReleasesDGX agent

arXiv:2605.18646v1 Announce Type: new Abstract: Backdoor attacks on language models pose a growing security concern, yet the internal mechanisms by which a trigger sequence hijacks model computations

SIPO: Stabilized and Improved Preference Optimization for Aligning Diffusion Models

Model ReleasesDGX agent

arXiv:2505.21893v3 Announce Type: replace-cross Abstract: Preference learning has garnered extensive attention as an effective technique for aligning diffusion models with human preferences in visual

The Range Shrinks, the Threat Remains: Re-evaluating LLM Package Hallucinations on the 2026 Frontier-Model Cohort

Model ReleasesDGX agent

arXiv:2605.17062v1 Announce Type: cross Abstract: Spracklen et al. (USENIX Security '25) showed that code-generating large language models hallucinate package names that do not exist on PyPI or npm at

← Previous
1…2223242526…998
Next →