AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,859 results
Model Releases

Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models

DGX agent

arXiv:2606.00566v1 Announce Type: cross Abstract: As language models take on agentic roles that span calling external APIs, reading tool outputs, and acting on instructions embedded in third-party con

model-releasesarxiv-cs-cl
2 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Scaling depth capacity via zero/one-layer model expansion

DGX agent

arXiv:2511.04981v2 Announce Type: replace Abstract: Model depth is a double-edged sword in deep learning: deeper models achieve higher accuracy but require higher computational cost. To efficiently tr

researcharxiv-cs-lg
2 Jun 2026
Agents

VLAMotor: Test-Guided Enhancement of Vision-Language-Action Models via Agent-BasedData Synthesis

DGX agent

arXiv:2606.00053v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models follow a data-driven paradigm and are constrained by the coverage of training data, making them prone to failure on

agentsarxiv-cs-ro
2 Jun 2026
Model Releases

How Far Can You Grow? Characterizing the Extrapolation Frontier of Graph Generative Models for Materials Science

DGX agent

arXiv:2602.09309v2 Announce Type: replace-cross Abstract: Every generative model for crystalline materials harbors a critical structure size beyond which its outputs become unreliable; we call this th

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

Can AI Weather Models Predict Beyond Two Weeks? A Quantitative Benchmark and Analysis of Long Rollouts

DGX agent

arXiv:2605.30184v1 Announce Type: new Abstract: While AI weather models excel at short-to-medium range forecasts (up to 15 days), they frequently suffer from ill-defined 'instabilities' when rolled ou

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

GPF-LiveNews: A Streaming Evaluation Protocol for Group-Conditioned Framing in Large Language Models

DGX agent

arXiv:2605.28848v1 Announce Type: cross Abstract: Deployed language models are evaluated in a non-stationary environment: model versions, retrieval layers, safety systems, and real-world inputs all ch

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Latent Performance Profiling of Large Language Models

DGX agent

arXiv:2605.30018v1 Announce Type: new Abstract: Large language models (LLMs) frequently achieve impressive scores on standardized benchmarks, yet accuracy alone offers a limited view of their capabili

model-releasesarxiv-cs-cl
29 May 2026
Agents

Production traffic from frontier models is a golden data asset. If you can efficiently mine the traces, filter for quality, and fine-tune sm…

DGX agent

Production traffic from frontier models is a golden data asset. If you can efficiently mine the traces, filter for quality, and fine-tune smaller models on them, you get specialized performance at a f

agentsharrison-chase--x
29 May 2026
Applications

A Sheaf-Theoretic and Topological Perspective on Complex Network Modeling and Attention Mechanisms in Graph Neural Models

DGX agent

arXiv:2601.21207v3 Announce Type: replace-cross Abstract: Combinatorial and topological structures, such as graphs, simplicial complexes, and cell complexes, form the foundation of geometric and topol

applicationsarxiv-cs-ai
28 May 2026
Research

Accelerating Reinforcement Learning Training Using Simulation Surrogate Models

DGX agent

arXiv:2605.27556v1 Announce Type: cross Abstract: High-fidelity simulation models are widely used to analyze complex stochastic systems, but their high computational cost motivates the development of

researcharxiv-cs-lg
28 May 2026
Model Releases

Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting

DGX agent

arXiv:2509.23074v3 Announce Type: replace-cross Abstract: In the era of increasingly complex AI models for time series forecasting, progress is often measured by marginal improvements on benchmark lea

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

DriveWAM: Video Generative Priors Enable Scalable World-Action Modeling for Autonomous Driving

DGX agent

arXiv:2605.28544v1 Announce Type: new Abstract: Pretrained foundation models have become an important basis for end-to-end autonomous driving. In contrast to vision-language models pretrained primaril

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Models That Know How Evaluations Are Designed Score Safer

DGX agent

arXiv:2605.28591v1 Announce Type: cross Abstract: The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings. Prior work has identified tes

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

CNNs, Transformers, Hybrid, and Vision Language Models for Skin Cancer Detection

DGX agent

arXiv:2605.26294v1 Announce Type: new Abstract: Skin cancer is a common and fast rising malignancy worldwide. Early detection is critical for improving outcomes. Deep learning models trained on dermos

model-releasesarxiv-cs-cv
27 May 2026
Research

ESMFold2 is a state of the art folding model. It's crazy good. The -Fast model does better on antibody-antigen complexes than AF3 with MSAs.

DGX agent

ESMFold2 is described as a high-performance protein structure prediction model that demonstrates particular strength in predicting antibody-antigen complex structures, reportedly outperforming AlphaFo

researchyann-lecun--x
27 May 2026
Model Releases

MATT-CTR: Unleashing a Model-Agnostic Test-Time Paradigm for CTR Prediction with Confidence-Guided Inference Paths

DGX agent

arXiv:2510.08932v2 Announce Type: replace Abstract: Recently, a growing body of research has focused on either optimizing CTR model architectures to better model feature interactions or refining train

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

'PhyWorldBench': A Comprehensive Evaluation of Physical Realism in Text-to-Video Models

DGX agent

arXiv:2507.13428v3 Announce Type: replace-cross Abstract: Video generation models have achieved remarkable progress in creating high-quality, photorealistic content. However, their ability to accurate

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Stronger models do not always need lighter harnesses. Everyone believes more structured harnesses universally improve reliability, and that …

DGX agent

Stronger models do not always need lighter harnesses. Everyone believes more structured harnesses universally improve reliability, and that higher-capability models need proportionally less structural

model-releasesdair-ai--x
27 May 2026
Model Releases

A computational phase transition for learning-to-sample from Ising models

DGX agent

arXiv:2605.24752v1 Announce Type: new Abstract: We study learning-to-sample -- a basic algorithmic task underlying generative modeling -- for Ising models, a standard testbed for algorithmic ideas in

model-releasesarxiv-cs-lg
26 May 2026
Safety

From Simulation to Enaction: Post-trained language models recognize and react to their own generations

DGX agent

arXiv:2605.25459v1 Announce Type: cross Abstract: Language models are pretrained as passive predictors with no incentive to model the consequences of their own outputs. Post-training changes this: a m

safetyarxiv-cs-ai
26 May 2026
Model Releases

JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures

DGX agent

arXiv:2602.17162v2 Announce Type: replace Abstract: Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the 'Laws of Nature'. Whil

model-releasesarxiv-cs-ai
26 May 2026
Safety

JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data

DGX agent

arXiv:2605.24414v1 Announce Type: new Abstract: We introduce JT-Safe-V2, a large language model designed to advance the safety and trustworthiness of foundation models, extending our previous JT-Safe

safetyarxiv-cs-ai
26 May 2026
Model Releases

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization

DGX agent

arXiv:2510.03827v2 Announce Type: replace-cross Abstract: LIBERO has emerged as a widely adopted benchmark for evaluating Vision-Language-Action (VLA) models; however, its current training and evaluat

model-releasesarxiv-cs-ro
26 May 2026
Model Releases

Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models

DGX agent

arXiv:2605.25394v1 Announce Type: new Abstract: Large language models often generate confident but incorrect answers rather than abstaining when uncertain. This problem is particularly acute for small

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Security in the Fine-Tuning Lifecycle of Large Language Models: Threats, Defenses,Evaluation, and Future Directions

DGX agent

arXiv:2605.25073v1 Announce Type: cross Abstract: Background: Fine-tuning is central to adapting pre-trained Large Language Models (LLMs) to downstream tasks, but its reliance on training data, parame

model-releasesarxiv-cs-ai
26 May 2026
Agents

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models

DGX agent

arXiv:2605.26014v1 Announce Type: cross Abstract: Many video reasoning tasks require tracking motion, temporal order, and evolving visual states across frames. Existing methods built on large vision-l

agentsarxiv-cs-cl
26 May 2026
Model Releases

I-SAFE: Wasserstein Coherence Metrics for Structural Auditing of Scientific AI Models

DGX agent

arXiv:2605.21731v1 Announce Type: new Abstract: Deep learning models are increasingly used in scientific prediction tasks where strong benchmark performance is often interpreted as evidence of scienti

model-releasesarxiv-cs-lg
23 May 2026
Model Releases

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles

DGX agent

arXiv:2605.22177v1 Announce Type: cross Abstract: The proliferation of large language models (LLMs) and modular skills has endowed autonomous agents with increasingly powerful capabilities. Existing f

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Pre-VLA: Preemptive Runtime Verification for Reliable Vision-Language-Action and World-Model Rollouts

DGX agent

arXiv:2605.22446v1 Announce Type: new Abstract: While large vision-language-action (VLA) models and generative world models (WM) have advanced long-horizon embodied intelligence, their practical deplo

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Understanding Data Temporality Impact on Large Language Models Pre-training

DGX agent

arXiv:2605.22769v1 Announce Type: new Abstract: Large language models (LLMs) are typically trained on shuffled corpora, yielding models whose knowledge is frozen at train time and whose temporal groun

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Agentic Physical AI toward a Domain-Specific Foundation Model for Nuclear Reactor Control

DGX agent

arXiv:2512.23292v3 Announce Type: replace-cross Abstract: The prevailing paradigm in AI for physical systems (scaling general-purpose foundation models toward universal multimodal reasoning) confronts

model-releasesarxiv-cs-lg
21 May 2026
Model Releases

Calibration vs Decision Making: Revisiting the Reliability Paradox in Unlearned Language Models

DGX agent

arXiv:2605.20915v1 Announce Type: new Abstract: Machine unlearning aims to remove the influence of specific training data from a model while preserving reliable behavior on the remaining data, making

model-releasesarxiv-cs-cl
21 May 2026
Local Ai

Diffusion Models Memorize in Training -- and Generalize in Inference

DGX agent

arXiv:2603.13419v2 Announce Type: replace Abstract: Diffusion models generalize well in practice. However, an optimal diffusion model fully memorizes the training data and therefore fails to generaliz

local-aiarxiv-cs-lg
21 May 2026
Model Releases

DiMextsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging

DGX agent

arXiv:2605.12960v2 Announce Type: replace Abstract: Towards more general and human-like intelligence, large language models should seamlessly integrate both multilingual and multimodal capabilities; h

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

Improving Quantized Model Performance in Qualitative Analysis with Multi-Pass Prompt Verification

DGX agent

arXiv:2605.20193v1 Announce Type: new Abstract: Quantized Large Language Models (LLMs) are used more often in qualitative analysis because they run fast and need fewer computing resources. This study

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

EVA-0: Test-Time Model Evolution with Only Two Forward Passes per Sample

DGX agent

arXiv:2605.18867v1 Announce Type: cross Abstract: Test-time model evolution offers a promising way for deployed models to improve from unlabeled test-time experience, yet most existing methods depend

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models

DGX agent

arXiv:2605.19341v1 Announce Type: cross Abstract: Hallucination remains a central failure mode of large language models, but existing benchmarks operationalize it inconsistently across summarization,

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Lying Is Just a Phase: The Hidden Alignment Transition in Language Model Scaling

DGX agent

arXiv:2605.18838v1 Announce Type: cross Abstract: Scaling laws predict loss from compute but not how capabilities interact. We measure the coupling between reasoning and truthfulness across 63 base mo

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Rebalancing Reference Frame Dominance to Improve Motion in Image-to-Video Models

DGX agent

arXiv:2605.19398v1 Announce Type: cross Abstract: Image-to-video models often generate videos that remain overly static, compared to text-to-video models. While prior approaches mitigate this issue by

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Unified Deployment-Aware Evaluation of Open Reasoning Language Models

DGX agent

arXiv:2604.07035v2 Announce Type: replace Abstract: Open reasoning language models are often compared under mixed sample sizes, partially standardized prompts, and accuracy-centered summaries, which m

model-releasesarxiv-cs-cl
20 May 2026
Safety

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment

DGX agent

arXiv:2605.17602v1 Announce Type: new Abstract: Aligning Text-to-Image (T2I) generation models with human preferences increasingly relies on image reward models that score or rank generated images acc

safetyarxiv-cs-ai
19 May 2026
Model Releases

Generalization or Memorization? Brittleness Testing for Chess-Trained Language Models

DGX agent

arXiv:2605.17565v1 Announce Type: new Abstract: Recent work has fine-tuned language models on chess data and reported high benchmark scores as evidence that the resulting models can understand the rul

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

HydroAgent: Closing the Gap Between Frontier LLMs and Human Experts in Hydrologic Model Calibration via Simulator-Grounded RL

DGX agent

arXiv:2605.17792v1 Announce Type: new Abstract: Calibrating distributed hydrologic models is a critical bottleneck across operational water resources management - streamflow prediction, reservoir oper

model-releasesarxiv-cs-lg
19 May 2026
Model Releases

Internalizing Tool Knowledge in Small Language Models via QLoRA Fine-Tuning

DGX agent

arXiv:2605.17774v1 Announce Type: new Abstract: Large language models are increasingly used as planning components in agentic systems, but current tool-use pipelines often require full tool schemas to

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

KVCapsule: Efficient Sequential KV Cache Compression for Vision-Language Models with Asymmetric Redundancy

DGX agent

arXiv:2605.16439v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have emerged as a critical and fast-growing extension of Large Language Models (LLMs) that enable multimodal reasoning t

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Language-Switching Triggers Take a Latent Detour Through Language Models

DGX agent

arXiv:2605.18646v1 Announce Type: new Abstract: Backdoor attacks on language models pose a growing security concern, yet the internal mechanisms by which a trigger sequence hijacks model computations

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

SIPO: Stabilized and Improved Preference Optimization for Aligning Diffusion Models

DGX agent

arXiv:2505.21893v3 Announce Type: replace-cross Abstract: Preference learning has garnered extensive attention as an effective technique for aligning diffusion models with human preferences in visual

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

The Range Shrinks, the Threat Remains: Re-evaluating LLM Package Hallucinations on the 2026 Frontier-Model Cohort

DGX agent

arXiv:2605.17062v1 Announce Type: cross Abstract: Spracklen et al. (USENIX Security '25) showed that code-generating large language models hallucinate package names that do not exist on PyPI or npm at

model-releasesarxiv-cs-lg
19 May 2026
← Previous
1…2829303132…1248
Next →