AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,841 results
2 Jun 2026

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models

Model ReleasesDGX agent

arXiv:2606.00793v1 Announce Type: new Abstract: Recent advancements in video-based world models have demonstrated an unprecedented ability to synthesize high-fidelity visual sequences. However, a fund

Model Multiplicity and Predictive Arbitrariness in Recidivism Risk Assessment

SafetyDGX agent

arXiv:2606.02198v1 Announce Type: new Abstract: Prediction tasks over individual futures, which are inherently noisy, often admit multiple similarly accurate models. When these models produce differen

Rethinking Scientific Modeling: Toward Physically Consistent and Simulation-Executable Programmatic Generation

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2602.07083v2 Announce Type: replace-cross Abstract: Structural modeling is a fundamental component of computational engineering science, in which even minor physical inconsistencies or specifica

RPCASSM: Robust PCA State Space Model For Infrared Small Target Detection

Model ReleasesDGX agent

arXiv:2606.01689v1 Announce Type: cross Abstract: The detection and segmentation of infrared small targets have important application significance in the fields of surveillance and security, maritime

T1: Tool-integrated Verification for Test-time Compute Scaling in Small Language Models

Model ReleasesDGX agent

arXiv:2504.04718v2 Announce Type: replace-cross Abstract: Recent studies have demonstrated that test-time compute scaling effectively improves the performance of small language models (sLMs). However,

Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?

Model ReleasesDGX agent

arXiv:2606.01247v1 Announce Type: new Abstract: Humans can reproduce the viewpoint specified by a target image through active head and body motion, yet spatial intelligence in foundation models has la

1 Jun 2026

Diversity Matters: Revisiting Test-Time Compute in Vision-Language Models

ResearchDGX agent

arXiv:2605.30713v1 Announce Type: cross Abstract: Test-time compute (TTC) strategies have emerged as a lightweight approach to boost reasoning in large language models (LLMs). However, their applicati

MADS: Model-Aware Diverse Core Set Selection for Instruction Tuning

Model ReleasesDGX agent

arXiv:2605.30857v1 Announce Type: new Abstract: Instruction fine-tuning is employed to enhance the instruction-following ability of large language models (LLMs). As the amount of instruction fine-tuni

Skill Availability and Presentation Granularity in Large-Language-Model Agents: A Controlled SkillsBench Study

Model ReleasesDGX agent

arXiv:2605.31408v1 Announce Type: cross Abstract: Skill documents provide procedural knowledge to large-language-model agents at inference time. This article studies whether the presentation granulari

UniScale: Adaptive Unified Inference Scaling via Online Joint Optimization of Model Routing and Test-Time Scaling

ApplicationsDGX agent

arXiv:2605.30898v1 Announce Type: new Abstract: In real-world deployments of large language models (LLMs), balancing inference quality and computational cost has become a central challenge. Existing a

Your Multimodal Speech Model Says I Have a Face for Radio

Model ReleasesDGX agent

arXiv:2605.30472v1 Announce Type: new Abstract: As large neural models have become better at language tasks, researchers are increasingly building multi- and omnimodal models that handle more modaliti

29 May 2026

AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Crystalline Materials

Model ReleasesDGX agent

arXiv:2510.04704v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown promising potential in scientific research, enabling tasks ranging from knowledge retrieval to propert

Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation

Model ReleasesDGX agent

arXiv:2605.28830v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are increasingly deployed in safety-critical applications, robust content moderation becomes essential. We present a c

Beyond Accuracy: Are Time Series Foundation Models Well-Calibrated?

ResearchDGX agent

arXiv:2510.16060v2 Announce Type: replace-cross Abstract: The recent development of foundation models for time series data has generated considerable interest in using such models across a variety of

CosmicFish-HRM: Adaptive Reasoning via Hierarchical Recurrent Mechanisms in Compact Language Models

Model ReleasesDGX agent

arXiv:2605.28919v1 Announce Type: cross Abstract: Large language models have achieved strong reasoning capabilities, though often at the cost of massive parameter counts and expensive inference. In th

DenseSteer: Steering Small Language Models towards Dense Math Reasoning

Model ReleasesDGX agent

arXiv:2605.29247v1 Announce Type: new Abstract: Large language models (LLMs) demonstrate strong chain-of-thought (CoT) reasoning abilities, while smaller models (<= 3B parameters) significantly underp

Label-Free Reinforcement Learning via Cross-Model Entropy

Model ReleasesDGX agent

arXiv:2605.29009v1 Announce Type: cross Abstract: Post-training large language models with reinforcement learning is bottlenecked by the reward signal. Existing approaches require either ground-truth

Opt-Verifier: Unleashing the Power of LLMs for Optimization Modeling via Dual-Side Verification

ResearchDGX agent

arXiv:2605.29556v1 Announce Type: new Abstract: Building mathematical optimization models is critical in operations research (OR), while it requires substantial human expertise. Recent advancements ha

Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought

Model ReleasesDGX agent

arXiv:2603.05488v4 Announce Type: replace-cross Abstract: We provide evidence of performative chain-of-thought (CoT) in reasoning models, where a model becomes strongly confident in its final answer,

Why Far Looks Up: Probing Spatial Representation in Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.30161v1 Announce Type: new Abstract: Vision-language models (VLMs) achieve strong performance on spatial reasoning benchmarks, yet it remains unclear whether this reflects structured 3D und

28 May 2026

A Simple State Space Model Excels at Multivariate Time Series Classification

Model ReleasesDGX agent

arXiv:2605.27406v1 Announce Type: new Abstract: Structured state space models (SSMs) have recently emerged as a promising foundation for sequence modeling, with Mamba-based architectures demonstrating

Evaluating Local Explainability Metrics for Machine Learning Models on Tabular Data

Model ReleasesDGX agent

arXiv:2605.27618v1 Announce Type: new Abstract: Despite the wide use of explainability techniques to attempt to understand the behavior of Artificial Intelligence (AI), the generated explanations may

Evaluation of AI Ethics Tools in Language Models: A Developers' Perspective Case Study

TutorialsDGX agent

arXiv:2512.15791v2 Announce Type: replace-cross Abstract: In Artificial Intelligence (AI), language models have gained significant importance due to the widespread adoption of systems capable of simul

FedMPT: Federated Multi-label Prompt Tuning of Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.28347v1 Announce Type: new Abstract: Multi-Label Recognition (MLR) based on Vision-Language Models (VLMs) aims to leverage their pre-trained knowledge to better adapt complex recognition sc

Learning Correlated Reward Models: Statistical Barriers and Opportunities

TutorialsDGX agent

arXiv:2510.15839v2 Announce Type: replace Abstract: Random Utility Models (RUMs) are a classical framework for modeling user preferences and play a key role in reward modeling for Reinforcement Learni

PetroBench: A Benchmark for Large Language Models in Petroleum Engineering

Model ReleasesDGX agent

arXiv:2605.28032v1 Announce Type: new Abstract: Large Language Models are increasingly applied in the petroleum industry, highlighting the need for a domain-specific evaluation framework. This study d

SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding

Model ReleasesDGX agent

arXiv:2601.21666v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) are a major focus of recent AI research. However, most prior work focuses on static image understanding, wh

When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models

Model ReleasesDGX agent

arXiv:2605.27851v1 Announce Type: new Abstract: Safety benchmark scores provide incomplete evidence of deployment readiness: aligned language models often adhere to rigid rules even when a situational

When Interpretability Is Unequally Distributed: Fairness in Hybrid Interpretable Models

Model ReleasesDGX agent

arXiv:2605.28626v1 Announce Type: new Abstract: Hybrid interpretable models combine a transparent component with a black-box model by assigning some examples to the former and deferring the rest to th

27 May 2026

Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders

ResearchDGX agent

arXiv:2605.27354v1 Announce Type: cross Abstract: Model internals encode rich information about how a large language model (LLM) processes its training data; however, post-training data engineering la

ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks — by Artificial Analysis and IBM

Model ReleasesDGX agent

ITBench-AA is a new benchmark developed by Artificial Analysis and IBM that evaluates frontier AI models on agentic enterprise IT tasks, with results showing that current leading models score below 50

Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation

Model ReleasesDGX agent

arXiv:2511.14993v3 Announce Type: replace-cross Abstract: This report introduces Kandinsky 5.0, a family of state-of-the-art foundation models for high-resolution image and 10-second video synthesis.

Learning to Predict Future-Aligned Research Proposals with Language Models

Model ReleasesDGX agent

arXiv:2603.27146v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to assist ideation in research, but evaluating the quality of LLM-generated research proposals re

Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation

ResearchDGX agent

arXiv:2507.16116v2 Announce Type: replace Abstract: The rapid advancement of video diffusion models has been hindered by fundamental limitations in temporal modeling, particularly the rigid synchroniz

Recon: Reconstruction-Guided Reasoning Synthesis for User Modeling

ResearchDGX agent

arXiv:2605.26969v1 Announce Type: cross Abstract: User modeling aims to use language models (LMs) to mimic an individual's behavior from a corpus of past context-action pairs (e.g., conversation turns

Self-Verified Distillation: Your Language Model Is Secretly Its Own Synthetic Data Pipeline

Model ReleasesDGX agent

arXiv:2605.26132v1 Announce Type: new Abstract: Can post-trained large language models (LLMs) further improve themselves using only unlabeled prompts, without external teachers or feedback from tools?

SpatialBench: Is Your Spatial Foundation Model an All-Round Player?

Model ReleasesDGX agent

arXiv:2605.27367v1 Announce Type: new Abstract: While spatial foundation models have demonstrated impressive performance on standard datasets, a critical question remains: are they truly all-round pla

26 May 2026

Frequency Matters: Fast Model-Agnostic Data Curation for Pruning and Quantization

ResearchDGX agent

arXiv:2603.16105v3 Announce Type: replace-cross Abstract: Post-training model compression is essential for enhancing the portability of Large Language Models (LLMs) while preserving their performance.

How Well Do Models Follow Their Constitutions?

Model ReleasesDGX agent

arXiv:2605.24229v1 Announce Type: new Abstract: Frontier AI developers now train models against long written behavioral specifications, such as Anthropic's constitution (Anthropic, 2025a) and OpenAI's

25 May 2026

An open source model has returned to #1 on the 3D Design leaderboard by Design Arena. Kimi K2.6 has reached the top of the leaderboard for 3…

Model ReleasesDGX agent

An open source model has returned to #1 on the 3D Design leaderboard by Design Arena. Kimi K2.6 has reached the top of the leaderboard for 3D Design, ahead of models 10X more expensive like Opus 4.7 b

Complete-muE: Optimal Hyperparameter Transfer and Scaling for MoE Models

Model ReleasesDGX agent

arXiv:2605.23893v1 Announce Type: new Abstract: We propose Complete-muE, a framework which targets hyperparameter transfer across dense FFN and any Mixture-of-Experts (MoE) setups in transformer block

22 May 2026

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models

Model ReleasesDGX agent

arXiv:2605.21573v1 Announce Type: new Abstract: We introduce Lens, a 3.8B-parameter T2I model that achieves performance competitive with, and in several cases surpassing, state-of-the-art models with

Reinforcing VLAs in Task-Agnostic World Models

SafetyDGX agent

arXiv:2605.12334v2 Announce Type: replace Abstract: Post-training Vision-Language-Action (VLA) models via reinforcement learning (RL) in learned world models has emerged as an effective strategy to ad

21 May 2026

Toxic Subword Pruning for Dialogue Response Generation on Large Language Models

Model ReleasesDGX agent

arXiv:2410.04155v2 Announce Type: replace Abstract: How to defend large language models (LLMs) from generating toxic content is an important research area. Yet, most research focused on various model

20 May 2026

Beyond Prediction Accuracy: Target-Space Recovery Profiles for Evaluating Model-Brain Alignment

Model ReleasesDGX agent

arXiv:2605.20127v1 Announce Type: cross Abstract: Artificial vision models are often evaluated against the human visual cortex by measuring how accurately their internal representations predict brain

Conflict-Free Replicated Data Types for Neural Network Model Merging: A Two-Layer Architecture Enabling CRDT-Compliant Model Merging Across 26 Strategies

ApplicationsDGX agent

arXiv:2605.19373v1 Announce Type: cross Abstract: All 26 neural network merge strategies we tested including weight averaging, SLERP, TIES, DARE, Fisher merging, and evolutionary approaches -- fail th

Dynamic Model Merging Made Slim

Model ReleasesDGX agent

arXiv:2605.18904v1 Announce Type: cross Abstract: Model merging enables the reuse of fine-tuned models without joint training or access to original data. Dynamic merging further improves flexibility b

Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance

Model ReleasesDGX agent

arXiv:2512.23461v2 Announce Type: replace-cross Abstract: Reward models (RMs) are essential in reinforcement learning from human feedback (RLHF) to align large language models (LLMs) with human values

Flash PD-SSM: Memory-Optimized Structured Sparse State-Space Models

ResearchDGX agent

arXiv:2605.19150v1 Announce Type: cross Abstract: State-space models (SSMs) face a fundamental trade-off between efficiency and expressivity that is mainly dictated by the structure of the model's tra

19 May 2026

GVGAI-LLM: Evaluating Large Language Model Agents with Infinite Games

Model ReleasesDGX agent

arXiv:2508.08501v3 Announce Type: replace Abstract: We introduce GVGAI-LLM, a video game benchmark for evaluating the reasoning and problem-solving capabilities of large language models (LLMs). Built

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation

Model ReleasesDGX agent

arXiv:2410.13846v3 Announce Type: replace-cross Abstract: Scaling language models to handle longer contexts introduces substantial memory challenges due to the growing cost of key-value (KV) caches. M

Noise2Params: Unification and Parameter Determination from Noise via a Probabilistic Event Camera Model

Model ReleasesDGX agent

arXiv:2605.16317v1 Announce Type: new Abstract: Accurate, unified models for event cameras (ECs) remain elusive, hampering calibration and algorithm design. We develop a foundational probabilistic mod

The Token Games: Evaluating Language Model Reasoning with Puzzle Duels

ResearchDGX agent

arXiv:2602.17831v2 Announce Type: replace Abstract: Evaluating the reasoning capabilities of Large Language Models is increasingly challenging as models improve. Human curation of hard questions is hi

Threats to Arabic Handwriting Recognition: Investigating Black-Box Adversarial Attacks on embedded ConvNet models

Model ReleasesDGX agent

arXiv:2605.18058v1 Announce Type: new Abstract: Arabic handwriting recognition (AHR) has made significant progress with deep learning models. AHR research has largely focused on performance, with secu

Venom: A PyTorch Generative Modeling Toolkit

TutorialsDGX agent

arXiv:2605.17605v1 Announce Type: new Abstract: Modern generative modeling has grown into a broad collection of related but often separately implemented paradigms, including denoising diffusion models

WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform

Model ReleasesDGX agent

arXiv:2605.17912v1 Announce Type: cross Abstract: World models have emerged as a central paradigm for embodied intelligence, enabling agents to predict action-conditioned future and reason about envir

18 May 2026

A Cross-Modal Prompt Injection Attack against Large Vision-Language Models with Image-Only Perturbation

Model ReleasesDGX agent

arXiv:2605.16090v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) have emerged as a powerful paradigm for multimodal intelligence, but their growing deployment also expands the at

A numerical study into neural network surrogate model performance for uncertainty propagation

ResearchDGX agent

arXiv:2605.16078v1 Announce Type: cross Abstract: Neural network surrogate models have emerged as a promising approach to model solution fields for a wide variety of boundary value problems encountere

A Unified View of Score-Based and Drifting Models

TutorialsDGX agent

arXiv:2603.07514v3 Announce Type: replace-cross Abstract: Drifting models train one-step generators by optimizing a kernel-induced mean-shift discrepancy between the data and model distributions, with

Feedback World Model Enables Precise Guidance of Diffusion Policy

Model ReleasesDGX agent

arXiv:2605.15705v1 Announce Type: cross Abstract: World models aim to improve robotic decision making by predicting the consequences of actions. However, in practice, their predictions often become un

← Previous
1…2930313233…998
Next →