AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,860 results
Tutorials

Learning Correlated Reward Models: Statistical Barriers and Opportunities

DGX agent

arXiv:2510.15839v2 Announce Type: replace Abstract: Random Utility Models (RUMs) are a classical framework for modeling user preferences and play a key role in reward modeling for Reinforcement Learni

tutorialsarxiv-cs-lg
28 May 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

PetroBench: A Benchmark for Large Language Models in Petroleum Engineering

DGX agent

arXiv:2605.28032v1 Announce Type: new Abstract: Large Language Models are increasingly applied in the petroleum industry, highlighting the need for a domain-specific evaluation framework. This study d

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding

DGX agent

arXiv:2601.21666v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) are a major focus of recent AI research. However, most prior work focuses on static image understanding, wh

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models

DGX agent

arXiv:2605.27851v1 Announce Type: new Abstract: Safety benchmark scores provide incomplete evidence of deployment readiness: aligned language models often adhere to rigid rules even when a situational

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

When Interpretability Is Unequally Distributed: Fairness in Hybrid Interpretable Models

DGX agent

arXiv:2605.28626v1 Announce Type: new Abstract: Hybrid interpretable models combine a transparent component with a black-box model by assigning some examples to the former and deferring the rest to th

model-releasesarxiv-cs-lg
28 May 2026
Research

Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders

DGX agent

arXiv:2605.27354v1 Announce Type: cross Abstract: Model internals encode rich information about how a large language model (LLM) processes its training data; however, post-training data engineering la

researcharxiv-cs-ai
27 May 2026
Model Releases

ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks — by Artificial Analysis and IBM

DGX agent

ITBench-AA is a new benchmark developed by Artificial Analysis and IBM that evaluates frontier AI models on agentic enterprise IT tasks, with results showing that current leading models score below 50

model-releaseshugging-face
27 May 2026
Model Releases

Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation

DGX agent

arXiv:2511.14993v3 Announce Type: replace-cross Abstract: This report introduces Kandinsky 5.0, a family of state-of-the-art foundation models for high-resolution image and 10-second video synthesis.

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Learning to Predict Future-Aligned Research Proposals with Language Models

DGX agent

arXiv:2603.27146v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to assist ideation in research, but evaluating the quality of LLM-generated research proposals re

model-releasesarxiv-cs-cl
27 May 2026
Research

Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation

DGX agent

arXiv:2507.16116v2 Announce Type: replace Abstract: The rapid advancement of video diffusion models has been hindered by fundamental limitations in temporal modeling, particularly the rigid synchroniz

researcharxiv-cs-cv
27 May 2026
Research

Recon: Reconstruction-Guided Reasoning Synthesis for User Modeling

DGX agent

arXiv:2605.26969v1 Announce Type: cross Abstract: User modeling aims to use language models (LMs) to mimic an individual's behavior from a corpus of past context-action pairs (e.g., conversation turns

researcharxiv-cs-ai
27 May 2026
Model Releases

Self-Verified Distillation: Your Language Model Is Secretly Its Own Synthetic Data Pipeline

DGX agent

arXiv:2605.26132v1 Announce Type: new Abstract: Can post-trained large language models (LLMs) further improve themselves using only unlabeled prompts, without external teachers or feedback from tools?

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

SpatialBench: Is Your Spatial Foundation Model an All-Round Player?

DGX agent

arXiv:2605.27367v1 Announce Type: new Abstract: While spatial foundation models have demonstrated impressive performance on standard datasets, a critical question remains: are they truly all-round pla

model-releasesarxiv-cs-cv
27 May 2026
Research

Frequency Matters: Fast Model-Agnostic Data Curation for Pruning and Quantization

DGX agent

arXiv:2603.16105v3 Announce Type: replace-cross Abstract: Post-training model compression is essential for enhancing the portability of Large Language Models (LLMs) while preserving their performance.

researcharxiv-cs-ai
26 May 2026
Model Releases

How Well Do Models Follow Their Constitutions?

DGX agent

arXiv:2605.24229v1 Announce Type: new Abstract: Frontier AI developers now train models against long written behavioral specifications, such as Anthropic's constitution (Anthropic, 2025a) and OpenAI's

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

An open source model has returned to #1 on the 3D Design leaderboard by Design Arena. Kimi K2.6 has reached the top of the leaderboard for 3…

DGX agent

An open source model has returned to #1 on the 3D Design leaderboard by Design Arena. Kimi K2.6 has reached the top of the leaderboard for 3D Design, ahead of models 10X more expensive like Opus 4.7 b

model-releaseskimi-moonshot--x
25 May 2026
Model Releases

Complete-muE: Optimal Hyperparameter Transfer and Scaling for MoE Models

DGX agent

arXiv:2605.23893v1 Announce Type: new Abstract: We propose Complete-muE, a framework which targets hyperparameter transfer across dense FFN and any Mixture-of-Experts (MoE) setups in transformer block

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models

DGX agent

arXiv:2605.21573v1 Announce Type: new Abstract: We introduce Lens, a 3.8B-parameter T2I model that achieves performance competitive with, and in several cases surpassing, state-of-the-art models with

model-releasesarxiv-cs-cv
22 May 2026
Safety

Reinforcing VLAs in Task-Agnostic World Models

DGX agent

arXiv:2605.12334v2 Announce Type: replace Abstract: Post-training Vision-Language-Action (VLA) models via reinforcement learning (RL) in learned world models has emerged as an effective strategy to ad

safetyarxiv-cs-ai
22 May 2026
Model Releases

Toxic Subword Pruning for Dialogue Response Generation on Large Language Models

DGX agent

arXiv:2410.04155v2 Announce Type: replace Abstract: How to defend large language models (LLMs) from generating toxic content is an important research area. Yet, most research focused on various model

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

Beyond Prediction Accuracy: Target-Space Recovery Profiles for Evaluating Model-Brain Alignment

DGX agent

arXiv:2605.20127v1 Announce Type: cross Abstract: Artificial vision models are often evaluated against the human visual cortex by measuring how accurately their internal representations predict brain

model-releasesarxiv-cs-ai
20 May 2026
Applications

Conflict-Free Replicated Data Types for Neural Network Model Merging: A Two-Layer Architecture Enabling CRDT-Compliant Model Merging Across 26 Strategies

DGX agent

arXiv:2605.19373v1 Announce Type: cross Abstract: All 26 neural network merge strategies we tested including weight averaging, SLERP, TIES, DARE, Fisher merging, and evolutionary approaches -- fail th

applicationsarxiv-cs-ai
20 May 2026
Model Releases

Dynamic Model Merging Made Slim

DGX agent

arXiv:2605.18904v1 Announce Type: cross Abstract: Model merging enables the reuse of fine-tuned models without joint training or access to original data. Dynamic merging further improves flexibility b

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance

DGX agent

arXiv:2512.23461v2 Announce Type: replace-cross Abstract: Reward models (RMs) are essential in reinforcement learning from human feedback (RLHF) to align large language models (LLMs) with human values

model-releasesarxiv-cs-ai
20 May 2026
Research

Flash PD-SSM: Memory-Optimized Structured Sparse State-Space Models

DGX agent

arXiv:2605.19150v1 Announce Type: cross Abstract: State-space models (SSMs) face a fundamental trade-off between efficiency and expressivity that is mainly dictated by the structure of the model's tra

researcharxiv-cs-ai
20 May 2026
Model Releases

GVGAI-LLM: Evaluating Large Language Model Agents with Infinite Games

DGX agent

arXiv:2508.08501v3 Announce Type: replace Abstract: We introduce GVGAI-LLM, a video game benchmark for evaluating the reasoning and problem-solving capabilities of large language models (LLMs). Built

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation

DGX agent

arXiv:2410.13846v3 Announce Type: replace-cross Abstract: Scaling language models to handle longer contexts introduces substantial memory challenges due to the growing cost of key-value (KV) caches. M

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Noise2Params: Unification and Parameter Determination from Noise via a Probabilistic Event Camera Model

DGX agent

arXiv:2605.16317v1 Announce Type: new Abstract: Accurate, unified models for event cameras (ECs) remain elusive, hampering calibration and algorithm design. We develop a foundational probabilistic mod

model-releasesarxiv-cs-cv
19 May 2026
Research

The Token Games: Evaluating Language Model Reasoning with Puzzle Duels

DGX agent

arXiv:2602.17831v2 Announce Type: replace Abstract: Evaluating the reasoning capabilities of Large Language Models is increasingly challenging as models improve. Human curation of hard questions is hi

researcharxiv-cs-ai
19 May 2026
Model Releases

Threats to Arabic Handwriting Recognition: Investigating Black-Box Adversarial Attacks on embedded ConvNet models

DGX agent

arXiv:2605.18058v1 Announce Type: new Abstract: Arabic handwriting recognition (AHR) has made significant progress with deep learning models. AHR research has largely focused on performance, with secu

model-releasesarxiv-cs-cv
19 May 2026
Tutorials

Venom: A PyTorch Generative Modeling Toolkit

DGX agent

arXiv:2605.17605v1 Announce Type: new Abstract: Modern generative modeling has grown into a broad collection of related but often separately implemented paradigms, including denoising diffusion models

tutorialsarxiv-cs-lg
19 May 2026
Model Releases

WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform

DGX agent

arXiv:2605.17912v1 Announce Type: cross Abstract: World models have emerged as a central paradigm for embodied intelligence, enabling agents to predict action-conditioned future and reason about envir

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

A Cross-Modal Prompt Injection Attack against Large Vision-Language Models with Image-Only Perturbation

DGX agent

arXiv:2605.16090v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) have emerged as a powerful paradigm for multimodal intelligence, but their growing deployment also expands the at

model-releasesarxiv-cs-cv
18 May 2026
Research

A numerical study into neural network surrogate model performance for uncertainty propagation

DGX agent

arXiv:2605.16078v1 Announce Type: cross Abstract: Neural network surrogate models have emerged as a promising approach to model solution fields for a wide variety of boundary value problems encountere

researcharxiv-cs-lg
18 May 2026
Tutorials

A Unified View of Score-Based and Drifting Models

DGX agent

arXiv:2603.07514v3 Announce Type: replace-cross Abstract: Drifting models train one-step generators by optimizing a kernel-induced mean-shift discrepancy between the data and model distributions, with

tutorialsarxiv-cs-ai
18 May 2026
Model Releases

Feedback World Model Enables Precise Guidance of Diffusion Policy

DGX agent

arXiv:2605.15705v1 Announce Type: cross Abstract: World models aim to improve robotic decision making by predicting the consequences of actions. However, in practice, their predictions often become un

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

Genome-Factory: A Library for Tuning, Deploying, and Interpreting Genomic Foundation Models

DGX agent

arXiv:2509.12266v2 Announce Type: replace-cross Abstract: We introduce Genome-Factory, the first integrated Python library for tuning, deploying, and interpreting genomic foundation models. Our core c

model-releasesarxiv-cs-lg
18 May 2026
Model Releases

Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels

DGX agent

arXiv:2605.15208v1 Announce Type: cross Abstract: Large Language Models are routinely compressed via post-training quantization to reduce inference costs and memory footprint for cloud and edge deploy

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

A Large Language Model Based Pipeline for Review of Systems Entity Recognition from Clinical Notes

DGX agent

arXiv:2506.11067v3 Announce Type: replace Abstract: Objective: Develop a cost-effective, large language model (LLM)-based pipeline for automatically extracting Review of Systems (ROS) entities from cl

model-releasesarxiv-cs-cl
15 May 2026
Applications

An Interpretable Latency Model for Speculative Decoding in LLM Serving

DGX agent

arXiv:2605.15051v1 Announce Type: new Abstract: Speculative decoding (SD) accelerates large language model (LLM) inference by using a smaller draft model to propose multiple tokens that are verified b

applicationsarxiv-cs-lg
15 May 2026
Model Releases

Darwin Family: MRI-Trust-Weighted Evolutionary Merging for Training-Free Scaling of Language-Model Reasoning

DGX agent

arXiv:2605.14386v1 Announce Type: cross Abstract: We present Darwin Family, a framework for training-free evolutionary merging of large language models via gradient-free weight-space recombination. We

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models

DGX agent

arXiv:2511.08565v3 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly operate in social contexts, motivating analysis of how they express and shift moral judgments. In th

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

NeuroMambaLLM: Dynamic Graph Learning of fMRI Functional Connectivity in Autistic Brains Using Mamba and Language Model Reasoning

DGX agent

arXiv:2602.13770v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated strong semantic reasoning across multimodal domains. However, their integration with graph-base

model-releasesarxiv-cs-lg
15 May 2026
Model Releases

On the Cultural Anachronism and Temporal Reasoning in Vision Language Models

DGX agent

arXiv:2605.15071v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly applied to cultural heritage materials, from digital archives to educational platforms. This work ident

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Tokenizer Fertility and Zero-Shot Performance of Foundation Models on Ukrainian Legal Text: A Comparative Study

DGX agent

arXiv:2605.14890v1 Announce Type: new Abstract: Foundation models tokenize Ukrainian legal text with vastly different efficiency, yet no systematic comparison exists for this domain. We benchmark seve

model-releasesarxiv-cs-cl
15 May 2026
Model Releases

AttenA+: Rectifying Action Inequality in Robotic Foundation Models

DGX agent

arXiv:2605.13548v1 Announce Type: cross Abstract: Existing robotic foundation models, while powerful, are predicated on an implicit assumption of temporal homogeneity: treating all actions as equally

model-releasesarxiv-cs-ai
14 May 2026
Research

Behavioral Geometric Supervision Aligns Video Foundation Models with Human Social Perception

DGX agent

arXiv:2510.01502v2 Announce Type: replace-cross Abstract: Current video foundation models, including the strongest self-supervised models such as V-JEPA2, fail to capture how humans organize social in

researcharxiv-cs-cv
14 May 2026
Model Releases

Language Model Goal Selection Differs from Humans' in a Self-Directed Learning Task

DGX agent

arXiv:2603.03295v2 Announce Type: replace-cross Abstract: Whether in agentic workflows, social studies, or chat settings, large language models (LLMs) are increasingly being asked to replace humans in

model-releasesarxiv-cs-ai
14 May 2026
← Previous
1…3738394041…1248
Next →