AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
48,543 results
Research

Predicting Large Model Test Losses with a Noisy Quadratic System

DGX agent

arXiv:2605.09154v1 Announce Type: new Abstract: We introduce a predictive model that estimates the pre-training loss of large models from model size (N), batch size (B) and number of weight updates (K

researcharxiv-cs-lg
12 May 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Chain-based Distillation for Effective Initialization of Variable-Sized Small Language Models

DGX agent

arXiv:2605.07783v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong performance but remain costly to deploy in resource-constrained settings. Training small language models (SL

model-releasesarxiv-cs-cl
11 May 2026
Hardware

LLMSpace: Carbon Footprint Modeling for Large Language Model Inference on LEO Satellites

DGX agent

arXiv:2605.05615v2 Announce Type: replace Abstract: Large language models (LLMs) impose rapidly growing energy demands, creating an emerging energy and carbon crisis driven by large-scale inference. S

hardwarearxiv-cs-lg
11 May 2026
Model Releases

EvoLM: Self-Evolving Language Models through Co-Evolved Discriminative Rubrics

DGX agent

arXiv:2605.03871v1 Announce Type: new Abstract: Language models encode substantial evaluative knowledge from pretraining, yet current post-training methods rely on external supervision (human annotati

model-releasesarxiv-cs-ai
7 May 2026
Model Releases

StoryAlign: Evaluating and Training Reward Models for Story Generation

DGX agent

arXiv:2605.04831v1 Announce Type: new Abstract: Story generation aims to automatically produce coherent, structured, and engaging narratives. Although large language models (LLMs) have significantly a

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

Deep Time Series Models: A Comprehensive Survey and Benchmark

DGX agent

arXiv:2407.13278v3 Announce Type: replace Abstract: Time series, characterized by a sequence of data points organized in a discrete-time order, are ubiquitous in real-world scenarios. Unlike other dat

model-releasesarxiv-cs-lg
5 May 2026
Research

Enforcing tail calibration when training probabilistic forecast models

DGX agent

arXiv:2506.13687v2 Announce Type: replace-cross Abstract: Probabilistic forecasts are typically obtained using state-of-the-art statistical and machine learning models, with model parameters estimated

researcharxiv-cs-lg
5 May 2026
Model Releases

TRIP-Evaluate: An Open Multimodal Benchmark for Evaluating Large Models in Transportation

DGX agent

arXiv:2605.00907v1 Announce Type: new Abstract: Large language models (LLMs) and multimodal large models (MLLMs) are increasingly used for transportation tasks such as regulation question answering, t

model-releasesarxiv-cs-cv
5 May 2026
Agents

Pragmos: A Process Agentic Modeling System

DGX agent

arXiv:2604.27311v1 Announce Type: cross Abstract: The advent of Large Language Models (LLMs) has significantly transformed tasks across Software Engineering. In the context of Business Process Managem

agentsarxiv-cs-ai
1 May 2026
Safety

A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

DGX agent

arXiv:2510.08049v3 Announce Type: replace-cross Abstract: Although Large Language Models (LLMs) exhibit advanced reasoning ability, conventional alignment remains largely dominated by outcome reward m

safetyarxiv-cs-ai
30 Apr 2026
Research

Fitting Large Nonlinear Mixed Effects Models Using Variational Expectation Maximization

DGX agent

arXiv:2604.26160v1 Announce Type: cross Abstract: Nonlinear Mixed Effects models (NLME) models are widely used in pharmacometrics and related fields to analyze hierarchical and longitudinal data. Howe

researcharxiv-cs-lg
30 Apr 2026
Model Releases

reward-lens: A Mechanistic Interpretability Library for Reward Models

DGX agent

arXiv:2604.26130v1 Announce Type: cross Abstract: Every RLHF-trained language model is shaped by a reward model, yet the mechanistic interpretability toolkit -- logit lens, direct logit attribution, a

model-releasesarxiv-cs-ai
30 Apr 2026
Research

Cornserve: A Distributed Serving System for Any-to-Any Multimodal Models

DGX agent

arXiv:2603.12118v2 Announce Type: replace Abstract: Any-to-Any models are an emerging class of multimodal models that accept combinations of multimodal data (e.g., text, image, video, audio) as input

researcharxiv-cs-lg
29 Apr 2026
Model Releases

Benchmarking Pathology Foundation Models for Breast Cancer Survival Prediction

DGX agent

arXiv:2604.24679v1 Announce Type: new Abstract: Pathology foundation models (PFMs) have recently emerged as powerful pretrained encoders for computational pathology, enabling transfer learning across

model-releasesarxiv-cs-cv
28 Apr 2026
Research

RoboECC: Multi-Factor-Aware Edge-Cloud Collaborative Deployment for VLA Models

DGX agent

arXiv:2603.20711v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models are mainstream in embodied intelligence but face high inference costs. Edge-Cloud Collaborative (ECC) depl

researcharxiv-cs-lg
28 Apr 2026
Model Releases

DistortBench: Benchmarking Vision Language Models on Image Distortion Identification

DGX agent

arXiv:2604.19966v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used in settings where sensitivity to low-level image degradations matters, including content moderatio

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

Foundation Models in Biomedical Imaging: Turning Hype into Reality

DGX agent

arXiv:2512.15808v2 Announce Type: replace-cross Abstract: Foundation models (FMs) are driving a prominent shift in biomedical imaging from task-specific models to unified backbone models for diverse t

model-releasesarxiv-cs-ai
23 Apr 2026
Research

Hallucination Early Detection in Diffusion Models

DGX agent

arXiv:2604.20354v1 Announce Type: new Abstract: Text-to-Image generation has seen significant advancements in output realism with the advent of diffusion models. However, diffusion models encounter di

researcharxiv-cs-cv
23 Apr 2026
Model Releases

3D Foundation Model for Generalizable Disease Detection in Head Computed Tomography

DGX agent

arXiv:2502.02779v3 Announce Type: replace-cross Abstract: Head computed tomography (CT) imaging is a widely-used imaging modality with multitudes of medical indications, particularly in assessing path

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

An Empirical Study of Multi-Generation Sampling for Jailbreak Detection in Large Language Models

DGX agent

arXiv:2604.18775v1 Announce Type: new Abstract: Detecting jailbreak behaviour in large language models remains challenging, particularly when strongly aligned models produce harmful outputs only rarel

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

Can Continual Pre-training Bridge the Performance Gap between General-purpose and Specialized Language Models in the Medical Domain?

DGX agent

arXiv:2604.19394v1 Announce Type: new Abstract: This paper narrows the performance gap between small, specialized models and significantly larger general-purpose models through domain adaptation via c

model-releasesarxiv-cs-cl
22 Apr 2026
Safety

How does the optimizer implicitly bias the model merging loss landscape?

DGX agent

arXiv:2510.04686v2 Announce Type: replace-cross Abstract: Model merging combines independent solutions with different capabilities into a single one while maintaining the same inference cost. Two popu

safetyarxiv-cs-ai
22 Apr 2026
Model Releases

Micro Language Models Enable Instant Responses

DGX agent

arXiv:2604.19642v1 Announce Type: new Abstract: Edge devices such as smartwatches and smart glasses cannot continuously run even the smallest 100M-1B parameter language models due to power and compute

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

OmniMouse: Scaling properties of multi-modal, multi-task Brain Models on 150B Neural Tokens

DGX agent

arXiv:2604.18827v1 Announce Type: cross Abstract: Scaling data and artificial neural networks has transformed AI, driving breakthroughs in language and vision. Whether similar principles apply to mode

model-releasesarxiv-cs-ai
22 Apr 2026
Research

An Exploration of Mamba for Speech Self-Supervised Models

DGX agent

arXiv:2506.12606v2 Announce Type: replace Abstract: While Mamba has demonstrated strong performance in language modeling, its potential as a speech self-supervised learning (SSL) model remains underex

researcharxiv-cs-cl
21 Apr 2026
Research

Joint Distillation for Fast Likelihood Evaluation and Sampling in Flow-based Models

DGX agent

arXiv:2512.02636v3 Announce Type: replace-cross Abstract: Log-likelihood evaluation enables important capabilities in generative models, including model comparison, certain fine-tuning objectives, and

researcharxiv-cs-cv
21 Apr 2026
Model Releases

MMErroR: A Benchmark for Erroneous Reasoning in Vision-Language Models

DGX agent

arXiv:2601.03331v2 Announce Type: replace Abstract: Recent advances in Vision-Language Models (VLMs) have improved performance in multi-modal learning, raising the question of whether these models tru

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

MoCo: A One-Stop Shop for Model Collaboration Research

DGX agent

arXiv:2601.21257v2 Announce Type: replace Abstract: Advancing beyond single monolithic language models (LMs), recent research increasingly recognizes the importance of model collaboration, where multi

safetyarxiv-cs-cl
21 Apr 2026
Applications

PARM: Pipeline-Adapted Reward Model

DGX agent

arXiv:2604.18327v1 Announce Type: cross Abstract: Reward models (RMs) are central to aligning large language models (LLMs) with human preferences, powering RLHF and advanced decoding strategies. While

applicationsarxiv-cs-cl
21 Apr 2026
Applications

Foundation Models in Robotics: A Comprehensive Review of Methods, Models, Datasets, Challenges and Future Research Directions

DGX agent

arXiv:2604.15395v1 Announce Type: new Abstract: Over the recent years, the field of robotics has been undergoing a transformative paradigm shift from fixed, single-task, domain-specific solutions towa

applicationsarxiv-cs-ro
20 Apr 2026
Research

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation

DGX agent

arXiv:2510.09065v2 Announce Type: replace-cross Abstract: We introduce MMAudioSep, a generative model for video/text-queried sound separation that is founded on a pretrained video-to-audio model. By l

researcharxiv-cs-cv
20 Apr 2026
Model Releases

Constrained Decoding for Safe Robot Navigation Foundation Models

DGX agent

arXiv:2509.01728v4 Announce Type: replace-cross Abstract: Recent advances in the development of robotic foundation models have led to promising end-to-end and general-purpose capabilities in robotic s

model-releasesarxiv-cs-lg
17 Apr 2026
Research

Constraint-based Pre-training: From Structured Constraints to Scalable Model Initialization

DGX agent

arXiv:2604.14769v1 Announce Type: new Abstract: The pre-training and fine-tuning paradigm has become the dominant approach for model adaptation. However, conventional pre-training typically yields mod

researcharxiv-cs-lg
17 Apr 2026
Local Ai

EEGDM: Learning EEG Representation with Latent Diffusion Model

DGX agent

arXiv:2508.20705v3 Announce Type: replace Abstract: Recent advances in self-supervised learning for EEG representation have largely relied on masked reconstruction, where models are trained to recover

local-aiarxiv-cs-lg
17 Apr 2026
Tutorials

Diffusion Language Models for Speech Recognition

DGX agent

arXiv:2604.14001v1 Announce Type: new Abstract: Diffusion language models have recently emerged as a leading alternative to standard language models, due to their ability for bidirectional attention a

tutorialsarxiv-cs-cl
16 Apr 2026
Model Releases

Working Notes on Late Interaction Dynamics: Analyzing Targeted Behaviors of Late Interaction Models

DGX agent

arXiv:2603.26259v2 Announce Type: replace-cross Abstract: While Late Interaction models exhibit strong retrieval performance, many of their underlying dynamics remain understudied, potentially hiding

model-releasesarxiv-cs-cl
16 Apr 2026
Research

Causal Fingerprints of AI Generative Models

DGX agent

arXiv:2509.15406v2 Announce Type: replace Abstract: AI generative models leave implicit traces in their generated images, which are commonly referred to as model fingerprints and are exploited for sou

researcharxiv-cs-cv
15 Apr 2026
Model Releases

SeedPrints: Fingerprints Can Even Tell Which Seed Your Large Language Model Was Trained From

DGX agent

arXiv:2509.26404v2 Announce Type: replace-cross Abstract: Fingerprinting Large Language Models (LLMs)is essential for provenance verification and model attribution. Existing fingerprinting methods are

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music

DGX agent

arXiv:2604.10905v1 Announce Type: cross Abstract: We present Audio Flamingo Next (AF-Next), the next-generation and most capable large audio-language model in the Audio Flamingo series, designed to ad

model-releasesarxiv-cs-ai
14 Apr 2026
Safety

DeepFleet: Multi-Agent Foundation Models for Mobile Robots

DGX agent

arXiv:2508.08574v3 Announce Type: replace Abstract: We introduce DeepFleet, a suite of foundation models designed to support coordination and planning for large-scale mobile robot fleets. These models

safetyarxiv-cs-ro
14 Apr 2026
Model Releases

General-purpose LLMs as Models of Human Driver Behavior: The Case of Simplified Merging

DGX agent

arXiv:2604.09609v1 Announce Type: new Abstract: Human behavior models are essential as behavior references and for simulating human agents in virtual safety assessment of automated vehicles (AVs), yet

model-releasesarxiv-cs-ai
14 Apr 2026
Applications

Integrating SAINT with Tree-Based Models: A Case Study in Employee Attrition Prediction

DGX agent

arXiv:2604.10337v1 Announce Type: new Abstract: Employee attrition presents a major challenge for organizations, increasing costs and reducing productivity. Predicting attrition accurately enables pro

applicationsarxiv-cs-lg
14 Apr 2026
Research

Introspective Diffusion Language Models

DGX agent

arXiv:2604.11035v1 Announce Type: new Abstract: Diffusion language models promise parallel generation, yet still lag behind autoregressive (AR) models in quality. We stem this gap to a failure of intr

researcharxiv-cs-ai
14 Apr 2026
Model Releases

Pioneer Agent: Continual Improvement of Small Language Models in Production

DGX agent

arXiv:2604.09791v1 Announce Type: new Abstract: Small language models are attractive for production deployment due to their low cost, fast inference, and ease of specialization. However, adapting them

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

StyleBench: Evaluating thinking styles in Large Language Models

DGX agent

arXiv:2509.20868v2 Announce Type: replace-cross Abstract: Structured reasoning can improve the inference performance of large language models (LLMs), but it also introduces computational cost and cont

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

Teaching Language Models How to Code Like Learners: Conversational Serialization for Student Simulation

DGX agent

arXiv:2604.10720v1 Announce Type: new Abstract: Artificial models that simulate how learners act and respond within educational systems are a promising tool for evaluating tutoring strategies and feed

model-releasesarxiv-cs-ai
14 Apr 2026
Research

2D or 3D: Who Governs Salience in VLA Models? -- Tri-Stage Token Pruning Framework with Modality Salience Awareness

DGX agent

arXiv:2604.09244v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as the mainstream of embodied intelligence. Recent VLA models have expanded their input modalities fr

researcharxiv-cs-cv
13 Apr 2026
Model Releases

BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning

DGX agent

arXiv:2604.09378v1 Announce Type: cross Abstract: Agent ecosystems increasingly rely on installable skills to extend functionality, and some skills bundle learned model artifacts as part of their exec

model-releasesarxiv-cs-ai
13 Apr 2026
← Previous
1…1011121314…1012
Next →