AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
Human
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
61+ results
15 Apr 2026

Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models

Model ReleasesDGX agent

arXiv:2604.12391v1 Announce Type: cross Abstract: In this paper, we present Chain-of-Models Pre-Training (CoM-PT), a novel performance-lossless training acceleration method for vision foundation model

Modeling Co-Pilots for Text-to-Model Translation

AgentsDGX agent

arXiv:2604.12955v1 Announce Type: new Abstract: There is growing interest in leveraging large language models (LLMs) for text-to-model translation and optimization tasks. This paper aims to advance th

15 Jul 2026

Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs

DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2607.11997v1 Announce Type: cross Abstract: Multi-task model merging combines separately trained expert models into a single model that handles all tasks without co-training. Standard practice m

17 Apr 2026

Zero-Shot Retail Theft Detection via Orchestrated Vision Models: A Model-Agnostic, Cost-Effective Alternative to Trained Single-Model Systems

Model ReleasesDGX agent

arXiv:2604.14846v1 Announce Type: new Abstract: Retail theft costs the global economy over 100 billion annually, yet existing AI-based detection systems require expensive custom model training on prop

Large Language Models to Enhance Business Process Modeling: Past, Present, and Future Trends

ApplicationsDGX agent

arXiv:2604.14034v1 Announce Type: cross Abstract: Recent advances in Generative Artificial Intelligence, particularly Large Language Models (LLMs), have stimulated growing interest in automating or as

26 May 2026

InfiFPO: Implicit Model Fusion via Preference Optimization in Large Language Models

Model ReleasesDGX agent

arXiv:2505.13878v3 Announce Type: replace-cross Abstract: Model fusion combines multiple Large Language Models (LLMs) with different strengths into a more powerful, integrated model through lightweigh

11 Aug 2026

One Adapter Pair per Model: A Universal Activation Interface for Language Models

Model ReleasesDGX agent

arXiv:2608.09521v1 Announce Type: new Abstract: Activation-based tools are usually tied to one model's native hidden space, requiring probes, sparse autoencoders, and natural-language interpreters to

Cross-Model Humor Preference Modeling with Cards Against Humanity

Model ReleasesDGX agent

arXiv:2608.07481v1 Announce Type: cross Abstract: This paper investigates whether one large language model can approximate the humor preferences of another in a controlled Cards Against Humanity-style

Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models

Model ReleasesDGX agent

arXiv:2608.09696v1 Announce Type: new Abstract: Predicting the answer to interventional ``what if'' questions --- the outcome of an action never taken --- requires a mechanistic, causal model, not a c

9 Jun 2026

Model Multiplicity for Adversarial Detection in Small Language Model Training on Edge Devices

Model ReleasesDGX agent

arXiv:2606.07857v1 Announce Type: cross Abstract: The rise of edge-based machine learning has enabled distributed adaptation of language models across mobile and IoT devices, offering privacy preserva

20 May 2026

Exploring and Developing a Pre-Model Safeguard with Draft Models

SafetyDGX agent

arXiv:2605.19321v1 Announce Type: cross Abstract: Large Language Model (LLM) alignment remains vulnerable to jailbreak attacks that elicit unsafe responses, motivating pre-model and post-model guards.

30 Jun 2026

Assessing the Business Process Modeling Competences of Large Language Models

Model ReleasesDGX agent

arXiv:2601.21787v2 Announce Type: replace-cross Abstract: The creation of Business Process Model and Notation (BPMN) models is a complex and time-consuming task requiring both domain knowledge and pro

Model Merging to Evolution: Parameter Space Exploration for Expert Models

Model ReleasesDGX agent

arXiv:2606.28373v1 Announce Type: cross Abstract: Model merging integrates the capabilities of multiple expert models to create strong models for multiple tasks without additional training, thereby re

Alternative Graph Neural Networks: Synergizing GEV Models and Deep Learning for Travel Mode Choice Modeling

Model ReleasesDGX agent

arXiv:2509.07123v2 Announce Type: replace-cross Abstract: Generalized extreme value models capture dependence among choice alternatives in discrete choice modeling, but require this dependence to be p

OptiMUS-0.3: Using Large Language Models to Model and Solve Optimization Problems at Scale

Model ReleasesDGX agent

arXiv:2407.19633v4 Announce Type: replace Abstract: Optimization problems are pervasive in sectors from manufacturing and distribution to healthcare. However, most such problems are still solved heuri

24 Jun 2026

Beyond the Autoregressive Horizon: A Comprehensive Survey of Diffusion Models, World Modelling, and State Space Models for Code

ResearchDGX agent

arXiv:2606.23690v1 Announce Type: cross Abstract: Autoregressive (AR) language models have driven significant progress in automated software engineering, enabling powerful code generation and assistan

BIM-Edit: Benchmarking Large Language Models for IFC-Based Building Information Modeling

Model ReleasesDGX agent

arXiv:2606.20146v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly applied to computer-aided design (CAD) to generate design artifacts from textual instructions. In engi

12 May 2026

Threat Modelling using Domain-Adapted Language Models: Empirical Evaluation and Insights

ResearchDGX agent

arXiv:2605.10808v1 Announce Type: cross Abstract: Large Language Models(LLMs) are increasingly explored for cybersecurity applications such as vulnerability detection. In the domain of threat modellin

Tabular Foundation Model for Generative Modelling

Model ReleasesDGX agent

arXiv:2605.09424v1 Announce Type: new Abstract: Generative modelling is a demanding test of foundation models, because it requires robust, holistic representation learning for a given data modality, r

30 Apr 2026

Do Larger Models Really Win in Drug Discovery? A Benchmark Assessment of Model Scaling in AI-Driven Molecular Property and Activity Prediction

Model ReleasesDGX agent

arXiv:2604.26498v1 Announce Type: new Abstract: The rapid growth of molecular foundation models and general-purpose large language models has encouraged a scale-centric view of artificial intelligence

30 Jul 2026

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models

Model ReleasesDGX agent

arXiv:2607.26173v1 Announce Type: new Abstract: Alignment training, model organisms, and toy models are usually treated as separate research areas. But projects in all three frequently use supervised

Mergeable Model-Side Aggregation States for Long-Context Language Models

Model ReleasesDGX agent

arXiv:2607.26448v1 Announce Type: new Abstract: A known limitation of long-context language models is their increasingly unreliable performance in non-additive, set-based aggregation as context length

24 Jul 2026

Learn2Zinc: Fine-tuning Small Language Models for Text-to-Model Translation in MiniZinc

Model ReleasesDGX agent

arXiv:2607.20456v1 Announce Type: cross Abstract: Large language models excel at code generation for mainstream programming languages but struggle with rare, domain-specific languages such as MiniZinc

What, Where, and How: Disentangling the Roles of Task, Language, and Model in Code Model Representations

Model ReleasesDGX agent

arXiv:2607.21491v1 Announce Type: new Abstract: Do independently trained language models come to represent the same thing in the same way? We answer for code, extending a recently introduced concept-c

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving

Model ReleasesDGX agent

arXiv:2607.20988v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models augmented with world modeling represent a promising paradigm for end-to-end autonomous driving. While pixel-level

28 May 2026

Continual Model Routing in Evolving Model Hubs

Model ReleasesDGX agent

arXiv:2605.28577v1 Announce Type: new Abstract: AI model hubs provide access to a rapidly growing collection of powerful pre-trained models, enabling off-the-shelf mixture-of-experts systems with diff

23 Apr 2026

Text to model via SysML: Automated generation of dynamical system computational models from unstructured natural language text via enhanced System Modeling Language diagrams

ResearchDGX agent

arXiv:2507.06803v3 Announce Type: replace-cross Abstract: This paper contributes to speeding up the design and deployment of engineering dynamical systems by proposing a strategy for exploiting domain

14 Apr 2026

Bridging What the Model Thinks and How It Speaks: Self-Aware Speech Language Models for Expressive Speech Generation

Model ReleasesDGX agent

arXiv:2604.11424v1 Announce Type: new Abstract: Speech Language Models (SLMs) exhibit strong semantic understanding, yet their generated speech often sounds flat and fails to convey expressive intent,

Reason Only When Needed: Efficient Generative Reward Modeling via Model-Internal Uncertainty

ResearchDGX agent

arXiv:2604.10072v1 Announce Type: new Abstract: Recent advancements in the Generative Reward Model (GRM) have demonstrated its potential to enhance the reasoning abilities of LLMs through Chain-of-Tho

31 Jul 2026

Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models

Model ReleasesDGX agent

arXiv:2607.27506v1 Announce Type: new Abstract: Language and embedding models used in RAG systems are conventionally assumed to require large-scale pretraining and explicit grounding supervision. We p

Unifying Adversarially Robust Model Experts in Vision-Language Models

SafetyDGX agent

arXiv:2607.27897v1 Announce Type: new Abstract: Vision-language models (VLMs), such as CLIP, are vulnerable to adversarial attacks, posing a serious problem for real-life applications and deployment.

28 Jul 2026

Foundation Models and Fine-Tuning: Toward a New Generation of Models for Time Series Forecasting

ResearchDGX agent

arXiv:2607.23146v1 Announce Type: new Abstract: Inspired by recent breakthroughs in large language models for natural language processing, foundation models have emerged as a promising paradigm for ze

BERT-based Models vs. Large Language Models for Low-Resource Named Entity Recognition: A Comparative Study on Marathi

Model ReleasesDGX agent

arXiv:2607.23344v1 Announce Type: new Abstract: Named Entity Recognition (NER) for low-resource languages such as Marathi remains a challenging task due to limited annotated resources and linguistic c

8 Jul 2026

Base Models Know How to Reason, Thinking Models Learn When

TutorialsDGX agent

arXiv:2510.07364v4 Announce Type: replace Abstract: What do thinking language models learn during training that their base models lack? We first present an unsupervised method that discovers a model's

7 Jul 2026

WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling

Model ReleasesDGX agent

arXiv:2607.03461v1 Announce Type: new Abstract: World models aim to capture environment dynamics in ways that support perception, reasoning, and action, and have recently become a central direction in

Evaluating and Understanding Model Editing for Medical Vision Language Models

Model ReleasesDGX agent

arXiv:2607.05310v1 Announce Type: new Abstract: Model editing promises a fast, targeted way to correct post-deployment mistakes in medical vision-language models (VLMs) without costly retraining. Howe

1 Jul 2026

Knowledge Distillation from Large Reasoning Models to Compact Student Models: A Case Study on the John O Bryan Mathematics Competition

Model ReleasesDGX agent

arXiv:2606.31048v1 Announce Type: cross Abstract: This paper investigates knowledge distillation from a large reasoning model (DeepSeek-R1) to a compact student model (Qwen2.5-7B). Using historical pr

How Can AI Find My Model? A Model-Finding Experimental Study Considering Data Formats, Embeddings, and Retrieval Strategies

ResearchDGX agent

arXiv:2606.30846v1 Announce Type: new Abstract: Discovering simulation models for reuse remains a fundamental challenge in Modeling and Simulation (M&S). When many models coexist, identifying those th

Rethinking Foundation Model Collaboration: Enhancing Specialized Models through Proxy Task Reasoning

ResearchDGX agent

arXiv:2606.31157v1 Announce Type: new Abstract: Foundation models are increasingly integrated into embodied intelligence systems, but directly assigning them structured prediction tasks requires preci

26 Jun 2026

When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models

SafetyDGX agent

arXiv:2606.27288v1 Announce Type: new Abstract: Multi-model LLM systems such as routing, voting, cascades, fusion, and mixture-of-agents are used to beat single-model accuracy. We show that their gain

10 Jun 2026

Beyond Model Size: Probing the Gaps in Visual in-Context Learning by Training a Tiny Model

Model ReleasesDGX agent

arXiv:2606.10905v1 Announce Type: new Abstract: Visual in-Context Learning (VICL) aims at making progress towards adaptive vision models, that can -- based on a few examples -- adapt to a new task at

5 Jun 2026

An Embarrassingly Simple Detector for Model Extraction Attacks in Large Language Model API Traffic

Model ReleasesDGX agent

arXiv:2606.05725v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed through hosted APIs, making model extraction a practical threat to model ownership and service

2 Jun 2026

IMWM: Intuition Models Complement World Models for Latent Planning

ResearchDGX agent

arXiv:2606.01626v1 Announce Type: new Abstract: Planning with a learned latent world model is a promising route to control from raw pixels, but a strong world model alone is not enough. We show this e

18 May 2026

SAE-RNA: A Sparse Autoencoder Model for Interpreting RNA Language Model Representations

ResearchDGX agent

arXiv:2510.02734v2 Announce Type: replace-cross Abstract: Deep learning, particularly with the advancement of Large Language Models, has transformed biomolecular modeling, with protein language models

Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training

Model ReleasesDGX agent

arXiv:2602.00747v2 Announce Type: replace-cross Abstract: Determining an effective data mixture is a key factor in Large Language Model (LLM) pre-training, where models must balance general competence

6 May 2026

When Should a Language Model Trust Itself? Same-Model Self-Verification as a Conditional Confidence Signal

Model ReleasesDGX agent

arXiv:2605.02915v1 Announce Type: new Abstract: Same-model self-verification, prompting a model to audit its own predicted answer, is a plausible confidence signal for selective prediction, but its pr

27 Apr 2026

Pre-trained Large Language Models Learn Hidden Markov Models In-context

TutorialsDGX agent

arXiv:2506.07298v3 Announce Type: replace-cross Abstract: Hidden Markov Models (HMMs) are foundational tools for modeling sequential data with latent Markovian structure, yet fitting them to real-worl

10 Apr 2026

Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization

Model ReleasesDGX agent

arXiv:2604.07343v1 Announce Type: cross Abstract: Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a cen

10 Aug 2026

Multi-Level Modeling of Large Language Model Inference Latency and Energy via Hybrid Analytical--Machine-Learning Predictors

Model ReleasesDGX agent

arXiv:2608.06723v1 Announce Type: cross Abstract: The rapid scaling of Large Language Models (LLMs) has significantly increased computational cost, energy consumption, and inference latency, making ac

Beyond Foundation Models: Dimension-Aware Neural Architecture Search with Small-Data Representation Models for Cryocooler Lifetime Prediction

Model ReleasesDGX agent

arXiv:2608.06993v1 Announce Type: cross Abstract: Large-scale pretrained time-series models achieve strong results through large-scale pretraining and task-agnostic representation learning, but they r

Transformers Struggle to Use Their Emergent World Models: Revisiting the Tower of Hanoi, and the Illusion of Thinking

Model ReleasesDGX agent

arXiv:2608.07077v1 Announce Type: new Abstract: The Tower of Hanoi is a simple planning puzzle that in prior work has proven challenging for large reasoning models (LRMs). Current models solve the sta

9 Jul 2026

Protocol Models: Scaling Decentralized Training with Communication-Efficient Model Parallelism

Model ReleasesDGX agent

arXiv:2506.01260v2 Announce Type: replace Abstract: Scaling models has led to significant advancements in deep learning, but training these models in decentralized settings remains challenging due to

8 Jun 2026

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling

SafetyDGX agent

arXiv:2507.06419v3 Announce Type: replace Abstract: Reward modeling (RM), which captures human preferences to align large language models (LLMs), is increasingly employed in tasks such as model finetu

4 Jun 2026

Speculative Thinking: Enhancing Small-Model Reasoning with Large Model Guidance at Inference Time

Model ReleasesDGX agent

arXiv:2504.12329v2 Announce Type: replace-cross Abstract: Recent advances leverage post-training to enhance model reasoning performance, which typically requires costly training pipelines and still su

7 May 2026

Cross-Model Consistency of Feature Importance in Electrospinning: Separating Robust from Model-Dependent Features

Model ReleasesDGX agent

arXiv:2605.04905v1 Announce Type: new Abstract: Electrospinning is a highly sensitive fabrication process in which small variations in operating parameters can significantly influence fiber morphology

5 May 2026

Model-Dowser: Data-Free Importance Probing to Mitigate Catastrophic Forgetting in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2602.04509v4 Announce Type: replace Abstract: Fine-tuning Multimodal Large Language Models (MLLMs) on task-specific data is an effective way to improve performance on downstream applications. Ho

12 Aug 2026

On Solomonoff Induction in Large Language Models and the Limits of Self-Improving: The Singularity Is Not Near Without Symbolic Model Synthesis

Model ReleasesDGX agent

arXiv:2601.05280v3 Announce Type: replace-cross Abstract: On the one hand, the question of whether large language models (LLMs) are Solomonoff induction estimators has become an explicit question at t

What Iterated Self-Feeding Probes of Language Models Measure, and a test that separates the construction from the model

AgentsDGX agent

arXiv:2608.10986v1 Announce Type: new Abstract: A growing class of methods probes a language model by feeding it its own output: self-consistency, iterated refinement, agentic loops. We ask what such

7 Aug 2026

Cross-Architecture Steering Transfer in Language Models: A Systematic Empirical Study

Model ReleasesDGX agent

arXiv:2608.05164v1 Announce Type: new Abstract: Independently trained large language models may develop shared internal representations of semantic concepts despite architectural differences -- but wh

4 Aug 2026

AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling

Model ReleasesDGX agent

arXiv:2608.02602v1 Announce Type: new Abstract: Language remains an outlier in generative modeling: while images, video, and audio are increasingly modeled in continuous latent spaces, text generation

← Previous
1
Next →
48,543 results
← Previous
123…810
Next →