AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
Human
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
49+ results
Model Releases

Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models

DGX agent

arXiv:2604.12391v1 Announce Type: cross Abstract: In this paper, we present Chain-of-Models Pre-Training (CoM-PT), a novel performance-lossless training acceleration method for vision foundation model

model-releasesarxiv-cs-ai
15 Apr 2026
DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs

DGX agent

arXiv:2607.11997v1 Announce Type: cross Abstract: Multi-task model merging combines separately trained expert models into a single model that handles all tasks without co-training. Standard practice m

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Zero-Shot Retail Theft Detection via Orchestrated Vision Models: A Model-Agnostic, Cost-Effective Alternative to Trained Single-Model Systems

DGX agent

arXiv:2604.14846v1 Announce Type: new Abstract: Retail theft costs the global economy over 100 billion annually, yet existing AI-based detection systems require expensive custom model training on prop

model-releasesarxiv-cs-cv
17 Apr 2026
Model Releases

InfiFPO: Implicit Model Fusion via Preference Optimization in Large Language Models

DGX agent

arXiv:2505.13878v3 Announce Type: replace-cross Abstract: Model fusion combines multiple Large Language Models (LLMs) with different strengths into a more powerful, integrated model through lightweigh

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

One Adapter Pair per Model: A Universal Activation Interface for Language Models

DGX agent

arXiv:2608.09521v1 Announce Type: new Abstract: Activation-based tools are usually tied to one model's native hidden space, requiring probes, sparse autoencoders, and natural-language interpreters to

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Model Multiplicity for Adversarial Detection in Small Language Model Training on Edge Devices

DGX agent

arXiv:2606.07857v1 Announce Type: cross Abstract: The rise of edge-based machine learning has enabled distributed adaptation of language models across mobile and IoT devices, offering privacy preserva

model-releasesarxiv-cs-ai
9 Jun 2026
Safety

Exploring and Developing a Pre-Model Safeguard with Draft Models

DGX agent

arXiv:2605.19321v1 Announce Type: cross Abstract: Large Language Model (LLM) alignment remains vulnerable to jailbreak attacks that elicit unsafe responses, motivating pre-model and post-model guards.

safetyarxiv-cs-ai
20 May 2026
Model Releases

Assessing the Business Process Modeling Competences of Large Language Models

DGX agent

arXiv:2601.21787v2 Announce Type: replace-cross Abstract: The creation of Business Process Model and Notation (BPMN) models is a complex and time-consuming task requiring both domain knowledge and pro

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Model Merging to Evolution: Parameter Space Exploration for Expert Models

DGX agent

arXiv:2606.28373v1 Announce Type: cross Abstract: Model merging integrates the capabilities of multiple expert models to create strong models for multiple tasks without additional training, thereby re

model-releasesarxiv-cs-ai
30 Jun 2026
Research

Beyond the Autoregressive Horizon: A Comprehensive Survey of Diffusion Models, World Modelling, and State Space Models for Code

DGX agent

arXiv:2606.23690v1 Announce Type: cross Abstract: Autoregressive (AR) language models have driven significant progress in automated software engineering, enabling powerful code generation and assistan

researcharxiv-cs-ai
24 Jun 2026
Research

Threat Modelling using Domain-Adapted Language Models: Empirical Evaluation and Insights

DGX agent

arXiv:2605.10808v1 Announce Type: cross Abstract: Large Language Models(LLMs) are increasingly explored for cybersecurity applications such as vulnerability detection. In the domain of threat modellin

researcharxiv-cs-ai
12 May 2026
Model Releases

Do Larger Models Really Win in Drug Discovery? A Benchmark Assessment of Model Scaling in AI-Driven Molecular Property and Activity Prediction

DGX agent

arXiv:2604.26498v1 Announce Type: new Abstract: The rapid growth of molecular foundation models and general-purpose large language models has encouraged a scale-centric view of artificial intelligence

model-releasesarxiv-cs-lg
30 Apr 2026
Model Releases

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models

DGX agent

arXiv:2607.26173v1 Announce Type: new Abstract: Alignment training, model organisms, and toy models are usually treated as separate research areas. But projects in all three frequently use supervised

model-releasesarxiv-cs-lg
30 Jul 2026
Applications

Large Language Models to Enhance Business Process Modeling: Past, Present, and Future Trends

DGX agent

arXiv:2604.14034v1 Announce Type: cross Abstract: Recent advances in Generative Artificial Intelligence, particularly Large Language Models (LLMs), have stimulated growing interest in automating or as

applicationsarxiv-cs-ai
17 Apr 2026
Model Releases

Learn2Zinc: Fine-tuning Small Language Models for Text-to-Model Translation in MiniZinc

DGX agent

arXiv:2607.20456v1 Announce Type: cross Abstract: Large language models excel at code generation for mainstream programming languages but struggle with rare, domain-specific languages such as MiniZinc

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Alternative Graph Neural Networks: Synergizing GEV Models and Deep Learning for Travel Mode Choice Modeling

DGX agent

arXiv:2509.07123v2 Announce Type: replace-cross Abstract: Generalized extreme value models capture dependence among choice alternatives in discrete choice modeling, but require this dependence to be p

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

Continual Model Routing in Evolving Model Hubs

DGX agent

arXiv:2605.28577v1 Announce Type: new Abstract: AI model hubs provide access to a rapidly growing collection of powerful pre-trained models, enabling off-the-shelf mixture-of-experts systems with diff

model-releasesarxiv-cs-ai
28 May 2026
Research

Text to model via SysML: Automated generation of dynamical system computational models from unstructured natural language text via enhanced System Modeling Language diagrams

DGX agent

arXiv:2507.06803v3 Announce Type: replace-cross Abstract: This paper contributes to speeding up the design and deployment of engineering dynamical systems by proposing a strategy for exploiting domain

researcharxiv-cs-ai
23 Apr 2026
Agents

Modeling Co-Pilots for Text-to-Model Translation

DGX agent

arXiv:2604.12955v1 Announce Type: new Abstract: There is growing interest in leveraging large language models (LLMs) for text-to-model translation and optimization tasks. This paper aims to advance th

agentsarxiv-cs-ai
15 Apr 2026
Model Releases

Bridging What the Model Thinks and How It Speaks: Self-Aware Speech Language Models for Expressive Speech Generation

DGX agent

arXiv:2604.11424v1 Announce Type: new Abstract: Speech Language Models (SLMs) exhibit strong semantic understanding, yet their generated speech often sounds flat and fails to convey expressive intent,

model-releasesarxiv-cs-cl
14 Apr 2026
Model Releases

Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models

DGX agent

arXiv:2607.27506v1 Announce Type: new Abstract: Language and embedding models used in RAG systems are conventionally assumed to require large-scale pretraining and explicit grounding supervision. We p

model-releasesarxiv-cs-cl
31 Jul 2026
Safety

Unifying Adversarially Robust Model Experts in Vision-Language Models

DGX agent

arXiv:2607.27897v1 Announce Type: new Abstract: Vision-language models (VLMs), such as CLIP, are vulnerable to adversarial attacks, posing a serious problem for real-life applications and deployment.

safetyarxiv-cs-cv
31 Jul 2026
Research

Foundation Models and Fine-Tuning: Toward a New Generation of Models for Time Series Forecasting

DGX agent

arXiv:2607.23146v1 Announce Type: new Abstract: Inspired by recent breakthroughs in large language models for natural language processing, foundation models have emerged as a promising paradigm for ze

researcharxiv-cs-lg
28 Jul 2026
Model Releases

What, Where, and How: Disentangling the Roles of Task, Language, and Model in Code Model Representations

DGX agent

arXiv:2607.21491v1 Announce Type: new Abstract: Do independently trained language models come to represent the same thing in the same way? We answer for code, extending a recently introduced concept-c

model-releasesarxiv-cs-cl
24 Jul 2026
Tutorials

Base Models Know How to Reason, Thinking Models Learn When

DGX agent

arXiv:2510.07364v4 Announce Type: replace Abstract: What do thinking language models learn during training that their base models lack? We first present an unsupervised method that discovers a model's

tutorialsarxiv-cs-ai
8 Jul 2026
Model Releases

WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling

DGX agent

arXiv:2607.03461v1 Announce Type: new Abstract: World models aim to capture environment dynamics in ways that support perception, reasoning, and action, and have recently become a central direction in

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

Knowledge Distillation from Large Reasoning Models to Compact Student Models: A Case Study on the John O Bryan Mathematics Competition

DGX agent

arXiv:2606.31048v1 Announce Type: cross Abstract: This paper investigates knowledge distillation from a large reasoning model (DeepSeek-R1) to a compact student model (Qwen2.5-7B). Using historical pr

model-releasesarxiv-cs-ai
1 Jul 2026
Safety

When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models

DGX agent

arXiv:2606.27288v1 Announce Type: new Abstract: Multi-model LLM systems such as routing, voting, cascades, fusion, and mixture-of-agents are used to beat single-model accuracy. We show that their gain

safetyarxiv-cs-ai
26 Jun 2026
Model Releases

Beyond Model Size: Probing the Gaps in Visual in-Context Learning by Training a Tiny Model

DGX agent

arXiv:2606.10905v1 Announce Type: new Abstract: Visual in-Context Learning (VICL) aims at making progress towards adaptive vision models, that can -- based on a few examples -- adapt to a new task at

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

An Embarrassingly Simple Detector for Model Extraction Attacks in Large Language Model API Traffic

DGX agent

arXiv:2606.05725v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed through hosted APIs, making model extraction a practical threat to model ownership and service

model-releasesarxiv-cs-cl
5 Jun 2026
Research

IMWM: Intuition Models Complement World Models for Latent Planning

DGX agent

arXiv:2606.01626v1 Announce Type: new Abstract: Planning with a learned latent world model is a promising route to control from raw pixels, but a strong world model alone is not enough. We show this e

researcharxiv-cs-lg
2 Jun 2026
Research

SAE-RNA: A Sparse Autoencoder Model for Interpreting RNA Language Model Representations

DGX agent

arXiv:2510.02734v2 Announce Type: replace-cross Abstract: Deep learning, particularly with the advancement of Large Language Models, has transformed biomolecular modeling, with protein language models

researcharxiv-cs-ai
18 May 2026
Model Releases

Tabular Foundation Model for Generative Modelling

DGX agent

arXiv:2605.09424v1 Announce Type: new Abstract: Generative modelling is a demanding test of foundation models, because it requires robust, holistic representation learning for a given data modality, r

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

When Should a Language Model Trust Itself? Same-Model Self-Verification as a Conditional Confidence Signal

DGX agent

arXiv:2605.02915v1 Announce Type: new Abstract: Same-model self-verification, prompting a model to audit its own predicted answer, is a plausible confidence signal for selective prediction, but its pr

model-releasesarxiv-cs-cl
6 May 2026
Tutorials

Pre-trained Large Language Models Learn Hidden Markov Models In-context

DGX agent

arXiv:2506.07298v3 Announce Type: replace-cross Abstract: Hidden Markov Models (HMMs) are foundational tools for modeling sequential data with latent Markovian structure, yet fitting them to real-worl

tutorialsarxiv-cs-ai
27 Apr 2026
Research

Reason Only When Needed: Efficient Generative Reward Modeling via Model-Internal Uncertainty

DGX agent

arXiv:2604.10072v1 Announce Type: new Abstract: Recent advancements in the Generative Reward Model (GRM) have demonstrated its potential to enhance the reasoning abilities of LLMs through Chain-of-Tho

researcharxiv-cs-cl
14 Apr 2026
Model Releases

Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization

DGX agent

arXiv:2604.07343v1 Announce Type: cross Abstract: Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a cen

model-releasesarxiv-cs-lg
10 Apr 2026
Model Releases

Cross-Model Humor Preference Modeling with Cards Against Humanity

DGX agent

arXiv:2608.07481v1 Announce Type: cross Abstract: This paper investigates whether one large language model can approximate the humor preferences of another in a controlled Cards Against Humanity-style

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Multi-Level Modeling of Large Language Model Inference Latency and Energy via Hybrid Analytical--Machine-Learning Predictors

DGX agent

arXiv:2608.06723v1 Announce Type: cross Abstract: The rapid scaling of Large Language Models (LLMs) has significantly increased computational cost, energy consumption, and inference latency, making ac

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

Mergeable Model-Side Aggregation States for Long-Context Language Models

DGX agent

arXiv:2607.26448v1 Announce Type: new Abstract: A known limitation of long-context language models is their increasingly unreliable performance in non-additive, set-based aggregation as context length

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving

DGX agent

arXiv:2607.20988v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models augmented with world modeling represent a promising paradigm for end-to-end autonomous driving. While pixel-level

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Protocol Models: Scaling Decentralized Training with Communication-Efficient Model Parallelism

DGX agent

arXiv:2506.01260v2 Announce Type: replace Abstract: Scaling models has led to significant advancements in deep learning, but training these models in decentralized settings remains challenging due to

model-releasesarxiv-cs-lg
9 Jul 2026
Research

How Can AI Find My Model? A Model-Finding Experimental Study Considering Data Formats, Embeddings, and Retrieval Strategies

DGX agent

arXiv:2606.30846v1 Announce Type: new Abstract: Discovering simulation models for reuse remains a fundamental challenge in Modeling and Simulation (M&S). When many models coexist, identifying those th

researcharxiv-cs-ai
1 Jul 2026
Model Releases

OptiMUS-0.3: Using Large Language Models to Model and Solve Optimization Problems at Scale

DGX agent

arXiv:2407.19633v4 Announce Type: replace Abstract: Optimization problems are pervasive in sectors from manufacturing and distribution to healthcare. However, most such problems are still solved heuri

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

BIM-Edit: Benchmarking Large Language Models for IFC-Based Building Information Modeling

DGX agent

arXiv:2606.20146v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly applied to computer-aided design (CAD) to generate design artifacts from textual instructions. In engi

model-releasesarxiv-cs-ai
24 Jun 2026
Safety

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling

DGX agent

arXiv:2507.06419v3 Announce Type: replace Abstract: Reward modeling (RM), which captures human preferences to align large language models (LLMs), is increasingly employed in tasks such as model finetu

safetyarxiv-cs-cl
8 Jun 2026
Model Releases

Speculative Thinking: Enhancing Small-Model Reasoning with Large Model Guidance at Inference Time

DGX agent

arXiv:2504.12329v2 Announce Type: replace-cross Abstract: Recent advances leverage post-training to enhance model reasoning performance, which typically requires costly training pipelines and still su

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training

DGX agent

arXiv:2602.00747v2 Announce Type: replace-cross Abstract: Determining an effective data mixture is a key factor in Large Language Model (LLM) pre-training, where models must balance general competence

model-releasesarxiv-cs-ai
18 May 2026
← Previous
1
Next →
48,543 results
← Previous
123…1012
Next →