AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
49+ results
Model Releases

Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models

DGX agent

arXiv:2604.12391v1 Announce Type: cross Abstract: In this paper, we present Chain-of-Models Pre-Training (CoM-PT), a novel performance-lossless training acceleration method for vision foundation model

model-releasesarxiv-cs-ai
15 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs

DGX agent

arXiv:2607.11997v1 Announce Type: cross Abstract: Multi-task model merging combines separately trained expert models into a single model that handles all tasks without co-training. Standard practice m

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Zero-Shot Retail Theft Detection via Orchestrated Vision Models: A Model-Agnostic, Cost-Effective Alternative to Trained Single-Model Systems

DGX agent

arXiv:2604.14846v1 Announce Type: new Abstract: Retail theft costs the global economy over 100 billion annually, yet existing AI-based detection systems require expensive custom model training on prop

model-releasesarxiv-cs-cv
17 Apr 2026
Model Releases

InfiFPO: Implicit Model Fusion via Preference Optimization in Large Language Models

DGX agent

arXiv:2505.13878v3 Announce Type: replace-cross Abstract: Model fusion combines multiple Large Language Models (LLMs) with different strengths into a more powerful, integrated model through lightweigh

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

One Adapter Pair per Model: A Universal Activation Interface for Language Models

DGX agent

arXiv:2608.09521v1 Announce Type: new Abstract: Activation-based tools are usually tied to one model's native hidden space, requiring probes, sparse autoencoders, and natural-language interpreters to

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Model Multiplicity for Adversarial Detection in Small Language Model Training on Edge Devices

DGX agent

arXiv:2606.07857v1 Announce Type: cross Abstract: The rise of edge-based machine learning has enabled distributed adaptation of language models across mobile and IoT devices, offering privacy preserva

model-releasesarxiv-cs-ai
9 Jun 2026
Safety

Exploring and Developing a Pre-Model Safeguard with Draft Models

DGX agent

arXiv:2605.19321v1 Announce Type: cross Abstract: Large Language Model (LLM) alignment remains vulnerable to jailbreak attacks that elicit unsafe responses, motivating pre-model and post-model guards.

safetyarxiv-cs-ai
20 May 2026
Model Releases

Assessing the Business Process Modeling Competences of Large Language Models

DGX agent

arXiv:2601.21787v2 Announce Type: replace-cross Abstract: The creation of Business Process Model and Notation (BPMN) models is a complex and time-consuming task requiring both domain knowledge and pro

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Model Merging to Evolution: Parameter Space Exploration for Expert Models

DGX agent

arXiv:2606.28373v1 Announce Type: cross Abstract: Model merging integrates the capabilities of multiple expert models to create strong models for multiple tasks without additional training, thereby re

model-releasesarxiv-cs-ai
30 Jun 2026
Research

Beyond the Autoregressive Horizon: A Comprehensive Survey of Diffusion Models, World Modelling, and State Space Models for Code

DGX agent

arXiv:2606.23690v1 Announce Type: cross Abstract: Autoregressive (AR) language models have driven significant progress in automated software engineering, enabling powerful code generation and assistan

researcharxiv-cs-ai
24 Jun 2026
Research

Threat Modelling using Domain-Adapted Language Models: Empirical Evaluation and Insights

DGX agent

arXiv:2605.10808v1 Announce Type: cross Abstract: Large Language Models(LLMs) are increasingly explored for cybersecurity applications such as vulnerability detection. In the domain of threat modellin

researcharxiv-cs-ai
12 May 2026
Model Releases

Do Larger Models Really Win in Drug Discovery? A Benchmark Assessment of Model Scaling in AI-Driven Molecular Property and Activity Prediction

DGX agent

arXiv:2604.26498v1 Announce Type: new Abstract: The rapid growth of molecular foundation models and general-purpose large language models has encouraged a scale-centric view of artificial intelligence

model-releasesarxiv-cs-lg
30 Apr 2026
Model Releases

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models

DGX agent

arXiv:2607.26173v1 Announce Type: new Abstract: Alignment training, model organisms, and toy models are usually treated as separate research areas. But projects in all three frequently use supervised

model-releasesarxiv-cs-lg
30 Jul 2026
Applications

Large Language Models to Enhance Business Process Modeling: Past, Present, and Future Trends

DGX agent

arXiv:2604.14034v1 Announce Type: cross Abstract: Recent advances in Generative Artificial Intelligence, particularly Large Language Models (LLMs), have stimulated growing interest in automating or as

applicationsarxiv-cs-ai
17 Apr 2026
Model Releases

Learn2Zinc: Fine-tuning Small Language Models for Text-to-Model Translation in MiniZinc

DGX agent

arXiv:2607.20456v1 Announce Type: cross Abstract: Large language models excel at code generation for mainstream programming languages but struggle with rare, domain-specific languages such as MiniZinc

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Alternative Graph Neural Networks: Synergizing GEV Models and Deep Learning for Travel Mode Choice Modeling

DGX agent

arXiv:2509.07123v2 Announce Type: replace-cross Abstract: Generalized extreme value models capture dependence among choice alternatives in discrete choice modeling, but require this dependence to be p

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

Continual Model Routing in Evolving Model Hubs

DGX agent

arXiv:2605.28577v1 Announce Type: new Abstract: AI model hubs provide access to a rapidly growing collection of powerful pre-trained models, enabling off-the-shelf mixture-of-experts systems with diff

model-releasesarxiv-cs-ai
28 May 2026
Research

Text to model via SysML: Automated generation of dynamical system computational models from unstructured natural language text via enhanced System Modeling Language diagrams

DGX agent

arXiv:2507.06803v3 Announce Type: replace-cross Abstract: This paper contributes to speeding up the design and deployment of engineering dynamical systems by proposing a strategy for exploiting domain

researcharxiv-cs-ai
23 Apr 2026
Agents

Modeling Co-Pilots for Text-to-Model Translation

DGX agent

arXiv:2604.12955v1 Announce Type: new Abstract: There is growing interest in leveraging large language models (LLMs) for text-to-model translation and optimization tasks. This paper aims to advance th

agentsarxiv-cs-ai
15 Apr 2026
Model Releases

Bridging What the Model Thinks and How It Speaks: Self-Aware Speech Language Models for Expressive Speech Generation

DGX agent

arXiv:2604.11424v1 Announce Type: new Abstract: Speech Language Models (SLMs) exhibit strong semantic understanding, yet their generated speech often sounds flat and fails to convey expressive intent,

model-releasesarxiv-cs-cl
14 Apr 2026
Model Releases

Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models

DGX agent

arXiv:2607.27506v1 Announce Type: new Abstract: Language and embedding models used in RAG systems are conventionally assumed to require large-scale pretraining and explicit grounding supervision. We p

model-releasesarxiv-cs-cl
31 Jul 2026
Safety

Unifying Adversarially Robust Model Experts in Vision-Language Models

DGX agent

arXiv:2607.27897v1 Announce Type: new Abstract: Vision-language models (VLMs), such as CLIP, are vulnerable to adversarial attacks, posing a serious problem for real-life applications and deployment.

safetyarxiv-cs-cv
31 Jul 2026
Research

Foundation Models and Fine-Tuning: Toward a New Generation of Models for Time Series Forecasting

DGX agent

arXiv:2607.23146v1 Announce Type: new Abstract: Inspired by recent breakthroughs in large language models for natural language processing, foundation models have emerged as a promising paradigm for ze

researcharxiv-cs-lg
28 Jul 2026
Model Releases

What, Where, and How: Disentangling the Roles of Task, Language, and Model in Code Model Representations

DGX agent

arXiv:2607.21491v1 Announce Type: new Abstract: Do independently trained language models come to represent the same thing in the same way? We answer for code, extending a recently introduced concept-c

model-releasesarxiv-cs-cl
24 Jul 2026
Tutorials

Base Models Know How to Reason, Thinking Models Learn When

DGX agent

arXiv:2510.07364v4 Announce Type: replace Abstract: What do thinking language models learn during training that their base models lack? We first present an unsupervised method that discovers a model's

tutorialsarxiv-cs-ai
8 Jul 2026
Model Releases

WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling

DGX agent

arXiv:2607.03461v1 Announce Type: new Abstract: World models aim to capture environment dynamics in ways that support perception, reasoning, and action, and have recently become a central direction in

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

Knowledge Distillation from Large Reasoning Models to Compact Student Models: A Case Study on the John O Bryan Mathematics Competition

DGX agent

arXiv:2606.31048v1 Announce Type: cross Abstract: This paper investigates knowledge distillation from a large reasoning model (DeepSeek-R1) to a compact student model (Qwen2.5-7B). Using historical pr

model-releasesarxiv-cs-ai
1 Jul 2026
Safety

When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models

DGX agent

arXiv:2606.27288v1 Announce Type: new Abstract: Multi-model LLM systems such as routing, voting, cascades, fusion, and mixture-of-agents are used to beat single-model accuracy. We show that their gain

safetyarxiv-cs-ai
26 Jun 2026
Model Releases

Beyond Model Size: Probing the Gaps in Visual in-Context Learning by Training a Tiny Model

DGX agent

arXiv:2606.10905v1 Announce Type: new Abstract: Visual in-Context Learning (VICL) aims at making progress towards adaptive vision models, that can -- based on a few examples -- adapt to a new task at

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

An Embarrassingly Simple Detector for Model Extraction Attacks in Large Language Model API Traffic

DGX agent

arXiv:2606.05725v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed through hosted APIs, making model extraction a practical threat to model ownership and service

model-releasesarxiv-cs-cl
5 Jun 2026
Research

IMWM: Intuition Models Complement World Models for Latent Planning

DGX agent

arXiv:2606.01626v1 Announce Type: new Abstract: Planning with a learned latent world model is a promising route to control from raw pixels, but a strong world model alone is not enough. We show this e

researcharxiv-cs-lg
2 Jun 2026
Research

SAE-RNA: A Sparse Autoencoder Model for Interpreting RNA Language Model Representations

DGX agent

arXiv:2510.02734v2 Announce Type: replace-cross Abstract: Deep learning, particularly with the advancement of Large Language Models, has transformed biomolecular modeling, with protein language models

researcharxiv-cs-ai
18 May 2026
Model Releases

Tabular Foundation Model for Generative Modelling

DGX agent

arXiv:2605.09424v1 Announce Type: new Abstract: Generative modelling is a demanding test of foundation models, because it requires robust, holistic representation learning for a given data modality, r

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

When Should a Language Model Trust Itself? Same-Model Self-Verification as a Conditional Confidence Signal

DGX agent

arXiv:2605.02915v1 Announce Type: new Abstract: Same-model self-verification, prompting a model to audit its own predicted answer, is a plausible confidence signal for selective prediction, but its pr

model-releasesarxiv-cs-cl
6 May 2026
Tutorials

Pre-trained Large Language Models Learn Hidden Markov Models In-context

DGX agent

arXiv:2506.07298v3 Announce Type: replace-cross Abstract: Hidden Markov Models (HMMs) are foundational tools for modeling sequential data with latent Markovian structure, yet fitting them to real-worl

tutorialsarxiv-cs-ai
27 Apr 2026
Research

Reason Only When Needed: Efficient Generative Reward Modeling via Model-Internal Uncertainty

DGX agent

arXiv:2604.10072v1 Announce Type: new Abstract: Recent advancements in the Generative Reward Model (GRM) have demonstrated its potential to enhance the reasoning abilities of LLMs through Chain-of-Tho

researcharxiv-cs-cl
14 Apr 2026
Model Releases

Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization

DGX agent

arXiv:2604.07343v1 Announce Type: cross Abstract: Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a cen

model-releasesarxiv-cs-lg
10 Apr 2026
Model Releases

Cross-Model Humor Preference Modeling with Cards Against Humanity

DGX agent

arXiv:2608.07481v1 Announce Type: cross Abstract: This paper investigates whether one large language model can approximate the humor preferences of another in a controlled Cards Against Humanity-style

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Multi-Level Modeling of Large Language Model Inference Latency and Energy via Hybrid Analytical--Machine-Learning Predictors

DGX agent

arXiv:2608.06723v1 Announce Type: cross Abstract: The rapid scaling of Large Language Models (LLMs) has significantly increased computational cost, energy consumption, and inference latency, making ac

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

Mergeable Model-Side Aggregation States for Long-Context Language Models

DGX agent

arXiv:2607.26448v1 Announce Type: new Abstract: A known limitation of long-context language models is their increasingly unreliable performance in non-additive, set-based aggregation as context length

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving

DGX agent

arXiv:2607.20988v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models augmented with world modeling represent a promising paradigm for end-to-end autonomous driving. While pixel-level

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Protocol Models: Scaling Decentralized Training with Communication-Efficient Model Parallelism

DGX agent

arXiv:2506.01260v2 Announce Type: replace Abstract: Scaling models has led to significant advancements in deep learning, but training these models in decentralized settings remains challenging due to

model-releasesarxiv-cs-lg
9 Jul 2026
Research

How Can AI Find My Model? A Model-Finding Experimental Study Considering Data Formats, Embeddings, and Retrieval Strategies

DGX agent

arXiv:2606.30846v1 Announce Type: new Abstract: Discovering simulation models for reuse remains a fundamental challenge in Modeling and Simulation (M&S). When many models coexist, identifying those th

researcharxiv-cs-ai
1 Jul 2026
Model Releases

OptiMUS-0.3: Using Large Language Models to Model and Solve Optimization Problems at Scale

DGX agent

arXiv:2407.19633v4 Announce Type: replace Abstract: Optimization problems are pervasive in sectors from manufacturing and distribution to healthcare. However, most such problems are still solved heuri

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

BIM-Edit: Benchmarking Large Language Models for IFC-Based Building Information Modeling

DGX agent

arXiv:2606.20146v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly applied to computer-aided design (CAD) to generate design artifacts from textual instructions. In engi

model-releasesarxiv-cs-ai
24 Jun 2026
Safety

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling

DGX agent

arXiv:2507.06419v3 Announce Type: replace Abstract: Reward modeling (RM), which captures human preferences to align large language models (LLMs), is increasingly employed in tasks such as model finetu

safetyarxiv-cs-cl
8 Jun 2026
Model Releases

Speculative Thinking: Enhancing Small-Model Reasoning with Large Model Guidance at Inference Time

DGX agent

arXiv:2504.12329v2 Announce Type: replace-cross Abstract: Recent advances leverage post-training to enhance model reasoning performance, which typically requires costly training pipelines and still su

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training

DGX agent

arXiv:2602.00747v2 Announce Type: replace-cross Abstract: Determining an effective data mixture is a key factor in Large Language Model (LLM) pre-training, where models must balance general competence

model-releasesarxiv-cs-ai
18 May 2026
← Previous
1
Next →
48,543 results
← Previous
123…1012
Next →