AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,840 results
Model Releases

Finding the Minimal Parameter Budget for Implicit Reasoning: A Data Complexity Driven Scaling Law for Language Models

DGX agent

arXiv:2504.03635v4 Announce Type: replace Abstract: Reasoning is a core capability of language models (LMs), yet it remains unclear how much model capacity is necessary to support reasoning during pre

model-releasesarxiv-cs-ai
2 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Interpretable Modeling of Driver Attention Shifts with a Vision--Language Model

DGX agent

arXiv:2508.05852v2 Announce Type: replace Abstract: Driver gaze is commonly modeled as a spatial heatmap, but heatmaps alone are difficult for humans to interpret because they do not explain which roa

safetyarxiv-cs-cv
2 Jun 2026
Model Releases

RynnVLA-002: A Unified Vision-Language-Action and World Model

DGX agent

arXiv:2511.17502v3 Announce Type: replace Abstract: We introduce RynnVLA-002, a unified Vision-Language-Action (VLA) and world model. The world model leverages action and visual inputs to predict futu

model-releasesarxiv-cs-ro
2 Jun 2026
Model Releases

Uncovering Competency Gaps in Large Language Models and Their Benchmarks

DGX agent

arXiv:2512.20638v2 Announce Type: replace-cross Abstract: The evaluation of large language models relies heavily on standardized benchmarks. These benchmarks provide useful aggregated metrics, but can

model-releasesarxiv-cs-ai
2 Jun 2026
Agents

Dreaming Of Others: Latent Teammate Modeling In World Models For Multi-Agent Reinforcement Learning

DGX agent

arXiv:2605.31361v1 Announce Type: cross Abstract: In cooperative multi-agent reinforcement learning (MARL), agents must coordinate with partners whose internal policies and intentions are not directly

agentsarxiv-cs-ai
1 Jun 2026
Model Releases

EvoDefense: Co-Evolving Black-Box Defense with Large Language Models

DGX agent

arXiv:2605.31140v1 Announce Type: cross Abstract: Large Language Models (LLMs) remain highly vulnerable to diverse attacks, particularly in black-box settings where the internals of target models are

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

How Trustpilot built a real-time architecture for data enrichment using Gemma

DGX agent

Processing millions of user reviews in real-time, under strict latency and cost constraints, is no easy task. Trustpilot has been doing exactly that with custom machine learning since long before larg

model-releasesgoogle-cloud-ai
1 Jun 2026
Model Releases

Chess-World-Model: A 10M-Game Benchmark for Exact State Tracking from Chess Move Sequences

DGX agent

arXiv:2605.30100v1 Announce Type: new Abstract: World models require state tracking, which is the ability to maintain a correct latent state across action sequences. Existing benchmarks are often synt

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Do Physics Foundation Models Learn Generalizable Physics? A Bias-Aware Benchmark Across Physical Regimes and Distribution Shifts

DGX agent

arXiv:2605.29283v1 Announce Type: cross Abstract: Recent physics foundation models claim general spatiotemporal forecasting ability, yet their evaluations often collapse performance into a single aver

model-releasesarxiv-cs-ai
29 May 2026
Safety

Draft-OPD: On-Policy Distillation for Speculative Draft Models

DGX agent

arXiv:2605.29343v1 Announce Type: new Abstract: Speculative decoding accelerates large language model inference by pairing a target model with a lightweight draft model whose proposed tokens are verif

safetyarxiv-cs-cl
29 May 2026
Model Releases

RightNow-Arabic-0.5B-Turbo: An Open Sub-1B Arabic Language Model via Vocabulary Injection and Edge-First Deployment

DGX agent

arXiv:2605.28827v1 Announce Type: new Abstract: Open Arabic large language models split into two classes: sub-1B multilingual models that treat Arabic as an afterthought (Qwen2.5-0.5B, Falcon-H1-0.5B)

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Why Specialist Models Still Matter: A Heterogeneous Multi-Agent Paradigm for Medical Artificial Intelligence

DGX agent

arXiv:2605.29744v1 Announce Type: new Abstract: The impressive performance of generalist large language models (LLMs) such as GPT and Claude in healthcare raises a critical question: will domain-speci

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Code as a Weapon: A Consensus-Labeled Prompt Bank for Measuring Coding-Model Compliance with Malicious-Code Requests

DGX agent

arXiv:2605.28734v1 Announce Type: cross Abstract: A general-purpose language model that answers a harmful question returns text; a coding model that complies with a malicious request can return a work

model-releasesarxiv-cs-cl
28 May 2026
Research

Continual Learning in Modern Hopfield Networks with an Application to Diffusion Models

DGX agent

arXiv:2605.27975v1 Announce Type: new Abstract: Generative models, including diffusion models, are increasingly used as foundation models and adapted through sequential fine-tuning, making continual l

researcharxiv-cs-lg
28 May 2026
Model Releases

Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning

DGX agent

arXiv:2605.28277v1 Announce Type: new Abstract: Whether large language models (LLMs) construct internal spatial world models from pure-text descriptions remains contested, and whether such capabilitie

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Towards Faithful Agentic XAI: A Verification Method and an Open-World Benchmark for Better Model Faithfulness

DGX agent

arXiv:2605.27879v1 Announce Type: new Abstract: Explainable AI (XAI) helps users interpret model behavior and identify potential faults. Agentic XAI systems use Large Language Models (LLMs) to make ex

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Benchmarking Convolutional, Transformer, Hybrid, and Vision Language Models for Multi Disease Retinal Screening

DGX agent

arXiv:2605.26283v1 Announce Type: new Abstract: Modern deep learning offers powerful tools for automated retinal screening, but it remains unclear how different visual model families compare in realis

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

Black-box Membership Inference Attacks on the Pre-training Data of Image-generation Models

DGX agent

arXiv:2605.27020v1 Announce Type: cross Abstract: The rapid advancement of diffusion-based image generation models has raised serious concerns regarding potential copyright and privacy infringements i

model-releasesarxiv-cs-ai
27 May 2026
Safety

Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization

DGX agent

arXiv:2605.26282v1 Announce Type: new Abstract: Model-based reinforcement learning (RL) can be effectively supported at scale through the use of world models. However, in practice, scaling such approa

safetyarxiv-cs-lg
27 May 2026
Model Releases

Benchmarking Patent Embeddings: A Multi-Task Evaluation of 22 Models Across Retrieval, Classification, and Clustering

DGX agent

arXiv:2605.24297v1 Announce Type: cross Abstract: Which fine-tuning signals improve patent embedding models, and do gains transfer across patent landscapes? We benchmark 22 embedding models, from 22M-

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

From Model Scaling to System Scaling: Scaling the Harness in Agentic AI

DGX agent

arXiv:2605.26112v1 Announce Type: new Abstract: This paper studies the next major bottleneck in agentic AI as system scaling, not only model scaling: the design of auditable, persistent, modular, and

model-releasesarxiv-cs-ai
26 May 2026
Research

Mimir: Large-scale Multilingual Concept Modeling

DGX agent

arXiv:2605.25263v1 Announce Type: cross Abstract: Current language modeling approaches are built around tokens. Text corpora are split into tokens, and models are trained by performing computations on

researcharxiv-cs-ai
26 May 2026
Model Releases

SomaliBench Eval: Measuring English-to-Somali Refusal Gaps in Open-Weight Language Models

DGX agent

arXiv:2605.25420v1 Announce Type: cross Abstract: Large language model safety evaluation remains heavily English-centered, leaving low-resource languages under-measured even when models are deployed g

model-releasesarxiv-cs-ai
26 May 2026
Agents

Why We Need World Models for AGI: Where LLMs Fail and How World Models May Outperform

DGX agent

arXiv:2605.23972v1 Announce Type: new Abstract: Large language models achieve strong performance in language generation and knowledge-intensive tasks, yet remain limited in settings requiring causal r

agentsarxiv-cs-ai
26 May 2026
Model Releases

A Comparative Evaluation of Structural Topic Models and BERTopic for Short, Open-Ended Survey Responses

DGX agent

arXiv:2605.23093v1 Announce Type: new Abstract: Topic modeling in applied psychology increasingly spans two methodological traditions: probabilistic bag-of-words models and newer embedding-based appro

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models

DGX agent

arXiv:2605.23238v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as economic agents in marketplaces, auctions, and bidding settings. Anticipating their behavior i

model-releasesarxiv-cs-ai
25 May 2026
Agents

Latent Cache Flow: Model-to-Model Communication Without Text

DGX agent

arXiv:2605.22863v1 Announce Type: new Abstract: LLM agents today communicate via text, which incurs considerable latency and information loss due to the need to autoregressively decode the sharer mode

agentsarxiv-cs-lg
25 May 2026
Local Ai

The Attribution Contract: Feature Attribution for Generative Language Models

DGX agent

arXiv:2605.23080v1 Announce Type: new Abstract: Feature attribution methods promise to identify which input features matter for a model output. In generative language models, however, it is often uncl

local-aiarxiv-cs-lg
25 May 2026
Model Releases

Provable Joint Decontamination for Benchmarking Multiple Large Language Models

DGX agent

arXiv:2605.21543v1 Announce Type: new Abstract: Benchmark data contamination has become a central challenge in LLM evaluation: when evaluation examples appear in the training data of one or more audit

model-releasesarxiv-cs-lg
23 May 2026
Model Releases

Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild

DGX agent

arXiv:2605.22064v1 Announce Type: new Abstract: Hy-MT2 is a family of fast-thinking multilingual translation models designed for complex real-world scenarios. It includes three model sizes: 1.8B, 7B,

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Chronicle: A Multimodal Foundation Model for Joint Language and Time Series Understanding

DGX agent

arXiv:2605.20268v1 Announce Type: cross Abstract: Real-world time series come with text: metadata, descriptions, news, reports. Yet time series foundation models process numerical sequences in isolati

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

Memory Grafting: Scaling Language Model Pre-training via Offline Conditional Memory

DGX agent

arXiv:2605.20948v1 Announce Type: new Abstract: Scaling conditional memory offers a promising way to increase language-model capacity, but existing methods such as Engram learn large memory tables fro

model-releasesarxiv-cs-cl
21 May 2026
Hardware

Towards Multi-Model LLM Schedulers: Empirical Insights into Offloading and Preemption

DGX agent

arXiv:2605.19593v1 Announce Type: new Abstract: Modern deployments of Large Language Models (LLMs) increasingly require serving multiple models with diverse architectures, sizes, and specialization on

hardwarearxiv-cs-ai
20 May 2026
Model Releases

Unlocking the Potential of Continual Model Merging: An ODE Perspective

DGX agent

arXiv:2605.19409v1 Announce Type: cross Abstract: Continual Model Merging (CMM) enables rapid customization of foundation models across sequentially arriving tasks, offering a scalable alternative to

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Extending Pretrained 10-Second ECG Foundation Models to Longer Horizons

DGX agent

arXiv:2605.16975v1 Announce Type: cross Abstract: Electrocardiogram (ECG) foundation models pretrained on typical diagnostic 10-second ECG segments, have demonstrated strong transferability across a r

model-releasesarxiv-cs-ai
19 May 2026
Research

Foundation Models for Credit Risk Prediction: A Game Changer?

DGX agent

arXiv:2605.18147v1 Announce Type: new Abstract: Predictive models play a pivotal role in credit risk management, guiding critical decisions through accurate estimation of default probabilities and los

researcharxiv-cs-lg
19 May 2026
Model Releases

GenTS: A Comprehensive Benchmark Library for Generative Time Series Models

DGX agent

arXiv:2605.17804v1 Announce Type: new Abstract: Generative models have demonstrated remarkable potential in time series analysis tasks, like synthesis, forecasting, imputation, etc. However, offering

model-releasesarxiv-cs-lg
19 May 2026
Model Releases

Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency

DGX agent

arXiv:2605.18732v1 Announce Type: cross Abstract: While scaling laws govern aggregate large language model performance, no scaling law has linked factual recall to both model size and training-data co

model-releasesarxiv-cs-ai
19 May 2026
Industry

Seems a lot of autoregressive models will be converted to diffusion models

DGX agent

Emad Mostaque, CEO of Stability AI, suggests that many autoregressive models may transition to or be replaced by diffusion-based approaches. This reflects speculation or prediction about a potential s

industryemad-mostaque--x
19 May 2026
Model Releases

Beyond Mode-Seeking RL: Trajectory-Balance Post-Training for Diffusion Language Models

DGX agent

arXiv:2605.13935v1 Announce Type: cross Abstract: Diffusion language models are a promising alternative to autoregressive models, yet post-training methods for them largely adapt reward-maximizing obj

model-releasesarxiv-cs-cl
15 May 2026
Model Releases

Hidden State Poisoning Attacks against Mamba-based Language Models

DGX agent

arXiv:2601.01972v4 Announce Type: cross Abstract: State space models (SSMs) like Mamba offer efficient alternatives to Transformer-based language models, with linear time complexity. Yet, their advers

model-releasesarxiv-cs-ai
15 May 2026
Safety

Slot-MPC: Goal-Conditioned Model Predictive Control with Object-Centric Representations

DGX agent

arXiv:2605.14937v1 Announce Type: cross Abstract: Predictive world models enable agents to model scene dynamics and reason about the consequences of their actions. Inspired by human perception, object

safetyarxiv-cs-ai
15 May 2026
Model Releases

Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling

DGX agent

arXiv:2605.13062v1 Announce Type: new Abstract: Recent image editing models have achieved remarkable progress in instruction following, multimodal understanding, and complex visual editing. However, e

model-releasesarxiv-cs-cv
14 May 2026
Model Releases

Understanding and Accelerating the Training of Masked Diffusion Language Models

DGX agent

arXiv:2605.13026v1 Announce Type: cross Abstract: Masked diffusion models (MDMs) have emerged as a promising alternative to autoregressive models (ARMs) for language modeling. However, MDMs are known

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

When is Warmstarting Effective for Scaling Language Models?

DGX agent

arXiv:2605.13405v1 Announce Type: new Abstract: Model growth from a given checkpoint aims to accelerate training of a larger model, offering potential resource savings. Despite recent interest, warmst

model-releasesarxiv-cs-lg
14 May 2026
Model Releases

Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models

DGX agent

arXiv:2605.11887v1 Announce Type: new Abstract: Large language models have achieved remarkable capabilities across diverse tasks, yet their internal decision-making processes remain largely opaque, li

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

READ: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language Modeling

DGX agent

arXiv:2312.06950v3 Announce Type: replace-cross Abstract: Fully fine-tuning pretrained large-scale transformer models has become a popular paradigm for video-language modeling tasks, such as temporal

model-releasesarxiv-cs-cl
13 May 2026
Research

A Single-Layer Model Can Do Language Modeling

DGX agent

arXiv:2605.10643v1 Announce Type: new Abstract: Modern language models scale depth by stacking layers, each holding its own state - a per-layer KV cache in transformers, a per-layer matrix in Mamba, G

researcharxiv-cs-cl
12 May 2026
← Previous
1…2223242526…1247
Next →