AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,515 results
24 Jun 2026

MambaRaw: Selective State Space Modeling for Efficient 4K Raw Image Reconstruction

Model ReleasesDGX agent

arXiv:2606.24479v1 Announce Type: new Abstract: In-camera JPEG previews are ubiquitous in raw image formats and provide an sRGB reference at negligible storage cost. Although existing metadata-based r

MedBench v5: A Dynamic, Process-Oriented, and Hallucination-Aware Benchmark for Clinical Multimodal Models

Model ReleasesDGX agent

arXiv:2606.24155v1 Announce Type: new Abstract: Existing medical AI benchmarks lack process visibility, atomic skill evaluation, and integrated hallucination detection. We introduce MedBench v5, a red

On the Stability of Prompt Ranking in Large Language Model Evaluation

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.24381v1 Announce Type: cross Abstract: Prompt-based interaction has become a dominant paradigm for using large language models (LLMs), where multiple candidate prompts are evaluated and the

Pigeonholing: Bad prompts hurt models to collapse and make mistakes

ResearchDGX agent

arXiv:2606.24267v1 Announce Type: cross Abstract: While in-context learning is generally shown to be effective in Large Language Models (LLMs), bad contexts can cause performance degradation and mode

Reinforcement Learning Towards Broadly and Persistently Beneficial Models

SafetyDGX agent

arXiv:2606.24014v1 Announce Type: new Abstract: As AI systems are deployed across increasingly diverse and high-stakes settings, model alignment must generalize beyond the tasks and domains seen durin

S1-Omni-Image: A Unified Model for Scientific Image Understanding, Generation, and Editing

ResearchDGX agent

arXiv:2606.24441v1 Announce Type: new Abstract: We present S1-Omni-Image, an open-weight unified multimodal model for scientific image understanding, generation, and editing. Unlike general-purpose im

ScaleToT: Generalizing Structured LLM Reasoning for Billion-Scale Low-Activity User Modeling

SafetyDGX agent

arXiv:2606.24605v1 Announce Type: new Abstract: Accurate user modeling often depends on rich interaction histories, which are unavailable for billions of low-activity users. Large Language Models (LLM

This is the strongest ARC-AGI-2 performance to date by an open-source model.

Model ReleasesDGX agent

This is the strongest ARC-AGI-2 performance to date by an open-source model. GLM-5.2 from @Zai_org on ARC-AGI (Verified) - ARC-AGI-2: 22.8%, 0.25 - ARC-AGI-1: 77.0%, 0.19 Performance is comparable wit

We have a new version of GPT-5.5 Instant for you, and it's much more fun to talk to. Our most-used model is now better at understanding the …

Model ReleasesDGX agent

We have a new version of GPT-5.5 Instant for you, and it's much more fun to talk to. Our most-used model is now better at understanding the intent behind a question and adapting its response according

When Preferences Fail to Become Incentives: A Utility-Behavior Gap in Large Language Models

SafetyDGX agent

arXiv:2606.22974v2 Announce Type: replace Abstract: Recent work on preference elicitation in large language models (LLMs) has demonstrated that, when given a series of choices between two outcomes, LL

23 Jun 2026

BadDreamer: Transferable Backdoor Attacks against Video World Models for Autonomous Driving

SafetyDGX agent

arXiv:2606.21172v1 Announce Type: new Abstract: Video world models are increasingly used in autonomous driving to forecast future scene evolution and provide future-aware spatio-temporal representatio

Certified World Models: Predictability Across Configuration, Horizon, and Resolution

ResearchDGX agent

arXiv:2606.13092v2 Announce Type: replace Abstract: Scale buys interpolation; structure buys certifiable transfer. A world model's average error does not say whether a particular rollout can be truste

Detail++: Training-Free Detail Enhancer for T2I Diffusion Models

Model ReleasesDGX agent

arXiv:2507.17853v3 Announce Type: replace Abstract: Recent advances in text-to-image (T2I) generation have led to impressive visual results. However, these models still face significant challenges whe

Does RoPE Prevent or Degrade Retrieval Heads? A Mechanistic Analysis Across Model Families

Model ReleasesDGX agent

arXiv:2606.21249v1 Announce Type: new Abstract: Retrieval heads, attention heads that copy information from earlier context to the current position, have been proposed as the mechanistic substrate for

dVLA-RL: Reinforcement Learning over Denoising Trajectories for Discrete Diffusion Vision-Language-Action Models

SafetyDGX agent

arXiv:2606.23623v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have established a powerful paradigm for generalist robotic manipulation by grounding control into the semantic reas

HiL-ResRL: A Model-Agnostic Finetuning Adapter via Human-in-the-loop Residual Reinforcement Learning

SafetyDGX agent

arXiv:2606.22860v1 Announce Type: new Abstract: Recent advancements in generative imitation learning have significantly propelled the field of robotic manipulation. However, the majority of existing m

Large Language Model-Assisted Cleaning of Report-Derived Labels in a Large-Scale Chest CT Dataset

Model ReleasesDGX agent

arXiv:2606.22382v1 Announce Type: cross Abstract: Purpose: To evaluate whether large language model (LLM)-assisted label cleaning can identify label-report discordance in CT-RATE, a large-scale public

Mistral debuts OCR 4, a model featuring structured document extraction with bounding boxes, block classification, and inline confidence scores, in 170 languages (Mistral AI Blog)

Model ReleasesDGX agent

Mistral AI Blog: Mistral debuts OCR 4, a model featuring structured document extraction with bounding boxes, block classification, and inline confidence scores, in 170 languages — Today, we're releasi

NeuroShield: A Device-Agnostic Foundation Model for EEG Authentication

ResearchDGX agent

arXiv:2606.20673v1 Announce Type: cross Abstract: A central challenge in EEG authentication is that models are typically tied to the acquisition settings in which they are trained. In particular, vari

OGD4All: A Framework for Accessible Interaction with Geospatial Open Government Data Based on Large Language Models

Model ReleasesDGX agent

arXiv:2602.00012v3 Announce Type: replace Abstract: We present OGD4All, a transparent, auditable, and reproducible framework based on Large Language Models (LLMs) to enhance citizens' interaction with

Prompting Diffusion Models for Zero-Shot Instance Segmentation

ResearchDGX agent

arXiv:2606.22660v1 Announce Type: new Abstract: Several disruptive research directions have recently emerged in computer vision, including foundation models achieving previously unseen zero-shot perfo

Quantum Convolutional Neural Networks for Groundwater Heat Plume Prediction: A Surrogate Modeling Approach

ResearchDGX agent

arXiv:2606.23411v1 Announce Type: cross Abstract: Quantum machine learning methods are increasingly explored for modeling complex environmental systems, including groundwater heat plume dynamics. In t

Reinforcement learning to improve large language model-based automated code compliance systems

Model ReleasesDGX agent

arXiv:2606.22402v1 Announce Type: cross Abstract: Large language model (LLM)-based approaches for automated code compliance (ACC) of building regulations are prone to generating incorrect and hallucin

Test-Time Alignment of Text-to-Image Diffusion Models via Null-Text Embedding Optimisation

SafetyDGX agent

arXiv:2511.20889v2 Announce Type: replace Abstract: Test-time alignment (TTA) aims to adapt models to specific rewards during inference. However, existing methods tend to either under-optimise or over

Zero-Shot Vision-Language Models for Classroom Engagement Recognition: A Benchmark Study of Prompt Sensitivity and Cross-Dataset Generalization

Model ReleasesDGX agent

arXiv:2606.21861v1 Announce Type: new Abstract: Automated classroom engagement recognition holds substantial promise for scalable learning analytics, yet the suitability of modern Vision-Language Mode

22 Jun 2026

My parallel agent side-project today was having Claude Code port the new Moebius image pinpointing model to ONNX in order to run it entirely…

Model ReleasesDGX agent

My parallel agent side-project today was having Claude Code port the new Moebius image pinpointing model to ONNX in order to run it entirely in the browser https://simonwillison.net/2026/Jun/22/portin

Sakana AI launches Fugu, a multi-agent orchestration system accessible through a single model API, claiming Fugu Ultra matches Fable and Mythos on benchmarks (Carl Franzen/VentureBeat)

Model ReleasesDGX agent

Carl Franzen / VentureBeat: Sakana AI launches Fugu, a multi-agent orchestration system accessible through a single model API, claiming Fugu Ultra matches Fable and Mythos on benchmarks — Last night,

21 Jun 2026

Let’s go open models! ❤️

Local AiDGX agent

Ollama, an open-source platform for running large language models locally, announced support or enthusiasm for open models on X (formerly Twitter). The post likely promotes the benefits of open-source

11 Jun 2026

Adaptive Multi-Resolution Procedural Knowledge Compression for Large Language Models

Model ReleasesDGX agent

arXiv:2606.12203v1 Announce Type: new Abstract: Large language models (LLMs) are widely used to tackle complex tasks with autonomous workflows. Recently, reusable natural language skills have emerged

Bridging the Morphology Gap: Adapting VLA Models to Dexterous Manipulation via Intent-Conditioned Fine-Tuning

Model ReleasesDGX agent

arXiv:2606.12109v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated remarkable zero-shot generalization in robotic manipulation, yet the vast majority of pre-traine

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards

ResearchDGX agent

arXiv:2511.16672v4 Announce Type: replace Abstract: Recent advances in large multimodal models (LMMs) have enabled impressive reasoning and perception abilities, yet most existing training pipelines s

PermDoRA -- Understanding Adapter Interference in Language Models: Limits of Parameter-Space Geometry

Model ReleasesDGX agent

arXiv:2606.11262v1 Announce Type: cross Abstract: Access control in large language models (LLMs) requires modular mechanisms to enable domain-specific behavior without retraining or cross-domain inter

Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.12412v1 Announce Type: cross Abstract: Vision-language models (VLMs) project images into hundreds to thousands of visual tokens, making decoder inference expensive in both attention computa

Sparse probes and murky physics: a case study of interpretability challenges in a foundation model for continuum dynamics

Model ReleasesDGX agent

arXiv:2606.11657v1 Announce Type: cross Abstract: Generative AI emulators are increasingly used in scientific domains where we already have strong theory, benchmarks, and physical intuition. This rais

Task-Aligned Stability Analysis of Vision-Language Models for Autonomous Driving Hazard Detection

Model ReleasesDGX agent

arXiv:2606.11889v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used for scene understanding in autonomous driving, but robustness analysis often relies on task-agnost

The model subsidies will eventually end and this workflow of “creating loops that will prompt your agents” will result in massive amounts of…

Model ReleasesDGX agent

The model subsidies will eventually end and this workflow of “creating loops that will prompt your agents” will result in massive amounts of code that’s not well understood that you will have to pay l

Towards Data-free and Training-free Compression for Speech Foundation Models Using Parameter Clustering

Model ReleasesDGX agent

arXiv:2606.11836v1 Announce Type: cross Abstract: This paper presents a novel data-free and training-free compression approach for speech foundation models using channelwise clustering via k-means. Mo

10 Jun 2026

A Continuous-Time Markov Chain Framework for Insertion Language Models

ResearchDGX agent

arXiv:2606.10199v1 Announce Type: cross Abstract: Insertion Language Models (ILMs) offer several advantages over left-to-right generation and mask-based generation. However, existing formulations of i

Benchmarking stereo reconstruction for 3D printable Martian terrain models

Model ReleasesDGX agent

arXiv:2606.10364v1 Announce Type: new Abstract: Reconstructing printable 3D models from Mars rover imagery is challenging because Martian terrain is low-texture, irregular, and partially observed. We

CITRAS-FM: Tiny Time Series Foundation Model for Covariate-Informed Zero-Shot Forecasting

Model ReleasesDGX agent

arXiv:2606.10798v1 Announce Type: new Abstract: Pretrained time series foundation models (TSFMs) have enabled zero-shot forecasting on unseen target series. However, existing TSFMs often incur high co

Conditional Vendi Score: Prompt-Aware Diversity Evaluation for Generative AI Models and LLMs

SafetyDGX agent

arXiv:2411.02817v2 Announce Type: replace-cross Abstract: Generative models guided by text prompts are widely evaluated for fidelity and prompt alignment, yet their ability to produce outputs remains

Density Field State Space Models: 1-Bit Distillation, Efficient Inference, and Knowledge Organization in Mamba-2

Local AiDGX agent

arXiv:2606.10932v1 Announce Type: new Abstract: We present Density Field State Space Models (DF-SSM), a framework for compressing SSMs to a 1-bit scaffold with int8 low-rank correction. Applied to Mam

Does Reasoning Preserve Alignment? On the Trustworthiness of Large Reasoning Models

SafetyDGX agent

arXiv:2606.11046v1 Announce Type: new Abstract: Instruction-tuned LLMs are increasingly converted into reasoning models through post-training to improve multi-step task performance. This conversion is

Few-step Generative Models as Lossy Compression

ResearchDGX agent

arXiv:2606.10450v1 Announce Type: new Abstract: DiffC provides a principled way to reuse pre-trained diffusion models for lossy compression, but its encoding and decoding procedures remain slow becaus

Going with the Flow: Koopman Behavioral Models as Pseudo Planners for Visuo-Motor Dexterity

SafetyDGX agent

arXiv:2602.07413v3 Announce Type: replace Abstract: Contemporary visuo-motor dexterity models often rely on expressive policy classes with diffusion and transformer backbones to achieve strong perform

I took Andrej Karpathy's LLM Council concept to the next level (Docker, MCP, Skill, Search, local (Ollama)/cloud model support and much more)

Local AiDGX agent

A developer expanded on Andrej Karpathy's LLM Council concept by implementing an enhanced system with Docker containerization, Model Context Protocol (MCP) integration, skill modules, web search capab

MedFeat: Model-Aware and Explainability-Driven Feature Engineering with LLMs for Clinical Tabular Prediction

SafetyDGX agent

arXiv:2603.02221v2 Announce Type: replace-cross Abstract: In clinical tabular prediction, classical machine learning models with feature engineering often outperform neural methods. LLMs are increasin

Meet DiffusionGemma! An experimental open model that explores a fast approach to text generation, released under an Apache 2.0 license. Movi…

Model ReleasesDGX agent

Meet DiffusionGemma! An experimental open model that explores a fast approach to text generation, released under an Apache 2.0 license. Moving beyond sequential, token-by-token processes to generate e

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models

SafetyDGX agent

arXiv:2606.11167v1 Announce Type: new Abstract: Full-duplex spoken dialogue models can listen and speak simultaneously, making them a promising architecture for natural conversation. However, current

OpenRTLSet: A Fully Open-Source Dataset for Large Language Model-based Verilog Module Design

Model ReleasesDGX agent

arXiv:2606.10285v1 Announce Type: new Abstract: OpenRTLSet introduces the largest fully open-source dataset for hardware design, offering over 131,000 diverse Verilog code samples to the research comm

PRISM: Parallel Residual Iterative Sequence Model

ResearchDGX agent

arXiv:2602.10796v3 Announce Type: replace Abstract: Generative sequence modeling faces a fundamental tension between the expressivity of Transformers and the efficiency of linear sequence models. Exis

ReasonAlloc: Hierarchical Decoding-Time KV Cache Budget Allocation for Reasoning Models

Model ReleasesDGX agent

arXiv:2606.11164v1 Announce Type: new Abstract: Long chain-of-thought (CoT) trajectories in large language model (LLM) reasoning cause severe inference bottlenecks due to rapid key-value (KV) cache gr

Sample Where You Struggle: Sharpening Base Model Reasoning via Entropy-Guided Power Sampling

Model ReleasesDGX agent

arXiv:2606.09926v1 Announce Type: cross Abstract: Sampling from the sequence-level power distribution p^alpha elicits RL-level reasoning from base language models without any parameter updates, but th

SSR-Merge: Subspace Signal Routing for Training-Free LoRA Merging in Diffusion Models

Model ReleasesDGX agent

arXiv:2606.10617v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) merging can efficiently combine diverse generative capabilities from multiple trained LoRAs for a diffusion model. However, e

the cost of fable is going to make smart model routing impossible to ignore

AgentsDGX agent

This post discusses how the pricing model of Fable (likely an AI/LLM service) creates economic incentives that make intelligent routing between different AI models a necessary optimization strategy ra

The Model Lab vs Agent Lab distinction is one of the clearest frameworks I've seen for understanding where AI value actually lives right now…

AgentsDGX agent

The Model Lab vs Agent Lab distinction is one of the clearest frameworks I've seen for understanding where AI value actually lives right now. TL;DR from @latentspacepod: • Model Labs compete on capabi

When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models

SafetyDGX agent

arXiv:2606.10740v1 Announce Type: new Abstract: Failures in multi-turn reasoning models are largely invisible to terminal-score evaluation. A model can lock onto an unsafe stance early in a long dialo

WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling

ResearchDGX agent

arXiv:2512.14614v2 Announce Type: replace Abstract: This paper presents WorldPlay, a streaming video diffusion model that enables real-time, interactive world modeling with long-term geometric consist

Wrote up my initial impressions of Claude Fable 5 - it has a big model smell: slow, expensive and capable of crunching through pretty much e…

Model ReleasesDGX agent

Wrote up my initial impressions of Claude Fable 5 - it has a big model smell: slow, expensive and capable of crunching through pretty much everything I threw at it https://simonwillison.net/2026/Jun/9

9 Jun 2026

3D Oral Modelling with Improved Vertex Distribution Using Matching-Based Learning

ResearchDGX agent

arXiv:2606.07907v1 Announce Type: cross Abstract: In our previous work, a deep learning-based framework for 3D intraoral reconstruction was proposed. The model directly predicts explicit 3D point clou

← Previous
1…105106107108109…1009
Next →