AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,965
  • Agents7,446
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,131
  • Local Ai4,857
  • Model Releases23,360
  • Research19,834
  • Safety13,174
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,965
  • Agents7,446
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,131
  • Local Ai4,857
  • Model Releases23,360
  • Research19,834
  • Safety13,174
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

86,965Total entries
1Added by human
86,964Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,458 results
23 May 2026

VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation

Model ReleasesDGX agent

arXiv:2605.22368v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed for software engineering, constructing high-quality benchmarks is crucial for evaluating not j

22 May 2026

Evaluating Commercial AI Chatbots as News Intermediaries

Model ReleasesDGX agent

arXiv:2605.22785v1 Announce Type: new Abstract: AI chatbots are rapidly shaping how people encounter the news, yet no prior study has systematically measured how accurately these systems, with their p

Hallucination as Commitment Failure: Larger LLMs Misfire Despite Knowing the Answer

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2605.22007v1 Announce Type: new Abstract: Hallucination is often viewed as a direct consequence of missing knowledge: a model answers incorrectly when the correct answer is absent from its gener

HealthCraft: A Reinforcement Learning Safety Environment for Emergency Medicine

Model ReleasesDGX agent

arXiv:2605.21496v1 Announce Type: cross Abstract: Frontier language models are being deployed into clinical workflows faster than the infrastructure to evaluate them safely. Static medical-QA benchmar

Hypergraph as Language

Local AiDGX agent

arXiv:2605.21858v1 Announce Type: new Abstract: Large language models (LLMs) have recently shown strong potential in modeling relational structures. However, existing approaches remain fundamentally g

img2vid: ComfyUI doesn't find the spatial upscaler

Local AiDGX agent

This Reddit post discusses a common ComfyUI issue where users cannot locate spatial upscaler models for img2vid workflows. The problem typically stems from placing spatial upscaler models in the wrong

Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements

Model ReleasesDGX agent

arXiv:2605.22079v1 Announce Type: new Abstract: Large language models (LLMs) are widely used to generate structured outputs such as JSON, SQL, and code, yet public resources remain limited for evaluat

MTP means Multi Token Prediction. It's a speculative decoding technique that can result in large inference speedups in many cases. 1. Update…

Model ReleasesDGX agent

MTP means Multi Token Prediction. It's a speculative decoding technique that can result in large inference speedups in many cases. 1. Update to LM Studio 0.4.14 2. Download a model that supports MTP l

MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation

Model ReleasesDGX agent

arXiv:2505.17123v3 Announce Type: replace Abstract: Recent advances in Large Language Models (LLMs) have shown promising results in complex reasoning tasks. However, current evaluations predominantly

RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator

Model ReleasesDGX agent

arXiv:2605.21748v1 Announce Type: new Abstract: As interactive LLM-based applications are created and refined, model developers need to evaluate the quality of generated text along many possible axes.

Specialization Beats Scale: A Strategic Variable Most AI Procurement Decisions Overlook

ToolsDGX agent

This article argues that specialized AI models often outperform larger, general-purpose models for specific use cases, challenging the common procurement assumption that bigger is always better. It li

Structural Anchor Pruning: Training-Free Multi-Vector Compression for Visual Document Retrieval

Model ReleasesDGX agent

arXiv:2601.20107v2 Announce Type: replace-cross Abstract: Recent Vision-Language Models (e.g., ColPali) enable fine-grained Visual Document Retrieval (VDR) but incur prohibitive multi-vector index sto

Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance

Model ReleasesDGX agent

arXiv:2605.22202v1 Announce Type: new Abstract: In this paper, we show that high-performing embedding models organize their embedding spaces in a consistent way. We evaluate 25 contemporary embedding

VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction

Model ReleasesDGX agent

arXiv:2602.13294v3 Announce Type: replace Abstract: Evaluating whether Multimodal Large Language Models (MLLMs) genuinely reason about physical dynamics remains challenging. Most existing benchmarks r

21 May 2026

DASH: Fast Differentiable Architecture Search for Hybrid Attention in Minutes on a Single GPU

Model ReleasesDGX agent

arXiv:2605.20936v1 Announce Type: cross Abstract: Hybrid attention architectures are becoming an increasingly important paradigm for improving LLM inference efficiency while preserving model quality,

Datasette Agent

Model ReleasesDGX agent

We just announced the first release of Datasette Agent, a new extensible AI assistant for Datasette. I've been working on my LLM Python library for just over three years now, and Datasette Agent repre

Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs

ResearchDGX agent

arXiv:2605.20382v1 Announce Type: new Abstract: Language models are trained to follow instructions, but they are also powerful pattern completers. What happens when these two objectives conflict? We c

Explainability Methods for Hardware Trojan Detection: A Systematic Comparison

Model ReleasesDGX agent

arXiv:2601.18696v4 Announce Type: replace Abstract: Hardware trojans are malicious circuits which compromise the functionality and security of an integrated circuit (IC). These circuits are manufactur

GraphRAG on Consumer Hardware: Benchmarking Local LLMs for Healthcare EHR Schema Retrieval

Model ReleasesDGX agent

arXiv:2605.20815v1 Announce Type: new Abstract: Graph-based Retrieval Augmented Generation (GraphRAG) extends retrieval-augmented generation to support structured reasoning over complex corpora, but i

JUDO: A Juxtaposed Domain-Oriented Multimodal Reasoner for Industrial Anomaly QA

Model ReleasesDGX agent

arXiv:2605.20284v1 Announce Type: new Abstract: Industrial anomaly detection has been significantly advanced by Large Multimodal Models (LMMs), enabling diverse human instructions beyond detection, pa

Leveraging LLMs for Grammar Adaptation: A Study on Metamodel-Grammar Co-Evolution

Model ReleasesDGX agent

arXiv:2605.21465v1 Announce Type: new Abstract: In model-driven engineering, metamodel evolution leads to the need to adapt corresponding grammars to maintain consistency, which typically requires ted

MONET: A Massive, Open, Non-redundant and Enriched Text-to-image dataset

Model ReleasesDGX agent

arXiv:2605.21272v1 Announce Type: new Abstract: Training large text-to-image models requires high-quality, curated datasets with diverse content and detailed captions. Yet the cost and complexity of c

Qwen 3.7 Max now available on Vercel AI Gateway

Model ReleasesDGX agent

Vercel has announced the availability of Qwen 3.7 Max, a language model, through its Vercel AI Gateway platform. This integration allows developers to access and use Qwen 3.7 Max alongside other AI mo

Sample Complexity of Transfer Learning: An Optimal Transport Approach

ResearchDGX agent

arXiv:2605.20545v1 Announce Type: cross Abstract: Transfer learning is an essential technique for many machine learning/AI models of complex structures such as large language models and generative AI.

SMoA: Spectrum Modulation Adapter for Parameter-Efficient Fine-Tuning

Model ReleasesDGX agent

arXiv:2605.21147v1 Announce Type: cross Abstract: As the number of model parameters increases, parameter-efficient fine-tuning (PEFT) has become the go-to choice for tailoring pre-trained large langua

The Yes-Man Syndrome: Benchmarking Abstention in Embodied Robotic Agents

Model ReleasesDGX agent

arXiv:2605.20544v1 Announce Type: cross Abstract: Vision-language models (VLMs) are used as high-level planners for embodied agents, translating natural language instructions and visual observations i

20 May 2026

Announcing the release of Stable Audio 3!

Local AiDGX agent

Stability AI announced the launch of Stable Audio 3, a family of three AI music models and one audio-based special effects model. Most of these releases are 'open weight' models trained on licensed tr

Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos

Model ReleasesDGX agent

arXiv:2605.18984v1 Announce Type: new Abstract: Recent video generative models have greatly improved the realism of AI-generated videos, yet their outputs still exhibit artifacts such as temporal inco

CogScale: Scalable Benchmark for Sequence Processing

Model ReleasesDGX agent

arXiv:2605.19758v1 Announce Type: new Abstract: The ability to maintain and manipulate information over time is a fundamental aspect of living beings and Artificial Intelligence. While modern models h

Detecting Fluent Optimization-Based Adversarial Prompts via Sequential Entropy Changes

Model ReleasesDGX agent

arXiv:2605.19966v1 Announce Type: cross Abstract: Optimization-based adversarial suffixes can jailbreak aligned large language models (LLMs) while remaining fluent, weakening static and windowed perpl

Distributionally Robust Control via Stein Variational Inference for Contact-Rich Manipulation

Model ReleasesDGX agent

arXiv:2605.19029v1 Announce Type: new Abstract: Reliable robotic manipulation requires control policies that can accurately represent and adapt to uncertainty arising from contact-rich interactions. M

EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs

Model ReleasesDGX agent

arXiv:2605.19559v1 Announce Type: cross Abstract: The rapid development of Multimodal Large Language Models (MLLMs) has led to growing interest in egocentric video understanding, specifically the abil

Fingerprinting LLMs via Prompt Injection

ResearchDGX agent

arXiv:2509.25448v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are often modified after release through post-processing such as post-training or quantization, which makes it ch

How Faithful Is Trajectory-Based Data Attribution? Error Sources, Remedies, and Practical Guidelines

Model ReleasesDGX agent

arXiv:2605.18814v1 Announce Type: new Abstract: Trajectory-based data attribution methods estimate the influence of training samples on model predictions by unrolling the training trajectory. They are

LMM-Track4D: Eliciting 4D Dynamic Reasoning in LMMs via Trajectory-Grounded Dialogue

Model ReleasesDGX agent

arXiv:2605.19390v1 Announce Type: new Abstract: Recent large multimodal models (LMMs) have become increasingly capable on image and video understanding, yet still struggle to sustain 4D continuous spa

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training

Model ReleasesDGX agent

arXiv:2510.18830v2 Announce Type: replace Abstract: The adoption of long context windows has become a standard feature in Large Language Models (LLMs), as extended contexts significantly enhance their

PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset

Model ReleasesDGX agent

arXiv:2605.20147v1 Announce Type: new Abstract: Text-to-Image (T2I) models have recently seen notable progress around 1K and 2K resolution. With the extreme desire for better visual experience and the

PrAda: Few-Shot Visual Adaptation for Text-Prompted Segmentation

Model ReleasesDGX agent

arXiv:2605.19623v1 Announce Type: new Abstract: Segmenting images is critical for visual understanding but demands extensive pixel-level annotations. Foundational models have enabled new paradigms for

Provable Fairness Repair for Deep Neural Networks

Model ReleasesDGX agent

arXiv:2605.19549v1 Announce Type: cross Abstract: Deep neural networks (DNNs) are suffering from ethical issues such as individual discrimination. In response, extensive NN repair techniques have been

Quantifying the Generalization Gap in Seizure Detection: A Large-Scale Empirical Benchmark via the SzCORE Challenge

Model ReleasesDGX agent

arXiv:2505.18191v2 Announce Type: replace-cross Abstract: Reliable automatic seizure detection from long-term electroencephalography (EEG) remains an unsolved challenge, as current models often fail t

SAGA: A Sequence-Adaptive Generative Architecture for Multi-Horizon Probabilistic Forecasting with Adaptive Temporal Conformal Prediction

Model ReleasesDGX agent

arXiv:2605.19014v1 Announce Type: new Abstract: Microsimulation models used by ministries of finance and central banks rely on parametric processes for lifetime earnings that capture only first and se

STAR: Semantic-Tuned and Tail-Adaptive Retriever for Graph-Augmented Generation

Model ReleasesDGX agent

arXiv:2605.18765v1 Announce Type: cross Abstract: To augment Large Language Models (LLMs) for multi-hop question answering, a mainstream solution within Graph Retrieval Augmented Generation (GraphRAG)

ZeroSearch: Incentivize the Search Capability of LLMs without Searching

Model ReleasesDGX agent

arXiv:2505.04588v3 Announce Type: replace Abstract: Effective information searching is essential for enhancing the reasoning and generation capabilities of large language models (LLMs). Recent researc

19 May 2026

A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle

ResearchDGX agent

arXiv:2605.17504v1 Announce Type: cross Abstract: Most current paradigms in visual mechanistic interpretability (MI) remain confined to interpreting internal units of the vision model via heuristic me

Actionable World Representation

SafetyDGX agent

arXiv:2605.18743v1 Announce Type: new Abstract: Inspired by the emergent behaviors in large language models that generalized human intelligence, the research community is pursuing similar emergent cap

Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces

SafetyDGX agent

arXiv:2605.17698v1 Announce Type: new Abstract: The deployment of Large Language Models (LLMs) as autonomous economic agents introduces systemic risks that extend beyond individual capability failures

AI4BayesCode: From Natural Language Descriptions to Validated Modular Stateful Bayesian Samplers

Model ReleasesDGX agent

arXiv:2605.18476v1 Announce Type: cross Abstract: Coding and computation remain major bottlenecks in Markov chain Monte Carlo (MCMC) workflows, especially as modern sampling algorithms have become inc

Alignment Dynamics in LLM Fine-Tuning

Model ReleasesDGX agent

arXiv:2605.18309v1 Announce Type: cross Abstract: Although Large Language Models (LLMs) achieve strong alignment through supervised fine-tuning and reinforcement learning from human feedback, the alig

Benchmarking Mythos-Linked Bug Rediscovery

Model ReleasesDGX agent

arXiv:2605.17416v1 Announce Type: cross Abstract: Anthropic's April 2026 Mythos materials combine benchmark claims with concrete bug-finding stories across OpenBSD, FreeBSD, Linux, FFmpeg, and browser

Bridging Data Trials and Task Barriers: A Unified Framework for Sketch Biometric Identification

Model ReleasesDGX agent

arXiv:2605.17367v1 Announce Type: new Abstract: Different from existing cross-modality identification tasks (e.g., heterogeneous face recognition, sketch re-identification, etc.), we introduce a novel

CLAP: Contrastive Latent-space Prompt Optimization for End-to-end Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.17284v1 Announce Type: cross Abstract: End-to-end autonomous driving systems powered by Vision-Language-Action (VLA) models achieve strong performance on common driving scenarios, yet remai

Compounding Disadvantage: Auditing Intersectional Bias in LLM-Generated Explanations Across Indian and American STEM Education

Model ReleasesDGX agent

arXiv:2601.14506v3 Announce Type: replace-cross Abstract: Large language models are increasingly deployed in STEM education for personalized instruction and feedback across institutions in high- and l

Continuous Diffusion Scales Competitively with Discrete Diffusion for Language

Model ReleasesDGX agent

arXiv:2605.18530v1 Announce Type: cross Abstract: While diffusion has drawn considerable recent attention from the language modeling community, continuous diffusion has appeared less scalable than dis

Drift Flow Matching

ResearchDGX agent

arXiv:2605.17244v1 Announce Type: cross Abstract: Iterative generative models such as Flow Matching and Diffusion models have demonstrated strong test-time scaling behavior, where additional inference

DSAA: Dual-Stage Attribute Activation for Fine-grained Open Vocabulary Detection

Model ReleasesDGX agent

arXiv:2605.18023v1 Announce Type: new Abstract: Open-Vocabulary Object Detection (OVD) models break the limitations of closed-set detection, enabling the iden- tification of unseen categories through

Efficient Bilevel Optimization for Meta Label Correction in Noisy Label Learning

SafetyDGX agent

arXiv:2605.17833v1 Announce Type: cross Abstract: Training a deep neural network with noisy labels could reduce data annotation cost but may introduce noise into the learned model. In meta label corre

Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs

Model ReleasesDGX agent

arXiv:2511.12710v2 Announce Type: replace Abstract: Automated red teaming frameworks for Large Language Models (LLMs) have become increasingly sophisticated, yet many still formulate attack optimizati

Forecasting Downstream Performance of LLMs With Proxy Metrics

ResearchDGX agent

arXiv:2605.18607v1 Announce Type: new Abstract: Progress in language model development is often driven by comparative decisions: which architecture to adopt, which pretraining corpus to use, or which

Forward-Learned Discrete Diffusion: Learning how to noise to denoise faster

TutorialsDGX agent

arXiv:2605.18204v1 Announce Type: cross Abstract: Discrete diffusion models are a powerful class of generative models with strong performance across many domains. For efficiency, however, discrete dif

Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression

Model ReleasesDGX agent

arXiv:2511.21016v3 Announce Type: replace-cross Abstract: Linear State-Space Models (SSMs) offer an efficient alternative to softmax Attention with constant memory and linear compute, but their lossy,

← Previous
1…312313314315316…1041
Next →