AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,429
  • Agents7,559
  • Applications5,412
  • Concepts5
  • Hardware1,837
  • Industry6,170
  • Local Ai4,935
  • Model Releases23,918
  • Research20,125
  • Safety13,371
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,429
  • Agents7,559
  • Applications5,412
  • Concepts5
  • Hardware1,837
  • Industry6,170
  • Local Ai4,935
  • Model Releases23,918
  • Research20,125
  • Safety13,371
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

Content type
AllBlog
88,429Total entries
1Added by human
88,428Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,648 results
Model Releases

MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation

DGX agent

arXiv:2505.17123v3 Announce Type: replace Abstract: Recent advances in Large Language Models (LLMs) have shown promising results in complex reasoning tasks. However, current evaluations predominantly

model-releasesarxiv-cs-cl
22 May 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator

DGX agent

arXiv:2605.21748v1 Announce Type: new Abstract: As interactive LLM-based applications are created and refined, model developers need to evaluate the quality of generated text along many possible axes.

model-releasesarxiv-cs-cl
22 May 2026
Tools

Specialization Beats Scale: A Strategic Variable Most AI Procurement Decisions Overlook

DGX agent

This article argues that specialized AI models often outperform larger, general-purpose models for specific use cases, challenging the common procurement assumption that bigger is always better. It li

toolshugging-face
22 May 2026
Model Releases

Structural Anchor Pruning: Training-Free Multi-Vector Compression for Visual Document Retrieval

DGX agent

arXiv:2601.20107v2 Announce Type: replace-cross Abstract: Recent Vision-Language Models (e.g., ColPali) enable fine-grained Visual Document Retrieval (VDR) but incur prohibitive multi-vector index sto

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance

DGX agent

arXiv:2605.22202v1 Announce Type: new Abstract: In this paper, we show that high-performing embedding models organize their embedding spaces in a consistent way. We evaluate 25 contemporary embedding

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction

DGX agent

arXiv:2602.13294v3 Announce Type: replace Abstract: Evaluating whether Multimodal Large Language Models (MLLMs) genuinely reason about physical dynamics remains challenging. Most existing benchmarks r

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

DASH: Fast Differentiable Architecture Search for Hybrid Attention in Minutes on a Single GPU

DGX agent

arXiv:2605.20936v1 Announce Type: cross Abstract: Hybrid attention architectures are becoming an increasingly important paradigm for improving LLM inference efficiency while preserving model quality,

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

Datasette Agent

DGX agent

We just announced the first release of Datasette Agent, a new extensible AI assistant for Datasette. I've been working on my LLM Python library for just over three years now, and Datasette Agent repre

model-releasessimon-willison
21 May 2026
Research

Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs

DGX agent

arXiv:2605.20382v1 Announce Type: new Abstract: Language models are trained to follow instructions, but they are also powerful pattern completers. What happens when these two objectives conflict? We c

researcharxiv-cs-cl
21 May 2026
Model Releases

Explainability Methods for Hardware Trojan Detection: A Systematic Comparison

DGX agent

arXiv:2601.18696v4 Announce Type: replace Abstract: Hardware trojans are malicious circuits which compromise the functionality and security of an integrated circuit (IC). These circuits are manufactur

model-releasesarxiv-cs-lg
21 May 2026
Model Releases

GraphRAG on Consumer Hardware: Benchmarking Local LLMs for Healthcare EHR Schema Retrieval

DGX agent

arXiv:2605.20815v1 Announce Type: new Abstract: Graph-based Retrieval Augmented Generation (GraphRAG) extends retrieval-augmented generation to support structured reasoning over complex corpora, but i

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

JUDO: A Juxtaposed Domain-Oriented Multimodal Reasoner for Industrial Anomaly QA

DGX agent

arXiv:2605.20284v1 Announce Type: new Abstract: Industrial anomaly detection has been significantly advanced by Large Multimodal Models (LMMs), enabling diverse human instructions beyond detection, pa

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

Leveraging LLMs for Grammar Adaptation: A Study on Metamodel-Grammar Co-Evolution

DGX agent

arXiv:2605.21465v1 Announce Type: new Abstract: In model-driven engineering, metamodel evolution leads to the need to adapt corresponding grammars to maintain consistency, which typically requires ted

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

MONET: A Massive, Open, Non-redundant and Enriched Text-to-image dataset

DGX agent

arXiv:2605.21272v1 Announce Type: new Abstract: Training large text-to-image models requires high-quality, curated datasets with diverse content and detailed captions. Yet the cost and complexity of c

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

Qwen 3.7 Max now available on Vercel AI Gateway

DGX agent

Vercel has announced the availability of Qwen 3.7 Max, a language model, through its Vercel AI Gateway platform. This integration allows developers to access and use Qwen 3.7 Max alongside other AI mo

model-releasesvercel-blog
21 May 2026
Research

Sample Complexity of Transfer Learning: An Optimal Transport Approach

DGX agent

arXiv:2605.20545v1 Announce Type: cross Abstract: Transfer learning is an essential technique for many machine learning/AI models of complex structures such as large language models and generative AI.

researcharxiv-cs-lg
21 May 2026
Model Releases

SMoA: Spectrum Modulation Adapter for Parameter-Efficient Fine-Tuning

DGX agent

arXiv:2605.21147v1 Announce Type: cross Abstract: As the number of model parameters increases, parameter-efficient fine-tuning (PEFT) has become the go-to choice for tailoring pre-trained large langua

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

The Yes-Man Syndrome: Benchmarking Abstention in Embodied Robotic Agents

DGX agent

arXiv:2605.20544v1 Announce Type: cross Abstract: Vision-language models (VLMs) are used as high-level planners for embodied agents, translating natural language instructions and visual observations i

model-releasesarxiv-cs-cv
21 May 2026
Local Ai

Announcing the release of Stable Audio 3!

DGX agent

Stability AI announced the launch of Stable Audio 3, a family of three AI music models and one audio-based special effects model. Most of these releases are 'open weight' models trained on licensed tr

local-air-stablediffusion
20 May 2026
Model Releases

Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos

DGX agent

arXiv:2605.18984v1 Announce Type: new Abstract: Recent video generative models have greatly improved the realism of AI-generated videos, yet their outputs still exhibit artifacts such as temporal inco

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

CogScale: Scalable Benchmark for Sequence Processing

DGX agent

arXiv:2605.19758v1 Announce Type: new Abstract: The ability to maintain and manipulate information over time is a fundamental aspect of living beings and Artificial Intelligence. While modern models h

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Detecting Fluent Optimization-Based Adversarial Prompts via Sequential Entropy Changes

DGX agent

arXiv:2605.19966v1 Announce Type: cross Abstract: Optimization-based adversarial suffixes can jailbreak aligned large language models (LLMs) while remaining fluent, weakening static and windowed perpl

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Distributionally Robust Control via Stein Variational Inference for Contact-Rich Manipulation

DGX agent

arXiv:2605.19029v1 Announce Type: new Abstract: Reliable robotic manipulation requires control policies that can accurately represent and adapt to uncertainty arising from contact-rich interactions. M

model-releasesarxiv-cs-ro
20 May 2026
Model Releases

EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs

DGX agent

arXiv:2605.19559v1 Announce Type: cross Abstract: The rapid development of Multimodal Large Language Models (MLLMs) has led to growing interest in egocentric video understanding, specifically the abil

model-releasesarxiv-cs-ai
20 May 2026
Research

Fingerprinting LLMs via Prompt Injection

DGX agent

arXiv:2509.25448v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are often modified after release through post-processing such as post-training or quantization, which makes it ch

researcharxiv-cs-cl
20 May 2026
Model Releases

How Faithful Is Trajectory-Based Data Attribution? Error Sources, Remedies, and Practical Guidelines

DGX agent

arXiv:2605.18814v1 Announce Type: new Abstract: Trajectory-based data attribution methods estimate the influence of training samples on model predictions by unrolling the training trajectory. They are

model-releasesarxiv-cs-lg
20 May 2026
Model Releases

LMM-Track4D: Eliciting 4D Dynamic Reasoning in LMMs via Trajectory-Grounded Dialogue

DGX agent

arXiv:2605.19390v1 Announce Type: new Abstract: Recent large multimodal models (LMMs) have become increasingly capable on image and video understanding, yet still struggle to sustain 4D continuous spa

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training

DGX agent

arXiv:2510.18830v2 Announce Type: replace Abstract: The adoption of long context windows has become a standard feature in Large Language Models (LLMs), as extended contexts significantly enhance their

model-releasesarxiv-cs-cl
20 May 2026
Model Releases

PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset

DGX agent

arXiv:2605.20147v1 Announce Type: new Abstract: Text-to-Image (T2I) models have recently seen notable progress around 1K and 2K resolution. With the extreme desire for better visual experience and the

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

PrAda: Few-Shot Visual Adaptation for Text-Prompted Segmentation

DGX agent

arXiv:2605.19623v1 Announce Type: new Abstract: Segmenting images is critical for visual understanding but demands extensive pixel-level annotations. Foundational models have enabled new paradigms for

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

Provable Fairness Repair for Deep Neural Networks

DGX agent

arXiv:2605.19549v1 Announce Type: cross Abstract: Deep neural networks (DNNs) are suffering from ethical issues such as individual discrimination. In response, extensive NN repair techniques have been

model-releasesarxiv-cs-lg
20 May 2026
Model Releases

Quantifying the Generalization Gap in Seizure Detection: A Large-Scale Empirical Benchmark via the SzCORE Challenge

DGX agent

arXiv:2505.18191v2 Announce Type: replace-cross Abstract: Reliable automatic seizure detection from long-term electroencephalography (EEG) remains an unsolved challenge, as current models often fail t

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

SAGA: A Sequence-Adaptive Generative Architecture for Multi-Horizon Probabilistic Forecasting with Adaptive Temporal Conformal Prediction

DGX agent

arXiv:2605.19014v1 Announce Type: new Abstract: Microsimulation models used by ministries of finance and central banks rely on parametric processes for lifetime earnings that capture only first and se

model-releasesarxiv-cs-lg
20 May 2026
Model Releases

STAR: Semantic-Tuned and Tail-Adaptive Retriever for Graph-Augmented Generation

DGX agent

arXiv:2605.18765v1 Announce Type: cross Abstract: To augment Large Language Models (LLMs) for multi-hop question answering, a mainstream solution within Graph Retrieval Augmented Generation (GraphRAG)

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

ZeroSearch: Incentivize the Search Capability of LLMs without Searching

DGX agent

arXiv:2505.04588v3 Announce Type: replace Abstract: Effective information searching is essential for enhancing the reasoning and generation capabilities of large language models (LLMs). Recent researc

model-releasesarxiv-cs-cl
20 May 2026
Research

A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle

DGX agent

arXiv:2605.17504v1 Announce Type: cross Abstract: Most current paradigms in visual mechanistic interpretability (MI) remain confined to interpreting internal units of the vision model via heuristic me

researcharxiv-cs-ai
19 May 2026
Safety

Actionable World Representation

DGX agent

arXiv:2605.18743v1 Announce Type: new Abstract: Inspired by the emergent behaviors in large language models that generalized human intelligence, the research community is pursuing similar emergent cap

safetyarxiv-cs-ai
19 May 2026
Safety

Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces

DGX agent

arXiv:2605.17698v1 Announce Type: new Abstract: The deployment of Large Language Models (LLMs) as autonomous economic agents introduces systemic risks that extend beyond individual capability failures

safetyarxiv-cs-lg
19 May 2026
Model Releases

AI4BayesCode: From Natural Language Descriptions to Validated Modular Stateful Bayesian Samplers

DGX agent

arXiv:2605.18476v1 Announce Type: cross Abstract: Coding and computation remain major bottlenecks in Markov chain Monte Carlo (MCMC) workflows, especially as modern sampling algorithms have become inc

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Alignment Dynamics in LLM Fine-Tuning

DGX agent

arXiv:2605.18309v1 Announce Type: cross Abstract: Although Large Language Models (LLMs) achieve strong alignment through supervised fine-tuning and reinforcement learning from human feedback, the alig

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Benchmarking Mythos-Linked Bug Rediscovery

DGX agent

arXiv:2605.17416v1 Announce Type: cross Abstract: Anthropic's April 2026 Mythos materials combine benchmark claims with concrete bug-finding stories across OpenBSD, FreeBSD, Linux, FFmpeg, and browser

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Bridging Data Trials and Task Barriers: A Unified Framework for Sketch Biometric Identification

DGX agent

arXiv:2605.17367v1 Announce Type: new Abstract: Different from existing cross-modality identification tasks (e.g., heterogeneous face recognition, sketch re-identification, etc.), we introduce a novel

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

CLAP: Contrastive Latent-space Prompt Optimization for End-to-end Autonomous Driving

DGX agent

arXiv:2605.17284v1 Announce Type: cross Abstract: End-to-end autonomous driving systems powered by Vision-Language-Action (VLA) models achieve strong performance on common driving scenarios, yet remai

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Compounding Disadvantage: Auditing Intersectional Bias in LLM-Generated Explanations Across Indian and American STEM Education

DGX agent

arXiv:2601.14506v3 Announce Type: replace-cross Abstract: Large language models are increasingly deployed in STEM education for personalized instruction and feedback across institutions in high- and l

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

Continuous Diffusion Scales Competitively with Discrete Diffusion for Language

DGX agent

arXiv:2605.18530v1 Announce Type: cross Abstract: While diffusion has drawn considerable recent attention from the language modeling community, continuous diffusion has appeared less scalable than dis

model-releasesarxiv-cs-ai
19 May 2026
Research

Drift Flow Matching

DGX agent

arXiv:2605.17244v1 Announce Type: cross Abstract: Iterative generative models such as Flow Matching and Diffusion models have demonstrated strong test-time scaling behavior, where additional inference

researcharxiv-cs-ai
19 May 2026
Model Releases

DSAA: Dual-Stage Attribute Activation for Fine-grained Open Vocabulary Detection

DGX agent

arXiv:2605.18023v1 Announce Type: new Abstract: Open-Vocabulary Object Detection (OVD) models break the limitations of closed-set detection, enabling the iden- tification of unseen categories through

model-releasesarxiv-cs-cv
19 May 2026
Safety

Efficient Bilevel Optimization for Meta Label Correction in Noisy Label Learning

DGX agent

arXiv:2605.17833v1 Announce Type: cross Abstract: Training a deep neural network with noisy labels could reduce data annotation cost but may introduce noise into the learned model. In meta label corre

safetyarxiv-cs-ai
19 May 2026
← Previous
1…399400401402403…1326
Next →