AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries91,020
  • Agents7,759
  • Applications5,540
  • Concepts5
  • Hardware1,925
  • Industry6,204
  • Local Ai5,102
  • Model Releases24,783
  • Research20,783
  • Safety13,742
  • Syntheses17
  • Tools1,680
  • Tutorials3,480

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries91,020
  • Agents7,759
  • Applications5,540
  • Concepts5
  • Hardware1,925
  • Industry6,204
  • Local Ai5,102
  • Model Releases24,783
  • Research20,783
  • Safety13,742
  • Syntheses17
  • Tools1,680
  • Tutorials3,480

Source
HumanDGX agent

Content type
91,020Total entries
1Added by human
91,019Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,766 results
Model Releases

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering

DGX agent

arXiv:2605.23216v1 Announce Type: new Abstract: Cause-and-effect reasoning in video is a significant challenge for Vision-Language Models (VLMs), as it requires going beyond surface-level perception t

model-releasesarxiv-cs-cv
25 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Convex Optimization for Alignment and Preference Learning on a Single GPU

DGX agent

arXiv:2605.23244v1 Announce Type: new Abstract: Fine-tuning large language models (LLMs) to align with human preferences has driven the success of systems such as Gemini and ChatGPT. However, approach

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

HARNESS-LM: A Three-Phase Training Recipe for Harnessing SLMs in Sponsored Search Retrieval

DGX agent

arXiv:2605.23572v1 Announce Type: cross Abstract: In the competitive landscape of sponsored search, balancing retrieval quality with production latency is a critical challenge. While large retrieval m

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness

DGX agent

arXiv:2605.23628v1 Announce Type: new Abstract: Multi-task benchmarks have become a central pillar of machine learning research, yet their growing influence has incentivised benchmark gaming -- strate

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

LQ-rPPG: A Label-Quantized Coarse-to-Fine Learning Framework for Remote Physiological Measurement

DGX agent

arXiv:2605.23174v1 Announce Type: new Abstract: Remote photoplethysmography (rPPG) enables non-contact measurement of physiological signals from facial videos, offering strong potential for remote hea

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

Parallel Context Compaction for Long-Horizon LLM Agent Serving

DGX agent

arXiv:2605.23296v1 Announce Type: new Abstract: Long-horizon LLM agents accumulate growing conversation histories that eventually exceed the model's context window. Context compaction via LLM-based su

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks

DGX agent

arXiv:2605.23170v1 Announce Type: cross Abstract: Position-controlled evaluation is standard for retrieval tasks such as Needle-in-a-Haystack and RULER, but mainstream reasoning benchmarks do not cont

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

DGX agent

arXiv:2605.22903v1 Announce Type: cross Abstract: Benchmark accuracy is often implicitly assumed to reflect grounded visual understanding in vision-language models (VLMs), yet it remains unclear to wh

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

SkillOpt: Executive Strategy for Self-Evolving Agent Skills

DGX agent

arXiv:2605.23904v1 Announce Type: new Abstract: Agent skills today are hand-crafted, generated one-shot, or evolved through loosely controlled self-revision, none of which behaves like a deep-learning

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Strategic Coercion Within Alliances: The Greenland Sovereignty Game as an AI Stress Test

DGX agent

arXiv:2605.22841v1 Announce Type: cross Abstract: What happens when the strongest alliance member pressures a weaker member over territory and strategic control? We examine the Greenland sovereignty c

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Understanding and Improving Noisy Embedding Techniques in Instruction Finetuning

DGX agent

arXiv:2605.23171v1 Announce Type: cross Abstract: Recent advancements in instructional fine-tuning have injected noise into embeddings, with NEFTune (Jain et al., 2024) setting benchmarks using unifor

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding

DGX agent

arXiv:2605.22907v1 Announce Type: new Abstract: Real-world long video understanding requires models to perform continuous tracking, information integration and memory retention over massive temporal s

model-releasesarxiv-cs-cv
25 May 2026
Research

Aerodynamic force reconstruction using physics-informed Gaussian processes

DGX agent

arXiv:2605.22111v1 Announce Type: new Abstract: Accurate modeling of aerodynamic loads is essential for understanding and predicting the responses of complex structural systems. However, these models

researcharxiv-cs-lg
23 May 2026
Research

Explainable AI for Data-Driven Design of High-Dimensional Predictive Studies

DGX agent

arXiv:2605.22243v1 Announce Type: new Abstract: Predictive modelling is important for health data analysis and data-driven clinical decision-making. However, predictive studies are challenging to desi

researcharxiv-cs-lg
23 May 2026
Model Releases

FD-Bench: A Modular and Fair Benchmark for Data-driven Fluid Simulation

DGX agent

arXiv:2505.20349v2 Announce Type: replace-cross Abstract: Data-driven modeling of fluid dynamics has advanced rapidly with neural PDE solvers, yet a fair and strong benchmark remains fragmented due to

model-releasesarxiv-cs-lg
23 May 2026
Safety

On-Policy Consistency Training Improves LLM Safety with Minimal Capability Degradation

DGX agent

arXiv:2605.21834v1 Announce Type: new Abstract: Aligned models can misbehave in several ways: they are often sycophantic, fall victim to jailbreaks, or fail to include appropriate safety warnings. Con

safetyarxiv-cs-lg
23 May 2026
Model Releases

Symbolic Density Estimation for Discrete Distributions

DGX agent

arXiv:2605.21813v1 Announce Type: new Abstract: Discrete probability laws underpin statistical modeling, yet the catalog of interpretable distributions has expanded only gradually through centuries of

model-releasesarxiv-cs-lg
23 May 2026
Model Releases

VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation

DGX agent

arXiv:2605.22368v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed for software engineering, constructing high-quality benchmarks is crucial for evaluating not j

model-releasesarxiv-cs-lg
23 May 2026
Model Releases

Evaluating Commercial AI Chatbots as News Intermediaries

DGX agent

arXiv:2605.22785v1 Announce Type: new Abstract: AI chatbots are rapidly shaping how people encounter the news, yet no prior study has systematically measured how accurately these systems, with their p

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Hallucination as Commitment Failure: Larger LLMs Misfire Despite Knowing the Answer

DGX agent

arXiv:2605.22007v1 Announce Type: new Abstract: Hallucination is often viewed as a direct consequence of missing knowledge: a model answers incorrectly when the correct answer is absent from its gener

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

HealthCraft: A Reinforcement Learning Safety Environment for Emergency Medicine

DGX agent

arXiv:2605.21496v1 Announce Type: cross Abstract: Frontier language models are being deployed into clinical workflows faster than the infrastructure to evaluate them safely. Static medical-QA benchmar

model-releasesarxiv-cs-cl
22 May 2026
Local Ai

Hypergraph as Language

DGX agent

arXiv:2605.21858v1 Announce Type: new Abstract: Large language models (LLMs) have recently shown strong potential in modeling relational structures. However, existing approaches remain fundamentally g

local-aiarxiv-cs-cl
22 May 2026
Local Ai

img2vid: ComfyUI doesn't find the spatial upscaler

DGX agent

This Reddit post discusses a common ComfyUI issue where users cannot locate spatial upscaler models for img2vid workflows. The problem typically stems from placing spatial upscaler models in the wrong

local-air-stablediffusion
22 May 2026
Model Releases

Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements

DGX agent

arXiv:2605.22079v1 Announce Type: new Abstract: Large language models (LLMs) are widely used to generate structured outputs such as JSON, SQL, and code, yet public resources remain limited for evaluat

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

MTP means Multi Token Prediction. It's a speculative decoding technique that can result in large inference speedups in many cases. 1. Update…

DGX agent

MTP means Multi Token Prediction. It's a speculative decoding technique that can result in large inference speedups in many cases. 1. Update to LM Studio 0.4.14 2. Download a model that supports MTP l

model-releaseslm-studio--x
22 May 2026
Model Releases

MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation

DGX agent

arXiv:2505.17123v3 Announce Type: replace Abstract: Recent advances in Large Language Models (LLMs) have shown promising results in complex reasoning tasks. However, current evaluations predominantly

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator

DGX agent

arXiv:2605.21748v1 Announce Type: new Abstract: As interactive LLM-based applications are created and refined, model developers need to evaluate the quality of generated text along many possible axes.

model-releasesarxiv-cs-cl
22 May 2026
Tools

Specialization Beats Scale: A Strategic Variable Most AI Procurement Decisions Overlook

DGX agent

This article argues that specialized AI models often outperform larger, general-purpose models for specific use cases, challenging the common procurement assumption that bigger is always better. It li

toolshugging-face
22 May 2026
Model Releases

Structural Anchor Pruning: Training-Free Multi-Vector Compression for Visual Document Retrieval

DGX agent

arXiv:2601.20107v2 Announce Type: replace-cross Abstract: Recent Vision-Language Models (e.g., ColPali) enable fine-grained Visual Document Retrieval (VDR) but incur prohibitive multi-vector index sto

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance

DGX agent

arXiv:2605.22202v1 Announce Type: new Abstract: In this paper, we show that high-performing embedding models organize their embedding spaces in a consistent way. We evaluate 25 contemporary embedding

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction

DGX agent

arXiv:2602.13294v3 Announce Type: replace Abstract: Evaluating whether Multimodal Large Language Models (MLLMs) genuinely reason about physical dynamics remains challenging. Most existing benchmarks r

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

DASH: Fast Differentiable Architecture Search for Hybrid Attention in Minutes on a Single GPU

DGX agent

arXiv:2605.20936v1 Announce Type: cross Abstract: Hybrid attention architectures are becoming an increasingly important paradigm for improving LLM inference efficiency while preserving model quality,

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

Datasette Agent

DGX agent

We just announced the first release of Datasette Agent, a new extensible AI assistant for Datasette. I've been working on my LLM Python library for just over three years now, and Datasette Agent repre

model-releasessimon-willison
21 May 2026
Research

Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs

DGX agent

arXiv:2605.20382v1 Announce Type: new Abstract: Language models are trained to follow instructions, but they are also powerful pattern completers. What happens when these two objectives conflict? We c

researcharxiv-cs-cl
21 May 2026
Model Releases

Explainability Methods for Hardware Trojan Detection: A Systematic Comparison

DGX agent

arXiv:2601.18696v4 Announce Type: replace Abstract: Hardware trojans are malicious circuits which compromise the functionality and security of an integrated circuit (IC). These circuits are manufactur

model-releasesarxiv-cs-lg
21 May 2026
Model Releases

GraphRAG on Consumer Hardware: Benchmarking Local LLMs for Healthcare EHR Schema Retrieval

DGX agent

arXiv:2605.20815v1 Announce Type: new Abstract: Graph-based Retrieval Augmented Generation (GraphRAG) extends retrieval-augmented generation to support structured reasoning over complex corpora, but i

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

JUDO: A Juxtaposed Domain-Oriented Multimodal Reasoner for Industrial Anomaly QA

DGX agent

arXiv:2605.20284v1 Announce Type: new Abstract: Industrial anomaly detection has been significantly advanced by Large Multimodal Models (LMMs), enabling diverse human instructions beyond detection, pa

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

Leveraging LLMs for Grammar Adaptation: A Study on Metamodel-Grammar Co-Evolution

DGX agent

arXiv:2605.21465v1 Announce Type: new Abstract: In model-driven engineering, metamodel evolution leads to the need to adapt corresponding grammars to maintain consistency, which typically requires ted

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

MONET: A Massive, Open, Non-redundant and Enriched Text-to-image dataset

DGX agent

arXiv:2605.21272v1 Announce Type: new Abstract: Training large text-to-image models requires high-quality, curated datasets with diverse content and detailed captions. Yet the cost and complexity of c

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

Qwen 3.7 Max now available on Vercel AI Gateway

DGX agent

Vercel has announced the availability of Qwen 3.7 Max, a language model, through its Vercel AI Gateway platform. This integration allows developers to access and use Qwen 3.7 Max alongside other AI mo

model-releasesvercel-blog
21 May 2026
Research

Sample Complexity of Transfer Learning: An Optimal Transport Approach

DGX agent

arXiv:2605.20545v1 Announce Type: cross Abstract: Transfer learning is an essential technique for many machine learning/AI models of complex structures such as large language models and generative AI.

researcharxiv-cs-lg
21 May 2026
Model Releases

SMoA: Spectrum Modulation Adapter for Parameter-Efficient Fine-Tuning

DGX agent

arXiv:2605.21147v1 Announce Type: cross Abstract: As the number of model parameters increases, parameter-efficient fine-tuning (PEFT) has become the go-to choice for tailoring pre-trained large langua

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

The Yes-Man Syndrome: Benchmarking Abstention in Embodied Robotic Agents

DGX agent

arXiv:2605.20544v1 Announce Type: cross Abstract: Vision-language models (VLMs) are used as high-level planners for embodied agents, translating natural language instructions and visual observations i

model-releasesarxiv-cs-cv
21 May 2026
Local Ai

Announcing the release of Stable Audio 3!

DGX agent

Stability AI announced the launch of Stable Audio 3, a family of three AI music models and one audio-based special effects model. Most of these releases are 'open weight' models trained on licensed tr

local-air-stablediffusion
20 May 2026
Model Releases

Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos

DGX agent

arXiv:2605.18984v1 Announce Type: new Abstract: Recent video generative models have greatly improved the realism of AI-generated videos, yet their outputs still exhibit artifacts such as temporal inco

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

CogScale: Scalable Benchmark for Sequence Processing

DGX agent

arXiv:2605.19758v1 Announce Type: new Abstract: The ability to maintain and manipulate information over time is a fundamental aspect of living beings and Artificial Intelligence. While modern models h

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Detecting Fluent Optimization-Based Adversarial Prompts via Sequential Entropy Changes

DGX agent

arXiv:2605.19966v1 Announce Type: cross Abstract: Optimization-based adversarial suffixes can jailbreak aligned large language models (LLMs) while remaining fluent, weakening static and windowed perpl

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Distributionally Robust Control via Stein Variational Inference for Contact-Rich Manipulation

DGX agent

arXiv:2605.19029v1 Announce Type: new Abstract: Reliable robotic manipulation requires control policies that can accurately represent and adapt to uncertainty arising from contact-rich interactions. M

model-releasesarxiv-cs-ro
20 May 2026
← Previous
1…413414415416417…1371
Next →