AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,993
  • Agents7,449
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,136
  • Local Ai4,859
  • Model Releases23,375
  • Research19,835
  • Safety13,176
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,993
  • Agents7,449
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,136
  • Local Ai4,859
  • Model Releases23,375
  • Research19,835
  • Safety13,176
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
86,993Total entries
1Added by human
86,992Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
51,106 results
Research

Aerodynamic force reconstruction using physics-informed Gaussian processes

DGX agent

arXiv:2605.22111v1 Announce Type: new Abstract: Accurate modeling of aerodynamic loads is essential for understanding and predicting the responses of complex structural systems. However, these models

researcharxiv-cs-lg
23 May 2026
Research
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Explainable AI for Data-Driven Design of High-Dimensional Predictive Studies

DGX agent

arXiv:2605.22243v1 Announce Type: new Abstract: Predictive modelling is important for health data analysis and data-driven clinical decision-making. However, predictive studies are challenging to desi

researcharxiv-cs-lg
23 May 2026
Model Releases

FD-Bench: A Modular and Fair Benchmark for Data-driven Fluid Simulation

DGX agent

arXiv:2505.20349v2 Announce Type: replace-cross Abstract: Data-driven modeling of fluid dynamics has advanced rapidly with neural PDE solvers, yet a fair and strong benchmark remains fragmented due to

model-releasesarxiv-cs-lg
23 May 2026
Safety

On-Policy Consistency Training Improves LLM Safety with Minimal Capability Degradation

DGX agent

arXiv:2605.21834v1 Announce Type: new Abstract: Aligned models can misbehave in several ways: they are often sycophantic, fall victim to jailbreaks, or fail to include appropriate safety warnings. Con

safetyarxiv-cs-lg
23 May 2026
Model Releases

Symbolic Density Estimation for Discrete Distributions

DGX agent

arXiv:2605.21813v1 Announce Type: new Abstract: Discrete probability laws underpin statistical modeling, yet the catalog of interpretable distributions has expanded only gradually through centuries of

model-releasesarxiv-cs-lg
23 May 2026
Model Releases

VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation

DGX agent

arXiv:2605.22368v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed for software engineering, constructing high-quality benchmarks is crucial for evaluating not j

model-releasesarxiv-cs-lg
23 May 2026
Model Releases

Evaluating Commercial AI Chatbots as News Intermediaries

DGX agent

arXiv:2605.22785v1 Announce Type: new Abstract: AI chatbots are rapidly shaping how people encounter the news, yet no prior study has systematically measured how accurately these systems, with their p

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Hallucination as Commitment Failure: Larger LLMs Misfire Despite Knowing the Answer

DGX agent

arXiv:2605.22007v1 Announce Type: new Abstract: Hallucination is often viewed as a direct consequence of missing knowledge: a model answers incorrectly when the correct answer is absent from its gener

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

HealthCraft: A Reinforcement Learning Safety Environment for Emergency Medicine

DGX agent

arXiv:2605.21496v1 Announce Type: cross Abstract: Frontier language models are being deployed into clinical workflows faster than the infrastructure to evaluate them safely. Static medical-QA benchmar

model-releasesarxiv-cs-cl
22 May 2026
Local Ai

Hypergraph as Language

DGX agent

arXiv:2605.21858v1 Announce Type: new Abstract: Large language models (LLMs) have recently shown strong potential in modeling relational structures. However, existing approaches remain fundamentally g

local-aiarxiv-cs-cl
22 May 2026
Model Releases

Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements

DGX agent

arXiv:2605.22079v1 Announce Type: new Abstract: Large language models (LLMs) are widely used to generate structured outputs such as JSON, SQL, and code, yet public resources remain limited for evaluat

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation

DGX agent

arXiv:2505.17123v3 Announce Type: replace Abstract: Recent advances in Large Language Models (LLMs) have shown promising results in complex reasoning tasks. However, current evaluations predominantly

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator

DGX agent

arXiv:2605.21748v1 Announce Type: new Abstract: As interactive LLM-based applications are created and refined, model developers need to evaluate the quality of generated text along many possible axes.

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Structural Anchor Pruning: Training-Free Multi-Vector Compression for Visual Document Retrieval

DGX agent

arXiv:2601.20107v2 Announce Type: replace-cross Abstract: Recent Vision-Language Models (e.g., ColPali) enable fine-grained Visual Document Retrieval (VDR) but incur prohibitive multi-vector index sto

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance

DGX agent

arXiv:2605.22202v1 Announce Type: new Abstract: In this paper, we show that high-performing embedding models organize their embedding spaces in a consistent way. We evaluate 25 contemporary embedding

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction

DGX agent

arXiv:2602.13294v3 Announce Type: replace Abstract: Evaluating whether Multimodal Large Language Models (MLLMs) genuinely reason about physical dynamics remains challenging. Most existing benchmarks r

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

DASH: Fast Differentiable Architecture Search for Hybrid Attention in Minutes on a Single GPU

DGX agent

arXiv:2605.20936v1 Announce Type: cross Abstract: Hybrid attention architectures are becoming an increasingly important paradigm for improving LLM inference efficiency while preserving model quality,

model-releasesarxiv-cs-cl
21 May 2026
Research

Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs

DGX agent

arXiv:2605.20382v1 Announce Type: new Abstract: Language models are trained to follow instructions, but they are also powerful pattern completers. What happens when these two objectives conflict? We c

researcharxiv-cs-cl
21 May 2026
Model Releases

Explainability Methods for Hardware Trojan Detection: A Systematic Comparison

DGX agent

arXiv:2601.18696v4 Announce Type: replace Abstract: Hardware trojans are malicious circuits which compromise the functionality and security of an integrated circuit (IC). These circuits are manufactur

model-releasesarxiv-cs-lg
21 May 2026
Model Releases

GraphRAG on Consumer Hardware: Benchmarking Local LLMs for Healthcare EHR Schema Retrieval

DGX agent

arXiv:2605.20815v1 Announce Type: new Abstract: Graph-based Retrieval Augmented Generation (GraphRAG) extends retrieval-augmented generation to support structured reasoning over complex corpora, but i

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

JUDO: A Juxtaposed Domain-Oriented Multimodal Reasoner for Industrial Anomaly QA

DGX agent

arXiv:2605.20284v1 Announce Type: new Abstract: Industrial anomaly detection has been significantly advanced by Large Multimodal Models (LMMs), enabling diverse human instructions beyond detection, pa

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

Leveraging LLMs for Grammar Adaptation: A Study on Metamodel-Grammar Co-Evolution

DGX agent

arXiv:2605.21465v1 Announce Type: new Abstract: In model-driven engineering, metamodel evolution leads to the need to adapt corresponding grammars to maintain consistency, which typically requires ted

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

MONET: A Massive, Open, Non-redundant and Enriched Text-to-image dataset

DGX agent

arXiv:2605.21272v1 Announce Type: new Abstract: Training large text-to-image models requires high-quality, curated datasets with diverse content and detailed captions. Yet the cost and complexity of c

model-releasesarxiv-cs-cv
21 May 2026
Research

Sample Complexity of Transfer Learning: An Optimal Transport Approach

DGX agent

arXiv:2605.20545v1 Announce Type: cross Abstract: Transfer learning is an essential technique for many machine learning/AI models of complex structures such as large language models and generative AI.

researcharxiv-cs-lg
21 May 2026
Model Releases

SMoA: Spectrum Modulation Adapter for Parameter-Efficient Fine-Tuning

DGX agent

arXiv:2605.21147v1 Announce Type: cross Abstract: As the number of model parameters increases, parameter-efficient fine-tuning (PEFT) has become the go-to choice for tailoring pre-trained large langua

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

The Yes-Man Syndrome: Benchmarking Abstention in Embodied Robotic Agents

DGX agent

arXiv:2605.20544v1 Announce Type: cross Abstract: Vision-language models (VLMs) are used as high-level planners for embodied agents, translating natural language instructions and visual observations i

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos

DGX agent

arXiv:2605.18984v1 Announce Type: new Abstract: Recent video generative models have greatly improved the realism of AI-generated videos, yet their outputs still exhibit artifacts such as temporal inco

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

CogScale: Scalable Benchmark for Sequence Processing

DGX agent

arXiv:2605.19758v1 Announce Type: new Abstract: The ability to maintain and manipulate information over time is a fundamental aspect of living beings and Artificial Intelligence. While modern models h

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Detecting Fluent Optimization-Based Adversarial Prompts via Sequential Entropy Changes

DGX agent

arXiv:2605.19966v1 Announce Type: cross Abstract: Optimization-based adversarial suffixes can jailbreak aligned large language models (LLMs) while remaining fluent, weakening static and windowed perpl

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Distributionally Robust Control via Stein Variational Inference for Contact-Rich Manipulation

DGX agent

arXiv:2605.19029v1 Announce Type: new Abstract: Reliable robotic manipulation requires control policies that can accurately represent and adapt to uncertainty arising from contact-rich interactions. M

model-releasesarxiv-cs-ro
20 May 2026
Model Releases

EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs

DGX agent

arXiv:2605.19559v1 Announce Type: cross Abstract: The rapid development of Multimodal Large Language Models (MLLMs) has led to growing interest in egocentric video understanding, specifically the abil

model-releasesarxiv-cs-ai
20 May 2026
Research

Fingerprinting LLMs via Prompt Injection

DGX agent

arXiv:2509.25448v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are often modified after release through post-processing such as post-training or quantization, which makes it ch

researcharxiv-cs-cl
20 May 2026
Model Releases

How Faithful Is Trajectory-Based Data Attribution? Error Sources, Remedies, and Practical Guidelines

DGX agent

arXiv:2605.18814v1 Announce Type: new Abstract: Trajectory-based data attribution methods estimate the influence of training samples on model predictions by unrolling the training trajectory. They are

model-releasesarxiv-cs-lg
20 May 2026
Model Releases

LMM-Track4D: Eliciting 4D Dynamic Reasoning in LMMs via Trajectory-Grounded Dialogue

DGX agent

arXiv:2605.19390v1 Announce Type: new Abstract: Recent large multimodal models (LMMs) have become increasingly capable on image and video understanding, yet still struggle to sustain 4D continuous spa

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training

DGX agent

arXiv:2510.18830v2 Announce Type: replace Abstract: The adoption of long context windows has become a standard feature in Large Language Models (LLMs), as extended contexts significantly enhance their

model-releasesarxiv-cs-cl
20 May 2026
Model Releases

PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset

DGX agent

arXiv:2605.20147v1 Announce Type: new Abstract: Text-to-Image (T2I) models have recently seen notable progress around 1K and 2K resolution. With the extreme desire for better visual experience and the

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

PrAda: Few-Shot Visual Adaptation for Text-Prompted Segmentation

DGX agent

arXiv:2605.19623v1 Announce Type: new Abstract: Segmenting images is critical for visual understanding but demands extensive pixel-level annotations. Foundational models have enabled new paradigms for

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

Provable Fairness Repair for Deep Neural Networks

DGX agent

arXiv:2605.19549v1 Announce Type: cross Abstract: Deep neural networks (DNNs) are suffering from ethical issues such as individual discrimination. In response, extensive NN repair techniques have been

model-releasesarxiv-cs-lg
20 May 2026
Model Releases

Quantifying the Generalization Gap in Seizure Detection: A Large-Scale Empirical Benchmark via the SzCORE Challenge

DGX agent

arXiv:2505.18191v2 Announce Type: replace-cross Abstract: Reliable automatic seizure detection from long-term electroencephalography (EEG) remains an unsolved challenge, as current models often fail t

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

SAGA: A Sequence-Adaptive Generative Architecture for Multi-Horizon Probabilistic Forecasting with Adaptive Temporal Conformal Prediction

DGX agent

arXiv:2605.19014v1 Announce Type: new Abstract: Microsimulation models used by ministries of finance and central banks rely on parametric processes for lifetime earnings that capture only first and se

model-releasesarxiv-cs-lg
20 May 2026
Model Releases

STAR: Semantic-Tuned and Tail-Adaptive Retriever for Graph-Augmented Generation

DGX agent

arXiv:2605.18765v1 Announce Type: cross Abstract: To augment Large Language Models (LLMs) for multi-hop question answering, a mainstream solution within Graph Retrieval Augmented Generation (GraphRAG)

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

ZeroSearch: Incentivize the Search Capability of LLMs without Searching

DGX agent

arXiv:2505.04588v3 Announce Type: replace Abstract: Effective information searching is essential for enhancing the reasoning and generation capabilities of large language models (LLMs). Recent researc

model-releasesarxiv-cs-cl
20 May 2026
Research

A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle

DGX agent

arXiv:2605.17504v1 Announce Type: cross Abstract: Most current paradigms in visual mechanistic interpretability (MI) remain confined to interpreting internal units of the vision model via heuristic me

researcharxiv-cs-ai
19 May 2026
Safety

Actionable World Representation

DGX agent

arXiv:2605.18743v1 Announce Type: new Abstract: Inspired by the emergent behaviors in large language models that generalized human intelligence, the research community is pursuing similar emergent cap

safetyarxiv-cs-ai
19 May 2026
Safety

Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces

DGX agent

arXiv:2605.17698v1 Announce Type: new Abstract: The deployment of Large Language Models (LLMs) as autonomous economic agents introduces systemic risks that extend beyond individual capability failures

safetyarxiv-cs-lg
19 May 2026
Model Releases

AI4BayesCode: From Natural Language Descriptions to Validated Modular Stateful Bayesian Samplers

DGX agent

arXiv:2605.18476v1 Announce Type: cross Abstract: Coding and computation remain major bottlenecks in Markov chain Monte Carlo (MCMC) workflows, especially as modern sampling algorithms have become inc

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Alignment Dynamics in LLM Fine-Tuning

DGX agent

arXiv:2605.18309v1 Announce Type: cross Abstract: Although Large Language Models (LLMs) achieve strong alignment through supervised fine-tuning and reinforcement learning from human feedback, the alig

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Benchmarking Mythos-Linked Bug Rediscovery

DGX agent

arXiv:2605.17416v1 Announce Type: cross Abstract: Anthropic's April 2026 Mythos materials combine benchmark claims with concrete bug-finding stories across OpenBSD, FreeBSD, Linux, FFmpeg, and browser

model-releasesarxiv-cs-ai
19 May 2026
← Previous
1…321322323324325…1065
Next →