AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries92,387
  • Agents7,863
  • Applications5,605
  • Concepts5
  • Hardware1,963
  • Industry6,238
  • Local Ai5,173
  • Model Releases25,258
  • Research21,121
  • Safety13,950
  • Syntheses17
  • Tools1,680
  • Tutorials3,514

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries92,387
  • Agents7,863
  • Applications5,605
  • Concepts5
  • Hardware1,963
  • Industry6,238
  • Local Ai5,173
  • Model Releases25,258
  • Research21,121
  • Safety13,950
  • Syntheses17
  • Tools1,680
  • Tutorials3,514

Source
HumanDGX agent

Content type
92,387Total entries
1Added by human
92,386Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
66,911 results
Applications

Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts

DGX agent

arXiv:2509.21892v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) models typically fix the number of activated experts k at both training and inference. However, real-world deployment

applicationsarxiv-cs-ai
12 May 2026
Research

ELF: Embedded Language Flows

AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2605.10938v1 Announce Type: cross Abstract: Diffusion and flow-based models have become the de facto approaches for generating continuous data, e.g., in domains such as images and videos. Their

researcharxiv-cs-ai
12 May 2026
Model Releases

ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room

DGX agent

arXiv:2505.22919v3 Announce Type: replace Abstract: Existing benchmarks for evaluating the clinical reasoning capabilities of large language models (LLMs) often lack a clear definition of 'clinical re

model-releasesarxiv-cs-cl
12 May 2026
Safety

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment

DGX agent

arXiv:2601.21484v2 Announce Type: replace Abstract: Reinforcement Learning (RL) post-training alignment for language models is effective, but also costly and unstable in practice, owing to its complic

safetyarxiv-cs-lg
12 May 2026
Research

Factual recall in linear associative memories: sharp asymptotics and mechanistic insights

DGX agent

arXiv:2605.10795v1 Announce Type: cross Abstract: Large language models demonstrate remarkable ability in factual recall, yet the fundamental limits of storing and retrieving input--output association

researcharxiv-cs-lg
12 May 2026
Model Releases

Follow the Mean: Reference-Guided Flow Matching

DGX agent

arXiv:2605.10302v1 Announce Type: new Abstract: Existing approaches to controllable generation typically rely on fine-tuning, auxiliary networks, or test-time search. We show that flow matching admits

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs

DGX agent

arXiv:2605.08905v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable success on reasoning benchmarks through Reinforcement Learning with Verifiable Rewards (RLVR), exc

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Frame In, Frame Out: Measuring Framing Bias in LLM-Generated News Summaries

DGX agent

arXiv:2505.05406v2 Announce Type: replace Abstract: News headlines and summaries shape how events are interpreted through selective emphasis and omission, a phenomenon commonly referred to as framing.

model-releasesarxiv-cs-cl
12 May 2026
Safety

Frequency Adapter with SAM for Generalized Medical Image Segmentation

DGX agent

arXiv:2605.09925v1 Announce Type: new Abstract: Medical image segmentation is a critical task in computer-aided diagnosis and treatment planning. However, deep learning models often struggle to genera

safetyarxiv-cs-cv
12 May 2026
Model Releases

Generating Symmetric Materials using Latent Flow Matching

DGX agent

arXiv:2605.10115v1 Announce Type: new Abstract: Tackling the task of materials generation, we aim to enhance the previously proposed All-atom Diffusion Transformer (ADiT) by introducing SymADiT, a sym

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

GenMed: A Pairwise Generative Reformulation of Medical Diagnostic Tasks

DGX agent

arXiv:2605.10645v1 Announce Type: new Abstract: Data-driven medical AI is traditionally formulated as a discriminative mapping from input X to output Y via a learned function f, which does not general

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

GONE: Structural Knowledge Unlearning via Neighborhood-Expanded Distribution Shaping

DGX agent

arXiv:2603.12275v1 Announce Type: cross Abstract: Unlearning knowledge is a pressing and challenging task in Large Language Models (LLMs) because of their unprecedented capability to memorize and dige

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

GravityGraphSAGE: Link Prediction in Directed Attributed Graphs

DGX agent

arXiv:2605.09408v1 Announce Type: new Abstract: Link prediction (inferring missing or future connections between nodes in a graph) is a fundamental problem in network science with widespread applicati

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

HiDrive: A Closed-Loop Benchmark for High-Level Autonomous Driving

DGX agent

arXiv:2605.09972v1 Announce Type: cross Abstract: End-to-end autonomous driving has witnessed rapid progress, yet existing benchmarks are increasingly saturated, with state-of-the-art models achieving

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Higher-Order Equilibrium Tracking for EM-Compressible Online Estimation

DGX agent

arXiv:2605.08864v1 Announce Type: new Abstract: We study online estimation in latent-variable models by recasting the problem as tracking a moving empirical equilibrium. Standard online EM and stochas

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

HOME-KGQA: A Benchmark Dataset for Multimodal Knowledge Graph Question Answering on Household Daily Activities

DGX agent

arXiv:2605.09348v1 Announce Type: cross Abstract: Large Language Models (LLMs) provide flexible natural language processing capabilities, while knowledge graphs (KGs) offer explicit and structured kno

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

How NVIDIA engineers and researchers build with Codex

DGX agent

NVIDIA engineers and researchers utilize OpenAI's Codex, a large language model trained on code, to accelerate software development and improve productivity across their engineering workflows. The art

model-releasesopenai
12 May 2026
Local Ai

I built ForgePilot: a Codex-style desktop workspace for Ollama with tools, MCP, web research, and document support

DGX agent

ForgePilot is a desktop workspace application designed for Ollama that combines local language model capabilities with development tools, including support for Model Context Protocol (MCP), web resear

local-air-ollama
12 May 2026
Model Releases

Incremental Multilingual Text2Cypher with Adapter Combination

DGX agent

arXiv:2601.16097v2 Announce Type: replace Abstract: Large Language Models enable users to access database using natural language interfaces using tools like Text2SQL, Text2SPARQL, and Text2Cypher, whi

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs

DGX agent

arXiv:2605.10267v1 Announce Type: new Abstract: In industrial procurement, an LLM answer is useful only if it survives a standards check: recommended material must match operating condition, every par

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Instruction Adherence in Coding Agent Configuration Files: A Factorial Study of Four File-Structure Variables

DGX agent

arXiv:2605.10039v1 Announce Type: cross Abstract: Frontier coding agents read configuration files (CLAUDE.md, AGENTS.md, Cursor Rules) at session start and are expected to follow the conventions insid

model-releasesarxiv-cs-cl
12 May 2026
Research

Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs

DGX agent

arXiv:2605.10633v1 Announce Type: cross Abstract: Fine-tuning Large Language Models (LLMs) on benign narrow data can sometimes induce broad harmful behaviors, a vulnerability termed emergent misalignm

researcharxiv-cs-ai
12 May 2026
Safety

Iterative Critique-and-Routing Controller for Multi-Agent Systems with Heterogeneous LLMs

DGX agent

arXiv:2605.08686v1 Announce Type: new Abstract: Multi-agent large language model (LLM) systems often rely on a controller to coordinate a pool of heterogeneous models, yet existing controllers are typ

safetyarxiv-cs-ai
12 May 2026
Model Releases

ITLC at SemEval-2026 Task 11: Normalization and Deterministic Parsing for Formal Reasoning in LLMs

DGX agent

arXiv:2603.02676v2 Announce Type: replace-cross Abstract: Large language models suffer from content effects in reasoning tasks, particularly in multi-lingual contexts. We introduce a novel method that

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

JODA: Composable Joint Dynamics for Articulated Objects

DGX agent

arXiv:2605.09954v1 Announce Type: cross Abstract: Articulated objects used in simulation and embodied AI are typically specified by geometry and kinematic structure, but lack the fine-grained dynamica

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

KAN Text to Vision? The Exploration of Kolmogorov-Arnold Networks for Multi-Scale Sequence-Based Pose Animation from Sign Language Notation

DGX agent

arXiv:2605.09572v1 Announce Type: cross Abstract: Sign language production from symbolic notation offers a scalable route to accessible sign animation. We present KANMultiSign, a multi-scale sequence

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

KARMA-MV: A Benchmark for Causal Question Answering on Music Videos

DGX agent

arXiv:2605.08175v1 Announce Type: cross Abstract: While significant progress has been made in Video Question Answering and cross-modal understanding, causal reasoning about how visual dynamics drive m

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

KEPIL: Knowledge-Enhanced Prompt-Image Learning for Prompt-Robust Disease Detection

DGX agent

arXiv:2605.09132v1 Announce Type: new Abstract: Vision--language models (VLMs) show promise for clinical decision support in radiology because they enable joint reasoning over radiological images and

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs

DGX agent

arXiv:2602.21198v2 Announce Type: replace-cross Abstract: Embodied LLMs endow robots with high-level task reasoning, but they cannot reflect on what went wrong or why, turning deployment into a sequen

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Learning to Perceive 'Where': Spatial Pretext Tasks for Robust Self-Supervised Learning

DGX agent

arXiv:2605.09963v1 Announce Type: new Abstract: Existing self-supervised learning (SSL) methods primarily learn object-invariant representations but often neglect the spatial structure and relationshi

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

LightAVSeg: Lightweight Audio-Visual Segmentation

DGX agent

arXiv:2605.08805v1 Announce Type: new Abstract: Audio-Visual Segmentation (AVS) targets pixel level localization of sounding emitting objects in videos. However, existing models rely on dense cross-mo

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM Systems

DGX agent

arXiv:2605.08305v1 Announce Type: cross Abstract: Large Language Model (LLM) systems have been the frontier of AI in many application domains, leading to new challenges and opportunities for hyperpara

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks

DGX agent

arXiv:2605.08437v1 Announce Type: cross Abstract: Existing benchmarks for legal AI focus primarily on tasks where LLMs must produce legal arguments or documents, yet the capacity to judge such argumen

model-releasesarxiv-cs-ai
12 May 2026
Safety

MapFormer: Self-Supervised Learning of Cognitive Maps with Input-Dependent Positional Embeddings

DGX agent

arXiv:2511.19279v4 Announce Type: replace-cross Abstract: A cognitive map is an internal model which encodes the abstract relationships among entities in the world, giving humans and animals the flexi

safetyarxiv-cs-cl
12 May 2026
Model Releases

MathConstraint: Automated Generation of Verified Combinatorial Reasoning Instances for LLMs

DGX agent

arXiv:2605.08498v1 Announce Type: cross Abstract: We introduce MathConstraint, a hard, adaptive benchmark for evaluating the combinatorial reasoning capabilities of LLMs. We combine constraint satisfa

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics

DGX agent

arXiv:2602.02561v2 Announce Type: replace-cross Abstract: While the ecosystem of Lean and Mathlib has enjoyed celebrated success in formal mathematical reasoning with the help of large language models

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests

DGX agent

arXiv:2510.22170v2 Announce Type: replace Abstract: Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure o

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Medical Incident Causal Factors and Preventive Measures Generation Using Tag-based Example Selection in Few-shot Learning

DGX agent

arXiv:2605.10025v1 Announce Type: cross Abstract: In high-stakes domains such as healthcare, the reliability of Large Language Models (LLMs) is critical, particularly when generating clinical insights

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All?

DGX agent

arXiv:2605.08816v1 Announce Type: new Abstract: In the animal kingdom, mirror self-recognition is a canonical probe of higher-order cognition, emerging only in some species. We ask whether an analogou

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI

DGX agent

arXiv:2605.08678v1 Announce Type: new Abstract: Modern AI progress has been driven by ML methods that are generalizable across settings and scalable to larger regimes. As large language models demonst

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

MMVIAD: Multi-view Multi-task Video Understanding for Industrial Anomaly Detection

DGX agent

arXiv:2605.10833v1 Announce Type: cross Abstract: Industrial anomaly detection is critical for manufacturing quality control, yet existing datasets mainly focus on static images or sparse views, which

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Muon-OGD: Muon-based Spectral Orthogonal Gradient Projection for LLM Continual Learning

DGX agent

arXiv:2605.08949v1 Announce Type: new Abstract: A central challenge in continual learning for large language models (LLMs) is catastrophic forgetting, where adapting to new tasks can substantially deg

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks

DGX agent

arXiv:2605.10639v1 Announce Type: new Abstract: The rapid adoption of LLMs in both research and industry highlights the challenges of deploying them safely and reveals a gap in the systematic evaluati

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Nested Slice Sampling: Vectorized Nested Sampling for GPU-Accelerated Inference

DGX agent

arXiv:2601.23252v2 Announce Type: replace-cross Abstract: Model comparison and calibrated uncertainty quantification often require integrating over parameters, but scalable inference can be challengin

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning

DGX agent

arXiv:2605.09727v1 Announce Type: cross Abstract: A central challenge in reinforcement learning (RL) is to learn models that generalize beyond the tasks on which they are trained, a goal traditionally

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Optimized Culprit Identification Using Mobilenet and Attention Mechanisms

DGX agent

arXiv:2605.08169v1 Announce Type: cross Abstract: Automated culprit identification in surveillance systems is a critical task that requires high accuracy along with computational efficiency for real-t

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Personalized Alignment Revisited: The Necessity and Sufficiency of User Diversity

DGX agent

arXiv:2605.09119v1 Announce Type: cross Abstract: Personalized alignment aims to adapt large language models to heterogeneous user preferences, yet the precise theoretical conditions for its statistic

model-releasesarxiv-cs-ai
12 May 2026
Safety

Personalizing LLMs with Binary Feedback: A Preference-Corrected Optimization Framework

DGX agent

arXiv:2605.10043v1 Announce Type: cross Abstract: Large Language Model (LLM) personalization aims to align model behaviors with individual user preferences. Existing methods often focus on isolated us

safetyarxiv-cs-ai
12 May 2026
← Previous
1…551552553554555…1394
Next →