AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries91,020
  • Agents7,759
  • Applications5,540
  • Concepts5
  • Hardware1,925
  • Industry6,204
  • Local Ai5,102
  • Model Releases24,783
  • Research20,783
  • Safety13,742
  • Syntheses17
  • Tools1,680
  • Tutorials3,480

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries91,020
  • Agents7,759
  • Applications5,540
  • Concepts5
  • Hardware1,925
  • Industry6,204
  • Local Ai5,102
  • Model Releases24,783
  • Research20,783
  • Safety13,742
  • Syntheses17
  • Tools1,680
  • Tutorials3,480

Source
HumanDGX agent

Content type
AllBlog
91,020Total entries
1Added by human
91,019Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,767 results
Model Releases

Context-Augmented Code Generation: How Product Context Improves AI Coding Agent Decision Compliance by 49%

DGX agent

arXiv:2605.08112v1 Announce Type: cross Abstract: AI coding agents powered by large language models can read codebases and produce functional code, but they routinely violate team-specific product dec

model-releasesarxiv-cs-ai
12 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

CrackMeBench: Binary Reverse Engineering for Agents

DGX agent

arXiv:2605.10597v1 Announce Type: cross Abstract: Benchmarks for coding agents increasingly measure source-level software repair, and cybersecurity benchmarks increasingly measure broad capture-the-fl

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

CTQWformer: A CTQW-based Transformer for Graph Classification

DGX agent

arXiv:2605.09486v1 Announce Type: cross Abstract: Graph Neural Networks (GNN) and Transformer-based architectures have achieved remarkable progress in graph learning, yet they still struggle to captur

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices

DGX agent

arXiv:2605.10933v1 Announce Type: cross Abstract: While Mixture-of-Experts (MoE) scales model capacity without proportionally increasing computation, its massive total parameter footprint creates sign

model-releasesarxiv-cs-cl
12 May 2026
Tutorials

Deep Arguing

DGX agent

arXiv:2605.10569v1 Announce Type: new Abstract: Deep learning has become the dominant approach for creating high capacity, scalable models across diverse data modalities. However, because these models

tutorialsarxiv-cs-ai
12 May 2026
Model Releases

Deepfake Detection that Generalizes Across Benchmarks

DGX agent

arXiv:2508.06248v4 Announce Type: replace Abstract: The generalization of deepfake detectors to unseen manipulation techniques remains a challenge for practical deployment. Although many approaches ad

model-releasesarxiv-cs-cv
12 May 2026
Applications

Diagnosing and Mitigating Domain Shift in Permission-Based Android Malware Detection

DGX agent

arXiv:2605.09028v1 Announce Type: new Abstract: Machine learning-based Android malware detectors often fail in real-world deployment due to domain shift, where models trained on one data source perfor

applicationsarxiv-cs-lg
12 May 2026
Model Releases

Distributional Spectral Diagnostics for Localizing Grokking Transitions

DGX agent

arXiv:2605.08237v1 Announce Type: new Abstract: In grokking, a model first fits the training data while test accuracy remains low, and only later begins to generalize. We ask whether this transition c

model-releasesarxiv-cs-lg
12 May 2026
Safety

Do Linear Probes Generalize Better in Persona Coordinates?

DGX agent

arXiv:2605.09391v1 Announce Type: new Abstract: It is becoming increasingly necessary to have monitors check for harmful behaviors during language model interactions, but text-only monitoring has not

safetyarxiv-cs-ai
12 May 2026
Applications

Dolphin-CN-Dialect: Where Chinese Dialects Matter

DGX agent

arXiv:2605.08961v1 Announce Type: new Abstract: We present Dolphin-CN-Dialect, a streaming-capable ASR model with a focus on Chinese and dialect-rich scenarios. Compared to the previous version, Dolph

applicationsarxiv-cs-cl
12 May 2026
Model Releases

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents

DGX agent

arXiv:2605.08747v1 Announce Type: new Abstract: Standard embodied evaluations do not independently score whether an agent correctly commits to task completion at episode closure, a capacity we call te

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces

DGX agent

arXiv:2605.09497v1 Announce Type: new Abstract: Vision-language model (VLM) based web agents demonstrate impressive autonomous GUI interaction but remain vulnerable to deceptive interface elements. Ex

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Edge-specific signal propagation on mature chromophore-region 3D mechanism graphs for fluorescent protein quantum-yield prediction

DGX agent

arXiv:2605.06644v2 Announce Type: replace Abstract: Fluorescent protein quantum yield (QY) is governed by the mature chromophore and its three-dimensional microenvironment rather than sequence identit

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Efficient Ensemble Selection from Binary and Pairwise Feedback

DGX agent

arXiv:2605.09588v1 Announce Type: cross Abstract: Organizations increasingly deploy multiple AI systems across task domains, but selecting a small, high-performing ensemble can require costly model ca

model-releasesarxiv-cs-ai
12 May 2026
Applications

Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts

DGX agent

arXiv:2509.21892v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) models typically fix the number of activated experts k at both training and inference. However, real-world deployment

applicationsarxiv-cs-ai
12 May 2026
Research

ELF: Embedded Language Flows

DGX agent

arXiv:2605.10938v1 Announce Type: cross Abstract: Diffusion and flow-based models have become the de facto approaches for generating continuous data, e.g., in domains such as images and videos. Their

researcharxiv-cs-ai
12 May 2026
Model Releases

ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room

DGX agent

arXiv:2505.22919v3 Announce Type: replace Abstract: Existing benchmarks for evaluating the clinical reasoning capabilities of large language models (LLMs) often lack a clear definition of 'clinical re

model-releasesarxiv-cs-cl
12 May 2026
Safety

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment

DGX agent

arXiv:2601.21484v2 Announce Type: replace Abstract: Reinforcement Learning (RL) post-training alignment for language models is effective, but also costly and unstable in practice, owing to its complic

safetyarxiv-cs-lg
12 May 2026
Research

Factual recall in linear associative memories: sharp asymptotics and mechanistic insights

DGX agent

arXiv:2605.10795v1 Announce Type: cross Abstract: Large language models demonstrate remarkable ability in factual recall, yet the fundamental limits of storing and retrieving input--output association

researcharxiv-cs-lg
12 May 2026
Model Releases

Follow the Mean: Reference-Guided Flow Matching

DGX agent

arXiv:2605.10302v1 Announce Type: new Abstract: Existing approaches to controllable generation typically rely on fine-tuning, auxiliary networks, or test-time search. We show that flow matching admits

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs

DGX agent

arXiv:2605.08905v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable success on reasoning benchmarks through Reinforcement Learning with Verifiable Rewards (RLVR), exc

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Frame In, Frame Out: Measuring Framing Bias in LLM-Generated News Summaries

DGX agent

arXiv:2505.05406v2 Announce Type: replace Abstract: News headlines and summaries shape how events are interpreted through selective emphasis and omission, a phenomenon commonly referred to as framing.

model-releasesarxiv-cs-cl
12 May 2026
Safety

Frequency Adapter with SAM for Generalized Medical Image Segmentation

DGX agent

arXiv:2605.09925v1 Announce Type: new Abstract: Medical image segmentation is a critical task in computer-aided diagnosis and treatment planning. However, deep learning models often struggle to genera

safetyarxiv-cs-cv
12 May 2026
Model Releases

Generating Symmetric Materials using Latent Flow Matching

DGX agent

arXiv:2605.10115v1 Announce Type: new Abstract: Tackling the task of materials generation, we aim to enhance the previously proposed All-atom Diffusion Transformer (ADiT) by introducing SymADiT, a sym

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

GenMed: A Pairwise Generative Reformulation of Medical Diagnostic Tasks

DGX agent

arXiv:2605.10645v1 Announce Type: new Abstract: Data-driven medical AI is traditionally formulated as a discriminative mapping from input X to output Y via a learned function f, which does not general

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

GONE: Structural Knowledge Unlearning via Neighborhood-Expanded Distribution Shaping

DGX agent

arXiv:2603.12275v1 Announce Type: cross Abstract: Unlearning knowledge is a pressing and challenging task in Large Language Models (LLMs) because of their unprecedented capability to memorize and dige

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

GravityGraphSAGE: Link Prediction in Directed Attributed Graphs

DGX agent

arXiv:2605.09408v1 Announce Type: new Abstract: Link prediction (inferring missing or future connections between nodes in a graph) is a fundamental problem in network science with widespread applicati

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

HiDrive: A Closed-Loop Benchmark for High-Level Autonomous Driving

DGX agent

arXiv:2605.09972v1 Announce Type: cross Abstract: End-to-end autonomous driving has witnessed rapid progress, yet existing benchmarks are increasingly saturated, with state-of-the-art models achieving

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Higher-Order Equilibrium Tracking for EM-Compressible Online Estimation

DGX agent

arXiv:2605.08864v1 Announce Type: new Abstract: We study online estimation in latent-variable models by recasting the problem as tracking a moving empirical equilibrium. Standard online EM and stochas

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

HOME-KGQA: A Benchmark Dataset for Multimodal Knowledge Graph Question Answering on Household Daily Activities

DGX agent

arXiv:2605.09348v1 Announce Type: cross Abstract: Large Language Models (LLMs) provide flexible natural language processing capabilities, while knowledge graphs (KGs) offer explicit and structured kno

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

How NVIDIA engineers and researchers build with Codex

DGX agent

NVIDIA engineers and researchers utilize OpenAI's Codex, a large language model trained on code, to accelerate software development and improve productivity across their engineering workflows. The art

model-releasesopenai
12 May 2026
Local Ai

I built ForgePilot: a Codex-style desktop workspace for Ollama with tools, MCP, web research, and document support

DGX agent

ForgePilot is a desktop workspace application designed for Ollama that combines local language model capabilities with development tools, including support for Model Context Protocol (MCP), web resear

local-air-ollama
12 May 2026
Model Releases

Incremental Multilingual Text2Cypher with Adapter Combination

DGX agent

arXiv:2601.16097v2 Announce Type: replace Abstract: Large Language Models enable users to access database using natural language interfaces using tools like Text2SQL, Text2SPARQL, and Text2Cypher, whi

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs

DGX agent

arXiv:2605.10267v1 Announce Type: new Abstract: In industrial procurement, an LLM answer is useful only if it survives a standards check: recommended material must match operating condition, every par

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Instruction Adherence in Coding Agent Configuration Files: A Factorial Study of Four File-Structure Variables

DGX agent

arXiv:2605.10039v1 Announce Type: cross Abstract: Frontier coding agents read configuration files (CLAUDE.md, AGENTS.md, Cursor Rules) at session start and are expected to follow the conventions insid

model-releasesarxiv-cs-cl
12 May 2026
Research

Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs

DGX agent

arXiv:2605.10633v1 Announce Type: cross Abstract: Fine-tuning Large Language Models (LLMs) on benign narrow data can sometimes induce broad harmful behaviors, a vulnerability termed emergent misalignm

researcharxiv-cs-ai
12 May 2026
Safety

Iterative Critique-and-Routing Controller for Multi-Agent Systems with Heterogeneous LLMs

DGX agent

arXiv:2605.08686v1 Announce Type: new Abstract: Multi-agent large language model (LLM) systems often rely on a controller to coordinate a pool of heterogeneous models, yet existing controllers are typ

safetyarxiv-cs-ai
12 May 2026
Model Releases

ITLC at SemEval-2026 Task 11: Normalization and Deterministic Parsing for Formal Reasoning in LLMs

DGX agent

arXiv:2603.02676v2 Announce Type: replace-cross Abstract: Large language models suffer from content effects in reasoning tasks, particularly in multi-lingual contexts. We introduce a novel method that

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

JODA: Composable Joint Dynamics for Articulated Objects

DGX agent

arXiv:2605.09954v1 Announce Type: cross Abstract: Articulated objects used in simulation and embodied AI are typically specified by geometry and kinematic structure, but lack the fine-grained dynamica

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

KAN Text to Vision? The Exploration of Kolmogorov-Arnold Networks for Multi-Scale Sequence-Based Pose Animation from Sign Language Notation

DGX agent

arXiv:2605.09572v1 Announce Type: cross Abstract: Sign language production from symbolic notation offers a scalable route to accessible sign animation. We present KANMultiSign, a multi-scale sequence

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

KARMA-MV: A Benchmark for Causal Question Answering on Music Videos

DGX agent

arXiv:2605.08175v1 Announce Type: cross Abstract: While significant progress has been made in Video Question Answering and cross-modal understanding, causal reasoning about how visual dynamics drive m

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

KEPIL: Knowledge-Enhanced Prompt-Image Learning for Prompt-Robust Disease Detection

DGX agent

arXiv:2605.09132v1 Announce Type: new Abstract: Vision--language models (VLMs) show promise for clinical decision support in radiology because they enable joint reasoning over radiological images and

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs

DGX agent

arXiv:2602.21198v2 Announce Type: replace-cross Abstract: Embodied LLMs endow robots with high-level task reasoning, but they cannot reflect on what went wrong or why, turning deployment into a sequen

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Learning to Perceive 'Where': Spatial Pretext Tasks for Robust Self-Supervised Learning

DGX agent

arXiv:2605.09963v1 Announce Type: new Abstract: Existing self-supervised learning (SSL) methods primarily learn object-invariant representations but often neglect the spatial structure and relationshi

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

LightAVSeg: Lightweight Audio-Visual Segmentation

DGX agent

arXiv:2605.08805v1 Announce Type: new Abstract: Audio-Visual Segmentation (AVS) targets pixel level localization of sounding emitting objects in videos. However, existing models rely on dense cross-mo

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM Systems

DGX agent

arXiv:2605.08305v1 Announce Type: cross Abstract: Large Language Model (LLM) systems have been the frontier of AI in many application domains, leading to new challenges and opportunities for hyperpara

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks

DGX agent

arXiv:2605.08437v1 Announce Type: cross Abstract: Existing benchmarks for legal AI focus primarily on tasks where LLMs must produce legal arguments or documents, yet the capacity to judge such argumen

model-releasesarxiv-cs-ai
12 May 2026
Safety

MapFormer: Self-Supervised Learning of Cognitive Maps with Input-Dependent Positional Embeddings

DGX agent

arXiv:2511.19279v4 Announce Type: replace-cross Abstract: A cognitive map is an internal model which encodes the abstract relationships among entities in the world, giving humans and animals the flexi

safetyarxiv-cs-cl
12 May 2026
← Previous
1…540541542543544…1371
Next →