AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,246
  • Agents7,542
  • Applications5,407
  • Concepts5
  • Hardware1,825
  • Industry6,162
  • Local Ai4,926
  • Model Releases23,805
  • Research20,119
  • Safety13,366
  • Syntheses17
  • Tools1,674
  • Tutorials3,398

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,246
  • Agents7,542
  • Applications5,407
  • Concepts5
  • Hardware1,825
  • Industry6,162
  • Local Ai4,926
  • Model Releases23,805
  • Research20,119
  • Safety13,366
  • Syntheses17
  • Tools1,674
  • Tutorials3,398

Source
HumanDGX agent

88,246Total entries
1Added by human
88,245Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,499 results
12 May 2026

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment

SafetyDGX agent

arXiv:2601.21484v2 Announce Type: replace Abstract: Reinforcement Learning (RL) post-training alignment for language models is effective, but also costly and unstable in practice, owing to its complic

Factual recall in linear associative memories: sharp asymptotics and mechanistic insights

ResearchDGX agent

arXiv:2605.10795v1 Announce Type: cross Abstract: Large language models demonstrate remarkable ability in factual recall, yet the fundamental limits of storing and retrieving input--output association

Follow the Mean: Reference-Guided Flow Matching

Model ReleasesDGX agent

arXiv:2605.10302v1 Announce Type: new Abstract: Existing approaches to controllable generation typically rely on fine-tuning, auxiliary networks, or test-time search. We show that flow matching admits

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs

Model ReleasesDGX agent

arXiv:2605.08905v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable success on reasoning benchmarks through Reinforcement Learning with Verifiable Rewards (RLVR), exc

Frame In, Frame Out: Measuring Framing Bias in LLM-Generated News Summaries

Model ReleasesDGX agent

arXiv:2505.05406v2 Announce Type: replace Abstract: News headlines and summaries shape how events are interpreted through selective emphasis and omission, a phenomenon commonly referred to as framing.

Frequency Adapter with SAM for Generalized Medical Image Segmentation

SafetyDGX agent

arXiv:2605.09925v1 Announce Type: new Abstract: Medical image segmentation is a critical task in computer-aided diagnosis and treatment planning. However, deep learning models often struggle to genera

Generating Symmetric Materials using Latent Flow Matching

Model ReleasesDGX agent

arXiv:2605.10115v1 Announce Type: new Abstract: Tackling the task of materials generation, we aim to enhance the previously proposed All-atom Diffusion Transformer (ADiT) by introducing SymADiT, a sym

GenMed: A Pairwise Generative Reformulation of Medical Diagnostic Tasks

Model ReleasesDGX agent

arXiv:2605.10645v1 Announce Type: new Abstract: Data-driven medical AI is traditionally formulated as a discriminative mapping from input X to output Y via a learned function f, which does not general

GONE: Structural Knowledge Unlearning via Neighborhood-Expanded Distribution Shaping

Model ReleasesDGX agent

arXiv:2603.12275v1 Announce Type: cross Abstract: Unlearning knowledge is a pressing and challenging task in Large Language Models (LLMs) because of their unprecedented capability to memorize and dige

GravityGraphSAGE: Link Prediction in Directed Attributed Graphs

Model ReleasesDGX agent

arXiv:2605.09408v1 Announce Type: new Abstract: Link prediction (inferring missing or future connections between nodes in a graph) is a fundamental problem in network science with widespread applicati

HiDrive: A Closed-Loop Benchmark for High-Level Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.09972v1 Announce Type: cross Abstract: End-to-end autonomous driving has witnessed rapid progress, yet existing benchmarks are increasingly saturated, with state-of-the-art models achieving

Higher-Order Equilibrium Tracking for EM-Compressible Online Estimation

Model ReleasesDGX agent

arXiv:2605.08864v1 Announce Type: new Abstract: We study online estimation in latent-variable models by recasting the problem as tracking a moving empirical equilibrium. Standard online EM and stochas

HOME-KGQA: A Benchmark Dataset for Multimodal Knowledge Graph Question Answering on Household Daily Activities

Model ReleasesDGX agent

arXiv:2605.09348v1 Announce Type: cross Abstract: Large Language Models (LLMs) provide flexible natural language processing capabilities, while knowledge graphs (KGs) offer explicit and structured kno

How NVIDIA engineers and researchers build with Codex

Model ReleasesDGX agent

NVIDIA engineers and researchers utilize OpenAI's Codex, a large language model trained on code, to accelerate software development and improve productivity across their engineering workflows. The art

I built ForgePilot: a Codex-style desktop workspace for Ollama with tools, MCP, web research, and document support

Local AiDGX agent

ForgePilot is a desktop workspace application designed for Ollama that combines local language model capabilities with development tools, including support for Model Context Protocol (MCP), web resear

Incremental Multilingual Text2Cypher with Adapter Combination

Model ReleasesDGX agent

arXiv:2601.16097v2 Announce Type: replace Abstract: Large Language Models enable users to access database using natural language interfaces using tools like Text2SQL, Text2SPARQL, and Text2Cypher, whi

IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs

Model ReleasesDGX agent

arXiv:2605.10267v1 Announce Type: new Abstract: In industrial procurement, an LLM answer is useful only if it survives a standards check: recommended material must match operating condition, every par

Instruction Adherence in Coding Agent Configuration Files: A Factorial Study of Four File-Structure Variables

Model ReleasesDGX agent

arXiv:2605.10039v1 Announce Type: cross Abstract: Frontier coding agents read configuration files (CLAUDE.md, AGENTS.md, Cursor Rules) at session start and are expected to follow the conventions insid

Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs

ResearchDGX agent

arXiv:2605.10633v1 Announce Type: cross Abstract: Fine-tuning Large Language Models (LLMs) on benign narrow data can sometimes induce broad harmful behaviors, a vulnerability termed emergent misalignm

Iterative Critique-and-Routing Controller for Multi-Agent Systems with Heterogeneous LLMs

SafetyDGX agent

arXiv:2605.08686v1 Announce Type: new Abstract: Multi-agent large language model (LLM) systems often rely on a controller to coordinate a pool of heterogeneous models, yet existing controllers are typ

ITLC at SemEval-2026 Task 11: Normalization and Deterministic Parsing for Formal Reasoning in LLMs

Model ReleasesDGX agent

arXiv:2603.02676v2 Announce Type: replace-cross Abstract: Large language models suffer from content effects in reasoning tasks, particularly in multi-lingual contexts. We introduce a novel method that

JODA: Composable Joint Dynamics for Articulated Objects

Model ReleasesDGX agent

arXiv:2605.09954v1 Announce Type: cross Abstract: Articulated objects used in simulation and embodied AI are typically specified by geometry and kinematic structure, but lack the fine-grained dynamica

KAN Text to Vision? The Exploration of Kolmogorov-Arnold Networks for Multi-Scale Sequence-Based Pose Animation from Sign Language Notation

Model ReleasesDGX agent

arXiv:2605.09572v1 Announce Type: cross Abstract: Sign language production from symbolic notation offers a scalable route to accessible sign animation. We present KANMultiSign, a multi-scale sequence

KARMA-MV: A Benchmark for Causal Question Answering on Music Videos

Model ReleasesDGX agent

arXiv:2605.08175v1 Announce Type: cross Abstract: While significant progress has been made in Video Question Answering and cross-modal understanding, causal reasoning about how visual dynamics drive m

KEPIL: Knowledge-Enhanced Prompt-Image Learning for Prompt-Robust Disease Detection

Model ReleasesDGX agent

arXiv:2605.09132v1 Announce Type: new Abstract: Vision--language models (VLMs) show promise for clinical decision support in radiology because they enable joint reasoning over radiological images and

Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs

Model ReleasesDGX agent

arXiv:2602.21198v2 Announce Type: replace-cross Abstract: Embodied LLMs endow robots with high-level task reasoning, but they cannot reflect on what went wrong or why, turning deployment into a sequen

Learning to Perceive 'Where': Spatial Pretext Tasks for Robust Self-Supervised Learning

Model ReleasesDGX agent

arXiv:2605.09963v1 Announce Type: new Abstract: Existing self-supervised learning (SSL) methods primarily learn object-invariant representations but often neglect the spatial structure and relationshi

LightAVSeg: Lightweight Audio-Visual Segmentation

Model ReleasesDGX agent

arXiv:2605.08805v1 Announce Type: new Abstract: Audio-Visual Segmentation (AVS) targets pixel level localization of sounding emitting objects in videos. However, existing models rely on dense cross-mo

LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM Systems

Model ReleasesDGX agent

arXiv:2605.08305v1 Announce Type: cross Abstract: Large Language Model (LLM) systems have been the frontier of AI in many application domains, leading to new challenges and opportunities for hyperpara

Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks

Model ReleasesDGX agent

arXiv:2605.08437v1 Announce Type: cross Abstract: Existing benchmarks for legal AI focus primarily on tasks where LLMs must produce legal arguments or documents, yet the capacity to judge such argumen

MapFormer: Self-Supervised Learning of Cognitive Maps with Input-Dependent Positional Embeddings

SafetyDGX agent

arXiv:2511.19279v4 Announce Type: replace-cross Abstract: A cognitive map is an internal model which encodes the abstract relationships among entities in the world, giving humans and animals the flexi

MathConstraint: Automated Generation of Verified Combinatorial Reasoning Instances for LLMs

Model ReleasesDGX agent

arXiv:2605.08498v1 Announce Type: cross Abstract: We introduce MathConstraint, a hard, adaptive benchmark for evaluating the combinatorial reasoning capabilities of LLMs. We combine constraint satisfa

MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics

Model ReleasesDGX agent

arXiv:2602.02561v2 Announce Type: replace-cross Abstract: While the ecosystem of Lean and Mathlib has enjoyed celebrated success in formal mathematical reasoning with the help of large language models

Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests

Model ReleasesDGX agent

arXiv:2510.22170v2 Announce Type: replace Abstract: Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure o

Medical Incident Causal Factors and Preventive Measures Generation Using Tag-based Example Selection in Few-shot Learning

Model ReleasesDGX agent

arXiv:2605.10025v1 Announce Type: cross Abstract: In high-stakes domains such as healthcare, the reliability of Large Language Models (LLMs) is critical, particularly when generating clinical insights

Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All?

Model ReleasesDGX agent

arXiv:2605.08816v1 Announce Type: new Abstract: In the animal kingdom, mirror self-recognition is a canonical probe of higher-order cognition, emerging only in some species. We ask whether an analogou

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI

Model ReleasesDGX agent

arXiv:2605.08678v1 Announce Type: new Abstract: Modern AI progress has been driven by ML methods that are generalizable across settings and scalable to larger regimes. As large language models demonst

MMVIAD: Multi-view Multi-task Video Understanding for Industrial Anomaly Detection

Model ReleasesDGX agent

arXiv:2605.10833v1 Announce Type: cross Abstract: Industrial anomaly detection is critical for manufacturing quality control, yet existing datasets mainly focus on static images or sparse views, which

Muon-OGD: Muon-based Spectral Orthogonal Gradient Projection for LLM Continual Learning

Model ReleasesDGX agent

arXiv:2605.08949v1 Announce Type: new Abstract: A central challenge in continual learning for large language models (LLMs) is catastrophic forgetting, where adapting to new tasks can substantially deg

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks

Model ReleasesDGX agent

arXiv:2605.10639v1 Announce Type: new Abstract: The rapid adoption of LLMs in both research and industry highlights the challenges of deploying them safely and reveals a gap in the systematic evaluati

Nested Slice Sampling: Vectorized Nested Sampling for GPU-Accelerated Inference

Model ReleasesDGX agent

arXiv:2601.23252v2 Announce Type: replace-cross Abstract: Model comparison and calibrated uncertainty quantification often require integrating over parameters, but scalable inference can be challengin

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.09727v1 Announce Type: cross Abstract: A central challenge in reinforcement learning (RL) is to learn models that generalize beyond the tasks on which they are trained, a goal traditionally

Optimized Culprit Identification Using Mobilenet and Attention Mechanisms

Model ReleasesDGX agent

arXiv:2605.08169v1 Announce Type: cross Abstract: Automated culprit identification in surveillance systems is a critical task that requires high accuracy along with computational efficiency for real-t

Personalized Alignment Revisited: The Necessity and Sufficiency of User Diversity

Model ReleasesDGX agent

arXiv:2605.09119v1 Announce Type: cross Abstract: Personalized alignment aims to adapt large language models to heterogeneous user preferences, yet the precise theoretical conditions for its statistic

Personalizing LLMs with Binary Feedback: A Preference-Corrected Optimization Framework

SafetyDGX agent

arXiv:2605.10043v1 Announce Type: cross Abstract: Large Language Model (LLM) personalization aims to align model behaviors with individual user preferences. Existing methods often focus on isolated us

PiCA: Pivot-Based Credit Assignment for Search Agentic Reinforcement Learning

AgentsDGX agent

arXiv:2605.09287v1 Announce Type: new Abstract: Large Language Model (LLM)-based search agents trained with reinforcement learning (RL) have significantly improved the performance of knowledge-intensi

Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead

Model ReleasesDGX agent

arXiv:2507.23009v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have achieved remarkable results on a range of standardized tests originally designed to assess human cognitive a

Prediction Bottlenecks Don't Discover Causal Structure (But Here's What They Actually Do)

Model ReleasesDGX agent

arXiv:2605.09169v1 Announce Type: cross Abstract: A Mamba state-space model trained only for next-step prediction appears to recover Granger-causal structure through a simple readout S = |W_{out} W_{i

PRIM: Meta-Learned Bayesian Root Cause Analysis

Model ReleasesDGX agent

arXiv:2605.08786v1 Announce Type: new Abstract: Root cause analysis (RCA) in complex systems is challenging due to error propagation across multiple variables, the need for structural causal knowledge

Quantifying Concentration Phenomena of Mean-Field Transformers in the Low-Temperature Regime

Model ReleasesDGX agent

arXiv:2605.10931v1 Announce Type: cross Abstract: Transformers with self-attention modules as their core components have become an integral architecture in modern large language and foundation models.

Re^2Math: Benchmarking Theorem Retrieval in Research-Level Mathematics

Model ReleasesDGX agent

arXiv:2605.09012v1 Announce Type: new Abstract: Large language models are increasingly capable at closed-world mathematical reasoning, but research assistance also requires source-grounded use of the

Rennala MVR: Improved Time Complexity for Parallel Stochastic Optimization via Momentum-Based Variance Reduction

Model ReleasesDGX agent

arXiv:2605.08871v1 Announce Type: cross Abstract: Large-scale machine learning models are trained on clusters of machines that exhibit heterogeneous performance due to hardware variability, network de

Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?

Model ReleasesDGX agent

arXiv:2605.10848v1 Announce Type: cross Abstract: Does a lexical retriever suffice as large language models (LLMs) become more capable in an agentic loop? This question naturally arises when building

Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs

Model ReleasesDGX agent

arXiv:2605.10094v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models show strong potential for general-purpose robotic manipulation, yet their closed-loop reliability often degrades u

RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement

Model ReleasesDGX agent

arXiv:2605.09730v1 Announce Type: new Abstract: Iterative self-refinement is a popular inference-time reliability technique, but its effectiveness in code-mode tool use depends heavily on the structur

Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMs

Model ReleasesDGX agent

arXiv:2509.02372v3 Announce Type: replace-cross Abstract: Large Language Models have become critical to modern software development, but their reliance on uncurated web-scale datasets for training int

Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories

SafetyDGX agent

arXiv:2605.08936v1 Announce Type: new Abstract: Large Reasoning Models possess remarkable capabilities for self-correction in general domain; however, they frequently struggle to recover from unsafe r

Separate First, Fuse Later: Mitigating Cross-Modal Interference in Audio-Visual LLMs Reasoning with Modality-Specific Chain-of-Thought

Model ReleasesDGX agent

arXiv:2605.09906v1 Announce Type: new Abstract: Audio and vision provide complementary evidence for audio-visual question answering, yet current audio-visual large language models may suffer from cros

Simultaneous Long-tailed Recognition and Multi-modal Fusion for Highly Imbalanced Multi-modal Data

Model ReleasesDGX agent

arXiv:2605.10498v1 Announce Type: cross Abstract: Long-tailed distributions in class-imbalanced data present a fundamental challenge for deep learning models, which tend to be biased toward majority c

Skill-R1: Agent Skill Evolution via Reinforcement Learning

SafetyDGX agent

arXiv:2605.09359v1 Announce Type: cross Abstract: Agentic large language models often rely on skills, reusable natural language procedures that guide planning, action, and tool use. In practice, skill

← Previous
1…416417418419420…1059
Next →