AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries91,020
  • Agents7,759
  • Applications5,540
  • Concepts5
  • Hardware1,925
  • Industry6,204
  • Local Ai5,102
  • Model Releases24,783
  • Research20,783
  • Safety13,742
  • Syntheses17
  • Tools1,680
  • Tutorials3,480

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries91,020
  • Agents7,759
  • Applications5,540
  • Concepts5
  • Hardware1,925
  • Industry6,204
  • Local Ai5,102
  • Model Releases24,783
  • Research20,783
  • Safety13,742
  • Syntheses17
  • Tools1,680
  • Tutorials3,480

Source
HumanDGX agent

Content type
AllBlog
91,020Total entries
1Added by human
91,019Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,767 results
Model Releases

MathConstraint: Automated Generation of Verified Combinatorial Reasoning Instances for LLMs

DGX agent

arXiv:2605.08498v1 Announce Type: cross Abstract: We introduce MathConstraint, a hard, adaptive benchmark for evaluating the combinatorial reasoning capabilities of LLMs. We combine constraint satisfa

model-releasesarxiv-cs-ai
12 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics

DGX agent

arXiv:2602.02561v2 Announce Type: replace-cross Abstract: While the ecosystem of Lean and Mathlib has enjoyed celebrated success in formal mathematical reasoning with the help of large language models

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests

DGX agent

arXiv:2510.22170v2 Announce Type: replace Abstract: Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure o

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Medical Incident Causal Factors and Preventive Measures Generation Using Tag-based Example Selection in Few-shot Learning

DGX agent

arXiv:2605.10025v1 Announce Type: cross Abstract: In high-stakes domains such as healthcare, the reliability of Large Language Models (LLMs) is critical, particularly when generating clinical insights

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All?

DGX agent

arXiv:2605.08816v1 Announce Type: new Abstract: In the animal kingdom, mirror self-recognition is a canonical probe of higher-order cognition, emerging only in some species. We ask whether an analogou

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI

DGX agent

arXiv:2605.08678v1 Announce Type: new Abstract: Modern AI progress has been driven by ML methods that are generalizable across settings and scalable to larger regimes. As large language models demonst

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

MMVIAD: Multi-view Multi-task Video Understanding for Industrial Anomaly Detection

DGX agent

arXiv:2605.10833v1 Announce Type: cross Abstract: Industrial anomaly detection is critical for manufacturing quality control, yet existing datasets mainly focus on static images or sparse views, which

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Muon-OGD: Muon-based Spectral Orthogonal Gradient Projection for LLM Continual Learning

DGX agent

arXiv:2605.08949v1 Announce Type: new Abstract: A central challenge in continual learning for large language models (LLMs) is catastrophic forgetting, where adapting to new tasks can substantially deg

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks

DGX agent

arXiv:2605.10639v1 Announce Type: new Abstract: The rapid adoption of LLMs in both research and industry highlights the challenges of deploying them safely and reveals a gap in the systematic evaluati

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Nested Slice Sampling: Vectorized Nested Sampling for GPU-Accelerated Inference

DGX agent

arXiv:2601.23252v2 Announce Type: replace-cross Abstract: Model comparison and calibrated uncertainty quantification often require integrating over parameters, but scalable inference can be challengin

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning

DGX agent

arXiv:2605.09727v1 Announce Type: cross Abstract: A central challenge in reinforcement learning (RL) is to learn models that generalize beyond the tasks on which they are trained, a goal traditionally

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Optimized Culprit Identification Using Mobilenet and Attention Mechanisms

DGX agent

arXiv:2605.08169v1 Announce Type: cross Abstract: Automated culprit identification in surveillance systems is a critical task that requires high accuracy along with computational efficiency for real-t

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Personalized Alignment Revisited: The Necessity and Sufficiency of User Diversity

DGX agent

arXiv:2605.09119v1 Announce Type: cross Abstract: Personalized alignment aims to adapt large language models to heterogeneous user preferences, yet the precise theoretical conditions for its statistic

model-releasesarxiv-cs-ai
12 May 2026
Safety

Personalizing LLMs with Binary Feedback: A Preference-Corrected Optimization Framework

DGX agent

arXiv:2605.10043v1 Announce Type: cross Abstract: Large Language Model (LLM) personalization aims to align model behaviors with individual user preferences. Existing methods often focus on isolated us

safetyarxiv-cs-ai
12 May 2026
Agents

PiCA: Pivot-Based Credit Assignment for Search Agentic Reinforcement Learning

DGX agent

arXiv:2605.09287v1 Announce Type: new Abstract: Large Language Model (LLM)-based search agents trained with reinforcement learning (RL) have significantly improved the performance of knowledge-intensi

agentsarxiv-cs-ai
12 May 2026
Model Releases

Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead

DGX agent

arXiv:2507.23009v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have achieved remarkable results on a range of standardized tests originally designed to assess human cognitive a

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Prediction Bottlenecks Don't Discover Causal Structure (But Here's What They Actually Do)

DGX agent

arXiv:2605.09169v1 Announce Type: cross Abstract: A Mamba state-space model trained only for next-step prediction appears to recover Granger-causal structure through a simple readout S = |W_{out} W_{i

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

PRIM: Meta-Learned Bayesian Root Cause Analysis

DGX agent

arXiv:2605.08786v1 Announce Type: new Abstract: Root cause analysis (RCA) in complex systems is challenging due to error propagation across multiple variables, the need for structural causal knowledge

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Quantifying Concentration Phenomena of Mean-Field Transformers in the Low-Temperature Regime

DGX agent

arXiv:2605.10931v1 Announce Type: cross Abstract: Transformers with self-attention modules as their core components have become an integral architecture in modern large language and foundation models.

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Re^2Math: Benchmarking Theorem Retrieval in Research-Level Mathematics

DGX agent

arXiv:2605.09012v1 Announce Type: new Abstract: Large language models are increasingly capable at closed-world mathematical reasoning, but research assistance also requires source-grounded use of the

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Rennala MVR: Improved Time Complexity for Parallel Stochastic Optimization via Momentum-Based Variance Reduction

DGX agent

arXiv:2605.08871v1 Announce Type: cross Abstract: Large-scale machine learning models are trained on clusters of machines that exhibit heterogeneous performance due to hardware variability, network de

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?

DGX agent

arXiv:2605.10848v1 Announce Type: cross Abstract: Does a lexical retriever suffice as large language models (LLMs) become more capable in an agentic loop? This question naturally arises when building

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs

DGX agent

arXiv:2605.10094v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models show strong potential for general-purpose robotic manipulation, yet their closed-loop reliability often degrades u

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement

DGX agent

arXiv:2605.09730v1 Announce Type: new Abstract: Iterative self-refinement is a popular inference-time reliability technique, but its effectiveness in code-mode tool use depends heavily on the structur

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMs

DGX agent

arXiv:2509.02372v3 Announce Type: replace-cross Abstract: Large Language Models have become critical to modern software development, but their reliance on uncurated web-scale datasets for training int

model-releasesarxiv-cs-ai
12 May 2026
Safety

Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories

DGX agent

arXiv:2605.08936v1 Announce Type: new Abstract: Large Reasoning Models possess remarkable capabilities for self-correction in general domain; however, they frequently struggle to recover from unsafe r

safetyarxiv-cs-ai
12 May 2026
Model Releases

Separate First, Fuse Later: Mitigating Cross-Modal Interference in Audio-Visual LLMs Reasoning with Modality-Specific Chain-of-Thought

DGX agent

arXiv:2605.09906v1 Announce Type: new Abstract: Audio and vision provide complementary evidence for audio-visual question answering, yet current audio-visual large language models may suffer from cros

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Simultaneous Long-tailed Recognition and Multi-modal Fusion for Highly Imbalanced Multi-modal Data

DGX agent

arXiv:2605.10498v1 Announce Type: cross Abstract: Long-tailed distributions in class-imbalanced data present a fundamental challenge for deep learning models, which tend to be biased toward majority c

model-releasesarxiv-cs-ai
12 May 2026
Safety

Skill-R1: Agent Skill Evolution via Reinforcement Learning

DGX agent

arXiv:2605.09359v1 Announce Type: cross Abstract: Agentic large language models often rely on skills, reusable natural language procedures that guide planning, action, and tool use. In practice, skill

safetyarxiv-cs-ai
12 May 2026
Model Releases

SmartEval: A Benchmark for Evaluating LLM-Generated Smart Contracts from Natural Language Specifications

DGX agent

arXiv:2605.09610v1 Announce Type: cross Abstract: We introduce SmartEval, a benchmark for systematically evaluating the quality of Solidity smart contracts generated by large language models (LLMs) fr

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Source or It Didn't Happen: A Multi-Agent Framework for Citation Hallucination Detection

DGX agent

arXiv:2605.08583v1 Announce Type: new Abstract: Large language models are increasingly used in scientific writing, yet they can fabricate citation-shaped references that appear plausible but fail bibl

model-releasesarxiv-cs-cl
12 May 2026
Research

Supersampling Stable Diffusion and More: An Approach for Interpolating Neural Networks Using Common Interpolation Methods

DGX agent

arXiv:2605.08698v1 Announce Type: new Abstract: Stable Diffusion (SD) has evolved DDPM (Denoising Diffusion Probabilistic Model) based image generation significantly by denoising in latent space inste

researcharxiv-cs-cv
12 May 2026
Research

Task-Aware Calibration: Provably Optimal Decoding in LLMs

DGX agent

arXiv:2605.10202v1 Announce Type: cross Abstract: LLM decoding often relies on the model's predictive distribution to generate an output. Consequently, misalignment with respect to the true generating

researcharxiv-cs-cl
12 May 2026
Tutorials

The finite expression method for turbulent dynamics with high-order moment recovery

DGX agent

arXiv:2605.10687v1 Announce Type: new Abstract: Turbulent dynamical systems are characterized by nonlinear interactions and stochastic effects that generate coupled statistical quantities, such as non

tutorialsarxiv-cs-lg
12 May 2026
Safety

The Value of Mechanistic Priors in Sequential Decision Making

DGX agent

arXiv:2605.10018v1 Announce Type: new Abstract: Hybrid mechanistic models, physical priors with learned residuals, promise to reduce the data required for good decisions, but have no computable criter

safetyarxiv-cs-lg
12 May 2026
Model Releases

TINS: Test-time ID-prototype-separated Negative Semantics Learning for OOD Detection

DGX agent

arXiv:2605.10756v1 Announce Type: new Abstract: Vision-language models enable OOD detection by comparing image alignment with ID labels and negative semantics. Existing negative-label-based methods ma

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

TRACER: Verifiable Generative Provenance for Multimodal Tool-Using Agents

DGX agent

arXiv:2605.09934v1 Announce Type: new Abstract: Multimodal large language models increasingly solve vision-centric tasks by calling external tools for visual inspection, OCR, retrieval, calculation, a

model-releasesarxiv-cs-cl
12 May 2026
Safety

Training with Harnesses: On-Policy Harness Self-Distillation for Complex Reasoning

DGX agent

arXiv:2605.08741v1 Announce Type: new Abstract: Inference-time harnesses substantially improve large language models on complex reasoning tasks. However, the intrinsic capabilities of the underlying m

safetyarxiv-cs-cl
12 May 2026
Model Releases

TrajPrism: A Multi-Task Benchmark for Language-Grounded Urban Trajectory Understanding

DGX agent

arXiv:2605.10782v1 Announce Type: new Abstract: Urban mobility is naturally expressed both as trajectories in space and as natural-language descriptions of travel intent, constraints, and preferences.

model-releasesarxiv-cs-ai
12 May 2026
Applications

TSNN: A Non-parametric and Interpretable Framework for Traffic Time Series Forecasting

DGX agent

arXiv:2605.09208v1 Announce Type: new Abstract: Although many complex models were proposed to analyze time series data, some studies have demonstrated remarkable performance with simpler structures. A

applicationsarxiv-cs-lg
12 May 2026
Model Releases

Uni-Synergy: Bridging Understanding and Generation for Personalized Reasoning via Co-operative Reinforcement Learning

DGX agent

arXiv:2605.10445v1 Announce Type: new Abstract: Unified Multimodal Models (UMMs) excel in general tasks but struggle to bridge the gap between personalized understanding and generation. Prior works la

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Unifying Scientific Communication: Fine-Grained Correspondence Across Scientific Media

DGX agent

arXiv:2605.05831v2 Announce Type: replace Abstract: The communication of scientific knowledge has become increasingly multimodal, spanning text, visuals, and speech through materials such as research

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Urban-ImageNet: A Large-Scale Multi-Modal Dataset and Evaluation Framework for Urban Space Perception

DGX agent

arXiv:2605.09936v1 Announce Type: new Abstract: We present Urban-ImageNet, a large-scale multi-modal dataset and evaluation benchmark for urban space perception from user-generated social media imager

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

WATCH: Wide-Area Archaeological Site Tracking for Change Detection

DGX agent

arXiv:2605.08160v1 Announce Type: cross Abstract: Monitoring archaeological sites at scale is vital for protecting cultural heritage, yet pinpointing when disturbances occur remains difficult because

model-releasesarxiv-cs-ai
12 May 2026
Research

What Concepts Lie Within? Detecting and Suppressing Risky Content in Diffusion Transformers

DGX agent

arXiv:2605.10180v1 Announce Type: new Abstract: The rise of text-to-image (T2I) models has increasingly raised concerns regarding the generation of risky content, such as sexual, violent, and copyrigh

researcharxiv-cs-cv
12 May 2026
Model Releases

What Parameter Golf taught us about AI-assisted research

DGX agent

Parameter Golf brought together 1,000+ participants and 2,000+ submissions to explore AI-assisted machine learning research, coding agents, quantization, and novel model design under strict constraint

model-releasesopenai
12 May 2026
Model Releases

What’s new in Microsoft Foundry | April 2026

DGX agent

April brings Foundry Local GA for local AI development, GPT-5.5 model support with Tier 5 and Tier 6 default quota in Microsoft Foundry, new tracing paths for Microsoft Agent Framework and hosted agen

model-releasesmicrosoft-foundry
12 May 2026
Model Releases

When Adaptation Fails: A Gradient-Based Diagnosis of Collapsed Gating in Vision-Language Prompt Learning

DGX agent

arXiv:2605.09549v1 Announce Type: new Abstract: Adaptive prompting mechanisms have been proposed to enhance vision-language models by dynamically tailoring prompts to inputs. However, in frozen few-sh

model-releasesarxiv-cs-lg
12 May 2026
← Previous
1…541542543544545…1371
Next →