AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,573
  • Agents7,495
  • Applications5,364
  • Concepts5
  • Hardware1,813
  • Industry6,149
  • Local Ai4,892
  • Model Releases23,569
  • Research19,967
  • Safety13,263
  • Syntheses17
  • Tools1,674
  • Tutorials3,365

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,573
  • Agents7,495
  • Applications5,364
  • Concepts5
  • Hardware1,813
  • Industry6,149
  • Local Ai4,892
  • Model Releases23,569
  • Research19,967
  • Safety13,263
  • Syntheses17
  • Tools1,674
  • Tutorials3,365

Source
HumanDGX agent

Content type
87,573Total entries
1Added by human
87,572Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
51,512 results
Model Releases

Convergence of Steepest Descent and Adam under Non-Uniform Smoothness

DGX agent

arXiv:2605.30648v1 Announce Type: new Abstract: Recent work has analyzed the convergence of first-order methods under non-uniform smoothness assumptions that better model the loss landscape in machine

model-releasesarxiv-cs-lg
1 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents

DGX agent

arXiv:2605.30590v1 Announce Type: cross Abstract: Two clinical AI systems can score nearly identically on coverage-based rubrics yet behave radically differently when their patient inputs change: one

safetyarxiv-cs-ai
1 Jun 2026
Model Releases

DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories

DGX agent

arXiv:2602.10809v2 Announce Type: replace Abstract: Existing multimodal retrieval systems excel at semantic matching but implicitly assume that query-image relevance can be measured in isolation. This

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

Generalistic or Specific Embeddings, Which is Better? An Empirical Study on Search for Clinical Coding in Non-English Languages

DGX agent

arXiv:2605.30529v1 Announce Type: cross Abstract: Sentence-embedding models for semantic search are overwhelmingly developed and evaluated on English corpora. When applied to clinical retrieval in oth

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Generating Reports or Repeating Templates? Measuring and Mitigating Template Collapse in 3D CT Report Generation

DGX agent

arXiv:2605.30984v1 Announce Type: cross Abstract: Modern 3D medical vision-language models (VLMs) can generate fluent radiology-style text while exhibit critically low pathology detection and output d

model-releasesarxiv-cs-ai
1 Jun 2026
Tutorials

Learning a Zeroth-Order Optimizer for Fine-Tuning LLMs

DGX agent

arXiv:2510.00419v2 Announce Type: replace Abstract: Zeroth-order optimizers have recently emerged as an attractive approach for fine-tuning large language models (LLMs), as they avoid backpropagation

tutorialsarxiv-cs-lg
1 Jun 2026
Model Releases

MIMO: Multilingual Information Retrieval via Monolingual Objectives

DGX agent

arXiv:2605.31171v1 Announce Type: cross Abstract: Multilingual Information Retrieval (MLIR) reflects real-world search environments in which queries and relevant documents may appear in different lang

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

NeUQI: Near-Optimal Uniform Quantization Parameter Initialization for Low-Bit LLMs

DGX agent

arXiv:2505.17595v4 Announce Type: replace-cross Abstract: Large language models (LLMs) achieve impressive performance across domains but face significant challenges when deployed on consumer-grade GPU

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference

DGX agent

arXiv:2510.07651v2 Announce Type: replace-cross Abstract: Large language models (LLMs) with extended context windows enable powerful applications but impose significant memory overhead, as caching all

model-releasesarxiv-cs-ai
1 Jun 2026
Safety

REAL: Regression-Aware Reinforcement Learning for LLM-as-a-Judge

DGX agent

arXiv:2603.17145v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed as automated evaluators that assign numeric scores to model outputs, a paradigm known a

safetyarxiv-cs-ai
1 Jun 2026
Research

Rethinking Sparse Mixture of Experts from a Unified Perspective

DGX agent

arXiv:2503.22996v3 Announce Type: replace Abstract: Sparse Mixture of Experts (SMoE) models scale the capacity of models while maintaining constant computational overhead. SMoE methods fall into two c

researcharxiv-cs-cl
1 Jun 2026
Model Releases

SAW-Bench: Learning Situated Awareness in the Real World

DGX agent

arXiv:2602.16682v2 Announce Type: replace Abstract: A core aspect of human perception is situated awareness, the ability to relate ourselves to the surrounding physical environment and reason over pos

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

Scaling Multi-Hop Training Data via Graph-Constrained Path Selection

DGX agent

arXiv:2605.31238v1 Announce Type: new Abstract: Endowing large language models with compositional reasoning over specialized documents requires multi-hop training data at scale, where such data rarely

model-releasesarxiv-cs-cl
1 Jun 2026
Tutorials

Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?

DGX agent

arXiv:2605.30557v1 Announce Type: cross Abstract: Spatial reasoning is a fundamental capability for vision-language models (VLMs) deployed in real-world environments. However, visual observations are

tutorialsarxiv-cs-ai
1 Jun 2026
Safety

SemStruct: Contextualizing Semantic Embeddings with Structural Information for Schema Matching

DGX agent

arXiv:2605.30729v1 Announce Type: new Abstract: Schema matching is a fundamental step in integrating heterogeneous data sources. While Pre-trained Language Models (PLMs) have revolutionized this task

safetyarxiv-cs-lg
1 Jun 2026
Tutorials

Spectral Reach: Understanding Neural Scaling as Progress into the Spectral Tail

DGX agent

arXiv:2605.31244v1 Announce Type: new Abstract: Neural scaling laws describe predictable power-law relationships between model size, dataset size, compute, and performance. While these laws guide the

tutorialsarxiv-cs-lg
1 Jun 2026
Model Releases

The Surface You Test Is Not the Surface That Breaks

DGX agent

arXiv:2605.30454v1 Announce Type: cross Abstract: Tool-augmented LLM agents are vulnerable to prompt injection: a third party who controls part of the agent's context can plant instructions that the a

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

UniDial-EvalKit: A Unified Toolkit for Evaluating Multi-Faceted Conversational Abilities

DGX agent

arXiv:2603.23160v2 Announce Type: replace Abstract: Benchmarking large language models (LLMs) and agents in multi-turn interactive scenarios is essential for understanding their practical capabilities

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification

DGX agent

arXiv:2509.24901v4 Announce Type: replace-cross Abstract: Although probing frozen models has become a standard evaluation paradigm, self-supervised learning in audio defaults to fine-tuning when pursu

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

Active Learning for Machine Learning Driven Molecular Dynamics

DGX agent

arXiv:2509.17208v3 Announce Type: replace Abstract: Machine-learned coarse-grained (CG) potentials are fast, but degrade over time when simulations reach under-sampled bio-molecular conformations, and

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

AgentCVR: Active Multi-Agent Cross-Video Reasoning via Script-Simulated Reinforcement Learning

DGX agent

arXiv:2605.29643v1 Announce Type: new Abstract: Cross-Video Reasoning (CVR) has emerged as a critical frontier in multimodal intelligence, requiring models to retrieve, align, and aggregate evidence d

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

Beyond English and Evasion: A Human-Annotated Multi-Domain Benchmark for High-Stakes LLM Safety Evaluation in Chinese

DGX agent

arXiv:2605.29667v1 Announce Type: new Abstract: When Large Language Models (LLMs) are deployed in Chinese-language settings, a troubling pattern emerges: safety systems that work well in English break

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

CalArena: A Large-Scale Post-Hoc Calibration Benchmark

DGX agent

arXiv:2605.30188v1 Announce Type: cross Abstract: Reliable probability estimates are critical in many machine learning applications, yet modern classifiers are often poorly calibrated. Post-hoc calibr

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Connecting Independently Trained Modes via Layer-Wise Connectivity

DGX agent

arXiv:2505.02604v5 Announce Type: replace Abstract: Empirical studies have shown that continuous low-loss paths can be constructed between independently trained neural network models. This phenomenon,

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Deep Adaptive Dimension Reduction for Bayesian Inference in Inverse Problems

DGX agent

arXiv:2605.29373v1 Announce Type: new Abstract: Solving high-dimensional PDE-governed inverse problems is often challenging due to complex non-Gaussian posterior distributions, expensive forward model

model-releasesarxiv-cs-lg
29 May 2026
Tutorials

Deep Psychovisual Image Representations

DGX agent

arXiv:2605.29260v1 Announce Type: new Abstract: Psychovisual models suggest human vision decouples low-level feature extraction from higher cognition by first forming intermediate abstractions. In con

tutorialsarxiv-cs-cv
29 May 2026
Model Releases

DynSess: Dynamic Session-Level Evaluation and Optimization Framework for Role-Playing Agents

DGX agent

arXiv:2605.29256v1 Announce Type: cross Abstract: Role-playing with large language models is fundamentally a session-level task, requiring agents to sustain character identity and interaction quality

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

ESPO: Early-Stopping Proximal Policy Optimization

DGX agent

arXiv:2605.29860v1 Announce Type: cross Abstract: When a large language model under reinforcement learning commits a wrong reasoning step early in a trajectory, standard algorithms force it to keep ge

model-releasesarxiv-cs-ai
29 May 2026
Research

From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges

DGX agent

arXiv:2601.08654v2 Announce Type: replace-cross Abstract: Rubric-based text evaluation increasingly uses large language models (LLMs) as scalable judges, but aligning frozen black-box models with huma

researcharxiv-cs-ai
29 May 2026
Safety

GAP3D: Generative Alignment of VLM Latents to Patch-Level Embeddings for 3D Generation

DGX agent

arXiv:2605.28995v1 Announce Type: new Abstract: Recent approaches integrating vision-language models (VLMs) as prompt encoders for generative model conditioning typically rely on expensive end-to-end

safetyarxiv-cs-cv
29 May 2026
Model Releases

GEO-Bench: Benchmarking Ranking Manipulation in Generative Engine Optimization

DGX agent

arXiv:2605.29107v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly rank products, documents, and recommendations for user queries, which makes manipulating these rankings a gr

model-releasesarxiv-cs-ai
29 May 2026
Tutorials

Geometry Matters: 3D Foundation Priors for Learning Semantic Correspondence

DGX agent

arXiv:2605.30093v1 Announce Type: new Abstract: Foundation features from self-supervised vision models and text-to-image diffusion models have proven effective for semantic correspondence estimation.

tutorialsarxiv-cs-cv
29 May 2026
Model Releases

GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents

DGX agent

arXiv:2605.29668v1 Announce Type: new Abstract: LLM agents acting in structured environments fail in operational rather than conversational ways, and reliability depends on procedural knowledge of the

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

How Far Ahead Do LLMs Plan? Uncovering the Latent Horizon in Chain-of-Thought Reasoning

DGX agent

arXiv:2602.02103v2 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) reasoning has become a central mechanism for eliciting multi-step reasoning in Large Language Models (LLMs). Yet recent

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

LiveSVG: Zero-Shot SVG Animation via Video Generation

DGX agent

arXiv:2605.30174v1 Announce Type: new Abstract: We introduce LiveSVG, a zero-shot approach for generating Scalable Vector Graphics (SVG) animations using video diffusion models. Current SVG animation

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

Minimal Prompt Perturbations Lead to Code Vulnerabilities: Prompt Fragility and Hidden-State Signals in Coding LLMs

DGX agent

arXiv:2605.29737v1 Announce Type: cross Abstract: LLM-based coding assistants are seeing rapid adoption, offering substantial gains in developer productivity. As organizations increasingly ship code t

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions

DGX agent

arXiv:2605.29738v1 Announce Type: cross Abstract: Legal NLP benchmarks overwhelmingly evaluate a single language or aggregate tasks that differ fundamentally across jurisdictions, making cross-lingual

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs

DGX agent

arXiv:2605.29300v1 Announce Type: cross Abstract: Recent Large Audio-Language Models (LALMs) have demonstrated promising abilities in understanding musical content. However, whether their responses ar

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

NaRA: Noise-Aware LoRA for Parameter-Efficient Fine-Tuning of Diffusion LLMs

DGX agent

arXiv:2605.29716v1 Announce Type: new Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising non-autoregressive generative paradigm. Given the prohibitive computational cost of

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

NICE: A Theory-Grounded Diagnostic Benchmark for Social Intelligence of LLMs

DGX agent

arXiv:2605.29685v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly applied in social contexts such as emotional companionship and customer service, measuring their social

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems

DGX agent

arXiv:2605.29676v1 Announce Type: new Abstract: Large language models in Agentic AI systems consume tool schemas and execution results and emit tool invocations as structured data. The default languag

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields

DGX agent

arXiv:2605.29833v1 Announce Type: new Abstract: As multimodal language models play an increasingly important role in scientific research, materials science offers a critical testbed due to its interdi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

On the Construction and Implications of Low-Loss Valleys in LoRA-based Bayesian Inference

DGX agent

arXiv:2605.29580v1 Announce Type: new Abstract: While parameter-efficient fine-tuning methods like low-rank adaptation (LoRA) are standard for large language models, principled estimation of epistemic

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Overcoming Forgetting in LLM Fine-Tuning with Evolution Strategies

DGX agent

arXiv:2605.30148v1 Announce Type: cross Abstract: Evolution Strategies (ES) has recently emerged as a competitive alternative to reinforcement learning (RL) for large language model (LLM) fine-tuning,

model-releasesarxiv-cs-ai
29 May 2026
Safety

PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning

DGX agent

arXiv:2605.29582v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown promise as educational tutors, yet effective tutoring requires more than solving problems: it must provide pro

safetyarxiv-cs-cl
29 May 2026
Model Releases

Return-to-Go Is More Than a Number: Q-Guided Alignment for Return-Conditioned Supervised Learning

DGX agent

arXiv:2605.29028v1 Announce Type: cross Abstract: Conditioned Sequence Models (CSMs) learn policies by treating return-to-go (RTG) as a control signal. However, existing CSMs often treat the RTGs as s

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Robust and Efficient Guardrails with Latent Reasoning

DGX agent

arXiv:2605.29068v1 Announce Type: new Abstract: Maintaining the safety of large language models (LLMs) is crucial as they are increasingly deployed in real-world applications. Existing safety guardrai

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

SciIntBench: Measuring LLM Compliance with Research Integrity Norms Under Adversarial Framing

DGX agent

arXiv:2605.29468v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to support scientific work, but it is unclear whether they uphold responsible conduct of research (

model-releasesarxiv-cs-ai
29 May 2026
← Previous
1…362363364365366…1074
Next →