AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,419
  • Agents7,559
  • Applications5,412
  • Concepts5
  • Hardware1,837
  • Industry6,170
  • Local Ai4,934
  • Model Releases23,909
  • Research20,125
  • Safety13,371
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,419
  • Agents7,559
  • Applications5,412
  • Concepts5
  • Hardware1,837
  • Industry6,170
  • Local Ai4,934
  • Model Releases23,909
  • Research20,125
  • Safety13,371
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

Content type
AllBlog
88,419Total entries
1Added by human
88,418Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,638 results
Safety

Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuning

DGX agent

arXiv:2603.07445v2 Announce Type: replace Abstract: Large language models (LLMs) often require fine-tuning (FT) to perform well on downstream tasks, but FT can induce safety-alignment drift even when

safetyarxiv-cs-cl
4 Jun 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Knowledge Index of Noah's Ark

DGX agent

arXiv:2606.05104v1 Announce Type: new Abstract: Knowledge benchmarks for LLMs face three issues: scaling-driven designs that do not operationalize disciplinary representativeness; flat-payment annotat

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

LLMs + Persona-Plug = Personalized LLMs

DGX agent

arXiv:2409.11901v2 Announce Type: replace Abstract: Personalization plays a critical role in numerous language tasks and applications, since users with the same requirements may prefer diverse outputs

model-releasesarxiv-cs-cl
4 Jun 2026
Research

Making Expert Reasoning Learnable with Self-Distillation

DGX agent

arXiv:2602.02405v2 Announce Type: replace-cross Abstract: Improving the reasoning capabilities of large language models (LLMs) typically relies either on the model's ability to sample a correct soluti

researcharxiv-cs-ai
4 Jun 2026
Model Releases

Multi-SPIN: Multi-Access Speculative Inference for Cooperative Token Generation at the Edge

DGX agent

arXiv:2606.04581v1 Announce Type: cross Abstract: Speculative inference (SPIN) was originally developed as an efficient architecture to accelerate Large Language Models (LLMs). In this work, we propos

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

@nvidia @nebiustf Setup guide: http://hermes-agent.nousresearch.com/docs/guides/run-nemotron-3-ultra-free Sign up for Nous Portal: http://po…

DGX agent

This is a setup guide for running Nemotron-3 Ultra, NVIDIA's open-source language model, through Nous Research's platform. The guide directs users to sign up for the Nous Portal and access documentati

model-releasesnous-research--x
4 Jun 2026
Model Releases

SAM 3D: 3Dfy Anything in Images

DGX agent

arXiv:2511.16624v2 Announce Type: replace-cross Abstract: We present SAM 3D, a generative model for visually grounded 3D object reconstruction, predicting geometry, texture, and layout from a single i

model-releasesarxiv-cs-ai
4 Jun 2026
Research

Spectral Scaling Laws of Muon

DGX agent

arXiv:2606.04058v1 Announce Type: cross Abstract: Orthonormalized update rules have rapidly become a leading choice of optimizer for training large language models, with recent open-source state-of-th

researcharxiv-cs-ai
4 Jun 2026
Model Releases

VAMPS: Visual-Assisted Mathematical Problem Solving Benchmark

DGX agent

arXiv:2606.04244v1 Announce Type: new Abstract: Multimodal large language models are increasingly capable of complex reasoning, yet their performance often degrades when they must externalize a proble

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

VCIFBench: Evaluating Complex Instruction Following for Video Understanding

DGX agent

arXiv:2606.04588v1 Announce Type: new Abstract: Multimodal large language models have made rapid progress in video understanding, yet existing benchmarks largely rely on simple prompts and provide lim

model-releasesarxiv-cs-cl
4 Jun 2026
Research

AI Rater Discrimination Depends on Scoring Protocol in Complex Clinical Decision-Making

DGX agent

arXiv:2606.03198v1 Announce Type: cross Abstract: Clinical AI evaluation increasingly delegates scoring to large language models (LLMs) acting as AI raters, yet their scoring behavior across evaluatio

researcharxiv-cs-ai
3 Jun 2026
Model Releases

An Asymptotic Theory of Chain-of-Thought in In-Context Learning

DGX agent

arXiv:2606.03217v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning has become a widely used mechanism for eliciting multi-step reasoning in large language models by generating intermed

model-releasesarxiv-cs-lg
3 Jun 2026
Research

Anomalies in Multivariate Time Series Benchmarks Are Mostly Univariate

DGX agent

arXiv:2606.02670v1 Announce Type: cross Abstract: Many recent multivariate time series anomaly detection (MT-SAD) models incorporate cross-channel modeling, under the implicit assumption that the stru

researcharxiv-cs-ai
3 Jun 2026
Model Releases

ATLAS: A Large-Scale Evaluation Benchmark for Adversarial LiDAR Perception

DGX agent

arXiv:2606.02924v1 Announce Type: new Abstract: Autonomous driving perception is typically evaluated on clean benchmark data, yet real-world deployment requires robustness to rare, structured, and pot

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

AUDITFLOW: Executable Symbolic Environments for Structured Financial Reporting Verification

DGX agent

arXiv:2606.03031v1 Announce Type: new Abstract: Structured financial audit verification is difficult for language-model agents because correctness depends on structured evidence rather than text alone

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Beyond Encoder Accumulation: Measuring Encoder Roles in Multi-Encoder VLMs

DGX agent

arXiv:2606.03879v1 Announce Type: cross Abstract: As foundation models scale toward fusing more heterogeneous visual streams, understanding how diverse encoders interact under joint training becomes a

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Causal Neural Probabilistic Circuits

DGX agent

arXiv:2603.01372v2 Announce Type: replace-cross Abstract: Concept Bottleneck Models (CBMs) enhance the interpretability of end-to-end neural networks by introducing a layer of concepts and predicting

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

DECA: Decentralizing Block-Wise Adam for Efficient LLM Full-Parameter Fine-Tuning on Non-IID Data

DGX agent

arXiv:2606.03209v1 Announce Type: new Abstract: Fine-tuning large language models (LLMs) in privacy-sensitive and resource-constrained environments remains challenging. Since training data are often d

model-releasesarxiv-cs-lg
3 Jun 2026
Research

DiffUNet^2: Bidirectional Prediction, Probabilistic Generation and Collaborative Visual Discovery for Scientific Data

DGX agent

arXiv:2606.03926v1 Announce Type: cross Abstract: Modeling temporal evolution is important to analyzing and reasoning about scientific phenomena, yet most machine learning methods provide deterministi

researcharxiv-cs-lg
3 Jun 2026
Safety

Effect of Demographic Bias on Skin Lesion Classification

DGX agent

arXiv:2606.03214v1 Announce Type: new Abstract: In this study, we evaluate the performance of skin lesion classification using ResNet-based convolutional models, focusing on the impact of demographic

safetyarxiv-cs-ai
3 Jun 2026
Model Releases

From Script to Semantics: Prompting Strategies for African NLI

DGX agent

arXiv:2606.03304v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evaluated in multilingual settings, yet their inference behavior in low-resource African languages remains

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

GTBench: A Curriculum-Grounded Benchmark for Evaluating LLMs as Mathematical Research Assistants in Graph Theory

DGX agent

arXiv:2606.03144v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as self-study assistants in technical disciplines, yet their reliability as mathematical reasoning as

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning

DGX agent

arXiv:2602.06960v3 Announce Type: replace-cross Abstract: Large reasoning models achieve strong performance by scaling inference-time chain-of-thought, but this paradigm suffers from quadratic cost, c

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Let the Dynamics Flow: Stable Flow Matching Dynamical Systems

DGX agent

arXiv:2606.03834v1 Announce Type: new Abstract: Flow matching has recently emerged as a powerful approach for imitation learning, enabling scalable, expressive, and multimodal motion policies. However

model-releasesarxiv-cs-ro
3 Jun 2026
Model Releases

Multilingual Unlearning in LLMs: Transfer, Dynamics, and Reversibility

DGX agent

arXiv:2606.03291v1 Announce Type: new Abstract: Large language models (LLMs) can memorize sensitive facts, motivating unlearning methods that remove targeted knowledge without costly retraining. Howev

model-releasesarxiv-cs-cl
3 Jun 2026
Agents

MUSE: A Unified Agentic Harness for MLLMs

DGX agent

arXiv:2606.03005v1 Announce Type: cross Abstract: Despite rapid progress, multimodal large language models (MLLMs) still fail on tasks that humans solve effortlessly, such as navigating a grid maze fr

agentsarxiv-cs-ai
3 Jun 2026
Model Releases

SCOPE: Real-Time Natural Language Camera Agent at the Edge

DGX agent

arXiv:2606.02951v1 Announce Type: cross Abstract: Deploying language-driven agents in robotics requires evaluations that reflect real-world task demands: natural-language instructions with reproducibl

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection

DGX agent

arXiv:2606.03305v1 Announce Type: new Abstract: Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment. Statistical tools for de

model-releasesarxiv-cs-ai
3 Jun 2026
Research

A Direct Approach for Handling Contextual Bandits with Latent State Dynamics

DGX agent

arXiv:2604.08149v2 Announce Type: replace Abstract: We consider a linear contextual bandit model where contexts and rewards are governed by a finite hidden Markov chain. We first revisit the simplifie

researcharxiv-cs-lg
2 Jun 2026
Model Releases

A Local Perturbation Theory for Cross-Domain Interference and Recovery in Multi-Domain RL

DGX agent

arXiv:2606.02398v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training improves large language models (LLMs) on individual domains such as mathematical reasoning, code generation,

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

ACON: Optimizing Context Compression for Long-horizon LLM Agents

DGX agent

arXiv:2510.00615v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as agents in dynamic real-world environments, where success depends on maintaining precise re

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation

DGX agent

arXiv:2606.00987v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have shown strong visual understanding and language-guided grounding abilities, yet their capacity for multi-temp

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

APEX-SQL: Talking to the data via Agentic Exploration for Text-to-SQL

DGX agent

arXiv:2602.16720v2 Announce Type: replace-cross Abstract: Text-to-SQL systems powered by Large Language Models have excelled on academic benchmarks but struggle in complex enterprise environments. The

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

ATLAS: Agentic Test-time Learning-to-Allocate Scaling

DGX agent

arXiv:2606.01667v1 Announce Type: new Abstract: Test-time scaling has become a major way to improve large language model reasoning, but its orchestration has remained designer-engineered: a fixed samp

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages

DGX agent

arXiv:2606.00154v1 Announce Type: cross Abstract: Recent advancements in multimodal large language models (MLLMs) have achieved remarkable progress in multimodal reasoning and code generation, catalyz

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Can LLMs Reason Structurally? Benchmarking via the Lens of Data Structures

DGX agent

arXiv:2505.24069v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are deployed on increasingly complex tasks that require multi-step decision-making. Understanding their algorithm

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

CityTrajBench: A Unified Benchmark for City-Scale Vehicle Trajectory Generation

DGX agent

arXiv:2606.02287v1 Announce Type: cross Abstract: Urban trajectory generation is a fundamental task for transportation simulation, urban planning, and mobility analytics. However, systematic compariso

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Connecting AI agents with unstructured data using Google Cloud Storage MCP Servers

DGX agent

Google Cloud Storage (GCS) is a foundational component of the modern agentic tech stack and the preferred home for unstructured data at scale. As enterprises deploy agents in production, the critical

model-releasesgoogle-cloud-ai
2 Jun 2026
Model Releases

Connecting the Dots: Benchmarking Reflective Memory in Long-Horizon Dialogue

DGX agent

arXiv:2606.01223v1 Announce Type: cross Abstract: Despite substantial progress in long-context modeling, existing benchmarks remain confined to factual memory for explicit recall, failing to measure t

model-releasesarxiv-cs-ai
2 Jun 2026
Safety

Consistency Training while Mitigating Obfuscation via Rate Matching

DGX agent

arXiv:2606.02211v1 Announce Type: cross Abstract: Large language models are often influenced by extraneous input features, such as cues revealing a user's preferred answer. Consistency training reduce

safetyarxiv-cs-ai
2 Jun 2026
Model Releases

Consistent and Distinctive: LLM Benchmark Efficiency via Maximum Independent Set Prompt Selection on Similarity Graphs

DGX agent

arXiv:2606.01400v1 Announce Type: cross Abstract: Evaluating large language models (LLMs) across comprehensive benchmarks is expensive and time-consuming. We propose a graph-based prompt selection fra

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Decision-Focused On-Policy Learning for Contextual Linear Optimization with Partial Feedback

DGX agent

arXiv:2606.01081v1 Announce Type: new Abstract: Decision-focused learning (DFL) trains predictive models by optimizing downstream decision quality rather than standalone prediction accuracy. For conte

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

Density-Aware Translation of Spurious Correlations in Zero-Shot VLMs

DGX agent

arXiv:2606.01710v1 Announce Type: new Abstract: Vision-Language models (VLMs), such as CLIP, achieve powerful zero-shot classification. However, their predictions remain sensitive to spurious correlat

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs

DGX agent

arXiv:2606.00477v1 Announce Type: new Abstract: Unified multimodal models (UMMs) have emerged as a promising paradigm for general-purpose multimodal intelligence. As they are deployed in real-world ap

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

FineVerify: Scaling Test-Time Compute with Fine-Grained Self-Verification for Agentic Search

DGX agent

arXiv:2606.00660v1 Announce Type: new Abstract: Agentic search requires language model agents to explore many sources and answer complex information-seeking questions. Scaling test-time compute is a p

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

FVSpec: Real-World Property-Based Tests as Lean Challenges

DGX agent

arXiv:2606.01008v1 Announce Type: cross Abstract: We present a benchmark for evaluating AI models and agents on real-world formal software verification tasks. We first scrape 11,039 property-based tes

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

GIRL-DETR: Gradient-Isolated Reinforcement Learning for Video Moment Retrieval

DGX agent

arXiv:2606.00775v1 Announce Type: cross Abstract: Video Moment Retrieval (VMR) task requires accurately localizing temporal boundaries aligned with natural language queries, but many models suffer fro

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications

DGX agent

arXiv:2606.00750v1 Announce Type: new Abstract: Recent advances in visual language models have enabled autonomous agents for complex reasoning, tool use, and document understanding. However, existing

model-releasesarxiv-cs-cl
2 Jun 2026
← Previous
1…394395396397398…1326
Next →