AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,419
  • Agents7,559
  • Applications5,412
  • Concepts5
  • Hardware1,837
  • Industry6,170
  • Local Ai4,934
  • Model Releases23,909
  • Research20,125
  • Safety13,371
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,419
  • Agents7,559
  • Applications5,412
  • Concepts5
  • Hardware1,837
  • Industry6,170
  • Local Ai4,934
  • Model Releases23,909
  • Research20,125
  • Safety13,371
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

Content type
AllBlog
88,419Total entries
1Added by human
88,418Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,638 results
Model Releases

Identifying High-Confidence Social Biases in LLMs for Trustworthy Conversational Tutoring Agents

DGX agent

arXiv:2606.01584v1 Announce Type: cross Abstract: Conversational tutoring agents have been shown to improve learning engagement and student outcomes, and large language models (LLMs) are increasingly

model-releasesarxiv-cs-ai
2 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Investigating and Alleviating Harm Amplification in LLM Interactions

DGX agent

arXiv:2606.02423v1 Announce Type: new Abstract: Large language models (LLMs) can serve as helpful assistants, yet they can equally function as harm amplifiers that enable malicious users to achieve ha

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

'I've Seen How This Goes': Characterizing Diversity via Progressive Conditional Surprise

DGX agent

arXiv:2606.01811v1 Announce Type: cross Abstract: Measuring the diversity of creative outputs is central to evaluating post-training mode collapse, comparing decoding strategies, and quantifying creat

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts

DGX agent

arXiv:2606.02404v1 Announce Type: new Abstract: Frontier model evaluations are shifting from foundational capabilities (e.g., instruction following and reasoning) toward compositional, agentic ones, b

model-releasesarxiv-cs-cl
2 Jun 2026
Research

LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding

DGX agent

arXiv:2602.23881v2 Announce Type: replace-cross Abstract: Speculative decoding accelerates autoregressive large language model (LLM) inference by using a lightweight draft model to propose candidate t

researcharxiv-cs-cl
2 Jun 2026
Model Releases

LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services

DGX agent

arXiv:2512.07436v3 Announce Type: replace Abstract: Recent advances in large reasoning models LRMs have enabled agentic search systems to perform complex multi-step reasoning across multiple sources.

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

MM-Snowball: Evaluating and Mitigating Hallucination Snowballing in Multimodal Multi-Turn Dialogue

DGX agent

arXiv:2606.00622v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) demonstrate remarkable visual understanding, yet their reliability in interactive settings is severely undermin

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification

DGX agent

arXiv:2606.01476v1 Announce Type: cross Abstract: On-Policy Distillation (OPD) trains a student model on its own generative trajectories under dense token-level feedback from a stronger teacher, mitig

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

On the Evaluation of Spiking Neural Network Configurations for Network Intrusion Detection

DGX agent

arXiv:2606.01442v1 Announce Type: cross Abstract: Network intrusion detection is a core component of modern cybersecurity infrastructure, yet the deep learning models that dominate the field are compu

model-releasesarxiv-cs-ai
2 Jun 2026
Local Ai

One-Shot Crowd Counting With Density Guidance For Scene Adaptation

DGX agent

arXiv:2602.07955v2 Announce Type: replace Abstract: Crowd scenes captured by cameras at different locations vary greatly, and existing crowd models have limited generalization for unseen surveillance

local-aiarxiv-cs-cv
2 Jun 2026
Model Releases

Optimal Regularization for Performative Learning

DGX agent

arXiv:2510.12249v2 Announce Type: replace Abstract: In performative learning, the data distribution reacts to the deployed model - for example, because strategic users adapt their features to game it

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

PaSBench-Video: A Streaming Video Benchmark for Proactive Safety Warning

DGX agent

arXiv:2606.02443v1 Announce Type: cross Abstract: Between the first visible sign of danger and the moment an accident occurs, there is often a window where intervention remains possible. Video-capable

model-releasesarxiv-cs-ai
2 Jun 2026
Research

Pinterest Canvas: Large-Scale Image Generation at Pinterest

DGX agent

arXiv:2603.06453v2 Announce Type: replace Abstract: While recent image generation models demonstrate a remarkable ability to handle a wide variety of image generation tasks, this flexibility makes the

researcharxiv-cs-cv
2 Jun 2026
Model Releases

Product-Aware Deep Autoencoders for Robust Process Monitoring in Multi-Product Cyber-Physical Systems

DGX agent

arXiv:2606.00052v1 Announce Type: new Abstract: As Industry 4.0 accelerates the integration of Cyber-Physical Systems (CPS) in manufacturing, robust anomaly detection has become critical for ensuring

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning

DGX agent

arXiv:2606.00963v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) exhibit emerging spatial reasoning capabilities, yet they remain unreliable on tasks requiring precise spatial understan

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning

DGX agent

arXiv:2512.07795v2 Announce Type: replace Abstract: Benchmark scores for LLM reasoning systems are reported as single numbers, yet the same model, strategy, and task can produce meaningfully different

model-releasesarxiv-cs-ai
2 Jun 2026
Research

Reconstructing Content via Collaborative Attention to Improve Multimodal Embedding Quality

DGX agent

arXiv:2603.01471v2 Announce Type: replace-cross Abstract: Multimodal embedding models, rooted in multimodal large language models (MLLMs), have yielded significant performance improvements across dive

researcharxiv-cs-lg
2 Jun 2026
Model Releases

Resonant Context Anchoring: Decoupling Attention Routing and Signal Gain at Inference Time

DGX agent

arXiv:2606.01923v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently exhibit 'contextual disregard' when faced with input evidence that conflicts with their internal parametric memo

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

SDR: Set-Distance Rewards for Radiology Report Generation

DGX agent

arXiv:2606.00440v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has rapidly advanced reasoning in vision--language models. However, for chest X-ray report generation, th

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Suppressing Forgery-Specific Shortcuts for Generalizable Deepfake Detection

DGX agent

arXiv:2606.01843v1 Announce Type: cross Abstract: Deepfake detection suffers from poor generalization across forgery methods, as existing models tend to rely on spurious method-specific shortcuts that

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

TCAR-Gen: Temporal Graph Retrieval with Evidence Fusion for Knowledge-Grounded Generation

DGX agent

arXiv:2606.00029v1 Announce Type: cross Abstract: Retrieval-augmented generation systems struggle with temporal reasoning and evidence fusion when answering complex questions over historical criminal

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

TECCI: Tricky Edits of Collected and Curated Images

DGX agent

arXiv:2606.01213v1 Announce Type: cross Abstract: Despite tremendous recent progress, current text-guided image editing methods still struggle with many aspects of editing involving instruction follow

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing

DGX agent

arXiv:2606.02184v1 Announce Type: cross Abstract: These names do not exist. Elena Vasquez and Marcus Chen have appeared as volcano experts, astronauts, thriller protagonists, podcast hosts, and academ

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue

DGX agent

arXiv:2606.01901v1 Announce Type: cross Abstract: We introduce the Image Reconstruction Game, a fully automated benchmark in which a vision-language model issues corrective instructions to an image ge

model-releasesarxiv-cs-ai
2 Jun 2026
Agents

Unsupervised Collaborative Domain Adaptation for Driving Scene Parsing

DGX agent

arXiv:2606.01818v1 Announce Type: new Abstract: Reliable driving scene parsing is a fundamental capability for autonomous vehicles operating in open and dynamic driving environments. However, adapting

agentsarxiv-cs-cv
2 Jun 2026
Model Releases

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

DGX agent

arXiv:2602.16763v2 Announce Type: replace Abstract: Artificial intelligence benchmarks are an important mechanism for measuring model progress and guiding deployment decisions. However, benchmarks qui

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

WorldCoder-Bench: Benchmarking Physically Grounded 3D World Synthesis

DGX agent

arXiv:2606.01869v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly asked not only to write static interfaces, but to construct executable interactive worlds from natural lan

model-releasesarxiv-cs-ai
2 Jun 2026
Agents

Before Parc Ferme: RL-Time Pruning for Efficient Embodied LLMs in Autonomous Driving

DGX agent

arXiv:2605.31256v1 Announce Type: new Abstract: Embodied Large Language Models (LLMs) are increasingly used as reasoning modules in robotic control pipelines to improve human-robot interaction, but th

agentsarxiv-cs-ro
1 Jun 2026
Model Releases

BlueFin: Benchmarking LLM Agents on Financial Spreadsheets

DGX agent

arXiv:2605.30907v1 Announce Type: cross Abstract: We present BlueFin, a benchmark that tasks large language model (LLM) agents with synthesis, manipulation, and comprehension tasks over spreadsheet wo

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

CoFiDA-M: Concept-Aware Feature Modulation for Cross-Domain Adaptation with Image-Only Inference

DGX agent

arXiv:2605.31591v1 Announce Type: new Abstract: Models for AI-based skin cancer screening suffer a severe performance drop when shifting from expert dermoscopic (source) images to consumer-grade clini

model-releasesarxiv-cs-cv
1 Jun 2026
Research

Consolidating Rewarded Perturbations for LLM Post-Training

DGX agent

arXiv:2605.31494v1 Announce Type: new Abstract: Post-training of language models is commonly framed as a sample-score-update loop implemented by gradient descent. A recent line of work, exemplified by

researcharxiv-cs-cl
1 Jun 2026
Model Releases

CSULoRA: Closest Safe Update Low-Rank Adaptation

DGX agent

arXiv:2605.30640v1 Announce Type: cross Abstract: Low-rank adaptation has become a standard method for parameter-efficient fine-tuning of large language models, but even small amounts of unsafe or adv

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Effective Reasoning Chains Reduce Intrinsic Dimensionality

DGX agent

arXiv:2602.09276v2 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) reasoning and its variants have substantially improved the performance of language models on complex reasoning tasks, y

model-releasesarxiv-cs-ai
1 Jun 2026
Tutorials

Eigenvectors of Experts are Training-free Non-collapsing Routers

DGX agent

arXiv:2605.30992v1 Announce Type: new Abstract: Sparse Mixture of Experts (SMoE) architectures improve the training efficiency of Large Language Models (LLMs) by routing input tokens to a selected sub

tutorialsarxiv-cs-lg
1 Jun 2026
Model Releases

EUDAIMONIA: Evaluating Undesirable Dynamics in AI

DGX agent

arXiv:2605.30654v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as conversational partners for companionship, emotional disclosure, and interpersonal advice, but t

model-releasesarxiv-cs-ai
1 Jun 2026
Research

Evaluating using Mock Tool Calls to Quarantine Untrusted Prompt Inputs

DGX agent

arXiv:2605.30521v1 Announce Type: new Abstract: Large language models must frequently process untrusted inputs, such as judging an answer from another model or running tasks like spam and harm classif

researcharxiv-cs-cl
1 Jun 2026
Model Releases

FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection

DGX agent

arXiv:2605.31349v1 Announce Type: cross Abstract: Hateful meme detection remains a formidable challenge for vision-language models, as existing benchmarks are structurally observational - confounding

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs

DGX agent

arXiv:2511.18760v2 Announce Type: replace Abstract: Informal mathematics has been central to modern large language model (LLM) reasoning, offering flexibility and efficient construction of arguments.

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

LegSegNet: A Public Deep Learning System for Lower Extremity CT Tissue Segmentation and Quantification

DGX agent

arXiv:2605.30829v1 Announce Type: new Abstract: Lower extremity computed tomography (CT) contains clinically relevant information for body composition analysis, sarcopenia assessment, and musculoskele

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft

DGX agent

arXiv:2605.30931v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and action generation. However, their ability to susta

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Nemotron 3 Ultra: Frontier smart. 5X faster. 30% cheaper. 💚💚💚

DGX agent

Nemotron 3 Ultra is NVIDIA's latest language model featuring significant improvements in speed (5X faster) and cost efficiency (30% cheaper) compared to previous versions, positioning it as a frontier

model-releasesjeremy-howard--x
1 Jun 2026
Model Releases

PhyDrawGen: Physically Grounded Diagram Generation from Natural Language

DGX agent

arXiv:2605.30512v1 Announce Type: new Abstract: Generating physics diagrams from text requires strict adherence to physical laws. While current generative models produce visually plausible outputs, th

model-releasesarxiv-cs-ai
1 Jun 2026
Research

Post-Training LLMs as Better Decision-Making Agents: A Regret-Minimization Approach

DGX agent

arXiv:2511.04393v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as 'agents' for decision-making (DM) in interactive and dynamic environments. Yet, since they

researcharxiv-cs-ai
1 Jun 2026
Model Releases

PRISM: Progressive Reasoning through Iterative Slot Memory for Vision

DGX agent

arXiv:2605.30942v1 Announce Type: new Abstract: Modern vision models process images in a single feed-forward pass, which limits their ability to recover missing evidence or refine uncertain representa

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

Re-examining Low Rank adaptation for private LLM fine-tuning

DGX agent

arXiv:2510.01137v3 Announce Type: replace Abstract: Privacy is a central concern when fine-tuning large language models (LLMs) on sensitive data, and differentially private stochastic gradient descent

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

Safe Equilibrium Policy Optimization for Strategic Agent Policies

DGX agent

arXiv:2605.30854v1 Announce Type: cross Abstract: Language models fine-tuned with reinforcement learning typically optimize for task reward, ignoring multi-agent strategic structure. Because these age

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

SERA: Soft-Verified Efficient Repository Agents

DGX agent

arXiv:2601.20789v3 Announce Type: replace Abstract: Open-weight coding agents should hold a fundamental advantage over closed-source systems because they can specialize to private codebases, encoding

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Smaller and Faster 3DGS via Post-Training Dictionary Learning

DGX agent

arXiv:2605.30396v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) is a promising neural scene representation for real-time rendering, but trained models often suffer from large memory foo

model-releasesarxiv-cs-lg
1 Jun 2026
← Previous
1…395396397398399…1326
Next →