AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries91,020
  • Agents7,759
  • Applications5,540
  • Concepts5
  • Hardware1,925
  • Industry6,204
  • Local Ai5,102
  • Model Releases24,783
  • Research20,783
  • Safety13,742
  • Syntheses17
  • Tools1,680
  • Tutorials3,480

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries91,020
  • Agents7,759
  • Applications5,540
  • Concepts5
  • Hardware1,925
  • Industry6,204
  • Local Ai5,102
  • Model Releases24,783
  • Research20,783
  • Safety13,742
  • Syntheses17
  • Tools1,680
  • Tutorials3,480

Source
HumanDGX agent

Content type
91,020Total entries
1Added by human
91,019Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,766 results
Model Releases

The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing

DGX agent

arXiv:2606.02184v1 Announce Type: cross Abstract: These names do not exist. Elena Vasquez and Marcus Chen have appeared as volcano experts, astronauts, thriller protagonists, podcast hosts, and academ

model-releasesarxiv-cs-lg
2 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue

DGX agent

arXiv:2606.01901v1 Announce Type: cross Abstract: We introduce the Image Reconstruction Game, a fully automated benchmark in which a vision-language model issues corrective instructions to an image ge

model-releasesarxiv-cs-ai
2 Jun 2026
Agents

Unsupervised Collaborative Domain Adaptation for Driving Scene Parsing

DGX agent

arXiv:2606.01818v1 Announce Type: new Abstract: Reliable driving scene parsing is a fundamental capability for autonomous vehicles operating in open and dynamic driving environments. However, adapting

agentsarxiv-cs-cv
2 Jun 2026
Model Releases

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

DGX agent

arXiv:2602.16763v2 Announce Type: replace Abstract: Artificial intelligence benchmarks are an important mechanism for measuring model progress and guiding deployment decisions. However, benchmarks qui

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

WorldCoder-Bench: Benchmarking Physically Grounded 3D World Synthesis

DGX agent

arXiv:2606.01869v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly asked not only to write static interfaces, but to construct executable interactive worlds from natural lan

model-releasesarxiv-cs-ai
2 Jun 2026
Agents

Before Parc Ferme: RL-Time Pruning for Efficient Embodied LLMs in Autonomous Driving

DGX agent

arXiv:2605.31256v1 Announce Type: new Abstract: Embodied Large Language Models (LLMs) are increasingly used as reasoning modules in robotic control pipelines to improve human-robot interaction, but th

agentsarxiv-cs-ro
1 Jun 2026
Model Releases

BlueFin: Benchmarking LLM Agents on Financial Spreadsheets

DGX agent

arXiv:2605.30907v1 Announce Type: cross Abstract: We present BlueFin, a benchmark that tasks large language model (LLM) agents with synthesis, manipulation, and comprehension tasks over spreadsheet wo

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

CoFiDA-M: Concept-Aware Feature Modulation for Cross-Domain Adaptation with Image-Only Inference

DGX agent

arXiv:2605.31591v1 Announce Type: new Abstract: Models for AI-based skin cancer screening suffer a severe performance drop when shifting from expert dermoscopic (source) images to consumer-grade clini

model-releasesarxiv-cs-cv
1 Jun 2026
Research

Consolidating Rewarded Perturbations for LLM Post-Training

DGX agent

arXiv:2605.31494v1 Announce Type: new Abstract: Post-training of language models is commonly framed as a sample-score-update loop implemented by gradient descent. A recent line of work, exemplified by

researcharxiv-cs-cl
1 Jun 2026
Model Releases

CSULoRA: Closest Safe Update Low-Rank Adaptation

DGX agent

arXiv:2605.30640v1 Announce Type: cross Abstract: Low-rank adaptation has become a standard method for parameter-efficient fine-tuning of large language models, but even small amounts of unsafe or adv

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Effective Reasoning Chains Reduce Intrinsic Dimensionality

DGX agent

arXiv:2602.09276v2 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) reasoning and its variants have substantially improved the performance of language models on complex reasoning tasks, y

model-releasesarxiv-cs-ai
1 Jun 2026
Tutorials

Eigenvectors of Experts are Training-free Non-collapsing Routers

DGX agent

arXiv:2605.30992v1 Announce Type: new Abstract: Sparse Mixture of Experts (SMoE) architectures improve the training efficiency of Large Language Models (LLMs) by routing input tokens to a selected sub

tutorialsarxiv-cs-lg
1 Jun 2026
Model Releases

EUDAIMONIA: Evaluating Undesirable Dynamics in AI

DGX agent

arXiv:2605.30654v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as conversational partners for companionship, emotional disclosure, and interpersonal advice, but t

model-releasesarxiv-cs-ai
1 Jun 2026
Research

Evaluating using Mock Tool Calls to Quarantine Untrusted Prompt Inputs

DGX agent

arXiv:2605.30521v1 Announce Type: new Abstract: Large language models must frequently process untrusted inputs, such as judging an answer from another model or running tasks like spam and harm classif

researcharxiv-cs-cl
1 Jun 2026
Model Releases

FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection

DGX agent

arXiv:2605.31349v1 Announce Type: cross Abstract: Hateful meme detection remains a formidable challenge for vision-language models, as existing benchmarks are structurally observational - confounding

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs

DGX agent

arXiv:2511.18760v2 Announce Type: replace Abstract: Informal mathematics has been central to modern large language model (LLM) reasoning, offering flexibility and efficient construction of arguments.

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

LegSegNet: A Public Deep Learning System for Lower Extremity CT Tissue Segmentation and Quantification

DGX agent

arXiv:2605.30829v1 Announce Type: new Abstract: Lower extremity computed tomography (CT) contains clinically relevant information for body composition analysis, sarcopenia assessment, and musculoskele

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft

DGX agent

arXiv:2605.30931v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and action generation. However, their ability to susta

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Nemotron 3 Ultra: Frontier smart. 5X faster. 30% cheaper. 💚💚💚

DGX agent

Nemotron 3 Ultra is NVIDIA's latest language model featuring significant improvements in speed (5X faster) and cost efficiency (30% cheaper) compared to previous versions, positioning it as a frontier

model-releasesjeremy-howard--x
1 Jun 2026
Model Releases

PhyDrawGen: Physically Grounded Diagram Generation from Natural Language

DGX agent

arXiv:2605.30512v1 Announce Type: new Abstract: Generating physics diagrams from text requires strict adherence to physical laws. While current generative models produce visually plausible outputs, th

model-releasesarxiv-cs-ai
1 Jun 2026
Research

Post-Training LLMs as Better Decision-Making Agents: A Regret-Minimization Approach

DGX agent

arXiv:2511.04393v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as 'agents' for decision-making (DM) in interactive and dynamic environments. Yet, since they

researcharxiv-cs-ai
1 Jun 2026
Model Releases

PRISM: Progressive Reasoning through Iterative Slot Memory for Vision

DGX agent

arXiv:2605.30942v1 Announce Type: new Abstract: Modern vision models process images in a single feed-forward pass, which limits their ability to recover missing evidence or refine uncertain representa

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

Re-examining Low Rank adaptation for private LLM fine-tuning

DGX agent

arXiv:2510.01137v3 Announce Type: replace Abstract: Privacy is a central concern when fine-tuning large language models (LLMs) on sensitive data, and differentially private stochastic gradient descent

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

Safe Equilibrium Policy Optimization for Strategic Agent Policies

DGX agent

arXiv:2605.30854v1 Announce Type: cross Abstract: Language models fine-tuned with reinforcement learning typically optimize for task reward, ignoring multi-agent strategic structure. Because these age

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

SERA: Soft-Verified Efficient Repository Agents

DGX agent

arXiv:2601.20789v3 Announce Type: replace Abstract: Open-weight coding agents should hold a fundamental advantage over closed-source systems because they can specialize to private codebases, encoding

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Smaller and Faster 3DGS via Post-Training Dictionary Learning

DGX agent

arXiv:2605.30396v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) is a promising neural scene representation for real-time rendering, but trained models often suffer from large memory foo

model-releasesarxiv-cs-lg
1 Jun 2026
Research

Speculative Decoding Across Languages

DGX agent

arXiv:2605.30580v1 Announce Type: new Abstract: Speculative decoding has become a crucial component of large language model (LLM) inference, enabling faster generation by drafting multiple tokens and

researcharxiv-cs-cl
1 Jun 2026
Model Releases

Symbolic Intermediaries as a Linguistic-Numerical Interface for LLM-Driven Geometric Reasoning

DGX agent

arXiv:2505.17607v3 Announce Type: replace Abstract: Large Language Models (LLMs) display reasoning capabilities over linguistic and symbolic objects but have limited capabilities to directly interpret

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Triaging Threats to Specialized Guardrails

DGX agent

arXiv:2605.30693v1 Announce Type: cross Abstract: Building robust safety guardrails is essential for deploying Large Language Models across diverse real-world applications. However, this goal remains

model-releasesarxiv-cs-cl
1 Jun 2026
Safety

Who Gets Credit or Blame? Attributing Accountability in Modern AI Systems

DGX agent

arXiv:2506.00175v5 Announce Type: replace-cross Abstract: Modern AI systems are typically developed through multiple stages-pretraining, fine-tuning rounds, and subsequent adaptation or alignment, whe

safetyarxiv-cs-ai
1 Jun 2026
Model Releases

Apertus LLM Family Expansion via Distillation and Quantization

DGX agent

arXiv:2605.29128v1 Announce Type: new Abstract: The wide adoption of LLMs has led to their use in great variety of applications and scenarios, such as chatbot assistants and data annotation, creating

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

CLUBench: A Clustering Benchmark

DGX agent

arXiv:2605.29933v1 Announce Type: new Abstract: Clustering is a fundamental problem in data science with a long-standing research history, yielding numerous insightful algorithms. Despite this progres

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

ConMoE: Expert-Pool Consolidation via Prototype Reassignment for MoE Compression

DGX agent

arXiv:2605.29350v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) language models reduce per-token computation but still require storing and serving all experts, making deployment memory-intens

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

DiffSpot: Can VLMs Spot Fine-Grained Visual Differences in Web Interfaces?

DGX agent

arXiv:2605.29615v1 Announce Type: cross Abstract: Vision-language models (VLMs) have made strong progress on high-level image-text alignment, yet their ability to perceive subtle visual differences re

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

EVL-ECG: Efficient ECG Interpretation With Multi-Aspect Heterogeneous Knowledge Distillation

DGX agent

arXiv:2605.29977v1 Announce Type: new Abstract: High-fidelity ECG interpretation is increasingly reliant on massive foundation models, yet their deployment in clinical edge-care remains hindered by ex

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

Evolutionary Dynamics of Cooperation in Next-Generation LLM Agent Systems: A Cross-Provider Empirical Extension

DGX agent

arXiv:2605.29874v1 Announce Type: cross Abstract: Do next-generation LLM agents inherit the cooperative biases documented in their predecessors, or does scale and provider diversity reshape equilibriu

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Fairness Beyond Demographics: Optimizing Performance Across Appearance-Based Hidden Cohorts in Medical Imaging

DGX agent

arXiv:2605.29827v1 Announce Type: new Abstract: Medical image analysis models can exhibit performance disparities across patient subgroups, threatening clinical safety and fairness. Existing methods t

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

Gram: Assessing sabotage propensities via automated alignment auditing

DGX agent

arXiv:2605.30322v1 Announce Type: cross Abstract: We introduce Gram, an automated alignment auditing framework to assess the propensity of AI agents to engage in sabotage. We evaluate Gemini models ac

model-releasesarxiv-cs-ai
29 May 2026
Research

Horizon Activation Mapping for Neural Networks in Time Series Forecasting

DGX agent

arXiv:2601.02094v4 Announce Type: replace Abstract: Neural networks for time series forecasting have relied on error metrics and architecture-specific interpretability approaches for model selection t

researcharxiv-cs-lg
29 May 2026
Model Releases

Knowledge Offloading: Decomposing LLMs into Sparse Backbones and Memory Modules

DGX agent

arXiv:2605.29075v1 Announce Type: new Abstract: LLMs encode both general capabilities and domain-specific knowledge in a single set of parameters. We ask whether this capacity can be reorganized: keep

model-releasesarxiv-cs-lg
29 May 2026
Research

LFQ: Logit-aware Final-block Quantization for Boosting the Generation Quality of Low-Bit Quantized LLMs

DGX agent

arXiv:2605.29756v1 Announce Type: new Abstract: As large language models continue to scale, low-bit weight-only post-training quantization (PTQ) offers a practical solution to their memory-efficient d

researcharxiv-cs-ai
29 May 2026
Safety

LLUMI: Improving LLM Writing Assistance for Mental Health Support with Online Community Feedback

DGX agent

arXiv:2605.30273v1 Announce Type: cross Abstract: Large language models (LLMs) show promise in generating supportive responses for mental health queries, but improving their usefulness, empathy, and s

safetyarxiv-cs-ai
29 May 2026
Model Releases

LUMINA: A Multi-Vendor Mammography Benchmark with Energy Harmonization Protocol

DGX agent

arXiv:2603.14644v3 Announce Type: replace-cross Abstract: Publicly available full-field digital mammography (FFDM) datasets remain limited in size, clinical annotations, and vendor diversity, hinderin

model-releasesarxiv-cs-cv
29 May 2026
Research

Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels

DGX agent

arXiv:2605.29800v1 Announce Type: new Abstract: LLM-as-a-judge panels aggregate votes from multiple models, with the expectation that diverse models yield more reliable evaluations. We develop a frame

researcharxiv-cs-cl
29 May 2026
Model Releases

On-Policy Replay for Continual Supervised Fine-Tuning

DGX agent

arXiv:2605.29495v1 Announce Type: new Abstract: Continual supervised fine-tuning (SFT) is the de facto recipe for adapting large language models (LLMs) to a stream of downstream tasks, but it suffers

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster-Based Distillation

DGX agent

arXiv:2605.29829v1 Announce Type: new Abstract: Leveraging Large Language Models (LLMs) to automatically formulate and solve optimization problems from natural language has emerged as an efficient par

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Personalized Turn-Level User Conversation Satisfaction Benchmark

DGX agent

arXiv:2605.29711v1 Announce Type: cross Abstract: User satisfaction with AI assistants is highly personalized: the same response may satisfy one user but disappoint another depending on what each user

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

PRAIB: Peer Review AI Benchmark of Behaviour of LLM-Assisted Reviewing

DGX agent

arXiv:2605.29815v1 Announce Type: new Abstract: The growing number of submitted papers has motivated the exploration of Large Language Models (LLMs) as a means to support and augment the peer review p

model-releasesarxiv-cs-ai
29 May 2026
← Previous
1…410411412413414…1371
Next →