AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
All
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,620 results
Model Releases

Open-World Evaluations for Measuring Frontier AI Capabilities

DGX agent

arXiv:2605.20520v1 Announce Type: new Abstract: Benchmark-based evaluation remains important for tracking frontier AI progress. But it can both overstate and understate deployed capability because it

model-releasesarxiv-cs-ai
22 May 2026
Model Releases
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind

DGX agent

arXiv:2605.20423v1 Announce Type: new Abstract: Large Language Models (LLMs) perform well on many language tasks, but their Theory of Mind (ToM) reasoning is still uneven in complex social settings. E

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

OSS: Open Suturing Skills Vision-Based Assessment Challenge 2024-2025

DGX agent

arXiv:2605.22200v1 Announce Type: new Abstract: Achieving high levels of surgical skill through effective training is essential for optimal patient outcomes. Automated, data-driven skill assessment ho

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

PartCo: Part-Level Correspondence Priors Enhance Category Discovery

DGX agent

arXiv:2509.22769v2 Announce Type: replace Abstract: Generalized Category Discovery (GCD) aims to identify both known and novel categories within unlabeled data by leveraging a set of labeled examples

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?

DGX agent

arXiv:2605.22109v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are increasingly deployed in human-facing roles where personality perception is critical, yet existing benchm

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Physiology and Anatomy Aware Inverse Inference of Myocardial Infarction for Cardiac Digital Twin

DGX agent

arXiv:2605.22044v1 Announce Type: new Abstract: Accurate localization of myocardial infarction is essential for risk stratification. While LGE-MRI remains the gold standard, it is resource-intensive.

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Polite on the Surface, Wrong in Practice: A Curated Dataset for Fixing Honorific Failures in Multilingual Bangla Generation

DGX agent

arXiv:2605.22487v1 Announce Type: new Abstract: Recent advances in Multilingual Large Language Models (MLLMs) have significantly enhanced cross-lingual conversational capabilities, yet modeling cultur

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Pre-VLA: Preemptive Runtime Verification for Reliable Vision-Language-Action and World-Model Rollouts

DGX agent

arXiv:2605.22446v1 Announce Type: new Abstract: While large vision-language-action (VLA) models and generative world models (WM) have advanced long-horizon embodied intelligence, their practical deplo

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

ProcBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents

DGX agent

arXiv:2605.20251v2 Announce Type: cross Abstract: Existing benchmarks for LLM coding agents primarily evaluate final outcomes. While useful for measuring overall capability, these metrics provide limi

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

PromptNCE: Pointwise Mutual Information Predictions Using Only LLMs and Contrastive Estimation Prompts

DGX agent

arXiv:2605.21776v1 Announce Type: new Abstract: Estimating mutual information from text usually requires training a task-specific critic, which limits its use in low-data settings. We ask whether larg

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Putnam 2025 Problems in Rocq using Opus 4.6 and Rocq-MCP

DGX agent

arXiv:2603.20405v2 Announce Type: replace-cross Abstract: We report on an experiment in which Claude Opus~4.6, equipped with a suite of Model Context Protocol (MCP) tools for the Rocq proof assistant,

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems

DGX agent

arXiv:2510.13910v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) mitigates key limitations of Large Language Models (LLMs)-such as factual errors, outdated knowledge, and hallu

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator

DGX agent

arXiv:2605.21748v1 Announce Type: new Abstract: As interactive LLM-based applications are created and refined, model developers need to evaluate the quality of generated text along many possible axes.

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents

DGX agent

arXiv:2605.22148v1 Announce Type: cross Abstract: Self-evolving skill libraries, pioneered by Voyager, let frozen LLM agents accumulate reusable knowledge without weight updates, yet recent evaluation

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation

DGX agent

arXiv:2605.20204v1 Announce Type: cross Abstract: LLM-based user simulation is the primary mechanism for end-to-end agent evaluation, yet simulated users are poor proxies for real humans: unconstraine

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

Reddit stock fell 5%+ on Friday after Meta launched a standalone app for online forums, called Forum; Reddit's stock is now down almost 40% this year (Jonathan Vanian/CNBC)

DGX agent

Jonathan Vanian / CNBC: Reddit stock fell 5%+ on Friday after Meta launched a standalone app for online forums, called Forum; Reddit's stock is now down almost 40% this year — Reddit shares fell about

model-releasestechmeme
22 May 2026
Model Releases

Reflective Prompt Tuning through Language Model Function-Calling

DGX agent

arXiv:2605.21781v1 Announce Type: new Abstract: Large language models (LLMs) have become increasingly capable of following instructions and complex reasoning, making prompting a flexible interface for

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Residual Skill Optimization for Text-to-SQL Ensembles

DGX agent

arXiv:2605.21792v1 Announce Type: new Abstract: Text-to-SQL ensembles improve over single-candidate generation by drawing multiple SQL candidates and selecting one, but their effectiveness is bounded

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Rethinking Noise-Robust Training for Frozen Vision Foundation Models: A Cross-Dataset Benchmark with a Case Study of Small-Loss Failure

DGX agent

arXiv:2605.22591v1 Announce Type: new Abstract: Frozen Vision Foundation Models (VFMs) with lightweight classification heads are increasingly used in medical imaging because they offer efficient and r

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

SADGE: Structure and Appearance Domain Gap Estimation of Synthetic and Real Data

DGX agent

arXiv:2605.22467v1 Announce Type: new Abstract: We propose SADGE, a quantitative similarity metric that predicts the performance of synthetic image datasets for common computer vision tasks without do

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching

DGX agent

arXiv:2605.21788v1 Announce Type: new Abstract: Zero-shot 3D visual grounding requires localizing objects in unstructured environments from free-form natural language. Recent vision-language model (VL

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals

DGX agent

arXiv:2605.21919v1 Announce Type: new Abstract: Assessing progress toward the Sustainable Development Goals (SDGs) requires multi-step reasoning over visual cues, contextual knowledge, and development

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Security Document Classification with a Fine-Tuned Local Large Language Model: Benchmark Data and an Open-Source System

DGX agent

arXiv:2605.20368v1 Announce Type: cross Abstract: Organizations that scan documents for sensitive information face a practical problem. Cloud services require data to be sent to external infrastructur

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

Seeing the Poem: Image-Semantic Detection of AI-Generated Modern Chinese Poetry with MLLMs

DGX agent

arXiv:2605.22654v1 Announce Type: new Abstract: Previous detection studies have shown that LLMs cannot be effectively used as detectors, but these studies have not addressed modern Chinese poetry. Mor

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

SegGuidedNet: Sub-Region-Aware Attention Supervision for Interpretable Brain Tumor Segmentation

DGX agent

arXiv:2605.22572v1 Announce Type: new Abstract: Accurate segmentation of brain tumour sub-regions from multi-parametric MRI is critical for treatment planning yet remains challenging due to morphologi

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Seizure-Semiology-Suite (S3): A Clinically Multimodal Dataset, Benchmark, and Models for Seizure Semiology Understanding

DGX agent

arXiv:2605.21852v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in general video understanding, their capacity to interpret invo

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

SFN-YOLO: Towards Free-Range Poultry Detection via Scale-aware Fusion Networks

DGX agent

arXiv:2509.17086v2 Announce Type: replace Abstract: Detecting and localizing poultry is essential for advancing smart poultry farming. Despite the progress of detection-centric methods, challenges per

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

SiameseNorm: Breaking the Barrier to Reconciling Pre/Post-Norm

DGX agent

arXiv:2602.08064v2 Announce Type: replace-cross Abstract: The long-standing tension between Pre- and Post-Norm remains an open problem in Transformer architecture, reflecting a fundamental trade-off b

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control

DGX agent

arXiv:2511.07820v3 Announce Type: replace-cross Abstract: Despite the rise of billion-parameter foundation models trained across thousands of GPUs, similar scaling gains have not been shown for humano

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation

DGX agent

arXiv:2605.22536v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have made rapid progress in spatial intelligence, yet existing spatial reasoning benchmarks largely assume pr

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving

DGX agent

arXiv:2512.10719v2 Announce Type: replace Abstract: End-to-end autonomous driving methods built on vision language models (VLMs) have undergone rapid development driven by their universal visual under

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

SpaceX now launches more rockets than every other country combined. This has never happened before in the history of spaceflight.

DGX agent

SpaceX has achieved a milestone where it launches more rockets than all other countries combined, marking an unprecedented moment in spaceflight history. This represents a significant shift in the com

model-releaseselon-musk--x
22 May 2026
Model Releases

Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action

DGX agent

arXiv:2605.22283v1 Announce Type: new Abstract: We introduce SOMA, the Spatial Memory framework for Out-of-Vision Manipulation in Vision-Language-Action (VLA) models. Most existing VLAs implicitly ass

model-releasesarxiv-cs-ro
22 May 2026
Model Releases

Steins;Gate Drive: Semantic Safety Arbitration over Structured Futures for Latency-Decoupled LLM Planning

DGX agent

arXiv:2605.22456v1 Announce Type: new Abstract: Cloud-hosted LLM driver agents provide useful semantic judgments, but their inference latency exceeds stepwise vehicle-control windows. Learned world mo

model-releasesarxiv-cs-ro
22 May 2026
Model Releases

Structural Anchor Pruning: Training-Free Multi-Vector Compression for Visual Document Retrieval

DGX agent

arXiv:2601.20107v2 Announce Type: replace-cross Abstract: Recent Vision-Language Models (e.g., ColPali) enable fine-grained Visual Document Retrieval (VDR) but incur prohibitive multi-vector index sto

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance

DGX agent

arXiv:2605.22202v1 Announce Type: new Abstract: In this paper, we show that high-performing embedding models organize their embedding spaces in a consistent way. We evaluate 25 contemporary embedding

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Sub-exponential Growth Dynamics in Complex Systems: A Piecewise Power-Law Model for the Diffusion of New Words and Names

DGX agent

arXiv:2511.04106v5 Announce Type: replace-cross Abstract: The diffusion of ideas and language in society has conventionally been described by S-shaped models, such as the logistic curve. However, the

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

SURGE: An Event-Centric Social Media Sentiment Time Series Benchmark with Interaction Structure

DGX agent

arXiv:2605.21198v1 Announce Type: cross Abstract: Public events on social media generate large volumes of discussion whose collective dynamics carry direct value for opinion forecasting and crisis res

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

Survive or Collapse: The Asymmetric Roles of Data Gating and Reward Grounding in Self-Play RL

DGX agent

arXiv:2605.22217v1 Announce Type: cross Abstract: Self-play reinforcement learning trains language models on their own generated tasks, co-evolving a proposer and solver without human labels. Recent s

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Tackle CSM in JPEG Steganalysis with Data Adaptation

DGX agent

arXiv:2605.21523v1 Announce Type: cross Abstract: Steganalysis models excel on benchmark datasets but struggle in the wild when analyzed images are produced by a processing pipeline unseen during trai

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Targeting Clause Type Distributions: a Picklock for Random Satisfiability Problems

DGX agent

arXiv:2605.20328v1 Announce Type: cross Abstract: Optimization problems such as the NP-complete 3-SAT provide an important benchmark for the difficult task of finding ground-states in strongly correla

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

Teaching AI Through Benchmark Construction: QuestBench as a Course-Based Practice for Accountable Knowledge Work

DGX agent

arXiv:2605.21413v2 Announce Type: new Abstract: As AI becomes part of everyday learning, many courses teach students to use it mainly as a productivity tool: how to prompt, search, summarize, write, c

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

Teaching Language Models to Forecast Research Success Through Comparative Idea Evaluation

DGX agent

arXiv:2605.21491v1 Announce Type: cross Abstract: As language models accelerate scientific research by automating hypothesis generation and implementation, a new bottleneck emerges: evaluating and fil

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Temporal Coding as a Substrate for Sensorimotor Object Inference: A Spiking Reinterpretation of Thousand Brains Architecture

DGX agent

arXiv:2605.22206v1 Announce Type: cross Abstract: The Thousand Brains Theory (TBT) and its open-source Monty framework model object recognition through sensorimotor inference -- identifying objects by

model-releasesarxiv-cs-ro
22 May 2026
Model Releases

The Blueprint: How Movix fills a gap in dental skills with specialized agentic AI

DGX agent

Welcome to The Blueprint, a regular feature where we highlight how Google Cloud customers are tackling unique and common challenges across industries using the latest AI and cloud technologies. We hop

model-releasesgoogle-cloud-ai
22 May 2026
Model Releases

The Download: coding’s future, the ‘Steroid Olympics,’ and AI-driven science

DGX agent

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Anthropic’s Code with Claude showed off coding’s future—whethe

model-releasesmit-tech-review
22 May 2026
Model Releases

Thermo-VL: Extending Vision-Language Models to Thermal Infrared Perception

DGX agent

arXiv:2605.21882v1 Announce Type: new Abstract: Vision-language models (VLMs) often fail under low illumination because their visual grounding is learned predominantly from RGB imagery, whereas therma

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Though I will continue to insist that the path we were on would have been much clearer if o3 had been called GPT-5 and GPT-5 had been called…

DGX agent

Ethan Mollick comments on OpenAI's naming convention for their AI models, specifically critiquing the decision to call the o3 model 'o3' rather than 'GPT-5,' suggesting the naming scheme creates confu

model-releasesethan-mollick--x
22 May 2026
← Previous
1…278279280281282…472
Next →