AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,595 results
22 May 2026

OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind

Model ReleasesDGX agent

arXiv:2605.20423v1 Announce Type: new Abstract: Large Language Models (LLMs) perform well on many language tasks, but their Theory of Mind (ToM) reasoning is still uneven in complex social settings. E

OSS: Open Suturing Skills Vision-Based Assessment Challenge 2024-2025

Model ReleasesDGX agent

arXiv:2605.22200v1 Announce Type: new Abstract: Achieving high levels of surgical skill through effective training is essential for optimal patient outcomes. Automated, data-driven skill assessment ho

PartCo: Part-Level Correspondence Priors Enhance Category Discovery

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2509.22769v2 Announce Type: replace Abstract: Generalized Category Discovery (GCD) aims to identify both known and novel categories within unlabeled data by leveraging a set of labeled examples

Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?

Model ReleasesDGX agent

arXiv:2605.22109v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are increasingly deployed in human-facing roles where personality perception is critical, yet existing benchm

Physiology and Anatomy Aware Inverse Inference of Myocardial Infarction for Cardiac Digital Twin

Model ReleasesDGX agent

arXiv:2605.22044v1 Announce Type: new Abstract: Accurate localization of myocardial infarction is essential for risk stratification. While LGE-MRI remains the gold standard, it is resource-intensive.

Polite on the Surface, Wrong in Practice: A Curated Dataset for Fixing Honorific Failures in Multilingual Bangla Generation

Model ReleasesDGX agent

arXiv:2605.22487v1 Announce Type: new Abstract: Recent advances in Multilingual Large Language Models (MLLMs) have significantly enhanced cross-lingual conversational capabilities, yet modeling cultur

Pre-VLA: Preemptive Runtime Verification for Reliable Vision-Language-Action and World-Model Rollouts

Model ReleasesDGX agent

arXiv:2605.22446v1 Announce Type: new Abstract: While large vision-language-action (VLA) models and generative world models (WM) have advanced long-horizon embodied intelligence, their practical deplo

ProcBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents

Model ReleasesDGX agent

arXiv:2605.20251v2 Announce Type: cross Abstract: Existing benchmarks for LLM coding agents primarily evaluate final outcomes. While useful for measuring overall capability, these metrics provide limi

PromptNCE: Pointwise Mutual Information Predictions Using Only LLMs and Contrastive Estimation Prompts

Model ReleasesDGX agent

arXiv:2605.21776v1 Announce Type: new Abstract: Estimating mutual information from text usually requires training a task-specific critic, which limits its use in low-data settings. We ask whether larg

Putnam 2025 Problems in Rocq using Opus 4.6 and Rocq-MCP

Model ReleasesDGX agent

arXiv:2603.20405v2 Announce Type: replace-cross Abstract: We report on an experiment in which Claude Opus~4.6, equipped with a suite of Model Context Protocol (MCP) tools for the Rocq proof assistant,

RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems

Model ReleasesDGX agent

arXiv:2510.13910v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) mitigates key limitations of Large Language Models (LLMs)-such as factual errors, outdated knowledge, and hallu

RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator

Model ReleasesDGX agent

arXiv:2605.21748v1 Announce Type: new Abstract: As interactive LLM-based applications are created and refined, model developers need to evaluate the quality of generated text along many possible axes.

Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents

Model ReleasesDGX agent

arXiv:2605.22148v1 Announce Type: cross Abstract: Self-evolving skill libraries, pioneered by Voyager, let frozen LLM agents accumulate reusable knowledge without weight updates, yet recent evaluation

RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation

Model ReleasesDGX agent

arXiv:2605.20204v1 Announce Type: cross Abstract: LLM-based user simulation is the primary mechanism for end-to-end agent evaluation, yet simulated users are poor proxies for real humans: unconstraine

Reddit stock fell 5%+ on Friday after Meta launched a standalone app for online forums, called Forum; Reddit's stock is now down almost 40% this year (Jonathan Vanian/CNBC)

Model ReleasesDGX agent

Jonathan Vanian / CNBC: Reddit stock fell 5%+ on Friday after Meta launched a standalone app for online forums, called Forum; Reddit's stock is now down almost 40% this year — Reddit shares fell about

Reflective Prompt Tuning through Language Model Function-Calling

Model ReleasesDGX agent

arXiv:2605.21781v1 Announce Type: new Abstract: Large language models (LLMs) have become increasingly capable of following instructions and complex reasoning, making prompting a flexible interface for

Residual Skill Optimization for Text-to-SQL Ensembles

Model ReleasesDGX agent

arXiv:2605.21792v1 Announce Type: new Abstract: Text-to-SQL ensembles improve over single-candidate generation by drawing multiple SQL candidates and selecting one, but their effectiveness is bounded

Rethinking Noise-Robust Training for Frozen Vision Foundation Models: A Cross-Dataset Benchmark with a Case Study of Small-Loss Failure

Model ReleasesDGX agent

arXiv:2605.22591v1 Announce Type: new Abstract: Frozen Vision Foundation Models (VFMs) with lightweight classification heads are increasingly used in medical imaging because they offer efficient and r

SADGE: Structure and Appearance Domain Gap Estimation of Synthetic and Real Data

Model ReleasesDGX agent

arXiv:2605.22467v1 Announce Type: new Abstract: We propose SADGE, a quantitative similarity metric that predicts the performance of synthetic image datasets for common computer vision tasks without do

SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching

Model ReleasesDGX agent

arXiv:2605.21788v1 Announce Type: new Abstract: Zero-shot 3D visual grounding requires localizing objects in unstructured environments from free-form natural language. Recent vision-language model (VL

SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals

Model ReleasesDGX agent

arXiv:2605.21919v1 Announce Type: new Abstract: Assessing progress toward the Sustainable Development Goals (SDGs) requires multi-step reasoning over visual cues, contextual knowledge, and development

Security Document Classification with a Fine-Tuned Local Large Language Model: Benchmark Data and an Open-Source System

Model ReleasesDGX agent

arXiv:2605.20368v1 Announce Type: cross Abstract: Organizations that scan documents for sensitive information face a practical problem. Cloud services require data to be sent to external infrastructur

Seeing the Poem: Image-Semantic Detection of AI-Generated Modern Chinese Poetry with MLLMs

Model ReleasesDGX agent

arXiv:2605.22654v1 Announce Type: new Abstract: Previous detection studies have shown that LLMs cannot be effectively used as detectors, but these studies have not addressed modern Chinese poetry. Mor

SegGuidedNet: Sub-Region-Aware Attention Supervision for Interpretable Brain Tumor Segmentation

Model ReleasesDGX agent

arXiv:2605.22572v1 Announce Type: new Abstract: Accurate segmentation of brain tumour sub-regions from multi-parametric MRI is critical for treatment planning yet remains challenging due to morphologi

Seizure-Semiology-Suite (S3): A Clinically Multimodal Dataset, Benchmark, and Models for Seizure Semiology Understanding

Model ReleasesDGX agent

arXiv:2605.21852v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in general video understanding, their capacity to interpret invo

SFN-YOLO: Towards Free-Range Poultry Detection via Scale-aware Fusion Networks

Model ReleasesDGX agent

arXiv:2509.17086v2 Announce Type: replace Abstract: Detecting and localizing poultry is essential for advancing smart poultry farming. Despite the progress of detection-centric methods, challenges per

SiameseNorm: Breaking the Barrier to Reconciling Pre/Post-Norm

Model ReleasesDGX agent

arXiv:2602.08064v2 Announce Type: replace-cross Abstract: The long-standing tension between Pre- and Post-Norm remains an open problem in Transformer architecture, reflecting a fundamental trade-off b

SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control

Model ReleasesDGX agent

arXiv:2511.07820v3 Announce Type: replace-cross Abstract: Despite the rise of billion-parameter foundation models trained across thousands of GPUs, similar scaling gains have not been shown for humano

SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation

Model ReleasesDGX agent

arXiv:2605.22536v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have made rapid progress in spatial intelligence, yet existing spatial reasoning benchmarks largely assume pr

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving

Model ReleasesDGX agent

arXiv:2512.10719v2 Announce Type: replace Abstract: End-to-end autonomous driving methods built on vision language models (VLMs) have undergone rapid development driven by their universal visual under

SpaceX now launches more rockets than every other country combined. This has never happened before in the history of spaceflight.

Model ReleasesDGX agent

SpaceX has achieved a milestone where it launches more rockets than all other countries combined, marking an unprecedented moment in spaceflight history. This represents a significant shift in the com

Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action

Model ReleasesDGX agent

arXiv:2605.22283v1 Announce Type: new Abstract: We introduce SOMA, the Spatial Memory framework for Out-of-Vision Manipulation in Vision-Language-Action (VLA) models. Most existing VLAs implicitly ass

Steins;Gate Drive: Semantic Safety Arbitration over Structured Futures for Latency-Decoupled LLM Planning

Model ReleasesDGX agent

arXiv:2605.22456v1 Announce Type: new Abstract: Cloud-hosted LLM driver agents provide useful semantic judgments, but their inference latency exceeds stepwise vehicle-control windows. Learned world mo

Structural Anchor Pruning: Training-Free Multi-Vector Compression for Visual Document Retrieval

Model ReleasesDGX agent

arXiv:2601.20107v2 Announce Type: replace-cross Abstract: Recent Vision-Language Models (e.g., ColPali) enable fine-grained Visual Document Retrieval (VDR) but incur prohibitive multi-vector index sto

Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance

Model ReleasesDGX agent

arXiv:2605.22202v1 Announce Type: new Abstract: In this paper, we show that high-performing embedding models organize their embedding spaces in a consistent way. We evaluate 25 contemporary embedding

Sub-exponential Growth Dynamics in Complex Systems: A Piecewise Power-Law Model for the Diffusion of New Words and Names

Model ReleasesDGX agent

arXiv:2511.04106v5 Announce Type: replace-cross Abstract: The diffusion of ideas and language in society has conventionally been described by S-shaped models, such as the logistic curve. However, the

SURGE: An Event-Centric Social Media Sentiment Time Series Benchmark with Interaction Structure

Model ReleasesDGX agent

arXiv:2605.21198v1 Announce Type: cross Abstract: Public events on social media generate large volumes of discussion whose collective dynamics carry direct value for opinion forecasting and crisis res

Survive or Collapse: The Asymmetric Roles of Data Gating and Reward Grounding in Self-Play RL

Model ReleasesDGX agent

arXiv:2605.22217v1 Announce Type: cross Abstract: Self-play reinforcement learning trains language models on their own generated tasks, co-evolving a proposer and solver without human labels. Recent s

Tackle CSM in JPEG Steganalysis with Data Adaptation

Model ReleasesDGX agent

arXiv:2605.21523v1 Announce Type: cross Abstract: Steganalysis models excel on benchmark datasets but struggle in the wild when analyzed images are produced by a processing pipeline unseen during trai

Targeting Clause Type Distributions: a Picklock for Random Satisfiability Problems

Model ReleasesDGX agent

arXiv:2605.20328v1 Announce Type: cross Abstract: Optimization problems such as the NP-complete 3-SAT provide an important benchmark for the difficult task of finding ground-states in strongly correla

Teaching AI Through Benchmark Construction: QuestBench as a Course-Based Practice for Accountable Knowledge Work

Model ReleasesDGX agent

arXiv:2605.21413v2 Announce Type: new Abstract: As AI becomes part of everyday learning, many courses teach students to use it mainly as a productivity tool: how to prompt, search, summarize, write, c

Teaching Language Models to Forecast Research Success Through Comparative Idea Evaluation

Model ReleasesDGX agent

arXiv:2605.21491v1 Announce Type: cross Abstract: As language models accelerate scientific research by automating hypothesis generation and implementation, a new bottleneck emerges: evaluating and fil

Temporal Coding as a Substrate for Sensorimotor Object Inference: A Spiking Reinterpretation of Thousand Brains Architecture

Model ReleasesDGX agent

arXiv:2605.22206v1 Announce Type: cross Abstract: The Thousand Brains Theory (TBT) and its open-source Monty framework model object recognition through sensorimotor inference -- identifying objects by

The Blueprint: How Movix fills a gap in dental skills with specialized agentic AI

Model ReleasesDGX agent

Welcome to The Blueprint, a regular feature where we highlight how Google Cloud customers are tackling unique and common challenges across industries using the latest AI and cloud technologies. We hop

The Download: coding’s future, the ‘Steroid Olympics,’ and AI-driven science

Model ReleasesDGX agent

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Anthropic’s Code with Claude showed off coding’s future—whethe

Thermo-VL: Extending Vision-Language Models to Thermal Infrared Perception

Model ReleasesDGX agent

arXiv:2605.21882v1 Announce Type: new Abstract: Vision-language models (VLMs) often fail under low illumination because their visual grounding is learned predominantly from RGB imagery, whereas therma

Though I will continue to insist that the path we were on would have been much clearer if o3 had been called GPT-5 and GPT-5 had been called…

Model ReleasesDGX agent

Ethan Mollick comments on OpenAI's naming convention for their AI models, specifically critiquing the decision to call the o3 model 'o3' rather than 'GPT-5,' suggesting the naming scheme creates confu

Token-Level LLM Collaboration via FusionRoute

Model ReleasesDGX agent

arXiv:2601.05106v4 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit strengths across diverse domains. However, achieving strong performance across these domains with a singl

Tokenization with Split Trees

Model ReleasesDGX agent

arXiv:2605.22705v1 Announce Type: new Abstract: We introduce Tokenization with Split Trees (ToaST), a subword tokenization method that directly optimizes compression under a new recursive inference pr

Towards Clinically Interpretable Ophthalmic VQA via Spatially-Grounded Lesion Evidence

Model ReleasesDGX agent

arXiv:2605.22414v1 Announce Type: new Abstract: Visual Question Answering (VQA) holds great promise for clinical support, particularly in ophthalmology, where retinal fundus photography is essential f

Towards Selection of Large Multimodal Models as Engines for Burned-in Protected Health Information Detection in Medical Images

Model ReleasesDGX agent

arXiv:2511.02014v2 Announce Type: replace Abstract: The detection of Protected Health Information (PHI) in medical imaging is critical for safeguarding patient privacy and ensuring compliance with reg

Training-Trajectory-Aware Token Selection

Model ReleasesDGX agent

arXiv:2601.10348v2 Announce Type: replace Abstract: Efficient distillation is a key pathway for converting expensive reasoning capability into deployable efficiency, yet in the frontier regime where t

Transcription and Recognition of Italian Parliamentary Speeches Using Vision-Language Models

Model ReleasesDGX agent

arXiv:2603.28103v2 Announce Type: replace-cross Abstract: Parliamentary proceedings represent a rich yet challenging resource for computational analysis, particularly when preserved only as scanned hi

TransitLM: A Large-Scale Dataset and Benchmark for Map-Free Transit Route Generation

Model ReleasesDGX agent

arXiv:2605.22355v1 Announce Type: new Abstract: Public transit route planning traditionally depends on structured map infrastructure and complex routing engines, and no existing dataset supports train

Transporting Task Vectors across Different Architectures without Training

Model ReleasesDGX agent

arXiv:2602.12952v2 Announce Type: replace-cross Abstract: Adapting large pre-trained models to downstream tasks often produces task-specific parameter updates that are expensive to relearn for every m

Two is better than one: A Collapse-free Multi-Reward RLIF Training Framework

Model ReleasesDGX agent

arXiv:2605.22620v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning ability of LLMs, but often depends on external supervis

Ultra-High-Definition Image Quality Assessment via Graph Representation Learning

Model ReleasesDGX agent

arXiv:2605.22192v1 Announce Type: new Abstract: Blind image quality assessment (BIQA) for ultrahighdefinition (UHD) images remains challenging because native-resolution inference is computationally ex

Understanding Data Temporality Impact on Large Language Models Pre-training

Model ReleasesDGX agent

arXiv:2605.22769v1 Announce Type: new Abstract: Large language models (LLMs) are typically trained on shuffled corpora, yielding models whose knowledge is frozen at train time and whose temporal groun

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation

Model ReleasesDGX agent

arXiv:2605.21611v1 Announce Type: new Abstract: We introduce spatially grounded contextual image generation, a controllable image generation task that reframes the conditioning paradigm. Instead of su

VDE Bench: Evaluating The Capability of Image Editing Models to Modify Visual Documents

Model ReleasesDGX agent

arXiv:2602.00122v2 Announce Type: replace Abstract: In recent years, image editing models have made significant progress, enabling users to manipulate visual content in a flexible and interactive mann

← Previous
1…221222223224225…377
Next →