AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
All
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,612 results
Model Releases

ProcBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents

DGX agent

arXiv:2605.20251v2 Announce Type: cross Abstract: Existing benchmarks for LLM coding agents primarily evaluate final outcomes. While useful for measuring overall capability, these metrics provide limi

model-releasesarxiv-cs-ai
22 May 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

PromptNCE: Pointwise Mutual Information Predictions Using Only LLMs and Contrastive Estimation Prompts

DGX agent

arXiv:2605.21776v1 Announce Type: new Abstract: Estimating mutual information from text usually requires training a task-specific critic, which limits its use in low-data settings. We ask whether larg

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Putnam 2025 Problems in Rocq using Opus 4.6 and Rocq-MCP

DGX agent

arXiv:2603.20405v2 Announce Type: replace-cross Abstract: We report on an experiment in which Claude Opus~4.6, equipped with a suite of Model Context Protocol (MCP) tools for the Rocq proof assistant,

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems

DGX agent

arXiv:2510.13910v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) mitigates key limitations of Large Language Models (LLMs)-such as factual errors, outdated knowledge, and hallu

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator

DGX agent

arXiv:2605.21748v1 Announce Type: new Abstract: As interactive LLM-based applications are created and refined, model developers need to evaluate the quality of generated text along many possible axes.

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents

DGX agent

arXiv:2605.22148v1 Announce Type: cross Abstract: Self-evolving skill libraries, pioneered by Voyager, let frozen LLM agents accumulate reusable knowledge without weight updates, yet recent evaluation

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation

DGX agent

arXiv:2605.20204v1 Announce Type: cross Abstract: LLM-based user simulation is the primary mechanism for end-to-end agent evaluation, yet simulated users are poor proxies for real humans: unconstraine

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

Reddit stock fell 5%+ on Friday after Meta launched a standalone app for online forums, called Forum; Reddit's stock is now down almost 40% this year (Jonathan Vanian/CNBC)

DGX agent

Jonathan Vanian / CNBC: Reddit stock fell 5%+ on Friday after Meta launched a standalone app for online forums, called Forum; Reddit's stock is now down almost 40% this year — Reddit shares fell about

model-releasestechmeme
22 May 2026
Model Releases

Reflective Prompt Tuning through Language Model Function-Calling

DGX agent

arXiv:2605.21781v1 Announce Type: new Abstract: Large language models (LLMs) have become increasingly capable of following instructions and complex reasoning, making prompting a flexible interface for

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Residual Skill Optimization for Text-to-SQL Ensembles

DGX agent

arXiv:2605.21792v1 Announce Type: new Abstract: Text-to-SQL ensembles improve over single-candidate generation by drawing multiple SQL candidates and selecting one, but their effectiveness is bounded

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Rethinking Noise-Robust Training for Frozen Vision Foundation Models: A Cross-Dataset Benchmark with a Case Study of Small-Loss Failure

DGX agent

arXiv:2605.22591v1 Announce Type: new Abstract: Frozen Vision Foundation Models (VFMs) with lightweight classification heads are increasingly used in medical imaging because they offer efficient and r

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

SADGE: Structure and Appearance Domain Gap Estimation of Synthetic and Real Data

DGX agent

arXiv:2605.22467v1 Announce Type: new Abstract: We propose SADGE, a quantitative similarity metric that predicts the performance of synthetic image datasets for common computer vision tasks without do

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching

DGX agent

arXiv:2605.21788v1 Announce Type: new Abstract: Zero-shot 3D visual grounding requires localizing objects in unstructured environments from free-form natural language. Recent vision-language model (VL

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals

DGX agent

arXiv:2605.21919v1 Announce Type: new Abstract: Assessing progress toward the Sustainable Development Goals (SDGs) requires multi-step reasoning over visual cues, contextual knowledge, and development

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Security Document Classification with a Fine-Tuned Local Large Language Model: Benchmark Data and an Open-Source System

DGX agent

arXiv:2605.20368v1 Announce Type: cross Abstract: Organizations that scan documents for sensitive information face a practical problem. Cloud services require data to be sent to external infrastructur

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

Seeing the Poem: Image-Semantic Detection of AI-Generated Modern Chinese Poetry with MLLMs

DGX agent

arXiv:2605.22654v1 Announce Type: new Abstract: Previous detection studies have shown that LLMs cannot be effectively used as detectors, but these studies have not addressed modern Chinese poetry. Mor

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

SegGuidedNet: Sub-Region-Aware Attention Supervision for Interpretable Brain Tumor Segmentation

DGX agent

arXiv:2605.22572v1 Announce Type: new Abstract: Accurate segmentation of brain tumour sub-regions from multi-parametric MRI is critical for treatment planning yet remains challenging due to morphologi

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Seizure-Semiology-Suite (S3): A Clinically Multimodal Dataset, Benchmark, and Models for Seizure Semiology Understanding

DGX agent

arXiv:2605.21852v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in general video understanding, their capacity to interpret invo

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

SFN-YOLO: Towards Free-Range Poultry Detection via Scale-aware Fusion Networks

DGX agent

arXiv:2509.17086v2 Announce Type: replace Abstract: Detecting and localizing poultry is essential for advancing smart poultry farming. Despite the progress of detection-centric methods, challenges per

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

SiameseNorm: Breaking the Barrier to Reconciling Pre/Post-Norm

DGX agent

arXiv:2602.08064v2 Announce Type: replace-cross Abstract: The long-standing tension between Pre- and Post-Norm remains an open problem in Transformer architecture, reflecting a fundamental trade-off b

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control

DGX agent

arXiv:2511.07820v3 Announce Type: replace-cross Abstract: Despite the rise of billion-parameter foundation models trained across thousands of GPUs, similar scaling gains have not been shown for humano

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation

DGX agent

arXiv:2605.22536v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have made rapid progress in spatial intelligence, yet existing spatial reasoning benchmarks largely assume pr

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving

DGX agent

arXiv:2512.10719v2 Announce Type: replace Abstract: End-to-end autonomous driving methods built on vision language models (VLMs) have undergone rapid development driven by their universal visual under

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

SpaceX now launches more rockets than every other country combined. This has never happened before in the history of spaceflight.

DGX agent

SpaceX has achieved a milestone where it launches more rockets than all other countries combined, marking an unprecedented moment in spaceflight history. This represents a significant shift in the com

model-releaseselon-musk--x
22 May 2026
Model Releases

Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action

DGX agent

arXiv:2605.22283v1 Announce Type: new Abstract: We introduce SOMA, the Spatial Memory framework for Out-of-Vision Manipulation in Vision-Language-Action (VLA) models. Most existing VLAs implicitly ass

model-releasesarxiv-cs-ro
22 May 2026
Model Releases

Steins;Gate Drive: Semantic Safety Arbitration over Structured Futures for Latency-Decoupled LLM Planning

DGX agent

arXiv:2605.22456v1 Announce Type: new Abstract: Cloud-hosted LLM driver agents provide useful semantic judgments, but their inference latency exceeds stepwise vehicle-control windows. Learned world mo

model-releasesarxiv-cs-ro
22 May 2026
Model Releases

Structural Anchor Pruning: Training-Free Multi-Vector Compression for Visual Document Retrieval

DGX agent

arXiv:2601.20107v2 Announce Type: replace-cross Abstract: Recent Vision-Language Models (e.g., ColPali) enable fine-grained Visual Document Retrieval (VDR) but incur prohibitive multi-vector index sto

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance

DGX agent

arXiv:2605.22202v1 Announce Type: new Abstract: In this paper, we show that high-performing embedding models organize their embedding spaces in a consistent way. We evaluate 25 contemporary embedding

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Sub-exponential Growth Dynamics in Complex Systems: A Piecewise Power-Law Model for the Diffusion of New Words and Names

DGX agent

arXiv:2511.04106v5 Announce Type: replace-cross Abstract: The diffusion of ideas and language in society has conventionally been described by S-shaped models, such as the logistic curve. However, the

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

SURGE: An Event-Centric Social Media Sentiment Time Series Benchmark with Interaction Structure

DGX agent

arXiv:2605.21198v1 Announce Type: cross Abstract: Public events on social media generate large volumes of discussion whose collective dynamics carry direct value for opinion forecasting and crisis res

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

Survive or Collapse: The Asymmetric Roles of Data Gating and Reward Grounding in Self-Play RL

DGX agent

arXiv:2605.22217v1 Announce Type: cross Abstract: Self-play reinforcement learning trains language models on their own generated tasks, co-evolving a proposer and solver without human labels. Recent s

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Tackle CSM in JPEG Steganalysis with Data Adaptation

DGX agent

arXiv:2605.21523v1 Announce Type: cross Abstract: Steganalysis models excel on benchmark datasets but struggle in the wild when analyzed images are produced by a processing pipeline unseen during trai

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Targeting Clause Type Distributions: a Picklock for Random Satisfiability Problems

DGX agent

arXiv:2605.20328v1 Announce Type: cross Abstract: Optimization problems such as the NP-complete 3-SAT provide an important benchmark for the difficult task of finding ground-states in strongly correla

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

Teaching AI Through Benchmark Construction: QuestBench as a Course-Based Practice for Accountable Knowledge Work

DGX agent

arXiv:2605.21413v2 Announce Type: new Abstract: As AI becomes part of everyday learning, many courses teach students to use it mainly as a productivity tool: how to prompt, search, summarize, write, c

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

Teaching Language Models to Forecast Research Success Through Comparative Idea Evaluation

DGX agent

arXiv:2605.21491v1 Announce Type: cross Abstract: As language models accelerate scientific research by automating hypothesis generation and implementation, a new bottleneck emerges: evaluating and fil

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Temporal Coding as a Substrate for Sensorimotor Object Inference: A Spiking Reinterpretation of Thousand Brains Architecture

DGX agent

arXiv:2605.22206v1 Announce Type: cross Abstract: The Thousand Brains Theory (TBT) and its open-source Monty framework model object recognition through sensorimotor inference -- identifying objects by

model-releasesarxiv-cs-ro
22 May 2026
Model Releases

The Blueprint: How Movix fills a gap in dental skills with specialized agentic AI

DGX agent

Welcome to The Blueprint, a regular feature where we highlight how Google Cloud customers are tackling unique and common challenges across industries using the latest AI and cloud technologies. We hop

model-releasesgoogle-cloud-ai
22 May 2026
Model Releases

The Download: coding’s future, the ‘Steroid Olympics,’ and AI-driven science

DGX agent

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Anthropic’s Code with Claude showed off coding’s future—whethe

model-releasesmit-tech-review
22 May 2026
Model Releases

Thermo-VL: Extending Vision-Language Models to Thermal Infrared Perception

DGX agent

arXiv:2605.21882v1 Announce Type: new Abstract: Vision-language models (VLMs) often fail under low illumination because their visual grounding is learned predominantly from RGB imagery, whereas therma

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Though I will continue to insist that the path we were on would have been much clearer if o3 had been called GPT-5 and GPT-5 had been called…

DGX agent

Ethan Mollick comments on OpenAI's naming convention for their AI models, specifically critiquing the decision to call the o3 model 'o3' rather than 'GPT-5,' suggesting the naming scheme creates confu

model-releasesethan-mollick--x
22 May 2026
Model Releases

Token-Level LLM Collaboration via FusionRoute

DGX agent

arXiv:2601.05106v4 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit strengths across diverse domains. However, achieving strong performance across these domains with a singl

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Tokenization with Split Trees

DGX agent

arXiv:2605.22705v1 Announce Type: new Abstract: We introduce Tokenization with Split Trees (ToaST), a subword tokenization method that directly optimizes compression under a new recursive inference pr

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Towards Clinically Interpretable Ophthalmic VQA via Spatially-Grounded Lesion Evidence

DGX agent

arXiv:2605.22414v1 Announce Type: new Abstract: Visual Question Answering (VQA) holds great promise for clinical support, particularly in ophthalmology, where retinal fundus photography is essential f

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Towards Selection of Large Multimodal Models as Engines for Burned-in Protected Health Information Detection in Medical Images

DGX agent

arXiv:2511.02014v2 Announce Type: replace Abstract: The detection of Protected Health Information (PHI) in medical imaging is critical for safeguarding patient privacy and ensuring compliance with reg

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Training-Trajectory-Aware Token Selection

DGX agent

arXiv:2601.10348v2 Announce Type: replace Abstract: Efficient distillation is a key pathway for converting expensive reasoning capability into deployable efficiency, yet in the frontier regime where t

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Transcription and Recognition of Italian Parliamentary Speeches Using Vision-Language Models

DGX agent

arXiv:2603.28103v2 Announce Type: replace-cross Abstract: Parliamentary proceedings represent a rich yet challenging resource for computational analysis, particularly when preserved only as scanned hi

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

TransitLM: A Large-Scale Dataset and Benchmark for Map-Free Transit Route Generation

DGX agent

arXiv:2605.22355v1 Announce Type: new Abstract: Public transit route planning traditionally depends on structured map infrastructure and complex routing engines, and no existing dataset supports train

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Transporting Task Vectors across Different Architectures without Training

DGX agent

arXiv:2602.12952v2 Announce Type: replace-cross Abstract: Adapting large pre-trained models to downstream tasks often produces task-specific parameter updates that are expensive to relearn for every m

model-releasesarxiv-cs-cv
22 May 2026
← Previous
1…277278279280281…472
Next →