AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
All
85,136Total entries
1Added by human
85,135Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,770 results
Model Releases

BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders

DGX agent

arXiv:2605.30162v1 Announce Type: new Abstract: Biosecurity evaluations of language models typically ask whether models produce hazardous output. This paper asks a complementary question: when a model

model-releasesarxiv-cs-ai
29 May 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Boston Children’s uses AI to unlock new diagnoses

DGX agent

Boston Children's Hospital has implemented AI technology to improve diagnostic accuracy and identify rare or complex medical conditions in pediatric patients that might otherwise go undiagnosed. The a

model-releasesopenai
29 May 2026
Model Releases

BrahmicTokenizer-131K: An Indic-Capable Drop-In Replacement for o200k_base

DGX agent

arXiv:2605.29379v1 Announce Type: new Abstract: We present BrahmicTokenizer-131K, a 131,072-vocabulary byte-level BPE tokenizer that closes the Brahmic compression gap at the 131K-vocabulary class whi

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Brain-IT-VQA: From Brain Signals to Answers

DGX agent

arXiv:2605.29588v1 Announce Type: cross Abstract: Decoding visual content from fMRI signals recorded while a person views images, and specifically answering questions about the seen images, is a long-

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Bridging the Semantic Gap for Categorical Data Clustering via Large Language Models

DGX agent

arXiv:2601.01162v3 Announce Type: replace-cross Abstract: Qualitative data are widespread in domains such as healthcare, marketing, and bioinformatics, where clustering offers a fundamental tool for p

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Building and Road Recognition in Dense Urban Informal Settlements: A Dataset and Benchmark

DGX agent

arXiv:2605.29856v1 Announce Type: new Abstract: As a widespread form of informal settlements, urban villages present significant challenges for sustainable urban development and governance. Precise ma

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

BullingerDB: A Dataset for Handwritten Text Recognition and Writer Retrieval

DGX agent

arXiv:2605.30235v1 Announce Type: new Abstract: We present BullingerDB, a large-scale benchmark dataset for historical document analysis based on the correspondence of Heinrich Bullinger (1504-1575).

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

CalArena: A Large-Scale Post-Hoc Calibration Benchmark

DGX agent

arXiv:2605.30188v1 Announce Type: cross Abstract: Reliable probability estimates are critical in many machine learning applications, yet modern classifiers are often poorly calibrated. Post-hoc calibr

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Can AI Weather Models Predict Beyond Two Weeks? A Quantitative Benchmark and Analysis of Long Rollouts

DGX agent

arXiv:2605.30184v1 Announce Type: new Abstract: While AI weather models excel at short-to-medium range forecasts (up to 15 days), they frequently suffer from ill-defined 'instabilities' when rolled ou

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Casual as an Anchor: Resolving Supervision Misalignment in Formality Transfer Dataset

DGX agent

arXiv:2605.29365v1 Announce Type: new Abstract: Formality transfer is commonly framed as a symmetric bidirectional task between informal and formal registers. We argue that this framing conceals a sup

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Certified Causal Defense with Generalizable Robustness

DGX agent

arXiv:2408.15451v3 Announce Type: replace Abstract: While machine learning models have proven effective across various scenarios, it is widely acknowledged that many models are vulnerable to adversari

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

ChatGPT diagnosed 40 million people with a disease that was invented as a joke. Not a real disease. Not a misunderstood disease. A completel…

DGX agent

ChatGPT diagnosed 40 million people with a disease that was invented as a joke. Not a real disease. Not a misunderstood disease. A completely fictional condition with a fake name, fake papers, and fak

model-releasesgary-marcus--x
29 May 2026
Model Releases

Chess-World-Model: A 10M-Game Benchmark for Exact State Tracking from Chess Move Sequences

DGX agent

arXiv:2605.30100v1 Announce Type: new Abstract: World models require state tracking, which is the ability to maintain a correct latent state across action sequences. Existing benchmarks are often synt

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Citation-Closure Retrieval and Per-Rule Attribution for Real-World Regulatory Compliance Question Answering

DGX agent

arXiv:2605.29742v1 Announce Type: new Abstract: Deploying Large Language Models (LLMs) for regulatory compliance demands rigorous traceability via comprehensive citations across multi-tiered authority

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

CityGen: Structure-Guided City-Style Synthesis for Cross-City Autonomous Driving

DGX agent

arXiv:2605.29935v1 Announce Type: cross Abstract: Autonomous driving systems are commonly trained and evaluated within limited geographic regions, which hinders their scalability when deployed in new

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Claude really can roleplay an economist. I love this little comment Claude made after some robustness checks on the paper it wrote: 'On a 1–…

DGX agent

Claude really can roleplay an economist. I love this little comment Claude made after some robustness checks on the paper it wrote: 'On a 1–10 identification scale, I'd now put the paper at about 4.5

model-releasesethan-mollick--x
29 May 2026
Model Releases

Cloud CISO Perspectives: How to build an AI-ready security program for the public sector

DGX agent

Welcome to the second Cloud CISO Perspectives for May 2026. Today, Usman Chaudhary, Field CISO, Google Public Sector, offers a guide for CISOs protecting government agencies and critical infrastructur

model-releasesgoogle-cloud-ai
29 May 2026
Model Releases

CLUBench: A Clustering Benchmark

DGX agent

arXiv:2605.29933v1 Announce Type: new Abstract: Clustering is a fundamental problem in data science with a long-standing research history, yielding numerous insightful algorithms. Despite this progres

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

CodeEvolve: an open source evolutionary coding agent for algorithmic discovery and optimization

DGX agent

arXiv:2510.14150v5 Announce Type: replace Abstract: We introduce CodeEvolve, an open-source framework that couples large language models with island-based evolutionary search for end-to-end algorithmi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Combating Data Laundering in LLM Training

DGX agent

arXiv:2604.01904v2 Announce Type: replace-cross Abstract: Data rights owners can detect unauthorized data use in large language model (LLM) training by querying with proprietary samples. Often, superi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Command A+ sets a new high for Cohere's machine translation capabilities. Opening a clear gap over open source peers Mistral Medium 3.5, Dee…

DGX agent

Command A+ sets a new high for Cohere's machine translation capabilities. Opening a clear gap over open source peers Mistral Medium 3.5, DeepSeek, & OpenAI's gpt-oss, as well as Claude Opus 4.6. A+ al

model-releasescohere--x
29 May 2026
Model Releases

CommunityFact: A Dynamic, Multilingual, Multi-domain Benchmark for Misinformation Detection in the Wild

DGX agent

arXiv:2605.30241v1 Announce Type: new Abstract: Misinformation verification increasingly occurs in public, fast-moving, and multilingual online settings, where static benchmarks provide an incomplete

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Comparative Evaluation of Machine Translation Systems on Images with Text

DGX agent

arXiv:2605.29476v1 Announce Type: new Abstract: This work presents a comparative evaluation of machine translation systems applied to images containing textual information, a task that lies at the int

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

COMPOSE: Composing Future Theorems from Citations and Formal Structure

DGX agent

arXiv:2605.30333v1 Announce Type: new Abstract: A plausible future mathematical claim must satisfy two constraints: it should follow the direction of prior work and respect the formal dependencies tha

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Composing Non-Conjugate Factor Graphs with Closed-Form Variational Inference

DGX agent

arXiv:2605.29467v1 Announce Type: cross Abstract: Stacking probabilistic building blocks into deeper architectures typically breaks closed-form inference. We show that closed-form inference can be pre

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

ConMoE: Expert-Pool Consolidation via Prototype Reassignment for MoE Compression

DGX agent

arXiv:2605.29350v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) language models reduce per-token computation but still require storing and serving all experts, making deployment memory-intens

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Connecting Independently Trained Modes via Layer-Wise Connectivity

DGX agent

arXiv:2505.02604v5 Announce Type: replace Abstract: Empirical studies have shown that continuous low-loss paths can be constructed between independently trained neural network models. This phenomenon,

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Contrastive Representation Regularization for Vision-Language-Action Models

DGX agent

arXiv:2510.01711v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have shown strong capabilities in robot manipulation by leveraging rich representations from pre-trained V

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Converted, Not Equivalent: Benchmarking Codebase Conversion via Observational Equivalence

DGX agent

arXiv:2605.29054v1 Announce Type: cross Abstract: Coding agents increasingly act as codebase-scale collaborators that can assist with codebase conversion, but this progress has exposed a critical weak

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation

DGX agent

arXiv:2605.30000v1 Announce Type: new Abstract: Front-end web code has become a core product surface for every frontier LLM release, yet evaluating these interactive applications at development speed

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

CosmicFish-HRM: Adaptive Reasoning via Hierarchical Recurrent Mechanisms in Compact Language Models

DGX agent

arXiv:2605.28919v1 Announce Type: cross Abstract: Large language models have achieved strong reasoning capabilities, though often at the cost of massive parameter counts and expensive inference. In th

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective

DGX agent

arXiv:2502.03805v2 Announce Type: replace Abstract: Large language models have revolutionized natural language processing but face significant challenges of high storage and runtime costs, due to the

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

CrystalXRD-Bench: Benchmarking Vision-Language Models for XRD Peak Indexing Across Diverse Crystalline Materials

DGX agent

arXiv:2605.29446v1 Announce Type: new Abstract: Miller-index identification from powder XRD patterns requires capabilities untested by existing multimodal benchmarks: the model must read a narrow peak

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Deep Adaptive Dimension Reduction for Bayesian Inference in Inverse Problems

DGX agent

arXiv:2605.29373v1 Announce Type: new Abstract: Solving high-dimensional PDE-governed inverse problems is often challenging due to complex non-Gaussian posterior distributions, expensive forward model

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Demystifying Scientific Problem-Solving in LLMs by Probing Knowledge and Reasoning

DGX agent

arXiv:2508.19202v3 Announce Type: replace Abstract: Scientific problem solving poses unique challenges for LLMs, requiring both deep domain knowledge and the ability to apply such knowledge through co

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

DenseSteer: Steering Small Language Models towards Dense Math Reasoning

DGX agent

arXiv:2605.29247v1 Announce Type: new Abstract: Large language models (LLMs) demonstrate strong chain-of-thought (CoT) reasoning abilities, while smaller models (<= 3B parameters) significantly underp

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Developer's guide to Gemini Enterprise and A2UI integration

DGX agent

If you've built a chatbot, you know this conversation: User: 'Book a table for two tomorrow at 7pm.' Agent: 'Okay, for what day?' User: 'Tomorrow.' Agent: 'What time?' A date picker would have ended t

model-releasesgoogle-cloud-ai
29 May 2026
Model Releases

Dial HEALTHDIAL for Advice: A Multilingual and Multi-Parallel Spoken Dialogue Dataset for Knowledge-Grounded Information Seeking

DGX agent

arXiv:2605.30107v1 Announce Type: new Abstract: Creating spoken dialogue datasets is methodologically challenging, and these challenges are amplified when the goal is to build multilingual, multi-para

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Differentiable Belief-based Opponent Shaping

DGX agent

arXiv:2605.29042v1 Announce Type: new Abstract: Human coordination often relies on the ability to influence the beliefs of others through strategic action. In multi-agent reinforcement learning, oppon

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

DiffSpot: Can VLMs Spot Fine-Grained Visual Differences in Web Interfaces?

DGX agent

arXiv:2605.29615v1 Announce Type: cross Abstract: Vision-language models (VLMs) have made strong progress on high-level image-text alignment, yet their ability to perceive subtle visual differences re

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Diffusion-based learning framework for Constrained Nonconvex Optimization with Weighted Bootstrapped Refinement

DGX agent

arXiv:2502.10330v4 Announce Type: replace Abstract: Recent advances in diffusion models show promising potential to accelerate nonconvex problem solving by leveraging their multimodality. However, mos

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Diffusion differentiable resampling

DGX agent

arXiv:2512.10401v3 Announce Type: replace-cross Abstract: This paper is concerned with differentiable resampling in the context of sequential Monte Carlo (e.g., particle filtering). Drawing on reparam

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

DirectorBench: Diagnosing Long-Form Video Generation with Personalized Multi-Agent Evaluation

DGX agent

arXiv:2605.30090v1 Announce Type: new Abstract: Long-form video generation is rapidly moving from short, single-scene synthesis toward minute-long, multi-shot creation with narrative structure, cinema

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Dissecting the Black Box: Circuit-Level Analysis of LLM Vulnerability Detection

DGX agent

arXiv:2605.29901v1 Announce Type: cross Abstract: Large language models (LLMs) can detect software vulnerabilities, but how do they actually identify vulnerable code? We address this question using me

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

DMC-CF: Dynamic Multimodal CounterFactual QA benchmark for Causal Reasoning

DGX agent

arXiv:2605.29339v1 Announce Type: new Abstract: With the rapid advancement of multimodal large language models (MLLMs), models have demonstrated increasingly powerful multimodal capabilities. However,

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

Do Physics Foundation Models Learn Generalizable Physics? A Bias-Aware Benchmark Across Physical Regimes and Distribution Shifts

DGX agent

arXiv:2605.29283v1 Announce Type: cross Abstract: Recent physics foundation models claim general spatiotemporal forecasting ability, yet their evaluations often collapse performance into a single aver

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark

DGX agent

arXiv:2605.30027v1 Announce Type: new Abstract: Multimodal documents contain diverse elements, such as tables, figures, and layouts, which can complicate retrieval tasks. While current approaches typi

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

Document Parsing + Gemini 🔥 Excited to collaborate with the Google team on this, here's to many more!

DGX agent

Document Parsing + Gemini 🔥 Excited to collaborate with the Google team on this, here's to many more! The team at @llama_index built an awesome template using LlamaParse and the new Managed Agents in

model-releasesjerry-liu--x
29 May 2026
← Previous
1…243244245246247…475
Next →