AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,620 results
29 May 2026

AI can give researchers the freedom to pursue “crazier” ideas. For Terence Tao, AI creates more room to experiment, test unexpected paths, a…

Model ReleasesDGX agent

AI tools are enabling researchers, including renowned mathematician Terence Tao, to explore unconventional and high-risk ideas by handling routine computational tasks and verification work. This techn

AI startup Shift launches a free home cleaning service in NYC to record first-person video with a camera-equipped cap and use it to train robots (Robert Hart/The Verge)

Model ReleasesDGX agent

Robert Hart / The Verge: AI startup Shift launches a free home cleaning service in NYC to record first-person video with a camera-equipped cap and use it to train robots — Shift says a ‘magic hat’ wi

Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2605.29396v1 Announce Type: new Abstract: Safety alignment for large language models (LLMs) aims to reduce harmful or unsafe behavior while preserving general utility. However, recent findings r

Alignment-Guided Score Matching for Text-to-Image Alignment in Diffusion Models

Model ReleasesDGX agent

arXiv:2605.30038v1 Announce Type: cross Abstract: Diffusion models generate highly realistic images but often struggle with precise text-image alignment. While recent post-training methods improve ali

AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation

Model ReleasesDGX agent

arXiv:2512.01334v2 Announce Type: replace Abstract: Text-guided image-to-video generation has made substantial progress, yet it still struggles to execute text-specified edits that require substantial

AlloyDB Hot Standby: Faster failovers, consistent performance

Model ReleasesDGX agent

AlloyDB for PostgreSQL is a fully managed, PostgreSQL-compatible database service designed for the most demanding enterprise workloads. It combines the best of PostgreSQL with the power of Google, del

AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training

Model ReleasesDGX agent

arXiv:2605.29664v1 Announce Type: cross Abstract: Pipeline parallelism is essential for large-scale model training, but existing asynchronous approaches often degrade convergence due to parameter mism

An End-to-End PyTorch Interface for Differentiable PDE Solvers: A RANS Model-Correction Study

Model ReleasesDGX agent

arXiv:2605.28858v1 Announce Type: cross Abstract: This work presents an end-to-end strategy for solving inverse problems constrained by Partial Differential Equations within a fully differentiable Mac

Another proof point for the open-weights thesis. From @RampLabs: 'If we built this again, we'd lean more on open-weight models.' Ramp pointe…

Model ReleasesDGX agent

Another proof point for the open-weights thesis. From @RampLabs: 'If we built this again, we'd lean more on open-weight models.' Ramp pointed 10K agents at their own backend. Kimi K2.6 and DeepSeek V4

Anthropic's run-rate revenue hits $47 billion

Model ReleasesDGX agent

The most interesting thing about Anthropic's 65B Series H announcement is this line (emphasis mine): Since our Series G in February, adoption has continued to grow across global enterprise customers,

Apertus LLM Family Expansion via Distillation and Quantization

Model ReleasesDGX agent

arXiv:2605.29128v1 Announce Type: new Abstract: The wide adoption of LLMs has led to their use in great variety of applications and scenarios, such as chatbot assistants and data annotation, creating

Architecture-Sensitive Supervised Fine-Tuning for Screen-Conditioned Action Prediction: A PiSAR Benchmark

Model ReleasesDGX agent

arXiv:2605.29400v1 Announce Type: new Abstract: We benchmark three supervised fine-tuned models against frontier zero-shot baselines on a 661-row held-out slice of PiSAR (Persona, intent, Screen, Acti

Are LLMs Socially Adaptive? Contrasting Belief Evolution in Large Language Models and Humans

Model ReleasesDGX agent

arXiv:2410.10398v3 Announce Type: replace-cross Abstract: As large language models (LLMs) increasingly engage in complex social interactions, ensuring that their behaviors align with human ethical pri

AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Crystalline Materials

Model ReleasesDGX agent

arXiv:2510.04704v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown promising potential in scientific research, enabling tasks ranging from knowledge retrieval to propert

AttuneBench: A Conversation-Based Benchmark for LLM Emotional Intelligence

Model ReleasesDGX agent

arXiv:2605.21739v2 Announce Type: replace Abstract: Emotional intelligence (EI), the ability to perceive, understand, and respond appropriately to others' emotional states, is central to human communi

Audio Deepfake Detection with Half-Truth Localisation Using Cross-Attentive Feature Fusion

Model ReleasesDGX agent

arXiv:2605.29531v1 Announce Type: cross Abstract: Audio deepfake detection is well-studied as a binary problem, but partially manipulated speech, where a short synthesised segment is spliced into an o

Auditing Training-Free 3D Shape Retrieval with Diffused Geodesic Moments

Model ReleasesDGX agent

arXiv:2605.29004v1 Announce Type: new Abstract: Reported retrieval scores for training-free shape descriptors conflate local signal design, normalization, aggregation, codebook fitting, and metric cho

AutoSizer: Automatic Sizing of Analog and Mixed-Signal Circuits via Large Language Model (LLM) Agents

Model ReleasesDGX agent

arXiv:2602.02849v2 Announce Type: replace Abstract: The design of Analog and Mixed-Signal (AMS) integrated circuits remains heavily reliant on expert knowledge, with transistor sizing a major bottlene

Balancing Multimodal Learning through Label Space Reshaping

Model ReleasesDGX agent

arXiv:2605.28869v1 Announce Type: cross Abstract: Multimodal learning often suffers from modality imbalance, where modalities that converge faster dominate optimization while others remain undertraine

Bandit Algorithms for Deep Brain Stimulation

Model ReleasesDGX agent

arXiv:2601.12699v2 Announce Type: replace Abstract: Deep Brain Stimulation (DBS) is an effective treatment for Parkinson's disease, but conventional fixed-parameter stimulation can reduce battery life

Battery-Sim-Agent: Leveraging LLM-Agent for Inverse Battery Parameter Estimation

Model ReleasesDGX agent

arXiv:2605.29560v1 Announce Type: new Abstract: Parameterizing high-fidelity 'digital twins' of batteries is a critical yet challenging inverse problem that hinders the pace of battery innovation. Pre

'Be My Cheese?': Cultural Nuance Benchmarking for Machine Translation in Multilingual LLMs

Model ReleasesDGX agent

arXiv:2602.04729v2 Announce Type: replace Abstract: We present a large-scale human evaluation benchmark for assessing cultural localisation in machine translation produced by state-of-the-art multilin

Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset

Model ReleasesDGX agent

arXiv:2605.29462v1 Announce Type: cross Abstract: The emergence of Large Vision-Language Models (LVLMs) has substantially expanded model capabilities beyond text-only understanding, enabling unified i

Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting

Model ReleasesDGX agent

arXiv:2509.23571v3 Announce Type: replace-cross Abstract: As cyber threats continue to grow in scale and sophistication, blue team defenders increasingly require advanced tools to proactively detect a

Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation

Model ReleasesDGX agent

arXiv:2605.28830v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are increasingly deployed in safety-critical applications, robust content moderation becomes essential. We present a c

Benchmarking Positional Encoding Strategies for Transformer-Based EEG Foundation Models

Model ReleasesDGX agent

arXiv:2605.29754v1 Announce Type: new Abstract: Electroencephalography (EEG) is a widely used non-invasive technique for measuring brain activity in brain-computer interface (BCI) applications. Superv

Benchmarking Single-Factor Physical Video-to-Audio Generation

Model ReleasesDGX agent

arXiv:2605.30339v1 Announce Type: new Abstract: Generative video-to-audio (V2A) models produce highly plausible soundtracks, but it remains unclear whether they capture the underlying physical process

BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents

Model ReleasesDGX agent

arXiv:2605.29225v1 Announce Type: new Abstract: Self-evolving agents improve over time by reflecting on past failures, but existing evaluation is limited in two ways: it measures only task scores, lea

Beyond English and Evasion: A Human-Annotated Multi-Domain Benchmark for High-Stakes LLM Safety Evaluation in Chinese

Model ReleasesDGX agent

arXiv:2605.29667v1 Announce Type: new Abstract: When Large Language Models (LLMs) are deployed in Chinese-language settings, a troubling pattern emerges: safety systems that work well in English break

Beyond Recall: Behavioral Specification as an Interpretive Layer for AI Personalization

Model ReleasesDGX agent

arXiv:2605.28969v1 Announce Type: cross Abstract: If an AI agent makes decisions on a person's behalf, those decisions must align with its user. We introduce representational accuracy to measure how f

BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders

Model ReleasesDGX agent

arXiv:2605.30162v1 Announce Type: new Abstract: Biosecurity evaluations of language models typically ask whether models produce hazardous output. This paper asks a complementary question: when a model

Boston Children’s uses AI to unlock new diagnoses

Model ReleasesDGX agent

Boston Children's Hospital has implemented AI technology to improve diagnostic accuracy and identify rare or complex medical conditions in pediatric patients that might otherwise go undiagnosed. The a

BrahmicTokenizer-131K: An Indic-Capable Drop-In Replacement for o200k_base

Model ReleasesDGX agent

arXiv:2605.29379v1 Announce Type: new Abstract: We present BrahmicTokenizer-131K, a 131,072-vocabulary byte-level BPE tokenizer that closes the Brahmic compression gap at the 131K-vocabulary class whi

Brain-IT-VQA: From Brain Signals to Answers

Model ReleasesDGX agent

arXiv:2605.29588v1 Announce Type: cross Abstract: Decoding visual content from fMRI signals recorded while a person views images, and specifically answering questions about the seen images, is a long-

Bridging the Semantic Gap for Categorical Data Clustering via Large Language Models

Model ReleasesDGX agent

arXiv:2601.01162v3 Announce Type: replace-cross Abstract: Qualitative data are widespread in domains such as healthcare, marketing, and bioinformatics, where clustering offers a fundamental tool for p

Building and Road Recognition in Dense Urban Informal Settlements: A Dataset and Benchmark

Model ReleasesDGX agent

arXiv:2605.29856v1 Announce Type: new Abstract: As a widespread form of informal settlements, urban villages present significant challenges for sustainable urban development and governance. Precise ma

BullingerDB: A Dataset for Handwritten Text Recognition and Writer Retrieval

Model ReleasesDGX agent

arXiv:2605.30235v1 Announce Type: new Abstract: We present BullingerDB, a large-scale benchmark dataset for historical document analysis based on the correspondence of Heinrich Bullinger (1504-1575).

CalArena: A Large-Scale Post-Hoc Calibration Benchmark

Model ReleasesDGX agent

arXiv:2605.30188v1 Announce Type: cross Abstract: Reliable probability estimates are critical in many machine learning applications, yet modern classifiers are often poorly calibrated. Post-hoc calibr

Can AI Weather Models Predict Beyond Two Weeks? A Quantitative Benchmark and Analysis of Long Rollouts

Model ReleasesDGX agent

arXiv:2605.30184v1 Announce Type: new Abstract: While AI weather models excel at short-to-medium range forecasts (up to 15 days), they frequently suffer from ill-defined 'instabilities' when rolled ou

Casual as an Anchor: Resolving Supervision Misalignment in Formality Transfer Dataset

Model ReleasesDGX agent

arXiv:2605.29365v1 Announce Type: new Abstract: Formality transfer is commonly framed as a symmetric bidirectional task between informal and formal registers. We argue that this framing conceals a sup

Certified Causal Defense with Generalizable Robustness

Model ReleasesDGX agent

arXiv:2408.15451v3 Announce Type: replace Abstract: While machine learning models have proven effective across various scenarios, it is widely acknowledged that many models are vulnerable to adversari

ChatGPT diagnosed 40 million people with a disease that was invented as a joke. Not a real disease. Not a misunderstood disease. A completel…

Model ReleasesDGX agent

ChatGPT diagnosed 40 million people with a disease that was invented as a joke. Not a real disease. Not a misunderstood disease. A completely fictional condition with a fake name, fake papers, and fak

Chess-World-Model: A 10M-Game Benchmark for Exact State Tracking from Chess Move Sequences

Model ReleasesDGX agent

arXiv:2605.30100v1 Announce Type: new Abstract: World models require state tracking, which is the ability to maintain a correct latent state across action sequences. Existing benchmarks are often synt

Citation-Closure Retrieval and Per-Rule Attribution for Real-World Regulatory Compliance Question Answering

Model ReleasesDGX agent

arXiv:2605.29742v1 Announce Type: new Abstract: Deploying Large Language Models (LLMs) for regulatory compliance demands rigorous traceability via comprehensive citations across multi-tiered authority

CityGen: Structure-Guided City-Style Synthesis for Cross-City Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.29935v1 Announce Type: cross Abstract: Autonomous driving systems are commonly trained and evaluated within limited geographic regions, which hinders their scalability when deployed in new

Claude really can roleplay an economist. I love this little comment Claude made after some robustness checks on the paper it wrote: 'On a 1–…

Model ReleasesDGX agent

Claude really can roleplay an economist. I love this little comment Claude made after some robustness checks on the paper it wrote: 'On a 1–10 identification scale, I'd now put the paper at about 4.5

Cloud CISO Perspectives: How to build an AI-ready security program for the public sector

Model ReleasesDGX agent

Welcome to the second Cloud CISO Perspectives for May 2026. Today, Usman Chaudhary, Field CISO, Google Public Sector, offers a guide for CISOs protecting government agencies and critical infrastructur

CLUBench: A Clustering Benchmark

Model ReleasesDGX agent

arXiv:2605.29933v1 Announce Type: new Abstract: Clustering is a fundamental problem in data science with a long-standing research history, yielding numerous insightful algorithms. Despite this progres

CodeEvolve: an open source evolutionary coding agent for algorithmic discovery and optimization

Model ReleasesDGX agent

arXiv:2510.14150v5 Announce Type: replace Abstract: We introduce CodeEvolve, an open-source framework that couples large language models with island-based evolutionary search for end-to-end algorithmi

Combating Data Laundering in LLM Training

Model ReleasesDGX agent

arXiv:2604.01904v2 Announce Type: replace-cross Abstract: Data rights owners can detect unauthorized data use in large language model (LLM) training by querying with proprietary samples. Often, superi

Command A+ sets a new high for Cohere's machine translation capabilities. Opening a clear gap over open source peers Mistral Medium 3.5, Dee…

Model ReleasesDGX agent

Command A+ sets a new high for Cohere's machine translation capabilities. Opening a clear gap over open source peers Mistral Medium 3.5, DeepSeek, & OpenAI's gpt-oss, as well as Claude Opus 4.6. A+ al

CommunityFact: A Dynamic, Multilingual, Multi-domain Benchmark for Misinformation Detection in the Wild

Model ReleasesDGX agent

arXiv:2605.30241v1 Announce Type: new Abstract: Misinformation verification increasingly occurs in public, fast-moving, and multilingual online settings, where static benchmarks provide an incomplete

Comparative Evaluation of Machine Translation Systems on Images with Text

Model ReleasesDGX agent

arXiv:2605.29476v1 Announce Type: new Abstract: This work presents a comparative evaluation of machine translation systems applied to images containing textual information, a task that lies at the int

COMPOSE: Composing Future Theorems from Citations and Formal Structure

Model ReleasesDGX agent

arXiv:2605.30333v1 Announce Type: new Abstract: A plausible future mathematical claim must satisfy two constraints: it should follow the direction of prior work and respect the formal dependencies tha

Composing Non-Conjugate Factor Graphs with Closed-Form Variational Inference

Model ReleasesDGX agent

arXiv:2605.29467v1 Announce Type: cross Abstract: Stacking probabilistic building blocks into deeper architectures typically breaks closed-form inference. We show that closed-form inference can be pre

ConMoE: Expert-Pool Consolidation via Prototype Reassignment for MoE Compression

Model ReleasesDGX agent

arXiv:2605.29350v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) language models reduce per-token computation but still require storing and serving all experts, making deployment memory-intens

Connecting Independently Trained Modes via Layer-Wise Connectivity

Model ReleasesDGX agent

arXiv:2505.02604v5 Announce Type: replace Abstract: Empirical studies have shown that continuous low-loss paths can be constructed between independently trained neural network models. This phenomenon,

Contrastive Representation Regularization for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2510.01711v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have shown strong capabilities in robot manipulation by leveraging rich representations from pre-trained V

Converted, Not Equivalent: Benchmarking Codebase Conversion via Observational Equivalence

Model ReleasesDGX agent

arXiv:2605.29054v1 Announce Type: cross Abstract: Coding agents increasingly act as codebase-scale collaborators that can assist with codebase conversion, but this progress has exposed a critical weak

Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation

Model ReleasesDGX agent

arXiv:2605.30000v1 Announce Type: new Abstract: Front-end web code has become a core product surface for every frontier LLM release, yet evaluating these interactive applications at development speed

← Previous
1…192193194195196…377
Next →