AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,332 results
23 Apr 2026

That browser version of LiteParse is available here: https://simonw.github.io/liteparse/ Detailed notes on how I built it using Claude Code …

Model ReleasesDGX agent

That browser version of LiteParse is available here: https://simonw.github.io/liteparse/ Detailed notes on how I built it using Claude Code (running for an hour) on my blog: https://simonwillison.net/

The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language Models

Model ReleasesDGX agent

arXiv:2604.20225v1 Announce Type: new Abstract: Evaluating the multilingual and multicultural capabilities of Large Language Models (LLMs) is essential for their global utility. However, current bench

The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2603.29025v2 Announce Type: replace-cross Abstract: Large language models systematically fail when a salient surface cue conflicts with an unstated feasibility constraint. We study this through

The OpenHands Software Agent SDK: A Composable and Extensible Foundation for Production Agents

Model ReleasesDGX agent

arXiv:2511.03690v2 Announce Type: replace-cross Abstract: Agents are now used widely in the process of software development, but building production-ready software engineering agents is a complex task

The Ratchet Effect in Silico through Interaction-Driven Cumulative Intelligence in Large Language Models

Model ReleasesDGX agent

arXiv:2507.21166v2 Announce Type: replace-cross Abstract: Human intelligence scales through cumulative cultural evolution (CCE), a ratchet process in which innovations are retained against entropic dr

The Role and Relationship of Initialization and Densification in 3D Gaussian Splatting

Model ReleasesDGX agent

arXiv:2603.20714v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has become the method of choice for photo-realistic 3D reconstruction of scenes, due to being able to efficiently and a

ThermoQA: A Three-Tier Benchmark for Evaluating Thermodynamic Reasoning in Large Language Models

Model ReleasesDGX agent

arXiv:2604.19758v1 Announce Type: new Abstract: We present ThermoQA, a benchmark of 293 open-ended engineering thermodynamics problems in three tiers: property lookups (110 Q), component analysis (101

Things have been degrading super fast in Claude Code. I still use Claude Code, but my default is now Codex. I still prefer Opus models for c…

Model ReleasesDGX agent

Things have been degrading super fast in Claude Code. I still use Claude Code, but my default is now Codex. I still prefer Opus models for coding, and so I will try again with the fixes. I appreciate

Three shifts in the AI stack today: - OpenAI ships GPT-5.5 as a mid-cycle drop before its IPO - Hugging Face's ML Intern beats Claude Code a…

Model ReleasesDGX agent

Three shifts in the AI stack today: - OpenAI ships GPT-5.5 as a mid-cycle drop before its IPO - Hugging Face's ML Intern beats Claude Code and Codex on research - Pliny used Claude Opus 4.7 to jailbre

tldr: claude code changed some harness settings which degraded perf. these small harness tweaks can matter a lot! 1. default reasoning high …

Model ReleasesDGX agent

tldr: claude code changed some harness settings which degraded perf. these small harness tweaks can matter a lot! 1. default reasoning high -> medium 2. bug that accidentally evicted thinking blocks o

To Know is to Construct: Schema-Constrained Generation for Agent Memory

Model ReleasesDGX agent

arXiv:2604.20117v1 Announce Type: new Abstract: Constructivist epistemology argues that knowledge is actively constructed rather than passively copied. Despite the generative nature of Large Language

Tokenised Flow Matching for Hierarchical Simulation Based Inference

Model ReleasesDGX agent

arXiv:2604.20723v1 Announce Type: cross Abstract: The cost of simulator evaluations is a key practical bottleneck for Simulation Based Inference (SBI). In hierarchical settings with shared global para

Top 10 uses for Codex at work

Model ReleasesDGX agent

Codex is OpenAI's AI model designed to understand and generate code, helping developers automate programming tasks and improve productivity. This guide from OpenAI likely outlines ten practical workpl

Toward Safe Autonomous Robotic Endovascular Interventions using World Models

Model ReleasesDGX agent

arXiv:2604.20151v1 Announce Type: cross Abstract: Autonomous mechanical thrombectomy (MT) presents substantial challenges due to highly variable vascular geometries and the requirements for accurate,

Towards Event-Aware Forecasting in DeFi: Insights from On-chain Automated Market Maker Protocols

Model ReleasesDGX agent

arXiv:2604.20374v1 Announce Type: new Abstract: Automated Market Makers (AMMs), as a core infrastructure of decentralized finance (DeFi), uniquely drive on-chain asset pricing through a deterministic

Towards High-Quality Machine Translation for Kokborok: A Low-Resource Tibeto-Burman Language of Northeast India

Model ReleasesDGX agent

arXiv:2604.19778v1 Announce Type: new Abstract: We present KokborokMT, a high-quality neural machine translation (NMT) system for Kokborok (ISO 639-3), a Tibeto-Burman language spoken primarily in Tri

Towards Secure Logging: Characterizing and Benchmarking Logging Code Security Issues with LLMs

Model ReleasesDGX agent

arXiv:2604.20211v1 Announce Type: cross Abstract: Logging code plays an important role in software systems by recording key events and behaviors, which are essential for debugging and monitoring. Howe

Trajectory2Task: Training Robust Tool-Calling Agents with Synthesized Yet Verifiable Data for Complex User Intents

Model ReleasesDGX agent

arXiv:2601.20144v3 Announce Type: replace Abstract: Tool-calling agents are increasingly deployed in real-world customer-facing workflows. Yet most studies on tool-calling agents focus on idealized se

UCCL-Zip: Lossless Compression Supercharged GPU Communication

Model ReleasesDGX agent

arXiv:2604.17172v2 Announce Type: replace-cross Abstract: The rapid growth of large language models (LLMs) has made GPU communication a critical bottleneck. While prior work reduces communication volu

Understanding the Staged Dynamics of Transformers in Learning Latent Structure

Model ReleasesDGX agent

arXiv:2511.19328v2 Announce Type: replace Abstract: Language modeling has shown us that transformers can discover latent structure from context, but the dynamics of how they acquire different componen

UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval

Model ReleasesDGX agent

arXiv:2604.20318v1 Announce Type: new Abstract: Composed image retrieval, multi-turn composed image retrieval, and composed video retrieval all share a common paradigm: composing the reference visual

Unplugging completely! No WiFi and zero notifications. A great way to get deep focus on a project. Here is a walkthrough showing how to run …

Model ReleasesDGX agent

Unplugging completely! No WiFi and zero notifications. A great way to get deep focus on a project. Here is a walkthrough showing how to run Gemma 4 (26B A4B) fully offline with LM Studio & OpenCode to

Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales

Model ReleasesDGX agent

arXiv:2604.20682v1 Announce Type: new Abstract: We present a systematic empirical study of transformer compression through over 40 experiments on GPT-2 (124M parameters) and Mistral 7B (7.24B paramete

Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback

Model ReleasesDGX agent

arXiv:2604.20210v1 Announce Type: cross Abstract: Individual differences in vibrotactile perception underscore the growing importance of personalization as haptic feedback becomes more prevalent in in

Video-ToC: Video Tree-of-Cue Reasoning

Model ReleasesDGX agent

arXiv:2604.20473v1 Announce Type: new Abstract: Existing Video Large Language Models (Video LLMs) struggle with complex video understanding, exhibiting limited reasoning capabilities and potential hal

Waiting at Superchargers is rare, but when it happens, customers should be able to plan with confidence. Superchargers are the only fast cha…

Model ReleasesDGX agent

Waiting at Superchargers is rare, but when it happens, customers should be able to plan with confidence. Superchargers are the only fast chargers with predictive wait times and we're now using vehicle

Weaviate 1.37 Release

Model ReleasesDGX agent

This release introduces the built-in MCP Server, Extensible Tokenizers, Diversity Search (MMR), and Query Profiling as previews, along with Incremental Backups, Gemini audio support for multi2vec-goog

WebGen-R1: Incentivizing Large Language Models to Generate Functional and Aesthetic Websites with Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.20398v1 Announce Type: new Abstract: While Large Language Models (LLMs) excel at function-level code generation, project-level tasks such as generating functional and visually aesthetic mul

We're the top open-weights model on Design Arena!

Model ReleasesDGX agent

We're the top open-weights model on Design Arena! BREAKING: Kimi K2.6 takes 1st overall of open weights models on Design Arena! Kimi K2.6 is in the same performance band as Claude Opus 4.7 - while est

We’ve been looking into recent reports around Claude Code quality issues, and just published a post-mortem on what we found.

Model ReleasesDGX agent

We’ve been looking into recent reports around Claude Code quality issues, and just published a post-mortem on what we found. Over the past month, some of you reported Claude Code's quality had slipped

White-Basilisk: A Hybrid Model for Code Vulnerability Detection

Model ReleasesDGX agent

arXiv:2507.08540v5 Announce Type: replace-cross Abstract: The proliferation of software vulnerabilities presents a significant challenge to cybersecurity, necessitating more effective detection method

Why AI-Generated Text Detection Fails: Evidence from Explainable AI Beyond Benchmark Accuracy

Model ReleasesDGX agent

arXiv:2603.23146v2 Announce Type: replace-cross Abstract: The widespread adoption of Large Language Models (LLMs) has made the detection of AI-Generated text a pressing and complex challenge. Although

WildFireVQA: A Large-Scale Radiometric Thermal VQA Benchmark for Aerial Wildfire Monitoring

Model ReleasesDGX agent

arXiv:2604.20190v1 Announce Type: new Abstract: Wildfire monitoring requires timely, actionable situational awareness from airborne platforms, yet existing aerial visual question answering (VQA) bench

Wordle 1,768 6/6 🟩⬛⬛🟩🟩 🟩⬛⬛🟩🟩 🟩⬛🟩🟩🟩 🟩⬛🟩🟩🟩 🟩⬛🟩🟩🟩 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This post documents a Wordle game result (puzzle #1,768) where the player solved it in 6 attempts, the maximum allowed before losing. The emoji grid shows the progression of letter feedback across gue

WorkflowGen:an adaptive workflow generation mechanism driven by trajectory experience

Model ReleasesDGX agent

arXiv:2604.19756v1 Announce Type: cross Abstract: Large language model (LLM) agents often suffer from high reasoning overhead, excessive token consumption, unstable execution, and inability to reuse p

wow the codex app is unrecognizable… almost like it shouldve been Atlas the whole time

Model ReleasesDGX agent

wow the codex app is unrecognizable… almost like it shouldve been Atlas the whole time With GPT-5.5, Codex now gets more of the job done across the browser, files, docs, and your computer. We've expan

X-PCR: A Benchmark for Cross-modality Progressive Clinical Reasoning in Ophthalmic Diagnosis

Model ReleasesDGX agent

arXiv:2604.20350v1 Announce Type: new Abstract: Despite significant progress in Multi-modal Large Language Models (MLLMs), their clinical reasoning capacity for multi-modal diagnosis remains largely u

Yutori launches Delegate to turn AI agents into proactive web workers

Model ReleasesDGX agent

Yutori Inc., a startup that develops autonomous artificial intelligence agents for everyday knowledge-work tasks on the web, today announced the launch of Delegate, which allows users to offload and “

22 Apr 2026

3D Foundation Model for Generalizable Disease Detection in Head Computed Tomography

Model ReleasesDGX agent

arXiv:2502.02779v3 Announce Type: replace-cross Abstract: Head computed tomography (CT) imaging is a widely-used imaging modality with multitudes of medical indications, particularly in assessing path

a bunch here where I’m saying ok Garry’s kinda right?! 👀…in some ways :) we’re making this loop much easier to close out of the box soon If…

Model ReleasesDGX agent

a bunch here where I’m saying ok Garry’s kinda right?! 👀…in some ways :) we’re making this loop much easier to close out of the box soon If more people get into evals & traces to ground self-improving

A Controlled Benchmark of Visual State-Space Backbones with Domain-Shift and Boundary Analysis for Remote-Sensing Segmentation

Model ReleasesDGX agent

arXiv:2604.18721v1 Announce Type: cross Abstract: Visual state-space models (SSMs) are increasingly promoted as efficient alternatives to Vision Transformers, yet their practical advantages remain unc

A Functionality-Grounded Benchmark for Evaluating Web Agents in E-commerce Domains

Model ReleasesDGX agent

arXiv:2508.15832v2 Announce Type: replace-cross Abstract: Web agents have shown great promise in performing many tasks on ecommerce website. To assess their capabilities, several benchmarks have been

A-MAR: Agent-based Multimodal Art Retrieval for Fine-Grained Artwork Understanding

Model ReleasesDGX agent

arXiv:2604.19689v1 Announce Type: new Abstract: Understanding artworks requires multi-step reasoning over visual content and cultural, historical, and stylistic context. While recent multimodal large

AD-Copilot: A Vision-Language Assistant for Industrial Anomaly Detection via Visual In-context Comparison

Model ReleasesDGX agent

arXiv:2603.13779v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have achieved impressive success in natural visual understanding, yet they consistently underperform

Agent-GWO: Collaborative Agents for Dynamic Prompt Optimization in Large Language Models

Model ReleasesDGX agent

arXiv:2604.18612v1 Announce Type: cross Abstract: Large Language Models (LLMs) have demonstrated strong capabilities in complex reasoning tasks, while recent prompting strategies such as Chain-of-Thou

Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs

Model ReleasesDGX agent

arXiv:2604.18576v2 Announce Type: replace Abstract: We present BLF (Bayesian Linguistic Forecaster), an agentic system for binary forecasting that achieves state-of-the-art performance on the Forecast

[AINews] OpenAI launches GPT-Image-2

Model ReleasesDGX agent

OpenAI has launched GPT-Image-2, an advancement in their image generation capabilities. The model likely represents improvements over previous versions in areas such as image quality, prompt understan

Air-Know: Arbiter-Calibrated Knowledge-Internalizing Robust Network for Composed Image Retrieval

Model ReleasesDGX agent

arXiv:2604.19386v1 Announce Type: new Abstract: Composed Image Retrieval (CIR) has attracted significant attention due to its flexible multimodal query method, yet its development is severely constrai

Alibaba launches Qwen3.6-27B, an open-weight dense model with 27B parameters, saying it surpasses Qwen3.5-397B-A17B on major coding benchmarks (Qwen)

Model ReleasesDGX agent

Qwen: Alibaba launches Qwen3.6-27B, an open-weight dense model with 27B parameters, saying it surpasses Qwen3.5-397B-A17B on major coding benchmarks — · 4226 words · QwenTeam丨Translations:.体中文 — HUGGI

AlignCultura: Towards Culturally Aligned Large Language Models?

Model ReleasesDGX agent

arXiv:2604.19016v1 Announce Type: new Abstract: Cultural alignment in Large Language Models (LLMs) is essential for producing contextually aware, respectful, and trustworthy outputs. Without it, model

All of the AI models have preferred names. If you asked Claude 4.5 for a software developer, you are going to get Marcus Chen. Wizards are m…

Model ReleasesDGX agent

All of the AI models have preferred names. If you asked Claude 4.5 for a software developer, you are going to get Marcus Chen. Wizards are mostly named Aldric. Space pilots are Kira from Claude, Mara

An Efficient Black-Box Reduction from Online Learning to Multicalibration, and a New Route to Phi-Regret Minimization

Model ReleasesDGX agent

arXiv:2604.19592v1 Announce Type: new Abstract: We give a Gordon-Greenwald-Marks (GGM) style black-box reduction from online learning to online multicalibration. Concretely, we show that to achieve hi

An Empirical Study of Multi-Generation Sampling for Jailbreak Detection in Large Language Models

Model ReleasesDGX agent

arXiv:2604.18775v1 Announce Type: new Abstract: Detecting jailbreak behaviour in large language models remains challenging, particularly when strongly aligned models produce harmful outputs only rarel

An Experimental Characterization of Mechanical Layer Jamming Systems

Model ReleasesDGX agent

arXiv:2511.07882v2 Announce Type: replace Abstract: Organisms in nature, such as Cephalopods and Pachyderms, exploit stiffness modulation to achieve amazing dexterity in the control of their appendage

An undergraduate used AI assistants to rewrite leaked source code for Claude Code in a different language, highlighting the uncertainty over copyright and AI (Meaghan Tobin/New York Times)

Model ReleasesDGX agent

Meaghan Tobin / New York Times: An undergraduate used AI assistants to rewrite leaked source code for Claude Code in a different language, highlighting the uncertainty over copyright and AI — Artifici

Analytical Extraction of Conditional Sobol' Indices via Basis Decomposition of Polynomial Chaos Expansions

Model ReleasesDGX agent

arXiv:2604.19165v1 Announce Type: cross Abstract: In uncertainty quantification, evaluating sensitivity measures under specific conditions (i.e., conditional Sobol' indices) is essential for systems w

... and Anthropic reverted this change. Claude Code is now part of Pro, as per the Pricing page. Important note on the growth hack: Anthropi…

Model ReleasesDGX agent

... and Anthropic reverted this change. Claude Code is now part of Pro, as per the Pricing page. Important note on the growth hack: Anthropic advertises safety and integrity as their values. A 'fake d

Announcing Spanner Omni: Your infrastructure, Google’s innovation

Model ReleasesDGX agent

Today, we announced the preview of Spanner Omni, a downloadable version of Spanner, that expands its industry-leading distributed database capabilities beyond Google Cloud. This enables enterprises to

Anthropic have now reverted the change to their pricing page, but I've not yet seen any official announcement about what their policies are …

Model ReleasesDGX agent

Anthropic have now reverted the change to their pricing page, but I've not yet seen any official announcement about what their policies are going to be going forward. More on my blog: https://simonwil

Anthropic investigates unauthorized access to restricted Claude Mythos AI model

Model ReleasesDGX agent

Anthropic PBC is investigating a report that unauthorized users accessed Claude Mythos, the next-level artificial intelligence model the company says is powerful enough to enable dangerous cyberattack

← Previous
1…318319320321322…373
Next →