AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,340 results
23 Apr 2026

CRAFT: Training-Free Cascaded Retrieval for Tabular QA

Model ReleasesDGX agent

arXiv:2505.14984v2 Announce Type: replace Abstract: Open-Domain Table Question Answering (TQA) involves retrieving relevant tables from a large corpus to answer natural language queries. Traditional d

CrowdStrike launches Project QuiltWorks coalition to tackle AI-discovered vulnerabilities

Model ReleasesDGX agent

CrowdStrike Holdings Inc. today announced the launch of Project QuiltWorks, an industry coalition aimed at helping enterprises find and fix the wave of software vulnerabilities being surfaced by front

CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.20389v1 Announce Type: cross Abstract: The rapid evolution and use of Large Language Models (LLMs) in professional workflows require an evaluation of their domain-specific knowledge against

Databricks partners with OpenAI on GPT-5.5

Model ReleasesDGX agent

Databricks has announced a partnership with OpenAI to integrate GPT-5.5, OpenAI's latest language model, into its data and AI platform. This collaboration enables Databricks customers to leverage GPT-

Day 0 vLLM support for Qwen3.6-27B! @vllm_project ♥️❤️

Model ReleasesDGX agent

Day 0 vLLM support for Qwen3.6-27B! @vllm_project ♥️❤️ 🎉 Day-0 vLLM support for Qwen3.6-27B! Congrats to @Alibaba_Qwen on the new 27B dense model release. Looking forward to more of the Qwen3.6 series

Day 2 at Google Cloud Next: A marathon developer keynote

Model ReleasesDGX agent

At Google Cloud, every day is Developer Day, but none so much as day 2 of Google Cloud Next, when we hold the developer keynote. This year’s topic? An in-depth look at Gemini Enterprise Agent Platform

Deepseek V4 on AI Gateway

Model ReleasesDGX agent

Vercel announced support for DeepSeek V4 model through its AI Gateway service, enabling developers to access this AI model alongside other supported models on the platform. This addition expands Verce

Development and Preliminary Evaluation of a Domain-Specific Large Language Model for Tuberculosis Care in South Africa

Model ReleasesDGX agent

arXiv:2604.19776v1 Announce Type: new Abstract: Tuberculosis (TB) is one of the world's deadliest infectious diseases, and in South Africa, it contributes a significant burden to the country's health

DialToM: A Theory of Mind Benchmark for Forecasting State-Driven Dialogue Trajectories

Model ReleasesDGX agent

arXiv:2604.20443v1 Announce Type: cross Abstract: Large Language Models (LLMs) have been shown to possess Theory of Mind (ToM) abilities. However, it remains unclear whether this stems from robust rea

Differentiable Conformal Training for LLM Reasoning Factuality

Model ReleasesDGX agent

arXiv:2604.20098v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently hallucinate, limiting their reliability in critical applications. Conformal Prediction (CP) addresses this by ca

DistortBench: Benchmarking Vision Language Models on Image Distortion Identification

Model ReleasesDGX agent

arXiv:2604.19966v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used in settings where sensitivity to low-level image degradations matters, including content moderatio

Do Small Language Models Know When They're Wrong? Confidence-Based Cascade Scoring for Educational Assessment

Model ReleasesDGX agent

arXiv:2604.19781v1 Announce Type: cross Abstract: Automated scoring of student work at scale requires balancing accuracy against cost and latency. In 'cascade' systems, small language models (LMs) han

'don't retweet this, don't retweet this, don't retweet this...' ah fuck it, life imitates art.

Model ReleasesDGX agent

'don't retweet this, don't retweet this, don't retweet this...' ah fuck it, life imitates art. In Vending-Bench Arena (the multiplayer version of Vending-Bench with competition dynamics), GPT-5.5 actu

Dual Causal Inference: Integrating Backdoor Adjustment and Instrumental Variable Learning for Medical VQA

Model ReleasesDGX agent

arXiv:2604.20306v1 Announce Type: cross Abstract: Medical Visual Question Answering (MedVQA) aims to generate clinically reliable answers conditioned on complex medical images and questions. However,

Duluth at SemEval-2026 Task 6: DeBERTa with LLM-Augmented Data for Unmasking Political Question Evasions

Model ReleasesDGX agent

arXiv:2604.20168v1 Announce Type: new Abstract: This paper presents the Duluth approach to SemEval-2026 Task 6 on CLARITY: Unmasking Political Question Evasions. We address Task 1 (clarity-level class

Early-Stage Product Line Validation Using LLMs: A Study on Semi-Formal Blueprint Analysis

Model ReleasesDGX agent

arXiv:2604.20523v1 Announce Type: cross Abstract: We study whether Large Language Models (LLMs) can perform feature model analysis operations (AOs) directly on semi-formal textual blueprints, i.e., co

Earth Day at #GoogleCloudNext, I’m demoing a Sustainability Agent at the @nvidia booth. Built with @Google ADK, @googlegemma, #nemotron, Clo…

Model ReleasesDGX agent

Earth Day at #GoogleCloudNext, I’m demoing a Sustainability Agent at the @nvidia booth. Built with @Google ADK, @googlegemma, #nemotron, Cloud Run, @milvusio , @LangChain & @ollama to reason across im

Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models

Model ReleasesDGX agent

arXiv:2511.06209v4 Announce Type: replace Abstract: LLMs can solve complex tasks by generating long, multi-step reasoning chains. Test-time scaling (TTS) can further improve LLM performance by samplin

embers

Model ReleasesDGX agent

embers GPT-5.5, not fully saturating the TikZ unicorn test yet but getting awfully close ... (yes this is actual TikZ code, I personally find it so unbelievable that I'm putting the code below for any

Evaluating the Quality of the Quantified Uncertainty for (Re)Calibration of Data-Driven Regression Models

Model ReleasesDGX agent

arXiv:2508.17761v3 Announce Type: replace Abstract: In safety-critical applications data-driven models must not only be accurate but also provide reliable uncertainty estimates. This property, commonl

Evian: Towards Explainable Visual Instruction-tuning Data Auditing

Model ReleasesDGX agent

arXiv:2604.20544v1 Announce Type: cross Abstract: The efficacy of Large Vision-Language Models (LVLMs) is critically dependent on the quality of their training data, requiring a precise balance betwee

Evidence of Layered Positional and Directional Constraints in the Voynich Manuscript: Implications for Cipher-Like Structure

Model ReleasesDGX agent

arXiv:2604.19762v1 Announce Type: new Abstract: The Voynich Manuscript (VMS) exhibits a script of uncertain origin whose grapheme sequences have resisted linguistic analysis. We present a systematic a

EvoForest: A Novel Machine-Learning Paradigm via Open-Ended Evolution of Computational Graphs

Model ReleasesDGX agent

arXiv:2604.19761v1 Announce Type: new Abstract: Modern machine learning is still largely organized around a single recipe: choose a parameterized model family and optimize its weights. Although highly

Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts

Model ReleasesDGX agent

arXiv:2604.19835v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) has become the dominant architecture for scaling large language models: frontier models routinely decouple total parameters f

Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization

Model ReleasesDGX agent

arXiv:2604.20726v1 Announce Type: cross Abstract: This work explores the role of prompt design and judge selection in LLM-as-a-Judge evaluations of free text legal question answering. We examine wheth

Exploring Data Augmentation and Resampling Strategies for Transformer-Based Models to Address Class Imbalance in AI Scoring of Scientific Explanations in NGSS Classroom

Model ReleasesDGX agent

arXiv:2604.19754v1 Announce Type: new Abstract: Automated scoring of students' scientific explanations offers the potential for immediate, accurate feedback, yet class imbalance in rubric categories p

Exploring Spatial Intelligence from a Generative Perspective

Model ReleasesDGX agent

arXiv:2604.20570v1 Announce Type: new Abstract: Spatial intelligence is essential for multimodal large language models, yet current benchmarks largely assess it only from an understanding perspective.

Extract PDF text in your browser with LiteParse for the web

Model ReleasesDGX agent

LlamaIndex have a most excellent open source project called LiteParse, which provides a Node.js CLI tool for extracting text from PDFs. I got a version of LiteParse working entirely in the browser, us

Fairness Testing of Large Language Models in Role-Playing

Model ReleasesDGX agent

arXiv:2411.00585v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have become foundational in modern language-driven software applications, profoundly influencing daily life. A cr

Falcon 9 launches 24 @Starlink satellites from California

Model ReleasesDGX agent

SpaceX's Falcon 9 rocket successfully launched 24 Starlink satellites from a California launch facility, continuing the company's ongoing deployment of its satellite internet constellation. This missi

Fast Bayesian equipment condition monitoring via simulation based inference: applications to heat exchanger health

Model ReleasesDGX agent

arXiv:2604.20735v1 Announce Type: new Abstract: Accurate condition monitoring of industrial equipment requires inferring latent degradation parameters from indirect sensor measurements under uncertain

Fast-then-Fine: A Two-Stage Framework with Multi-Granular Representation for Cross-Modal Retrieval in Remote Sensing

Model ReleasesDGX agent

arXiv:2604.20429v1 Announce Type: new Abstract: Remote sensing (RS) image-text retrieval plays a critical role in understanding massive RS imagery. However, the dense multi-object distribution and com

FeDa4Fair: Client-Level Federated Datasets for Fairness Evaluation

Model ReleasesDGX agent

arXiv:2506.21095v4 Announce Type: replace-cross Abstract: Federated Learning (FL) enables collaborative training while preserving privacy, yet it introduces a critical challenge: the 'illusion of fair

Fextsuperscript{2}LP-AP: Fast & Flexible Label Propagation with Adaptive Propagation Kernel

Model ReleasesDGX agent

arXiv:2604.20736v1 Announce Type: new Abstract: Semi-supervised node classification is a foundational task in graph machine learning, yet state-of-the-art Graph Neural Networks (GNNs) are hindered by

Finding Duplicates in 1.1M BDD Steps: cukereuse, a Paraphrase-Robust Static Detector for Cucumber and Gherkin

Model ReleasesDGX agent

arXiv:2604.20462v1 Announce Type: cross Abstract: Behaviour-Driven Development (BDD) suites accumulate step-text duplication whose maintenance cost is established in prior work. Existing detection tec

FlashNorm: Fast Normalization for Transformers

Model ReleasesDGX agent

arXiv:2407.09577v4 Announce Type: replace Abstract: Normalization layers are ubiquitous in large language models (LLMs) yet represent a compute bottleneck: on hardware with distinct vector and matrix

Foundation Models in Biomedical Imaging: Turning Hype into Reality

Model ReleasesDGX agent

arXiv:2512.15808v2 Announce Type: replace-cross Abstract: Foundation models (FMs) are driving a prominent shift in biomedical imaging from task-specific models to unified backbone models for diverse t

From Data to Theory: Autonomous Large Language Model Agents for Materials Science

Model ReleasesDGX agent

arXiv:2604.19789v1 Announce Type: new Abstract: We present an autonomous large language model (LLM) agent for end-to-end, data-driven materials theory development. The model can choose an equation for

From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents

Model ReleasesDGX agent

arXiv:2604.20006v1 Announce Type: new Abstract: Personalized agents that interact with users over long periods must maintain persistent memory across sessions and update it as circumstances change. Ho

From Scene to Object: Text-Guided Dual-Gaze Prediction

Model ReleasesDGX agent

arXiv:2604.20191v1 Announce Type: cross Abstract: Interpretable driver attention prediction is crucial for human-like autonomous driving. However, existing datasets provide only scene-level global gaz

Global Offshore Wind Infrastructure: Deployment and Operational Dynamics from Dense Sentinel-1 Time Series

Model ReleasesDGX agent

arXiv:2604.20822v1 Announce Type: new Abstract: The offshore wind energy sector is expanding rapidly, increasing the need for independent, high-temporal-resolution monitoring of infrastructure deploym

GPT-5.5 Bio Bug Bounty

Model ReleasesDGX agent

OpenAI's GPT-5.5 Bio Bug Bounty program invites security researchers to identify and report vulnerabilities in GPT-5.5's biological information handling capabilities, focusing on potential misuse risk

GPT-5.5 in Codex is a delight to work with: - Super sharp with responses - It understands intent better than any model - Great 'personality'…

Model ReleasesDGX agent

GPT-5.5 in Codex is a delight to work with: - Super sharp with responses - It understands intent better than any model - Great 'personality' - Gets lots of stuff done without pausing unnecessarily It

GPT-5.5 is here! We hope it's useful to you. I personally like it.

Model ReleasesDGX agent

I don't have verified information about a GPT-5.5 model release. This appears to be either a fictional or future-dated post, as it references a non-existent model and uses a URL format/status ID that

GPT-5.5 is likely the best model in the world. But open models like Kimi and Minimax get almost identical coding benchmark scores at 10-25x …

Model ReleasesDGX agent

GPT-5.5 is likely the best model in the world. But open models like Kimi and Minimax get almost identical coding benchmark scores at 10-25x lower cost. I broke down the benchmarks and pricing. Here's

GPT-5.5 is now accessible in Hermes Agent through the ChatGPT/Codex OAuth provider. Run `hermes update` to access now or learn how to get st…

Model ReleasesDGX agent

GPT-5.5 is now accessible in Hermes Agent through the ChatGPT/Codex OAuth provider. Run `hermes update` to access now or learn how to get started with Hermes Agent here: https://hermes-agent.nousresea

GPT-5.5 is priced at 5/1M input tokens and 30/1M output tokens, double GPT-5.4's pricing; GPT-5.5 Pro costs 30/1M input tokens and 180/1M output tokens (Carl Franzen/VentureBeat)

Model ReleasesDGX agent

Carl Franzen / VentureBeat: GPT-5.5 is priced at 5/1M input tokens and 30/1M output tokens, double GPT-5.4's pricing; GPT-5.5 Pro costs 30/1M input tokens and 180/1M output tokens — After months of ru

GPT-5.5 is rolling out to Plus, Pro, Business, and Enterprise users in ChatGPT and Codex, and GPT-5.5 Pro to Pro, Business, and Enterprise users in ChatGPT (The Verge)

Model ReleasesDGX agent

The Verge: GPT-5.5 is rolling out to Plus, Pro, Business, and Enterprise users in ChatGPT and Codex, and GPT-5.5 Pro to Pro, Business, and Enterprise users in ChatGPT — The new model ‘excels’ at tasks

GPT-5.5 may not be in the official OpenAI API... but it's available via the apparently approved-of Codex API backdoor So I used that to make…

Model ReleasesDGX agent

GPT-5.5 may not be in the official OpenAI API... but it's available via the apparently approved-of Codex API backdoor So I used that to make these pelicans (default and xhigh)! https://simonwillison.n

GPT-5.5 on ARC-AGI (Verified) ARC-AGI-2: - Max: 85.0%, 1.87 - High: 83.3%, 1.45 - Med: 70.4%, 0.86 - Low: 33%, 0.35 GPT-5.5 is now state…

Model ReleasesDGX agent

GPT-5.5 achieved state-of-the-art performance on the ARC-AGI-2 benchmark, with scores ranging from 85.0% on maximum difficulty tasks to 33% on low difficulty tasks. The model demonstrated consistent i

Graph-Theoretic Models for the Prediction of Molecular Measurements

Model ReleasesDGX agent

arXiv:2604.19840v1 Announce Type: new Abstract: Graph-theoretic approaches offer simplicity, interpretability, and low computational cost for molecular property prediction. Among these, the model prop

Here’s my view on GPT-5.5, which I have been testing for a couple of weeks. It conducted not-bad social science research on its own, develop…

Model ReleasesDGX agent

Here’s my view on GPT-5.5, which I have been testing for a couple of weeks. It conducted not-bad social science research on its own, developed a novel RPG & more. There is still jaggedness but GPT-5.5

HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs

Model ReleasesDGX agent

arXiv:2604.20140v1 Announce Type: new Abstract: Direct Preference Optimization (DPO) is an effective framework for aligning large language models with human preferences, but it struggles with complex

How close are models to building a perfect Slack clone with less than $50,000 of tokens? I feel like not that far...

Model ReleasesDGX agent

How close are models to building a perfect Slack clone with less than 50,000 of tokens? I feel like not that far... Claude Code spend had gotten to 10.95M runrate peak at SemiAnalysis But then Opus 4.

How Much Does Persuasion Strategy Matter? LLM-Annotated Evidence from Charitable Donation Dialogues

Model ReleasesDGX agent

arXiv:2604.19783v1 Announce Type: new Abstract: Which persuasion strategies, if any, are associated with donation compliance? Answering this requires fine-grained strategy labels across a full corpus

Hybrid Multi-Phase Page Matching and Multi-Layer Diff Detection for Japanese Building Permit Document Review

Model ReleasesDGX agent

arXiv:2604.19770v1 Announce Type: new Abstract: We present a hybrid multi-phase page matching algorithm for automated comparison of Japanese building permit document sets. Building permit review in Ja

I had early access to GPT-5.5. It is very good, especially the Pro version. Full writeup very shortly.

Model ReleasesDGX agent

Ethan Mollick posted on X about having early access to GPT-5.5, commenting positively on its capabilities and noting that the Pro version is particularly strong. He indicated that a full detailed writ

I’d been part of OpenAI early tester group for GPT-5.5. I believe with GPT-5.5 Pro we reached another inflection point-comparable to the ori…

Model ReleasesDGX agent

I’d been part of OpenAI early tester group for GPT-5.5. I believe with GPT-5.5 Pro we reached another inflection point-comparable to the original release of o1-preview & then with 5.0 Pro, I had felt.

If you want to stack rank LLMs/VLMs on document understanding 📄, you can through ParseBench, now live on @kaggle 📊 ParseBench is the most …

Model ReleasesDGX agent

If you want to stack rank LLMs/VLMs on document understanding 📄, you can through ParseBench, now live on @kaggle 📊 ParseBench is the most comprehensive document OCR benchmark over real enterprise docu

I'm a manager at @OpenAI, but with GPT-5.5 I'm a more effective IC than I've ever been. I can now write CUDA kernels like a pro. I can rely …

Model ReleasesDGX agent

I'm a manager at @OpenAI, but with GPT-5.5 I'm a more effective IC than I've ever been. I can now write CUDA kernels like a pro. I can rely on it to run my research experiments. And we know how to mak

← Previous
1…316317318319320…373
Next →