AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,332 results
23 Apr 2026

I had early access to GPT-5.5. It is very good, especially the Pro version. Full writeup very shortly.

Model ReleasesDGX agent

Ethan Mollick posted on X about having early access to GPT-5.5, commenting positively on its capabilities and noting that the Pro version is particularly strong. He indicated that a full detailed writ

I’d been part of OpenAI early tester group for GPT-5.5. I believe with GPT-5.5 Pro we reached another inflection point-comparable to the ori…

Model ReleasesDGX agent

I’d been part of OpenAI early tester group for GPT-5.5. I believe with GPT-5.5 Pro we reached another inflection point-comparable to the original release of o1-preview & then with 5.0 Pro, I had felt.

If you want to stack rank LLMs/VLMs on document understanding 📄, you can through ParseBench, now live on @kaggle 📊 ParseBench is the most …


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

If you want to stack rank LLMs/VLMs on document understanding 📄, you can through ParseBench, now live on @kaggle 📊 ParseBench is the most comprehensive document OCR benchmark over real enterprise docu

I'm a manager at @OpenAI, but with GPT-5.5 I'm a more effective IC than I've ever been. I can now write CUDA kernels like a pro. I can rely …

Model ReleasesDGX agent

I'm a manager at @OpenAI, but with GPT-5.5 I'm a more effective IC than I've ever been. I can now write CUDA kernels like a pro. I can rely on it to run my research experiments. And we know how to mak

IMPACT-CYCLE: A Contract-Based Multi-Agent System for Claim-Level Supervisory Correction of Long-Video Semantic Memory

Model ReleasesDGX agent

arXiv:2604.20136v1 Announce Type: cross Abstract: Correcting errors in long-video understanding is disproportionately costly: existing multimodal pipelines produce opaque, end-to-end outputs that expo

important (and very jakub-coded) jakub quote:

Model ReleasesDGX agent

important (and very jakub-coded) jakub quote: OpenAI Unveils GPT-5.5. Company Says Expect a Faster Model Release Pace 👀 OpenAI: 'We see pretty significant improvements in the short term, extremely sig

🚨 In a new court filing (below), Clippers owner Steve Ballmer dismisses @pablofindsout as “gossip” from a “former talking head and televisi…

Model ReleasesDGX agent

🚨 In a new court filing (below), Clippers owner Steve Ballmer dismisses @pablofindsout as “gossip” from a “former talking head and television personality.” Here is an excerpt from the federal whistleb

In ChatGPT, full-stack inference improvements enable a more capable model at faster speed. This efficiency is a game-changer for GPT-5.5 Pro…

Model ReleasesDGX agent

In ChatGPT, full-stack inference improvements enable a more capable model at faster speed. This efficiency is a game-changer for GPT-5.5 Pro, now a much more practical option for demanding tasks, and

Infection-Reasoner: A Compact Vision-Language Model for Wound Infection Classification with Evidence-Grounded Clinical Reasoning

Model ReleasesDGX agent

arXiv:2604.19937v1 Announce Type: cross Abstract: Assessing chronic wound infection from photographs is challenging because visual appearance varies across wound etiologies, anatomical locations, and

Instagram launches Instants, an app for sharing disappearing photos, in Italy and Spain, after rolling out an Instants feature in its main app in some regions (Sydney Bradley/Business Insider)

Model ReleasesDGX agent

Sydney Bradley / Business Insider: Instagram launches Instants, an app for sharing disappearing photos, in Italy and Spain, after rolling out an Instants feature in its main app in some regions — - In

Interesting, OpenAI just released a free healthcare version of ChatGPT-5.4 for clinicians that beat specialty-matched physicians with unlimi…

Model ReleasesDGX agent

Interesting, OpenAI just released a free healthcare version of ChatGPT-5.4 for clinicians that beat specialty-matched physicians with unlimited time + web access on a benchmark of real & hard clinical

Intersectional Fairness in Large Language Models

Model ReleasesDGX agent

arXiv:2604.20677v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in socially sensitive settings, raising concerns about fairness and biases, particularly across i

Introducing GPT-5.5 A new class of intelligence for real work and powering agents, built to understand complex goals, use tools, check its w…

Model ReleasesDGX agent

Introducing GPT-5.5 A new class of intelligence for real work and powering agents, built to understand complex goals, use tools, check its work, and carry more tasks through to completion. It marks a

I've been previewing this in Codex for a few weeks - it's very good! Had some great results from it having it run security reviews against c…

Model ReleasesDGX agent

I've been previewing this in Codex for a few weeks - it's very good! Had some great results from it having it run security reviews against code written using other models Introducing GPT-5.5 A new cla

IVY-FAKE: A Unified Explainable Framework and Benchmark for Image and Video AIGC Detection

Model ReleasesDGX agent

arXiv:2506.00979v5 Announce Type: replace-cross Abstract: The rapid development of Artificial Intelligence Generated Content (AIGC) techniques has enabled the creation of high-quality synthetic conten

KANMixer: a minimal KAN-centered mixer for long-term time series forecasting

Model ReleasesDGX agent

arXiv:2508.01575v2 Announce Type: replace Abstract: Long-term time series forecasting (LTSF) underpins critical applications from energy management to weather prediction, yet achieving reliable multi-

Kimi K2.6 becomes the #1 open model on MathArena!

Model ReleasesDGX agent

Kimi K2.6 achieved the top ranking on MathArena, a benchmark for evaluating mathematical problem-solving capabilities in open-source language models. This announcement highlights the model's superior

Knapsack Optimization-based Schema Linking for LLM-based Text-to-SQL Generation

Model ReleasesDGX agent

arXiv:2502.12911v3 Announce Type: replace Abstract: Generating SQLs from user queries is a long-standing challenge, where the accuracy of initial schema linking significantly impacts subsequent SQL ge

Knowledge Capsules: Structured Nonparametric Memory Units for LLMs

Model ReleasesDGX agent

arXiv:2604.20487v1 Announce Type: cross Abstract: Large language models (LLMs) encode knowledge in parametric weights, making it costly to update or extend without retraining. Retrieval-augmented gene

KoALa-Bench: Evaluating Large Audio Language Models on Korean Speech Understanding and Faithfulness

Model ReleasesDGX agent

arXiv:2604.19782v1 Announce Type: cross Abstract: Recent advances in large audio language models (LALMs) have enabled multilingual speech understanding. However, benchmarks for evaluating LALMs remain

KOCO-BENCH: Can Large Language Models Leverage Domain Knowledge in Software Development?

Model ReleasesDGX agent

arXiv:2601.13240v2 Announce Type: replace-cross Abstract: Large language models (LLMs) excel at general programming but struggle with domain-specific software development, necessitating domain special

Large Language Models Meet Biomedical Knowledge Graphs for Mechanistically Grounded Therapeutic Prioritization

Model ReleasesDGX agent

arXiv:2604.19815v1 Announce Type: new Abstract: Drug repurposing is often framed as a candidate identification task, but existing approaches provide limited guidance for distinguishing biologically pl

Large Language Models Outperform Humans in Fraud Detection and Resistance to Motivated Investor Pressure

Model ReleasesDGX agent

arXiv:2604.20652v1 Announce Type: new Abstract: Large language models trained on human feedback may suppress fraud warnings when investors arrive already persuaded of a fraudulent opportunity. We test

Last night was the biggest disaster in the history of Tesla. Let me walk you through what actually happened on that earnings call, because t…

Model ReleasesDGX agent

Last night was the biggest disaster in the history of Tesla. Let me walk you through what actually happened on that earnings call, because the headlines are doing you a disservice: Elon Musk got on th

Last week, we launched Gemini 3.1 TTS, our latest and best text-to-speech model. This new model introduces [awe] audio tags, an intuitive wa…

Model ReleasesDGX agent

Last week, we launched Gemini 3.1 TTS, our latest and best text-to-speech model. This new model introduces [awe] audio tags, an intuitive way to guide vocal style, pace, and delivery. Here are some ti

Latent Stochastic Interpolants

Model ReleasesDGX agent

arXiv:2506.02276v2 Announce Type: replace Abstract: Stochastic Interpolants (SI) is a powerful framework for generative modeling, capable of flexibly transforming between two probability distributions

LayerTracer: A Joint Task-Particle and Vulnerable-Layer Analysis framework for Arbitrary Large Language Model Architectures

Model ReleasesDGX agent

arXiv:2604.20556v1 Announce Type: cross Abstract: Currently, Large Language Models (LLMs) feature a diversified architectural landscape, including traditional Transformer, GateDeltaNet, and Mamba. How

Learning to Evolve: A Self-Improving Framework for Multi-Agent Systems via Textual Parameter Graph Optimization

Model ReleasesDGX agent

arXiv:2604.20714v1 Announce Type: new Abstract: Designing and optimizing multi-agent systems (MAS) is a complex, labor-intensive process of 'Agent Engineering.' Existing automatic optimization methods

Learning When Not to Decide: A Framework for Overcoming Factual Presumptuousness in AI Adjudication

Model ReleasesDGX agent

arXiv:2604.19895v1 Announce Type: new Abstract: A well-known limitation of AI systems is presumptuousness: the tendency of AI systems to provide confident answers when information may be lacking. This

Less Languages, Less Tokens: An Efficient Unified Logic Cross-lingual Chain-of-Thought Reasoning Framework

Model ReleasesDGX agent

arXiv:2604.20090v1 Announce Type: new Abstract: Cross-lingual chain-of-thought (XCoT) with self-consistency markedly enhances multilingual reasoning, yet existing methods remain costly due to extensiv

LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent

Model ReleasesDGX agent

arXiv:2604.17931v2 Announce Type: replace Abstract: Reinforcement Learning (RL) has emerged as a powerful training paradigm for LLM-based agents. However, scaling agentic RL for deep research remains

LLAMADRS: Evaluating Open-Source LLMs on Real Clinical Interviews--To Reason or Not to Reason?

Model ReleasesDGX agent

arXiv:2501.03624v2 Announce Type: replace-cross Abstract: Large language models (LLMs) excel on many NLP benchmarks, but their behavior on real-world, semi-structured prediction remains underexplored.

LLM-guided phase diagram construction through high-throughput experimentation

Model ReleasesDGX agent

arXiv:2604.20304v1 Announce Type: cross Abstract: Constructing phase diagrams for multicomponent alloys requires extensive experimental measurements and is a time-consuming task. Here we investigate w

looks like new Pareto frontiers across everything: - Context: 400K context in Codex and a 1M in API - API Pricing: 5/m input and 30/m outp…

Model ReleasesDGX agent

looks like new Pareto frontiers across everything: - Context: 400K context in Codex and a 1M in API - API Pricing: 5/m input and 30/m output tokens. - Codex improved its own inference speed 20% lol -

LoRA-FA: Efficient and Effective Low Rank Representation Fine-tuning

Model ReleasesDGX agent

arXiv:2308.03303v2 Announce Type: replace Abstract: Fine-tuning large language models (LLMs) is crucial for improving their performance on downstream tasks, but full-parameter fine-tuning (Full-FT) is

MambaLiteUNet: Cross-Gated Adaptive Feature Fusion for Robust Skin Lesion Segmentation

Model ReleasesDGX agent

arXiv:2604.20286v1 Announce Type: cross Abstract: Recent segmentation models have demonstrated promising efficiency by aggressively reducing parameter counts and computational complexity. However, the

MAPRPose: Mask-Aware Proposal and Amodal Refinement for Multi-Object 6D Pose Estimation

Model ReleasesDGX agent

arXiv:2604.20650v1 Announce Type: new Abstract: 6D object pose estimation in cluttered scenes remains challenging due to severe occlusion and sensor noise. We propose MAPRPose, a two-stage framework t

Measuring the Machine: Evaluating Generative AI as Pluralist Sociotechical Systems

Model ReleasesDGX agent

arXiv:2604.20545v1 Announce Type: new Abstract: In measurement theory, instruments do not simply record reality; they help constitute what is observed. The same holds for generative AI evaluation: ben

Meta-Tool: Efficient Few-Shot Tool Adaptation for Small Language Models

Model ReleasesDGX agent

arXiv:2604.20148v1 Announce Type: cross Abstract: Can small language models achieve strong tool-use performance without complex adaptation mechanisms? This paper investigates this question through Met

MetaboNet: The Largest Publicly Available Consolidated Dataset for Type 1 Diabetes Management

Model ReleasesDGX agent

arXiv:2601.11505v2 Announce Type: replace-cross Abstract: Progress in Type 1 Diabetes (T1D) algorithm development is limited by the fragmentation and lack of standardization across existing T1D manage

Microsoft launches ‘vibe working’ in Word, Excel, and PowerPoint

Model ReleasesDGX agent

Microsoft is rolling out a new Agent Mode inside Office apps like Word, Excel, and PowerPoint this week. Previously described by Microsoft as 'vibe working,' the Agent Mode is a more powerful version

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models

Model ReleasesDGX agent

arXiv:2604.19809v1 Announce Type: new Abstract: We introduce MIRROR, a benchmark comprising eight experiments across four metacognitive levels that evaluates whether large language models can use self

MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror

Model ReleasesDGX agent

arXiv:2604.14785v2 Announce Type: replace Abstract: Recent progress in Multimodal Large Language Models (MLLMs) has demonstrated remarkable advances in perception and reasoning, suggesting their poten

Mitigating Hallucinations in Large Vision-Language Models without Performance Degradation

Model ReleasesDGX agent

arXiv:2604.20366v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) exhibit powerful generative capabilities but frequently produce hallucinations that compromise output reliability.

Mitigating Prompt-Induced Cognitive Biases in General-Purpose AI for Software Engineering

Model ReleasesDGX agent

arXiv:2604.16756v2 Announce Type: replace-cross Abstract: Prompt-induced cognitive biases are changes in a general-purpose AI (GPAI) system's decisions caused solely by biased wording in the input (e.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design

Model ReleasesDGX agent

arXiv:2412.14590v2 Announce Type: replace Abstract: Quantization has become one of the most effective methodologies to compress LLMs into smaller size. However, the existing quantization solutions sti

MLG-Stereo: ViT Based Stereo Matching with Multi-Stage Local-Global Enhancement

Model ReleasesDGX agent

arXiv:2604.20393v1 Announce Type: new Abstract: With the development of deep learning, ViT-based stereo matching methods have made significant progress due to their remarkable robustness and zero-shot

Model Capability Assessment and Safeguards for Biological Weaponization

Model ReleasesDGX agent

arXiv:2604.19811v1 Announce Type: cross Abstract: AI leaders and safety reports increasingly warn that advances in model reasoning may enable biological misuse, including by low-expertise users, while

MSLAU-Net: A Hybrid CNN-Transformer Network for Medical Image Segmentation

Model ReleasesDGX agent

arXiv:2505.18823v2 Announce Type: replace Abstract: Accurate medical image segmentation allows for the precise delineation of anatomical structures and pathological regions, which is essential for tre

Mythos and the Unverified Cage: Z3-Based Pre-Deployment Verification for Frontier-Model Sandbox Infrastructure

Model ReleasesDGX agent

arXiv:2604.20496v1 Announce Type: cross Abstract: The April 2026 Claude Mythos sandbox escape exposed a critical weakness in frontier AI containment: the infrastructure surrounding advanced models rem

New in the Codex app: - GPT-5.5 - Browser control - Sheets & Slides - Docs & PDFs - OS-wide dictation - Auto-review mode Enjoy!

Model ReleasesDGX agent

The Codex app now includes several new features: GPT-5.5 integration, browser control capabilities, support for Google Sheets and Slides, document and PDF handling, OS-wide dictation functionality, an

'Newspaper Eat' Means 'Not Tasty': A Taxonomy and Benchmark for Coded Language in Real-World Chinese Online Reviews

Model ReleasesDGX agent

arXiv:2601.19932v2 Announce Type: replace Abstract: Coded language is an important part of human communication. It refers to cases where users intentionally encode meaning so that the surface text dif

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model

Model ReleasesDGX agent

arXiv:2604.20806v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) have made substantial advances in reasoning tasks at the Olympiad level. Nevertheless, current Olympiad-level mul

On Bayesian Softmax-Gated Mixture-of-Experts Models

Model ReleasesDGX agent

arXiv:2604.20551v1 Announce Type: cross Abstract: Mixture-of-experts models provide a flexible framework for learning complex probabilistic input-output relationships by combining multiple expert mode

One day while testing GPT-5.5, I had my first taste of AGI. We had a branch with hundreds of visual and front-end changes, plus complex refa…

Model ReleasesDGX agent

One day while testing GPT-5.5, I had my first taste of AGI. We had a branch with hundreds of visual and front-end changes, plus complex refactors. At the same time, main had changed a lot too. Conflic

🔹 One Prompt → 100-page PDF report +cited dataset + 30-page executive PPT + 20 financial charts

Model ReleasesDGX agent

Moonshot's Kimi AI demonstrated capabilities to generate comprehensive business reports from a single prompt, including a 100-page PDF with citations, a 30-page executive PowerPoint presentation, and

ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence

Model ReleasesDGX agent

arXiv:2604.20719v1 Announce Type: cross Abstract: Omnimodal Notation Processing (ONP) represents a unique frontier for omnimodal AI due to the rigorous, multi-dimensional alignment required across aud

🚨 OpenAI just launched GPT-5.5. The OpenAI team was nice enough to give me early access over the last several weeks, and I just want to fla…

Model ReleasesDGX agent

🚨 OpenAI just launched GPT-5.5. The OpenAI team was nice enough to give me early access over the last several weeks, and I just want to flag: there is a certain class of models (one that we’re hitting

OpenAI launches GPT-5.5, designed to handle complex tasks with minimal guidance; the model will be used to power the company's upcoming 'super app' (Rachel Metz/Bloomberg)

Model ReleasesDGX agent

Rachel Metz / Bloomberg: OpenAI launches GPT-5.5, designed to handle complex tasks with minimal guidance; the model will be used to power the company's upcoming “super app” — OpenAI is introducing an

OpenAI releases GPT-5.5 with advanced math, coding capabilities

Model ReleasesDGX agent

OpenAI Group PBC today launched a new large language model that is significantly better than its predecessors at solving math problems and writing code. GPT-5.5 is rolling out a week after rival Anthr

← Previous
1…316317318319320…373
Next →