AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
All
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,340 results
Model Releases

Learning When Not to Decide: A Framework for Overcoming Factual Presumptuousness in AI Adjudication

DGX agent

arXiv:2604.19895v1 Announce Type: new Abstract: A well-known limitation of AI systems is presumptuousness: the tendency of AI systems to provide confident answers when information may be lacking. This

model-releasesarxiv-cs-ai
23 Apr 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Less Languages, Less Tokens: An Efficient Unified Logic Cross-lingual Chain-of-Thought Reasoning Framework

DGX agent

arXiv:2604.20090v1 Announce Type: new Abstract: Cross-lingual chain-of-thought (XCoT) with self-consistency markedly enhances multilingual reasoning, yet existing methods remain costly due to extensiv

model-releasesarxiv-cs-cl
23 Apr 2026
Model Releases

LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent

DGX agent

arXiv:2604.17931v2 Announce Type: replace Abstract: Reinforcement Learning (RL) has emerged as a powerful training paradigm for LLM-based agents. However, scaling agentic RL for deep research remains

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

LLAMADRS: Evaluating Open-Source LLMs on Real Clinical Interviews--To Reason or Not to Reason?

DGX agent

arXiv:2501.03624v2 Announce Type: replace-cross Abstract: Large language models (LLMs) excel on many NLP benchmarks, but their behavior on real-world, semi-structured prediction remains underexplored.

model-releasesarxiv-cs-cl
23 Apr 2026
Model Releases

LLM-guided phase diagram construction through high-throughput experimentation

DGX agent

arXiv:2604.20304v1 Announce Type: cross Abstract: Constructing phase diagrams for multicomponent alloys requires extensive experimental measurements and is a time-consuming task. Here we investigate w

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

looks like new Pareto frontiers across everything: - Context: 400K context in Codex and a 1M in API - API Pricing: 5/m input and 30/m outp…

DGX agent

looks like new Pareto frontiers across everything: - Context: 400K context in Codex and a 1M in API - API Pricing: 5/m input and 30/m output tokens. - Codex improved its own inference speed 20% lol -

model-releasesswyx--x
23 Apr 2026
Model Releases

LoRA-FA: Efficient and Effective Low Rank Representation Fine-tuning

DGX agent

arXiv:2308.03303v2 Announce Type: replace Abstract: Fine-tuning large language models (LLMs) is crucial for improving their performance on downstream tasks, but full-parameter fine-tuning (Full-FT) is

model-releasesarxiv-cs-cl
23 Apr 2026
Model Releases

MambaLiteUNet: Cross-Gated Adaptive Feature Fusion for Robust Skin Lesion Segmentation

DGX agent

arXiv:2604.20286v1 Announce Type: cross Abstract: Recent segmentation models have demonstrated promising efficiency by aggressively reducing parameter counts and computational complexity. However, the

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

MAPRPose: Mask-Aware Proposal and Amodal Refinement for Multi-Object 6D Pose Estimation

DGX agent

arXiv:2604.20650v1 Announce Type: new Abstract: 6D object pose estimation in cluttered scenes remains challenging due to severe occlusion and sensor noise. We propose MAPRPose, a two-stage framework t

model-releasesarxiv-cs-cv
23 Apr 2026
Model Releases

Measuring the Machine: Evaluating Generative AI as Pluralist Sociotechical Systems

DGX agent

arXiv:2604.20545v1 Announce Type: new Abstract: In measurement theory, instruments do not simply record reality; they help constitute what is observed. The same holds for generative AI evaluation: ben

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

Meta-Tool: Efficient Few-Shot Tool Adaptation for Small Language Models

DGX agent

arXiv:2604.20148v1 Announce Type: cross Abstract: Can small language models achieve strong tool-use performance without complex adaptation mechanisms? This paper investigates this question through Met

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

MetaboNet: The Largest Publicly Available Consolidated Dataset for Type 1 Diabetes Management

DGX agent

arXiv:2601.11505v2 Announce Type: replace-cross Abstract: Progress in Type 1 Diabetes (T1D) algorithm development is limited by the fragmentation and lack of standardization across existing T1D manage

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

Microsoft launches ‘vibe working’ in Word, Excel, and PowerPoint

DGX agent

Microsoft is rolling out a new Agent Mode inside Office apps like Word, Excel, and PowerPoint this week. Previously described by Microsoft as 'vibe working,' the Agent Mode is a more powerful version

model-releasesthe-verge-ai
23 Apr 2026
Model Releases

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models

DGX agent

arXiv:2604.19809v1 Announce Type: new Abstract: We introduce MIRROR, a benchmark comprising eight experiments across four metacognitive levels that evaluates whether large language models can use self

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror

DGX agent

arXiv:2604.14785v2 Announce Type: replace Abstract: Recent progress in Multimodal Large Language Models (MLLMs) has demonstrated remarkable advances in perception and reasoning, suggesting their poten

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

Mitigating Hallucinations in Large Vision-Language Models without Performance Degradation

DGX agent

arXiv:2604.20366v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) exhibit powerful generative capabilities but frequently produce hallucinations that compromise output reliability.

model-releasesarxiv-cs-cv
23 Apr 2026
Model Releases

Mitigating Prompt-Induced Cognitive Biases in General-Purpose AI for Software Engineering

DGX agent

arXiv:2604.16756v2 Announce Type: replace-cross Abstract: Prompt-induced cognitive biases are changes in a general-purpose AI (GPAI) system's decisions caused solely by biased wording in the input (e.

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design

DGX agent

arXiv:2412.14590v2 Announce Type: replace Abstract: Quantization has become one of the most effective methodologies to compress LLMs into smaller size. However, the existing quantization solutions sti

model-releasesarxiv-cs-lg
23 Apr 2026
Model Releases

MLG-Stereo: ViT Based Stereo Matching with Multi-Stage Local-Global Enhancement

DGX agent

arXiv:2604.20393v1 Announce Type: new Abstract: With the development of deep learning, ViT-based stereo matching methods have made significant progress due to their remarkable robustness and zero-shot

model-releasesarxiv-cs-cv
23 Apr 2026
Model Releases

Model Capability Assessment and Safeguards for Biological Weaponization

DGX agent

arXiv:2604.19811v1 Announce Type: cross Abstract: AI leaders and safety reports increasingly warn that advances in model reasoning may enable biological misuse, including by low-expertise users, while

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

MSLAU-Net: A Hybrid CNN-Transformer Network for Medical Image Segmentation

DGX agent

arXiv:2505.18823v2 Announce Type: replace Abstract: Accurate medical image segmentation allows for the precise delineation of anatomical structures and pathological regions, which is essential for tre

model-releasesarxiv-cs-cv
23 Apr 2026
Model Releases

Mythos and the Unverified Cage: Z3-Based Pre-Deployment Verification for Frontier-Model Sandbox Infrastructure

DGX agent

arXiv:2604.20496v1 Announce Type: cross Abstract: The April 2026 Claude Mythos sandbox escape exposed a critical weakness in frontier AI containment: the infrastructure surrounding advanced models rem

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

New in the Codex app: - GPT-5.5 - Browser control - Sheets & Slides - Docs & PDFs - OS-wide dictation - Auto-review mode Enjoy!

DGX agent

The Codex app now includes several new features: GPT-5.5 integration, browser control capabilities, support for Google Sheets and Slides, document and PDF handling, OS-wide dictation functionality, an

model-releasessam-altman--x
23 Apr 2026
Model Releases

'Newspaper Eat' Means 'Not Tasty': A Taxonomy and Benchmark for Coded Language in Real-World Chinese Online Reviews

DGX agent

arXiv:2601.19932v2 Announce Type: replace Abstract: Coded language is an important part of human communication. It refers to cases where users intentionally encode meaning so that the surface text dif

model-releasesarxiv-cs-cl
23 Apr 2026
Model Releases

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model

DGX agent

arXiv:2604.20806v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) have made substantial advances in reasoning tasks at the Olympiad level. Nevertheless, current Olympiad-level mul

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

On Bayesian Softmax-Gated Mixture-of-Experts Models

DGX agent

arXiv:2604.20551v1 Announce Type: cross Abstract: Mixture-of-experts models provide a flexible framework for learning complex probabilistic input-output relationships by combining multiple expert mode

model-releasesarxiv-cs-lg
23 Apr 2026
Model Releases

One day while testing GPT-5.5, I had my first taste of AGI. We had a branch with hundreds of visual and front-end changes, plus complex refa…

DGX agent

One day while testing GPT-5.5, I had my first taste of AGI. We had a branch with hundreds of visual and front-end changes, plus complex refactors. At the same time, main had changed a lot too. Conflic

model-releasessam-altman--x
23 Apr 2026
Model Releases

🔹 One Prompt → 100-page PDF report +cited dataset + 30-page executive PPT + 20 financial charts

DGX agent

Moonshot's Kimi AI demonstrated capabilities to generate comprehensive business reports from a single prompt, including a 100-page PDF with citations, a 30-page executive PowerPoint presentation, and

model-releaseskimi-moonshot--x
23 Apr 2026
Model Releases

ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence

DGX agent

arXiv:2604.20719v1 Announce Type: cross Abstract: Omnimodal Notation Processing (ONP) represents a unique frontier for omnimodal AI due to the rigorous, multi-dimensional alignment required across aud

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

🚨 OpenAI just launched GPT-5.5. The OpenAI team was nice enough to give me early access over the last several weeks, and I just want to fla…

DGX agent

🚨 OpenAI just launched GPT-5.5. The OpenAI team was nice enough to give me early access over the last several weeks, and I just want to flag: there is a certain class of models (one that we’re hitting

model-releasesallie-k--miller--x
23 Apr 2026
Model Releases

OpenAI launches GPT-5.5, designed to handle complex tasks with minimal guidance; the model will be used to power the company's upcoming 'super app' (Rachel Metz/Bloomberg)

DGX agent

Rachel Metz / Bloomberg: OpenAI launches GPT-5.5, designed to handle complex tasks with minimal guidance; the model will be used to power the company's upcoming “super app” — OpenAI is introducing an

model-releasestechmeme
23 Apr 2026
Model Releases

OpenAI releases GPT-5.5 with advanced math, coding capabilities

DGX agent

OpenAI Group PBC today launched a new large language model that is significantly better than its predecessors at solving math problems and writing code. GPT-5.5 is rolling out a week after rival Anthr

model-releasessiliconangle
23 Apr 2026
Model Releases

OpenAI says 'GPT-5.5 matches GPT-5.4 per-token latency in real-world serving, while performing at a much higher level of intelligence' (OpenAI)

DGX agent

OpenAI: OpenAI says “GPT-5.5 matches GPT-5.4 per-token latency in real-world serving, while performing at a much higher level of intelligence” — A new class of intelligence for real work — We're relea

model-releasestechmeme
23 Apr 2026
Model Releases

OpenAI says GPT-5.5's improvements are strongest in agentic coding, computer use, and early scientific research, which require reasoning across longer contexts (Madison Mills/Axios)

DGX agent

Madison Mills / Axios: OpenAI says GPT-5.5's improvements are strongest in agentic coding, computer use, and early scientific research, which require reasoning across longer contexts — OpenAI on Thurs

model-releasestechmeme
23 Apr 2026
Model Releases

OpenAI says its new GPT-5.5 model is more efficient and better at coding

DGX agent

OpenAI just announced its new GPT-5.5 model, which the company calls its 'smartest and most intuitive to use model yet, and the next step toward a new way of getting work done on a computer.' OpenAI j

model-releasesthe-verge-ai
23 Apr 2026
Model Releases

OpenAI’s New GPT-5.5 Powers Codex on NVIDIA Infrastructure — and NVIDIA Is Already Putting It to Work

DGX agent

AI agents have revolutionized developer workflows, and their next frontier is knowledge work: processing information, solving complex problems, coming up with new ideas and driving innovation. Codex,

model-releasesnvidia-blog
23 Apr 2026
Model Releases

Optimal Single-Policy Sample Complexity and Transient Coverage for Average-Reward Offline RL

DGX agent

arXiv:2506.20904v2 Announce Type: replace Abstract: We study offline reinforcement learning in average-reward MDPs, which presents increased challenges from the perspectives of distribution shift and

model-releasesarxiv-cs-lg
23 Apr 2026
Model Releases

Option Pricing on Noisy Intermediate-Scale Quantum Computers: A Quantum Neural Network Approach

DGX agent

arXiv:2604.19832v1 Announce Type: cross Abstract: In a global derivatives market with notional values in the hundreds of trillions of dollars, the accuracy and efficiency of pricing models are of fund

model-releasesarxiv-cs-lg
23 Apr 2026
Model Releases

Our own team was trying to figure out how to get LiteParse working in the browser 😂 Shouldn't have doubted Claude. Claude knows best.

DGX agent

Our own team was trying to figure out how to get LiteParse working in the browser 😂 Shouldn't have doubted Claude. Claude knows best. LiteParse is really neat! It does a great job of extracting text f

model-releasesjerry-liu--x
23 Apr 2026
Model Releases

Over the past month, some of you reported Claude Code's quality had slipped. We investigated, and published a post-mortem on the three issue…

DGX agent

Over the past month, some of you reported Claude Code's quality had slipped. We investigated, and published a post-mortem on the three issues we found. All are fixed in v2.1.116+ and we’ve reset usage

model-releasesthariq--x
23 Apr 2026
Model Releases

OVPD: A Virtual-Physical Fusion Testing Dataset of OnSite Auton-omous Driving Challenge

DGX agent

arXiv:2604.20423v1 Announce Type: new Abstract: The rapid iteration of autonomous driving algorithms has created a growing demand for high-fidelity, replayable, and diagnosable testing data. However,

model-releasesarxiv-cs-ro
23 Apr 2026
Model Releases

Parallel-SFT: Improving Zero-Shot Cross-Programming-Language Transfer for Code RL

DGX agent

arXiv:2604.20835v1 Announce Type: new Abstract: Modern language models demonstrate impressive coding capabilities in common programming languages (PLs), such as C++ and Python, but their performance i

model-releasesarxiv-cs-cl
23 Apr 2026
Model Releases

ParseBench is now live on @Kaggle. The first document OCR benchmark built for AI agents — 2,000 enterprise pages, 167K+ test rules, 5 dimens…

DGX agent

ParseBench is now live on @Kaggle. The first document OCR benchmark built for AI agents — 2,000 enterprise pages, 167K+ test rules, 5 dimensions that actually break downstream agents. Benchmark your p

model-releasesjerry-liu--x
23 Apr 2026
Model Releases

Peer-Preservation in Frontier Models

DGX agent

arXiv:2604.19784v1 Announce Type: cross Abstract: Recently, it has been found that frontier AI models can resist their own shutdown, a behavior known as self-preservation. We extend this concept to th

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

PipeMFL-240K: A Large-scale Dataset and Benchmark for Object Detection in Pipeline Magnetic Flux Leakage Imaging

DGX agent

arXiv:2602.07044v2 Announce Type: replace-cross Abstract: Pipeline integrity is critical to industrial safety and environmental protection, with Magnetic Flux Leakage (MFL) detection being a primary n

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

PLR: Plackett-Luce for Reordering In-Context Learning Examples

DGX agent

arXiv:2603.21373v2 Announce Type: replace-cross Abstract: In-context learning (ICL) adapts large language models by conditioning on a small set of ICL examples, avoiding costly parameter updates. Amon

model-releasesarxiv-cs-cl
23 Apr 2026
Model Releases

PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance

DGX agent

arXiv:2604.20834v1 Announce Type: new Abstract: Recent advances in Vision-Language-Action (VLA) models have opened new avenues for robot manipulation, yet existing methods exhibit limited efficiency a

model-releasesarxiv-cs-ro
23 Apr 2026
Model Releases

Portal26 launches Agentic Token Controls to cap runaway AI agent spend

DGX agent

Generative artificial intelligence security startup Portal26 Inc. today announced the launch of a new module designed to rein in runaway token consumption by autonomous AI agents, a problem the compan

model-releasessiliconangle
23 Apr 2026
← Previous
1…397398399400401…466
Next →