AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,563 results
9 Jul 2026

The Key to Going Linear: Analysis-Driven Transformer Linearization

Model ReleasesDGX agent

arXiv:2607.07706v1 Announce Type: new Abstract: The quadratic cost of causal self-attention severely bottlenecks long-context transformer inference. While numerous post hoc linearization pipelines exi

The metrics discussion at OpenAI is a little confusing to me. I appreciate the clarification about bad benchmarks, but they spent a lot of m…

Model ReleasesDGX agent

The metrics discussion at OpenAI is a little confusing to me. I appreciate the clarification about bad benchmarks, but they spent a lot of money developing a very good benchmark of autonomous model ab

The new GPT-5.6 family: Luna, Terra, Sol

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

OpenAI's latest flagship model hit general availability this morning, and comes in three sizes: Luna, Terra, and Sol (from smallest to largest). The new models are priced per 1M input/output tokens as

The Solutions team @MistralAI is hiring globally! If you’re entrepreneurial, hands-on, and want to shape how enterprises adopt AI, let’s tal…

Model ReleasesDGX agent

The Solutions team @MistralAI is hiring globally! If you’re entrepreneurial, hands-on, and want to shape how enterprises adopt AI, let’s talk. Apply: https://mistral.ai/careers/?utm_source=linkedin&ut

The upcoming wave of SpaceXAI Grok updates is insane Grok 4.5: The 1.5T foundation model is being refined almost daily, and its context wind…

Model ReleasesDGX agent

The upcoming wave of SpaceXAI Grok updates is insane Grok 4.5: The 1.5T foundation model is being refined almost daily, and its context window is expected to jump to 1M tokens, possibly as soon as nex

Think Big, Search Small: Where Capacity Matters in Hierarchical Search Agents?

Model ReleasesDGX agent

arXiv:2607.07548v1 Announce Type: new Abstract: Large language model based search agents increasingly adopt multi-agent architectures in which a main agent decomposes a complex question into sub-queri

Thinking Ahead: Foresight Intelligence in MLLMs and World Model

Model ReleasesDGX agent

arXiv:2511.18735v3 Announce Type: replace-cross Abstract: In this work, we define Foresight Intelligence as the capability to anticipate and interpret future events-an ability essential for applicatio

This is confusing... Claude Cowork was a purposefully more secure (& slightly weaker) alternative to Code designed so non-coders couldn't ca…

Model ReleasesDGX agent

This is confusing... Claude Cowork was a purposefully more secure (& slightly weaker) alternative to Code designed so non-coders couldn't cause too much trouble. I got that. But I don't understand wha

TimEE: End-to-end Time Series Classification via In-Context Learning

Model ReleasesDGX agent

arXiv:2607.07500v1 Announce Type: cross Abstract: Time series classification (TSC) is dominated by a two-stage paradigm: train a feature encoder -- either from scratch on the target dataset or via pre

To be frank, at @QodoAI we were cautious about the last few GPT-5.x upgrades. GPT-5.6 is a clear-cut upgrade. Better code-review quality, fe…

Model ReleasesDGX agent

To be frank, at @QodoAI we were cautious about the last few GPT-5.x upgrades. GPT-5.6 is a clear-cut upgrade. Better code-review quality, fewer tokens, lower latency. @OpenAI cooked here. obviously th

TRACE-Seg3D: Counterfactual Context Auditing For Robust 3D Glioma Segmentation Under Institutional Shift

Model ReleasesDGX agent

arXiv:2607.07038v1 Announce Type: new Abstract: Medical image segmentation models can achieve strong benchmark performance while remaining sensitive to scanner, protocol, and institutional variation.

Tree-of-Thoughts Reasoning for Text-to-Image In-Context Learning

Model ReleasesDGX agent

arXiv:2607.07117v1 Announce Type: cross Abstract: In text-to-image in-context learning (T2I-ICL), a model has to infer a latent compositional pattern from fewshot demonstrations for generating a query

Try @Grok 4.5!

Model ReleasesDGX agent

Try @Grok 4.5! Grok 4.5 is the top non-Anthropic model on AA-Briefcase, combining frontier agentic knowledge work capabilities with leading cost and time-efficiency Yesterday @SpaceXAI released Grok 4

Try it now in Devin Cloud, Devin Desktop, and Devin CLI. Read more at https://devin.ai/blog/gpt-5-6

Model ReleasesDGX agent

Cognition AI announced new features or capabilities for Devin (their AI software engineer tool) that are now available across Devin Cloud, Devin Desktop, and Devin CLI, with additional details availab

Two-Stage Multi-Modal Fusion with Adaptive Alignment for Action Quality Assessment

Model ReleasesDGX agent

arXiv:2607.07438v1 Announce Type: new Abstract: Action Quality Assessment (AQA) aims to evaluate how well a person performs a movement, which is essential in applications such as sports scoring, skill

u guys were clowning on @greptile but turns out they were just the inspo for @openai to go so, so much harder https://x.com/JangLawrenceK/st…

Model ReleasesDGX agent

u guys were clowning on @greptile but turns out they were just the inspo for @openai to go so, so much harder https://x.com/JangLawrenceK/status/2075204015890325703 guys I just cancelled my Claude pla

Unraveling Machine Behavior by Multi-Level Bias Analysis and Detection: Methodology and Application to Computer Vision

Model ReleasesDGX agent

arXiv:2607.07236v1 Announce Type: new Abstract: This study investigates the presence and propagation of bias within Neural Networks through a comprehensive multi-level analysis spanning the learned la

VCDP: Variation-Conditioned Distributional Proxy Learning for Semi-Supervised Medical Image Segmentation

Model ReleasesDGX agent

arXiv:2607.07416v1 Announce Type: new Abstract: Semi-supervised 3D medical image segmentation reduces the need for dense voxel-level annotations by exploiting unlabeled volumes. Although existing meth

Video2Reaction: Mapping Video to Audience Reaction Distribution in the Wild

Model ReleasesDGX agent

arXiv:2607.06875v1 Announce Type: new Abstract: Understanding and forecasting audience reactions to video content are crucial for improving content creation, recommendation systems, and media analysis

Watch Ollama's @jmorgan and @peterfenton on @CNBC with @dee_bosa at 12pm PT / 3pm ET to discuss 'The Post-Frontier Era' and 'Open-Source AI'…

Model ReleasesDGX agent

Watch Ollama's @jmorgan and @peterfenton on @CNBC with @dee_bosa at 12pm PT / 3pm ET to discuss 'The Post-Frontier Era' and 'Open-Source AI's Breakout Moment' Benchmark’s @peterfenton says 90%+ of tok

We built a plugin that traces every Claude Code session straight into LangSmith. Three commands, one JSON block, and every message, tool cal…

Model ReleasesDGX agent

We built a plugin that traces every Claude Code session straight into LangSmith. Three commands, one JSON block, and every message, tool call, and subagent run shows up as an inspectable trace. Setup

We comprehensively benchmarked GPT-5.6 on document understanding. At a high-level there's no change between GPT-5.6 Sol and GPT-5.5 in terms…

Model ReleasesDGX agent

We comprehensively benchmarked GPT-5.6 on document understanding. At a high-level there's no change between GPT-5.6 Sol and GPT-5.5 in terms of performance over tables, text, charts, layout, and more.

What Predicts Correctness in Text-to-SQL? A Selective-Prediction Study

Model ReleasesDGX agent

arXiv:2607.06799v1 Announce Type: cross Abstract: Evaluating uncertainty in AI-generated SQL queries requires estimating whether a query is correct, where correct means it executes to the same result

What's on My Network? Using Large Language Models to Identify Real-World IoT Devices at Scale

Model ReleasesDGX agent

arXiv:2510.13817v2 Announce Type: replace Abstract: The growth of IoT devices in shared environments has outpaced our ability to identify them, posing urgent risks to privacy, safety, and accountabili

Where to Intervene? Benchmarking Fairness-Aware Learning on Differentially Private Synthetic Tabular Data

Model ReleasesDGX agent

arXiv:2607.07471v1 Announce Type: cross Abstract: Machine learning models are increasingly deployed in high-stakes domains, raising concerns about both privacy and fairness. Differential Privacy (DP)

Wordle 1,845 5/6 ⬛⬛⬛⬛🟨 ⬛⬛🟨⬛🟨 🟨⬛🟨🟨⬛ ⬛🟨⬛🟨🟩 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This post documents a Wordle game result (puzzle #1,845) played by Anthropic, showing the progression of guesses through color-coded tile feedback until reaching the correct five-letter word solution

8 Jul 2026

1/3 Best-of-N leaves $$ on the table by not accounting for variance in task difficulty. We built budget-aware execution: turn the dial on co…

Model ReleasesDGX agent

1/3 Best-of-N leaves $$ on the table by not accounting for variance in task difficulty. We built budget-aware execution: turn the dial on compute or speed, while keeping quality constant, to save cost

2/3 By building a reliable early stopping mechanism, we could apply cascading (save up to 44% compute by not running unnecessary rollouts ) …

Model ReleasesDGX agent

2/3 By building a reliable early stopping mechanism, we could apply cascading (save up to 44% compute by not running unnecessary rollouts ) or parallel execution (up to 25% faster by sparing wait time

A good voice model should be enjoyable to talk to, and GPT-Live is a great conversationalist with a more natural and defined personality tha…

Model ReleasesDGX agent

GPT-Live is OpenAI's voice model designed to be an engaging conversational partner with natural speech and a distinct personality. The model prioritizes making interactions enjoyable for users through

A VLM-Enhanced Framework for Comprehensive Traffic Sign Condition Assessment Integrating Daytime Visual Performance and Nighttime Retroreflectivity Evaluation

Model ReleasesDGX agent

arXiv:2607.06478v1 Announce Type: new Abstract: Traffic signs are crucial components of road safety, serving as visual tools under all lighting conditions. The Manual on Uniform Traffic Control Device

AbICL: In-Context Learning for Antigen-Specific Antibody Affinity Ranking

Model ReleasesDGX agent

arXiv:2607.05846v1 Announce Type: cross Abstract: Accurate ranking of antibody candidates according to their binding affinity is essential for therapeutic antibody discovery. However, existing methods

Added Grok 4.5 to the Harbor Town Playable gallery. Here is it's 'build me a procedurally generated 3D simulation showing the evolution of a…

Model ReleasesDGX agent

Added Grok 4.5 to the Harbor Town Playable gallery. Here is it's 'build me a procedurally generated 3D simulation showing the evolution of a harbor town from 3000 BCE to 3000 AD, it should look beauti

aiAuthZ: Off-Host, Identity-Bound Authorization for AI Agents

Model ReleasesDGX agent

arXiv:2607.05518v1 Announce Type: cross Abstract: AI agents issue tool calls on the basis of text they cannot verify, so any party who controls part of the context can forge the appearance of authorit

AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language Models

Model ReleasesDGX agent

arXiv:2607.06485v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly deployed on infrared (IR) remote sensing imagery in security-critical settings, yet their adversarial r

An Experimental Design Approach to Evaluating Agentic AI's Autonomous Model Discovery

Model ReleasesDGX agent

arXiv:2607.06413v1 Announce Type: cross Abstract: Large language model coding agents increasingly perform open-ended data modeling and analysis. These agents are stochastic and adaptive, and therefore

Announcing Robostral Navigate, our first model for embodied navigation: an 8B robotics navigation model that guides robots to autonomously p…

Model ReleasesDGX agent

Announcing Robostral Navigate, our first model for embodied navigation: an 8B robotics navigation model that guides robots to autonomously perform tasks specified with natural language. Single RGB cam

APVI-SLAM: Real-Time Acoustic-Pressure-Visual-Inertial Localization and Photorealistic Mapping System in Complex Underwater Environment

Model ReleasesDGX agent

arXiv:2607.06222v1 Announce Type: new Abstract: Extreme subsea environments often cause severe feature de-gradation and estimator divergence in underwater visual-inertial SLAM. Although sensors like D

Arkenstone Defense launches with $35M to help startups sell to the Pentagon

Model ReleasesDGX agent

Federal contracting software startup Arkenstone Defense Inc. formally launched today with 35 million in new funding to take on the compliance, security and workforce setup that keeps many commercial t

Artificial Analysis assessment

Model ReleasesDGX agent

Artificial Analysis assessment SpaceXAI just released Grok 4.5, and it ranks #4 on GDPval-AA v2 with an Elo of 1543 - behind only the latest Claude releases from Anthropic on real-world agentic knowle

ArtisanCAD: An Industrial-Level CAD Agent with Expert-Grounded Knowledge Distillation

Model ReleasesDGX agent

arXiv:2607.05750v1 Announce Type: new Abstract: Computer-aided design (CAD) for industrial components requires long-horizon procedural modeling, robust feature dependencies, editable parametric geomet

As coding models improve, evals need to become harder, fairer, and more trustworthy. Better benchmarks help the field understand real progre…

Model ReleasesDGX agent

As AI coding models advance in capability, current evaluation benchmarks must evolve to remain challenging and meaningful measures of progress. OpenAI argues that improved benchmarks need to be harder

AtomBench: A Benchmarking Framework for Generative Crystal Reconstruction Models in Conventional Superconductors

Model ReleasesDGX agent

arXiv:2510.16165v2 Announce Type: replace Abstract: A key question in benchmarking generative crystal reconstruction models is how the amount and type of crystallographic information provided to a gen

Auditing of Unlearning Algorithms

Model ReleasesDGX agent

arXiv:2607.05898v1 Announce Type: new Abstract: Evaluating whether unlearning algorithms truly remove training data influence remains an open challenge. We propose a practical auditor that computes da

Auto-DSM Under the Lens: A Black-Box Evaluation Framework for LLM-Based DSM Generation

Model ReleasesDGX agent

arXiv:2607.05985v1 Announce Type: new Abstract: This paper presents a black-box evaluation framework to systematically assess the ability of Large Language Models (LLMs) to generate Design Structure M

BabyVision: Visual Reasoning Beyond Language

Model ReleasesDGX agent

arXiv:2601.06521v2 Announce Type: replace-cross Abstract: While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic

BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension

Model ReleasesDGX agent

arXiv:2607.05614v1 Announce Type: cross Abstract: Document comprehension is a challenging yet impactful task for Multimodal Large Language Models, especially as these systems see growing adoption in r

Benchmarking KV-Cache Optimizations across Task Quality and System Performance for Long-Context Serving

Model ReleasesDGX agent

arXiv:2607.05399v1 Announce Type: cross Abstract: Large language model serving is increasingly limited by KV-cache growth under long-context workloads, yet existing KV-cache compression techniques are

Benchmarking the Robustness of Autonomous Driving to Environmental Illusions: A Lane Perception Perspective

Model ReleasesDGX agent

arXiv:2607.05783v1 Announce Type: new Abstract: Environmental illusions (eg., shadows, reflections, and tire marks) are naturally existing yet overlooked phenomena in real-world driving environments.

Beyond Heuristic Tuning: Power-Calibrated LLM Watermarking

Model ReleasesDGX agent

arXiv:2607.05694v1 Announce Type: cross Abstract: Logit-based watermarking is a widely used mechanism for identifying LLM generated content, yet its effectiveness is governed by a fundamental trade-of

Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents

Model ReleasesDGX agent

arXiv:2510.19771v4 Announce Type: replace Abstract: LLM-based agents are increasingly moving towards proactivity: rather than awaiting instruction, they exercise agency to anticipate user needs and so

Beyond Refusal: A Same-Lineage Study of Aligned and Abliterated LLMs for Vulnerability Analysis

Model ReleasesDGX agent

arXiv:2607.05842v1 Announce Type: cross Abstract: Large language model (LLM)-assisted software security operates at a difficult boundary: the vulnerability-analysis terminology needed for legitimate c

Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents

Model ReleasesDGX agent

arXiv:2607.05775v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly evaluated on their ability to use tools, plan multi-step tasks, coordinate with other agents, and ope

BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech

Model ReleasesDGX agent

arXiv:2607.06054v1 Announce Type: cross Abstract: Off-the-shelf TTS systems are poorly adapted to Taiwanese Mandarin. Their accent defaults to other Mandarin variants, their tokenizers over-segment co

Box Agent uses @LangChain’s Deep Agents harness to bring specialized agents into the enterprise content platform.w @NVIDIA + LangChain’s wor…

Model ReleasesDGX agent

Box Agent uses @LangChain’s Deep Agents harness to bring specialized agents into the enterprise content platform.w @NVIDIA + LangChain’s work with Nemotron 3 Ultra reinforces where AI is headed: open,

Breaking Structural Isolation: Scalable Graph Clustering via Community-Aware Sampling and Structural Entropy

Model ReleasesDGX agent

arXiv:2607.05469v1 Announce Type: cross Abstract: Unsupervised graph clustering is a fundamental technique for uncovering underlying semantic patterns in large-scale networks. Although Graph Contrasti

Building and connecting a production-ready ecommerce MCP server using Amazon Bedrock AgentCore and Mistral AI Studio

Model ReleasesDGX agent

In this post, you build and connect that server end to end. You will implement MCP tools, set up two-layer JSON Web Token (JWT) authentication, deploy with AWS Cloud Development Kit (AWS CDK), and con

Busy couple days ahead! Today, @nvidia Nemotron 3 Ultra support for @LangChain Deep Agents at a fraction of the cost -w/ Quick Start Prompt …

Model ReleasesDGX agent

Busy couple days ahead! Today, @nvidia Nemotron 3 Ultra support for @LangChain Deep Agents at a fraction of the cost -w/ Quick Start Prompt Tomorrow is Wikimania🔥 @hwchase17 is chatting with with @Bra

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models

Model ReleasesDGX agent

arXiv:2607.06534v1 Announce Type: new Abstract: Existing 3D scene-grounded Large Language Models (3D-LLMs) focus on answering questions grounded in simplified single-room 3D scenes, lacking the abilit

CANONIC: Governance Is Compilation

Model ReleasesDGX agent

arXiv:2607.05410v1 Announce Type: cross Abstract: We present CANONIC: governed intelligence that compiles digital artifacts into an evidence ledger at scale. Large language models generate prose faste

ChatGPT’s upgraded voice mode is better at shutting up

Model ReleasesDGX agent

OpenAI is overhauling ChatGPT's voice mode with a new model that it says is more like 'talking to another person.' The new GPT-Live-1 is designed to interrupt you less and will also wait for you to co

← Previous
1…9293949596…377
Next →