AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,555 results
8 Jul 2026

Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding

Model ReleasesDGX agent

arXiv:2607.05722v1 Announce Type: new Abstract: We introduce Nemotron-Labs-Diffusion, a tri-mode language model (LM) that unifies AR, diffusion, and self-speculation decoding within a single architect

New in Claude Code: /checkup Run /checkup to: 1. Clean up unused skills/MCPs/plugins and save context 2. Dedup your local CLAUDE.md against …

Model ReleasesDGX agent

New in Claude Code: /checkup Run /checkup to: 1. Clean up unused skills/MCPs/plugins and save context 2. Dedup your local CLAUDE.md against the checked in CLAUDE.md 3. Break up root CLAUDE.md into nes

Nice stats on usage of open models across OpenCode. GLM-5.2 is still underrated, but one of the models that has really surprised me on agent…


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

Nice stats on usage of open models across OpenCode. GLM-5.2 is still underrated, but one of the models that has really surprised me on agentic tasks is deepseek-v4-flash. Extremely cheap and effective

nope! OpenAI would not have the nerve.

Model ReleasesDGX agent

nope! OpenAI would not have the nerve. Hey @GaryMarcus, quick question. Did you get months of early access to OpenAI's GPT-5.6 Sol like every AI influencer on X apparently did? Asking because their re

Not another demo or benchmark. @ShoucongChen is a senior member of our technical staff. A real project, scoped at 1-month. Delivered in 4 da…

Model ReleasesDGX agent

Not another demo or benchmark. @ShoucongChen is a senior member of our technical staff. A real project, scoped at 1-month. Delivered in 4 days with GLM5.2 Fast. The best devs deserve >400 t/sec. Take

NVIDIA Nemotron Achieves Benchmark-Leading Performance With LangChain Deep Agents Harness

Model ReleasesDGX agent

NVIDIA Nemotron 3 Ultra is offering leading performance at lower cost than top closed models with the largest and most widely adopted AI agent orchestration platform. LangChain tuned its Deep Agents h

Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis

Model ReleasesDGX agent

arXiv:2602.00846v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) struggle with alignment due to the limitations of existing reward models (RMs), which are predominantly vis

Onnes: A Physics-Grounded Multi-Agent LLM Simulator for Cryogenic Fault Diagnosis in Quantum Computing Infrastructure

Model ReleasesDGX agent

arXiv:2607.05805v1 Announce Type: new Abstract: Dilution refrigerators are the enabling infrastructure of superconducting quantum computers, yet their fault diagnosis is still dominated by threshold a

- OpenAI continually outperforms Anthropic models on computer use (with Claude, I’m surprised when it works, with Codex, I expect it to work…

Model ReleasesDGX agent

- OpenAI continually outperforms Anthropic models on computer use (with Claude, I’m surprised when it works, with Codex, I expect it to work) - when I prompt with Fable 5.5, I feel like I’m motivating

OpenAI launches GPT-Live, a new generation of voice models built on a full-duplex architecture, meaning they can listen and speak at the same time (OpenAI)

Model ReleasesDGX agent

OpenAI: OpenAI launches GPT-Live, a new generation of voice models built on a full-duplex architecture, meaning they can listen and speak at the same time — A new generation of voice models for natura

OpenAI launches GPT-Live voice model series ahead of broad GPT-5.6 release

Model ReleasesDGX agent

OpenAI Group PBC today introduced GPT-Live, a family of artificial intelligence models optimized to process spoken instructions. The model series will power ChatGPT’s voice mode. Additionally, OpenAI

OpenAI says it found widespread task issues in SWE-Bench Pro, estimates ~30% of tasks are broken, and retracts its earlier recommendation to adopt the benchmark (OpenAI)

Model ReleasesDGX agent

OpenAI: OpenAI says it found widespread task issues in SWE-Bench Pro, estimates ~30% of tasks are broken, and retracts its earlier recommendation to adopt the benchmark — Through a detailed audit, we

Optimized Adaptive Loop Filter in Versatile Video Coding

Model ReleasesDGX agent

arXiv:2607.05737v1 Announce Type: new Abstract: In the Versatile Video Coding~(VVC) standard, adaptive loop filter~(ALF), including Geometry transformation-based Adaptive Loop Filter~(GALF) and Cross

OrchardBench: A Physically-Grounded, GPU-Parallel Apple-Orchard Simulation Benchmark for Agricultural Robotics

Model ReleasesDGX agent

arXiv:2607.06337v1 Announce Type: cross Abstract: Robotic tree-fruit harvesting is a flagship problem for agricultural automation, but progress is bottlenecked by the cost and irreproducibility of fie

Our audit of SWE-Bench Pro found that a meaningful share of public tasks contain issues that can distort results. Some correct solutions fai…

Model ReleasesDGX agent

Our audit of SWE-Bench Pro found that a meaningful share of public tasks contain issues that can distort results. Some correct solutions fail because of hidden requirements, contradictory instructions

Own the frontier: 'In our evals, Nemotron 3 Ultra with a tuned LangChain Deep Agents harness provides advanced agent performance at a much l…

Model ReleasesDGX agent

Own the frontier: 'In our evals, Nemotron 3 Ultra with a tuned LangChain Deep Agents harness provides advanced agent performance at a much lower inference cost. The main takeaway is that agent perform

Parameter-Free Encoders Remain Viable for RDB Foundation Models

Model ReleasesDGX agent

arXiv:2607.05476v1 Announce Type: new Abstract: Given a relational database (RDB) storing heterogeneous tabular information, how can we predict missing (or future) values in some target column of inte

Partial Symmetry Detection for 3D Geometry using Contrastive Learning with Geodesic Point Cloud Patches

Model ReleasesDGX agent

arXiv:2312.08230v2 Announce Type: replace Abstract: Detecting partial extrinsic symmetry in 3D geometry is a fundamental yet persistent challenge in computer vision and graphics, critical for tasks ra

PatchOptic for Shared-State LLM Workflows with Projected Views and Verified Structured Updates

Model ReleasesDGX agent

arXiv:2607.05483v1 Announce Type: cross Abstract: Agentic workflows often operate over shared, structured state. Because LLM context windows are limited, each model invocation is typically shown only

PCBWorld: A Benchmark Environment for Engine-Grounded PCB Design Automation

Model ReleasesDGX agent

arXiv:2607.05915v1 Announce Type: new Abstract: PCB routing is the task of connecting the nets of a board with copper traces under strict design rules, yet learning-based methods still lag behind rule

PIPBench: A Profile-Inclusive Framework for Personalized Image Generation Evaluation

Model ReleasesDGX agent

arXiv:2607.06440v1 Announce Type: new Abstract: Recent text-to-image models such as DALLE-3 excel at following diverse prompts yet remain blind to individual aesthetic preferences. We study personaliz

Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability

Model ReleasesDGX agent

arXiv:2607.06196v1 Announce Type: new Abstract: Current AI safety evaluation and benchmarking frameworks predominantly rely on Western-centric culture-agnostic defaults that mask critical regional law

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

Model ReleasesDGX agent

arXiv:2607.05992v1 Announce Type: cross Abstract: Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heav

PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails

Model ReleasesDGX agent

arXiv:2607.05910v1 Announce Type: cross Abstract: Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an intrinsic property of an image. Rea

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents

Model ReleasesDGX agent

arXiv:2607.06008v1 Announce Type: new Abstract: Large language model (LLM) agents have shown strong performance in long-horizon tasks that require planning, tool use, and interaction with external env

Population-Level Profiling of DSM-5 Depressive Symptoms Among Self-Reported ADHD and ASD Users on Twitter: An Exploratory Study Using Advanced NLP and Statistical Analysis

Model ReleasesDGX agent

arXiv:2607.05626v1 Announce Type: new Abstract: Background: Depression frequently co-occurs with ADHD and autism spectrum disorder (ASD), but population-level differences in symptom expression between

Privilege and confidentiality in generative AI workflows

Model ReleasesDGX agent

arXiv:2607.05479v1 Announce Type: cross Abstract: Generative AI (GenAI) systems store and process client data in three distinct ways: in the model's parameters through training and memorisation, in th

Prompt-Adapter Context Routing for Parameter-Efficient Multi-Shot Long Video Extrapolation

Model ReleasesDGX agent

arXiv:2607.06481v1 Announce Type: cross Abstract: We present PACR-Video, a parameter-efficient framework for multi-shot long video extrapolation that preserves recurring entities, scene structure, vis

Prompt: ANNALS — a living kingdom in a single file Working title behavior: the app titles itself per seed — 'The Annals of Vaelmere', 'The A…

Model ReleasesDGX agent

Prompt: ANNALS — a living kingdom in a single file Working title behavior: the app titles itself per seed — 'The Annals of Vaelmere', 'The Annals of Osterholt' — because the central conceit is that yo

Qdrant Beats Elastic’s DiskBBQ at 2x Throughput, Half the Latency, and 1/3 the Compute

Model ReleasesDGX agent

TL;DR Elastic recently published a benchmark claiming that their proprietary, disk-based index (dubbed “DiskBBQ”) delivers up to 7x higher throughput than Qdrant when deployed on nodes with network-at

Quantifying Frontier LLM Capabilities for Container Sandbox Escape

Model ReleasesDGX agent

arXiv:2603.02277v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly act as autonomous agents, using tools to execute code, read and write files, and access networks, cr

Really excited to partner with @nvidia on the NemoClaw Deep Agents Blueprint Deep Agents is a fully open source agent harness that we are tu…

Model ReleasesDGX agent

Really excited to partner with @nvidia on the NemoClaw Deep Agents Blueprint Deep Agents is a fully open source agent harness that we are tuning to make perform incredibly well with open models Introd

Refiant goes where rivals only promised with a 10 million-token AI model

Model ReleasesDGX agent

Artificial intelligence optimization startup Refiant Inc. today launched Protea, a suite of long-context AI models led by a 10 million-token context window that the company says ranks among the larges

Rewriting Bun in Rust

Model ReleasesDGX agent

Rewriting Bun in Rust Jarred Sumner has been promising this blog post (since May 9th) about his Zig to Rust rewrite of Bun for significantly longer than it took him to finish the rewrite. Honestly, it

RFHNet: Relational and Frequency-Aware Hashing Network for Large-Scale Fine-Grained Food Image Retrieval

Model ReleasesDGX agent

arXiv:2607.06148v1 Announce Type: new Abstract: Fine-grained food image retrieval is a key task in computational gastronomy, with applications in food traceability, dietary monitoring, and smart cater

RPAM: A Principled Metric for Evaluating Associations in Language Models with High Predictive Validity in Downstream Outputs

Model ReleasesDGX agent

arXiv:2607.05679v1 Announce Type: cross Abstract: Language models (LMs) exhibit problematic biases, such as stereotypes. Effectively analyzing and mitigating such biases requires accurate and generali

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications

Model ReleasesDGX agent

arXiv:2607.06411v1 Announce Type: cross Abstract: Developers increasingly delegate real maintenance work to product-grade coding agents, and many state tasks in their native language, in the style of

SafeImpute: Reliable Clinical Data Imputation via Conformal Selection

Model ReleasesDGX agent

arXiv:2607.05613v1 Announce Type: new Abstract: Clinical care often relies on key laboratory indicators, yet real-world patient visits are sparse and tests are ordered irregularly, leading to pervasiv

SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place Recognition

Model ReleasesDGX agent

arXiv:2509.25723v4 Announce Type: replace Abstract: Visual Place Recognition (VPR) requires robust retrieval of geotagged images despite large appearance, viewpoint, and environmental variation. Prior

Scientific Code Search at Scale: A Multi-Domain Dataset and Benchmark

Model ReleasesDGX agent

arXiv:2607.05443v1 Announce Type: cross Abstract: Scientists increasingly rely on open-source tools to support their research workflows, yet discovering relevant software among over 600 million GitHub

Self-Review Reinforcement Learning (SRRL) with Cross-Episode Memory and Policy Distillation

Model ReleasesDGX agent

arXiv:2607.05541v1 Announce Type: cross Abstract: Reinforcement Learning is commonly used to train large language models using environmental feedback. In applied settings, the environment usually prov

Self-Routing: Parameter-Free Expert Routing from Hidden States

Model ReleasesDGX agent

arXiv:2604.00421v2 Announce Type: replace Abstract: Mixture-of-Experts (MoE) layers increase model capacity by activating only a small subset of experts per token, and typically rely on a learned rout

Separating signal from noise in coding evaluations

Model ReleasesDGX agent

This OpenAI article discusses methods for distinguishing meaningful performance indicators from random variation when evaluating AI coding capabilities. It likely covers evaluation methodologies, stat

SEVRA-BENCH: Social Engineering of Vulnerabilities in Review Agents

Model ReleasesDGX agent

arXiv:2606.13757v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed in automated code-review systems, where their approvals can determine which code is mer

Social 3D Scene Graphs: Modeling Human Actions and Relations for Interactive Service Robots

Model ReleasesDGX agent

arXiv:2509.24966v2 Announce Type: replace Abstract: Understanding how people interact with their surroundings and each other is essential for enabling robots to act in socially compliant and context-a

SpaceXAI launches Grok 4.5, its first model built in partnership with Cursor, designed to 'handle difficult, long-running' legal, finance, and coding tasks (Carmen Arroyo/Bloomberg)

Model ReleasesDGX agent

Carmen Arroyo / Bloomberg: SpaceXAI launches Grok 4.5, its first model built in partnership with Cursor, designed to “handle difficult, long-running” legal, finance, and coding tasks — SpaceXAI has un

SpaceXAI’s newest AI model Grok 4.5 dramatically undercuts Anthropic and OpenAI on price

Model ReleasesDGX agent

Elon Musk’s SpaceXAI Corp. has released a new model called Grok 4.5, in what is its first major launch since it went public a few weeks earlier. In a blog post earlier today, the company said Grok 4.5

SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation

Model ReleasesDGX agent

arXiv:2607.05721v1 Announce Type: new Abstract: Uncertainty estimation is essential not only for the trustworthy deployment of large language models (LLMs) but also as a foundation for self-refinement

Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders

Model ReleasesDGX agent

arXiv:2508.16560v4 Announce Type: replace-cross Abstract: Sparse Autoencoders (SAEs) extract features from LLM internal activations, meant to correspond to interpretable concepts. A core SAE training

Spider 2.0-AIFunc: Extending Real-World Text-to-SQL to AI-Native SQL Workflows

Model ReleasesDGX agent

arXiv:2607.06229v1 Announce Type: cross Abstract: Major cloud data platforms now expose large language model capabilities as native SQL functions, enabling analysts to perform classification, filterin

Statistically Meaningful Geometry and Gauge Symmetry Breaking: A Geometric Foundation for Scientific Discovery and Intelligence Emergence

Model ReleasesDGX agent

arXiv:2607.05436v1 Announce Type: new Abstract: The rapid scaling of over-parameterized machine learning architectures, particularly LLMs, raises a profound crisis: do these systems exhibit genuine in

StepShield: When, Not Whether to Intervene on Rogue Agents

Model ReleasesDGX agent

arXiv:2601.22136v2 Announce Type: replace-cross Abstract: Agent safety benchmarks measure whether a monitor detects harm, not when. Yet timing is the difference between intervention and autopsy. We in

Structured Data Extraction from Real Estate Documents using Clustering, Classification, and Large Language Models

Model ReleasesDGX agent

arXiv:2607.06012v1 Announce Type: new Abstract: Real estate property listings expose structured metadata through the API. Still, the richest property-level information (i.e., legal status, structural

super interesting - and a reminder that language models model language and not ideas per se. also a stark example of how easily LLMs can (as…

Model ReleasesDGX agent

super interesting - and a reminder that language models model language and not ideas per se. also a stark example of how easily LLMs can (as a consequence) be influenced by propaganda. Asking ChatGPT

Superhuman launches Docs, merging writing, AI and data for document collaboration

Model ReleasesDGX agent

Superhuman Inc., the company formerly known as Grammarly, today announced the launch of Docs, a product that enables multiple users to collaborate on writing using artificial intelligence. Docs contai

Supervised Reward Inference

Model ReleasesDGX agent

arXiv:2502.18447v2 Announce Type: replace Abstract: Existing approaches to reward inference typically assume that humans provide demonstrations according to specific behavior models. However, humans o

The anthropomorphization of this one is off the charts. Why does @AnthropicAI do this? It is intellectually lazy. Unscientific.

Model ReleasesDGX agent

The anthropomorphization of this one is off the charts. Why does @AnthropicAI do this? It is intellectually lazy. Unscientific. New Anthropic research: A global workspace in language models. Of everyt

The Balkanization of Execution-Security Research for AI Coding Agents: Isolation, Access Control, and Time-of-Check-to-Time-of-Use Vulnerabilities

Model ReleasesDGX agent

arXiv:2607.05743v1 Announce Type: cross Abstract: AI coding agents now read repositories, call tools, and execute shell commands with limited human oversight, and a fast-growing body of work studies w

The Granularity Paradox: How Temporal Disaggregation Inflates In-Sample Fit and Compounds Out-of-Sample Error

Model ReleasesDGX agent

arXiv:2607.05450v1 Announce Type: cross Abstract: This paper explores the 'Granularity Paradox' in time-series forecasting, wherein finer temporal disaggregation (e.g., Monthly to Weekly/Daily) improv

The next generation of ChatGPT Voice is here. Livestream starts at 10am PT. https://openai.com/live/

Model ReleasesDGX agent

OpenAI announced a livestream event showcasing the next generation of ChatGPT Voice, scheduled to begin at 10am PT. The announcement was made via OpenAI's official X (formerly Twitter) account, direct

← Previous
1…9495969798…376
Next →