AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
All
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,566 results
Model Releases

New in Claude Code: /checkup Run /checkup to: 1. Clean up unused skills/MCPs/plugins and save context 2. Dedup your local CLAUDE.md against …

DGX agent

New in Claude Code: /checkup Run /checkup to: 1. Clean up unused skills/MCPs/plugins and save context 2. Dedup your local CLAUDE.md against the checked in CLAUDE.md 3. Break up root CLAUDE.md into nes

model-releasesboris-cherny--x
8 Jul 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Nice stats on usage of open models across OpenCode. GLM-5.2 is still underrated, but one of the models that has really surprised me on agent…

DGX agent

Nice stats on usage of open models across OpenCode. GLM-5.2 is still underrated, but one of the models that has really surprised me on agentic tasks is deepseek-v4-flash. Extremely cheap and effective

model-releasesdair-ai--x
8 Jul 2026
Model Releases

nope! OpenAI would not have the nerve.

DGX agent

nope! OpenAI would not have the nerve. Hey @GaryMarcus, quick question. Did you get months of early access to OpenAI's GPT-5.6 Sol like every AI influencer on X apparently did? Asking because their re

model-releasesgary-marcus--x
8 Jul 2026
Model Releases

Not another demo or benchmark. @ShoucongChen is a senior member of our technical staff. A real project, scoped at 1-month. Delivered in 4 da…

DGX agent

Not another demo or benchmark. @ShoucongChen is a senior member of our technical staff. A real project, scoped at 1-month. Delivered in 4 days with GLM5.2 Fast. The best devs deserve >400 t/sec. Take

model-releasesfireworks-ai--x
8 Jul 2026
Model Releases

NVIDIA Nemotron Achieves Benchmark-Leading Performance With LangChain Deep Agents Harness

DGX agent

NVIDIA Nemotron 3 Ultra is offering leading performance at lower cost than top closed models with the largest and most widely adopted AI agent orchestration platform. LangChain tuned its Deep Agents h

model-releasesnvidia-blog
8 Jul 2026
Model Releases

Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis

DGX agent

arXiv:2602.00846v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) struggle with alignment due to the limitations of existing reward models (RMs), which are predominantly vis

model-releasesarxiv-cs-cl
8 Jul 2026
Model Releases

Onnes: A Physics-Grounded Multi-Agent LLM Simulator for Cryogenic Fault Diagnosis in Quantum Computing Infrastructure

DGX agent

arXiv:2607.05805v1 Announce Type: new Abstract: Dilution refrigerators are the enabling infrastructure of superconducting quantum computers, yet their fault diagnosis is still dominated by threshold a

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

- OpenAI continually outperforms Anthropic models on computer use (with Claude, I’m surprised when it works, with Codex, I expect it to work…

DGX agent

- OpenAI continually outperforms Anthropic models on computer use (with Claude, I’m surprised when it works, with Codex, I expect it to work) - when I prompt with Fable 5.5, I feel like I’m motivating

model-releasesallie-k--miller--x
8 Jul 2026
Model Releases

OpenAI launches GPT-Live, a new generation of voice models built on a full-duplex architecture, meaning they can listen and speak at the same time (OpenAI)

DGX agent

OpenAI: OpenAI launches GPT-Live, a new generation of voice models built on a full-duplex architecture, meaning they can listen and speak at the same time — A new generation of voice models for natura

model-releasestechmeme
8 Jul 2026
Model Releases

OpenAI launches GPT-Live voice model series ahead of broad GPT-5.6 release

DGX agent

OpenAI Group PBC today introduced GPT-Live, a family of artificial intelligence models optimized to process spoken instructions. The model series will power ChatGPT’s voice mode. Additionally, OpenAI

model-releasessiliconangle
8 Jul 2026
Model Releases

OpenAI says it found widespread task issues in SWE-Bench Pro, estimates ~30% of tasks are broken, and retracts its earlier recommendation to adopt the benchmark (OpenAI)

DGX agent

OpenAI: OpenAI says it found widespread task issues in SWE-Bench Pro, estimates ~30% of tasks are broken, and retracts its earlier recommendation to adopt the benchmark — Through a detailed audit, we

model-releasestechmeme
8 Jul 2026
Model Releases

Optimized Adaptive Loop Filter in Versatile Video Coding

DGX agent

arXiv:2607.05737v1 Announce Type: new Abstract: In the Versatile Video Coding~(VVC) standard, adaptive loop filter~(ALF), including Geometry transformation-based Adaptive Loop Filter~(GALF) and Cross

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

OrchardBench: A Physically-Grounded, GPU-Parallel Apple-Orchard Simulation Benchmark for Agricultural Robotics

DGX agent

arXiv:2607.06337v1 Announce Type: cross Abstract: Robotic tree-fruit harvesting is a flagship problem for agricultural automation, but progress is bottlenecked by the cost and irreproducibility of fie

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

Our audit of SWE-Bench Pro found that a meaningful share of public tasks contain issues that can distort results. Some correct solutions fai…

DGX agent

Our audit of SWE-Bench Pro found that a meaningful share of public tasks contain issues that can distort results. Some correct solutions fail because of hidden requirements, contradictory instructions

model-releasesopenai--x
8 Jul 2026
Model Releases

Own the frontier: 'In our evals, Nemotron 3 Ultra with a tuned LangChain Deep Agents harness provides advanced agent performance at a much l…

DGX agent

Own the frontier: 'In our evals, Nemotron 3 Ultra with a tuned LangChain Deep Agents harness provides advanced agent performance at a much lower inference cost. The main takeaway is that agent perform

model-releasesharrison-chase--x
8 Jul 2026
Model Releases

Parameter-Free Encoders Remain Viable for RDB Foundation Models

DGX agent

arXiv:2607.05476v1 Announce Type: new Abstract: Given a relational database (RDB) storing heterogeneous tabular information, how can we predict missing (or future) values in some target column of inte

model-releasesarxiv-cs-lg
8 Jul 2026
Model Releases

Partial Symmetry Detection for 3D Geometry using Contrastive Learning with Geodesic Point Cloud Patches

DGX agent

arXiv:2312.08230v2 Announce Type: replace Abstract: Detecting partial extrinsic symmetry in 3D geometry is a fundamental yet persistent challenge in computer vision and graphics, critical for tasks ra

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

PatchOptic for Shared-State LLM Workflows with Projected Views and Verified Structured Updates

DGX agent

arXiv:2607.05483v1 Announce Type: cross Abstract: Agentic workflows often operate over shared, structured state. Because LLM context windows are limited, each model invocation is typically shown only

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

PCBWorld: A Benchmark Environment for Engine-Grounded PCB Design Automation

DGX agent

arXiv:2607.05915v1 Announce Type: new Abstract: PCB routing is the task of connecting the nets of a board with copper traces under strict design rules, yet learning-based methods still lag behind rule

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

PIPBench: A Profile-Inclusive Framework for Personalized Image Generation Evaluation

DGX agent

arXiv:2607.06440v1 Announce Type: new Abstract: Recent text-to-image models such as DALLE-3 excel at following diverse prompts yet remain blind to individual aesthetic preferences. We study personaliz

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability

DGX agent

arXiv:2607.06196v1 Announce Type: new Abstract: Current AI safety evaluation and benchmarking frameworks predominantly rely on Western-centric culture-agnostic defaults that mask critical regional law

model-releasesarxiv-cs-cl
8 Jul 2026
Model Releases

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

DGX agent

arXiv:2607.05992v1 Announce Type: cross Abstract: Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heav

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails

DGX agent

arXiv:2607.05910v1 Announce Type: cross Abstract: Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an intrinsic property of an image. Rea

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents

DGX agent

arXiv:2607.06008v1 Announce Type: new Abstract: Large language model (LLM) agents have shown strong performance in long-horizon tasks that require planning, tool use, and interaction with external env

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

Population-Level Profiling of DSM-5 Depressive Symptoms Among Self-Reported ADHD and ASD Users on Twitter: An Exploratory Study Using Advanced NLP and Statistical Analysis

DGX agent

arXiv:2607.05626v1 Announce Type: new Abstract: Background: Depression frequently co-occurs with ADHD and autism spectrum disorder (ASD), but population-level differences in symptom expression between

model-releasesarxiv-cs-cl
8 Jul 2026
Model Releases

Privilege and confidentiality in generative AI workflows

DGX agent

arXiv:2607.05479v1 Announce Type: cross Abstract: Generative AI (GenAI) systems store and process client data in three distinct ways: in the model's parameters through training and memorisation, in th

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

Prompt-Adapter Context Routing for Parameter-Efficient Multi-Shot Long Video Extrapolation

DGX agent

arXiv:2607.06481v1 Announce Type: cross Abstract: We present PACR-Video, a parameter-efficient framework for multi-shot long video extrapolation that preserves recurring entities, scene structure, vis

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

Prompt: ANNALS — a living kingdom in a single file Working title behavior: the app titles itself per seed — 'The Annals of Vaelmere', 'The A…

DGX agent

Prompt: ANNALS — a living kingdom in a single file Working title behavior: the app titles itself per seed — 'The Annals of Vaelmere', 'The Annals of Osterholt' — because the central conceit is that yo

model-releasesethan-mollick--x
8 Jul 2026
Model Releases

Qdrant Beats Elastic’s DiskBBQ at 2x Throughput, Half the Latency, and 1/3 the Compute

DGX agent

TL;DR Elastic recently published a benchmark claiming that their proprietary, disk-based index (dubbed “DiskBBQ”) delivers up to 7x higher throughput than Qdrant when deployed on nodes with network-at

model-releasesqdrant
8 Jul 2026
Model Releases

Quantifying Frontier LLM Capabilities for Container Sandbox Escape

DGX agent

arXiv:2603.02277v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly act as autonomous agents, using tools to execute code, read and write files, and access networks, cr

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

Really excited to partner with @nvidia on the NemoClaw Deep Agents Blueprint Deep Agents is a fully open source agent harness that we are tu…

DGX agent

Really excited to partner with @nvidia on the NemoClaw Deep Agents Blueprint Deep Agents is a fully open source agent harness that we are tuning to make perform incredibly well with open models Introd

model-releasesharrison-chase--x
8 Jul 2026
Model Releases

Refiant goes where rivals only promised with a 10 million-token AI model

DGX agent

Artificial intelligence optimization startup Refiant Inc. today launched Protea, a suite of long-context AI models led by a 10 million-token context window that the company says ranks among the larges

model-releasessiliconangle
8 Jul 2026
Model Releases

Rewriting Bun in Rust

DGX agent

Rewriting Bun in Rust Jarred Sumner has been promising this blog post (since May 9th) about his Zig to Rust rewrite of Bun for significantly longer than it took him to finish the rewrite. Honestly, it

model-releasessimon-willison
8 Jul 2026
Model Releases

RFHNet: Relational and Frequency-Aware Hashing Network for Large-Scale Fine-Grained Food Image Retrieval

DGX agent

arXiv:2607.06148v1 Announce Type: new Abstract: Fine-grained food image retrieval is a key task in computational gastronomy, with applications in food traceability, dietary monitoring, and smart cater

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

RPAM: A Principled Metric for Evaluating Associations in Language Models with High Predictive Validity in Downstream Outputs

DGX agent

arXiv:2607.05679v1 Announce Type: cross Abstract: Language models (LMs) exhibit problematic biases, such as stereotypes. Effectively analyzing and mitigating such biases requires accurate and generali

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications

DGX agent

arXiv:2607.06411v1 Announce Type: cross Abstract: Developers increasingly delegate real maintenance work to product-grade coding agents, and many state tasks in their native language, in the style of

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

SafeImpute: Reliable Clinical Data Imputation via Conformal Selection

DGX agent

arXiv:2607.05613v1 Announce Type: new Abstract: Clinical care often relies on key laboratory indicators, yet real-world patient visits are sparse and tests are ordered irregularly, leading to pervasiv

model-releasesarxiv-cs-lg
8 Jul 2026
Model Releases

SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place Recognition

DGX agent

arXiv:2509.25723v4 Announce Type: replace Abstract: Visual Place Recognition (VPR) requires robust retrieval of geotagged images despite large appearance, viewpoint, and environmental variation. Prior

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

Scientific Code Search at Scale: A Multi-Domain Dataset and Benchmark

DGX agent

arXiv:2607.05443v1 Announce Type: cross Abstract: Scientists increasingly rely on open-source tools to support their research workflows, yet discovering relevant software among over 600 million GitHub

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

Self-Review Reinforcement Learning (SRRL) with Cross-Episode Memory and Policy Distillation

DGX agent

arXiv:2607.05541v1 Announce Type: cross Abstract: Reinforcement Learning is commonly used to train large language models using environmental feedback. In applied settings, the environment usually prov

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

Self-Routing: Parameter-Free Expert Routing from Hidden States

DGX agent

arXiv:2604.00421v2 Announce Type: replace Abstract: Mixture-of-Experts (MoE) layers increase model capacity by activating only a small subset of experts per token, and typically rely on a learned rout

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

Separating signal from noise in coding evaluations

DGX agent

This OpenAI article discusses methods for distinguishing meaningful performance indicators from random variation when evaluating AI coding capabilities. It likely covers evaluation methodologies, stat

model-releasesopenai
8 Jul 2026
Model Releases

SEVRA-BENCH: Social Engineering of Vulnerabilities in Review Agents

DGX agent

arXiv:2606.13757v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed in automated code-review systems, where their approvals can determine which code is mer

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

Social 3D Scene Graphs: Modeling Human Actions and Relations for Interactive Service Robots

DGX agent

arXiv:2509.24966v2 Announce Type: replace Abstract: Understanding how people interact with their surroundings and each other is essential for enabling robots to act in socially compliant and context-a

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

SpaceXAI launches Grok 4.5, its first model built in partnership with Cursor, designed to 'handle difficult, long-running' legal, finance, and coding tasks (Carmen Arroyo/Bloomberg)

DGX agent

Carmen Arroyo / Bloomberg: SpaceXAI launches Grok 4.5, its first model built in partnership with Cursor, designed to “handle difficult, long-running” legal, finance, and coding tasks — SpaceXAI has un

model-releasestechmeme
8 Jul 2026
Model Releases

SpaceXAI’s newest AI model Grok 4.5 dramatically undercuts Anthropic and OpenAI on price

DGX agent

Elon Musk’s SpaceXAI Corp. has released a new model called Grok 4.5, in what is its first major launch since it went public a few weeks earlier. In a blog post earlier today, the company said Grok 4.5

model-releasessiliconangle
8 Jul 2026
Model Releases

SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation

DGX agent

arXiv:2607.05721v1 Announce Type: new Abstract: Uncertainty estimation is essential not only for the trustworthy deployment of large language models (LLMs) but also as a foundation for self-refinement

model-releasesarxiv-cs-cl
8 Jul 2026
Model Releases

Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders

DGX agent

arXiv:2508.16560v4 Announce Type: replace-cross Abstract: Sparse Autoencoders (SAEs) extract features from LLM internal activations, meant to correspond to interpretable concepts. A core SAE training

model-releasesarxiv-cs-ai
8 Jul 2026
← Previous
1…118119120121122…471
Next →