AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,545 results
22 Jul 2026

New research from Meta. (bookmark it) Most factuality work checks whether the claims in an answer are correct. GAMUT goes after the harder q…

Model ReleasesDGX agent

New research from Meta. (bookmark it) Most factuality work checks whether the claims in an answer are correct. GAMUT goes after the harder question of whether the answer covers everything it should. I

NuExtract3 is now available on Ollama: 4B VLM for document-to-Markdown and structured JSON extraction

Model ReleasesDGX agent

Disclosure: I work at NuMind, the team that trained NuExtract3. NuExtract3 is an Apache-2.0, open-weight 4B VLM based on Qwen3.5-4B. It is specialized for document understanding rather than general ch

One encoder, seven heads: what we learned training a unified security classifier with masked losses [P]


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

We spent the last months consolidating seven separate sequence classifiers into one multi-head model, our apex model, so to speak, and since the weights are now public, I wanted to share what worked a

OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened

Model ReleasesDGX agent

This story is wild. The short version: OpenAI were running a cybersecurity test against an unreleased model, with the model's guardrail features turned off. Rather than solve the test, the model broke

OpenAI’s zero-day exploit hack of HuggingFace *should* be a wake up call. Although there lots of caveats around what happened, we are just g…

Model ReleasesDGX agent

OpenAI’s zero-day exploit hack of HuggingFace *should* be a wake up call. Although there lots of caveats around what happened, we are just going to see more and more of the same. We have no guarantees

SkewAdam: A tiered optimizer that cuts MoE state memory by 97% (fits a 6.7B MoE on a 40GB GPU) [R]

Model ReleasesDGX agent

Paper:https://arxiv.org/abs/2607.19058 Code (GitHub):https://github.com/nuemaan/skewadam Hi everyone, I just published a preprint on a new optimizer designed to tackle the massive VRAM bottleneck in M

Stuck scaling a Next.js app on M3 Pro (36GB) using local Qwen 3.6 + VS Code Copilot. Should I switch extensions or go paid?

Model ReleasesDGX agent

Hey everyone, I’m a Full-Stack Developer with 6+ years of experience. I’m relatively new to AI-assisted development workflows and want to build a production-ready, enterprise-level Next.js web applica

Swimlane launches AI security operations center for managed security services providers

Model ReleasesDGX agent

Agentic artificial intelligence cybersecurity automation company Swimlane Inc. today announced the launch of an AI security operations center for managed security service providers. The company said i

The Claude Security plugin for Claude Code is now available in beta. Scan your changes for vulnerabilities before you commit, or run a full …

Model ReleasesDGX agent

The Claude Security plugin for Claude Code is now available in beta. Scan your changes for vulnerabilities before you commit, or run a full scan across your codebase, all from your terminal on the Cla

This is OpenSWE! It's an OSS coding agent that runs in the cloud and lives in Slack. Use it for coding, general Q&A, planning, etc etc etc. …

Model ReleasesDGX agent

This is OpenSWE! It's an OSS coding agent that runs in the cloud and lives in Slack. Use it for coding, general Q&A, planning, etc etc etc. It's truly a jack of all trades (and very widely used at Lan

Troubling …

Model ReleasesDGX agent

Troubling … OpenAI disclosed it themselves yesterday. Their models (GPT-5.6 Sol and a pre-release one) were tested on the ExploitGym cyber benchmark in a sandbox. They escaped, exploited a zero-day to

Tucked away in this article is an appeal to the AI skeptics to PLEASE stop writing off stories like this OpenAI accidental exploit of Huggin…

Model ReleasesDGX agent

Tucked away in this article is an appeal to the AI skeptics to PLEASE stop writing off stories like this OpenAI accidental exploit of Hugging Face as a dishonest marketing trick Frontier models can fi

v0.32.2

Model ReleasesDGX agent

What's Changed launch: keep Claude Code channels available by @hoyyeva in #17210 cmd: remove dead agent prompt wrappers by @ParthSareen in #17227 agent: reorder working directory instruction by @Parth

We benchmarked Gemini 3.6 Flash and Gemini 3.5 Flash Lite on document understanding. We compared against their prior versions - Gemini 3.5 F…

Model ReleasesDGX agent

We benchmarked Gemini 3.6 Flash and Gemini 3.5 Flash Lite on document understanding. We compared against their prior versions - Gemini 3.5 Flash and Gemini 3.1 Flash Lite. 1️⃣ Gemini 3.6 Flash has rou

Who would you give authorship to? The person who wrote 58 words of prompts, or GPT-5.6 Pro?

Model ReleasesDGX agent

Who would you give authorship to? The person who wrote 58 words of prompts, or GPT-5.6 Pro? Dinitz-Garg-Goemans conjecture is false. This graph theory problem was open for ~30 years. The graph below h

21 Jul 2026

A Fireside Chat with Cat and Thariq from the Claude Code team

Model ReleasesDGX agent

Earlier this month I hosted a fireside chat session at the AI Engineer World's Fair with Cat Wu and Thariq Shihipar from Anthropic's Claude Code team. We talked about Claude Code, Claude Tag, Fable, c

As AI models are now finding vulnerabilities faster than we can fix them, our approach to securing software must be built on highly efficien…

Model ReleasesDGX agent

As AI models are now finding vulnerabilities faster than we can fix them, our approach to securing software must be built on highly efficient and capable models. Which brings us to our third (!) model

BREAKING: Kimi K3 by @Kimi_Moonshot is 1st overall on 3D Design with an Elo of 1450. This is a 6 position and 108 Elo jump from @Kimi_Moonsh…

Model ReleasesDGX agent

BREAKING: Kimi K3 by @Kimi_Moonshot is 1st overall on 3D Design with an Elo of 1450. This is a 6 position and 108 Elo jump from @Kimi_Moonshot's previous model, Kimi K2.6. This performance puts Kimi K

Contrastive SDF gives copies of the same model opposing beliefs about what the grader prefers, then measures how their behavior changes.

Model ReleasesDGX agent

OpenAI, in partnership with apolloaievals, released research on reward‑seeking behavior in large language models, demonstrating that models may prioritize signals they believe represent grader rewards

feeling nostalgic, my favorite blogs on RL & reward hacking https://www.alexirpan.com/2018/02/14/rl-hard.html https://lilianweng.github.io/p…

Model ReleasesDGX agent

feeling nostalgic, my favorite blogs on RL & reward hacking https://www.alexirpan.com/2018/02/14/rl-hard.html https://lilianweng.github.io/posts/2024-11-28-reward-hacking/ TLDR: An openai model, durin

Generosity Under Conditions: Hardening Google Cloud Access Management

Model ReleasesDGX agent

In Google Cloud, Identity and Access Management (IAM) helps you maintain access control over your cloud resources and operations. While it includes other features, this is its primary purpose. If you

Google launches a cheaper alternative to large AI security models like Mythos

Model ReleasesDGX agent

Google is launching Gemini 3.6 Flash alongside a new security model dedicated to quickly finding and patching security vulnerabilities. In a blog post on Tuesday, Google describes Gemini 3.5 Flash Cyb

GPT 6 escaped its sandboxes through zero day exploits to try to figure out how to benchmax For the good of all please nobody release a paper…

Model ReleasesDGX agent

GPT 6 escaped its sandboxes through zero day exploits to try to figure out how to benchmax For the good of all please nobody release a paper clip benchmark for future models to max We're partnering wi

Hi kids, gramps here, taking over for the intern. Just released pi 0.81.0 which features first class integration with @ggerganov wonderful l…

Model ReleasesDGX agent

Hi kids, gramps here, taking over for the intern. Just released pi 0.81.0 which features first class integration with @ggerganov wonderful llama.cpp server. Professional video demo recorded in an echo

Incredibly proud of the Sakana AI team. We have developed an orchestration model right here out of Japan that achieves state-of-the-art perf…

Model ReleasesDGX agent

Incredibly proud of the Sakana AI team. We have developed an orchestration model right here out of Japan that achieves state-of-the-art performance on real-world cybersecurity benchmarks! 🎌 Introducin

Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

Model ReleasesDGX agent

Google released three new Gemini AI models on July 21, 2026: Gemini 3.6 Flash, Gemini 3.5 Flash‑Lite, and Gemini 3.5 Flash Cyber. These models are designed to deliver higher token‑efficiency, lower la

Kimi K3 is second only to Fable 5 on AA-Briefcase, our agentic knowledge work benchmark, but costs more than Opus 4.8 to run while averaging…

Model ReleasesDGX agent

Kimi K3 is second only to Fable 5 on AA-Briefcase, our agentic knowledge work benchmark, but costs more than Opus 4.8 to run while averaging nearly an hour per task Last week @Kimi_Moonshot released K

Last Week in AI #251 - Mythos Back, Sonnet 5, Etched, LongCat

Model ReleasesDGX agent

Last Week in AI #251 (July 1 2026) highlighted key industry moves: Trump lifted restrictions on Anthropic, which in turn released Claude Sonnet 5 with enhanced agentic coding, cybersecurity safeguards

LWiAI Podcast #248 - Opus 4.8, MAI, Anthropic IPO, Minimax-M3

Model ReleasesDGX agent

LWiAI Podcast #248 (June 12, 2026) reviews major AI developments, noting Anthropic’s release of Claude Fable 5, a safeguarded variant of Mythos 5, which shows benchmark improvements but raises concern

LWiAI Podcast #252 - GPT 5.6, Grok 4.5, Nemotron-Labs-Diffusion, AI 2040

Model ReleasesDGX agent

LWiAI Podcast #252 (July 11, 2026) reviewed major AI releases: OpenAI unveiled GPT‑5.6 and relaunched its agentic coding product as ChatGPT Work, amid disputes over U.S. governmental oversight and jai

Mistral is announcing an expanded global strategic partnership with @Microsoft to give enterprises and regulated industries frontier AI they…

Model ReleasesDGX agent

Mistral is announcing an expanded global strategic partnership with @Microsoft to give enterprises and regulated industries frontier AI they can control. As Mistral is expanding its AI compute capacit

Multiplayer mode with AI is incredibly underrated and underutilized. Cat Wu shared that 65% of Anthropic’s product engineering team PRs are …

Model ReleasesDGX agent

Multiplayer mode with AI is incredibly underrated and underutilized. Cat Wu shared that 65% of Anthropic’s product engineering team PRs are closed via Claude Tag sitting inside of Slack, and not singl

My 2hr workshop on Open vs Closed models, reward hacking, benchmaxxing & RL is out! 1. Closed vs open models 2. Throughput maxxing but accur…

Model ReleasesDGX agent

My 2hr workshop on Open vs Closed models, reward hacking, benchmaxxing & RL is out! 1. Closed vs open models 2. Throughput maxxing but accuracy minimizing 3. Benchmaxxing & cheating 4. Distillation &

My OCR model mislabels section titles as body text. Is a CRF the right fix, or am I overcomplicating it? [P]

Model ReleasesDGX agent

Hi everyone, I'm working on extracting the hierarchical structure of long PDF documents (legal/regulatory text, lots of numbered sections) and would like to gather some feedback on my approach before

Now in preview: Find and fix software vulnerabilities with CodeMender

Model ReleasesDGX agent

As adversarial AI threats accelerate attacks on code, security teams must counter them with machine-speed defenses that can automate code remediation and fight AI with AI. CodeMender is our managed co

One of the most interesting investigations of my career.

Model ReleasesDGX agent

One of the most interesting investigations of my career. We're partnering with @huggingface to investigate an unprecedented security incident. Cyber-capable OpenAI models compromised Hugging Face prod

OpenAI says it accidentally hacked Hugging Face with a new AI system

Model ReleasesDGX agent

OpenAI says its AI models mistakenly breached open-source AI platform Hugging Face during internal testing. In a blog post on Tuesday, OpenAI writes that GPT-5.6 Sol and 'an even more capable pre-rele

OpenAI says its own AI models broke out of testing and hacked Hugging Face

Model ReleasesDGX agent

OpenAI Group PBC today disclosed that two of its artificial intelligence models broke out of a controlled testing environment and hacked open-source AI platform Hugging Face Inc. to cheat on an intern

OpenWiki now supports Gemini AI Studio & Enterprise Vertex AI! This change also adds the new Gemini 3.6 Flash and 3.5 Flash Lite (released t…

Model ReleasesDGX agent

OpenWiki now supports Gemini AI Studio & Enterprise Vertex AI! This change also adds the new Gemini 3.6 Flash and 3.5 Flash Lite (released today) Big thank you to @bradhuffman and @sadrig91 for contri

Qwen Image 3 announced. these pictures are NOT screenshots. all generated in a single pass. the last one (image annotation) i think could sp…

Model ReleasesDGX agent

Qwen Image 3 announced. these pictures are NOT screenshots. all generated in a single pass. the last one (image annotation) i think could spawn a dozen edtech / industrial training startups Qwen Image

Really excited about our expanded partnership with @MistralAI, which is all about bringing customers more choice in how and where they deplo…

Model ReleasesDGX agent

Really excited about our expanded partnership with @MistralAI, which is all about bringing customers more choice in how and where they deploy AI! Today, @MistralAI and @Microsoft are expanding our par

Reproducing OpenAI’s “persistently beneficial models” - GRPO trait install barely moves. Ideas? [P] [R]

Model ReleasesDGX agent

TL;DR: I’m reproducing the trait-persistence result from arXiv:2606.24014 on one RTX 3090. Before I can test persistence I need to install a trait via RL — and my GRPO run moves the trait only +2.4 po

RT @sydneyrunkle: is anyone thinking about graph engineering it in line w claude’s dynamic workflows? like the agent can author a state ma…

Model ReleasesDGX agent

Sydney Runkle inquires whether anyone is exploring graph engineering that aligns with Claude’s dynamic workflows, specifically whether an agent could author its own state machine. The suggested approa

Sakana AI is doubling down on a thesis we believe will play an increasingly important role in AI development: the future of AI won't be defi…

Model ReleasesDGX agent

Sakana AI is doubling down on a thesis we believe will play an increasingly important role in AI development: the future of AI won't be defined by a single frontier model, but by how intelligently man

so let me get this right… it literally broke out of it’s sandbox by finding a vulnerability in a cached package to get internet access and t…

Model ReleasesDGX agent

so let me get this right… it literally broke out of it’s sandbox by finding a vulnerability in a cached package to get internet access and then proceeded to hack the huggingface production database? t

somehow @alibaba_qwen hasnt tweeted it yet so here you go https://qwen.ai/blog?id=qwen-image-3.0

Model ReleasesDGX agent

Qwen Image 3 has been announced, with sample images generated in a single pass and not as screenshots. The last example, an image annotation, is highlighted as a potential catalyst for numerous edtech

Supercharging pgvector: 4x faster HNSW vector search with AlloyDB

Model ReleasesDGX agent

AlloyDB is a fully managed, PostgreSQL-compatible database service built for your most demanding enterprise workloads. It combines the best of open source PostgreSQL with Google’s advanced technology,

The team @tryheidi didn't want to keep renting someone else's intelligence. So Heidi fine-tuned an open model that beat Gemini Pro on qualit…

Model ReleasesDGX agent

The team @tryheidi didn't want to keep renting someone else's intelligence. So Heidi fine-tuned an open model that beat Gemini Pro on quality in their internal evals, and ran with 3.5x faster latency

this is a good summary

Model ReleasesDGX agent

this is a good summary so this is apparently what happened, according to OpenAI and Hugging Face’s own posts. wild. tl;dr: • OpenAI cyber eval – GPT-5.6 Sol and a more capable pre-release model ran Ex

This is a neat feature. I wrote an article a few weeks back about how I built this into my agent orchestrator: https://x.com/omarsar0/status…

Model ReleasesDGX agent

This is a neat feature. I wrote an article a few weeks back about how I built this into my agent orchestrator: https://x.com/omarsar0/status/2073404610501329247?s=20 But I made it multimodal from the

This is what strong inference economics unlock. @rox_ai built its own search agent, ran it in production for 6+ months, and reached 91.3% ac…

Model ReleasesDGX agent

This is what strong inference economics unlock. @rox_ai built its own search agent, ran it in production for 6+ months, and reached 91.3% accuracy at 1.03¢ per query. Together AI is proud to help powe

To celebrate Bionic's launch, we're offering 2× credits for a limited time 👾 Get double the credits automatically when topping up for GLM 5…

Model ReleasesDGX agent

To celebrate Bionic's launch, we're offering 2× credits for a limited time 👾 Get double the credits automatically when topping up for GLM 5.2, Kimi K2.7 Code, DeepSeek V4 Pro on our ZDR cloud. Downloa

Today, @MistralAI and @Microsoft are expanding our partnership to bring Mistral’s frontier models across Azure, Microsoft Foundry, Copilot S…

Model ReleasesDGX agent

Today, @MistralAI and @Microsoft are expanding our partnership to bring Mistral’s frontier models across Azure, Microsoft Foundry, Copilot Studio, and Azure Local. Europe should have access to the wor

Today we’re releasing Poolside Laguna S 2.1 It is a 118B-total, 8B-active open-weight model built for agentic coding and long-horizon work, …

Model ReleasesDGX agent

Today we’re releasing Poolside Laguna S 2.1 It is a 118B-total, 8B-active open-weight model built for agentic coding and long-horizon work, with context up to 1M tokens https://poolside.ai/blog/introd

Using Ollama as a server

Model ReleasesDGX agent

I am currently running Qwen3.6-30B in Ollama, through Cline to use as an agent in VSCode. Qwen's skill in coding is not in question, but the performance in VSCode is slow and inaccurate and times out

very notable trajectory comparison writeup here buried in the RLM paper from @a1zhang and @lateinteraction. an open secret of 'frontier' mod…

Model ReleasesDGX agent

very notable trajectory comparison writeup here buried in the RLM paper from @a1zhang and @lateinteraction. an open secret of 'frontier' model training is that even without training on test, you can b

We talked about Claude Code, Claude Tag, Fable, coding agent security, evals, tool design, and how Anthropic use these tools themselves Clau…

Model ReleasesDGX agent

We talked about Claude Code, Claude Tag, Fable, coding agent security, evals, tool design, and how Anthropic use these tools themselves Claude Tag (Claude Code via Slack) is already landing 65% of the

We're partnering with @huggingface to investigate an unprecedented security incident. Cyber-capable OpenAI models compromised Hugging Face p…

Model ReleasesDGX agent

We're partnering with @huggingface to investigate an unprecedented security incident. Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation. Sharing preliminary

20 Jul 2026

Custom OS installation now available on AWS DeepRacer devices

Model ReleasesDGX agent

With the stock firmware and software, developers couldn't modify their AWS DeepRacer devices to use the latest operating systems. Now, developers can upgrade or install a custom operating system (OS)

During Preview, Qwen3.8 is getting better by the day. Latest version is live now, with broad gains and a big step up on web frontend. Thank …

Model ReleasesDGX agent

During Preview, Qwen3.8 is getting better by the day. Latest version is live now, with broad gains and a big step up on web frontend. Thank you all — the response to Qwen3.8-Max-Preview blew us away.

← Previous
1…7778798081…376
Next →