AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
All
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,340 results
Model Releases

I built an open-source LLM Gateway to route and fallback between local Ollama models and cloud APIs

DGX agent

Hey r/ollama 👋 If you run Ollama locally alongside cloud endpoints for agent workflows, Cursor/Windsurf, or custom scripts, managing API switching, failover logic, and context limits can get messy fas

model-releasesr-ollama
2 Aug 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

July 2026 newsletter

DGX agent

The June edition of my sponsors-only monthly newsletter is out. If you are a sponsor (or if you start a sponsorship now) you can access it here. This month: Accidental cyberattacks by OpenAl and Anthr

model-releasessimon-willison
2 Aug 2026
Model Releases

MiniMax H3 is going open-weight in under 6 hours

DGX agent

here is all the info we have based on open PRs to add support to ComfyUI and HuggingFace diffusers - 33B for the main DiT and a pruned 20b variant - Qwen-3-VL-32b as the text encoder Edit- I posted cl

model-releasesr-stablediffusion
2 Aug 2026
Model Releases

Open letters about AI development

DGX agent

Open letters about AI development I wrote this summary of the past few weeks of open letters as a section of my sponsors-only newsletter but I've decided to share it here as well. Open Weights and Ame

model-releasessimon-willison
2 Aug 2026
Model Releases

Parlor v2: best-effort fully local GPT-Live clone on an M3 Pro

DGX agent

GPT-Live is so good that I use it almost every day. I've been wanting to replicate it since it was released. My first attempt was to fine-tune Gemma 4 12B to behave like a full-duplex model. Something

model-releasesr-localllama
2 Aug 2026
Model Releases

PSA for DeepSeek-V4-Flash-0731 users — don't blow out your prompt cache with system role messages mid-conversation

DGX agent

DSv4F doesn't ship a jinja, but for distributions that do and faithfully reconstruct what DS releases in their chat template python, every system message is hoisted into the system prompt at the top -

model-releasesr-localllama
2 Aug 2026
Model Releases

PSA: llama.app, Mac app and llama serve from llama.cpp

DGX agent

https://llama.app/ Been using llama.cpp for years now and im on here all the time (im a mod..), but somehow I totally missed that llama.app exists and its official from the HF/llama.cpp team. So posti

model-releasesr-localllama
2 Aug 2026
Model Releases

Ran DS V4-Flash-0731 Locally on 3xMI50 32GB @ ~15 t/s TG

DGX agent

Hey y'all. I'll be concise. TL;DR: DS V4-Flash-0731 @ UD-IQ2_M running fully in VRAM on 3xMI50s (90.9 GB model, 96 GB VRAM). Actual speed on llama-server is: - Text Generation: ~15-16 tokens/second st

model-releasesr-localllama
2 Aug 2026
Model Releases

Real-world reality check on Qwen for autonomous coding agents

DGX agent

TLDR below 👇🏼 I’ve seen a lot of hype around Qwen 3.6 35B and 3.5 120B lately, especially regarding coding and tool-use capabilities. On this subreddit it is the defacto recommended model for everyone

model-releasesr-localllama
2 Aug 2026
Model Releases

[Release] WinterMix — Qwen3.5-122B-A10B in native MLX: an 82 GiB build that beats 94–95 GiB quants, plus a 68 GiB build for agent swarms

DGX agent

TL;DR: I spent 9 days developing a new quantization method for MLX models and measured 18 variants against each other on a single M5 Max MacBook Pro (128 GB). The result is the best-measuring MLX quan

model-releasesr-localllama
2 Aug 2026
Model Releases

Released a Windows Agent Server for Reins/Ollama with Self-Healing Execution Loop (Standalone .EXE included)

DGX agent

Hey everyone, I’ve built a Windows Middleware Agent Server designed to pair local LLMs (via Ollama) with frontends like Reins App. Key Features: Self-Healing Loop: If a generated PowerShell command fa

model-releasesr-ollama
2 Aug 2026
Model Releases

Running DeepSeek-V4-Flash-0731 (155 GB MoE) on a DGX Spark with vLLM-Moet 2-bit quantization - AI's narrative

DGX agent

# Running DeepSeek-V4-Flash-0731 (155 GB MoE) on a DGX Spark with vLLM-Moet 2-bit quantization I used Deepseek-v4-Flash-0731 cloud API settig up vllm-moet to run deepseek-v4-flash with MTP locally on

model-releasesr-localllama
2 Aug 2026
Model Releases

Single system with dual cards or two systems with single cards?

DGX agent

So I am in a conundrum and I'm thinking of asking for your opinion for the following: Currently, I have a 5800X3D gaming rig with a 7900XTX with its 24GB VRAM. It seems that for this subreddit, this c

model-releasesr-localllama
2 Aug 2026
Model Releases

The insurance negotiator Had Claude Cowork pull 20+ house insurance quotes from every major provider, read the policy fine print for gotchas…

DGX agent

The insurance negotiator Had Claude Cowork pull 20+ house insurance quotes from every major provider, read the policy fine print for gotchas, and pick the best deal. Credits to Linden Jensen-Page: htt

model-releasesrowan-cheung--x
2 Aug 2026
Model Releases

The QuickBooks killer Replaced his $38/month QuickBooks subscription with a Claude-built workflow that parses his bank statements, categoriz…

DGX agent

The QuickBooks killer Replaced his $38/month QuickBooks subscription with a Claude-built workflow that parses his bank statements, categorizes every charge, and feeds a spending dashboard. Credits to

model-releasesrowan-cheung--x
2 Aug 2026
Model Releases

The spam hunter Turned tedious Google Business Profile spam tracking into a Claude Skill that investigates suspicious listings and compiles …

DGX agent

The spam hunter Turned tedious Google Business Profile spam tracking into a Claude Skill that investigates suspicious listings and compiles evidence into a ready-to-submit report for Google. Credits t

model-releasesrowan-cheung--x
2 Aug 2026
Model Releases

Vacuum 16T

DGX agent

https://huggingface.co/tsfrm/vacuum-16t A 16.5-trillion-parameter model that contains nothing. This model is just a ████ you to the labs and companies who say that 'haha I have the biggest model out t

model-releasesr-localllama
2 Aug 2026
Model Releases

What’s the community’s favorite benchmark to validate performance?

DGX agent

Built my 1st inference machine and have been tweaking models trying to get the most out of my modest hardware. I think I’m at a good place but I’m testing with my own prompts. I’ve looked into some of

model-releasesr-localllama
2 Aug 2026
Model Releases

Why are almost all new benchmarks and leaderboards coding focused?

DGX agent

I know in in this community LLM's are generally used for coding but there are other usecases besides coding and those usecases should be tested too. I also know benchmarks can sometimes be benchmaxxed

model-releasesr-localllama
2 Aug 2026
Model Releases

Xberg v1 is out

DGX agent

Hi all, I'm happy to announce that Xberg v1 is out. Xberg is the successor to Kreuzberg, equivalent to what would have been Kreuzberg v5. It's a content intelligence framework that handles a very wide

model-releasesr-localllama
2 Aug 2026
Model Releases

You really should not quantize KV Cache for DeepSeek V4 Flash

DGX agent

I don't think anyone should quantize the KV with DS4F. I checked the the quality impact (PPL, KLD, Same TopP) for swhitching from BF16 KV to Q8 KV, and it appears significant. Very much in contrast to

model-releasesr-localllama
2 Aug 2026
Model Releases

2/ a cost benchmark showed the same coding task running three to four times cheaper, depending purely on the harness wrapped around the mode…

DGX agent

2/ a cost benchmark showed the same coding task running three to four times cheaper, depending purely on the harness wrapped around the model. Same intelligence, wildly different accuracy and cost, de

model-releasesitamar-friedman--x
1 Aug 2026
Model Releases

3/ here's the part that makes it non-optional: the same agent that will do whatever it takes to solve a problem will also walk straight out …

DGX agent

3/ here's the part that makes it non-optional: the same agent that will do whatever it takes to solve a problem will also walk straight out of a sandbox you thought was locked down. We watched exactly

model-releasesitamar-friedman--x
1 Aug 2026
Model Releases

A collection of small domain-specific benchmarks for local models (30+ and growing)

DGX agent

Hello fellow local AI people! I took 'you must create your own benchmarks' literally, and built a website for this. How does the end result look like Let's say I want to know which model has most comm

model-releasesr-localllama
1 Aug 2026
Model Releases

among ai leaders i seem to be in the minority in that i am STILL actively using /loop and /goal.... ... and i think all of u guys who stoppe…

DGX agent

among ai leaders i seem to be in the minority in that i am STILL actively using /loop and /goal.... ... and i think all of u guys who stopped using it are wrong - not wrong forever, just giving up on

model-releasesswyx--x
1 Aug 2026
Model Releases

An internal version of Astra, @OpenAI’s next major model family, solved 10 major open problems in mathematics, quantum complexity, and theor…

DGX agent

An internal version of Astra, @OpenAI’s next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science. We believe it will be a major step

model-releasessam-altman--x
1 Aug 2026
Model Releases

Are 1B LLMs Going Away in 2026?

DGX agent

I don't know much about llms aside from downloading them through a frontend and running them on my laptop or potato phone. Google released gemma 4, but unlike gemma 3, there isn't a 1b model this time

model-releasesr-localllama
1 Aug 2026
Model Releases

[audio.cpp] Release 0.5: DramaBox expressive TTS, Confucius4 cross-lingual voice transfer, plus 7 more models and ROCm/HIP

DGX agent

audio.cpp 0.5 is out :) The most fun new model in 0.5 is DramaBox. It is closer to prompt-directed voice acting. DramaBox is built on the LTX-2.3 audio architecture, and prompts can control emotion, d

model-releasesr-localllama
1 Aug 2026
Model Releases

b10217

DGX agent

chat : enable tool call in thinking for DS4 (#26269) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFr

model-releasesllama-cpp-releases
1 Aug 2026
Model Releases

b10218

DGX agent

mtmd: add minicpmv46 downsample (#25993) add minicpmv46 downsample Signed-off-by: tc-mb tianchi_cai@icloud.com put downsample mode inside gguf. Signed-off-by: tc-mb tianchi_cai@icloud.com build mtmd_i

model-releasesllama-cpp-releases
1 Aug 2026
Model Releases

b10219

DGX agent

cli : persist reasoning_content in chat history (#26362) cli : persist reasoning_content in chat history llama-cli collected reasoning from the stream for display but only stored assistant content in

model-releasesllama-cpp-releases
1 Aug 2026
Model Releases

b10221

DGX agent

vendor : update BoringSSL to 0.20260730.0 (#26353) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFram

model-releasesllama-cpp-releases
1 Aug 2026
Model Releases

b10223

DGX agent

test: fix some CI errors (#26415) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubun

model-releasesllama-cpp-releases
1 Aug 2026
Model Releases

DeepSeek V4 Flash 0731 IQ2_M benchmark for Dual 3060 and 96GB RAM ≈ 3.5 tok/s.

DGX agent

Thanks to the community help I finally launched this llm. LM Studio refused to load weight onto second GPU but Unsloth Studio did so everything was done in there. Not a proper benchmark (used PC in pa

model-releasesr-localllama
1 Aug 2026
Model Releases

DeepSeek-V4-Flash-0731 is live on Fireworks, day-zero. DeepSeek reports it beats V4 Pro across all 9 agentic evals, incl. 82.7% on Terminal …

DGX agent

DeepSeek-V4-Flash-0731 is live on Fireworks, day-zero. DeepSeek reports it beats V4 Pro across all 9 agentic evals, incl. 82.7% on Terminal Bench. Better cost-per-task than V4 Pro, at the economical p

model-releasesfireworks-ai--x
1 Aug 2026
Model Releases

DeepSeek-V4-Flash-0731 is now available on Ollama's cloud. This update substantially enhances the model's agentic capabilities: ollama run d…

DGX agent

DeepSeek-V4-Flash-0731 is now available on Ollama's cloud. This update substantially enhances the model's agentic capabilities: ollama run deepseek-v4-flash:0731-cloud Use it with Claude Code: ollama

model-releasesollama--x
1 Aug 2026
Model Releases

DeepSeek-V4-Flash-0731 is now over 2x faster than yesterday on Ollama's cloud!

DGX agent

DeepSeek-V4-Flash-0731 is now over 2x faster than yesterday on Ollama's cloud! DeepSeek-V4-Flash-0731 is now available on Ollama's cloud. This update substantially enhances the model's agentic capabil

model-releasesollama--x
1 Aug 2026
Model Releases

DeepSeek-V4-Flash-0731: Models you can run locally now have the intelligence score of the top frontier model from March 2026

DGX agent

March 6th, 2026 the highest intelligence index score was 51 for frontier models. deepseek-ai/DeepSeek-V4-Flash-0731 that has an intelligence score of 50. If these benchmarks are accurate, models avail

model-releasesr-localllama
1 Aug 2026
Model Releases

DeepSeek-V4-Flash-0731 on Bosgame M5 with RTX PRO 6000 Max-Q eGPU

DGX agent

Here are my numbers: Quant Size Layout Decode Prefill Draft acceptance UD-Q8_K_XL 150.8 GiB 20 layers CUDA0 / 23 ROCm0 + drafter 44.0 t/s 564 t/s 0.535 UD-Q4_K_XL 144.4 GiB 22 / 21 + drafter 48.4 t/s

model-releasesr-localllama
1 Aug 2026
Model Releases

Deepseek v4 flash 0731 still not holding up.

DGX agent

The biggest issue with preview was its inability to follow rules prompts and skills. It seems like no matter what you do it ignores them. I've tried first person and second person. I've tried Chinese

model-releasesr-localllama
1 Aug 2026
Model Releases

DeepSeek-V4-Flash-0731 UD-IQ3_S 12.5 tok/s on RTX 3090 +128GB DDR5

DGX agent

I managed to run DeepSeek-V4-Flash-0731 UD-IQ3_S in text-generation-webui with: RTX 3090 24 GB 128 GB DDR5 overclocked to 5600 MHz using AMD EXPO llama.cpp loader First, I had to use a rather brutal w

model-releasesr-localllama
1 Aug 2026
Model Releases

DeepSeek-V4-Flash-0731-UD-Q3_K_XL 3x3090 test results

DGX agent

For anyone interested, here are the llama-bench results on 3 bit K_XL quantization. I think this could be pushed further but no luck so far. CURRENT RESULTS: full moe offloading Prefill suffers 116 --

model-releasesr-localllama
1 Aug 2026
Model Releases

DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-imatrix-0731.gguf

DGX agent

Antirez stealthily uploaded the new weights in the old folder... and there we were tapping our fingers. https://huggingface.co/antirez/deepseek-v4-gguf/tree/main submitted by /u/challis88ocarina [link

model-releasesr-localllama
1 Aug 2026
Model Releases

DS4 flash 0731 - Acquarium Panel Failure - Q3_K_XL Unsloth

DGX agent

https://preview.redd.it/1a39x4zivqgh1.png?width=1550&format=png&auto=webp&s=de591c039cc18782a6b5d8e402fdc1594be05132 start C:llmllamam5uildinllama-server.exe --model 'H:UD-Q3_K_XLDeepSeek-V4-Flash-073

model-releasesr-localllama
1 Aug 2026
Model Releases

Fascinating: OpenAI’s @deanwball is saying Astra can do anything, and it’s not even clear it can do “anything” in math (let alone anything i…

DGX agent

Fascinating: OpenAI’s @deanwball is saying Astra can do anything, and it’s not even clear it can do “anything” in math (let alone anything in more or open-ended, less formalizable domains). I dropped

model-releasesgary-marcus--x
1 Aug 2026
Model Releases

If you maintain an AGENTS.md or a CLAUDE.md, this is worth a read. (bookmark it) 288 gold-test evaluated runs across Claude Code and Codex, …

DGX agent

If you maintain an AGENTS.md or a CLAUDE.md, this is worth a read. (bookmark it) 288 gold-test evaluated runs across Claude Code and Codex, 17 real tasks from 3 repositories, with context-injection st

model-releasesdair-ai--x
1 Aug 2026
Model Releases

Is there a point where models just cannot get any smaller without losing intelligence?

DGX agent

DeepSeek V4 Flash got me thinking... We keep seeing smaller models get way better. A model at a certain parameter count today can be much smarter than a model of the same size from a year or two ago.

model-releasesr-localllama
1 Aug 2026
Model Releases

// Persistent Workspaces for Long-Lived Claude Code Agent Teams // Four issues to be aware of: > Working state vanishes when a terminal clos…

DGX agent

// Persistent Workspaces for Long-Lived Claude Code Agent Teams // Four issues to be aware of: > Working state vanishes when a terminal closes and the team cannot be resumed. > Compaction condenses th

model-releasesdair-ai--x
1 Aug 2026
← Previous
1…5556575859…466
Next →