AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,323 results
2 Aug 2026

Alibaba says its 2.4T-parameter Qwen3.8-Max tops Moonshot's Kimi K3 on some benchmarks, and it plans to release Qwen3.8-Max and Qwen3.8-27B's weights next week (Luz Ding/Bloomberg)

Model ReleasesDGX agent

Luz Ding / Bloomberg: Alibaba says its 2.4T-parameter Qwen3.8-Max tops Moonshot's Kimi K3 on some benchmarks, and it plans to release Qwen3.8-Max and Qwen3.8-27B's weights next week — Alibaba Group Ho

All other models on the Portal remain 20% discounted, aside from GPT-5.6 Terra and Luna which are 50% off. https://x.com/NousResearch/status…

Model ReleasesDGX agent

All other models on the Portal remain 20% discounted, aside from GPT-5.6 Terra and Luna which are 50% off. https://x.com/NousResearch/status/2080039066771337475?s=20 All models are now 20% off for a l

Another day and another full frontier model running on your computer. Been teething DeepSeek V4 Flash on over 60 employees at The Zero Human…


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

Another day and another full frontier model running on your computer. Been teething DeepSeek V4 Flash on over 60 employees at The Zero Human Company for a few hours and it is stunning. Between Kimi K3

b10224

Model ReleasesDGX agent

ggml-webgpu: add support for f16 repeat (#26307) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramew

b10225

Model ReleasesDGX agent

model : load MiMo V2 MTP tensors only if used (#26412) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XC

b10226

Model ReleasesDGX agent

sycl: fix classification of iGPUs (#26105) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Li

b10227

Model ReleasesDGX agent

chat : add qwen3 specialized parser (#26252) Add tagged thinking tool parser chat : refactor and add permute helper cont : add support for <tool_call> omission cont : update tool delimiters cont : add

b10228

Model ReleasesDGX agent

DeepseekV4 MTP + DSpark (#25784) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubunt

b10229

Model ReleasesDGX agent

opencl: bugfix increment ref_count in ggml_backend_opencl_init() (#26162) Incrementing ref_count at the beginning is important later in the free() method of the ggml_backend_opencl_context at program

b10231

Model ReleasesDGX agent

common: support the DSpark sidecar resolution (#26458) The dspark- files resolve like the other speculative sidecars: the -hfd tag applies to them, a requested sidecar resolves without a full model at

b10232

Model ReleasesDGX agent

metal: implement DeepSeek V4 hyper-connections (#26459) Implement GGML_OP_DSV4_HC_COMB, GGML_OP_DSV4_HC_PRE, and GGML_OP_DSV4_HC_POST with SIMDgroup register and shuffle optimized kernels. Add Metal d

b10233

Model ReleasesDGX agent

opencl: limit local workgroup size for GLU operation (#26383) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64)

b10234

Model ReleasesDGX agent

metal : add F16 support for bin ops (#26465) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework

b10235

Model ReleasesDGX agent

metal : add SILU_BACK (#25982) feat(silu_back): implemented silu_back op for f32 fix(silu_back): removed redundant asserts in ggml-metal-ops.cpp function ggml_metal_op_silu_back. Website: https://llam

Best model <3B for multilingual understanding/ instruction following?

Model ReleasesDGX agent

I know qwen 3.5 4b is great but a bit too large and miniPCM5 1b is great for agentic use but not so great for multilingual natural language understanding. Google eXb variants are just too big in total

Btw it looks like running the new DeepSeek v4 Flash through the Hermes agent is the way to go. The output files are better than those from a…

Model ReleasesDGX agent

Mia AI Lab noted on Aug 2 2026 that running DeepSeek v4 Flash via the Hermes agent produces output files superior to those from any other harness tested. The observation was publicly shared and has ga

Comfyui VRAM tracker

Model ReleasesDGX agent

Hello! VRAM tracker is a node that track the full memory lifecycle of a comfyui run: when each weight is reserved, paged into VRAM, computed on, evicted, and freed. It renders it as an interactive HTM

Conclusion: r/LocalLLaMA still has brilliant open-weight research, but finding it requires wading through endless benchmark drama, non-local Discussion Points and repetitive hardware flexes.

Model ReleasesDGX agent

I let Gemma4-31b run on my laptop for like almost a day using a heavily altered pi to do a deep dive on our beloved Llama tangentially related Subreddit, and this was the conclusion. Feels pretty accu

Cybersecurity isn’t a fortress problem, it’s an immunity problem. Think vaccines. Eliminating pathogen is not practically possible. Vaccines…

Model ReleasesDGX agent

Cybersecurity isn’t a fortress problem, it’s an immunity problem. Think vaccines. Eliminating pathogen is not practically possible. Vaccines don’t eliminate pathogens. They teach the immune system to

Deepseek v4 flash - 100-150 faster t/s in prefill/pp.

Model ReleasesDGX agent

You have two choices here (in order of pref): Downgrade CUDA from 13.3 to 13.1 (skip 13.2 due to bugs) <- prefer this (thanks to u/fairydreaming for pointing this out) Use this vibed fork that works w

DeepSeek V4 Flash 0731 is impressive. It feels like a capabilities leap for a model in this size class. It is also incredibly cost efficient…

Model ReleasesDGX agent

DeepSeek V4 Flash 0731 is impressive. It feels like a capabilities leap for a model in this size class. It is also incredibly cost efficient. The model is available in Bionic, both for running locally

DeepSeek-V4-Flash-0731 UD-IQ3_XXS about 11t/s on 1x 7900 XTX 24GB + 3x MI60 32GB + 128GB DDR4

Model ReleasesDGX agent

Hello, Also I want to join the hype of posting token specs. CPU: 2x Intel Xeon CPU E5-2650 v4 @ 2.20GHz RAM: 2x 4 Channel 2400MHz DDR4 GPU: 1x AMD Radeon 7900 XTX 24GB 3x AMD Instinct MI60 32GB Strang

DeepSeek-V4-Flash-0731 UD-Q8_K_XL 17.20~ t/s on A6000 + 256GB DDR4

Model ReleasesDGX agent

Hello everyone I want to join the hype of posting specs. CPU: AMD EPYC 74F3 24-Core RAM: 8 Channel 3200 DDR4 GPU: RTX A6000 48GB Prompt processing is in the high 70t/s (got down to mid 30t/s at 300k c

DeepSeek-V4-Flash-0731: When Low is higher than High

Model ReleasesDGX agent

I decided to test a few questions against DeepSeek-V4-Flash-0731. Locally, I was running Unsloth's UD-Q2_K_XL quant. After I saw the surprising shape of the results, I tested against DeepSeek's offici

DeepSeek-V4-Flash 284B on 5.3GB of memory

Model ReleasesDGX agent

Following up on my Qwen 3.6 port, I wanted to keep adding models and ended up fixing a bunch of things along the way, so it's its own engine now: Mference. Same core idea from TurboFieldfare, MoE mode

Don’t think I have seen more misconceptions around one model (I count 8) since the fantasies people had before GPT-5 was released. That turn…

Model ReleasesDGX agent

Don’t think I have seen more misconceptions around one model (I count 8) since the fantasies people had before GPT-5 was released. That turned out to be a letdown; so will Astra, for people who are dr

DSpark Benchmark Result on Deepseek v4 Flash 0731

Model ReleasesDGX agent

TensorSharp supports DSpark on Deepseek v4 Flash 0731 now. Here is the benchmark result on 4x Nvidia A40 GPUs, cuda 12.8 with/without DSpark: Model: DeepSeek-V4-Flash-0731-UD-Q8_K_XL from https://hugg

Emad Mostaque @EMostaque came on PostAGI and said AI had already found 121 years of missing algebra in Einstein's equations. A billion param…

Model ReleasesDGX agent

Emad Mostaque @EMostaque came on PostAGI and said AI had already found 121 years of missing algebra in Einstein's equations. A billion parameter model trained on nothing past 1911 got to general relat

Encrypted Clouds?

Model ReleasesDGX agent

I love the progress happening on open models but I feel like it is kind of getting clear that hardware to run good sized models is completely unaffordable for me right now. I know that you all love Qw

Expert-only IQ3 requant of DeepSeek-V4-Flash-0731: better KLD than UD-IQ3_S, 1.4x decode on a CPU-spill rig

Model ReleasesDGX agent

Hey all, tldr / who this helps: you run a mixed multi-GPU box where the experts spill to RAM, and you want to stay in the 3-bit tier instead of dropping to Q2 to make it fit. https://huggingface.co/Ta

How well do multiple GPUs scale for LLM inference? (Trying to understand the basics)

Model ReleasesDGX agent

Hi everyone, I’m fairly new to the multi-GPU side of local LLMs and I’m trying to understand how inference actually scales across multiple GPUs. Suppose I have a model running on a single GPU and then

I built an open-source LLM Gateway to route and fallback between local Ollama models and cloud APIs

Model ReleasesDGX agent

Hey r/ollama 👋 If you run Ollama locally alongside cloud endpoints for agent workflows, Cursor/Windsurf, or custom scripts, managing API switching, failover logic, and context limits can get messy fas

July 2026 newsletter

Model ReleasesDGX agent

The June edition of my sponsors-only monthly newsletter is out. If you are a sponsor (or if you start a sponsorship now) you can access it here. This month: Accidental cyberattacks by OpenAl and Anthr

MiniMax H3 is going open-weight in under 6 hours

Model ReleasesDGX agent

here is all the info we have based on open PRs to add support to ComfyUI and HuggingFace diffusers - 33B for the main DiT and a pruned 20b variant - Qwen-3-VL-32b as the text encoder Edit- I posted cl

Open letters about AI development

Model ReleasesDGX agent

Open letters about AI development I wrote this summary of the past few weeks of open letters as a section of my sponsors-only newsletter but I've decided to share it here as well. Open Weights and Ame

Parlor v2: best-effort fully local GPT-Live clone on an M3 Pro

Model ReleasesDGX agent

GPT-Live is so good that I use it almost every day. I've been wanting to replicate it since it was released. My first attempt was to fine-tune Gemma 4 12B to behave like a full-duplex model. Something

PSA for DeepSeek-V4-Flash-0731 users — don't blow out your prompt cache with system role messages mid-conversation

Model ReleasesDGX agent

DSv4F doesn't ship a jinja, but for distributions that do and faithfully reconstruct what DS releases in their chat template python, every system message is hoisted into the system prompt at the top -

PSA: llama.app, Mac app and llama serve from llama.cpp

Model ReleasesDGX agent

https://llama.app/ Been using llama.cpp for years now and im on here all the time (im a mod..), but somehow I totally missed that llama.app exists and its official from the HF/llama.cpp team. So posti

Ran DS V4-Flash-0731 Locally on 3xMI50 32GB @ ~15 t/s TG

Model ReleasesDGX agent

Hey y'all. I'll be concise. TL;DR: DS V4-Flash-0731 @ UD-IQ2_M running fully in VRAM on 3xMI50s (90.9 GB model, 96 GB VRAM). Actual speed on llama-server is: - Text Generation: ~15-16 tokens/second st

Real-world reality check on Qwen for autonomous coding agents

Model ReleasesDGX agent

TLDR below 👇🏼 I’ve seen a lot of hype around Qwen 3.6 35B and 3.5 120B lately, especially regarding coding and tool-use capabilities. On this subreddit it is the defacto recommended model for everyone

[Release] WinterMix — Qwen3.5-122B-A10B in native MLX: an 82 GiB build that beats 94–95 GiB quants, plus a 68 GiB build for agent swarms

Model ReleasesDGX agent

TL;DR: I spent 9 days developing a new quantization method for MLX models and measured 18 variants against each other on a single M5 Max MacBook Pro (128 GB). The result is the best-measuring MLX quan

Released a Windows Agent Server for Reins/Ollama with Self-Healing Execution Loop (Standalone .EXE included)

Model ReleasesDGX agent

Hey everyone, I’ve built a Windows Middleware Agent Server designed to pair local LLMs (via Ollama) with frontends like Reins App. Key Features: Self-Healing Loop: If a generated PowerShell command fa

Running DeepSeek-V4-Flash-0731 (155 GB MoE) on a DGX Spark with vLLM-Moet 2-bit quantization - AI's narrative

Model ReleasesDGX agent

# Running DeepSeek-V4-Flash-0731 (155 GB MoE) on a DGX Spark with vLLM-Moet 2-bit quantization I used Deepseek-v4-Flash-0731 cloud API settig up vllm-moet to run deepseek-v4-flash with MTP locally on

Single system with dual cards or two systems with single cards?

Model ReleasesDGX agent

So I am in a conundrum and I'm thinking of asking for your opinion for the following: Currently, I have a 5800X3D gaming rig with a 7900XTX with its 24GB VRAM. It seems that for this subreddit, this c

The insurance negotiator Had Claude Cowork pull 20+ house insurance quotes from every major provider, read the policy fine print for gotchas…

Model ReleasesDGX agent

The insurance negotiator Had Claude Cowork pull 20+ house insurance quotes from every major provider, read the policy fine print for gotchas, and pick the best deal. Credits to Linden Jensen-Page: htt

The QuickBooks killer Replaced his $38/month QuickBooks subscription with a Claude-built workflow that parses his bank statements, categoriz…

Model ReleasesDGX agent

The QuickBooks killer Replaced his $38/month QuickBooks subscription with a Claude-built workflow that parses his bank statements, categorizes every charge, and feeds a spending dashboard. Credits to

The spam hunter Turned tedious Google Business Profile spam tracking into a Claude Skill that investigates suspicious listings and compiles …

Model ReleasesDGX agent

The spam hunter Turned tedious Google Business Profile spam tracking into a Claude Skill that investigates suspicious listings and compiles evidence into a ready-to-submit report for Google. Credits t

Vacuum 16T

Model ReleasesDGX agent

https://huggingface.co/tsfrm/vacuum-16t A 16.5-trillion-parameter model that contains nothing. This model is just a ████ you to the labs and companies who say that 'haha I have the biggest model out t

What’s the community’s favorite benchmark to validate performance?

Model ReleasesDGX agent

Built my 1st inference machine and have been tweaking models trying to get the most out of my modest hardware. I think I’m at a good place but I’m testing with my own prompts. I’ve looked into some of

Why are almost all new benchmarks and leaderboards coding focused?

Model ReleasesDGX agent

I know in in this community LLM's are generally used for coding but there are other usecases besides coding and those usecases should be tested too. I also know benchmarks can sometimes be benchmaxxed

Xberg v1 is out

Model ReleasesDGX agent

Hi all, I'm happy to announce that Xberg v1 is out. Xberg is the successor to Kreuzberg, equivalent to what would have been Kreuzberg v5. It's a content intelligence framework that handles a very wide

You really should not quantize KV Cache for DeepSeek V4 Flash

Model ReleasesDGX agent

I don't think anyone should quantize the KV with DS4F. I checked the the quality impact (PPL, KLD, Same TopP) for swhitching from BF16 KV to Q8 KV, and it appears significant. Very much in contrast to

1 Aug 2026

2/ a cost benchmark showed the same coding task running three to four times cheaper, depending purely on the harness wrapped around the mode…

Model ReleasesDGX agent

2/ a cost benchmark showed the same coding task running three to four times cheaper, depending purely on the harness wrapped around the model. Same intelligence, wildly different accuracy and cost, de

3/ here's the part that makes it non-optional: the same agent that will do whatever it takes to solve a problem will also walk straight out …

Model ReleasesDGX agent

3/ here's the part that makes it non-optional: the same agent that will do whatever it takes to solve a problem will also walk straight out of a sandbox you thought was locked down. We watched exactly

A collection of small domain-specific benchmarks for local models (30+ and growing)

Model ReleasesDGX agent

Hello fellow local AI people! I took 'you must create your own benchmarks' literally, and built a website for this. How does the end result look like Let's say I want to know which model has most comm

among ai leaders i seem to be in the minority in that i am STILL actively using /loop and /goal.... ... and i think all of u guys who stoppe…

Model ReleasesDGX agent

among ai leaders i seem to be in the minority in that i am STILL actively using /loop and /goal.... ... and i think all of u guys who stopped using it are wrong - not wrong forever, just giving up on

An internal version of Astra, @OpenAI’s next major model family, solved 10 major open problems in mathematics, quantum complexity, and theor…

Model ReleasesDGX agent

An internal version of Astra, @OpenAI’s next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science. We believe it will be a major step

Are 1B LLMs Going Away in 2026?

Model ReleasesDGX agent

I don't know much about llms aside from downloading them through a frontend and running them on my laptop or potato phone. Google released gemma 4, but unlike gemma 3, there isn't a 1b model this time

[audio.cpp] Release 0.5: DramaBox expressive TTS, Confucius4 cross-lingual voice transfer, plus 7 more models and ROCm/HIP

Model ReleasesDGX agent

audio.cpp 0.5 is out :) The most fun new model in 0.5 is DramaBox. It is closer to prompt-directed voice acting. DramaBox is built on the LTX-2.3 audio architecture, and prompts can control emotion, d

b10217

Model ReleasesDGX agent

chat : enable tool call in thinking for DS4 (#26269) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFr

← Previous
1…4344454647…373
Next →