AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
550 results
Model Releases

Deepseek V4 Flash just hit Colibri, does anyone have numbers?

DGX agent

I'm mosty interested in 128-192GB VRAM with 128-256GB RAM to spare, so SSD streaming is basically not even necessary. Seems only FP4 is supported, so older hardware will likely be slow - no Unsloth GG

model-releasesr-localllama
5 Aug 2026
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

DeepSeek-V4-Flash on SM89 4x48gb 4090s with DSpark

DGX agent

https://github.com/yhfgyyf/vllm-deepseek-v4-sm89 I couldn't believe that someone actually got vLLM working with this particular set of GPUs, but here it is. The video is from right after I got it work

model-releasesr-localllama
5 Aug 2026
Model Releases

I remember a time when 'flash' meant 32B

DGX agent

I mean, Deepseek V4 Flash is an absolutely fantastic model, even though I can't run it on my machine it's so fascinating to see how it performs. Knowing that potentially it could be run at home is rea

model-releasesr-localllama
5 Aug 2026
Model Releases

I updated my localy run benchmark with DeepSeek V4 Flash 0731

DGX agent

It's the purple cluster on the top left (the good corner...) I'm running the MXFP4 version from Bartoswski with Dspark at 1K t/s prefill and 90 t/s gen (average). I tried different sampling params, yo

model-releasesr-localllama
5 Aug 2026
Model Releases

Inkling-Small 276B-A12B at ~2.9 tok/s on <10gb memory

DGX agent

A follow up to the launch of Mference, it now supports and runs Inkling-Small 276B-A12B. Inkling-Small (Thinking Machines, Apache 2.0), from the pipenetwork/Inkling-Small-MLX-4bit conversion: 276B tot

model-releasesr-localllama
5 Aug 2026
Model Releases

LFM2.5-2.6B on a OnePlus 13 at 17 tok/s ~ Pure CPU

DGX agent

As you all know the model is 2.69B parameters with a 128K context window and purpose-built for multi-step agent workflows. What you are seeing is the Q4_K_M GGUF running on my own inference engine bui

model-releasesr-localllama
5 Aug 2026
Model Releases

MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp

DGX agent

TensorSharp's MoE CPU-offload feature has been merged into main. Here is the parameters description of this feature: Mixture-of-Experts CPU offload: --n-cpu-moe <N> | -ncmoe <N> Keep the routed MoE ex

model-releasesr-localllama
5 Aug 2026
Model Releases

Prime Agent - a new coding harness surpassing Codex/CC/PI

DGX agent

Prime Agent is an open-source coding and research agent for general and long-running work. A self-improving RLM harness for coding and long-running autonomous tasks. Designed to be both token-efficien

model-releasesr-localllama
5 Aug 2026
Model Releases

PSA Update CUDA from 13.2 to 13.3 to solve DeepSeek V4 Flash 0731 Looping Problem!

DGX agent

So one of yall mentioned that cuda 13.1 or 13.2 is broken for unsloth so I looked in to it, and they were right. I had 13.2 installed, after I switched to 13.3 no more looping!!! Before the cuda updat

model-releasesr-localllama
5 Aug 2026
Model Releases

Qwen Developers' responses from their recent Twitter/X AMA

DGX agent

Questions & Responses(in BOLD) below. Favorite question(s) moved to end of the thread with combined responses(removed duplicates). Be optimistic folks. I'm sure we're getting other models too apart fr

model-releasesr-localllama
5 Aug 2026
Model Releases

Qwen3-TTS voice cloning is now in mainline llama.cpp — the old demo finally became real support

DGX agent

People may remember the Qwen3-TTS llama.cpp demo from a few months ago. That PR said it probably wouldn’t be merged because llama.cpp was missing some of the graph and API pieces it needed. A new impl

model-releasesr-localllama
5 Aug 2026
Model Releases

Scenema Audio Comes to ComfyUI, Runs on 8GB VRAM

DGX agent

Hey everyone! Scenema Audio is now a native ComfyUI custom node. Same model that powers scenema.ai now quantized so it fits on 8GB VRAM. When we first released it a few months ago as an API and Docker

model-releasesr-localllama
5 Aug 2026
Model Releases

Stable Diffusion might actually be remembered in the history books, and I don’t think that’s an overstatement

DGX agent

Hear me out before you roll your eyes. We tend to only recognize turning points in hindsight. Nobody in 1993 thought the Mosaic browser would be a history book moment, but the web is. I think Stable D

model-releasesr-stablediffusion
5 Aug 2026
Model Releases

Sulphur 3 is looking for funding

DGX agent

Hello, I'm the guy who made Sulphur 2. With the recent release of a certain video model, we are looking to mobilize and train Sulphur 3 on this new model. We are targeting $10,000 USD. This certain ne

model-releasesr-stablediffusion
5 Aug 2026
Model Releases

The Google Bug Hunters Team admitted to me that they cannot fundamentally patch prompt engineering bypasses in Gemini

DGX agent

Hello everyone Yesterday I gave a report on Gemini bugs and the techniques I learned on Gemini so far with the Engineering Prompt and interestingly today I got a very interesting and controversial ans

model-releasesr-chatgpt
5 Aug 2026
Model Releases

Thinking of buying more DRAM right now...

DGX agent

So I'm looking at https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF and I realize my 128GB of DRAM just isn't cutting it for this (incredibly powerful) model. If only I had another 64GB, I th

model-releasesr-localllama
5 Aug 2026
Model Releases

Utilize a nvidia gpu and amd gpu together for 2 different ai models?

DGX agent

We run a local model instance in our company that the dev we hired built for us. We're a trade business and we want to further use our on hand hardware for it. The specs given we have is a 5090 gpu wi

model-releasesr-localllama
5 Aug 2026
Model Releases

Watch a local Ollama's qwen3:8b turn one English question into a 9-node investigation graph - planned, admitted by a deterministic gate, and run live in the browser (open source, MIT)

DGX agent

The video is one real run, not a mock-up: grapharc go 'why did checkout latency spike at 09:14 UTC?' --model ollama/qwen3:8b A local 8B model proposes the graph → triage fanning out into four parallel

model-releasesr-ollama
5 Aug 2026
Model Releases

Xiaomi-Robotics-1: New robotics model released

DGX agent

Xiaomi-Robotics-1 is a robot foundation model trained on over 100K hours of real-world manipulation trajectories. It is a Vision-Language-Action (VLA) model engineered for out-of-the-box mobile manipu

model-releasesr-localllama
5 Aug 2026
Model Releases

A 2.6B model with tool calling and 128K context now runs at 30 tok/s on a phone

DGX agent

Liquid AI released LFM2.5-2.6B today, and this might be more relevant to local AI than another massive model most people cannot run. The model is only 2.69B parameters, has 128K context, supports tool

model-releasesr-localllama
4 Aug 2026
Model Releases

A llama.cpp PR caches “hot” MoE experts on the GPU — 33 → 56 tok/s reported with 8GB VRAM

DGX agent

A new llama.cpp PR (#26563) adds a heatmap that tracks which MoE experts are used most often. Instead of keeping every expert on the GPU or offloading all of them, it caches the frequently selected ex

model-releasesr-localllama
4 Aug 2026
Model Releases

Company approved 128GB Mac for research proposal, best model?

DGX agent

I‘m doing a research proposal at my company about running local LLMs to replace daily coding models. Qwen 3.6 27B (or 3.8 potentially) is widely seen as the best model in that 20-60GB space, is that s

model-releasesr-localllama
4 Aug 2026
Model Releases

Decrease the power limit of your 5090 to at least 480W - the performance penalty for inference is negligible.

DGX agent

I run my inference machine in the living room, so noise and heat output are a significant concern. Ran a quick test using my daily driver model (Qwen 3.6-27b) and at 480W, the card outputs only 2.1% l

model-releasesr-localllama
4 Aug 2026
Model Releases

[Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]

DGX agent

First of all, obviously I took some help from AI to type this post and this is the topic that enabled me to accomplish all that: https://old.reddit.com/r/LocalLLaMA/comments/1veow4b/deepseek_v4flash_2

model-releasesr-localllama
4 Aug 2026
Model Releases

DeepSeek V4 Flash 0731 (Q4) now reaches 1,328 tok/s prefill and ~29 tok/s decode on one RTX PRO 6000

DGX agent

I've been working on speeding up DeepSeek-V4-Flash-0731 in Krasis and have now got the long-prompt prefill quite a bit faster on a single RTX PRO 6000 96GB. These are timing-disabled internal Krasis r

model-releasesr-localllama
4 Aug 2026
Model Releases

Deepseek V4 flash 0731 ranks #21 on Agent Arena

DGX agent

https://preview.redd.it/522fsdwvtdhh1.png?width=1200&format=png&auto=webp&s=6a6cf7a467514167a8193029dbd20fb3a9ba4f6c It ranks lower than both Sonnet 4.6 and Luna. I'd wager Luna costs in the same ball

model-releasesr-localllama
4 Aug 2026
Model Releases

Deepseek V4 Flash 2-bit quant is the first model I can run locally that achieves 100% in this SQL benchmark

DGX agent

I really like to use this one SQL benchmark when testing new models. I had another post some time ago with my benchmarks, but I decided to post a new one because of how well Deepseek did. I like the b

model-releasesr-localllama
4 Aug 2026
Model Releases

DeepSeek v4 Flash vs. Qwen3.6-27B, 3.5-122B, and Gemma 4 31B Benchmark

DGX agent

Just wanted to share my agentic coding benchmark run of DSv4F 0731 at both High and Low reasoning efforts (not Max)... I ran a 109-question subset of Aider Polyglot (the JS/C++/Python languages), base

model-releasesr-localllama
4 Aug 2026
Model Releases

GPT-OSS has turned one year old today!

DGX agent

It is one of the best local models ever released, in both 20B and 120B versions. I always come back to it, especially the 120B version. Its only competition is, in my opinion, Qwen 3.5 122B, but that

model-releasesr-localllama
4 Aug 2026
Model Releases

I benchmarked the 4 models I had pulled. The 1.1GB one beat the 2GB one at math and lost badly at extraction.

DGX agent

152 generations, deterministic grading (exact number/string/JSON/regex), no LLM judge on my 16GB laptop. task type | deepseek-r1:1.5b (1.1GB) | llama3.2:3b (2.0GB) | gemma:2b | codellama 7b arithmetic

model-releasesr-ollama
4 Aug 2026
Model Releases

I built a DwarfStar-inspired Vulkan/Metal inference engine for Qwen3.6-35B-A3B on 16 GB machines

DGX agent

Disclosure: I’m the author and maintainer of QuarkStar. I built QuarkStar, a small native inference engine inspired by Antirez’s DwarfStar. QuarkStar currently supports: Qwen3.6-35B-A3B, using the sam

model-releasesr-localllama
4 Aug 2026
Model Releases

inclusionAI/Ling-3.0-flash · Hugging Face

DGX agent

The Ling-3.0-flash MoE is now open-weighted at 124B A5B params. I know the original announcements were before the Kimi K3, DeepSeek-V4-Flash and Qwen3.8 hype, but this model might still have a good ni

model-releasesr-localllama
4 Aug 2026
Model Releases

inclusionAI/Ling-3.0-flash weights are up on Hugging Face — MIT, BF16 plus an official FP8

DGX agent

Went public in the last few minutes, both repos ungated. Ling-3.0-flash, BF16, 24 shards, ~255GB Ling-3.0-flash-fp8, official FP8, ~128GB 127.5B total, they quote 5.1B active. What jumped out at me in

model-releasesr-localllama
4 Aug 2026
Model Releases

Is LM Studio abandoning their core product?

DGX agent

Some of you may be aware that a few weeks ago, LM Studio announced a new agent, Bionic. This is pretty much an agentic harness for both local models and paid cloud models. But most aren't aware that L

model-releasesr-localllama
4 Aug 2026
Model Releases

Kimi K3 full model running on 16x GB10 cluster at 20+tps

DGX agent

Kimi K3 full model running on 16x GB10 cluster at 20+tps average (llama-benchy coherent corpus) 38tps peak, 750tps prefill. This is the first run of full k3 with dspark on my cluster. I will be doing

model-releasesr-localllama
4 Aug 2026
Model Releases

LFM2.5-2.6B is out

DGX agent

Released today, with emphasis on agentic capabilities. I really like their models for simple, high volume tasks ('summarize these gazillion documents') and their 8b-a1b was my go-to for certain tasks

model-releasesr-localllama
4 Aug 2026
Model Releases

Llama.cpp PR 8% speed boost

DGX agent

Llama.cpp currently uses cpu based sampling for user with mtp enabled. The PR moves sampling to the gpu, which on a 5090 boasts an 8% increase in tok/s for qwen3.6:35b. I tested it on my P40 and obser

model-releasesr-localllama
4 Aug 2026
Model Releases

Pc build limitations

DGX agent

Here's the build I managed to scrape together System Specifications: CPU: Intel Core i7-7700K Motherboard: ASUS ROG Strix Z270-E Gaming RAM: 32GB Corsair Vengeance DDR4-3000 Storage: 1TB Crucial P5 Pl

model-releasesr-ollama
4 Aug 2026
Model Releases

Probably the best way to run DS4 flash on a mac right now (192gb+ vram)

DGX agent

Found this quant, so thought I would share, since its the best I've found so far for running on my mac (m3 ultra). It's got dspark/mtp support so runs faster than anything else I've tried. The tok/s o

model-releasesr-localllama
4 Aug 2026
Model Releases

Why are Chinese models better* at Frontend than the western top labs?

DGX agent

I use A LOT both openAI and Anthropic products. When I need some frontend work (pure web dev) (or answer that feel less verbose and more to the point) I use Anthropic. For multimodality openAI feels b

model-releasesr-localllama
4 Aug 2026
Model Releases

Why are Gamers so incredibly hostile to AI? Is it just a tiny vocal minority that spreads such toxic vitriol online?

DGX agent

It's more accurate to say that many highly engaged online gamers are hostile to AI, not that 'gamers' as a whole are. Gaming is a huge community with hundreds of millions of people, and opinions vary

model-releasesr-chatgpt
4 Aug 2026
Model Releases

AI9Stars released G9v3-39A5B

DGX agent

AI9Stars has released G9v3-39A5B an open weights language model designed to deliver even stronger reasoning capabilities than ai9stars/G9v3-3B with its 39B and 5 active experts. It is released under t

model-releasesr-localllama
3 Aug 2026
Model Releases

Broken Image Generator v2.0

DGX agent

Does anyone have an idea, when OpenAI will fix ChatGPT's glitched images? Since the 2.0 image generator was released, there is ugly glitches when you use generated image as a reference or just change

model-releasesr-chatgpt
3 Aug 2026
Model Releases

'Data center in a Box (on Wheels)' 256Gb VRAM/512Gb RAM AI Server 6-8 Month Operational Review, Stability Write Up, Benchmarks

DGX agent

I've been out of these forums for awhile but I figured I would provide a formal update on how this has been going now that it has some operation time under its belt, just to put the information out th

model-releasesr-localllama
3 Aug 2026
Model Releases

DeepSeek V4 Flash 0731 - Happy Numbers (700pp/18tg) and Thoughts

DGX agent

Originally, I was only getting around 140pp/s and about 21tg/s, but the config with -b 8192 -ub 8192 --cpu-moe is vastly superior, let's say 700pp/s and 18tg/s in the most relevant range. Test System:

model-releasesr-localllama
3 Aug 2026
Model Releases

DeepSeek V4-Flash (284B MoE) at 33 tok/s single / 68 tok/s aggregate on 2× RTX 3090 + a used quad-Xeon DDR4 server — full config

DGX agent

Ran DeepSeek V4-Flash-0731 — the full official checkpoint, not a re-quant — on commodity used hardware. Sharing because I couldn't find anyone else publishing Ampere results for this engine. Why bothe

model-releasesr-localllama
3 Aug 2026
Model Releases

Döner Bench DeepSeek-V4-Flash IQ2_XS running on a single RTX 3090

DGX agent

https://preview.redd.it/3zcvpbds14hh1.png?width=1911&format=png&auto=webp&s=a79aafb71eeca97638da93d2591902631e897fd5 I tried a test similar to the recent model-quant comparisons, but this time I focus

model-releasesr-localllama
3 Aug 2026
Model Releases

How to run big models on old hardware 30B at 22 tok/s on 6GB GPU and 16GB RAM

DGX agent

I have been working on this tool for months and there are a lot of new functionalities and tests that are going to be released in the next few weeks! The goal of the tool is to allow community members

model-releasesr-ollama
3 Aug 2026
← Previous
123456…12
Next →