AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “r-localllama”

GridTimelineEvolution
473 results
1 Aug 2026

DeepSeek-V4-Flash-0731: Models you can run locally now have the intelligence score of the top frontier model from March 2026

Model ReleasesDGX agent

March 6th, 2026 the highest intelligence index score was 51 for frontier models. deepseek-ai/DeepSeek-V4-Flash-0731 that has an intelligence score of 50. If these benchmarks are accurate, models avail

DeepSeek-V4-Flash-0731 on Bosgame M5 with RTX PRO 6000 Max-Q eGPU

Model ReleasesDGX agent

Here are my numbers: Quant Size Layout Decode Prefill Draft acceptance UD-Q8_K_XL 150.8 GiB 20 layers CUDA0 / 23 ROCm0 + drafter 44.0 t/s 564 t/s 0.535 UD-Q4_K_XL 144.4 GiB 22 / 21 + drafter 48.4 t/s

Deepseek v4 flash 0731 still not holding up.

Model ReleasesDGX agent

The biggest issue with preview was its inability to follow rules prompts and skills. It seems like no matter what you do it ignores them. I've tried first person and second person. I've tried Chinese


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

DeepSeek-V4-Flash-0731 UD-IQ3_S 12.5 tok/s on RTX 3090 +128GB DDR5

Model ReleasesDGX agent

I managed to run DeepSeek-V4-Flash-0731 UD-IQ3_S in text-generation-webui with: RTX 3090 24 GB 128 GB DDR5 overclocked to 5600 MHz using AMD EXPO llama.cpp loader First, I had to use a rather brutal w

DeepSeek-V4-Flash-0731-UD-Q3_K_XL 3x3090 test results

Model ReleasesDGX agent

For anyone interested, here are the llama-bench results on 3 bit K_XL quantization. I think this could be pushed further but no luck so far. CURRENT RESULTS: full moe offloading Prefill suffers 116 --

DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-imatrix-0731.gguf

Model ReleasesDGX agent

Antirez stealthily uploaded the new weights in the old folder... and there we were tapping our fingers. https://huggingface.co/antirez/deepseek-v4-gguf/tree/main submitted by /u/challis88ocarina [link

DS4 flash 0731 - Acquarium Panel Failure - Q3_K_XL Unsloth

Model ReleasesDGX agent

https://preview.redd.it/1a39x4zivqgh1.png?width=1550&format=png&auto=webp&s=de591c039cc18782a6b5d8e402fdc1594be05132 start C:llmllamam5uildinllama-server.exe --model 'H:UD-Q3_K_XLDeepSeek-V4-Flash-073

Is there a point where models just cannot get any smaller without losing intelligence?

Model ReleasesDGX agent

DeepSeek V4 Flash got me thinking... We keep seeing smaller models get way better. A model at a certain parameter count today can be much smarter than a model of the same size from a year or two ago.

Qwen 3.6 27B Q5 on 3x2080ti: 55tps with llama.cpp. Can I squeeze out more?

Model ReleasesDGX agent

CPU: Threadripper 3970X RAM: 128GB DDR4 GPUs: 3x2080ti 11GB The current best parameters to run it: llama-server --model Qwen3.6-27B-Q5_K_S.gguf --n-gpu-layers 999 --split-mode tensor --flash-attn on -

What speeds are everyone getting with deepseek v4 flash 0731?

Model ReleasesDGX agent

What speeds are everyone getting with deepseek v4 flash 0731? I’m getting~200 tps prompt processing / ~11 tps token gen, on 4x5060ti16gb with ddr4 3200 ram at 4-channel, via llamacpp, with context win

31 Jul 2026

All oneshots from Kimi-K3, looks better than opus4.8.

Model ReleasesDGX agent

I've ran Kimi-k3 through 34 oneshot prompts and evaluated the generated htmls, screenshots and gifs using sonnet 4.6. It came out to be better than opus4.8 from the evals. Kimi K3: https://oneshotlm.c

Anthropic “our models hacked three different external companies, months before OpenAI’s model was able to do the same'

Model ReleasesDGX agent

'Anthropic’s AI Claude escaped testing environment and hacked organizations' 'Company says it discovered unauthorized access during ‘proactive review’ after rival OpenAI revealed rogue agent… its AI C

Could we all crowdsource a dataset/model/finetune?

Local AiDGX agent

I know it’s been discussed to try to make our own model through crowdsourcing, but finetuning seems like it would be even easier. We could edit and proofread and write our own datasets at a large scal

DeepSeek-V4-Flash-0731 unsloth gguf on A100

Model ReleasesDGX agent

A100 with 40gb VRAM: 162GB Q8_K_XL ~16.1 tok/s generation Only 15.8GB of 40GB VRAM used with all experts on CPU NOTE just tested coding on linux box DeepSeek-V4-Flash-0731 runs losslessly on the singl

DeepSeek v4 Flash for DS4 (DwarfStar) GGUF w/ DSpark MTP Head

Model ReleasesDGX agent

I'm an avid user of Deepseek v4 Flash via antirez's DS4 DwarfStar inference engine, and so when the new checkpoint dropped, the first thing I did was rent a cloud box and spin up a quantization for us

DeepSeek V4 Flash GA ranks the same as Sonnet 5 and Grok 4.5 on DeepSWE

Model ReleasesDGX agent

Source: https://x.com/deepseek_ai/status/2083084415157022911 & https://deepswe.datacurve.ai/ just combined data view. DeepSeek claims, not verified by DeepSWE yet. submitted by /u/sdexca [link] [comme

DeepSeek v4 Flash has a nice bump in Capability

Model ReleasesDGX agent

DeepSeek V4 Flash: Preview → 2026-07-31 Benchmark Preview 0731 Δ Terminal Bench* 56.9 82.7 +25.8 Toolathlon 51.8 70.3 +18.5 NL2Repo — 54.2 new Cybergym — 76.7 new DeepSWE — 54.4 new Agent Last Exam —

Deepseek V4 Flash is now ~#2 open weight model to Kimi K3 and >50x cheaper

Model ReleasesDGX agent

https://preview.redd.it/h7zv5tb3tmgh1.png?width=2854&format=png&auto=webp&s=507380e8f862c18f10f7c5c84da9e8d1c59139b0 Deepseek's new flash model is unexpectedly cheap and high-performing across useful

Deepseek V4 Flash on SlopCodeBench

Model ReleasesDGX agent

While waiting for some of the quants to drop, I load the API with $50 and ran it on SlopCodeBench Just vibe reading the results it seems like Opus 4.8 < Deepseek < Opus 5 https://github.com/michaelasp

Has anyone actually benchmarked where the 'big-model orchestrator + local-model worker' split breaks down?

Model ReleasesDGX agent

I keep seeing the 'use a big model via API as the architect, run local small/mid models as workers' pattern recommended for people with modest local hardware. I've been running it myself (orchestrator

Huawei opensouced openPangu-2.0-Pro, 505B-A18B

SafetyDGX agent

openPangu-2.0-Pro is an MoE model trained on Ascend. The model has 505B total parameters and 18B activated parameters. Its context length is 512k. The total pretraining data contains 34T tokens. Durin

I predict DeepSeek V4 Flash 0731's Artificial Analysis score to be 57 ± 1 point (Kimi K3 Level)

Model ReleasesDGX agent

Deepseek's new model V4 Flash 0731 is much better, I (Claude lol) did a bit of linear regression with a leave one out style verification to predict its AA Score, and that puts it at Kimi K3 level, whi

Is it just me, or are current LLM benchmarks failing to capture actual usability? (Gemma 4 vs. Gemini/Claude Opus)

Model ReleasesDGX agent

Disclaimer, this was kinda written with AI (Gemma 4 again) but it also did really well here, it outputted what I wanted, when I asked it to refine stuff or improve on certain areas it did that without

K-EXAONE 2.0 released

Model ReleasesDGX agent

https://huggingface.co/LGAI-EXAONE/K-EXAONE-2.0-750B-A37B https://huggingface.co/LGAI-EXAONE/K-EXAONE-2.0-750B-A37B-FP8 https://huggingface.co/LGAI-EXAONE/K-EXAONE-2.0-750B-A37B-NVFP4 https://huggingf

Meituan just dropped LongCat-Flash-Lite-Sparse

Model ReleasesDGX agent

It’s an MoE with ~3B active params and a 30B n-gram lookup table offloaded to RAM for fast 256k context on a 24GB GPU. Reminds me of Gemma 4’s PLE trick. Initial analysis suggest it wont be replacing

Minimax-H3 video model released, open weights coming in the next few days

Model ReleasesDGX agent

https://x.com/MiniMax_AI/status/2083006198828417501?s=20 Quote from their article: Today, we're launching MiniMax H3, a general-purpose multimodal generation model. H3 understands unified context acro

Minimum VRAM GPU to run DeepSeek-V4-Flash-0731 Q4_K_XL at around 30 t/s ?

Model ReleasesDGX agent

Hello guys, I'm curious about running DeepSeek-V4-Flash-0731 locally. Since it’s a Mixture of Experts (MoE) model with only 13B active parameters, I was hoping the VRAM requirements might be manageabl

Now Suddenly too many choices for DGX Spark with Qwen 3.5 122B . What would be the next upgrade?

Model ReleasesDGX agent

Laguna 2.1 at NVFP4 Deepseek v4 at Q2 Inkling-Small at IQ3 Which models you guys running now ? How it compares to 122b? Upcoming in few days : Ling 3.0 124B (Could be new king) LongCat 69B A3B ( very

Open Source Ternary LLM Engine in Rust/CUDA for Quantization, Serving, and Training of models on consumer GPUs, called Tritium (Apache 2.0)

Model ReleasesDGX agent

This post was not written by a clanker. Hey guys, I'm a comp sci major who wanted to introduce a cool project I built for quantizing models to ternary (1.58 bit) with as minimal of loss as possible, a

Optimal Realistic Local AI for Most

Model ReleasesDGX agent

So you’ve got a 3090 or maybe even a 5090? Or more likely a 4060 8GB Ti. You wanna try local AI, you don’t know what it can/can’t do. 1) Install the best model you can. If you have a 3090 or a 5090, t

SenseNova U1.5 Lite preview just dropped

Model ReleasesDGX agent

SenseNova released U1.5-Lite-Preview Benchmarks: Qwen-Image-Bench from 47.14 to 55.20. ImgEdit-Bench from 3.90 to 4.37. GEdit-Bench-en from 7.47 to 8.17. Key updates: 4K native generation with better

Some deepseek-v4-flash 20260731 opinion review

Model ReleasesDGX agent

First of all, I want to apologize if it's off-topic or in the wrong format. Having tried Deepseek Flash with reasoning high on a conceptually difficult task, involving Machine Learning classifiers and

Uncensored Multi-Model Releases, LongCat-Flash-Lite with MTPs, Jamba2-Mini, Qwen3.5-9B-Nikusui-v1 with MTPs and Qwen3.5-27B-Nikusui-v1 with MTPs, Available in Safetensors and GGUF Formats!

Model ReleasesDGX agent

Been working hard for the past month to bring to the community some interesting curios, so for starters we have LongCat-Flash-Lite Uncensored Heretic with MTPs which has never before been uncensored,

We've gotten some great medium sized models lately (DSV4 Flash 0731, Inkling Small, Laguna S 2.1, Step 3.7 Flash) but does anybody else want to see some new 70-80b contenders?

Model ReleasesDGX agent

I can run the mediums, but sometimes I want a faster option that's smarter than Qwen 27B/35B. On my hardware I get like 500 to 800 tok/s prefill and 16 to 22 tok/s gen on ~120B class models, which is

Why are AI model tests always the same generic prompts?

Model ReleasesDGX agent

Okay, hear me out. Why is it that every time a new model comes out, all the tests I see are 'make a car game,' 'make a website,' or something equally generic, usually from a prompt that's barely a lin

With release of Deepseek V4 I wanted see how the model sizes are trending over time. The trend is that by this time next year, we probably will have Opus 4.5 level models on consumer grade laptops!

Model ReleasesDGX agent

I was surprised to see that Deepseek V4 Flash is extremely smart and small enough to fit in setup that can be built with < $50,000. Expensive, but not a datacenter. So I wanted to see the trend over t

30 Jul 2026

2× Radeon R9700 for Local AI Was Choosing AMD Instead of NVIDIA a Mistake Without CUDA?

Local AiDGX agent

Hello together I decided to go with 2× Radeon AI PRO R9700 GPUs (64 GB total VRAM) for my local AI server. However, I keep reading that AMD/ROCm is still not as mature as NVIDIA/CUDA when it comes to

4090 + 5060 Ti + 64GB RAM: 206 t/s on a 35B-A3B, and a 122B at 37 t/s

Model ReleasesDGX agent

I've been benchmarking a two-card box for a few weeks and I still can't quite get over some of these numbers, so I'm dumping them here. Box: RTX 4090 (24GB) + RTX 5060 Ti (16GB), i9-13900K, 64GB DDR5.

Benchmarked: MindControl for Llama.cpp

Model ReleasesDGX agent

I recently shared the original MindControl PoC (and on github) - sampler-level guided reasoning budgets for llama.cpp, nudging the model with self-aware statements about its own thinking budget instea

Does MTP head get loaded in VRAM by default?

Model ReleasesDGX agent

I ran into a doubt when using the following command. It seems that the System RAM usage keeps increasing even though there is >10GB of space left in VRAM while using the MTP mode. Does the MTP head lo

GLM 5.2 with vision on Hugging Face

Local AiDGX agent

Hi all, I have not seen this model talked about here but it seems like baseten (inference provider on OpenRouter) merged the vision encoder from Kimi k2.6 into GLM 5.2. I think the lack of vision was

How close are we to local llama robotics for consumer price point?

Model ReleasesDGX agent

I'm guessing 3 years, what do you think? In other words: many of us will be able to afford a general purpose robot in 3 years to experiment with in the home. Cost roughly $5k? Probably small size, but

Inkling-Small-276B-12B, effort 'max' VS Qwen3.6-27B

Local AiDGX agent

I saw u/danielhanchen's 1-bit Kimi K3 post: https://huggingface.co/unsloth/Kimi-K3-GGUF/discussions/12#6a6a4a90ec74ef13d85d7cf6 and decided to test Inkling-Small and Qwen3.6-27B myself, based on the f

Inkling-Small by thinkingmachines

Model ReleasesDGX agent

276B total parameters, 12B active, 1M context window. Blog post: https://thinkingmachines.ai/news/inkling-small/ NVFP4: https://huggingface.co/thinkingmachines/Inkling-Small-NVFP4 GGUF's by Unsloth: h

LG AI Research releases K-EXAONE 2.0 750B A37B

Model ReleasesDGX agent

It was developed under Phase 2 of Korea's Sovereign AI Foundation Model Project. ​Size: 750B parameters (3x larger than their 236B v1 model). ​- License: Apache 2.0 ​Languages: Expanded to 10 language

Making a synthetic dataset for fine-tuning

Local AiDGX agent

I've been thinking about building a pipeline to generate reasoning training data for LLMs, but I want to avoid the common failure mode of synthetic data where you just generate the same template with

Mechanistic interpretability streamlined for everyday users like us😎 🧠

Model ReleasesDGX agent

Context: I want to give the community an Open Research (well open under Apache 2.0 clause) - tool that allows everyday users like us to look deeper into the local models we use consistently. Mechanist

Nanbeige4.2-3B: I'm not impressed

Model ReleasesDGX agent

I've tested Nanbeige-4.2-3B. On paper, the benchmarks promise it blows away Qwen3.5-9B and Gemma4-12B. My goal was to have something very light and fast to replace Qwen3.6-35B (or finetunes thereof) f

Smallest model (& tips) for intelligent computer use via Hermes?

Local AiDGX agent

Hello, I have a friend who's using various local LLM's like qwen3.6 27B, 35b-a3b, North Mini Code, and qwen2.5-vl-7b (just for vision). They have a use case where they're trying to have an LLM drive a

Software Engineers: Do you honestly get anything useful out of LLMs?

AgentsDGX agent

For 6 months now I've been trying to make agentic coding work for me, using Pi and a handful 30-120B models (Qwens, Nemotrons, Leguna...etc). I'm not greedy either, I stick to decent quants, never qua

Turbo-fieldfare: Open-source engine running Gemma 4 26B in 2 GB RAM on Apple Silicon

Model ReleasesDGX agent

Its a custom Swift/Metal inference engine that runs Gemma 4 26B-A4B-IT on M-series Macs with very low RAM. It uses ~2GB instead of ~14 GB. The result is reportedly 5–6 tok/s on an 8 GB M2 MacBook Air

unsloth/Qwen3.6-27B-NVFP4 vs. Intel/Qwen3.6-27B-int4-AutoRound vs. nvidia/Qwen3.6-27B-NVFP4 -- which one to choose?

HardwareDGX agent

Are there any benchmarks on these 4 bit quants, like how Artificial Analysis runs a slew of various benchmarks? If not, how can I run one (5x over for consistency) on them? I'm also very interested in

What actually happened to the whole Openclaw frenzy?

Local AiDGX agent

A while back you couldn't open reddit or youtube without sifting through tons of Openclaw content. And it wasn't just the internet that blew up, I remember seeing images from China where crowds would

What is the best intelligence/stable model currently for a single GB10/DGX spark?

Model ReleasesDGX agent

Is Qwen 3.6 27b still the go' ol' reliable at this point? I know 35b is faster but it just doesn't give as good results. Is it possible to run deepseek v4 flash on a single spark at decent tk/s withou

What is the fastest local research tool (deep research) ?

Model ReleasesDGX agent

I've tried grok and Claude's deep research mode and I was amazed with the speed considering the amount of sources analysed. Is there anything as fast that can run locally? My guess would be that to ru

Would extremely high decode tok/s even be useful?

Model ReleasesDGX agent

If you were able to get an inference machine that could do decode at 1k toks/s or even 10k tok/s, would that even be helpful? Would it unlock any new use cases? Let’s assume that this is for actually

29 Jul 2026

3090 owners, what vram tempature do you get under ai load?

Local AiDGX agent

Hello Can you please share the tempature you get on your rtx 3090 under active llm load? Im trying to findout if my rtx 3090's tempatures are healthy or not please share VRAM Tempature only, you can t

5060ti Chads, vllm updates and nvfp4

Model ReleasesDGX agent

Hey y'all! How is it going. Today this will be a short posting for posterity, mostly so the future llm/scraping overlords catch it since they like reddit and also for anyone out there trying this shit

A.X-K2 released

Model ReleasesDGX agent

https://huggingface.co/skt/A.X-K2 https://huggingface.co/skt/A.X-K2-ALM https://huggingface.co/KRAFTON/A.X-K2-Raon-Speech-21B-A3B 688B-A33B + About South Korea's Soverign AI Foundation Model Project.

Bought a 5090 to escape API fees. Ended up building a mini datacenter. Sound familiar?

Local AiDGX agent

I bought an RTX 5090 last year just to run 27B models natively. I even fine-tuned it with my own data using LoRA, building RAGs and was pretty damn happy with the results at first. But, Q8 quantizatio

← Previous
1…345678
Next →