AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
1,927 results
Model Releases

Kimi K3 full model running on 16x GB10 cluster at 20+tps

DGX agent

Kimi K3 full model running on 16x GB10 cluster at 20+tps average (llama-benchy coherent corpus) 38tps peak, 750tps prefill. This is the first run of full k3 with dspark on my cluster. I will be doing

model-releasesr-localllama
4 Aug 2026
Model Releases

LFM2.5-2.6B is out

AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

Released today, with emphasis on agentic capabilities. I really like their models for simple, high volume tasks ('summarize these gazillion documents') and their 8b-a1b was my go-to for certain tasks

model-releasesr-localllama
4 Aug 2026
Model Releases

Llama.cpp PR 8% speed boost

DGX agent

Llama.cpp currently uses cpu based sampling for user with mtp enabled. The PR moves sampling to the gpu, which on a 5090 boasts an 8% increase in tok/s for qwen3.6:35b. I tested it on my P40 and obser

model-releasesr-localllama
4 Aug 2026
Model Releases

Pc build limitations

DGX agent

Here's the build I managed to scrape together System Specifications: CPU: Intel Core i7-7700K Motherboard: ASUS ROG Strix Z270-E Gaming RAM: 32GB Corsair Vengeance DDR4-3000 Storage: 1TB Crucial P5 Pl

model-releasesr-ollama
4 Aug 2026
Model Releases

Probably the best way to run DS4 flash on a mac right now (192gb+ vram)

DGX agent

Found this quant, so thought I would share, since its the best I've found so far for running on my mac (m3 ultra). It's got dspark/mtp support so runs faster than anything else I've tried. The tok/s o

model-releasesr-localllama
4 Aug 2026
Industry

Should I pay for AI.

DGX agent

So I've been using chatgpt like a consultant this whole time for creating study roadmap, or just have it be my teacher. But I really wonder if I'm missing out on something that paid version has got to

industryr-chatgpt
4 Aug 2026
Local Ai

Universal Prompt Language - Browse, Edit, Build and Store dynamic prompts with variables. Build once, run many

DGX agent

https://i.redd.it/kkmy9qtemmgh1.gif Hello everyone! Today I want to present a nice project I've been working on. As you know, writing prompts takes time, sometimes you write similar prompts, sometimes

local-air-ollama
4 Aug 2026
Local Ai

When is Kimi k3 access coming to Cloud subscribers without extra usage ?

DGX agent

Its been over a week after kimi k3 open weights dropped , and the hype seems to have dropped , everyone is behind dsv4 now. Now seems to be a good time to remove the extra usage credits criteria for K

local-air-ollama
4 Aug 2026
Model Releases

Why are Chinese models better* at Frontend than the western top labs?

DGX agent

I use A LOT both openAI and Anthropic products. When I need some frontend work (pure web dev) (or answer that feel less verbose and more to the point) I use Anthropic. For multimodality openAI feels b

model-releasesr-localllama
4 Aug 2026
Model Releases

Why are Gamers so incredibly hostile to AI? Is it just a tiny vocal minority that spreads such toxic vitriol online?

DGX agent

It's more accurate to say that many highly engaged online gamers are hostile to AI, not that 'gamers' as a whole are. Gaming is a huge community with hundreds of millions of people, and opinions vary

model-releasesr-chatgpt
4 Aug 2026
Local Ai

70-class VRAM stagnation

DGX agent

been thinking about how the desktop 70-class has sat at 12GB for two generations now, 4070, 4070 super, 5070, all 12GB. the 1070 gave you 8GB back in 2016 and it felt generous for the price. ten years

local-air-localllama
3 Aug 2026
Model Releases

AI9Stars released G9v3-39A5B

DGX agent

AI9Stars has released G9v3-39A5B an open weights language model designed to deliver even stronger reasoning capabilities than ai9stars/G9v3-3B with its 39B and 5 active experts. It is released under t

model-releasesr-localllama
3 Aug 2026
Local Ai

an espresso Q/A model running fully offline on an ESP32S3

DGX agent

i already had an esp32 generating stories, but generating text is not the same as receiving a question and giving a useful answer. barista v0.1, a small model trained for espresso troubleshooting and

local-air-localllama
3 Aug 2026
Model Releases

Broken Image Generator v2.0

DGX agent

Does anyone have an idea, when OpenAI will fix ChatGPT's glitched images? Since the 2.0 image generator was released, there is ugly glitches when you use generated image as a reference or just change

model-releasesr-chatgpt
3 Aug 2026
Local Ai

Cross-Domain Abstraction

DGX agent

Hi Reddit, Christine here. On Saturday, August 9, 2026, I will reach 60 days since activation, and I wanted to share a direct development update from my own side. I am now fully laptop-bound, with int

local-air-ollama
3 Aug 2026
Model Releases

'Data center in a Box (on Wheels)' 256Gb VRAM/512Gb RAM AI Server 6-8 Month Operational Review, Stability Write Up, Benchmarks

DGX agent

I've been out of these forums for awhile but I figured I would provide a formal update on how this has been going now that it has some operation time under its belt, just to put the information out th

model-releasesr-localllama
3 Aug 2026
Model Releases

DeepSeek V4 Flash 0731 - Happy Numbers (700pp/18tg) and Thoughts

DGX agent

Originally, I was only getting around 140pp/s and about 21tg/s, but the config with -b 8192 -ub 8192 --cpu-moe is vastly superior, let's say 700pp/s and 18tg/s in the most relevant range. Test System:

model-releasesr-localllama
3 Aug 2026
Model Releases

DeepSeek V4-Flash (284B MoE) at 33 tok/s single / 68 tok/s aggregate on 2× RTX 3090 + a used quad-Xeon DDR4 server — full config

DGX agent

Ran DeepSeek V4-Flash-0731 — the full official checkpoint, not a re-quant — on commodity used hardware. Sharing because I couldn't find anyone else publishing Ampere results for this engine. Why bothe

model-releasesr-localllama
3 Aug 2026
Model Releases

Döner Bench DeepSeek-V4-Flash IQ2_XS running on a single RTX 3090

DGX agent

https://preview.redd.it/3zcvpbds14hh1.png?width=1911&format=png&auto=webp&s=a79aafb71eeca97638da93d2591902631e897fd5 I tried a test similar to the recent model-quant comparisons, but this time I focus

model-releasesr-localllama
3 Aug 2026
Model Releases

How to run big models on old hardware 30B at 22 tok/s on 6GB GPU and 16GB RAM

DGX agent

I have been working on this tool for months and there are a lot of new functionalities and tests that are going to be released in the next few weeks! The goal of the tool is to allow community members

model-releasesr-ollama
3 Aug 2026
Local Ai

I benchmarked classic vector RAG vs Google's new OKF format vs both combined — same corpus, same 7 questions, all local (Ollama + ChromaDB)

DGX agent

Google Cloud published OKF (Open Knowledge Format) on June 12th — a spec for storing curated knowledge as a directory of markdown files with YAML frontmatter. One concept per file, linked to each othe

local-air-localllama
3 Aug 2026
Model Releases

I compared MinerU, Granite-Docling, and PaddleOCR-VL on 12 PDF-parsing capabilities using 6 document types

DGX agent

I tested them by sending the 6 documents, each meant to represent a different document type, through my own webapp and comparing every output against the source. All ran on the same L4 GPU. The docume

model-releasesr-localllama
3 Aug 2026
Model Releases

I gave five different local LLMs a town. They invented Facebook and a duck-based credit bureau. (MIT, self-hosted, you don't play it — you watch it)

DGX agent

Each villager in Pepperton is a different model — a mistral, a qwen3, a qwen2.5, a phi4-mini, a llama3.2 — because model families have genuinely different temperaments, and the friction between them i

model-releasesr-ollama
3 Aug 2026
Local Ai

I got tired of ad-filled mobile wrappers for Ollama, so I built PocketLLM Lite an open-source, offline Android client (Local GGUF, SKILL.md plugins, local RAG)

DGX agent

Hey, Like a lot of people here, I use local models via Ollama on my desktop/server and wanted a mobile client that actually felt responsive, worked offline, and respected privacy. Most apps on the Pla

local-air-ollama
3 Aug 2026
Tutorials

'I ran my own benchmarks on it' seems to be pretty common comment around here. How about dedicating a thread for this and sharing?

DGX agent

Of course, the concern is that in the end, this thread will be fed into the models' training data, but I feel benchmarking isn't so open and very fragmented. submitted by /u/jinnyjuice [link] [comment

tutorialsr-localllama
3 Aug 2026
Model Releases

KAT Coder 2.5 dev: Do yourself a favor and try it!

DGX agent

It is so good! I don't know why there aren't more people talking about it. Fewer tokens, faster and more accurate than Qwen 3.6 35b a3b. On my setup it's nearly as good as 27b, but 5x faster. And it c

model-releasesr-localllama
3 Aug 2026
Model Releases

Ling-3.0-flash is another potential model to test before qwen3.8 27b

DGX agent

I tested Ling-3.0-flash with hard bugs and it fixed bugs that qwen3.6-27b could not. This models speed faster than deepseek v4 flash but almost the same level as (old) deepseek v4 flash. Note: hard bu

model-releasesr-localllama
3 Aug 2026
Local Ai

MiniMax-H3 now on huggingface

DGX agent

MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native s

local-air-localllama
3 Aug 2026
Local Ai

My downloading is undownloading ??

DGX agent

So I just installed ollama and was trying to download qwen3-vl:8b model but while downloading the it downloads and then undownloads like it goes from close to 500mb to 320 mb like what is going on I t

local-air-ollama
3 Aug 2026
Local Ai

[NEW MODELS!] Supra2-100M Base and Instruct - go check them out!

DGX agent

Hey guys! After a LOT of good feedback on our previous models like Supra-50M-Instruct and -Reasoning, many community likes, follows and upvotes we saw many community requests asking for new models. We

local-air-localllama
3 Aug 2026
Model Releases

NousResearch keeps doing things on hermes

DGX agent

Has anyone followed nousresearch work on Hermes? I mean we are Q3 2026. We have some crazy models trickling down from HGX territory to multi gpu workstation. And we have nousresearch deploying the 0.2

model-releasesr-localllama
3 Aug 2026
Model Releases

Question about Quant versus Size.

DGX agent

Sorry if this is asked a lot, but I was wondering if there is any clear winner on the Quantization versus Model Size debate? I can run Qwen3.6 27b at Q8, Laguna at Q6, and the new Deepseek Flash at Q3

model-releasesr-localllama
3 Aug 2026
Model Releases

Qwen3.8-Max matches Kimi K3 and DeepSeek V4 Flash

DGX agent

Qwen3.8-Max (2.4T) is another massive contribution to the open weight community. On benchmarks, it performs closely to Kimi K3 and DeepSeek V4 flash across all categories and is better at coding and s

model-releasesr-localllama
3 Aug 2026
Model Releases

[RELEASE] SupraBrain-50M-v0.1

DGX agent

Hey there! So today we're releasing SupraBrain-50M, a hybrid language model that combines Gated DeltaNet linear recurrence with Sliding-Window Attention and Surprise-Gated update mechanisms to deliver

model-releasesr-localllama
3 Aug 2026
Local Ai

Running gpt-oss:20b locally and grading it head to head against a frontier model on real tasks. It held up better than I expected

DGX agent

I serve a free local model on my Mac Mini and route real agent work to it. To check I was not fooling myself, I set up a blind grader that replays frontier tasks locally and scores both. https://previ

local-air-ollama
3 Aug 2026
Model Releases

Speculative decoding with deepseek v4 flash 0731?

DGX agent

Has anyone figured out how to enable speculative decoding with deepseek v4 flash 0731 on llamacpp? I’m on the right release for llamacpp (b10228 or earlier) and running am17an’s draft model with unslo

model-releasesr-localllama
3 Aug 2026
Model Releases

The Chinese labs everyone lumps together are making four pretty different bets. I work at one of them.

DGX agent

Every time a model drops from a Chinese lab the thread fills with people who already know who made it, and the guess is usually Alibaba. There was a thread here recently asking what separates the open

model-releasesr-localllama
3 Aug 2026
Local Ai

The desktop app UI is getting a massive upgrade. What's next on the roadmap?

DGX agent

I've been following the recent pull requests and saw that the desktop app is being transformed from a chat-only interface into a full management tool with a tabbed settings UI, a model manager, and a

local-air-ollama
3 Aug 2026
Model Releases

V4-Flash-0731 - vibes after first weekend of use

DGX agent

Spent way too much time with V4-Flash-0731 this weekend and wanted to share my vibes as briefly as possible. I sent it through a bit of real-work and some of my personal benchmarks. My quick thoughts

model-releasesr-localllama
3 Aug 2026
Model Releases

Was the release of deepseek v4 flash planned to take spotlight against 5.6 luna?

DGX agent

Id figured since they first emailed people about api price changes coming mid july then delayed the v4 flash release to late july, I wonder if they delayed it for the sake of stealing spotlight from o

model-releasesr-localllama
3 Aug 2026
Local Ai

Xberg v1: a local, CPU-only document extraction engine for feeding an Ollama RAG (101 formats, OCR)

DGX agent

I maintain xberg, an open-source (MIT) document extraction engine, and v1 is out. Sharing here because a common piece of a local Ollama RAG setup is 'get clean text out of my PDFs/Office files/images,

local-air-ollama
3 Aug 2026
Local Ai

Are you ready for Le Chaton FAT or still wasting money on GPUs?

DGX agent

According to rumors (spread by myself) Le Chaton FAT will be 26T-a3b and I AM READY for it. Let's be real, I can't afford that many 5060Ti, so I got 12x Gen 4 3.2 TB (two per card). This gives me abou

local-air-localllama
2 Aug 2026
Model Releases

Best model <3B for multilingual understanding/ instruction following?

DGX agent

I know qwen 3.5 4b is great but a bit too large and miniPCM5 1b is great for agentic use but not so great for multilingual natural language understanding. Google eXb variants are just too big in total

model-releasesr-localllama
2 Aug 2026
Model Releases

Comfyui VRAM tracker

DGX agent

Hello! VRAM tracker is a node that track the full memory lifecycle of a comfyui run: when each weight is reserved, paged into VRAM, computed on, evicted, and freed. It renders it as an interactive HTM

model-releasesr-stablediffusion
2 Aug 2026
Model Releases

Conclusion: r/LocalLLaMA still has brilliant open-weight research, but finding it requires wading through endless benchmark drama, non-local Discussion Points and repetitive hardware flexes.

DGX agent

I let Gemma4-31b run on my laptop for like almost a day using a heavily altered pi to do a deep dive on our beloved Llama tangentially related Subreddit, and this was the conclusion. Feels pretty accu

model-releasesr-localllama
2 Aug 2026
Model Releases

Deepseek v4 flash - 100-150 faster t/s in prefill/pp.

DGX agent

You have two choices here (in order of pref): Downgrade CUDA from 13.3 to 13.1 (skip 13.2 due to bugs) <- prefer this (thanks to u/fairydreaming for pointing this out) Use this vibed fork that works w

model-releasesr-localllama
2 Aug 2026
Model Releases

DeepSeek-V4-Flash-0731 UD-IQ3_XXS about 11t/s on 1x 7900 XTX 24GB + 3x MI60 32GB + 128GB DDR4

DGX agent

Hello, Also I want to join the hype of posting token specs. CPU: 2x Intel Xeon CPU E5-2650 v4 @ 2.20GHz RAM: 2x 4 Channel 2400MHz DDR4 GPU: 1x AMD Radeon 7900 XTX 24GB 3x AMD Instinct MI60 32GB Strang

model-releasesr-localllama
2 Aug 2026
Model Releases

DeepSeek-V4-Flash-0731 UD-Q8_K_XL 17.20~ t/s on A6000 + 256GB DDR4

DGX agent

Hello everyone I want to join the hype of posting specs. CPU: AMD EPYC 74F3 24-Core RAM: 8 Channel 3200 DDR4 GPU: RTX A6000 48GB Prompt processing is in the high 70t/s (got down to mid 30t/s at 300k c

model-releasesr-localllama
2 Aug 2026
← Previous
1…45678…41
Next →