AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “hardware”

GridTimelineEvolution
268 results
Model Releases

100% Local RAG Without Internet and Without Ollama

DGX agent

Build a 100% offline fast Retrieval Augmented Generation (RAG) system that runs without an internet connection, without cloud APIs, without OpenAI/Ollama Published a video where you can build a fully

model-releasesr-ollama
7 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

llama.cpp PR reports up to 169% faster quantized-KV decode at 118K context on Intel Battlemage from one SYCL kernel switch

DGX agent

A fresh llama.cpp PR (#26689) changes what looks like a tiny SYCL FlashAttention dispatch decision. With a quantized KV cache ('q4_0' / 'q8_0'), decode was being sent through the VEC kernel. On the au

model-releasesr-localllama
7 Aug 2026
Model Releases

Serving Deepseek v4 Flash 0731 on 2x DGX Spark — 5-7 GB OS headroom, what would you do to lower VRAM usage and increase OS available RAM?

DGX agent

Hey all, I'm serving DSv4Flash 0731 on a cluster of 2x DGX Sparks but am running into constant issues with having almost no RAM (unified memory) left for the OS/cache and I'd love to hear the communit

model-releasesr-localllama
7 Aug 2026
Model Releases

Auto-fit vs tuned MoE offload: 564 → 1330 pp tok/s, unchanged decode (Qwen3.6-35B-A3B Q6 / RTX 3090)

DGX agent

TL;DR: On a Qwen3.6-35B-A3B Q6 setup sized for 64K context on a 24GB RTX 3090, spilling eight MoE expert layers to CPU freed enough VRAM to increase -b from 512 to 1024 and -ub from 128 to 512. Prompt

model-releasesr-localllama
6 Aug 2026
Model Releases

Unsloth's Gemma 4 mmproj silently broke vision & audio on newer llama.cpp builds — anyone else hit this?

DGX agent

So I had been building ScreenMind, kinda like local ai desktop assistant that uses Gemma 4 for screen analysis, voice memo transcription, and meeting transcription — all through llama-server. Everythi

model-releasesr-localllama
6 Aug 2026
Model Releases

Building a Fully Local PDF Read-Aloud & PDF-to-Audiobook Desktop App with Kokoro 82M, Qwen, and llama.cpp

DGX agent

Hey everyone, I’ve been building Speechfony - a desktop app for reading PDFs (and EPUBs) with offline text-to-speech. Open a document, listen sentence-by-sentence with highlighting, or export selected

model-releasesr-localllama
5 Aug 2026
Local Ai

I built xSignalBot: an auto-reply bot for Signal that answers with a local LLM via Ollama — zero cloud, zero cost

DGX agent

Disclaimer: I'm the developer of this project — sharing because it might be useful to others running self-hosted AI. (Full transparency, as Reddit's self-promo etiquette expects.) xSignalBot is an ope

local-air-ollama
5 Aug 2026
Model Releases

Inkling-Small 276B-A12B at ~2.9 tok/s on <10gb memory

DGX agent

A follow up to the launch of Mference, it now supports and runs Inkling-Small 276B-A12B. Inkling-Small (Thinking Machines, Apache 2.0), from the pipenetwork/Inkling-Small-MLX-4bit conversion: 276B tot

model-releasesr-localllama
5 Aug 2026
Model Releases

A 2.6B model with tool calling and 128K context now runs at 30 tok/s on a phone

DGX agent

Liquid AI released LFM2.5-2.6B today, and this might be more relevant to local AI than another massive model most people cannot run. The model is only 2.69B parameters, has 128K context, supports tool

model-releasesr-localllama
4 Aug 2026
Model Releases

[Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]

DGX agent

First of all, obviously I took some help from AI to type this post and this is the topic that enabled me to accomplish all that: https://old.reddit.com/r/LocalLLaMA/comments/1veow4b/deepseek_v4flash_2

model-releasesr-localllama
4 Aug 2026
Model Releases

DeepSeek V4 Flash 0731 (Q4) now reaches 1,328 tok/s prefill and ~29 tok/s decode on one RTX PRO 6000

DGX agent

I've been working on speeding up DeepSeek-V4-Flash-0731 in Krasis and have now got the long-prompt prefill quite a bit faster on a single RTX PRO 6000 96GB. These are timing-disabled internal Krasis r

model-releasesr-localllama
4 Aug 2026
Local Ai

70-class VRAM stagnation

DGX agent

been thinking about how the desktop 70-class has sat at 12GB for two generations now, 4070, 4070 super, 5070, all 12GB. the 1070 gave you 8GB back in 2016 and it felt generous for the price. ten years

local-air-localllama
3 Aug 2026
Model Releases

Döner Bench DeepSeek-V4-Flash IQ2_XS running on a single RTX 3090

DGX agent

https://preview.redd.it/3zcvpbds14hh1.png?width=1911&format=png&auto=webp&s=a79aafb71eeca97638da93d2591902631e897fd5 I tried a test similar to the recent model-quant comparisons, but this time I focus

model-releasesr-localllama
3 Aug 2026
Local Ai

I got tired of ad-filled mobile wrappers for Ollama, so I built PocketLLM Lite an open-source, offline Android client (Local GGUF, SKILL.md plugins, local RAG)

DGX agent

Hey, Like a lot of people here, I use local models via Ollama on my desktop/server and wanted a mobile client that actually felt responsive, worked offline, and respected privacy. Most apps on the Pla

local-air-ollama
3 Aug 2026
Local Ai

Xberg v1: a local, CPU-only document extraction engine for feeding an Ollama RAG (101 formats, OCR)

DGX agent

I maintain xberg, an open-source (MIT) document extraction engine, and v1 is out. Sharing here because a common piece of a local Ollama RAG setup is 'get clean text out of my PDFs/Office files/images,

local-air-ollama
3 Aug 2026
Model Releases

Parlor v2: best-effort fully local GPT-Live clone on an M3 Pro

DGX agent

GPT-Live is so good that I use it almost every day. I've been wanting to replicate it since it was released. My first attempt was to fine-tune Gemma 4 12B to behave like a full-duplex model. Something

model-releasesr-localllama
2 Aug 2026
Model Releases

[Release] WinterMix — Qwen3.5-122B-A10B in native MLX: an 82 GiB build that beats 94–95 GiB quants, plus a 68 GiB build for agent swarms

DGX agent

TL;DR: I spent 9 days developing a new quantization method for MLX models and measured 18 variants against each other on a single M5 Max MacBook Pro (128 GB). The result is the best-measuring MLX quan

model-releasesr-localllama
2 Aug 2026
Model Releases

Running DeepSeek-V4-Flash-0731 (155 GB MoE) on a DGX Spark with vLLM-Moet 2-bit quantization - AI's narrative

DGX agent

# Running DeepSeek-V4-Flash-0731 (155 GB MoE) on a DGX Spark with vLLM-Moet 2-bit quantization I used Deepseek-v4-Flash-0731 cloud API settig up vllm-moet to run deepseek-v4-flash with MTP locally on

model-releasesr-localllama
2 Aug 2026
Model Releases

DeepSeek-V4-Flash-0731: Models you can run locally now have the intelligence score of the top frontier model from March 2026

DGX agent

March 6th, 2026 the highest intelligence index score was 51 for frontier models. deepseek-ai/DeepSeek-V4-Flash-0731 that has an intelligence score of 50. If these benchmarks are accurate, models avail

model-releasesr-localllama
1 Aug 2026
Model Releases

Using an AMD V620 workstation card for ComfyUI - success

DGX agent

A few weeks ago I posted about if it was worth using a V620 for Comfyui, and was told it likely wouldn't work, at least in Windows 11. And if it did, it would be far too slow and unusable. I decided t

model-releasesr-stablediffusion
31 Jul 2026
Model Releases

Quantizing Kimi K3 (2.8T A50B) to GGUF ourselves - Q3_K_S works, 1.1 TB on disk

DGX agent

we're experimenting with our own dynamic GGUF quants of kimi k3, made from the original weights with our llama.cpp fork. Q3_K_S is done and works 1114.76 GiB on disk. Q1 and Q2 are in progress, result

model-releasesr-localllama
29 Jul 2026
Model Releases

DeepSeek V4 Flash, up to 32 tok/s on AMD Ryzen AI MAX+ 395

DGX agent

Hey fellow llamas. we have something new for Strix Halo owners we thought would be useful to share. i'll keep it short: We were able to fit DeepSeek V4 Flash plus its speculative draft on a single Ryz

model-releasesr-localllama
28 Jul 2026
Model Releases

I got Kimi-k3 running.....

DGX agent

Results: prompt eval: 40 tokens / 97.5s → 0.41 tok/s eval: 400 tokens / 1769.9s → 0.23 tok/s total: 440 tokens / 1867s (31 min) Prompt: 'Write a C++ function that reverses a linked list in place. Expl

model-releasesr-localllama
28 Jul 2026
Local Ai

K2Lab: Standalone(ish) Krea2 bbox style prompting and lora containment

DGX agent

I've been digging into Krea2 to see if there's any way to condition inference in specific regions of pixel -> latent space to implement bbox type prompting in order to apply multiple simultaneous char

local-air-stablediffusion
28 Jul 2026
Model Releases

Kimi K3 weights drop today. We're deploying on A100s, H200s and B300s this week and the A100 math is already rough

DGX agent

tldr; we are going to host K3 on A100s (yes, thats correct, we'll try to see if it holds up), H200s & B300s - expect results for A100s & H200s this week while we setup the B300 cluster this weekend &

model-releasesr-localllama
27 Jul 2026
Local Ai

Ornith-397B running at Q4 on a single RTX PRO 6000 Blackwell 96GB - 2,354 tok/s prefill, ~20–24 tok/s decode

DGX agent

I've been building Krasis, an MoE-focused runtime for streaming big models through limited VRAM on NVIDIA consumer/workstation GPUs, and I think this is the most interesting result so far: Ornith-1.0-

local-air-localllama
27 Jul 2026
Model Releases

Qwen3.6-27B speculative decoding gets better on heavier quants

DGX agent

I finished the speed leg of my spec-decode benchmarking for Qwen3.6-27B, main algorithms across quants. Overall: the heavier the quant, the more spec-decode buys you (10 of 10 speculative configs rank

model-releasesr-localllama
27 Jul 2026
Local Ai

Will prices finally go down?

DGX agent

I am seeing more and more videos as posts about how OpenAI is in complete financial ruin, Anthropic isn't much better. Their expenses go with the revenue they make etc etc. Meta made big investments i

local-air-localllama
26 Jul 2026
Model Releases

CachyLLama: llama.cpp fork with persistent SSD-backed KV caching for local agent workflows

DGX agent

If you run local agentic coding harnesses (Aider, Claude Code, etc.), prompt evaluation usually eats up most of your execution time. Every turn re-evaluates thousands of identical prefix tokens_system

model-releasesr-localllama
25 Jul 2026
Model Releases

I released Inflect v2: two ultra-tiny complete TTS models under 4M and 10M parameters

DGX agent

I’ve spent the past month trying to find the point where an extremely small TTS model stops feeling like a size experiment and starts feeling genuinely useful. Today I’m releasing Inflect v2, with two

model-releasesr-localllama
25 Jul 2026
Local Ai

I spent a year building a free SDXL & Anima trainer that runs on my 12 GB GPU — here's what came out of it

DGX agent

A little over a year ago I got frustrated trying to fine-tune SDXL on my RTX 3060. Every option either forced lower resolution, locked away important settings behind massive config files, or needed a

local-air-stablediffusion
25 Jul 2026
Local Ai

LFM 2.5 230M running at 1440 tok/s in-browser through a custom backend

DGX agent

Everything runs through WebGPU, in-browser or in electron/tauri apps. It's fully portable and supports either Nvidia and Apple Silicon (Metal). The actual kernels are optimized for the specific hardwa

local-air-localllama
25 Jul 2026
Model Releases

MI50 power curve tests

DGX agent

tests done power limiting the GPU on LACT - real power usage varies wildy at 20W it ranges from 25W to 56W same behavior happens on every setting prompt for the test runs: https://github.com/lukesdevl

model-releasesr-localllama
25 Jul 2026
Model Releases

CachyLLama’s: llama.cpp fork with persistent KV cache that makes long local-agent sessions much less painful

DGX agent

I’m not affiliated with this project, but I’ve been running it recently and I’m surprised it hasn’t received more attention here: https://github.com/fewtarius/CachyLLama CachyLLama is a fork of llama.

model-releasesr-localllama
24 Jul 2026
Model Releases

Getting the most out of MTP

DGX agent

If you want to get the most out of MTP. You have to run some tests / benchmarks to do so. Turning it on with defaults will get improvements, but for many models and card combinations, you are leaving

model-releasesr-localllama
24 Jul 2026
Model Releases

[Paper] Statistically-Lossless Quantization of Large Language Models

DGX agent

Model quantization has become essential for efficient large language model deployment, yet existing approaches involve clear trade-offs: methods such as GPTQ and AWQ achieve practical compression but

model-releasesr-localllama
24 Jul 2026
Local Ai

Spent two weeks on a kernel that benchmarked 29x faster. End to end it's maybe 6-10%, and it's not even wired in yet.

DGX agent

I've been building a C99 inference engine from scratch (no Python, no BLAS, just gcc and make) that runs BitNet's ternary models on CPU. A few weeks ago I got obsessed with the matmul kernel - wrote a

local-air-localllama
24 Jul 2026
Model Releases

Deepseek V4 Flash ~105 t/s on two Nvidia 4090d 48G (ada) in vLLM

DGX agent

TLDR: I (with the help of AI) re-implemented every Blackwell-only kernel (DeepGEMM, FlashInfer sparse-MLA, block-scaled FP8) in Triton, because they simply don't exist for sm89. The performance is 2-3

model-releasesr-localllama
23 Jul 2026
Local Ai

I run GLM-4.5-Air (110B) on 16Gb ram consumer machine and Qwen3-30B at 20 tok/s

DGX agent

In the past few months I’ve experimenting heavily and tortured my old 2016 Desktop PC to run the biggest Local LLM I can fit. I documented the whole process and research and I’ve published a repositor

local-air-ollama
23 Jul 2026
Model Releases

Trained a 32B FLUX.2 LoRA on a 24GB AMD 7900 XTX, native ROCm on Windows — full guide + patches

DGX agent

TL;DR: Everyone says QLoRA past ~13B is dead on a 24GB card. I got the full 32B FLUX.2 dev transformer QLoRA-training resident on the GPU on a 7900 XTX under native ROCm on Windows (no ZLUDA, no CUDA

model-releasesr-stablediffusion
23 Jul 2026
Model Releases

SkewAdam: A tiered optimizer that cuts MoE state memory by 97% (fits a 6.7B MoE on a 40GB GPU) [R]

DGX agent

Paper:https://arxiv.org/abs/2607.19058 Code (GitHub):https://github.com/nuemaan/skewadam Hi everyone, I just published a preprint on a new optimizer designed to tackle the massive VRAM bottleneck in M

model-releasesr-machinelearning
22 Jul 2026
Model Releases

Stuck scaling a Next.js app on M3 Pro (36GB) using local Qwen 3.6 + VS Code Copilot. Should I switch extensions or go paid?

DGX agent

Hey everyone, I’m a Full-Stack Developer with 6+ years of experience. I’m relatively new to AI-assisted development workflows and want to build a production-ready, enterprise-level Next.js web applica

model-releasesr-ollama
22 Jul 2026
Model Releases

NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B on 2x3090s

DGX agent

I managed to get this model working on 2x 3090s with full 262k ctx and N=4, if anyone is interested to try it, thanks to this quant: https://huggingface.co/danielrmay/NVIDIA-Nemotron-Labs-3-Puzzle-75B

model-releasesr-localllama
16 Jul 2026
Model Releases

I built a new attention mechanism (wave field) — runs 128K context where standard attention OOMs, 80+ tok/s on laptop CPU

DGX agent

Hey r/LocalLLaMA — solo researcher here. I built a new attention architecture and want independent testers. Wave Field LLM replaces O(N²) dot-product attention with FFT wave convolution on a field. Tr

model-releasesr-localllama
15 Jul 2026
Local Ai

I built a 100% local, CPU-only voice loop for Ollama — talk to your models hands-free (Silero VAD + Parakeet STT + Supertonic TTS 3)

DGX agent

A developer created a fully local, CPU-based voice interface for Ollama that enables hands-free conversation with AI models by combining three open-source components: Silero VAD (voice activity detect

local-air-ollama
11 Jun 2026
Model Releases

Running Gemma 4 QAT 12B on an 8GB GPU at 16k context — measured the KV-cache tradeoffs

DGX agent

This post discusses running Google's Gemma 4 QAT (Quantized Aware Training) 12B model on a GPU with 8GB of memory while maintaining a 16k token context window. The author likely shares performance ben

model-releasesr-ollama
11 Jun 2026
Local Ai

What is the best open sourced image model?

DGX agent

The best open-source image generation models in 2026 include FLUX.1 [schnell], Stable Diffusion 3.5 Large, HiDream-I1-Full, SANA-Sprint 1.6B, and HunyuanImage-3.0 . FLUX.1 [dev] holds the crown for ph

local-air-stablediffusion
10 Jun 2026
Local Ai

'Testing LCM on a GTX 750 Ti 4GB: Surprisingly Usable for Low-VRAM AI Image Generation'

DGX agent

This post documents testing Latent Consistency Models (LCM) on a GTX 750 Ti graphics card with 4GB of VRAM, demonstrating that this older, lower-end GPU can still run AI image generation models with a

local-air-stablediffusion
8 Jun 2026
← Previous
123456
Next →