AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “hardware”

GridTimelineEvolution
268 results
CompaniesToolsTechniques

Each lane shows up to 8 recent matching entries, ordered from earlier to later. Tracks load separately to keep the 75,000+ entry wiki fast.

Companies

CompanyAnthropic8 recent entries
22 Jul 2026Stuck scaling a Next.js app on M3 Pro (36GB) using local Qwen 3.6 + VS Code Copilot. Should I switch extensions or go paid?

Hey everyone, I’m a Full-Stack Developer with 6+ years of experience. I’m relatively new to AI-assisted development workflows and want to build a production-ready, enterprise-level Next.js web applica

→25 Jul 2026CachyLLama: llama.cpp fork with persistent SSD-backed KV caching for local agent workflows

If you run local agentic coding harnesses (Aider, Claude Code, etc.), prompt evaluation usually eats up most of your execution time. Every turn re-evaluates thousands of identical prefix tokens_system

3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
→26 Jul 2026Will prices finally go down?

I am seeing more and more videos as posts about how OpenAI is in complete financial ruin, Anthropic isn't much better. Their expenses go with the revenue they make etc etc. Meta made big investments i

→28 Jul 2026DeepSeek V4 Flash, up to 32 tok/s on AMD Ryzen AI MAX+ 395

Hey fellow llamas. we have something new for Strix Halo owners we thought would be useful to share. i'll keep it short: We were able to fit DeepSeek V4 Flash plus its speculative draft on a single Ryz

→30 Jul 2026P.A.I. — Sleek Native Desktop AI Overlayer for Local Ollama Models 🤖⚡

Greetings Community! 👋 I hope everyone is doing well! I'm Tauhid — Senior EEE student from a Bangladeshi University Today I'd like to share an open-source project I’ve been developing called P.A.I. (P

→31 Jul 2026Experience sharing: How do you use your local models and for what kind of tasks?

Here is my experience, which I would like to share with you and I also would like to hear your thoughts and valuable tips&tricks. Hardware: Mac Mini M4 (32GB Unified Memory) Model Server: Ollama Orche

→2 Aug 2026Encrypted Clouds?

I love the progress happening on open models but I feel like it is kind of getting clear that hardware to run good sized models is completely unaffordable for me right now. I know that you all love Qw

→12 Aug 2026Is the future of AI selling hardware for Open Source/Models?

I’m not super knowledgeable of the entire AI industry, but as we see this industry grow and the Cold War that is happening between the US and China on AI development, I can’t help but notice what US c

CompanyOpenAI8 recent entries
31 Jul 2026Open Source Ternary LLM Engine in Rust/CUDA for Quantization, Serving, and Training of models on consumer GPUs, called Tritium (Apache 2.0)

This post was not written by a clanker. Hey guys, I'm a comp sci major who wanted to introduce a cool project I built for quantizing models to ternary (1.58 bit) with as minimal of loss as possible, a

→2 Aug 2026Encrypted Clouds?

I love the progress happening on open models but I feel like it is kind of getting clear that hardware to run good sized models is completely unaffordable for me right now. I know that you all love Qw

→4 Aug 2026Why are Gamers so incredibly hostile to AI? Is it just a tiny vocal minority that spreads such toxic vitriol online?

It's more accurate to say that many highly engaged online gamers are hostile to AI, not that 'gamers' as a whole are. Gaming is a huge community with hundreds of millions of people, and opinions vary

→5 Aug 2026Inkling-Small 276B-A12B at ~2.9 tok/s on <10gb memory

A follow up to the launch of Mference, it now supports and runs Inkling-Small 276B-A12B. Inkling-Small (Thinking Machines, Apache 2.0), from the pipenetwork/Inkling-Small-MLX-4bit conversion: 276B tot

→6 Aug 2026I ported vLLM's serving stack to C++20: 66 MiB binary, no Python at inference, output checked token-for-token against vLLM

I'm the author, so discount the enthusiasm accordingly. This is an unaffiliated community port, not endorsed by the vLLM project, which it uses to verify its correctness. What started it: I love vLLM,

→7 Aug 2026Serving Deepseek v4 Flash 0731 on 2x DGX Spark — 5-7 GB OS headroom, what would you do to lower VRAM usage and increase OS available RAM?

Hey all, I'm serving DSv4Flash 0731 on a cluster of 2x DGX Sparks but am running into constant issues with having almost no RAM (unified memory) left for the OS/cache and I'd love to hear the communit

→11 Aug 2026I built a weird, low-power llama.cpp server using an Intel N100 + RTX 5060Ti

Everything started with the sudden death of my old ASRock J1900. While looking for the perfect ITX replacement, I stumbled upon the Chinese CW-NAS-ADLN-K motherboard, which looked perfect on paper: In

→12 Aug 2026Is the future of AI selling hardware for Open Source/Models?

I’m not super knowledgeable of the entire AI industry, but as we see this industry grow and the Cold War that is happening between the US and China on AI development, I can’t help but notice what US c

CompanyGoogle3 recent entries
13 Apr 2026I benchmarked Gemma4:e4b vs Gemma3:27B vs GPT-4o-mini vs Gemini 2.5 Flash on a Mac Mini M4 Pro 24gb — full results

A Reddit user on r/ollama conducted a hands-on benchmark comparing Gemma4:e4b (Google's compact ~4.5B effective-parameter edge model) against Gemma3:27B, GPT-4o-mini, and Gemini 2.5 Flash, all run or

→17 Apr 2026Ernie Image Turbo is not bad at all (Using INT8 quant and Gemini for prompt enhancement, RTX 30 series GPU with low vram)

Ernie Image Turbo is a text-to-image generation model that can run efficiently on consumer-grade hardware like RTX 30 series GPUs with limited VRAM by using INT8 quantization. The post discusses techn

→10 Aug 2026Need real world ML problems to evaluate my educational ML tools

I'm a retired platform engineer, coding mainly in Rust, and involved with a ML study group. I developed a ML programming language (alternative to Python, Colab) to help me learn (and teach) ML concept

CompanyMeta8 recent entries
7 Aug 2026Echo Dot 2 can run 28M LLM at decent speed

Code and instructions available here: https://github.com/albertoZurini/echo-dot-2-playground Hello there! After a few days of experimenting I was able to get a completely local voice pipeline running

→7 Aug 2026A llama.cpp PR makes Q2_0 3.0–3.6x faster on x86 CPUs, 8B decode goes 2.39 → 8.20 tok/s

I was going through the current llama.cpp CPU PRs and #26348 stood out because this isn't the usual +5% kernel optimization. It adds an x86 VNNI implementation for the Q2_0 × Q8_0 dot product, and the

→8 Aug 2026Extremely slow DSpark draft model performance (1-2 t/s) with DeepSeek-V4-Flash on llama-server compared to MTP?

Hey everyone, I could use some advice on setting up speculative decoding correctly with llama-server. My Hardware: GPUs: RTX 4090 + RTX 6000 Pro (120GB total VRAM) RAM: 32GB I am currently testing the

→8 Aug 2026enabling PCI-E p2p for consumer Nvidia cards will yield you more than you think

Disclaimer - no LLM was used to write this post/note As larger post about my setup will come later, want to give heads-up to folks who use VLLM and >= 2 GPUs. So I have pretty meaty server (8 channel

→9 Aug 2026KLQ: Training-free measured rotation quantization. Beats all training-free rotation-based quantization methods on W4A4KV4-bits. Llama 3.2 1B KLQ-quantized beats SpinQuant and gets close to ReSpinQuant without GPTQ/LDLQ rounding.

First of all, I'm not a lab, this was a solo summer research project that finally culminated into the github repo and the writeup. The repo includes a much deeper dive with methods, findings about qua

→11 Aug 2026I built a weird, low-power llama.cpp server using an Intel N100 + RTX 5060Ti

Everything started with the sudden death of my old ASRock J1900. While looking for the perfect ITX replacement, I stumbled upon the Chinese CW-NAS-ADLN-K motherboard, which looked perfect on paper: In

→11 Aug 2026DeepSeek V4 Flash 0731 at 27+ t/s decode on Strix Halo — Vulkan + DSpark full guide

Been benchmarking DSv4 Flash 0731 on a Flow Z13 (Ryzen AI MAX+ 395, Radeon 8060S / gfx1151, 128GB LPDDR5X) for the past week. Figured I'd share what actually works and what doesn't — there are a lot o

→12 Aug 2026Is the future of AI selling hardware for Open Source/Models?

I’m not super knowledgeable of the entire AI industry, but as we see this industry grow and the Cold War that is happening between the US and China on AI development, I can’t help but notice what US c

CompanyMistral7 recent entries
11 Apr 2026What's model should I run?

A Reddit discussion from the r/ollama community where a user seeks advice on which AI language model to run locally using Ollama. Responses likely include hardware-based recommendations (such as RAM a

→12 Apr 2026Any models?

A Reddit post on r/ollama where a community member asks about model availability or recommendations for use with the Ollama local AI runtime. The discussion likely covers which open-source models (suc

→13 Apr 2026I’m looking for advice on setting up a local AI model that can generate Word reports automatically.

This r/ollama thread discusses community advice on configuring a locally-run AI model (via Ollama) to automatically generate Word documents or reports, covering topics such as model selection, scripti

→5 Jun 2026What are the most capable LLM models I can run on my laptop?

A discussion on r/ollama exploring which high-performance LLM models can be effectively run locally on standard laptop hardware , likely covering model size comparisons, hardware requirements, and per

→23 Jul 2026Trained a 32B FLUX.2 LoRA on a 24GB AMD 7900 XTX, native ROCm on Windows — full guide + patches

TL;DR: Everyone says QLoRA past ~13B is dead on a 24GB card. I got the full 32B FLUX.2 dev transformer QLoRA-training resident on the GPU on a 7900 XTX under native ROCm on Windows (no ZLUDA, no CUDA

→27 Jul 2026Kimi K3 weights drop today. We're deploying on A100s, H200s and B300s this week and the A100 math is already rough

tldr; we are going to host K3 on A100s (yes, thats correct, we'll try to see if it holds up), H200s & B300s - expect results for A100s & H200s this week while we setup the B300 cluster this weekend &

→28 Jul 2026I got Kimi-k3 running.....

Results: prompt eval: 40 tokens / 97.5s → 0.41 tok/s eval: 400 tokens / 1769.9s → 0.23 tok/s total: 440 tokens / 1867s (31 min) Prompt: 'Write a C++ function that reverses a linked list in place. Expl

CompanyxAI2 recent entries
15 Apr 2026need a grok/neno-banana like img2img generation colab cell

A r/StableDiffusion thread where a user seeks a Google Colab notebook cell that replicates the img2img generation style or workflow of tools like 'grok' or 'neno-banana' — likely referring to specific

→26 Jul 2026Will prices finally go down?

I am seeing more and more videos as posts about how OpenAI is in complete financial ruin, Anthropic isn't much better. Their expenses go with the revenue they make etc etc. Meta made big investments i

CompanyDeepSeek8 recent entries
8 Aug 2026Extremely slow DSpark draft model performance (1-2 t/s) with DeepSeek-V4-Flash on llama-server compared to MTP?

Hey everyone, I could use some advice on setting up speculative decoding correctly with llama-server. My Hardware: GPUs: RTX 4090 + RTX 6000 Pro (120GB total VRAM) RAM: 32GB I am currently testing the

→8 Aug 2026enabling PCI-E p2p for consumer Nvidia cards will yield you more than you think

Disclaimer - no LLM was used to write this post/note As larger post about my setup will come later, want to give heads-up to folks who use VLLM and >= 2 GPUs. So I have pretty meaty server (8 channel

→10 Aug 2026Need real world ML problems to evaluate my educational ML tools

I'm a retired platform engineer, coding mainly in Rust, and involved with a ML study group. I developed a ML programming language (alternative to Python, Colab) to help me learn (and teach) ML concept

→10 Aug 2026DeepSeek V4 Flash 0731 is the ‘killer app’ that is going to sell A LOT of DGX Sparks

Having a ‘Killer Application’ that everyone wants to use helps sell hardware, plain and simple. DeepSeek V4 Flash 0731 isn’t an app of course, but I think it’s going to be the major catalyst for getti

→11 Aug 2026We quantized DeepSeek V4 0731 and benchmarked it against popular quants on 8× RTX 5090

We converted the model from the original safetensors and found two issues. The first one made our quantization fail several times, the second one does not fail at all, it just quietly ruins the base 1

→11 Aug 2026I ran Muse Glimmer @ 1M context - All tests passed.

Heeeey all! I just completed some fun tests with Muse Glimmer, I thought I'd let you know. In fact, the summary below was written by Muse itself! I ran a 2× DGX Spark cluster and got Meta's day-old Mu

→11 Aug 2026I gave DeepSeek V4 Flash basic vision by training a 40M connector on 100K examples

I wanted to find out whether a huge text-only MoE could be given basic vision without retraining the language model itself. The short answer is yes. I froze DeepSeek V4 Flash and a 417M-parameter Moon

→11 Aug 2026DeepSeek V4 Flash 0731 at 27+ t/s decode on Strix Halo — Vulkan + DSpark full guide

Been benchmarking DSv4 Flash 0731 on a Flow Z13 (Ryzen AI MAX+ 395, Radeon 8060S / gfx1151, 128GB LPDDR5X) for the past week. Figured I'd share what actually works and what doesn't — there are a lot o

CompanyNVIDIA8 recent entries
10 Aug 2026Need real world ML problems to evaluate my educational ML tools

I'm a retired platform engineer, coding mainly in Rust, and involved with a ML study group. I developed a ML programming language (alternative to Python, Colab) to help me learn (and teach) ML concept

→10 Aug 2026Introducing Muse Glimmer: an open-weight model optimized for always-on local agent workflows

Hi r/LocalLLaMA 👋 Today we’re excited to release Muse Glimmer, a 30B open-weight model built specifically for local agent workflows. We’re releasing the weights to the community under a permissive Apa

→10 Aug 2026I compared GGUF quants of Qwen3.6 27B to NVFP4, AWQ, AutoRound, and FP8

There's an interactive chart and some extra data in the blog post if you're interested. There are plenty of KL-divergence benchmarks for GGUF models, but most of them compare one GGUF quant against an

→10 Aug 2026DeepSeek V4 Flash 0731 is the ‘killer app’ that is going to sell A LOT of DGX Sparks

Having a ‘Killer Application’ that everyone wants to use helps sell hardware, plain and simple. DeepSeek V4 Flash 0731 isn’t an app of course, but I think it’s going to be the major catalyst for getti

→11 Aug 2026Nvidia Nemo Switchyard

https://github.com/NVIDIA-NeMo/Switchyard Finally an open source LLM router. An alternative to openrouter fusion and Sakana Fugu. Doesn't look like it does exactly what Sakana Fugu does according to i

→11 Aug 2026I ran Muse Glimmer @ 1M context - All tests passed.

Heeeey all! I just completed some fun tests with Muse Glimmer, I thought I'd let you know. In fact, the summary below was written by Muse itself! I ran a 2× DGX Spark cluster and got Meta's day-old Mu

→11 Aug 2026I built a weird, low-power llama.cpp server using an Intel N100 + RTX 5060Ti

Everything started with the sudden death of my old ASRock J1900. While looking for the perfect ITX replacement, I stumbled upon the Chinese CW-NAS-ADLN-K motherboard, which looked perfect on paper: In

→12 Aug 2026Is the future of AI selling hardware for Open Source/Models?

I’m not super knowledgeable of the entire AI industry, but as we see this industry grow and the Cold War that is happening between the US and China on AI development, I can’t help but notice what US c