AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
Human
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “hardware”

GridTimelineEvolution
61+ results
12 Aug 2026

Is the future of AI selling hardware for Open Source/Models?

HardwareDGX agent

I’m not super knowledgeable of the entire AI industry, but as we see this industry grow and the Cold War that is happening between the US and China on AI development, I can’t help but notice what US c

What do you guys do for GPU Kernels?

HardwareDGX agent

I'm trying to figure out GPU Kernel optimization on older hardware like SM80(ampere) . Is there tools you guys use? Or frameworks? Im waiting for this framework https://www.reddit.com/r/LocalLLaMA/com

26 Jul 2026

100B Models on Cheap Hardware: how realistic and limitations

Model ReleasesDGX agent
DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

There is a lot of buzz around running 100B parameter models on cheap local hardware using ternary (1.58-bit) quantization like Microsoft's BitNet architecture. The theoretical hardware shortcuts are i

Understanding GPU Inference Workloads [D]

HardwareDGX agent

Hey everyone, I have been looking into how people source compute for their Inference workloads (and in general). I wanted to understand some specific pain points here. If you've used online services l

14 Apr 2026

MoBo, CPU, RAM suggestion | I'm terrible at guessing hardware | no gpu

HardwareDGX agent

This Reddit thread from r/ollama features a user seeking community recommendations for motherboard, CPU, and RAM components to build a system optimized for running Ollama locally without a dedicated G

IC LoRAs for LTX2 have so much potential - you can train SOTA control video capabilities on potato hardware - 4 examples w/ links below

Local AiDGX agent

IC-LoRA (In-Context LoRA) enables conditioning video generation on reference video frames at inference time, allowing fine-grained video-to-video control on top of a text-to-video base model. Unlike t

what's the best place to buy GPU server?

HardwareDGX agent

This Reddit thread from r/ollama discusses community recommendations for purchasing or renting GPU servers to run Ollama and local LLMs. It likely covers options ranging from dedicated GPU server prov

13 Apr 2026

Looking for people with different hardware to help benchmark local LLM behavioral reliability

Model ReleasesDGX agent

A Reddit post in the r/ollama community seeking volunteers with diverse hardware setups to participate in a collaborative effort to benchmark the **behavioral reliability** of locally-run large langua

3 Aug 2026

How to run big models on old hardware 30B at 22 tok/s on 6GB GPU and 16GB RAM

Model ReleasesDGX agent

I have been working on this tool for months and there are a lot of new functionalities and tests that are going to be released in the next few weeks! The goal of the tool is to allow community members

'Data center in a Box (on Wheels)' 256Gb VRAM/512Gb RAM AI Server 6-8 Month Operational Review, Stability Write Up, Benchmarks

Model ReleasesDGX agent

I've been out of these forums for awhile but I figured I would provide a formal update on how this has been going now that it has some operation time under its belt, just to put the information out th

DeepSeek V4-Flash (284B MoE) at 33 tok/s single / 68 tok/s aggregate on 2× RTX 3090 + a used quad-Xeon DDR4 server — full config

Model ReleasesDGX agent

Ran DeepSeek V4-Flash-0731 — the full official checkpoint, not a re-quant — on commodity used hardware. Sharing because I couldn't find anyone else publishing Ampere results for this engine. Why bothe

11 Apr 2026

Musicvideo on local Hardware

Local AiDGX agent

A Reddit post on r/StableDiffusion in which a user shares a music video created entirely using AI-generated imagery produced on their own local hardware, likely using tools such as Stable Diffusion wi

My 2026 Ollama Setup Guide: What Actually Works Best for Daily Use on Consumer Hardware

Local AiDGX agent

This r/ollama community post is a practitioner's guide sharing personal, real-world experience running Ollama on everyday consumer hardware in 2026, covering which models, quantization settings, and c

7 Jun 2026

Tired of bloated UIs for Ollama? Built a minimal IDE that infers your hardware and just works

Local AiDGX agent

A developer created a minimal IDE for Ollama that automatically detects hardware capabilities and provides a streamlined user interface without unnecessary complexity. The tool aims to simplify the us

6 May 2026

OpenAI phone leaks: a push toward AI first hardware.

IndustryDGX agent

OpenAI is fast-tracking development of its first AI-focused smartphone with mass production targeted for early 2027 , marking the company's entry into consumer hardware . The device, positioned as an

21 Apr 2026

Anyone here using ai agent orchestration software to control multiple hermes agents? I'm retired and have some extra hardware

AgentsDGX agent

This Reddit post from r/ollama asks the community about AI agent orchestration software for managing multiple Hermes agents, posted by a retired individual with available hardware resources. The post

9 Aug 2026

Open-weight video gen that actually delivers. Five days with MiniMax H3 on local hardware.

Local AiDGX agent

H3 weights went live on HuggingFace August 3rd and I started pulling them immediately. An omni-modal video model with native stereo audio in the same forward pass, where audio can actually drive the v

300b on 32gb MoE-streaming findings + optimisations

HardwareDGX agent

The past week I've been running DSv4 inference on my laptop by keeping everything RAM-resident except the MXFP4-experts (since expert pool is ~147GB and won't fit) TL;DR - read speed is the limiter mo

I Turned My Underused Gaming Laptop Into a Local AI Workstation

Local AiDGX agent

TL;DR: I am building a Windows-first local AI setup for people who want to try local LLMs without spending days choosing models, setting up Ollama, Docker, WSL, Open WebUI, agents, and tool permission

Open Model: Google Weather Next 2

HardwareDGX agent

I am not a meteorologist, but I just read a very interesting article: https://arstechnica.com/science/2026/08/deepminds-hurricane-model-bought-forecasters-an-extra-day/ In a paper published on Thursda

28 May 2026

[Guide] How to securely run ComfyUI on Windows (Docker>WSL2) [RTX 3090, logic can be applied to other hardware]

TutorialsDGX agent

ComfyUI is a node-based interface for Stable Diffusion that can be run on Windows using Docker and WSL2 (Windows Subsystem for Linux 2) for improved security and performance. This guide provides instr

12 May 2026

Use Case: Invoice processing with local LLM - Which LLM and hardware requirements?

Local AiDGX agent

This discussion explores using Ollama to run large language models locally for invoice processing while maintaining control over data. The thread likely addresses selecting appropriate smaller models

17 Apr 2026

Best Ollama model for n8n workflows (RAG, file handling, reasoning) + hardware requirements?

Local AiDGX agent

Models like Qwen3 and Llama 3.2 are commonly used for n8n RAG workflows , with selection depending on use case requirements. Running Ollama models locally requires at least 16 GB of RAM on your device

8 Aug 2026

Quick survey (2 min) on trust in hardware specs for open-source models

Local AiDGX agent

Hi everyone, I'm a systems analysis student researching a problem a lot of you probably know well: how much you actually trust the published VRAM/RAM requirements for open-source models before trying

2 Aug 2026

Conclusion: r/LocalLLaMA still has brilliant open-weight research, but finding it requires wading through endless benchmark drama, non-local Discussion Points and repetitive hardware flexes.

Model ReleasesDGX agent

I let Gemma4-31b run on my laptop for like almost a day using a heavily altered pi to do a deep dive on our beloved Llama tangentially related Subreddit, and this was the conclusion. Feels pretty accu

I pushed Kimi K3 onto one CPU with 8 GB of RAM

HardwareDGX agent

I deployed K3 on 32 H100s at work a couple of weeks ago and then got annoyed that there was no way to poke at it on my own machine. So I wrote an inference engine for it in C99. Nothing clever going o

3 May 2026

torch-nvenc-compress: GPU NVENC silicon as a PCIe bandwidth multiplier — PCA + pure-ctypes Video Codec SDK wrapper. Parallel-path overlap measured at 67% of theoretical max on a real GEMM + encode workload. [P]

HardwareDGX agent

torch-nvenc-compress is a Python library that leverages GPU NVENC (NVIDIA's hardware video encoding) to optimize PCIe bandwidth utilization by compressing data during transfer. The project implements

4 Aug 2026

Why are Gamers so incredibly hostile to AI? Is it just a tiny vocal minority that spreads such toxic vitriol online?

Model ReleasesDGX agent

It's more accurate to say that many highly engaged online gamers are hostile to AI, not that 'gamers' as a whole are. Gaming is a huge community with hundreds of millions of people, and opinions vary

Has anyone been working on a solid setup for DSV4F on x2+ R9700s?

HardwareDGX agent

I'm hoping that one of you guys has been working on an inference engine or has somehow found improvements to running DSV4F on RDNA4 multi-GPU setups. I am currently building a custom inference engine

11 Aug 2026

DeepSeek V4 Flash 0731 at 27+ t/s decode on Strix Halo — Vulkan + DSpark full guide

Model ReleasesDGX agent

Been benchmarking DSv4 Flash 0731 on a Flow Z13 (Ryzen AI MAX+ 395, Radeon 8060S / gfx1151, 128GB LPDDR5X) for the past week. Figured I'd share what actually works and what doesn't — there are a lot o

I ran Muse Glimmer @ 1M context - All tests passed.

Model ReleasesDGX agent

Heeeey all! I just completed some fun tests with Muse Glimmer, I thought I'd let you know. In fact, the summary below was written by Muse itself! I ran a 2× DGX Spark cluster and got Meta's day-old Mu

Nvidia Nemo Switchyard

HardwareDGX agent

https://github.com/NVIDIA-NeMo/Switchyard Finally an open source LLM router. An alternative to openrouter fusion and Sakana Fugu. Doesn't look like it does exactly what Sakana Fugu does according to i

5 Jun 2026

What are the most capable LLM models I can run on my laptop?

Local AiDGX agent

A discussion on r/ollama exploring which high-performance LLM models can be effectively run locally on standard laptop hardware , likely covering model size comparisons, hardware requirements, and per

29 May 2026

What is the best used or refurbished laptop with GPU for open source Imege generation?

HardwareDGX agent

This Reddit discussion in the StableDiffusion community addresses recommendations for affordable, used or refurbished laptops equipped with GPUs suitable for running open-source image generation model

15 Apr 2026

Cline with Ollama on a RTX4090 (24GRAM) and i9 with 64 GRAM

Local AiDGX agent

This Reddit post from r/ollama discusses a user's experience running Cline (an AI coding agent) with Ollama on a high-end local hardware setup consisting of an NVIDIA RTX 4090 with 24GB VRAM and an In

Roop Unleashed 4.3.1 not fully utilizing RTX 5070 Ti / 5080X (Low GPU/RAM usage)

HardwareDGX agent

This Reddit thread from r/StableDiffusion discusses a performance issue where Roop Unleashed version 4.3.1 — a face-swapping tool commonly used alongside Stable Diffusion — fails to fully utilize the

Running a 31B model locally made me realize how insane LLM infra actually is

Local AiDGX agent

A Reddit post from r/ollama in which a user shares their experience running a 31B parameter model locally using Ollama, reflecting on the surprisingly demanding hardware and infrastructure requirement

7 Aug 2026

Echo Dot 2 can run 28M LLM at decent speed

Model ReleasesDGX agent

Code and instructions available here: https://github.com/albertoZurini/echo-dot-2-playground Hello there! After a few days of experimenting I was able to get a completely local voice pipeline running

LabyrinthBench: a local-focused, judge-free LLM benchmark that measures context recall under interference for multi-step agentic tasks.

Model ReleasesDGX agent

LabyrinthBench measures the thing that actually kills long agent runs — whether a model can still use what it learned twenty turns ago — deterministically, with no LLM judge, on your own hardware, wit

5 Aug 2026

RTX 3090 MiniMax H3 Speed Comparison: FP8 Scaled vs INT8 ConvRot (W8A8)

HardwareDGX agent

Setup: GPU: RTX 3090 24GB RAM: 32GB ComfyUI 0.30.0 PyTorch 2.13.0+cu130 CUDA 13.0 SageAttention enabled (sageattention-2.2.0+cu130torch2.10.0andhigher.post6-cp310-abi3-win_amd64) Spectrum node + Euler

31 Jul 2026

Open Source Ternary LLM Engine in Rust/CUDA for Quantization, Serving, and Training of models on consumer GPUs, called Tritium (Apache 2.0)

Model ReleasesDGX agent

This post was not written by a clanker. Hey guys, I'm a comp sci major who wanted to introduce a cool project I built for quantizing models to ternary (1.58 bit) with as minimal of loss as possible, a

30 Jul 2026

I built ganfs: A Python package that uses GANs to automate feature selection for high-dimensional datasets. (No domain expert required) [P] [R]

HardwareDGX agent

Hey everyone, I recently open-sourced a new Python package called ganfs (Generative Adversarial Network Feature Selection), and I wanted to share it with the community. The Problem: Selecting the best

unsloth/Qwen3.6-27B-NVFP4 vs. Intel/Qwen3.6-27B-int4-AutoRound vs. nvidia/Qwen3.6-27B-NVFP4 -- which one to choose?

HardwareDGX agent

Are there any benchmarks on these 4 bit quants, like how Artificial Analysis runs a slew of various benchmarks? If not, how can I run one (5x over for consistency) on them? I'm also very interested in

28 Jul 2026

Are single GPU research still published in ML/DL and its applications nowadays? Which are the most notable recent ones? [D]

HardwareDGX agent

ML research is progressing at breakneck speed where frontier labs in both academia and industry have access to considerably large computes (GPUs). Where do small labs or independent researchers go in

27 Jul 2026

Built & Trained a Transformer from Scratch in Pure PyTorch for English-to-Tamil Machine Translation [Math + Code Breakdown] [P]

HardwareDGX agent

Hi everyone! 👋 I built and trained the complete Transformer architecture from scratch using pure PyTorch (`torch.nn` primitives) based on the original 'Attention Is All You Need' paper. I trained the

25 Jul 2026

I tried making a cinematic action trailer using Krea 2 + LTX 2.3

HardwareDGX agent

I wanted to challenge myself and see how far I could push Krea 2 and LTX 2.3, so I decided to create a short cinematic action trailer. It ended up being one of the most enjoyable AI projects I've work

Mobile Offline LLMs: What do you use them for?

Model ReleasesDGX agent

I've spent the last year or so playing around with open source MLX and GGUF models on iPhone hardware. Given the limitations in memory, GPU/CPU/ANE, and in turn the context window I've been trying to

SVDQuant + native INT8/W4A4 for Krea 2 on ComfyUI — up to 2x faster, works on any modern NVIDIA GPU

HardwareDGX agent

Quantized Krea 2 Turbo checkpoints for ComfyUI, up to 2x faster and about a third smaller than the usual FP8 version — no calibration dataset, no quality cliff. How to use it (short version): clone th

21 Jul 2026

Looking for feedback on my GPU-accelerated Snake AI project [P]

HardwareDGX agent

I've been building an AI that learns to play the classic Snake game through reinforcement learning. The goal is to reach high scores while keeping training time as low as possible. The current version

20 Jun 2026

An open handbook on LLM inference at scale (GPU internals, KV cache, batching, vLLM/SGLang/TensorRT-LLM) [P]

HardwareDGX agent

This handbook covers the technical aspects of running large language models efficiently at scale, focusing on GPU optimization techniques including GPU internals, key-value (KV) cache management, batc

9 Jun 2026

Haven't seen much about the Nvidia Cosmos 3 video model that dropped, what's up with that?

HardwareDGX agent

Nvidia Cosmos 3 is an open physical AI foundation model built on a mixture-of-transformers architecture that combines vision reasoning, world generation, and action prediction for reasoning, simulatio

8 Jun 2026

Xiaomi just claimed 1,000+ tps on a 1T model using a standard 8-GPU server

HardwareDGX agent

Xiaomi achieved over 1,000 tokens per second output from a 1 trillion-parameter model using a single standard 8-GPU commodity node through extreme model-system codesign . The approach combines FP4 qua

6 Jun 2026

Mac mini M4 vs Pc with Nvidia 5060 8gb for ai workloads?

HardwareDGX agent

The Mac mini M4 uses unified memory architecture where CPU and GPU share a single 24GB memory pool, while the RTX 5060 has dedicated VRAM. Mac mini M4 is preferred for large model inference (70B param

1 Jun 2026

Nvidia releases Cosmos3-Super-Image2Video . 64B parametres

HardwareDGX agent

Cosmos3-Super-Image2Video is a 64B model for temporally coherent image-to-video generation . NVIDIA's Cosmos platform is designed to accelerate Physical AI development by enabling machines to understa

Nvidia releasesCosmos3-Super-Text2Image model . 64 billion paramteres

HardwareDGX agent

Cosmos3-Super-Text2Image is a 64 billion parameter omnimodal world model capable of generating high-quality images from text inputs as part of NVIDIA's Cosmos 3 foundation model platform. The model is

25 May 2026

Testing ZIT and Flux-1 with 'NVIDIA PiD — Pixel Diffusion Decoder'

HardwareDGX agent

NVIDIA's Pixel Diffusion Decoder (PiD) is an open-source decoder that replaces VAE decoders in image generation pipelines without retraining, producing sharper fine details and textures through a lear

21 May 2026

I got tired of API limits, so I hooked up OpenClaw to an unlimited Qwen3.6:35b backend on a full H100 for $1.6/hr (Demo)

HardwareDGX agent

This post describes setting up OpenClaw with a self-hosted Qwen 3.6:35b language model backend running on an H100 GPU for approximately $1.60 per hour, eliminating API rate limits. The user shares the

18 May 2026

🧬 flux-genotype: A self-evolving AI kernel that runs on CPU with Ollama — mutates its own architecture

Local AiDGX agent

Flux-genotype is a self-evolving AI kernel designed to run on CPU hardware using Ollama, featuring the capability to mutate and adapt its own architecture dynamically. This project demonstrates an app

16 May 2026

Reduce your GPU power limit

HardwareDGX agent

Setting a GPU power limit using nvidia-smi reduces heat output by approximately 20% with only a 5-8% inference speed loss. Undervolting the GPU can reduce power consumption by 5-15% with zero performa

5 May 2026

No GPU utilization

HardwareDGX agent

Ollama not using GPU is commonly diagnosed by running `ollama ps`—if it shows 100% CPU, Ollama isn't detecting the GPU. The issue typically has multiple causes, with common ones including driver probl

← Previous
1
Next →
268 results
← Previous
123…5
Next →