MI25 for 80-100€ worth it?
seems to be about as good as a vega 56 with 16Gb of VRAM, is it worth it? (don’t want to deal with NVIDIA drivers on Linux, already have an rx6650xt and might simply use vulkan for llamacpp inference)
Knowledge catalogue
seems to be about as good as a vega 56 with 16Gb of VRAM, is it worth it? (don’t want to deal with NVIDIA drivers on Linux, already have an rx6650xt and might simply use vulkan for llamacpp inference)
This PR should be ready for testing now. I tested with a very small (8B params) sub-model extracted from the original one. Appreciate if someone can test with the bigger model. GGUF(for testing) from
Running across 2 clusters using llama.cpp over RPC too. Both clusters are not enough to hold everything in memory, so main cluster still partially offloads to run. Goal will be to get all the GPUs in
I am wondering if anyone can give opinion on if Ollama Cloud pro or max plans are worth it. Id be looking to use it with Kimi K3, Qwen 3.8 and Deepseek v4flash for now. Wondering if it would be better
Basically what the title says. For me, it would always crash and burn trying to use tensor split. Apparently, there's some bug where GPU memory gets corrupted with the default microbatch (512) or high
Hi everyone, I'm a systems analysis student researching a problem a lot of you probably know well: how much you actually trust the published VRAM/RAM requirements for open-source models before trying
I compared Qwen 35B-A3B MoE against Qwen 27B dense on a series of local coding-maintenance tasks. On my R9700/llama.cpp setup, the MoE model generated about 3.9× faster (~116 vs ~30 tok/s), but the co
I've been tuning my new Radeon AI Pro R9700, and figured that this would be useful information for people who are trying to optimise their setups. I'm pretty happy with these results and looking forwa
Not the biggest or shiniest, but it's mine From gaming machine inference on the original llama models, to a 4x RTX 6000 Pro Max Q + 4x 3090s local AI cluster. Pictures are in reverse chronological ord
MiniMax-H3 in ComfyUI 0.30.0, RTX 4080 16 GB (224 W cap). int8 DiT + int8 Qwen3-VL-32B text encoder. Turbo LoRA @ 0.9, euler + simple, 6 steps, no CFG. MiniMaxH3ReferenceToVideo with 3 reference image
Looking for V100 users to share your config and it's performance. GPU: Tesla V100 PCIE 32Gb Qwen3.6 27B Q4_K_M + Q8_0 MTP 128K context length Pi coding agent llama.cpp model preset: [*] spec-default =
Build a 100% offline fast Retrieval Augmented Generation (RAG) system that runs without an internet connection, without cloud APIs, without OpenAI/Ollama Published a video where you can build a fully
Follow-up to my original Spectrum MiniMax H3 post: https://www.reddit.com/r/StableDiffusion/comments/1vf1ze3/spectrum_acceleration_for_minimax_h3_in_comfyui/ In that first post, I released the MiniMax
I was going through the current llama.cpp CPU PRs and #26348 stood out because this isn't the usual +5% kernel optimization. It adds an x86 VNNI implementation for the Q2_0 × Q8_0 dot product, and the
I have not been successful with management to get funding for local resources despite bringing forth solid arguments about data sovereignty and related architectures. What actually succeeded in gettin
Or is there any reason why I feel like model output quality seems to be better when I use higher micro-batch values (ub) in llama-cpp? I don't really have any hard numbers or anything (just running th
Press My earlier prediction that Tesla would buy them completely missed the mark. With AMD focusing heavily on the enterprise side, the idea of consumer-facing hot-swappable AI model chips looks prett
Is anyone here successfully running DeepSeek-V4-Flash-0731 locally with vLLM, especially on AMD MI325X? My setup: GPU: 1x AMD Instinct MI325X Model: deepseek-ai/DeepSeek-V4-Flash-0731 vLLM: 0.26.0 ROC
today was the last day of my subscription on ollama cloud, to be honest it was a great price value for me and with GLM 5.2 and Deepseek V4 Pro i was able to Vibe code my custom woocomerce shop with mu
Code and instructions available here: https://github.com/albertoZurini/echo-dot-2-playground Hello there! After a few days of experimenting I was able to get a completely local voice pipeline running
Hello, For the past few days I have been benchmarking Gemma 4 26b QAT UD Q4_K_XL extensively versus Bartowski's Q4_K_L. While QAT is certainly very effective and reducing memory consumption versus the
Hey everyone, I just wanted to share my journey here for some motivation. Three years ago, I saw the sudden spike in AI and realized it was the future of tech. My goal at the time was to be an indie g
There are already plenty of different extensions for voice input, but all I found required having a second server running. I wanted something super simplistic: launching local STT server just for my p
Since now we have kimi k3 and next week we are getting Qwen 3.8 Max and also soon V4 pro Deepseek. I am curious if the old power house like Kimi 2.6/7 code and GLM.5.2 are all that relevant. especiall
LabyrinthBench measures the thing that actually kills long agent runs — whether a model can still use what it learned twenty turns ago — deterministically, with no LLM judge, on your own hardware, wit
LFM2.5-2.6B is a new tiny model by LiquidAI, with benchmarks that put it head to head with much larger models. I've run llama-perplexity on many model GGUF quants, crossed with many KV cache quants, t
A fresh llama.cpp PR (#26689) changes what looks like a tiny SYCL FlashAttention dispatch decision. With a quantized KV cache ('q4_0' / 'q8_0'), decode was being sent through the VEC kernel. On the au
I swear AA is not the bipartisan they so claim. An open source mode (Qwen 3.8 max) was number 1 on the agentic index, then they just so happen to launch 'v4.1.1' of their index in which they just adju
Hi there! I've messaged the Ollama support team 4 times with no response in 3 weeks. This is getting ridiculous. Does anyone have any recommendations as to how I should seek support? I don't want to i
High-performance inference of NVIDIA's Parakeet TDT 0.6B V2 English transcription model, in the browser. Check out the live demo: https://parakeet.narcotic.sh/ A fully custom, dependancy-free implemen
I'm considering replacing a single RTX 3090 with two ASRock AMD Pro R9700s for about 2900 new out of the door. That would move me from 24GB to 64GB VRAM. Yes yes, CUDA/ROCm, but the real problem is po
I run the following on a 5090 and have been okay with its performance, it does most things somewhere 80-100 t/s, though that can slow down at full 262k context - more like 40 t/s at times. I use it pr
Hey all, I'm serving DSv4Flash 0731 on a cluster of 2x DGX Sparks but am running into constant issues with having almost no RAM (unified memory) left for the OS/cache and I'd love to hear the communit
📝 Introduction We present Wan-Animate-2, a novel end-to-end character animation framework that directly consumes driving videos in a redesigned Diffusion Transformer, which achieves high-fidelity moti
For a long time now, the most popular posts on LocalLLaMA have been either about using LLM in the cloud or about politics. I suspect that people using local models are about 10% now. You can say that
Following up on yesterday's post about running everyone's faves on 2 x 16gb cards while maximizing performance and KV. Previous post data used abandoned Cu130 VLLM image. Stats here are done on cu129-
https://preview.redd.it/kihat320ashh1.png?width=1672&format=png&auto=webp&s=a7ccc40ba3fb229ac7ebf57e8e6a314e0ee45646 Hi r/StableDiffusion! u/New-Requirement1419 -> dacongya (Head of H3 Researcher) u/A
A small update to Sir Shortoken. Sir Shortoken already had Quick, Balanced, Deep, Bullets, and Aggressive Bullets. I wanted something between Bullets and normal prose. So I added LELP-S+ (Less English
TL;DR: On a Qwen3.6-35B-A3B Q6 setup sized for 64K context on a 24GB RTX 3090, spilling eight MoE expert layers to CPU freed enough VRAM to increase -b from 512 to 1024 and -ub from 128 to 512. Prompt
Hi all. These are my system specs: dual xeon e5 2696 v2 , 160gb DDR3 ram ECC(1600mhz), 3 gpus: 3060 12gb, p100 16gb, 3050 6gb. And a 400gb nvme sdd RAID0, 3000 mb/s. The model is Deepseek-flash-0731 U
Looking for best current solutions for combining cloud models and local models seamlessly inside a harness' orchestration Edit: Right now, we don't have harnesses (that I'm aware of) that are blending
Up until now, I’d always assumed it was basically just transcribing what I said, feeding the text into the model, and reading the response back. So out of curiosity, I asked if it could actually tell
First of all, my setup: Ryzen 9 5950x DDR4 3200Mhz 64gb (2x32) Dual 3090s, no NVLINK Runtime: llama.cpp Nvidia Drivers 610 Windows 11 25H2 Qwen 3.6 27B Q8 I've been using llama-server with --split-mod
J'ai consacré beaucoup de temps à l'optimisation de DeepSeek-V4-Flash-0731 GGUF sur une seule RTX 3090. Mon exigence absolue pour chaque configuration était la suivante : Le modèle doit rester utilisa
So I can either pull the trigger on a 128gb AI max+ 395 laptop or wait for RTX Spark for LLMs. Maybe I get it now and the price of the spark is super high so it's a good purchase or maybe the Spark sh
https://preview.redd.it/o6ik6qboeohh1.png?width=1134&format=png&auto=webp&s=4016f26c50c1d93bd3d0c7e880e9b55a2d75310f I have been running Qwen3.6 27b for a little while (mostly coding tasks) and recent
Just came across this coding benchmark: SciCode Artificialanalysis.ai reports a ranking which contradicts the feeling we've towards those models in real life coding. Is Gemma 4 really that good, or a
I cancelled my pro plan ealier because I wanted to use new Deepseek v4 flash 0731 which was available on Openrouter through API only (not yet on ollama cloud at the time). The old Deepseek v4 flash/pr
As for me, I own a system with an RTX 5090, Ryzen 9 9950X3D2, and 64 GB of DDR5. Every time I see research come out with a new way to train AI, I immediately think to try it on my system to see the re
Seen a ton of posts today about the DeepSeek API price hike. Half the feed is doom-posting, the other half is explaining basic GPU economics. Honestly, I get the cost side. Sub-cent tokens were never
[Fully open source under GPL3, made from the ground up for use with local models, no subscriptions, no corporate backing] When i first started this, it was meant to be a fully lightweight, extremely m
I'm the author, so discount the enthusiasm accordingly. This is an unaffiliated community port, not endorsed by the vLLM project, which it uses to verify its correctness. What started it: I love vLLM,
I built this because the existing benchmarks were using random data and with MTP content types can vary a lot on what performance you see. 5% or more with content types. BetterBench is designed to hav
Link to the article: KV Cache Quantization Benchmarks: KVarN, Precision Tail KLD benchmarks with BeeLlama.cpp v0.4.0, fork of llama.cpp with more KV cache quantization options. Models: Qwen 3.6 27B Q5
Salve a tutti. Sto valutando di mettere dei modelli locali, magari su LM studio o altro software se mi spiegate il perché da utilizzare sia come PT sia per Soc L2/L3. Ho un PC con 128 GB RAM ddr5 8gb
NVIDIA Nemotron Parse 2.0 transforms document images into structured, machine-readable representations with text, layout classes, bounding boxes, and reading-order information. Given a Red, Green, Blu
nvidias nemotron omni is open weights and it sees, hears and reasons. theres already a 4bit mlx quant on hugging face but only the text backbone loads with standard mlx tooling. the model card says it
🐦⬛ Magpie-TTS Multilingual 🦜 Nemotron Speech Streaming EN 0.6B 🦜 Nemotron-3.5 ASR Streaming 🦜 Parakeet CTC 1.1B 🦜 Parakeet TDT 0.6B v3 🥦 NanoCodec Merged PR https://huggingface.co/nvidia/magpie_tts_m
GGUFs here: https://huggingface.co/ReadyArt/gemma-4-31B-it-scotoma-2-GGUF Disclaimer: By slop, we are specifically talking about specific tics with the model(sentence structures), but this doesn't inc
I love to see these impressive models coming out that compete with the giants from companies like Z.ai, Moonshot, Alibaba, etc. A win for the open source/weight community is always welcome. While I am