CompanyAnthropic6 recent entries12 Apr 2026KIV: 1M token context window on a RTX 4070 (12GB VRAM), no retraining, drop-in HuggingFace cache replacement - Works with any model that uses DynamicCache [P]KIV is a project shared on r/MachineLearning presenting a drop-in replacement for HuggingFace's `DynamicCache` that enables up to 1 million token context windows on consumer hardware with only 12GB of→21 Apr 2026I built a Free OpenSource CLI coding agent specifically for 8k context windows and local LLMs.A free, open-source CLI coding agent that runs locally using Ollama and can read, write, and edit files while running shell commands through natural language. The tool supports Claude API as an altern
CompanyOpenAI8 recent entries10 Apr 2026How to get ChatGPT to do substantially more work.I was unable to retrieve the specific Reddit post at the URL provided (reddit.com/r/ChatGPT/comments/1shtuw2), and the search results did not surface its content. The post may be removed, private, ...→10 Apr 2026Grok for youThe specific Reddit post (r/ChatGPT, post ID `1shi3fn`, titled 'Grok for you') was not directly retrievable or indexed in search results. Based on the available context from surrounding community d...→10 Apr 2026Err... what's going on?I was unable to retrieve the specific Reddit post from the URL provided, and my search did not return relevant results for that thread. Reddit content is often not indexed in real time, and I canno...→11 Apr 2026what happenedI was unable to directly access the specific Reddit post at the URL provided (r/ChatGPT post ID `1si4tvq`), and the web search did not return that specific post in its results. Reddit posts are oft...→13 Apr 2026Is it possible to rank in both SERP and AI Overviews?Yes, it is possible to rank in both traditional SERP results and Google AI Overviews simultaneously. Google pulls content sources from its index before determining rankings or SERP features, meaning t→18 May 2026Witchcraft, fast local semantic search on top of SQLite [P]Witchcraft is a Rust reimplementation of Stanford's XTR-Warp semantic search engine that uses a single-file SQLite database for storage, enabling client-side deployment. The system operates completely→24 Jul 2026I built an open-source multi-agent SDLC harness that beats a cold Claude Code run on large repos, by learning the repo once. Real benchmarks (incl. where it loses) inside. [P]Built an open-source AI coding agent that was 7%–75% cheaper than a cold 'claude -p' run on 6/6 well-localized tasks across repositories up to ~82k LOC. The biggest difference: Cold agent: 6.83, 207 t→7 Aug 2026Serving Deepseek v4 Flash 0731 on 2x DGX Spark — 5-7 GB OS headroom, what would you do to lower VRAM usage and increase OS available RAM?Hey all, I'm serving DSv4Flash 0731 on a cluster of 2x DGX Sparks but am running into constant issues with having almost no RAM (unified memory) left for the OS/cache and I'd love to hear the communit
CompanyGoogle2 recent entries24 Jul 2026I built an open-source multi-agent SDLC harness that beats a cold Claude Code run on large repos, by learning the repo once. Real benchmarks (incl. where it loses) inside. [P]Built an open-source AI coding agent that was 7%–75% cheaper than a cold 'claude -p' run on 6/6 well-localized tasks across repositories up to ~82k LOC. The biggest difference: Cold agent: 6.83, 207 t→2 Aug 2026Vacuum 16Thttps://huggingface.co/tsfrm/vacuum-16t A 16.5-trillion-parameter model that contains nothing. This model is just a ████ you to the labs and companies who say that 'haha I have the biggest model out t
CompanyMeta3 recent entries28 Jul 2026DeepSeek V4 Flash, up to 32 tok/s on AMD Ryzen AI MAX+ 395Hey fellow llamas. we have something new for Strix Halo owners we thought would be useful to share. i'll keep it short: We were able to fit DeepSeek V4 Flash plus its speculative draft on a single Ryz→2 Aug 2026Expert-only IQ3 requant of DeepSeek-V4-Flash-0731: better KLD than UD-IQ3_S, 1.4x decode on a CPU-spill rigHey all, tldr / who this helps: you run a mixed multi-GPU box where the experts spill to RAM, and you want to stay in the 3-bit tier instead of dropping to Q2 to make it fit. https://huggingface.co/Ta→7 Aug 2026Serving Deepseek v4 Flash 0731 on 2x DGX Spark — 5-7 GB OS headroom, what would you do to lower VRAM usage and increase OS available RAM?Hey all, I'm serving DSv4Flash 0731 on a cluster of 2x DGX Sparks but am running into constant issues with having almost no RAM (unified memory) left for the OS/cache and I'd love to hear the communit
CompanyMistral1 recent entries6 Aug 2026nvidia/NVIDIA-Nemotron-Parse-2.0 · Hugging FaceNVIDIA Nemotron Parse 2.0 transforms document images into structured, machine-readable representations with text, layout classes, bounding boxes, and reading-order information. Given a Red, Green, Blu
CompanyxAI4 recent entries10 Apr 2026Grok for youThe specific Reddit post (r/ChatGPT, post ID `1shi3fn`, titled 'Grok for you') was not directly retrievable or indexed in search results. Based on the available context from surrounding community d...→10 Apr 2026Err... what's going on?I was unable to retrieve the specific Reddit post from the URL provided, and my search did not return relevant results for that thread. Reddit content is often not indexed in real time, and I canno...→24 Jul 2026I built an open-source multi-agent SDLC harness that beats a cold Claude Code run on large repos, by learning the repo once. Real benchmarks (incl. where it loses) inside. [P]Built an open-source AI coding agent that was 7%–75% cheaper than a cold 'claude -p' run on 6/6 well-localized tasks across repositories up to ~82k LOC. The biggest difference: Cold agent: 6.83, 207 t→2 Aug 2026Vacuum 16Thttps://huggingface.co/tsfrm/vacuum-16t A 16.5-trillion-parameter model that contains nothing. This model is just a ████ you to the labs and companies who say that 'haha I have the biggest model out t
CompanyDeepSeek8 recent entries9 May 2026DeepSeek V4 paper full version is out, FP4 QAT details and stability tricks [D]DeepSeek released the full technical report for DeepSeek-V4 on April 24, 2026, titled 'DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence.' The paper details FP4 quantization-awa→28 Jul 2026DeepSeek V4 Flash, up to 32 tok/s on AMD Ryzen AI MAX+ 395Hey fellow llamas. we have something new for Strix Halo owners we thought would be useful to share. i'll keep it short: We were able to fit DeepSeek V4 Flash plus its speculative draft on a single Ryz→1 Aug 2026DeepSeek-V4-Flash-0731: Models you can run locally now have the intelligence score of the top frontier model from March 2026March 6th, 2026 the highest intelligence index score was 51 for frontier models. deepseek-ai/DeepSeek-V4-Flash-0731 that has an intelligence score of 50. If these benchmarks are accurate, models avail→2 Aug 2026Vacuum 16Thttps://huggingface.co/tsfrm/vacuum-16t A 16.5-trillion-parameter model that contains nothing. This model is just a ████ you to the labs and companies who say that 'haha I have the biggest model out t→2 Aug 2026Expert-only IQ3 requant of DeepSeek-V4-Flash-0731: better KLD than UD-IQ3_S, 1.4x decode on a CPU-spill rigHey all, tldr / who this helps: you run a mixed multi-GPU box where the experts spill to RAM, and you want to stay in the 3-bit tier instead of dropping to Q2 to make it fit. https://huggingface.co/Ta→3 Aug 2026'Data center in a Box (on Wheels)' 256Gb VRAM/512Gb RAM AI Server 6-8 Month Operational Review, Stability Write Up, BenchmarksI've been out of these forums for awhile but I figured I would provide a formal update on how this has been going now that it has some operation time under its belt, just to put the information out th→6 Aug 2026How come artificialanalysis.ai ranks Gemma4 above Qwen3.6 27b in SciCodeJust came across this coding benchmark: SciCode Artificialanalysis.ai reports a ranking which contradicts the feeling we've towards those models in real life coding. Is Gemma 4 really that good, or a →7 Aug 2026Serving Deepseek v4 Flash 0731 on 2x DGX Spark — 5-7 GB OS headroom, what would you do to lower VRAM usage and increase OS available RAM?Hey all, I'm serving DSv4Flash 0731 on a cluster of 2x DGX Sparks but am running into constant issues with having almost no RAM (unified memory) left for the OS/cache and I'd love to hear the communit
CompanyNVIDIA7 recent entries10 Apr 2026[D] 60% MatMul Performance Bug in cuBLAS on RTX 5090 [D]A bug was identified in NVIDIA's cuBLAS library where `cublasSgemmStridedBatched` dispatches the same suboptimal `cutlass_80_simt_sgemm_128x32_8x5` kernel for every batched FP32 workload from 256×...→13 Apr 2026Is an nvidia DGK Spark or similar worth it?This Reddit thread on r/ollama discusses whether the NVIDIA DGX Spark — powered by the GB10 Grace Blackwell Superchip and delivering 1 petaFLOP of performance — is a worthwhile investment for running →25 Jul 2026PSA: DO NOT use Intel consumer platforms for multi-GPU setupsSince a lot more people are trying to build their own multi-GPU machines, I thought I should help to prevent a common mistake people make with building multi-GPU machines. Which is using an Intel cons→28 Jul 2026DeepSeek V4 Flash, up to 32 tok/s on AMD Ryzen AI MAX+ 395Hey fellow llamas. we have something new for Strix Halo owners we thought would be useful to share. i'll keep it short: We were able to fit DeepSeek V4 Flash plus its speculative draft on a single Ryz→3 Aug 2026'Data center in a Box (on Wheels)' 256Gb VRAM/512Gb RAM AI Server 6-8 Month Operational Review, Stability Write Up, BenchmarksI've been out of these forums for awhile but I figured I would provide a formal update on how this has been going now that it has some operation time under its belt, just to put the information out th→6 Aug 2026nvidia/NVIDIA-Nemotron-Parse-2.0 · Hugging FaceNVIDIA Nemotron Parse 2.0 transforms document images into structured, machine-readable representations with text, layout classes, bounding boxes, and reading-order information. Given a Red, Green, Blu→7 Aug 2026Serving Deepseek v4 Flash 0731 on 2x DGX Spark — 5-7 GB OS headroom, what would you do to lower VRAM usage and increase OS available RAM?Hey all, I'm serving DSv4Flash 0731 on a cluster of 2x DGX Sparks but am running into constant issues with having almost no RAM (unified memory) left for the OS/cache and I'd love to hear the communit