CompanyAnthropic8 recent entries31 Jul 2026What's your local AI coding setup on a MacBook Pro M4?I've spent the last couple of days trying different setups (Ollama, Continue, Claude Code, Gemini CLI, OpenRouter...) and at this point I feel like I've spent more time configuring tools than actually→31 Jul 2026Experience sharing: How do you use your local models and for what kind of tasks?Here is my experience, which I would like to share with you and I also would like to hear your thoughts and valuable tips&tricks. Hardware: Mac Mini M4 (32GB Unified Memory) Model Server: Ollama Orche
CompanyOpenAI8 recent entries30 Jul 2026LG AI Research releases K-EXAONE 2.0 750B A37BIt was developed under Phase 2 of Korea's Sovereign AI Foundation Model Project. Size: 750B parameters (3x larger than their 236B v1 model). - License: Apache 2.0 Languages: Expanded to 10 language→30 Jul 20262 images + 1 prompt > expected outputHi, I'm trying to replicate a thing locally, that I can do on ChatGPT. What I want is to give a local AI two reference images (a face and a background item) and a prompt about the composition of the p→4 Aug 2026Why are Gamers so incredibly hostile to AI? Is it just a tiny vocal minority that spreads such toxic vitriol online?It's more accurate to say that many highly engaged online gamers are hostile to AI, not that 'gamers' as a whole are. Gaming is a huge community with hundreds of millions of people, and opinions vary →4 Aug 2026GPT-OSS has turned one year old today!It is one of the best local models ever released, in both 20B and 120B versions. I always come back to it, especially the 120B version. Its only competition is, in my opinion, Qwen 3.5 122B, but that →5 Aug 2026Qwen Developers' responses from their recent Twitter/X AMAQuestions & Responses(in BOLD) below. Favorite question(s) moved to end of the thread with combined responses(removed duplicates). Be optimistic folks. I'm sure we're getting other models too apart fr→8 Aug 2026Claude Code in 9 lines pythonI was wondering what a minimal coding agent implementation would look like that can be used like Claude Code or Codex Not feature-by-feature of course but basically stripping everything out that is no→11 Aug 2026Introducing Unsloth Desktop appHi LocalLlama, we're super excited to release Unsloth Desktop today! 🦥 It's the first desktop app that enables you to run and train models locally. Open-source. Available on Mac, Windows, and Linux Su→11 Aug 2026Can Gemma and Qwen models catch hallucinations by looking at their own logprobs?Hi! I'm really obsessed with LLM hallucinations for the last 6 days 😭 I started by designing system prompts to attack hallucinations but failed, obviously. Now I tried reading logprobs and... I think
CompanyGoogle8 recent entries14 Apr 2026AI tool to analyze a video and generate a prompt?This Reddit thread from r/StableDiffusion discusses the community's search for AI tools capable of analyzing an existing video and automatically generating a descriptive text prompt from it — essentia→4 Jun 2026Best Visual Reasoning Model in 2026 (Including APIs) [D]Gemini 3.1 Pro and Gemini 3-Pro lead visual reasoning benchmarks , with GPT-5.2, Kimi-K2.5, and GPT-5.2-Pro following . A 2026 evaluation benchmarked 15 leading multimodal models on visual reasoning a→10 Jun 2026Ideogram4 vs Flux.2 Dev vs GPT Image 2 vs Nano Banana ProA practical comparison of Ideogram 4.0, Nano Banana Pro, and GPT Image 2 across text rendering, design, photorealism, and real creator workflows. Ideogram 4.0 is the first open-weight challenger to ra→22 Jul 2026NuExtract3 is now available on Ollama: 4B VLM for document-to-Markdown and structured JSON extractionDisclosure: I work at NuMind, the team that trained NuExtract3. NuExtract3 is an Apache-2.0, open-weight 4B VLM based on Qwen3.5-4B. It is specialized for document understanding rather than general ch→24 Jul 2026I built an open-source multi-agent SDLC harness that beats a cold Claude Code run on large repos, by learning the repo once. Real benchmarks (incl. where it loses) inside. [P]Built an open-source AI coding agent that was 7%–75% cheaper than a cold 'claude -p' run on 6/6 well-localized tasks across repositories up to ~82k LOC. The biggest difference: Cold agent: 6.83, 207 t→26 Jul 202623 Gemma4-E4B models compared with abliterlitics: the most downloaded one is also the most brokenThis is our biggest comparison yet. We've taken 23 Gemma 4 E4B models from huggingface and ran them through the abliterlitics gauntlet. We also have a new abliterlitics discord, feel free to jump on a→31 Jul 2026What's your local AI coding setup on a MacBook Pro M4?I've spent the last couple of days trying different setups (Ollama, Continue, Claude Code, Gemini CLI, OpenRouter...) and at this point I feel like I've spent more time configuring tools than actually→10 Aug 2026Need real world ML problems to evaluate my educational ML toolsI'm a retired platform engineer, coding mainly in Rust, and involved with a ML study group. I developed a ML programming language (alternative to Python, Colab) to help me learn (and teach) ML concept
CompanyMeta8 recent entries31 Jul 2026Uncensored Multi-Model Releases, LongCat-Flash-Lite with MTPs, Jamba2-Mini, Qwen3.5-9B-Nikusui-v1 with MTPs and Qwen3.5-27B-Nikusui-v1 with MTPs, Available in Safetensors and GGUF Formats!Been working hard for the past month to bring to the community some interesting curios, so for starters we have LongCat-Flash-Lite Uncensored Heretic with MTPs which has never before been uncensored, →1 Aug 2026A collection of small domain-specific benchmarks for local models (30+ and growing)Hello fellow local AI people! I took 'you must create your own benchmarks' literally, and built a website for this. How does the end result look like Let's say I want to know which model has most comm→4 Aug 2026Why are Gamers so incredibly hostile to AI? Is it just a tiny vocal minority that spreads such toxic vitriol online?It's more accurate to say that many highly engaged online gamers are hostile to AI, not that 'gamers' as a whole are. Gaming is a huge community with hundreds of millions of people, and opinions vary →7 Aug 2026Echo Dot 2 can run 28M LLM at decent speedCode and instructions available here: https://github.com/albertoZurini/echo-dot-2-playground Hello there! After a few days of experimenting I was able to get a completely local voice pipeline running →8 Aug 2026Showoff Saturday: Local 4x 6000 Pro (multi-year progression)Not the biggest or shiniest, but it's mine From gaming machine inference on the original llama models, to a 4x RTX 6000 Pro Max Q + 4x 3090s local AI cluster. Pictures are in reverse chronological ord→10 Aug 2026Needle 2: 14MB agentic LLM for phones, wearables, smart home and robots.Hey LocalLlaMa, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots a→10 Aug 2026Achievable 253 t/s - unsloth/Muse Glimmer 30B UD-Q5_K_M on a 5090Benchmarked Muse Glimmer 30B on my RTX 5090 (32GB), 262k context, UD-Q5_K_M + dflash-kquant + mmproj. Workload Stock master + DFlash ngram-simple PR #26842 + DFlash Code patch 78 t/s 57 t/s 220-253 t/→12 Aug 2026What unique, custom QOL upgrades have you given your local agents?Warning: Kinda long post. If you don't like reading, please skip for your own sanity. Also, I've got nothing to sell, just a tinkerer, so I just want to share ideas and learn from you guys too. When I
CompanyMistral6 recent entries13 Apr 2026Ollama / Mistral with MCP to MempalaceThis Reddit post from r/ollama discusses integrating Ollama-served Mistral with MemPalace — a free, locally-run AI memory system — via the Model Context Protocol (MCP). MemPalace runs entirely on a us→13 Apr 2026I’m looking for advice on setting up a local AI model that can generate Word reports automatically.This r/ollama thread discusses community advice on configuring a locally-run AI model (via Ollama) to automatically generate Word documents or reports, covering topics such as model selection, scripti→14 Apr 2026Agents in Ollama and LangflowThis Reddit post from r/ollama likely discusses how to build and run AI agents locally by combining Ollama — which handles local model serving to keep data private — with Langflow's visual, drag-and-d→4 May 2026OneTrainer now supports Ernie LoRAOneTrainer is a one-stop solution for all diffusion training needs. The tool now supports the Ernie Image model, which can be trained using LoRA (Low-Rank Adaptation) methods. This adds support for tr→4 Aug 2026GPT-OSS has turned one year old today!It is one of the best local models ever released, in both 20B and 120B versions. I always come back to it, especially the 120B version. Its only competition is, in my opinion, Qwen 3.5 122B, but that →5 Aug 2026Prime Agent - a new coding harness surpassing Codex/CC/PIPrime Agent is an open-source coding and research agent for general and long-running work. A self-improving RLM harness for coding and long-running autonomous tasks. Designed to be both token-efficien
CompanyxAI7 recent entries10 Apr 2026How many of you have actually stopped using GPT and switched to something else?A Reddit thread in r/ChatGPT asked users whether they had actually stopped using ChatGPT and switched to competing AI tools, reflecting a broader wave of user dissatisfaction driven by perceived de...→15 Apr 2026need a grok/neno-banana like img2img generation colab cellA r/StableDiffusion thread where a user seeks a Google Colab notebook cell that replicates the img2img generation style or workflow of tools like 'grok' or 'neno-banana' — likely referring to specific→5 May 2026Parllama -- a terminal UI for Ollama model management and multi-provider LLM chatParllama is a TUI (Text UI) application designed for easy management and use of Ollama-based LLMs that also works with major cloud-provided LLMs. It provides core model management features including f→24 Jul 2026I built an open-source multi-agent SDLC harness that beats a cold Claude Code run on large repos, by learning the repo once. Real benchmarks (incl. where it loses) inside. [P]Built an open-source AI coding agent that was 7%–75% cheaper than a cold 'claude -p' run on 6/6 well-localized tasks across repositories up to ~82k LOC. The biggest difference: Cold agent: 6.83, 207 t→25 Jul 2026Old Coder Needs help with New AI Development and wants to get up to speed to understand it all.Hi Guys, I'm an old coder and DBA that has been in the field for almost 40 years. More and more the jobs I was doing for work are being taken over by AI and the need for my type of work is diminishing→30 Jul 2026What is the fastest local research tool (deep research) ?I've tried grok and Claude's deep research mode and I was amazed with the speed considering the amount of sources analysed. Is there anything as fast that can run locally? My guess would be that to ru→11 Aug 20261 Day in and I feel okay saying Muse-Glimmer-30B finally beats 3.6-27B for the size in some use-casesA few things right off the bat: it reasons very efficiently. Like Grok 4.5 levels of efficient thinking it quantizes very well. My first few tests with iq3_xxs were better than Qwen/Gemma behaved at t
CompanyDeepSeek8 recent entries5 Aug 2026Qwen Developers' responses from their recent Twitter/X AMAQuestions & Responses(in BOLD) below. Favorite question(s) moved to end of the thread with combined responses(removed duplicates). Be optimistic folks. I'm sure we're getting other models too apart fr→7 Aug 2026Anyone running DeepSeek-V4-Flash-0731 on MI325X with vLLM? Mine is behaving completely brokenIs anyone here successfully running DeepSeek-V4-Flash-0731 locally with vLLM, especially on AMD MI325X? My setup: GPU: 1x AMD Instinct MI325X Model: deepseek-ai/DeepSeek-V4-Flash-0731 vLLM: 0.26.0 ROC→9 Aug 2026Doom Loop: Anyone Else Having DeepSeek v4 Flash 0731 Issues on ollama cloud?Am I the only one having issues with DeepSeek V4 Flash? It gets stuck in a loop, as if it can't call the tools, and keeps repeating the same things endlessly without moving forward. Is it a poorly wri→9 Aug 2026DeepSeek-V4-Flash-0731 Q8_K_XL sometimes stops mid-task in OpenCode - anyone else seeing this?Hey everyone, I've been experimenting with the new DeepSeek-V4-Flash-0731 release locally using the Unsloth Studio Q8_K_XL GGUF with OpenCode. Overall, it's been working really well, but I've noticed →10 Aug 2026Need real world ML problems to evaluate my educational ML toolsI'm a retired platform engineer, coding mainly in Rust, and involved with a ML study group. I developed a ML programming language (alternative to Python, Colab) to help me learn (and teach) ML concept→10 Aug 2026Best open-source harness like Claude Code?Avid claude code user here looking to do equivalent things with local models. Just want to plug in something like Qwen and have the interface be 1:1 with claude code. Any suggestion? submitted by /u/N→11 Aug 2026DeepSeek-V4-Flash-0731 (284B MoE) at 75 tok/s on 2× DGX Spark — full recipe, 11 gotchas, reboot-proof cluster, Codex CLI integrationSpent two nights getting deepseek-ai/DeepSeek-V4-Flash-0731 (284B MoE, 13B active, native FP4/FP8, 1M context) running production-grade on two DGX Sparks connected by one QSFP DAC cable. Everything — →12 Aug 2026Idea for a deepseek-v4-flash-0731 backed automated research workflow to be leveraged via qwen3.6/3.8 27b for difficult tasks that require highly technical, not easy to find information.Sometimes you have tasks that are outside of your expertise and the idea is this workflow automation could be leveraged to manage to have local AI figure it out using research from his workflow gather
CompanyNVIDIA8 recent entries6 Aug 2026nvidias nemotron omni only loads its text half on a mac, so i wrote the vision and audio towers in mlxnvidias nemotron omni is open weights and it sees, hears and reasons. theres already a 4bit mlx quant on hugging face but only the text backbone loads with standard mlx tooling. the model card says it→7 Aug 2026Echo Dot 2 can run 28M LLM at decent speedCode and instructions available here: https://github.com/albertoZurini/echo-dot-2-playground Hello there! After a few days of experimenting I was able to get a completely local voice pipeline running →9 Aug 2026I Turned My Underused Gaming Laptop Into a Local AI WorkstationTL;DR: I am building a Windows-first local AI setup for people who want to try local LLMs without spending days choosing models, setting up Ollama, Docker, WSL, Open WebUI, agents, and tool permission→10 Aug 2026Need real world ML problems to evaluate my educational ML toolsI'm a retired platform engineer, coding mainly in Rust, and involved with a ML study group. I developed a ML programming language (alternative to Python, Colab) to help me learn (and teach) ML concept→10 Aug 2026Introducing Muse Glimmer: an open-weight model optimized for always-on local agent workflowsHi r/LocalLLaMA 👋 Today we’re excited to release Muse Glimmer, a 30B open-weight model built specifically for local agent workflows. We’re releasing the weights to the community under a permissive Apa→10 Aug 2026I compared GGUF quants of Qwen3.6 27B to NVFP4, AWQ, AutoRound, and FP8There's an interactive chart and some extra data in the blog post if you're interested. There are plenty of KL-divergence benchmarks for GGUF models, but most of them compare one GGUF quant against an→11 Aug 2026Introducing Unsloth Desktop appHi LocalLlama, we're super excited to release Unsloth Desktop today! 🦥 It's the first desktop app that enables you to run and train models locally. Open-source. Available on Mac, Windows, and Linux Su→12 Aug 2026Tested Nemotron 3.5 Lightning locally on coding, Hermes Agent and agentic workRan the model with quants (Q5) and MTP by bartowski with llama.cpp server. It takes ~24GB ram running on M5 Pro with 48GB at about 65t/s. On some tasks it was quite the overthinker. Overall, the quali