CompanyAnthropic8 recent entries27 Jul 2026Give any Ollama-compatible client session memory + a shared knowledge wiki by swapping the chat URLHey folks — I built ContextMemory, an open-source agentic context gateway for apps that already talk to LLMs. The idea is simple: keep your existing POST /api/chat client (Ollama wire format), point i→29 Jul 2026ChatGPT Made Me Cry TonightSorry if flair is wrong. I decided to finally get a ChatGPT subscription after some conversations with it about health issues with my dog. I've only used AI for coding work, primarily Claude, but I fe
CompanyOpenAI8 recent entries31 Jul 2026DeepSeek v4 Flash for DS4 (DwarfStar) GGUF w/ DSpark MTP HeadI'm an avid user of Deepseek v4 Flash via antirez's DS4 DwarfStar inference engine, and so when the new checkpoint dropped, the first thing I did was rent a cloud box and spin up a quantization for us→4 Aug 2026Why are Gamers so incredibly hostile to AI? Is it just a tiny vocal minority that spreads such toxic vitriol online?It's more accurate to say that many highly engaged online gamers are hostile to AI, not that 'gamers' as a whole are. Gaming is a huge community with hundreds of millions of people, and opinions vary →5 Aug 2026Stable Diffusion might actually be remembered in the history books, and I don’t think that’s an overstatementHear me out before you roll your eyes. We tend to only recognize turning points in hindsight. Nobody in 1993 thought the Mosaic browser would be a history book moment, but the web is. I think Stable D→5 Aug 2026Qwen Developers' responses from their recent Twitter/X AMAQuestions & Responses(in BOLD) below. Favorite question(s) moved to end of the thread with combined responses(removed duplicates). Be optimistic folks. I'm sure we're getting other models too apart fr→6 Aug 2026I ported vLLM's serving stack to C++20: 66 MiB binary, no Python at inference, output checked token-for-token against vLLMI'm the author, so discount the enthusiasm accordingly. This is an unaffiliated community port, not endorsed by the vLLM project, which it uses to verify its correctness. What started it: I love vLLM,→11 Aug 2026Writing formats I can no longer read because AI has beaten them to deathI don’t even care anymore whether these posts are *actually* AI-generated. The problem is that there’s now a very specific style of internet writing that instantly makes my brain refuse to continue re→11 Aug 2026Using the different models in different industries - your experience?Hi all, I'm seeing so many conversations from people discussing how they're 'using the models wrong' and 'don't use Sol max as high is enough for you', yet all the conversations lack the nuance of wha→11 Aug 2026OpenAI just launched a cybersecurity model that answers 95% of advanced threat queries. And Meta put a frontier model on your laptop. Same day.Something happened today that I think most people are going to miss because there are two separate stories and neither one is getting the full picture. OpenAI expanded Daybreak. If you haven't heard o
CompanyGoogle8 recent entries25 Jul 2026Im back from gemini. GPT is astronomically better again.like 8 months ago I was tinkering with both and gemini was so much better i went with that. I had a project at work that I needed an AI to sift through a manual and schematic for and no matter what, g→26 Jul 202623 Gemma4-E4B models compared with abliterlitics: the most downloaded one is also the most brokenThis is our biggest comparison yet. We've taken 23 Gemma 4 E4B models from huggingface and ran them through the abliterlitics gauntlet. We also have a new abliterlitics discord, feel free to jump on a→31 Jul 2026What's your local AI coding setup on a MacBook Pro M4?I've spent the last couple of days trying different setups (Ollama, Continue, Claude Code, Gemini CLI, OpenRouter...) and at this point I feel like I've spent more time configuring tools than actually→2 Aug 2026Vacuum 16Thttps://huggingface.co/tsfrm/vacuum-16t A 16.5-trillion-parameter model that contains nothing. This model is just a ████ you to the labs and companies who say that 'haha I have the biggest model out t→3 Aug 2026Was the release of deepseek v4 flash planned to take spotlight against 5.6 luna?Id figured since they first emailed people about api price changes coming mid july then delayed the v4 flash release to late july, I wonder if they delayed it for the sake of stealing spotlight from o→5 Aug 2026Stable Diffusion might actually be remembered in the history books, and I don’t think that’s an overstatementHear me out before you roll your eyes. We tend to only recognize turning points in hindsight. Nobody in 1993 thought the Mosaic browser would be a history book moment, but the web is. I think Stable D→9 Aug 2026CyberKimi just dropped strong results on one of ExploitBench’s hardest V8 bugs , points away from MythosHey everyone ! Quick share from the cyber + local LLM side of things that I found interesting. During this week’s hacker summer camp, an AI researcher and reverse malware engineer veteran 'lordx64' on→10 Aug 2026Need real world ML problems to evaluate my educational ML toolsI'm a retired platform engineer, coding mainly in Rust, and involved with a ML study group. I developed a ML programming language (alternative to Python, Colab) to help me learn (and teach) ML concept
CompanyMeta8 recent entries4 Aug 2026Why are Gamers so incredibly hostile to AI? Is it just a tiny vocal minority that spreads such toxic vitriol online?It's more accurate to say that many highly engaged online gamers are hostile to AI, not that 'gamers' as a whole are. Gaming is a huge community with hundreds of millions of people, and opinions vary →5 Aug 2026Qwen3-TTS voice cloning is now in mainline llama.cpp — the old demo finally became real supportPeople may remember the Qwen3-TTS llama.cpp demo from a few months ago. That PR said it probably wouldn’t be merged because llama.cpp was missing some of the graph and API pieces it needed. A new impl→6 Aug 2026Dual 3090 setup: 400 pp t/s to 1600 pp t/s on Qwen 3.6 27B... with slightly lower tps.First of all, my setup: Ryzen 9 5950x DDR4 3200Mhz 64gb (2x32) Dual 3090s, no NVLINK Runtime: llama.cpp Nvidia Drivers 610 Windows 11 25H2 Qwen 3.6 27B Q8 I've been using llama-server with --split-mod→7 Aug 2026Qwen 3.6 27B flags/settings in llama.cppI run the following on a 5090 and have been okay with its performance, it does most things somewhere 80-100 t/s, though that can slow down at full 262k context - more like 40 t/s at times. I use it pr→7 Aug 2026Echo Dot 2 can run 28M LLM at decent speedCode and instructions available here: https://github.com/albertoZurini/echo-dot-2-playground Hello there! After a few days of experimenting I was able to get a completely local voice pipeline running →8 Aug 2026Showoff Saturday: Local 4x 6000 Pro (multi-year progression)Not the biggest or shiniest, but it's mine From gaming machine inference on the original llama models, to a 4x RTX 6000 Pro Max Q + 4x 3090s local AI cluster. Pictures are in reverse chronological ord→11 Aug 2026OpenAI just launched a cybersecurity model that answers 95% of advanced threat queries. And Meta put a frontier model on your laptop. Same day.Something happened today that I think most people are going to miss because there are two separate stories and neither one is getting the full picture. OpenAI expanded Daybreak. If you haven't heard o→12 Aug 2026What unique, custom QOL upgrades have you given your local agents?Warning: Kinda long post. If you don't like reading, please skip for your own sanity. Also, I've got nothing to sell, just a tinkerer, so I just want to share ideas and learn from you guys too. When I
CompanyMistral4 recent entries23 Jul 2026Trained a 32B FLUX.2 LoRA on a 24GB AMD 7900 XTX, native ROCm on Windows — full guide + patchesTL;DR: Everyone says QLoRA past ~13B is dead on a 24GB card. I got the full 32B FLUX.2 dev transformer QLoRA-training resident on the GPU on a 7900 XTX under native ROCm on Windows (no ZLUDA, no CUDA →23 Jul 2026DeepSeek Founder’s 4-hour investor meeting: DeepSeek is prioritizing AGI over user growth and commercialisationA Chinese article compiled 52 remarks from Liang Wenfeng’s four-hour investor meeting. I’ve summarised the most important ones below. DeepSeek has one central objective: AGI. This is not the time to m→23 Jul 2026Arcee AI has spoken out against the ban on open Chinese models in USThis is rather counterintuitive, since banning Chinese models would benefit them the most. Jensen Huang is also against the ban, although the interests here are more obvious. Do you think that if Arce→29 Jul 2026'Uncensored' LLMs are measurably more optimistic than their base modelsHi. Many people think uncensored models are basically the same model that just doesn't refuse, but... I was recently checking whether uncensored models would give me better answers for stock market pr
CompanyxAI5 recent entries10 Apr 2026HappyHorse is from Alibaba ATH, not Grok / Veo 3.2 / Wan 2.7 / Seedance 2HappyHorse-1.0 is an AI video generation model that anonymously appeared on the Artificial Analysis Video Arena leaderboard in early April 2026, where both its V1 and V2 versions rapidly climbed to...→10 Apr 2026Got early access to a real-time interactive video model, here's what I foundI was unable to retrieve the specific Reddit post at the provided URL through my search. The post (reddit.com/r/StableDiffusion/comments/1shxmfk) did not surface in the search results, and I cannot...→10 Apr 2026Bad news on Happy Horse from twitterHappyHorse-1.0 is a pseudonymous AI video generation model that appeared on April 7, 2026, topping the Artificial Analysis Video Arena leaderboard in both text-to-video and image-to-video (no audio...→25 Jul 2026Old Coder Needs help with New AI Development and wants to get up to speed to understand it all.Hi Guys, I'm an old coder and DBA that has been in the field for almost 40 years. More and more the jobs I was doing for work are being taken over by AI and the need for my type of work is diminishing→2 Aug 2026Vacuum 16Thttps://huggingface.co/tsfrm/vacuum-16t A 16.5-trillion-parameter model that contains nothing. This model is just a ████ you to the labs and companies who say that 'haha I have the biggest model out t
CompanyDeepSeek8 recent entries7 Aug 2026Qwen 3.6 27B flags/settings in llama.cppI run the following on a 5090 and have been okay with its performance, it does most things somewhere 80-100 t/s, though that can slow down at full 262k context - more like 40 t/s at times. I use it pr→8 Aug 2026Is anyone else finding DeepSeek-V4-Flash unreliable for non-coding tasks?(I am not a native speaker, written by myself, so please bear with me) I really want to like DeepSeek-V4-Flash-0731. But it has serious flaws that don't align with the high score on intelligence bench→9 Aug 2026Underestimated budget solution: radeon 780m iGPUThere are so many posts where people complaining about high prices and asking for solution <= 1000 EUR. So, there is one solution to consider: PC/mini PC/laptop on Ryzen 7 260/Ryzen 9 8945HX/etc CPU w→10 Aug 2026Need real world ML problems to evaluate my educational ML toolsI'm a retired platform engineer, coding mainly in Rust, and involved with a ML study group. I developed a ML programming language (alternative to Python, Colab) to help me learn (and teach) ML concept→10 Aug 2026DeepSeek V4 Flash 0731 is the ‘killer app’ that is going to sell A LOT of DGX SparksHaving a ‘Killer Application’ that everyone wants to use helps sell hardware, plain and simple. DeepSeek V4 Flash 0731 isn’t an app of course, but I think it’s going to be the major catalyst for getti→11 Aug 2026We quantized DeepSeek V4 0731 and benchmarked it against popular quants on 8× RTX 5090We converted the model from the original safetensors and found two issues. The first one made our quantization fail several times, the second one does not fail at all, it just quietly ruins the base 1→11 Aug 2026I gave DeepSeek V4 Flash basic vision by training a 40M connector on 100K examplesI wanted to find out whether a huge text-only MoE could be given basic vision without retraining the language model itself. The short answer is yes. I froze DeepSeek V4 Flash and a 417M-parameter Moon→12 Aug 2026DeepSeek V4 Flash 0731 uncensored (jailbreak pt2)Since lot's of people were sceptical or whatever, heres how to uncensor / jailbreak V4 flash and proof. No it is not lead on whatever, first prompt, first try, every time. Put this in System message:
CompanyNVIDIA8 recent entries6 Aug 2026Get AI max+ 395 laptop or wait for rtx spark?So I can either pull the trigger on a 128gb AI max+ 395 laptop or wait for RTX Spark for LLMs. Maybe I get it now and the price of the spark is super high so it's a good purchase or maybe the Spark sh→6 Aug 2026Dual 3090 setup: 400 pp t/s to 1600 pp t/s on Qwen 3.6 27B... with slightly lower tps.First of all, my setup: Ryzen 9 5950x DDR4 3200Mhz 64gb (2x32) Dual 3090s, no NVLINK Runtime: llama.cpp Nvidia Drivers 610 Windows 11 25H2 Qwen 3.6 27B Q8 I've been using llama-server with --split-mod→7 Aug 2026Echo Dot 2 can run 28M LLM at decent speedCode and instructions available here: https://github.com/albertoZurini/echo-dot-2-playground Hello there! After a few days of experimenting I was able to get a completely local voice pipeline running →8 Aug 2026Has anyone here fiddled with TPUs for inference ?I discovered recently that Google uses their own TPUs, like tiny ASIC cards like the toy ones that existed for bitcoin. And while it sounds inefficient the fact they use thousands of them because...th→9 Aug 2026I Turned My Underused Gaming Laptop Into a Local AI WorkstationTL;DR: I am building a Windows-first local AI setup for people who want to try local LLMs without spending days choosing models, setting up Ollama, Docker, WSL, Open WebUI, agents, and tool permission→10 Aug 2026Need real world ML problems to evaluate my educational ML toolsI'm a retired platform engineer, coding mainly in Rust, and involved with a ML study group. I developed a ML programming language (alternative to Python, Colab) to help me learn (and teach) ML concept→10 Aug 2026DeepSeek V4 Flash 0731 is the ‘killer app’ that is going to sell A LOT of DGX SparksHaving a ‘Killer Application’ that everyone wants to use helps sell hardware, plain and simple. DeepSeek V4 Flash 0731 isn’t an app of course, but I think it’s going to be the major catalyst for getti→11 Aug 2026Rumored 50-series Super refresh bumps everything +50% VRAMleaked Super specs have the 5070 Ti and 5080 going 16GB to 24GB and the 5070 to 18GB, thanks to the new 3GB GDDR7 modules. 24GB on a Ti-class card is actually the number people here have been waiting