CompanyAnthropic6 recent entries23 Jul 2026Model 'distillation' accusations are getting way overblown at this pointThe news about Anthropic settling a class action lawsuit for 1.5B over training data isn't just a legal headache for them, it's a massive warning sign for engineering teams relying entirely on closed →23 Jul 2026GPT-5.5 Scores 10.6% on ActiveVision, Humans Hit 96.1% [R]The interesting finding from a new [arXiv paper](https://arxiv.org/abs/2607.16165) isn't that a frontier vision model failed a new benchmark, that happens weekly, but the specific shape of the failure
CompanyOpenAI6 recent entries13 Apr 2026Sam Altman's Molotov attack suspect listed the names of other AI CEOs and investors in a 'last warning' note, the feds saidA Texas man, 20-year-old Daniel Moreno-Gama, was charged with attempted murder after allegedly throwing a Molotov cocktail at OpenAI CEO Sam Altman's San Francisco home on April 10, 2026, with surveil→24 May 2026Scientists invented a fake disease. AI told people it was realResearchers from the University of Gothenburg invented a fake disease called 'bixonimania,' a fictional skin condition supposedly caused by screen time. Multiple AI chatbots including Google's Gemini,→21 Jul 2026Reproducing OpenAI’s “persistently beneficial models” - GRPO trait install barely moves. Ideas? [P] [R]TL;DR: I’m reproducing the trait-persistence result from arXiv:2606.24014 on one RTX 3090. Before I can test persistence I need to install a trait via RL — and my GRPO run moves the trait only +2.4 po→23 Jul 2026I trained a 0.5M model on 1B tokens of Fineweb-edu dataset.Hi everyone, About a month ago I publish my very first research paper on my neural network architecture called Silia. You can look at the model here: https://huggingface.co/Srijan-Srivastava/Silia-v2 →23 Jul 2026DeepSeek Founder’s 4-hour investor meeting: DeepSeek is prioritizing AGI over user growth and commercialisationA Chinese article compiled 52 remarks from Liang Wenfeng’s four-hour investor meeting. I’ve summarised the most important ones below. DeepSeek has one central objective: AGI. This is not the time to m→11 Aug 2026I built a weird, low-power llama.cpp server using an Intel N100 + RTX 5060TiEverything started with the sudden death of my old ASRock J1900. While looking for the perfect ITX replacement, I stumbled upon the Chinese CW-NAS-ADLN-K motherboard, which looked perfect on paper: In
CompanyGoogle3 recent entries15 Apr 2026How much harder is it these days to get into a PhD program without having a high ranking degree for UG? [D]This Reddit discussion thread from r/MachineLearning explores the growing challenges faced by applicants from non-elite undergraduate institutions when applying to PhD programs in machine learning and→24 May 2026Scientists invented a fake disease. AI told people it was realResearchers from the University of Gothenburg invented a fake disease called 'bixonimania,' a fictional skin condition supposedly caused by screen time. Multiple AI chatbots including Google's Gemini,→2 Aug 2026Vacuum 16Thttps://huggingface.co/tsfrm/vacuum-16t A 16.5-trillion-parameter model that contains nothing. This model is just a ████ you to the labs and companies who say that 'haha I have the biggest model out t
CompanyMeta6 recent entries15 Apr 2026[P] Added 8 Indian languages to Chatterbox TTS via LoRA — 1.4% of parameters, no phoneme engineering [P]A community researcher shared on r/MachineLearning how they extended Chatterbox TTS — Resemble AI's open-source, 500M-parameter model — to support 8 Indian languages using LoRA (Low-Rank Adaptation), →23 Jul 2026Model 'distillation' accusations are getting way overblown at this pointThe news about Anthropic settling a class action lawsuit for 1.5B over training data isn't just a legal headache for them, it's a massive warning sign for engineering teams relying entirely on closed →9 Aug 2026KLQ: Training-free measured rotation quantization. Beats all training-free rotation-based quantization methods on W4A4KV4-bits. Llama 3.2 1B KLQ-quantized beats SpinQuant and gets close to ReSpinQuant without GPTQ/LDLQ rounding.First of all, I'm not a lab, this was a solo summer research project that finally culminated into the github repo and the writeup. The repo includes a much deeper dive with methods, findings about qua→10 Aug 2026Needle 2: 14MB agentic LLM for phones, wearables, smart home and robots.Hey LocalLlaMa, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots a→11 Aug 2026I built a weird, low-power llama.cpp server using an Intel N100 + RTX 5060TiEverything started with the sudden death of my old ASRock J1900. While looking for the perfect ITX replacement, I stumbled upon the Chinese CW-NAS-ADLN-K motherboard, which looked perfect on paper: In→12 Aug 2026New Muse-Glimmer-30B SoTA Quants - hopefully a new lineup :)Hey Folks, I've been making quants for a while - recently I took a short break to get into hardcore research (submitted my first EMNLP paper during it!). Along the way, I built up a little arsenal of
CompanyMistral2 recent entries23 Jul 2026Trained a 32B FLUX.2 LoRA on a 24GB AMD 7900 XTX, native ROCm on Windows — full guide + patchesTL;DR: Everyone says QLoRA past ~13B is dead on a 24GB card. I got the full 32B FLUX.2 dev transformer QLoRA-training resident on the GPU on a 7900 XTX under native ROCm on Windows (no ZLUDA, no CUDA →23 Jul 2026DeepSeek Founder’s 4-hour investor meeting: DeepSeek is prioritizing AGI over user growth and commercialisationA Chinese article compiled 52 remarks from Liang Wenfeng’s four-hour investor meeting. I’ve summarised the most important ones below. DeepSeek has one central objective: AGI. This is not the time to m
CompanyxAI1 recent entries2 Aug 2026Vacuum 16Thttps://huggingface.co/tsfrm/vacuum-16t A 16.5-trillion-parameter model that contains nothing. This model is just a ████ you to the labs and companies who say that 'haha I have the biggest model out t
CompanyDeepSeek8 recent entries23 Jul 2026DeepSeek Founder’s 4-hour investor meeting: DeepSeek is prioritizing AGI over user growth and commercialisationA Chinese article compiled 52 remarks from Liang Wenfeng’s four-hour investor meeting. I’ve summarised the most important ones below. DeepSeek has one central objective: AGI. This is not the time to m→30 Jul 2026Nanbeige4.2-3B: I'm not impressedI've tested Nanbeige-4.2-3B. On paper, the benchmarks promise it blows away Qwen3.5-9B and Gemma4-12B. My goal was to have something very light and fast to replace Qwen3.6-35B (or finetunes thereof) f→2 Aug 2026Vacuum 16Thttps://huggingface.co/tsfrm/vacuum-16t A 16.5-trillion-parameter model that contains nothing. This model is just a ████ you to the labs and companies who say that 'haha I have the biggest model out t→2 Aug 2026[Release] WinterMix — Qwen3.5-122B-A10B in native MLX: an 82 GiB build that beats 94–95 GiB quants, plus a 68 GiB build for agent swarmsTL;DR: I spent 9 days developing a new quantization method for MLX models and measured 18 variants against each other on a single M5 Max MacBook Pro (128 GB). The result is the best-measuring MLX quan→3 Aug 2026The Chinese labs everyone lumps together are making four pretty different bets. I work at one of them.Every time a model drops from a Chinese lab the thread fills with people who already know who made it, and the guess is usually Alibaba. There was a thread here recently asking what separates the open→6 Aug 2026How many people in this sub try to train their own AI from scratch on their systems just for fun and to test out techniques from research papers?As for me, I own a system with an RTX 5090, Ryzen 9 9950X3D2, and 64 GB of DDR5. Every time I see research come out with a new way to train AI, I immediately think to try it on my system to see the re→9 Aug 2026endless-frontier/BigBang-v1 - qwen 3.5 finetunestable bench https://huggingface.co/bartowski/endless-frontier_BigBang-v1-GGUF I'm downloading this model only because Bartowski converted it to .gguf, so it might be interesting. Doubts : The headline→10 Aug 20261M context with 17 GB model in 24 GB VRAM: 'for the first time I was able to load a context of almost 1M tokens and extract 7 needles from various parts of the text'https://preview.redd.it/xxjh11f38jih1.png?width=1852&format=png&auto=webp&s=76850ed51e29a8bc86c2ca718d4320075eed4363 Just wanted to share a user report that I found to be very interesting. Some person
CompanyNVIDIA8 recent entries6 Jun 2026Built a fully-local paper-RAG across 2× 1080 Ti + a 3090. Three Ollama gotchas that each cost me a day.A developer documented their experience building a fully-local paper Retrieval-Augmented Generation (RAG) system using two NVIDIA GTX 1080 Ti GPUs and one RTX 3090, sharing three significant challenge→23 Jul 2026Trained a 32B FLUX.2 LoRA on a 24GB AMD 7900 XTX, native ROCm on Windows — full guide + patchesTL;DR: Everyone says QLoRA past ~13B is dead on a 24GB card. I got the full 32B FLUX.2 dev transformer QLoRA-training resident on the GPU on a 7900 XTX under native ROCm on Windows (no ZLUDA, no CUDA →24 Jul 2026Nvidia releases Qwen-Image-Flash'The NVIDIA Qwen-Image-Flash model generates images from text prompts using a four-step, DMD2-distilled version of Qwen/Qwen-Image. The distillation used DMD2 from NVIDIA FastGen, NVIDIA Model Optimiz→27 Jul 2026Built & Trained a Transformer from Scratch in Pure PyTorch for English-to-Tamil Machine Translation [Math + Code Breakdown] [P]Hi everyone! 👋 I built and trained the complete Transformer architecture from scratch using pure PyTorch (`torch.nn` primitives) based on the original 'Attention Is All You Need' paper. I trained the →1 Aug 2026Github repo to learn the OPD/OPSD and how they perform compared to GRPO, on a consumer grade GPU [P]I am trying to learn concepts like On Policy Distillation (OPD), On Policy Self Distillation (OPSD) and how do they compare to RL algorithms like GRPO. There are a lot of papers on this, but because o→9 Aug 2026[2606.05682] Beyond Output Matching: Preserving Internal Geometry in NVFP4 LLM DistillationDemand for low-precision inference, including NVFP4-based approaches, has grown as large language models are increasingly deployed in latency and cost constrained production environments. Quantization→11 Aug 2026Nvidia Nemo Switchyardhttps://github.com/NVIDIA-NeMo/Switchyard Finally an open source LLM router. An alternative to openrouter fusion and Sakana Fugu. Doesn't look like it does exactly what Sakana Fugu does according to i→11 Aug 2026I built a weird, low-power llama.cpp server using an Intel N100 + RTX 5060TiEverything started with the sudden death of my old ASRock J1900. While looking for the perfect ITX replacement, I stumbled upon the Chinese CW-NAS-ADLN-K motherboard, which looked perfect on paper: In