GLM 5.2 and ik_llama.ccp
Running GLM-5.2 (the new glm-dsa arch), Unsloth UD-Q4_K_XL, on a 4-socket Xeon E7-8880 v4 box with 1TB RAM and a single RTX 3060 12GB. ik_llama.cpp, experts on CPU (--cpu-moe), 24 attention layers on
Knowledge catalogue
Running GLM-5.2 (the new glm-dsa arch), Unsloth UD-Q4_K_XL, on a 4-socket Xeon E7-8880 v4 box with 1TB RAM and a single RTX 3060 12GB. ik_llama.cpp, experts on CPU (--cpu-moe), 24 attention layers on
I ran DeepSeek V4 Flash through Claude Code, OpenCode and Pi on my own benchmark, and the quality came out basically the same across all three while the time and tokens spent was wildly different. Cla
Qwen3.6-27B: SFT vs continued pre-training vs RL? I’m interested in adapting Qwen3.6-27B, but I’m increasingly unsure whether conventional SFT/LoRA is the best route if the goal is to add a capability
Jamie John / Financial Times: How AI companies are targeting the education market, including making free or cut-price tailored learning tools in partnership with schools and edtech startups — Anthropi
I am trying to decide on buying the ollama pro subscription. But their usage policy is vague as hell. I don't mind the 'GPU usage time' but how much GPU time do I actually get? I still don't seem to f
Wall Street Journal: How US companies flipped from “tokenmaxxing” to “thrift-maxxing”, mixing cheaper Chinese models with OpenAI and Anthropic, threatening the labs' IPO valuations — Companies big and
Hugging Face CEO calls for ‘radical transparency’ after ‘unprecedented’ OpenAI hack https://techcrunch.com/2026/07/26/hugging-face-ceo-calls-for-radical-transparency-after-unprecedented-openai-hack/?u
Most graph-based LLM interfaces use a canvas as a visual layer over what is still a linear chat. I wanted the graph itself to determine what Ollama receives. ThoughtDAG has one rule: wires are the con
This was my Bachelor's Final Project: implementing YOLO26n inference completely from scratch using ARM64 Assembly Language and C, without relying on existing inference frameworks. The goal was to unde
I'm a software engineer who mainly builds softwaes/applications, and I'm starting to work on machine learning projects. Since ML workloads often require GPUs, I know services like Google Colab and Kag
important thread; don’t read just the first tweet (which has a caveat in the second). In case you think the problem of science slop is hypothetical: Here's the president of OpenAI retweeting wrong sci
It continues to boggle the mind how many people, who otherwise seem to have a functioning intellect, appear to lose all cognitive capacity when it comes to thinking about actions that impact the compa
Andrej Karpathy, a prominent advocate for open-source AI and a co-founder of OpenAI, appears to have removed Anthropic from his X bio, suggesting he may have left the company. Karpathy joined Anthropi
What quantization or storage size tiers should we expect further down the line? This site suggests q4 and q8 are coming. But I don't know how to translate that to storage size for local inference. htt
Kimi K3 will be here on Monday, but how do you know you’re ready to scale it in production? With our latest inference platform updates you can: 1/ Run shadow traffic to see how a new model performs on
Hey r/LocalLLaMA — maintainer here, obviously biased. OpenSmith is an open-source Python tracing tool for LLM pipelines. The idea: drop u/trace on any function, run opensmith ui, get a full local dash
NOTE -> I expect answer from people who actually have experience and strong understanding of these. please give something beneficial. I'm building a SaaS platform in Sri Lanka that handles documents a
I am setting up codex+ollama+qwen3.6:27b for my hobby coding project on a windows pc with rtx5090 - earlier i tried to setup vllm in Ubuntu container but couldn’t get that to work - now using ollama,
New research from NVIDIA. Does AdamW have a scale ceiling? This work claims yes, and shows where it sits. At batch sizes up to 100M tokens for next-token prediction, SOAP and Muon maintain training st
https://preview.redd.it/umcg41g2kkfh1.png?width=502&format=png&auto=webp&s=7b1ed545cd1c7840a461101bd02d78ab4a9cb40d I tried to create an Ollama account to purchase a subscription a few months ago. But
Ollama is proud to sign @satyanadella's letter. Our mission from day one has been to make open models accessible to every developer to unlock the next frontier in America and across the globe. Open-we
I have been running some experiments with smaller open-weight LLMs on multiple-choice questions of Swedish medical licensing exams. On a dataset called MedQA-SWE, GPT-4 scored 84% accuracy in 2024 and
Large language models (LLMs) are transforming recommender systems from matching co-occurrence patterns in historical behavior toward reasoning about the intent that drives it. RecGPT-V1 pioneered this
“Pelican on a bicycle” LLM benchmark by @simonw 2024 (https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle/) is now “Call of Duty” 2026… Claude Opus 5 one-shotted this game. EVERYTHING you see
How does it work? I read that the limts were more than what Opencode Go offers and subscribed to the 20USD Pro plan. But In practice based on how the usage bar fills up it seems like the 5 hour window
Recent discussions about open-sourcing make me feel that I should go back and revisit these important open-source works in representation learning that pushed the field forward and eventually made vis
https://preview.redd.it/qoogv5m1skfh1.png?width=688&format=png&auto=webp&s=3f06fb28ceec3cd3ab95bcebd0c71374332aac9b Built a simple desktop GUI (Tkinter) that supports multiple local TTS engines (curre
Simply untrue. The Economist asked Musk about a wide range of things, including AI and robotics. Watch it here for yourself: https://www.economist.com/insider/the-insider/an-interview-with-elon-musk?u
Wall Street Journal: Sources: Nvidia is in talks to provide a ~$250B backstop for OpenAI as part of a 10 GW data center project that SoftBank is developing in Ohio — Project would be one of the larges
--> .abbel-fig { display: block; text-align: center; margin: 2.4em 0; line-height: 1.4; max-width: 100%; } .abbel-fig img { display: block; margin: 0.65em auto 0; height: auto; max-width: 100%; } /* I
“The first autonomous agent cyberattack is an unprecedented event. It deserves an unprecedented response!” In the spirit of transparency, here’s what I asked @OpenAI: • Radical transparency: let’s rel
The Top AI Papers of the Week (July 20 - July 26): - GAMUT - PRO-LONG - Harness Handbook - From Memory to Skills - Progressive Disclosure - Global Workspace in LLMs - Structured Output Collapses Diver
This blog is full of design choices as you build your Agentic systems. Critical choices to be made, evals need to be upgraded if the underlying intelligence is upgraded. Well thought out blog—new stre
This part is spot on: > 'Overall, we found that we were over-constraining Claude Code...while these constraints were once needed to avoid worst case scenarios, we have since found we can delete many o
training to benchmarks ≠ getting to AGI This suggests that Opus 5 's gain on ARC-AGI-3 was the result of specific training to improve on that eval, and not a generalized increase in abstract reasoning
Hey everyone, I have been looking into how people source compute for their Inference workloads (and in general). I wanted to understand some specific pain points here. If you've used online services l
There are a few reasons why problems from International Mathematical Olympiad function as a good benchmark for LLMs: - The problems are new, not included in the training data of any model - Hard math
At Bitwize, we've been building Logue, a native macOS app for AI meeting notes and writing, and we just open-sourced it (MIT). We're sharing it here because the whole point is that it runs 100% on-dev
We turned the browser into the inference server. @RunAnywhereAI ( @runanywhere/web ) runs models client-side with WebGPU + WASM SIMD, stores them in OPFS and generates with 0 outbound bytes. No backen
I have tried ollama locally on my gaming PC (5070 with 12bg of VRAM) It works pretty nice on qwen2.5-coder:7b (and 14b) So I decided to take a step further, and install an openWebUI instance on my hom
I am seeing more and more videos as posts about how OpenAI is in complete financial ruin, Anthropic isn't much better. Their expenses go with the revenue they make etc etc. Meta made big investments i
Qwen3.6-27b is fantastic! It makes me wonder if there's a hard ceiling to smaller sized models. Do you guys think the ceiling of intelligence for smaller models will be constrained by factors like par
Yes, open-source / open-weight models are important for a healthy AI ecosystem. That's how we can verify things, check claims, and keep up outside the closed labs. Plus, it gives us the freedom to run
3 year lurker, now i finally got my server up and running. dont know which model to choose. llama.cpp or vllm, what makes more sense? mainly single user with maybe 2-3 more additional users in family,
Hi everyone, I have a system which I have been using for gaming, R7 5700X, 48GB DDR4, RTX3090 24GB. But I want to use it for Local AI to reduce my reliance on cloud AI providers (mainly usage limits -
Financial Times: A look at China's bid to build an alternative global order in AI by making open models widely available and training people in developing countries to use them — Beijing makes most am
Financial Times: A profile of Yang Zhilin, founder of Moonshot AI, which faced early doubts over revenue and model capabilities before its Kimi K3 delivered a “DeepSeek moment” — Known as ‘Yang the ge
All anyone has to do is go back 15 years and look at all the ridiculous predictions Elmo has made about Mars, self driving cars, his ridiculous Boring Company. No one should take anything he says seri
AI hardware competition entered a sharper phase this week as Advanced Micro Devices Inc. used its flagship AI event to argue it isn’t merely chasing Nvidia Corp. — it intends to lead the market outrig
As AI expands into robotics and other real-world applications, the focus is shifting toward integrated platforms that simplify development and deployment. Open ecosystems and standardized architecture
The market for data center GPUs is evolving beyond individual chip specifications into a contest over fully integrated rack-scale systems. As AI workloads scale into the gigawatt range, buyers increas
Anthropic PBC today rolled out a large language model called Claude Opus 5 to its chatbot service and developer platform. The company says the LLM approaches the output quality of its top-end Mythos 5
Twenty years on, AWS EC2 compute is meeting demand shaped by agentic AI, physical AI, and customers pushing general-purpose cloud infrastructure into new territory its earliest architects never antici
Cuda and Vulkan Benchmark: TensorSharp vs. llama.cpp I would like to share my latest open source local Unsloth (GGUF) LLM inference engine and applications. It supports many models from Unsloth, like
I apologize that this question has been asked in various flavors over time, but I couldn't find anything in the posts before that matches the options I have. I have a PC and a mac, both available in m
but they won’t. AI safety experts say OpenAI’s rogue models may mean the company has already blown past its own internal red lines. That would mean it should pause development until it creates better
If you run local agentic coding harnesses (Aider, Claude Code, etc.), prompt evaluation usually eats up most of your execution time. Every turn re-evaluates thousands of identical prefix tokens_system
AMD’s AI accelerator software stack has progressed sharply, moving from a 0 % chance of catching Nvidia’s CUDA moat in early 2025 to a “great chance” of success by July 2026 after leadership changes a
Claude Opus 5 generated an entire game demo from Scratch, with no external assets used—Matt Shumer posted a video of the AI‑created gameplay on Twitter. The clip, narrated by Shumer as “AI games are g
Cohere has proudly signed on to this letter. The importance of open-source models to the AI ecosystem cannot be understated. We believe everyone, in every country, should have control over the technol