Are 1B LLMs Going Away in 2026?
I don't know much about llms aside from downloading them through a frontend and running them on my laptop or potato phone. Google released gemma 4, but unlike gemma 3, there isn't a 1b model this time
Knowledge catalogue
I don't know much about llms aside from downloading them through a frontend and running them on my laptop or potato phone. Google released gemma 4, but unlike gemma 3, there isn't a 1b model this time
audio.cpp 0.5 is out :) The most fun new model in 0.5 is DramaBox. It is closer to prompt-directed voice acting. DramaBox is built on the LTX-2.3 audio architecture, and prompts can control emotion, d
chat : enable tool call in thinking for DS4 (#26269) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFr
mtmd: add minicpmv46 downsample (#25993) add minicpmv46 downsample Signed-off-by: tc-mb tianchi_cai@icloud.com put downsample mode inside gguf. Signed-off-by: tc-mb tianchi_cai@icloud.com build mtmd_i
cli : persist reasoning_content in chat history (#26362) cli : persist reasoning_content in chat history llama-cli collected reasoning from the stream for display but only stored assistant content in
vendor : update BoringSSL to 0.20260730.0 (#26353) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFram
test: fix some CI errors (#26415) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubun
Banning open source model, will 'remove capabilities for the defenders' When OpenAI's models breached Huggingface, a Chinese open model is what cleaned up the mess. Anthropic's Fable 5 refused, so Hug
I'm currently using the gemma4:31b-cloud, and it is pretty good, but sometimes it gets confused. Are there more free tier models out there I should try out? or is gemma4 the cap of free tier cloud mod
Release: datasette-apps 0.2a0 Changes that improve Datasette Apps when created and edited using Datasette Agent: New app_debug() tool allowing agent to open an app (invisibly) and test it using JavaSc
Thanks to the community help I finally launched this llm. LM Studio refused to load weight onto second GPU but Unsloth Studio did so everything was done in there. Not a proper benchmark (used PC in pa
DeepSeek-V4-Flash-0731 is live on Fireworks, day-zero. DeepSeek reports it beats V4 Pro across all 9 agentic evals, incl. 82.7% on Terminal Bench. Better cost-per-task than V4 Pro, at the economical p
DeepSeek-V4-Flash-0731 is now available on Ollama's cloud. This update substantially enhances the model's agentic capabilities: ollama run deepseek-v4-flash:0731-cloud Use it with Claude Code: ollama
DeepSeek-V4-Flash-0731 is now over 2x faster than yesterday on Ollama's cloud! DeepSeek-V4-Flash-0731 is now available on Ollama's cloud. This update substantially enhances the model's agentic capabil
March 6th, 2026 the highest intelligence index score was 51 for frontier models. deepseek-ai/DeepSeek-V4-Flash-0731 that has an intelligence score of 50. If these benchmarks are accurate, models avail
Here are my numbers: Quant Size Layout Decode Prefill Draft acceptance UD-Q8_K_XL 150.8 GiB 20 layers CUDA0 / 23 ROCm0 + drafter 44.0 t/s 564 t/s 0.535 UD-Q4_K_XL 144.4 GiB 22 / 21 + drafter 48.4 t/s
The biggest issue with preview was its inability to follow rules prompts and skills. It seems like no matter what you do it ignores them. I've tried first person and second person. I've tried Chinese
I managed to run DeepSeek-V4-Flash-0731 UD-IQ3_S in text-generation-webui with: RTX 3090 24 GB 128 GB DDR5 overclocked to 5600 MHz using AMD EXPO llama.cpp loader First, I had to use a rather brutal w
For anyone interested, here are the llama-bench results on 3 bit K_XL quantization. I think this could be pushed further but no luck so far. CURRENT RESULTS: full moe offloading Prefill suffers 116 --
Antirez stealthily uploaded the new weights in the old folder... and there we were tapping our fingers. https://huggingface.co/antirez/deepseek-v4-gguf/tree/main submitted by /u/challis88ocarina [link
https://preview.redd.it/1a39x4zivqgh1.png?width=1550&format=png&auto=webp&s=de591c039cc18782a6b5d8e402fdc1594be05132 start C:llmllamam5uildinllama-server.exe --model 'H:UD-Q3_K_XLDeepSeek-V4-Flash-073
Fascinating: OpenAI’s @deanwball is saying Astra can do anything, and it’s not even clear it can do “anything” in math (let alone anything in more or open-ended, less formalizable domains). I dropped
A few people asked for this after the bandwidth thread, so here it is on its own instead of buried in a comment. The dense rule was simple: every token reads every weight, so tokens/sec ≈ bandwidth ÷
On July 24, @swyx announced that he had started work on 'forge agents' and outlined four new features for SmolForge: customizable skins and spritesheet animations. He also referenced an upcoming blog
I am trying to learn concepts like On Policy Distillation (OPD), On Policy Self Distillation (OPSD) and how do they compare to RL algorithms like GRPO. There are a lot of papers on this, but because o
> Google was too nervous to release it and DeepMind was blocked from shipping products that could disrupt Google. bookmark for the next vc that asks you 'what if <incumbent> builds this?' @_chenglou I
great write up, evals are hard! here are 2 broad buckets we use to evaluate agents: 1. Measure the State of the World 2. Agent as a Judge on the Trajectory 1. Measure the state of the environment befo
Grok Build can do almost anything you can think of http://X.ai/cli Most people seriously underestimate what Grok Build can do They assume an AI coding agent is only useful for building apps or writing
Hot take on OpenAI’s Astra: - Obviously impressive - But math is different from most other problems in that it is more amenable to to formal verification and synthetic data. How well it works in open-
Ciao a tutti ho appena comprato 4 Huawei duo 96 GB a poco più di 4mila euro qualcuno ha già utilizzato CANN consigli da darmi? E la prima volta che utilizzo questo framework e non so proprio da dove i
GitHub: https://github.com/Loann110/Vao2 Hello everyone, I’m currently developing Vao2, an open-source application that brings together news, YouTube channels, GitHub repositories and more (to be adde
If Leopold had read this on June 26 and trimmed his bets accordingly, SALP would not have melted down. I laid everything out. https://open.substack.com/pub/garymarcus/p/the-month-generative-ai-lost-it
If you maintain an AGENTS.md or a CLAUDE.md, this is worth a read. (bookmark it) 288 gold-test evaluated runs across Claude Code and Codex, 17 real tasks from 3 repositories, with context-injection st
In a world where building AI applications is getting easier every day, the biggest moat won’t be the application itself. It will be a deep understanding of your users. The winners will be the companie
New York Times: Inside Larry Ellison's debt-fueled push to turn Oracle into an AI juggernaut by aligning with Trump, backing Project Stargate, and partnering with OpenAI — the first full day of the se
DeepSeek V4 Flash got me thinking... We keep seeing smaller models get way better. A model at a certain parameter count today can be much smarter than a model of the same size from a year or two ago.
Kimi K3 has set a new bar for OSS model intelligence! 2.8T params, 1M context, OpenAI-compatible API. Complete guide to running Kimi K3 on Together AI 👇 👏👏 @Kimi_Moonshot 👏👏 https://www.together.ai/bl
**Kimi K3 is Moonshot AI’s 2.8‑trillion‑parameter open‑weight language model—the largest ever released—designed for frontier tasks such as long‑horizon coding and deep reasoning.** Its architecture us
Hello all - i just got a new mac mini with 24GB of RAM and wanting to run local AI for Home Assistant and Hermes. I have been struggling to find a snapy model that will work with my machine. Currently
OpenAI really ought give @HuggingFace a $100M grant, much as @ClementDelangue is asking. And anyone who wants to understand the attack from HF’s side should read this. Great walkthrough; A+ for visual
Our Twitch streams with @Cohere_Labs are #1 in the tech category⚡️ Cohere Labs' free ML Summer School series has brought together viewers from 20+ countries to learn about everything from NLP in LLMs
// Persistent Workspaces for Long-Lived Claude Code Agent Teams // Four issues to be aware of: > Working state vanishes when a terminal closes and the team cannot be resumed. > Compaction condenses th
Wondering if it's possible to actually produce ai art works that are indistinguishable from authentic medieval woodcut illustrations like the one attached. All the AI attempts I've seen at re creating
at openai, many people hook their chatgpt up to slack. people really don't like when a coworker's chatgpt contacts them asking for help with a task, even when they'd be perfectly happy doing that same
CPU: Threadripper 3970X RAM: 128GB DDR4 GPUs: 3x2080ti 11GB The current best parameters to run it: llama-server --model Qwen3.6-27B-Q5_K_S.gguf --n-gpu-layers 999 --split-mode tensor --flash-attn on -
so cool to just see this pop up on my feed. hermes has built such an organic and creative community. there are so many bells and whistles to explore and being able to find those through natural langua
so much for the “general” part of general intelligence I have tried many times to get ChatGPT, Claude, Grok or Gemini to write scripts for my YouTube videos. It is still a complete failure. For one th
Some things to note: 1) AI is getting very good at math and science. 2) Two years ago LLMs could not do basic math consistently 3) According to @polynoamial this cost less than $2000 in current API co
Berber Jin / Wall Street Journal: Sources detail how OpenAI fell behind Anthropic in revenue growth and valuation after prioritizing consumer chatbots and flashy side projects over coding tools — Bets
“Stochastic parrots” is not my term (it’s @emilymbender’s). But a lot of people today commenting on it are confused. To some extent (though I don’t think it’s a perfect metaphor, and have said that be
Ten advances in mathematics and theoretical computer science A few days ago it was Anthropic discovering cryptographic weaknesses with Claude using Mythos Preview, spending 100,000 on tokens and with
The “AGI-is-near” community keeps committing the same logical fallacy over and over; I have seen it at least half a dozen times today alone. Every time there’s an advance, I see the same error. Here’s
TLDR: There is no one prompting trick that will result in Krea 2 Turbo giving you exactly the style you want and across the whole image. Instead, if you are trying to achieve styles without the use of
Kyle Alspach / CRN: ThreatLocker raised a $190M Series F led by Elephant as it looks to extend its zero-trust enterprise security platform to protect against AI-related risks — The cybersecurity vendo
So upfront I'll admit that this is a ChatGPT summary of my chat with it about this idea i had, but this post wouldn't exist any other way, so... I’m trying to get a better understanding of how VRAM is
DeepSeek V4 Flash is now available on InferX, and it’s free to use. We’re continuing to add GPU capacity as demand grows. While we’re bringing additional capacity online, you may occasionally see high
Vivid depiction of a dark scenario for Nvidia. “Indeed, CEO Jensen Huang has been so aggressive in funding the AI ecosystem that the company is functioning almost like a bank; with annual free cash fl
Wake me when Astra solves a significant open-world problem that doesn’t revolve around formal verification. Or at least fixes poor @skdh’s video problems. I have tried many times to get ChatGPT, Claud
We are going to see a lot of vertically focused AI native companies accelerate. Routers, open-source models and specialized post-training enabled by companies like @FireworksAI_HQ have all made dramat
What speeds are everyone getting with deepseek v4 flash 0731? I’m getting~200 tps prompt processing / ~11 tps token gen, on 4x5060ti16gb with ddr4 3200 ram at 4-channel, via llamacpp, with context win