AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
Human
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
1,927 results
10 Aug 2026

Introducing Muse Glimmer: an open-weight model optimized for always-on local agent workflows

Model ReleasesDGX agent

Hi r/LocalLLaMA 👋 Today we’re excited to release Muse Glimmer, a 30B open-weight model built specifically for local agent workflows. We’re releasing the weights to the community under a permissive Apa

MiniMax H3 with a 4B or 8B text encoder instead of the 32B: update, the voice matches now

Local AiDGX agent

MiniMax H3 loads a 32B text encoder, 15.7 GB, just to turn your prompt into a conditioning tensor. I replaced it with a Qwen3-VL 4B or 8B plus a learned map into the same space. Same DiT, same VAEs, s

Most Civitai 10$ checkpoint's are scams. Don't fall for it.

Model ReleasesDGX agent

If you read -> huge claims + AI like generated presentation + no negative comment AND '$10 to download on my patreon/whatever' = they're scammers. Period. 1- Anyone leaving a negative comment or tiny

DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Motif-Technologies/Motif-3 official realese

Model ReleasesDGX agent

Motif-Technologies is one of the tech company participated South Korea's AI Foundation Model project.(독파모) Upstage(Solar Series), LG AI Research(EXAONE Series), and SKT(A.X Series) are the competitors

Muse Glimmer ACTUALLY fits on a single RTX 3090

Model ReleasesDGX agent

I did some testing this morning, and I was surprised to find that Muse Glimmer actually comfortably fits on a single RTX 3090 with full context + DFlash + mmproj at Q4_K_XL, unlike Qwen3.6-27B and Gem

Muse Glimmer on 1/2 AMD v620

Model ReleasesDGX agent

Hey. Just tried it on my old ass gpus 😄 Surprisingly Tensor Split is working on 2 gpus almost doubling PP (wonder how it will work with 4 gpus) Q6 — 1 GPU llama-server --model <MODEL_DIR>/Muse-Glimmer

Native Long Video Understanding Models locally?

Model ReleasesDGX agent

I've been building a personal project and wanted to check with the community on multi-modal inputs since I can't find a lot of material around this online. Ultimately I'm trying to build something tha

Need real world ML problems to evaluate my educational ML tools

Model ReleasesDGX agent

I'm a retired platform engineer, coding mainly in Rust, and involved with a ML study group. I developed a ML programming language (alternative to Python, Colab) to help me learn (and teach) ML concept

Needle 2: 14MB agentic LLM for phones, wearables, smart home and robots.

Model ReleasesDGX agent

Hey LocalLlaMa, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots a

omlab/VLX-Seek-1.5-10B · Hugging Face

Local AiDGX agent

VLX-Seek-1.5-10B VLX-Seek-1.5-10B is the open-source 10B model in the VLX-Seek 1.5 family, designed for fine-grained perception and visual grounding in embodied scenarios. It targets practical setting

Playing with physics is so cool in Minimax H3

Model ReleasesDGX agent

Prompt: 'integrated_multimodal_description: [Shot 1] Live-action, ultra-realistic first-person footage at night on a rainy city street, filmed with authentic handheld smartphone qualities. The phone i

Please Share Your Experience About Muse Glimmer

Model ReleasesDGX agent

I have a classic test for local LLM's. I asked for 8 ball pool game with only one HTML file and Muse Glimmer spend 21k Token(I m using full context so 128k) and only created a 220 lines of HTML and sa

RAG-art: Build Your Own Art Expert with ollama

Local AiDGX agent

I built myself a personal AI art history assistant https://github.com/lololerigolo60/RAG-art/tree/main I love art history but I have way too many books, PDFs, and notes scattered everywhere. So I buil

Rätt kontakt för rätt person.

Local AiDGX agent

Söker en riktigt vass programmerare – jag har ett projekt jag tror kan bli stort. Jag letar efter en extremt kunnig utvecklare som vill hoppa på ett projekt från ett tidigt skede. Jag kan inte avslöja

Running Qwen 3.5 35B A3B-Q8_0 gguf on a cheap radeon 7600 at 18 token/s

Model ReleasesDGX agent

I also have 64 gb ddr4 ryzen 5600 Using llama.cpp Ubuntu distro Settings are as follows --n-gpu-layers 999 --n-cpu-moe 37 --no-mmap -ctk q8_0 -ctv q8_0 -fa 1 -c 9000 submitted by /u/Sweaty_Perception6

Summary of Takeaways from the Minimax AMA

Model ReleasesDGX agent

Summary from https://www.reddit.com/r/StableDiffusion/comments/1vh9rtw/ama_minimax_h3_team_ask_us_anything_about_our/ This summary was compiled with AI but cross-checked manually by me for accuracy. I

Tested Muse Glimmer locally on coding with OpenCode & agentic work

Model ReleasesDGX agent

Ran the model with quants (Q4) by Unsloth with latest (build from master) llama.cpp server. It takes ~20GB ram running on M5 Pro with 48GB at about 17t/s. Didn't do any reasoning loops/overthinking. O

What Characters Minimax H3 knows - American Edition

Local AiDGX agent

As promised, the first Batch of Characters that Minimax knows - American knowdledge Edition. Hope this helps the Community. Workflow for this was simple: This is the Prompt: Brad Pitt integrated_multi

What Characters Minimax H3 knows - Part 2 - Videogames

Local AiDGX agent

Here is the second Edition, Videogames. Workflow is the same as in the first Part, its pretty simple: This is the Prompt: Brad Pitt integrated_multimodal_description: [Shot 1] Live-action, contemporar

Why Speculative Decoding went mature in 2026?

Local AiDGX agent

Spec-dec has been a thing for a while, in fact, it's wasn't an idea that was born for LLM inference. E.g. Uber's https://github.com/uber/submitqueue applied it to a merge queue. Apple & GDM had been r

Word doc cleaning

Model ReleasesDGX agent

I have been trying to parse word docs for use with llama3.1:8b in Ollama. I only need the text - Even if I cut and past into a text editor weird characters seem to stick around which break llama/Ollam

9 Aug 2026

24 GB of VRAM is not really 24 GB for a local LLM. Here is the worksheet I use

Model ReleasesDGX agent

I kept seeing model file size compared directly with the number printed on the GPU box. That misses several memory buckets. A simple planning model is: usable capacity = advertised VRAM x 0.90 total t

[2606.05682] Beyond Output Matching: Preserving Internal Geometry in NVFP4 LLM Distillation

Model ReleasesDGX agent

Demand for low-precision inference, including NVFP4-based approaches, has grown as large language models are increasingly deployed in latency and cost constrained production environments. Quantization

300b on 32gb MoE-streaming findings + optimisations

HardwareDGX agent

The past week I've been running DSv4 inference on my laptop by keeping everything RAM-resident except the MXFP4-experts (since expert pool is ~147GB and won't fit) TL;DR - read speed is the limiter mo

AMD llama.cpp: reducing MTP buffer overhead gave me 64K → 149K context for Qwen 27B

Model ReleasesDGX agent

Available context length with and without the patch: Model: QWEN 27B ROCm stock patched Vulkan stock patched IQ4_XS Pure, single 16GB GPU 19.456 76.032 68,352 78,592 Q6_K_L on 16GB + 12GB 64,256 149,2

Anyone already used a model imported directly in the ollama cloud

Local AiDGX agent

Ollama allons you to import model but have you ever tried doing so ? Like running model imported from hugging face or you own model ? Any use case you wanna share ? Very curious about that submitted b

Best Embedding + Reranking Model

Model ReleasesDGX agent

What Local Embedding + Reranking Models are you guys running for RAG? I went down this rabbit hole because I wanted a Embedding Model + Reranker for a Translation Memory Server. Essentially, given X p

Cloud Usage Limits

Local AiDGX agent

Former Ollama Cloud $20 dollar plan holder look at returning. How's the state of the usage ATM? It was in a dire state when I left a few months ago. Is it still very limited with Mid sized models? M3,

CyberKimi just dropped strong results on one of ExploitBench’s hardest V8 bugs , points away from Mythos

Model ReleasesDGX agent

Hey everyone ! Quick share from the cyber + local LLM side of things that I found interesting. During this week’s hacker summer camp, an AI researcher and reverse malware engineer veteran 'lordx64' on

DeepSeek V4 Flash 0731 hits 82.7% on Terminal-Bench 2.1 in an independent public-harness run (445 trials)

Model ReleasesDGX agent

Disclosure: I’m the author of Ante. DeepSeek recently reported an 82.7% score on Terminal-Bench 2.1 for DeepSeek V4 Flash 0731. Its evaluation used “DeepSeek Harness minimal mode,” which hasn’t been r

DeepSeek v4 Flash 0731 locally on CPU

Model ReleasesDGX agent

After seeing the benchmark results for the full release of DS v4 Flash 0731, I replaced my 2 x 16GB DDR4 ram sticks with 2 x 32GB DDR4 ram sticks to get a max supported of 128 GB RAM, in hope to be ab

DeepSeek-V4-Flash-0731 Q8_K_XL sometimes stops mid-task in OpenCode - anyone else seeing this?

Model ReleasesDGX agent

Hey everyone, I've been experimenting with the new DeepSeek-V4-Flash-0731 release locally using the Unsloth Studio Q8_K_XL GGUF with OpenCode. Overall, it's been working really well, but I've noticed

Doom Loop: Anyone Else Having DeepSeek v4 Flash 0731 Issues on ollama cloud?

Model ReleasesDGX agent

Am I the only one having issues with DeepSeek V4 Flash? It gets stuck in a loop, as if it can't call the tools, and keeps repeating the same things endlessly without moving forward. Is it a poorly wri

endless-frontier/BigBang-v1 - qwen 3.5 finetunes

Model ReleasesDGX agent

table bench https://huggingface.co/bartowski/endless-frontier_BigBang-v1-GGUF I'm downloading this model only because Bartowski converted it to .gguf, so it might be interesting. Doubts : The headline

I Turned My Underused Gaming Laptop Into a Local AI Workstation

Local AiDGX agent

TL;DR: I am building a Windows-first local AI setup for people who want to try local LLMs without spending days choosing models, setting up Ollama, Docker, WSL, Open WebUI, agents, and tool permission

It took two years, but we finally have a 'local Sora'

Local AiDGX agent

Who remembers when OpenAI previewed Sora two years ago and the quality felt unreal? We had never seen anything like it. Back then, Sora 1 didn't even generate audio and was heavily censored. Prompt: i

KLQ: Training-free measured rotation quantization. Beats all training-free rotation-based quantization methods on W4A4KV4-bits. Llama 3.2 1B KLQ-quantized beats SpinQuant and gets close to ReSpinQuant without GPTQ/LDLQ rounding.

Model ReleasesDGX agent

First of all, I'm not a lab, this was a solo summer research project that finally culminated into the github repo and the writeup. The repo includes a much deeper dive with methods, findings about qua

Lophius: A workbench for language model research, from the creator of Heretic

Local AiDGX agent

Hi folks, I hate slop as much as you do, so instead of starting with 'The Problem', I'll just cut to the chase: I just published Lophius, which is the culmination of more than two years of fighting wi

M3 16GB running Ollama (Qwen 9B) is extremely slow (10-12 mins per task). Am I doing something wrong?

Model ReleasesDGX agent

Hey everyone, I constantly see high praise for M3 and M4 Macs for local LLM inference, even the base/16GB models. However, my experience has been quite different, and I'm trying to figure out if I hav

Memory Bandwidth problems with Intel Sapphire Rapids

Model ReleasesDGX agent

I have a Xeon w7-3465 and 4 sticks of RDIMM DDR5-4800 with a theoretical max bandwidth of 153GB/s. I am trying to run DeepSeek-V4-Flash-0731 as it is an MoE and the weights are in MXFP4, so I should r

[NEW MODEL] SupraElegans-500K

Model ReleasesDGX agent

*SupraLabs released a new experimental model!* SupraElegans-500K is a ~500,000-parameter causal language model built around a sparse, signed, recurrent neural graph. No Transformer, no attention mecha

Open Model: Google Weather Next 2

HardwareDGX agent

I am not a meteorologist, but I just read a very interesting article: https://arstechnica.com/science/2026/08/deepminds-hurricane-model-bought-forecasters-an-extra-day/ In a paper published on Thursda

Open-weight video gen that actually delivers. Five days with MiniMax H3 on local hardware.

Local AiDGX agent

H3 weights went live on HuggingFace August 3rd and I started pulling them immediately. An omni-modal video model with native stereo audio in the same forward pass, where audio can actually drive the v

The Gemma team will host a special event on August 20

Model ReleasesDGX agent

Tweet by u/hackerllama Could be copium, but I would love to see Gemma 4.1 there with unified audio input for all model sizes perhaps even up to 120B, much improved tool calling (even with the latest t

Trustfactor in training data?

Local AiDGX agent

Would it be possible and make sense to add metadata to training data e.g. a trustfactor (0.0 - 1.0)? For example: the older data is the less trustworthy it is. And data after 2022 gets less trustworth

Underestimated budget solution: radeon 780m iGPU

Model ReleasesDGX agent

There are so many posts where people complaining about high prices and asking for solution <= 1000 EUR. So, there is one solution to consider: PC/mini PC/laptop on Ryzen 7 260/Ryzen 9 8945HX/etc CPU w

Updated benchmark: Deepseek V4 Flash on SlopCodeBench (local)

Model ReleasesDGX agent

Howdy - I posted a benchmark here - https://www.reddit.com/r/LocalLLaMA/comments/1vbtiy7/deepseek_v4_flash_on_slopcodebench/ This was using the hosted API - since then I've been playing around with qu

8 Aug 2026

any reasonably fast public benchmarks I should run quants of deepseek flash 0731 on?

Model ReleasesDGX agent

I have various quants of this model and am curious how they perform. can anyone recommend which benchmark would be a good test case for quantization effects? Maybe that can be completed with about 1 m

Anyone else amped up over Qwen 3.8?

Model ReleasesDGX agent

I’ve been using 3.6 27B Q4, and that quant is fast on an M5. The code has been average, but consistently “good enough.” And, after a year, I can see home LLMs being served at home much like streaming

Building a budget 32GB → 48GB VRAM home AI server: 2-3x RX 9060 XT 16GB vs RTX 5060 Ti 16GB, AM5 vs used EPYC?

Model ReleasesDGX agent

I’m planning a dedicated home AI server, mainly for local LLM inference, agents/tool use, Docker services, and eventually larger MoE models with CPU offload. My plan is to start with 2x 16GB GPUs = 32

Building a zero-dependency C inference engine for BitNet (1.58-bit) - lessons from hitting 36 tok/s on a Xeon CPU

Local AiDGX agent

Over the past few months I have been building a CPU-first inference engine from scratch in pure C99 (no Python, no CUDA, no BLAS, just GCC and make). The focus has been running 1.58-bit ternary models

Claude Code in 9 lines python

Model ReleasesDGX agent

I was wondering what a minimal coding agent implementation would look like that can be used like Claude Code or Codex Not feature-by-feature of course but basically stripping everything out that is no

DeepSeek V4 Flash 0731 appreciation post

Model ReleasesDGX agent

I’m running DSV4F 0731 on dual spark, and honestly… wow. It’s an absolute workhorse, and the benchmarks are real. Everyday tasks with Hermes agent? Effortless. Coding tasks with OpenCode? I’m genuinel

enabling PCI-E p2p for consumer Nvidia cards will yield you more than you think

Model ReleasesDGX agent

Disclaimer - no LLM was used to write this post/note As larger post about my setup will come later, want to give heads-up to folks who use VLLM and >= 2 GPUs. So I have pretty meaty server (8 channel

Extremely slow DSpark draft model performance (1-2 t/s) with DeepSeek-V4-Flash on llama-server compared to MTP?

Model ReleasesDGX agent

Hey everyone, I could use some advice on setting up speculative decoding correctly with llama-server. My Hardware: GPUs: RTX 4090 + RTX 6000 Pro (120GB total VRAM) RAM: 32GB I am currently testing the

Has anyone here fiddled with TPUs for inference ?

Local AiDGX agent

I discovered recently that Google uses their own TPUs, like tiny ASIC cards like the toy ones that existed for bitcoin. And while it sounds inefficient the fact they use thousands of them because...th

I tested a fresh GitHub download → Ollama → first local coding-agent task (72 seconds, no cloud API)

Model ReleasesDGX agent

I’m building DesktopLab, an open-source local-first control plane for development agents. I recorded the setup boundary that most agent demos skip: DesktopLab detects the host, proposes the supported

IDE with Locall LLMs?

Local AiDGX agent

What IDE are you using. its another problem area for me . I usually use VSCode , but with local llms I have not found an extension which works optimally VSCode CoPilot chat with Ollama: CoPilot bloats

Is anyone else finding DeepSeek-V4-Flash unreliable for non-coding tasks?

Model ReleasesDGX agent

(I am not a native speaker, written by myself, so please bear with me) I really want to like DeepSeek-V4-Flash-0731. But it has serious flaws that don't align with the high score on intelligence bench

Kimi K3 (Unsloth) IQ2-XXS from 711GB down to 478GB!!! Only Multi-language was removed to trim the size

Model ReleasesDGX agent

Firstly a big thanks to the poster 'hellohazine', he basically only removed the multi-lingual fat of the model and just kept the English language intact. It is the exact model, and the rest of the mod

← Previous
1234…33
Next →