AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
Human
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
1,927 results
Local Ai

Best Local LLMs - August 2026

DGX agent

Wowee!! Just when you thought it couldn't get better for open weight models, we probably have had our best period yet!?!?! Models that rival the closed frontier, Opus level models on non-insane hardwa

local-air-localllama
10 Aug 2026
Model Releases

Best open-source harness like Claude Code?

DGX agent

Avid claude code user here looking to do equivalent things with local models. Just want to plug in something like Qwen and have the interface be 1:1 with claude code. Any suggestion? submitted by /u/N

DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
model-releasesr-localllama
10 Aug 2026
Model Releases

Chat UIs with native audio input for multimodal models?

DGX agent

I've been running Gemma 4 E4B with oMLX and I can't find any chat interfaces that directly send the audio file to the model instead of running the audio through a separate STT layer. I can confirm the

model-releasesr-localllama
10 Aug 2026
Model Releases

Comparing how Cline, Kilo, and Qwen Code handle long-task context/state (and why context loops keep happening)

DGX agent

I've been comparing Cline / Kilo / Qwen Code lately since they all handle long-task state differently. Cline: has Focus Chain, a markdown file kept outside the conversation that gets reinjected on a c

model-releasesr-localllama
10 Aug 2026
Model Releases

DeepSeek V4 Flash 0731 is the ‘killer app’ that is going to sell A LOT of DGX Sparks

DGX agent

Having a ‘Killer Application’ that everyone wants to use helps sell hardware, plain and simple. DeepSeek V4 Flash 0731 isn’t an app of course, but I think it’s going to be the major catalyst for getti

model-releasesr-localllama
10 Aug 2026
Model Releases

GPT 5.6 Sol High and X-High (Web Chat) feels severely nerfed since 08/06/2026 update

DGX agent

GPT-5.6 Sol High and X-High (Web Chat) feels severely nerfed since 08/06/2026 update (https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt) I use it for a pretty complex Unreal Engine 5 project (

model-releasesr-chatgpt
10 Aug 2026
Local Ai

How to prevent LLM to act like a robot/assistant?

DGX agent

I'm playing with a conversational agent I made using either api/generate or api/chats. In both case I do ask him to not ask follow up question, to not act like an assistant, etc. Either from a system

local-air-ollama
10 Aug 2026
Model Releases

I asked OPUS 5 to make a video about what it's like to be an LLM

DGX agent

Full prompt I gave to Claude Opus 5: can you use whatever resources you like, and python, to generate a short 'youtube poop' video and render it using ffmpeg ? can you put more of a personal spin on i

model-releasesr-chatgpt
10 Aug 2026
Model Releases

I compared GGUF quants of Qwen3.6 27B to NVFP4, AWQ, AutoRound, and FP8

DGX agent

There's an interactive chart and some extra data in the blog post if you're interested. There are plenty of KL-divergence benchmarks for GGUF models, but most of them compare one GGUF quant against an

model-releasesr-localllama
10 Aug 2026
Model Releases

I trained a 1B-parameter LLM from scratch on 20B tokens for about $200

DGX agent

A few months ago, I had the idea of making a LLM from scratch as a personal project (for learning and partly for improving my resume). Since I learned a lot from other posts on here over the past year

model-releasesr-localllama
10 Aug 2026
Local Ai

I trained an open-source realism LoRA for MiniMax H3 - it makes generated people actually look real (weights inside)

DGX agent

Update : New version is ready and online , should be much better, fully functionnal on ComfyUI, and you can find before/after here : https://huggingface.co/fal/MiniMax-H3-Realism-People-LoRA/blob/main

local-air-stablediffusion
10 Aug 2026
Model Releases

inclusionAI/Ling-3.0-tiny · 8B A1.3B MoE· Hugging Face

DGX agent

Looks like the Ling team open weighted a much smaller version of the Ling-3.0-flash they open weighted a few days ago. It's 8B params with 1.3B active, and seems to fall between the 4B and 8-12B Qwen

model-releasesr-localllama
10 Aug 2026
Model Releases

Introducing Muse Glimmer: an open-weight model optimized for always-on local agent workflows

DGX agent

Hi r/LocalLLaMA 👋 Today we’re excited to release Muse Glimmer, a 30B open-weight model built specifically for local agent workflows. We’re releasing the weights to the community under a permissive Apa

model-releasesr-localllama
10 Aug 2026
Local Ai

MiniMax H3 with a 4B or 8B text encoder instead of the 32B: update, the voice matches now

DGX agent

MiniMax H3 loads a 32B text encoder, 15.7 GB, just to turn your prompt into a conditioning tensor. I replaced it with a Qwen3-VL 4B or 8B plus a learned map into the same space. Same DiT, same VAEs, s

local-air-stablediffusion
10 Aug 2026
Model Releases

Most Civitai 10$ checkpoint's are scams. Don't fall for it.

DGX agent

If you read -> huge claims + AI like generated presentation + no negative comment AND '$10 to download on my patreon/whatever' = they're scammers. Period. 1- Anyone leaving a negative comment or tiny

model-releasesr-stablediffusion
10 Aug 2026
Model Releases

Motif-Technologies/Motif-3 official realese

DGX agent

Motif-Technologies is one of the tech company participated South Korea's AI Foundation Model project.(독파모) Upstage(Solar Series), LG AI Research(EXAONE Series), and SKT(A.X Series) are the competitors

model-releasesr-localllama
10 Aug 2026
Model Releases

Muse Glimmer ACTUALLY fits on a single RTX 3090

DGX agent

I did some testing this morning, and I was surprised to find that Muse Glimmer actually comfortably fits on a single RTX 3090 with full context + DFlash + mmproj at Q4_K_XL, unlike Qwen3.6-27B and Gem

model-releasesr-localllama
10 Aug 2026
Model Releases

Muse Glimmer on 1/2 AMD v620

DGX agent

Hey. Just tried it on my old ass gpus 😄 Surprisingly Tensor Split is working on 2 gpus almost doubling PP (wonder how it will work with 4 gpus) Q6 — 1 GPU llama-server --model <MODEL_DIR>/Muse-Glimmer

model-releasesr-localllama
10 Aug 2026
Model Releases

Native Long Video Understanding Models locally?

DGX agent

I've been building a personal project and wanted to check with the community on multi-modal inputs since I can't find a lot of material around this online. Ultimately I'm trying to build something tha

model-releasesr-localllama
10 Aug 2026
Model Releases

Need real world ML problems to evaluate my educational ML tools

DGX agent

I'm a retired platform engineer, coding mainly in Rust, and involved with a ML study group. I developed a ML programming language (alternative to Python, Colab) to help me learn (and teach) ML concept

model-releasesr-localllama
10 Aug 2026
Model Releases

Needle 2: 14MB agentic LLM for phones, wearables, smart home and robots.

DGX agent

Hey LocalLlaMa, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots a

model-releasesr-localllama
10 Aug 2026
Local Ai

omlab/VLX-Seek-1.5-10B · Hugging Face

DGX agent

VLX-Seek-1.5-10B VLX-Seek-1.5-10B is the open-source 10B model in the VLX-Seek 1.5 family, designed for fine-grained perception and visual grounding in embodied scenarios. It targets practical setting

local-air-localllama
10 Aug 2026
Model Releases

Playing with physics is so cool in Minimax H3

DGX agent

Prompt: 'integrated_multimodal_description: [Shot 1] Live-action, ultra-realistic first-person footage at night on a rainy city street, filmed with authentic handheld smartphone qualities. The phone i

model-releasesr-stablediffusion
10 Aug 2026
Model Releases

Please Share Your Experience About Muse Glimmer

DGX agent

I have a classic test for local LLM's. I asked for 8 ball pool game with only one HTML file and Muse Glimmer spend 21k Token(I m using full context so 128k) and only created a 220 lines of HTML and sa

model-releasesr-localllama
10 Aug 2026
Local Ai

RAG-art: Build Your Own Art Expert with ollama

DGX agent

I built myself a personal AI art history assistant https://github.com/lololerigolo60/RAG-art/tree/main I love art history but I have way too many books, PDFs, and notes scattered everywhere. So I buil

local-air-ollama
10 Aug 2026
Local Ai

Rätt kontakt för rätt person.

DGX agent

Söker en riktigt vass programmerare – jag har ett projekt jag tror kan bli stort. Jag letar efter en extremt kunnig utvecklare som vill hoppa på ett projekt från ett tidigt skede. Jag kan inte avslöja

local-air-ollama
10 Aug 2026
Model Releases

Running Qwen 3.5 35B A3B-Q8_0 gguf on a cheap radeon 7600 at 18 token/s

DGX agent

I also have 64 gb ddr4 ryzen 5600 Using llama.cpp Ubuntu distro Settings are as follows --n-gpu-layers 999 --n-cpu-moe 37 --no-mmap -ctk q8_0 -ctv q8_0 -fa 1 -c 9000 submitted by /u/Sweaty_Perception6

model-releasesr-localllama
10 Aug 2026
Model Releases

Summary of Takeaways from the Minimax AMA

DGX agent

Summary from https://www.reddit.com/r/StableDiffusion/comments/1vh9rtw/ama_minimax_h3_team_ask_us_anything_about_our/ This summary was compiled with AI but cross-checked manually by me for accuracy. I

model-releasesr-stablediffusion
10 Aug 2026
Model Releases

Tested Muse Glimmer locally on coding with OpenCode & agentic work

DGX agent

Ran the model with quants (Q4) by Unsloth with latest (build from master) llama.cpp server. It takes ~20GB ram running on M5 Pro with 48GB at about 17t/s. Didn't do any reasoning loops/overthinking. O

model-releasesr-localllama
10 Aug 2026
Local Ai

What Characters Minimax H3 knows - American Edition

DGX agent

As promised, the first Batch of Characters that Minimax knows - American knowdledge Edition. Hope this helps the Community. Workflow for this was simple: This is the Prompt: Brad Pitt integrated_multi

local-air-stablediffusion
10 Aug 2026
Local Ai

What Characters Minimax H3 knows - Part 2 - Videogames

DGX agent

Here is the second Edition, Videogames. Workflow is the same as in the first Part, its pretty simple: This is the Prompt: Brad Pitt integrated_multimodal_description: [Shot 1] Live-action, contemporar

local-air-stablediffusion
10 Aug 2026
Local Ai

Why Speculative Decoding went mature in 2026?

DGX agent

Spec-dec has been a thing for a while, in fact, it's wasn't an idea that was born for LLM inference. E.g. Uber's https://github.com/uber/submitqueue applied it to a merge queue. Apple & GDM had been r

local-air-localllama
10 Aug 2026
Model Releases

Word doc cleaning

DGX agent

I have been trying to parse word docs for use with llama3.1:8b in Ollama. I only need the text - Even if I cut and past into a text editor weird characters seem to stick around which break llama/Ollam

model-releasesr-ollama
10 Aug 2026
Model Releases

24 GB of VRAM is not really 24 GB for a local LLM. Here is the worksheet I use

DGX agent

I kept seeing model file size compared directly with the number printed on the GPU box. That misses several memory buckets. A simple planning model is: usable capacity = advertised VRAM x 0.90 total t

model-releasesr-ollama
9 Aug 2026
Model Releases

[2606.05682] Beyond Output Matching: Preserving Internal Geometry in NVFP4 LLM Distillation

DGX agent

Demand for low-precision inference, including NVFP4-based approaches, has grown as large language models are increasingly deployed in latency and cost constrained production environments. Quantization

model-releasesr-localllama
9 Aug 2026
Hardware

300b on 32gb MoE-streaming findings + optimisations

DGX agent

The past week I've been running DSv4 inference on my laptop by keeping everything RAM-resident except the MXFP4-experts (since expert pool is ~147GB and won't fit) TL;DR - read speed is the limiter mo

hardwarer-localllama
9 Aug 2026
Model Releases

AMD llama.cpp: reducing MTP buffer overhead gave me 64K → 149K context for Qwen 27B

DGX agent

Available context length with and without the patch: Model: QWEN 27B ROCm stock patched Vulkan stock patched IQ4_XS Pure, single 16GB GPU 19.456 76.032 68,352 78,592 Q6_K_L on 16GB + 12GB 64,256 149,2

model-releasesr-localllama
9 Aug 2026
Local Ai

Anyone already used a model imported directly in the ollama cloud

DGX agent

Ollama allons you to import model but have you ever tried doing so ? Like running model imported from hugging face or you own model ? Any use case you wanna share ? Very curious about that submitted b

local-air-ollama
9 Aug 2026
Model Releases

Best Embedding + Reranking Model

DGX agent

What Local Embedding + Reranking Models are you guys running for RAG? I went down this rabbit hole because I wanted a Embedding Model + Reranker for a Translation Memory Server. Essentially, given X p

model-releasesr-localllama
9 Aug 2026
Local Ai

Cloud Usage Limits

DGX agent

Former Ollama Cloud $20 dollar plan holder look at returning. How's the state of the usage ATM? It was in a dire state when I left a few months ago. Is it still very limited with Mid sized models? M3,

local-air-ollama
9 Aug 2026
Model Releases

CyberKimi just dropped strong results on one of ExploitBench’s hardest V8 bugs , points away from Mythos

DGX agent

Hey everyone ! Quick share from the cyber + local LLM side of things that I found interesting. During this week’s hacker summer camp, an AI researcher and reverse malware engineer veteran 'lordx64' on

model-releasesr-ollama
9 Aug 2026
Model Releases

DeepSeek V4 Flash 0731 hits 82.7% on Terminal-Bench 2.1 in an independent public-harness run (445 trials)

DGX agent

Disclosure: I’m the author of Ante. DeepSeek recently reported an 82.7% score on Terminal-Bench 2.1 for DeepSeek V4 Flash 0731. Its evaluation used “DeepSeek Harness minimal mode,” which hasn’t been r

model-releasesr-localllama
9 Aug 2026
Model Releases

DeepSeek v4 Flash 0731 locally on CPU

DGX agent

After seeing the benchmark results for the full release of DS v4 Flash 0731, I replaced my 2 x 16GB DDR4 ram sticks with 2 x 32GB DDR4 ram sticks to get a max supported of 128 GB RAM, in hope to be ab

model-releasesr-localllama
9 Aug 2026
Model Releases

DeepSeek-V4-Flash-0731 Q8_K_XL sometimes stops mid-task in OpenCode - anyone else seeing this?

DGX agent

Hey everyone, I've been experimenting with the new DeepSeek-V4-Flash-0731 release locally using the Unsloth Studio Q8_K_XL GGUF with OpenCode. Overall, it's been working really well, but I've noticed

model-releasesr-localllama
9 Aug 2026
Model Releases

Doom Loop: Anyone Else Having DeepSeek v4 Flash 0731 Issues on ollama cloud?

DGX agent

Am I the only one having issues with DeepSeek V4 Flash? It gets stuck in a loop, as if it can't call the tools, and keeps repeating the same things endlessly without moving forward. Is it a poorly wri

model-releasesr-ollama
9 Aug 2026
Model Releases

endless-frontier/BigBang-v1 - qwen 3.5 finetunes

DGX agent

table bench https://huggingface.co/bartowski/endless-frontier_BigBang-v1-GGUF I'm downloading this model only because Bartowski converted it to .gguf, so it might be interesting. Doubts : The headline

model-releasesr-localllama
9 Aug 2026
Local Ai

I Turned My Underused Gaming Laptop Into a Local AI Workstation

DGX agent

TL;DR: I am building a Windows-first local AI setup for people who want to try local LLMs without spending days choosing models, setting up Ollama, Docker, WSL, Open WebUI, agents, and tool permission

local-air-ollama
9 Aug 2026
Local Ai

It took two years, but we finally have a 'local Sora'

DGX agent

Who remembers when OpenAI previewed Sora two years ago and the quality felt unreal? We had never seen anything like it. Back then, Sora 1 didn't even generate audio and was heavily censored. Prompt: i

local-air-stablediffusion
9 Aug 2026
← Previous
1234…41
Next →