AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
1,437 results
Model Releases

Tested in Coding: BF16 Muse Glimmer vs BF16 Qwen3.6 27B

DGX agent

I'm guessing that many people have been waiting for this comparison. For clarity, both models are running at full FP16 KV-cache. Due to VRAM limitations, Muse Glimmer is running full 262,144 context,

model-releasesr-localllama
11 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Need real world ML problems to evaluate my educational ML tools

DGX agent

I'm a retired platform engineer, coding mainly in Rust, and involved with a ML study group. I developed a ML programming language (alternative to Python, Colab) to help me learn (and teach) ML concept

model-releasesr-localllama
10 Aug 2026
Local Ai

omlab/VLX-Seek-1.5-10B · Hugging Face

DGX agent

VLX-Seek-1.5-10B VLX-Seek-1.5-10B is the open-source 10B model in the VLX-Seek 1.5 family, designed for fine-grained perception and visual grounding in embodied scenarios. It targets practical setting

local-air-localllama
10 Aug 2026
Model Releases

nvidias nemotron omni only loads its text half on a mac, so i wrote the vision and audio towers in mlx

DGX agent

nvidias nemotron omni is open weights and it sees, hears and reasons. theres already a 4bit mlx quant on hugging face but only the text backbone loads with standard mlx tooling. the model card says it

model-releasesr-localllama
6 Aug 2026
Model Releases

[Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]

DGX agent

First of all, obviously I took some help from AI to type this post and this is the topic that enabled me to accomplish all that: https://old.reddit.com/r/LocalLLaMA/comments/1veow4b/deepseek_v4flash_2

model-releasesr-localllama
4 Aug 2026
Local Ai

Hugging Face CEO says China is winning the AI race and dominating on open models

DGX agent

This is something that was spoken here and there, and now it is like writing on the wall. The main additional point is that China has created an independent supply chain. Starting from raw materials a

local-air-localllama
4 Aug 2026
Model Releases

Is LM Studio abandoning their core product?

DGX agent

Some of you may be aware that a few weeks ago, LM Studio announced a new agent, Bionic. This is pretty much an agentic harness for both local models and paid cloud models. But most aren't aware that L

model-releasesr-localllama
4 Aug 2026
Model Releases

KAT Coder 2.5 dev: Do yourself a favor and try it!

DGX agent

It is so good! I don't know why there aren't more people talking about it. Fewer tokens, faster and more accurate than Qwen 3.6 35b a3b. On my setup it's nearly as good as 27b, but 5x faster. And it c

model-releasesr-localllama
3 Aug 2026
Model Releases

DeepSeek-V4-Flash 284B on 5.3GB of memory

DGX agent

Following up on my Qwen 3.6 port, I wanted to keep adding models and ended up fixing a bunch of things along the way, so it's its own engine now: Mference. Same core idea from TurboFieldfare, MoE mode

model-releasesr-localllama
2 Aug 2026
Model Releases

How well do multiple GPUs scale for LLM inference? (Trying to understand the basics)

DGX agent

Hi everyone, I’m fairly new to the multi-GPU side of local LLMs and I’m trying to understand how inference actually scales across multiple GPUs. Suppose I have a model running on a single GPU and then

model-releasesr-localllama
2 Aug 2026
Model Releases

Real-world reality check on Qwen for autonomous coding agents

DGX agent

TLDR below 👇🏼 I’ve seen a lot of hype around Qwen 3.6 35B and 3.5 120B lately, especially regarding coding and tool-use capabilities. On this subreddit it is the defacto recommended model for everyone

model-releasesr-localllama
2 Aug 2026
Tutorials

There's no 'one weird trick” for prompting Krea 2 art styles—just many guidelines [WF included]

DGX agent

TLDR: There is no one prompting trick that will result in Krea 2 Turbo giving you exactly the style you want and across the whole image. Instead, if you are trying to achieve styles without the use of

tutorialsr-stablediffusion
1 Aug 2026
Model Releases

Benchmarked: MindControl for Llama.cpp

DGX agent

I recently shared the original MindControl PoC (and on github) - sampler-level guided reasoning budgets for llama.cpp, nudging the model with self-aware statements about its own thinking budget instea

model-releasesr-localllama
30 Jul 2026
Model Releases

How Kimi K3 Engineered Its Way to the Frontier [R]

DGX agent

Kimi K3 by Moonshot reached the frontier as an open-weight model. Artificial Analysis ranks it fourth of 580 models, behind only Claude Opus 5, Fable 5, and GPT-5.6 Sol. Moonshot released more than th

model-releasesr-machinelearning
30 Jul 2026
Model Releases

Nanbeige4.2-3B: I'm not impressed

DGX agent

I've tested Nanbeige-4.2-3B. On paper, the benchmarks promise it blows away Qwen3.5-9B and Gemma4-12B. My goal was to have something very light and fast to replace Qwen3.6-35B (or finetunes thereof) f

model-releasesr-localllama
30 Jul 2026
Model Releases

DeepSeek V4 Flash, up to 32 tok/s on AMD Ryzen AI MAX+ 395

DGX agent

Hey fellow llamas. we have something new for Strix Halo owners we thought would be useful to share. i'll keep it short: We were able to fit DeepSeek V4 Flash plus its speculative draft on a single Ryz

model-releasesr-localllama
28 Jul 2026
Local Ai

What 'task oriented' models are folks running on N100 MiniPCs with 16GB of RAM and no GPU?

DGX agent

By 'task oriented', I dont really mean agentic, I mean no deep coding ability, no need for conversation. More things like classification, identification, simple interaction with web apps and APIs, etc

local-air-localllama
28 Jul 2026
Local Ai

Unexpected use of local llm

DGX agent

I was refreshing my youtube and found out my favourite reviewer uploaded a battery test of 78 smartphones: https://youtu.be/MpgUFrsIWSQ the author said they started using robotic arm to simulate a per

local-air-localllama
27 Jul 2026
Model Releases

We could really use Qwen3.8 in 27B, 35B, 122B and 397B sizes

DGX agent

Instead of 2T+ models, continuing to release highly capable small to medium size LLMs would really help to keep this community vibrant. Hardly anyone can even dream of running the recent 1.5-2T+ beast

model-releasesr-localllama
27 Jul 2026
Model Releases

You can now fine-tune my 3.96M-parameter TTS on your own voice or language

DGX agent

When I released Inflect v2 last week, I thought most people would ask whether a TTS model this small actually sounded decent. Instead, I kept getting two questions: “Can I train it on my own voice?” “

model-releasesr-localllama
27 Jul 2026
Local Ai

I implemented the YOLO26n model inference from scratch using ARM64 Assembly Language (No framework) [P]

DGX agent

This was my Bachelor's Final Project: implementing YOLO26n inference completely from scratch using ARM64 Assembly Language and C, without relying on existing inference frameworks. The goal was to unde

local-air-machinelearning
26 Jul 2026
Model Releases

Nvidia releases Qwen-Image-Flash

DGX agent

'The NVIDIA Qwen-Image-Flash model generates images from text prompts using a four-step, DMD2-distilled version of Qwen/Qwen-Image. The distillation used DMD2 from NVIDIA FastGen, NVIDIA Model Optimiz

model-releasesr-stablediffusion
24 Jul 2026
Local Ai

Absurd claim: the distilled model outperforms the originals

DGX agent

As an AI community of LLM experts, are we really going to stay silent while US officials make absurd claims to push anti-consumer laws? Not only does the release timeline between Fable and K3 make hig

local-air-localllama
23 Jul 2026
Model Releases

microsoft/Fara1.5-27B · Hugging Face

DGX agent

Fara1.5-27B is a multimodal computer use agent (CUA) for web browsers, from Microsoft Research AI Frontiers. It observes the browser through screenshots and acts on the user's behalf by emitting struc

model-releasesr-localllama
22 Jul 2026
Local Ai

OpenCodeRAG - RAG for OpenCode via locally hosted models

DGX agent

OpenCodeRAG is a local embedding service using FastAPI and SentenceTransformer, paired with a Node.js plugin that integrates RAG tools with Qdrant vector database via YAML configuration. It provides a

local-air-ollama
21 Jun 2026
Local Ai

PixlStash 1.3: grid loading speed, JoyCaption and bulk tag selections with your chosen model

DGX agent

PixlStash 1.3 is a Python-based image management and tagging web app that improves grid loading performance and introduces JoyCaption integration for AI-powered image captioning. The update adds suppo

local-air-stablediffusion
25 May 2026
Local Ai

I Built a local Stable Diffusion GUI specifically for older GPUs (GTX 1060). Features Zero-Copy ADetailer, URL Model Downloader, and real-time VRAM monitoring.

DGX agent

This post describes a custom Stable Diffusion graphical interface optimized for older graphics cards, specifically the GTX 1060, incorporating features like zero-copy ADetailer (a detail enhancement t

local-air-stablediffusion
23 May 2026
Local Ai

Built a Chrome extension that talks directly with your local Ollama models

DGX agent

A Chrome extension that communicates directly with local Ollama instances without sending data to external servers , enabling users to submit messages through a popup interface with responses streamin

local-air-ollama
13 May 2026
Local Ai

Apple Removes 256GB M3 Ultra Mac Studio Model From Online Store

DGX agent

Apple has removed its 256GB M3 Ultra Mac Studio from sale, limiting the machine to 96GB of unified memory , following the removal of the 512GB configuration in March . The removal is likely due to a g

local-air-localllama
9 May 2026
Local Ai

Thoth v3.21.0 - Buddy Companion, Model Picker Improvements, and Stronger Linux Startup

DGX agent

Thoth v3.21.0 is a local-first AI assistant with integrated tools including a personal knowledge graph, voice, vision, shell, and browser automation capabilities. This release introduces Buddy Compani

local-air-ollama
7 May 2026
Model Releases

Claude Desktop 3P Gateway

DGX agent

Claude Desktop has a third-party inference feature that lets you replace Anthropic's API with any model provider, including a local AI model running entirely on your machine. This feature can be activ

model-releasesr-ollama
6 May 2026
Model Releases

GPT-5.5 Instant is starting to roll out in ChatGPT.

DGX agent

OpenAI is rolling out GPT-5.5 Instant to all ChatGPT users as the new default model, replacing GPT-5.3 Instant . The model produces 52.5% fewer hallucinated claims than its predecessor on high-stakes

model-releasesr-chatgpt
5 May 2026
Model Releases

GitHub: ComfyUI SenseNova U1 Released – Anyone Got It Working Yet for ComfyUI?

DGX agent

SenseNova U1 is a new series of native multimodal models that unifies multimodal understanding, reasoning, and generation within a monolithic architecture, marking a fundamental paradigm shift from mo

model-releasesr-stablediffusion
4 May 2026
Local Ai

Most accurate Al model for generating videos from images while preserving text?

DGX agent

This post likely discusses the challenge of generating videos from images while preserving text, as text and fine details in AI-generated videos often appear garbled or distorted. Based on the subredd

local-air-stablediffusion
2 May 2026
Model Releases

Comparison of low Steps, Klein 9b x Z image turbo x Ernie Turbo x Qwen 2512 8 Steps

DGX agent

This r/StableDiffusion post presents a community-driven visual comparison of several modern, fast text-to-image diffusion models — Flux.2 Klein 9B (a smaller, faster distillation of Flux.2 Dev availab

model-releasesr-stablediffusion
15 Apr 2026
Model Releases

Is Gemma 4 26B MoE or 31B good as an MCP agent for coding with Xcode?

DGX agent

This r/ollama discussion explores the suitability of Google's Gemma 4 models — specifically the 26B Mixture of Experts (MoE) and 31B Dense variants — as MCP (Model Context Protocol) agents for coding

model-releasesr-ollama
15 Apr 2026
Research

ClawBench: Can AI Agents Complete Everyday Online Tasks? 153 tasks, 144 live websites, best model at 33.3% [R]

DGX agent

ClawBench is a benchmark of 153 everyday web tasks spanning 144 live platforms across 15 categories — from completing purchases and booking appointments to submitting job applications. Unlike existing

researchr-machinelearning
14 Apr 2026
Local Ai

Flux 2 Klein 9B produces absolutely awful and ugly skin textures

DGX agent

This r/StableDiffusion post discusses a widely noted quality issue with the FLUX.2 Klein 9B model, where users report that it produces poor skin textures in human portraits — the base model has a majo

local-air-stablediffusion
14 Apr 2026
Local Ai

Codex with Voiden

DGX agent

'Voiden' doesn't appear in any search results as a known model or tool in the Ollama ecosystem. Based on the Reddit source and the broader context of the r/ollama community, this post likely discusses

local-air-ollama
13 Apr 2026
Local Ai

Got early access to a real-time interactive video model, here's what I found

DGX agent

I was unable to retrieve the specific Reddit post at the provided URL through my search. The post (reddit.com/r/StableDiffusion/comments/1shxmfk) did not surface in the search results, and I cannot...

local-air-stablediffusion
10 Apr 2026
Local Ai

LiquidAI/LFM2.5-VL-3B · Hugging Face

DGX agent

LFM2.5-VL-3B is a multimodal variant of LFM2.5, a family of hybrid models designed for on-device deployment. It builds on LFM2-VL-3B with further mid- and post-training. LFM2.5-VL-3B can process both

local-air-localllama
12 Aug 2026
Model Releases

Tested Nemotron 3.5 Lightning locally on coding, Hermes Agent and agentic work

DGX agent

Ran the model with quants (Q5) and MTP by bartowski with llama.cpp server. It takes ~24GB ram running on M5 Pro with 48GB at about 65t/s. On some tasks it was quite the overthinker. Overall, the quali

model-releasesr-localllama
12 Aug 2026
Model Releases

Achievable 253 t/s - unsloth/Muse Glimmer 30B UD-Q5_K_M on a 5090

DGX agent

Benchmarked Muse Glimmer 30B on my RTX 5090 (32GB), 262k context, UD-Q5_K_M + dflash-kquant + mmproj. Workload Stock master + DFlash ngram-simple PR #26842 + DFlash Code patch 78 t/s 57 t/s 220-253 t/

model-releasesr-localllama
10 Aug 2026
Model Releases

CyberKimi just dropped strong results on one of ExploitBench’s hardest V8 bugs , points away from Mythos

DGX agent

Hey everyone ! Quick share from the cyber + local LLM side of things that I found interesting. During this week’s hacker summer camp, an AI researcher and reverse malware engineer veteran 'lordx64' on

model-releasesr-ollama
9 Aug 2026
Model Releases

KLQ: Training-free measured rotation quantization. Beats all training-free rotation-based quantization methods on W4A4KV4-bits. Llama 3.2 1B KLQ-quantized beats SpinQuant and gets close to ReSpinQuant without GPTQ/LDLQ rounding.

DGX agent

First of all, I'm not a lab, this was a solo summer research project that finally culminated into the github repo and the writeup. The repo includes a much deeper dive with methods, findings about qua

model-releasesr-localllama
9 Aug 2026
Model Releases

~45% lower MiniMax H3 sampler time with new Spectrum settings — degree 1 works surprisingly well (v0.1.8)

DGX agent

Follow-up to my original Spectrum MiniMax H3 post: https://www.reddit.com/r/StableDiffusion/comments/1vf1ze3/spectrum_acceleration_for_minimax_h3_in_comfyui/ In that first post, I released the MiniMax

model-releasesr-stablediffusion
7 Aug 2026
Model Releases

Am I just hallucinating

DGX agent

Or is there any reason why I feel like model output quality seems to be better when I use higher micro-batch values (ub) in llama-cpp? I don't really have any hard numbers or anything (just running th

model-releasesr-localllama
7 Aug 2026
Model Releases

Serving Deepseek v4 Flash 0731 on 2x DGX Spark — 5-7 GB OS headroom, what would you do to lower VRAM usage and increase OS available RAM?

DGX agent

Hey all, I'm serving DSv4Flash 0731 on a cluster of 2x DGX Sparks but am running into constant issues with having almost no RAM (unified memory) left for the OS/cache and I'd love to hear the communit

model-releasesr-localllama
7 Aug 2026
← Previous
1…678910…30
Next →