AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
1,437 results
Model Releases

Tesla V100 Qwen3.6 27B Performance

DGX agent

Looking for V100 users to share your config and it's performance. GPU: Tesla V100 PCIE 32Gb Qwen3.6 27B Q4_K_M + Q8_0 MTP 128K context length Pi coding agent llama.cpp model preset: [*] spec-default =

model-releasesr-localllama
8 Aug 2026
Model Releases

Anyone running DeepSeek-V4-Flash-0731 on MI325X with vLLM? Mine is behaving completely broken

AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

Is anyone here successfully running DeepSeek-V4-Flash-0731 locally with vLLM, especially on AMD MI325X? My setup: GPU: 1x AMD Instinct MI325X Model: deepseek-ai/DeepSeek-V4-Flash-0731 vLLM: 0.26.0 ROC

model-releasesr-localllama
7 Aug 2026
Model Releases

How come artificialanalysis.ai ranks Gemma4 above Qwen3.6 27b in SciCode

DGX agent

Just came across this coding benchmark: SciCode Artificialanalysis.ai reports a ranking which contradicts the feeling we've towards those models in real life coding. Is Gemma 4 really that good, or a

model-releasesr-localllama
6 Aug 2026
Model Releases

DeepSeek v4 Flash vs. Qwen3.6-27B, 3.5-122B, and Gemma 4 31B Benchmark

DGX agent

Just wanted to share my agentic coding benchmark run of DSv4F 0731 at both High and Low reasoning efforts (not Max)... I ran a 109-question subset of Aider Polyglot (the JS/C++/Python languages), base

model-releasesr-localllama
4 Aug 2026
Model Releases

GPT-OSS has turned one year old today!

DGX agent

It is one of the best local models ever released, in both 20B and 120B versions. I always come back to it, especially the 120B version. Its only competition is, in my opinion, Qwen 3.5 122B, but that

model-releasesr-localllama
4 Aug 2026
Model Releases

AI9Stars released G9v3-39A5B

DGX agent

AI9Stars has released G9v3-39A5B an open weights language model designed to deliver even stronger reasoning capabilities than ai9stars/G9v3-3B with its 39B and 5 active experts. It is released under t

model-releasesr-localllama
3 Aug 2026
Model Releases

Döner Bench DeepSeek-V4-Flash IQ2_XS running on a single RTX 3090

DGX agent

https://preview.redd.it/3zcvpbds14hh1.png?width=1911&format=png&auto=webp&s=a79aafb71eeca97638da93d2591902631e897fd5 I tried a test similar to the recent model-quant comparisons, but this time I focus

model-releasesr-localllama
3 Aug 2026
Model Releases

Question about Quant versus Size.

DGX agent

Sorry if this is asked a lot, but I was wondering if there is any clear winner on the Quantization versus Model Size debate? I can run Qwen3.6 27b at Q8, Laguna at Q6, and the new Deepseek Flash at Q3

model-releasesr-localllama
3 Aug 2026
Model Releases

The Chinese labs everyone lumps together are making four pretty different bets. I work at one of them.

DGX agent

Every time a model drops from a Chinese lab the thread fills with people who already know who made it, and the guess is usually Alibaba. There was a thread here recently asking what separates the open

model-releasesr-localllama
3 Aug 2026
Local Ai

Making a synthetic dataset for fine-tuning

DGX agent

I've been thinking about building a pipeline to generate reasoning training data for LLMs, but I want to avoid the common failure mode of synthetic data where you just generate the same template with

local-air-localllama
30 Jul 2026
Model Releases

Mechanistic interpretability streamlined for everyday users like us😎 🧠

DGX agent

Context: I want to give the community an Open Research (well open under Apache 2.0 clause) - tool that allows everyday users like us to look deeper into the local models we use consistently. Mechanist

model-releasesr-localllama
30 Jul 2026
Model Releases

I built a tool to actually test which weights matter before quantizing, instead of guessing (Qwen3.6-27B, 3 builds: Bedrock/Tightrope/Gambit)

DGX agent

Most quantization works like this: pick a bit depth, apply it everywhere, maybe let imatrix take a rough guess at what matters, ship it. Most don't check which specific weight groups can take a hit an

model-releasesr-localllama
28 Jul 2026
Model Releases

Ollama Cloud Quota Benchmark

DGX agent

Recently I bought an Ollama Cloud sub and accidently spent my whole 5h quota upon using DeepSeek V4 Pro... but why? isnt it supposed to be a cheap model? Youd think there would be a correlation betwee

model-releasesr-ollama
25 Jul 2026
Model Releases

Kwaipilot/KAT-Coder-V2.5-Dev · Hugging Face

DGX agent

from kwaipilot: Following the release of KAT-Coder-V2.5 in July, we are pleased to release the open-weight version KAT-Coder-V2.5-Dev, an MOE model with a total parameter count of 35B and 3B activated

model-releasesr-localllama
23 Jul 2026
Model Releases

Running Qwen 3.6 35B MoE (Q4_K_M) on a Zeus (Xiaomi 12 Pro, 12GB RAM)

DGX agent

Shoutout to this awesome guy - https://www.reddit.com/r/LLM/s/IDUyU3v9ap Thanks to his project, BigMoeOnEdge https://github.com/Helldez/BigMoeOnEdge, I managed to successfully run a 35B MoE model on j

model-releasesr-localllama
23 Jul 2026
Model Releases

Running Gemma 4 QAT 12B on an 8GB GPU at 16k context — measured the KV-cache tradeoffs

DGX agent

This post discusses running Google's Gemma 4 QAT (Quantized Aware Training) 12B model on a GPU with 8GB of memory while maintaining a 16k token context window. The author likely shares performance ben

model-releasesr-ollama
11 Jun 2026
Model Releases

Mistral-7B v0.3 at 128K in llama.cpp: 22,657 → 13,235 MiB live VRAM with ≤0.004 PPL drift

DGX agent

Mistral-7B v0.3 model achieves significant memory optimization when running at 128K context length in llama.cpp, reducing live VRAM usage from 22,657 MiB to 13,235 MiB while maintaining minimal perfor

model-releasesr-ollama
26 May 2026
Local Ai

img2vid: ComfyUI doesn't find the spatial upscaler

DGX agent

This Reddit post discusses a common ComfyUI issue where users cannot locate spatial upscaler models for img2vid workflows. The problem typically stems from placing spatial upscaler models in the wrong

local-air-stablediffusion
22 May 2026
Local Ai

Announcing the release of Stable Audio 3!

DGX agent

Stability AI announced the launch of Stable Audio 3, a family of three AI music models and one audio-based special effects model. Most of these releases are 'open weight' models trained on licensed tr

local-air-stablediffusion
20 May 2026
Local Ai

Ring-2.6-1T Open sourced today! Soooo looking forward to trying it on Ollama!

DGX agent

Ring-2.6-1T is a trillion-parameter flagship reasoning model designed for real-world complex task scenarios, now available as an open-source model. The model features about 63B activated parameters pe

local-air-ollama
14 May 2026
Model Releases

A compilation of the open-source LoRAs for LTX 2.3 - released in May

DGX agent

LTX-2.3 is an open-source video generation model released in January 2026 that supports LoRA fine-tuning for customizing styles, characters, and use cases. The Reddit post compiles available LTX-2.3 m

model-releasesr-stablediffusion
13 May 2026
Local Ai

kimi k2.6 is now 'available' on ollama cloud

DGX agent

Kimi K2.6 is an open-source model featuring advanced coding, long-horizon execution, and agent swarm capabilities that is now available via Ollama Cloud . The model excels in coding and agentic tools

local-air-ollama
20 Apr 2026
Model Releases

Was happy with Gemma 4 Cloud, but had to change due to API Errors, GLM 5.1 spends a lot more ressources

DGX agent

A user reported satisfaction with Gemma 4 Cloud but switched to GLM 5.1 due to API errors, noting that the alternative model consumes significantly more resources. The post likely discusses the perfor

model-releasesr-ollama
20 Apr 2026
Model Releases

Ernie Image Turbo is not bad at all (Using INT8 quant and Gemini for prompt enhancement, RTX 30 series GPU with low vram)

DGX agent

Ernie Image Turbo is a text-to-image generation model that can run efficiently on consumer-grade hardware like RTX 30 series GPUs with limited VRAM by using INT8 quantization. The post discusses techn

model-releasesr-stablediffusion
17 Apr 2026
Local Ai

gguf import with vision / mmproj, not working!

DGX agent

This r/ollama thread addresses a known compatibility issue where importing multimodal GGUF models into Ollama using a separate mmproj (multimodal projector) file fails to enable vision capabilities. O

local-air-ollama
15 Apr 2026
Hardware

GPU stays sometimes at 100% usage even when done replying. Is it normal?

DGX agent

This r/ollama post addresses a commonly reported behavior where Ollama's GPU usage remains at or near 100% even after a model has finished generating a response. Certain models appear to 'hang' after

hardwarer-ollama
15 Apr 2026
Local Ai

ZIB ZIT hand over?

DGX agent

This r/StableDiffusion Reddit thread likely discusses the topic of hand generation quality when using ZIB and ZIT — two AI image generation models from the Z-Image ecosystem used in Stable Diffusion a

local-air-stablediffusion
15 Apr 2026
Model Releases

ERNIE Image released

DGX agent

ERNIE Image is an open-source text-to-image generation model developed by Baidu, built on a single-stream Diffusion Transformer (DiT) paired with a lightweight Prompt Enhancer that expands brief user

model-releasesr-stablediffusion
14 Apr 2026
Model Releases

I benchmarked Gemma4:e4b vs Gemma3:27B vs GPT-4o-mini vs Gemini 2.5 Flash on a Mac Mini M4 Pro 24gb — full results

DGX agent

A Reddit user on r/ollama conducted a hands-on benchmark comparing Gemma4:e4b (Google's compact ~4.5B effective-parameter edge model) against Gemma3:27B, GPT-4o-mini, and Gemini 2.5 Flash, all run or

model-releasesr-ollama
13 Apr 2026
Model Releases

Benchmark Your Local LLMs in 3 Commands

DGX agent

This r/ollama post describes a streamlined method for performance-testing locally-running large language models using the Ollama framework, achievable with just three terminal commands. It likely intr

model-releasesr-ollama
12 Apr 2026
Model Releases

I built a VS Code extension that cuts my Claude API bill to ~$5/day

DGX agent

A developer shared on r/ollama how they built a custom VS Code extension that routes Claude API requests through locally-run models via Ollama, dramatically reducing cloud API costs to approximately $

model-releasesr-ollama
12 Apr 2026
Local Ai

Hallucination problem

DGX agent

A Reddit thread in r/ollama where a user reports experiencing AI hallucination issues when running local language models through Ollama. The discussion likely covers symptoms such as models generating

local-air-ollama
11 Apr 2026
Model Releases

“Maybe I should try Chat again after only using Claude for a while”. First response:

DGX agent

I was unable to retrieve the specific Reddit thread content from that URL through my search. Reddit threads often require direct access to load user-generated content, and the search did not return...

model-releasesr-chatgpt
11 Apr 2026
Local Ai

OpenClaw + Ollama + gemma4:26b is fast in raw Ollama, but first heavy OpenClaw turns are extremely slow or hit idle timeout

DGX agent

Users running OpenClaw with `gemma4:26b` via Ollama encounter significantly slow or timed-out first turns in a session, even though the model responds quickly when queried directly through raw Olla...

local-air-ollama
11 Apr 2026
Local Ai

Cannot search pdf document using WebUI and Ollama

DGX agent

Users in the Ollama/Open WebUI community commonly report being unable to query PDF documents via the Open WebUI interface when using a locally hosted Ollama backend, with the model failing to recog...

local-air-ollama
10 Apr 2026
Local Ai

Is the ASUS ROG Flow Z13 with 128GB of Unified Memory (AMD Strix Halo) a good option to run large LLMs (70B+)?

DGX agent

The ASUS ROG Flow Z13 (2025) with AMD Ryzen AI Max+ 395 (Strix Halo) and 128GB of unified LPDDR5X memory is a capable portable option for running large LLMs locally, with ASUS officially stating it...

local-air-ollama
10 Apr 2026
Local Ai

LTX 2.3 - Image + Audio + Video ControlNet (IC-LoRA) to Video

DGX agent

LTX-2.3 is a DiT-based audio-video foundation model from Lightricks that generates synchronized video and audio within a single model pass, representing a significant upgrade over LTX-2 with improv...

local-air-stablediffusion
10 Apr 2026
Local Ai

Mac mini M4 48GB

DGX agent

The Mac Mini M4 Pro with 48GB unified memory is a popular choice in the local AI community for running large language models via Ollama, as its Apple Silicon architecture makes all 48GB of RAM dire...

local-air-ollama
10 Apr 2026
Local Ai

What Your Local LLM Actually Sees: Debugging Ollama Traffic in Quarkus with mitmproxy

DGX agent

This tutorial demonstrates how to use mitmproxy to inspect the actual HTTP traffic sent from a Quarkus application to a local Ollama model via its OpenAI-compatible endpoint, revealing the real JSO...

local-air-ollama
10 Apr 2026
Local Ai

Tried running LLMs locally to save API costs… ended up waiting 13 minutes for ONE response 🤡

DGX agent

A Reddit post in r/ollama describes a user's experience attempting to run LLMs locally via Ollama to avoid cloud API costs, only to encounter severely degraded performance — waiting 13 minutes for ...

local-air-ollama
9 Apr 2026
Model Releases

DeepSeek V4 Flash 0731 is the ‘killer app’ that is going to sell A LOT of DGX Sparks

DGX agent

Having a ‘Killer Application’ that everyone wants to use helps sell hardware, plain and simple. DeepSeek V4 Flash 0731 isn’t an app of course, but I think it’s going to be the major catalyst for getti

model-releasesr-localllama
10 Aug 2026
Model Releases

inclusionAI/Ling-3.0-tiny · 8B A1.3B MoE· Hugging Face

DGX agent

Looks like the Ling team open weighted a much smaller version of the Ling-3.0-flash they open weighted a few days ago. It's 8B params with 1.3B active, and seems to fall between the 4B and 8-12B Qwen

model-releasesr-localllama
10 Aug 2026
Model Releases

DeepSeek-V4-Flash-0731 Q8_K_XL sometimes stops mid-task in OpenCode - anyone else seeing this?

DGX agent

Hey everyone, I've been experimenting with the new DeepSeek-V4-Flash-0731 release locally using the Unsloth Studio Q8_K_XL GGUF with OpenCode. Overall, it's been working really well, but I've noticed

model-releasesr-localllama
9 Aug 2026
Model Releases

enabling PCI-E p2p for consumer Nvidia cards will yield you more than you think

DGX agent

Disclaimer - no LLM was used to write this post/note As larger post about my setup will come later, want to give heads-up to folks who use VLLM and >= 2 GPUs. So I have pretty meaty server (8 channel

model-releasesr-localllama
8 Aug 2026
Model Releases

Qwen3.6 27B + 35B on vLLM, single R9700 (gfx1201)

DGX agent

I've been tuning my new Radeon AI Pro R9700, and figured that this would be useful information for people who are trying to optimise their setups. I'm pretty happy with these results and looking forwa

model-releasesr-localllama
8 Aug 2026
Model Releases

I ported vLLM's serving stack to C++20: 66 MiB binary, no Python at inference, output checked token-for-token against vLLM

DGX agent

I'm the author, so discount the enthusiasm accordingly. This is an unaffiliated community port, not endorsed by the vLLM project, which it uses to verify its correctness. What started it: I love vLLM,

model-releasesr-localllama
6 Aug 2026
Model Releases

Inkling-Small 276B-A12B at ~2.9 tok/s on <10gb memory

DGX agent

A follow up to the launch of Mference, it now supports and runs Inkling-Small 276B-A12B. Inkling-Small (Thinking Machines, Apache 2.0), from the pipenetwork/Inkling-Small-MLX-4bit conversion: 276B tot

model-releasesr-localllama
5 Aug 2026
Model Releases

MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp

DGX agent

TensorSharp's MoE CPU-offload feature has been merged into main. Here is the parameters description of this feature: Mixture-of-Experts CPU offload: --n-cpu-moe <N> | -ncmoe <N> Keep the routed MoE ex

model-releasesr-localllama
5 Aug 2026
← Previous
1…89101112…30
Next →