AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
Human
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
1,931 results
6 Aug 2026

nvidias nemotron omni only loads its text half on a mac, so i wrote the vision and audio towers in mlx

Model ReleasesDGX agent

nvidias nemotron omni is open weights and it sees, hears and reasons. theres already a 4bit mlx quant on hugging face but only the text backbone loads with standard mlx tooling. the model card says it

🟩 NVIDIA's whole speech stack just went local. ASR + TTS + codec, quantized to GGUF, running on-device via NeMo-Speech.cpp

Model ReleasesDGX agent

🐦‍⬛ Magpie-TTS Multilingual 🦜 Nemotron Speech Streaming EN 0.6B 🦜 Nemotron-3.5 ASR Streaming 🦜 Parakeet CTC 1.1B 🦜 Parakeet TDT 0.6B v3 🥦 NanoCodec Merged PR https://huggingface.co/nvidia/magpie_tts_m

Scotoma-2: Gemma4, but with less annoying slop and better writing.

Model Releases
DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

GGUFs here: https://huggingface.co/ReadyArt/gemma-4-31B-it-scotoma-2-GGUF Disclaimer: By slop, we are specifically talking about specific tics with the model(sentence structures), but this doesn't inc

The death of SLMs?

Model ReleasesDGX agent

I love to see these impressive models coming out that compete with the giants from companies like Z.ai, Moonshot, Alibaba, etc. A win for the open source/weight community is always welcome. While I am

They almost catched up on Frontier performance, so now catching up on prices

Model ReleasesDGX agent

Users also report that the free version was significantly downgraded after the release of the new models this is very important for us when considering local hosting. A lot of people decided not to bu

Unsloth's Gemma 4 mmproj silently broke vision & audio on newer llama.cpp builds — anyone else hit this?

Model ReleasesDGX agent

So I had been building ScreenMind, kinda like local ai desktop assistant that uses Gemma 4 for screen analysis, voice memo transcription, and meeting transcription — all through llama-server. Everythi

5 Aug 2026

40% speedup of MoE training with faster megakernel, by cursor, of all people (for B200s)

Local AiDGX agent

daily reminder not to trust benchmarks and run it yourself. claimed e2e speedup is ~40%, forwards are ~140% faster I would wager that compared to a naive kernel anyone can write it's more in the range

A ultra-lightweight mini agent - zero framework and local/ollama first

Model ReleasesDGX agent

https://github.com/mohsinkaleem/agent-mini.git A minimal 3k lines, local-first AI agent you can actually understand and extend. Optimized for smaller local models like qwen 3.6 4b or 9b pip install ag

Agent memory layers don't need an LLM deciding what to remember

Local AiDGX agent

Most agent memory setups run a model call on the way in. Something reads the turn, decides whether it's worth keeping, rewrites it into a 'memory', tags it with a type and an importance score. That's

Anyone interested in building a harness-only benchmark?

Model ReleasesDGX agent

There are a lot of LLM benchmarks but few, if any, harness benchmarks. I am thinking this would be a really good community project to build one. End goal: a leaderboard of harness performance (multipl

Best experience with MCP UE 5.8

Local AiDGX agent

Hey. I'm very limited with brain capacity and I'm new to this. If you have experience, tell me which local model fits the best, what makes best blueprints etc. For example I want to spawn enemies usin

Bro, I need you to FOCUS.

Model ReleasesDGX agent

Come on guys, we could get deepseek cheaper with more usage directly with them if we wanted. WHERE IS KIMI K3 WITHOUT EXTRA USAGE NEEDED !!? This is pissing me off like crazy submitted by /u/Other_Che

Building a Fully Local PDF Read-Aloud & PDF-to-Audiobook Desktop App with Kokoro 82M, Qwen, and llama.cpp

Model ReleasesDGX agent

Hey everyone, I’ve been building Speechfony - a desktop app for reading PDFs (and EPUBs) with offline text-to-speech. Open a document, listen sentence-by-sentence with highlighting, or export selected

DeepSeek V4 Flash 0731 at 10–17 t/s (nothink) on MacBook M5 Pro **64GB***, partly via SSD streaming

Model ReleasesDGX agent

Inspired by a post from u/giveen I motivated claude (no patinence on my side to work through everything myself) to help me get DS running on my MacBook M5 Pro 64GB and it exceeded my expectations.. be

Deepseek V4 Flash just hit Colibri, does anyone have numbers?

Model ReleasesDGX agent

I'm mosty interested in 128-192GB VRAM with 128-256GB RAM to spare, so SSD streaming is basically not even necessary. Seems only FP4 is supported, so older hardware will likely be slow - no Unsloth GG

DeepSeek-V4-Flash on SM89 4x48gb 4090s with DSpark

Model ReleasesDGX agent

https://github.com/yhfgyyf/vllm-deepseek-v4-sm89 I couldn't believe that someone actually got vLLM working with this particular set of GPUs, but here it is. The video is from right after I got it work

How do I run Ollama on Termux?

Local AiDGX agent

Aight, I've given myself a challenge. I want to run Ollama on termux, Seen plenty of tutorials but I'm still confused. Plus, this phone is relatively old, so it might not work. Last time it failed so

I built xSignalBot: an auto-reply bot for Signal that answers with a local LLM via Ollama — zero cloud, zero cost

Local AiDGX agent

Disclaimer: I'm the developer of this project — sharing because it might be useful to others running self-hosted AI. (Full transparency, as Reddit's self-promo etiquette expects.) xSignalBot is an ope

I remember a time when 'flash' meant 32B

Model ReleasesDGX agent

I mean, Deepseek V4 Flash is an absolutely fantastic model, even though I can't run it on my machine it's so fascinating to see how it performs. Knowing that potentially it could be run at home is rea

I took a local OCR model's accuracy from 60% to 99%

Local AiDGX agent

I built a local OCR pipeline a few days ago, and it turned into a surprisingly interesting experiment—taking accuracy from around 60% to 99%. I wrote a short blog about what worked, what failed, and t

I updated my localy run benchmark with DeepSeek V4 Flash 0731

Model ReleasesDGX agent

It's the purple cluster on the top left (the good corner...) I'm running the MXFP4 version from Bartoswski with Dspark at 1K t/s prefill and 90 t/s gen (average). I tried different sampling params, yo

Inkling-Small 276B-A12B at ~2.9 tok/s on <10gb memory

Model ReleasesDGX agent

A follow up to the launch of Mference, it now supports and runs Inkling-Small 276B-A12B. Inkling-Small (Thinking Machines, Apache 2.0), from the pipenetwork/Inkling-Small-MLX-4bit conversion: 276B tot

Intern S2 Mobius

Local AiDGX agent

A Qwen3.5-35B derived model with an interesting architectural difference that results in larger throughput and less token consumption (allegedly): https://huggingface.co/internlm/Intern-S2-Mobius subm

LFM2.5-2.6B on a OnePlus 13 at 17 tok/s ~ Pure CPU

Model ReleasesDGX agent

As you all know the model is 2.69B parameters with a 128K context window and purpose-built for multi-step agent workflows. What you are seeing is the Q4_K_M GGUF running on my own inference engine bui

Millions of people finally have a friend

Local AiDGX agent

Because of ollama millions of people finally have a friend that can keep up with them and at the very least tolerate their word salad. (Not to mention that people are writing and reading more than eve

MiniMax are issuing takedowns on decensor/explicit H3 LoRAs

Local AiDGX agent

I saw that someone had uploaded an experimental decensor LoRA for H3 on HuggingFace earlier today. Shortly afterwards, in the discussions, a MiniMax employee issued a warning that if it was not remove

MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp

Model ReleasesDGX agent

TensorSharp's MoE CPU-offload feature has been merged into main. Here is the parameters description of this feature: Mixture-of-Experts CPU offload: --n-cpu-moe <N> | -ncmoe <N> Keep the routed MoE ex

Ollama on mac mini/studio?

Local AiDGX agent

I don't own a mac Right now, i want to buy one but its main job will be to host a Ollama or similar software to give me usable AI models in my network, since I don't want to pay for Cloude, GitHub cop

Prime Agent - a new coding harness surpassing Codex/CC/PI

Model ReleasesDGX agent

Prime Agent is an open-source coding and research agent for general and long-running work. A self-improving RLM harness for coding and long-running autonomous tasks. Designed to be both token-efficien

Projects created using OLLAMA

Local AiDGX agent

Hi! I'm a new user of Ollama here. What are some GitHub projects that I can look at and play with to see how people work with the Ollama library across various programming languages? I wouldn't want t

PSA Update CUDA from 13.2 to 13.3 to solve DeepSeek V4 Flash 0731 Looping Problem!

Model ReleasesDGX agent

So one of yall mentioned that cuda 13.1 or 13.2 is broken for unsloth so I looked in to it, and they were right. I had 13.2 installed, after I switched to 13.3 no more looping!!! Before the cuda updat

Qwen Developers' responses from their recent Twitter/X AMA

Model ReleasesDGX agent

Questions & Responses(in BOLD) below. Favorite question(s) moved to end of the thread with combined responses(removed duplicates). Be optimistic folks. I'm sure we're getting other models too apart fr

Qwen3-TTS voice cloning is now in mainline llama.cpp — the old demo finally became real support

Model ReleasesDGX agent

People may remember the Qwen3-TTS llama.cpp demo from a few months ago. That PR said it probably wouldn’t be merged because llama.cpp was missing some of the graph and API pieces it needed. A new impl

RTX 3090 MiniMax H3 Speed Comparison: FP8 Scaled vs INT8 ConvRot (W8A8)

HardwareDGX agent

Setup: GPU: RTX 3090 24GB RAM: 32GB ComfyUI 0.30.0 PyTorch 2.13.0+cu130 CUDA 13.0 SageAttention enabled (sageattention-2.2.0+cu130torch2.10.0andhigher.post6-cp310-abi3-win_amd64) Spectrum node + Euler

Scenema Audio Comes to ComfyUI, Runs on 8GB VRAM

Model ReleasesDGX agent

Hey everyone! Scenema Audio is now a native ComfyUI custom node. Same model that powers scenema.ai now quantized so it fits on 8GB VRAM. When we first released it a few months ago as an API and Docker

Stable Diffusion might actually be remembered in the history books, and I don’t think that’s an overstatement

Model ReleasesDGX agent

Hear me out before you roll your eyes. We tend to only recognize turning points in hindsight. Nobody in 1993 thought the Mosaic browser would be a history book moment, but the web is. I think Stable D

Sulphur 3 is looking for funding

Model ReleasesDGX agent

Hello, I'm the guy who made Sulphur 2. With the recent release of a certain video model, we are looking to mobilize and train Sulphur 3 on this new model. We are targeting $10,000 USD. This certain ne

The Google Bug Hunters Team admitted to me that they cannot fundamentally patch prompt engineering bypasses in Gemini

Model ReleasesDGX agent

Hello everyone Yesterday I gave a report on Gemini bugs and the techniques I learned on Gemini so far with the Engineering Prompt and interestingly today I got a very interesting and controversial ans

Thinking of buying more DRAM right now...

Model ReleasesDGX agent

So I'm looking at https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF and I realize my 128GB of DRAM just isn't cutting it for this (incredibly powerful) model. If only I had another 64GB, I th

Unexpected Crash-Looping of Multi-GPU Box: Caught by Hand-Scrutinized Logs - A Tale of Overridden Keep-Alive Policy and Eviction Thrashing

Local AiDGX agent

I've been freelancing for over a decade now, and I can't stress enough the importance of thorough investigation when dealing with strange software behaviors. Recently, I ran into an issue where my mul

Utilize a nvidia gpu and amd gpu together for 2 different ai models?

Model ReleasesDGX agent

We run a local model instance in our company that the dev we hired built for us. We're a trade business and we want to further use our on hand hardware for it. The specs given we have is a 5090 gpu wi

Watch a local Ollama's qwen3:8b turn one English question into a 9-node investigation graph - planned, admitted by a deterministic gate, and run live in the browser (open source, MIT)

Model ReleasesDGX agent

The video is one real run, not a mock-up: grapharc go 'why did checkout latency spike at 09:14 UTC?' --model ollama/qwen3:8b A local 8B model proposes the graph → triage fanning out into four parallel

Xiaomi-Robotics-1: New robotics model released

Model ReleasesDGX agent

Xiaomi-Robotics-1 is a robot foundation model trained on over 100K hours of real-world manipulation trajectories. It is a Vision-Language-Action (VLA) model engineered for out-of-the-box mobile manipu

4 Aug 2026

A 2.6B model with tool calling and 128K context now runs at 30 tok/s on a phone

Model ReleasesDGX agent

Liquid AI released LFM2.5-2.6B today, and this might be more relevant to local AI than another massive model most people cannot run. The model is only 2.69B parameters, has 128K context, supports tool

A llama.cpp PR caches “hot” MoE experts on the GPU — 33 → 56 tok/s reported with 8GB VRAM

Model ReleasesDGX agent

A new llama.cpp PR (#26563) adds a heatmap that tracks which MoE experts are used most often. Instead of keeping every expert on the GPU or offloading all of them, it caches the frequently selected ex

Call Sam and tell him AGI is here

IndustryDGX agent

I jest, but I'm seriously impressed. Tonight I played with Codex for the first time. I had ChatGPT play two games on addictinggames.com Bow Man 2, and The Idiot Test, two really old browser games. Wha

Company approved 128GB Mac for research proposal, best model?

Model ReleasesDGX agent

I‘m doing a research proposal at my company about running local LLMs to replace daily coding models. Qwen 3.6 27B (or 3.8 potentially) is widely seen as the best model in that 20-60GB space, is that s

Decrease the power limit of your 5090 to at least 480W - the performance penalty for inference is negligible.

Model ReleasesDGX agent

I run my inference machine in the living room, so noise and heat output are a significant concern. Ran a quick test using my daily driver model (Qwen 3.6-27b) and at 480W, the card outputs only 2.1% l

[Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]

Model ReleasesDGX agent

First of all, obviously I took some help from AI to type this post and this is the topic that enabled me to accomplish all that: https://old.reddit.com/r/LocalLLaMA/comments/1veow4b/deepseek_v4flash_2

DeepSeek V4 Flash 0731 (Q4) now reaches 1,328 tok/s prefill and ~29 tok/s decode on one RTX PRO 6000

Model ReleasesDGX agent

I've been working on speeding up DeepSeek-V4-Flash-0731 in Krasis and have now got the long-prompt prefill quite a bit faster on a single RTX PRO 6000 96GB. These are timing-disabled internal Krasis r

Deepseek V4 flash 0731 ranks #21 on Agent Arena

Model ReleasesDGX agent

https://preview.redd.it/522fsdwvtdhh1.png?width=1200&format=png&auto=webp&s=6a6cf7a467514167a8193029dbd20fb3a9ba4f6c It ranks lower than both Sonnet 4.6 and Luna. I'd wager Luna costs in the same ball

Deepseek V4 Flash 2-bit quant is the first model I can run locally that achieves 100% in this SQL benchmark

Model ReleasesDGX agent

I really like to use this one SQL benchmark when testing new models. I had another post some time ago with my benchmarks, but I decided to post a new one because of how well Deepseek did. I like the b

DeepSeek v4 Flash vs. Qwen3.6-27B, 3.5-122B, and Gemma 4 31B Benchmark

Model ReleasesDGX agent

Just wanted to share my agentic coding benchmark run of DSv4F 0731 at both High and Low reasoning efforts (not Max)... I ran a 109-question subset of Aider Polyglot (the JS/C++/Python languages), base

GitHub - john-paul-ruf/spirit-guides: A reflective companion for self-guided inner work — AI-powered guides lead introspective sessions drawn from therapy, spirituality, and philosophy. Desktop app (Electron + React) with local markdown storage

Local AiDGX agent

I would love some feedback on this work in progress. Create, dialog with, mash up, and evolve spirit guides based on real concepts spanning psychological, philosophical, and religious origins. Honestl

GPT-OSS has turned one year old today!

Model ReleasesDGX agent

It is one of the best local models ever released, in both 20B and 120B versions. I always come back to it, especially the 120B version. Its only competition is, in my opinion, Qwen 3.5 122B, but that

Has anyone been working on a solid setup for DSV4F on x2+ R9700s?

HardwareDGX agent

I'm hoping that one of you guys has been working on an inference engine or has somehow found improvements to running DSV4F on RDNA4 multi-GPU setups. I am currently building a custom inference engine

Help with a prompt on Synthid

Local AiDGX agent

First I want to say how evil Synthid sounds. Did the creator want it to be like the Sith Lord or something?! For text generation the claim is it chooses word outputs that can be watermarked easily? Ma

Hugging Face CEO says China is winning the AI race and dominating on open models

Local AiDGX agent

This is something that was spoken here and there, and now it is like writing on the wall. The main additional point is that China has created an independent supply chain. Starting from raw materials a

I added a verify-before-load safety check for Ollama models

Local AiDGX agent

I maintain llm-checker, and I’ve added structural model-file validation for Ollama. Ollama stores downloaded models as local blobs. If one is truncated, malformed, or has invalid internal offsets, you

I benchmarked the 4 models I had pulled. The 1.1GB one beat the 2GB one at math and lost badly at extraction.

Model ReleasesDGX agent

152 generations, deterministic grading (exact number/string/JSON/regex), no LLM judge on my 16GB laptop. task type | deepseek-r1:1.5b (1.1GB) | llama3.2:3b (2.0GB) | gemma:2b | codellama 7b arithmetic

← Previous
123456…33
Next →