AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
Human
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
1,927 results
12 Aug 2026

According to AMD, Arm, and Microsoft, agentic AI could push CPU-to-GPU ratios from 1:4 to even1:1

Local AiDGX agent

In OCP APAC 2026, Tai AMD SVP of compute and enterprise AI said agents don't cut GPU demand but they just pile on a whole extra layer of orchestration, retrieval, and tool-calling work that runs on CP

Hidden Reasoning from Claude and GPT are Decoded, and it is interesting

Model ReleasesDGX agent

Yesteday a paper showed a gap that allows to see 100% of the reasoning tokens form ALL Claude and GPT models Stealing Reasoning Traces from Proprietary LLM APIs. check it out, they have published lots

How to do clean uninstall of chatgpt desktop app on windows?

Local AiDGX agent

ChatGPT desktop app will not download images even after a complete reinstall I am on Windows 11 and the Download button in the ChatGPT desktop app does nothing when I try to download generated images.

DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Idea for a deepseek-v4-flash-0731 backed automated research workflow to be leveraged via qwen3.6/3.8 27b for difficult tasks that require highly technical, not easy to find information.

Model ReleasesDGX agent

Sometimes you have tasks that are outside of your expertise and the idea is this workflow automation could be leveraged to manage to have local AI figure it out using research from his workflow gather

Is the future of AI selling hardware for Open Source/Models?

HardwareDGX agent

I’m not super knowledgeable of the entire AI industry, but as we see this industry grow and the Cold War that is happening between the US and China on AI development, I can’t help but notice what US c

LiquidAI/LFM2.5-VL-3B · Hugging Face

Local AiDGX agent

LFM2.5-VL-3B is a multimodal variant of LFM2.5, a family of hybrid models designed for on-device deployment. It builds on LFM2-VL-3B with further mid- and post-training. LFM2.5-VL-3B can process both

New Muse-Glimmer-30B SoTA Quants - hopefully a new lineup :)

Model ReleasesDGX agent

Hey Folks, I've been making quants for a while - recently I took a short break to get into hardcore research (submitted my first EMNLP paper during it!). Along the way, I built up a little arsenal of

Qwen 3.8 2.4T is out , no 27b today RIP.

Model ReleasesDGX agent

i was hoping they would release both today but the big one just dropped now , and it says its been listed since 5 hours ago on huggingface. I guess the real date for 27b is https://modelscope.cn/model

RAG for regular users?

Local AiDGX agent

One of the reasons I got into local LLMs was the possibility of getting answers using my own documents and books (a few hundreds) instead of having to search through them manually. However since I'm n

Tested Nemotron 3.5 Lightning locally on coding, Hermes Agent and agentic work

Model ReleasesDGX agent

Ran the model with quants (Q5) and MTP by bartowski with llama.cpp server. It takes ~24GB ram running on M5 Pro with 48GB at about 65t/s. On some tasks it was quite the overthinker. Overall, the quali

What do you guys do for GPU Kernels?

HardwareDGX agent

I'm trying to figure out GPU Kernel optimization on older hardware like SM80(ampere) . Is there tools you guys use? Or frameworks? Im waiting for this framework https://www.reddit.com/r/LocalLLaMA/com

What unique, custom QOL upgrades have you given your local agents?

Model ReleasesDGX agent

Warning: Kinda long post. If you don't like reading, please skip for your own sanity. Also, I've got nothing to sell, just a tinkerer, so I just want to share ideas and learn from you guys too. When I

11 Aug 2026

1 Day in and I feel okay saying Muse-Glimmer-30B finally beats 3.6-27B for the size in some use-cases

Model ReleasesDGX agent

A few things right off the bat: it reasons very efficiently. Like Grok 4.5 levels of efficient thinking it quantizes very well. My first few tests with iq3_xxs were better than Qwen/Gemma behaved at t

10 year garbage card for local llms

Model ReleasesDGX agent

Hello everyone! ​I like dumb things. I like working with weak computers and microcontrollers. I like the simplicity and low electricity usage. Simply put, the efficiency of a 'dumb' PC. ​The first tim

12GB VRAM gang, what's our plan?

Model ReleasesDGX agent

Seems like we're limited to qwen finetuned MoEs for now. Looking at the current landscape - focus seems to be on dense models (muse glimmer 30b, qwen 3.8 27b) for smaller setups. Is upgrading to 24GB

Anyone Using (Koreas) 'Solar Open 2' (250B, 15B) Model?

Model ReleasesDGX agent

I just heard of this model. Seems to be a competitor to DeepSeek V4 Flash. About the same size and active parameters. Anyone tested it compared to V4 Flash? Link: https://huggingface.co/upstage/Solar-

Can Gemma and Qwen models catch hallucinations by looking at their own logprobs?

Model ReleasesDGX agent

Hi! I'm really obsessed with LLM hallucinations for the last 6 days 😭 I started by designing system prompts to attack hallucinations but failed, obviously. Now I tried reading logprobs and... I think

DeepSeek-V4-Flash-0731 (284B MoE) at 75 tok/s on 2× DGX Spark — full recipe, 11 gotchas, reboot-proof cluster, Codex CLI integration

Model ReleasesDGX agent

Spent two nights getting deepseek-ai/DeepSeek-V4-Flash-0731 (284B MoE, 13B active, native FP4/FP8, 1M context) running production-grade on two DGX Sparks connected by one QSFP DAC cable. Everything —

DeepSeek V4 Flash 0731 at 27+ t/s decode on Strix Halo — Vulkan + DSpark full guide

Model ReleasesDGX agent

Been benchmarking DSv4 Flash 0731 on a Flow Z13 (Ryzen AI MAX+ 395, Radeon 8060S / gfx1151, 128GB LPDDR5X) for the past week. Figured I'd share what actually works and what doesn't — there are a lot o

DeepSeek-V4-Flash acting as my Linux sysadmin

Model ReleasesDGX agent

I'm very happy with some Linux admin tasks I'm throwing at a locally running DeepSeek. My request was simple, check why 'samples' folder is taking more and more space on one of the machines on my LAN,

I built a weird, low-power llama.cpp server using an Intel N100 + RTX 5060Ti

Model ReleasesDGX agent

Everything started with the sudden death of my old ASRock J1900. While looking for the perfect ITX replacement, I stumbled upon the Chinese CW-NAS-ADLN-K motherboard, which looked perfect on paper: In

I gave DeepSeek V4 Flash basic vision by training a 40M connector on 100K examples

Model ReleasesDGX agent

I wanted to find out whether a huge text-only MoE could be given basic vision without retraining the language model itself. The short answer is yes. I froze DeepSeek V4 Flash and a 417M-parameter Moon

I put Gemma 4 E4B and E2B into an e-reader so I can ask my weird questions and share my thoughts in private directly in app.

Model ReleasesDGX agent

Here's how it works in the app: Framework: Runs on LiteRT-LM (like Google's AI Edge). Models: Downloads either the E2B (~2.5 GB) or E4B (~3.6 GB) INT4 quantized models directly from ungated litert-com

I ran Muse Glimmer @ 1M context - All tests passed.

Model ReleasesDGX agent

Heeeey all! I just completed some fun tests with Muse Glimmer, I thought I'd let you know. In fact, the summary below was written by Muse itself! I ran a 2× DGX Spark cluster and got Meta's day-old Mu

I tested the CMP170HX

Model ReleasesDGX agent

Lots of rumor and misinfo bouncing around, so I put some of these old mining cards to the test. I used 4 of the 8GB cards, set to 64GB each. Lots of models fit entirely on a single card, and you can a

I will be parting with my 4x Spark Cluster.

Local AiDGX agent

Laid off then my partner of 10 years said he's leaving, have to move, etc... I will post the r/hardwareswap link when I make it. I'm willing to add some incentive for r/LocalLLaMA folks. I will also a

I will pay 1 million for a developer to build this

IndustryDGX agent

I WILL PAY 1 MILLION FOR THIS 😭 I will literally pay 1,000,000 to the developer who fixes this. ChatGPT macOS: ⌘C text → ⌘Tab back to ChatGPT → ⌘V …and the search box ISN’T FOCUSED. 🤦‍♂️ Why do I have

Introducing Unsloth Desktop app

Model ReleasesDGX agent

Hi LocalLlama, we're super excited to release Unsloth Desktop today! 🦥 It's the first desktop app that enables you to run and train models locally. Open-source. Available on Mac, Windows, and Linux Su

Llama-CPP Parallel Agents --> fine for decode, but one agent's prefill will grind all other agents to a halt

Model ReleasesDGX agent

Testing with 3-5 agents. Decode performance is superb, however if one performs a web search and needs to process a few thousand tokens, ALL other agents will grind to a halt: I've tried tuning a littl

[llama.cpp PR #26608] Ling-3.0 support (unmerged)

Model ReleasesDGX agent

aetherbird has done some great work getting Ling-3.0 to work in llama.cpp. The architecture is generally identical to deepseekv2. I recently added a microscopic 40 line PR to his that adds support for

Local Benchmark : Muse Glimmer 30B vs Qwen 3.6 27B vs Gemma4 31B (and many other models and finetunes)

Model ReleasesDGX agent

Needs a lot of requests compared to Qwen (almost twice) and Gemma (almost x3). Final score is fine, even though it is 'not a coding model' https://wonderrico.github.io/local_llm_benchmark/benchmark-ma

Luth-2: New State-of-the-Art French Small Language Models

Model ReleasesDGX agent

Hey everyone, Today we release Luth-2-0.8B and Luth2-2-2B, two non-reasoning models that set a new state of the art for French across a wide variety of tasks for their size 🚀 A few notable scores on F

MiniMax-H3: ~38 GB less VRAM with Runtime LoRA Bypass — DoRA Dynamic LoRA Loader v1.0.39

Local AiDGX agent

GitHub: https://github.com/xmarre/ComfyUI-DoRA-Dynamic-LoRA-Loader Release v1.0.39: https://github.com/xmarre/ComfyUI-DoRA-Dynamic-LoRA-Loader/releases/tag/v1.0.39 Also available through ComfyUI Manag

Muse-Glimmer 30B Hits ~280 t/s in Real Production Coding

Model ReleasesDGX agent

These numbers were captured during a real feature implementation task in Next.js and Nest.js (adding a theme switching system across components). The structural predictability of UI/state refactoring

Nvidia Nemo Switchyard

HardwareDGX agent

https://github.com/NVIDIA-NeMo/Switchyard Finally an open source LLM router. An alternative to openrouter fusion and Sakana Fugu. Doesn't look like it does exactly what Sakana Fugu does according to i

Observations on Muse-Glimmer reasoning traces being noticeably different from qwen / gemma models and questions for you guys

Model ReleasesDGX agent

Just downloaded the model, UD-Q5_K_XL quant, asked it to generate a long story to test out reasoning and speed with dflash (super fast btw, ~ 90 to 160 tok/s on a 5090 depending on task) and was surpr

OpenAI just launched a cybersecurity model that answers 95% of advanced threat queries. And Meta put a frontier model on your laptop. Same day.

Model ReleasesDGX agent

Something happened today that I think most people are going to miss because there are two separate stories and neither one is getting the full picture. OpenAI expanded Daybreak. If you haven't heard o

Psychological methods really work. Has anyone tried to encouraging and instilling confidence to GPT?

Model ReleasesDGX agent

Oddly enough, encouragement actually influences the performance of not only Claude but also GPT and other AIs. While Claude was working on a complex problem related to the Riemann Hypothesis, the Anth

Rumored 50-series Super refresh bumps everything +50% VRAM

Local AiDGX agent

leaked Super specs have the 5070 Ti and 5080 going 16GB to 24GB and the 5070 to 18GB, thanks to the new 3GB GDDR7 modules. 24GB on a Ti-class card is actually the number people here have been waiting

Tested in Coding: BF16 Muse Glimmer vs BF16 Qwen3.6 27B

Model ReleasesDGX agent

I'm guessing that many people have been waiting for this comparison. For clarity, both models are running at full FP16 KV-cache. Due to VRAM limitations, Muse Glimmer is running full 262,144 context,

The small open weight models are scarier in AI development

Model ReleasesDGX agent

Imagine if your everyday laptop could run an AI model smart enough to take care of 90% of your work—totally private, lightning fast, and completely free of monthly fees. That is the exact tipping poin

Today it was apparently my turn

Model ReleasesDGX agent

I’ve been using ChatGPT for about two months now after giving up on Claude and Gemini and I became a true zealot. I’m now using Pro, and despite having the personality, set to default, never had any p

Toy project: a chat title model that fits in 5 MiB of ram

Model ReleasesDGX agent

Not even sure if I'm allowed to post this, what with the 'completely/primarily LLM generated copy' rule (the post itself is fine, but the repo/model I'm sharing definitely is, whoops) and the whole li

Using the different models in different industries - your experience?

IndustryDGX agent

Hi all, I'm seeing so many conversations from people discussing how they're 'using the models wrong' and 'don't use Sol max as high is enough for you', yet all the conversations lack the nuance of wha

We quantized DeepSeek V4 0731 and benchmarked it against popular quants on 8× RTX 5090

Model ReleasesDGX agent

We converted the model from the original safetensors and found two issues. The first one made our quantization fail several times, the second one does not fail at all, it just quietly ruins the base 1

Writing formats I can no longer read because AI has beaten them to death

TutorialsDGX agent

I don’t even care anymore whether these posts are *actually* AI-generated. The problem is that there’s now a very specific style of internet writing that instantly makes my brain refuse to continue re

10 Aug 2026

1M context with 17 GB model in 24 GB VRAM: 'for the first time I was able to load a context of almost 1M tokens and extract 7 needles from various parts of the text'

Model ReleasesDGX agent

https://preview.redd.it/xxjh11f38jih1.png?width=1852&format=png&auto=webp&s=76850ed51e29a8bc86c2ca718d4320075eed4363 Just wanted to share a user report that I found to be very interesting. Some person

Achievable 253 t/s - unsloth/Muse Glimmer 30B UD-Q5_K_M on a 5090

Model ReleasesDGX agent

Benchmarked Muse Glimmer 30B on my RTX 5090 (32GB), 262k context, UD-Q5_K_M + dflash-kquant + mmproj. Workload Stock master + DFlash ngram-simple PR #26842 + DFlash Code patch 78 t/s 57 t/s 220-253 t/

Best Local LLMs - August 2026

Local AiDGX agent

Wowee!! Just when you thought it couldn't get better for open weight models, we probably have had our best period yet!?!?! Models that rival the closed frontier, Opus level models on non-insane hardwa

Best open-source harness like Claude Code?

Model ReleasesDGX agent

Avid claude code user here looking to do equivalent things with local models. Just want to plug in something like Qwen and have the interface be 1:1 with claude code. Any suggestion? submitted by /u/N

Chat UIs with native audio input for multimodal models?

Model ReleasesDGX agent

I've been running Gemma 4 E4B with oMLX and I can't find any chat interfaces that directly send the audio file to the model instead of running the audio through a separate STT layer. I can confirm the

Comparing how Cline, Kilo, and Qwen Code handle long-task context/state (and why context loops keep happening)

Model ReleasesDGX agent

I've been comparing Cline / Kilo / Qwen Code lately since they all handle long-task state differently. Cline: has Focus Chain, a markdown file kept outside the conversation that gets reinjected on a c

DeepSeek V4 Flash 0731 is the ‘killer app’ that is going to sell A LOT of DGX Sparks

Model ReleasesDGX agent

Having a ‘Killer Application’ that everyone wants to use helps sell hardware, plain and simple. DeepSeek V4 Flash 0731 isn’t an app of course, but I think it’s going to be the major catalyst for getti

GPT 5.6 Sol High and X-High (Web Chat) feels severely nerfed since 08/06/2026 update

Model ReleasesDGX agent

GPT-5.6 Sol High and X-High (Web Chat) feels severely nerfed since 08/06/2026 update (https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt) I use it for a pretty complex Unreal Engine 5 project (

How to prevent LLM to act like a robot/assistant?

Local AiDGX agent

I'm playing with a conversational agent I made using either api/generate or api/chats. In both case I do ask him to not ask follow up question, to not act like an assistant, etc. Either from a system

I asked OPUS 5 to make a video about what it's like to be an LLM

Model ReleasesDGX agent

Full prompt I gave to Claude Opus 5: can you use whatever resources you like, and python, to generate a short 'youtube poop' video and render it using ffmpeg ? can you put more of a personal spin on i

I compared GGUF quants of Qwen3.6 27B to NVFP4, AWQ, AutoRound, and FP8

Model ReleasesDGX agent

There's an interactive chart and some extra data in the blog post if you're interested. There are plenty of KL-divergence benchmarks for GGUF models, but most of them compare one GGUF quant against an

I trained a 1B-parameter LLM from scratch on 20B tokens for about $200

Model ReleasesDGX agent

A few months ago, I had the idea of making a LLM from scratch as a personal project (for learning and partly for improving my resume). Since I learned a lot from other posts on here over the past year

I trained an open-source realism LoRA for MiniMax H3 - it makes generated people actually look real (weights inside)

Local AiDGX agent

Update : New version is ready and online , should be much better, fully functionnal on ComfyUI, and you can find before/after here : https://huggingface.co/fal/MiniMax-H3-Realism-People-LoRA/blob/main

inclusionAI/Ling-3.0-tiny · 8B A1.3B MoE· Hugging Face

Model ReleasesDGX agent

Looks like the Ling team open weighted a much smaller version of the Ling-3.0-flash they open weighted a few days ago. It's 8B params with 1.3B active, and seems to fall between the 4B and 8-12B Qwen

← Previous
123…33
Next →