AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
1,446 results
Model Releases

DeepSeek v4 Flash has a nice bump in Capability

DGX agent

DeepSeek V4 Flash: Preview → 2026-07-31 Benchmark Preview 0731 Δ Terminal Bench* 56.9 82.7 +25.8 Toolathlon 51.8 70.3 +18.5 NL2Repo — 54.2 new Cybergym — 76.7 new DeepSWE — 54.4 new Agent Last Exam —

model-releasesr-localllama
31 Jul 2026
Model Releases

Deepseek V4 Flash on SlopCodeBench

AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

While waiting for some of the quants to drop, I load the API with $50 and ran it on SlopCodeBench Just vibe reading the results it seems like Opus 4.8 < Deepseek < Opus 5 https://github.com/michaelasp

model-releasesr-localllama
31 Jul 2026
Model Releases

Meituan just dropped LongCat-Flash-Lite-Sparse

DGX agent

It’s an MoE with ~3B active params and a 30B n-gram lookup table offloaded to RAM for fast 256k context on a 24GB GPU. Reminds me of Gemma 4’s PLE trick. Initial analysis suggest it wont be replacing

model-releasesr-localllama
31 Jul 2026
Model Releases

SenseNova U1.5 Lite preview just dropped

DGX agent

SenseNova released U1.5-Lite-Preview Benchmarks: Qwen-Image-Bench from 47.14 to 55.20. ImgEdit-Bench from 3.90 to 4.37. GEdit-Bench-en from 7.47 to 8.17. Key updates: 4K native generation with better

model-releasesr-localllama
31 Jul 2026
Model Releases

Will ollama upgrade Deepseek V4 Flash on cloud?

DGX agent

https://preview.redd.it/vxl4zslewigh1.png?width=1435&format=png&auto=webp&s=5a419870ca0cb13076be6c9ff4ef33d8177d5eda New version is 25% better than previous one and is near GLM-5.2 quality submitted b

model-releasesr-ollama
31 Jul 2026
Local Ai

GLM 5.2 with vision on Hugging Face

DGX agent

Hi all, I have not seen this model talked about here but it seems like baseten (inference provider on OpenRouter) merged the vision encoder from Kimi k2.6 into GLM 5.2. I think the lack of vision was

local-air-localllama
30 Jul 2026
Model Releases

How to uninstall Ollama Claude code

DGX agent

pretty simple I did ollama launch claude, found out its slow asf, wanted to delete It and Idk how I have no idra if its the same as normal Claude Code uninstall or if it's a Little bit of a different

model-releasesr-ollama
30 Jul 2026
Model Releases

Inkling-Small by thinkingmachines

DGX agent

276B total parameters, 12B active, 1M context window. Blog post: https://thinkingmachines.ai/news/inkling-small/ NVFP4: https://huggingface.co/thinkingmachines/Inkling-Small-NVFP4 GGUF's by Unsloth: h

model-releasesr-localllama
30 Jul 2026
Model Releases

Turbo-fieldfare: Open-source engine running Gemma 4 26B in 2 GB RAM on Apple Silicon

DGX agent

Its a custom Swift/Metal inference engine that runs Gemma 4 26B-A4B-IT on M-series Macs with very low RAM. It uses ~2GB instead of ~14 GB. The result is reportedly 5–6 tok/s on an 8 GB M2 MacBook Air

model-releasesr-localllama
30 Jul 2026
Model Releases

What is the fastest local research tool (deep research) ?

DGX agent

I've tried grok and Claude's deep research mode and I was amazed with the speed considering the amount of sources analysed. Is there anything as fast that can run locally? My guess would be that to ru

model-releasesr-localllama
30 Jul 2026
Model Releases

ChatGPT Made Me Cry Tonight

DGX agent

Sorry if flair is wrong. I decided to finally get a ChatGPT subscription after some conversations with it about health issues with my dog. I've only used AI for coding work, primarily Claude, but I fe

model-releasesr-chatgpt
29 Jul 2026
Model Releases

dropped 4k on a spark, am I crazy?

DGX agent

Saw that the Asus Ascent 1tb was going for $3,950 from a few sources, couldn't stop thinking about it, finally just went ahead and did it. Am I completely insane? Will I regret this? I can't imagine t

model-releasesr-localllama
29 Jul 2026
Local Ai

Everyone posts day-one impressions. What's still in your stack a month later?

DGX agent

Day one threads are the least useful thing we produce here and we produce a lot of them. Model drops, forty people run their favourite prompt, half say it's the best thing ever and half say benchmaxxe

local-air-localllama
29 Jul 2026
Model Releases

First Kimi K3 results on home lab ~ 4t/s

DGX agent

I've got better results than expected for 768gb DDR5 and 2x5090. Using fork https://github.com/pwilkin/llama.cpp/tree/kimi-k3-text and https://huggingface.co/GrEarl/Kimi-K3-GGUF Q2_K quant. Prefill sp

model-releasesr-localllama
29 Jul 2026
Model Releases

PSA: llama.cpp now loads MTP tensors by default for any draft-mtp arch, even with MTP disabled

DGX agent

If your GGUF has MTP/NextN tensors baked in (GLM-5.2, hy_v3, qwen35moe, step35, etc.), recent llama.cpp builds load them by default — even if you never pass --spec-type draft-mtp. Before, they were sk

model-releasesr-localllama
29 Jul 2026
Local Ai

Vendor-agnostic ML inference on production edge devices [R]

DGX agent

I work on PostSlate, a video editing tool, and this comes out of our own work. We run ML models on-device, face detection and embedding among other things, which means we can't assume anything about t

local-air-machinelearning
29 Jul 2026
Model Releases

Building a training dataset: pulling and restoring stills from video sources

DGX agent

I was looking for a tool to help me train a character lora from an old movie (think 1980's low-budget movie). The digital transfer was low-quality; modern upscales exist and they are horrible. So I wa

model-releasesr-stablediffusion
28 Jul 2026
Local Ai

I Just Built an Open-Source Alternative to Flora AI

DGX agent

Hey, I just made a node editor for generating assets using AI. It’s something similar to ComfyUI, but it’s more simple and abstract. You can say it’s the open-source alternative to Flora AI. It’s curr

local-air-stablediffusion
28 Jul 2026
Local Ai

Manga Coloring Tool 2

DGX agent

Hey everyone! 👋 I'm excited to announce the official release of Manga Coloring Tool 2.0, a completely free, local, open-source web application designed to colorize manga pages and chapters effortlessl

local-air-stablediffusion
28 Jul 2026
Model Releases

NeurIPS 2026 Reviewer: AI-Generated Rebuttals (and Paper) [D]

DGX agent

One of the papers I reviewed has what seems to be entirely LLM-generated rebuttals, and the original paper is also clearly LLM-generated, with Claude-speak everywhere. While the authors acknowledge LL

model-releasesr-machinelearning
28 Jul 2026
Model Releases

spec: add DSpark speculative decoding by wjinxu · Pull Request #25173 · ggml-org/llama.cpp

DGX agent

It's time to experiment using DSpark! Please share your stats(pp/tg improvements). DSpark related stuff to check: DeepSpec - a deepseek-ai Collection DeepSeek-V4 with DSpark - DeepSeek-V4-Pro-DSpark &

model-releasesr-localllama
28 Jul 2026
Model Releases

ThinkingCap-Qwen3.6-27B warrants a look

DGX agent

It has only been two days since I move 100% from Qwen3.5-27B F16 to ThinkingCap-Qwen3.6-27B F16. Where I was getting tps in 30-40 range (depending on the size of the context), I am definitely getting

model-releasesr-localllama
28 Jul 2026
Local Ai

Zuck's opinion: The AI Future Is for Everyone

DGX agent

’Tis the season of AI open letters and manifestos, apparently. Mark Zuckerberg has now entered the debate over the future of AI with a WSJ op-ed published today - and frankly, his position is much mor

local-air-localllama
28 Jul 2026
Local Ai

Nvidia CEO Jensen Huang defends Open Source AI by saying distillation is fundamental to learning

DGX agent

Nvidia CEO Jensen Huang “Distillation - learning from AI, learning from other people, and learning from other sources of knowledge, is fundamental to intelligence. We are constantly learning from one

local-air-localllama
27 Jul 2026
Model Releases

Qwen3.6-27B speculative decoding gets better on heavier quants

DGX agent

I finished the speed leg of my spec-decode benchmarking for Qwen3.6-27B, main algorithms across quants. Overall: the heavier the quant, the more spec-decode buys you (10 of 10 speculative configs rank

model-releasesr-localllama
27 Jul 2026
Local Ai

RX 9060 XT 16GB vs RTX 5060 Ti 16GB for local AI — worth the price difference?

DGX agent

Hey everyone! I’m building a PC to run AI models locally and can’t decide between the RX 9060 XT 16GB and the RTX 5060 Ti 16GB. Both have the same amount of VRAM, but is AMD actually a solid choice fo

local-air-ollama
27 Jul 2026
Model Releases

16 bit better than lower quants for Qwen3.6-27B

DGX agent

I am writing a fairly complex C++ windows MFC application. I have a few 3090s and can run F16 Qwen3.6-27B with 256K context and MTP. The quality of code is exceptional with this quant vs its lower qua

model-releasesr-localllama
26 Jul 2026
Model Releases

Harness showdown: Claude Code vs OpenCode vs Pi with DeepSeek V4 Flash

DGX agent

I ran DeepSeek V4 Flash through Claude Code, OpenCode and Pi on my own benchmark, and the quality came out basically the same across all three while the time and tokens spent was wildly different. Cla

model-releasesr-localllama
26 Jul 2026
Model Releases

I want to use AI coding agents for machine learning projects [D]

DGX agent

I'm a software engineer who mainly builds softwaes/applications, and I'm starting to work on machine learning projects. Since ML workloads often require GPUs, I know services like Google Colab and Kag

model-releasesr-machinelearning
26 Jul 2026
Model Releases

Local-first LLM pipeline tracer — @trace on any function, dashboard at localhost. Feedback welcome.

DGX agent

Hey r/LocalLLaMA — maintainer here, obviously biased. OpenSmith is an open-source Python tracing tool for LLM pipelines. The idea: drop u/trace on any function, run opensmith ui, get a full local dash

model-releasesr-localllama
26 Jul 2026
Model Releases

Need help with setup

DGX agent

I am setting up codex+ollama+qwen3.6:27b for my hobby coding project on a windows pc with rtx5090 - earlier i tried to setup vllm in Ubuntu container but couldn’t get that to work - now using ollama,

model-releasesr-ollama
26 Jul 2026
Model Releases

DKV: Open-source KV-cache compression framework for local LLM inference (CLI + technical report)

DGX agent

Hi everyone! Over the past five months I've been working on DKV (DifferentialKV), an open-source project exploring KV-cache compression for long-context local LLM inference. The goal is to reduce KV-c

model-releasesr-localllama
25 Jul 2026
Model Releases

Im back from gemini. GPT is astronomically better again.

DGX agent

like 8 months ago I was tinkering with both and gemini was so much better i went with that. I had a project at work that I needed an AI to sift through a manual and schematic for and no matter what, g

model-releasesr-chatgpt
25 Jul 2026
Model Releases

Launching ComfyUI with a Blank Canvas (StabilityMatrix)

DGX agent

Hey everyone, I’m posting this question in the subreddit because I haven’t been able to figure it out with the help of AI. I’ve asked ChatGPT and Gemini, but their answers are all over the place. So,

model-releasesr-stablediffusion
25 Jul 2026
Model Releases

Extened garlic to run Qwen3.5 35B A3B float8 at 55 tok/s on RTX 5060 Ti

DGX agent

In a previous post (https://www.reddit.com/r/LocalLLaMA/comments/1utefpr/running_qwen3_30b_a3b_at_50_toks_on_rtx_5060_ti/) there seemed to be great demand for bringing in Qwen3.5 35B. Some Gated Delta

model-releasesr-localllama
24 Jul 2026
Model Releases

I built a compiler that turns computation graphs into the weights of a vanilla transformer — no training anywhere [P]

DGX agent

I've been chasing the question of what algorithms a transformer can actually express -- separate from what it can learn. So I built a compiler: define a computation graph in ordinary Python, and it pr

model-releasesr-machinelearning
24 Jul 2026
Tutorials

The 'distillation' claim is just ridiculous in nature

DGX agent

Even if China was distilling from US models (assuming all accusations are true), nothing about it makes it illegal. It is like saying you distilled knowledge from your professor in colleges and now he

tutorialsr-localllama
24 Jul 2026
Model Releases

Apple M5 isn't making full use of its matmul cores yet

DGX agent

At the moment MLX (and Llama.cpp for Macs) run 16bit activations everywhere. Despite this, the M5 generation silicon actually does support INT8 activations - it actually allows w4a8 d_type. It's just

model-releasesr-localllama
23 Jul 2026
Model Releases

I built an open-source RAG chatbot starter that runs fully locally with Ollama (FastAPI + ChromaDB)

DGX agent

I kept re-wiring the same RAG plumbing on every project, so I turned it into a clean starter and open-sourced it. Upload a PDF, ask questions, and get answers with page-level source citations. It runs

model-releasesr-ollama
23 Jul 2026
Local Ai

Interesting reasoning by phi4-mini-reasoning

DGX agent

https://preview.redd.it/joxxvyzhf1fh1.png?width=2386&format=png&auto=webp&s=a55c09d11eb080b83d3b66b59013a09118913c9d Freshly installed, just asked it 'who are you'. Why does this happen? 😄 submitted b

local-air-ollama
23 Jul 2026
Hardware

Looking for feedback on my GPU-accelerated Snake AI project [P]

DGX agent

I've been building an AI that learns to play the classic Snake game through reinforcement learning. The goal is to reach high scores while keeping training time as low as possible. The current version

hardwarer-machinelearning
21 Jul 2026
Model Releases

cuda: extract Q1_0 elements via __byte_perm by dfriehs · Pull Request #25628 · ggml-org/llama.cpp

DGX agent

I don't have the ability to access Reddit posts or browse specific URLs. To provide you with an accurate factual summary for your knowledge base, I would need either: 1. The actual content/text from t

model-releasesr-localllama
16 Jul 2026
Model Releases

Qwen3.5 122B-A10B · ROCmFP4 iMatrix

DGX agent

Hola Strix and AMD stacker frendios. Read the Lineage and Credits, this uses charlie12345/ROCmFPX, won't work on native llama.cpp yet. 122B total · 10B active · 60.70 GiB · 28.50 tok/s MTP-off · BF16

model-releasesr-localllama
16 Jul 2026
Model Releases

ggml-zendnn : add Q8_0 quantization support by z-sachin · Pull Request #23414 · ggml-org/llama.cpp

DGX agent

Benchmark Results Benchmark configuration: threads = 96 type_k = bf16 type_v = bf16 Llama-3.1-8B-Instruct Q8_0 Prompt Size GGML_CPU_Q8_0 t/s ZenDNN_Q8_0 t/s Gain 256 472.28 730.87 54.75% 512 450.86 83

model-releasesr-localllama
15 Jul 2026
Model Releases

Hermes on Android (Graphene OS)

DGX agent

https://youtu.be/oxpGq5FITgA?si=nkHWLReGCDYe7QfL I got Hermes running in the native Debian Terminal in Graphene OS and its really slick. Voice dictation works amazingly. Im using a remote Hermes gatew

model-releasesr-localllama
15 Jul 2026
Model Releases

Recent llama.cpp updates for SYCL/Intel

DGX agent

Some fixes & boost(pp) for SYCL/Intel. Merged PRs: [SYCL] Flash Attention with XMX engine via oneDNN graph API (SDPA) on KV f16 for Xe2 ; Qwen3.6-27b-Q8_0 prefill speed up x1.21 at p=512 and x4.26 at

model-releasesr-localllama
15 Jul 2026
Research

A slightly improved DVD-JEPA demo [P]

DGX agent

This post likely presents an enhanced demonstration of DVD-JEPA, a video variant of the Joint Embedding Predictive Architecture model. The JEPA framework has been extended to video tasks (V-JEPA) , an

researchr-machinelearning
21 Jun 2026
Model Releases

The “dead internet theory” in action: In World of Warcraft, a server without humans has appeared - instead, 1,800 DeepSeek-based bots are playing there. The bots behave like regular players: they chat, level up characters, run dungeons, and even fight each other.

DGX agent

A World of Warcraft server has become populated entirely by approximately 1,800 AI bots based on DeepSeek, which engage in typical player activities including chatting, leveling characters, running du

model-releasesr-chatgpt
21 Jun 2026
← Previous
1…1819202122…31
Next →