AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
Human
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
1,931 results
26 Jul 2026

Question regarding Ollama Cloud Metering.

Local AiDGX agent

How does it work? I read that the limts were more than what Opencode Go offers and subscribed to the 20USD Pro plan. But In practice based on how the usage bar fills up it seems like the 5 hour window

Simple desktop GUI for multiple local TTS models (Tkinter)

Model ReleasesDGX agent

https://preview.redd.it/qoogv5m1skfh1.png?width=688&format=png&auto=webp&s=3f06fb28ceec3cd3ab95bcebd0c71374332aac9b Built a simple desktop GUI (Tkinter) that supports multiple local TTS engines (curre

Understanding GPU Inference Workloads [D]

HardwareDGX agent

Hey everyone, I have been looking into how people source compute for their Inference workloads (and in general). I wanted to understand some specific pain points here. If you've used online services l

DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

We compared different LLMs on IMO 2026 [R]

Model ReleasesDGX agent

There are a few reasons why problems from International Mathematical Olympiad function as a good benchmark for LLMs: - The problems are new, not included in the training data of any model - Hard math

We open-sourced Logue — a privacy-first macOS meeting-notes + writing app that runs on-device (MLX, Apple Silicon) entirely

Local AiDGX agent

At Bitwize, we've been building Logue, a native macOS app for AI meeting notes and writing, and we just open-sourced it (MIT). We're sharing it here because the whole point is that it runs 100% on-dev

what am I doing wrong? (Ollama + openWebUI)

Local AiDGX agent

I have tried ollama locally on my gaming PC (5070 with 12bg of VRAM) It works pretty nice on qwen2.5-coder:7b (and 14b) So I decided to take a step further, and install an openWebUI instance on my hom

Will prices finally go down?

Local AiDGX agent

I am seeing more and more videos as posts about how OpenAI is in complete financial ruin, Anthropic isn't much better. Their expenses go with the revenue they make etc etc. Meta made big investments i

Will small model intelligence be limited by parameter count?

Model ReleasesDGX agent

Qwen3.6-27b is fantastic! It makes me wonder if there's a hard ceiling to smaller sized models. Do you guys think the ceiling of intelligence for smaller models will be constrained by factors like par

25 Jul 2026

4x 3090, 96gb vram what Model to drive Hermes?

Model ReleasesDGX agent

3 year lurker, now i finally got my server up and running. dont know which model to choose. llama.cpp or vllm, what makes more sense? mainly single user with maybe 2-3 more additional users in family,

5700, 48GB RAM and a 3090 24Gb. Best OS and framework/model?

Model ReleasesDGX agent

Hi everyone, I have a system which I have been using for gaming, R7 5700X, 48GB DDR4, RTX3090 24GB. But I want to use it for Local AI to reduce my reliance on cloud AI providers (mainly usage limits -

Benchmarks: TensorSharp vs. llama.cpp

Model ReleasesDGX agent

Cuda and Vulkan Benchmark: TensorSharp vs. llama.cpp I would like to share my latest open source local Unsloth (GGUF) LLM inference engine and applications. It supports many models from Unsloth, like

Best C++ Local Model? (July 24th 2026 Edition :-P)

Model ReleasesDGX agent

I apologize that this question has been asked in various flavors over time, but I couldn't find anything in the posts before that matches the options I have. I have a PC and a mac, both available in m

CachyLLama: llama.cpp fork with persistent SSD-backed KV caching for local agent workflows

Model ReleasesDGX agent

If you run local agentic coding harnesses (Aider, Claude Code, etc.), prompt evaluation usually eats up most of your execution time. Every turn re-evaluates thousands of identical prefix tokens_system

Deepseek V4 flash - Hy3 or is Qwen3.6 27B still the most solid for agentic/coding?

Model ReleasesDGX agent

I understand that the laguna model is either still buggy or potentially benchmaxxed. So I’d like to know for people who really tested, are DS flash or Hy3 really better in your usecase? submitted by /

DKV: Open-source KV-cache compression framework for local LLM inference (CLI + technical report)

Model ReleasesDGX agent

Hi everyone! Over the past five months I've been working on DKV (DifferentialKV), an open-source project exploring KV-cache compression for long-context local LLM inference. The goal is to reduce KV-c

Getting a second GPU in addition to my RTX3090

Model ReleasesDGX agent

Hello, I've been learning how to use local LLMs for a year or so on my workstation, using a RTX3090. Current setup : - i5 12400 - 64gb RAM - RTX 3090 - OS : Fedora KDE workstation I'm using LMStudio t

Help me complete my AI collection

Model ReleasesDGX agent

I’m building the ultimate AI tool vault, but every great collection has a few missing pieces. Note: I will react to every comment AI's currently installed: Qwen3.5-0.8B-UD-Q4_K_XL.gguf(classification)

I released Inflect v2: two ultra-tiny complete TTS models under 4M and 10M parameters

Model ReleasesDGX agent

I’ve spent the past month trying to find the point where an extremely small TTS model stops feeling like a size experiment and starts feeling genuinely useful. Today I’m releasing Inflect v2, with two

I spent a year building a free SDXL & Anima trainer that runs on my 12 GB GPU — here's what came out of it

Local AiDGX agent

A little over a year ago I got frustrated trying to fine-tune SDXL on my RTX 3060. Every option either forced lower resolution, locked away important settings behind massive config files, or needed a

I spent months testing whether ChatGPT can create a consistent 100-page comic. This is the result.

IndustryDGX agent

About 2.5 years ago I tried making an AI comic. It failed. Characters changed, environments drifted, and every page needed manual editing. So I started over with one simple rule: One prompt = one fini

I tried making a cinematic action trailer using Krea 2 + LTX 2.3

HardwareDGX agent

I wanted to challenge myself and see how far I could push Krea 2 and LTX 2.3, so I decided to create a short cinematic action trailer. It ended up being one of the most enjoyable AI projects I've work

Im back from gemini. GPT is astronomically better again.

Model ReleasesDGX agent

like 8 months ago I was tinkering with both and gemini was so much better i went with that. I had a project at work that I needed an AI to sift through a manual and schematic for and no matter what, g

Is it worth getting 128GB MacBook Pro? Will it ever be comparable to today’s frontier models for coding?

Model ReleasesDGX agent

I am a long time iOS app developer. In the last year I have been using Cursor+Claude/others to assist with app development. I am concerned that the current low pricing will disappear eventually. I am

Is this real ? Qwen3.6:27b with 128k context fit in 24Gb VRAM ?

Model ReleasesDGX agent

https://preview.redd.it/yw41s1jikefh1.png?width=1942&format=png&auto=webp&s=3a180ae6443c1db9f7b0ce621533a4b2aa553921 Hi, I've been running Ollama on my Unraid server since the llama2 era. I use to be

Kimi Linear 48B A3B?

Model ReleasesDGX agent

Just noticed this exists, 1M context MOE with 48B par seems just like what Ive been looking for - it runs pretty damn fast too compared to Qwen 3.6 35B. after some testing it seems capable of producin

Krea 2 Identity Edit running in Forge Neo (extension port) — works great, GitHub release soon

Local AiDGX agent

I ported Krea 2 Identity Edit support to Forge Neo as an extension — the instruction-based, identity-preserving edit LoRA that until now required ComfyUI + custom nodes. It replicates the full dual-co

Launching ComfyUI with a Blank Canvas (StabilityMatrix)

Model ReleasesDGX agent

Hey everyone, I’m posting this question in the subreddit because I haven’t been able to figure it out with the help of AI. I’ve asked ChatGPT and Gemini, but their answers are all over the place. So,

LFM 2.5 230M running at 1440 tok/s in-browser through a custom backend

Local AiDGX agent

Everything runs through WebGPU, in-browser or in electron/tauri apps. It's fully portable and supports either Nvidia and Apple Silicon (Metal). The actual kernels are optimized for the specific hardwa

Llama.cpp now has full MCP support!

Model ReleasesDGX agent

After a long and grueling effort spearheaded by ngxson, llama.cpp now fully supports MCP for all protocols. Over-the-web HTTP servers were already supported in the client (since they don't require any

Local alternative to Kling AI 3.0 Motion Control (ComfyUI, 16GB VRAM)

Local AiDGX agent

Hi everyone, I'm looking for a local alternative to Kling AI 3.0 Motion Control that I can run in ComfyUI. What I'm specifically looking for is a model or workflow that allows me to: - Control charact

MI50 power curve tests

Model ReleasesDGX agent

tests done power limiting the GPU on LACT - real power usage varies wildy at 20W it ranges from 25W to 56W same behavior happens on every setting prompt for the test runs: https://github.com/lukesdevl

Microsoft's website shows OpenAI as one of the signatories of the open weight AI letter

Local AiDGX agent

https://preview.redd.it/a24z80gr6afh1.png?width=1181&format=png&auto=webp&s=4a844ebe2319eb6230dbdc63c9caf492bed5ff47 So, this came up on: https://www.microsoft.com/en-us/corporate-responsibility/topic

Mobile Offline LLMs: What do you use them for?

Model ReleasesDGX agent

I've spent the last year or so playing around with open source MLX and GGUF models on iPhone hardware. Given the limitations in memory, GPU/CPU/ANE, and in turn the context window I've been trying to

Modelos do ollama cloud perdendo qualidade?

Local AiDGX agent

A algumas semanas, percebi algo diferente, aparentemente modelos de qualidade, em especial o GLM 5.1 ficou mais burro, e começou a mandar caracteres em mandarim para mim sem eu nem usar eles, e isto n

My ChatGPT created a rally speech in response to counting letters.

AgentsDGX agent

I was exploring how to prompt the voice chat so that it can correctly count the number of “e”s in seventeen (which it fails most of the time). At the beginning of a new chat, instead of counting “e”s

Old Coder Needs help with New AI Development and wants to get up to speed to understand it all.

Local AiDGX agent

Hi Guys, I'm an old coder and DBA that has been in the field for almost 40 years. More and more the jobs I was doing for work are being taken over by AI and the need for my type of work is diminishing

Ollama Cloud Quota Benchmark

Model ReleasesDGX agent

Recently I bought an Ollama Cloud sub and accidently spent my whole 5h quota upon using DeepSeek V4 Pro... but why? isnt it supposed to be a cheap model? Youd think there would be a correlation betwee

Ollama Qwen3.6:35b randomly stops outputting tokens

Model ReleasesDGX agent

RTX 4070, 32gb system ram, Linux. NVIDIA-SMI 610.43.03, KMD Version: 610.43.03, CUDA UMD Version: 13.3 Systemd service modifications: [Service] Environment='OLLAMA_HOST=0.0.0.0:11434' Environment='OLL

OrangePi AI Studio Pro - Qwen3.5-122B-A10B

Local AiDGX agent

https://preview.redd.it/wbq8ullnbafh1.png?width=1409&format=png&auto=webp&s=e6d2fe2b1c87c724bc64003c25f917dcee53260f I finally got round to tweaking this, with a bit of help from GLM5.2. The trick to

PSA: DO NOT use Intel consumer platforms for multi-GPU setups

Local AiDGX agent

Since a lot more people are trying to build their own multi-GPU machines, I thought I should help to prevent a common mistake people make with building multi-GPU machines. Which is using an Intel cons

SVDQuant + native INT8/W4A4 for Krea 2 on ComfyUI — up to 2x faster, works on any modern NVIDIA GPU

HardwareDGX agent

Quantized Krea 2 Turbo checkpoints for ComfyUI, up to 2x faster and about a third smaller than the usual FP8 version — no calibration dataset, no quality cliff. How to use it (short version): clone th

Who ONLY use local models?

Local AiDGX agent

Please be honest. I would love to hear about guys really dedicated to local AI and who really reject subscriptions (especially to openai and anthropic). What do you use your model for? submitted by /u

24 Jul 2026

[audio.cpp] Release 0.4: Higgs Audio v3 TTS 4B (10x real time)+ Fish Audio S2 Pro in C++/GGML, full GGUF loading, Q8 speed and VRAM gains

Model ReleasesDGX agent

audio.cpp again :) Release 0.4 is out. The headline this time is new high-quality TTS coverage plus GGUF becoming a first-class across the project. What’s new: Added Higgs Audio v3 TTS 4B, Fish Audio

[BIG DATASET RELEASE] - SupraLabs/reasoning-corpus-4K-5M-v1 - Train your tiny SLMs to think!

Local AiDGX agent

https://preview.redd.it/b7ybs7nqx5fh1.png?width=3440&format=png&auto=webp&s=e6aaaa15cbe59debaae1ebb7fcd708167e86dc35 Hey r/LocalLLaMA ! We are back and we have something really amazing today. Our big

CachyLLama’s: llama.cpp fork with persistent KV cache that makes long local-agent sessions much less painful

Model ReleasesDGX agent

I’m not affiliated with this project, but I’ve been running it recently and I’m surprised it hasn’t received more attention here: https://github.com/fewtarius/CachyLLama CachyLLama is a fork of llama.

Can LLMs solve mazes?

Model ReleasesDGX agent

https://reddit.com/link/1v5rvuq/video/bgmwc754i9fh1/player My goal was to create a benchmark to measure the spatial awareness and memory of models. Eventually, I came up with the simple idea of a maze

Does Ollama Cloud prompt caching even work?

Local AiDGX agent

I launched a new session and tasked GLM to create an implementation plan for a spec. The plan was on the bigger side, about 5k lines. I started with 0% 5h used and ended with 80% used. A few more twea

Explain love in one sentence...

Local AiDGX agent

Love is an active commitment of deep affection where you find genuine joy in prioritizing someone else’s well-being as deeply as or beyond your own needs. VS Love is the deep, enduring connection betw

Extened garlic to run Qwen3.5 35B A3B float8 at 55 tok/s on RTX 5060 Ti

Model ReleasesDGX agent

In a previous post (https://www.reddit.com/r/LocalLLaMA/comments/1utefpr/running_qwen3_30b_a3b_at_50_toks_on_rtx_5060_ti/) there seemed to be great demand for bringing in Qwen3.5 35B. Some Gated Delta

Fizgig Krea 2 training features update

SafetyDGX agent

https://github.com/shootthesound/Fizgig Intelligent trainer - Per-image loss tracking with self-adapting training runs — every image gets its own verdict (easy / suspect / stuck / exhausted) and its o

FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence

Local AiDGX agent

Introducing FLUX 3. One multi-modal model for Image, Video, Audio and Action-Prediction. Creations are truer to life in every kind of style. Blog Post : https://bfl.ai/blog/flux-3 submitted by /u/pmtt

For the first time, ChatGPT asked me a question instead of writing a wall of text

TutorialsDGX agent

I copy-pasted this cooking recipe and accidentally hit enter before adding the actual prompt. Normally this would result in ChatGPT interpreting the text, trying to guess what I need, and giving a wal

Getting the most out of MTP

Model ReleasesDGX agent

If you want to get the most out of MTP. You have to run some tests / benchmarks to do so. Turning it on with defaults will get improvements, but for many models and card combinations, you are leaving

Honest take on Laguna S2.1 and its uses (from actual use)

Model ReleasesDGX agent

So I've taken some time to actually test laguna on a few of my own projects. I wanted to share as I feel most peoples comments at this point have just been about getting it running or saying it doesnt

Hugging Face releases The Stack v3 – largest open code dataset yet

Local AiDGX agent

From Anton Lozhkov on 𝕏: https://x.com/anton_lozhkov/status/2080254608639701222 Two ways in: stack-v3-train - near-deduplicated, quality-filtered, PII-redacted, contents inline. Point load_dataset at

I built a compiler that turns computation graphs into the weights of a vanilla transformer — no training anywhere [P]

Model ReleasesDGX agent

I've been chasing the question of what algorithms a transformer can actually express -- separate from what it can learn. So I built a compiler: define a computation graph in ordinary Python, and it pr

I built an open-source multi-agent SDLC harness that beats a cold Claude Code run on large repos, by learning the repo once. Real benchmarks (incl. where it loses) inside. [P]

Model ReleasesDGX agent

Built an open-source AI coding agent that was 7%–75% cheaper than a cold 'claude -p' run on 6/6 well-localized tasks across repositories up to ~82k LOC. The biggest difference: Cold agent: 6.83, 207 t

Is corruption the lobbying against Open weights?

Local AiDGX agent

Like, reading things like Anthropic 'donated' to some people with the condition of lobbying against Chinese LLMs.. it's that right? It feels nothing like freedom but at the same time it's said 'out lo

Is everyone training a single model together, based on the principle of the Tor network?

Local AiDGX agent

I just had a thought while scrolling. No idea if this already exists. What if users trained an AI model together—a bit like the Tor network or Bitcoin mining back in the day? - Participants download a

It appears that the anti opensource AI lobby is far outgunned already

Local AiDGX agent

The earlier post on this subreddit by 20+ companies signing the petition including Microsoft, Meta, Nvidia, YC (https://www.microsoft.com/en-us/corporate-responsibility/topics/open-weight/) etc plus t

← Previous
1…89101112…33
Next →