AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “r-localllama”

GridTimelineEvolution
469 results
26 Jul 2026

Simple desktop GUI for multiple local TTS models (Tkinter)

Model ReleasesDGX agent

https://preview.redd.it/qoogv5m1skfh1.png?width=688&format=png&auto=webp&s=3f06fb28ceec3cd3ab95bcebd0c71374332aac9b Built a simple desktop GUI (Tkinter) that supports multiple local TTS engines (curre

We open-sourced Logue — a privacy-first macOS meeting-notes + writing app that runs on-device (MLX, Apple Silicon) entirely

Local AiDGX agent

At Bitwize, we've been building Logue, a native macOS app for AI meeting notes and writing, and we just open-sourced it (MIT). We're sharing it here because the whole point is that it runs 100% on-dev

Will prices finally go down?

Local AiDGX agent

I am seeing more and more videos as posts about how OpenAI is in complete financial ruin, Anthropic isn't much better. Their expenses go with the revenue they make etc etc. Meta made big investments i


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Will small model intelligence be limited by parameter count?

Model ReleasesDGX agent

Qwen3.6-27b is fantastic! It makes me wonder if there's a hard ceiling to smaller sized models. Do you guys think the ceiling of intelligence for smaller models will be constrained by factors like par

25 Jul 2026

4x 3090, 96gb vram what Model to drive Hermes?

Model ReleasesDGX agent

3 year lurker, now i finally got my server up and running. dont know which model to choose. llama.cpp or vllm, what makes more sense? mainly single user with maybe 2-3 more additional users in family,

5700, 48GB RAM and a 3090 24Gb. Best OS and framework/model?

Model ReleasesDGX agent

Hi everyone, I have a system which I have been using for gaming, R7 5700X, 48GB DDR4, RTX3090 24GB. But I want to use it for Local AI to reduce my reliance on cloud AI providers (mainly usage limits -

Benchmarks: TensorSharp vs. llama.cpp

Model ReleasesDGX agent

Cuda and Vulkan Benchmark: TensorSharp vs. llama.cpp I would like to share my latest open source local Unsloth (GGUF) LLM inference engine and applications. It supports many models from Unsloth, like

CachyLLama: llama.cpp fork with persistent SSD-backed KV caching for local agent workflows

Model ReleasesDGX agent

If you run local agentic coding harnesses (Aider, Claude Code, etc.), prompt evaluation usually eats up most of your execution time. Every turn re-evaluates thousands of identical prefix tokens_system

Deepseek V4 flash - Hy3 or is Qwen3.6 27B still the most solid for agentic/coding?

Model ReleasesDGX agent

I understand that the laguna model is either still buggy or potentially benchmaxxed. So I’d like to know for people who really tested, are DS flash or Hy3 really better in your usecase? submitted by /

DKV: Open-source KV-cache compression framework for local LLM inference (CLI + technical report)

Model ReleasesDGX agent

Hi everyone! Over the past five months I've been working on DKV (DifferentialKV), an open-source project exploring KV-cache compression for long-context local LLM inference. The goal is to reduce KV-c

Getting a second GPU in addition to my RTX3090

Model ReleasesDGX agent

Hello, I've been learning how to use local LLMs for a year or so on my workstation, using a RTX3090. Current setup : - i5 12400 - 64gb RAM - RTX 3090 - OS : Fedora KDE workstation I'm using LMStudio t

Help me complete my AI collection

Model ReleasesDGX agent

I’m building the ultimate AI tool vault, but every great collection has a few missing pieces. Note: I will react to every comment AI's currently installed: Qwen3.5-0.8B-UD-Q4_K_XL.gguf(classification)

I released Inflect v2: two ultra-tiny complete TTS models under 4M and 10M parameters

Model ReleasesDGX agent

I’ve spent the past month trying to find the point where an extremely small TTS model stops feeling like a size experiment and starts feeling genuinely useful. Today I’m releasing Inflect v2, with two

Is it worth getting 128GB MacBook Pro? Will it ever be comparable to today’s frontier models for coding?

Model ReleasesDGX agent

I am a long time iOS app developer. In the last year I have been using Cursor+Claude/others to assist with app development. I am concerned that the current low pricing will disappear eventually. I am

Kimi Linear 48B A3B?

Model ReleasesDGX agent

Just noticed this exists, 1M context MOE with 48B par seems just like what Ive been looking for - it runs pretty damn fast too compared to Qwen 3.6 35B. after some testing it seems capable of producin

LFM 2.5 230M running at 1440 tok/s in-browser through a custom backend

Local AiDGX agent

Everything runs through WebGPU, in-browser or in electron/tauri apps. It's fully portable and supports either Nvidia and Apple Silicon (Metal). The actual kernels are optimized for the specific hardwa

Llama.cpp now has full MCP support!

Model ReleasesDGX agent

After a long and grueling effort spearheaded by ngxson, llama.cpp now fully supports MCP for all protocols. Over-the-web HTTP servers were already supported in the client (since they don't require any

MI50 power curve tests

Model ReleasesDGX agent

tests done power limiting the GPU on LACT - real power usage varies wildy at 20W it ranges from 25W to 56W same behavior happens on every setting prompt for the test runs: https://github.com/lukesdevl

Microsoft's website shows OpenAI as one of the signatories of the open weight AI letter

Local AiDGX agent

https://preview.redd.it/a24z80gr6afh1.png?width=1181&format=png&auto=webp&s=4a844ebe2319eb6230dbdc63c9caf492bed5ff47 So, this came up on: https://www.microsoft.com/en-us/corporate-responsibility/topic

Mobile Offline LLMs: What do you use them for?

Model ReleasesDGX agent

I've spent the last year or so playing around with open source MLX and GGUF models on iPhone hardware. Given the limitations in memory, GPU/CPU/ANE, and in turn the context window I've been trying to

Old Coder Needs help with New AI Development and wants to get up to speed to understand it all.

Local AiDGX agent

Hi Guys, I'm an old coder and DBA that has been in the field for almost 40 years. More and more the jobs I was doing for work are being taken over by AI and the need for my type of work is diminishing

Ollama Qwen3.6:35b randomly stops outputting tokens

Model ReleasesDGX agent

RTX 4070, 32gb system ram, Linux. NVIDIA-SMI 610.43.03, KMD Version: 610.43.03, CUDA UMD Version: 13.3 Systemd service modifications: [Service] Environment='OLLAMA_HOST=0.0.0.0:11434' Environment='OLL

OrangePi AI Studio Pro - Qwen3.5-122B-A10B

Local AiDGX agent

https://preview.redd.it/wbq8ullnbafh1.png?width=1409&format=png&auto=webp&s=e6d2fe2b1c87c724bc64003c25f917dcee53260f I finally got round to tweaking this, with a bit of help from GLM5.2. The trick to

PSA: DO NOT use Intel consumer platforms for multi-GPU setups

Local AiDGX agent

Since a lot more people are trying to build their own multi-GPU machines, I thought I should help to prevent a common mistake people make with building multi-GPU machines. Which is using an Intel cons

Who ONLY use local models?

Local AiDGX agent

Please be honest. I would love to hear about guys really dedicated to local AI and who really reject subscriptions (especially to openai and anthropic). What do you use your model for? submitted by /u

24 Jul 2026

[audio.cpp] Release 0.4: Higgs Audio v3 TTS 4B (10x real time)+ Fish Audio S2 Pro in C++/GGML, full GGUF loading, Q8 speed and VRAM gains

Model ReleasesDGX agent

audio.cpp again :) Release 0.4 is out. The headline this time is new high-quality TTS coverage plus GGUF becoming a first-class across the project. What’s new: Added Higgs Audio v3 TTS 4B, Fish Audio

[BIG DATASET RELEASE] - SupraLabs/reasoning-corpus-4K-5M-v1 - Train your tiny SLMs to think!

Local AiDGX agent

https://preview.redd.it/b7ybs7nqx5fh1.png?width=3440&format=png&auto=webp&s=e6aaaa15cbe59debaae1ebb7fcd708167e86dc35 Hey r/LocalLLaMA ! We are back and we have something really amazing today. Our big

CachyLLama’s: llama.cpp fork with persistent KV cache that makes long local-agent sessions much less painful

Model ReleasesDGX agent

I’m not affiliated with this project, but I’ve been running it recently and I’m surprised it hasn’t received more attention here: https://github.com/fewtarius/CachyLLama CachyLLama is a fork of llama.

Can LLMs solve mazes?

Model ReleasesDGX agent

https://reddit.com/link/1v5rvuq/video/bgmwc754i9fh1/player My goal was to create a benchmark to measure the spatial awareness and memory of models. Eventually, I came up with the simple idea of a maze

Extened garlic to run Qwen3.5 35B A3B float8 at 55 tok/s on RTX 5060 Ti

Model ReleasesDGX agent

In a previous post (https://www.reddit.com/r/LocalLLaMA/comments/1utefpr/running_qwen3_30b_a3b_at_50_toks_on_rtx_5060_ti/) there seemed to be great demand for bringing in Qwen3.5 35B. Some Gated Delta

FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence

Local AiDGX agent

Introducing FLUX 3. One multi-modal model for Image, Video, Audio and Action-Prediction. Creations are truer to life in every kind of style. Blog Post : https://bfl.ai/blog/flux-3 submitted by /u/pmtt

Getting the most out of MTP

Model ReleasesDGX agent

If you want to get the most out of MTP. You have to run some tests / benchmarks to do so. Turning it on with defaults will get improvements, but for many models and card combinations, you are leaving

Honest take on Laguna S2.1 and its uses (from actual use)

Model ReleasesDGX agent

So I've taken some time to actually test laguna on a few of my own projects. I wanted to share as I feel most peoples comments at this point have just been about getting it running or saying it doesnt

Hugging Face releases The Stack v3 – largest open code dataset yet

Local AiDGX agent

From Anton Lozhkov on 𝕏: https://x.com/anton_lozhkov/status/2080254608639701222 Two ways in: stack-v3-train - near-deduplicated, quality-filtered, PII-redacted, contents inline. Point load_dataset at

Is corruption the lobbying against Open weights?

Local AiDGX agent

Like, reading things like Anthropic 'donated' to some people with the condition of lobbying against Chinese LLMs.. it's that right? It feels nothing like freedom but at the same time it's said 'out lo

It appears that the anti opensource AI lobby is far outgunned already

Local AiDGX agent

The earlier post on this subreddit by 20+ companies signing the petition including Microsoft, Meta, Nvidia, YC (https://www.microsoft.com/en-us/corporate-responsibility/topics/open-weight/) etc plus t

More than 20 companies including NVIDIA, Meta, Microsoft, Palantir, and Hugging Face have signed a letter urging policymakers to avoid premature restrictions on open weight models.

Local AiDGX agent

The Open Letter was initiated by Microsoft and published today: “Open Weights and American AI Leadership”. It argues against broad or premature restrictions on open-weight models and explicitly says p

[Paper] Statistically-Lossless Quantization of Large Language Models

Model ReleasesDGX agent

Model quantization has become essential for efficient large language model deployment, yet existing approaches involve clear trade-offs: methods such as GPTQ and AWQ achieve practical compression but

Spent two weeks on a kernel that benchmarked 29x faster. End to end it's maybe 6-10%, and it's not even wired in yet.

Local AiDGX agent

I've been building a C99 inference engine from scratch (no Python, no BLAS, just gcc and make) that runs BitNet's ternary models on CPU. A few weeks ago I got obsessed with the matmul kernel - wrote a

swiss-ai/Apertus-v1.5 70B/8B

Model ReleasesDGX agent

https://huggingface.co/swiss-ai/Apertus-v1.5-70B https://huggingface.co/swiss-ai/Apertus-v1.5-8B Apertus 1.5 is a family of 8B and 70B parameter language models designed to advance the state of multil

The 'distillation' claim is just ridiculous in nature

TutorialsDGX agent

Even if China was distilling from US models (assuming all accusations are true), nothing about it makes it illegal. It is like saying you distilled knowledge from your professor in colleges and now he

What's the last model trained on human-data only?

Local AiDGX agent

From my understanding, most current LLMs are trained on trillions and trillions of tokens of mostly AI-generated data. Are there any recent models that are trained purely (or as close as possible) on

Zagreus-0.4B-por a small open source language model for Portuguese

Model ReleasesDGX agent

mii-llm, an open source AI lab, released Zagreus-0.4B-por, a compact bilingual Portuguese–English language model pretrained entirely from scratch. The model has approximately 400 million parameters an

23 Jul 2026

A caveman qwen3.6 27B

Local AiDGX agent

Just saw this on huggingface: https://huggingface.co/ProCreations/grug-27b The benchmarks claim that it's quite a bit better than qwen3.6 27B original and that they reduced the amount of necessary tok

Absurd claim: the distilled model outperforms the originals

Local AiDGX agent

As an AI community of LLM experts, are we really going to stay silent while US officials make absurd claims to push anti-consumer laws? Not only does the release timeline between Fable and K3 make hig

AI9Stars released G9v3-3B

Model ReleasesDGX agent

AI9Stars has released G9v3-3B an open weights language model designed to deliver strong reasoning capabilities within a lightweight 3 billion parameter size. It is released under the Apache 2.0 licens

Apple M5 isn't making full use of its matmul cores yet

Model ReleasesDGX agent

At the moment MLX (and Llama.cpp for Macs) run 16bit activations everywhere. Despite this, the M5 generation silicon actually does support INT8 activations - it actually allows w4a8 d_type. It's just

Arcee AI has spoken out against the ban on open Chinese models in US

Model ReleasesDGX agent

This is rather counterintuitive, since banning Chinese models would benefit them the most. Jensen Huang is also against the ban, although the interests here are more obvious. Do you think that if Arce

contrib: allow all AI-generated code in general by ngxson · Pull Request #26012 · ggml-org/llama.cpp

Model ReleasesDGX agent

Having read some merged PRs in the past, I know that they were fully written by Claude Code (or similar), so this basically fixes the delusion. But at the same time, we might start seeing more AI slop

CPU-only inference on a Celeron N5095 SBC: 6 models from 0.6B to 8B, benchmarked

Model ReleasesDGX agent

I wanted to know how cheap you can go and still run local models, so I ran Ollama CPU-only on a Youyeetoo X1S. It's a single-board x86 machine with a Celeron N5095 (Jasper Lake, 4C/4T, 15W), 16GB of R

DeepSeek Founder’s 4-hour investor meeting: DeepSeek is prioritizing AGI over user growth and commercialisation

Model ReleasesDGX agent

A Chinese article compiled 52 remarks from Liang Wenfeng’s four-hour investor meeting. I’ve summarised the most important ones below. DeepSeek has one central objective: AGI. This is not the time to m

Deepseek V4 Flash ~105 t/s on two Nvidia 4090d 48G (ada) in vLLM

Model ReleasesDGX agent

TLDR: I (with the help of AI) re-implemented every Blackwell-only kernel (DeepGEMM, FlashInfer sparse-MLA, block-scaled FP8) in Triton, because they simply don't exist for sm89. The performance is 2-3

FYI You dont need expensive networking for multi-node gpu. 30t/s laguna Q2_K_XL (39.7GB) on 2x4060+1x4060 using a $20 usb->ethernet.

Model ReleasesDGX agent

Turns out a regular ethernet cable between 2 nodes can run laguna UD-Q2_K_XL (39.7GB) using a direct point to point network. Interestingly on `nvidia-smi dmon -s pucvmet -d 2`, the inter/intra gpu tra

I Made a Local Huggingface On My NAS

Local AiDGX agent

https://preview.redd.it/u8alj38wr0fh1.png?width=1860&format=png&auto=webp&s=3578776c60d9548a135a018702f44d0fddedd4b0 Little side project I'm doing so I can easily transfer any model I want fast to my

I trained a 0.5M model on 1B tokens of Fineweb-edu dataset.

Model ReleasesDGX agent

Hi everyone, About a month ago I publish my very first research paper on my neural network architecture called Silia. You can look at the model here: https://huggingface.co/Srijan-Srivastava/Silia-v2

inclusionAI/LLaDA2.2-flash · Hugging Face

Model ReleasesDGX agent

LLaDA2.2-flash is an agent-oriented diffusion language model in the LLaDA2 series. By introducing Levenshtein Editing (with DELETE and INSERT control tokens) to diffusion language modeling, it represe

Kwaipilot/KAT-Coder-V2.5-Dev · Hugging Face

Model ReleasesDGX agent

from kwaipilot: Following the release of KAT-Coder-V2.5 in July, we are pleased to release the open-weight version KAT-Coder-V2.5-Dev, an MOE model with a total parameter count of 35B and 3B activated

Laguna-S-2.1 'thinking forever' loops seem to be a quantization artifact

Model ReleasesDGX agent

If you're running Laguna S 2.1 on llama.cpp and hitting thinking loops because it won't close its </think> tags, you might want to look at your quant before you spend too much time tweaking settings.

Model 'distillation' accusations are getting way overblown at this point

Model ReleasesDGX agent

The news about Anthropic settling a class action lawsuit for 1.5B over training data isn't just a legal headache for them, it's a massive warning sign for engineering teams relying entirely on closed

MoE models around A2B

Model ReleasesDGX agent

There's a bunch of small MoE with around 1B active params, like LFM2.5 8B A1B and Granite 4.0h 7B A1B; and then there are models with 3B+ like Qwen 3.x ~30B A3B and Gemma 4 26B A4B, but those are alre

← Previous
1…5678
Next →