AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
All
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “r-localllama”

GridTimelineEvolution
469 results
Local Ai

Who ONLY use local models?

DGX agent

Please be honest. I would love to hear about guys really dedicated to local AI and who really reject subscriptions (especially to openai and anthropic). What do you use your model for? submitted by /u

local-air-localllama
25 Jul 2026
Model Releases

[audio.cpp] Release 0.4: Higgs Audio v3 TTS 4B (10x real time)+ Fish Audio S2 Pro in C++/GGML, full GGUF loading, Q8 speed and VRAM gains

Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

audio.cpp again :) Release 0.4 is out. The headline this time is new high-quality TTS coverage plus GGUF becoming a first-class across the project. What’s new: Added Higgs Audio v3 TTS 4B, Fish Audio

model-releasesr-localllama
24 Jul 2026
Local Ai

[BIG DATASET RELEASE] - SupraLabs/reasoning-corpus-4K-5M-v1 - Train your tiny SLMs to think!

DGX agent

https://preview.redd.it/b7ybs7nqx5fh1.png?width=3440&format=png&auto=webp&s=e6aaaa15cbe59debaae1ebb7fcd708167e86dc35 Hey r/LocalLLaMA ! We are back and we have something really amazing today. Our big

local-air-localllama
24 Jul 2026
Model Releases

CachyLLama’s: llama.cpp fork with persistent KV cache that makes long local-agent sessions much less painful

DGX agent

I’m not affiliated with this project, but I’ve been running it recently and I’m surprised it hasn’t received more attention here: https://github.com/fewtarius/CachyLLama CachyLLama is a fork of llama.

model-releasesr-localllama
24 Jul 2026
Model Releases

Can LLMs solve mazes?

DGX agent

https://reddit.com/link/1v5rvuq/video/bgmwc754i9fh1/player My goal was to create a benchmark to measure the spatial awareness and memory of models. Eventually, I came up with the simple idea of a maze

model-releasesr-localllama
24 Jul 2026
Model Releases

Extened garlic to run Qwen3.5 35B A3B float8 at 55 tok/s on RTX 5060 Ti

DGX agent

In a previous post (https://www.reddit.com/r/LocalLLaMA/comments/1utefpr/running_qwen3_30b_a3b_at_50_toks_on_rtx_5060_ti/) there seemed to be great demand for bringing in Qwen3.5 35B. Some Gated Delta

model-releasesr-localllama
24 Jul 2026
Local Ai

FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence

DGX agent

Introducing FLUX 3. One multi-modal model for Image, Video, Audio and Action-Prediction. Creations are truer to life in every kind of style. Blog Post : https://bfl.ai/blog/flux-3 submitted by /u/pmtt

local-air-localllama
24 Jul 2026
Model Releases

Getting the most out of MTP

DGX agent

If you want to get the most out of MTP. You have to run some tests / benchmarks to do so. Turning it on with defaults will get improvements, but for many models and card combinations, you are leaving

model-releasesr-localllama
24 Jul 2026
Model Releases

Honest take on Laguna S2.1 and its uses (from actual use)

DGX agent

So I've taken some time to actually test laguna on a few of my own projects. I wanted to share as I feel most peoples comments at this point have just been about getting it running or saying it doesnt

model-releasesr-localllama
24 Jul 2026
Local Ai

Hugging Face releases The Stack v3 – largest open code dataset yet

DGX agent

From Anton Lozhkov on 𝕏: https://x.com/anton_lozhkov/status/2080254608639701222 Two ways in: stack-v3-train - near-deduplicated, quality-filtered, PII-redacted, contents inline. Point load_dataset at

local-air-localllama
24 Jul 2026
Local Ai

Is corruption the lobbying against Open weights?

DGX agent

Like, reading things like Anthropic 'donated' to some people with the condition of lobbying against Chinese LLMs.. it's that right? It feels nothing like freedom but at the same time it's said 'out lo

local-air-localllama
24 Jul 2026
Local Ai

It appears that the anti opensource AI lobby is far outgunned already

DGX agent

The earlier post on this subreddit by 20+ companies signing the petition including Microsoft, Meta, Nvidia, YC (https://www.microsoft.com/en-us/corporate-responsibility/topics/open-weight/) etc plus t

local-air-localllama
24 Jul 2026
Local Ai

More than 20 companies including NVIDIA, Meta, Microsoft, Palantir, and Hugging Face have signed a letter urging policymakers to avoid premature restrictions on open weight models.

DGX agent

The Open Letter was initiated by Microsoft and published today: “Open Weights and American AI Leadership”. It argues against broad or premature restrictions on open-weight models and explicitly says p

local-air-localllama
24 Jul 2026
Model Releases

[Paper] Statistically-Lossless Quantization of Large Language Models

DGX agent

Model quantization has become essential for efficient large language model deployment, yet existing approaches involve clear trade-offs: methods such as GPTQ and AWQ achieve practical compression but

model-releasesr-localllama
24 Jul 2026
Local Ai

Spent two weeks on a kernel that benchmarked 29x faster. End to end it's maybe 6-10%, and it's not even wired in yet.

DGX agent

I've been building a C99 inference engine from scratch (no Python, no BLAS, just gcc and make) that runs BitNet's ternary models on CPU. A few weeks ago I got obsessed with the matmul kernel - wrote a

local-air-localllama
24 Jul 2026
Model Releases

swiss-ai/Apertus-v1.5 70B/8B

DGX agent

https://huggingface.co/swiss-ai/Apertus-v1.5-70B https://huggingface.co/swiss-ai/Apertus-v1.5-8B Apertus 1.5 is a family of 8B and 70B parameter language models designed to advance the state of multil

model-releasesr-localllama
24 Jul 2026
Tutorials

The 'distillation' claim is just ridiculous in nature

DGX agent

Even if China was distilling from US models (assuming all accusations are true), nothing about it makes it illegal. It is like saying you distilled knowledge from your professor in colleges and now he

tutorialsr-localllama
24 Jul 2026
Local Ai

What's the last model trained on human-data only?

DGX agent

From my understanding, most current LLMs are trained on trillions and trillions of tokens of mostly AI-generated data. Are there any recent models that are trained purely (or as close as possible) on

local-air-localllama
24 Jul 2026
Model Releases

Zagreus-0.4B-por a small open source language model for Portuguese

DGX agent

mii-llm, an open source AI lab, released Zagreus-0.4B-por, a compact bilingual Portuguese–English language model pretrained entirely from scratch. The model has approximately 400 million parameters an

model-releasesr-localllama
24 Jul 2026
Local Ai

A caveman qwen3.6 27B

DGX agent

Just saw this on huggingface: https://huggingface.co/ProCreations/grug-27b The benchmarks claim that it's quite a bit better than qwen3.6 27B original and that they reduced the amount of necessary tok

local-air-localllama
23 Jul 2026
Local Ai

Absurd claim: the distilled model outperforms the originals

DGX agent

As an AI community of LLM experts, are we really going to stay silent while US officials make absurd claims to push anti-consumer laws? Not only does the release timeline between Fable and K3 make hig

local-air-localllama
23 Jul 2026
Model Releases

AI9Stars released G9v3-3B

DGX agent

AI9Stars has released G9v3-3B an open weights language model designed to deliver strong reasoning capabilities within a lightweight 3 billion parameter size. It is released under the Apache 2.0 licens

model-releasesr-localllama
23 Jul 2026
Model Releases

Apple M5 isn't making full use of its matmul cores yet

DGX agent

At the moment MLX (and Llama.cpp for Macs) run 16bit activations everywhere. Despite this, the M5 generation silicon actually does support INT8 activations - it actually allows w4a8 d_type. It's just

model-releasesr-localllama
23 Jul 2026
Model Releases

Arcee AI has spoken out against the ban on open Chinese models in US

DGX agent

This is rather counterintuitive, since banning Chinese models would benefit them the most. Jensen Huang is also against the ban, although the interests here are more obvious. Do you think that if Arce

model-releasesr-localllama
23 Jul 2026
Model Releases

contrib: allow all AI-generated code in general by ngxson · Pull Request #26012 · ggml-org/llama.cpp

DGX agent

Having read some merged PRs in the past, I know that they were fully written by Claude Code (or similar), so this basically fixes the delusion. But at the same time, we might start seeing more AI slop

model-releasesr-localllama
23 Jul 2026
Model Releases

CPU-only inference on a Celeron N5095 SBC: 6 models from 0.6B to 8B, benchmarked

DGX agent

I wanted to know how cheap you can go and still run local models, so I ran Ollama CPU-only on a Youyeetoo X1S. It's a single-board x86 machine with a Celeron N5095 (Jasper Lake, 4C/4T, 15W), 16GB of R

model-releasesr-localllama
23 Jul 2026
Model Releases

DeepSeek Founder’s 4-hour investor meeting: DeepSeek is prioritizing AGI over user growth and commercialisation

DGX agent

A Chinese article compiled 52 remarks from Liang Wenfeng’s four-hour investor meeting. I’ve summarised the most important ones below. DeepSeek has one central objective: AGI. This is not the time to m

model-releasesr-localllama
23 Jul 2026
Model Releases

Deepseek V4 Flash ~105 t/s on two Nvidia 4090d 48G (ada) in vLLM

DGX agent

TLDR: I (with the help of AI) re-implemented every Blackwell-only kernel (DeepGEMM, FlashInfer sparse-MLA, block-scaled FP8) in Triton, because they simply don't exist for sm89. The performance is 2-3

model-releasesr-localllama
23 Jul 2026
Model Releases

FYI You dont need expensive networking for multi-node gpu. 30t/s laguna Q2_K_XL (39.7GB) on 2x4060+1x4060 using a $20 usb->ethernet.

DGX agent

Turns out a regular ethernet cable between 2 nodes can run laguna UD-Q2_K_XL (39.7GB) using a direct point to point network. Interestingly on `nvidia-smi dmon -s pucvmet -d 2`, the inter/intra gpu tra

model-releasesr-localllama
23 Jul 2026
Local Ai

I Made a Local Huggingface On My NAS

DGX agent

https://preview.redd.it/u8alj38wr0fh1.png?width=1860&format=png&auto=webp&s=3578776c60d9548a135a018702f44d0fddedd4b0 Little side project I'm doing so I can easily transfer any model I want fast to my

local-air-localllama
23 Jul 2026
Model Releases

I trained a 0.5M model on 1B tokens of Fineweb-edu dataset.

DGX agent

Hi everyone, About a month ago I publish my very first research paper on my neural network architecture called Silia. You can look at the model here: https://huggingface.co/Srijan-Srivastava/Silia-v2

model-releasesr-localllama
23 Jul 2026
Model Releases

inclusionAI/LLaDA2.2-flash · Hugging Face

DGX agent

LLaDA2.2-flash is an agent-oriented diffusion language model in the LLaDA2 series. By introducing Levenshtein Editing (with DELETE and INSERT control tokens) to diffusion language modeling, it represe

model-releasesr-localllama
23 Jul 2026
Model Releases

Kwaipilot/KAT-Coder-V2.5-Dev · Hugging Face

DGX agent

from kwaipilot: Following the release of KAT-Coder-V2.5 in July, we are pleased to release the open-weight version KAT-Coder-V2.5-Dev, an MOE model with a total parameter count of 35B and 3B activated

model-releasesr-localllama
23 Jul 2026
Model Releases

Laguna-S-2.1 'thinking forever' loops seem to be a quantization artifact

DGX agent

If you're running Laguna S 2.1 on llama.cpp and hitting thinking loops because it won't close its </think> tags, you might want to look at your quant before you spend too much time tweaking settings.

model-releasesr-localllama
23 Jul 2026
Model Releases

Model 'distillation' accusations are getting way overblown at this point

DGX agent

The news about Anthropic settling a class action lawsuit for 1.5B over training data isn't just a legal headache for them, it's a massive warning sign for engineering teams relying entirely on closed

model-releasesr-localllama
23 Jul 2026
Model Releases

MoE models around A2B

DGX agent

There's a bunch of small MoE with around 1B active params, like LFM2.5 8B A1B and Granite 4.0h 7B A1B; and then there are models with 3B+ like Qwen 3.x ~30B A3B and Gemma 4 26B A4B, but those are alre

model-releasesr-localllama
23 Jul 2026
Local Ai

PaddlePaddle/HPD-Parsing · Hugging Face

DGX agent

HPD-Parsing: Hierarchical Parallel Document Parsing We introduce HPD-Parsing, a lightweight (1B) and high-throughput document parsing model built on a Hierarchical Parallel Decoding paradigm. Unified

local-air-localllama
23 Jul 2026
Model Releases

[Paper] SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

DGX agent

Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlappe

model-releasesr-localllama
23 Jul 2026
Model Releases

PSA on Laguna S-2.1 - Use the updated chat template and GGUF

DGX agent

Link to their official GGUF repo: https://huggingface.co/poolside/Laguna-S-2.1-GGUF/tree/main All the GGUFs received this fix 5ish hours ago - correct yarn_attn_factor to 1.0 (llama.cpp derives mscale

model-releasesr-localllama
23 Jul 2026
Model Releases

Running Qwen 3.6 35B MoE (Q4_K_M) on a Zeus (Xiaomi 12 Pro, 12GB RAM)

DGX agent

Shoutout to this awesome guy - https://www.reddit.com/r/LLM/s/IDUyU3v9ap Thanks to his project, BigMoeOnEdge https://github.com/Helldez/BigMoeOnEdge, I managed to successfully run a 35B MoE model on j

model-releasesr-localllama
23 Jul 2026
Model Releases

🇦🇹 Austria is rolling out a government AI-platform using Mistral models and Open WebUI

DGX agent

This is a surprisingly large real-world deployment: 'GovGPT' is part of Austria’s Public AI initiative, running on sovereign infrastructure (in their BRZ - federal datacenter) with Mistral open-weight

model-releasesr-localllama
22 Jul 2026
Model Releases

Cactus Hybrid: We taught Gemma 4 to know when it's wrong

DGX agent

Hey HN, Henry & Roman here from Cactus. A small, on-device model is fast and private, but sometimes wrong, but frontier models are getting expensive pretty fast. So, we post-trained Gemma 4 E2B post-t

model-releasesr-localllama
22 Jul 2026
Model Releases

Genesis-Science-1 (GS1), 1T open-weight model later this year from Arcee AI

DGX agent

Today the Department of Energy (DOE) and Arcee AI announced the development of Genesis-Science-1 (GS1), an open model for scientific research. This is a joint effort to bring advanced AI into scientif

model-releasesr-localllama
22 Jul 2026
Model Releases

Instead of panicking about the Hugging Face attack, people need to start questioning OpenAI's insecure sandboxes.

DGX agent

One thing I noticed in American politics, whenever the government wants to push unpopular actions or laws, they often introduce fear to convince the public to support them. This is actually how i view

model-releasesr-localllama
22 Jul 2026
Model Releases

microsoft/Fara1.5-27B · Hugging Face

DGX agent

Fara1.5-27B is a multimodal computer use agent (CUA) for web browsers, from Microsoft Research AI Frontiers. It observes the browser through screenshots and acts on the user's behalf by emitting struc

model-releasesr-localllama
22 Jul 2026
Model Releases

MindControl - llama.cpp fork to guide the reasoning process via injection during sampling

DGX agent

The primary driver of this project is that I'd become frustrated with the reasoning behavior of smaller local models such as Qwen3.6-27B (i believe particularly at lower temperatures, and where system

model-releasesr-localllama
22 Jul 2026
Model Releases

cuda: extract Q1_0 elements via __byte_perm by dfriehs · Pull Request #25628 · ggml-org/llama.cpp

DGX agent

I don't have the ability to access Reddit posts or browse specific URLs. To provide you with an accurate factual summary for your knowledge base, I would need either: 1. The actual content/text from t

model-releasesr-localllama
16 Jul 2026
Model Releases

NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B on 2x3090s

DGX agent

I managed to get this model working on 2x 3090s with full 262k ctx and N=4, if anyone is interested to try it, thanks to this quant: https://huggingface.co/danielrmay/NVIDIA-Nemotron-Labs-3-Puzzle-75B

model-releasesr-localllama
16 Jul 2026
← Previous
1…78910
Next →