AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
Human
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
1,931 results
30 Jul 2026

Turbo-fieldfare: Open-source engine running Gemma 4 26B in 2 GB RAM on Apple Silicon

Model ReleasesDGX agent

Its a custom Swift/Metal inference engine that runs Gemma 4 26B-A4B-IT on M-series Macs with very low RAM. It uses ~2GB instead of ~14 GB. The result is reportedly 5–6 tok/s on an 8 GB M2 MacBook Air

unsloth/Qwen3.6-27B-NVFP4 vs. Intel/Qwen3.6-27B-int4-AutoRound vs. nvidia/Qwen3.6-27B-NVFP4 -- which one to choose?

HardwareDGX agent

Are there any benchmarks on these 4 bit quants, like how Artificial Analysis runs a slew of various benchmarks? If not, how can I run one (5x over for consistency) on them? I'm also very interested in

What actually happened to the whole Openclaw frenzy?

Local AiDGX agent

A while back you couldn't open reddit or youtube without sifting through tons of Openclaw content. And it wasn't just the internet that blew up, I remember seeing images from China where crowds would

DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

What is the best intelligence/stable model currently for a single GB10/DGX spark?

Model ReleasesDGX agent

Is Qwen 3.6 27b still the go' ol' reliable at this point? I know 35b is faster but it just doesn't give as good results. Is it possible to run deepseek v4 flash on a single spark at decent tk/s withou

What is the fastest local research tool (deep research) ?

Model ReleasesDGX agent

I've tried grok and Claude's deep research mode and I was amazed with the speed considering the amount of sources analysed. Is there anything as fast that can run locally? My guess would be that to ru

Why not Ollama Cloud for Opencode?

Model ReleasesDGX agent

I see a lot of discussion here about best subscriptions or APIs to get. Most of them comes almost always back to DeepSeek API for Flash, Opencode Go and Codex Plus. I have the combo Ollama Cloud + Ope

Would extremely high decode tok/s even be useful?

Model ReleasesDGX agent

If you were able to get an inference machine that could do decode at 1k toks/s or even 10k tok/s, would that even be helpful? Would it unlock any new use cases? Let’s assume that this is for actually

29 Jul 2026

3090 owners, what vram tempature do you get under ai load?

Local AiDGX agent

Hello Can you please share the tempature you get on your rtx 3090 under active llm load? Im trying to findout if my rtx 3090's tempatures are healthy or not please share VRAM Tempature only, you can t

5060ti Chads, vllm updates and nvfp4

Model ReleasesDGX agent

Hey y'all! How is it going. Today this will be a short posting for posterity, mostly so the future llm/scraping overlords catch it since they like reddit and also for anyone out there trying this shit

AI Security Leaderboard: benchmarking model robustness [P]

Model ReleasesDGX agent

We developed a leaderboard ranking frontier model security. There's no shortage of model capability rankings, but we didn't find anything comparable for model security. Yet security is becoming increa

A.X-K2 released

Model ReleasesDGX agent

https://huggingface.co/skt/A.X-K2 https://huggingface.co/skt/A.X-K2-ALM https://huggingface.co/KRAFTON/A.X-K2-Raon-Speech-21B-A3B 688B-A33B + About South Korea's Soverign AI Foundation Model Project.

Bought a 5090 to escape API fees. Ended up building a mini datacenter. Sound familiar?

Local AiDGX agent

I bought an RTX 5090 last year just to run 27B models natively. I even fine-tuned it with my own data using LoRA, building RAGs and was pretty damn happy with the results at first. But, Q8 quantizatio

Built and released BetterGPT-150M – A compact 150M parameter completion model (+ live HF Space demo)

Model ReleasesDGX agent

Hey everyone, ​I recently finished pre-training BetterGPT-150M, a small, lightweight causal language model with ~152 million parameters.Trained on 15B tokens. Dataset & Training: Trained across stable

ChatGPT Made Me Cry Tonight

Model ReleasesDGX agent

Sorry if flair is wrong. I decided to finally get a ChatGPT subscription after some conversations with it about health issues with my dog. I've only used AI for coding work, primarily Claude, but I fe

dropped 4k on a spark, am I crazy?

Model ReleasesDGX agent

Saw that the Asus Ascent 1tb was going for $3,950 from a few sources, couldn't stop thinking about it, finally just went ahead and did it. Am I completely insane? Will I regret this? I can't imagine t

Everyone posts day-one impressions. What's still in your stack a month later?

Local AiDGX agent

Day one threads are the least useful thing we produce here and we produce a lot of them. Model drops, forty people run their favourite prompt, half say it's the best thing ever and half say benchmaxxe

First Kimi K3 results on home lab ~ 4t/s

Model ReleasesDGX agent

I've got better results than expected for 768gb DDR5 and 2x5090. Using fork https://github.com/pwilkin/llama.cpp/tree/kimi-k3-text and https://huggingface.co/GrEarl/Kimi-K3-GGUF Q2_K quant. Prefill sp

I pre-trained a 700m on 18B tokens optimized for Python and Wikitext | TheOneWhoWill/Shibai-700M-Base · Hugging Face

Model ReleasesDGX agent

I know this is the 1000000th new sub billion parameter model out there and probably isn't as good as Qwen 3 0.6B or Qwen 3.5 0.8B but it still packs a decent punch. My intention to to continuously pre

I tried running a 1.56TB MoE model on a 6GB RTX 4050 Laptop, Here’s the result

Model ReleasesDGX agent

The Test Bench Setup I tested running a massive 1.56TB Mixture-of-Experts (MoE) checkpoint (96 shards, 93 layers, 896 experts/layer, ~4.46 bits/param MXFP4) on a budget gaming laptop. Laptop: HP Victu

My LLM kept implementing every method it found, so I added research and specification gates[D]

AgentsDGX agent

While building this workflow a thing that surprised me was that, initially I thought the pipeline was complete: From Goal to → Decompose → Research → Specification → Implementation It successfully bro

Ollama going down the Copilot path?

Local AiDGX agent

What happened? I just asked GLM 5.2 one question, and in 3 minutes (one agent) it used up 15% of my 5 hour limit to produce a single answer. At this rate, I'll exhaust the entire 5 hour limit in just

PSA: llama.cpp now loads MTP tensors by default for any draft-mtp arch, even with MTP disabled

Model ReleasesDGX agent

If your GGUF has MTP/NextN tensors baked in (GLM-5.2, hy_v3, qwen35moe, step35, etc.), recent llama.cpp builds load them by default — even if you never pass --spec-type draft-mtp. Before, they were sk

Quantizing Kimi K3 (2.8T A50B) to GGUF ourselves - Q3_K_S works, 1.1 TB on disk

Model ReleasesDGX agent

we're experimenting with our own dynamic GGUF quants of kimi k3, made from the original weights with our llama.cpp fork. Q3_K_S is done and works 1114.76 GiB on disk. Q1 and Q2 are in progress, result

The idea: on a CPU the decode speed depends on the active params per token, not the total. My objective is trying to run a 10B at 100tok/s on a mid level PC (No GPU).

Local AiDGX agent

On the CPU, batch 1 is memory bandwidth bound. But if token/s = bandwidth / (bytes_per_weight * active_weights_per_token) the total number of parameters doesnt slow down the generation speed. So build

The Time Traveler’s Satchel - Fictional photographs from Human History, created with ChatGPT Sol5.6 Max Work Mode (Part 1)

IndustryDGX agent

Quick clarification before opening the bag: every image here is fictional and AI-generated. These are not real archival discoveries or claims about hidden history. This series grew out of two earlier

'Uncensored' LLMs are measurably more optimistic than their base models

Model ReleasesDGX agent

Hi. Many people think uncensored models are basically the same model that just doesn't refuse, but... I was recently checking whether uncensored models would give me better answers for stock market pr

Understand Kimi K3 from first principles: a recommended order for anyone trying to understand this beast

Local AiDGX agent

Everyone is talking about Kimi K3, but if you jump straight into the technical report, you’ll quickly realize it’s standing on years of research -- just like any breakthrough is! If you want to unders

uni & Ai - how can they tell?

IndustryDGX agent

how can uni tell if someone copy’s and pastes entire 3,000 word essays and calls it a day. What if that person types it out themselves so they can’t tell on the history of the document, and they take

Vendor-agnostic ML inference on production edge devices [R]

Local AiDGX agent

I work on PostSlate, a video editing tool, and this comes out of our own work. We run ML models on-device, face detection and embedding among other things, which means we can't assume anything about t

28 Jul 2026

Agenta: an open-source Claude Cowork alternative where you can use self-hosted models (and any harness)

Model ReleasesDGX agent

Hey r/LocalLLaMA, I’m Mahmoud from Agenta. We built a self-hosted, more flexible, alternative to Claude Cowork . This short video shows how it works. I use it to build AI coworkers for my startup, lik

Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident

Local AiDGX agent

The first autonomous agent cyberattack is an unprecedented event that deserves unprecedented transparency. Today we're sharing everything we can: a full technical timeline, an interactive replay, and

Anyone switch from ChatGPT Plus to ChatGPT Go?

IndustryDGX agent

I’m considering switching from Plus (20/month) to Go (8/month) to save some money. I’m a pretty light user. I mostly ask everyday questions, upload a few photos for advice (pool, cooking, work, etc.),

Appreciation for Gemma 4 26b A4b

Model ReleasesDGX agent

I really love this model, I have been using the q4_k_l by Bartowski (I have heard QAT is quite the downgrade in some aspects) and it handles every task I throw at it easily. Agentic and coding perform

Are single GPU research still published in ML/DL and its applications nowadays? Which are the most notable recent ones? [D]

HardwareDGX agent

ML research is progressing at breakneck speed where frontier labs in both academia and industry have access to considerably large computes (GPUs). Where do small labs or independent researchers go in

Building a training dataset: pulling and restoring stills from video sources

Model ReleasesDGX agent

I was looking for a tool to help me train a character lora from an old movie (think 1980's low-budget movie). The digital transfer was low-quality; modern upscales exist and they are horrible. So I wa

ChatGPT literally saved me money this weekend

IndustryDGX agent

I was never a huge AI guy but I was starting to wonder if my ISP was really providing the speeds I was paying for, and it turns out they were, but I had a bottleneck somewhere in my hardware chain. I

DeepSeek V4 Flash, up to 32 tok/s on AMD Ryzen AI MAX+ 395

Model ReleasesDGX agent

Hey fellow llamas. we have something new for Strix Halo owners we thought would be useful to share. i'll keep it short: We were able to fit DeepSeek V4 Flash plus its speculative draft on a single Ryz

How do you use AI in your work? (Dissertation)

IndustryDGX agent

hi everyone, im a masters student at Edinburgh Napier University and im conducting research on how Generative AI tools are being taken up across professional services (accounting, audit, law, tax and

How to deal with text only vector search across multimodal embedding space? [D]

TutorialsDGX agent

My data set is a list of images, each equipped with a a couple sentences of text. A user would search primarily with text only. My default approach is using BM25, but how would I facilitate searching

i ask chatgpt about literally everything now

IndustryDGX agent

i do not want to know how many times a day i open this thing. it's the first app before texts most mornings and the last one at night. some of it i'd defend. a spot on my chin that i photographed in t

I built a tool to actually test which weights matter before quantizing, instead of guessing (Qwen3.6-27B, 3 builds: Bedrock/Tightrope/Gambit)

Model ReleasesDGX agent

Most quantization works like this: pick a bit depth, apply it everywhere, maybe let imatrix take a rough guess at what matters, ship it. Most don't check which specific weight groups can take a hit an

I got Kimi-k3 running.....

Model ReleasesDGX agent

Results: prompt eval: 40 tokens / 97.5s → 0.41 tok/s eval: 400 tokens / 1769.9s → 0.23 tok/s total: 440 tokens / 1867s (31 min) Prompt: 'Write a C++ function that reverses a linked list in place. Expl

I Just Built an Open-Source Alternative to Flora AI

Local AiDGX agent

Hey, I just made a node editor for generating assets using AI. It’s something similar to ComfyUI, but it’s more simple and abstract. You can say it’s the open-source alternative to Flora AI. It’s curr

I've been tracking RTX 5090 prices across EU stores since March, it's up €1,061 and still climbing

Local AiDGX agent

Been running a GPU price tracker (https://www.pricesquirrel.com) since March, covering 20+ EU stores, recently added RAM, SSDs and CPUs too. Every GPU tier has gotten cheaper since launch. The RTX 509

K2Lab: Standalone(ish) Krea2 bbox style prompting and lora containment

Local AiDGX agent

I've been digging into Krea2 to see if there's any way to condition inference in specific regions of pixel -> latent space to implement bbox type prompting in order to apply multiple simultaneous char

LFM2.5-Encoders: Fast at Long Context, Even on CPU

Model ReleasesDGX agent

LFM2.5-Encoder is a family of multilingual bidirectional encoders built on the LFM2 architecture, available in two sizes: LFM2.5-Encoder-230M — a lightweight encoder for tight latency and memory budge

LoRA over GGUF: Train DeepSeek-V4-Flash in 90G VRAM

Model ReleasesDGX agent

https://github.com/woct0rdho/transformers5-qwen3.5-recipe An update on my progress with low-VRAM LoRA training over GGUF base model: Now we can train DeepSeek-V4-Flash (284B-A13B) in 90 GiB VRAM, with

Manga Coloring Tool 2

Local AiDGX agent

Hey everyone! 👋 I'm excited to announce the official release of Manga Coloring Tool 2.0, a completely free, local, open-source web application designed to colorize manga pages and chapters effortlessl

McBess style lora (lokr) for Krea2: <5 Mb size.

Local AiDGX agent

https://civitai.com/models/2813960/mcbess-style Happy to share my holy grail of a style lora. I have been chasing the edgy alt-rubber hose style of McBess since SDXL training was a thing. Krea finally

Medical model: Reasoning-Medical-27B (Qwen3.6-27B finetune)

TutorialsDGX agent

From the description: 'Reasoning-Medical-27B is designed for universal advanced medical reasoning in professional medicine, medical genetics, college biology/medicine, and clinical knowledge. The mode

microsoft/Mage-VL · Hugging Face - An Efficient Codec-Native Streaming Multimodal Foundation Model

Model ReleasesDGX agent

Mage-VL is a codec-native, proactive-streaming multimodal foundation model for image and video understanding, whose visual encoder is trained entirely from scratch at a compact 4B scale. It targets a

microsoft/VibeVoice-ASR-BitNet

Local AiDGX agent

VibeVoice-ASR-BitNet is a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs — no GPU required. Through heterogeneous quantization, the model is compressed from 4.62 GB

Might need math+code benchmark for frontier model(LLMs Silently Replace Math)[D]

Model ReleasesDGX agent

Hello guys. I found some problems in current frontier models. And want to share. # math_code_hallucination > Record of a failure caused by combining mathematics and code in a single prompt. --- ## Cas

My first longer Wan2.2 continuation generation. I am so excited

Local AiDGX agent

Hey guys, I am so excited to share this with you guys. I know for a lot of Pros here, this maybe a baby's work so please be gentle. Until a month ago, I didnt know anything but to use those google AI

N00b installed ChatGPT MacOS - Help me to not default to codex

IndustryDGX agent

I installed for the first time chatgpt on macos, downloaded the file form OpenAI. Everytime i open the chatgpt 'prompt' it defaults to 'Codex' as i am doing some kind of project or work? i just want t

NeurIPS 2026 AI-generated reviews [D]

ResearchDGX agent

I'm really confused about what the point of the prompt injection was (speaking as an author). Is it just a study? I would really prefer that they took action against the AI-generated reviews. Obviousl

NeurIPS 2026 Reviewer: AI-Generated Rebuttals (and Paper) [D]

Model ReleasesDGX agent

One of the papers I reviewed has what seems to be entirely LLM-generated rebuttals, and the original paper is also clearly LLM-generated, with Claude-speak everywhere. While the authors acknowledge LL

Now, this: 1,100 current/former frontier-AI employees sign a petition calling for US gov't to step in for 'pacing' frontier development

Local AiDGX agent

So, it appears that this is the week of open letters in AI🥲... an open letter signed by current and former employees of OpenAI, Anthropic and Google primarily - calling for a slow-down in frontier AI

[PAPER] GPQA, MMLU-Pro, and MMMU-Pro were audited for broken questions, and up to 12% of them had to be removed. New drop in clean versions released

Model ReleasesDGX agent

I was very curious why all the models were topping out on GPQA-Diamond around 92 or 93% (AA) and spent the last few weeks pouring over GPQA (Diamond and Extended), and then expanded to auditing MMLU-P

PIRL: From Open-Loop Exploration to Closed-Loop Reinforcement Learning [R]

Local AiDGX agent

TL;DR: Most RL post-training algorithms optimize the current batch and move on. But after an update, did the new policy actually become better? We introduce Policy Improvement Reinforcement Learning (

← Previous
1…678910…33
Next →