AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “engineering”

GridTimelineEvolution
153 results
Model Releases

Recent llama.cpp updates for SYCL/Intel

DGX agent

Some fixes & boost(pp) for SYCL/Intel. Merged PRs: [SYCL] Flash Attention with XMX engine via oneDNN graph API (SDPA) on KV f16 for Xe2 ; Qwen3.6-27b-Q8_0 prefill speed up x1.21 at p=512 and x4.26 at

model-releasesr-localllama
15 Jul 2026
Research

Building a Custom Drones MuJoCo Environment [P]

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

This post likely covers the process of creating a custom drone simulation environment using MuJoCo, a physics engine commonly used in machine learning research. The project involves leveraging MuJoCo'

researchr-machinelearning
6 Jun 2026
Industry

Companies Are Using Reddit to Manipulate ChatGPT and Google AI Search

DGX agent

Peptide and hormone replacement therapy companies have been systematically posting to r/biohackers to manipulate AI-generated search answers on ChatGPT and Google through a practice called AI-engine o

industryr-chatgpt
3 Jun 2026
Local Ai

Running real-time 1080p video generation and editing on your own (Dreamverse OSS release)

DGX agent

Dreamverse is an AI video generation engine that produces 30 seconds of 1080p video in 4.5 seconds on a single GPU , making it significantly faster than existing systems like Sora. It provides a creat

local-air-stablediffusion
27 May 2026
Research

Witchcraft, fast local semantic search on top of SQLite [P]

DGX agent

Witchcraft is a Rust reimplementation of Stanford's XTR-Warp semantic search engine that uses a single-file SQLite database for storage, enabling client-side deployment. The system operates completely

researchr-machinelearning
18 May 2026
Hardware

CUDA V.13?

DGX agent

Ollama's MLX engine runs on NVIDIA GPUs via CUDA v13 on Windows and Linux. Users have reported that Ollama crashes on RTX 3060 with cuda_v13 in versions 0.13.0 through 0.15.6, while deleting the cuda_

hardwarer-ollama
2 May 2026
Local Ai

Which prompt and model do you think could be used to recreate this image as closely as possible? Do you think z-image or z-image turbo would work?

DGX agent

This Reddit post from r/StableDiffusion asks the community for recommendations on which prompt engineering techniques and AI image generation models (specifically comparing z-image and z-image turbo v

local-air-stablediffusion
26 Apr 2026
Local Ai

GLM 5.1 Feels very very very Slow on Ollama Cloud :(

DGX agent

A Reddit post discussing performance issues with GLM-5.1 when running through Ollama Cloud. GLM-5.1 is Z.AI's next-generation flagship model for agentic engineering, with significantly stronger coding

local-air-ollama
23 Apr 2026
Industry

A neat trick for better image outputs

DGX agent

A Reddit post from the r/ChatGPT community sharing a user-discovered technique for improving the quality of AI-generated images in ChatGPT, likely involving prompt engineering strategies or workflow a

industryr-chatgpt
15 Apr 2026
Research

Hosting Live session for sub 10ms retrieval by Moss (YC backed) [N]

DGX agent

Moss is a YC-backed high-performance runtime for real-time semantic search that delivers sub-10ms lookups, instant index updates, and zero infrastructure overhead, running where the agent lives — clou

researchr-machinelearning
15 Apr 2026
Local Ai

OllamaGPT Teaser

DGX agent

OllamaGPT is a self-hosted, ChatGPT-style web interface that runs locally using Ollama as its backend LLM engine, requiring no API keys or external services. It is designed for developers and power us

local-air-ollama
15 Apr 2026
Local Ai

Correct me if I’m wrong: Ollama can’t fine tune like Unsloth Studio

DGX agent

Ollama is a local inference engine designed for running pre-built LLMs on your own machine, and it does not include fine-tuning capabilities — this distinction is correct. Unsloth (and its Unsloth Stu

local-air-ollama
14 Apr 2026
Industry

Why do GPT responses feel exhausting even when they’re right?

DGX agent

A Reddit thread from r/ChatGPT exploring the user experience phenomenon where GPT responses can feel cognitively draining or over-engineered even when technically accurate — touching on issues like ve

industryr-chatgpt
14 Apr 2026
Local Ai

GLM 5.1 on Ollama

DGX agent

GLM-5.1 is Z.AI's next-generation flagship model for agentic engineering, featuring significantly stronger coding capabilities than its predecessor and achieving state-of-the-art performance on SWE-Be

local-air-ollama
13 Apr 2026
Hardware

TurboOCR: 270–1200 img/s OCR with Paddle + TensorRT (C++/CUDA, FP16) [P]

DGX agent

TurboOCR is a high-performance OCR project that combines PaddleOCR with NVIDIA TensorRT, implemented in C++ and CUDA, achieving throughput of 270–1,200 images per second using FP16 half-precision infe

hardwarer-machinelearning
13 Apr 2026
Local Ai

glm 5.1 is doing well

DGX agent

GLM-5.1 is Z.ai's next-generation flagship model for agentic engineering, built on a 754-billion parameter Mixture-of-Experts architecture with 40 billion active parameters per token, a 200,000-tok...

local-air-ollama
10 Apr 2026
Local Ai

Use the Same Model Across Ollama, LM Studio, Jan, and your Favorite Local AI Apps

DGX agent

Local AI tools such as Ollama, LM Studio, and Jan all rely on the same underlying inference engine (llama.cpp) and support compatible model formats (primarily GGUF), meaning a single downloaded mod...

local-air-ollama
9 Apr 2026
Model Releases

I built a weird, low-power llama.cpp server using an Intel N100 + RTX 5060Ti

DGX agent

Everything started with the sudden death of my old ASRock J1900. While looking for the perfect ITX replacement, I stumbled upon the Chinese CW-NAS-ADLN-K motherboard, which looked perfect on paper: In

model-releasesr-localllama
11 Aug 2026
Model Releases

I gave DeepSeek V4 Flash basic vision by training a 40M connector on 100K examples

DGX agent

I wanted to find out whether a huge text-only MoE could be given basic vision without retraining the language model itself. The short answer is yes. I froze DeepSeek V4 Flash and a 417M-parameter Moon

model-releasesr-localllama
11 Aug 2026
Model Releases

Serving Deepseek v4 Flash 0731 on 2x DGX Spark — 5-7 GB OS headroom, what would you do to lower VRAM usage and increase OS available RAM?

DGX agent

Hey all, I'm serving DSv4Flash 0731 on a cluster of 2x DGX Sparks but am running into constant issues with having almost no RAM (unified memory) left for the OS/cache and I'd love to hear the communit

model-releasesr-localllama
7 Aug 2026
Model Releases

Qwen Developers' responses from their recent Twitter/X AMA

DGX agent

Questions & Responses(in BOLD) below. Favorite question(s) moved to end of the thread with combined responses(removed duplicates). Be optimistic folks. I'm sure we're getting other models too apart fr

model-releasesr-localllama
5 Aug 2026
Model Releases

DeepSeek v4 Flash vs. Qwen3.6-27B, 3.5-122B, and Gemma 4 31B Benchmark

DGX agent

Just wanted to share my agentic coding benchmark run of DSv4F 0731 at both High and Low reasoning efforts (not Max)... I ran a 109-question subset of Aider Polyglot (the JS/C++/Python languages), base

model-releasesr-localllama
4 Aug 2026
Model Releases

'Data center in a Box (on Wheels)' 256Gb VRAM/512Gb RAM AI Server 6-8 Month Operational Review, Stability Write Up, Benchmarks

DGX agent

I've been out of these forums for awhile but I figured I would provide a formal update on how this has been going now that it has some operation time under its belt, just to put the information out th

model-releasesr-localllama
3 Aug 2026
Model Releases

[Release] WinterMix — Qwen3.5-122B-A10B in native MLX: an 82 GiB build that beats 94–95 GiB quants, plus a 68 GiB build for agent swarms

DGX agent

TL;DR: I spent 9 days developing a new quantization method for MLX models and measured 18 variants against each other on a single M5 Max MacBook Pro (128 GB). The result is the best-measuring MLX quan

model-releasesr-localllama
2 Aug 2026
Model Releases

Is it just me, or are current LLM benchmarks failing to capture actual usability? (Gemma 4 vs. Gemini/Claude Opus)

DGX agent

Disclaimer, this was kinda written with AI (Gemma 4 again) but it also did really well here, it outputted what I wanted, when I asked it to refine stuff or improve on certain areas it did that without

model-releasesr-localllama
31 Jul 2026
Local Ai

Built a local-first workflow automation platform around Ollama - now at v0.11.0

DGX agent

I've been working on an open-source workflow automation platform over the past few months, with Ollama as one of the first-class providers rather than treating it as an afterthought. The goal wasn't t

local-air-ollama
27 Jul 2026
Local Ai

I implemented the YOLO26n model inference from scratch using ARM64 Assembly Language (No framework) [P]

DGX agent

This was my Bachelor's Final Project: implementing YOLO26n inference completely from scratch using ARM64 Assembly Language and C, without relying on existing inference frameworks. The goal was to unde

local-air-machinelearning
26 Jul 2026
Model Releases

CachyLLama: llama.cpp fork with persistent SSD-backed KV caching for local agent workflows

DGX agent

If you run local agentic coding harnesses (Aider, Claude Code, etc.), prompt evaluation usually eats up most of your execution time. Every turn re-evaluates thousands of identical prefix tokens_system

model-releasesr-localllama
25 Jul 2026
Model Releases

Cactus Hybrid: We taught Gemma 4 to know when it's wrong

DGX agent

Hey HN, Henry & Roman here from Cactus. A small, on-device model is fast and private, but sometimes wrong, but frontier models are getting expensive pretty fast. So, we post-trained Gemma 4 E2B post-t

model-releasesr-localllama
22 Jul 2026
Model Releases

Kimi K3 in the next few hours. Deepseek V4 GA later in the week. New Liquid models. New Mistral models sometime this month. And some rumours suggest GLM 5.5 is coming in August. Openweight AI is eating good.

DGX agent

dam bois we eating good this week ngl, The velocity of the open_weight ecosystem right now is hitting a point where proprietary, closed-source APIs are losing their leverage on compute intelligence. W

model-releasesr-localllama
14 Jul 2026
Agents

How to make Codex (or any agent) do your work without any instructions (it learns by watching you!). Open-source

DGX agent

This Reddit post from r/ChatGPT discusses a technique for enabling OpenAI's Codex (or similar AI coding agents) to autonomously replicate a user's workflow by observing and learning from their actions

agentsr-chatgpt
14 Apr 2026
Industry

I'm using ChatGPT to create an alchemy game where every combination is possible!

DGX agent

A Reddit user shares their project of building an AI-powered alchemy game inspired by titles like Little Alchemy, using ChatGPT to dynamically generate the result of every possible element combination

industryr-chatgpt
14 Apr 2026
Research

[D] Will Google’s TurboQuant algorithm hurt AI demand for memory chips? [D]

DGX agent

This r/MachineLearning discussion centers on Google's TurboQuant, a training-free KV cache compression algorithm released in March 2026 that compresses cache storage from 16 bits down to 3 bits with m

researchr-machinelearning
12 Apr 2026
Research

Educational PyTorch repo for distributed training from scratch: DP, FSDP, TP, FSDP+TP, and PP [P]

DGX agent

This Reddit post shares an educational PyTorch repository designed to teach distributed training techniques from the ground up, covering Data Parallelism (DP), Fully Sharded Data Parallel (FSDP), Tens

researchr-machinelearning
12 Apr 2026
Research

FlashAttention (FA1–FA4) in PyTorch - educational implementations focused on algorithmic differences [P]

DGX agent

This r/MachineLearning post presents educational PyTorch implementations of FlashAttention versions 1 through 4, designed to highlight the key algorithmic differences across each iteration rather than

researchr-machinelearning
11 Apr 2026
Research

What if your HNSW index stored 3-bit embeddings instead of float32? [R]

DGX agent

A research paper (arXiv:2601.11557) proposes replacing the dominant 'HNSW + float32 + cosine similarity' vector database stack with an information-theoretic alternative that uses Maximally Informat...

researchr-machinelearning
11 Apr 2026
Research

[D] 60% MatMul Performance Bug in cuBLAS on RTX 5090 [D]

DGX agent

A bug was identified in NVIDIA's cuBLAS library where `cublasSgemmStridedBatched` dispatches the same suboptimal `cutlass_80_simt_sgemm_128x32_8x5` kernel for every batched FP32 workload from 256×...

researchr-machinelearning
10 Apr 2026
Research

Anyone have an S3-compatible store that actually saturates H100s without the AWS egress tax? [R]

DGX agent

A r/MachineLearning discussion explores the challenge of finding S3-compatible object storage that can both saturate H100 GPU bandwidth and avoid AWS's egress fees during high-throughput AI trainin...

researchr-machinelearning
9 Apr 2026
Research

[P] Building a LLM from scratch with Mary Shelley's 'Frankenstein' (on Kaggle)

DGX agent

A beginner-friendly tutorial demonstrating how to build a ~3.2M parameter LLM from scratch using Mary Shelley's *Frankenstein* as the sole training corpus, designed to run on Kaggle's free GPU in u...

researchr-machinelearning
8 Apr 2026
Model Releases

Hidden Reasoning from Claude and GPT are Decoded, and it is interesting

DGX agent

Yesteday a paper showed a gap that allows to see 100% of the reasoning tokens form ALL Claude and GPT models Stealing Reasoning Traces from Proprietary LLM APIs. check it out, they have published lots

model-releasesr-localllama
12 Aug 2026
Local Ai

RAG for regular users?

DGX agent

One of the reasons I got into local LLMs was the possibility of getting answers using my own documents and books (a few hundreds) instead of having to search through them manually. However since I'm n

local-air-localllama
12 Aug 2026
Model Releases

I put Gemma 4 E4B and E2B into an e-reader so I can ask my weird questions and share my thoughts in private directly in app.

DGX agent

Here's how it works in the app: Framework: Runs on LiteRT-LM (like Google's AI Edge). Models: Downloads either the E2B (~2.5 GB) or E4B (~3.6 GB) INT4 quantized models directly from ungated litert-com

model-releasesr-localllama
11 Aug 2026
Model Releases

I ran Muse Glimmer @ 1M context - All tests passed.

DGX agent

Heeeey all! I just completed some fun tests with Muse Glimmer, I thought I'd let you know. In fact, the summary below was written by Muse itself! I ran a 2× DGX Spark cluster and got Meta's day-old Mu

model-releasesr-localllama
11 Aug 2026
Model Releases

[llama.cpp PR #26608] Ling-3.0 support (unmerged)

DGX agent

aetherbird has done some great work getting Ling-3.0 to work in llama.cpp. The architecture is generally identical to deepseekv2. I recently added a microscopic 40 line PR to his that adds support for

model-releasesr-localllama
11 Aug 2026
Local Ai

Best Local LLMs - August 2026

DGX agent

Wowee!! Just when you thought it couldn't get better for open weight models, we probably have had our best period yet!?!?! Models that rival the closed frontier, Opus level models on non-insane hardwa

local-air-localllama
10 Aug 2026
Model Releases

Comparing how Cline, Kilo, and Qwen Code handle long-task context/state (and why context loops keep happening)

DGX agent

I've been comparing Cline / Kilo / Qwen Code lately since they all handle long-task state differently. Cline: has Focus Chain, a markdown file kept outside the conversation that gets reinjected on a c

model-releasesr-localllama
10 Aug 2026
Local Ai

I Turned My Underused Gaming Laptop Into a Local AI Workstation

DGX agent

TL;DR: I am building a Windows-first local AI setup for people who want to try local LLMs without spending days choosing models, setting up Ollama, Docker, WSL, Open WebUI, agents, and tool permission

local-air-ollama
9 Aug 2026
Model Releases

Showoff Saturday: Local 4x 6000 Pro (multi-year progression)

DGX agent

Not the biggest or shiniest, but it's mine From gaming machine inference on the original llama models, to a 4x RTX 6000 Pro Max Q + 4x 3090s local AI cluster. Pictures are in reverse chronological ord

model-releasesr-localllama
8 Aug 2026
← Previous
1234
Next →