AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
84,630 results
Model Releases

b10227

DGX agent

chat : add qwen3 specialized parser (#26252) Add tagged thinking tool parser chat : refactor and add permute helper cont : add support for <tool_call> omission cont : update tool delimiters cont : add

model-releasesllama-cpp-releases
2 Aug 2026
Model Releases

b10228

DGX agent

DeepseekV4 MTP + DSpark (#25784) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubunt

model-releasesllama-cpp-releases
X Post
Paper
YouTube
Reddit
GitHub
2 Aug 2026
Model Releases

b10229

DGX agent

opencl: bugfix increment ref_count in ggml_backend_opencl_init() (#26162) Incrementing ref_count at the beginning is important later in the free() method of the ggml_backend_opencl_context at program

model-releasesllama-cpp-releases
2 Aug 2026
Model Releases

b10231

DGX agent

common: support the DSpark sidecar resolution (#26458) The dspark- files resolve like the other speculative sidecars: the -hfd tag applies to them, a requested sidecar resolves without a full model at

model-releasesllama-cpp-releases
2 Aug 2026
Model Releases

b10232

DGX agent

metal: implement DeepSeek V4 hyper-connections (#26459) Implement GGML_OP_DSV4_HC_COMB, GGML_OP_DSV4_HC_PRE, and GGML_OP_DSV4_HC_POST with SIMDgroup register and shuffle optimized kernels. Add Metal d

model-releasesllama-cpp-releases
2 Aug 2026
Model Releases

b10233

DGX agent

opencl: limit local workgroup size for GLU operation (#26383) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64)

model-releasesllama-cpp-releases
2 Aug 2026
Model Releases

b10234

DGX agent

metal : add F16 support for bin ops (#26465) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework

model-releasesllama-cpp-releases
2 Aug 2026
Model Releases

b10235

DGX agent

metal : add SILU_BACK (#25982) feat(silu_back): implemented silu_back op for f32 fix(silu_back): removed redundant asserts in ggml-metal-ops.cpp function ggml_metal_op_silu_back. Website: https://llam

model-releasesllama-cpp-releases
2 Aug 2026
Model Releases

Best model <3B for multilingual understanding/ instruction following?

DGX agent

I know qwen 3.5 4b is great but a bit too large and miniPCM5 1b is great for agentic use but not so great for multilingual natural language understanding. Google eXb variants are just too big in total

model-releasesr-localllama
2 Aug 2026
Model Releases

Btw it looks like running the new DeepSeek v4 Flash through the Hermes agent is the way to go. The output files are better than those from a…

DGX agent

Mia AI Lab noted on Aug 2 2026 that running DeepSeek v4 Flash via the Hermes agent produces output files superior to those from any other harness tested. The observation was publicly shared and has ga

model-releasesnous-research--x
2 Aug 2026
Safety

Checkmate: you can’t take the harness (which is typically in large part symbolic) away from the neural model without giving up performance. …

DGX agent

Checkmate: you can’t take the harness (which is typically in large part symbolic) away from the neural model without giving up performance. HUGE victory for neurosymbolic AI, straight from @AnthropicA

safetygary-marcus--x
2 Aug 2026
Industry

Chinese VC firms are rushing to raise new funds after three years of record-low fundraising, amid renewed enthusiasm for China's tech, AI, and robotics sectors (Eleanor Olcott/Financial Times)

DGX agent

Eleanor Olcott / Financial Times: Chinese VC firms are rushing to raise new funds after three years of record-low fundraising, amid renewed enthusiasm for China's tech, AI, and robotics sectors — Mana

industrytechmeme
2 Aug 2026
Model Releases

Comfyui VRAM tracker

DGX agent

Hello! VRAM tracker is a node that track the full memory lifecycle of a comfyui run: when each weight is reserved, paged into VRAM, computed on, evicted, and freed. It renders it as an interactive HTM

model-releasesr-stablediffusion
2 Aug 2026
Model Releases

Conclusion: r/LocalLLaMA still has brilliant open-weight research, but finding it requires wading through endless benchmark drama, non-local Discussion Points and repetitive hardware flexes.

DGX agent

I let Gemma4-31b run on my laptop for like almost a day using a heavily altered pi to do a deep dive on our beloved Llama tangentially related Subreddit, and this was the conclusion. Feels pretty accu

model-releasesr-localllama
2 Aug 2026
Model Releases

Cybersecurity isn’t a fortress problem, it’s an immunity problem. Think vaccines. Eliminating pathogen is not practically possible. Vaccines…

DGX agent

Cybersecurity isn’t a fortress problem, it’s an immunity problem. Think vaccines. Eliminating pathogen is not practically possible. Vaccines don’t eliminate pathogens. They teach the immune system to

model-releasesfireworks-ai--x
2 Aug 2026
Model Releases

Deepseek v4 flash - 100-150 faster t/s in prefill/pp.

DGX agent

You have two choices here (in order of pref): Downgrade CUDA from 13.3 to 13.1 (skip 13.2 due to bugs) <- prefer this (thanks to u/fairydreaming for pointing this out) Use this vibed fork that works w

model-releasesr-localllama
2 Aug 2026
Model Releases

DeepSeek V4 Flash 0731 is impressive. It feels like a capabilities leap for a model in this size class. It is also incredibly cost efficient…

DGX agent

DeepSeek V4 Flash 0731 is impressive. It feels like a capabilities leap for a model in this size class. It is also incredibly cost efficient. The model is available in Bionic, both for running locally

model-releaseslm-studio--x
2 Aug 2026
Model Releases

DeepSeek-V4-Flash-0731 UD-IQ3_XXS about 11t/s on 1x 7900 XTX 24GB + 3x MI60 32GB + 128GB DDR4

DGX agent

Hello, Also I want to join the hype of posting token specs. CPU: 2x Intel Xeon CPU E5-2650 v4 @ 2.20GHz RAM: 2x 4 Channel 2400MHz DDR4 GPU: 1x AMD Radeon 7900 XTX 24GB 3x AMD Instinct MI60 32GB Strang

model-releasesr-localllama
2 Aug 2026
Model Releases

DeepSeek-V4-Flash-0731 UD-Q8_K_XL 17.20~ t/s on A6000 + 256GB DDR4

DGX agent

Hello everyone I want to join the hype of posting specs. CPU: AMD EPYC 74F3 24-Core RAM: 8 Channel 3200 DDR4 GPU: RTX A6000 48GB Prompt processing is in the high 70t/s (got down to mid 30t/s at 300k c

model-releasesr-localllama
2 Aug 2026
Model Releases

DeepSeek-V4-Flash-0731: When Low is higher than High

DGX agent

I decided to test a few questions against DeepSeek-V4-Flash-0731. Locally, I was running Unsloth's UD-Q2_K_XL quant. After I saw the surprising shape of the results, I tested against DeepSeek's offici

model-releasesr-localllama
2 Aug 2026
Model Releases

DeepSeek-V4-Flash 284B on 5.3GB of memory

DGX agent

Following up on my Qwen 3.6 port, I wanted to keep adding models and ended up fixing a bunch of things along the way, so it's its own engine now: Mference. Same core idea from TurboFieldfare, MoE mode

model-releasesr-localllama
2 Aug 2026
Model Releases

Don’t think I have seen more misconceptions around one model (I count 8) since the fantasies people had before GPT-5 was released. That turn…

DGX agent

Don’t think I have seen more misconceptions around one model (I count 8) since the fantasies people had before GPT-5 was released. That turned out to be a letdown; so will Astra, for people who are dr

model-releasesgary-marcus--x
2 Aug 2026
Model Releases

DSpark Benchmark Result on Deepseek v4 Flash 0731

DGX agent

TensorSharp supports DSpark on Deepseek v4 Flash 0731 now. Here is the benchmark result on 4x Nvidia A40 GPUs, cuda 12.8 with/without DSpark: Model: DeepSeek-V4-Flash-0731-UD-Q8_K_XL from https://hugg

model-releasesr-localllama
2 Aug 2026
Model Releases

Emad Mostaque @EMostaque came on PostAGI and said AI had already found 121 years of missing algebra in Einstein's equations. A billion param…

DGX agent

Emad Mostaque @EMostaque came on PostAGI and said AI had already found 121 years of missing algebra in Einstein's equations. A billion parameter model trained on nothing past 1911 got to general relat

model-releaseselon-musk--x
2 Aug 2026
Model Releases

Encrypted Clouds?

DGX agent

I love the progress happening on open models but I feel like it is kind of getting clear that hardware to run good sized models is completely unaffordable for me right now. I know that you all love Qw

model-releasesr-localllama
2 Aug 2026
Safety

exactly. math isn’t done. not at all.

DGX agent

exactly. math isn’t done. not at all. I don’t think being critical of the amazing work AI is doing in pure math is fair to @OpenAI until I can start to say why I feel it’s not yet at the level of our

safetygary-marcus--x
2 Aug 2026
Model Releases

Expert-only IQ3 requant of DeepSeek-V4-Flash-0731: better KLD than UD-IQ3_S, 1.4x decode on a CPU-spill rig

DGX agent

Hey all, tldr / who this helps: you run a mixed multi-GPU box where the experts spill to RAM, and you want to stay in the 3-bit tier instead of dropping to Q2 to make it fit. https://huggingface.co/Ta

model-releasesr-localllama
2 Aug 2026
Applications

Experts say US law is unprepared for rogue AI agents and models, as recent OpenAI and Anthropic incidents raise questions over legal liability and repercussions (Lily Hay Newman/Wired)

DGX agent

Lily Hay Newman / Wired: Experts say US law is unprepared for rogue AI agents and models, as recent OpenAI and Anthropic incidents raise questions over legal liability and repercussions — Both major A

applicationstechmeme
2 Aug 2026
Agents

Fascinating to see @ClementDelangue, CEO of @huggingface speaking on @FaceTheNation. Excellent points and solid advocacy around the benefits…

DGX agent

Fascinating to see @ClementDelangue, CEO of @huggingface speaking on @FaceTheNation. Excellent points and solid advocacy around the benefits of open AI models, which helped him defend against a rogue

agentsclem-delangue--x
2 Aug 2026
Local Ai

Got '403 Your account is currently unavailable' on Ollama Max after 2 weeks — No reply from support in 3 days. Anyone else?

DGX agent

Hey everyone, I'm reaching out to see if anyone else has run into a similar issue or if there's an alternative way to reach their team. I subscribed to the Ollama Max plan on June 17th. Everything was

local-air-ollama
2 Aug 2026
Hardware

Hermes Agent is now dramatically more efficient, especially for smaller/weaker/local models! With the help of @nvidia's Nemo Relay and sever…

DGX agent

Hermes Agent is now dramatically more efficient, especially for smaller/weaker/local models! With the help of @nvidia's Nemo Relay and several other strategies Hermes was able to identify a ton of opt

hardwarenous-research--x
2 Aug 2026
Agents

How do you test your setup?

DGX agent

We all have been there, tinkering around with models is fun but we rarely do it with research precision and issues are often subtle and hard to reproduce. There are a lot of benchmarks but running the

agentsr-localllama
2 Aug 2026
Model Releases

How well do multiple GPUs scale for LLM inference? (Trying to understand the basics)

DGX agent

Hi everyone, I’m fairly new to the multi-GPU side of local LLMs and I’m trying to understand how inference actually scales across multiple GPUs. Suppose I have a model running on a single GPU and then

model-releasesr-localllama
2 Aug 2026
Hardware

Hugging Face CEO Clément Delague says “AI is actually an opportunity to fix a lot of the cybersecurity problems” because his company used Nv…

DGX agent

Hugging Face CEO Clément Delague says “AI is actually an opportunity to fix a lot of the cybersecurity problems” because his company used Nvidia’s version of a Chinese open model to defend itself agai

hardwareclem-delangue--x
2 Aug 2026
Local Ai

I built a self-hosted studio that turns one reference photo into a curated, captioned, trained and tested LoRA — one browser tab, open source, MIT

DGX agent

I shared this tool here a week ago and the feedback shaped a big new version, so here's the full tour of what it does today. Screenshots of every screen: github.com/perfectgf/lora-dataset-studio — plu

local-air-stablediffusion
2 Aug 2026
Model Releases

I built an open-source LLM Gateway to route and fallback between local Ollama models and cloud APIs

DGX agent

Hey r/ollama 👋 If you run Ollama locally alongside cloud endpoints for agent workflows, Cursor/Windsurf, or custom scripts, managing API switching, failover logic, and context limits can get messy fas

model-releasesr-ollama
2 Aug 2026
Hardware

I pushed Kimi K3 onto one CPU with 8 GB of RAM

DGX agent

I deployed K3 on 32 H100s at work a couple of weeks ago and then got annoyed that there was no way to poke at it on my own machine. So I wrote an inference engine for it in C99. Nothing clever going o

hardwarer-localllama
2 Aug 2026
Agents

I wish I had found this sooner. Nous Research launched a FREE Hermes agent Skills Hub. 90,000+ community skills across 200+ categories. Skil…

DGX agent

I wish I had found this sooner. Nous Research launched a FREE Hermes agent Skills Hub. 90,000+ community skills across 200+ categories. Skills from OpenAI, Anthropic, HuggingFace & more. Thank me late

agentsnous-research--x
2 Aug 2026
Industry

Instead of limiting the progress of AI companies or preventing them from releasing models as proposed in the bipartisan AI Kill Switch Act, …

DGX agent

Instead of limiting the progress of AI companies or preventing them from releasing models as proposed in the bipartisan AI Kill Switch Act, Hugging Face CEO Clément Delague says he would rather see Co

industryclem-delangue--x
2 Aug 2026
Applications

Is paying artists enough to convince them to embrace AI?

DGX agent

Illustrators have spent years sounding the alarm about generative artificial intelligence startups training their models on artists' work without permission. They've pointed out how the practice is ta

applicationsthe-verge-ai
2 Aug 2026
Local Ai

Is training a WAN Lora going to fix the identity shift in I2V?

DGX agent

Using a character Lora image from Krea2 as my first frame but even for a subtle camera move there is identity loss immediately. Will training a character Lora for Wan2.2 help fix it? I've never traine

local-air-stablediffusion
2 Aug 2026
Local Ai

It looks like MiniMax H3 uses Qwen3-VL-32B as Text Encoder and has has a split Transformer

DGX agent

I'm trying to get more hints from this PR in comfyui github https://github.com/Comfy-Org/ComfyUI/pull/15210 but from what I've gathered so far it looks: It uses Qwen3-VL-32B as the encoder (50 layers

local-air-stablediffusion
2 Aug 2026
Agents

It's not time to slow down but to accelerate! The recent AI-powered cyberattacks have everyone talking about the risks of AI. We should. But…

DGX agent

It's not time to slow down but to accelerate! The recent AI-powered cyberattacks have everyone talking about the risks of AI. We should. But let's not lose sight of the bigger picture! If we work hard

agentsclem-delangue--x
2 Aug 2026
Model Releases

July 2026 newsletter

DGX agent

The June edition of my sponsors-only monthly newsletter is out. If you are a sponsor (or if you start a sponsorship now) you can access it here. This month: Accidental cyberattacks by OpenAl and Anthr

model-releasessimon-willison
2 Aug 2026
Safety

LLMs can know a task is impossible and still optimize it anyway. Ask whether to walk or drive to a car wash 50 meters away, and some models …

DGX agent

LLMs can know a task is impossible and still optimize it anyway. Ask whether to walk or drive to a car wash 50 meters away, and some models focus on distance while missing that the car itself must rea

safetygary-marcus--x
2 Aug 2026
Model Releases

MiniMax H3 is going open-weight in under 6 hours

DGX agent

here is all the info we have based on open PRs to add support to ComfyUI and HuggingFace diffusers - 33B for the main DiT and a pruned 20b variant - Qwen-3-VL-32b as the text encoder Edit- I posted cl

model-releasesr-stablediffusion
2 Aug 2026
Research

More on the pelican on the bicycle test from @simonw: https://simonwillison.net/2025/Jun/6/six-months-in-llms/ I uploaded the source here so…

DGX agent

More on the pelican on the bicycle test from @simonw: https://simonwillison.net/2025/Jun/6/six-months-in-llms/ I uploaded the source here so it's playable in the browser, forkable etc. https://karpath

researchkarpathy--x
2 Aug 2026
Research

New research from Google DeepMind. (bookmark it) SkillSmith treats model weights as an additional modality the LLM reads natively. The augme…

DGX agent

New research from Google DeepMind. (bookmark it) SkillSmith treats model weights as an additional modality the LLM reads natively. The augmented model ingests existing prefix weights alongside rich te

researchdair-ai--x
2 Aug 2026
← Previous
1…173174175176177…1764
Next →