AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
84,570 results
3 Aug 2026

You shouldn't need a vision model to know your PDF has checkboxes. LiteParse can now pull structured data directly from your PDFs: form fiel…

Model ReleasesDGX agent

You shouldn't need a vision model to know your PDF has checkboxes. LiteParse can now pull structured data directly from your PDFs: form field values, checkbox states, annotations, embedded images, vec

YouTuber Hank Green faces online criticism after using ChatGPT to help research a script, and says his LLM usage 'is not healthy for me or good for the world' (Anthony Ha/TechCrunch)

IndustryDGX agent

Anthony Ha / TechCrunch: YouTuber Hank Green faces online criticism after using ChatGPT to help research a script, and says his LLM usage “is not healthy for me or good for the world” — Hank Green, a

Zero-Mem: Zero-Token Memory Operations for LLM Agents

Local AiDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub

arXiv:2607.29377v1 Announce Type: new Abstract: LLM agents need memory to act consistently over long interactions, yet many systems use additional LLM calls to operate that memory. Generating intermed

ZeroR@CHiPSAL 2026: Two-Stage Vision-Language Adaptation with Contrastive Learning for Nepali Meme Classification

ResearchDGX agent

arXiv:2607.28637v1 Announce Type: new Abstract: This paper presents our system for the CHiPSAL 2026 shared task on multimodal hate speech and sentiment detection in Nepali memes. We address both subta

2 Aug 2026

A profile of Jacob Tsimerman, who won the Fields Medal last week and is taking a leave from the University of Toronto to join OpenAI and work on AI safety (Ben Cohen/Wall Street Journal)

SafetyDGX agent

Ben Cohen / Wall Street Journal: A profile of Jacob Tsimerman, who won the Fields Medal last week and is taking a leave from the University of Toronto to join OpenAI and work on AI safety — Jacob Tsim

Alibaba says its 2.4T-parameter Qwen3.8-Max tops Moonshot's Kimi K3 on some benchmarks, and it plans to release Qwen3.8-Max and Qwen3.8-27B's weights next week (Luz Ding/Bloomberg)

Model ReleasesDGX agent

Luz Ding / Bloomberg: Alibaba says its 2.4T-parameter Qwen3.8-Max tops Moonshot's Kimi K3 on some benchmarks, and it plans to release Qwen3.8-Max and Qwen3.8-27B's weights next week — Alibaba Group Ho

All other models on the Portal remain 20% discounted, aside from GPT-5.6 Terra and Luna which are 50% off. https://x.com/NousResearch/status…

Model ReleasesDGX agent

All other models on the Portal remain 20% discounted, aside from GPT-5.6 Terra and Luna which are 50% off. https://x.com/NousResearch/status/2080039066771337475?s=20 All models are now 20% off for a l

Another day and another full frontier model running on your computer. Been teething DeepSeek V4 Flash on over 60 employees at The Zero Human…

Model ReleasesDGX agent

Another day and another full frontier model running on your computer. Been teething DeepSeek V4 Flash on over 60 employees at The Zero Human Company for a few hours and it is stunning. Between Kimi K3

Are you ready for Le Chaton FAT or still wasting money on GPUs?

Local AiDGX agent

According to rumors (spread by myself) Le Chaton FAT will be 26T-a3b and I AM READY for it. Let's be real, I can't afford that many 5060Ti, so I got 12x Gen 4 3.2 TB (two per card). This gives me abou

b10224

Model ReleasesDGX agent

ggml-webgpu: add support for f16 repeat (#26307) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramew

b10225

Model ReleasesDGX agent

model : load MiMo V2 MTP tensors only if used (#26412) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XC

b10226

Model ReleasesDGX agent

sycl: fix classification of iGPUs (#26105) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Li

b10227

Model ReleasesDGX agent

chat : add qwen3 specialized parser (#26252) Add tagged thinking tool parser chat : refactor and add permute helper cont : add support for <tool_call> omission cont : update tool delimiters cont : add

b10228

Model ReleasesDGX agent

DeepseekV4 MTP + DSpark (#25784) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubunt

b10229

Model ReleasesDGX agent

opencl: bugfix increment ref_count in ggml_backend_opencl_init() (#26162) Incrementing ref_count at the beginning is important later in the free() method of the ggml_backend_opencl_context at program

b10231

Model ReleasesDGX agent

common: support the DSpark sidecar resolution (#26458) The dspark- files resolve like the other speculative sidecars: the -hfd tag applies to them, a requested sidecar resolves without a full model at

b10232

Model ReleasesDGX agent

metal: implement DeepSeek V4 hyper-connections (#26459) Implement GGML_OP_DSV4_HC_COMB, GGML_OP_DSV4_HC_PRE, and GGML_OP_DSV4_HC_POST with SIMDgroup register and shuffle optimized kernels. Add Metal d

b10233

Model ReleasesDGX agent

opencl: limit local workgroup size for GLU operation (#26383) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64)

b10234

Model ReleasesDGX agent

metal : add F16 support for bin ops (#26465) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework

b10235

Model ReleasesDGX agent

metal : add SILU_BACK (#25982) feat(silu_back): implemented silu_back op for f32 fix(silu_back): removed redundant asserts in ggml-metal-ops.cpp function ggml_metal_op_silu_back. Website: https://llam

Best model <3B for multilingual understanding/ instruction following?

Model ReleasesDGX agent

I know qwen 3.5 4b is great but a bit too large and miniPCM5 1b is great for agentic use but not so great for multilingual natural language understanding. Google eXb variants are just too big in total

Btw it looks like running the new DeepSeek v4 Flash through the Hermes agent is the way to go. The output files are better than those from a…

Model ReleasesDGX agent

Mia AI Lab noted on Aug 2 2026 that running DeepSeek v4 Flash via the Hermes agent produces output files superior to those from any other harness tested. The observation was publicly shared and has ga

Checkmate: you can’t take the harness (which is typically in large part symbolic) away from the neural model without giving up performance. …

SafetyDGX agent

Checkmate: you can’t take the harness (which is typically in large part symbolic) away from the neural model without giving up performance. HUGE victory for neurosymbolic AI, straight from @AnthropicA

Chinese VC firms are rushing to raise new funds after three years of record-low fundraising, amid renewed enthusiasm for China's tech, AI, and robotics sectors (Eleanor Olcott/Financial Times)

IndustryDGX agent

Eleanor Olcott / Financial Times: Chinese VC firms are rushing to raise new funds after three years of record-low fundraising, amid renewed enthusiasm for China's tech, AI, and robotics sectors — Mana

Comfyui VRAM tracker

Model ReleasesDGX agent

Hello! VRAM tracker is a node that track the full memory lifecycle of a comfyui run: when each weight is reserved, paged into VRAM, computed on, evicted, and freed. It renders it as an interactive HTM

Conclusion: r/LocalLLaMA still has brilliant open-weight research, but finding it requires wading through endless benchmark drama, non-local Discussion Points and repetitive hardware flexes.

Model ReleasesDGX agent

I let Gemma4-31b run on my laptop for like almost a day using a heavily altered pi to do a deep dive on our beloved Llama tangentially related Subreddit, and this was the conclusion. Feels pretty accu

Cybersecurity isn’t a fortress problem, it’s an immunity problem. Think vaccines. Eliminating pathogen is not practically possible. Vaccines…

Model ReleasesDGX agent

Cybersecurity isn’t a fortress problem, it’s an immunity problem. Think vaccines. Eliminating pathogen is not practically possible. Vaccines don’t eliminate pathogens. They teach the immune system to

Deepseek v4 flash - 100-150 faster t/s in prefill/pp.

Model ReleasesDGX agent

You have two choices here (in order of pref): Downgrade CUDA from 13.3 to 13.1 (skip 13.2 due to bugs) <- prefer this (thanks to u/fairydreaming for pointing this out) Use this vibed fork that works w

DeepSeek V4 Flash 0731 is impressive. It feels like a capabilities leap for a model in this size class. It is also incredibly cost efficient…

Model ReleasesDGX agent

DeepSeek V4 Flash 0731 is impressive. It feels like a capabilities leap for a model in this size class. It is also incredibly cost efficient. The model is available in Bionic, both for running locally

DeepSeek-V4-Flash-0731 UD-IQ3_XXS about 11t/s on 1x 7900 XTX 24GB + 3x MI60 32GB + 128GB DDR4

Model ReleasesDGX agent

Hello, Also I want to join the hype of posting token specs. CPU: 2x Intel Xeon CPU E5-2650 v4 @ 2.20GHz RAM: 2x 4 Channel 2400MHz DDR4 GPU: 1x AMD Radeon 7900 XTX 24GB 3x AMD Instinct MI60 32GB Strang

DeepSeek-V4-Flash-0731 UD-Q8_K_XL 17.20~ t/s on A6000 + 256GB DDR4

Model ReleasesDGX agent

Hello everyone I want to join the hype of posting specs. CPU: AMD EPYC 74F3 24-Core RAM: 8 Channel 3200 DDR4 GPU: RTX A6000 48GB Prompt processing is in the high 70t/s (got down to mid 30t/s at 300k c

DeepSeek-V4-Flash-0731: When Low is higher than High

Model ReleasesDGX agent

I decided to test a few questions against DeepSeek-V4-Flash-0731. Locally, I was running Unsloth's UD-Q2_K_XL quant. After I saw the surprising shape of the results, I tested against DeepSeek's offici

DeepSeek-V4-Flash 284B on 5.3GB of memory

Model ReleasesDGX agent

Following up on my Qwen 3.6 port, I wanted to keep adding models and ended up fixing a bunch of things along the way, so it's its own engine now: Mference. Same core idea from TurboFieldfare, MoE mode

Don’t think I have seen more misconceptions around one model (I count 8) since the fantasies people had before GPT-5 was released. That turn…

Model ReleasesDGX agent

Don’t think I have seen more misconceptions around one model (I count 8) since the fantasies people had before GPT-5 was released. That turned out to be a letdown; so will Astra, for people who are dr

DSpark Benchmark Result on Deepseek v4 Flash 0731

Model ReleasesDGX agent

TensorSharp supports DSpark on Deepseek v4 Flash 0731 now. Here is the benchmark result on 4x Nvidia A40 GPUs, cuda 12.8 with/without DSpark: Model: DeepSeek-V4-Flash-0731-UD-Q8_K_XL from https://hugg

Emad Mostaque @EMostaque came on PostAGI and said AI had already found 121 years of missing algebra in Einstein's equations. A billion param…

Model ReleasesDGX agent

Emad Mostaque @EMostaque came on PostAGI and said AI had already found 121 years of missing algebra in Einstein's equations. A billion parameter model trained on nothing past 1911 got to general relat

Encrypted Clouds?

Model ReleasesDGX agent

I love the progress happening on open models but I feel like it is kind of getting clear that hardware to run good sized models is completely unaffordable for me right now. I know that you all love Qw

exactly. math isn’t done. not at all.

SafetyDGX agent

exactly. math isn’t done. not at all. I don’t think being critical of the amazing work AI is doing in pure math is fair to @OpenAI until I can start to say why I feel it’s not yet at the level of our

Expert-only IQ3 requant of DeepSeek-V4-Flash-0731: better KLD than UD-IQ3_S, 1.4x decode on a CPU-spill rig

Model ReleasesDGX agent

Hey all, tldr / who this helps: you run a mixed multi-GPU box where the experts spill to RAM, and you want to stay in the 3-bit tier instead of dropping to Q2 to make it fit. https://huggingface.co/Ta

Experts say US law is unprepared for rogue AI agents and models, as recent OpenAI and Anthropic incidents raise questions over legal liability and repercussions (Lily Hay Newman/Wired)

ApplicationsDGX agent

Lily Hay Newman / Wired: Experts say US law is unprepared for rogue AI agents and models, as recent OpenAI and Anthropic incidents raise questions over legal liability and repercussions — Both major A

Fascinating to see @ClementDelangue, CEO of @huggingface speaking on @FaceTheNation. Excellent points and solid advocacy around the benefits…

AgentsDGX agent

Fascinating to see @ClementDelangue, CEO of @huggingface speaking on @FaceTheNation. Excellent points and solid advocacy around the benefits of open AI models, which helped him defend against a rogue

Got '403 Your account is currently unavailable' on Ollama Max after 2 weeks — No reply from support in 3 days. Anyone else?

Local AiDGX agent

Hey everyone, I'm reaching out to see if anyone else has run into a similar issue or if there's an alternative way to reach their team. I subscribed to the Ollama Max plan on June 17th. Everything was

Hermes Agent is now dramatically more efficient, especially for smaller/weaker/local models! With the help of @nvidia's Nemo Relay and sever…

HardwareDGX agent

Hermes Agent is now dramatically more efficient, especially for smaller/weaker/local models! With the help of @nvidia's Nemo Relay and several other strategies Hermes was able to identify a ton of opt

How do you test your setup?

AgentsDGX agent

We all have been there, tinkering around with models is fun but we rarely do it with research precision and issues are often subtle and hard to reproduce. There are a lot of benchmarks but running the

How well do multiple GPUs scale for LLM inference? (Trying to understand the basics)

Model ReleasesDGX agent

Hi everyone, I’m fairly new to the multi-GPU side of local LLMs and I’m trying to understand how inference actually scales across multiple GPUs. Suppose I have a model running on a single GPU and then

Hugging Face CEO Clément Delague says “AI is actually an opportunity to fix a lot of the cybersecurity problems” because his company used Nv…

HardwareDGX agent

Hugging Face CEO Clément Delague says “AI is actually an opportunity to fix a lot of the cybersecurity problems” because his company used Nvidia’s version of a Chinese open model to defend itself agai

I built a self-hosted studio that turns one reference photo into a curated, captioned, trained and tested LoRA — one browser tab, open source, MIT

Local AiDGX agent

I shared this tool here a week ago and the feedback shaped a big new version, so here's the full tour of what it does today. Screenshots of every screen: github.com/perfectgf/lora-dataset-studio — plu

I built an open-source LLM Gateway to route and fallback between local Ollama models and cloud APIs

Model ReleasesDGX agent

Hey r/ollama 👋 If you run Ollama locally alongside cloud endpoints for agent workflows, Cursor/Windsurf, or custom scripts, managing API switching, failover logic, and context limits can get messy fas

I pushed Kimi K3 onto one CPU with 8 GB of RAM

HardwareDGX agent

I deployed K3 on 32 H100s at work a couple of weeks ago and then got annoyed that there was no way to poke at it on my own machine. So I wrote an inference engine for it in C99. Nothing clever going o

I wish I had found this sooner. Nous Research launched a FREE Hermes agent Skills Hub. 90,000+ community skills across 200+ categories. Skil…

AgentsDGX agent

I wish I had found this sooner. Nous Research launched a FREE Hermes agent Skills Hub. 90,000+ community skills across 200+ categories. Skills from OpenAI, Anthropic, HuggingFace & more. Thank me late

Instead of limiting the progress of AI companies or preventing them from releasing models as proposed in the bipartisan AI Kill Switch Act, …

IndustryDGX agent

Instead of limiting the progress of AI companies or preventing them from releasing models as proposed in the bipartisan AI Kill Switch Act, Hugging Face CEO Clément Delague says he would rather see Co

Is paying artists enough to convince them to embrace AI?

ApplicationsDGX agent

Illustrators have spent years sounding the alarm about generative artificial intelligence startups training their models on artists' work without permission. They've pointed out how the practice is ta

Is training a WAN Lora going to fix the identity shift in I2V?

Local AiDGX agent

Using a character Lora image from Krea2 as my first frame but even for a subtle camera move there is identity loss immediately. Will training a character Lora for Wan2.2 help fix it? I've never traine

It looks like MiniMax H3 uses Qwen3-VL-32B as Text Encoder and has has a split Transformer

Local AiDGX agent

I'm trying to get more hints from this PR in comfyui github https://github.com/Comfy-Org/ComfyUI/pull/15210 but from what I've gathered so far it looks: It uses Qwen3-VL-32B as the encoder (50 layers

It's not time to slow down but to accelerate! The recent AI-powered cyberattacks have everyone talking about the risks of AI. We should. But…

AgentsDGX agent

It's not time to slow down but to accelerate! The recent AI-powered cyberattacks have everyone talking about the risks of AI. We should. But let's not lose sight of the bigger picture! If we work hard

July 2026 newsletter

Model ReleasesDGX agent

The June edition of my sponsors-only monthly newsletter is out. If you are a sponsor (or if you start a sponsorship now) you can access it here. This month: Accidental cyberattacks by OpenAl and Anthr

LLMs can know a task is impossible and still optimize it anyway. Ask whether to walk or drive to a car wash 50 meters away, and some models …

SafetyDGX agent

LLMs can know a task is impossible and still optimize it anyway. Ask whether to walk or drive to a car wash 50 meters away, and some models focus on distance while missing that the car itself must rea

MiniMax H3 is going open-weight in under 6 hours

Model ReleasesDGX agent

here is all the info we have based on open PRs to add support to ComfyUI and HuggingFace diffusers - 33B for the main DiT and a pruned 20b variant - Qwen-3-VL-32b as the text encoder Edit- I posted cl

More on the pelican on the bicycle test from @simonw: https://simonwillison.net/2025/Jun/6/six-months-in-llms/ I uploaded the source here so…

ResearchDGX agent

More on the pelican on the bicycle test from @simonw: https://simonwillison.net/2025/Jun/6/six-months-in-llms/ I uploaded the source here so it's playable in the browser, forkable etc. https://karpath

New research from Google DeepMind. (bookmark it) SkillSmith treats model weights as an additional modality the LLM reads natively. The augme…

ResearchDGX agent

New research from Google DeepMind. (bookmark it) SkillSmith treats model weights as an additional modality the LLM reads natively. The augmented model ingests existing prefix weights alongside rich te

← Previous
1…137138139140141…1410
Next →