AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
1,437 results
Model Releases

Hidden Reasoning from Claude and GPT are Decoded, and it is interesting

DGX agent

Yesteday a paper showed a gap that allows to see 100% of the reasoning tokens form ALL Claude and GPT models Stealing Reasoning Traces from Proprietary LLM APIs. check it out, they have published lots

model-releasesr-localllama
12 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

I put Gemma 4 E4B and E2B into an e-reader so I can ask my weird questions and share my thoughts in private directly in app.

DGX agent

Here's how it works in the app: Framework: Runs on LiteRT-LM (like Google's AI Edge). Models: Downloads either the E2B (~2.5 GB) or E4B (~3.6 GB) INT4 quantized models directly from ungated litert-com

model-releasesr-localllama
11 Aug 2026
Model Releases

Needle 2: 14MB agentic LLM for phones, wearables, smart home and robots.

DGX agent

Hey LocalLlaMa, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots a

model-releasesr-localllama
10 Aug 2026
Model Releases

24 GB of VRAM is not really 24 GB for a local LLM. Here is the worksheet I use

DGX agent

I kept seeing model file size compared directly with the number printed on the GPU box. That misses several memory buckets. A simple planning model is: usable capacity = advertised VRAM x 0.90 total t

model-releasesr-ollama
9 Aug 2026
Model Releases

They almost catched up on Frontier performance, so now catching up on prices

DGX agent

Users also report that the free version was significantly downgraded after the release of the new models this is very important for us when considering local hosting. A lot of people decided not to bu

model-releasesr-localllama
6 Aug 2026
Model Releases

Unsloth's Gemma 4 mmproj silently broke vision & audio on newer llama.cpp builds — anyone else hit this?

DGX agent

So I had been building ScreenMind, kinda like local ai desktop assistant that uses Gemma 4 for screen analysis, voice memo transcription, and meeting transcription — all through llama-server. Everythi

model-releasesr-localllama
6 Aug 2026
Model Releases

Sulphur 3 is looking for funding

DGX agent

Hello, I'm the guy who made Sulphur 2. With the recent release of a certain video model, we are looking to mobilize and train Sulphur 3 on this new model. We are targeting $10,000 USD. This certain ne

model-releasesr-stablediffusion
5 Aug 2026
Model Releases

Is it just me, or are current LLM benchmarks failing to capture actual usability? (Gemma 4 vs. Gemini/Claude Opus)

DGX agent

Disclaimer, this was kinda written with AI (Gemma 4 again) but it also did really well here, it outputted what I wanted, when I asked it to refine stuff or improve on certain areas it did that without

model-releasesr-localllama
31 Jul 2026
Model Releases

Memory bandwidth, not VRAM size, sets your tokens/sec — here's the arithmetic

DGX agent

Every week someone asks which card to buy and the thread turns into people naming GPUs they happen to own. There's an actual calculation behind it, it takes two numbers off the spec sheet, and it pred

model-releasesr-ollama
30 Jul 2026
Model Releases

90 agentic bakeoff runs: ThinkingCap vs Fable Fusion vs stock Qwen3.6-27B

DGX agent

Last week someone here said ThinkingCap and Fable Fusion 'really do beat the OG' for agentic work, so I ran it: 6 self-grading tasks, 5 reps, 3 models, 90 isolated runs. Tooling, since that's half the

model-releasesr-localllama
26 Jul 2026
Model Releases

We compared different LLMs on IMO 2026 [R]

DGX agent

There are a few reasons why problems from International Mathematical Olympiad function as a good benchmark for LLMs: - The problems are new, not included in the training data of any model - Hard math

model-releasesr-machinelearning
26 Jul 2026
Model Releases

NuExtract3 is now available on Ollama: 4B VLM for document-to-Markdown and structured JSON extraction

DGX agent

Disclosure: I work at NuMind, the team that trained NuExtract3. NuExtract3 is an Apache-2.0, open-weight 4B VLM based on Qwen3.5-4B. It is specialized for document understanding rather than general ch

model-releasesr-ollama
22 Jul 2026
Local Ai

OpenCode + Ollama + MCP

DGX agent

I installed OpenCode and an Ollama model (qwen3.5) sucessfully connected the model respond in OpenCode but doesn't find My MCP server, i Made one using fastMCP other models like bigPickle and openai m

local-air-ollama
22 Jul 2026
Hardware

Open Model: Google Weather Next 2

DGX agent

I am not a meteorologist, but I just read a very interesting article: https://arstechnica.com/science/2026/08/deepminds-hurricane-model-bought-forecasters-an-extra-day/ In a paper published on Thursda

hardwarer-localllama
9 Aug 2026
Local Ai

Quick survey (2 min) on trust in hardware specs for open-source models

DGX agent

Hi everyone, I'm a systems analysis student researching a problem a lot of you probably know well: how much you actually trust the published VRAM/RAM requirements for open-source models before trying

local-air-ollama
8 Aug 2026
Local Ai

Wan-Animate-2: Pushing the Application Boundaries of Character Animation Models

DGX agent

📝 Introduction We present Wan-Animate-2, a novel end-to-end character animation framework that directly consumes driving videos in a redesigned Diffusion Transformer, which achieves high-fidelity moti

local-air-localllama
7 Aug 2026
Local Ai

When did Ollama become so cool with non self hosted models?

DGX agent

It feels like Ollama’s brand identity has shifted. It used to be centered on self‑hosted models, but now most of the new releases seem to be API‑based services with far more emphasis on cloud integrat

local-air-ollama
2 Aug 2026
Local Ai

What's the last model trained on human-data only?

DGX agent

From my understanding, most current LLMs are trained on trillions and trillions of tokens of mostly AI-generated data. Are there any recent models that are trained purely (or as close as possible) on

local-air-localllama
24 Jul 2026
Agents

Training a harness for model-agnostic and task-environment-agnostic capability improvements with PyTorch-like framework [P]

DGX agent

I worked on this project (https://github.com/workofart/harness-training) for the past few months to reframe 'Agent-driven Self-improving Harness' to 'Harness Training'. The idea is simple, the harness

agentsr-machinelearning
20 Jul 2026
Local Ai

What is the best open sourced image model?

DGX agent

The best open-source image generation models in 2026 include FLUX.1 [schnell], Stable Diffusion 3.5 Large, HiDream-I1-Full, SANA-Sprint 1.6B, and HunyuanImage-3.0 . FLUX.1 [dev] holds the crown for ph

local-air-stablediffusion
10 Jun 2026
Local Ai

Model test: SenseNova U1 vs GPT Image2 vs Nano Banana in Infographic generation

DGX agent

A comparative test of three AI image generation models—SenseNova U1, GPT Image 2, and Nano Banana 2—evaluating their performance on infographic generation tasks. GPT Image 2 proved more reliable for e

local-air-stablediffusion
26 May 2026
Local Ai

Need advice: Best ComfyUl workflow for texturing a 3D model from 4 orthographic views using reference images?

DGX agent

A Reddit user seeks guidance on configuring ComfyUI workflows to texture 3D models using orthographic reference images from multiple angles. The discussion likely covers best practices for using Stabl

local-air-stablediffusion
19 May 2026
Industry

New ChatGPT update removes image library, further buries model picker

DGX agent

OpenAI removed the Images shortcut from ChatGPT's left sidebar, moving it to a new Library section that is not yet available in all countries. The update also buried the model picker, requiring users

industryr-chatgpt
30 Apr 2026
Industry

The new image model makes all images suspect

DGX agent

ChatGPT Images 2.0 generates highly realistic images that are difficult to distinguish from authentic photographs, raising concerns about misinformation and deepfakes. The new model features improved

industryr-chatgpt
23 Apr 2026
Local Ai

Ollama broken on Apple M5 + macOS 26 — Metal shader crash on every model (500 error)

DGX agent

Ollama consistently fails to run any model on Apple M5 hardware running macOS 26, with the Metal backend failing to initialize and the runner process terminating with a 500 Internal Server Error. The

local-air-ollama
15 Apr 2026
Model Releases

Introducing Unsloth Desktop app

DGX agent

Hi LocalLlama, we're super excited to release Unsloth Desktop today! 🦥 It's the first desktop app that enables you to run and train models locally. Open-source. Available on Mac, Windows, and Linux Su

model-releasesr-localllama
11 Aug 2026
Local Ai

MiniMax-H3: ~38 GB less VRAM with Runtime LoRA Bypass — DoRA Dynamic LoRA Loader v1.0.39

DGX agent

GitHub: https://github.com/xmarre/ComfyUI-DoRA-Dynamic-LoRA-Loader Release v1.0.39: https://github.com/xmarre/ComfyUI-DoRA-Dynamic-LoRA-Loader/releases/tag/v1.0.39 Also available through ComfyUI Manag

local-air-stablediffusion
11 Aug 2026
Model Releases

Muse Glimmer ACTUALLY fits on a single RTX 3090

DGX agent

I did some testing this morning, and I was surprised to find that Muse Glimmer actually comfortably fits on a single RTX 3090 with full context + DFlash + mmproj at Q4_K_XL, unlike Qwen3.6-27B and Gem

model-releasesr-localllama
10 Aug 2026
Model Releases

endless-frontier/BigBang-v1 - qwen 3.5 finetunes

DGX agent

table bench https://huggingface.co/bartowski/endless-frontier_BigBang-v1-GGUF I'm downloading this model only because Bartowski converted it to .gguf, so it might be interesting. Doubts : The headline

model-releasesr-localllama
9 Aug 2026
Local Ai

Lophius: A workbench for language model research, from the creator of Heretic

DGX agent

Hi folks, I hate slop as much as you do, so instead of starting with 'The Problem', I'll just cut to the chase: I just published Lophius, which is the culmination of more than two years of fighting wi

local-air-localllama
9 Aug 2026
Model Releases

Echo Dot 2 can run 28M LLM at decent speed

DGX agent

Code and instructions available here: https://github.com/albertoZurini/echo-dot-2-playground Hello there! After a few days of experimenting I was able to get a completely local voice pipeline running

model-releasesr-localllama
7 Aug 2026
Model Releases

I remember a time when 'flash' meant 32B

DGX agent

I mean, Deepseek V4 Flash is an absolutely fantastic model, even though I can't run it on my machine it's so fascinating to see how it performs. Knowing that potentially it could be run at home is rea

model-releasesr-localllama
5 Aug 2026
Local Ai

We applied BitNet-style ternary quantization to a super-resolution transformer. The whole model is 668 KB gzipped and runs in the browser.

DGX agent

Everyone's been doing 1.58-bit for LLMs, so we tried it on a vision transformer: Swin2SR (lightweight ×2 variant, 1.01M params), quantized so every weight is −1, 0, or +1 with a small per-group scale

local-air-stablediffusion
31 Jul 2026
Model Releases

LG AI Research releases K-EXAONE 2.0 750B A37B

DGX agent

It was developed under Phase 2 of Korea's Sovereign AI Foundation Model Project. ​Size: 750B parameters (3x larger than their 236B v1 model). ​- License: Apache 2.0 ​Languages: Expanded to 10 language

model-releasesr-localllama
30 Jul 2026
Model Releases

Appreciation for Gemma 4 26b A4b

DGX agent

I really love this model, I have been using the q4_k_l by Bartowski (I have heard QAT is quite the downgrade in some aspects) and it handles every task I throw at it easily. Agentic and coding perform

model-releasesr-localllama
28 Jul 2026
Tutorials

Medical model: Reasoning-Medical-27B (Qwen3.6-27B finetune)

DGX agent

From the description: 'Reasoning-Medical-27B is designed for universal advanced medical reasoning in professional medicine, medical genetics, college biology/medicine, and clinical knowledge. The mode

tutorialsr-localllama
28 Jul 2026
Model Releases

Small context windows + knowledge graphs: the serialization format alone doubled my multi-hop accuracy (benchmarked 10 formats)

DGX agent

Running local models means every token counts — an 8K or 16K window fills up fast when you're stuffing graph context into prompts for RAG. I benchmarked 10 graph serialization formats (JSON, GraphML,

model-releasesr-localllama
27 Jul 2026
Model Releases

Benchmarks: TensorSharp vs. llama.cpp

DGX agent

Cuda and Vulkan Benchmark: TensorSharp vs. llama.cpp I would like to share my latest open source local Unsloth (GGUF) LLM inference engine and applications. It supports many models from Unsloth, like

model-releasesr-localllama
25 Jul 2026
Model Releases

Getting the most out of MTP

DGX agent

If you want to get the most out of MTP. You have to run some tests / benchmarks to do so. Turning it on with defaults will get improvements, but for many models and card combinations, you are leaving

model-releasesr-localllama
24 Jul 2026
Local Ai

Ideogram 4.0 Just Open Sourced!

DGX agent

Ideogram released version 4.0 of its text-to-image model as an open-weight model with native 2K resolution, bounding box control, and improved text rendering. The model weights are available on GitHub

local-air-stablediffusion
3 Jun 2026
Local Ai

OpenStudio - Hybrid local/cloud (openrouter) AI router

DGX agent

OpenStudio is a hybrid AI router that combines local model inference with cloud-based model access through OpenRouter, a unified API providing access to hundreds of AI models through a single endpoint

local-air-ollama
25 May 2026
Model Releases

Why there is no cloud version for Qwen 3.6 27/35B?

DGX agent

The Qwen 3.6-27B and 35B models are designed as open-weight models that developers can run locally on their own hardware without requiring cloud services. Alibaba released a separate cloud-only produc

model-releasesr-ollama
25 Apr 2026
Safety

easyaligner: Forced alignment with GPU acceleration and flexible text normalization (compatible with all w2v2 models on HF Hub) [P]

DGX agent

easyaligner is a forced alignment library designed to be performant and easy to use , leveraging GPU acceleration to align audio with text transcriptions. The tool supports flexible text normalization

safetyr-machinelearning
18 Apr 2026
Research

Paper from Kimi: Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter [R]

DGX agent

Mooncake is the serving platform for Kimi developed by Moonshot AI, featuring a KVCache-centric disaggregated architecture that separates prefill and decoding clusters while leveraging underutilized C

researchr-machinelearning
18 Apr 2026
Hardware

IA local con NVIDIA RTX PRO™ 4000 Blackwell 16GB GDDR7

DGX agent

This Reddit post from the r/ollama community discusses running local AI/LLM workloads using the NVIDIA RTX PRO 4000 Blackwell GPU via Ollama, a framework for running large language models locally. The

hardwarer-ollama
14 Apr 2026
Model Releases

Gemma:26b thinking issue in openWebUI

DGX agent

This r/ollama thread discusses user-reported issues with the Gemma 4 26B (a Mixture of Experts model) and its 'thinking' mode when used through Open WebUI. Key problems include the model getting stuck

model-releasesr-ollama
13 Apr 2026
Local Ai

Need help to download from civitai in China

DGX agent

This Reddit thread from r/StableDiffusion addresses the challenge faced by users in China trying to access and download models from Civitai, which may be restricted or slow due to network limitations

local-air-stablediffusion
13 Apr 2026
Model Releases

Ace Step 1.5 XL ComfyUI automation workflow without lama for generating random tags using qwen, generate song and then give it a rating by using waveform analysis

DGX agent

This is a ComfyUI automation workflow for the ACE-Step 1.5 XL music generation model that uses Qwen (a language model text encoder) to randomly generate music tags/captions without requiring LAMA, ...

model-releasesr-stablediffusion
10 Apr 2026
← Previous
1…56789…30
Next →