AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “r-ollama”

GridTimelineEvolution
427 results
CompaniesToolsTechniques

Each lane shows up to 8 recent matching entries, ordered from earlier to later. Tracks load separately to keep the 75,000+ entry wiki fast.

Companies

CompanyAnthropic8 recent entries
30 Jul 2026How to uninstall Ollama Claude code

pretty simple I did ollama launch claude, found out its slow asf, wanted to delete It and Idk how I have no idra if its the same as normal Claude Code uninstall or if it's a Little bit of a different

→31 Jul 2026What's your local AI coding setup on a MacBook Pro M4?

I've spent the last couple of days trying different setups (Ollama, Continue, Claude Code, Gemini CLI, OpenRouter...) and at this point I feel like I've spent more time configuring tools than actually

→
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
31 Jul 2026
Experience sharing: How do you use your local models and for what kind of tasks?

Here is my experience, which I would like to share with you and I also would like to hear your thoughts and valuable tips&tricks. Hardware: Mac Mini M4 (32GB Unified Memory) Model Server: Ollama Orche

→2 Aug 2026I built an open-source LLM Gateway to route and fallback between local Ollama models and cloud APIs

Hey r/ollama 👋 If you run Ollama locally alongside cloud endpoints for agent workflows, Cursor/Windsurf, or custom scripts, managing API switching, failover logic, and context limits can get messy fas

→5 Aug 2026Watch a local Ollama's qwen3:8b turn one English question into a 9-node investigation graph - planned, admitted by a deterministic gate, and run live in the browser (open source, MIT)

The video is one real run, not a mock-up: grapharc go 'why did checkout latency spike at 09:14 UTC?' --model ollama/qwen3:8b A local 8B model proposes the graph → triage fanning out into four parallel

→6 Aug 2026An Update to Sir Shortoken: Introducing LELP-S+ (Less English, Less Prose)

A small update to Sir Shortoken. Sir Shortoken already had Quick, Balanced, Deep, Bullets, and Aggressive Bullets. I wanted something between Bullets and normal prose. So I added LELP-S+ (Less English

→9 Aug 2026M3 16GB running Ollama (Qwen 9B) is extremely slow (10-12 mins per task). Am I doing something wrong?

Hey everyone, I constantly see high praise for M3 and M4 Macs for local LLM inference, even the base/16GB models. However, my experience has been quite different, and I'm trying to figure out if I hav

→9 Aug 2026CyberKimi just dropped strong results on one of ExploitBench’s hardest V8 bugs , points away from Mythos

Hey everyone ! Quick share from the cyber + local LLM side of things that I found interesting. During this week’s hacker summer camp, an AI researcher and reverse malware engineer veteran 'lordx64' on

CompanyOpenAI8 recent entries
24 Jul 2026No sé nada de Ollama, ni programación ni idea, pero estoy creando un agente evolutivo

Con ayuda de ChatGPT y con el modelo de Ollama, Qwen3:14b estoy creando un agente que corre local y tiene la iniciativa para pensar, investigar, aprender, generar propuestas y esperar mi autorización

→26 Jul 2026I built an open-source Ollama canvas where the wires are the actual context

Most graph-based LLM interfaces use a canvas as a visual layer over what is still a linear chat. I wanted the graph itself to determine what Ollama receives. ThoughtDAG has one rule: wires are the con

→27 Jul 2026Give any Ollama-compatible client session memory + a shared knowledge wiki by swapping the chat URL

Hey folks — I built ContextMemory, an open-source agentic context gateway for apps that already talk to LLMs. The idea is simple: keep your existing POST /api/chat client (Ollama wire format), point i

→30 Jul 2026P.A.I. — Sleek Native Desktop AI Overlayer for Local Ollama Models 🤖⚡

Greetings Community! 👋 I hope everyone is doing well! I'm Tauhid — Senior EEE student from a Bangladeshi University Today I'd like to share an open-source project I’ve been developing called P.A.I. (P

→30 Jul 20262 images + 1 prompt > expected output

Hi, I'm trying to replicate a thing locally, that I can do on ChatGPT. What I want is to give a local AI two reference images (a face and a background item) and a prompt about the composition of the p

→1 Aug 2026V4 flash vs V4 Flash (0731). Guys, new DeepSeek V4 Flash(0731) is now free on InferX

DeepSeek V4 Flash is now available on InferX, and it’s free to use. We’re continuing to add GPU capacity as demand grows. While we’re bringing additional capacity online, you may occasionally see high

→2 Aug 2026I built an open-source LLM Gateway to route and fallback between local Ollama models and cloud APIs

Hey r/ollama 👋 If you run Ollama locally alongside cloud endpoints for agent workflows, Cursor/Windsurf, or custom scripts, managing API switching, failover logic, and context limits can get messy fas

→5 Aug 2026Watch a local Ollama's qwen3:8b turn one English question into a 9-node investigation graph - planned, admitted by a deterministic gate, and run live in the browser (open source, MIT)

The video is one real run, not a mock-up: grapharc go 'why did checkout latency spike at 09:14 UTC?' --model ollama/qwen3:8b A local 8B model proposes the graph → triage fanning out into four parallel

CompanyGoogle8 recent entries
13 Apr 2026I benchmarked Gemma4:e4b vs Gemma3:27B vs GPT-4o-mini vs Gemini 2.5 Flash on a Mac Mini M4 Pro 24gb — full results

A Reddit user on r/ollama conducted a hands-on benchmark comparing Gemma4:e4b (Google's compact ~4.5B effective-parameter edge model) against Gemma3:27B, GPT-4o-mini, and Gemini 2.5 Flash, all run or

→16 Apr 2026Google’s DeepMind just released new 4B and 27B MedGemma models!

Google DeepMind released MedGemma, a collection of medical vision-language foundation models based on Gemma 3 in 4B and 27B parameter sizes, demonstrating advanced medical understanding and reasoning

→22 Jul 2026NuExtract3 is now available on Ollama: 4B VLM for document-to-Markdown and structured JSON extraction

Disclosure: I work at NuMind, the team that trained NuExtract3. NuExtract3 is an Apache-2.0, open-weight 4B VLM based on Qwen3.5-4B. It is specialized for document understanding rather than general ch

→23 Jul 2026I built an open-source RAG chatbot starter that runs fully locally with Ollama (FastAPI + ChromaDB)

I kept re-wiring the same RAG plumbing on every project, so I turned it into a clean starter and open-sourced it. Upload a PDF, ask questions, and get answers with page-level source citations. It runs

→31 Jul 2026What's your local AI coding setup on a MacBook Pro M4?

I've spent the last couple of days trying different setups (Ollama, Continue, Claude Code, Gemini CLI, OpenRouter...) and at this point I feel like I've spent more time configuring tools than actually

→2 Aug 2026I built an open-source LLM Gateway to route and fallback between local Ollama models and cloud APIs

Hey r/ollama 👋 If you run Ollama locally alongside cloud endpoints for agent workflows, Cursor/Windsurf, or custom scripts, managing API switching, failover logic, and context limits can get messy fas

→6 Aug 2026An Update to Sir Shortoken: Introducing LELP-S+ (Less English, Less Prose)

A small update to Sir Shortoken. Sir Shortoken already had Quick, Balanced, Deep, Bullets, and Aggressive Bullets. I wanted something between Bullets and normal prose. So I added LELP-S+ (Less English

→9 Aug 2026CyberKimi just dropped strong results on one of ExploitBench’s hardest V8 bugs , points away from Mythos

Hey everyone ! Quick share from the cyber + local LLM side of things that I found interesting. During this week’s hacker summer camp, an AI researcher and reverse malware engineer veteran 'lordx64' on

CompanyMeta8 recent entries
25 May 2026Android app Ollama Talk now on playstore

Ollama Talk is an Android app that connects to an Ollama server, enabling conversations with AI models like Llama and Mistral . The app features an intuitive chat interface with real-time conversation

→26 May 2026A tool to get Claude Code-style reliability from fully local models

Ollama exposes an Anthropic-compatible Messages endpoint , allowing developers to run powerful open-source AI models locally with no API costs and pair them with Claude Code for a capable local AI cod

→5 Jun 2026What are the most capable LLM models I can run on my laptop?

A discussion on r/ollama exploring which high-performance LLM models can be effectively run locally on standard laptop hardware , likely covering model size comparisons, hardware requirements, and per

→21 Jun 2026Your data changes and your multi-hop RAG goes stale? This one updates with embed-and-append -> open-weights Llama-3.3-70B, your own vLLM endpoint, no graph rebuild

This post discusses a solution for keeping multi-hop retrieval-augmented generation (RAG) systems updated when data changes, using an embed-and-append approach with the open-weights Llama-3.3-70B mode

→26 Jul 2026100B Models on Cheap Hardware: how realistic and limitations

There is a lot of buzz around running 100B parameter models on cheap local hardware using ternary (1.58-bit) quantization like Microsoft's BitNet architecture. The theoretical hardware shortcuts are i

→30 Jul 2026P.A.I. — Sleek Native Desktop AI Overlayer for Local Ollama Models 🤖⚡

Greetings Community! 👋 I hope everyone is doing well! I'm Tauhid — Senior EEE student from a Bangladeshi University Today I'd like to share an open-source project I’ve been developing called P.A.I. (P

→3 Aug 2026I gave five different local LLMs a town. They invented Facebook and a duck-based credit bureau. (MIT, self-hosted, you don't play it — you watch it)

Each villager in Pepperton is a different model — a mistral, a qwen3, a qwen2.5, a phi4-mini, a llama3.2 — because model families have genuinely different temperaments, and the friction between them i

→4 Aug 2026Pc build limitations

Here's the build I managed to scrape together System Specifications: CPU: Intel Core i7-7700K Motherboard: ASUS ROG Strix Z270-E Gaming RAM: 32GB Corsair Vengeance DDR4-3000 Storage: 1TB Crucial P5 Pl

CompanyMistral8 recent entries
13 Apr 2026I’m looking for advice on setting up a local AI model that can generate Word reports automatically.

This r/ollama thread discusses community advice on configuring a locally-run AI model (via Ollama) to automatically generate Word documents or reports, covering topics such as model selection, scripti

→14 Apr 2026Agents in Ollama and Langflow

This Reddit post from r/ollama likely discusses how to build and run AI agents locally by combining Ollama — which handles local model serving to keep data private — with Langflow's visual, drag-and-d

→12 May 2026Uncensored LLM

Uncensored LLMs are architectures that have been modified or fine-tuned to remove standard safety alignment layers (guardrails) that limit a model's ability to discuss sensitive topics. Ollama offers

→25 May 2026Android app Ollama Talk now on playstore

Ollama Talk is an Android app that connects to an Ollama server, enabling conversations with AI models like Llama and Mistral . The app features an intuitive chat interface with real-time conversation

→26 May 2026Mistral-7B v0.3 at 128K in llama.cpp: 22,657 → 13,235 MiB live VRAM with ≤0.004 PPL drift

Mistral-7B v0.3 model achieves significant memory optimization when running at 128K context length in llama.cpp, reducing live VRAM usage from 22,657 MiB to 13,235 MiB while maintaining minimal perfor

→5 Jun 2026What are the most capable LLM models I can run on my laptop?

A discussion on r/ollama exploring which high-performance LLM models can be effectively run locally on standard laptop hardware , likely covering model size comparisons, hardware requirements, and per

→25 Jul 2026Is this real ? Qwen3.6:27b with 128k context fit in 24Gb VRAM ?

https://preview.redd.it/yw41s1jikefh1.png?width=1942&format=png&auto=webp&s=3a180ae6443c1db9f7b0ce621533a4b2aa553921 Hi, I've been running Ollama on my Unraid server since the llama2 era. I use to be

→3 Aug 2026I gave five different local LLMs a town. They invented Facebook and a duck-based credit bureau. (MIT, self-hosted, you don't play it — you watch it)

Each villager in Pepperton is a different model — a mistral, a qwen3, a qwen2.5, a phi4-mini, a llama3.2 — because model families have genuinely different temperaments, and the friction between them i

CompanyxAI2 recent entries
5 May 2026Parllama -- a terminal UI for Ollama model management and multi-provider LLM chat

Parllama is a TUI (Text UI) application designed for easy management and use of Ollama-based LLMs that also works with major cloud-provided LLMs. It provides core model management features including f

→4 Aug 2026I added a verify-before-load safety check for Ollama models

I maintain llm-checker, and I’ve added structural model-file validation for Ollama. Ollama stores downloaded models as local blobs. If one is truncated, malformed, or has invalid internal offsets, you

CompanyDeepSeek8 recent entries
2 Aug 2026I built an open-source LLM Gateway to route and fallback between local Ollama models and cloud APIs

Hey r/ollama 👋 If you run Ollama locally alongside cloud endpoints for agent workflows, Cursor/Windsurf, or custom scripts, managing API switching, failover logic, and context limits can get messy fas

→4 Aug 2026I benchmarked the 4 models I had pulled. The 1.1GB one beat the 2GB one at math and lost badly at extraction.

152 generations, deterministic grading (exact number/string/JSON/regex), no LLM judge on my 16GB laptop. task type | deepseek-r1:1.5b (1.1GB) | llama3.2:3b (2.0GB) | gemma:2b | codellama 7b arithmetic

→5 Aug 2026Bro, I need you to FOCUS.

Come on guys, we could get deepseek cheaper with more usage directly with them if we wanted. WHERE IS KIMI K3 WITHOUT EXTRA USAGE NEEDED !!? This is pissing me off like crazy submitted by /u/Other_Che

→6 Aug 2026How is Deepseek v4 flash 0731 running on Ollama cloud?

I cancelled my pro plan ealier because I wanted to use new Deepseek v4 flash 0731 which was available on Openrouter through API only (not yet on ollama cloud at the time). The old Deepseek v4 flash/pr

→6 Aug 2026An Update to Sir Shortoken: Introducing LELP-S+ (Less English, Less Prose)

A small update to Sir Shortoken. Sir Shortoken already had Quick, Balanced, Deep, Bullets, and Aggressive Bullets. I wanted something between Bullets and normal prose. So I added LELP-S+ (Less English

→7 Aug 2026Cancelling my subscription also it was great

today was the last day of my subscription on ollama cloud, to be honest it was a great price value for me and with GLM 5.2 and Deepseek V4 Pro i was able to Vibe code my custom woocomerce shop with mu

→8 Aug 2026Ollama Cloud reviews

I am wondering if anyone can give opinion on if Ollama Cloud pro or max plans are worth it. Id be looking to use it with Kimi K3, Qwen 3.8 and Deepseek v4flash for now. Wondering if it would be better

→9 Aug 2026Doom Loop: Anyone Else Having DeepSeek v4 Flash 0731 Issues on ollama cloud?

Am I the only one having issues with DeepSeek V4 Flash? It gets stuck in a loop, as if it can't call the tools, and keeps repeating the same things endlessly without moving forward. Is it a poorly wri

CompanyNVIDIA8 recent entries
6 Jun 2026Mac mini M4 vs Pc with Nvidia 5060 8gb for ai workloads?

The Mac mini M4 uses unified memory architecture where CPU and GPU share a single 24GB memory pool, while the RTX 5060 has dedicated VRAM. Mac mini M4 is preferred for large model inference (70B param

→6 Jun 2026I built a small Windows tool to monitor and manage Ollama more easily

A tiny Windows system tray tool that monitors local Ollama runtime with quick visual feedback about status, resource usage, and models . The app uses color-coded tray icons for quick status checks and

→6 Jun 2026Built a fully-local paper-RAG across 2× 1080 Ti + a 3090. Three Ollama gotchas that each cost me a day.

A developer documented their experience building a fully-local paper Retrieval-Augmented Generation (RAG) system using two NVIDIA GTX 1080 Ti GPUs and one RTX 3090, sharing three significant challenge

→27 Jul 2026Unable to get GPU Passthrough working - Docker

Setup as follows: Proxmox -> Debian -> Docker -> Ollama. Other containers work. Compose file contains gpu device. Does it need nvidia runtime or any other options? If someone could provide an example

→27 Jul 2026RX 9060 XT 16GB vs RTX 5060 Ti 16GB for local AI — worth the price difference?

Hey everyone! I’m building a PC to run AI models locally and can’t decide between the RX 9060 XT 16GB and the RTX 5060 Ti 16GB. Both have the same amount of VRAM, but is AMD actually a solid choice fo

→1 Aug 2026What's currently the 'smartest' LLM to use on 8GB vram and 16 RAM and same thing for 8 VRAM and 64 RAM?

Been trying to find something that actually handles my workload well instead of just being 'fine.' Started on Qwen 2.5 7B, moved to Qwen 3 8B, and right now I'm using Nemotron 3 Ultra (the big 550B on

→8 Aug 2026I tested a fresh GitHub download → Ollama → first local coding-agent task (72 seconds, no cloud API)

I’m building DesktopLab, an open-source local-first control plane for development agents. I recorded the setup boundary that most agent demos skip: DesktopLab detects the host, proposes the supported

→9 Aug 2026I Turned My Underused Gaming Laptop Into a Local AI Workstation

TL;DR: I am building a Windows-first local AI setup for people who want to try local LLMs without spending days choosing models, setting up Ollama, Docker, WSL, Open WebUI, agents, and tool permission