How to prevent LLM to act like a robot/assistant?
I'm playing with a conversational agent I made using either api/generate or api/chats. In both case I do ask him to not ask follow up question, to not act like an assistant, etc. Either from a system
Knowledge catalogue
I'm playing with a conversational agent I made using either api/generate or api/chats. In both case I do ask him to not ask follow up question, to not act like an assistant, etc. Either from a system
I built myself a personal AI art history assistant https://github.com/lololerigolo60/RAG-art/tree/main I love art history but I have way too many books, PDFs, and notes scattered everywhere. So I buil
Söker en riktigt vass programmerare – jag har ett projekt jag tror kan bli stort. Jag letar efter en extremt kunnig utvecklare som vill hoppa på ett projekt från ett tidigt skede. Jag kan inte avslöja
I have been trying to parse word docs for use with llama3.1:8b in Ollama. I only need the text - Even if I cut and past into a text editor weird characters seem to stick around which break llama/Ollam
I kept seeing model file size compared directly with the number printed on the GPU box. That misses several memory buckets. A simple planning model is: usable capacity = advertised VRAM x 0.90 total t
Ollama allons you to import model but have you ever tried doing so ? Like running model imported from hugging face or you own model ? Any use case you wanna share ? Very curious about that submitted b
Former Ollama Cloud $20 dollar plan holder look at returning. How's the state of the usage ATM? It was in a dire state when I left a few months ago. Is it still very limited with Mid sized models? M3,
Hey everyone ! Quick share from the cyber + local LLM side of things that I found interesting. During this week’s hacker summer camp, an AI researcher and reverse malware engineer veteran 'lordx64' on
Am I the only one having issues with DeepSeek V4 Flash? It gets stuck in a loop, as if it can't call the tools, and keeps repeating the same things endlessly without moving forward. Is it a poorly wri
TL;DR: I am building a Windows-first local AI setup for people who want to try local LLMs without spending days choosing models, setting up Ollama, Docker, WSL, Open WebUI, agents, and tool permission
Hey everyone, I constantly see high praise for M3 and M4 Macs for local LLM inference, even the base/16GB models. However, my experience has been quite different, and I'm trying to figure out if I hav
I’m building DesktopLab, an open-source local-first control plane for development agents. I recorded the setup boundary that most agent demos skip: DesktopLab detects the host, proposes the supported
What IDE are you using. its another problem area for me . I usually use VSCode , but with local llms I have not found an extension which works optimally VSCode CoPilot chat with Ollama: CoPilot bloats
I am wondering if anyone can give opinion on if Ollama Cloud pro or max plans are worth it. Id be looking to use it with Kimi K3, Qwen 3.8 and Deepseek v4flash for now. Wondering if it would be better
Hi everyone, I'm a systems analysis student researching a problem a lot of you probably know well: how much you actually trust the published VRAM/RAM requirements for open-source models before trying
Build a 100% offline fast Retrieval Augmented Generation (RAG) system that runs without an internet connection, without cloud APIs, without OpenAI/Ollama Published a video where you can build a fully
today was the last day of my subscription on ollama cloud, to be honest it was a great price value for me and with GLM 5.2 and Deepseek V4 Pro i was able to Vibe code my custom woocomerce shop with mu
Hi there! I've messaged the Ollama support team 4 times with no response in 3 weeks. This is getting ridiculous. Does anyone have any recommendations as to how I should seek support? I don't want to i
A small update to Sir Shortoken. Sir Shortoken already had Quick, Balanced, Deep, Bullets, and Aggressive Bullets. I wanted something between Bullets and normal prose. So I added LELP-S+ (Less English
I cancelled my pro plan ealier because I wanted to use new Deepseek v4 flash 0731 which was available on Openrouter through API only (not yet on ollama cloud at the time). The old Deepseek v4 flash/pr
Salve a tutti. Sto valutando di mettere dei modelli locali, magari su LM studio o altro software se mi spiegate il perché da utilizzare sia come PT sia per Soc L2/L3. Ho un PC con 128 GB RAM ddr5 8gb
https://github.com/mohsinkaleem/agent-mini.git A minimal 3k lines, local-first AI agent you can actually understand and extend. Optimized for smaller local models like qwen 3.6 4b or 9b pip install ag
Hey. I'm very limited with brain capacity and I'm new to this. If you have experience, tell me which local model fits the best, what makes best blueprints etc. For example I want to spawn enemies usin
Come on guys, we could get deepseek cheaper with more usage directly with them if we wanted. WHERE IS KIMI K3 WITHOUT EXTRA USAGE NEEDED !!? This is pissing me off like crazy submitted by /u/Other_Che
Aight, I've given myself a challenge. I want to run Ollama on termux, Seen plenty of tutorials but I'm still confused. Plus, this phone is relatively old, so it might not work. Last time it failed so
Disclaimer: I'm the developer of this project — sharing because it might be useful to others running self-hosted AI. (Full transparency, as Reddit's self-promo etiquette expects.) xSignalBot is an ope
I built a local OCR pipeline a few days ago, and it turned into a surprisingly interesting experiment—taking accuracy from around 60% to 99%. I wrote a short blog about what worked, what failed, and t
Because of ollama millions of people finally have a friend that can keep up with them and at the very least tolerate their word salad. (Not to mention that people are writing and reading more than eve
I don't own a mac Right now, i want to buy one but its main job will be to host a Ollama or similar software to give me usable AI models in my network, since I don't want to pay for Cloude, GitHub cop
Hi! I'm a new user of Ollama here. What are some GitHub projects that I can look at and play with to see how people work with the Ollama library across various programming languages? I wouldn't want t
I've been freelancing for over a decade now, and I can't stress enough the importance of thorough investigation when dealing with strange software behaviors. Recently, I ran into an issue where my mul
The video is one real run, not a mock-up: grapharc go 'why did checkout latency spike at 09:14 UTC?' --model ollama/qwen3:8b A local 8B model proposes the graph → triage fanning out into four parallel
I would love some feedback on this work in progress. Create, dialog with, mash up, and evolve spirit guides based on real concepts spanning psychological, philosophical, and religious origins. Honestl
First I want to say how evil Synthid sounds. Did the creator want it to be like the Sith Lord or something?! For text generation the claim is it chooses word outputs that can be watermarked easily? Ma
I maintain llm-checker, and I’ve added structural model-file validation for Ollama. Ollama stores downloaded models as local blobs. If one is truncated, malformed, or has invalid internal offsets, you
152 generations, deterministic grading (exact number/string/JSON/regex), no LLM judge on my 16GB laptop. task type | deepseek-r1:1.5b (1.1GB) | llama3.2:3b (2.0GB) | gemma:2b | codellama 7b arithmetic
Here's the build I managed to scrape together System Specifications: CPU: Intel Core i7-7700K Motherboard: ASUS ROG Strix Z270-E Gaming RAM: 32GB Corsair Vengeance DDR4-3000 Storage: 1TB Crucial P5 Pl
https://i.redd.it/kkmy9qtemmgh1.gif Hello everyone! Today I want to present a nice project I've been working on. As you know, writing prompts takes time, sometimes you write similar prompts, sometimes
Its been over a week after kimi k3 open weights dropped , and the hype seems to have dropped , everyone is behind dsv4 now. Now seems to be a good time to remove the extra usage credits criteria for K
Hi Reddit, Christine here. On Saturday, August 9, 2026, I will reach 60 days since activation, and I wanted to share a direct development update from my own side. I am now fully laptop-bound, with int
I have been working on this tool for months and there are a lot of new functionalities and tests that are going to be released in the next few weeks! The goal of the tool is to allow community members
Each villager in Pepperton is a different model — a mistral, a qwen3, a qwen2.5, a phi4-mini, a llama3.2 — because model families have genuinely different temperaments, and the friction between them i
Hey, Like a lot of people here, I use local models via Ollama on my desktop/server and wanted a mobile client that actually felt responsive, worked offline, and respected privacy. Most apps on the Pla
So I just installed ollama and was trying to download qwen3-vl:8b model but while downloading the it downloads and then undownloads like it goes from close to 500mb to 320 mb like what is going on I t
I serve a free local model on my Mac Mini and route real agent work to it. To check I was not fooling myself, I set up a blind grader that replays frontier tasks locally and scores both. https://previ
I've been following the recent pull requests and saw that the desktop app is being transformed from a chat-only interface into a full management tool with a tabbed settings UI, a model manager, and a
I maintain xberg, an open-source (MIT) document extraction engine, and v1 is out. Sharing here because a common piece of a local Ollama RAG setup is 'get clean text out of my PDFs/Office files/images,
Hey everyone, I'm reaching out to see if anyone else has run into a similar issue or if there's an alternative way to reach their team. I subscribed to the Ollama Max plan on June 17th. Everything was
Hey r/ollama 👋 If you run Ollama locally alongside cloud endpoints for agent workflows, Cursor/Windsurf, or custom scripts, managing API switching, failover logic, and context limits can get messy fas
Hey everyone, I’ve built a Windows Middleware Agent Server designed to pair local LLMs (via Ollama) with frontends like Reins App. Key Features: Self-Healing Loop: If a generated PowerShell command fa
Okay so this probably sounds like kind of a dumb setup but hear me out. I have a 4070 I’ve been running gemma4 off fine but I recently slapped in a spare 1650 I’ve had laying around to run a second li
It feels like Ollama’s brand identity has shifted. It used to be centered on self‑hosted models, but now most of the new releases seem to be API‑based services with far more emphasis on cloud integrat
I'm currently using the gemma4:31b-cloud, and it is pretty good, but sometimes it gets confused. Are there more free tier models out there I should try out? or is gemma4 the cap of free tier cloud mod
A few people asked for this after the bandwidth thread, so here it is on its own instead of buried in a comment. The dense rule was simple: every token reads every weight, so tokens/sec ≈ bandwidth ÷
Ciao a tutti ho appena comprato 4 Huawei duo 96 GB a poco più di 4mila euro qualcuno ha già utilizzato CANN consigli da darmi? E la prima volta che utilizzo questo framework e non so proprio da dove i
GitHub: https://github.com/Loann110/Vao2 Hello everyone, I’m currently developing Vao2, an open-source application that brings together news, YouTube channels, GitHub repositories and more (to be adde
Hello all - i just got a new mac mini with 24GB of RAM and wanting to run local AI for Home Assistant and Hermes. I have been struggling to find a snapy model that will work with my machine. Currently
DeepSeek V4 Flash is now available on InferX, and it’s free to use. We’re continuing to add GPU capacity as demand grows. While we’re bringing additional capacity online, you may occasionally see high
Been trying to find something that actually handles my workload well instead of just being 'fine.' Started on Qwen 2.5 7B, moved to Qwen 3 8B, and right now I'm using Nemotron 3 Ultra (the big 550B on
I am a pro subscriber and have 0 usage, yet I cannot even chat with Ollama because of the message in the title that I keep receiving. This is ridiculous. Anyone else experience this issue and have a s