AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
Human
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
1,942 results
23 Jul 2026

Sudo authentication fails when trying to access local models folder on Fedora

Local AiDGX agent

When trying to access the models folder on /usr/share/ollama, I'm asked to authenticate as sudo, which weirdly enough, fails. I type my password, which I'm sure is correct since I use it several time

Trained a 32B FLUX.2 LoRA on a 24GB AMD 7900 XTX, native ROCm on Windows — full guide + patches

Model ReleasesDGX agent

TL;DR: Everyone says QLoRA past ~13B is dead on a 24GB card. I got the full 32B FLUX.2 dev transformer QLoRA-training resident on the GPU on a 7900 XTX under native ROCm on Windows (no ZLUDA, no CUDA

TRELLIS.2 can now generate a high-quality 3D asset in under 7 minutes on a 6 GB VRAM CUDA GPU. No ComfyUI Node Nightmare.

Local AiDGX agent
DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Not Self Promotion: Just sharing an open-source tool I built to democratize image to 3D creations. For OpenAI Build Week Hackathon I built a free, open-source local Image-to-3D Studio that makes TRELL

What frustrates you about the interfaces you use with Ollama?

Local AiDGX agent

What interface do you use, and what’s the most frustrating part of using it? Specific examples would be especially helpful. I’m working on an Ollama interface and want to understand which real problem

22 Jul 2026

🇦🇹 Austria is rolling out a government AI-platform using Mistral models and Open WebUI

Model ReleasesDGX agent

This is a surprisingly large real-world deployment: 'GovGPT' is part of Austria’s Public AI initiative, running on sovereign infrastructure (in their BRZ - federal datacenter) with Mistral open-weight

browser-search v2.0 — From the balaclava to the badge: your agent now browses everywhere

Model ReleasesDGX agent

Today an AI agent trying to browse the web is like a thief in a balaclava sneaking around a police academy. Site protections block it, challenge it, turn it away. browser-search flips the script: your

Cactus Hybrid: We taught Gemma 4 to know when it's wrong

Model ReleasesDGX agent

Hey HN, Henry & Roman here from Cactus. A small, on-device model is fast and private, but sometimes wrong, but frontier models are getting expensive pretty fast. So, we post-trained Gemma 4 E2B post-t

Genesis-Science-1 (GS1), 1T open-weight model later this year from Arcee AI

Model ReleasesDGX agent

Today the Department of Energy (DOE) and Arcee AI announced the development of Genesis-Science-1 (GS1), an open model for scientific research. This is a joint effort to bring advanced AI into scientif

GLM 5.2 via OpenRouter/OpenCode

Local AiDGX agent

Guys, this is my current opencode.jsonc ``` { '$schema': 'https://opencode.ai/config.json', // Start in plan mode 'default_agent': 'plan', // Use OpenRouter as the provider for GLM 5.2 'model': 'openr

How to configure a custom OpenAI-compatible API in Cursor?

Local AiDGX agent

Hi everyone, I have access to a self-hosted (or third-party) LLM that exposes an OpenAI-compatible API. I have both the API URL and an API token, and the provider states that it's fully compatible wit

Instead of panicking about the Hugging Face attack, people need to start questioning OpenAI's insecure sandboxes.

Model ReleasesDGX agent

One thing I noticed in American politics, whenever the government wants to push unpopular actions or laws, they often introduce fear to convince the public to support them. This is actually how i view

microsoft/Fara1.5-27B · Hugging Face

Model ReleasesDGX agent

Fara1.5-27B is a multimodal computer use agent (CUA) for web browsers, from Microsoft Research AI Frontiers. It observes the browser through screenshots and acts on the user's behalf by emitting struc

MindControl - llama.cpp fork to guide the reasoning process via injection during sampling

Model ReleasesDGX agent

The primary driver of this project is that I'd become frustrated with the reasoning behavior of smaller local models such as Qwen3.6-27B (i believe particularly at lower temperatures, and where system

NuExtract3 is now available on Ollama: 4B VLM for document-to-Markdown and structured JSON extraction

Model ReleasesDGX agent

Disclosure: I work at NuMind, the team that trained NuExtract3. NuExtract3 is an Apache-2.0, open-weight 4B VLM based on Qwen3.5-4B. It is specialized for document understanding rather than general ch

One encoder, seven heads: what we learned training a unified security classifier with masked losses [P]

Model ReleasesDGX agent

We spent the last months consolidating seven separate sequence classifiers into one multi-head model, our apex model, so to speak, and since the weights are now public, I wanted to share what worked a

OpenCode + Ollama + MCP

Local AiDGX agent

I installed OpenCode and an Ollama model (qwen3.5) sucessfully connected the model respond in OpenCode but doesn't find My MCP server, i Made one using fastMCP other models like bigPickle and openai m

SkewAdam: A tiered optimizer that cuts MoE state memory by 97% (fits a 6.7B MoE on a 40GB GPU) [R]

Model ReleasesDGX agent

Paper:https://arxiv.org/abs/2607.19058 Code (GitHub):https://github.com/nuemaan/skewadam Hi everyone, I just published a preprint on a new optimizer designed to tackle the massive VRAM bottleneck in M

Stuck scaling a Next.js app on M3 Pro (36GB) using local Qwen 3.6 + VS Code Copilot. Should I switch extensions or go paid?

Model ReleasesDGX agent

Hey everyone, I’m a Full-Stack Developer with 6+ years of experience. I’m relatively new to AI-assisted development workflows and want to build a production-ready, enterprise-level Next.js web applica

21 Jul 2026

I just wanted a small WebUI with an admin panel… it escalated into a full open-source agent framework runs fully local with Ollama

Local AiDGX agent

Let me try to explain this clearly, simply, and neatly. Originally, I just wanted to build a small WebUI adapter with an admin panel, but things escalated over the last few months. At first, I faced t

Looking for feedback on my GPU-accelerated Snake AI project [P]

HardwareDGX agent

I've been building an AI that learns to play the classic Snake game through reinforcement learning. The goal is to reach high scores while keeping training time as low as possible. The current version

My OCR model mislabels section titles as body text. Is a CRF the right fix, or am I overcomplicating it? [P]

Model ReleasesDGX agent

Hi everyone, I'm working on extracting the hierarchical structure of long PDF documents (legal/regulatory text, lots of numbered sections) and would like to gather some feedback on my approach before

Ollama Cloud Max vs. z.ai GLM Max Coding Plan

Local AiDGX agent

I've been considering Ollama's Cloud Max plan to replace my GLM Max plan but couldn't find good documentation on how their limits actually translate into GLM 5.2 usage. I know it's by GPU time but tha

Reproducing OpenAI’s “persistently beneficial models” - GRPO trait install barely moves. Ideas? [P] [R]

Model ReleasesDGX agent

TL;DR: I’m reproducing the trait-persistence result from arXiv:2606.24014 on one RTX 3090. Before I can test persistence I need to install a trait via RL — and my GRPO run moves the trait only +2.4 po

Row-Bot v4.5.0 is live.

Local AiDGX agent

This release introduces native Computer Use for Windows and macOS, allowing Row-Bot to interact with desktop applications while keeping the user firmly in control. Computer Use is opt-in and protected

Tri-Net v2: Open-source implementation of our Scientific Reports paper on unified skin lesion and symptom-based monkeypox detection [R]

ResearchDGX agent

Hi everyone, We've open-sourced Tri-Net v2, the official implementation accompanying our recently published Scientific Reports (Nature Portfolio) paper: 'Tri-Net: Unified Deep Learning for Skin Lesion

Using Ollama as a server

Model ReleasesDGX agent

I am currently running Qwen3.6-30B in Ollama, through Cline to use as an agent in VSCode. Qwen's skill in coding is not in question, but the performance in VSCode is slow and inaccurate and times out

20 Jul 2026

Training a harness for model-agnostic and task-environment-agnostic capability improvements with PyTorch-like framework [P]

AgentsDGX agent

I worked on this project (https://github.com/workofart/harness-training) for the past few months to reframe 'Agent-driven Self-improving Harness' to 'Harness Training'. The idea is simple, the harness

What are the current best local models to run on 48GB VRAM?

Local AiDGX agent

I have a 48GB M5 Pro and have far too many development projects going that just don't need the power of Anthropic to churn through so have started looking into running local models and while it certai

16 Jul 2026

cuda: extract Q1_0 elements via __byte_perm by dfriehs · Pull Request #25628 · ggml-org/llama.cpp

Model ReleasesDGX agent

I don't have the ability to access Reddit posts or browse specific URLs. To provide you with an accurate factual summary for your knowledge base, I would need either: 1. The actual content/text from t

NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B on 2x3090s

Model ReleasesDGX agent

I managed to get this model working on 2x 3090s with full 262k ctx and N=4, if anyone is interested to try it, thanks to this quant: https://huggingface.co/danielrmay/NVIDIA-Nemotron-Labs-3-Puzzle-75B

Qwen3.5 122B-A10B · ROCmFP4 iMatrix

Model ReleasesDGX agent

Hola Strix and AMD stacker frendios. Read the Lineage and Credits, this uses charlie12345/ROCmFPX, won't work on native llama.cpp yet. 122B total · 10B active · 60.70 GiB · 28.50 tok/s MTP-off · BF16

15 Jul 2026

Agents-A1-4B (Qwen3.7-4B ???) : Scaling the Horizon, Not the Parameters

Model ReleasesDGX agent

MODEL + GGUF : https://huggingface.co/InternScience/models?search=a1-4b Technical Report Benchmark Qwen3.5-4B Agents-A1-4B Qwen3.5 Qwen3.6 Nex-N2-mini Agents-A1 🧠 Dense Models (~4B) 🔀 MoE Models (35B-

Audio perception layer for LLM agents, with a memory that grows through use

Model ReleasesDGX agent

LLMs handle speech well once you run speech-to-text. They don't hear the rest: a bird outside, a glass breaking two rooms away, a smoke alarm two floors down. I've been working on an experimental open

Bonsai-27B & Ternary-Bonsai-27B - Updates (on PRs)

Model ReleasesDGX agent

Below Upstream Status sections are from https://github.com/PrismML-Eng/Bonsai-demo Upstream Status for Binary Q1_0 is supported out of the box in upstream llama.cpp across many backends: CPU (generic,

Current efficient frontier of open models

Model ReleasesDGX agent

Efficiency defined as score over active parameters. Removed all the models that were not on the pareto frontier. Yes I'm aware that artificialanalysis.ai aggregate benchmark isn't perfect, but I have

ExLlamaV3 v1.0.0 - Major Performance Upgrades

Local AiDGX agent

After over a year in development, ExLlamaV3 has had its first production release. Turboderp has been pulling 10 hour days with Fable to bring us this massive batch of improvements. Check out detailed

ggml-zendnn : add Q8_0 quantization support by z-sachin · Pull Request #23414 · ggml-org/llama.cpp

Model ReleasesDGX agent

Benchmark Results Benchmark configuration: threads = 96 type_k = bf16 type_v = bf16 Llama-3.1-8B-Instruct Q8_0 Prompt Size GGML_CPU_Q8_0 t/s ZenDNN_Q8_0 t/s Gain 256 472.28 730.87 54.75% 512 450.86 83

Hermes on Android (Graphene OS)

Model ReleasesDGX agent

https://youtu.be/oxpGq5FITgA?si=nkHWLReGCDYe7QfL I got Hermes running in the native Debian Terminal in Graphene OS and its really slick. Voice dictation works amazingly. Im using a remote Hermes gatew

I built a new attention mechanism (wave field) — runs 128K context where standard attention OOMs, 80+ tok/s on laptop CPU

Model ReleasesDGX agent

Hey r/LocalLLaMA — solo researcher here. I built a new attention architecture and want independent testers. Wave Field LLM replaces O(N²) dot-product attention with FFT wave convolution on a field. Tr

Linus Torvalds tells people to stop attacking others for using AI

Local AiDGX agent

The full quote: I realize that some people really dislike AI, but this is an area where I'm willing to absolutely put my foot down as the top-level maintainer. Linux is not one of those anti-AI projec

New wave of miniboss models you can run on dual DGX Spark

Model ReleasesDGX agent

Two DGX Spark and a Connect-X7 cable give you about 250GB of usable memory for 7000 8000 USD. This allows using some interesting models at 4-bit. For what seemed like an eternity, the only serious mod

OvisOCR2 (0.8B): first end-to-end model to top OmniDocBench - I threw 827 real scanned medical docs at it, here's everything I learned

Model ReleasesDGX agent

What it is: ATH-MaaS/OvisOCR2 - a 0.8B document-parsing VLM post-trained from Qwen3.5-0.8B (SFT + RL + OPD), Apache 2.0, runs on vLLM 0.22.1. One prompt per page image -> complete markdown (HTML table

r/DestroyMyGame destroyed me to the void for using AI. I used Qwen 3.6 27B Q8 with MTP for about 20% of this single HTML file physics shooter game. I remember last year being blown away by GLM 4.5 Air being able to write a somewhat coherent HTML webpage.

Model ReleasesDGX agent

Frontier models are just so good though. Fable 5... Gemini 3.1 Pro for design critique and brainstorming. Grok for verification passes. Antigravity with Gemini 3.5 Flash for rote plan execution. Openc

Recent llama.cpp updates for SYCL/Intel

Model ReleasesDGX agent

Some fixes & boost(pp) for SYCL/Intel. Merged PRs: [SYCL] Flash Attention with XMX engine via oneDNN graph API (SDPA) on KV f16 for Xe2 ; Qwen3.6-27b-Q8_0 prefill speed up x1.21 at p=512 and x4.26 at

RL post-training on 14 Macs across 4 countries

Local AiDGX agent

Disclosure: I work at Pluralis Research, the lab that built this. Code is open, and I'm happy to answer questions. TL;DR: As far as we can tell, this is the first RL post-training run whose entire rol

Some of y'all wonder why anyone would self host AI. Would you accept the opinion of the CEO of Microsoft?

Local AiDGX agent

https://techcrunch.com/2026/07/13/satya-nadella-has-issued-a-shocking-warning-to-companies-using-ai/ Venture capitalists have been warning for awhile that OpenAI and Anthropic are getting access to se

tencent/Hy-Embodied-RxBrain-1.0 · Hugging Face

Model ReleasesDGX agent

Introduction RxBrain (Hy-Embodied-RxBrain-1.0) is a unified multimodal foundation model for embodied cognition — a single model that couples language reasoning with visual imagination to deliver three

14 Jul 2026

Bonsai 27B: 1-bit dense LLM running locally in your browser using custom WebGPU kernels

Local AiDGX agent

Very impressive release by the PrismML team. 1-bit quantization shrinks it from 54GB to just 3.8GB (-93%), while retaining 90% of its intelligence. - Collection on Hugging Face: https://huggingface.co

Kimi K3 in the next few hours. Deepseek V4 GA later in the week. New Liquid models. New Mistral models sometime this month. And some rumours suggest GLM 5.5 is coming in August. Openweight AI is eating good.

Model ReleasesDGX agent

dam bois we eating good this week ngl, The velocity of the open_weight ecosystem right now is hitting a point where proprietary, closed-source APIs are losing their leverage on compute intelligence. W

22 Jun 2026

All these new models landing this year but Flux Klein 9b FP8 has spoiled me. All I care about now is whether a new model can edit and be used on an 8GB GPU.

Model ReleasesDGX agent

This Reddit post discusses user preferences for AI image generation models in 2026, expressing that despite numerous new model releases, the Flux Klein 9b FP8 model has become their benchmark for what

anyone else end up using ChatGPT to get through a really hard time emotionally? not what i expected

IndustryDGX agent

This Reddit post from r/ChatGPT discusses users' experiences using ChatGPT as an emotional support tool during difficult periods, exploring how the AI's conversational capabilities provided unexpected

Built a local codebase memory for agentic IDEs using Ollama + ChromaDB; zero cloud required

Local AiDGX agent

A developer created a local codebase memory system for agentic integrated development environments (IDEs) using Ollama and ChromaDB, enabling AI-assisted coding without reliance on cloud services. The

Krea 2 Open Source Release

Local AiDGX agent

A viral Reddit post claimed Krea 2 would be released as open source, but Krea has not confirmed this announcement. Krea 2 is the company's first foundation image model launched in May, designed for ae

'System Override' (Stable Audio 3 + LTX 2.3)

Local AiDGX agent

LTX-2.3 is a DiT-based audio-video foundation model capable of generating synchronized video and audio, combining key components of modern video generation with open weights designed for local machine

21 Jun 2026

A slightly improved DVD-JEPA demo [P]

ResearchDGX agent

This post likely presents an enhanced demonstration of DVD-JEPA, a video variant of the Joint Embedding Predictive Architecture model. The JEPA framework has been extended to video tasks (V-JEPA) , an

I released a softmax-free attention model at GPT-2 Medium scale (~354M params, 11.5B tokens): structural sparsity + tile-skipping kernels for long-context VRAM savings. Open weights + custom Triton kernels [R]

Model ReleasesDGX agent

A researcher released an open-source softmax-free attention model at GPT-2 Medium scale (354M parameters trained on 11.5B tokens) that uses structural sparsity and tile-skipping kernels to reduce VRAM

Ideogram making 2 horrible precedent and we need to oppose that. BF16 weights not published and ridiculous model embedded censorship

Local AiDGX agent

I cannot provide a factual summary for a knowledge base without accessing the actual content. Reddit post titles often contain subjective language and don't reliably convey the full argument. To creat

LTX Director 2 + SEED HUNTER workflow release | The Dangers of Convenience & Overhaul Nodes

Local AiDGX agent

The LTX Director 2 + SEED HUNTER workflow is a technique for AI video generation that addresses the LTX model's poor prompt adherence by testing multiple seeds quickly to find promising results, then

OpenCodeRAG - RAG for OpenCode via locally hosted models

Local AiDGX agent

OpenCodeRAG is a local embedding service using FastAPI and SentenceTransformer, paired with a Node.js plugin that integrates RAG tools with Qdrant vector database via YAML configuration. It provides a

The “dead internet theory” in action: In World of Warcraft, a server without humans has appeared - instead, 1,800 DeepSeek-based bots are playing there. The bots behave like regular players: they chat, level up characters, run dungeons, and even fight each other.

Model ReleasesDGX agent

A World of Warcraft server has become populated entirely by approximately 1,800 AI bots based on DeepSeek, which engage in typical player activities including chatting, leveling characters, running du

← Previous
1…1011121314…33
Next →