AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
Human
84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
84,660 results
26 Jul 2026

100B Models on Cheap Hardware: how realistic and limitations

Model ReleasesDGX agent

There is a lot of buzz around running 100B parameter models on cheap local hardware using ternary (1.58-bit) quantization like Microsoft's BitNet architecture. The theoretical hardware shortcuts are i

16 bit better than lower quants for Qwen3.6-27B

Model ReleasesDGX agent

I am writing a fairly complex C++ windows MFC application. I have a few 3090s and can run F16 Qwen3.6-27B with 256K context and MTP. The quality of code is exceptional with this quant vs its lower qua

23 Gemma4-E4B models compared with abliterlitics: the most downloaded one is also the most broken

Model ReleasesDGX agent

This is our biggest comparison yet. We've taken 23 Gemma 4 E4B models from huggingface and ran them through the abliterlitics gauntlet. We also have a new abliterlitics discord, feel free to jump on a

DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

90 agentic bakeoff runs: ThinkingCap vs Fable Fusion vs stock Qwen3.6-27B

Model ReleasesDGX agent

Last week someone here said ThinkingCap and Fable Fusion 'really do beat the OG' for agentic work, so I ran it: 6 self-grading tasks, 5 reps, 3 models, 90 isolated runs. Tooling, since that's half the

again, the log is the agent

AgentsDGX agent

The thread argues that applied AI has entered a distributed‑systems phase, with event‑driven architectures treating logs as “agents” that consume data rather than serve as endpoints. It notes that alm

ai-sage/GigaChat3.1-Audio-10B-A1.8B · Hugging Face

Local AiDGX agent

GigaChat Audio 10B is an audio-native LLM built on top of the GigaChat 3.1 Lightning text model. A Conformer speech encoder and a modality adapter feed audio embeddings directly into a Mixture-of-Expe

An Inside Look at the Relay Market Powering Token Resellers and Fraud

ToolsDGX agent

An Inside Look at the Relay Market Powering Token Resellers and Fraud Fascinating investigation by Matt Lenhard into the market that has grown up around reselling LLM tokens at a discount by pooling A

Announcing the Claude Code-compatible interface for our new Fugu-Ultra v1.1! 🐡 Put a dynamically coordinated team of frontier models to wor…

Model ReleasesDGX agent

Announcing the Claude Code-compatible interface for our new Fugu-Ultra v1.1! 🐡 Put a dynamically coordinated team of frontier models to work inside the coding workflow you already know. Instead of rel

b10141

Model ReleasesDGX agent

mtmd: fix android build (#26150) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubunt

BeeLlama.cpp v0.4.1: KVarN, KV precision tail, q2_0-q3_1 KV cache, improved support. KLD benchmarks: tail 1024 makes kvarn5 and q6_0 match q8_0, for much less VRAM

Model ReleasesDGX agent

TL;DR llama.cpp fork with more KV cache quantization features, with all claims supported by benchmarks: KVarN, KV cache precision tail, additional types of standard KV cache (q2_0-q3_1, q6_0, q6_1), a

BREAKING: A Redditor just discovered that shared Claude conversations have been showing up in public search results. The post has 4K upvotes…

Model ReleasesDGX agent

BREAKING: A Redditor just discovered that shared Claude conversations have been showing up in public search results. The post has 4K upvotes and hundreds of comments, so this is spreading fast. Here's

CEO of Hugging Face: 'In the spirit of transparency, here’s what I asked OpenAI'

Local AiDGX agent

clem 🤗 on 𝕏: https://x.com/ClementDelangue/status/2081056675558195657 • Radical transparency: let’s release the traces from the “rogue” agents so the entire research community can study what happened.

Do people building local LLM rigs track RTX Ada/workstation card prices, or just consumer cards like the 5090?

Local AiDGX agent

curious how people here approach buying high-end/workstation cards (RTX 6000 Ada, 5000 Ada, etc) for local LLM work, do you actively watch pricing/timing on these specifically, or is the consumer 5090

GLM 5.2 and ik_llama.ccp

Model ReleasesDGX agent

Running GLM-5.2 (the new glm-dsa arch), Unsloth UD-Q4_K_XL, on a 4-socket Xeon E7-8880 v4 box with 1TB RAM and a single RTX 3060 12GB. ik_llama.cpp, experts on CPU (--cpu-moe), 24 attention layers on

Harness showdown: Claude Code vs OpenCode vs Pi with DeepSeek V4 Flash

Model ReleasesDGX agent

I ran DeepSeek V4 Flash through Claude Code, OpenCode and Pi on my own benchmark, and the quality came out basically the same across all three while the time and tokens spent was wildly different. Cla

Has anyone compared pre-training, SFT/LoRA and reinforcement post-training on Qwen3.6-27B?

Model ReleasesDGX agent

Qwen3.6-27B: SFT vs continued pre-training vs RL? I’m interested in adapting Qwen3.6-27B, but I’m increasingly unsure whether conventional SFT/LoRA is the best route if the goal is to add a capability

How AI companies are targeting the education market, including making free or cut-price tailored learning tools in partnership with schools and edtech startups (Jamie John/Financial Times)

ApplicationsDGX agent

Jamie John / Financial Times: How AI companies are targeting the education market, including making free or cut-price tailored learning tools in partnership with schools and edtech startups — Anthropi

How much is the GPU usage?

Local AiDGX agent

I am trying to decide on buying the ollama pro subscription. But their usage policy is vague as hell. I don't mind the 'GPU usage time' but how much GPU time do I actually get? I still don't seem to f

How US companies flipped from 'tokenmaxxing' to 'thrift-maxxing', mixing cheaper Chinese models with OpenAI and Anthropic, threatening the labs' IPO valuations (Wall Street Journal)

IndustryDGX agent

Wall Street Journal: How US companies flipped from “tokenmaxxing” to “thrift-maxxing”, mixing cheaper Chinese models with OpenAI and Anthropic, threatening the labs' IPO valuations — Companies big and

Hugging Face CEO calls for ‘radical transparency’ after ‘unprecedented’ OpenAI hack https://techcrunch.com/2026/07/26/hugging-face-ceo-calls…

IndustryDGX agent

Hugging Face CEO calls for ‘radical transparency’ after ‘unprecedented’ OpenAI hack https://techcrunch.com/2026/07/26/hugging-face-ceo-calls-for-radical-transparency-after-unprecedented-openai-hack/?u

I built an open-source Ollama canvas where the wires are the actual context

Local AiDGX agent

Most graph-based LLM interfaces use a canvas as a visual layer over what is still a linear chat. I wanted the graph itself to determine what Ollama receives. ThoughtDAG has one rule: wires are the con

I implemented the YOLO26n model inference from scratch using ARM64 Assembly Language (No framework) [P]

Local AiDGX agent

This was my Bachelor's Final Project: implementing YOLO26n inference completely from scratch using ARM64 Assembly Language and C, without relying on existing inference frameworks. The goal was to unde

I want to use AI coding agents for machine learning projects [D]

Model ReleasesDGX agent

I'm a software engineer who mainly builds softwaes/applications, and I'm starting to work on machine learning projects. Since ML workloads often require GPUs, I know services like Google Colab and Kag

important thread; don’t read just the first tweet (which has a caveat in the second).

SafetyDGX agent

important thread; don’t read just the first tweet (which has a caveat in the second). In case you think the problem of science slop is hypothetical: Here's the president of OpenAI retweeting wrong sci

It continues to boggle the mind how many people, who otherwise seem to have a functioning intellect, appear to lose all cognitive capacity w…

HardwareDGX agent

It continues to boggle the mind how many people, who otherwise seem to have a functioning intellect, appear to lose all cognitive capacity when it comes to thinking about actions that impact the compa

Karparthy removed Anthropic from his bio

Local AiDGX agent

Andrej Karpathy, a prominent advocate for open-source AI and a co-founder of OpenAI, appears to have removed Anthropic from his X bio, suggesting he may have left the company. Karpathy joined Anthropi

Kimi K3 Tomorrow!!!

Local AiDGX agent

What quantization or storage size tiers should we expect further down the line? This site suggests q4 and q8 are coming. But I don't know how to translate that to storage size for local inference. htt

Kimi K3 will be here on Monday, but how do you know you’re ready to scale it in production? With our latest inference platform updates you c…

Model ReleasesDGX agent

Kimi K3 will be here on Monday, but how do you know you’re ready to scale it in production? With our latest inference platform updates you can: 1/ Run shadow traffic to see how a new model performs on

Local-first LLM pipeline tracer — @trace on any function, dashboard at localhost. Feedback welcome.

Model ReleasesDGX agent

Hey r/LocalLLaMA — maintainer here, obviously biased. OpenSmith is an open-source Python tracing tool for LLM pipelines. The idea: drop u/trace on any function, run opensmith ui, get a full local dash

Multi-Tenant SaaS: Which Architecture Would You Choose? [D]

ResearchDGX agent

NOTE -> I expect answer from people who actually have experience and strong understanding of these. please give something beneficial. I'm building a SaaS platform in Sri Lanka that handles documents a

Need help with setup

Model ReleasesDGX agent

I am setting up codex+ollama+qwen3.6:27b for my hobby coding project on a windows pc with rtx5090 - earlier i tried to setup vllm in Ubuntu container but couldn’t get that to work - now using ollama,

New research from NVIDIA. Does AdamW have a scale ceiling? This work claims yes, and shows where it sits. At batch sizes up to 100M tokens f…

Model ReleasesDGX agent

New research from NVIDIA. Does AdamW have a scale ceiling? This work claims yes, and shows where it sits. At batch sizes up to 100M tokens for next-token prediction, SOAP and Muon maintain training st

Ollama Cloud Phone Number Verification Issue

Local AiDGX agent

https://preview.redd.it/umcg41g2kkfh1.png?width=502&format=png&auto=webp&s=7b1ed545cd1c7840a461101bd02d78ab4a9cb40d I tried to create an Ollama account to purchase a subscription a few months ago. But

Ollama is proud to sign @satyanadella's letter. Our mission from day one has been to make open models accessible to every developer to unloc…

Local AiDGX agent

Ollama is proud to sign @satyanadella's letter. Our mission from day one has been to make open models accessible to every developer to unlock the next frontier in America and across the globe. Open-we

Open-weight 4B models approach o3-level medical question answering in Swedish [P]

Model ReleasesDGX agent

I have been running some experiments with smaller open-weight LLMs on multiple-choice questions of Swedish medical licensing exams. On a dataset called MedQA-SWE, GPT-4 scored 84% accuracy in 2024 and

[Paper] RecGPT-V3 Technical Report

Local AiDGX agent

Large language models (LLMs) are transforming recommender systems from matching co-occurrence patterns in historical behavior toward reasoning about the intent that drives it. RecGPT-V1 pioneered this

“Pelican on a bicycle” LLM benchmark by @simonw 2024 (https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle/) is now “Call of Duty” 20…

Model ReleasesDGX agent

“Pelican on a bicycle” LLM benchmark by @simonw 2024 (https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle/) is now “Call of Duty” 2026… Claude Opus 5 one-shotted this game. EVERYTHING you see

Question regarding Ollama Cloud Metering.

Local AiDGX agent

How does it work? I read that the limts were more than what Opencode Go offers and subscribed to the 20USD Pro plan. But In practice based on how the usage bar fills up it seems like the 5 hour window

Recent discussions about open-sourcing make me feel that I should go back and revisit these important open-source works in representation le…

ResearchDGX agent

Recent discussions about open-sourcing make me feel that I should go back and revisit these important open-source works in representation learning that pushed the field forward and eventually made vis

Simple desktop GUI for multiple local TTS models (Tkinter)

Model ReleasesDGX agent

https://preview.redd.it/qoogv5m1skfh1.png?width=688&format=png&auto=webp&s=3f06fb28ceec3cd3ab95bcebd0c71374332aac9b Built a simple desktop GUI (Tkinter) that supports multiple local TTS engines (curre

Simply untrue. The Economist asked Musk about a wide range of things, including AI and robotics. Watch it here for yourself: https://www.eco…

TutorialsDGX agent

Simply untrue. The Economist asked Musk about a wide range of things, including AI and robotics. Watch it here for yourself: https://www.economist.com/insider/the-insider/an-interview-with-elon-musk?u

Sources: Nvidia is in talks to provide a ~$250B backstop for OpenAI as part of a 10 GW data center project that SoftBank is developing in Ohio (Wall Street Journal)

HardwareDGX agent

Wall Street Journal: Sources: Nvidia is in talks to provide a ~$250B backstop for OpenAI as part of a 10 GW data center project that SoftBank is developing in Ohio — Project would be one of the larges

Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction

Model ReleasesDGX agent

--> .abbel-fig { display: block; text-align: center; margin: 2.4em 0; line-height: 1.4; max-width: 100%; } .abbel-fig img { display: block; margin: 0.65em auto 0; height: auto; max-width: 100%; } /* I

“The first autonomous agent cyberattack is an unprecedented event. It deserves an unprecedented response!”

AgentsDGX agent

“The first autonomous agent cyberattack is an unprecedented event. It deserves an unprecedented response!” In the spirit of transparency, here’s what I asked @OpenAI: • Radical transparency: let’s rel

The Top AI Papers of the Week (July 20 - July 26): - GAMUT - PRO-LONG - Harness Handbook - From Memory to Skills - Progressive Disclosure - …

ResearchDGX agent

The Top AI Papers of the Week (July 20 - July 26): - GAMUT - PRO-LONG - Harness Handbook - From Memory to Skills - Progressive Disclosure - Global Workspace in LLMs - Structured Output Collapses Diver

This blog is full of design choices as you build your Agentic systems. Critical choices to be made, evals need to be upgraded if the underly…

AgentsDGX agent

This blog is full of design choices as you build your Agentic systems. Critical choices to be made, evals need to be upgraded if the underlying intelligence is upgraded. Well thought out blog—new stre

This part is spot on: > 'Overall, we found that we were over-constraining Claude Code...while these constraints were once needed to avoid wo…

Model ReleasesDGX agent

This part is spot on: > 'Overall, we found that we were over-constraining Claude Code...while these constraints were once needed to avoid worst case scenarios, we have since found we can delete many o

training to benchmarks ≠ getting to AGI

SafetyDGX agent

training to benchmarks ≠ getting to AGI This suggests that Opus 5 's gain on ARC-AGI-3 was the result of specific training to improve on that eval, and not a generalized increase in abstract reasoning

Understanding GPU Inference Workloads [D]

HardwareDGX agent

Hey everyone, I have been looking into how people source compute for their Inference workloads (and in general). I wanted to understand some specific pain points here. If you've used online services l

We compared different LLMs on IMO 2026 [R]

Model ReleasesDGX agent

There are a few reasons why problems from International Mathematical Olympiad function as a good benchmark for LLMs: - The problems are new, not included in the training data of any model - Hard math

We open-sourced Logue — a privacy-first macOS meeting-notes + writing app that runs on-device (MLX, Apple Silicon) entirely

Local AiDGX agent

At Bitwize, we've been building Logue, a native macOS app for AI meeting notes and writing, and we just open-sourced it (MIT). We're sharing it here because the whole point is that it runs 100% on-dev

We turned the browser into the inference server. @RunAnywhereAI ( @runanywhere/web ) runs models client-side with WebGPU + WASM SIMD, stores…

AgentsDGX agent

We turned the browser into the inference server. @RunAnywhereAI ( @runanywhere/web ) runs models client-side with WebGPU + WASM SIMD, stores them in OPFS and generates with 0 outbound bytes. No backen

what am I doing wrong? (Ollama + openWebUI)

Local AiDGX agent

I have tried ollama locally on my gaming PC (5070 with 12bg of VRAM) It works pretty nice on qwen2.5-coder:7b (and 14b) So I decided to take a step further, and install an openWebUI instance on my hom

Will prices finally go down?

Local AiDGX agent

I am seeing more and more videos as posts about how OpenAI is in complete financial ruin, Anthropic isn't much better. Their expenses go with the revenue they make etc etc. Meta made big investments i

Will small model intelligence be limited by parameter count?

Model ReleasesDGX agent

Qwen3.6-27b is fantastic! It makes me wonder if there's a hard ceiling to smaller sized models. Do you guys think the ceiling of intelligence for smaller models will be constrained by factors like par

Yes, open-source / open-weight models are important for a healthy AI ecosystem. That's how we can verify things, check claims, and keep up o…

Model ReleasesDGX agent

Yes, open-source / open-weight models are important for a healthy AI ecosystem. That's how we can verify things, check claims, and keep up outside the closed labs. Plus, it gives us the freedom to run

25 Jul 2026

4x 3090, 96gb vram what Model to drive Hermes?

Model ReleasesDGX agent

3 year lurker, now i finally got my server up and running. dont know which model to choose. llama.cpp or vllm, what makes more sense? mainly single user with maybe 2-3 more additional users in family,

5700, 48GB RAM and a 3090 24Gb. Best OS and framework/model?

Model ReleasesDGX agent

Hi everyone, I have a system which I have been using for gaming, R7 5700X, 48GB DDR4, RTX3090 24GB. But I want to use it for Local AI to reduce my reliance on cloud AI providers (mainly usage limits -

A look at China's bid to build an alternative global order in AI by making open models widely available and training people in developing countries to use them (Financial Times)

IndustryDGX agent

Financial Times: A look at China's bid to build an alternative global order in AI by making open models widely available and training people in developing countries to use them — Beijing makes most am

A profile of Yang Zhilin, founder of Moonshot AI, which faced early doubts over revenue and model capabilities before its Kimi K3 delivered a 'DeepSeek moment' (Financial Times)

Model ReleasesDGX agent

Financial Times: A profile of Yang Zhilin, founder of Moonshot AI, which faced early doubts over revenue and model capabilities before its Kimi K3 delivered a “DeepSeek moment” — Known as ‘Yang the ge

← Previous
1…202203204205206…1411
Next →