AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
All
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,272 results
Model Releases

You don’t need new ways to talk to your agents, you need new ways for your agents to talk to you 🫵 (do you?) Introducing Remoko: your mobil…

DGX agent

You don’t need new ways to talk to your agents, you need new ways for your agents to talk to you 🫵 (do you?) Introducing Remoko: your mobile agent relay http://remoko.app I wanted a way for my long-ru

model-releasesyohei-nakajima--x
9 Aug 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Anthropic says auto mode will be the default in Claude Code for Pro, Max, Team plans, starting on Aug. 14, claiming it's good enough at catching harmful actions (Simon Willison/Simon Willison's Weblog)

DGX agent

Simon Willison / Simon Willison's Weblog: Anthropic says auto mode will be the default in Claude Code for Pro, Max, Team plans, starting on Aug. 14, claiming it's good enough at catching harmful actio

model-releasestechmeme
8 Aug 2026
Model Releases

any reasonably fast public benchmarks I should run quants of deepseek flash 0731 on?

DGX agent

I have various quants of this model and am curious how they perform. can anyone recommend which benchmark would be a good test case for quantization effects? Maybe that can be completed with about 1 m

model-releasesr-localllama
8 Aug 2026
Model Releases

Anyone else amped up over Qwen 3.8?

DGX agent

I’ve been using 3.6 27B Q4, and that quant is fast on an M5. The code has been average, but consistently “good enough.” And, after a year, I can see home LLMs being served at home much like streaming

model-releasesr-localllama
8 Aug 2026
Model Releases

Auto mode is now the default in Claude Code for Pro, Max, and Team plans

DGX agent

Auto mode is now the default in Claude Code for Pro, Max, and Team plans Anthropic are really confident in Claude Code's auto mode, to the point that they are making it the default setting for new ses

model-releasessimon-willison
8 Aug 2026
Model Releases

b10327

DGX agent

CUDA: fix thread/block count in quantized cpy kernel launches (#26731) CUDA: fix thread/block count in quantized cpy kernel launches tests: add uneven block count cpy case Website: https://llama.app m

model-releasesllama-cpp-releases
8 Aug 2026
Model Releases

b10328

DGX agent

server: add initial tool isolation support (via docker) (#26507) server: add initial tool isolation support (via docker) add docs adapt get_info py: fix type check cont separate tools_io_sandbox / too

model-releasesllama-cpp-releases
8 Aug 2026
Model Releases

b10329

DGX agent

server, ui: only offer a working directory when a tool reads it (#26762) The working directory chip showed up as soon as the server exposed any builtin tool, so a server started with just get_datetime

model-releasesllama-cpp-releases
8 Aug 2026
Model Releases

b10330

DGX agent

CUDA: fuse rms_norm + mul + rope (+ view + set_rows) (#26767) CUDA: fuse rms_norm + mul + rope (+ view + set_rows) tests: add broadcast weight case to rms_norm_mul_rope CUDA: check memory ranges befor

model-releasesllama-cpp-releases
8 Aug 2026
Model Releases

b10331

DGX agent

server: report the isolate working directory from get_info (#26773) server: report the isolate working directory from get_info Without an explicit cwd, get_info fell back to the server process working

model-releasesllama-cpp-releases
8 Aug 2026
Model Releases

Building a budget 32GB → 48GB VRAM home AI server: 2-3x RX 9060 XT 16GB vs RTX 5060 Ti 16GB, AM5 vs used EPYC?

DGX agent

I’m planning a dedicated home AI server, mainly for local LLM inference, agents/tool use, Docker services, and eventually larger MoE models with CPU offload. My plan is to start with 2x 16GB GPUs = 32

model-releasesr-localllama
8 Aug 2026
Model Releases

Claude Code in 9 lines python

DGX agent

I was wondering what a minimal coding agent implementation would look like that can be used like Claude Code or Codex Not feature-by-feature of course but basically stripping everything out that is no

model-releasesr-localllama
8 Aug 2026
Model Releases

Currently serving 200 tps+ output speed for DeepSeek-V4-Flash on Ollama's cloud with zero data retention (ZDR). Have an amazing weekend 🫡

DGX agent

Currently serving 200 tps+ output speed for DeepSeek-V4-Flash on Ollama's cloud with zero data retention (ZDR). Have an amazing weekend 🫡 DeepSeek-V4-Flash-0731 is now fully rolled out as the new defa

model-releasesollama--x
8 Aug 2026
Model Releases

DeepSeek V4 Flash 0731 appreciation post

DGX agent

I’m running DSV4F 0731 on dual spark, and honestly… wow. It’s an absolute workhorse, and the benchmarks are real. Everyday tasks with Hermes agent? Effortless. Coding tasks with OpenCode? I’m genuinel

model-releasesr-localllama
8 Aug 2026
Model Releases

enabling PCI-E p2p for consumer Nvidia cards will yield you more than you think

DGX agent

Disclaimer - no LLM was used to write this post/note As larger post about my setup will come later, want to give heads-up to folks who use VLLM and >= 2 GPUs. So I have pretty meaty server (8 channel

model-releasesr-localllama
8 Aug 2026
Model Releases

Extremely slow DSpark draft model performance (1-2 t/s) with DeepSeek-V4-Flash on llama-server compared to MTP?

DGX agent

Hey everyone, I could use some advice on setting up speculative decoding correctly with llama-server. My Hardware: GPUs: RTX 4090 + RTX 6000 Pro (120GB total VRAM) RAM: 32GB I am currently testing the

model-releasesr-localllama
8 Aug 2026
Model Releases

Firebird Launches CIS Region’s Largest AI Factory in Armenia

DGX agent

The global buildout of AI infrastructure reached a new milestone today — Firebird, an emerging AI cloud, launched the CIS region’s largest AI factory in Armenia, establishing a new AI computing hub po

model-releasesnvidia-blog
8 Aug 2026
Model Releases

I tested a fresh GitHub download → Ollama → first local coding-agent task (72 seconds, no cloud API)

DGX agent

I’m building DesktopLab, an open-source local-first control plane for development agents. I recorded the setup boundary that most agent demos skip: DesktopLab detects the host, proposes the supported

model-releasesr-ollama
8 Aug 2026
Model Releases

I use auto mode for everything and now that will be the default in Claude. Anthropic had to decide whether to prioritize maximization of hum…

DGX agent

I use auto mode for everything and now that will be the default in Claude. Anthropic had to decide whether to prioritize maximization of human control or minimization of risk, and it chose the latter.

model-releasesallie-k--miller--x
8 Aug 2026
Model Releases

Is anyone else finding DeepSeek-V4-Flash unreliable for non-coding tasks?

DGX agent

(I am not a native speaker, written by myself, so please bear with me) I really want to like DeepSeek-V4-Flash-0731. But it has serious flaws that don't align with the high score on intelligence bench

model-releasesr-localllama
8 Aug 2026
Model Releases

Kimi K3 (Unsloth) IQ2-XXS from 711GB down to 478GB!!! Only Multi-language was removed to trim the size

DGX agent

Firstly a big thanks to the poster 'hellohazine', he basically only removed the multi-lingual fat of the model and just kept the English language intact. It is the exact model, and the rest of the mod

model-releasesr-localllama
8 Aug 2026
Model Releases

LiteParse can now extract structured data from your PDF in milliseconds: ✅ checkbox states ✅ annotations ✅ vector graphics ✅ word-level boun…

DGX agent

LiteParse can now extract structured data from your PDF in milliseconds: ✅ checkbox states ✅ annotations ✅ vector graphics ✅ word-level bounding boxes It is the most comprehensive, accurate (and fast)

model-releasesjerry-liu--x
8 Aug 2026
Model Releases

model: support Longcat-Flash (need testing) by ngxson · Pull Request #19182 · ggml-org/llama.cpp

DGX agent

This PR should be ready for testing now. I tested with a very small (8B params) sub-model extracted from the original one. Appreciate if someone can test with the bigger model. GGUF(for testing) from

model-releasesr-localllama
8 Aug 2026
Model Releases

My first run of Kimi K3 locally.

DGX agent

Running across 2 clusters using llama.cpp over RPC too. Both clusters are not enough to hold everything in memory, so main cluster still partially offloads to run. Goal will be to get all the GPUs in

model-releasesr-localllama
8 Aug 2026
Model Releases

Now we have a timeline of the OpenAI accidental attack against Hugging Face

DGX agent

My comment on Now we have a timeline of the OpenAI accidental attack against Hugging Face — Hacker News.I think one of the most interesting details here might be tucked away in that first bulletin poi

model-releasessimon-willison
8 Aug 2026
Model Releases

Ollama Cloud reviews

DGX agent

I am wondering if anyone can give opinion on if Ollama Cloud pro or max plans are worth it. Id be looking to use it with Kimi K3, Qwen 3.8 and Deepseek v4flash for now. Wondering if it would be better

model-releasesr-ollama
8 Aug 2026
Model Releases

@OpenAI oo claude code has this now!!! need to try https://x.com/ClaudeDevs/status/2085817074816070014

DGX agent

@OpenAI oo claude code has this now!!! need to try https://x.com/ClaudeDevs/status/2085817074816070014 New in Claude Code: your sessions can now message each other. Instead of having to re-explain you

model-releasesswyx--x
8 Aug 2026
Model Releases

PSA for anyone with multiple V620's or other gfx1030 cards having problems making llama.cpp tensor split work -- set '-ub 384' and -b to a multiple of that depending on number of GPUs

DGX agent

Basically what the title says. For me, it would always crash and burn trying to use tensor split. Apparently, there's some bug where GPU memory gets corrupted with the default microbatch (512) or high

model-releasesr-localllama
8 Aug 2026
Model Releases

Qwen 35B-A3B MoE vs 27B dense in local coding tests: ~4× faster, much smaller quality gap than I expected

DGX agent

I compared Qwen 35B-A3B MoE against Qwen 27B dense on a series of local coding-maintenance tasks. On my R9700/llama.cpp setup, the MoE model generated about 3.9× faster (~116 vs ~30 tok/s), but the co

model-releasesr-localllama
8 Aug 2026
Model Releases

Qwen3.6 27B + 35B on vLLM, single R9700 (gfx1201)

DGX agent

I've been tuning my new Radeon AI Pro R9700, and figured that this would be useful information for people who are trying to optimise their setups. I'm pretty happy with these results and looking forwa

model-releasesr-localllama
8 Aug 2026
Model Releases

Showoff Saturday: Local 4x 6000 Pro (multi-year progression)

DGX agent

Not the biggest or shiniest, but it's mine From gaming machine inference on the original llama models, to a 4x RTX 6000 Pro Max Q + 4x 3090s local AI cluster. Pictures are in reverse chronological ord

model-releasesr-localllama
8 Aug 2026
Model Releases

Tesla V100 Qwen3.6 27B Performance

DGX agent

Looking for V100 users to share your config and it's performance. GPU: Tesla V100 PCIE 32Gb Qwen3.6 27B Q4_K_M + Q8_0 MTP 128K context length Pi coding agent llama.cpp model preset: [*] spec-default =

model-releasesr-localllama
8 Aug 2026
Model Releases

The reports of the demise of Google are greatly exaggerated. I wouldn't underestimate them

DGX agent

François Chollet commented that claims the demise of Google were greatly exaggerated, cautioning against undervaluation. According to a Polymarket report, Sergey Brin is expected to take direct oversi

model-releasesfrancois-chollet--x
8 Aug 2026
Model Releases

We analyzed DeepSeek V4 Flash and GPT-5.6 Luna on DeepSWE. A DeepSeek-first cascade with test-suite verification solved MORE tasks than Luna…

DGX agent

Researchers from TogetherAI analyzed DeepSeek V4 Flash and GPT‑5.6 Luna on the DeepSWE benchmark. The study found that employing a DeepSeek‑first cascade with test‑suite verification solved more tasks

model-releasestogether-ai--x
8 Aug 2026
Model Releases

Weirdly iirc stable diffusion (1.4) finished training around four years ago today too

DGX agent

Emad posted that Stable Diffusion v1.4 reached the end of its training cycle roughly four years before the post was published. Greg Brockman added that GPT‑4 similarly completed training around the sa

model-releasesemad-mostaque--x
8 Aug 2026
Model Releases

100% Local RAG Without Internet and Without Ollama

DGX agent

Build a 100% offline fast Retrieval Augmented Generation (RAG) system that runs without an internet connection, without cloud APIs, without OpenAI/Ollama Published a video where you can build a fully

model-releasesr-ollama
7 Aug 2026
Model Releases

~45% lower MiniMax H3 sampler time with new Spectrum settings — degree 1 works surprisingly well (v0.1.8)

DGX agent

Follow-up to my original Spectrum MiniMax H3 post: https://www.reddit.com/r/StableDiffusion/comments/1vf1ze3/spectrum_acceleration_for_minimax_h3_in_comfyui/ In that first post, I released the MiniMax

model-releasesr-stablediffusion
7 Aug 2026
Model Releases

A llama.cpp PR makes Q2_0 3.0–3.6x faster on x86 CPUs, 8B decode goes 2.39 → 8.20 tok/s

DGX agent

I was going through the current llama.cpp CPU PRs and #26348 stood out because this isn't the usual +5% kernel optimization. It adds an x86 VNNI implementation for the Q2_0 × Q8_0 dot product, and the

model-releasesr-localllama
7 Aug 2026
Model Releases

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval

DGX agent

arXiv:2608.05260v1 Announce Type: new Abstract: Contrastive vision-language models such as CLIP and BLIP are typically trained on short image captions, limiting their ability to retrieve images from d

model-releasesarxiv-cs-cv
7 Aug 2026
Model Releases

A Six-Dimensional Taxonomy of Post-Training Adaptation Techniques with Applications in AI Governance

DGX agent

arXiv:2608.06246v1 Announce Type: new Abstract: Post-training adaptation has become central to modern machine learning practice and includes techniques such as retraining, fine-tuning, parameter-effic

model-releasesarxiv-cs-lg
7 Aug 2026
Model Releases

A Two-Tier Perspective on Inference-Time Parallelism in Multi-Agent LLM Systems

DGX agent

arXiv:2608.05791v1 Announce Type: cross Abstract: Large language model (LLM)-driven multi-agent systems typically require multiple model invocations and complex coordination during inference, and thei

model-releasesarxiv-cs-ai
7 Aug 2026
Model Releases

A Unified Risk View of Uncertainty: Posterior Risk for Disentanglement and Evaluation Beyond Proxies

DGX agent

arXiv:2608.05995v1 Announce Type: new Abstract: Reliable uncertainty estimates are critical in safety-sensitive applications, where understanding the sources of predictive uncertainty is essential. Th

model-releasesarxiv-cs-lg
7 Aug 2026
Model Releases

Abstract Event Causal Rules: Induction and Application

DGX agent

arXiv:2608.05205v1 Announce Type: new Abstract: Event-centric intelligent analytical systems heavily depend on explicit causal event knowledge for risk early warning, decision-making support and narra

model-releasesarxiv-cs-ai
7 Aug 2026
Model Releases

Adapting Vision Foundation Models with Cascaded Semantics

DGX agent

arXiv:2608.05393v1 Announce Type: new Abstract: Prompt tuning, a leading parameter-efficient adaptation paradigm in NLP, has recently been extended to computer vision. Visual prompt tuning (VPT) adapt

model-releasesarxiv-cs-cv
7 Aug 2026
Model Releases

Adaptive-WAM: Quality-Guided Early-Exit Planning from Intermediate Video-Diffusion Features

DGX agent

arXiv:2608.06008v1 Announce Type: new Abstract: Large video diffusion models provide rich spatiotemporal priors for autonomous driving, but existing world-action models often inherit the cost of itera

model-releasesarxiv-cs-ro
7 Aug 2026
Model Releases

Afford-X: Generalizable and Slim Affordance Reasoning for Task-oriented Manipulation

DGX agent

arXiv:2503.03556v3 Announce Type: replace Abstract: Object affordance reasoning, the ability to infer object functionalities based on physical properties, is fundamental for task-oriented planning and

model-releasesarxiv-cs-cv
7 Aug 2026
Model Releases

After evaluating one of our upcoming models, Astra, we're treating it as our first 'critical' model for cybersecurity under our Preparedness…

DGX agent

After evaluating one of our upcoming models, Astra, we're treating it as our first 'critical' model for cybersecurity under our Preparedness Framework. This is a scenario we've planned for, and we're

model-releasesopenai--x
7 Aug 2026
Model Releases

Agentic self-driving microscopy benchmarks support qualification but do not necessarily generalize to unseen tasks

DGX agent

arXiv:2608.05266v1 Announce Type: new Abstract: Large language model agents are increasingly being developed to control a wide range of scientific characterization tools including microscopes and sync

model-releasesarxiv-cs-ai
7 Aug 2026
← Previous
1…2324252627…464
Next →