AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
83,832 results
8 Aug 2026

Building a zero-dependency C inference engine for BitNet (1.58-bit) - lessons from hitting 36 tok/s on a Xeon CPU

Local AiDGX agent

Over the past few months I have been building a CPU-first inference engine from scratch in pure C99 (no Python, no CUDA, no BLAS, just GCC and make). The focus has been running 1.58-bit ternary models

Claude Code in 9 lines python

Model ReleasesDGX agent

I was wondering what a minimal coding agent implementation would look like that can be used like Claude Code or Codex Not feature-by-feature of course but basically stripping everything out that is no

CPUs and the rise of neurosymbolic AI by @garymarcus (with implications for what Artificial Intelligence implies about natural intelligence …

SafetyDGX agent

CPUs and the rise of neurosymbolic AI by @garymarcus (with implications for what Artificial Intelligence implies about natural intelligence -- Gary's & my interest ever since we worked together 37 yea

Content type
AllBlogX PostPaperYouTubeRedditGitHub

Currently serving 200 tps+ output speed for DeepSeek-V4-Flash on Ollama's cloud with zero data retention (ZDR). Have an amazing weekend 🫡

Model ReleasesDGX agent

Currently serving 200 tps+ output speed for DeepSeek-V4-Flash on Ollama's cloud with zero data retention (ZDR). Have an amazing weekend 🫡 DeepSeek-V4-Flash-0731 is now fully rolled out as the new defa

DeepSeek V4 Flash 0731 appreciation post

Model ReleasesDGX agent

I’m running DSV4F 0731 on dual spark, and honestly… wow. It’s an absolute workhorse, and the benchmarks are real. Everyday tasks with Hermes agent? Effortless. Coding tasks with OpenCode? I’m genuinel

Don't miss the bit where OpenAI first found out they were responsible for the Hugging Face attack when they reached out to HF to get one of …

ToolsDGX agent

Don't miss the bit where OpenAI first found out they were responsible for the Hugging Face attack when they reached out to HF to get one of their credentials revoked and HF told them it had already be

enabling PCI-E p2p for consumer Nvidia cards will yield you more than you think

Model ReleasesDGX agent

Disclaimer - no LLM was used to write this post/note As larger post about my setup will come later, want to give heads-up to folks who use VLLM and >= 2 GPUs. So I have pretty meaty server (8 channel

Extremely slow DSpark draft model performance (1-2 t/s) with DeepSeek-V4-Flash on llama-server compared to MTP?

Model ReleasesDGX agent

Hey everyone, I could use some advice on setting up speculative decoding correctly with llama-server. My Hardware: GPUs: RTX 4090 + RTX 6000 Pro (120GB total VRAM) RAM: 32GB I am currently testing the

Firebird Launches CIS Region’s Largest AI Factory in Armenia

Model ReleasesDGX agent

The global buildout of AI infrastructure reached a new milestone today — Firebird, an emerging AI cloud, launched the CIS region’s largest AI factory in Armenia, establishing a new AI computing hub po

Forecasting the AI bubble: When scarcity turns to surplus

IndustryDGX agent

Artificial intelligence can be technologically transformative and still produce a capital bubble. Those two ideas are not in conflict. The bubble bursting does not require AI to fail. It only requires

Has anyone here fiddled with TPUs for inference ?

Local AiDGX agent

I discovered recently that Google uses their own TPUs, like tiny ASIC cards like the toy ones that existed for bitcoin. And while it sounds inefficient the fact they use thousands of them because...th

HUD mode Hermes stops being a window you switch to and becomes a layer over the app you're both working in. Or keep it around as a little bu…

AgentsDGX agent

HUD mode Hermes stops being a window you switch to and becomes a layer over the app you're both working in. Or keep it around as a little buddy agent. Ask it random things, drag it anywhere, it's your

I tested a fresh GitHub download → Ollama → first local coding-agent task (72 seconds, no cloud API)

Model ReleasesDGX agent

I’m building DesktopLab, an open-source local-first control plane for development agents. I recorded the setup boundary that most agent demos skip: DesktopLab detects the host, proposes the supported

I use auto mode for everything and now that will be the default in Claude. Anthropic had to decide whether to prioritize maximization of hum…

Model ReleasesDGX agent

I use auto mode for everything and now that will be the default in Claude. Anthropic had to decide whether to prioritize maximization of human control or minimization of risk, and it chose the latter.

IDE with Locall LLMs?

Local AiDGX agent

What IDE are you using. its another problem area for me . I usually use VSCode , but with local llms I have not found an extension which works optimally VSCode CoPilot chat with Ollama: CoPilot bloats

Imagine image 2.0, non-agentic yet, more to come in a week or two 💙

AgentsDGX agent

Imagine image 2.0, non-agentic yet, more to come in a week or two 💙 Announcing Imagine Image 2.0, our next generation image model with precision editing, crisp text rendering, improved factuality, and

Is anyone else finding DeepSeek-V4-Flash unreliable for non-coding tasks?

Model ReleasesDGX agent

(I am not a native speaker, written by myself, so please bear with me) I really want to like DeepSeek-V4-Flash-0731. But it has serious flaws that don't align with the high score on intelligence bench

Kimi K3 (Unsloth) IQ2-XXS from 711GB down to 478GB!!! Only Multi-language was removed to trim the size

Model ReleasesDGX agent

Firstly a big thanks to the poster 'hellohazine', he basically only removed the multi-lingual fat of the model and just kept the English language intact. It is the exact model, and the rest of the mod

LiteParse can now extract structured data from your PDF in milliseconds: ✅ checkbox states ✅ annotations ✅ vector graphics ✅ word-level boun…

Model ReleasesDGX agent

LiteParse can now extract structured data from your PDF in milliseconds: ✅ checkbox states ✅ annotations ✅ vector graphics ✅ word-level bounding boxes It is the most comprehensive, accurate (and fast)

MI25 for 80-100€ worth it?

Local AiDGX agent

seems to be about as good as a vega 56 with 16Gb of VRAM, is it worth it? (don’t want to deal with NVIDIA drivers on Linux, already have an rx6650xt and might simply use vulkan for llamacpp inference)

model: support Longcat-Flash (need testing) by ngxson · Pull Request #19182 · ggml-org/llama.cpp

Model ReleasesDGX agent

This PR should be ready for testing now. I tested with a very small (8B params) sub-model extracted from the original one. Appreciate if someone can test with the bigger model. GGUF(for testing) from

My first run of Kimi K3 locally.

Model ReleasesDGX agent

Running across 2 clusters using llama.cpp over RPC too. Both clusters are not enough to hold everything in memory, so main cluster still partially offloads to run. Goal will be to get all the GPUs in

Neat example here of the agents communicating purely through file names, including adding base64-encoded attachments and using 'zz' prefixes…

ToolsDGX agent

Neat example here of the agents communicating purely through file names, including adding base64-encoded attachments and using 'zz' prefixes to ensure their new message sorts to the bottom of the list

Now we have a timeline of the OpenAI accidental attack against Hugging Face

Model ReleasesDGX agent

My comment on Now we have a timeline of the OpenAI accidental attack against Hugging Face — Hacker News.I think one of the most interesting details here might be tucked away in that first bulletin poi

Ollama Cloud reviews

Model ReleasesDGX agent

I am wondering if anyone can give opinion on if Ollama Cloud pro or max plans are worth it. Id be looking to use it with Kimi K3, Qwen 3.8 and Deepseek v4flash for now. Wondering if it would be better

@OpenAI oo claude code has this now!!! need to try https://x.com/ClaudeDevs/status/2085817074816070014

Model ReleasesDGX agent

@OpenAI oo claude code has this now!!! need to try https://x.com/ClaudeDevs/status/2085817074816070014 New in Claude Code: your sessions can now message each other. Instead of having to re-explain you

OpenAI reveals upcoming Astra model may possess ‘critical’ hacking capabilities

IndustryDGX agent

OpenAI Group PBC today disclosed that one of its unreleased large language models may pose a significant cybersecurity risk. The algorithm, which is known as Astra, was first detailed last week. OpenA

PSA for anyone with multiple V620's or other gfx1030 cards having problems making llama.cpp tensor split work -- set '-ub 384' and -b to a multiple of that depending on number of GPUs

Model ReleasesDGX agent

Basically what the title says. For me, it would always crash and burn trying to use tensor split. Apparently, there's some bug where GPU memory gets corrupted with the default microbatch (512) or high

Quick survey (2 min) on trust in hardware specs for open-source models

Local AiDGX agent

Hi everyone, I'm a systems analysis student researching a problem a lot of you probably know well: how much you actually trust the published VRAM/RAM requirements for open-source models before trying

Qwen 35B-A3B MoE vs 27B dense in local coding tests: ~4× faster, much smaller quality gap than I expected

Model ReleasesDGX agent

I compared Qwen 35B-A3B MoE against Qwen 27B dense on a series of local coding-maintenance tasks. On my R9700/llama.cpp setup, the MoE model generated about 3.9× faster (~116 vs ~30 tok/s), but the co

Qwen3.6 27B + 35B on vLLM, single R9700 (gfx1201)

Model ReleasesDGX agent

I've been tuning my new Radeon AI Pro R9700, and figured that this would be useful information for people who are trying to optimise their setups. I'm pretty happy with these results and looking forwa

Showoff Saturday: Local 4x 6000 Pro (multi-year progression)

Model ReleasesDGX agent

Not the biggest or shiniest, but it's mine From gaming machine inference on the original llama models, to a 4x RTX 6000 Pro Max Q + 4x 3090s local AI cluster. Pictures are in reverse chronological ord

Spongebob MiniMax H3 test (4 x 5 seconds) turbo 6 steps 1344×768 (AI gen post)

Local AiDGX agent

MiniMax-H3 in ComfyUI 0.30.0, RTX 4080 16 GB (224 W cap). int8 DiT + int8 Qwen3-VL-32B text encoder. Turbo LoRA @ 0.9, euler + simple, 6 steps, no CFG. MiniMaxH3ReferenceToVideo with 3 reference image

Tesla V100 Qwen3.6 27B Performance

Model ReleasesDGX agent

Looking for V100 users to share your config and it's performance. GPU: Tesla V100 PCIE 32Gb Qwen3.6 27B Q4_K_M + Q8_0 MTP 128K context length Pi coding agent llama.cpp model preset: [*] spec-default =

The reports of the demise of Google are greatly exaggerated. I wouldn't underestimate them

Model ReleasesDGX agent

François Chollet commented that claims the demise of Google were greatly exaggerated, cautioning against undervaluation. According to a Polymarket report, Sergey Brin is expected to take direct oversi

We analyzed DeepSeek V4 Flash and GPT-5.6 Luna on DeepSWE. A DeepSeek-first cascade with test-suite verification solved MORE tasks than Luna…

Model ReleasesDGX agent

Researchers from TogetherAI analyzed DeepSeek V4 Flash and GPT‑5.6 Luna on the DeepSWE benchmark. The study found that employing a DeepSeek‑first cascade with test‑suite verification solved more tasks

Weirdly iirc stable diffusion (1.4) finished training around four years ago today too

Model ReleasesDGX agent

Emad posted that Stable Diffusion v1.4 reached the end of its training cycle roughly four years before the post was published. Greg Brockman added that GPT‑4 similarly completed training around the sa

When you move a model into production, you want the quality you evaluated to carry through the serving stack. @Kimi_Moonshot benchmarked Kim…

ApplicationsDGX agent

When you move a model into production, you want the quality you evaluated to carry through the serving stack. @Kimi_Moonshot benchmarked Kimi K3 across major inference providers, and Together AI ranke

You may have been told to watch this video about the OpenAI AI hack. You really should, even if you don't usually care about tech stuff. If …

ApplicationsDGX agent

You may have been told to watch this video about the OpenAI AI hack. You really should, even if you don't usually care about tech stuff. If nothing else, click this link to the 18 minutes in & see how

7 Aug 2026

100% Local RAG Without Internet and Without Ollama

Model ReleasesDGX agent

Build a 100% offline fast Retrieval Augmented Generation (RAG) system that runs without an internet connection, without cloud APIs, without OpenAI/Ollama Published a video where you can build a fully

~45% lower MiniMax H3 sampler time with new Spectrum settings — degree 1 works surprisingly well (v0.1.8)

Model ReleasesDGX agent

Follow-up to my original Spectrum MiniMax H3 post: https://www.reddit.com/r/StableDiffusion/comments/1vf1ze3/spectrum_acceleration_for_minimax_h3_in_comfyui/ In that first post, I released the MiniMax

A Bridge from Audio to Video: Phoneme-Viseme Alignment Allows Every Face to Speak Multiple Languages

SafetyDGX agent

arXiv:2510.06612v2 Announce Type: replace Abstract: Speech-driven talking face synthesis (TFS) focuses on generating lifelike facial animations from speech input. Current TFS models perform well in En

A Foundational EDM2-Based Generative Model for High-Resolution Synthetic Fetal Ultrasound Imaging from Open Datasets

ResearchDGX agent

arXiv:2608.05471v1 Announce Type: cross Abstract: Prenatal ultrasound imaging is key for assessing fetal health, but AI progress is limited by scarce, privacy-restricted, and hard-to-annotate datasets

A Lexical Analysis of online Reviews on Human-AI Interactions

ResearchDGX agent

arXiv:2511.13480v2 Announce Type: replace-cross Abstract: This study focuses on understanding the complex dynamics between humans and AI systems by analyzing user reviews. While previous research has

A llama.cpp PR makes Q2_0 3.0–3.6x faster on x86 CPUs, 8B decode goes 2.39 → 8.20 tok/s

Model ReleasesDGX agent

I was going through the current llama.cpp CPU PRs and #26348 stood out because this isn't the usual +5% kernel optimization. It adds an x86 VNNI implementation for the Q2_0 × Q8_0 dot product, and the

A lot of the questions we get from developers are about the concepts behind the API: TTFT, context windows, sampling, fine-tuning, quantizat…

TutorialsDGX agent

A lot of the questions we get from developers are about the concepts behind the API: TTFT, context windows, sampling, fine-tuning, quantization, deployment tradeoffs. We added Learn to the Together do

A Low-Power Wearable Respiratory Sensor for Non-Invasive Stress Monitoring

ResearchDGX agent

arXiv:2608.05697v1 Announce Type: cross Abstract: Respiration provides a continuously available window into physiological state and behavior. However, monitoring it outside controlled settings remains

A Master-Salve Robot Manipulator for Needle-Based Teleoperation in MRI Chamber

ResearchDGX agent

arXiv:2608.06354v1 Announce Type: new Abstract: We present a MR safe, master-slave robot manipulator for abdominal interventions in the MRI chamber. A human operated 2+1-DoF master controller manipula

A Multi-Layer System for Ultra-High-Resolution Static 360-Degree Telepresence

ResearchDGX agent

arXiv:2608.05570v1 Announce Type: cross Abstract: 360-degree video telepresence offers strong immersive potential but remains constrained by the limited resolution of current capture and display hardw

A neural operator view on U-Nets for inverse imaging problems

ResearchDGX agent

arXiv:2608.05839v1 Announce Type: cross Abstract: Deep neural networks have shown great empirical success in the solution of a wide variety of ill-posed inverse problems in imaging. Yet, very few work

A note on conditional PAC-efficient reasoning in large language model routing

ResearchDGX agent

arXiv:2512.03057v2 Announce Type: replace-cross Abstract: We study distribution-free risk control for model routing, motivated by large language model reasoning. We formalize pointwise conditional eff

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval

Model ReleasesDGX agent

arXiv:2608.05260v1 Announce Type: new Abstract: Contrastive vision-language models such as CLIP and BLIP are typically trained on short image captions, limiting their ability to retrieve images from d

A Reverse-BSDE Diffusion Sampler

ResearchDGX agent

arXiv:2505.06800v2 Announce Type: replace-cross Abstract: Diffusion-based generative models have renewed interest in stochastic differential equation methods for sampling from complex distributions. W

A Six-Dimensional Taxonomy of Post-Training Adaptation Techniques with Applications in AI Governance

Model ReleasesDGX agent

arXiv:2608.06246v1 Announce Type: new Abstract: Post-training adaptation has become central to modern machine learning practice and includes techniques such as retraining, fine-tuning, parameter-effic

A Study of ASR Adaptation and Representation Dimensionality Reduction in Persian Speech Emotion Recognition Using Whisper

ResearchDGX agent

arXiv:2608.05165v1 Announce Type: cross Abstract: Speech Emotion Recognition (SER) in low-resource languages remains a challenging problem due to limited labeled data. In this work, we study the use o

A Survey of Adversarial Efficiency Degradation for Vision Transformer by Exploiting Input-adaptive Optimization

ResearchDGX agent

arXiv:2608.05217v1 Announce Type: cross Abstract: Vision Transformers (ViTs) increasingly rely on input-adaptive inference, such as token pruning and early halting, to meet energy and latency budgets.

A System for Train Condition Monitoring and Structural Health Assessment of Rail Vehicles

SafetyDGX agent

arXiv:2608.05221v1 Announce Type: new Abstract: The ongoing digitalization of rail systems and the increasing use of artificial intelligence (AI) are fundamentally transforming the design, operation,

A Two-Tier Perspective on Inference-Time Parallelism in Multi-Agent LLM Systems

Model ReleasesDGX agent

arXiv:2608.05791v1 Announce Type: cross Abstract: Large language model (LLM)-driven multi-agent systems typically require multiple model invocations and complex coordination during inference, and thei

A Unified Causal Inference Framework for the Desirability of Outcome Ranking Paradigm in Benefit-Risk Evaluation

SafetyDGX agent

arXiv:2608.05244v1 Announce Type: cross Abstract: We developed a unified covariate-adjusted causal inference framework for estimating the desirability of outcome ranking (DOOR) probability for benefit

A Unified Framework for Trajectory Prediction with Explicit Planning and Reaction Decomposition

ResearchDGX agent

arXiv:2608.05673v1 Announce Type: new Abstract: Trajectory prediction has shifted toward structured formulations with explicit social modeling. However, existing methods inadequately distinguish the f

← Previous
1…5758596061…1398
Next →