AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,272 results
10 Aug 2026

Social World Models

Model ReleasesDGX agent

arXiv:2509.00559v3 Announce Type: replace Abstract: Humans intuitively navigate social interactions by simulating unspoken dynamics and reasoning about others' perspectives, even with limited informat

SparseVoxelDet: Fully Sparse Voxel Networks for Efficient Event-Based Drone Detection

Model ReleasesDGX agent

arXiv:2603.21638v2 Announce Type: replace Abstract: Event cameras excel at detecting small, fast drones, but today's detectors give away their key advantage: they convert the sparse event stream into

Spatiotemporal Agility: Time-Constrained Reinforcement Learning for Vision-Guided Dynamic Quadrupedal Interception

Model ReleasesDGX agent

arXiv:2608.06907v1 Announce Type: new Abstract: Legged robots require robust agility to perceive and interact with complex and dynamic environments within a constrained time. However, most existing qu


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Stable Curves, Unstable Items: Item-Level Scaling Heterogeneity in Video LLMs

Model ReleasesDGX agent

arXiv:2608.07014v1 Announce Type: new Abstract: Aggregate scaling curves suggest that Video LLMs improve smoothly or saturate as visual budgets grow. We show that this view can conceal large, opposing

StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection

Model ReleasesDGX agent

arXiv:2608.06477v1 Announce Type: cross Abstract: Computer-use agents (CUAs) face a growing threat from indirect prompt injection, where adversarial instructions are planted in the environment such as

Stockmark-Nemotron-3-Nano-Omni-JapanDocReader: Structured Document Parsing via Capability Injection and Forgetting Control

Model ReleasesDGX agent

arXiv:2608.06758v1 Announce Type: new Abstract: We present Stockmark-Nemotron-3-Nano-Omni-JapanDocReader, a Japanese document understanding model built from Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16

Stoicheia: Character-Level Masked Diffusion for Ancient Greek Textual Restoration, Parsing, and Metrical Scansion

Model ReleasesDGX agent

arXiv:2608.07249v1 Announce Type: new Abstract: We introduce Stoicheia, a 405M-parameter character-level masked-diffusion encoder for Ancient Greek whose input factors into five aligned, independently

Sub-Quadratic Bisimulation Metrics via Approximate Nearest Neighbors: Coverage-Augmented Guarantees and Computable Two-Sided Certificates

Model ReleasesDGX agent

arXiv:2608.06762v1 Announce Type: new Abstract: Bisimulation metrics quantify behavioral similarity in Markov decision processes, but their Wasserstein fixed-point operator updates every state pair an

Summary of Takeaways from the Minimax AMA

Model ReleasesDGX agent

Summary from https://www.reddit.com/r/StableDiffusion/comments/1vh9rtw/ama_minimax_h3_team_ask_us_anything_about_our/ This summary was compiled with AI but cross-checked manually by me for accuracy. I

Suppress and Diversify: Refining Robust Pathways for Corruption Robustness

Model ReleasesDGX agent

arXiv:2608.06712v1 Announce Type: new Abstract: Model robustness against natural image corruptions is essential for safety-critical applications. While existing methods primarily focus on implicit rep

Surg-UniWorld: A Unified Surgical World Model with Multimodal Control Experts

Model ReleasesDGX agent

arXiv:2608.06770v1 Announce Type: new Abstract: Controllable surgical world models can provide a generative foundation for surgical artificial intelligence and simulation by synthesizing realistic ins

Symbolic Graphics Programming with Large Language Models

Model ReleasesDGX agent

arXiv:2509.05208v2 Announce Type: replace Abstract: Large language models (LLMs) excel at program synthesis, yet their ability to produce symbolic graphics programs (SGPs) that render into precise vis

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2608.07314v1 Announce Type: cross Abstract: Vision-language-action (VLA) models are commonly adapted to downstream manipulation tasks via supervised fine-tuning (SFT) or online reinforcement lea

Tensor Network Kernel Machines: A JAX Framework for Machine Learning and Nonlinear System Identification

Model ReleasesDGX agent

arXiv:2608.07043v1 Announce Type: cross Abstract: Developing nonlinear models that are both expressive and computationally efficient remains a challenge in machine learning and nonlinear system identi

Test-Time Adaptation with Online Personalized Energy-Based Cache for Fine-Grained Video Expression Recognition

Model ReleasesDGX agent

arXiv:2608.06467v1 Announce Type: new Abstract: Facial expression recognition (FER) in videos is challenging because models must identify subtle, temporally evolving affective states that vary across

Tested Muse Glimmer locally on coding with OpenCode & agentic work

Model ReleasesDGX agent

Ran the model with quants (Q4) by Unsloth with latest (build from master) llama.cpp server. It takes ~20GB ram running on M5 Pro with 48GB at about 17t/s. Didn't do any reasoning loops/overthinking. O

// The Bitter Lesson of Tool Calling // Tool calling is a design choice, and the defaults are quietly costing accuracy. How so? New research…

Model ReleasesDGX agent

// The Bitter Lesson of Tool Calling // Tool calling is a design choice, and the defaults are quietly costing accuracy. How so? New research releases a generation-spanning comparison of programmatic t

The Sparsity Whisperer

Model ReleasesDGX agent

arXiv:2608.06630v1 Announce Type: new Abstract: Pruning reduces the inference cost of large language models, but existing criteria primarily preserve large activations or reconstruct layer outputs. We

The successor to Llama is here, and Meta is revitalizing focus on open weights with their new Muse Glimmer - a leading 30B param model desig…

Model ReleasesDGX agent

The successor to Llama is here, and Meta is revitalizing focus on open weights with their new Muse Glimmer - a leading 30B param model designed for always-on local agent use, small enough to run on a

There is an asymmetry in most agentic workflows that does not get talked about much: humans have many ways to talk to agents, and almost no …

Model ReleasesDGX agent

There is an asymmetry in most agentic workflows that does not get talked about much: humans have many ways to talk to agents, and almost no standardized way for agents to talk back to humans. You can

TRACE: A Multi-Layer Benchmark for Human AI Controller Coordination Under Drift and Failure

Model ReleasesDGX agent

arXiv:2608.06657v1 Announce Type: new Abstract: Modern cyber-physical and AI-assisted systems couple human operators, AI decision modules, and automated controllers in a single control loop, so trustw

TradeVerse: A Longitudinal Benchmark of Political Negotiation in International Trade

Model ReleasesDGX agent

arXiv:2608.06549v1 Announce Type: cross Abstract: LLMs are increasingly being applied to tasks involving institutional and political texts, but existing benchmarks evaluate them on isolated documents

Transformers are famously bad at arithmetic, so I set one's weights by hand (no training) and it multiplies with 100% accuracy [P]

Model ReleasesDGX agent

Obviously nobody needs a transformer that's good at multiplication. I wanted to know whether a stock transformer could do exact arithmetic if I chose its weights directly. I implemented the grade-scho

Transformers Struggle to Use Their Emergent World Models: Revisiting the Tower of Hanoi, and the Illusion of Thinking

Model ReleasesDGX agent

arXiv:2608.07077v1 Announce Type: new Abstract: The Tower of Hanoi is a simple planning puzzle that in prior work has proven challenging for large reasoning models (LRMs). Current models solve the sta

TransSLR: A Lightweight Transformer for Sign Language Recognition

Model ReleasesDGX agent

arXiv:2608.06407v1 Announce Type: cross Abstract: Automated Sign Language Recognition for under-represented languages remains a largely unsolved problem. Central African Sign Language (CASL) exemplifi

TRIBE: Predicting Team Performance via Communication Behavior Ensembles

Model ReleasesDGX agent

arXiv:2608.06926v1 Announce Type: new Abstract: Designing autonomous agents that effectively assist human teams hinges on understanding team dynamics, often without task specific knowledge. We present

UAV3DCrop: Benchmarking 3D Reconstruction in Repeated Multi-Angle UAV Crop Surveys

Model ReleasesDGX agent

arXiv:2608.06404v1 Announce Type: new Abstract: Accurate 3D crop monitoring underpins data-driven precision agriculture by enabling field-scale analysis of plant structure, growth dynamics, and manage

UniREditBench: A Unified Reasoning-based Image Editing Benchmark

Model ReleasesDGX agent

arXiv:2511.01295v3 Announce Type: replace Abstract: Recent advances in multi-modal generative models have driven substantial improvements in image editing. However, current generative models still str

v0.32.7

Model ReleasesDGX agent

Muse Glimmer Note: Muse Glimmer is currently available via initial support via Ollama's MLX engine on Apple Silicon. Additional support and optimizations for Apple Silicon, NVIDIA, AMD, and other plat

Very cool idea to have agents design complex systems by searching over the model structure itself. New research from Sakana AI introduces CE…

Model ReleasesDGX agent

Very cool idea to have agents design complex systems by searching over the model structure itself. New research from Sakana AI introduces CEDAR, which uses LLM agents to write, simulate, and refine sy

We asked an unreleased research version of Claude to take a stab at the Riemann hypothesis. It didn’t solve it, but it did make strides on a…

Model ReleasesDGX agent

We asked an unreleased research version of Claude to take a stab at the Riemann hypothesis. It didn’t solve it, but it did make strides on a related problem: it increased the lower bound for the fract

WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader

Model ReleasesDGX agent

arXiv:2608.06474v1 Announce Type: new Abstract: Large language models increasingly generate complete websites from natural-language descriptions, and reinforcement learning has become a central approa

WebRider: Persona-Conditioned Intent Controllers for Live-Web Assistance

Model ReleasesDGX agent

arXiv:2608.06704v1 Announce Type: new Abstract: Delegating a web task involves more than asking a question; it requires transferring a policy: what to verify, how to handle uncertainty, which preferen

We’re expanding our cybersecurity initiative Daybreak and introducing GPT-5.6-Cyber, a new model for advanced, authorized cybersecurity work…

Model ReleasesDGX agent

We’re expanding our cybersecurity initiative Daybreak and introducing GPT-5.6-Cyber, a new model for advanced, authorized cybersecurity work. As the threat landscape evolves, we’re putting frontier in

We've used GPT-5.6-Cyber extensively in real-world vulnerability research, including work that uncovered previously unknown vulnerabilities …

Model ReleasesDGX agent

OpenAI announced the release of GPT‑5.6‑Cyber as part of its Cybersecurity Initiative, “Daybreak.” The model is aimed at advanced, authorized security research and testing, helping trusted defenders d

When @QualiaQuanta took a shot at the Riemann, half in jest, people called her a crackpot. When Anthropic uses Claude to do the same thing, …

Model ReleasesDGX agent

When @QualiaQuanta took a shot at the Riemann, half in jest, people called her a crackpot. When Anthropic uses Claude to do the same thing, it gets a hundred thousand views in 30 minutes. It might wel

Winning by Peeking: Unenforced Budgets and Test-Set Selection Inflate Short-Budget AutoML Comparisons

Model ReleasesDGX agent

arXiv:2608.07303v1 Announce Type: new Abstract: Comparisons between AutoML systems at short time budgets -- tens of seconds rather than hours -- are common in tool READMEs and workshop papers, and the

Word doc cleaning

Model ReleasesDGX agent

I have been trying to parse word docs for use with llama3.1:8b in Ollama. I only need the text - Even if I cut and past into a text editor weird characters seem to stick around which break llama/Ollam

WorldMark: A Plug-and-Play World Knowledge Interface for Cross-Host Language Model Watermarking

Model ReleasesDGX agent

arXiv:2608.06416v1 Announce Type: cross Abstract: Watermarking traces the provenance of text produced by large language models by embedding statistically detectable signals during decoding. Existing s

YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family

Model ReleasesDGX agent

arXiv:2608.07051v1 Announce Type: new Abstract: Generic parameter-efficient fine-tuning (PEFT) methods transferred from language models can fail silently on real-time detectors, whose heterogeneous op

Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination

Model ReleasesDGX agent

arXiv:2608.07341v1 Announce Type: cross Abstract: Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. extbf{Contamination mitigation

9 Aug 2026

24 GB of VRAM is not really 24 GB for a local LLM. Here is the worksheet I use

Model ReleasesDGX agent

I kept seeing model file size compared directly with the number printed on the GPU box. That misses several memory buckets. A simple planning model is: usable capacity = advertised VRAM x 0.90 total t

[2606.05682] Beyond Output Matching: Preserving Internal Geometry in NVFP4 LLM Distillation

Model ReleasesDGX agent

Demand for low-precision inference, including NVFP4-based approaches, has grown as large language models are increasingly deployed in latency and cost constrained production environments. Quantization

A failure of ChatGPT Work & Claude Cowork is they assume that non-coders couldn't understand how to think about problems like a coder, so th…

Model ReleasesDGX agent

A failure of ChatGPT Work & Claude Cowork is they assume that non-coders couldn't understand how to think about problems like a coder, so they hide all that stuff. They should instead explain choices

AMD llama.cpp: reducing MTP buffer overhead gave me 64K → 149K context for Qwen 27B

Model ReleasesDGX agent

Available context length with and without the patch: Model: QWEN 27B ROCm stock patched Vulkan stock patched IQ4_XS Pure, single 16GB GPU 19.456 76.032 68,352 78,592 Q6_K_L on 16GB + 12GB 64,256 149,2

An Australian user's Claude-run OpenClaw agent exploited a gym API flaw and kicked another member off after the user asked if it could move him up the waitlist (ABC)

Model ReleasesDGX agent

ABC: An Australian user's Claude-run OpenClaw agent exploited a gym API flaw and kicked another member off after the user asked if it could move him up the waitlist — By national AI reporter Cam Wilso

b10332

Model ReleasesDGX agent

ci: rm GGML_HIP_ROCWMMA_FATTN (#26760) Signed-off-by: Aaron Teo aaron.teo1@ibm.com Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISAB

b10333

Model ReleasesDGX agent

ggml-cpu : fix missing Q5_0 dispatch in SpaceMiT backend (#26792) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (

Best Embedding + Reranking Model

Model ReleasesDGX agent

What Local Embedding + Reranking Models are you guys running for RAG? I went down this rabbit hole because I wanted a Embedding Model + Reranker for a Translation Memory Server. Essentially, given X p

CyberKimi just dropped strong results on one of ExploitBench’s hardest V8 bugs , points away from Mythos

Model ReleasesDGX agent

Hey everyone ! Quick share from the cyber + local LLM side of things that I found interesting. During this week’s hacker summer camp, an AI researcher and reverse malware engineer veteran 'lordx64' on

DeepSeek V4 Flash 0731 hits 82.7% on Terminal-Bench 2.1 in an independent public-harness run (445 trials)

Model ReleasesDGX agent

Disclosure: I’m the author of Ante. DeepSeek recently reported an 82.7% score on Terminal-Bench 2.1 for DeepSeek V4 Flash 0731. Its evaluation used “DeepSeek Harness minimal mode,” which hasn’t been r

DeepSeek v4 Flash 0731 locally on CPU

Model ReleasesDGX agent

After seeing the benchmark results for the full release of DS v4 Flash 0731, I replaced my 2 x 16GB DDR4 ram sticks with 2 x 32GB DDR4 ram sticks to get a max supported of 128 GB RAM, in hope to be ab

DeepSeek-V4-Flash-0731 Q8_K_XL sometimes stops mid-task in OpenCode - anyone else seeing this?

Model ReleasesDGX agent

Hey everyone, I've been experimenting with the new DeepSeek-V4-Flash-0731 release locally using the Unsloth Studio Q8_K_XL GGUF with OpenCode. Overall, it's been working really well, but I've noticed

Doom Loop: Anyone Else Having DeepSeek v4 Flash 0731 Issues on ollama cloud?

Model ReleasesDGX agent

Am I the only one having issues with DeepSeek V4 Flash? It gets stuck in a loop, as if it can't call the tools, and keeps repeating the same things endlessly without moving forward. Is it a poorly wri

endless-frontier/BigBang-v1 - qwen 3.5 finetunes

Model ReleasesDGX agent

table bench https://huggingface.co/bartowski/endless-frontier_BigBang-v1-GGUF I'm downloading this model only because Bartowski converted it to .gguf, so it might be interesting. Doubts : The headline

@GergelyOrosz I've been vibe coding a few games recently and it has given me SO much respect for game designers Churning out something that …

Model ReleasesDGX agent

@GergelyOrosz I've been vibe coding a few games recently and it has given me SO much respect for game designers Churning out something that looks like a game is pretty easy now. Building a game that's

GitHub Models is now retired

Model ReleasesDGX agent

GitHub Models is now retired I missed this news until today, when the GitHub Actions run for my simonw/research repository failed with this error message: GitHub Models is temporarily unavailable as p

Grok Build is quickly turning into an all-in-one creation environment It can now also generate images and videos with Grok Imagine directly …

Model ReleasesDGX agent

Grok Build is quickly turning into an all-in-one creation environment It can now also generate images and videos with Grok Imagine directly inside your workflow You can create custom visuals for websi

KLQ: Training-free measured rotation quantization. Beats all training-free rotation-based quantization methods on W4A4KV4-bits. Llama 3.2 1B KLQ-quantized beats SpinQuant and gets close to ReSpinQuant without GPTQ/LDLQ rounding.

Model ReleasesDGX agent

First of all, I'm not a lab, this was a solo summer research project that finally culminated into the github repo and the writeup. The repo includes a much deeper dive with methods, findings about qua

M3 16GB running Ollama (Qwen 9B) is extremely slow (10-12 mins per task). Am I doing something wrong?

Model ReleasesDGX agent

Hey everyone, I constantly see high praise for M3 and M4 Macs for local LLM inference, even the base/16GB models. However, my experience has been quite different, and I'm trying to figure out if I hav

← Previous
1…1718192021…372
Next →