AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,332 results
31 Jul 2026

DeepSeek V4 Flash 0731 in Hermes Agent and one prompt, took 32 minutes and cost 0.07$, this model is so cheap to the point where 2 dollars c…

Model ReleasesDGX agent

**DeepSeek V4 Flash 0731 Performance Test** On July 31 2026, a single prompt executed via the Hermes Agent on DeepSeek V4 Flash 0731 completed in 32 minutes and incurred an estimated cost of 0.07 USD.

DeepSeek-V4-Flash-0731 unsloth gguf on A100

Model ReleasesDGX agent

A100 with 40gb VRAM: 162GB Q8_K_XL ~16.1 tok/s generation Only 15.8GB of 40GB VRAM used with all experts on CPU NOTE just tested coding on linux box DeepSeek-V4-Flash-0731 runs losslessly on the singl

DeepSeek v4 Flash for DS4 (DwarfStar) GGUF w/ DSpark MTP Head

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

I'm an avid user of Deepseek v4 Flash via antirez's DS4 DwarfStar inference engine, and so when the new checkpoint dropped, the first thing I did was rent a cloud box and spin up a quantization for us

DeepSeek V4 Flash GA ranks the same as Sonnet 5 and Grok 4.5 on DeepSWE

Model ReleasesDGX agent

Source: https://x.com/deepseek_ai/status/2083084415157022911 & https://deepswe.datacurve.ai/ just combined data view. DeepSeek claims, not verified by DeepSWE yet. submitted by /u/sdexca [link] [comme

DeepSeek v4 Flash has a nice bump in Capability

Model ReleasesDGX agent

DeepSeek V4 Flash: Preview → 2026-07-31 Benchmark Preview 0731 Δ Terminal Bench* 56.9 82.7 +25.8 Toolathlon 51.8 70.3 +18.5 NL2Repo — 54.2 new Cybergym — 76.7 new DeepSWE — 54.4 new Agent Last Exam —

Deepseek V4 Flash is now ~#2 open weight model to Kimi K3 and >50x cheaper

Model ReleasesDGX agent

https://preview.redd.it/h7zv5tb3tmgh1.png?width=2854&format=png&auto=webp&s=507380e8f862c18f10f7c5c84da9e8d1c59139b0 Deepseek's new flash model is unexpectedly cheap and high-performing across useful

🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta! 🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now fa…

Model ReleasesDGX agent

🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta! 🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive perform

Deepseek V4 Flash on SlopCodeBench

Model ReleasesDGX agent

While waiting for some of the quants to drop, I load the API with $50 and ran it on SlopCodeBench Just vibe reading the results it seems like Opus 4.8 < Deepseek < Opus 5 https://github.com/michaelasp

DeepSeek V4 Flash scores 50 on the Artificial Analysis Intelligence Index, matching Gemini 3.6 Flash and up 10 points from the preview launch in April (Artificial Analysis)

Model ReleasesDGX agent

Artificial Analysis: DeepSeek V4 Flash scores 50 on the Artificial Analysis Intelligence Index, matching Gemini 3.6 Flash and up 10 points from the preview launch in April — DeepSeek V4 Flash 0731 is

Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillation

Model ReleasesDGX agent

arXiv:2606.13657v3 Announce Type: replace Abstract: On-policy distillation (OPD) has recently become a prominent post-training recipe by combining two desirable ingredients: on-policy student-generate

Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset

Model ReleasesDGX agent

arXiv:2607.27420v1 Announce Type: cross Abstract: Humanity's Last Exam (HLE) is widely used to evaluate frontier language models. HLE organizes its questions into eight subject-domain categories, whos

Divergence Decoding: Training-Free Capability Fusion

Model ReleasesDGX agent

arXiv:2607.27248v1 Announce Type: cross Abstract: While large language models excel in reasoning, these generalists often lack knowledge for specialized scientific domains. Conversely, domain models~(

DoTime: A Synthetic Benchmark Generator for Interventional and Counterfactual Time Series

Model ReleasesDGX agent

arXiv:2607.27263v1 Announce Type: new Abstract: Most benchmarks for causal inference over time series are observational, small, or domain-specific, leaving interventional and counterfactual estimation

DS@GT ARC at ImageCLEFmedical 2026: Architectural Diversity for Concept Detection and Foundation-Model Scaling for Caption Prediction in Medical Image Analysis

Model ReleasesDGX agent

arXiv:2607.27763v1 Announce Type: new Abstract: We describe the DS@GT submissions to the ImageCLEFmedical Caption 2026 challenge, which continues a long-running benchmark on the ROCOv2 dataset with tw

ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders

Model ReleasesDGX agent

arXiv:2607.27404v1 Announce Type: new Abstract: Existing benchmarks for electrocardiogram foundation models primarily evaluate downstream predictive performance, providing limited insight into whether

Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale

Model ReleasesDGX agent

arXiv:2607.28074v1 Announce Type: cross Abstract: Computer-use agents learn from what their actions change, so training one needs applications it can act on, break and reset. The applications that mat

EEG-EditBench: Probing Visual Information in EEG-Image Retrieval Models with Controlled Image Edits

Model ReleasesDGX agent

arXiv:2607.27857v1 Announce Type: new Abstract: Recent EEG-to-image retrieval models have achieved strong performance in identifying viewed images from semantically diverse candidates. Yet such succes

Efficient LLMs with AMP: Attention Heads and MLP Pruning

Model ReleasesDGX agent

arXiv:2504.21174v2 Announce Type: replace Abstract: Deep learning drives a new wave in computing systems and triggers the automation of increasingly complex problems. In particular, Large Language Mod

EgoGVAE: Ego-body Mesh Reconstruction via Guided Variational Autoencoder

Model ReleasesDGX agent

arXiv:2607.27755v1 Announce Type: new Abstract: We address the problem of recovering the full-body mesh from only the head pose. This task has become essential for various applications based on head-m

EHGCN: Hierarchical Euclidean-Hyperbolic Fusion via Motion-Aware GCN for Hybrid Event Stream Perception

Model ReleasesDGX agent

arXiv:2504.16616v4 Announce Type: replace Abstract: Event cameras, characterized by microsecond temporal resolution and very High Dynamic Range (HDR), emit high-speed event streams for perception task

EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents

Model ReleasesDGX agent

arXiv:2607.28229v1 Announce Type: new Abstract: The web is increasingly accessed by AI agents rather than humans. Every agent needs knowledge, especially in the life-sciences, where agentic pipelines

Epistemic diversity across language models mitigates knowledge collapse

Model ReleasesDGX agent

arXiv:2512.15011v3 Announce Type: replace Abstract: Artificial intelligence (AI) increasingly generates the very content used to train future AI systems. This feedback loop can degrade model quality,

eta-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

Model ReleasesDGX agent

arXiv:2607.28582v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains brittle in practice: making it work reli

Evidence-Ledger Adjudication for Claim-Evidence Traceability

Model ReleasesDGX agent

arXiv:2607.26512v1 Announce Type: new Abstract: AI agents can draft claims faster than authors can check whether the cited or retrieved evidence supports them. We study evidence-ledger adjudication: a

EvoCause: LLM-Guided Evolution of Causal Graphs for Root Cause Analysis

Model ReleasesDGX agent

arXiv:2607.27290v1 Announce Type: new Abstract: Modern telecommunication, cloud, and microservice systems emit correlated alarm cascades when components fail. Root cause analysis (RCA) aims to identif

Exact Action Values Are Not Enough: Rollout-Verified Reinforcement Fine-Tuning of a Reasoning Model for Multi-Zone VAV Control

Model ReleasesDGX agent

arXiv:2607.27914v1 Announce Type: new Abstract: Multi-zone variable-air-volume control must balance thermal comfort, indoor air quality, and electricity use across several continuous actuators. Model

Expected Survival-Time Bounds for Robust Optimization Over Time under Isotropic Gaussian Dynamics

Model ReleasesDGX agent

arXiv:2607.27280v1 Announce Type: cross Abstract: Robust Optimization Over Time (ROOT) is a recent branch of evolutionary dynamic optimization that seeks solutions capable of remaining effective acros

Experience sharing: How do you use your local models and for what kind of tasks?

Model ReleasesDGX agent

Here is my experience, which I would like to share with you and I also would like to hear your thoughts and valuable tips&tricks. Hardware: Mac Mini M4 (32GB Unified Memory) Model Server: Ollama Orche

Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation

Model ReleasesDGX agent

arXiv:2607.27372v1 Announce Type: cross Abstract: The deep learning revolution, kicked off by AlexNet, taught us that end-to-end training beats decomposing a problem into hand-designed stages. Generat

Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?

Model ReleasesDGX agent

arXiv:2607.26367v1 Announce Type: new Abstract: An important skill in theoretical physics is to recognize when a new problem can be transformed into a known model. We study this skill as an AI-agent t

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

Model ReleasesDGX agent

arXiv:2607.09306v3 Announce Type: replace Abstract: Behavioural auditing asks whether a language model behaves as it claims, but detection scores are reported without separating two targets: whether a

Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations

Model ReleasesDGX agent

arXiv:2607.28319v1 Announce Type: new Abstract: This work presents Fairness Pruning, a lightweight structural intervention method designed for the management and future mitigation of demographic bias

Fantastic Adaptive Taxonomies and How to Use Them

Model ReleasesDGX agent

arXiv:2607.16387v2 Announce Type: replace-cross Abstract: An agent system's execution traces record how it fails, and procedures that improve such a system without changing model weights (trajectory s

Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants

Model ReleasesDGX agent

arXiv:2607.26611v1 Announce Type: new Abstract: AI-assisted coding increasingly translates informal user intent into executable software, yet coding requests often contain ambiguities that recur in us

FinanceHarness: Autonomous Financial Deep Research Framework

Model ReleasesDGX agent

arXiv:2607.27853v1 Announce Type: new Abstract: Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research

Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents

Model ReleasesDGX agent

arXiv:2607.27275v1 Announce Type: new Abstract: Post-training quantization to 4-bit weights is widely reported to be nearly lossless. We test this claim for multi-turn, tool-calling agents, where it n

FlexiGrad: Adaptive Gradient Modulation for Hierarchical Fine-Grained Classification

Model ReleasesDGX agent

arXiv:2607.17563v2 Announce Type: replace Abstract: Many fine-grained recognition tasks contain hierarchical labels such as order, family and species. Although this supervision should be beneficial, j

FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing

Model ReleasesDGX agent

arXiv:2508.02092v3 Announce Type: replace-cross Abstract: Large language models represent significant investments in computation, data, and engineering expertise, making them extraordinarily valuable

From Expert Reduction to Behavioral Divergence: Tracing Numerical State through Sparse MoE Inference

Model ReleasesDGX agent

arXiv:2607.28097v1 Announce Type: new Abstract: Mathematically equivalent expert-reduction orders can produce observably different sparse-MoE executions. We isolate this effect in native DeepSeek-V4-F

From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models

Model ReleasesDGX agent

arXiv:2607.27654v1 Announce Type: new Abstract: Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of do

Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering

Model ReleasesDGX agent

arXiv:2607.28568v1 Announce Type: new Abstract: Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a

Generalization and Trade-off in Adversarial Training: An RKHS Perspective via Kernel Integral Operators

Model ReleasesDGX agent

arXiv:2607.27995v1 Announce Type: cross Abstract: Adversarial training has emerged as a powerful approach for protecting models against adversarial attacks in a broad range of real-world applications.

Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search

Model ReleasesDGX agent

arXiv:2607.27422v1 Announce Type: new Abstract: Good action rankings do not make a contrastive critic safe to maximize. These critics increasingly act as value-like objectives for best-of-K selection,

Google starts rolling out access to Gemini Spark for Google AI Pro subscribers to over 160 countries and adds a Chrome auto browse integration on desktop (Abner Li/9to5Google)

Model ReleasesDGX agent

Abner Li / 9to5Google: Google starts rolling out access to Gemini Spark for Google AI Pro subscribers to over 160 countries and adds a Chrome auto browse integration on desktop — Gemini Spark is getti

Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning

Model ReleasesDGX agent

arXiv:2607.27766v1 Announce Type: new Abstract: On-device in-context learning (ICL) relies on pre-inference retrieval to select demonstrations for useful context before downstream model inference. Thi

Graph Neural Multilevel Preconditioners for Iterative Solvers

Model ReleasesDGX agent

arXiv:2607.28456v1 Announce Type: cross Abstract: Solving large, sparse linear systems is a core task in scientific computing, and efficient iterative solvers rely critically on effective and robust p

Graph Neural Network Force Fields for Spin Dynamics in Metallic Magnets

Model ReleasesDGX agent

arXiv:2607.28537v1 Announce Type: cross Abstract: Metallic magnets exhibit complex spin dynamics governed by electronically generated interactions. Predictive simulations of such dynamics typically re

GuideSkill: Evolving Executable LLM Agent Skills for Guideline-Grounded Clinical Reasoning

Model ReleasesDGX agent

arXiv:2607.26160v1 Announce Type: new Abstract: Clinical practice guidelines (CPGs) encode diagnostic criteria, but LLM systems typically retrieve guideline text or absorb it through training rather t

GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference

Model ReleasesDGX agent

arXiv:2607.27694v1 Announce Type: cross Abstract: Low-bit quantization is essential for efficient LLM inference, and both rotation and fine-grained group quantization have shown individual promise. Ho

Hallucinations and Truth: A Comprehensive Accuracy Evaluation of RAG, LoRA and DoRA

Model ReleasesDGX agent

arXiv:2502.10497v2 Announce Type: replace Abstract: Recent advancements in Generative AI have significantly improved the efficiency and adaptability of natural language processing (NLP) systems, parti

HARGO: Heterogeneity-Aware Reward-Guided Optimization for RL Post-Training of LLMs on HPC Tasks

Model ReleasesDGX agent

arXiv:2607.28301v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) can equip large language models (LLMs) with domain knowledge for high-performance computing (HPC) tasks such as data race d

Has anyone actually benchmarked where the 'big-model orchestrator + local-model worker' split breaks down?

Model ReleasesDGX agent

I keep seeing the 'use a big model via API as the architect, run local small/mid models as workers' pattern recommended for people with modest local hardware. I've been running it myself (orchestrator

Hermes and Pi Agent led on the average cost per task, while Claude Code cost about 3.7x as much as Pi: - 0.39 Hermes Agent - 0.40 Pi Agent…

Model ReleasesDGX agent

Hermes and Pi Agent led on the average cost per task, while Claude Code cost about 3.7x as much as Pi: - 0.39 Hermes Agent - 0.40 Pi Agent - 0.47 Codex - 0.51 OpenCode - 0.54 Kimi Code - 1.47 Claude C

I have trained a model to predict my blood sugar [P]

Model ReleasesDGX agent

It's an encoder-only transformer that consumes past(blood glucose + carbs + insulin) and future(carbs + insulin) and predicts future blood glucose for the next 2 hours. Announced meals and boluses/bas

I predict DeepSeek V4 Flash 0731's Artificial Analysis score to be 57 ± 1 point (Kimi K3 Level)

Model ReleasesDGX agent

Deepseek's new model V4 Flash 0731 is much better, I (Claude lol) did a bit of linear regression with a leave one out style verification to predict its AA Score, and that puts it at Kimi K3 level, whi

i see your moore's law and i raise you 20x

Model ReleasesDGX agent

i see your moore's law and i raise you 20x GPT-5.4 full at xhigh scored 51, exactly where Luna max sits today. GPT-5.4 costs 2.50/15; Luna now costs 0.20/1.20. In other words, roughly four months late

I switched my https://agent.datasette.io instance to Luna (it was previously on Gemini 3.1 Flash-Lite - Luna is cheaper now) - you can sign …

Model ReleasesDGX agent

Simon Willison switched his Datasette Agent instance from Gemini 3.1 Flash‑Lite to GPT‑5.6 “Luna” after a recent 80% price drop. He reports the new model is significantly faster and automatically gene

IDP AutoOpt: Agent-Driven Optimization of Document Processing Pipeline Configurations

Model ReleasesDGX agent

arXiv:2607.26075v1 Announce Type: cross Abstract: We present IDP AutoOpt, an autonomous LLM agent that discovers high-performing configurations for intelligent document processing (IDP) pipelines. Tun

If LoRA is underperforming, don't reach for more expensive full parameter fine-tuning right away. We ran three cheap tests (data coverage, o…

Model ReleasesDGX agent

If LoRA is underperforming, don't reach for more expensive full parameter fine-tuning right away. We ran three cheap tests (data coverage, optimization, rank) to see if we could close the gap between

IFCMemoryBench: Evaluating Long-Term Memory of LLM-Based Agents in BIM Information Retrieval

Model ReleasesDGX agent

arXiv:2607.26072v1 Announce Type: cross Abstract: Long-term memory is becoming a core capability of LLM-based agents, but existing evaluations largely test conversational recall in open-domain or pers

← Previous
1…4546474849…373
Next →