AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,332 results
Model Releases

EvoCause: LLM-Guided Evolution of Causal Graphs for Root Cause Analysis

DGX agent

arXiv:2607.27290v1 Announce Type: new Abstract: Modern telecommunication, cloud, and microservice systems emit correlated alarm cascades when components fail. Root cause analysis (RCA) aims to identif

model-releasesarxiv-cs-lg
31 Jul 2026
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Exact Action Values Are Not Enough: Rollout-Verified Reinforcement Fine-Tuning of a Reasoning Model for Multi-Zone VAV Control

DGX agent

arXiv:2607.27914v1 Announce Type: new Abstract: Multi-zone variable-air-volume control must balance thermal comfort, indoor air quality, and electricity use across several continuous actuators. Model

model-releasesarxiv-cs-lg
31 Jul 2026
Model Releases

Expected Survival-Time Bounds for Robust Optimization Over Time under Isotropic Gaussian Dynamics

DGX agent

arXiv:2607.27280v1 Announce Type: cross Abstract: Robust Optimization Over Time (ROOT) is a recent branch of evolutionary dynamic optimization that seeks solutions capable of remaining effective acros

model-releasesarxiv-cs-lg
31 Jul 2026
Model Releases

Experience sharing: How do you use your local models and for what kind of tasks?

DGX agent

Here is my experience, which I would like to share with you and I also would like to hear your thoughts and valuable tips&tricks. Hardware: Mac Mini M4 (32GB Unified Memory) Model Server: Ollama Orche

model-releasesr-ollama
31 Jul 2026
Model Releases

Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation

DGX agent

arXiv:2607.27372v1 Announce Type: cross Abstract: The deep learning revolution, kicked off by AlexNet, taught us that end-to-end training beats decomposing a problem into hand-designed stages. Generat

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?

DGX agent

arXiv:2607.26367v1 Announce Type: new Abstract: An important skill in theoretical physics is to recognize when a new problem can be transformed into a known model. We study this skill as an AI-agent t

model-releasesarxiv-cs-ai
31 Jul 2026
Model Releases

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

DGX agent

arXiv:2607.09306v3 Announce Type: replace Abstract: Behavioural auditing asks whether a language model behaves as it claims, but detection scores are reported without separating two targets: whether a

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations

DGX agent

arXiv:2607.28319v1 Announce Type: new Abstract: This work presents Fairness Pruning, a lightweight structural intervention method designed for the management and future mitigation of demographic bias

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

Fantastic Adaptive Taxonomies and How to Use Them

DGX agent

arXiv:2607.16387v2 Announce Type: replace-cross Abstract: An agent system's execution traces record how it fails, and procedures that improve such a system without changing model weights (trajectory s

model-releasesarxiv-cs-ai
31 Jul 2026
Model Releases

Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants

DGX agent

arXiv:2607.26611v1 Announce Type: new Abstract: AI-assisted coding increasingly translates informal user intent into executable software, yet coding requests often contain ambiguities that recur in us

model-releasesarxiv-cs-ai
31 Jul 2026
Model Releases

FinanceHarness: Autonomous Financial Deep Research Framework

DGX agent

arXiv:2607.27853v1 Announce Type: new Abstract: Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents

DGX agent

arXiv:2607.27275v1 Announce Type: new Abstract: Post-training quantization to 4-bit weights is widely reported to be nearly lossless. We test this claim for multi-turn, tool-calling agents, where it n

model-releasesarxiv-cs-lg
31 Jul 2026
Model Releases

FlexiGrad: Adaptive Gradient Modulation for Hierarchical Fine-Grained Classification

DGX agent

arXiv:2607.17563v2 Announce Type: replace Abstract: Many fine-grained recognition tasks contain hierarchical labels such as order, family and species. Although this supervision should be beneficial, j

model-releasesarxiv-cs-cv
31 Jul 2026
Model Releases

FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing

DGX agent

arXiv:2508.02092v3 Announce Type: replace-cross Abstract: Large language models represent significant investments in computation, data, and engineering expertise, making them extraordinarily valuable

model-releasesarxiv-cs-ai
31 Jul 2026
Model Releases

From Expert Reduction to Behavioral Divergence: Tracing Numerical State through Sparse MoE Inference

DGX agent

arXiv:2607.28097v1 Announce Type: new Abstract: Mathematically equivalent expert-reduction orders can produce observably different sparse-MoE executions. We isolate this effect in native DeepSeek-V4-F

model-releasesarxiv-cs-lg
31 Jul 2026
Model Releases

From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models

DGX agent

arXiv:2607.27654v1 Announce Type: new Abstract: Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of do

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering

DGX agent

arXiv:2607.28568v1 Announce Type: new Abstract: Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

Generalization and Trade-off in Adversarial Training: An RKHS Perspective via Kernel Integral Operators

DGX agent

arXiv:2607.27995v1 Announce Type: cross Abstract: Adversarial training has emerged as a powerful approach for protecting models against adversarial attacks in a broad range of real-world applications.

model-releasesarxiv-cs-lg
31 Jul 2026
Model Releases

Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search

DGX agent

arXiv:2607.27422v1 Announce Type: new Abstract: Good action rankings do not make a contrastive critic safe to maximize. These critics increasingly act as value-like objectives for best-of-K selection,

model-releasesarxiv-cs-lg
31 Jul 2026
Model Releases

Google starts rolling out access to Gemini Spark for Google AI Pro subscribers to over 160 countries and adds a Chrome auto browse integration on desktop (Abner Li/9to5Google)

DGX agent

Abner Li / 9to5Google: Google starts rolling out access to Gemini Spark for Google AI Pro subscribers to over 160 countries and adds a Chrome auto browse integration on desktop — Gemini Spark is getti

model-releasestechmeme
31 Jul 2026
Model Releases

Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning

DGX agent

arXiv:2607.27766v1 Announce Type: new Abstract: On-device in-context learning (ICL) relies on pre-inference retrieval to select demonstrations for useful context before downstream model inference. Thi

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

Graph Neural Multilevel Preconditioners for Iterative Solvers

DGX agent

arXiv:2607.28456v1 Announce Type: cross Abstract: Solving large, sparse linear systems is a core task in scientific computing, and efficient iterative solvers rely critically on effective and robust p

model-releasesarxiv-cs-lg
31 Jul 2026
Model Releases

Graph Neural Network Force Fields for Spin Dynamics in Metallic Magnets

DGX agent

arXiv:2607.28537v1 Announce Type: cross Abstract: Metallic magnets exhibit complex spin dynamics governed by electronically generated interactions. Predictive simulations of such dynamics typically re

model-releasesarxiv-cs-lg
31 Jul 2026
Model Releases

GuideSkill: Evolving Executable LLM Agent Skills for Guideline-Grounded Clinical Reasoning

DGX agent

arXiv:2607.26160v1 Announce Type: new Abstract: Clinical practice guidelines (CPGs) encode diagnostic criteria, but LLM systems typically retrieve guideline text or absorb it through training rather t

model-releasesarxiv-cs-ai
31 Jul 2026
Model Releases

GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference

DGX agent

arXiv:2607.27694v1 Announce Type: cross Abstract: Low-bit quantization is essential for efficient LLM inference, and both rotation and fine-grained group quantization have shown individual promise. Ho

model-releasesarxiv-cs-lg
31 Jul 2026
Model Releases

Hallucinations and Truth: A Comprehensive Accuracy Evaluation of RAG, LoRA and DoRA

DGX agent

arXiv:2502.10497v2 Announce Type: replace Abstract: Recent advancements in Generative AI have significantly improved the efficiency and adaptability of natural language processing (NLP) systems, parti

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

HARGO: Heterogeneity-Aware Reward-Guided Optimization for RL Post-Training of LLMs on HPC Tasks

DGX agent

arXiv:2607.28301v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) can equip large language models (LLMs) with domain knowledge for high-performance computing (HPC) tasks such as data race d

model-releasesarxiv-cs-lg
31 Jul 2026
Model Releases

Has anyone actually benchmarked where the 'big-model orchestrator + local-model worker' split breaks down?

DGX agent

I keep seeing the 'use a big model via API as the architect, run local small/mid models as workers' pattern recommended for people with modest local hardware. I've been running it myself (orchestrator

model-releasesr-localllama
31 Jul 2026
Model Releases

Hermes and Pi Agent led on the average cost per task, while Claude Code cost about 3.7x as much as Pi: - 0.39 Hermes Agent - 0.40 Pi Agent…

DGX agent

Hermes and Pi Agent led on the average cost per task, while Claude Code cost about 3.7x as much as Pi: - 0.39 Hermes Agent - 0.40 Pi Agent - 0.47 Codex - 0.51 OpenCode - 0.54 Kimi Code - 1.47 Claude C

model-releasesnous-research--x
31 Jul 2026
Model Releases

I have trained a model to predict my blood sugar [P]

DGX agent

It's an encoder-only transformer that consumes past(blood glucose + carbs + insulin) and future(carbs + insulin) and predicts future blood glucose for the next 2 hours. Announced meals and boluses/bas

model-releasesr-machinelearning
31 Jul 2026
Model Releases

I predict DeepSeek V4 Flash 0731's Artificial Analysis score to be 57 ± 1 point (Kimi K3 Level)

DGX agent

Deepseek's new model V4 Flash 0731 is much better, I (Claude lol) did a bit of linear regression with a leave one out style verification to predict its AA Score, and that puts it at Kimi K3 level, whi

model-releasesr-localllama
31 Jul 2026
Model Releases

i see your moore's law and i raise you 20x

DGX agent

i see your moore's law and i raise you 20x GPT-5.4 full at xhigh scored 51, exactly where Luna max sits today. GPT-5.4 costs 2.50/15; Luna now costs 0.20/1.20. In other words, roughly four months late

model-releasessam-altman--x
31 Jul 2026
Model Releases

I switched my https://agent.datasette.io instance to Luna (it was previously on Gemini 3.1 Flash-Lite - Luna is cheaper now) - you can sign …

DGX agent

Simon Willison switched his Datasette Agent instance from Gemini 3.1 Flash‑Lite to GPT‑5.6 “Luna” after a recent 80% price drop. He reports the new model is significantly faster and automatically gene

model-releasessimon-willison--x
31 Jul 2026
Model Releases

IDP AutoOpt: Agent-Driven Optimization of Document Processing Pipeline Configurations

DGX agent

arXiv:2607.26075v1 Announce Type: cross Abstract: We present IDP AutoOpt, an autonomous LLM agent that discovers high-performing configurations for intelligent document processing (IDP) pipelines. Tun

model-releasesarxiv-cs-ai
31 Jul 2026
Model Releases

If LoRA is underperforming, don't reach for more expensive full parameter fine-tuning right away. We ran three cheap tests (data coverage, o…

DGX agent

If LoRA is underperforming, don't reach for more expensive full parameter fine-tuning right away. We ran three cheap tests (data coverage, optimization, rank) to see if we could close the gap between

model-releasesfireworks-ai--x
31 Jul 2026
Model Releases

IFCMemoryBench: Evaluating Long-Term Memory of LLM-Based Agents in BIM Information Retrieval

DGX agent

arXiv:2607.26072v1 Announce Type: cross Abstract: Long-term memory is becoming a core capability of LLM-based agents, but existing evaluations largely test conversational recall in open-domain or pers

model-releasesarxiv-cs-ai
31 Jul 2026
Model Releases

IFHierBench: Hierarchical Instruction Following for Large Language Models

DGX agent

arXiv:2607.27912v1 Announce Type: cross Abstract: Instruction-following ability is critical for deploying large language models in real-world applications, where downstream components depend on the ou

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

'Intelligence too cheap to meter' battle is on! Given that DeepSeek-V4-Flash-Preview is already great for agentic tasks, there is no doubt t…

DGX agent

'Intelligence too cheap to meter' battle is on! Given that DeepSeek-V4-Flash-Preview is already great for agentic tasks, there is no doubt this new checkpoint must be an absolute beast. 20+ point jump

model-releasesdair-ai--x
31 Jul 2026
Model Releases

Introducing Qwen-Audio-3.0-ASR-Flash: More context-aware. Stronger domain-term recognition. 🚀Our latest ASR model upgrades: • Context consi…

DGX agent

Introducing Qwen-Audio-3.0-ASR-Flash: More context-aware. Stronger domain-term recognition. 🚀Our latest ASR model upgrades: • Context consistency • Domain-term recognition • Custom hotwords • Speech p

model-releasesqwen--x
31 Jul 2026
Model Releases

Is it just me, or are current LLM benchmarks failing to capture actual usability? (Gemma 4 vs. Gemini/Claude Opus)

DGX agent

Disclaimer, this was kinda written with AI (Gemma 4 again) but it also did really well here, it outputted what I wanted, when I asked it to refine stuff or improve on certain areas it did that without

model-releasesr-localllama
31 Jul 2026
Model Releases

It’s been a busy couple of weeks! ICYMI, here’s the recap ⬇️ — Gemini Robotics 2 from @GoogleDeepmind brings whole-body intelligence to robo…

DGX agent

It’s been a busy couple of weeks! ICYMI, here’s the recap ⬇️ — Gemini Robotics 2 from @GoogleDeepmind brings whole-body intelligence to robots — Gemini 3.5 Flash-Lite is our fastest, most cost-effecti

model-releasesgoogle-ai--x
31 Jul 2026
Model Releases

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

DGX agent

arXiv:2607.27670v1 Announce Type: new Abstract: Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that creat

model-releasesarxiv-cs-cv
31 Jul 2026
Model Releases

K-EXAONE 2.0 released

DGX agent

https://huggingface.co/LGAI-EXAONE/K-EXAONE-2.0-750B-A37B https://huggingface.co/LGAI-EXAONE/K-EXAONE-2.0-750B-A37B-FP8 https://huggingface.co/LGAI-EXAONE/K-EXAONE-2.0-750B-A37B-NVFP4 https://huggingf

model-releasesr-localllama
31 Jul 2026
Model Releases

KAISEN: Reproducible Subgroup Fairness Auditing for Clinical Risk Models

DGX agent

arXiv:2607.28608v1 Announce Type: new Abstract: Clinical risk models routinely achieve strong aggregate performance while producing materially different error rates across patient subgroups. Audit pip

model-releasesarxiv-cs-lg
31 Jul 2026
Model Releases

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning

DGX agent

arXiv:2607.27610v1 Announce Type: new Abstract: Reinforcement learning (RL) finetuning significantly enhances the reasoning capabilities of large language models (LLMs), yet its effectiveness critical

model-releasesarxiv-cs-lg
31 Jul 2026
Model Releases

KernelGenBench: A Multi-Source and Multi-Chip Benchmark for LLM-based Kernel Generation

DGX agent

arXiv:2607.27231v1 Announce Type: cross Abstract: Large language models (LLMs) have significantly increased the demand for efficient accelerator kernels, but kernel development remains a highly specia

model-releasesarxiv-cs-lg
31 Jul 2026
Model Releases

Language Diversity: Evaluating Language Usage and AI Performance on African Languages in Digital Spaces

DGX agent

arXiv:2512.01557v3 Announce Type: replace Abstract: This study examines the digital representation of African languages and the challenges this presents for current language detection tools. We evalua

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation

DGX agent

arXiv:2607.27353v1 Announce Type: new Abstract: Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or

model-releasesarxiv-cs-cl
31 Jul 2026
← Previous
1…5758596061…466
Next →