AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,340 results
Model Releases

When Retrieval Helps and Distracts: Evaluating Evidence-Generating LLMs for Biomedical Claim Verification

DGX agent

arXiv:2608.01409v1 Announce Type: new Abstract: Biomedical fact-checking systems must do more than predict whether a claim is supported, contradicted, or unaddressed: they should also produce evidence

model-releasesarxiv-cs-cl
4 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Who Belongs in the Eval Set? A Capability-Taxonomy-Driven Pipeline for Curating Regression Eval Sets in Agent-Extensibility Platforms

DGX agent

arXiv:2608.01004v1 Announce Type: new Abstract: Platform teams hosting agent-extensibility surfaces face a regression-economics paradox: every onboarding customer ships an evaluation set tuned to thei

model-releasesarxiv-cs-lg
4 Aug 2026
Model Releases

Why are Chinese models better* at Frontend than the western top labs?

DGX agent

I use A LOT both openAI and Anthropic products. When I need some frontend work (pure web dev) (or answer that feel less verbose and more to the point) I use Anthropic. For multimodality openAI feels b

model-releasesr-localllama
4 Aug 2026
Model Releases

Why are Gamers so incredibly hostile to AI? Is it just a tiny vocal minority that spreads such toxic vitriol online?

DGX agent

It's more accurate to say that many highly engaged online gamers are hostile to AI, not that 'gamers' as a whole are. Gaming is a huge community with hundreds of millions of people, and opinions vary

model-releasesr-chatgpt
4 Aug 2026
Model Releases

Why Formal Monitors Fail: Attack Distribution Entropy as a Coverage Bound for LTL-Based LLM Agent Safety

DGX agent

arXiv:2608.01388v1 Announce Type: cross Abstract: Runtime safety monitors based on Linear Temporal Logic (LTL) and finite automata (FSA) are increasingly deployed to intercept unsafe tool-call sequenc

model-releasesarxiv-cs-lg
4 Aug 2026
Model Releases

Why Large Language Models Fail at Tabular Prediction

DGX agent

arXiv:2608.02412v1 Announce Type: new Abstract: Large language models (LLMs) have become the default tool for a remarkable range of tasks, yet they have had conspicuously little success at one of the

model-releasesarxiv-cs-lg
4 Aug 2026
Model Releases

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

DGX agent

arXiv:2608.02603v1 Announce Type: new Abstract: Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the appa

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

Writing-System-Level Tokenizer Adaptation for Byte-Level BPE

DGX agent

arXiv:2608.00582v1 Announce Type: new Abstract: Pretrained byte-level BPE tokenizers can segment underrepresented languages inefficiently. Replacing a tokenizer changes the meaning of nearly every tok

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding

DGX agent

arXiv:2608.00036v1 Announce Type: new Abstract: Real-world document tasks often ask professionals to answer questions from annual reports, regulations, clinical guidelines, and technical manuals that

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

XSPA: Crafting Imperceptible X-Shaped Sparse Adversarial Perturbations for Transferable Attacks on VLMs

DGX agent

arXiv:2603.28568v2 Announce Type: replace Abstract: Vision-language models (VLMs) share visual-textual representations across zero-shot classification, image captioning, and visual question answering

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

Z-PEFT: Zero-shot Backdoor Detection in Parameter-Efficient Fine-Tuning via Canonical Spectral Signatures

DGX agent

arXiv:2608.02271v1 Announce Type: new Abstract: Parameter-Efficient Fine-tuned (PEFT) models are frequently downloaded from open repositories by practitioners. This widespread practice creates a signi

model-releasesarxiv-cs-lg
4 Aug 2026
Model Releases

A Benchmark for Strategic Auditee Gaming Under Continuous Compliance Monitoring

DGX agent

arXiv:2605.06340v2 Announce Type: replace-cross Abstract: Continuous post-deployment compliance audits, mandated by emerging regulations such as the EU AI Act and Digital Services Act, create a class

model-releasesarxiv-cs-lg
3 Aug 2026
Model Releases

A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples

DGX agent

arXiv:2607.29122v1 Announce Type: new Abstract: Pixel-space diffusion models aim to learn an end-to-end generator directly over raw pixels. This is challenging because a single model must capture both

model-releasesarxiv-cs-cv
3 Aug 2026
Model Releases

A robust association between LLM use and scientific productivity: Assessing stopping-time selection

DGX agent

arXiv:2607.28968v1 Announce Type: cross Abstract: Renault, Bergeaud, and Bosquet (hereafter RBB) argue that dating LLM adoption as the first month in which an author's abstract is flagged induces a st

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

A Unified Benchmark of Deep Learning Models for Multi-task 3D Brain Tumor Segmentation from Magnetic Resonance Imaging

DGX agent

arXiv:2607.28858v1 Announce Type: cross Abstract: Automatic brain tumor segmentation from magnetic resonance imaging (MRI) has become a fundamental task in computer-assisted diagnosis, treatment plann

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

ActionParty: Multi-Subject Action Binding in Generative Video Games

DGX agent

arXiv:2604.02330v2 Announce Type: replace-cross Abstract: Recent advances in video diffusion have enabled the development of 'world models' capable of simulating interactive environments. However, the

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

Adaptive Policy Backbone via Shared Network

DGX agent

arXiv:2509.22310v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has achieved impressive results across domains, yet learning an optimal policy typically requires extensive intera

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

AFAIK the most significant breakthrough since 2017 besides scaling old ideas was broadening from base models into larger systems that incorp…

DGX agent

AFAIK the most significant breakthrough since 2017 besides scaling old ideas was broadening from base models into larger systems that incorporate symbol/manipulating entities like harnesses, tools, an

model-releasesgary-marcus--x
3 Aug 2026
Model Releases

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers

DGX agent

arXiv:2607.29626v1 Announce Type: new Abstract: As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly i

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

Agentic Harness for Real-World Compilers

DGX agent

arXiv:2603.20075v2 Announce Type: replace-cross Abstract: Compilers are critical to modern computing, yet fixing compiler bugs is difficult. While recent large language model (LLM) advancements enable

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

AI9Stars released G9v3-39A5B

DGX agent

AI9Stars has released G9v3-39A5B an open weights language model designed to deliver even stronger reasoning capabilities than ai9stars/G9v3-3B with its 39B and 5 active experts. It is released under t

model-releasesr-localllama
3 Aug 2026
Model Releases

Alibaba debuts Qwen3.8-Max model with 2.4T parameters

DGX agent

Alibaba Group Holding Ltd. today debuted a new addition to its Qwen series of open-source large language models. Qwen3.8-Max is the Chinese e-commerce giant’s most capable LLM to date. It features 2.4

model-releasessiliconangle
3 Aug 2026
Model Releases

All frontier AI models (including GPT-5.6 Sol, Fable 5, and Kimi K3) consistently fail to implement solvers for basic nonlinear PDEs correct…

DGX agent

All frontier AI models (including GPT-5.6 Sol, Fable 5, and Kimi K3) consistently fail to implement solvers for basic nonlinear PDEs correctly, often introducing significant errors in numerical stabil

model-releasesgary-marcus--x
3 Aug 2026
Model Releases

AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction

DGX agent

arXiv:2607.29549v1 Announce Type: new Abstract: Large language models have demonstrated strong mathematical problem-solving capabilities, yet reliably verifying their candidate answers remains challen

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

An internal version of our next major model produced 10 new results on long-standing open problems in mathematics and theoretical computer s…

DGX agent

An internal version of our next major model produced 10 new results on long-standing open problems in mathematics and theoretical computer science, using roughly $2,000 worth of tokens at GPT-5.6 Sol

model-releasesopenai--x
3 Aug 2026
Model Releases

Analytical and Bootstrap Confidence Intervals of Double Machine Learning: Simulation studies and an application to rural-urban difference in obesity prevalence

DGX agent

arXiv:2607.29456v1 Announce Type: cross Abstract: Double Machine Learning (DML) is a popular approach for treatment effect estimation in various settings, which allows a wide range of flexible machine

model-releasesarxiv-cs-lg
3 Aug 2026
Model Releases

Appreciate it! Now let's understand the world through the eyes of Qwen3.8. 🥳

DGX agent

Qwen3.8-Max from Alibaba’s Qwen team achieved second place in the Vision Arena benchmark, scoring 1,305 points. It trails only Claude Fable 5 (High), which leads by a slim 13‑point margin. The post un

model-releasesqwen--x
3 Aug 2026
Model Releases

ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation

DGX agent

arXiv:2607.29539v1 Announce Type: cross Abstract: Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs). While prior work has

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements

DGX agent

arXiv:2607.28661v1 Announce Type: new Abstract: Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matching? The financial domain, demanding

model-releasesarxiv-cs-cl
3 Aug 2026
Model Releases

Artificial Analysis: DeepSeek's V4-Flash costs 0.14/1M input and 0.28/1M output tokens, or 0.03 per test, far below Kimi K3's 0.86 and GPT-5.6 Sol's $1.86 (Eduardo Baptista/Reuters)

DGX agent

Eduardo Baptista / Reuters: Artificial Analysis: DeepSeek's V4-Flash costs 0.14/1M input and 0.28/1M output tokens, or 0.03 per test, far below Kimi K3's 0.86 and GPT-5.6 Sol's $1.86 — A version of Ch

model-releasestechmeme
3 Aug 2026
Model Releases

🚨ASI/AGI is imminent fans don’t realize that they ALREADY lost the argument around Astra. My argument (spelled out in detail in my Substack…

DGX agent

🚨ASI/AGI is imminent fans don’t realize that they ALREADY lost the argument around Astra. My argument (spelled out in detail in my Substack today) was that Astra was unlikely to be the dramatic leap f

model-releasesgary-marcus--x
3 Aug 2026
Model Releases

Ask anything, anonymously. Qwen3.8-Max has landed on Venice. Give it a try!

DGX agent

Qwen from Alibaba has released the Qwen 3.8‑Max model on the Venice platform, enabling users to ask questions anonymously. The announcement encourages users to try the new functionality immediately. T

model-releasesqwen--x
3 Aug 2026
Model Releases

Assessing the Generalization of Graph Neural Networks for Fault Location Across Increasing Distributed Energy Resource Penetration Levels

DGX agent

arXiv:2607.29293v1 Announce Type: new Abstract: Accurate fault location is critical for distribution network reliability. However, increasing distributed energy resource (DER) penetration complicates

model-releasesarxiv-cs-lg
3 Aug 2026
Model Releases

Autonomous Repair for Multi-Agent Systems via Monte-Carlo Tree Search

DGX agent

arXiv:2607.29055v1 Announce Type: cross Abstract: Multi-agent systems (MAS) are increasingly deployed to solve complex tasks. In case of incorrect or unsatisfactory outputs, users have to manually loc

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

b10236

DGX agent

metal: implement DSv4 Lightning Indexer (#25893) metal: implement F16 Lightning Indexer Implement GGML_OP_LIGHTNING_INDEXER for 128-dimensional, 64-head inputs with F32 queries and weights plus F16 ke

model-releasesllama-cpp-releases
3 Aug 2026
Model Releases

b10237

DGX agent

llama : MTP support for DeepSeek V3.2 (#26457) llama : MTP support for DeepSeek V3.2 model : no need to include MTP layers during DeepSeek V3.2 model type discovery Co-authored-by: Stanisław Szymczyk

model-releasesllama-cpp-releases
3 Aug 2026
Model Releases

b10238

DGX agent

model: MTP support for Qwen3-Next (#25589) mtp for qwen3nex fix for python type-check Fix to compute num_mtp from directly mtp layer define opt_num_mtp_layers in _QwenMtpMixin and fix some comments Fi

model-releasesllama-cpp-releases
3 Aug 2026
Model Releases

b10240

DGX agent

server: add notice for upcoming default port change 8080 --> 9931 (#26508) server: add notice for upcoming default port change 8080 --> 9931 add link to PR correct to 9931 Website: https://llama.app m

model-releasesllama-cpp-releases
3 Aug 2026
Model Releases

b10241

DGX agent

CUDA: Fix data-races when reusing SMEM in block_reduce (#26385) CUDA: Fix data-races when reusing block_reduce block_reduce currently doesn't resync after reading from SMEM, causing potential data-rac

model-releasesllama-cpp-releases
3 Aug 2026
Model Releases

b10242

DGX agent

CUDA: Add backend sampler for penalties sampler (#25262) sampling: enhance penalty handling in common_sampler_init Set default value for penalty_last_n based on model context if not specified. Ensure

model-releasesllama-cpp-releases
3 Aug 2026
Model Releases

b10243

DGX agent

llama : allocate indexer cache only in 'full' indexer layers (#26474) Co-authored-by: Stanisław Szymczyk sszymczy@gmail.com Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Appl

model-releasesllama-cpp-releases
3 Aug 2026
Model Releases

b10244

DGX agent

model: M3: Move MSA into a new memory implementation (#26338) Move MSA logic from llama-kv-cache into llama-kv-cache-msa cont : minor cont : ws fix Co-authored-by: Georgi Gerganov ggerganov@gmail.com

model-releasesllama-cpp-releases
3 Aug 2026
Model Releases

b10245

DGX agent

graph : fix unused input tensors in minimax m3 graph (#26519) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64)

model-releasesllama-cpp-releases
3 Aug 2026
Model Releases

b10246

DGX agent

opencl: route large q6_K lm_head to the flat GEMV (#26427) add a direct size condition for large weights; the original dimension condition is insufficient -- q6_K lm_head for gemma-4 E2B has [1536, 26

model-releasesllama-cpp-releases
3 Aug 2026
Model Releases

Benchmarking Frontier Large Language Models Against Official Crash Database Coding Using Police Crash Narratives

DGX agent

arXiv:2607.29064v1 Announce Type: cross Abstract: Police crash narratives contain information that may supplement structured crash databases, but manual review is labor-intensive and it remains unclea

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation

DGX agent

arXiv:2607.28801v1 Announce Type: cross Abstract: Benchmark datasets are central to evaluating Large Language Models (LLMs), yet they are typically conceived as monolithic tasks, obscuring substantial

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications

DGX agent

arXiv:2607.28840v1 Announce Type: new Abstract: Large language models are increasingly deployed in financial applications that combine retrieval, proprietary data, tool use, orchestration logic, monit

model-releasesarxiv-cs-cl
3 Aug 2026
Model Releases

Bootstrapping Self-Supervised Learning of Binary Classification Using Error Bounds: A Case Study on a Robotic Insertion Task

DGX agent

arXiv:2607.29640v1 Announce Type: new Abstract: Flexible manufacturing requires rapid deployment of solutions and minimal setup time to remain competitive. An essential attribute is the ability to con

model-releasesarxiv-cs-ro
3 Aug 2026
← Previous
1…5051525354…466
Next →