AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,323 results
Model Releases

Minimax-Optimal Semiparametric Contextual Dynamic Pricing with Multimodal Revenue

DGX agent

arXiv:2608.03142v1 Announce Type: cross Abstract: We study contextual dynamic pricing with arbitrary covariate sequences and bounded, possibly nonbinary purchase quantities. Demand follows a semiparam

model-releasesarxiv-cs-ai
5 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Mistral introduces Shieldstral to provide lightweight policy-aware moderation for AI models

DGX agent

French artificial intelligence startup Mistral AI SAS today introduced a lightweight multimodal safety artificial intelligence open-weight model that can classify outputs for AI models that outperform

model-releasessiliconangle
5 Aug 2026
Model Releases

MMLongBench-Doc-V2: A Corrected-Annotation, Semantics-Aware Revision of MMLongBench-Doc

DGX agent

arXiv:2608.03397v1 Announce Type: new Abstract: MMLongBench-Doc is a long-document QA benchmark of 1,082 questions over 135 PDFs. Two properties of it push measured scores away from the quantity they

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

Modeling Long-Term Memory and Temporal Attention Shifts for Video Salient Object Ranking with a New Benchmark

DGX agent

arXiv:2203.17257v2 Announce Type: replace Abstract: Salient Object Ranking (SOR) aims to estimate the relative saliency order among multiple salient objects. While SOR has been extensively studied in

model-releasesarxiv-cs-cv
5 Aug 2026
Model Releases

MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp

DGX agent

TensorSharp's MoE CPU-offload feature has been merged into main. Here is the parameters description of this feature: Mixture-of-Experts CPU offload: --n-cpu-moe <N> | -ncmoe <N> Keep the routed MoE ex

model-releasesr-localllama
5 Aug 2026
Model Releases

MoEGen: Mixture-of-Experts for Instance-Adaptive LoRA Generation

DGX agent

arXiv:2608.03275v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) enables efficient adaptation of large language models, but existing MoE-based PEFT methods typically improve capa

model-releasesarxiv-cs-cl
5 Aug 2026
Model Releases

Morphology-Aware Implicit Super-Resolution Network for Pathological Images

DGX agent

arXiv:2608.03664v1 Announce Type: new Abstract: Accurate diagnosis in Digital Pathology (DP) relies on high-resolution whole-slide images, yet clinical deployment is often limited by hardware costs. S

model-releasesarxiv-cs-cv
5 Aug 2026
Model Releases

MT-Web2Code: Benchmarking Coding Agents on Multi-Turn Regional Reconstruction and Localized Modification

DGX agent

arXiv:2608.03474v1 Announce Type: new Abstract: Recent advances in Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in web UI generation. However, existing benchmarks pre

model-releasesarxiv-cs-cv
5 Aug 2026
Model Releases

MultiCompose: Multi-Concept Personalized Composition with Per-Subject Attribute Binding

DGX agent

arXiv:2608.03708v1 Announce Type: new Abstract: Text-to-image diffusion models enable personalization of specific visual concepts from a small number of reference images. However, generating a single

model-releasesarxiv-cs-cv
5 Aug 2026
Model Releases

MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial Reasoning

DGX agent

arXiv:2608.03882v1 Announce Type: cross Abstract: Geospatial reasoning, i.e., computing distances, containment, and other spatial relations over real-world entities, is central to navigation and logis

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

MuRA: Multi-Rank Adaptation for Efficient and Effective Test-Time Vision-Language Generalization

DGX agent

arXiv:2608.03885v1 Announce Type: new Abstract: Vision-language models exhibit remarkable zero-shot capabilities but suffer significant performance degradation under distribution shifts. While test-ti

model-releasesarxiv-cs-cv
5 Aug 2026
Model Releases

NearID: Identity Representation Learning via Near-identity Distractors

DGX agent

arXiv:2604.01973v2 Announce Type: replace Abstract: When evaluating identity-focused tasks such as personalized generation and image editing, existing vision encoders entangle object identity with bac

model-releasesarxiv-cs-cv
5 Aug 2026
Model Releases

Neural Networks with Local Converging Inputs for Efficient Options Pricing Models

DGX agent

arXiv:2608.02778v1 Announce Type: new Abstract: We present a novel application of Neural Networks with Local Converging Inputs (NNLCI) to improve the efficiency of existing numerical methods for prici

model-releasesarxiv-cs-lg
5 Aug 2026
Model Releases

NOMADD: Numerical Optimization of Models Adapting to Data Drift

DGX agent

arXiv:2608.02845v1 Announce Type: new Abstract: Tabular model performance degrades when feature distributions change over time or the relationship between features and outcome variables change over ti

model-releasesarxiv-cs-lg
5 Aug 2026
Model Releases

Oilbird: Training-Free Speculative Decoding with Keys the Verifier Already Computes

DGX agent

arXiv:2608.03839v1 Announce Type: new Abstract: Training-free speculative decoding drafts by matching an exact suffix of the context against a pool of earlier context. That lookup misses correct draft

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet

DGX agent

arXiv:2608.03428v1 Announce Type: cross Abstract: Image based dietary assessment offers a scalable alternative to self reported food diaries, yet fine-grained food recognition remains challenging due

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

Omega-S: A Functional Resilience Index for LLM Fine-Tuning

DGX agent

arXiv:2608.03887v1 Announce Type: new Abstract: Fine-tuning a large language model on new data degrades what it previously learned. We present Omega-S, a drop-in penalty computed from the weight matri

model-releasesarxiv-cs-lg
5 Aug 2026
Model Releases

On the Implicit Flatness Bias of Sharpness-Aware Minimization: A Linear Stability Analysis with Quantitative Hyperparameter Bounds

DGX agent

arXiv:2608.03197v1 Announce Type: new Abstract: Sharpness-Aware Minimization (SAM) improves generalization by seeking parameters whose loss is robust to local adversarial perturbations, but the quanti

model-releasesarxiv-cs-lg
5 Aug 2026
Model Releases

On the missing benchmarks layer and a potential solution

DGX agent

arXiv:2608.02996v1 Announce Type: new Abstract: Latin America is missing a foundational layer for native AI development: the benchmark layer. The benchmark layer does two things no other layer can - i

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

On the missing data layer and a potential solution

DGX agent

arXiv:2608.02949v1 Announce Type: new Abstract: Latin America is missing two foundational layers of AI infrastructure: the dataset layer and the benchmark layer. This paper targets the dataset layer.

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

OncoTriad-QA: A Patient-Level Radiology-Pathology-Genomics Benchmark for Pan-Cancer Reasoning

DGX agent

arXiv:2608.02615v1 Announce Type: cross Abstract: Cancer diagnosis and characterization require integrating complementary evidence from radiology, pathology, genomics, and clinical metadata. However,

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

One Knob to Rule Them All: A Unified Optimal Transport View of Cold-Start Active Learning

DGX agent

arXiv:2608.03249v1 Announce Type: new Abstract: Cold-Start Active Learning (CSAL) aims to select a valuable subset from an unlabeled pool without any prior knowledge or human assistance. Existing meth

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

One-Point Contraction: Erasing Representational Separability toward Irreversible Deep Forgetting

DGX agent

arXiv:2507.07754v3 Announce Type: replace-cross Abstract: Machine unlearning is usually evaluated by what the classifier outputs: forget-set accuracy, confidence, membership-inference scores. We show

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

One-shotting a Raccoon Heist game using Claude Fable 5

DGX agent

Back in 2024 I tweeted screenshots of a game concept generated by GPT-3 and some concept 'art' created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5

model-releasessimon-willison
5 Aug 2026
Model Releases

Operationally Feasible Synthetic Power-Grid Scenarios via Learning the AC-Operable Joint Distribution

DGX agent

arXiv:2608.03878v1 Announce Type: new Abstract: Synthetic power-grid scenarios are essential for planning, resilience assessment, contingency analysis, and data-driven power-system applications. Recen

model-releasesarxiv-cs-lg
5 Aug 2026
Model Releases

PACE: Adaptive Budget Allocation for Time-Efficient Embodied Planning

DGX agent

arXiv:2608.03034v1 Announce Type: cross Abstract: Reasoning-enhanced large language models have achieved remarkable improvements in planning tasks, yet their deployment in embodied systems remains imp

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

Particle-based Generalised Stochastic Optimisation

DGX agent

arXiv:2608.02844v1 Announce Type: cross Abstract: We develop a class of diffusion-based stochastic particle optimisation methods for loss functions with intractable gradients. Specifically, we conside

model-releasesarxiv-cs-lg
5 Aug 2026
Model Releases

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs

DGX agent

arXiv:2608.04010v1 Announce Type: cross Abstract: Existing scaling strategies for Multimodal Large Language Models (MLLMs) typically expand either model parameters or sequential inference computation,

model-releasesarxiv-cs-cl
5 Aug 2026
Model Releases

PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents

DGX agent

arXiv:2608.04003v1 Announce Type: new Abstract: Recursive self-improvement requires agents to turn accumulated experience into better future behavior. Personal AI agents offer a concrete setting for s

model-releasesarxiv-cs-cl
5 Aug 2026
Model Releases

Pattern over Pixels: Measuring Pattern Completion Bias in Multimodal Code Generation

DGX agent

arXiv:2608.03691v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are increasingly used to translate webpage screenshots into front-end code, but repeated UI patterns may sway

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

PI-Mem: Pushing Long-Context Reasoning to 3.6M Tokens with Parallel-Iterative Memory

DGX agent

arXiv:2608.03048v1 Announce Type: cross Abstract: Long-context reasoning remains a critical bottleneck for large language models, as recent recurrent-memory approaches face two inherent challenges: se

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

Pin Once, Swap Light: Subspace-Aligned Centroid-Residual Training for Efficient Ultra-LoRA Serving

DGX agent

arXiv:2608.03579v1 Announce Type: cross Abstract: Modern multi-tenant Low-Rank Adapters (LoRAs) serving systems concurrently host tens to hundreds of LoRA adapters. Though powerful, this introduces a

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

Pingala: Prosody-Aware Decoding for Sanskrit Poetry Generation

DGX agent

arXiv:2603.24413v2 Announce Type: replace Abstract: Poetry generation in Sanskrit typically requires the verse to be semantically coherent and adhere to strict prosodic rules. In Sanskrit prosody, eve

model-releasesarxiv-cs-cl
5 Aug 2026
Model Releases

PLAN: Parallel Liquid-Inspired Approximation Network for Efficient Representation Learning in Flexible Job Shop Scheduling

DGX agent

arXiv:2608.03041v1 Announce Type: cross Abstract: Deep reinforcement learning (DRL) approaches for flexible job shop scheduling (FJSP) heavily rely on attention-centric architectures to achieve state-

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

Predictive Enhancement Calibration for Latent Breast MRI Virtual Contrast Enhancement

DGX agent

arXiv:2608.03612v1 Announce Type: cross Abstract: Virtual contrast enhancement (VCE) synthesizes enhanced breast MR images from pre-contrast acquisitions. Modern latent generators offer strong image p

model-releasesarxiv-cs-cv
5 Aug 2026
Model Releases

Prime Agent - a new coding harness surpassing Codex/CC/PI

DGX agent

Prime Agent is an open-source coding and research agent for general and long-running work. A self-improving RLM harness for coding and long-running autonomous tasks. Designed to be both token-efficien

model-releasesr-localllama
5 Aug 2026
Model Releases

PRISMA: Improving the Accuracy-Latency Frontier of Diffusion-based PDE Solvers Using Physics-Informed Spectral Attention

DGX agent

arXiv:2512.01370v2 Announce Type: replace-cross Abstract: Diffusion-based solvers for partial differential equations (PDEs) are often bottle-necked by slow gradient-based test-time optimization routin

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

Provably Learning Multi-Head Attention with Queries

DGX agent

arXiv:2608.03294v1 Announce Type: new Abstract: We study the problem of learning multi-head softmax attention from black-box input-output access. The learner may query arbitrary real-valued token sequ

model-releasesarxiv-cs-lg
5 Aug 2026
Model Releases

PSA Update CUDA from 13.2 to 13.3 to solve DeepSeek V4 Flash 0731 Looping Problem!

DGX agent

So one of yall mentioned that cuda 13.1 or 13.2 is broken for unsloth so I looked in to it, and they were right. I had 13.2 installed, after I switched to 13.3 no more looping!!! Before the cuda updat

model-releasesr-localllama
5 Aug 2026
Model Releases

Quantifying Hallucinations in Language Language Models on Medical Textbooks

DGX agent

arXiv:2603.09986v3 Announce Type: replace-cross Abstract: Hallucinations, the tendency for large language models to provide responses with factually incorrect and unsupported claims, is a serious prob

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

Quantization Effects on Biomedical LLM Reliability

DGX agent

arXiv:2608.03854v1 Announce Type: new Abstract: When decoder language models are used as classifiers, predicted class probabilities depend on implementation choices, including the prompt template, ver

model-releasesarxiv-cs-lg
5 Aug 2026
Model Releases

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding

DGX agent

arXiv:2608.02980v1 Announce Type: new Abstract: Large Multimodal Models (LMMs) have achieved remarkable success on images and short videos, yet scaling them to long videos remains challenging due to f

model-releasesarxiv-cs-cv
5 Aug 2026
Model Releases

Qwen Developers' responses from their recent Twitter/X AMA

DGX agent

Questions & Responses(in BOLD) below. Favorite question(s) moved to end of the thread with combined responses(removed duplicates). Be optimistic folks. I'm sure we're getting other models too apart fr

model-releasesr-localllama
5 Aug 2026
Model Releases

Qwen-Image-3.0-Pro is live on Qwen Cloud now! Try it out👇 https://www.qwencloud.com/models/qwen-image-3.0-pro?utm_content=g_20000001188

DGX agent

Qwen-Image-3.0-Pro is live on Qwen Cloud now! Try it out👇 https://www.qwencloud.com/models/qwen-image-3.0-pro?utm_content=g_20000001188 🔔 Qwen-Image-3.0 is now live on Qwen Cloud! Ranked #1 among Chin

model-releasesqwen--x
5 Aug 2026
Model Releases

Qwen3-TTS voice cloning is now in mainline llama.cpp — the old demo finally became real support

DGX agent

People may remember the Qwen3-TTS llama.cpp demo from a few months ago. That PR said it probably wouldn’t be merged because llama.cpp was missing some of the graph and API pieces it needed. A new impl

model-releasesr-localllama
5 Aug 2026
Model Releases

Qwen3.8-Max hits #2 in Image-to-WebDev Arena! It sees, it builds~😎

DGX agent

Qwen3.8-Max hits #2 in Image-to-WebDev Arena! It sees, it builds~😎 Exciting news: Qwen3.8-Max by @Alibaba_Qwen is #2 in Image-to-WebDev Arena! With 1,631 pts, it’s trailing only Claude Opus 5 (Max) by

model-releasesqwen--x
5 Aug 2026
Model Releases

Reachability Is Not Realization: Tracing the Sources of LLM Benchmark Gains

DGX agent

arXiv:2608.03219v1 Announce Type: new Abstract: Benchmark gains are often treated as evidence of greater LLM capability. Yet the same gain can reflect different changes in model behavior. A model may

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

Rectify Then Diffuse: Disentangling Concepts Before Denoising Trajectory Unfolds

DGX agent

arXiv:2608.03135v1 Announce Type: cross Abstract: Text-to-image diffusion models can generate individual concepts well, but they often omit or merge concepts incorrectly with multiple concepts. We tra

model-releasesarxiv-cs-ai
5 Aug 2026
← Previous
1…3940414243…466
Next →