AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,428 results
4 Aug 2026

Self-Improving Large Language Models via Progressive Experience Evolution

SafetyDGX agent

arXiv:2608.02139v1 Announce Type: new Abstract: Large language models (LLMs) capable of self-improvement require not only effective policy optimization, but also a principled mechanism for transformin

3 Aug 2026

A Unified Benchmark of Deep Learning Models for Multi-task 3D Brain Tumor Segmentation from Magnetic Resonance Imaging

Model ReleasesDGX agent

arXiv:2607.28858v1 Announce Type: cross Abstract: Automatic brain tumor segmentation from magnetic resonance imaging (MRI) has become a fundamental task in computer-assisted diagnosis, treatment plann

COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2607.28649v1 Announce Type: cross Abstract: COSI-Lab presents a multimodal, multi-sensor dataset of an interdisciplinary scientific workshop containing 32 academics at an international conferenc

Distilling Knowledge from Large Language Models into Lightweight Reinforcement Learning Agents for Autonomous Cyber Operations

Model ReleasesDGX agent

arXiv:2607.28826v1 Announce Type: new Abstract: Autonomous Cyber Operations (ACO) are increasingly important for defending enterprise networks as cyber threats continue to evolve in sophistication. AC

FBFM: A Training-Free Asynchronous Feedback Mechanism for Flow-Matching in World-Action Models Execution

TutorialsDGX agent

arXiv:2607.29235v1 Announce Type: cross Abstract: Although world-action models (WAMs) enhance long-horizon robot control by predicting visual evolution before acting, long-horizon reliability demands

Ling-3.0-flash is another potential model to test before qwen3.8 27b

Model ReleasesDGX agent

I tested Ling-3.0-flash with hard bugs and it fixed bugs that qwen3.6-27b could not. This models speed faster than deepseek v4 flash but almost the same level as (old) deepseek v4 flash. Note: hard bu

You shouldn't need a vision model to know your PDF has checkboxes. LiteParse can now pull structured data directly from your PDFs: form fiel…

Model ReleasesDGX agent

You shouldn't need a vision model to know your PDF has checkboxes. LiteParse can now pull structured data directly from your PDFs: form field values, checkbox states, annotations, embedded images, vec

2 Aug 2026

Instead of limiting the progress of AI companies or preventing them from releasing models as proposed in the bipartisan AI Kill Switch Act, …

IndustryDGX agent

Instead of limiting the progress of AI companies or preventing them from releasing models as proposed in the bipartisan AI Kill Switch Act, Hugging Face CEO Clément Delague says he would rather see Co

LLMs can know a task is impossible and still optimize it anyway. Ask whether to walk or drive to a car wash 50 meters away, and some models …

SafetyDGX agent

LLMs can know a task is impossible and still optimize it anyway. Ask whether to walk or drive to a car wash 50 meters away, and some models focus on distance while missing that the car itself must rea

31 Jul 2026

Benchmarking Foundation and Large Language Models for Few-Shot Medical Image Segmentation

Model ReleasesDGX agent

arXiv:2607.27856v1 Announce Type: new Abstract: Few-shot medical image segmentation (FS-MIS) aims to segment novel regions of interest (ROIs) from a few annotated support examples. Despite rapid progr

DeepSeek V4 Flash 0731 in Hermes Agent and one prompt, took 32 minutes and cost 0.07$, this model is so cheap to the point where 2 dollars c…

Model ReleasesDGX agent

**DeepSeek V4 Flash 0731 Performance Test** On July 31 2026, a single prompt executed via the Hermes Agent on DeepSeek V4 Flash 0731 completed in 32 minutes and incurred an estimated cost of 0.07 USD.

Eco3S: Complex Socio-Economic System Simulation via Agent-Based Models

SafetyDGX agent

arXiv:2607.26588v1 Announce Type: new Abstract: The rapid development of large language models (LLMs) has renewed interest in agent-based modeling (ABM). However, current LLM-based ABM research faces

I have trained a model to predict my blood sugar [P]

Model ReleasesDGX agent

It's an encoder-only transformer that consumes past(blood glucose + carbs + insulin) and future(carbs + insulin) and predicts future blood glucose for the next 2 hours. Announced meals and boluses/bas

ORCA-bench: How Ready Are Language Model Agents for Oncall?

Model ReleasesDGX agent

arXiv:2607.28545v1 Announce Type: new Abstract: Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics,

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

ResearchDGX agent

arXiv:2607.28362v1 Announce Type: new Abstract: We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The obstacle is representational: existi

Sources: OpenAI demoed a new 'Astra' AI model family to US policymakers and regulators this week, touting its improved abilities to complete long-running tasks (The Information)

Model ReleasesDGX agent

The Information: Sources: OpenAI demoed a new “Astra” AI model family to US policymakers and regulators this week, touting its improved abilities to complete long-running tasks — OpenAI is preparing t

(Towards) Scalable Reliable Automated Evaluation with Large Language Models

ResearchDGX agent

arXiv:2607.28282v1 Announce Type: new Abstract: Evaluating the quality and relevance of textual outputs from Large Language Models (LLMs) remains challenging and resource-intensive. Existing automated

We've gotten some great medium sized models lately (DSV4 Flash 0731, Inkling Small, Laguna S 2.1, Step 3.7 Flash) but does anybody else want to see some new 70-80b contenders?

Model ReleasesDGX agent

I can run the mediums, but sometimes I want a faster option that's smarter than Qwen 27B/35B. On my hardware I get like 500 to 800 tok/s prefill and 16 to 22 tok/s gen on ~120B class models, which is

30 Jul 2026

Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steering

SafetyDGX agent

arXiv:2607.26411v1 Announce Type: new Abstract: Unified multimodal models (UMMs) aim to integrate understanding and generation within a single architecture, yet it remains unclear whether these capabi

Equilibrium Training of Energy-Based Models with Parallel Trajectory Tempering

ResearchDGX agent

arXiv:2607.27077v1 Announce Type: new Abstract: Energy-Based Models (EBMs) provide an interpretable framework for generative modeling of scientific data, but poor Markov Chain Monte Carlo mixing often

GPT-5.6 found optimizations that 'reduced end-to-end serving costs by 20%' for OpenAI to serve that model Presumably that's billions of doll…

Model ReleasesDGX agent

GPT-5.6 found optimizations that 'reduced end-to-end serving costs by 20%' for OpenAI to serve that model Presumably that's billions of dollars a month in savings at this point? Codex analysed product

MedARC: Training-Free Adaptive Redundancy Compression of Visual Tokens for 3D Medical Vision-Language Models

Local AiDGX agent

arXiv:2607.26554v1 Announce Type: new Abstract: Integrating 3D medical images with vision-language models (VLMs) holds substantial promise for computer-aided diagnosis. However, volumetric images gene

one of the top cybersecurity models, post-trained from open weights by @depthfirstlabs on @FireworksAI_HQ long-horizon RL is as much an infr…

SafetyDGX agent

one of the top cybersecurity models, post-trained from open weights by @depthfirstlabs on @FireworksAI_HQ long-horizon RL is as much an infra problem as a research one: 100+ turn rollouts, async/pipel

WildShadowRemover: In-the-Wild Video Shadow Removal via Detail-Preserving Video Diffusion Models

Model ReleasesDGX agent

arXiv:2607.26203v1 Announce Type: new Abstract: Video shadow removal in the wild remains challenging due to complex illumination, diverse shadow appearances, and limited training data. Despite its imp

29 Jul 2026

Beyond Counts: A Distributional Robustness Margin For Pathology Foundation Models

Local AiDGX agent

arXiv:2607.25497v1 Announce Type: cross Abstract: Pathology foundation models are approaching clinical deployment, yet remain vulnerable to systematic non-biological variation across centres. Differen

BREAKING: Grok 4.5 just claimed the top spot on the new HighWalk Benchmark. The independent test measures how well AI models update real tec…

Model ReleasesDGX agent

BREAKING: Grok 4.5 just claimed the top spot on the new HighWalk Benchmark. The independent test measures how well AI models update real technical specifications from 46 Laravel commits — heavy on cod

Can Deep Generative Models Reproduce Non-Stationary Gaussian Random Fields?

ApplicationsDGX agent

arXiv:2607.25929v1 Announce Type: cross Abstract: Deep generative models (DGMs) are widely used for complex high-dimensional data and increasingly applied to spatial and spatio-temporal modeling. Thei

Diffusion Model-based Parameter Estimation in Dynamic Power Systems

Model ReleasesDGX agent

arXiv:2411.10431v3 Announce Type: replace Abstract: Parameter estimation, which represents a classical inverse problem, is often ill-posed as different parameter combinations can yield identical outpu

Dual-Domain Manifold Modeling for Hyperspectral Image Fusion

Model ReleasesDGX agent

arXiv:2607.25338v1 Announce Type: new Abstract: Achieving a coherent integration of spectral richness and spatial fidelity remains a central objective in hyperspectral image fusion. However, existing

Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact

Model ReleasesDGX agent

arXiv:2607.25589v1 Announce Type: cross Abstract: Medical-imaging AI benchmarks combine datasets, DICOM rendering, prompts, provider APIs, automated labels, statistical code, manuscripts, and reposito

Inspect India Evals: An Open Benchmarking Framework for Evaluating Large Language Models in the Indian Linguistic and Cultural Context

Model ReleasesDGX agent

arXiv:2607.25375v1 Announce Type: new Abstract: India is a vast nation of over 1.4 billion people, varied by hundreds of diverse and locally specific traditions and cultures and 22 officially recogniz

Medical world models in healthcare: foundations, applications, and challenges for trustworthy clinical translation

SafetyDGX agent

arXiv:2607.25242v1 Announce Type: new Abstract: Medical world models offer a framework for extending medical artificial intelligence beyond static prediction by representing evolving patient states an

Memory for Large Language Models

Model ReleasesDGX agent

arXiv:2607.25380v1 Announce Type: new Abstract: Memory has evolved into a foundational architectural dimension in large language models (LLMs), shifting from an implicit byproduct of computation to a

MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios

Model ReleasesDGX agent

arXiv:2607.25186v1 Announce Type: new Abstract: Background: Most medical large language model (LLM) benchmarks focus on examination knowledge or isolated tasks and may not reflect the longitudinal, mu

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2607.24957v1 Announce Type: new Abstract: We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Model

RDQ: Residual Distribution Quantization for Large Language Models

Model ReleasesDGX agent

arXiv:2607.10137v2 Announce Type: replace Abstract: Post-training quantization (PTQ) of large language models degrades sharply below 4-bit precision. We identify the root cause as residual stream dist

RRS-10K: A Multitask Vision-Language Model Benchmark for Rare Remote Sensing Image Interpretation

Model ReleasesDGX agent

arXiv:2607.24810v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved strong performance on general remote sensing tasks. However, their capability for rare scenes remains insuff

28 Jul 2026

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models

ResearchDGX agent

arXiv:2603.05147v2 Announce Type: replace Abstract: Current research on Vision-Language-Action (VLA) models predominantly focuses on enhancing generalization through reasoning techniques. While effect

Concept-based Visual Counterfactual Explanations with Diffusion Models

SafetyDGX agent

arXiv:2607.22544v1 Announce Type: new Abstract: Visual counterfactual explanations aim to answer 'what minimal change to this image would flip the model's prediction?', and are increasingly important

Development of Vision-Language Model-based GNSS Spoofing Detection for Autonomous Vehicle Navigation

Local AiDGX agent

arXiv:2607.23962v1 Announce Type: new Abstract: Autonomous vehicles (AVs) depend on Global Navigation Satellite Systems (GNSS) for localization and navigation, making them vulnerable to spoofing attac

Exact Evaluation of the Accuracy of Diffusion Models for Inverse Problems with Gaussian Data Distributions

ResearchDGX agent

arXiv:2507.07008v2 Announce Type: replace Abstract: Used as priors for Bayesian inverse problems, diffusion models have recently attracted considerable attention in the literature. Their flexibility a

Guiding Language Models to Be More Empathetic: Culturally Sensitive Mental Health Advice Generation Through Human-LLM Collaboration

Model ReleasesDGX agent

arXiv:2607.23538v1 Announce Type: new Abstract: Despite recent advances in large language models (LLMs), their ability to generate empathetic mental health counseling responses in low-resource languag

How OpenAI hacked HuggingFace. What we know. Hugging Face proved that open platforms and open models can still win those battles when the al…

Model ReleasesDGX agent

How OpenAI hacked HuggingFace. What we know. Hugging Face proved that open platforms and open models can still win those battles when the alternative is locked-down systems that refuse to assist their

Hybrid AI-Physical Modeling for Penetration Bias Correction in X-band InSAR DEMs: A Greenland Case Study

SafetyDGX agent

arXiv:2504.08909v2 Announce Type: replace Abstract: Digital elevation models derived from Interferometric Synthetic Aperture Radar (InSAR) data over glacial and snow-covered regions often exhibit syst

IJCB-AFMFR 2026: Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data

SafetyDGX agent

arXiv:2607.24422v1 Announce Type: new Abstract: This paper presents a summary of the Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data (AFMFR), held at the 2

LEDOM: Reverse Language Model

ResearchDGX agent

arXiv:2507.01335v4 Announce Type: replace-cross Abstract: Autoregressive language models are trained exclusively left-to-right. We explore the complementary factorization, training right-to-left at sc

LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models

Model ReleasesDGX agent

arXiv:2411.00918v5 Announce Type: replace-cross Abstract: Mixture of experts (MoE) architectures have become a cornerstone for scaling up and are a key component in most large language models such as

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models

HardwareDGX agent

arXiv:2607.23193v1 Announce Type: new Abstract: Existing token compression methods for omnimodal large language models typically rely on one modality to determine what to retain in the other. We show

Stability of AI Governance Systems: A Coupled Dynamics Model of Public Trust and Social Disruptions

Model ReleasesDGX agent

arXiv:2603.20248v2 Announce Type: replace-cross Abstract: AI systems are increasingly entrenched in public governance, yet scholarship lacks formal tools to determine when deviations of public trust i

TRE: Training-Free Hallucination Detection for Diffusion Language Models

Model ReleasesDGX agent

arXiv:2607.22661v1 Announce Type: new Abstract: Diffusion large language models (D-LLMs) have recently gained increasing attention, yet their reliability is significantly hindered by the hallucination

Two Regimes of Chain-of-Thought Unfaithfulness: Behavioral Detection Fails Where Models Are Wrong

Model ReleasesDGX agent

arXiv:2607.23458v1 Announce Type: new Abstract: Chain-of-thought (CoT) explanations support oversight only if they are faithful: the stated reasoning must actually produce the answer. Auditing black-b

We recently joined NVIDIA, Microsoft, and others in supporting open-weight AI. Today, we’re putting that belief into the product with model …

HardwareDGX agent

We recently joined NVIDIA, Microsoft, and others in supporting open-weight AI. Today, we’re putting that belief into the product with model choice on Replit, starting with Kimi K3. The future is the r

27 Jul 2026

A Defense of the Quadratic Model

Local AiDGX agent

arXiv:2607.21716v1 Announce Type: new Abstract: Due to the complexity of neural network loss landscapes, optimization theory is forced to rely on idealized models, and there is generally a tradeoff be

Current smallest usable coding model

Model ReleasesDGX agent

I've been seeing a lot of news about the latest gemma 4 and qwen 3.6 being really good and the current go-to models but those are out of reach for my GPU at the moment. With 4GB VRAM and 40 GB RAM, I

Deep Sigma Point Processes for RCS Modeling in Spaceborne SAR Imagery

ResearchDGX agent

arXiv:2607.21745v1 Announce Type: cross Abstract: Radar cross-section (RCS) modeling is foundational to advancing the utility and sensitivity of spaceborne radar systems. This study introduces a deep

Kimi K3 is live on Fireworks. Day 0, inference and training. US-hosted, and zero data retention. This is the first frontier open model in th…

Model ReleasesDGX agent

Kimi K3 is live on Fireworks. Day 0, inference and training. US-hosted, and zero data retention. This is the first frontier open model in the 3 trillion parameter class. It sports 1M context, native v

LLM-Based Visual Explanation Evaluation Framework for Assessing the Explainability of Facial Skin Disease Classification Models

Model ReleasesDGX agent

arXiv:2606.16794v2 Announce Type: replace Abstract: This study proposes a domain-specific LLM-based Visual Explanation Evaluation Framework for assessing visual attention explanations in facial skin d

NVIDIA Nemotron 3 Ultra Leads Open Models on Accuracy and Efficiency in Agentic RTL Coding

Model ReleasesDGX agent

NVIDIA’s Nemotron 3 Ultra, when paired with the ACE‑RTL agent, achieves a 97.1 % average pass rate on the CVDP benchmark across nine RTL task categories—surpassing GLM 5.2 and Kimi K2.6 while using up

Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On

ResearchDGX agent

arXiv:2607.21694v1 Announce Type: new Abstract: We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-purpose image editor, Oxygen-TryOn is

Phylogenetic signal in marine mammal and bird vocalizations captured by audio foundation models: the limited benefit of domain-specific pretraining

ResearchDGX agent

arXiv:2607.22458v1 Announce Type: new Abstract: Do learned audio embeddings encode structure that nobody told them to encode? We probe four large pretrained audio models (AST, CLAP, BEATs-bio and Bird

← Previous
1…7879808182…1008
Next →