AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,490 results
28 Apr 2026

Machine Learning and Deep Learning Models for Short Term Electricity Price Forecasting in Australia's National Electricity Market

Model ReleasesDGX agent

arXiv:2604.23908v1 Announce Type: new Abstract: Short term electricity price forecast is essential in competitive power markets, yet electricity price series exhibit high volatility, irregularity, and

Machine learning models for estimating counterfactuals in a single-arm inflammatory bowel disease study

ResearchDGX agent

arXiv:2604.23465v1 Announce Type: new Abstract: Single-arm trials accelerate study timelines by reducing the number of patients that must be recruited for a concurrent control group. However, these de

Nemotron-3-Nano-Omni-30B-A3B-Reasoning, New model?

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding for enterprise Q&A, summarization, transcription, and document intelligence, w

Nvidia launches Nemotron 3 Nano Omni, an open multimodal model with a 30B-A3B hybrid MoE architecture; the Nemotron 3 family saw 50M+ downloads in the past year (Kyt Dotson/SiliconANGLE)

Model ReleasesDGX agent

Kyt Dotson / SiliconANGLE: Nvidia launches Nemotron 3 Nano Omni, an open multimodal model with a 30B-A3B hybrid MoE architecture; the Nemotron 3 family saw 50M+ downloads in the past year — Nvidia Cor

NVILA: Efficient Frontier Visual Language Models

ResearchDGX agent

arXiv:2412.04468v3 Announce Type: replace Abstract: Visual language models (VLMs) have made significant advances in accuracy in recent years. However, their efficiency has received much less attention

OpenAI models, Codex, and Managed Agents come to AWS

Model ReleasesDGX agent

OpenAI announced the availability of its models, including Codex, and managed agent capabilities on Amazon Web Services (AWS) infrastructure. This integration enables AWS customers to access OpenAI's

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling

Model ReleasesDGX agent

arXiv:2410.05970v3 Announce Type: replace-cross Abstract: Multimodal document understanding is a challenging task to process and comprehend large amounts of textual and visual information. Recent adva

Seeing Is No Longer Believing: Frontier Image Generation Models, Synthetic Visual Evidence, and Real-World Risk

Model ReleasesDGX agent

arXiv:2604.24197v1 Announce Type: cross Abstract: Frontier image generation has moved from artistic synthesis toward synthetic visual evidence. Systems such as GPT Image 2, Nano Banana Pro, Nano Banan

Small Language Model Helps Resolve Semantic Ambiguity of LLM Prompt

ResearchDGX agent

arXiv:2604.23263v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly utilized in various complex reasoning tasks due to their excellent instruction following capability. How

Speech Enhancement Based on Drifting Models

Model ReleasesDGX agent

arXiv:2604.24199v1 Announce Type: cross Abstract: We propose Speech Enhancement based on Drifting Models (DriftSE), a novel generative framework that formulates denoising as an equilibrium problem. Ra

Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models

Model ReleasesDGX agent

arXiv:2511.08577v2 Announce Type: replace-cross Abstract: Improving reasoning abilities of Large Language Models (LLMs), especially under parameter constraints, is crucial for real-world applications.

Today we’re shipping Laguna M.1 and Laguna XS.2 – our first public models. We’re also shipping our agent harness and a preview product exper…

AgentsDGX agent

Today we’re shipping Laguna M.1 and Laguna XS.2 – our first public models. We’re also shipping our agent harness and a preview product experience. Both models were trained from scratch on our own stac

Unified Multi-Foundation-Model Slide Representation for Pan-Cancer Recognition and Text-Guided Tumor Localization

Local AiDGX agent

arXiv:2604.22846v1 Announce Type: new Abstract: The expanding ecosystem of pathology foundation models has produced powerful but fragmented tile-level representations, limiting their use in clinical t

27 Apr 2026

Can Large Language Models Adequately Perform Symbolic Reasoning Over Time Series?

Model ReleasesDGX agent

arXiv:2508.03963v4 Announce Type: replace Abstract: Uncovering hidden symbolic laws from time series data, as an aspiration dating back to Kepler's discovery of planetary motion, remains a core challe

FILTR: Extracting Topological Features from Pretrained 3D Models

Model ReleasesDGX agent

arXiv:2604.22334v1 Announce Type: new Abstract: Recent advances in pretraining 3D point cloud encoders (e.g., Point-BERT, Point-MAE) have produced powerful models, whose abilities are typically evalua

Fine-Grained Analysis of Shared Syntactic Mechanisms in Language Models

Local AiDGX agent

arXiv:2604.22166v1 Announce Type: new Abstract: While language models demonstrate sophisticated syntactic capabilities, the extent to which their internal mechanisms align with cross-constructional pr

From Interpretability to Performance: Optimizing Retrieval Heads for Long-Context Language Models

Model ReleasesDGX agent

arXiv:2601.11020v3 Announce Type: replace Abstract: Advances in mechanistic interpretability have identified special attention heads, known as retrieval heads, that are responsible for retrieving info

going to buy 2 rtx 6000 just because of the capabilities of local models becoming great! no more outsourcing of research to the api

Model ReleasesDGX agent

going to buy 2 rtx 6000 just because of the capabilities of local models becoming great! no more outsourcing of research to the api Xiaomi MiMo-V2.5 is now officially open-sourced! MIT License, suppor

Graph-to-Vision: Multi-graph Understanding and Reasoning using Vision-Language Models

Model ReleasesDGX agent

arXiv:2503.21435v3 Announce Type: replace Abstract: Recent advances in Vision-Language Models (VLMs) have shown promising capabilities in interpreting visualized graph data, offering a new perspective

MambaCSP: Hybrid-Attention State Space Models for Hardware-Efficient Channel State Prediction

Local AiDGX agent

arXiv:2604.21957v1 Announce Type: cross Abstract: Recent works have demonstrated that attention-based transformer and large language model (LLM) architectures can achieve strong channel state predicti

Mochi: Aligning Pre-training and Inference for Efficient Graph Foundation Models via Meta-Learning

ApplicationsDGX agent

arXiv:2604.22031v1 Announce Type: cross Abstract: We propose Mochi, a Graph Foundation Model that addresses task unification and training efficiency by adopting a meta-learning based training framewor

MTT-Bench: Predicting Social Dominance in Mice via Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2604.22492v1 Announce Type: cross Abstract: Understanding social dominance in animal behavior is critical for neuroscience and behavioral studies. In this work, we explore the capability of Mult

Multimodal Neural Operators for Real-Time Biomechanical Modelling of Traumatic Brain Injury

HardwareDGX agent

arXiv:2510.03248v3 Announce Type: replace-cross Abstract: Background: Traumatic brain injury modeling requires integrating volumetric neuroimaging, demographic parameters, and acquisition metadata. Fi

Nuclear Diffusion Models for Low-Rank Background Suppression in Videos

ApplicationsDGX agent

arXiv:2509.20886v2 Announce Type: replace Abstract: Video sequences often contain structured noise and background artifacts that obscure dynamic content, posing challenges for accurate analysis and re

Relaxation-Informed Training of Neural Network Surrogate Models

Model ReleasesDGX agent

arXiv:2604.22746v1 Announce Type: cross Abstract: ReLU neural networks trained as surrogate models can be embedded exactly in mixed-integer linear programs (MILPs), enabling global optimization over t

Score-based Membership Inference on Diffusion Models

Model ReleasesDGX agent

arXiv:2509.25003v2 Announce Type: replace-cross Abstract: Membership inference attacks (MIAs) against Diffusion Models (DMs) raise pressing privacy concerns by revealing whether a sample was part of t

SpikingBrain2.0: Brain-Inspired Foundation Models for Efficient Long-Context and Cross-Platform Inference

HardwareDGX agent

arXiv:2604.22575v1 Announce Type: new Abstract: Scaling context length is reshaping large-model development, yet full-attention Transformers suffer from prohibitive computation and inference bottlenec

Spontaneous Persuasion: An Audit of Model Persuasiveness in Everyday Conversations

ResearchDGX agent

arXiv:2604.22109v1 Announce Type: cross Abstract: Large language models (LLMs) possess strong persuasive capabilities that outperform humans in head-to-head comparisons. Users report consulting LLMs t

Top 3 trending models of the week on HF: @deepseek_ai @OpenAI & @Alibaba_Qwen!

IndustryDGX agent

This post highlights the three most popular models on Hugging Face during a given week, featuring DeepSeek AI, OpenAI, and Alibaba's Qwen models. The post was shared by Clem Delangue, CEO of Hugging F

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis

SafetyDGX agent

arXiv:2604.22225v1 Announce Type: new Abstract: While generative text-to-speech (TTS) models approach human-level quality, monolithic metrics fail to diagnose fine-grained acoustic artifacts or explai

Where Should LoRA Go? Component-Type Placement in Hybrid Language Models

ResearchDGX agent

arXiv:2604.22127v1 Announce Type: new Abstract: Hybrid language models that interleave attention with recurrent components are increasingly competitive with pure Transformers, yet standard LoRA practi

25 Apr 2026

New Dflash drafting model for the 27b Lets gooooo https://huggingface.co/z-lab/Qwen3.6-27B-DFlash

IndustryDGX agent

A new Dflash drafting model based on Qwen 3.6 with 27 billion parameters has been released on Hugging Face, available at z-lab/Qwen3.6-27B-DFlash. This model likely implements speculative decoding or

24 Apr 2026

500+ likes in 28 mins. On their way to be the fastest model ever to get to #1 trending on HF! https://huggingface.co/deepseek-ai/DeepSeek-V4…

Model ReleasesDGX agent

DeepSeek-V4 rapidly gained over 500 likes within 28 minutes on Hugging Face, demonstrating exceptional user engagement and positioning it as a strong contender to become the fastest model to reach #1

A-IC3: Learning-Guided Adaptive Inductive Generalization for Hardware Model Checking

Model ReleasesDGX agent

arXiv:2604.21688v1 Announce Type: cross Abstract: The IC3 algorithm represents the state-of-the-art (SOTA) hardware model checking technique, owing to its robust performance and scalability. A signifi

AITP: Traffic Accident Responsibility Allocation via Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2604.20878v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress in Traffic Accident Detection (TAD) and Traffic Accident Understanding (TAU).

Anthropic details Project Deal, a marketplace experiment where Claude models bought, sold, and negotiated personal belongings on behalf of Anthropic employees (Anthropic)

Model ReleasesDGX agent

Anthropic: Anthropic details Project Deal, a marketplace experiment where Claude models bought, sold, and negotiated personal belongings on behalf of Anthropic employees — At Anthropic, we're interest

🎉 Day-0 support for @deepseek_ai V4 Pro and Flash on vLLM — a new generation of DeepSeek model, purpose-built for tasks up to 1M tokens. Al…

Model ReleasesDGX agent

🎉 Day-0 support for @deepseek_ai V4 Pro and Flash on vLLM — a new generation of DeepSeek model, purpose-built for tasks up to 1M tokens. Alongside the release, we're publishing a first-principles walk

DeepSeek V4 Pro costs 1.74/1M input tokens and 3.48/1M output tokens, while V4 Flash costs 0.14/1M and 0.28/1M; both models are the cheapest in their class (Simon Willison/Simon Willison's Weblog)

Model ReleasesDGX agent

Simon Willison / Simon Willison's Weblog: DeepSeek V4 Pro costs 1.74/1M input tokens and 3.48/1M output tokens, while V4 Flash costs 0.14/1M and 0.28/1M; both models are the cheapest in their class —

Differentially Private Model Merging

Model ReleasesDGX agent

arXiv:2604.20985v1 Announce Type: cross Abstract: In machine learning applications, privacy requirements during inference or deployment time could change constantly due to varying policies, regulation

Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition

SafetyDGX agent

arXiv:2604.21276v1 Announce Type: cross Abstract: As pretrained large language models replace task-specific decoders in speech recognition, a critical question arises: do their text-derived priors mak

Federated Co-tuning Framework for Large and Small Language Models

Model ReleasesDGX agent

arXiv:2411.11707v3 Announce Type: replace-cross Abstract: By adapting Large Language Models (LLMs) to domain-specific tasks or enriching them with domain-specific knowledge, we can fully harness the c

Foveated Reasoning: Stateful, Action-based Visual Focusing for Vision-Language Models

ResearchDGX agent

arXiv:2604.21079v1 Announce Type: new Abstract: Vision-language models benefit from high-resolution images, but the increase in visual-token count incurs high compute overhead. Humans resolve this ten

FunduSegmenter: Leveraging the RETFound Foundation Model for Joint Optic Disc and Optic Cup Segmentation in Retinal Fundus Images

ResearchDGX agent

arXiv:2508.11354v3 Announce Type: replace-cross Abstract: Purpose: This study introduces the first adaptation of RETFound for joint optic disc (OD) and optic cup (OC) segmentation. RETFound is a well-

GPT-5.5 is now available on Perplexity for Max subscribers. GPT-5.5 is also rolling out as the default orchestration model in Computer for b…

Model ReleasesDGX agent

Perplexity has made GPT-5.5 available to Max subscribers and is rolling it out as the default orchestration model in Perplexity Computer. The announcement indicates expanded access to OpenAI's GPT-5.5

IRIS: Interpolative Renyi Iterative Self-play for Large Language Model Fine-Tuning

Model ReleasesDGX agent

arXiv:2604.20933v1 Announce Type: cross Abstract: Self-play fine-tuning enables large language models to improve beyond supervised fine-tuning without additional human annotations by contrasting annot

M-CARE: Standardized Clinical Case Reporting for AI Model Behavioral Disorders, with a 20-Case Atlas and Experimental Validation

ResearchDGX agent

arXiv:2604.20871v1 Announce Type: cross Abstract: We introduce M-CARE (Model Clinical Assessment and Reporting for Evaluation), a clinical case report framework for AI model behavioral disorders adapt

Machine Behavior in Relational Moral Dilemmas: Moral Rightness, Predicted Human Behavior, and Model Decisions

SafetyDGX agent

arXiv:2604.21871v1 Announce Type: new Abstract: Human moral judgment is context-dependent and modulated by interpersonal relationships. As large language models (LLMs) increasingly function as decisio

Multilinguality at the Edge: Developing Language Models for the Global South

ResearchDGX agent

arXiv:2604.21637v1 Announce Type: new Abstract: Where and how language models (LMs) are deployed determines who can benefit from them. However, there are several challenges that prevent effective depl

Nemobot Games: Crafting Strategic AI Gaming Agents for Interactive Learning with Large Language Models

Model ReleasesDGX agent

arXiv:2604.21896v1 Announce Type: new Abstract: This paper introduces a new paradigm for AI game programming, leveraging large language models (LLMs) to extend and operationalize Claude Shannon's taxo

PDGMM-VAE: A Variational Autoencoder with Adaptive Per-Dimension Gaussian Mixture Model Priors for Nonlinear ICA

ResearchDGX agent

arXiv:2603.23547v2 Announce Type: replace-cross Abstract: Independent component analysis is a core framework within blind source separation for recovering latent source signals from observed mixtures

SCM: Sleep-Consolidated Memory with Algorithmic Forgetting for Large Language Models

Model ReleasesDGX agent

arXiv:2604.20943v1 Announce Type: new Abstract: We present SCM (Sleep-Consolidated Memory), a research preview of a memory architecture for large language models that draws on neuroscientific principl

Toward Efficient Membership Inference Attacks against Federated Large Language Models: A Projection Residual Approach

Model ReleasesDGX agent

arXiv:2604.21197v1 Announce Type: new Abstract: Federated Large Language Models (FedLLMs) enable multiple parties to collaboratively fine-tune LLMs without sharing raw data, addressing challenges of l

Trust but Verify: Introducing DAVinCI -- A Framework for Dual Attribution and Verification in Claim Inference for Language Models

ApplicationsDGX agent

arXiv:2604.21193v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable fluency and versatility across a wide range of NLP tasks, yet they remain prone to factual ina

Unlocking the Power of Large Language Models for Multi-table Entity Matching

ResearchDGX agent

arXiv:2604.21238v1 Announce Type: new Abstract: Multi-table entity matching (MEM) addresses the limitations of dual-table approaches by enabling simultaneous identification of equivalent entities acro

23 Apr 2026

Co-Located Tests, Better AI Code: How Test Syntax Structure Affects Foundation Model Code Generation

ResearchDGX agent

arXiv:2604.19826v1 Announce Type: cross Abstract: AI coding assistants increasingly generate code alongside tests. How developers structure test code, whether inline with the implementation or in sepa

Convergent Evolution: How Different Language Models Learn Similar Number Representations

TutorialsDGX agent

arXiv:2604.20817v1 Announce Type: cross Abstract: Language models trained on natural text learn to represent numbers using periodic features with dominant periods at T=2, 5, 10. In this paper, we iden

Do Small Language Models Know When They're Wrong? Confidence-Based Cascade Scoring for Educational Assessment

Model ReleasesDGX agent

arXiv:2604.19781v1 Announce Type: cross Abstract: Automated scoring of student work at scale requires balancing accuracy against cost and latency. In 'cascade' systems, small language models (LMs) han

Explainable Speech Emotion Recognition: Weighted Attribute Fairness to Model Demographic Contributions to Social Bias

SafetyDGX agent

arXiv:2604.19763v1 Announce Type: cross Abstract: Speech Emotion Recognition (SER) systems have growing applications in sensitive domains such as mental health and education, where biased predictions

Handbook of Rough Set Extensions and Uncertainty Models

ResearchDGX agent

arXiv:2604.19794v1 Announce Type: new Abstract: Rough set theory models uncertainty by approximating target concepts through lower and upper sets induced by indiscernibility, or more generally, by gra

Improving clinical interpretability of linear neuroimaging models through feature whitening

ResearchDGX agent

arXiv:2604.20675v1 Announce Type: new Abstract: Linear models are widely used in computational neuroimaging to identify biomarkers associated with brain pathologies. However, interpreting the learned

← Previous
1…9293949596…1009
Next →