AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
HumanDGX agent

87,171Total entries
1Added by human
87,170Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,612 results
28 Apr 2026

MVIGER: Multi-View Variational Integration of Complementary Knowledge for Generative Recommender

ApplicationsDGX agent

arXiv:2408.08686v4 Announce Type: replace-cross Abstract: Language Models (LMs) have been widely used in recommender systems to incorporate textual information of items into item IDs, leveraging their

No Test Cases, No Problem: Distillation-Driven Code Generation for Scientific Workflows

Model ReleasesDGX agent

arXiv:2604.23106v1 Announce Type: cross Abstract: Existing multi-agent Large Language Model (LLM) frameworks for code generation typically use execution feedback and improve iteratively using Input/Ou

Nvidia introduces Nemotron 3 Nano Omni with vision and speech for powerful agentic AI use

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Nvidia Corp. today launched a powerful reasoning artificial intelligence model that unifies text, vision and speech, capable of acting as the “brains” of faster, smarter agentic AI applications. Dubbe

Patterns vs. Patients: Evaluating LLMs against Mental Health Professionals on Personality Disorder Diagnosis through First-Person Narratives

Model ReleasesDGX agent

arXiv:2512.20298v2 Announce Type: replace-cross Abstract: Growing reliance on LLMs for psychiatric self-assessment raises questions about their ability to interpret qualitative patient narratives. Thi

PoseX: AI Defeats Physics Approaches on Protein-Ligand Cross Docking

Model ReleasesDGX agent

arXiv:2505.01700v3 Announce Type: replace Abstract: Existing protein-ligand docking studies typically focus on the self-docking scenario, which is less practical in real applications. Moreover, some s

Progressive Approximation in Deep Residual Networks: Theory and Validation

Model ReleasesDGX agent

arXiv:2604.24154v1 Announce Type: cross Abstract: The Universal Approximation Theorem (UAT) guarantees universal function approximation but does not explain how residual models distribute approximatio

Protecting the Trace: A Principled Black-Box Approach Against Distillation Attacks

SafetyDGX agent

arXiv:2604.23238v1 Announce Type: cross Abstract: Frontier models push the boundaries of what is learnable at extreme computational costs, yet distillation via sampling reasoning traces exposes closed

Quantifying Divergence in Inter-LLM Communication Through API Retrieval and Ranking

SafetyDGX agent

arXiv:2604.22760v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly operate as autonomous agents that reason over external APIs to perform complex tasks. However, their reliabi

RealFin: How Well Do LLMs Reason About Finance When Users Leave Things Unsaid?

Model ReleasesDGX agent

arXiv:2602.07096v2 Announce Type: replace-cross Abstract: Reliable financial reasoning requires knowing not only how to answer, but also when an answer cannot be justified. In real financial practice,

Resource-Lean Lexicon Induction for German Dialects

Model ReleasesDGX agent

arXiv:2604.23824v1 Announce Type: new Abstract: Automatic induction of high-quality dictionaries is essential for building lexical resources, yet low-resource languages and dialects pose several chall

Self Knowledge Re-expression: A Fully Local Method for Adapting LLMs to Tasks Using Intrinsic Knowledge

Local AiDGX agent

arXiv:2604.22939v1 Announce Type: cross Abstract: While the next-token prediction (NTP) paradigm enables large language models (LLMs) to express their intrinsic knowledge, its sequential nature constr

SemiSAM-O1: How far can we push the boundary of annotation-efficient medical image segmentation?

ResearchDGX agent

arXiv:2604.24109v1 Announce Type: new Abstract: Semi-supervised learning (SSL) has become a promising solution to alleviate the annotation burden of deep learning-based medical image segmentation mode

SEVerA: Verified Synthesis of Self-Evolving Agents

Model ReleasesDGX agent

arXiv:2603.25111v2 Announce Type: replace Abstract: Recent advances have shown the effectiveness of self-evolving LLM agents on tasks such as program repair and scientific discovery. In this paradigm,

Skill Retrieval Augmentation for Agentic AI

Model ReleasesDGX agent

arXiv:2604.24594v1 Announce Type: cross Abstract: As large language models (LLMs) evolve into agentic problem solvers, they increasingly rely on external, reusable skills to handle tasks beyond their

Stabilizing Efficient Reasoning with Step-Level Advantage Selection

Model ReleasesDGX agent

arXiv:2604.24003v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong reasoning performance by allocating substantial computation at inference time, often generating long and ver

Switch Attention: Towards Dynamic and Fine-grained Hybrid Transformers

Model ReleasesDGX agent

arXiv:2603.26380v2 Announce Type: replace Abstract: The attention mechanism has been the core component in modern transformer architectures. However, the computation of standard full attention scales

Symmetric Equilibrium Propagation for Thermodynamic Diffusion Training

Model ReleasesDGX agent

arXiv:2604.23806v1 Announce Type: cross Abstract: The reverse process in score-based diffusion models is formally equivalent to overdamped Langevin dynamics in a time-dependent energy landscape. In ou

SynthPert: Enhancing LLM Biological Reasoning via Synthetic Reasoning Traces for Cellular Perturbation Prediction

Model ReleasesDGX agent

arXiv:2509.25346v2 Announce Type: replace Abstract: Predicting cellular responses to genetic perturbations represents a fundamental challenge in systems biology, critical for advancing therapeutic dis

Test-Time Adaptation for Unsupervised Combinatorial Optimization

Local AiDGX agent

arXiv:2601.21048v2 Announce Type: replace Abstract: Unsupervised neural combinatorial optimization (NCO) enables learning powerful solvers without access to ground-truth solutions. Existing approaches

The new LLM trained only on pre-1931 text is small enough that it can potentially run on device, so, with the right tools, you can get a ful…

ApplicationsDGX agent

The new LLM trained only on pre-1931 text is small enough that it can potentially run on device, so, with the right tools, you can get a fully vintage version of Siri, but from the era of Downton Abbe

The Price of Agreement: Measuring LLM Sycophancy in Agentic Financial Applications

Model ReleasesDGX agent

arXiv:2604.24668v1 Announce Type: new Abstract: Given the increased use of LLMs in financial systems today, it becomes important to evaluate the safety and robustness of such systems. One failure mode

Try now: https://www.together.ai/models/nvidia-nemotron-3-nano-omni#

Model ReleasesDGX agent

NVIDIA Nemotron-3 Nano Omni is now available to try through Together AI's platform, offering access to a compact multimodal model capable of processing both text and audio inputs. This announcement hi

Welcome to the agentic era: Public sector highlights and reflections from Next ‘26

Model ReleasesDGX agent

Welcome to the agentic era! Last week, leaders from our public sector customer and partner ecosystem took the stage at Google Cloud Next to share how they are leveraging AI and agents to scale their i

When VLMs 'Fix' Students: Identifying and Penalizing Over-Correction in the Evaluation of Multi-line Handwritten Math OCR

Model ReleasesDGX agent

arXiv:2604.22774v1 Announce Type: cross Abstract: Accurate transcription of handwritten mathematics is crucial for educational AI systems, yet current benchmarks fail to evaluate this capability prope

World-R1: Reinforcing 3D Constraints for Text-to-Video Generation

SafetyDGX agent

arXiv:2604.24764v1 Announce Type: new Abstract: Recent video foundation models demonstrate impressive visual synthesis but frequently suffer from geometric inconsistencies. While existing methods atte

27 Apr 2026

Bridging the Long-Tail Gap: Robust Retrieval-Augmented Relation Completion via Multi-Stage Paraphrase Infusion

Model ReleasesDGX agent

arXiv:2604.22261v1 Announce Type: new Abstract: Large language models (LLMs) struggle with relation completion (RC), both with and without retrieval-augmented generation (RAG), particularly when the r

Call-Chain-Aware LLM-Based Test Generation for Java Projects

Model ReleasesDGX agent

arXiv:2604.22046v1 Announce Type: cross Abstract: Large language models (LLMs) have recently shown strong potential for generating project-level unit tests. However, existing state-of-the-art approach

Causal Concept Graphs in LLM Latent Space for Stepwise Reasoning

ResearchDGX agent

arXiv:2603.10377v2 Announce Type: replace-cross Abstract: Sparse autoencoders can localize where concepts live in language models, but not how they interact during multi-step reasoning. We propose Cau

DocPrune:Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning

ResearchDGX agent

arXiv:2604.22281v1 Announce Type: new Abstract: Recent advances in vision-language models have demonstrated remarkable performance across diverse multi-modal tasks, including document question answeri

Flux 2 dev

Local AiDGX agent

FLUX 2 Dev is an open-weight, 32-billion-parameter AI model developed by Black Forest Labs for text-to-image generation and advanced image editing . The model is available in ComfyUI and Diffusers fra

How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks

Model ReleasesDGX agent

arXiv:2604.22750v1 Announce Type: new Abstract: The wide adoption of AI agents in complex human workflows is driving rapid growth in LLM token consumption. When agents are deployed on tasks that requi

Lately I've been having fun with running coding agents fully locally. The setup I landed on is: - Pi agent - Gemma 4 26B A4B - Server of cho…

Model ReleasesDGX agent

Lately I've been having fun with running coding agents fully locally. The setup I landed on is: - Pi agent - Gemma 4 26B A4B - Server of choice: LM Studio/Ollama/llama.cpp I wrote a step-by-step guide

Learning Coverage- and Power-Optimal Transmitter Placement from Building Maps: A Comparative Study of Direct and Indirect Neural Approaches

Model ReleasesDGX agent

arXiv:2604.22056v1 Announce Type: new Abstract: Optimal wireless transmitter placement is a central task in radio-network planning, yet exhaustive search becomes prohibitively expensive at scale. This

Multimodal Diffusion to Mutually Enhance Polarized Light and Low Resolution EBSD Data

TutorialsDGX agent

arXiv:2604.22212v1 Announce Type: cross Abstract: In spite of the utility of 3-D electron back-scattered diffraction (EBSD) microscopy, the data collection process can be time-consuming with serial-se

Optimal sequential decision-making for error propagation mitigation in digital twins

Model ReleasesDGX agent

arXiv:2604.22168v1 Announce Type: new Abstract: Here, we explore the problem of error propagation mitigation in modular digital twins as a sequential decision process. Building on a companion study th

Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning

ResearchDGX agent

arXiv:2604.22074v1 Announce Type: new Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) on chain-of-thought reasoning has become a standard part of language model post-training recipes.

Rethinking Token Pruning for Historical Screenshots in GUI Visual Agents: Semantic, Spatial, and Temporal Perspectives

ResearchDGX agent

arXiv:2603.26041v3 Announce Type: replace Abstract: In recent years, GUI visual agents built upon Multimodal Large Language Models (MLLMs) have demonstrated strong potential in navigation tasks. Howev

Running Qwen3.5-397B-A17B (4bit quants, 177 GB) on two DGX Sparks using llama.cpp with RPC and RDMA:

Model ReleasesDGX agent

This post documents a technical demonstration of running the large Qwen3.5-397B-A17B model across distributed hardware using llama.cpp with advanced networking protocols. The approach leverages 4-bit

Selective Rotary Position Embedding

ResearchDGX agent

arXiv:2511.17388v2 Announce Type: replace Abstract: Position information is essential for language modeling. In softmax transformers, Rotary Position Embeddings (extit{RoPE}) encode positions through

Shared Lexical Task Representations Explain Behavioral Variability In LLMs

ApplicationsDGX agent

arXiv:2604.22027v1 Announce Type: cross Abstract: One of the most common complaints about large language models (LLMs) is their prompt sensitivity -- that is, the fact that their ability to perform a

SpaMEM: Benchmarking Dynamic Spatial Reasoning via Perception-Memory Integration in Embodied Environments

Model ReleasesDGX agent

arXiv:2604.22409v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have advanced static visual--spatial reasoning, yet they often fail to preserve long-horizon spatial coherence

Spend Less, Fit Better: Budget-Efficient Scaling Law Fitting via Active Experiment Selection

Model ReleasesDGX agent

arXiv:2604.22753v1 Announce Type: new Abstract: Scaling laws are used to plan multi-million-dollar training runs, but fitting those laws can itself cost millions. In modern large-scale workflows, asse

StateX: Enhancing RNN Recall via Post-training State Expansion

ResearchDGX agent

arXiv:2509.22630v3 Announce Type: replace-cross Abstract: Recurrent neural networks (RNNs), such as linear attention and state-space models, have gained popularity due to their constant per-token comp

TabSCM: A practical Framework for Generating Realistic Tabular Data

SafetyDGX agent

arXiv:2604.22337v1 Announce Type: new Abstract: Most tabular-data generators match marginal statistics yet ignore causal structure, leading downstream models to learn spurious or unfair patterns. We p

The Shape of Adversarial Influence: Characterizing LLM Latent Spaces with Persistent Homology

ResearchDGX agent

arXiv:2505.20435v3 Announce Type: replace-cross Abstract: Existing interpretability methods for Large Language Models (LLMs) predominantly capture linear directions or isolated features. This overlook

Trying to make an Illustrious LoRA, does anyone know of a tool that can make manually editing .txt tag files easier? CivitAI's LoRA trainer service has a convenient GUI for editing tags, but I can't find anything like it locally.

Local AiDGX agent

This Reddit post discusses the challenge of manually editing tag files (.txt) when training a LoRA (Low-Rank Adaptation) model for Illustrious, noting that while CivitAI's LoRA trainer offers a conven

When Does LLM Self-Correction Help? A Control-Theoretic Markov Diagnostic and Verify-First Intervention

Model ReleasesDGX agent

arXiv:2604.22273v1 Announce Type: new Abstract: Iterative self-correction is widely used in agentic LLM systems, but when repeated refinement helps versus hurts remains unclear. We frame self-correcti

Wiggle and Go! System Identification for Zero-Shot Dynamic Rope Manipulation

Model ReleasesDGX agent

arXiv:2604.22102v1 Announce Type: cross Abstract: Many robotic tasks are unforgiving; a single mistake in a dynamic throw can lead to unacceptable delays or unrecoverable failure. To mitigate this, we

26 Apr 2026

Higher res figures (and summaries) in the LLM architecture gallery: https://sebastianraschka.com/llm-architecture-gallery/#card-deepseek-v4-…

Model ReleasesDGX agent

Sebastian Raschka has updated his LLM architecture gallery with higher resolution figures and improved summaries, including coverage of the DeepSeek V4 model architecture. This resource provides visua

25 Apr 2026

Balanced Performance Across Artistic Styles: More uniform quality across diverse aesthetic domains, effectively reducing style-dependent qua…

Model ReleasesDGX agent

Qwen's latest model improvements focus on achieving more consistent and uniform performance quality across different artistic styles and aesthetic domains, reducing the variability in output quality t

Quoting Romain Huet

Model ReleasesDGX agent

Since GPT-5.4, we’ve unified Codex and the main model into a single system, so there’s no separate coding line anymore. GPT-5.5 takes this further, with strong gains in agentic coding, computer use, a

24 Apr 2026

APCoTTA: Continual Test-Time Adaptation for Semantic Segmentation of Airborne LiDAR Point Clouds

Model ReleasesDGX agent

arXiv:2505.09971v3 Announce Type: replace Abstract: Airborne laser scanning (ALS) point cloud semantic segmentation is a fundamental task for large-scale 3D scene understanding. Fixed models deployed

ATATA: One Algorithm to Align Them All

SafetyDGX agent

arXiv:2601.11194v2 Announce Type: replace Abstract: We suggest a new multi-modal algorithm for joint inference of paired structurally aligned samples with Rectified Flow models. While some existing me

Beyond N-gram: Data-Aware X-GRAM Extraction for Efficient Embedding Parameter Scaling

Model ReleasesDGX agent

arXiv:2604.21724v1 Announce Type: new Abstract: Large token-indexed lookup tables provide a compute-decoupled scaling path, but their practical gains are often limited by poor parameter efficiency and

Calibeating Prediction-Powered Inference

ResearchDGX agent

arXiv:2604.21260v1 Announce Type: cross Abstract: We study semisupervised mean estimation with a small labeled sample, a large unlabeled sample, and a black-box prediction model whose output may be mi

Counterfactual Segmentation Reasoning: Diagnosing and Mitigating Pixel-Grounding Hallucination

Model ReleasesDGX agent

arXiv:2506.21546v4 Announce Type: replace-cross Abstract: Segmentation Vision-Language Models (VLMs) have significantly advanced grounded visual understanding, yet they remain prone to pixel-grounding

Cross-Domain Data Selection and Augmentation for Automatic Compliance Detection

ApplicationsDGX agent

arXiv:2604.21469v1 Announce Type: new Abstract: Automating the detection of regulatory compliance remains a challenging task due to the complexity and variability of legal texts. Models trained on one

Decoupled DiLoCo for Resilient Distributed Pre-training

Model ReleasesDGX agent

arXiv:2604.21428v1 Announce Type: new Abstract: Modern large-scale language model pre-training relies heavily on the single program multiple data (SPMD) paradigm, which requires tight coupling across

Deepseek v4 Pro

Model ReleasesDGX agent

DeepSeek-V4-Pro is a Mixture-of-Experts language model with 1.6 trillion total parameters and 49 billion activated per token, supporting a 1 million token context length. Released under the MIT Licens

Diplomatic cable: US State Department has ordered a global push to bring attention to what it says are efforts by Chinese companies to steal IP from US AI labs (Raphael Satter/Reuters)

Model ReleasesDGX agent

Raphael Satter / Reuters: Diplomatic cable: US State Department has ordered a global push to bring attention to what it says are efforts by Chinese companies to steal IP from US AI labs — The U.S. Sta

← Previous
1…360361362363364…1044
Next →