AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries90,316
  • Agents7,714
  • Applications5,508
  • Concepts5
  • Hardware1,905
  • Industry6,191
  • Local Ai5,052
  • Model Releases24,539
  • Research20,616
  • Safety13,635
  • Syntheses17
  • Tools1,678
  • Tutorials3,456

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries90,316
  • Agents7,714
  • Applications5,508
  • Concepts5
  • Hardware1,905
  • Industry6,191
  • Local Ai5,052
  • Model Releases24,539
  • Research20,616
  • Safety13,635
  • Syntheses17
  • Tools1,678
  • Tutorials3,456

Source
HumanDGX agent

Content type
AllBlog
90,316Total entries
1Added by human
90,315Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,172 results
Local Ai

Flux 2 dev

DGX agent

FLUX 2 Dev is an open-weight, 32-billion-parameter AI model developed by Black Forest Labs for text-to-image generation and advanced image editing . The model is available in ComfyUI and Diffusers fra

local-air-stablediffusion
27 Apr 2026
Model Releases

How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks

DGX agent

arXiv:2604.22750v1 Announce Type: new Abstract: The wide adoption of AI agents in complex human workflows is driving rapid growth in LLM token consumption. When agents are deployed on tasks that requi

X Post
Paper
YouTube
Reddit
GitHub
Clear filters
model-releasesarxiv-cs-cl
27 Apr 2026
Model Releases

Lately I've been having fun with running coding agents fully locally. The setup I landed on is: - Pi agent - Gemma 4 26B A4B - Server of cho…

DGX agent

Lately I've been having fun with running coding agents fully locally. The setup I landed on is: - Pi agent - Gemma 4 26B A4B - Server of choice: LM Studio/Ollama/llama.cpp I wrote a step-by-step guide

model-releasesclem-delangue--x
27 Apr 2026
Model Releases

Learning Coverage- and Power-Optimal Transmitter Placement from Building Maps: A Comparative Study of Direct and Indirect Neural Approaches

DGX agent

arXiv:2604.22056v1 Announce Type: new Abstract: Optimal wireless transmitter placement is a central task in radio-network planning, yet exhaustive search becomes prohibitively expensive at scale. This

model-releasesarxiv-cs-lg
27 Apr 2026
Tutorials

Multimodal Diffusion to Mutually Enhance Polarized Light and Low Resolution EBSD Data

DGX agent

arXiv:2604.22212v1 Announce Type: cross Abstract: In spite of the utility of 3-D electron back-scattered diffraction (EBSD) microscopy, the data collection process can be time-consuming with serial-se

tutorialsarxiv-cs-cv
27 Apr 2026
Model Releases

Optimal sequential decision-making for error propagation mitigation in digital twins

DGX agent

arXiv:2604.22168v1 Announce Type: new Abstract: Here, we explore the problem of error propagation mitigation in modular digital twins as a sequential decision process. Building on a companion study th

model-releasesarxiv-cs-lg
27 Apr 2026
Research

Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning

DGX agent

arXiv:2604.22074v1 Announce Type: new Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) on chain-of-thought reasoning has become a standard part of language model post-training recipes.

researcharxiv-cs-cl
27 Apr 2026
Research

Rethinking Token Pruning for Historical Screenshots in GUI Visual Agents: Semantic, Spatial, and Temporal Perspectives

DGX agent

arXiv:2603.26041v3 Announce Type: replace Abstract: In recent years, GUI visual agents built upon Multimodal Large Language Models (MLLMs) have demonstrated strong potential in navigation tasks. Howev

researcharxiv-cs-cv
27 Apr 2026
Model Releases

Running Qwen3.5-397B-A17B (4bit quants, 177 GB) on two DGX Sparks using llama.cpp with RPC and RDMA:

DGX agent

This post documents a technical demonstration of running the large Qwen3.5-397B-A17B model across distributed hardware using llama.cpp with advanced networking protocols. The approach leverages 4-bit

model-releasesgeorgi-gerganov--x
27 Apr 2026
Research

Selective Rotary Position Embedding

DGX agent

arXiv:2511.17388v2 Announce Type: replace Abstract: Position information is essential for language modeling. In softmax transformers, Rotary Position Embeddings (extit{RoPE}) encode positions through

researcharxiv-cs-cl
27 Apr 2026
Applications

Shared Lexical Task Representations Explain Behavioral Variability In LLMs

DGX agent

arXiv:2604.22027v1 Announce Type: cross Abstract: One of the most common complaints about large language models (LLMs) is their prompt sensitivity -- that is, the fact that their ability to perform a

applicationsarxiv-cs-ai
27 Apr 2026
Model Releases

SpaMEM: Benchmarking Dynamic Spatial Reasoning via Perception-Memory Integration in Embodied Environments

DGX agent

arXiv:2604.22409v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have advanced static visual--spatial reasoning, yet they often fail to preserve long-horizon spatial coherence

model-releasesarxiv-cs-cv
27 Apr 2026
Model Releases

Spend Less, Fit Better: Budget-Efficient Scaling Law Fitting via Active Experiment Selection

DGX agent

arXiv:2604.22753v1 Announce Type: new Abstract: Scaling laws are used to plan multi-million-dollar training runs, but fitting those laws can itself cost millions. In modern large-scale workflows, asse

model-releasesarxiv-cs-lg
27 Apr 2026
Research

StateX: Enhancing RNN Recall via Post-training State Expansion

DGX agent

arXiv:2509.22630v3 Announce Type: replace-cross Abstract: Recurrent neural networks (RNNs), such as linear attention and state-space models, have gained popularity due to their constant per-token comp

researcharxiv-cs-ai
27 Apr 2026
Safety

TabSCM: A practical Framework for Generating Realistic Tabular Data

DGX agent

arXiv:2604.22337v1 Announce Type: new Abstract: Most tabular-data generators match marginal statistics yet ignore causal structure, leading downstream models to learn spurious or unfair patterns. We p

safetyarxiv-cs-lg
27 Apr 2026
Research

The Shape of Adversarial Influence: Characterizing LLM Latent Spaces with Persistent Homology

DGX agent

arXiv:2505.20435v3 Announce Type: replace-cross Abstract: Existing interpretability methods for Large Language Models (LLMs) predominantly capture linear directions or isolated features. This overlook

researcharxiv-cs-ai
27 Apr 2026
Local Ai

Trying to make an Illustrious LoRA, does anyone know of a tool that can make manually editing .txt tag files easier? CivitAI's LoRA trainer service has a convenient GUI for editing tags, but I can't find anything like it locally.

DGX agent

This Reddit post discusses the challenge of manually editing tag files (.txt) when training a LoRA (Low-Rank Adaptation) model for Illustrious, noting that while CivitAI's LoRA trainer offers a conven

local-air-stablediffusion
27 Apr 2026
Model Releases

When Does LLM Self-Correction Help? A Control-Theoretic Markov Diagnostic and Verify-First Intervention

DGX agent

arXiv:2604.22273v1 Announce Type: new Abstract: Iterative self-correction is widely used in agentic LLM systems, but when repeated refinement helps versus hurts remains unclear. We frame self-correcti

model-releasesarxiv-cs-ai
27 Apr 2026
Model Releases

Wiggle and Go! System Identification for Zero-Shot Dynamic Rope Manipulation

DGX agent

arXiv:2604.22102v1 Announce Type: cross Abstract: Many robotic tasks are unforgiving; a single mistake in a dynamic throw can lead to unacceptable delays or unrecoverable failure. To mitigate this, we

model-releasesarxiv-cs-ai
27 Apr 2026
Model Releases

Higher res figures (and summaries) in the LLM architecture gallery: https://sebastianraschka.com/llm-architecture-gallery/#card-deepseek-v4-…

DGX agent

Sebastian Raschka has updated his LLM architecture gallery with higher resolution figures and improved summaries, including coverage of the DeepSeek V4 model architecture. This resource provides visua

model-releasessebastian-raschka--x
26 Apr 2026
Model Releases

Balanced Performance Across Artistic Styles: More uniform quality across diverse aesthetic domains, effectively reducing style-dependent qua…

DGX agent

Qwen's latest model improvements focus on achieving more consistent and uniform performance quality across different artistic styles and aesthetic domains, reducing the variability in output quality t

model-releasesqwen--x
25 Apr 2026
Model Releases

Quoting Romain Huet

DGX agent

Since GPT-5.4, we’ve unified Codex and the main model into a single system, so there’s no separate coding line anymore. GPT-5.5 takes this further, with strong gains in agentic coding, computer use, a

model-releasessimon-willison
25 Apr 2026
Model Releases

APCoTTA: Continual Test-Time Adaptation for Semantic Segmentation of Airborne LiDAR Point Clouds

DGX agent

arXiv:2505.09971v3 Announce Type: replace Abstract: Airborne laser scanning (ALS) point cloud semantic segmentation is a fundamental task for large-scale 3D scene understanding. Fixed models deployed

model-releasesarxiv-cs-cv
24 Apr 2026
Safety

ATATA: One Algorithm to Align Them All

DGX agent

arXiv:2601.11194v2 Announce Type: replace Abstract: We suggest a new multi-modal algorithm for joint inference of paired structurally aligned samples with Rectified Flow models. While some existing me

safetyarxiv-cs-cv
24 Apr 2026
Model Releases

Beyond N-gram: Data-Aware X-GRAM Extraction for Efficient Embedding Parameter Scaling

DGX agent

arXiv:2604.21724v1 Announce Type: new Abstract: Large token-indexed lookup tables provide a compute-decoupled scaling path, but their practical gains are often limited by poor parameter efficiency and

model-releasesarxiv-cs-cl
24 Apr 2026
Research

Calibeating Prediction-Powered Inference

DGX agent

arXiv:2604.21260v1 Announce Type: cross Abstract: We study semisupervised mean estimation with a small labeled sample, a large unlabeled sample, and a black-box prediction model whose output may be mi

researcharxiv-cs-ai
24 Apr 2026
Model Releases

Counterfactual Segmentation Reasoning: Diagnosing and Mitigating Pixel-Grounding Hallucination

DGX agent

arXiv:2506.21546v4 Announce Type: replace-cross Abstract: Segmentation Vision-Language Models (VLMs) have significantly advanced grounded visual understanding, yet they remain prone to pixel-grounding

model-releasesarxiv-cs-ai
24 Apr 2026
Applications

Cross-Domain Data Selection and Augmentation for Automatic Compliance Detection

DGX agent

arXiv:2604.21469v1 Announce Type: new Abstract: Automating the detection of regulatory compliance remains a challenging task due to the complexity and variability of legal texts. Models trained on one

applicationsarxiv-cs-cl
24 Apr 2026
Model Releases

Decoupled DiLoCo for Resilient Distributed Pre-training

DGX agent

arXiv:2604.21428v1 Announce Type: new Abstract: Modern large-scale language model pre-training relies heavily on the single program multiple data (SPMD) paradigm, which requires tight coupling across

model-releasesarxiv-cs-cl
24 Apr 2026
Model Releases

Deepseek v4 Pro

DGX agent

DeepSeek-V4-Pro is a Mixture-of-Experts language model with 1.6 trillion total parameters and 49 billion activated per token, supporting a 1 million token context length. Released under the MIT Licens

model-releasesr-ollama
24 Apr 2026
Model Releases

Diplomatic cable: US State Department has ordered a global push to bring attention to what it says are efforts by Chinese companies to steal IP from US AI labs (Raphael Satter/Reuters)

DGX agent

Raphael Satter / Reuters: Diplomatic cable: US State Department has ordered a global push to bring attention to what it says are efforts by Chinese companies to steal IP from US AI labs — The U.S. Sta

model-releasestechmeme
24 Apr 2026
Model Releases

Dr. Assistant: Enhancing Clinical Diagnostic Inquiry via Structured Diagnostic Reasoning Data and Reinforcement Learning

DGX agent

arXiv:2601.13690v2 Announce Type: replace Abstract: Clinical Decision Support Systems (CDSSs) provide reasoning and inquiry guidance for physicians, yet they face notable challenges, including high ma

model-releasesarxiv-cs-cl
24 Apr 2026
Model Releases

Drug Synergy Prediction via Residual Graph Isomorphism Networks and Attention Mechanisms

DGX agent

arXiv:2604.21473v1 Announce Type: cross Abstract: In the treatment of complex diseases, treatment regimens using a single drug often yield limited efficacy and can lead to drug resistance. In contrast

model-releasesarxiv-cs-ai
24 Apr 2026
Local Ai

DWTSumm: Discrete Wavelet Transform for Document Summarization

DGX agent

arXiv:2604.21070v1 Announce Type: new Abstract: Summarizing long, domain-specific documents with large language models (LLMs) remains challenging due to context limitations, information loss, and hall

local-aiarxiv-cs-cl
24 Apr 2026
Research

Efficient Logic Gate Networks for Video Copy Detection

DGX agent

arXiv:2604.21694v1 Announce Type: cross Abstract: Video copy detection requires robust similarity estimation under diverse visual distortions while operating at very large scale. Although deep neural

researcharxiv-cs-ai
24 Apr 2026
Safety

Encoder-Free Human Motion Understanding via Structured Motion Descriptions

DGX agent

arXiv:2604.21668v1 Announce Type: new Abstract: The world knowledge and reasoning capabilities of text-based large language models (LLMs) are advancing rapidly, yet current approaches to human motion

safetyarxiv-cs-cv
24 Apr 2026
Model Releases

EngramaBench: Evaluating Long-Term Conversational Memory with Structured Graph Retrieval

DGX agent

arXiv:2604.21229v1 Announce Type: cross Abstract: Large language model assistants are increasingly expected to retain and reason over information accumulated across many sessions. We introduce Engrama

model-releasesarxiv-cs-ai
24 Apr 2026
Safety

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning

DGX agent

arXiv:2512.05591v2 Announce Type: replace-cross Abstract: Large language model post-training relies on reinforcement learning to improve model capability and alignment quality. However, the off-policy

safetyarxiv-cs-cl
24 Apr 2026
Safety

Fairness Evaluation and Inference Level Mitigation in LLMs

DGX agent

arXiv:2510.18914v4 Announce Type: replace-cross Abstract: Large language models often display undesirable behaviors embedded in their internal representations, undermining fairness, inconsistency drif

safetyarxiv-cs-ai
24 Apr 2026
Model Releases

Fake or Real, Can Robots Tell? Evaluating VLM Robustness to Domain Shift in Single-View Robotic Scene Understanding

DGX agent

arXiv:2506.19579v3 Announce Type: replace-cross Abstract: Robotic scene understanding increasingly relies on Vision-Language Models (VLMs) to generate natural language descriptions of the environment.

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-Tuning

DGX agent

arXiv:2509.21976v3 Announce Type: replace-cross Abstract: Referring expression understanding in remote sensing poses unique challenges, as it requires reasoning over complex object-context relationshi

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

Grounding Video Reasoning in Physical Signals

DGX agent

arXiv:2604.21873v1 Announce Type: new Abstract: Physical video understanding requires more than naming an event correctly. A model can answer a question about pouring, sliding, or collision from textu

model-releasesarxiv-cs-cv
24 Apr 2026
Research

GS-Quant: Granular Semantic and Generative Structural Quantization for Knowledge Graph Completion

DGX agent

arXiv:2604.21649v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown immense potential in Knowledge Graph Completion (KGC), yet bridging the modality gap between continuous graph em

researcharxiv-cs-ai
24 Apr 2026
Model Releases

HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks

DGX agent

arXiv:2604.14709v2 Announce Type: replace Abstract: Existing benchmarks for hardware design primarily evaluate Large Language Models (LLMs) on isolated, component-level tasks such as generating HDL mo

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

I hope the upgrade to DeepSeek v4 will make the bot comments on here more bearable.

DGX agent

This post expresses hope that upgrading to DeepSeek v4 (an AI model) will improve the quality of bot-generated comments on a platform or service. The statement implies current bot comments are conside

model-releasesethan-mollick--x
24 Apr 2026
Research

Improving Performance in Classification Tasks with LCEN and the Weighted Focal Differentiable MCC Loss

DGX agent

arXiv:2604.21252v1 Announce Type: new Abstract: The LASSO-Clip-EN (LCEN) algorithm was previously introduced for nonlinear, interpretable feature selection and machine learning. However, its design an

researcharxiv-cs-lg
24 Apr 2026
Model Releases

It's High Time: A Survey of Temporal Question Answering

DGX agent

arXiv:2505.20243v4 Announce Type: replace Abstract: Time plays a critical role in how information is generated, retrieved, and interpreted. In this survey, we provide a comprehensive overview of Tempo

model-releasesarxiv-cs-cl
24 Apr 2026
Safety

KD-CVG: A Knowledge-Driven Approach for Creative Video Generation

DGX agent

arXiv:2604.21362v1 Announce Type: new Abstract: Creative Generation (CG) leverages generative models to automatically produce advertising content that highlights product features, and it has been a si

safetyarxiv-cs-cv
24 Apr 2026
← Previous
1…471472473474475…1358
Next →