AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “synthesis”

GridTimelineEvolution
2,818 results
14 Apr 2026

On the Effectiveness of Textual Prompting with Lightweight Fine-Tuning for SAM3 Remote Sensing Segmentation

SafetyDGX agent

arXiv:2512.15564v2 Announce Type: replace Abstract: Remote sensing (RS) image segmentation is constrained by the limited availability of annotated data and a gap between overhead imagery and natural i

Problem Reductions at Scale: Agentic Integration of Computationally Hard Problems

AgentsDGX agent

arXiv:2604.11535v1 Announce Type: new Abstract: Solving an NP-hard optimization problem often requires reformulating it for a specific solver -- quantum hardware, a commercial optimizer, or a domain h

Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.10949v1 Announce Type: cross Abstract: Unified multimodal models (UMMs) were designed to combine the reasoning ability of large language models (LLMs) with the generation capability of visi

Retrieval as Generation: A Unified Framework with Self-Triggered Information Planning

TutorialsDGX agent

arXiv:2604.11407v1 Announce Type: cross Abstract: We revisit retrieval-augmented generation (RAG) by embedding retrieval control directly into generation. Instead of treating retrieval as an external

RTMC: Step-Level Credit Assignment via Rollout Trees

AgentsDGX agent

arXiv:2604.11037v1 Announce Type: cross Abstract: Multi-step agentic reinforcement learning benefits from fine-grained credit assignment, yet existing approaches offer limited options: critic-free met

Subargument Argumentation Frameworks: Separating Direct Conflict from Structural Dependency

ResearchDGX agent

arXiv:2601.12038v3 Announce Type: replace Abstract: Dung's abstract argumentation frameworks model acceptability solely in terms of an attack relation, thereby conflating two conceptually distinct asp

SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning

TutorialsDGX agent

arXiv:2604.10228v1 Announce Type: new Abstract: Current multimodal models often suffer from shallow reasoning, leading to errors caused by incomplete or inconsistent thought processes. To address this

TaleDiffusion: Multi-Character Story Generation with Dialogue Rendering

ResearchDGX agent

arXiv:2509.04123v2 Announce Type: replace Abstract: Text-to-story visualization is challenging due to the need for consistent interaction among multiple characters across frames. Existing methods stru

TRACE: An Experiential Framework for Coherent Multi-hop Knowledge Graph Question Answering

TutorialsDGX agent

arXiv:2604.11193v1 Announce Type: new Abstract: Multi-hop Knowledge Graph Question Answering (KGQA) requires coherent reasoning across relational paths, yet existing methods often treat each reasoning

X-SYS: A Reference Architecture for Interactive Explanation Systems

ResearchDGX agent

arXiv:2602.12748v3 Announce Type: replace Abstract: The explainable AI (XAI) research community has proposed numerous technical methods, yet deploying explainability as systems remains challenging: In

13 Apr 2026

BlendFusion -- Scalable Synthetic Data Generation for Diffusion Model Training

ResearchDGX agent

arXiv:2604.09022v1 Announce Type: new Abstract: With the rapid adoption of diffusion models, synthetic data generation has emerged as a promising approach for addressing the growing demand for large-s

CatalogStitch: Dimension-Aware and Occlusion-Preserving Object Compositing for Catalog Image Generation

Model ReleasesDGX agent

arXiv:2604.08836v1 Announce Type: new Abstract: Generative object compositing methods have shown remarkable ability to seamlessly insert objects into scenes. However, when applied to real-world catalo

Dual Mamba for Node-Specific Representation Learning: Tackling Over-Smoothing with Selective State Space Modeling

ResearchDGX agent

arXiv:2511.06756v3 Announce Type: replace Abstract: Over-smoothing remains a fundamental challenge in deep Graph Neural Networks (GNNs), where repeated message passing causes node representations to b

ELT: Elastic Looped Transformers for Visual Generation

Model ReleasesDGX agent

arXiv:2604.09168v1 Announce Type: new Abstract: We introduce Elastic Looped Transformers (ELT), a highly parameter-efficient class of visual generative models based on a recurrent transformer architec

Gen-n-Val: Agentic Image Data Generation and Validation

AgentsDGX agent

arXiv:2506.04676v2 Announce Type: replace-cross Abstract: The data scarcity, label noise, and long-tailed category imbalance remain important and unresolved challenges in many computer vision tasks, s

Hitem3D 2.0: Multi-View Guided Native 3D Texture Generation

SafetyDGX agent

arXiv:2604.09231v1 Announce Type: new Abstract: Although recent advances have improved the quality of 3D texture generation, existing methods still struggle with incomplete texture coverage, cross-vie

Maybe hot take - I’ve read a bunch of RL for image generation papers over last few months and honestly it’s been pretty disappointing. All o…

TutorialsDGX agent

Maybe hot take - I’ve read a bunch of RL for image generation papers over last few months and honestly it’s been pretty disappointing. All of them are variations of GRPO and all of them are incrementa

Measurement-Consistent Langevin Corrector for Stabilizing Latent Diffusion Inverse Problem Solvers

ResearchDGX agent

arXiv:2601.04791v3 Announce Type: replace Abstract: While latent diffusion models (LDMs) have emerged as powerful priors for inverse problems, existing LDM-based solvers frequently suffer from instabi

OmniPrism: Learning Disentangled Visual Concept for Image Generation

TutorialsDGX agent

arXiv:2412.12242v2 Announce Type: replace-cross Abstract: Creative visual concept generation often draws inspiration from specific concepts in a reference image to produce relevant outcomes. However,

Overhang Tower: Resource-Rational Adaptation in Sequential Physical Planning

ResearchDGX agent

arXiv:2604.09072v1 Announce Type: new Abstract: Humans effortlessly navigate the physical world by predicting how objects behave under gravity and contact forces, yet how such judgments support sequen

RIRF: Reasoning Image Restoration Framework

AgentsDGX agent

arXiv:2604.09511v1 Announce Type: new Abstract: Universal image restoration (UIR) aims to recover clean images from diverse and unknown degradations using a unified model. Existing UIR methods primari

Training-free, Perceptually Consistent Low-Resolution Previews with High-Resolution Image for Efficient Workflows of Diffusion Models

TutorialsDGX agent

arXiv:2604.09227v1 Announce Type: cross Abstract: Image generative models have become indispensable tools to yield exquisite high-resolution (HR) images for everyone, ranging from general users to pro

V-CAGE: Vision-Closed-Loop Agentic Generation Engine for Robotic Manipulation

AgentsDGX agent

arXiv:2604.09036v1 Announce Type: new Abstract: Scaling Vision-Language-Action (VLA) models requires massive datasets that are both semantically coherent and physically feasible. However, existing sce

11 Apr 2026

struggling choosing one edit model from klein 9b or qwen 2511.

Model ReleasesDGX agent

This r/StableDiffusion thread discusses the community debate around choosing between FLUX.2 [klein] 9B and Qwen Image Edit 2511 as an image editing model, two strong open-source contenders in the spac

The one thing I still don't know how to do: TTS/singing a specific song but with a specific voice

TutorialsDGX agent

The specific Reddit post could not be retrieved from the search results. However, based on the context of the URL and related results, I can provide the following best-effort summary based on what ...

10 Apr 2026

Advanced inpaint/edit Klein/Qwen workflows

Model ReleasesDGX agent

A Reddit post on r/StableDiffusion discussing advanced ComfyUI workflows that combine the FLUX Klein and Qwen Image Edit models for precision inpainting and image editing tasks. FLUX Klein offers ...

AnomalyVFM -- Transforming Vision Foundation Models into Zero-Shot Anomaly Detectors

Model ReleasesDGX agent

arXiv:2601.20524v2 Announce Type: replace Abstract: Zero-shot anomaly detection aims to detect and localise abnormal regions in the image without access to any in-domain training images. While recent

Balanced Diffusion-Guided Fusion for Multimodal Remote Sensing Classification

TutorialsDGX agent

arXiv:2509.23310v3 Announce Type: replace Abstract: Deep learning-based techniques for the analysis of multimodal remote sensing data have become popular due to their ability to effectively integrate

CAMotion: A High-Quality Benchmark for Camouflaged Moving Object Detection in the Wild

Model ReleasesDGX agent

arXiv:2604.08287v1 Announce Type: new Abstract: Discovering camouflaged objects is a challenging task in computer vision due to the high similarity between camouflaged objects and their surroundings.

Distilling Specialized Orders for Visual Generation

ResearchDGX agent

arXiv:2504.17069v2 Announce Type: replace Abstract: Autoregressive (AR) image generators are becoming increasingly popular due to their ability to produce high-quality images and their scalability. Ty

DMin: Scalable Training Data Influence Estimation for Diffusion Models

ResearchDGX agent

arXiv:2412.08637v4 Announce Type: replace Abstract: Identifying the training data samples that most influence a generated image is a critical task in understanding diffusion models (DMs), yet existing

DP-DeGauss: Dynamic Probabilistic Gaussian Decomposition for Egocentric 4D Scene Reconstruction

ResearchDGX agent

arXiv:2604.07986v1 Announce Type: new Abstract: Egocentric video is crucial for next-generation 4D scene reconstruction, with applications in AR/VR and embodied AI. However, reconstructing dynamic fir

Evaluating Low-Light Image Enhancement Across Multiple Intensity Levels

Model ReleasesDGX agent

arXiv:2511.15496v2 Announce Type: replace Abstract: Imaging in low-light environments is challenging due to reduced scene radiance, which leads to elevated sensor noise and reduced color saturation. M

Face2Scene: Using Facial Degradation as an Oracle for Diffusion-Based Scene Restoration

TutorialsDGX agent

arXiv:2603.16570v2 Announce Type: replace Abstract: Recent advances in image restoration have enabled high-fidelity recovery of faces from degraded inputs using reference-based face restoration models

FIT: A Large-Scale Dataset for Fit-Aware Virtual Try-On

Model ReleasesDGX agent

arXiv:2604.08526v1 Announce Type: new Abstract: Given a person and a garment image, virtual try-on (VTO) aims to synthesize a realistic image of the person wearing the garment, while preserving their

Inference-Time Scaling of Diffusion Language Models via Trajectory Refinement

TutorialsDGX agent

arXiv:2507.08390v4 Announce Type: replace Abstract: Discrete diffusion models have recently emerged as strong alternatives to autoregressive language models, matching their performance through large-s

LumiCtrl : Learning Illuminant Prompts for Lighting Control in Personalized Text-to-Image Models

ResearchDGX agent

arXiv:2512.17489v2 Announce Type: replace Abstract: Text-to-image (T2I) models have demonstrated remarkable progress in creative image generation, yet they still lack precise control over scene illumi

Matrix Profile for Time-Series Anomaly Detection: A Reproducible Open-Source Benchmark on TSB-AD

Model ReleasesDGX agent

arXiv:2604.02445v2 Announce Type: replace Abstract: Matrix Profile (MP) methods are an interpretable and scalable family of distance-based methods for time-series anomaly detection, but strong benchma

Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models

Local AiDGX agent

arXiv:2601.04068v3 Announce Type: replace Abstract: Aligning text-to-video diffusion models with human preferences is crucial for generating high-quality videos. Existing Direct Preference Otimization

MoRight: Motion Control Done Right

TutorialsDGX agent

arXiv:2604.07348v1 Announce Type: cross Abstract: Generating motion-controlled videos--where user-specified actions drive physically plausible scene dynamics under freely chosen viewpoints--demands tw

MV-SAM3D: Adaptive Multi-View Fusion for Layout-Aware 3D Generation

ApplicationsDGX agent

arXiv:2603.11633v2 Announce Type: replace Abstract: Recent unified 3D generation models have made remarkable progress in producing high-quality 3D assets from a single image. Notably, layout-aware app

PhyEdit: Towards Real-World Object Manipulation via Physically-Grounded Image Editing

Model ReleasesDGX agent

arXiv:2604.07230v2 Announce Type: replace Abstract: Achieving physically accurate object manipulation in image editing is essential for its potential applications in interactive world models. However,

Physical Knot Classification Beyond Accuracy: A Benchmark and Diagnostic Study

Model ReleasesDGX agent

arXiv:2603.23286v3 Announce Type: replace Abstract: Physical knot classification is a fine-grained task in which the intended cue is rope crossing structure, but high accuracy may still come from appe

PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratch

Model ReleasesDGX agent

arXiv:2510.06670v2 Announce Type: replace Abstract: High-quality instruction data is critical for LLM alignment, yet existing open-source datasets often lack efficiency, requiring hundreds of thousand

PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards

Model ReleasesDGX agent

arXiv:2512.01236v2 Announce Type: replace Abstract: Personalized generation models for a single subject have demonstrated remarkable effectiveness, highlighting their significant potential. However, w

ReconPhys: Reconstruct Appearance and Physical Attributes from Single Video

ResearchDGX agent

arXiv:2604.07882v1 Announce Type: new Abstract: Reconstructing non-rigid objects with physical plausibility remains a significant challenge. Existing approaches leverage differentiable rendering for p

SciFigDetect: A Benchmark for AI-Generated Scientific Figure Detection

Model ReleasesDGX agent

arXiv:2604.08211v1 Announce Type: new Abstract: Modern multimodal generators can now produce scientific figures at near-publishable quality, creating a new challenge for visual forensics and research

Self-Improving 4D Perception via Self-Distillation

ResearchDGX agent

arXiv:2604.08532v1 Announce Type: new Abstract: Large-scale multi-view reconstruction models have made remarkable progress, but most existing approaches still rely on fully supervised training with gr

SHAPE: Stage-aware Hierarchical Advantage via Potential Estimation for LLM Reasoning

ResearchDGX agent

arXiv:2604.06636v1 Announce Type: cross Abstract: Process supervision has emerged as a promising approach for enhancing LLM reasoning, yet existing methods fail to distinguish meaningful progress from

Towards Hierarchical Multi-Step Reward Models for Enhanced Reasoning in Large Language Models

ResearchDGX agent

arXiv:2503.13551v5 Announce Type: replace Abstract: Recent studies show that Large Language Models (LLMs) achieve strong reasoning capabilities through supervised fine-tuning or reinforcement learning

9 Apr 2026

Studying Sutton and Barto's RL book and its connections to RL for LLMs (e.g., tool use, math reasoning, agents, and so on)? [D]

AgentsDGX agent

A Reddit discussion thread on r/MachineLearning in which practitioners explore how foundational concepts from Sutton and Barto's *Reinforcement Learning: An Introduction* — including MDPs, policy g...

8 Apr 2026

PS: I finally got around to trying out @randal_olson 's Tufte Test tool to prettify the benchmark plot. Great tool 👌! https://www.goodeyela…

Model ReleasesDGX agent

Sebastian Raschka (rasbt) used Randal Olson's Tufte Test tool, developed by Goodeye Labs, to improve the visual quality of a machine learning benchmark plot. The Tufte Test encodes seven of Tufte'...

12 Aug 2026

AlbumentationsX: One Augmentation Pipeline for Images and Related Annotations

ResearchDGX agent

arXiv:2608.11123v1 Announce Type: new Abstract: Augmentation can corrupt a training example when an image and its annotations receive different random changes. A crop must use the same coordinates for

CARE: Confidence-Aware Reasoning for Reliable Medical VQA

SafetyDGX agent

arXiv:2608.10964v1 Announce Type: cross Abstract: Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual qu

DreamOmni3: Scribble-based Editing and Generation

Model ReleasesDGX agent

arXiv:2512.22525v2 Announce Type: replace Abstract: Recently unified generation and editing models have achieved remarkable success with their impressive performance. These models rely mainly on text

DSAR: Dual-Stream Autoregressive Modeling of Temporal Cloth Dynamics for Photorealistic Animatable Avatars

TutorialsDGX agent

arXiv:2608.10500v1 Announce Type: new Abstract: Creating photorealistic and temporally coherent animatable human avatars from RGB videos remains challenging. Current methods struggle to capture realis

Embodied Multimodal Grounding for Open-Vocabulary Mobile Manipulation via Semantic 3D Gaussian Splatting

Local AiDGX agent

arXiv:2608.10756v1 Announce Type: cross Abstract: Embodied mobile manipulation requires language, visual observations, three-dimensional scene structure, and action feasibility to be aligned before ex

Fast and Memory-Efficient Wavelet Convolutions via I/O-Aware Reformulation

Model ReleasesDGX agent

arXiv:2608.10805v1 Announce Type: cross Abstract: Wavelet convolution (WTConv) has emerged as an increasingly popular drop-in replacement for standard convolutions, expanding a network's receptive fie

FoR-SALE: Frame of Reference-guided Spatial Adjustment in LLM-based Diffusion Editing

SafetyDGX agent

arXiv:2509.23452v2 Announce Type: replace-cross Abstract: Current text-to-image generation models, even state-of-the-art models, exhibit a significant performance gap when spatial expressions are desc

Introspective Attention Modulation for Safe Text-to-Image Generation

Model ReleasesDGX agent

arXiv:2607.14945v2 Announce Type: replace Abstract: State-of-the-art flow based text-to-image (T2I) models exhibit remarkable generative abilities but remain vulnerable to producing unsafe content. Pr

← Previous
1…2122232425…47
Next →