AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Research

DeltaSeg: Tiered Attention and Deep Delta Learning for Multi-Class Structural Defect Segmentation

DGX agent

arXiv:2604.18745v1 Announce Type: new Abstract: Automated segmentation of structural defects from visual inspection imagery remains challenging due to the diversity of damage types, extreme class imba

researcharxiv-cs-cv
22 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Denoising, Fast and Slow: Difficulty-Aware Adaptive Sampling for Image Generation

DGX agent

arXiv:2604.19141v1 Announce Type: new Abstract: Diffusion- and flow-based models usually allocate compute uniformly across space, updating all patches with the same timestep and number of function eva

safetyarxiv-cs-cv
22 Apr 2026
Model Releases

Detection of T-shirt Presentation Attacks in Face Recognition Systems

DGX agent

arXiv:2604.19365v1 Announce Type: new Abstract: Face recognition systems are often used for biometric authentication. Nevertheless, it is known that without any protective measures, face recognition s

model-releasesarxiv-cs-cv
22 Apr 2026
Safety

Diff-SBSR: Learning Multimodal Feature-Enhanced Diffusion Models for Zero-Shot Sketch-Based 3D Shape Retrieval

DGX agent

arXiv:2604.19135v1 Announce Type: new Abstract: This paper presents the first exploration of text-to-image diffusion models for zero-shot sketch-based 3D shape retrieval (ZS-SBSR). Existing sketch-bas

safetyarxiv-cs-cv
22 Apr 2026
Safety

DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval

DGX agent

arXiv:2604.19432v1 Announce Type: new Abstract: Vision foundation models have shown great promise for open-set 3D object retrieval (3DOR) through efficient adaptation to multi-view images. Leveraging

safetyarxiv-cs-cv
22 Apr 2026
Local Ai

Divide-and-Conquer Approach to Holistic Cognition in High-Similarity Contexts with Limited Data

DGX agent

arXiv:2604.19339v1 Announce Type: new Abstract: Ultra-fine-grained visual categorization (Ultra-FGVC) aims to classify highly similar subcategories within fine-grained objects using limited training s

local-aiarxiv-cs-cv
22 Apr 2026
Agents

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents

DGX agent

arXiv:2604.19264v1 Announce Type: new Abstract: Agentic multimodal models have garnered significant attention for their ability to leverage external tools to tackle complex tasks. However, it is obser

agentsarxiv-cs-cv
22 Apr 2026
Model Releases

DUALVISION: RGB-Infrared Multimodal Large Language Models for Robust Visual Reasoning

DGX agent

arXiv:2604.18829v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved impressive performance on visual perception and reasoning tasks with RGB imagery, yet they remain

model-releasesarxiv-cs-cv
22 Apr 2026
Model Releases

EfficientPENet: Real-Time Depth Completion from Sparse LiDAR via Lightweight Multi-Modal Fusion

DGX agent

arXiv:2604.18790v1 Announce Type: new Abstract: Depth completion from sparse LiDAR measurements and corresponding RGB images is a prerequisite for accurate 3D perception in robotic systems. Existing m

model-releasesarxiv-cs-cv
22 Apr 2026
Research

EgoMotion: Hierarchical Reasoning and Diffusion for Egocentric Vision-Language Motion Generation

DGX agent

arXiv:2604.19105v1 Announce Type: new Abstract: Faithfully modeling human behavior in dynamic environments is a foundational challenge for embodied intelligence. While conditional motion synthesis has

researcharxiv-cs-cv
22 Apr 2026
Safety

Ensembling Pruned Attention Heads For Uncertainty-Aware Efficient Transformers

DGX agent

arXiv:2510.18358v2 Announce Type: replace-cross Abstract: Uncertainty quantification (UQ) is essential for deploying deep neural networks in safety-critical settings. Although methods like Deep Ensemb

safetyarxiv-cs-cv
22 Apr 2026
Research

EPS: Efficient Patch Sampling for Video Overfitting in Deep Super-Resolution Model Training

DGX agent

arXiv:2411.16312v2 Announce Type: replace Abstract: Leveraging the overfitting property of deep neural networks (DNNs) is trending in video delivery systems to enhance video quality within bandwidth l

researcharxiv-cs-cv
22 Apr 2026
Applications

Evaluating Histogram Matching for Robust Deep learning-Based Grapevine Disease Detection

DGX agent

arXiv:2604.19510v1 Announce Type: new Abstract: Variability in illumination is a primary factor limiting deep learning robustness for field-based plant disease detection. This study evaluates Histogra

applicationsarxiv-cs-cv
22 Apr 2026
Research

Evaluation of Winning Solutions of 2025 Low Power Computer Vision Challenge

DGX agent

arXiv:2604.19054v1 Announce Type: new Abstract: The IEEE Low-Power Computer Vision Challenge (LPCVC) aims to promote the development of efficient vision models for edge devices, balancing accuracy wit

researcharxiv-cs-cv
22 Apr 2026
Agents

Explore Like Humans: Autonomous Exploration with Online SG-Memo Construction for Embodied Agents

DGX agent

arXiv:2604.19034v1 Announce Type: new Abstract: Constructing structured spatial memory is essential for enabling long-horizon reasoning in complex embodied navigation tasks. Current memory constructio

agentsarxiv-cs-cv
22 Apr 2026
Research

Face Anything: 4D Face Reconstruction from Any Image Sequence

DGX agent

arXiv:2604.19702v1 Announce Type: new Abstract: Accurate reconstruction and tracking of dynamic human faces from image sequences is challenging because non-rigid deformations, expression changes, and

researcharxiv-cs-cv
22 Apr 2026
Model Releases

Fast and Robust Diffusion Posterior Sampling for MR Image Reconstruction Using the Preconditioned Unadjusted Langevin Algorithm

DGX agent

arXiv:2512.05791v2 Announce Type: replace-cross Abstract: Purpose: The Unadjusted Langevin Algorithm (ULA) in combination with diffusion models can generate high quality MRI reconstructions with uncer

model-releasesarxiv-cs-cv
22 Apr 2026
Agents

Feasibility of Indoor Frame-Wise Lidar Semantic Segmentation via Distillation from Visual Foundation Model

DGX agent

arXiv:2604.18831v1 Announce Type: new Abstract: Frame-wise semantic segmentation of indoor lidar scans is a fundamental step toward higher-level 3D scene understanding and mapping applications. Howeve

agentsarxiv-cs-cv
22 Apr 2026
Research

Feature Perturbation Pool-based Fusion Network for Unified Multi-Class Industrial Defect Detection

DGX agent

arXiv:2604.19259v1 Announce Type: new Abstract: Multi-class defect detection constitutes a critical yet challenging task in industrial quality inspection, where existing approaches typically suffer fr

researcharxiv-cs-cv
22 Apr 2026
Model Releases

FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling

DGX agent

arXiv:2509.12052v3 Announce Type: replace Abstract: Current talking-head generation has gradually shifted from GAN-based methods to diffusion-based paradigms, achieving remarkable progress in visual f

model-releasesarxiv-cs-cv
22 Apr 2026
Safety

Framelet-Based Blind Image Restoration with Minimax Concave Regularization

DGX agent

arXiv:2604.19314v1 Announce Type: new Abstract: Recovering corrupted images is one of the most challenging problems in image processing. Among various restoration tasks, blind image deblurring has bee

safetyarxiv-cs-cv
22 Apr 2026
Local Ai

Generative Drifting for Conditional Medical Image Generation

DGX agent

arXiv:2604.19736v1 Announce Type: new Abstract: Conditional medical image generation plays an important role in many clinically relevant imaging tasks. However, existing methods still face a fundament

local-aiarxiv-cs-cv
22 Apr 2026
Research

Generative Texture Filtering

DGX agent

arXiv:2604.19039v1 Announce Type: new Abstract: We present a generative method for texture filtering, which exhibits surprisingly good performance and generalizability. Our core idea is to empower tex

researcharxiv-cs-cv
22 Apr 2026
Research

Geometry-Guided Self-Supervision for Ultra-Fine-Grained Recognition with Limited Data

DGX agent

arXiv:2604.19345v1 Announce Type: new Abstract: This paper investigates the intrinsic geometrical features of highly similar objects and introduces a general self-supervised framework called the Geome

researcharxiv-cs-cv
22 Apr 2026
Model Releases

GRAFT: Geometric Refinement and Fitting Transformer for Human Scene Reconstruction

DGX agent

arXiv:2604.19624v1 Announce Type: new Abstract: Reconstructing physically plausible 3D human-scene interactions (HSI) from a single image currently presents a trade-off: optimization based methods off

model-releasesarxiv-cs-cv
22 Apr 2026
Safety

Guiding Distribution Matching Distillation with Gradient-Based Reinforcement Learning

DGX agent

arXiv:2604.19009v1 Announce Type: cross Abstract: Diffusion distillation, exemplified by Distribution Matching Distillation (DMD), has shown great promise in few-step generation but often sacrifices q

safetyarxiv-cs-cv
22 Apr 2026
Model Releases

HarmoniDiff-RS: Training-Free Diffusion Harmonization for Satellite Image Composition

DGX agent

arXiv:2604.19392v1 Announce Type: new Abstract: Satellite image composition plays a critical role in remote sensing applications such as data augmentation, disaste simulation, and urban planning. We p

model-releasesarxiv-cs-cv
22 Apr 2026
Local Ai

HMR-Net: Hierarchical Modular Routing for Cross-Domain Object Detection in Aerial Images

DGX agent

arXiv:2604.18866v1 Announce Type: new Abstract: Despite advances in object detection, aerial imagery remains a challenging domain, as models often fail to generalize across variations in spatial resol

local-aiarxiv-cs-cv
22 Apr 2026
Model Releases

How Far Are Video Models from True Multimodal Reasoning?

DGX agent

arXiv:2604.19193v1 Announce Type: new Abstract: Despite remarkable progress toward general-purpose video models, a critical question remains unanswered: how far are these models from achieving true mu

model-releasesarxiv-cs-cv
22 Apr 2026
Applications

InHabit: Leveraging Image Foundation Models for Scalable 3D Human Placement

DGX agent

arXiv:2604.19673v1 Announce Type: new Abstract: Training embodied agents to understand 3D scenes as humans do requires large-scale data of people meaningfully interacting with diverse environments, ye

applicationsarxiv-cs-cv
22 Apr 2026
Research

IonMorphNet: Generalizable Learning of Ion Image Morphologies for Peak Picking in Mass Spectrometry Imaging

DGX agent

arXiv:2604.19369v1 Announce Type: new Abstract: Peak picking is a fundamental preprocessing step in Mass Spectrometry Imaging (MSI), where each sample is represented by hundreds to thousands of ion im

researcharxiv-cs-cv
22 Apr 2026
Tutorials

IR-Flow: Bridging Discriminative and Generative Image Restoration via Rectified Flow

DGX agent

arXiv:2604.19680v1 Announce Type: new Abstract: In image restoration, single-step discriminative mappings often lack fine details via expectation learning, whereas generative paradigms suffer from ine

tutorialsarxiv-cs-cv
22 Apr 2026
Research

KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering

DGX agent

arXiv:2601.11632v2 Announce Type: replace Abstract: Multi-modal Large Language Models (MLLMs) for Visual Question Answering (VQA) often suffer from dual limitations: knowledge hallucination and insuff

researcharxiv-cs-cv
22 Apr 2026
Safety

Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation

DGX agent

arXiv:2604.19234v1 Announce Type: new Abstract: Reinforcement learning, particularly Group Relative Policy Optimization (GRPO), has emerged as an effective framework for post-training visual generativ

safetyarxiv-cs-cv
22 Apr 2026
Model Releases

Lingua-SafetyBench: A Benchmark for Safety Evaluation of Multilingual Vision-Language Models

DGX agent

arXiv:2601.22737v2 Announce Type: replace Abstract: The robust safety of Vision-Language Large Models (VLLMs) against joint multilingual and multimodal threats remains severely underexplored. Current

model-releasesarxiv-cs-cv
22 Apr 2026
Local Ai

Localization-Guided Foreground Augmentation in Autonomous Driving

DGX agent

arXiv:2604.18940v1 Announce Type: new Abstract: Autonomous driving systems often degrade under adverse visibility conditions-such as rain, nighttime, or snow-where online scene geometry (e.g., lane di

local-aiarxiv-cs-cv
22 Apr 2026
Model Releases

LoViF 2026 Challenge on Real-World All-in-One Image Restoration: Methods and Results

DGX agent

arXiv:2604.19445v1 Announce Type: new Abstract: This paper presents a review for the LoViF Challenge on Real-World All-in-One Image Restoration. The challenge aimed to advance research on real-world a

model-releasesarxiv-cs-cv
22 Apr 2026
Agents

MAGICIAN: Efficient Long-Term Planning with Imagined Gaussians for Active Mapping

DGX agent

arXiv:2603.22650v2 Announce Type: replace Abstract: Active mapping aims to determine how an agent should move to efficiently reconstruct unknown environments. Most existing approaches rely on greedy n

agentsarxiv-cs-cv
22 Apr 2026
Research

Match-Any-Events: Zero-Shot Motion-Robust Feature Matching Across Wide Baselines for Event Cameras

DGX agent

arXiv:2604.18744v1 Announce Type: new Abstract: Event cameras have recently shown promising capabilities in instantaneous motion estimation due to their robustness to low light and fast motions. Howev

researcharxiv-cs-cv
22 Apr 2026
Research

MedFlowSeg: Flow Matching for Medical Image Segmentation with Frequency-Aware Attention

DGX agent

arXiv:2604.19675v1 Announce Type: new Abstract: Flow matching has recently emerged as a principled framework for learning continuous-time transport maps, enabling efficient deterministic generation wi

researcharxiv-cs-cv
22 Apr 2026
Local Ai

Memory Over Maps: 3D Object Localization Without Reconstruction

DGX agent

arXiv:2603.20530v2 Announce Type: replace-cross Abstract: Target localization is a prerequisite for embodied tasks such as navigation and manipulation. Conventional approaches rely on constructing exp

local-aiarxiv-cs-cv
22 Apr 2026
Safety

Mind2Drive: Predicting Driver Intentions from EEG in Real-world On-Road Driving

DGX agent

arXiv:2604.19368v1 Announce Type: new Abstract: Predicting driver intention from neurophysiological signals offers a promising pathway for enhancing proactive safety in advanced driver assistance syst

safetyarxiv-cs-cv
22 Apr 2026
Research

MiTA Attention: Efficient Fast-Weight Scaling via a Mixture of Top-k Activations

DGX agent

arXiv:2602.01219v4 Announce Type: replace-cross Abstract: The attention operator in Transformers can be viewed as a two-layer fast-weight MLP, whose weights are dynamically instantiated from input tok

researcharxiv-cs-cv
22 Apr 2026
Safety

Mitigating Long-Tail Bias via Prompt-Controlled Diffusion Augmentation

DGX agent

arXiv:2602.04749v2 Announce Type: replace Abstract: Long-tailed class imbalance remains a fundamental obstacle in semantic segmentation of high-resolution remote-sensing imagery, where dominant classe

safetyarxiv-cs-cv
22 Apr 2026
Safety

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation

DGX agent

arXiv:2604.19679v1 Announce Type: new Abstract: Recent advances in Diffusion Transformers (DiTs) have enabled high-quality joint audio-video generation, producing videos with synchronized audio within

safetyarxiv-cs-cv
22 Apr 2026
Research

MoonSeg3R: Monocular Online Zero-Shot Segment Anything in 3D with Reconstructive Foundation Priors

DGX agent

arXiv:2512.15577v2 Announce Type: replace Abstract: In this paper, we focus on online zero-shot monocular 3D instance segmentation, a novel practical setting where existing approaches fail to perform

researcharxiv-cs-cv
22 Apr 2026
Safety

MOSA: Motion-Guided Semantic Alignment for Dynamic Scene Graph Generation

DGX agent

arXiv:2604.19631v1 Announce Type: new Abstract: Dynamic Scene Graph Generation (DSGG) aims to structurally model objects and their dynamic interactions in video sequences for high-level semantic under

safetyarxiv-cs-cv
22 Apr 2026
Model Releases

MSDS: Deep Structural Similarity with Multiscale Representation

DGX agent

arXiv:2604.19159v1 Announce Type: new Abstract: Deep-feature-based perceptual similarity models have demonstrated strong alignment with human visual perception in Image Quality Assessment (IQA). Howev

model-releasesarxiv-cs-cv
22 Apr 2026
← Previous
1…222223224225226…261
Next →