AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
22 Apr 2026

AlignedCut: Visual Concepts Discovery on Brain-Guided Universal Feature Space

SafetyDGX agent

arXiv:2406.18344v2 Announce Type: replace Abstract: We study the intriguing connection between visual data, deep networks, and the brain. Our method creates a universal channel alignment by using brai

Allo{SR}^2: Rectifying One-Step Super-Resolution to Stay Real via Allomorphic Generative Flows

SafetyDGX agent

arXiv:2604.19238v1 Announce Type: new Abstract: Real-world image super-resolution (Real-SR) has been revolutionized by leveraging the powerful generative priors of large-scale diffusion and flow-based

An Object-Centered Data Acquisition Method for 3D Gaussian Splatting using Mobile Phones

Local AiDGX agent

arXiv:2604.19216v1 Announce Type: new Abstract: Data acquisition through mobile phones remains a challenge for 3D Gaussian Splatting (3DGS). In this work we target the object-centered scenario and ena


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model

Model ReleasesDGX agent

arXiv:2604.19747v1 Announce Type: new Abstract: Sparse-view 3D reconstruction is essential for modeling scenes from casual captures, but remain challenging for non-generative reconstruction. Existing

Attend what matters: Leveraging vision foundational models for breast cancer classification using mammograms

Local AiDGX agent

arXiv:2604.19350v1 Announce Type: new Abstract: Vision Transformers (exttt{ViT}) have become the architecture of choice for many computer vision tasks, yet their performance in computer-aided diagnost

Autonomous Skeletal Landmark Localization towards Agentic C-Arm Control

Local AiDGX agent

arXiv:2604.18740v1 Announce Type: new Abstract: Purpose: Automated C-arm positioning ensures timely treatment in patients requiring emergent interventions. When a conventional Deep Learning (DL) appro

BALTIC: A Benchmark and Cross-Domain Strategy for 3D Reconstruction Across Air and Underwater Domains Under Varying Illumination

Model ReleasesDGX agent

arXiv:2604.19133v1 Announce Type: new Abstract: Robust 3D reconstruction across varying environmental conditions remains a critical challenge for robotic perception, particularly when transitioning be

Benchmarking Vision Foundation Models for Domain-Generalizable Face Anti-Spoofing

ResearchDGX agent

arXiv:2604.19196v1 Announce Type: new Abstract: Face Anti-Spoofing (FAS) remains challenging due to the requirement for robust domain generalization across unseen environments. While recent trends lev

Bridging Foundation Models and ASTM Metallurgical Standards for Automated Grain Size Estimation from Microscopy Images

Model ReleasesDGX agent

arXiv:2604.18957v1 Announce Type: new Abstract: Extracting standardized metallurgical metrics from microscopy images remains challenging due to complex grain morphology and the data demands of supervi

Bridging Semantics and Geometry: A Decoupled LVLM-SAM Framework for Reasoning Segmentation in Optical Remote Sensing

SafetyDGX agent

arXiv:2512.19302v2 Announce Type: replace Abstract: Large Vision--Language Models (LVLMs) hold great promise for advancing optical remote sensing (RS) analysis, yet existing reasoning segmentation fra

CAHAL: Clinically Applicable resolution enHAncement for Low-resolution MRI scans

SafetyDGX agent

arXiv:2604.18781v1 Announce Type: new Abstract: Large-scale automated morphometric analysis of brain MRI is limited by the thick-slice, anisotropic acquisitions prevalent in routine clinical practice.

Can We Build Scene Graphs, Not Classify Them? FlowSG: Progressive Image-Conditioned Scene Graph Generation with Flow Matching

Local AiDGX agent

arXiv:2604.18623v1 Announce Type: new Abstract: Scene Graph Generation (SGG) unifies object localization and visual relationship reasoning by predicting boxes and subject-predicate-object triples. Yet

Centralized Copy-Paste: Enhanced Data Augmentation Strategy for Wildland Fire Semantic Segmentation

ResearchDGX agent

arXiv:2507.06321v2 Announce Type: replace Abstract: Collecting and annotating images for the purpose of training segmentation models is often cost prohibitive. In the domain of wildland fire science,

CityRAG: Stepping Into a City via Spatially-Grounded Video Generation

AgentsDGX agent

arXiv:2604.19741v1 Announce Type: new Abstract: We address the problem of generating a 3D-consistent, navigable environment that is spatially grounded: a simulation of a real location. Existing video

CLIPoint3D: Language-Grounded Few-Shot Unsupervised 3D Point Cloud Domain Adaptation

Model ReleasesDGX agent

arXiv:2602.20409v2 Announce Type: replace Abstract: Recent vision-language models (VLMs) such as CLIP demonstrate impressive cross-modal reasoning, extending beyond images to 3D perception. Yet, these

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation

Model ReleasesDGX agent

arXiv:2604.19636v1 Announce Type: new Abstract: Synthesizing human--object interaction (HOI) videos has broad practical value in e-commerce, digital advertising, and virtual marketing. However, curren

Colour Extraction Pipeline for Odonates using Computer Vision

ResearchDGX agent

arXiv:2604.18725v1 Announce Type: new Abstract: The correlation between insect morphological traits and climate has been documented in physiological studies, but such studies remain limited by the tim

Concept Inconsistency in Dermoscopic Concept Bottleneck Models: A Rough-Set Analysis of the Derm7pt Dataset

Model ReleasesDGX agent

arXiv:2604.19323v1 Announce Type: cross Abstract: Concept Bottleneck Models (CBMs) route predictions exclusively through a clinically grounded concept layer, binding interpretability to concept-label

ConvVitMamba: Efficient Multiscale Convolution, Transformer, and Mamba-Based Sequence modelling for Hyperspectral Image Classification

Model ReleasesDGX agent

arXiv:2604.18856v1 Announce Type: new Abstract: Hyperspectral image (HSI) classification remains challenging due to high spectral dimensionality, redundancy, and limited labeled data. Although convolu

CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers

SafetyDGX agent

arXiv:2604.19632v1 Announce Type: new Abstract: Graphic design images consist of multiple editable layers, such as text, background, and decorative elements, while most generative models produce raste

CrossPan: A Comprehensive Benchmark for Cross-Sequence Pancreas MRI Segmentation and Generalization

Model ReleasesDGX agent

arXiv:2604.18797v1 Announce Type: new Abstract: Automatic pancreas segmentation is fundamental to abdominal MRI analysis, yet deep learning models trained on one MRI sequence often fail catastrophical

Data Leakage Detection and De-duplication in Large Scale Geospatial Image Datasets

ResearchDGX agent

arXiv:2304.02296v2 Announce Type: replace Abstract: In our study, we conducted a comprehensive analysis of three widely used datasets in the domain of building footprint extraction using deep neural n

DDF2Pol: A Dual-Domain Feature Fusion Network for PolSAR Image Classification

Model ReleasesDGX agent

arXiv:2604.18853v1 Announce Type: new Abstract: This paper presents DDF2Pol, a lightweight dual-domain convolutional neural network for PolSAR image classification. The proposed architecture integrate

Deep sprite-based image models: An analysis

Model ReleasesDGX agent

arXiv:2604.19480v1 Announce Type: new Abstract: While foundation models drive steady progress in image segmentation and diffusion algorithms compose always more realistic images, the seemingly simple

DeltaSeg: Tiered Attention and Deep Delta Learning for Multi-Class Structural Defect Segmentation

ResearchDGX agent

arXiv:2604.18745v1 Announce Type: new Abstract: Automated segmentation of structural defects from visual inspection imagery remains challenging due to the diversity of damage types, extreme class imba

Denoising, Fast and Slow: Difficulty-Aware Adaptive Sampling for Image Generation

SafetyDGX agent

arXiv:2604.19141v1 Announce Type: new Abstract: Diffusion- and flow-based models usually allocate compute uniformly across space, updating all patches with the same timestep and number of function eva

Detection of T-shirt Presentation Attacks in Face Recognition Systems

Model ReleasesDGX agent

arXiv:2604.19365v1 Announce Type: new Abstract: Face recognition systems are often used for biometric authentication. Nevertheless, it is known that without any protective measures, face recognition s

Diff-SBSR: Learning Multimodal Feature-Enhanced Diffusion Models for Zero-Shot Sketch-Based 3D Shape Retrieval

SafetyDGX agent

arXiv:2604.19135v1 Announce Type: new Abstract: This paper presents the first exploration of text-to-image diffusion models for zero-shot sketch-based 3D shape retrieval (ZS-SBSR). Existing sketch-bas

DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval

SafetyDGX agent

arXiv:2604.19432v1 Announce Type: new Abstract: Vision foundation models have shown great promise for open-set 3D object retrieval (3DOR) through efficient adaptation to multi-view images. Leveraging

Divide-and-Conquer Approach to Holistic Cognition in High-Similarity Contexts with Limited Data

Local AiDGX agent

arXiv:2604.19339v1 Announce Type: new Abstract: Ultra-fine-grained visual categorization (Ultra-FGVC) aims to classify highly similar subcategories within fine-grained objects using limited training s

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents

AgentsDGX agent

arXiv:2604.19264v1 Announce Type: new Abstract: Agentic multimodal models have garnered significant attention for their ability to leverage external tools to tackle complex tasks. However, it is obser

DUALVISION: RGB-Infrared Multimodal Large Language Models for Robust Visual Reasoning

Model ReleasesDGX agent

arXiv:2604.18829v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved impressive performance on visual perception and reasoning tasks with RGB imagery, yet they remain

EfficientPENet: Real-Time Depth Completion from Sparse LiDAR via Lightweight Multi-Modal Fusion

Model ReleasesDGX agent

arXiv:2604.18790v1 Announce Type: new Abstract: Depth completion from sparse LiDAR measurements and corresponding RGB images is a prerequisite for accurate 3D perception in robotic systems. Existing m

EgoMotion: Hierarchical Reasoning and Diffusion for Egocentric Vision-Language Motion Generation

ResearchDGX agent

arXiv:2604.19105v1 Announce Type: new Abstract: Faithfully modeling human behavior in dynamic environments is a foundational challenge for embodied intelligence. While conditional motion synthesis has

Ensembling Pruned Attention Heads For Uncertainty-Aware Efficient Transformers

SafetyDGX agent

arXiv:2510.18358v2 Announce Type: replace-cross Abstract: Uncertainty quantification (UQ) is essential for deploying deep neural networks in safety-critical settings. Although methods like Deep Ensemb

EPS: Efficient Patch Sampling for Video Overfitting in Deep Super-Resolution Model Training

ResearchDGX agent

arXiv:2411.16312v2 Announce Type: replace Abstract: Leveraging the overfitting property of deep neural networks (DNNs) is trending in video delivery systems to enhance video quality within bandwidth l

Evaluating Histogram Matching for Robust Deep learning-Based Grapevine Disease Detection

ApplicationsDGX agent

arXiv:2604.19510v1 Announce Type: new Abstract: Variability in illumination is a primary factor limiting deep learning robustness for field-based plant disease detection. This study evaluates Histogra

Evaluation of Winning Solutions of 2025 Low Power Computer Vision Challenge

ResearchDGX agent

arXiv:2604.19054v1 Announce Type: new Abstract: The IEEE Low-Power Computer Vision Challenge (LPCVC) aims to promote the development of efficient vision models for edge devices, balancing accuracy wit

Explore Like Humans: Autonomous Exploration with Online SG-Memo Construction for Embodied Agents

AgentsDGX agent

arXiv:2604.19034v1 Announce Type: new Abstract: Constructing structured spatial memory is essential for enabling long-horizon reasoning in complex embodied navigation tasks. Current memory constructio

Face Anything: 4D Face Reconstruction from Any Image Sequence

ResearchDGX agent

arXiv:2604.19702v1 Announce Type: new Abstract: Accurate reconstruction and tracking of dynamic human faces from image sequences is challenging because non-rigid deformations, expression changes, and

Fast and Robust Diffusion Posterior Sampling for MR Image Reconstruction Using the Preconditioned Unadjusted Langevin Algorithm

Model ReleasesDGX agent

arXiv:2512.05791v2 Announce Type: replace-cross Abstract: Purpose: The Unadjusted Langevin Algorithm (ULA) in combination with diffusion models can generate high quality MRI reconstructions with uncer

Feasibility of Indoor Frame-Wise Lidar Semantic Segmentation via Distillation from Visual Foundation Model

AgentsDGX agent

arXiv:2604.18831v1 Announce Type: new Abstract: Frame-wise semantic segmentation of indoor lidar scans is a fundamental step toward higher-level 3D scene understanding and mapping applications. Howeve

Feature Perturbation Pool-based Fusion Network for Unified Multi-Class Industrial Defect Detection

ResearchDGX agent

arXiv:2604.19259v1 Announce Type: new Abstract: Multi-class defect detection constitutes a critical yet challenging task in industrial quality inspection, where existing approaches typically suffer fr

FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling

Model ReleasesDGX agent

arXiv:2509.12052v3 Announce Type: replace Abstract: Current talking-head generation has gradually shifted from GAN-based methods to diffusion-based paradigms, achieving remarkable progress in visual f

Framelet-Based Blind Image Restoration with Minimax Concave Regularization

SafetyDGX agent

arXiv:2604.19314v1 Announce Type: new Abstract: Recovering corrupted images is one of the most challenging problems in image processing. Among various restoration tasks, blind image deblurring has bee

Generative Drifting for Conditional Medical Image Generation

Local AiDGX agent

arXiv:2604.19736v1 Announce Type: new Abstract: Conditional medical image generation plays an important role in many clinically relevant imaging tasks. However, existing methods still face a fundament

Generative Texture Filtering

ResearchDGX agent

arXiv:2604.19039v1 Announce Type: new Abstract: We present a generative method for texture filtering, which exhibits surprisingly good performance and generalizability. Our core idea is to empower tex

Geometry-Guided Self-Supervision for Ultra-Fine-Grained Recognition with Limited Data

ResearchDGX agent

arXiv:2604.19345v1 Announce Type: new Abstract: This paper investigates the intrinsic geometrical features of highly similar objects and introduces a general self-supervised framework called the Geome

GRAFT: Geometric Refinement and Fitting Transformer for Human Scene Reconstruction

Model ReleasesDGX agent

arXiv:2604.19624v1 Announce Type: new Abstract: Reconstructing physically plausible 3D human-scene interactions (HSI) from a single image currently presents a trade-off: optimization based methods off

Guiding Distribution Matching Distillation with Gradient-Based Reinforcement Learning

SafetyDGX agent

arXiv:2604.19009v1 Announce Type: cross Abstract: Diffusion distillation, exemplified by Distribution Matching Distillation (DMD), has shown great promise in few-step generation but often sacrifices q

HarmoniDiff-RS: Training-Free Diffusion Harmonization for Satellite Image Composition

Model ReleasesDGX agent

arXiv:2604.19392v1 Announce Type: new Abstract: Satellite image composition plays a critical role in remote sensing applications such as data augmentation, disaste simulation, and urban planning. We p

HMR-Net: Hierarchical Modular Routing for Cross-Domain Object Detection in Aerial Images

Local AiDGX agent

arXiv:2604.18866v1 Announce Type: new Abstract: Despite advances in object detection, aerial imagery remains a challenging domain, as models often fail to generalize across variations in spatial resol

How Far Are Video Models from True Multimodal Reasoning?

Model ReleasesDGX agent

arXiv:2604.19193v1 Announce Type: new Abstract: Despite remarkable progress toward general-purpose video models, a critical question remains unanswered: how far are these models from achieving true mu

InHabit: Leveraging Image Foundation Models for Scalable 3D Human Placement

ApplicationsDGX agent

arXiv:2604.19673v1 Announce Type: new Abstract: Training embodied agents to understand 3D scenes as humans do requires large-scale data of people meaningfully interacting with diverse environments, ye

IonMorphNet: Generalizable Learning of Ion Image Morphologies for Peak Picking in Mass Spectrometry Imaging

ResearchDGX agent

arXiv:2604.19369v1 Announce Type: new Abstract: Peak picking is a fundamental preprocessing step in Mass Spectrometry Imaging (MSI), where each sample is represented by hundreds to thousands of ion im

IR-Flow: Bridging Discriminative and Generative Image Restoration via Rectified Flow

TutorialsDGX agent

arXiv:2604.19680v1 Announce Type: new Abstract: In image restoration, single-step discriminative mappings often lack fine details via expectation learning, whereas generative paradigms suffer from ine

KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering

ResearchDGX agent

arXiv:2601.11632v2 Announce Type: replace Abstract: Multi-modal Large Language Models (MLLMs) for Visual Question Answering (VQA) often suffer from dual limitations: knowledge hallucination and insuff

Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation

SafetyDGX agent

arXiv:2604.19234v1 Announce Type: new Abstract: Reinforcement learning, particularly Group Relative Policy Optimization (GRPO), has emerged as an effective framework for post-training visual generativ

Lingua-SafetyBench: A Benchmark for Safety Evaluation of Multilingual Vision-Language Models

Model ReleasesDGX agent

arXiv:2601.22737v2 Announce Type: replace Abstract: The robust safety of Vision-Language Large Models (VLLMs) against joint multilingual and multimodal threats remains severely underexplored. Current

Localization-Guided Foreground Augmentation in Autonomous Driving

Local AiDGX agent

arXiv:2604.18940v1 Announce Type: new Abstract: Autonomous driving systems often degrade under adverse visibility conditions-such as rain, nighttime, or snow-where online scene geometry (e.g., lane di

← Previous
1…177178179180181…209
Next →