AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
14 Apr 2026

AIM-Bench: Benchmarking and Improving Affective Image Manipulation via Fine-Grained Hierarchical Control

Model ReleasesDGX agent

arXiv:2604.10454v1 Announce Type: new Abstract: Affective Image Manipulation (AIM) aims to evoke specific emotions through targeted editing. Current image editing benchmarks primarily focus on object-

Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions

ResearchDGX agent

arXiv:2604.11730v1 Announce Type: new Abstract: Using behavioural science, health interventions focus on behaviour change by providing a framework to help patients acquire and maintain healthy habits

AmodalSVG: Amodal Image Vectorization via Semantic Layer Peeling

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.10940v1 Announce Type: new Abstract: We introduce AmodalSVG, a new framework for amodal image vectorization that produces semantically organized and geometrically complete SVG representatio

Analytical Modeling and Correction of Distance Error in Homography-Based Ground-Plane Mapping

ResearchDGX agent

arXiv:2604.10805v1 Announce Type: new Abstract: Accurate distance estimation from monocular cameras is essential for intelligent monitoring systems. In many deployments, image coordinates are mapped t

Anatomy-Informed Deep Learning for Abdominal Aortic Aneurysm Segmentation

ResearchDGX agent

arXiv:2604.10312v1 Announce Type: new Abstract: In CT angiography, the accurate segmentation of abdominal aortic aneurysms (AAAs) is difficult due to large anatomical variability, low-contrast vessel

Any 3D Scene is Worth 1K Tokens: 3D-Grounded Representation for Scene Generation at Scale

ResearchDGX agent

arXiv:2604.11331v1 Announce Type: new Abstract: 3D scene generation has long been dominated by 2D multi-view or video diffusion models. This is due not only to the lack of scene-level 3D latent repres

Are Pretrained Image Matchers Good Enough for SAR-Optical Satellite Registration?

ResearchDGX agent

arXiv:2604.10217v1 Announce Type: new Abstract: Cross-modal optical-SAR (Synthetic Aperture Radar) registration is a bottleneck for disaster-response via remote sensing, yet modern image matchers are

Are We Recognizing the Jaguar or Its Background? A Diagnostic Framework for Jaguar Re-Identification

Model ReleasesDGX agent

arXiv:2604.09690v1 Announce Type: new Abstract: Jaguar re-identification (re-ID) from citizen-science imagery can look strong on standard retrieval metrics while still relying on the wrong evidence, s

ArtiCAD: Articulated CAD Assembly Design via Multi-Agent Code Generation

AgentsDGX agent

arXiv:2604.10992v1 Announce Type: new Abstract: Parametric Computer-Aided Design (CAD) of articulated assemblies is essential for product development, yet generating these multi-part, movable models f

At FullTilt: Real-Time Open-Set 3D Macromolecule Detection Directly from Tilted 2D Projections

Local AiDGX agent

arXiv:2604.10766v1 Announce Type: new Abstract: Open-set 3D macromolecule detection in cryogenic electron tomography eliminates the need for target-specific model retraining. However, strict VRAM cons

Attention-Guided Dual-Stream Learning for Group Engagement Recognition: Fusing Transformer-Encoded Motion Dynamics with Scene Context via Adaptive Gating

ResearchDGX agent

arXiv:2604.10078v1 Announce Type: new Abstract: Student engagement is crucial for improving learning outcomes in group activities. Highly engaged students perform better both individually and contribu

Automatic Uncertainty-Aware Synthetic Data Bootstrapping for Historical Map Segmentation

ApplicationsDGX agent

arXiv:2511.15875v2 Announce Type: replace Abstract: The automated analysis of historical documents, particularly maps, has drastically benefited from advances in deep learning and its success across v

Autonomous Diffractometry Enabled by Visual Reinforcement Learning

SafetyDGX agent

arXiv:2604.11773v1 Announce Type: cross Abstract: Automation underpins progress across scientific and industrial disciplines. Yet, automating tasks requiring interpretation of abstract visual informat

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs

Model ReleasesDGX agent

arXiv:2604.10528v1 Announce Type: new Abstract: While Vision-Language Models (VLMs) demonstrate remarkable zero-shot recognition capabilities across a diverse spectrum of multimodal tasks, it yet rema

BEM: Training-Free Background Embedding Memory for False-Positive Suppression in Real-Time Fixed-Background Camera

ApplicationsDGX agent

arXiv:2604.11714v1 Announce Type: new Abstract: Pretrained detectors perform well on benchmarks but often suffer performance degradation in real-world deployments due to distribution gaps between trai

Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation

Model ReleasesDGX agent

arXiv:2504.14988v3 Announce Type: replace Abstract: Recent advancements in Large Vision-Language Models (LVLMs) have demonstrated remarkable multimodal perception capabilities, garnering significant a

Benchmarking Vision-Language Models under Contradictory Virtual Content Attacks in Augmented Reality

Model ReleasesDGX agent

arXiv:2604.05510v2 Announce Type: replace Abstract: Augmented reality (AR) has rapidly expanded over the past decade. As AR becomes increasingly integrated into daily life, its security and reliabilit

Beyond Model Design: Data-Centric Training and Self-Ensemble for Gaussian Color Image Denoising

Local AiDGX agent

arXiv:2604.11468v1 Announce Type: new Abstract: This paper presents our solution to the NTIRE 2026 Image Denoising Challenge (Gaussian color image denoising at fixed noise level sigma = 50). Rather th

Beyond Reconstruction: Reconstruction-to-Vector Diffusion for Hyperspectral Anomaly Detection

SafetyDGX agent

arXiv:2604.11390v1 Announce Type: new Abstract: While Hyperspectral Anomaly Detection (HAD) excels at identifying sparse targets in complex scenes, existing models remain trapped in a scalar 'reconstr

Bidirectional Cross-Attention Fusion of High-Res RGB and Low-Res HSI for Multimodal Automated Waste Sorting

Model ReleasesDGX agent

arXiv:2603.13941v2 Announce Type: replace Abstract: Growing waste streams and the transition to a circular economy require efficient automated waste sorting. In industrial settings, materials move on

Bidirectional Learning of Facial Action Units and Expressions via Structured Semantic Mapping across Heterogeneous Datasets

SafetyDGX agent

arXiv:2604.10541v1 Announce Type: new Abstract: Facial action unit (AU) detection and facial expression (FE) recognition can be jointly viewed as affective facial behavior tasks, representing fine-gra

Biomarker-Based Pretraining for Chagas Disease Screening in Electrocardiograms

ResearchDGX agent

arXiv:2604.09782v1 Announce Type: new Abstract: Chagas disease screening via ECGs is limited by scarce and noisy labels in existing datasets. We propose a biomarker-based pretraining approach, where a

BLPR: Robust License Plate Recognition under Viewpoint and Illumination Variations via Confidence-Driven VLM Fallback

ApplicationsDGX agent

arXiv:2604.09927v1 Announce Type: new Abstract: Robust license plate recognition in unconstrained environments remains a significant challenge, particularly in underrepresented regions with limited da

Bootstrapping Video Semantic Segmentation Model via Distillation-assisted Test-Time Adaptation

ApplicationsDGX agent

arXiv:2604.10950v1 Announce Type: new Abstract: Fully supervised Video Semantic Segmentation (VSS) relies heavily on densely annotated video data, limiting practical applicability. Alternatively, appl

Boxes2Pixels: Learning Defect Segmentation from Noisy SAM Masks

Model ReleasesDGX agent

arXiv:2604.11162v1 Announce Type: new Abstract: Accurate defect segmentation is critical for industrial inspection, yet dense pixel-level annotations are rarely available. A common workaround is to co

Brain-Grasp: Graph-based Saliency Priors for Improved fMRI-based Visual Brain Decoding

SafetyDGX agent

arXiv:2604.10617v1 Announce Type: cross Abstract: Recent progress in brain-guided image generation has improved the quality of fMRI-based reconstructions; however, fundamental challenges remain in pre

Bridging the RGB-IR Gap: Consensus and Discrepancy Modeling for Text-Guided Multispectral Detection

SafetyDGX agent

arXiv:2604.11234v1 Announce Type: new Abstract: Text-guided multispectral object detection uses text semantics to guide semantic-aware cross-modal interaction between RGB and IR for more robust percep

Byte-level generative predictions for forensics multimedia carving

ResearchDGX agent

arXiv:2604.11010v1 Announce Type: new Abstract: Digital forensic investigations often face significant challenges when recovering fragmented multimedia files that lack file system metadata. While trad

Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?

Model ReleasesDGX agent

arXiv:2511.21998v2 Announce Type: replace Abstract: Multi-modal Large Language Models (LLM) have advanced conversational abilities but struggle with providing live, interactive step-by-step guidance,

Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification?

ApplicationsDGX agent

arXiv:2601.06993v2 Announce Type: replace Abstract: Multi-modal large language models (MLLMs) exhibit strong general-purpose capabilities, yet still struggle on Fine-Grained Visual Classification (FGV

Catalyst: Out-of-Distribution Detection via Elastic Scaling

ResearchDGX agent

arXiv:2602.02409v2 Announce Type: replace Abstract: Out-of-distribution (OOD) detection is critical for the safe deployment of deep neural networks. State-of-the-art post-hoc methods typically derive

CDPR: Cross-modal Diffusion with Polarization for Reliable Monocular Depth Estimation

ApplicationsDGX agent

arXiv:2604.11097v1 Announce Type: new Abstract: Monocular depth estimation is a fundamental yet challenging task in computer vision, especially under complex conditions such as textureless surfaces, t

CityGuard: Graph-Aware Private Descriptors for Bias-Resilient Identity Search Across Urban Cameras

SafetyDGX agent

arXiv:2602.18047v3 Announce Type: replace Abstract: City-scale person re-identification across distributed cameras must handle severe appearance changes from viewpoint, occlusion, and domain shift whi

CoFusion: Multispectral and Hyperspectral Image Fusion via Spectral Coordinate Attention

Model ReleasesDGX agent

arXiv:2604.10584v1 Announce Type: new Abstract: Multispectral and Hyperspectral Image Fusion (MHIF) aims to reconstruct high-resolution images by integrating low-resolution hyperspectral images (LRHSI

Compact single-shot ranging and near-far imaging using metasurfaces

ResearchDGX agent

arXiv:2604.10037v1 Announce Type: cross Abstract: We present a metasurface imaging system capable of simultaneously capturing two images at close range (1-2~cm) and an additional image at long range (

ComSim: Building Scalable Real-World Robot Data Generation via Compositional Simulation

SafetyDGX agent

arXiv:2604.11386v1 Announce Type: cross Abstract: Recent advancements in foundational models, such as large language models and world models, have greatly enhanced the capabilities of robotics, enabli

ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectors

ResearchDGX agent

arXiv:2512.09056v3 Announce Type: replace Abstract: Object pose estimation is a fundamental task in computer vision and robotics, yet most methods require extensive, dataset-specific training. Concurr

Context-Aware Semantic Segmentation via Stage-Wise Attention

Model ReleasesDGX agent

arXiv:2601.11310v2 Announce Type: replace Abstract: Semantic ultra-high-resolution (UHR) image segmentation is essential in remote sensing applications such as aerial mapping and environmental monitor

Context Matters: Vision-Based Depression Detection Comparing Classical and Deep Approaches

SafetyDGX agent

arXiv:2604.10344v1 Announce Type: new Abstract: The classical approach to detecting depression from vision emphasizes interpretable features, such as facial expression, and classifiers such as the Sup

Continuous Adversarial Flow Models

TutorialsDGX agent

arXiv:2604.11521v1 Announce Type: cross Abstract: We propose continuous adversarial flow models, a type of continuous-time flow model trained with an adversarial objective. Unlike flow matching, which

Contour Refinement using Discrete Diffusion in Low Data Regime

SafetyDGX agent

arXiv:2602.05880v2 Announce Type: replace Abstract: Boundary detection of irregular and translucent objects is an important problem with applications in medical imaging, environmental monitoring and m

ConvFormer3D-TAP: Phase/Uncertainty-Aware Front-End Fusion for Cine CMR View Classification Pipelines

ResearchDGX agent

arXiv:2604.11389v1 Announce Type: new Abstract: Reliable recognition of standard cine cardiac MRI views is essential because each view determines which cardiac anatomy is visualized and which quantita

CoPS: Conditional Prompt Synthesis for Zero-Shot Anomaly Detection

SafetyDGX agent

arXiv:2508.03447v2 Announce Type: replace Abstract: Recently, large pre-trained vision-language models have shown remarkable performance in zero-shot anomaly detection (ZSAD). With fine-tuning on a si

Counting to Four is still a Chore for VLMs

Model ReleasesDGX agent

arXiv:2604.10039v1 Announce Type: new Abstract: Vision--language models (VLMs) have achieved impressive performance on complex multimodal reasoning tasks, yet they still fail on simple grounding skill

CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance

AgentsDGX agent

arXiv:2508.16644v4 Announce Type: replace Abstract: Diffusion models excel at photorealistic synthesis but struggle with precise object counts, especially in high-density settings. We introduce COUNTL

CPAM: Context-Preserving Adaptive Manipulation for Zero-Shot Real Image Editing

Model ReleasesDGX agent

arXiv:2506.18438v2 Announce Type: replace Abstract: Editing natural images using textual descriptions in text-to-image diffusion models remains a significant challenge, particularly in achieving consi

CraftGraffiti: Exploring Human Identity with Custom Graffiti Art via Facial-Preserving Diffusion Models

ApplicationsDGX agent

arXiv:2508.20640v2 Announce Type: replace Abstract: Preserving facial identity under extreme stylistic transformation remains a major challenge in generative art. In graffiti, a high-contrast, abstrac

CylinderDepth: Cylindrical Spatial Attention for Multi-View Consistent Self-Supervised Surround Depth Estimation

ResearchDGX agent

arXiv:2511.16428v3 Announce Type: replace Abstract: Self-supervised surround-view depth estimation enables dense, low-cost 3D perception with a 360{eg} field of view from multiple minimally overlappin

Dark-EvGS: Event Camera as an Eye for Radiance Field in the Dark

ResearchDGX agent

arXiv:2507.11931v2 Announce Type: replace Abstract: In low-light environments, conventional cameras often struggle to capture clear multi-view images of objects due to dynamic range limitations and mo

Data-Driven Automated Identification of Optimal Feature-Representative Images in Infrared Thermography Using Statistical and Morphological Metrics

Local AiDGX agent

arXiv:2604.09728v1 Announce Type: new Abstract: Infrared thermography (IRT) is a widely used non-destructive testing technique for detecting structural features such as subsurface defects. However, mo

Data-Efficient Semantic Segmentation of 3D Point Clouds via Open-Vocabulary Image Segmentation-based Pseudo-Labeling

Model ReleasesDGX agent

arXiv:2604.11007v1 Announce Type: new Abstract: Semantic segmentation of 3D point cloud scenes is a crucial task for various applications. In real-world scenarios, training segmentation models often f

DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning

SafetyDGX agent

arXiv:2503.11892v3 Announce Type: replace Abstract: Multimodal representation learning aims to capture both shared and complementary semantic information across multiple modalities. However, the intri

Decoding Ancient Oracle Bone Script via Generative Dictionary Retrieval

ResearchDGX agent

arXiv:2604.09668v1 Announce Type: cross Abstract: Understanding humanity's earliest writing systems is crucial for reconstructing civilization's origins, yet many ancient scripts remain undeciphered.

Decompose, Mix, Adapt: A Unified Framework for Parameter-Efficient Neural Network Recombination and Compression

Model ReleasesDGX agent

arXiv:2603.27383v2 Announce Type: replace Abstract: Parameter Recombination (PR) methods aim to efficiently compose the weights of a neural network for applications like Parameter-Efficient FineTuning

Decoupled Generative Modeling for Human-Object Interaction Synthesis

ResearchDGX agent

arXiv:2512.19049v2 Announce Type: replace Abstract: Synthesizing realistic human-object interaction (HOI) is essential for 3D computer vision and robotics, underpinning animation and embodied control.

Decoupled Similarity for Task-Aware Token Pruning in Large Vision-Language Models

ResearchDGX agent

arXiv:2604.11240v1 Announce Type: new Abstract: Token pruning has emerged as an effective approach to reduce the substantial computational overhead of Large Vision-Language Models (LVLMs) by discardin

DeepShapeMatchingKit: Accelerated Functional Map Solver and Shape Matching Pipelines Revisited

ResearchDGX agent

arXiv:2604.10377v1 Announce Type: new Abstract: Deep functional maps, leveraging learned feature extractors and spectral correspondence solvers, are fundamental to non-rigid 3D shape matching. Based o

DeepSketcher: Internalizing Visual Manipulation for Multimodal Reasoning

ResearchDGX agent

arXiv:2509.25866v2 Announce Type: replace Abstract: The 'thinking with images' paradigm represents a pivotal shift in the reasoning of Vision Language Models (VLMs), moving from text-dominant chain-of

Defending against Patch-Based and Texture-Based Adversarial Attacks with Spectral Decomposition

Local AiDGX agent

arXiv:2604.10715v1 Announce Type: new Abstract: Adversarial examples present significant challenges to the security of Deep Neural Network (DNN) applications. Specifically, there are patch-based and t

Degradation-Aware and Structure-Preserving Diffusion for Real-World Image Super-Resolution

Local AiDGX agent

arXiv:2604.11470v1 Announce Type: new Abstract: Real-world image super-resolution is particularly challenging for diffusion models because real degradations are complex, heterogeneous, and rarely mode

← Previous
1…194195196197198…207
Next →