AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
Safety

Autonomous Diffractometry Enabled by Visual Reinforcement Learning

DGX agent

arXiv:2604.11773v1 Announce Type: cross Abstract: Automation underpins progress across scientific and industrial disciplines. Yet, automating tasks requiring interpretation of abstract visual informat

safetyarxiv-cs-cv
14 Apr 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs

DGX agent

arXiv:2604.10528v1 Announce Type: new Abstract: While Vision-Language Models (VLMs) demonstrate remarkable zero-shot recognition capabilities across a diverse spectrum of multimodal tasks, it yet rema

model-releasesarxiv-cs-cv
14 Apr 2026
Applications

BEM: Training-Free Background Embedding Memory for False-Positive Suppression in Real-Time Fixed-Background Camera

DGX agent

arXiv:2604.11714v1 Announce Type: new Abstract: Pretrained detectors perform well on benchmarks but often suffer performance degradation in real-world deployments due to distribution gaps between trai

applicationsarxiv-cs-cv
14 Apr 2026
Model Releases

Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation

DGX agent

arXiv:2504.14988v3 Announce Type: replace Abstract: Recent advancements in Large Vision-Language Models (LVLMs) have demonstrated remarkable multimodal perception capabilities, garnering significant a

model-releasesarxiv-cs-cv
14 Apr 2026
Model Releases

Benchmarking Vision-Language Models under Contradictory Virtual Content Attacks in Augmented Reality

DGX agent

arXiv:2604.05510v2 Announce Type: replace Abstract: Augmented reality (AR) has rapidly expanded over the past decade. As AR becomes increasingly integrated into daily life, its security and reliabilit

model-releasesarxiv-cs-cv
14 Apr 2026
Local Ai

Beyond Model Design: Data-Centric Training and Self-Ensemble for Gaussian Color Image Denoising

DGX agent

arXiv:2604.11468v1 Announce Type: new Abstract: This paper presents our solution to the NTIRE 2026 Image Denoising Challenge (Gaussian color image denoising at fixed noise level sigma = 50). Rather th

local-aiarxiv-cs-cv
14 Apr 2026
Safety

Beyond Reconstruction: Reconstruction-to-Vector Diffusion for Hyperspectral Anomaly Detection

DGX agent

arXiv:2604.11390v1 Announce Type: new Abstract: While Hyperspectral Anomaly Detection (HAD) excels at identifying sparse targets in complex scenes, existing models remain trapped in a scalar 'reconstr

safetyarxiv-cs-cv
14 Apr 2026
Model Releases

Bidirectional Cross-Attention Fusion of High-Res RGB and Low-Res HSI for Multimodal Automated Waste Sorting

DGX agent

arXiv:2603.13941v2 Announce Type: replace Abstract: Growing waste streams and the transition to a circular economy require efficient automated waste sorting. In industrial settings, materials move on

model-releasesarxiv-cs-cv
14 Apr 2026
Safety

Bidirectional Learning of Facial Action Units and Expressions via Structured Semantic Mapping across Heterogeneous Datasets

DGX agent

arXiv:2604.10541v1 Announce Type: new Abstract: Facial action unit (AU) detection and facial expression (FE) recognition can be jointly viewed as affective facial behavior tasks, representing fine-gra

safetyarxiv-cs-cv
14 Apr 2026
Research

Biomarker-Based Pretraining for Chagas Disease Screening in Electrocardiograms

DGX agent

arXiv:2604.09782v1 Announce Type: new Abstract: Chagas disease screening via ECGs is limited by scarce and noisy labels in existing datasets. We propose a biomarker-based pretraining approach, where a

researcharxiv-cs-cv
14 Apr 2026
Applications

BLPR: Robust License Plate Recognition under Viewpoint and Illumination Variations via Confidence-Driven VLM Fallback

DGX agent

arXiv:2604.09927v1 Announce Type: new Abstract: Robust license plate recognition in unconstrained environments remains a significant challenge, particularly in underrepresented regions with limited da

applicationsarxiv-cs-cv
14 Apr 2026
Applications

Bootstrapping Video Semantic Segmentation Model via Distillation-assisted Test-Time Adaptation

DGX agent

arXiv:2604.10950v1 Announce Type: new Abstract: Fully supervised Video Semantic Segmentation (VSS) relies heavily on densely annotated video data, limiting practical applicability. Alternatively, appl

applicationsarxiv-cs-cv
14 Apr 2026
Model Releases

Boxes2Pixels: Learning Defect Segmentation from Noisy SAM Masks

DGX agent

arXiv:2604.11162v1 Announce Type: new Abstract: Accurate defect segmentation is critical for industrial inspection, yet dense pixel-level annotations are rarely available. A common workaround is to co

model-releasesarxiv-cs-cv
14 Apr 2026
Safety

Brain-Grasp: Graph-based Saliency Priors for Improved fMRI-based Visual Brain Decoding

DGX agent

arXiv:2604.10617v1 Announce Type: cross Abstract: Recent progress in brain-guided image generation has improved the quality of fMRI-based reconstructions; however, fundamental challenges remain in pre

safetyarxiv-cs-cv
14 Apr 2026
Safety

Bridging the RGB-IR Gap: Consensus and Discrepancy Modeling for Text-Guided Multispectral Detection

DGX agent

arXiv:2604.11234v1 Announce Type: new Abstract: Text-guided multispectral object detection uses text semantics to guide semantic-aware cross-modal interaction between RGB and IR for more robust percep

safetyarxiv-cs-cv
14 Apr 2026
Research

Byte-level generative predictions for forensics multimedia carving

DGX agent

arXiv:2604.11010v1 Announce Type: new Abstract: Digital forensic investigations often face significant challenges when recovering fragmented multimedia files that lack file system metadata. While trad

researcharxiv-cs-cv
14 Apr 2026
Model Releases

Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?

DGX agent

arXiv:2511.21998v2 Announce Type: replace Abstract: Multi-modal Large Language Models (LLM) have advanced conversational abilities but struggle with providing live, interactive step-by-step guidance,

model-releasesarxiv-cs-cv
14 Apr 2026
Applications

Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification?

DGX agent

arXiv:2601.06993v2 Announce Type: replace Abstract: Multi-modal large language models (MLLMs) exhibit strong general-purpose capabilities, yet still struggle on Fine-Grained Visual Classification (FGV

applicationsarxiv-cs-cv
14 Apr 2026
Research

Catalyst: Out-of-Distribution Detection via Elastic Scaling

DGX agent

arXiv:2602.02409v2 Announce Type: replace Abstract: Out-of-distribution (OOD) detection is critical for the safe deployment of deep neural networks. State-of-the-art post-hoc methods typically derive

researcharxiv-cs-cv
14 Apr 2026
Applications

CDPR: Cross-modal Diffusion with Polarization for Reliable Monocular Depth Estimation

DGX agent

arXiv:2604.11097v1 Announce Type: new Abstract: Monocular depth estimation is a fundamental yet challenging task in computer vision, especially under complex conditions such as textureless surfaces, t

applicationsarxiv-cs-cv
14 Apr 2026
Safety

CityGuard: Graph-Aware Private Descriptors for Bias-Resilient Identity Search Across Urban Cameras

DGX agent

arXiv:2602.18047v3 Announce Type: replace Abstract: City-scale person re-identification across distributed cameras must handle severe appearance changes from viewpoint, occlusion, and domain shift whi

safetyarxiv-cs-cv
14 Apr 2026
Model Releases

CoFusion: Multispectral and Hyperspectral Image Fusion via Spectral Coordinate Attention

DGX agent

arXiv:2604.10584v1 Announce Type: new Abstract: Multispectral and Hyperspectral Image Fusion (MHIF) aims to reconstruct high-resolution images by integrating low-resolution hyperspectral images (LRHSI

model-releasesarxiv-cs-cv
14 Apr 2026
Research

Compact single-shot ranging and near-far imaging using metasurfaces

DGX agent

arXiv:2604.10037v1 Announce Type: cross Abstract: We present a metasurface imaging system capable of simultaneously capturing two images at close range (1-2~cm) and an additional image at long range (

researcharxiv-cs-cv
14 Apr 2026
Safety

ComSim: Building Scalable Real-World Robot Data Generation via Compositional Simulation

DGX agent

arXiv:2604.11386v1 Announce Type: cross Abstract: Recent advancements in foundational models, such as large language models and world models, have greatly enhanced the capabilities of robotics, enabli

safetyarxiv-cs-cv
14 Apr 2026
Research

ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectors

DGX agent

arXiv:2512.09056v3 Announce Type: replace Abstract: Object pose estimation is a fundamental task in computer vision and robotics, yet most methods require extensive, dataset-specific training. Concurr

researcharxiv-cs-cv
14 Apr 2026
Model Releases

Context-Aware Semantic Segmentation via Stage-Wise Attention

DGX agent

arXiv:2601.11310v2 Announce Type: replace Abstract: Semantic ultra-high-resolution (UHR) image segmentation is essential in remote sensing applications such as aerial mapping and environmental monitor

model-releasesarxiv-cs-cv
14 Apr 2026
Safety

Context Matters: Vision-Based Depression Detection Comparing Classical and Deep Approaches

DGX agent

arXiv:2604.10344v1 Announce Type: new Abstract: The classical approach to detecting depression from vision emphasizes interpretable features, such as facial expression, and classifiers such as the Sup

safetyarxiv-cs-cv
14 Apr 2026
Tutorials

Continuous Adversarial Flow Models

DGX agent

arXiv:2604.11521v1 Announce Type: cross Abstract: We propose continuous adversarial flow models, a type of continuous-time flow model trained with an adversarial objective. Unlike flow matching, which

tutorialsarxiv-cs-cv
14 Apr 2026
Safety

Contour Refinement using Discrete Diffusion in Low Data Regime

DGX agent

arXiv:2602.05880v2 Announce Type: replace Abstract: Boundary detection of irregular and translucent objects is an important problem with applications in medical imaging, environmental monitoring and m

safetyarxiv-cs-cv
14 Apr 2026
Research

ConvFormer3D-TAP: Phase/Uncertainty-Aware Front-End Fusion for Cine CMR View Classification Pipelines

DGX agent

arXiv:2604.11389v1 Announce Type: new Abstract: Reliable recognition of standard cine cardiac MRI views is essential because each view determines which cardiac anatomy is visualized and which quantita

researcharxiv-cs-cv
14 Apr 2026
Safety

CoPS: Conditional Prompt Synthesis for Zero-Shot Anomaly Detection

DGX agent

arXiv:2508.03447v2 Announce Type: replace Abstract: Recently, large pre-trained vision-language models have shown remarkable performance in zero-shot anomaly detection (ZSAD). With fine-tuning on a si

safetyarxiv-cs-cv
14 Apr 2026
Model Releases

Counting to Four is still a Chore for VLMs

DGX agent

arXiv:2604.10039v1 Announce Type: new Abstract: Vision--language models (VLMs) have achieved impressive performance on complex multimodal reasoning tasks, yet they still fail on simple grounding skill

model-releasesarxiv-cs-cv
14 Apr 2026
Agents

CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance

DGX agent

arXiv:2508.16644v4 Announce Type: replace Abstract: Diffusion models excel at photorealistic synthesis but struggle with precise object counts, especially in high-density settings. We introduce COUNTL

agentsarxiv-cs-cv
14 Apr 2026
Model Releases

CPAM: Context-Preserving Adaptive Manipulation for Zero-Shot Real Image Editing

DGX agent

arXiv:2506.18438v2 Announce Type: replace Abstract: Editing natural images using textual descriptions in text-to-image diffusion models remains a significant challenge, particularly in achieving consi

model-releasesarxiv-cs-cv
14 Apr 2026
Applications

CraftGraffiti: Exploring Human Identity with Custom Graffiti Art via Facial-Preserving Diffusion Models

DGX agent

arXiv:2508.20640v2 Announce Type: replace Abstract: Preserving facial identity under extreme stylistic transformation remains a major challenge in generative art. In graffiti, a high-contrast, abstrac

applicationsarxiv-cs-cv
14 Apr 2026
Research

CylinderDepth: Cylindrical Spatial Attention for Multi-View Consistent Self-Supervised Surround Depth Estimation

DGX agent

arXiv:2511.16428v3 Announce Type: replace Abstract: Self-supervised surround-view depth estimation enables dense, low-cost 3D perception with a 360{eg} field of view from multiple minimally overlappin

researcharxiv-cs-cv
14 Apr 2026
Research

Dark-EvGS: Event Camera as an Eye for Radiance Field in the Dark

DGX agent

arXiv:2507.11931v2 Announce Type: replace Abstract: In low-light environments, conventional cameras often struggle to capture clear multi-view images of objects due to dynamic range limitations and mo

researcharxiv-cs-cv
14 Apr 2026
Local Ai

Data-Driven Automated Identification of Optimal Feature-Representative Images in Infrared Thermography Using Statistical and Morphological Metrics

DGX agent

arXiv:2604.09728v1 Announce Type: new Abstract: Infrared thermography (IRT) is a widely used non-destructive testing technique for detecting structural features such as subsurface defects. However, mo

local-aiarxiv-cs-cv
14 Apr 2026
Model Releases

Data-Efficient Semantic Segmentation of 3D Point Clouds via Open-Vocabulary Image Segmentation-based Pseudo-Labeling

DGX agent

arXiv:2604.11007v1 Announce Type: new Abstract: Semantic segmentation of 3D point cloud scenes is a crucial task for various applications. In real-world scenarios, training segmentation models often f

model-releasesarxiv-cs-cv
14 Apr 2026
Safety

DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning

DGX agent

arXiv:2503.11892v3 Announce Type: replace Abstract: Multimodal representation learning aims to capture both shared and complementary semantic information across multiple modalities. However, the intri

safetyarxiv-cs-cv
14 Apr 2026
Research

Decoding Ancient Oracle Bone Script via Generative Dictionary Retrieval

DGX agent

arXiv:2604.09668v1 Announce Type: cross Abstract: Understanding humanity's earliest writing systems is crucial for reconstructing civilization's origins, yet many ancient scripts remain undeciphered.

researcharxiv-cs-cv
14 Apr 2026
Model Releases

Decompose, Mix, Adapt: A Unified Framework for Parameter-Efficient Neural Network Recombination and Compression

DGX agent

arXiv:2603.27383v2 Announce Type: replace Abstract: Parameter Recombination (PR) methods aim to efficiently compose the weights of a neural network for applications like Parameter-Efficient FineTuning

model-releasesarxiv-cs-cv
14 Apr 2026
Research

Decoupled Generative Modeling for Human-Object Interaction Synthesis

DGX agent

arXiv:2512.19049v2 Announce Type: replace Abstract: Synthesizing realistic human-object interaction (HOI) is essential for 3D computer vision and robotics, underpinning animation and embodied control.

researcharxiv-cs-cv
14 Apr 2026
Research

Decoupled Similarity for Task-Aware Token Pruning in Large Vision-Language Models

DGX agent

arXiv:2604.11240v1 Announce Type: new Abstract: Token pruning has emerged as an effective approach to reduce the substantial computational overhead of Large Vision-Language Models (LVLMs) by discardin

researcharxiv-cs-cv
14 Apr 2026
Research

DeepShapeMatchingKit: Accelerated Functional Map Solver and Shape Matching Pipelines Revisited

DGX agent

arXiv:2604.10377v1 Announce Type: new Abstract: Deep functional maps, leveraging learned feature extractors and spectral correspondence solvers, are fundamental to non-rigid 3D shape matching. Based o

researcharxiv-cs-cv
14 Apr 2026
Research

DeepSketcher: Internalizing Visual Manipulation for Multimodal Reasoning

DGX agent

arXiv:2509.25866v2 Announce Type: replace Abstract: The 'thinking with images' paradigm represents a pivotal shift in the reasoning of Vision Language Models (VLMs), moving from text-dominant chain-of

researcharxiv-cs-cv
14 Apr 2026
Local Ai

Defending against Patch-Based and Texture-Based Adversarial Attacks with Spectral Decomposition

DGX agent

arXiv:2604.10715v1 Announce Type: new Abstract: Adversarial examples present significant challenges to the security of Deep Neural Network (DNN) applications. Specifically, there are patch-based and t

local-aiarxiv-cs-cv
14 Apr 2026
Local Ai

Degradation-Aware and Structure-Preserving Diffusion for Real-World Image Super-Resolution

DGX agent

arXiv:2604.11470v1 Announce Type: new Abstract: Real-world image super-resolution is particularly challenging for diffusion models because real degradations are complex, heterogeneous, and rarely mode

local-aiarxiv-cs-cv
14 Apr 2026
← Previous
1…243244245246247…259
Next →