AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Research

LEXIS: LatEnt ProXimal Interaction Signatures for 3D HOI from an Image

DGX agent

arXiv:2604.20800v1 Announce Type: new Abstract: Reconstructing 3D Human-Object Interaction from an RGB image is essential for perceptive systems. Yet, this remains challenging as it requires capturing

researcharxiv-cs-cv
23 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Agents

Lifecycle-Aware Federated Continual Learning in Mobile Autonomous Systems

DGX agent

arXiv:2604.20745v1 Announce Type: cross Abstract: Federated continual learning (FCL) allows distributed autonomous fleets to adapt collaboratively to evolving terrain types across extended mission lif

agentsarxiv-cs-cv
23 Apr 2026
Research

LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model

DGX agent

arXiv:2604.20796v1 Announce Type: new Abstract: We present LLaDA2.0-Uni, a unified discrete diffusion large language model (dLLM) that supports multimodal understanding and generation within a nativel

researcharxiv-cs-cv
23 Apr 2026
Research

Lucky High Dynamic Range Smartphone Imaging

DGX agent

arXiv:2604.19976v1 Announce Type: new Abstract: While the human eye can perceive an impressive twenty stops of dynamic range, smartphone camera sensors remain limited to about twelve stops despite dec

researcharxiv-cs-cv
23 Apr 2026
Model Releases

MAPRPose: Mask-Aware Proposal and Amodal Refinement for Multi-Object 6D Pose Estimation

DGX agent

arXiv:2604.20650v1 Announce Type: new Abstract: 6D object pose estimation in cluttered scenes remains challenging due to severe occlusion and sensor noise. We propose MAPRPose, a two-stage framework t

model-releasesarxiv-cs-cv
23 Apr 2026
Research

Maximum Likelihood Reconstruction for Multi-Look Digital Holography with Markov-Modeled Speckle Correlation

DGX agent

arXiv:2604.20154v1 Announce Type: cross Abstract: Multi-look acquisition is a widely used strategy for reducing speckle noise in coherent imaging systems such as digital holography. By acquiring multi

researcharxiv-cs-cv
23 Apr 2026
Tutorials

MD-Face: MoE-Enhanced Label-Free Disentangled Representation for Interactive Facial Attribute Editing

DGX agent

arXiv:2604.20317v1 Announce Type: new Abstract: GAN-based facial attribute editing is widely used in virtual avatars and social media but often suffers from attribute entanglement, where modifying one

tutorialsarxiv-cs-cv
23 Apr 2026
Model Releases

Mitigating Hallucinations in Large Vision-Language Models without Performance Degradation

DGX agent

arXiv:2604.20366v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) exhibit powerful generative capabilities but frequently produce hallucinations that compromise output reliability.

model-releasesarxiv-cs-cv
23 Apr 2026
Model Releases

MLG-Stereo: ViT Based Stereo Matching with Multi-Stage Local-Global Enhancement

DGX agent

arXiv:2604.20393v1 Announce Type: new Abstract: With the development of deep learning, ViT-based stereo matching methods have made significant progress due to their remarkable robustness and zero-shot

model-releasesarxiv-cs-cv
23 Apr 2026
Model Releases

MSLAU-Net: A Hybrid CNN-Transformer Network for Medical Image Segmentation

DGX agent

arXiv:2505.18823v2 Announce Type: replace Abstract: Accurate medical image segmentation allows for the precise delineation of anatomical structures and pathological regions, which is essential for tre

model-releasesarxiv-cs-cv
23 Apr 2026
Research

Object Referring-Guided Scanpath Prediction with Perception-Enhanced Vision-Language Models

DGX agent

arXiv:2604.20361v1 Announce Type: new Abstract: Object Referring-guided Scanpath Prediction (ORSP) aims to predict the human attention scanpath when they search for a specific target object in a visua

researcharxiv-cs-cv
23 Apr 2026
Research

On the Impact of Face Segmentation-Based Background Removal on Recognition and Morphing Attack Detection

DGX agent

arXiv:2604.20585v1 Announce Type: new Abstract: This study investigates the impact of face image background correction through segmentation on face recognition and morphing attack detection performanc

researcharxiv-cs-cv
23 Apr 2026
Research

Online CS-based SAR Edge-Mapping

DGX agent

arXiv:2604.19989v1 Announce Type: new Abstract: With modern defense applications increasingly relying on inexpensive, small Unmanned Aerial Vehicles (UAVs), a major challenge lies in designing intelli

researcharxiv-cs-cv
23 Apr 2026
Safety

OnSiteVRU: A High-Resolution Trajectory Dataset for High-Density Vulnerable Road Users

DGX agent

arXiv:2503.23365v2 Announce Type: replace Abstract: With the acceleration of urbanization and the growth of transportation demands, the safety of vulnerable road users (VRUs, such as pedestrians and c

safetyarxiv-cs-cv
23 Apr 2026
Research

Opportunistic Bone-Loss Screening from Routine Knee Radiographs Using a Multi-Task Deep Learning Framework with Sensitivity-Constrained Threshold Optimization

DGX agent

arXiv:2604.20268v1 Announce Type: new Abstract: Background: Osteoporosis and osteopenia are often undiagnosed until fragility fractures occur. Dual-energy X-ray absorptiometry (DXA) is the reference s

researcharxiv-cs-cv
23 Apr 2026
Research

Optimizing Data Augmentation for Real-Time Small UAV Detection: A Lightweight Context-Aware Approach

DGX agent

arXiv:2604.19999v1 Announce Type: new Abstract: Visual detection of Unmanned Aerial Vehicles (UAVs) is a critical task in surveillance systems due to their small physical size and environmental challe

researcharxiv-cs-cv
23 Apr 2026
Research

Pairing Regularization for Mitigating Many-to-One Collapse in GANs

DGX agent

arXiv:2604.20130v1 Announce Type: cross Abstract: Mode collapse remains a fundamental challenge in training generative adversarial networks (GANs). While existing works have primarily focused on inter

researcharxiv-cs-cv
23 Apr 2026
Research

ParetoSlider: Diffusion Models Post-Training for Continuous Reward Control

DGX agent

arXiv:2604.20816v1 Announce Type: cross Abstract: Reinforcement Learning (RL) post-training has become the standard for aligning generative models with human preferences, yet most methods rely on a si

researcharxiv-cs-cv
23 Apr 2026
Research

PASTA: A Patch-Agnostic Twofold-Stealthy Backdoor Attack on Vision Transformers

DGX agent

arXiv:2604.20047v1 Announce Type: new Abstract: Vision Transformers (ViTs) have achieved remarkable success across vision tasks, yet recent studies show they remain vulnerable to backdoor attacks. Exi

researcharxiv-cs-cv
23 Apr 2026
Research

PFGNet: A Fully Convolutional Frequency-Guided Peripheral Gating Network for Efficient Spatiotemporal Predictive Learning

DGX agent

arXiv:2602.20537v3 Announce Type: replace Abstract: Spatiotemporal predictive learning (STPL) aims to forecast future frames from past observations and is essential across a wide range of applications

researcharxiv-cs-cv
23 Apr 2026
Applications

Physics-informed Active Polarimetric 3D Imaging for Specular Surfaces

DGX agent

arXiv:2602.19470v2 Announce Type: replace Abstract: 3D imaging of specular surfaces remains challenging in real-world scenarios, such as in-line inspection or hand-held scanning, requiring fast and ac

applicationsarxiv-cs-cv
23 Apr 2026
Research

Physics-Informed Conditional Diffusion for Motion-Robust Retinal Temporal Laser Speckle Contrast Imaging

DGX agent

arXiv:2604.20594v1 Announce Type: new Abstract: Retinal laser speckle contrast imaging (LSCI) is a noninvasive optical modality for monitoring retinal blood flow dynamics. However, conventional tempor

researcharxiv-cs-cv
23 Apr 2026
Safety

ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards

DGX agent

arXiv:2604.20486v1 Announce Type: new Abstract: Training multimodal agents via reinforcement learning for knowledge-intensive visual reasoning is fundamentally hindered by the extreme sparsity of outc

safetyarxiv-cs-cv
23 Apr 2026
Research

R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs

DGX agent

arXiv:2604.20696v1 Announce Type: new Abstract: Large vision-language models (LVLMs) have demonstrated impressive performance in various multimodal understanding and reasoning tasks. However, they sti

researcharxiv-cs-cv
23 Apr 2026
Research

Random Walk on Point Clouds for Feature Detection

DGX agent

arXiv:2604.20474v1 Announce Type: new Abstract: The points on the point clouds that can entirely outline the shape of the model are of critical importance, as they serve as the foundation for numerous

researcharxiv-cs-cv
23 Apr 2026
Model Releases

RareSpot+: A Benchmark, Model, and Active Learning Framework for Small and Rare Wildlife in Aerial Imagery

DGX agent

arXiv:2604.20000v1 Announce Type: new Abstract: Automated wildlife monitoring from aerial imagery is vital for conservation but remains limited by two persistent challenges: the difficulty of detectin

model-releasesarxiv-cs-cv
23 Apr 2026
Model Releases

RefAerial: A Benchmark and Approach for Referring Detection in Aerial Images

DGX agent

arXiv:2604.20543v1 Announce Type: new Abstract: Referring detection refers to locate the target referred by natural languages, which has recently attracted growing research interests. However, existin

model-releasesarxiv-cs-cv
23 Apr 2026
Tutorials

Render-in-the-Loop: Vector Graphics Generation via Visual Self-Feedback

DGX agent

arXiv:2604.20730v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have shown promising capabilities in generating Scalable Vector Graphics (SVG) via direct code synthesis. Howev

tutorialsarxiv-cs-cv
23 Apr 2026
Model Releases

Rethinking Where to Edit: Task-Aware Localization for Instruction-Based Image Editing

DGX agent

arXiv:2604.20258v1 Announce Type: new Abstract: Instruction-based image editing (IIE) aims to modify images according to textual instructions while preserving irrelevant content. Despite recent advanc

model-releasesarxiv-cs-cv
23 Apr 2026
Model Releases

retinalysis-vascx: An explainable software toolbox for the extraction of retinal vascular biomarkers

DGX agent

arXiv:2602.08580v2 Announce Type: replace-cross Abstract: Automatic extraction of retinal vascular biomarkers from color fundus images (CFI) is crucial for large-scale studies of the retinal vasculatu

model-releasesarxiv-cs-cv
23 Apr 2026
Research

Retinex Meets Language: A Physics-Semantics-Guided Underwater Image Enhancement Network

DGX agent

arXiv:2603.07076v2 Announce Type: replace Abstract: Underwater images often suffer from severe degradation caused by light absorption and scattering, leading to color distortion, low contrast and redu

researcharxiv-cs-cv
23 Apr 2026
Model Releases

REVNET: Rotation-Equivariant Point Cloud Completion via Vector Neuron Anchor Transformer

DGX agent

arXiv:2601.08558v2 Announce Type: replace Abstract: Incomplete point clouds captured by 3D sensors often result in the loss of both geometric and semantic information. Most existing point cloud comple

model-releasesarxiv-cs-cv
23 Apr 2026
Research

Robust Principal Component Completion

DGX agent

arXiv:2603.25132v2 Announce Type: replace Abstract: Robust principal component analysis (RPCA) seeks a low-rank component and a sparse component from their summation. Yet, in many applications of inte

researcharxiv-cs-cv
23 Apr 2026
Safety

Rodrigues Network for Learning Robot Actions

DGX agent

arXiv:2506.02618v2 Announce Type: replace-cross Abstract: Understanding and predicting articulated actions is important in robot learning. However, common architectures such as MLPs and Transformers l

safetyarxiv-cs-cv
23 Apr 2026
Safety

Sampling-Aware Quantization for Diffusion Models

DGX agent

arXiv:2505.02242v2 Announce Type: replace Abstract: Diffusion models have recently emerged as the dominant approach in visual generation tasks. However, the lengthy denoising chains and the computatio

safetyarxiv-cs-cv
23 Apr 2026
Model Releases

SceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation

DGX agent

arXiv:2604.19907v1 Announce Type: new Abstract: Recent agentic frameworks for 3D scene synthesis have advanced realism and diversity by integrating heterogeneous generation and editing tools. These to

model-releasesarxiv-cs-cv
23 Apr 2026
Research

Secure Rate-Distortion-Perception: A Randomized Distributed Function Computation Approach for Realism

DGX agent

arXiv:2604.20245v1 Announce Type: cross Abstract: Fundamental rate-distortion-perception (RDP) trade-offs arise in applications requiring maintained perceptual quality of reconstructed data, such as n

researcharxiv-cs-cv
23 Apr 2026
Model Releases

SegEarth-OV3: Exploring SAM 3 for Open-Vocabulary Semantic Segmentation in Remote Sensing Images

DGX agent

arXiv:2512.08730v2 Announce Type: replace Abstract: Most existing methods for training-free open-vocabulary semantic segmentation are based on CLIP. While these approaches have made progress, they oft

model-releasesarxiv-cs-cv
23 Apr 2026
Research

Self-supervised pretraining for an iterative image size agnostic vision transformer

DGX agent

arXiv:2604.20392v1 Announce Type: new Abstract: Vision Transformers (ViTs) dominate self-supervised learning (SSL). While they have proven highly effective for large-scale pretraining, they are comput

researcharxiv-cs-cv
23 Apr 2026
Research

Semantic-Fast-SAM: Efficient Semantic Segmenter

DGX agent

arXiv:2604.20169v1 Announce Type: new Abstract: We propose Semantic-Fast-SAM (SFS), a semantic segmentation framework that combines the Fast Segment Anything model with a semantic labeling pipeline to

researcharxiv-cs-cv
23 Apr 2026
Applications

Semantic-guided Gaussian Splatting for High-Fidelity Underwater Scene Reconstruction

DGX agent

arXiv:2509.00800v3 Announce Type: replace Abstract: Accurate 3D reconstruction in degraded imaging conditions remains a key challenge in photogrammetry and neural rendering. In underwater environments

applicationsarxiv-cs-cv
23 Apr 2026
Model Releases

Semi-Supervised Flow Matching for Mosaiced and Panchromatic Fusion Imaging

DGX agent

arXiv:2604.20128v1 Announce Type: new Abstract: Fusing a low resolution (LR) mosaiced hyperspectral image (HSI) with a high resolution (HR) panchromatic (PAN) image offers a promising avenue for video

model-releasesarxiv-cs-cv
23 Apr 2026
Model Releases

SGAP-Gaze: Scene Grid Attention Based Point-of-Gaze Estimation Network for Driver Gaze

DGX agent

arXiv:2604.19888v1 Announce Type: new Abstract: Driver gaze estimation is essential for understanding the driver's situational awareness of surrounding traffic. Existing gaze estimation models use dri

model-releasesarxiv-cs-cv
23 Apr 2026
Research

SpaCeFormer: Fast Proposal-Free Open-Vocabulary 3D Instance Segmentation

DGX agent

arXiv:2604.20395v1 Announce Type: new Abstract: Open-vocabulary 3D instance segmentation is a core capability for robotics and AR/VR, but prior methods trade one bottleneck for another: multi-stage 2D

researcharxiv-cs-cv
23 Apr 2026
Research

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models

DGX agent

arXiv:2604.20705v1 Announce Type: new Abstract: Reinforcement learning (RL) with verifiable rewards (RLVR) has demonstrated the great potential of enhancing the reasoning abilities in multimodal large

researcharxiv-cs-cv
23 Apr 2026
Tutorials

Stability-Driven Motion Generation for Object-Guided Human-Human Co-Manipulation

DGX agent

arXiv:2604.20336v1 Announce Type: new Abstract: Co-manipulation requires multiple humans to synchronize their motions with a shared object while ensuring reasonable interactions, maintaining natural p

tutorialsarxiv-cs-cv
23 Apr 2026
Research

Structure-Augmented Standard Plane Detection with Temporal Aggregation in Blind-Sweep Fetal Ultrasound

DGX agent

arXiv:2604.20591v1 Announce Type: new Abstract: In low-resource settings, blind-sweep ultrasound provides a practical and accessible method for identifying fetal growth restriction. However, unlike fr

researcharxiv-cs-cv
23 Apr 2026
Model Releases

SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark

DGX agent

arXiv:2604.20319v1 Announce Type: new Abstract: Fine-grained spatiotemporal reasoning on surgical videos is critical, yet the capabilities of Multi-modal Large Language Models (MLLMs) in this domain r

model-releasesarxiv-cs-cv
23 Apr 2026
← Previous
1…220221222223224…261
Next →