Online CS-based SAR Edge-Mapping
arXiv:2604.19989v1 Announce Type: new Abstract: With modern defense applications increasingly relying on inexpensive, small Unmanned Aerial Vehicles (UAVs), a major challenge lies in designing intelli
Knowledge catalogue
arXiv:2604.19989v1 Announce Type: new Abstract: With modern defense applications increasingly relying on inexpensive, small Unmanned Aerial Vehicles (UAVs), a major challenge lies in designing intelli
arXiv:2503.23365v2 Announce Type: replace Abstract: With the acceleration of urbanization and the growth of transportation demands, the safety of vulnerable road users (VRUs, such as pedestrians and c
arXiv:2604.20268v1 Announce Type: new Abstract: Background: Osteoporosis and osteopenia are often undiagnosed until fragility fractures occur. Dual-energy X-ray absorptiometry (DXA) is the reference s
arXiv:2604.19999v1 Announce Type: new Abstract: Visual detection of Unmanned Aerial Vehicles (UAVs) is a critical task in surveillance systems due to their small physical size and environmental challe
arXiv:2604.20130v1 Announce Type: cross Abstract: Mode collapse remains a fundamental challenge in training generative adversarial networks (GANs). While existing works have primarily focused on inter
arXiv:2604.20816v1 Announce Type: cross Abstract: Reinforcement Learning (RL) post-training has become the standard for aligning generative models with human preferences, yet most methods rely on a si
arXiv:2604.20047v1 Announce Type: new Abstract: Vision Transformers (ViTs) have achieved remarkable success across vision tasks, yet recent studies show they remain vulnerable to backdoor attacks. Exi
arXiv:2602.20537v3 Announce Type: replace Abstract: Spatiotemporal predictive learning (STPL) aims to forecast future frames from past observations and is essential across a wide range of applications
arXiv:2602.19470v2 Announce Type: replace Abstract: 3D imaging of specular surfaces remains challenging in real-world scenarios, such as in-line inspection or hand-held scanning, requiring fast and ac
arXiv:2604.20594v1 Announce Type: new Abstract: Retinal laser speckle contrast imaging (LSCI) is a noninvasive optical modality for monitoring retinal blood flow dynamics. However, conventional tempor
arXiv:2604.20486v1 Announce Type: new Abstract: Training multimodal agents via reinforcement learning for knowledge-intensive visual reasoning is fundamentally hindered by the extreme sparsity of outc
arXiv:2604.20696v1 Announce Type: new Abstract: Large vision-language models (LVLMs) have demonstrated impressive performance in various multimodal understanding and reasoning tasks. However, they sti
arXiv:2604.20474v1 Announce Type: new Abstract: The points on the point clouds that can entirely outline the shape of the model are of critical importance, as they serve as the foundation for numerous
arXiv:2604.20000v1 Announce Type: new Abstract: Automated wildlife monitoring from aerial imagery is vital for conservation but remains limited by two persistent challenges: the difficulty of detectin
arXiv:2604.20543v1 Announce Type: new Abstract: Referring detection refers to locate the target referred by natural languages, which has recently attracted growing research interests. However, existin
arXiv:2604.20730v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have shown promising capabilities in generating Scalable Vector Graphics (SVG) via direct code synthesis. Howev
arXiv:2604.20258v1 Announce Type: new Abstract: Instruction-based image editing (IIE) aims to modify images according to textual instructions while preserving irrelevant content. Despite recent advanc
arXiv:2602.08580v2 Announce Type: replace-cross Abstract: Automatic extraction of retinal vascular biomarkers from color fundus images (CFI) is crucial for large-scale studies of the retinal vasculatu
arXiv:2603.07076v2 Announce Type: replace Abstract: Underwater images often suffer from severe degradation caused by light absorption and scattering, leading to color distortion, low contrast and redu
arXiv:2601.08558v2 Announce Type: replace Abstract: Incomplete point clouds captured by 3D sensors often result in the loss of both geometric and semantic information. Most existing point cloud comple
arXiv:2603.25132v2 Announce Type: replace Abstract: Robust principal component analysis (RPCA) seeks a low-rank component and a sparse component from their summation. Yet, in many applications of inte
arXiv:2506.02618v2 Announce Type: replace-cross Abstract: Understanding and predicting articulated actions is important in robot learning. However, common architectures such as MLPs and Transformers l
arXiv:2505.02242v2 Announce Type: replace Abstract: Diffusion models have recently emerged as the dominant approach in visual generation tasks. However, the lengthy denoising chains and the computatio
arXiv:2604.19907v1 Announce Type: new Abstract: Recent agentic frameworks for 3D scene synthesis have advanced realism and diversity by integrating heterogeneous generation and editing tools. These to
arXiv:2604.20245v1 Announce Type: cross Abstract: Fundamental rate-distortion-perception (RDP) trade-offs arise in applications requiring maintained perceptual quality of reconstructed data, such as n
arXiv:2512.08730v2 Announce Type: replace Abstract: Most existing methods for training-free open-vocabulary semantic segmentation are based on CLIP. While these approaches have made progress, they oft
arXiv:2604.20392v1 Announce Type: new Abstract: Vision Transformers (ViTs) dominate self-supervised learning (SSL). While they have proven highly effective for large-scale pretraining, they are comput
arXiv:2604.20169v1 Announce Type: new Abstract: We propose Semantic-Fast-SAM (SFS), a semantic segmentation framework that combines the Fast Segment Anything model with a semantic labeling pipeline to
arXiv:2509.00800v3 Announce Type: replace Abstract: Accurate 3D reconstruction in degraded imaging conditions remains a key challenge in photogrammetry and neural rendering. In underwater environments
arXiv:2604.20128v1 Announce Type: new Abstract: Fusing a low resolution (LR) mosaiced hyperspectral image (HSI) with a high resolution (HR) panchromatic (PAN) image offers a promising avenue for video
arXiv:2604.19888v1 Announce Type: new Abstract: Driver gaze estimation is essential for understanding the driver's situational awareness of surrounding traffic. Existing gaze estimation models use dri
arXiv:2604.20395v1 Announce Type: new Abstract: Open-vocabulary 3D instance segmentation is a core capability for robotics and AR/VR, but prior methods trade one bottleneck for another: multi-stage 2D
arXiv:2604.20705v1 Announce Type: new Abstract: Reinforcement learning (RL) with verifiable rewards (RLVR) has demonstrated the great potential of enhancing the reasoning abilities in multimodal large
arXiv:2604.20336v1 Announce Type: new Abstract: Co-manipulation requires multiple humans to synchronize their motions with a shared object while ensuring reasonable interactions, maintaining natural p
arXiv:2604.20591v1 Announce Type: new Abstract: In low-resource settings, blind-sweep ultrasound provides a practical and accessible method for identifying fetal growth restriction. However, unlike fr
arXiv:2604.20319v1 Announce Type: new Abstract: Fine-grained spatiotemporal reasoning on surgical videos is critical, yet the capabilities of Multi-modal Large Language Models (MLLMs) in this domain r
arXiv:2409.07609v2 Announce Type: replace-cross Abstract: Deploying adversarially robust machine learning systems requires continuous trade-offs between robustness, cost, and latency. We present an au
arXiv:2604.19829v1 Announce Type: new Abstract: Tactile graphics require careful expert validation before reaching blind and visually impaired (BVI) learners, yet existing datasets provide only coarse
arXiv:2603.20714v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has become the method of choice for photo-realistic 3D reconstruction of scenes, due to being able to efficiently and a
arXiv:2604.20123v1 Announce Type: new Abstract: In natural images, object skeletons are used to represent geometric shapes. However, even slight variations in pose or movement can cause noticeable cha
arXiv:2602.12755v2 Announce Type: replace Abstract: Diffusion-based image generators are promising priors for ill-posed inverse problems like sparse-view X-ray Computed Tomography (CT). As most studie
arXiv:2604.19923v1 Announce Type: new Abstract: We introduce UniCon3R (Unified Contact-aware 3D Reconstruction), a unified feed-forward framework for online human-scene 4D reconstruction from monocula
arXiv:2604.20318v1 Announce Type: new Abstract: Composed image retrieval, multi-turn composed image retrieval, and composed video retrieval all share a common paradigm: composing the reference visual
arXiv:2604.20473v1 Announce Type: new Abstract: Existing Video Large Language Models (Video LLMs) struggle with complex video understanding, exhibiting limited reasoning capabilities and potential hal
arXiv:2604.19945v1 Announce Type: new Abstract: In this paper, we investigate the problem of how to effectively master tool-use to solve complex visual reasoning tasks for Multimodal Large Language Mo
arXiv:2604.19858v1 Announce Type: new Abstract: We present Wan-Image, a unified visual generation system explicitly engineered to paradigm-shift image generation models from casual synthesizers into p
arXiv:2604.20213v1 Announce Type: new Abstract: Accurate segmentation of maxillary sinus in panoramic X-ray images is essential for dental diagnosis and surgical planning; however, this task remains r
arXiv:2604.20574v1 Announce Type: new Abstract: Purpose: Gaze-following, the task of inferring where individuals are looking, has been widely studied in computer vision, advancing research in visual a
arXiv:2604.20190v1 Announce Type: new Abstract: Wildfire monitoring requires timely, actionable situational awareness from airborne platforms, yet existing aerial visual question answering (VQA) bench
arXiv:2604.20289v1 Announce Type: new Abstract: Real-time world simulation is becoming a key infrastructure for scalable evaluation and online reinforcement learning of autonomous driving systems. Rec
arXiv:2604.20350v1 Announce Type: new Abstract: Despite significant progress in Multi-modal Large Language Models (MLLMs), their clinical reasoning capacity for multi-modal diagnosis remains largely u
arXiv:2604.18721v1 Announce Type: cross Abstract: Visual state-space models (SSMs) are increasingly promoted as efficient alternatives to Vision Transformers, yet their practical advantages remain unc
arXiv:2604.18988v1 Announce Type: new Abstract: Multimodal empathetic response generation (MERG) aims to generate emotionally engaging and empathetic responses based on users' multimodal contexts. Exi
arXiv:2604.19715v1 Announce Type: new Abstract: Distribution networks with high penetration of Distributed Energy Resources (DERs) increasingly rely on communication networks to coordinate grid-intera
arXiv:2604.18980v1 Announce Type: new Abstract: Reducing the number of Gaussian-tile pairs is one of the most promising approaches to improve 3D Gaussian Splatting (3D-GS) rendering speed on GPUs. How
arXiv:2510.14630v2 Announce Type: replace Abstract: We introduce Representation Tokenizer (RepTok), a generative modeling framework that represents an image using a single continuous latent token obta
arXiv:2604.19233v1 Announce Type: new Abstract: Deep learning-based object detectors have achieved remarkable success across numerous computer vision applications, yet they continue to struggle with s
arXiv:2604.18961v1 Announce Type: cross Abstract: This paper presents an AI-enabled cascaded hybrid vision/force control framework for tendon-driven aerial continuum manipulators based on constant-str
arXiv:2604.19386v1 Announce Type: new Abstract: Composed Image Retrieval (CIR) has attracted significant attention due to its flexible multimodal query method, yet its development is severely constrai
arXiv:2604.18713v1 Announce Type: new Abstract: Automated 3D segmentation of prostate lesions from biparametric MRI (bp-MRI) is essential for reliable algorithmic analysis, but achieving high precisio