Synthesis: Arxiv-Cs-Cv
Auto-generated synthesis of 874 entries about arxiv-cs-cv
Knowledge catalogue
Auto-generated synthesis of 874 entries about arxiv-cs-cv
arXiv:2608.10107v1 Announce Type: new Abstract: Emerging World-Action Models (WAMs) have demonstrated promising performance in autonomous driving by jointly modeling future driving scene evolution and
arXiv:2608.10203v1 Announce Type: new Abstract: Despite the success of convolutional neural networks in image classification tasks and their general application in multi-modal models, their susceptibi
arXiv:2608.10978v1 Announce Type: new Abstract: Optical music recognition (OMR) transcribes music scores into digital formats. While the field has advanced significantly on monophonic and piano-form s
arXiv:2608.10411v1 Announce Type: new Abstract: We present a theory of textured appearance of optically rough surfaces based on wave optics, emphasizing the role of texture for passive depth from focu
arXiv:2608.11205v1 Announce Type: new Abstract: Frechet distance has recently emerged as an effective distribution-level objective for generator post-training, complementing the conventional sample-le
arXiv:2608.11123v1 Announce Type: new Abstract: Augmentation can corrupt a training example when an image and its annotations receive different random changes. A crop must use the same coordinates for
arXiv:2608.09989v1 Announce Type: cross Abstract: There has been a tremendous amount of image processing and machine learning research to measure and classify disease progression from live optical coh
arXiv:2608.09993v1 Announce Type: cross Abstract: Registration between cone-beam computed tomography (CBCT) and intraoral scans (IOS) is essential for patient-specific surgical planning. However, disp
arXiv:2608.10744v1 Announce Type: new Abstract: 4D generation synthesizes dynamic 3D scenes from conditions such as text or images. Existing methods either reconstruct generated RGB videos with a sepa
arXiv:2608.10600v1 Announce Type: cross Abstract: Skill abstraction---the process of learning reusable and temporally extended behaviors---has emerged as a key paradigm for improving sample efficiency
arXiv:2608.10479v1 Announce Type: new Abstract: Latent diffusion models have recently advanced video frame interpolation by synthesizing intermediate frames between input images. However, handling lar
arXiv:2608.10680v1 Announce Type: new Abstract: Visible-infrared object detection relies on complementary RGB and thermal cues, but its performance is often degraded by cross-modal spatial misalignmen
arXiv:2608.11074v1 Announce Type: new Abstract: Evaluating detailed image captions from Vision-Language Models (VLMs) requires going beyond surface-level semantic similarity. Reference-based metrics (
arXiv:2608.11203v1 Announce Type: new Abstract: This paper presents a self-supervised representation learning framework for understanding 3D skeleton-based human motion in soccer, using future motion
arXiv:2608.10345v1 Announce Type: new Abstract: Free-viewpoint 3D scene media is increasingly important for immersive applications, yet practical capture often suffers from severe view sparsity and mo
arXiv:2608.11150v1 Announce Type: new Abstract: While 3D Gaussian Splatting (3DGS) has advanced open vocabulary scene understanding, existing methods remain confined to explicit queries. They struggle
arXiv:2608.10278v1 Announce Type: new Abstract: Spatial understanding is fundamental to embodied intelligence, underpinning applications such as robotic manipulation, embodied navigation, and autonomo
arXiv:2608.10677v1 Announce Type: new Abstract: Professionals across medicine, engineering, finance, manufacturing, and the sciences often make consequential decisions from charts. Existing chart benc
arXiv:2608.10712v1 Announce Type: new Abstract: 3D scene reconstruction, modeling, and rendering are highly relevant for numerous tasks, and 3D Gaussian splatting has become a standard choice in this
arXiv:2608.10885v1 Announce Type: new Abstract: Pulmonary nodule malignancy prediction typically depends on image-trained specialist deep learning (DL) models that require substantial annotated imagin
arXiv:2608.11093v1 Announce Type: cross Abstract: Cross-view feature matching aims to establish reliable correspondences across images with large viewpoint variations. Over the past decade, the field
arXiv:2608.10173v1 Announce Type: new Abstract: Most radiotherapy dose-prediction models use only CT images and anatomical structures, although intensity-modulated proton therapy (IMPT) dose also depe
arXiv:2512.22525v2 Announce Type: replace Abstract: Recently unified generation and editing models have achieved remarkable success with their impressive performance. These models rely mainly on text
arXiv:2608.10413v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for end-to-end autonomous driving by enabling unified reasoning across
arXiv:2608.10500v1 Announce Type: new Abstract: Creating photorealistic and temporally coherent animatable human avatars from RGB videos remains challenging. Current methods struggle to capture realis
arXiv:2608.10435v1 Announce Type: new Abstract: Diffusion models have been widely explored in protein backbone generation due to their powerful generation capabilities.However, in today's AI-driven bi
arXiv:2608.10796v1 Announce Type: new Abstract: Understanding both expressed and evoked emotions is critical for multimodal large language models (MLLMs) to achieve comprehensive affect-aware interact
arXiv:2509.26087v5 Announce Type: replace Abstract: In perception for automated vehicles, safety is critical not only for the driver but also for other agents in the scene, particularly vulnerable roa
arXiv:2608.10684v1 Announce Type: new Abstract: Multi-oriented text is ubiquitous in real-world scenes and remains a major challenge for scene text recognition (STR). Existing rotation-aware methods e
arXiv:2608.10756v1 Announce Type: cross Abstract: Embodied mobile manipulation requires language, visual observations, three-dimensional scene structure, and action feasibility to be aligned before ex
arXiv:2608.10343v1 Announce Type: new Abstract: While deep learning-based denoising has become widely adopted in low-dose CT, conventional models use generic architectures designed for natural images,
arXiv:2608.10801v1 Announce Type: new Abstract: Spatio-temporal PV data are essential for understanding adoption processes in off-grid regions, yet such data remain largely unavailable. Automated segm
arXiv:2608.11096v1 Announce Type: new Abstract: Learned image compression (LIC) has achieved impressive rate-distortion performance. However, existing methods remain highly vulnerable to packet loss,
arXiv:2507.21606v2 Announce Type: replace Abstract: The success of visual tracking has been largely driven by datasets with manual box annotations. However, these box annotations require tremendous hu
arXiv:2608.10764v1 Announce Type: new Abstract: Counterfactual video understanding evaluates whether models grasp physical and commonsense regularities. However, existing multiple-choice question (MCQ
arXiv:2506.11142v3 Announce Type: replace Abstract: Semi-supervised semantic segmentation (SSSS) faces persistent challenges in effectively leveraging unlabeled data, such as ineffective utilization o
arXiv:2608.10860v1 Announce Type: cross Abstract: World-action models (WAMs) predict the future to act better, but nearly all of them predict only RGB latents, trained purely for pixel reconstruction,
arXiv:2608.10396v1 Announce Type: new Abstract: Transforming table-form documents into machine-processable records requires recovering not only their visible content but also the multilevel structure
arXiv:2608.11076v1 Announce Type: new Abstract: Automated lesion segmentation in whole-body PET/CT imaging can assist clinicians with cancer detection, staging, and treatment planning across radiotrac
arXiv:2608.10317v1 Announce Type: new Abstract: We present TAR (Traffic Anomaly Reasoning) and TAR-Bench datasets, resources for training and evaluating video-language models beyond anomaly detection.
arXiv:2608.10602v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has recently enabled real-time novel view synthesis with impressive quality. However, it struggles to recover accurate surf
arXiv:2608.10426v1 Announce Type: new Abstract: Open-vocabulary remote sensing segmentation has recently emerged as a promising paradigm that enables pixel-level recognition of arbitrary categories sp
arXiv:2608.10886v1 Announce Type: new Abstract: Robots operating in human environments need memories that capture not only what objects exist and where, but also how people use them over time and how
arXiv:2608.10723v1 Announce Type: new Abstract: Vision Transformers underperform convolutional networks when training data is scarce, and distilling convolutional inductive biases from a CNN teacher i
arXiv:2603.25467v3 Announce Type: replace Abstract: Vision-Language Models (VLMs) are powerful open-set reasoners, yet their direct use as anomaly detectors in video surveillance is fragile: without c
arXiv:2608.10938v1 Announce Type: new Abstract: Despite substantial progress in visual localization, from scene coordinate regression to direct camera pose regression, achieving both robust generaliza
arXiv:2608.10995v1 Announce Type: new Abstract: Existing diffusion-based methods have recently made significant progress in image dehazing. However, they typically neglect the physics of haze formatio
arXiv:2512.05746v3 Announce Type: replace Abstract: Diffusion models have demonstrated significant applications in the field of image generation. However, their high computational and memory costs pos
arXiv:2608.11051v1 Announce Type: new Abstract: As robots increasingly operate in human-populated environments, anticipating human intentions is essential for enabling proactive and socially aware beh
arXiv:2608.10181v1 Announce Type: new Abstract: Computer vision saliency models predict where people will look, one map per image, and a billion-dollar predicted-attention industry sells those maps in
arXiv:2608.10001v1 Announce Type: cross Abstract: Continuous parameterization of medical data has emerged as a powerful paradigm for resolution-independent image representation. While Implicit Neural
arXiv:2608.10724v1 Announce Type: new Abstract: Multimodal object detection proves effective in remote sensing, especially the RGB-Infrared paradigm. The parallel feature extractors provide rich multi
arXiv:2607.14945v2 Announce Type: replace Abstract: State-of-the-art flow based text-to-image (T2I) models exhibit remarkable generative abilities but remain vulnerable to producing unsafe content. Pr
arXiv:2608.11135v1 Announce Type: new Abstract: Camouflaged object detection (COD) aims to segment objects that are visually concealed in their surroundings and has attracted increasing attention in r
arXiv:2608.10566v1 Announce Type: cross Abstract: How many directions does a neural representation use to encode a concept? A common answer repeatedly erases probe directions and reports the stopping
arXiv:2608.11077v1 Announce Type: new Abstract: Feed-forward Gaussian reconstruction has recently emerged as an efficient approach for driving scene reconstruction. However, prevailing LiDAR-based met
arXiv:2509.06191v2 Announce Type: replace-cross Abstract: Recent 3D generative models, which are capable of generating full object shapes from just a few images, now open up new opportunities in robot
arXiv:2608.10057v1 Announce Type: new Abstract: We introduce LEGO for advanced open-vocabulary scene understanding. Beyond basic concept recognition, its core innovation lies in capturing the intrinsi
arXiv:2608.10429v1 Announce Type: new Abstract: Deep learning models that synthesize PET from CT or MRI can reduce patient dose and scanner demand, but are typically optimized with global losses such