ShapeUP: Scalable Image-Conditioned 3D Editing
arXiv:2602.05676v2 Announce Type: replace Abstract: Recent advancements in 3D foundation models have enabled the generation of high-fidelity assets, yet precise 3D manipulation remains a significant c
Knowledge catalogue
arXiv:2602.05676v2 Announce Type: replace Abstract: Recent advancements in 3D foundation models have enabled the generation of high-fidelity assets, yet precise 3D manipulation remains a significant c
arXiv:2604.24000v1 Announce Type: cross Abstract: The Laplacian operator transforms the image into its Laplacian field, which usually is sparse and satisfies a stable distribution. On the other hand,
arXiv:2506.18493v2 Announce Type: replace Abstract: Customizing image generation remains a core challenge in controllable image synthesis. For single-concept generation, maintaining both identity pres
arXiv:2604.23996v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) has become a prevalent backbone for large vision-language models (VLMs), yet how modality-specific signals should guide expert
arXiv:2508.05417v3 Announce Type: replace Abstract: Slot Attention (SA) lies at the heart of mainstream Object-Centric Learning (OCL). Image features can be aggregated into object-level representation
arXiv:2604.23662v1 Announce Type: new Abstract: The increasing global deployment of solar photovoltaic (PV) systems needs robust, scalable, and automated inspection technologies capable of detecting a
arXiv:2511.17092v4 Announce Type: replace Abstract: Articulated objects are ubiquitous in daily environments, and their 3D reconstruction holds great significance across various fields. However, exist
arXiv:2604.23551v1 Announce Type: new Abstract: Reconstructing realistic underwater scenes from underwater video remains a meaningful yet challenging task in the multimedia domain. The inherent spatio
arXiv:2401.07669v2 Announce Type: replace Abstract: Adapting CLIP for videos has gained popularity due to its semantic and rich representation. While CLIP is a good starting point, it typically underg
arXiv:2604.23309v1 Announce Type: new Abstract: Remote sensing image change captioning (RSICC) aims to describe the difference between two remote sensing images. While recent methods have explored vid
arXiv:2402.11789v5 Announce Type: replace-cross Abstract: Anomaly localization in images -- identifying regions that deviate from normal patterns -- is vital in applications such as medical diagnosis
arXiv:2512.10959v2 Announce Type: replace Abstract: We introduce StereoSpace, a diffusion-based framework for monocular-to-stereo synthesis that models geometry purely through viewpoint conditioning,
arXiv:2601.20597v2 Announce Type: replace Abstract: Continual Text-to-Video Retrieval (CTVR) is a challenging multimodal continual learning setting, where models must incrementally learn new semantic
arXiv:2604.22899v1 Announce Type: new Abstract: Industrial anomaly detection based on RGB-3D multimodal data has emerged as a mainstream paradigm for intelligent quality inspection. However, existing
arXiv:2604.24459v1 Announce Type: new Abstract: Despite recent advances in text-to-image generation, models still struggle to accurately render prompt-specified text with correct spatial layout -- esp
arXiv:2602.19019v2 Announce Type: replace Abstract: Generative AI models pose a significant challenge to intellectual property (IP), as they can replicate unique artistic styles and concepts without a
arXiv:2604.01644v2 Announce Type: replace Abstract: Natural language provides an intuitive way to express spatial intent in geospatial applications. While existing localization methods often rely on d
arXiv:2604.24119v1 Announce Type: new Abstract: Topology reasoning is crucial for autonomous driving. Current methods primarily focus on instance-level learning for centerline detection, followed by a
arXiv:2604.24235v1 Announce Type: new Abstract: Touchless interaction with medical images is becoming increasingly important in the surgical field, where sterility and continuity of the operational wo
arXiv:2604.23094v1 Announce Type: new Abstract: The real-world adoption of portrait relighting is hindered by dataset domain gaps, camera sensitivity, and computational costs. We address these challen
arXiv:2601.02018v2 Announce Type: replace Abstract: Segment Anything Models (SAMs), known for their exceptional zero-shot segmentation performance, have garnered significant attention in the research
arXiv:2603.15941v2 Announce Type: replace Abstract: Automated diagnosis from chest computed tomography (CT) scans faces two persistent challenges in clinical deployment: distribution shift across acqu
arXiv:2603.11831v2 Announce Type: replace Abstract: The field of Computer-Aided Design (CAD) generation has made significant progress in recent years. Existing methods typically fall into two separate
arXiv:2604.23105v1 Announce Type: new Abstract: Deep learning drives major advances in autonomous driving (AD), where object detectors are central to perception. However, adversarial attacks pose sign
arXiv:2604.22904v1 Announce Type: cross Abstract: Gadoxetate disodium-enhanced MRI is essential for the detection and characterization of hepatocellular carcinoma. However, acquisition of the hepatobi
arXiv:2604.24763v1 Announce Type: new Abstract: Unified multimodal models typically rely on pretrained vision encoders and use separate visual representations for understanding and generation, creatin
arXiv:2507.04503v2 Announce Type: replace Abstract: Accurate localization using visual information is a critical yet challenging task, especially in urban environments where nearby buildings and const
arXiv:2604.23019v1 Announce Type: new Abstract: Accurate classification of tropical tree species from unoccupied aerial vehicle (UAV) imagery remains challenging due to high species diversity and stro
arXiv:2604.22846v1 Announce Type: new Abstract: The expanding ecosystem of pathology foundation models has produced powerful but fragmented tile-level representations, limiting their use in clinical t
arXiv:2604.23066v1 Announce Type: new Abstract: Urban flooding affects lives and infrastructure worldwide. Mapping inundation in complex urban environments from satellite imagery remains challenging d
arXiv:2604.23380v1 Announce Type: cross Abstract: Aligning denoising generative models with human preferences or verifiable rewards remains a key challenge. While policy-gradient online reinforcement
arXiv:2510.08618v2 Announce Type: replace-cross Abstract: Omni-modal large language models (OLLMs) offer a promising end-to-end solution for slide-enhanced speech recognition due to their inherent mul
arXiv:2604.23641v1 Announce Type: new Abstract: This paper introduces VDLF-Net, which attaches a compact VAE to a multi-scale CNN backbone. Latent vectors and softmax-gate support the backbone feature
arXiv:2604.22872v1 Announce Type: new Abstract: Autonomous vehicles (AVs) rely on real-time perception systems to understand road environments and ensure safe navigation. However, implementing reliabl
arXiv:2604.23799v1 Announce Type: new Abstract: Accurate whole-cell and nuclear segmentation is essential for precision pathology and spatial omics, yet routine hematoxylin and eosin (H&E) staining pr
arXiv:2512.07834v2 Announce Type: replace Abstract: Voxel art is a distinctive stylization widely used in games and digital media, yet automated generation from 3D meshes remains challenging due to co
arXiv:2604.23706v1 Announce Type: new Abstract: Histologic assessment of ulcerative colitis (UC) activity is an important endpoint in clinical trials and routine care, but manual grading with indices
arXiv:2604.22834v1 Announce Type: new Abstract: This paper presents webmcu-vision-web, a single-file, zero-install browser application for end-to-end TinyML vision model training and deployment on the
arXiv:2604.24718v1 Announce Type: new Abstract: Monocular RGB cameras mounted on drones are widely used for wildlife monitoring, yet most analytical pipelines remain confined to two-dimensional image
arXiv:2604.24764v1 Announce Type: new Abstract: Recent video foundation models demonstrate impressive visual synthesis but frequently suffer from geometric inconsistencies. While existing methods atte
arXiv:2604.23536v1 Announce Type: new Abstract: Diffusion models have achieved unprecedented success in text-aligned generation, largely driven by Classifier-Free Guidance (CFG). However, standard CFG
arXiv:2604.24479v1 Announce Type: new Abstract: Computer-Aided Design (CAD) models are defined by their construction history: a parametric recipe that encodes design intent. However, existing large-sc
arXiv:2604.23709v1 Announce Type: new Abstract: Single image dehazing is often constrained by a trade-off between restoration quality and computational efficiency. While efficient, CNN networks strugg
arXiv:2511.13211v2 Announce Type: replace Abstract: Despite recent advancements in 3D-text cross-modal alignment, existing state-of-the-art methods still struggle to align fine-grained textual semanti
arXiv:2604.22657v1 Announce Type: new Abstract: Accurate identification of individual farm animals in group-housed environments is a cornerstone of precision livestock management. However, current ind
arXiv:2512.13511v2 Announce Type: replace Abstract: Our objective is to build an embedding model that captures the nuanced relationship between a search query and candidate videos. We cover three aspe
arXiv:2604.22476v1 Announce Type: new Abstract: Disciplines such as business process management and process mining aid organizations by discovering insights about processes on the basis of recorded ev
arXiv:2602.23872v2 Announce Type: replace Abstract: To address the scale mismatch caused by large altitude variations in UAV visual place recognition, we propose a monocular vision-only altitude-adapt
arXiv:2604.22139v1 Announce Type: new Abstract: Reliable automated analysis of Optical Coherence Tomography (OCT) imaging is crucial for diagnosing retinal disorders but faces a critical barrier: the
arXiv:2604.22202v1 Announce Type: new Abstract: Symmetry detection is a fundamental problem in computer vision, and symmetries serve as powerful priors for downstream tasks. However, existing learning
arXiv:2604.22557v1 Announce Type: cross Abstract: The emergence of large-scale pretrained foundation models has transformed computer vision, enabling strong performance across diverse downstream tasks
arXiv:2510.10254v2 Announce Type: replace Abstract: Recent advances in large generative models have shown that simple autoregressive formulations, when scaled appropriately, can exhibit strong zero-sh
arXiv:2604.22280v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have emerged as a promising foundation for universal multimodal embeddings. Recent studies have shown that reas
arXiv:2604.22220v1 Announce Type: new Abstract: Digital image watermarking has advanced rapidly for copyright protection of generative AI, yet the comparatively limited progress in watermark attack te
arXiv:2604.22274v1 Announce Type: new Abstract: Open-vocabulary scene graph generation (SGG) aims to describe visual scenes with flexible and fine-grained relation phrases beyond a fixed predicate voc
arXiv:2604.22192v1 Announce Type: new Abstract: Chart-to-code generation demands strict visual precision and syntactic correctness from Vision-Language Models (VLMs). However, existing approaches are
arXiv:2604.21960v1 Announce Type: cross Abstract: Computed Tomography (CT) is a widely used imaging modality in medical and industrial applications. To limit radiation exposure and measurement time, t
arXiv:2604.22477v1 Announce Type: new Abstract: Neuron labeling assigns textual descriptions to internal units of deep networks. Existing approaches typically rely on highly activating examples, often
arXiv:2510.19592v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) demonstrate strong video understanding by attending to visual tokens relevant to textual queries. To direct
arXiv:2604.22331v1 Announce Type: new Abstract: This study analyses simulated and real-world implementations of depth-aware rover navigation, highlighting the transition from stereo vision to monocula