On the Uphill Battle of Image frequency Analysis
arXiv:2604.07563v1 Announce Type: new Abstract: This work is a follow up on the newly proposed clustering algorithm called The Inverse Square Mean Shift Algorithm. In this paper a special case of algo
Knowledge catalogue
arXiv:2604.07563v1 Announce Type: new Abstract: This work is a follow up on the newly proposed clustering algorithm called The Inverse Square Mean Shift Algorithm. In this paper a special case of algo
arXiv:2604.08031v1 Announce Type: cross Abstract: Most Human-Machine Interaction (HMI) research overlooks the maneuvering needs of passengers in autonomous driving (AD). Natural language offers an int
arXiv:2512.03532v2 Announce Type: replace Abstract: Generalizing open-vocabulary 3D instance segmentation (OV-3DIS) to diverse, unstructured, and mesh-free environments is crucial for robotics and AR/
arXiv:2604.08266v1 Announce Type: new Abstract: Leveraging the general world knowledge of Large Language Models (LLMs) holds significant promise for improving the ability of autonomous driving systems
arXiv:2604.08238v1 Announce Type: new Abstract: The increasing adaptation of vision models across domains, such as satellite imagery and medical scans, has raised an emerging privacy risk: models may
arXiv:2604.08110v1 Announce Type: new Abstract: Training-free open-vocabulary semantic segmentation(TF-OVSS) has recently attracted attention for its ability to perform dense prediction by leveraging
arXiv:2604.08461v1 Announce Type: new Abstract: Open-Vocabulary Segmentation (OVS) aims to segment image regions beyond predefined category sets by leveraging semantic descriptions. While CLIP based a
arXiv:2512.09665v2 Announce Type: replace Abstract: We address the problem of fair classification in settings where data is scarce and unbalanced across demographic groups. Such low-data regimes are c
arXiv:2602.06912v2 Announce Type: replace Abstract: Unsupervised segmentation from self-supervised ViT patches holds promise but lacks robustness: multi-object scenes confound saliency cues, and low-s
arXiv:2604.07901v1 Announce Type: new Abstract: 360 video object segmentation (360VOS) aims to predict temporally-consistent masks in 360 videos, offering full-scene coverage, benefiting applications,
arXiv:2604.07912v1 Announce Type: new Abstract: Finding parking consumes a disproportionate share of food delivery time, yet no system addresses precise parking-spot selection relative to merchant ent
arXiv:2604.08538v1 Announce Type: new Abstract: AI agents are changing the requirements for document parsing. What matters is semantic correctness: parsed output must preserve the structure and
arXiv:2506.17212v2 Announce Type: replace Abstract: Articulated objects are common in the real world, yet modeling their structure and motion remains a challenging task for 3D reconstruction methods.
arXiv:2604.07427v1 Announce Type: new Abstract: Modern text-to-image (T2I) models generate high-fidelity visuals but remain indifferent to individual user preferences. While existing reward models opt
arXiv:2604.08395v1 Announce Type: new Abstract: Recent advances in Vision-Language Models (VLMs) have greatly enhanced the integration of visual perception and linguistic reasoning, driving rapid prog
arXiv:2604.08503v1 Announce Type: new Abstract: Recent advances in generative video modeling, driven by large-scale datasets and powerful architectures, have yielded remarkable visual realism. However
arXiv:2604.07230v2 Announce Type: replace Abstract: Achieving physically accurate object manipulation in image editing is essential for its potential applications in interactive world models. However,
arXiv:2603.23286v3 Announce Type: replace Abstract: Physical knot classification is a fine-grained task in which the intended cue is rope crossing structure, but high accuracy may still come from appe
arXiv:2503.09640v2 Announce Type: replace-cross Abstract: Rendering realistic human-object interactions (HOIs) from sparse-view inputs is a challenging yet crucial task for various real-world applicat
arXiv:2503.24135v3 Announce Type: replace Abstract: Weakly supervised object localization (WSOL) methods allow training models to classify images and localize ROIs. WSOL only requires low-cost image-c
arXiv:2604.07779v1 Announce Type: new Abstract: Pathology foundation models (FMs) have become central to computational histopathology, offering strong transfer performance across a wide range of diagn
arXiv:2604.02073v2 Announce Type: replace Abstract: Universal multimodal embedding (UME) maps heterogeneous inputs into a shared retrieval space with a single model. Recent approaches improve UME by g
arXiv:2604.08340v1 Announce Type: new Abstract: While Vision-Language Models (VLMs) have achieved remarkable progress in static visual understanding, their deployment in complex 3D embodied environmen
arXiv:2604.08125v1 Announce Type: new Abstract: Human-like multimodal reaction generation is essential for natural group interactions between humans and embodied AI. However, existing approaches are l
arXiv:2604.08272v1 Announce Type: new Abstract: Deep image prior (DIP) is an unsupervised deep learning framework that has been successfully applied to a variety of inverse imaging problems. However,
arXiv:2502.02514v5 Announce Type: replace Abstract: Image AutoRegressive generation has emerged as a new powerful paradigm with image autoregressive models (IARs) matching state-of-the-art diffusion m
arXiv:2604.08037v1 Announce Type: cross Abstract: Talking-head generation has advanced rapidly with diffusion-based generative models, but training usually depends on centralized face-video and speech
arXiv:2512.18662v2 Announce Type: replace-cross Abstract: End-to-end (E2E) autonomous driving models that take only camera images as input and directly predict a future trajectory are appealing for th
arXiv:2512.01236v2 Announce Type: replace Abstract: Personalized generation models for a single subject have demonstrated remarkable effectiveness, highlighting their significant potential. However, w
arXiv:2604.08502v1 Announce Type: new Abstract: Class Activation Mapping (CAM) methods are widely used to generate visual explanations for deep learning classifiers in medical imaging. However, existi
arXiv:2512.06774v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has become a leading representation for high-fidelity 3D assets, yet protecting these assets via digital watermarking r
arXiv:2505.24848v4 Announce Type: replace Abstract: To enable egocentric contextual AI in always-on smart glasses, it is crucial to be able to keep a record of the user's interactions with the world,
arXiv:2505.24499v2 Announce Type: replace Abstract: Generating high-quality Scalable Vector Graphics (SVGs) is challenging for Large Language Models (LLMs), as it requires advanced reasoning for struc
arXiv:2604.07882v1 Announce Type: new Abstract: Reconstructing non-rigid objects with physical plausibility remains a significant challenge. Existing approaches leverage differentiable rendering for p
arXiv:2503.02537v4 Announce Type: replace Abstract: Diffusion models have achieved remarkable progress across various visual generation tasks. However, their performance significantly declines when ge
arXiv:2604.07884v1 Announce Type: new Abstract: High-fidelity generative models are increasingly needed in privacy-sensitive scenarios, where access to data is severely restricted due to regulatory an
arXiv:2604.07765v1 Announce Type: new Abstract: Earth Observation (EO) systems are essentially designed to support domain experts who often express their requirements through vague natural language ra
arXiv:2604.08282v1 Announce Type: new Abstract: Radar perception models are trained with different inputs, from range-Doppler spectra to sparse point clouds. Dense spectra are assumed to outperform sp
arXiv:2604.08536v1 Announce Type: new Abstract: We introduce RewardFlow, an inversion-free framework that steers pretrained diffusion and flow-matching models at inference time through multi-reward La
arXiv:2604.07774v1 Announce Type: cross Abstract: This paper focuses on embodied task planning, where an agent acquires visual observations from the environment and executes atomic actions to accompli
arXiv:2604.08034v1 Announce Type: new Abstract: Image registration is a fundamental task that aligns anatomical structures between images. While CNNs perform well, they lack rotation equivariance - a
arXiv:2505.17732v2 Announce Type: replace Abstract: Accurate, fast, and reliable 3D perception is essential for autonomous driving. Recently, bird's-eye view (BEV)-based perception approaches have eme
arXiv:2604.07890v1 Announce Type: new Abstract: Highly multiplexed microscopy enables rich spatial characterization of tissues at single-cell resolution, yet most analyses rely on two-dimensional sect
arXiv:2604.07994v1 Announce Type: new Abstract: Transformer-based approaches have revolutionized image super-resolution by modeling long-range dependencies. However, the quadratic computational comple
arXiv:2604.08542v1 Announce Type: new Abstract: This paper addresses the task of large-scale 3D scene reconstruction from long video sequences. Recent feed-forward reconstruction models have shown pro
arXiv:2604.08366v1 Announce Type: cross Abstract: Large-scale deep learning models for physical AI applications depend on diverse training data collection efforts. These models and correspondingly, th
arXiv:2604.07990v1 Announce Type: new Abstract: The convergence of 3D geometric perception and video synthesis has created an unprecedented demand for large-scale video data that is rich in both seman
arXiv:2604.08211v1 Announce Type: new Abstract: Modern multimodal generators can now produce scientific figures at near-publishable quality, creating a new challenge for visual forensics and research
arXiv:2604.03134v2 Announce Type: replace Abstract: Few-Shot Medical Image Segmentation (FSMIS) aims to segment novel object classes in medical images using only minimal annotated examples, addressing
arXiv:2604.08008v1 Announce Type: new Abstract: Retrieving rare and safety-critical driving scenarios from large-scale datasets is essential for building robust autonomous driving (AD) systems. As dat
arXiv:2604.08532v1 Announce Type: new Abstract: Large-scale multi-view reconstruction models have made remarkable progress, but most existing approaches still rely on fully supervised training with gr
arXiv:2604.08147v1 Announce Type: cross Abstract: Recent advances in audio-visual representation learning have shown the value of combining contrastive alignment with masked reconstruction. However, j
arXiv:2509.26036v3 Announce Type: replace Abstract: While Contrastive Language-Image Pretraining (CLIP) excels at zero-shot tasks by aligning image and text embeddings, its performance in few-shot cla
arXiv:2604.07936v1 Announce Type: new Abstract: Stain variability is a pervasive source of distribution shift and potential shortcut learning in renal pathology AI. We ask whether lupus nephritis glom
arXiv:2604.08544v1 Announce Type: cross Abstract: Robotic manipulation with deformable objects represents a data-intensive regime in embodied learning, where shape, contact, and topology co-evolve in
arXiv:2604.07477v1 Announce Type: new Abstract: For applications including facial identification, forensic analysis, photographic improvement, and medical imaging diagnostics, facial image deblurring
arXiv:2504.13378v2 Announce Type: replace-cross Abstract: Generating high-quality, photorealistic textures for 3D human avatars remains a fundamental yet challenging task in computer vision and multim
arXiv:2604.05933v3 Announce Type: replace Abstract: Ultrasound perception typically requires multiple scan views through probe movement to reduce diagnostic ambiguity, mitigate acoustic occlusions, an
arXiv:2512.23365v3 Announce Type: replace Abstract: The rapid progress of Multimodal Large Language Models (MLLMs) has unlocked the potential for enhanced 3D scene understanding and spatial reasoning.
arXiv:2604.07923v1 Announce Type: new Abstract: Dynamic urban environments are often captured by cameras placed at spatially separated locations with little or no view overlap. However, most existing