Data Pyramid for Embodied Manipulation
arXiv:2607.24744v1 Announce Type: cross Abstract: Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require d
Knowledge catalogue
arXiv:2607.24744v1 Announce Type: cross Abstract: Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require d
arXiv:2607.23921v1 Announce Type: new Abstract: Answer grounding in visual question answering aims to locate the region from a given natural language question associated with the visual content of an
arXiv:2501.13335v4 Announce Type: replace Abstract: We introduce a novel framework for modeling high-fidelity, animatable 3D human avatars from motion-blurred monocular video inputs. Motion blur is pr
arXiv:2607.24554v1 Announce Type: cross Abstract: Advancing multimodal retrieval-augmented generation (RAG) for complex document understanding presents a formidable dual dilemma of accuracy and effici
arXiv:2607.22770v1 Announce Type: cross Abstract: Although artificial intelligence (AI) has shown promising performance in several medical tasks, accurate dementia etiology diagnosis with AI remains c
arXiv:2607.24579v1 Announce Type: cross Abstract: When computing sub/super-level-set persistent homology (PH), the effect of noise may introduce millions of (short-lived) topological generators, prese
arXiv:2607.24159v1 Announce Type: cross Abstract: Generalizable robot manipulation requires policies that can anticipate how visual scenes evolve while executing language instructions. While recent Vi
arXiv:2607.23962v1 Announce Type: new Abstract: Autonomous vehicles (AVs) depend on Global Navigation Satellite Systems (GNSS) for localization and navigation, making them vulnerable to spoofing attac
arXiv:2607.22687v1 Announce Type: new Abstract: Foundation vision models trained on natural images transfer to medical tasks without domain pre-training, but volumetric classification requires aggrega
arXiv:2607.23464v1 Announce Type: cross Abstract: Deep learning dominates polarimetric synthetic aperture radar (PolSAR) image classification, with Mamba architectures serving as favorable backbones d
arXiv:2607.23070v1 Announce Type: new Abstract: Food segmentation is essential for applications such as intelligent catering, dietary assessment, and recommendation. However, existing benchmarks fail
arXiv:2607.23132v1 Announce Type: new Abstract: Assessing the severity of a traffic accident scenario is important to decide which emergency service to dispatch. Missing an ambulance dispatch on a ped
arXiv:2607.24721v1 Announce Type: new Abstract: With the growth of gaming, animation, and virtual reality industries, the demand for efficient generation of stylized 3D assets is rapidly increasing. H
arXiv:2607.22801v1 Announce Type: cross Abstract: Underwater image enhancement is challenged by spatially non-uniform, wavelength-dependent attenuation. Propagation distance and wavelength govern this
arXiv:2607.22705v1 Announce Type: new Abstract: Object-centric learning aims to represent scenes as objects whose properties can be reused in new combinations. Existing evaluations usually score segme
arXiv:2607.24210v1 Announce Type: new Abstract: In clinical oncology studies, metastatic cancer is commonly evaluated using 'Response Evaluation Criteria in Solid Tumors' (RECIST), in which the diamet
arXiv:2607.23994v1 Announce Type: new Abstract: In this work, we investigate a previously unexplored architectural dimension for infrared small target detection: the organization of effective receptiv
arXiv:2607.23908v1 Announce Type: new Abstract: High quality reference data remain a critical bottleneck for crop-type mapping at any spatial and temporal scale. Operational systems such as WorldCerea
arXiv:2607.24313v1 Announce Type: cross Abstract: Marine life monitoring is limited by strict energy constraints, poor underwater connectivity, and the high cost of transmitting raw multimodal data fr
arXiv:2406.13128v2 Announce Type: replace Abstract: Due to the intricate structure of vascular trees, minor segmentation errors can significantly alter connectivity patterns and increase variability i
arXiv:2602.00115v2 Announce Type: replace Abstract: This paper introduces a novel asynchronous, event-driven algorithm for real-time detection of small event clusters in event camera data. Similar to
arXiv:2607.24194v1 Announce Type: new Abstract: Online platforms increasingly rely on automated age estimation systems to enforce minimum-age policies. Focusing on vision-based models designed for thi
arXiv:2506.03168v2 Announce Type: replace Abstract: Amid the challenges posed by global population growth and climate change, traditional agricultural Internet of Things (IoT) systems is currently und
arXiv:2607.22734v1 Announce Type: new Abstract: Satellite-derived Land Surface Temperature (LST) provides spatially comprehensive data that ground stations cannot match. However, its utility is freque
arXiv:2602.14021v2 Announce Type: replace Abstract: Reconstructing and tracking dynamic 3D scenes is a fundamental challenge in computer vision. Existing methods typically decouple geometry from motio
arXiv:2607.24522v1 Announce Type: cross Abstract: While on-policy distillation (OPD) effectively addresses sparse rewards and exposure bias in large language model post-training, its extension to flow
arXiv:2607.22698v1 Announce Type: new Abstract: Perception under adverse weather remains a critical bottleneck for reliable autonomous driving, yet existing benchmarks lack the systematic multi-modal
arXiv:2607.22765v1 Announce Type: cross Abstract: Electron microscopy enables nanoscale cellular visualization but faces a trade-off between imaging resolution and acquisition speed. Existing learning
arXiv:2607.23542v1 Announce Type: new Abstract: Efficient border control is becoming a significant global challenge, mainly due to severe congestion and extended passenger waiting times. To mitigate t
arXiv:2607.22719v1 Announce Type: new Abstract: Multiplicative Gamma noise is a signal-dependent degradation in coherent imaging; synthetic aperture radar (SAR) despeckling is its most prominent real-
arXiv:2607.22847v1 Announce Type: new Abstract: Human gaze does more than point to visual targets; it serves as a subtle indicator of social intent within static images, whereas standard models typica
arXiv:2607.23917v1 Announce Type: new Abstract: We introduce a novel learning problem: decoding gaze into natural language descriptions of human goals across diverse visual tasks. Unlike prior work, w
arXiv:2508.09478v2 Announce Type: replace Abstract: In this work, we present GazeLT, a human visual attention integration-disintegration approach for long-tailed disease classification. A radiologist'
arXiv:2601.01856v3 Announce Type: replace Abstract: Feature-based anomaly detection is widely adopted in industrial inspection due to the strong representational power of large pre-trained vision enco
arXiv:2607.22733v1 Announce Type: new Abstract: We investigate whether a generative model can supply useful synthetic motor-imagery (MI) electroencephalography (EEG) trials that improve the accuracy o
arXiv:2607.22772v1 Announce Type: cross Abstract: Diffusion models provide strong generative capabilities for video compression at ultra-low bitrates. Existing diffusion-based video codecs adapt base
arXiv:2607.24403v1 Announce Type: new Abstract: Feed-forward 3D Gaussian Splatting (3DGS) enables scalable scene reconstruction without per-scene optimization, yet produces dense Gaussians that are co
arXiv:2607.23530v1 Announce Type: new Abstract: Bounding boxes are fundamental for object localization in visual detection tasks. Among them, oriented bounding boxes are widely used in visual detectio
arXiv:2607.23687v1 Announce Type: new Abstract: Parametric models of the human head are essential tools traditionally used in computer vision and graphics for animation, rendering, and reconstruction.
arXiv:2607.23657v1 Announce Type: new Abstract: Articulated portrait mesh estimation is fundamental to 3D understanding, avatar generation, and immersive interaction. Existing approaches primarily rel
arXiv:2503.14736v3 Announce Type: replace Abstract: Photorealistic and animatable hand avatars are essential for applications such as AR/VR, gaming, and telepresence. Recent 3D Gaussian Splatting (3DG
arXiv:2607.23861v1 Announce Type: new Abstract: We present DynHair, a novel method for tracking and modeling dynamic hair for human head avatars. From video input, we reconstruct a dynamic head avatar
arXiv:2607.23292v1 Announce Type: cross Abstract: Deep learning-based visual-infrared fused face detection models are increasingly deployed across a wide range of applications, yet they remain suscept
arXiv:2603.14807v3 Announce Type: replace Abstract: LLM-based agents have demonstrated impressive zero-shot performance in vision-language navigation (VLN) tasks. However, most zero-shot methods prima
arXiv:2607.24364v1 Announce Type: new Abstract: Predicting spatial gene expression from routine hematoxylin and eosin (H&E) images provides a practical complement to experimental spatial transcriptomi
arXiv:2607.22703v1 Announce Type: new Abstract: Prostate cancer diagnosis with multiparametric MRI (mpMRI) is commonly based on PI-RADS assessment or binary classification, which suffer from subjectiv
arXiv:2607.22830v1 Announce Type: new Abstract: In visual storytelling, human performances are central to creative intent and narrative meaning. However, preserving human identity and performance whil
arXiv:2607.24422v1 Announce Type: new Abstract: This paper presents a summary of the Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data (AFMFR), held at the 2
arXiv:2607.24140v1 Announce Type: new Abstract: Image inpainting aims to recover missing regions while preserving structural consistency. We propose a non-parametric method without network training ba
arXiv:2510.02913v2 Announce Type: replace Abstract: Vision-language models such as CLIP demonstrate impressive zero-shot generalization but remain highly vulnerable to adversarial attacks. Prior adver
arXiv:2607.24727v1 Announce Type: new Abstract: Background. Pediatric musculoskeletal trauma represents up to 18% of pediatric ED visits, yet diagnosis still depends on ionizing radiography. Cumulativ
arXiv:2508.13287v3 Announce Type: replace-cross Abstract: 3D Gaussian Splatting (3DGS) has recently gained popularity for efficient scene rendering by representing scenes as explicit sets of anisotrop
arXiv:2607.22780v1 Announce Type: cross Abstract: Faithful inverse rendering requires visibility and indirect radiance to explain secondary illumination and inter-reflection, yet rasterization-oriente
arXiv:2607.24431v1 Announce Type: new Abstract: Camera-only 4D occupancy forecasting enables autonomous vehicles to predict future 3D semantic scenes solely from historical multi-view images, which is
arXiv:2607.23078v1 Announce Type: new Abstract: Longitudinal medical imaging captures temporal evolution of lesions, yet extracting the underlying dynamical parameters governing this evolution remains
arXiv:2607.23371v1 Announce Type: cross Abstract: Vascular segmentation is a standard procedure for clinical diagnosis, yet the specific visual features determining model decisions remain poorly under
arXiv:2607.23588v1 Announce Type: new Abstract: Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high
arXiv:2607.22783v1 Announce Type: cross Abstract: Recent advances in conventional and learning-based image coding have increased the demand for benchmark datasets that support fine-grained assessment
arXiv:2607.23704v1 Announce Type: cross Abstract: The deployment of embodied agents in self-driving laboratories could accelerate scientific discovery, yet their reliability is constrained by the irre
arXiv:2603.25629v2 Announce Type: replace Abstract: While language reasoning models excel in many tasks, visual reasoning remains challenging for current large multimodal models (LMMs). As a result, m