Layering Virtual Try-On
arXiv:2607.22924v1 Announce Type: new Abstract: In the real world, fashion is about layering: adding a jacket over a shirt, or a sequence of adding and removing layers, rather than just a single-layer
Knowledge catalogue
arXiv:2607.22924v1 Announce Type: new Abstract: In the real world, fashion is about layering: adding a jacket over a shirt, or a sequence of adding and removing layers, rather than just a single-layer
arXiv:2607.24184v1 Announce Type: new Abstract: Infrared small target detection (IRSTD) is important for low-altitude perception, unmanned-system warning, and security monitoring. However, weak target
arXiv:2607.22789v1 Announce Type: cross Abstract: Tracheostomy requires precise localization of the tracheal incision site; however, conventional manual palpation is subjective and often unreliable, w
arXiv:2607.22803v1 Announce Type: cross Abstract: Recovering the 6-DoF pose of the knee bones from a plain radiograph, given the patient's segmented pre-operative CT, turns a routine low-dose image in
arXiv:2607.23488v1 Announce Type: cross Abstract: Text-to-image diffusion models expose many inference-time sampling parameters, including prompts, negative prompts, classifier-free guidance scales, a
arXiv:2606.16586v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) remain unreliable on fine-grained visual perception, even when high-resolution inputs preserve the necessar
arXiv:2607.23883v1 Announce Type: new Abstract: In this paper, we examine the difficulties of using standard techniques for medical image classification due to long-tailed distributions (wherein rarer
arXiv:2607.24135v1 Announce Type: new Abstract: Single-image self-supervised denoising replaces unavailable clean targets with surrogate targets constructed from noisy observations. Its effectiveness
arXiv:2607.24002v1 Announce Type: new Abstract: Low-light image enhancement (LLIE) aims to improve image quality and clarity in diverse and demanding low-illumination environments. However, existing d
arXiv:2607.22707v1 Announce Type: new Abstract: Single-image reflection removal aims to recover a clean transmission layer from one image captured through glass. We study an explicit decomposition pip
arXiv:2411.15469v3 Announce Type: replace Abstract: Continual Learning (CL) aims to equip AI models with the ability to learn a sequence of tasks over time, without forgetting previously learned knowl
arXiv:2607.23937v1 Announce Type: new Abstract: Few-step distilled diffusion models generate high-quality images quickly, but often lose per-prompt diversity, producing near-identical samples across r
arXiv:2607.23608v1 Announce Type: new Abstract: The Action Research Arm Test (ARAT) is a widely-used upper limb outcome measure in neurorehabilitation, but its ordinal scoring is subjective and suffer
arXiv:2607.24224v1 Announce Type: new Abstract: Multi-modality data from different sensors provides rich complementary information for 3D perception, becoming an essential component in reliable autono
arXiv:2607.24424v1 Announce Type: new Abstract: Vision-language models commonly project all tokens produced by a pretrained vision encoder into a large language model. However, final-layer features ca
arXiv:2607.23504v1 Announce Type: new Abstract: Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to maintain long-horizon visual history for trajectory consistency wh
arXiv:2607.22890v1 Announce Type: cross Abstract: Domain Randomization (DR) is a standard technique for closing the Sim-to-Real gap, yet traditional DR pipelines rely on classical computer graphics re
arXiv:2607.22773v1 Announce Type: cross Abstract: Objective: We evaluated whether metric 3D geometry of neurosurgical operative exposure can be recovered from standard monocular operating-microscope i
arXiv:2607.24729v1 Announce Type: new Abstract: We introduce MicroZoom, a generative framework for gigapixel image synthesis at the microscopic scale. Given a standard photograph and a sparse set of c
arXiv:2607.22702v1 Announce Type: new Abstract: Text-motion representation learning has advanced rapidly, with growing interest in multi person interactions for animation, AR/VR, and embodied AI. Thes
arXiv:2607.24407v1 Announce Type: new Abstract: Multimodal Large Language Models have made great progress in grounding tasks, yet existing methods still struggle to unify precise localization and comp
arXiv:2607.24665v1 Announce Type: new Abstract: Modern large language models scale successfully by pairing capacity growth with efficiency, keeping per-token and deployment costs under control as capa
arXiv:2607.22973v1 Announce Type: new Abstract: Millimeter-wave (mmWave) radar offers privacy-preserving and lighting-robust sensing for human motion reconstruction, but learning models that generaliz
arXiv:2607.23511v1 Announce Type: new Abstract: End-to-end autonomous driving systems commonly follow a cascaded two-stage pipeline where a perception stage compresses multi-modal sensor inputs into a
arXiv:2512.03522v3 Announce Type: replace-cross Abstract: Robots are often required to localize in environments with unknown object classes and semantic ambiguity. However, when performing global loca
arXiv:2607.24436v1 Announce Type: new Abstract: High-fidelity 3D generative modeling increasingly relies on the latent diffusion paradigm, where the reconstruction quality of the underlying 3D VAE bec
arXiv:2607.23451v1 Announce Type: new Abstract: Multi-modal object Re-Identification (ReID) aims to retrieve specific objects by integrating complementary information from multiple modalities. However
arXiv:2607.24302v1 Announce Type: new Abstract: Human mesh recovery (HMR) aims to recover 3D human meshes from images. Most existing HMR benchmarks and methods focus on either multi-person reconstruct
arXiv:2607.23979v1 Announce Type: new Abstract: Visual tire recognition serves as a core supporting technique for vehicle safety monitoring, autonomous driving perception and automated automotive main
arXiv:2607.23576v1 Announce Type: new Abstract: Conventional frame-based cameras face significant challenges in detecting objects under high-speed motion blur or in low-light environments. Neuromorphi
arXiv:2607.22738v1 Announce Type: cross Abstract: Current 3D generative models mostly produce a final surface: a visually strong but largely opaque mesh. Interactive 3D worlds need more than a surface
arXiv:2607.24495v1 Announce Type: new Abstract: Structured-light (SL) cameras power depth sensing in millions of devices, and recent neural SL decoding methods have substantially improved their depth
arXiv:2607.23122v1 Announce Type: cross Abstract: Ambient occlusion (AO) and soft shadows are critical visibility cues for spatial perception in real-time rendering. Hardware ray tracing provides a di
arXiv:2607.23844v1 Announce Type: new Abstract: High-resolution image and video diffusion models, including SD3, FLUX, and recent video diffusion transformers, have substantially improved generative q
arXiv:2607.23023v1 Announce Type: new Abstract: Recent advances in diffusion-based generative models have enabled real-time audio-driven avatar generation and unified audio-visual synthesis, providing
arXiv:2607.23193v1 Announce Type: new Abstract: Existing token compression methods for omnimodal large language models typically rely on one modality to determine what to retain in the other. We show
arXiv:2607.23855v1 Announce Type: cross Abstract: Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Des
arXiv:2601.08001v2 Announce Type: replace-cross Abstract: Tear film (TF) breakup is a key driver of understanding dry eye disease, yet estimating TF thickness and osmolarity from fluorescence (FL) ima
arXiv:2603.02063v2 Announce Type: replace Abstract: Although data generation is often straightforward, extracting information from data is more difficult. Object-centric representation learning can ex
arXiv:2607.23194v1 Announce Type: new Abstract: Scene Text Recognition (STR) models are trained almost exclusively on word crops of at most 25 characters, yet real deployments (signage, product labels
arXiv:2607.24703v1 Announce Type: new Abstract: Female pelvic diseases remain an under researched area characterized by often delayed diagnosis. While pelvic MRI offers superior soft-tissue contrast f
arXiv:2602.05387v2 Announce Type: replace Abstract: MRI provides superior soft tissue contrast without ionizing radiation; however, the absence of electron density information limits its direct use fo
arXiv:2607.23694v1 Announce Type: new Abstract: Efficient surgical segmentation empowers clinical diagnosis, intraoperative monitoring, and downstream robotic pipelines for reconstruction and simulati
arXiv:2607.23631v1 Announce Type: new Abstract: Gigapixel Whole-Slide Images (WSIs) present a fundamental computational bottleneck for vision-language models (VLMs) due to extreme sequence lengths. Ex
arXiv:2607.22726v1 Announce Type: new Abstract: Despite advances in Video Large Language Models (VLLMs) that have displayed promising outcomes in video understanding, the redundancy in the long-durati
arXiv:2607.23680v1 Announce Type: new Abstract: Semantic segmentation in agricultural imagery is often evaluated under in-domain protocols, yet practical deployment requires robustness to appearance p
arXiv:2411.00967v2 Announce Type: replace Abstract: The future of agriculture is intertwined with automation. Accurate fruit detection, yield estimation, and harvest time prediction are crucial for ef
arXiv:2607.24052v1 Announce Type: new Abstract: High-curvature regions in 3D point clouds encapsulate critical fine-grained geometric semantics yet exhibit a distinct long-tail sparsity in their spati
arXiv:2607.22963v1 Announce Type: cross Abstract: Synthetic aperture radar (SAR) image generation can mitigate data scarcity, but controllablegeneration under sparse observation angles remains difficu
arXiv:2607.24353v1 Announce Type: new Abstract: Text-to-image generation models can synthesize high-quality images from natural language descriptions, but their performance remains highly sensitive to
arXiv:2607.22755v1 Announce Type: cross Abstract: pyALDIC is an open-source Python implementation of augmented Lagrangian digital image correlation (AL-DIC) for full-field displacement and strain meas
arXiv:2607.24598v1 Announce Type: new Abstract: Video instance segmentation (VIS) requires models to detect, segment, and track object identities across frames, and most methods enforce temporal consi
arXiv:2607.23517v1 Announce Type: new Abstract: We present a real-time human-centric world model for upper-body interactive generation, aiming to synthesize coherent local world dynamics centered on a
arXiv:2607.24199v1 Announce Type: new Abstract: Understanding and complying with traffic regulations is a safety-critical requirement for autonomous driving, yet remains challenging due to the diversi
arXiv:2511.12940v2 Announce Type: replace Abstract: Recent advancements in video generation has shifted from bidirectional models for short videos to autoregressive ones for ultra long video generatio
arXiv:2607.24098v1 Announce Type: new Abstract: Referring video object segmentation (RVOS) requires segmenting a target specified by natural language throughout a video. Recent agentic approaches comb
arXiv:2607.24465v1 Announce Type: new Abstract: Model merging aims to combine multiple domain-specialized experts trained from a shared foundation model into a single multi-task model. Existing approa
arXiv:2607.23758v1 Announce Type: new Abstract: Large-scale road surface reconstruction supports high-definition mapping, autonomous-driving perception, annotation, and simulation. Existing road-speci
arXiv:2607.23468v1 Announce Type: new Abstract: Real-time 6-DoF object pose tracking is essential for many robotics applications, and several approaches exist. Yet even today's approaches remain unrel
arXiv:2607.23958v1 Announce Type: new Abstract: Point cloud denoising is essentially a geometric recovery task that aims to reconstruct the intrinsic structure of a smooth 2D Riemannian manifold embed