Dual-End Consistency Model
arXiv:2602.10764v2 Announce Type: replace Abstract: The slow iterative sampling nature remains a major bottleneck for the practical deployment of diffusion and flow-based generative models. While cons
Knowledge catalogue
arXiv:2602.10764v2 Announce Type: replace Abstract: The slow iterative sampling nature remains a major bottleneck for the practical deployment of diffusion and flow-based generative models. While cons
arXiv:2604.17542v1 Announce Type: new Abstract: Conventional test-time adaptation (TTA) approaches typically adapt the model using only a small fraction of test samples, often those with low-entropy p
arXiv:2604.17688v1 Announce Type: new Abstract: 3D human pose estimation is a classic and important research direction in the field of computer vision. In recent years, Transformer-based methods have
arXiv:2604.16987v1 Announce Type: new Abstract: The rapid evolution of video generation technologies poses a significant challenge to media forensics, as conventional detection methods often fail to g
arXiv:2604.16483v1 Announce Type: new Abstract: Concept erasure in Text-To-Image (T2I) diffusion models is vital for safe content generation, but existing inference-time methods face significant limit
arXiv:2604.17710v1 Announce Type: new Abstract: Zero-shot learning (ZSL) aims to recognize unseen classes without visual instances. However, existing methods usually assume clean labels, overlooking r
arXiv:2604.17969v1 Announce Type: new Abstract: Visual search in 3D environments requires embodied agents to actively explore their surroundings and acquire task-relevant evidence. However, existing v
arXiv:2604.18367v1 Announce Type: new Abstract: Early action prediction seeks to anticipate an action before it fully unfolds, but limited visual evidence makes this task especially challenging. We in
arXiv:2604.16893v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) has demonstrated remarkable effectiveness in improving the reasoning capabilities of large languag
arXiv:2604.16783v1 Announce Type: new Abstract: Vehicle trajectory prediction is central to highway perception, but deployment on roadside edge devices necessitates bounded, deterministic end-to-end l
arXiv:2604.17500v1 Announce Type: new Abstract: Scene text editing (STE) has achieved remarkable progress in accurately rendering target text through diffusion-based methods. However, we identify a cr
arXiv:2509.20360v3 Announce Type: replace Abstract: Recent advances in foundation models highlight a clear trend toward unification and scaling, showing emergent capabilities across diverse domains. W
arXiv:2604.17749v1 Announce Type: new Abstract: Understanding physical transformation processes is crucial for both human cognition and artificial intelligence systems, particularly from an egocentric
arXiv:2602.14122v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in vision-language understanding. Yet, human perception is inher
arXiv:2604.18167v1 Announce Type: new Abstract: Modern text-to-image (T2I) models amplify harmful societal biases, challenging their ethical deployment. We introduce an inference-time method that reli
arXiv:2604.17211v1 Announce Type: new Abstract: We present EmbodiedHead, a speech-driven talking-head framework that equips LLMs with real-time visual avatars for conversation. A practical embodied av
arXiv:2505.00986v2 Announce Type: replace-cross Abstract: Continual Test-time adaptation (CTTA) continuously adapts the deployed model on every incoming batch of data. While achieving optimal accuracy
arXiv:2511.12554v2 Announce Type: replace Abstract: Visual Emotion Analysis (VEA) aims to bridge the affective gap between visual content and human emotional responses. Despite its promise, progress i
arXiv:2604.18075v1 Announce Type: new Abstract: We investigate recently introduced domain-class incremental learning scenarios for vision-language models (VLMs). Recent works address this challenge us
arXiv:2604.18336v1 Announce Type: cross Abstract: Indoor robot navigation is often compromised by glass surfaces, which severely corrupt depth sensor measurements. While foundation models like Depth A
arXiv:2604.17233v1 Announce Type: new Abstract: Personalized image aesthetics assessment (PIAA) aims to predict an individual user's subjective rating of an image, which requires modeling user-specifi
arXiv:2501.12119v3 Announce Type: replace-cross Abstract: We introduce ENTIRE, a novel deep learning-based approach for fast and accurate volume rendering time prediction. Predicting rendering time is
arXiv:2604.16481v1 Announce Type: new Abstract: Large-scale text-to-image (T2I) diffusion models deliver remarkable visual fidelity but pose safety risks due to their capacity to reproduce undesirable
arXiv:2603.03692v2 Announce Type: replace Abstract: Classifier-Free Guidance (CFG) has established the foundation for guidance mechanisms in diffusion models, showing that well-designed guidance proxi
arXiv:2604.18320v1 Announce Type: new Abstract: Self-evolution of multimodal large language models (MLLMs) remains a critical challenge: pseudo-label-based methods suffer from progressive quality degr
arXiv:2604.17087v1 Announce Type: new Abstract: Recent Multimodal Large Language Models (MLLMs) have demonstrated strong performance on vision-language understanding tasks, yet their inference efficie
arXiv:2604.16528v1 Announce Type: new Abstract: Embryo selection is one of multiple crucial steps in in-vitro fertilization, commonly based on morphological assessment by clinical embryologists. Altho
arXiv:2504.04814v3 Announce Type: replace-cross Abstract: Trustworthy artificial intelligence (AI) is essential in healthcare, particularly for high-stakes tasks like medical image segmentation. Expla
arXiv:2604.17879v1 Announce Type: new Abstract: Camouflaged Object Detection is challenging due to the high degree of similarity between camouflaged objects and their surrounding backgrounds. Current
arXiv:2502.13637v2 Announce Type: replace Abstract: Human affordance learning investigates contextually relevant novel pose prediction such that the estimated pose represents a valid human action with
arXiv:2505.22226v2 Announce Type: replace Abstract: Recent theoretical advances reveal that the Hadamard product induces nonlinear representations and implicit high-dimensional mappings for the field
arXiv:2604.18168v1 Announce Type: new Abstract: Few-step generation has been a long-standing goal, with recent one-step generation methods exemplified by MeanFlow achieving remarkable results. Existin
arXiv:2604.16780v1 Announce Type: new Abstract: This paper presents FairNVT, a lightweight debiasing framework for pretrained transformer-based encoders that improves both representation and predictio
arXiv:2604.16522v1 Announce Type: new Abstract: This paper proposes a fast and online method for jointly performing 3D multi-object tracking and pose estimation using multiple monocular cameras. Our a
arXiv:2511.17171v4 Announce Type: replace Abstract: Predicting wildfire risk is a reasoning-intensive spatial problem that requires the integration of visual, climatic, and geographic factors to infer
arXiv:2604.17720v1 Announce Type: cross Abstract: Point-based Neural Networks (PNNs) have become a key approach for point cloud processing. However, a core operation in these models, Farthest Point Sa
arXiv:2604.17625v1 Announce Type: new Abstract: This paper introduces a novel methodology for generating fast and memory-efficient video continuations. Our method, dubbed FlowC2S, fine-tunes a pre-tra
arXiv:2603.19857v2 Announce Type: replace-cross Abstract: Recent Video-to-Audio (V2A) methods have achieved remarkable progress, enabling the synthesis of realistic, high-quality audio. However, they
arXiv:2604.17268v1 Announce Type: new Abstract: AI-generated imagery has reached near-photorealistic fidelity, yet this technology poses significant threats to information security and societal trust.
arXiv:2603.09721v2 Announce Type: replace Abstract: High-fidelity video generation remains challenging for diffusion models due to the difficulty of modeling complex spatio-temporal dynamics efficient
arXiv:2603.07690v2 Announce Type: replace Abstract: Streaming Visual Geometry Transformers such as StreamVGGT enable strong online 3D perception, but their KV-cache grows unbounded over long streams,
arXiv:2604.16800v1 Announce Type: new Abstract: Addressing the issues of severe noise and high frequency structural degradation in visible images under low-light conditions, this paper proposes a Near
arXiv:2604.17298v1 Announce Type: new Abstract: Video Scene Graph Generation aims to obtain structured semantic representations of objects and their relationships in videos for high-level understandin
arXiv:2604.17231v1 Announce Type: new Abstract: Unrecovered e-waste represents a significant economic loss. Hard disk drives (HDDs) comprise a valuable e-waste stream necessitating robotic disassembly
arXiv:2604.17455v1 Announce Type: new Abstract: Visual prompting has emerged as a powerful method for adapting pre-trained models to new domains without updating model parameters. However, existing pr
arXiv:2604.17110v1 Announce Type: new Abstract: Clinical AI development has traditionally followed a collaborative paradigm that depends on close interaction between clinicians and specialized AI team
arXiv:2604.16504v1 Announce Type: new Abstract: Manual digitisation of structured handwritten documents is slow and costly. We benchmark 17 leading frontier multi-modal large language models and open-
arXiv:2604.16462v1 Announce Type: new Abstract: High-resolution Multimodal Large Language Models (MLLMs) face prohibitive computational costs during inference due to the explosion of visual tokens. Ex
arXiv:2604.16758v1 Announce Type: new Abstract: We present a system for automated detection, localization, and scoring of arrow punctures on 40,cm indoor archery target faces, trained on only 48 annot
arXiv:2604.17721v1 Announce Type: new Abstract: We address the challenge of point cloud registration using color information, where traditional methods relying solely on geometric features often strug
arXiv:2604.17307v1 Announce Type: new Abstract: Detecting face forgeries using CLIP has recently emerged as a promising and increasingly popular research direction. Owing to its rich visual knowledge
arXiv:2604.16796v1 Announce Type: new Abstract: Generative semantic communication (SemCom) harnesses pretrained generative priors to improve the perceptual quality of wireless image transmission. Exis
arXiv:2604.16487v1 Announce Type: new Abstract: CLIP retrieval is typically framed as a pointwise similarity problem in a shared embedding space. While CLIP achieves strong global cross-modal alignmen
arXiv:2604.18260v1 Announce Type: new Abstract: Multimodal large language models have demonstrated remarkable capabilities in 2D vision, motivating their extension to 3D scene understanding. Recent st
arXiv:2604.17822v1 Announce Type: new Abstract: Class-Incremental Learning (CIL) aims to continuously acquire new categories while preserving previously learned knowledge. Recently, Contrastive Langua
arXiv:2602.16213v2 Announce Type: replace-cross Abstract: This paper introduces a novel approach to sea ice modeling using Graph Neural Networks (GNNs), utilizing the natural graph structure of sea ic
arXiv:2604.18047v1 Announce Type: new Abstract: Continuous Spatio-Temporal Video Super-Resolution (C-STVSR) aims to simultaneously enhance the spatial resolution and frame rate of videos by arbitrary
arXiv:2604.18037v1 Announce Type: new Abstract: Composed Image Retrieval (CIR) is a flexible image retrieval paradigm that enables users to accurately locate the target image through a multimodal quer
arXiv:2604.16823v1 Announce Type: new Abstract: Vision Transformer (ViT) has brought new breakthroughs to the field of image classification by introducing the self-attention mechanism and Graph Convol
arXiv:2507.05920v2 Announce Type: replace Abstract: State-of-the-art large multi-modal models (LMMs) face challenges when processing high-resolution images, as these inputs are converted into enormous