Asymmetric Flow Models
arXiv:2605.12964v1 Announce Type: new Abstract: Flow-based generation in high-dimensional spaces is difficult because velocity prediction requires modeling high-dimensional noise, even when data has s
Knowledge catalogue
arXiv:2605.12964v1 Announce Type: new Abstract: Flow-based generation in high-dimensional spaces is difficult because velocity prediction requires modeling high-dimensional noise, even when data has s
arXiv:2605.13381v1 Announce Type: new Abstract: As AI-generated synthetic images become increasingly realistic, Vision Transformers (ViTs) have emerged as a cornerstone of modern deepfake detection. H
arXiv:2605.13455v1 Announce Type: new Abstract: Synapses are densely packed submicron structures that dynamically reorganize during learning and memory formation. Longitudinal extit{in vivo} imaging o
arXiv:2510.01502v2 Announce Type: replace-cross Abstract: Current video foundation models, including the strongest self-supervised models such as V-JEPA2, fail to capture how humans organize social in
arXiv:2603.05582v2 Announce Type: replace-cross Abstract: The issue of algorithmic biases in deep learning has led to the development of various debiasing techniques, many of which perform complex tra
arXiv:2605.13794v1 Announce Type: cross Abstract: We present BlitzGS, a distributed 3DGS framework that reduces active Gaussian workload for fast city-scale reconstruction. BlitzGS manages this worklo
arXiv:2605.12560v1 Announce Type: cross Abstract: Improving patient outcomes depends on the prompt and accurate diagnosis of brain tumors, but manual MRI scan analysis is still time-consuming and unre
arXiv:2605.13059v1 Announce Type: new Abstract: Clinical diagnostic workups typically follow a modality escalation pathway: after initial clinical evaluation, clinicians begin with routine structural
arXiv:2605.13544v1 Announce Type: new Abstract: Fine-grained Vision-Language Pre-training (FVLP) demonstrates significant potential in 3D medical image understanding by aligning anatomy-level visual r
arXiv:2605.13675v1 Announce Type: new Abstract: Deep neural networks trained with different architectures, objectives, and datasets have been reported to converge on similar visual representations. Ho
arXiv:2605.12882v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have significantly advanced document understanding, yet current Doc-VQA evaluations score only the final answ
arXiv:2605.13306v1 Announce Type: new Abstract: Illuminant estimation aims to infer scene illumination from image measurements despite intrinsic ambiguities between surface reflectance and lighting. M
arXiv:2403.11247v3 Announce Type: replace Abstract: Recent work has shown that 3D Gaussian-based SLAM enables high-quality reconstruction, accurate pose estimation, and real-time rendering of scenes.
arXiv:2601.22868v3 Announce Type: replace Abstract: Anomaly detection usually assumes that abnormality is an intrinsic property of an observation. A defect is a defect, and a rare object is rare, rega
arXiv:2605.12650v1 Announce Type: new Abstract: Foundation diffusion models can generate photorealistic natural images, but adapting them to medical imaging remains challenging. In medical adaptation,
arXiv:2603.07433v2 Announce Type: replace-cross Abstract: Dynamic Data selection aims to accelerate training by prioritizing informative samples during online training. However, existing methods typic
arXiv:2605.12952v1 Announce Type: new Abstract: Grad-ECLIP is published at ICML 2024 and represents a new Transformer interpretation technical route (intermediate features-based). First, this paper de
arXiv:2605.13619v1 Announce Type: cross Abstract: Extended depth of field microscopy encodes axial information into a single acquisition through engineered point spread functions, but conventional and
arXiv:2605.13182v1 Announce Type: new Abstract: Diffusion-based models have shown strong performance in video super-resolution (VSR) and video frame interpolation (VFI). However, their role in the cou
arXiv:2605.12939v1 Announce Type: new Abstract: Recent diffusion- and flow-based VTON methods achieve strong results with pretrained generative models, but their reliance on multi-step sampling incurs
arXiv:2605.12649v1 Announce Type: new Abstract: Dataset distillation aims to synthesize a compact proxy dataset that is unreadable or non-raw from the original dataset for privacy protection and highl
arXiv:2605.12623v1 Announce Type: cross Abstract: Multilingual document understanding remains limited for low-resource languages due to scarce training data and model-based annotation pipelines that p
arXiv:2605.13179v1 Announce Type: new Abstract: The Engram module -- a hash-keyed, O(1) associative memory injected into Transformer layers -- was recently shown to improve large language model pretra
arXiv:2605.13349v1 Announce Type: new Abstract: Diffusion-based point editing methods have gained significant traction in image editing tasks due to their ability to manipulate image semantics and fin
arXiv:2605.12625v1 Announce Type: cross Abstract: Continuous-action policies trained on a single demonstrated trajectory per scene suffer from mode collapse: samples cluster around the demonstrated ma
arXiv:2605.13156v1 Announce Type: new Abstract: Vision-language models (VLMs) have demonstrated remarkable capabilities in bridging visual perception and natural language understanding, enabling a wid
arXiv:2605.13122v1 Announce Type: new Abstract: Instruction-based image editing (IIE) models have recently demonstrated strong capability in modifying specific image regions according to natural langu
arXiv:2605.13062v1 Announce Type: new Abstract: Recent image editing models have achieved remarkable progress in instruction following, multimodal understanding, and complex visual editing. However, e
arXiv:2509.13858v2 Announce Type: replace Abstract: Dataset distillation aims to synthesize a compact dataset from the original large-scale one, enabling highly efficient learning while preserving com
arXiv:2605.13041v1 Announce Type: new Abstract: With recent advances in embodied agents and AR devices, egocentric observations are readily available as input for real-world interactive online applica
arXiv:2511.17031v2 Announce Type: replace-cross Abstract: The rapidly growing computational demands of diffusion models for image generation have raised significant concerns about energy consumption a
arXiv:2605.13803v1 Announce Type: new Abstract: Video temporal grounding (VTG) takes an untrimmed video and a natural-language query as input and localizes the temporal moment that best matches the qu
arXiv:2602.22455v2 Announce Type: replace Abstract: We investigate the feasibility of using Multimodal Large Language Models (MLLMs) for real-time online episodic memory question answering. While clou
arXiv:2605.13402v1 Announce Type: new Abstract: Computing a minimum s-t cut in a graph is a solution to a wide range of computer vision problems, and is often done using the Boykov-Kolmogorov (BK) alg
arXiv:2605.13475v1 Announce Type: new Abstract: Federated Learning (FL) enables collaborative training of distributed clients while protecting privacy. To enhance generalization capability in FL, prot
arXiv:2605.13193v1 Announce Type: new Abstract: Fine-grained recognition in everyday life is often not a closed-book classification problem: when encountering unfamiliar objects, humans actively searc
arXiv:2605.13108v1 Announce Type: new Abstract: Face presentation attack detection (FacePAD) remains challenging under diverse spoofing representation, including 2D print and replay, 3D mask-based spo
arXiv:2602.10326v2 Announce Type: replace Abstract: Despite the remarkable success of sampling-based generative models such as flow matching, they can still produce samples of inconsistent or degraded
arXiv:2509.23056v2 Announce Type: replace Abstract: Remote sensing object detection is a critical technology for real-world applications such as natural resource monitoring, traffic management, and UA
arXiv:2603.26839v2 Announce Type: replace-cross Abstract: How do multimodal models solve visual spatial tasks -- through genuine planning, or through brute-force search in token space? We introduce ex
arXiv:2605.13151v1 Announce Type: new Abstract: Category-agnostic pose estimation (CAPE) aims to localize keypoints on query images from arbitrary categories, using only a few annotated support exampl
arXiv:2605.12778v1 Announce Type: cross Abstract: Recent advances in generative models have yielded impressive progress on motion in-betweening, allowing for more complex, varied, and realistic motion
arXiv:2605.13755v1 Announce Type: new Abstract: In recent years, autonomous driving has significantly in creased the demand for high-quality data to train 2D and 3D perception models for safety-critic
arXiv:2505.05376v3 Announce Type: replace Abstract: We propose a novel method that reconstructs hair strands directly from colorless 3D scans by leveraging multi-modal hair orientation extraction. Hai
arXiv:2602.17555v3 Announce Type: replace Abstract: Video reasoning requires a fine-grained understanding of the temporal dependencies and event-level relations between objects and events in videos. C
arXiv:2605.12957v1 Announce Type: new Abstract: Recent developments in generative models and large-scale datasets have substantially advanced 3D world generation, facilitating a broad range of domains
arXiv:2605.12919v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) is becoming a practical representation for novel view synthesis, but its growing adoption, together with rapid advances in
arXiv:2605.13632v1 Announce Type: cross Abstract: In this paper, we propose GTA-VLA(Guide, Think, Act), an interactive Vision-Language-Action (VLA) framework that enables spatially steerable embodied
arXiv:2602.07029v3 Announce Type: replace-cross Abstract: This work introduces the first closed-loop adaptive optics (AO) system capable of optically correcting aberrations in real-time without a guid
arXiv:2605.13664v1 Announce Type: new Abstract: Thermal-infrared (TIR) hyperspectral imagery (HSI) provides critical scene information for various applications. However, its practical utility is sever
arXiv:2605.13073v1 Announce Type: new Abstract: In-the-wild 3D Gaussian Splatting remains challenging due to transient distractors and illumination-induced cross-view appearance inconsistencies. Exist
arXiv:2605.13581v1 Announce Type: new Abstract: Hyperspectral image (HSI) restoration is crucial for reliable analysis, as real HSIs suffer from degradations like noise, blur, and resolution loss. How
arXiv:2605.12619v1 Announce Type: cross Abstract: The perceptual representations supporting our ability to recognize faces remain a computational mystery. Deep neural networks offer mechanistic hypoth
arXiv:2605.12967v1 Announce Type: new Abstract: The rapid advancement of generative AI has enabled the creation of highly realistic and diverse synthetic images, posing critical challenges for image p
arXiv:2605.13293v1 Announce Type: new Abstract: Boundary Representation (BRep) is the standard format for Computer-Aided Design (CAD), yet reconstructing high-quality BReps from single-view images rem
arXiv:2605.08320v1 Announce Type: cross Abstract: Monocular depth estimation (MDE) with self-supervised training approaches struggles in low-texture areas, where photometric losses may lead to ambiguo
arXiv:2601.22853v3 Announce Type: replace Abstract: Multimodal deep learning (MDL) has achieved remarkable success across various domains, yet its practical deployment is often hindered by incomplete
arXiv:2605.12725v1 Announce Type: new Abstract: Recent video anomaly detection research has expanded rapidly with an emphasis on general models of normality intended to work across many different scen
arXiv:2605.13813v1 Announce Type: new Abstract: Automated CT triage requires models that are simultaneously accurate across diverse pathologies and reliable under institutional shift. While Vision Tra
arXiv:2605.12772v1 Announce Type: new Abstract: Wu et al. (2026) showed that most frontier large language models (LLMs) recommend a sponsored, roughly twice-as-expensive flight when their system promp