Decoding Children's Gait Behavior
arXiv:2608.00371v1 Announce Type: new Abstract: We introduce a new problem domain for human action recognition: the fine-grained analysis of children's gait behaviors from standard RGB video. We speci
Knowledge catalogue
arXiv:2608.00371v1 Announce Type: new Abstract: We introduce a new problem domain for human action recognition: the fine-grained analysis of children's gait behaviors from standard RGB video. We speci
arXiv:2608.01761v1 Announce Type: new Abstract: End-to-end (E2E) autonomous driving algorithms require rigorous closed-loop validation in simulation environments offering high visual fidelity, strong
arXiv:2608.01848v1 Announce Type: new Abstract: Recent visual-text compression (VTC) methods, typified by DeepSeek-OCR, report impressive high token compression ratios for long-context modeling tasks
arXiv:2608.00835v1 Announce Type: cross Abstract: Fragile X Syndrome (FXS) is a neurodevelopmental disorder caused by reduced expression of fragile X mental retardation protein (FMRP), leading to disr
arXiv:2506.02976v4 Announce Type: replace Abstract: The MARIO challenge, held at MICCAI 2024, focused on advancing the automated detection and monitoring of age-related macular degeneration (AMD) thro
arXiv:2608.02092v1 Announce Type: new Abstract: Deep multimodal fusion for object detection has demonstrated good performance through mining modal characteristics. However, existing feature-level fusi
arXiv:2608.01827v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to
arXiv:2608.02099v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) has emerged as a leading technique for real-time novel view synthesis, yet existing 3DGS accelerators suffer from poor ar
arXiv:2608.02044v1 Announce Type: new Abstract: Tracking links observations of the same object through visual change, yet cannot by itself determine when the object is empty or filled, intact or cut.
arXiv:2608.02191v1 Announce Type: new Abstract: Although image deraining has advanced substantially, existing methods mainly focus on 2D image restoration. As spatial intelligence applications such as
arXiv:2608.01823v1 Announce Type: new Abstract: Hallucination remains a persistent challenge in generative super-resolution (GSR), where reconstructed results may contain visually plausible yet weakly
arXiv:2608.00078v1 Announce Type: new Abstract: Deploying convolutional neural networks generated by large language models (LLMs) on real mobile hardware requires more than GPU validation accuracy: IN
arXiv:2608.01343v1 Announce Type: new Abstract: The emergence of transformer-based deep learning models has brought unprecedented performance across various domains, particularly in natural language p
arXiv:2608.00554v1 Announce Type: cross Abstract: Dexterous object rotation is a sequential contact problem: each support, release, and re-contact decision must both produce the desired object motion,
arXiv:2608.02428v1 Announce Type: new Abstract: Forecasting future states from video sequences is a critical challenge for autonomous robotic systems and a fundamental objective of world modeling. Pri
arXiv:2608.00617v1 Announce Type: new Abstract: Many physical attributes are irreversible: ice melts but does not re-freeze, paper chars but does not un-burn. Do video generators respect this? We show
arXiv:2508.08831v2 Announce Type: replace-cross Abstract: Generating synthetic images that closely mimic those from real cameras is instrumental in training visual models and enabling end-to-end visuo
arXiv:2608.01985v1 Announce Type: new Abstract: Visual token pruning reduces the computational cost of Vision-Language Models (VLMs) by removing redundant visual tokens. The key is to learn a score th
arXiv:2608.00540v1 Announce Type: new Abstract: Tool-integrated vision-language agents have made remarkable progress on compositional and multi-step visual reasoning. Yet their outputs frequently exhi
arXiv:2405.06945v4 Announce Type: replace Abstract: Jointly recovering explicit surface geometry and high-quality appearance from multi-view images remains challenging. This capability is essential fo
arXiv:2608.00110v1 Announce Type: new Abstract: 3D scene understanding requires reasoning about entity existence, spatial layout, and object relations, yet RGB images alone often provide insufficient
arXiv:2607.15933v2 Announce Type: replace Abstract: The effectiveness of modern visual representation learning and autoregressive models critically depends on vector quantization (VQ), which discretiz
arXiv:2607.17999v2 Announce Type: replace-cross Abstract: Spatial understanding is crucial for foundation models (FMs), and maps have long helped humans organize and reason about geographic informatio
arXiv:2608.00536v1 Announce Type: new Abstract: Reinforcement learning (RL) for document parsing often relies on reference-based rewards rooted in edit distance (e.g., tree edit distance), yet it rema
arXiv:2608.00089v1 Announce Type: new Abstract: With rapid growth in the fields of empirical and computational aesthetics we have seen a vast increase in large image datasets annotated for aesthetics.
arXiv:2608.02396v1 Announce Type: new Abstract: Most evidence on the effectiveness of explainable artificial intelligence (XAI) attribution methods has been established on convolutional neural network
arXiv:2608.00548v1 Announce Type: new Abstract: Recent image-generation models and multimodal agents can produce high-quality visuals for increasingly complex visual communication tasks. Yet their ras
arXiv:2608.00486v1 Announce Type: new Abstract: Accurate prediction of object trajectories during manipulation is essential for closing the perception-action loop. Progress is limited on two fronts: a
arXiv:2603.00919v3 Announce Type: replace Abstract: Large language models (LLMs) have shown great promise for autonomous driving. However, discretizing numbers into tokens limits precise numerical rea
arXiv:2608.01338v1 Announce Type: new Abstract: High-definition (HD) maps are essential for autonomous driving systems. In constructing such maps, onboard multi-view camera images, standard-definition
arXiv:2608.00086v1 Announce Type: new Abstract: Brain tumor diagnosis is a time-sensitive process in which patients may wait weeks for a finalized pathology report. This problem motivates automated sy
arXiv:2608.02495v1 Announce Type: new Abstract: Despite the remarkable progress over the past decades, accurately identifying small objects remains challenging because of their insufficient visual cue
arXiv:2608.01178v1 Announce Type: new Abstract: We present DynActiveGS, a dynamic-aware active reconstruction framework based on 3D Gaussian Splatting (3DGS) for autonomous exploration in dynamic envi
arXiv:2608.01638v1 Announce Type: new Abstract: Egocentric visual grounding requires high-resolution inputs to localize small objects. However, scaling Multimodal Large Language Models to this domain
arXiv:2608.01452v1 Announce Type: cross Abstract: Dynamic manipulation is a critical capability for robots operating in complex and dynamic environments, where robots must interact with objects that a
arXiv:2608.00694v1 Announce Type: new Abstract: Event cameras offer microsecond-level temporal resolution and high dynamic range, potentially facilitating motion-blur-free panoramic imaging from fast
arXiv:2608.02474v1 Announce Type: new Abstract: Audio-driven video generation (A2V) has achieved promising progress in synthesizing temporally coherent and audio-visually aligned videos, yet its infer
arXiv:2601.17883v3 Announce Type: replace-cross Abstract: Electroencephalography (EEG) foundation models (FMs) have recently emerged as a promising paradigm for brain-computer interfaces, aiming to le
arXiv:2603.12147v2 Announce Type: replace Abstract: Egocentric video provides a natural modality for studying human behavior, but conventional visual understanding captures mainly observable scenes, o
arXiv:2608.00060v1 Announce Type: new Abstract: Here we introduce ELECTRIC (Evidential Learning-Enhanced CT Reconstruction via Iterative Correction), a physics-guided Bayesian formulation. An evidenti
arXiv:2608.00584v1 Announce Type: new Abstract: Recent advances in image generation and editing have made prompt quality a key bottleneck for e-commerce creatives. Vision-language models (VLMs) can ge
arXiv:2604.00933v2 Announce Type: replace Abstract: Text-to-image diffusion models achieve high visual fidelity, yet fine-grained affective control remains difficult because textual emotion cues often
arXiv:2608.00071v1 Announce Type: new Abstract: The rapid emergence of 3D CT foundation models has opened new avenues for predictive modeling from CT imaging, offering a compelling alternative to trad
arXiv:2608.01572v1 Announce Type: new Abstract: Autonomous driving (AD) systems have advanced rapidly over the past decade; however, robust perception under adverse weather conditions remains a major
arXiv:2608.01696v1 Announce Type: new Abstract: Player-centric ball action spotting requires temporally precise event detection together with actor attribution in crowded, partially observed multi-age
arXiv:2608.02284v1 Announce Type: new Abstract: Open-vocabulary segmentation identifies and segments objects from arbitrary textual descriptions. SAM 3 supports noun-phrase-guided segmentation and ach
arXiv:2608.02549v1 Announce Type: cross Abstract: Efficient and perceptually meaningful quality assessment is a fundamental requirement for image and video processing, compression, and streaming syste
arXiv:2608.01948v1 Announce Type: new Abstract: Long-horizon event-based action understanding remains underexplored because existing datasets largely comprise short, trimmed clips, while collecting na
arXiv:2608.00074v1 Announce Type: new Abstract: This paper presents a multimodal machine-learning framework for calibration monitoring, quality assessment, and adaptive acquisition support in archaeol
arXiv:2608.02289v1 Announce Type: new Abstract: Realistic and diverse trajectory generation is central to enabling higher levels of vehicle automation. While rule-based and classical learning-based me
arXiv:2608.01058v1 Announce Type: new Abstract: Artificial Intelligence is increasingly applied to surgical video analysis for phase segmentation, skill assessment, and workflow optimization. A key ch
arXiv:2608.01049v1 Announce Type: cross Abstract: World models have attracted significant attention for their ability to capture and predict the structure and dynamics of the physical world. In this e
arXiv:2608.01661v1 Announce Type: new Abstract: The challenge of fair deepfake detection (FDD) has attracted increasing attention. Existing fairness-enhanced detectors often suffer from suboptimal gen
arXiv:2608.01958v1 Announce Type: new Abstract: 4D Gaussian Splatting (4DGS) excels in dynamic 3D reconstruction and real-time novel view synthesis via efficient 4D Gaussian representations and parall
arXiv:2608.00053v1 Announce Type: cross Abstract: The Discrete Fourier Transform, the Discrete Cosine Transform, and their block-wise variants underpin most deployed image and video codecs. Their effe
arXiv:2608.01664v1 Announce Type: new Abstract: We present our ImageCLEF 2026 Multimodal Reasoning system for the Visual Multiple Choice Question Answering (Visual MCQ) and Visual Open Question Answer
arXiv:2608.00111v1 Announce Type: cross Abstract: Image restoration quality can be evaluated along three complementary facets: pixel-level fidelity, human perception, and downstream machine preference
arXiv:2608.01129v1 Announce Type: cross Abstract: Although recent robot perception research emphasizes training on data from diverse environments to improve generalization, most existing methods still
arXiv:2608.02483v1 Announce Type: new Abstract: Two active learning algorithms for hyperspectral image (HSI) classification are proposed that combine density-aware Fermat distances with Poisson-reweig
arXiv:2608.01663v1 Announce Type: new Abstract: Promptable segmentation foundation models (FMs) such as SAM3 and Medical SAM3 promise few-shot, interactively-specified segmentation for medical imaging