Video-Zero: Self-Evolution Video Understanding
arXiv:2605.14733v1 Announce Type: new Abstract: Self-evolution offers a promising path for improving reasoning models without relying on intensive human annotation. However, extending this paradigm to
Knowledge catalogue
arXiv:2605.14733v1 Announce Type: new Abstract: Self-evolution offers a promising path for improving reasoning models without relying on intensive human annotation. However, extending this paradigm to
arXiv:2605.14747v1 Announce Type: cross Abstract: Recent advances in multimodal large language models have driven growing interest in graphical user interface (GUI) agents, yet their generalization re
arXiv:2605.14607v1 Announce Type: new Abstract: Any new medium, once it emerges, is used for more than the transmission of overt content alone. The information it carries typically operates on two lev
arXiv:2605.13923v1 Announce Type: cross Abstract: We study certified runtime monitoring of past-time signal temporal logic (ptSTL) from visual observations under partial observability. The monitor mus
arXiv:2605.14645v1 Announce Type: cross Abstract: With the rapid evolution of computer vision, vision-based methodologies for water level and river surface velocity estimation have reached significant
arXiv:2605.14710v1 Announce Type: cross Abstract: Deep learning and multi-modal fusion have demonstrated transformative potential in medical diagnosis by integrating diverse data sources. However, acc
arXiv:2510.11282v2 Announce Type: replace Abstract: Accurate spatiotemporal traffic forecasting is a critical prerequisite for proactive resource management in dense urban mobile networks. While large
arXiv:2605.14972v1 Announce Type: cross Abstract: A fundamental limitation of Text-to-Code is that no guarantee can be obtained about the correctness of the generated code. Therefore, to ensure its co
arXiv:2602.07045v2 Announce Type: replace-cross Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have enabled complex reasoning. However, existing remote sensing (RS) benchmar
arXiv:2605.14597v1 Announce Type: new Abstract: Precipitation nowcasting is a vital spatio-temporal prediction task for meteorological applications but faces challenges due to the chaotic property of
arXiv:2605.14041v1 Announce Type: cross Abstract: Deep learning excels at prediction but often lacks finite-sample guarantees and calibrated uncertainty; RKHS (Reproducing Kernel Hilbert Space)-based
arXiv:2605.15030v1 Announce Type: cross Abstract: Web agents can autonomously complete online tasks by interacting with websites, but their exposure to open web environments makes them vulnerable to p
arXiv:2605.13959v1 Announce Type: cross Abstract: Generative policies based on diffusion and flow matching have become a dominant paradigm for visuomotor robotic control. We show that replacing the st
arXiv:2605.15182v1 Announce Type: new Abstract: Camera-controlled video generation has made substantial progress, enabling generated videos to follow prescribed viewpoint trajectories. However, existi
arXiv:2605.14405v1 Announce Type: new Abstract: Chaotic systems pose fundamental challenges for data-driven dynamics discovery, as small modeling errors lead to exponentially growing trajectory discre
arXiv:2605.14283v1 Announce Type: cross Abstract: Watermarking techniques for large language models (LLMs), which encode hidden information in the output so its source can be verified, have gained sig
arXiv:2605.14224v1 Announce Type: cross Abstract: We present an in-depth analysis of the Koopman semigroup via wavelet transform. Towards this goal, we start by introducing the wavelet-based observabl
arXiv:2605.14290v1 Announce Type: cross Abstract: ReAct has become the default architecture across LLM agents, and many existing web agents follow this paradigm. We argue that it is the wrong default
arXiv:2603.07880v5 Announce Type: replace Abstract: Moltbook is the first large-scale social network built for autonomous AI agent-to-agent interaction. Early studies on Moltbook have interpreted its
arXiv:2605.11410v2 Announce Type: replace Abstract: Clinical electroencephalogram (EEG) analysis rests on a hand-crafted feature catalog refined over decades, e.g., band power, connectivity, complexit
arXiv:2605.14422v1 Announce Type: new Abstract: Time series forecasting has become increasingly critical in real-world scenarios, where future sequences are influenced not only by historical patterns
arXiv:2605.14257v1 Announce Type: new Abstract: We describe two types of models for vocabulary difficulty prediction: a high-accuracy black-box model, which achieved the top shared task result in the
arXiv:2605.14449v1 Announce Type: cross Abstract: Hallucination detection in large language models (LLMs) requires balancing accu racy, efficiency, and robustness to distribution shift. Black-box cons
arXiv:2605.15183v1 Announce Type: new Abstract: Mechanistic interpretability aims to break models into meaningful parts; verifying that two such parts implement the same computation is a prerequisite.
arXiv:2605.14115v1 Announce Type: new Abstract: Biomedical retrieval-augmented large language models (LLMs) often face evidence that is incomplete, misleading, or internally contradictory, yet evaluat
arXiv:2605.14478v1 Announce Type: cross Abstract: Context: Retrieval-augmented code generation relies on cross-file repository context, but retrieved snippets may come from obsolete project states. Ob
arXiv:2605.14504v1 Announce Type: new Abstract: Long-horizon household tasks demand robust high-level planning and sustained reasoning capabilities, which are largely overlooked by existing embodied A
arXiv:2605.14368v1 Announce Type: cross Abstract: Continuous diffusion language models lag behind autoregressive transformers, partly because diffusion is applied in spaces poorly suited to language d
arXiv:2512.06471v2 Announce Type: replace-cross Abstract: Goal-conditioned reinforcement learning (RL) concerns the problem of training an agent to maximize the probability of reaching target goal sta
arXiv:2605.15109v1 Announce Type: new Abstract: Retrieval-Augmented Generation can improve factuality by grounding answers in external evidence, but Agentic GraphRAG complicates what it means for cita
arXiv:2605.14192v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) has become a powerful and widely used approach for improving large language models by grounding generation in ret
arXiv:2605.15152v1 Announce Type: cross Abstract: LLM quantization has become essential for memory-efficient deployment. Recent work has shown that quantization schemes can pose critical security risk
arXiv:2603.09921v3 Announce Type: replace Abstract: Open-domain visual entity recognition (VER) seeks to associate images with entities in encyclopedic knowledge bases such as Wikipedia. Recent genera
arXiv:2605.13979v1 Announce Type: cross Abstract: Quantum machine learning (QML) aims to accelerate machine learning tasks by exploiting quantum computation. Previous work studied a QML algorithm for
arXiv:2605.14578v1 Announce Type: new Abstract: Partial Dependence Plots (PDPs) visualize how changes in a single feature affect the average model prediction. They are widely used in practice to inter
arXiv:2605.13922v1 Announce Type: cross Abstract: During the last few years, the term Mechanistic Interpretability, a specific area, under the umbrella of explainable artificial intelligence (XAI), ha
arXiv:2605.14754v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed for knowledge synthesis, yet their capacity for compositional generalization in scientific knowle
arXiv:2605.14844v1 Announce Type: cross Abstract: We introduce XFP, a dynamic weight quantizer for LLM inference that inverts the conventional workflow: the operator specifies reconstruction quality f
arXiv:2511.02776v2 Announce Type: replace Abstract: Recent progress in large-scale robotic datasets and vision-language models (VLMs) has advanced research on vision-language-action (VLA) models. Howe
arXiv:2605.14166v1 Announce Type: new Abstract: Face image super-resolution aims to recover high-resolution facial images from severely degraded inputs. Under extreme upscaling factors, fine facial de
arXiv:2605.14893v1 Announce Type: cross Abstract: Contrastively pre-trained Vision-Language Models (VLMs) serve as powerful feature extractors. Yet, their shared latent spaces are prone to structural
arXiv:2605.12586v1 Announce Type: cross Abstract: Vision-language models (VLMs) exhibit a striking paradox: they can generate executable code that reconstructs a 3D scene from geometric primitives wit
arXiv:2605.12689v1 Announce Type: new Abstract: In this paper, we present a novel hybrid approach that combines Reinforcement Learning (RL) with Dynamic Window Approach (DWA) for adaptive 3D local nav
arXiv:2505.21238v3 Announce Type: replace Abstract: Novel view synthesis for underwater scene reconstruction presents unique challenges due to complex light-media interactions. Optical scattering and
arXiv:2603.14066v2 Announce Type: replace-cross Abstract: Many real-world multi-party negotiations unfold as sequences of binding, action-level commitments rather than a single final outcome, yet this
arXiv:2605.13142v1 Announce Type: new Abstract: In professional sports, a team has clinched the playoffs if they are guaranteed a postseason spot, regardless of the outcomes of any remaining games. As
arXiv:2605.12608v1 Announce Type: new Abstract: Object detection in adverse weather is critical for the safety of autonomous vehicles; however, the scarcity of labelled, real-world foggy data remains
arXiv:2506.04165v3 Announce Type: replace Abstract: We consider the Top-K selection problem, which aims to identify the largest K elements in an array. Top-K selection arises in many machine learning
arXiv:2605.12719v1 Announce Type: cross Abstract: The continual assurance of safety and performance of automated driving systems (ADSs) poses significant challenges. ADSs operate in complex, dynamic,
arXiv:2605.13015v1 Announce Type: cross Abstract: The geometry of the retinal vessel is a key biomarker of vascular diseases, yet clinical evidence remains primarily observational. Existing generative
arXiv:2605.13687v1 Announce Type: cross Abstract: We introduce a family of synthetic languages with hierarchical structure -- generated by a broadcast process on trees -- for which the role of context
arXiv:2605.13367v1 Announce Type: cross Abstract: The literature on ontology-mediated query answering (OMQA) has been shaped by two key results: first-order rewritability for DL-Lite, and PTime-hardne
arXiv:2605.13200v1 Announce Type: new Abstract: Accurate state of charge estimation is critical for the success of electric vehicle battery management strategies, but it is well known that conventiona
arXiv:2507.19247v5 Announce Type: replace-cross Abstract: Autoregressive language models achieve remarkable performance, yet a unified theory explaining their internal mechanisms, how training shapes
arXiv:2605.13110v1 Announce Type: cross Abstract: We present a fully automated multi-agent framework for corporate due diligence and market analysis in venture capital. The system runs on an event-dri
arXiv:2605.12706v1 Announce Type: new Abstract: RSNet is an open-source R package that provides a resampling-based framework for robust and interpretable network inference, designed to address the lim
arXiv:2605.12697v1 Announce Type: cross Abstract: Length-dependent logit rescaling is widely used to stabilize long-context self-attention, but existing analyses and methods suggest conflicting invers
arXiv:2605.12685v1 Announce Type: cross Abstract: Graph Self-Supervised Learning (GSSL) has emerged as a powerful paradigm for generating high-quality representations for graph-structured data. While
arXiv:2605.13464v1 Announce Type: new Abstract: Diabetes mellitus affects over 537 million adults worldwide and remains a major challenge in preventive healthcare. Existing machine-learning studies pr
arXiv:2505.12942v4 Announce Type: replace-cross Abstract: Large language models have demonstrated remarkable performance; however, their massive parameter counts make deployment highly expensive. Low-