TIP: Token Importance in On-Policy Distillation
arXiv:2604.14084v1 Announce Type: new Abstract: On-policy knowledge distillation (OPD) trains a student on its own rollouts under token-level supervision from a teacher. Not all token positions matter
Knowledge catalogue
arXiv:2604.14084v1 Announce Type: new Abstract: On-policy knowledge distillation (OPD) trains a student on its own rollouts under token-level supervision from a teacher. Not all token positions matter
arXiv:2604.13368v1 Announce Type: new Abstract: Fine-tuning large language models (LLMs) aims to adapt pre-trained models to specific tasks using relatively small and domain-specific datasets. Among P
arXiv:2602.21627v3 Announce Type: replace Abstract: This paper presents a new unified approach to semantic segmentation in both images and videos by using language modeling to output the masks as sequ
arXiv:2604.13787v1 Announce Type: new Abstract: Large Language Models (LLMs) enhance their problem-solving capability by utilizing external tools. However, in open-world scenarios with massive and evo
arXiv:2604.13519v1 Announce Type: new Abstract: Tool calling has greatly expanded the practical utility of large language models (LLMs) by enabling them to interact with external applications. As LLM
arXiv:2603.15620v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models excel in static manipulation but struggle in dynamic environments with moving targets. This performance gap prim
arXiv:2604.14013v1 Announce Type: cross Abstract: We promote in this paper the processing of radar data in the frequency domain to achieve higher robustness against noise and structural errors, especi
arXiv:2604.13186v1 Announce Type: new Abstract: Unsafe surgical care is a critical health concern, often linked to limitations in surgeon experience, skills, and situational awareness. Integrating pat
arXiv:2604.13322v1 Announce Type: new Abstract: Raveling, the loss of aggregates, is a major form of asphalt pavement surface distress, especially on highways. While research has shown that machine le
arXiv:2604.14069v1 Announce Type: new Abstract: Human-Object Interaction (HOI) detection is a longstanding computer vision problem concerned with predicting the interaction between humans and objects.
arXiv:2604.14074v1 Announce Type: new Abstract: Semantic Multi-Object Tracking (SMOT) extends multi-object tracking with semantic outputs such as video summaries, instance-level captions, and interact
arXiv:2604.13552v1 Announce Type: new Abstract: Large language models (LLMs) demonstrate strong reasoning capabilities, but their performance often degrades under distribution shift. Existing test-tim
arXiv:2604.05478v2 Announce Type: replace-cross Abstract: Immune checkpoint inhibitors (ICIs) have transformed cancer therapy; yet substantial proportion of patients exhibit intrinsic or acquired resi
arXiv:2604.14116v1 Announce Type: cross Abstract: While Large Language Models (LLMs) have empowered AI research agents to perform isolated scientific tasks, automating complex, real-world workflows, s
arXiv:2601.10245v2 Announce Type: replace-cross Abstract: Multi-step reasoning tasks like mathematical problem solving are vulnerable to cascading failures, where a single incorrect step leads to comp
arXiv:2601.07422v2 Announce Type: replace Abstract: Despite their impressive capabilities, large language models (LLMs) frequently generate hallucinations. Previous work shows that their internal stat
arXiv:2505.18232v3 Announce Type: replace-cross Abstract: The deployment of large language models (LLMs) is largely hindered by their large number of parameters. Structural pruning has emerged as a pr
arXiv:2604.13565v1 Announce Type: new Abstract: Ultra-high-resolution (UHR) remote sensing imagery couples kilometer-scale context with query-critical evidence that may occupy only a few pixels. Such
arXiv:2604.13822v1 Announce Type: new Abstract: MLLM-based GUI agents have demonstrated strong capabilities in complex user interface interaction tasks. However, long-horizon scenarios remain challeng
arXiv:2604.14113v1 Announce Type: cross Abstract: GUI grounding, which localizes interface elements from screenshots given natural language queries, remains challenging for small icons and dense layou
arXiv:2604.14089v1 Announce Type: new Abstract: We present UMI-3D, a multimodal extension of the Universal Manipulation Interface (UMI) for robust and scalable data collection in embodied manipulation
arXiv:2603.08639v2 Announce Type: replace Abstract: Ensuring trustworthiness in open-world visual recognition requires models that are interpretable, fair, and robust to distribution shifts. Yet moder
arXiv:2604.13383v1 Announce Type: new Abstract: Ambient Lighting Normalization (ALN) aims to restore images degraded by complex, spatially varying illumination conditions. Existing methods, such as IF
arXiv:2511.23332v2 Announce Type: replace Abstract: Instruction-driven segmentation in remote sensing generates masks from guidance, offering great potential for accessible and generalizable applicati
arXiv:2604.13470v1 Announce Type: new Abstract: We prove that conditional diffusion models whose reverse kernels are finite Gaussian mixtures with ReLU-network logits can approximate suitably regular
arXiv:2604.13197v1 Announce Type: new Abstract: Process reward models (PRMs) provide fine-grained reward signals along the reasoning process, but training reliable PRMs often requires step annotations
arXiv:2604.13584v1 Announce Type: new Abstract: We present UNRIO, an uncertainty-aware radar-inertial odometry system that estimates ego-velocity directly from raw mmWave radar IQ signals rather than
arXiv:2604.13928v1 Announce Type: new Abstract: Industrial time-series data from real production environments exhibits substantially higher complexity than commonly used benchmark datasets, primarily
arXiv:2604.13988v1 Announce Type: new Abstract: Objective: Investigate whether hypnogram 'realism' can be used to guide an unsupervised method for handling arbitrary types of signal degradation in mob
arXiv:2604.13502v1 Announce Type: new Abstract: Social Determinants of Health (SDOH) refer to environmental, behavioral, and social conditions that influence how individuals live, work, and age. SDOH
arXiv:2604.13309v1 Announce Type: new Abstract: In this paper we present a novel visual servoing framework to control a robotic manipulator in the configuration space by using purely natural visual fe
arXiv:2604.06484v2 Announce Type: replace Abstract: Cultural values are expressed not only through language but also through visual scenes and everyday social practices. Yet existing evaluations of cu
arXiv:2604.13323v1 Announce Type: new Abstract: Many robot planning tasks require satisfaction of one or more constraints throughout the entire trajectory. For geometric constraints, manifold-constrai
arXiv:2604.13596v1 Announce Type: new Abstract: Instance-level object segmentation across disparate egocentric and exocentric views is a fundamental challenge in visual understanding, critical for app
arXiv:2604.13425v1 Announce Type: new Abstract: Video chroma-lux editing, which aims to modify illumination and color while preserving structural and temporal fidelity, remains a significant challenge
arXiv:2512.14234v2 Announce Type: replace Abstract: Human communication is inherently multimodal and social: words, prosody, and body language jointly carry intent. Yet most prior systems model human
arXiv:2604.13687v1 Announce Type: cross Abstract: Glitches frequently contaminate data in gravitational-wave detectors, complicating the observation and analysis of astrophysical signals. This work in
arXiv:2604.13654v1 Announce Type: new Abstract: Vision-and-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) represents a pivotal challenge in embodied artificial intelligence, focused on ena
arXiv:2603.08486v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) face safety misalignment, where visual inputs enable harmful outputs. To address this, existing methods req
arXiv:2506.01247v2 Announce Type: replace Abstract: Steering vision foundation models at test time, without updating foundation-model weights or using labeled target data, is a desirable yet challengi
arXiv:2604.02486v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have achieved impressive performance across a wide range of multimodal tasks. However, they often fail on tasks
arXiv:2604.13660v1 Announce Type: new Abstract: In Deepfake Detection (DFD) tasks, researchers proposed two types of MLLM-based methods: complementary combination with small DFD detectors, or static f
arXiv:2604.13204v1 Announce Type: new Abstract: The motion planning problem requires finding a collision-free path between start and goal configurations in high-dimensional, cluttered spaces. Recent l
arXiv:2604.13318v1 Announce Type: cross Abstract: Autonomous web agents powered by large language models (LLMs) have shown promise in completing complex browser tasks, yet they still struggle with lon
arXiv:2604.13610v1 Announce Type: new Abstract: In computer vision, a prevailing method for quantifying dataset bias is to train a model to distinguish between datasets. High classification accuracy i
arXiv:2604.13349v1 Announce Type: new Abstract: Communication in Large Language Model (LLM)-based multi-agent systems is moving beyond discrete tokens to preserve richer context. Recent work such as L
arXiv:2503.23137v2 Announce Type: replace-cross Abstract: Understanding humor-particularly when it involves complex, contradictory narratives that require comparative reasoning-remains a significant c
arXiv:2604.13776v1 Announce Type: cross Abstract: Watermarking is becoming the default mechanism for AI content authentication, with governance policies and frameworks referencing it as infrastructure
arXiv:2604.13321v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) struggle with tasks that require reasoning about 2D object orientation in images, as documented in prior work.
arXiv:2604.13403v1 Announce Type: new Abstract: In-context learning (ICL) enables models to adapt to new tasks via inference-time demonstrations. Despite its success in large language models, the exte
arXiv:2604.13438v1 Announce Type: new Abstract: Privacy concerns in LLMs have led to the rapidly growing need to enforce a data's 'right to be forgotten'. Machine unlearning addresses precisely this t
arXiv:2603.24959v2 Announce Type: replace Abstract: Wireless bioelectronic interfaces are increasingly used to control tissue-engineered biohybrid robotic systems. However, a unifying engineering fram
arXiv:2603.26259v2 Announce Type: replace-cross Abstract: While Late Interaction models exhibit strong retrieval performance, many of their underlying dynamics remain understudied, potentially hiding
arXiv:2604.13055v1 Announce Type: new Abstract: Today's evolving labor markets rely increasingly on recommender systems for hiring, talent management, and workforce analytics, with natural language pr
arXiv:2511.04671v2 Announce Type: replace-cross Abstract: Human videos are a scalable source of training data for robot learning. However, humans and robots significantly differ in embodiment, making
arXiv:2604.13556v1 Announce Type: new Abstract: Cross-layer key-value (KV) compression has been found to be effective in efficient inference of large language models (LLMs). Although they reduce the m
arXiv:2511.05757v3 Announce Type: replace-cross Abstract: We introduce a differentiable framework for zero-shot adaptive control over parametric families of nonlinear dynamical systems. Our approach i
arXiv:2512.09953v2 Announce Type: replace-cross Abstract: Machine unlearning aims to remove the influence of specific data points from a trained model to satisfy privacy, copyright, and safety require
arXiv:2604.13568v1 Announce Type: new Abstract: Wideband spectrum sensing for low-altitude monitoring is critical yet challenging due to heterogeneous protocols,large bandwidths, and non-stationary SN
arXiv:2604.12027v1 Announce Type: new Abstract: Recently, the robotics community has regained interest in radar-based perception and state estimation. A 2D imaging radar provides dense 360deg informat