Mean Flow Policy Optimization
arXiv:2604.14698v1 Announce Type: new Abstract: Diffusion models have recently emerged as expressive policy representations for online reinforcement learning (RL). However, their iterative generative
Knowledge catalogue
arXiv:2604.14698v1 Announce Type: new Abstract: Diffusion models have recently emerged as expressive policy representations for online reinforcement learning (RL). However, their iterative generative
arXiv:2604.15190v1 Announce Type: cross Abstract: Simulating group-level user behavior enables scalable counterfactual evaluation of merchant strategies without costly online experiments. However, bui
arXiv:2604.14249v1 Announce Type: new Abstract: We introduce Metric-Aware Principal Component Analysis (MAPCA), a unified framework for scale-invariant representation learning based on the generalised
arXiv:2512.05024v3 Announce Type: replace-cross Abstract: As generative AI models are increasingly used to simulate real-world systems, quantifying the ``sim-to-real'' gap is critical. For each input
arXiv:2604.14808v1 Announce Type: new Abstract: Machine unlearning for large language models (LLMs) aims to remove targeted knowledge while preserving general capability. In this paper, we recast LLM
arXiv:2604.14986v1 Announce Type: new Abstract: Safe and efficient assistive planning for visually impaired scenarios remains challenging, since existing methods struggle with multi-objective optimiza
arXiv:2509.23468v3 Announce Type: replace-cross Abstract: Effectively integrating diverse sensory modalities is crucial for robotic manipulation. However, the typical approach of feature concatenation
arXiv:2601.15488v2 Announce Type: replace Abstract: Large Language Models (LLMs) exhibit social biases, which can lead to harmful stereotypes and unfair outcomes. We propose extbf{Multi-Persona Thinki
arXiv:2604.14908v1 Announce Type: new Abstract: We study downlink beam and rate adaptation in a multi-user mmWave MISO system where multiple base stations (BSs), each using analog beamforming from fin
arXiv:2509.06921v2 Announce Type: replace-cross Abstract: Cybersecurity demands both rapid pattern recognition and deliberative reasoning, yet purely neural or purely symbolic approaches each address
arXiv:2604.14706v1 Announce Type: new Abstract: Recent advances in 3D Gaussian Splatting (3DGS) have enabled highly efficient and photorealistic novel view synthesis. However, segmenting objects accur
arXiv:2604.14595v1 Announce Type: new Abstract: This position paper argues that recent progress with diversity in NLP is disproportionately concentrated on a small number of areas surrounding fairness
arXiv:2603.13933v2 Announce Type: replace Abstract: Ensuring the safety and compliance of large language models (LLMs) is of paramount importance. However, existing LLM safety datasets often rely on a
arXiv:2604.14243v1 Announce Type: new Abstract: Real-world decision-making systems operate in environments where state transitions depend not only on the agent's actions, but also on extbf{exogenous f
arXiv:2601.08310v2 Announce Type: replace Abstract: Recent Large Reasoning Models (LRMs) achieve strong performance by leveraging long-form Chain-of-Thought (CoT) reasoning, but uniformly applying ove
arXiv:2505.20761v3 Announce Type: replace Abstract: While the performance of machine learning systems has experienced significant improvement in recent years, relatively little attention has been paid
arXiv:2603.13683v2 Announce Type: replace Abstract: Although debiased large language models (LLMs) excel at handling known or low-bias prompts, they often fail on unfamiliar and high-bias prompts. We
arXiv:2509.21823v2 Announce Type: replace Abstract: Reward is critical to the evaluation and training of large language models (LLMs). However, existing rule-based or model-based reward methods strugg
arXiv:2604.14352v1 Announce Type: cross Abstract: Online A/B testing at scale relies on proxy metrics -- short-term, easily-measured signals used in place of slow-moving long-term outcomes. When the p
arXiv:2604.14634v1 Announce Type: new Abstract: Multiple choice evaluation is widely used for benchmarking large language models, yet near ceiling accuracy in low option settings can be sustained by s
arXiv:2604.14175v1 Announce Type: new Abstract: We present a unified system addressing both Subtask 3 (answer generation) and Subtask 4 (evidence sentence alignment) of the ArchEHR-QA Shared Task. For
arXiv:2604.15281v1 Announce Type: new Abstract: 3D policy learning promises superior generalization and cross-embodiment transfer, but progress has been hindered by training instabilities and severe o
arXiv:2604.15308v1 Announce Type: new Abstract: High-level autonomous driving requires motion planners capable of modeling multimodal future uncertainties while remaining robust in closed-loop interac
arXiv:2604.14951v1 Announce Type: cross Abstract: Tool learning with foundation models aims to endow AI systems with the ability to invoke external resources -- such as APIs, computational utilities,
arXiv:2604.14888v1 Announce Type: new Abstract: Recent advances in vision language models (VLMs) offer reasoning capabilities, yet how these unfold and integrate visual and textual information remains
arXiv:2502.05740v2 Announce Type: replace-cross Abstract: Cancer surgery is a key treatment for gastrointestinal (GI) cancers, a group of cancers that account for more than 35% of cancer-related death
arXiv:2604.14265v1 Announce Type: new Abstract: We study behavior-regularized reinforcement learning (RL), where regularization toward a reference distribution (the dataset in offline RL or the base m
arXiv:2604.14910v1 Announce Type: new Abstract: Achieving high-fidelity generation in extremely few sampling steps has long been a central goal of generative modeling. Existing approaches largely rely
arXiv:2604.15201v1 Announce Type: new Abstract: As reinforcement learning (RL) deployments expand into safety-critical domains, existing evaluation methods fail to systematically identify hazards aris
arXiv:2604.14353v1 Announce Type: new Abstract: Localization of autonomous mobile robots (AMRs) in enclosed or semi-enclosed environments such as offices, hotels, hospitals, indoor parking facilities,
arXiv:2509.12833v2 Announce Type: replace Abstract: Projection-based safety filters, which modify unsafe actions by mapping them to the closest safe alternative, are widely used to enforce safety cons
arXiv:2604.14373v1 Announce Type: new Abstract: Rural environmental risks are shaped by place-based conditions (e.g., housing quality, road access, land-surface patterns), yet standard vulnerability i
arXiv:2604.14474v1 Announce Type: new Abstract: Traditional esports scouting workflows rely heavily on manual video review and aggregate performance metrics, which often fail to capture the nuanced de
arXiv:2604.14163v1 Announce Type: new Abstract: Maritime distress communications transmitted over very high frequency (VHF) radio are safety-critical voice messages used to report emergencies at sea.
arXiv:2603.27833v2 Announce Type: cross Abstract: In this work, we first prove that the separation principle holds for communication-constrained LQR problems under i.i.d. zero-mean disturbances with a
arXiv:2604.14672v1 Announce Type: new Abstract: Large language models (LLMs) are being increasingly used in urban planning, but since gendered space theory highlights how gender hierarchies are embedd
arXiv:2604.14379v1 Announce Type: cross Abstract: Reinforcement learning (RL) has emerged as a powerful tool for aligning diffusion models with human preferences, typically by optimizing a single rewa
arXiv:2604.14631v1 Announce Type: new Abstract: Effective code generation requires both model capability and a problem representation that carefully structures how models reason and plan. Existing app
arXiv:2604.14629v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have shown remarkable capabilities in joint vision-language understanding, but their large scale poses significant challen
arXiv:2604.14834v1 Announce Type: new Abstract: Recent advancements in whole-body control through deep reinforcement learning have enabled humanoid robots to achieve remarkable progress in real-world
arXiv:2604.14414v1 Announce Type: new Abstract: Turn-level metrics are widely used to evaluate properties of multi-turn human-LLM conversations, from safety and sycophancy to dialogue quality. However
arXiv:2604.14176v1 Announce Type: new Abstract: Generalized Category Discovery (GCD) leverages labeled data to categorize unlabeled samples from known or unknown classes. Most previous methods jointly
arXiv:2604.14807v1 Announce Type: cross Abstract: The rapid integration of large language models (LLMs) into everyday workflows has transformed how individuals perform cognitive tasks such as writing,
arXiv:2604.14197v1 Announce Type: new Abstract: Large language model (LLM) performance depends heavily on prompt design, yet prompt construction is often described and applied inconsistently. Our purp
“The sharpest drop came from people who used the model for direct answers, not from those who used it more like a hint system, which suggests the real issue is not AI exposure itself but replacing eff
arXiv:2512.03048v4 Announce Type: replace-cross Abstract: Static content-based AI value alignment is insufficient for robust alignment under capability scaling, distributional shift, and increasing au
arXiv:2603.18373v2 Announce Type: replace Abstract: When VLMs answer correctly, do they genuinely rely on visual information or exploit language shortcuts? We introduce the Tri-Layer Diagnostic Framew
arXiv:2604.14237v1 Announce Type: new Abstract: Transistor topology optimization is a critical step in standard cell design, directly dictating diffusion sharing efficiency and downstream routability.
arXiv:2506.09457v3 Announce Type: replace Abstract: Direct Alignment Algorithms (DAAs), such as Direct Preference Optimization (DPO) and Simple Preference Optimization (SimPO), have emerged as efficie
arXiv:2511.14178v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have demonstrated significant potential in real-world robotic manipulation. However, pre-trained VLA policies st
arXiv:2604.13715v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) enable general audio understanding and demonstrate remarkable performance across various audio tasks. However, the
arXiv:2604.13488v1 Announce Type: new Abstract: Autonomous Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) enable digital automation on end-user devices. Whil
arXiv:2604.14787v1 Announce Type: cross Abstract: Network Digital Twins (NDTs) enable safe what-if analysis for 6G cloud-edge infrastructures, but adoption is often limited by fragmented workflows fro
arXiv:2604.14209v1 Announce Type: new Abstract: As deep neural networks are deployed in safety-critical domains such as autonomous driving and medical diagnosis, stakeholders need explanations that ar
arXiv:2604.15074v1 Announce Type: new Abstract: This paper presents a two-stage trajectory planning framework for a multi-UAV rigid-payload cascaded transportation system, aiming to address planning c
arXiv:2511.07412v2 Announce Type: replace Abstract: Developing embodied AI for intelligent surgical systems requires safe, controllable environments for continual learning and evaluation. However, saf
arXiv:2604.14967v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) extends Large Vision-Language Models (LVLMs) with external visual knowledge. However, existing visual RAG systems t
arXiv:2604.15196v1 Announce Type: new Abstract: We propose a novel hierarchical spatiotemporal vector quantization framework for unsupervised skeleton-based temporal action segmentation. We first intr
arXiv:2604.15221v1 Announce Type: cross Abstract: We propose a framework for vision-based human pose estimation and motion prediction that gives conformal prediction guarantees for certifiably safe hu
arXiv:2604.14548v1 Announce Type: cross Abstract: As speech language models (SLMs) transition from personal devices into shared, multi-user environments, their responses must account for far more than