Positive-Only Drifting Policy Optimization
arXiv:2604.16519v1 Announce Type: new Abstract: In the field of online reinforcement learning (RL), traditional Gaussian policies and flow-based methods are often constrained by their unimodal express
Knowledge catalogue
arXiv:2604.16519v1 Announce Type: new Abstract: In the field of online reinforcement learning (RL), traditional Gaussian policies and flow-based methods are often constrained by their unimodal express
arXiv:2511.23170v5 Announce Type: replace Abstract: Contrastive vision-language pre-training frameworks such as CLIP have demonstrated impressive zero-shot performance across a range of vision-languag
arXiv:2604.17163v1 Announce Type: new Abstract: We propose PPEDCRF, a calibrated selective perturbation framework that protects background-based location privacy in released video frames against galle
arXiv:2604.17338v1 Announce Type: cross Abstract: Unlike code completion, debugging requires localizing faults and applying targeted edits. We observe that frontier LLMs often regenerate correct but o
arXiv:2604.16505v1 Announce Type: new Abstract: The selection of the optimal embryo for transfer is a critical yet challenging step in in vitro fertilization (IVF), primarily due to its reliance on th
arXiv:2604.18085v1 Announce Type: new Abstract: Matrix-level low-rank compression is a promising way to reduce the cost of large language models, but running compression and evaluating the resulting m
arXiv:2604.18316v1 Announce Type: cross Abstract: The most common cause of dementia is Alzheimer disease, a progressive neurodegenerative disorder affecting older adults that gradually impairs memory,
arXiv:2508.20751v2 Announce Type: replace Abstract: Recent advancements highlight the importance of GRPO-based reinforcement learning methods and benchmarking in enhancing text-to-image (T2I) generati
arXiv:2506.13674v3 Announce Type: replace Abstract: Parameter-Efficient Fine-Tuning (PEFT) methods have become crucial for rapidly adapting large language models (LLMs) to downstream tasks. Prefix-Tun
arXiv:2511.07329v3 Announce Type: replace-cross Abstract: It introduces FractalNet, a fractal-inspired computational architectures for advanced large language model analysis that mainly challenges mod
arXiv:2604.16334v1 Announce Type: new Abstract: The use of Deep Neural Network based systems in the real world is growing. They have achieved state-of-the-art performance on many image, speech and tex
arXiv:2508.05132v2 Announce Type: replace Abstract: As medical LLMs transition to clinical deployment, assessing their ethical reasoning capability becomes critical. While achieving high accuracy on k
arXiv:2604.17670v1 Announce Type: new Abstract: We introduce Prior-Fitted Functional Flows, a generative foundation model for pharmacokinetics that enables zero-shot population synthesis and individua
arXiv:2604.16909v1 Announce Type: new Abstract: As large language models (LLMs) evolve from conversational assistants into agents capable of handling complex tasks, they are increasingly deployed in h
arXiv:2604.18354v1 Announce Type: new Abstract: Emotion plays a pivotal role in shaping negotiation outcomes, influencing trust, cooperation, and long-term relationships. Developing negotiation dialog
arXiv:2601.15220v2 Announce Type: replace Abstract: We identify a novel phenomenon in language models: benign fine-tuning of frontier models can lead to privacy collapse. We find that diverse, subtle
arXiv:2604.16523v1 Announce Type: new Abstract: This paper proposes a novel privacy-preserving semantic segmentation method that can use independent keys for each client and image. In the proposed met
arXiv:2510.16054v2 Announce Type: replace-cross Abstract: When users submit queries to Large Language Models (LLMs), their prompts can often contain sensitive data, forcing a difficult choice: Send th
arXiv:2510.18109v4 Announce Type: replace-cross Abstract: Evaluating the usefulness of data before purchase is essential when obtaining data for high-quality machine learning models, yet both model bu
arXiv:2604.17476v1 Announce Type: cross Abstract: Multi-user virtual reality enables immersive interaction. However, rendering avatars for numerous participants on each headset incurs prohibitive comp
arXiv:2505.14412v2 Announce Type: replace-cross Abstract: Effective prompt engineering remains a central challenge in fully harnessing the capabilities of LLMs. While well-designed prompts can dramati
arXiv:2604.17290v1 Announce Type: new Abstract: LLMs are widely used for code generation and mathematical reasoning tasks where they are required to generate structured output. They either need to rea
arXiv:2604.01348v2 Announce Type: replace Abstract: Test-time scaling has emerged as an effective way to improve language models on challenging reasoning tasks. However, most existing methods treat ea
arXiv:2604.17957v1 Announce Type: new Abstract: Process Reward Models (PRMs) have emerged as a powerful tool for providing step-level feedback when evaluating the reasoning of Large Language Models (L
arXiv:2509.26278v4 Announce Type: replace-cross Abstract: Most existing approaches formulate action quality assessment and skill proficiency estimation as discriminative prediction tasks, typically pr
arXiv:2604.17715v1 Announce Type: cross Abstract: Recent advances in large language models for test case generation have improved branch coverage via prompt-engineered mutations. However, they still l
arXiv:2604.18459v1 Announce Type: new Abstract: Visual agents operating in the wild must respond to queries precisely when sufficient evidence first appears in a video stream, a critical capability th
arXiv:2508.10531v3 Announce Type: replace Abstract: Modifications to test-time sampling have emerged as an important extension to diffusion algorithms, with the goal of biasing the generative process
arXiv:2604.17126v1 Announce Type: new Abstract: Vision-language models enable open-vocabulary object grounding through natural language queries, under the implicit assumption that semantically equival
arXiv:2604.17920v1 Announce Type: new Abstract: Synthetic Aperture Radar (SAR) plays a critical role in maritime surveillance, yet deep learning for SAR analysis is limited by the lack of pixel-level
arXiv:2604.18444v1 Announce Type: cross Abstract: Zero-shot vision-language models (VLMs) have shown promise for chest radiograph classification, but their performance is often limited by confounding
arXiv:2406.08334v2 Announce Type: replace-cross Abstract: Memory pressure has emerged as a dominant constraint in scaling the training of large language models (LLMs), particularly in resource-constra
arXiv:2604.16889v1 Announce Type: new Abstract: Existing feature-interpretation pipelines typically operate on uniformly sampled units, but only a small fraction of cross-layer transcoder (CLT) featur
arXiv:2510.08047v2 Announce Type: replace-cross Abstract: Robust ASR under domain shift is crucial because real-world systems encounter unseen accents and domains with limited labeled data. Although p
arXiv:2603.04531v2 Announce Type: replace Abstract: Tactile dexterous manipulation is essential to automating complex household tasks, yet learning effective control policies remains a challenge. Whil
arXiv:2508.02750v2 Announce Type: replace Abstract: This review presents a comprehensive survey and benchmark of pulse shape discrimination (PSD) algorithms for radiation detection, classifying nearly
arXiv:2601.22012v2 Announce Type: replace Abstract: Catastrophic forgetting in continual learning is often measured at the performance or last-layer representation level, overlooking the underlying me
arXiv:2206.14234v3 Announce Type: replace-cross Abstract: In deterministic optimization, it is typically assumed that all problem parameters are fixed and known. In practice, however, some parameters
arXiv:2604.16858v1 Announce Type: new Abstract: Image Quality Assessment (IQA) models are increasingly deployed as perceptual critics to guide generative models and image restoration. This role demand
arXiv:2604.16779v1 Announce Type: cross Abstract: Quantum feature maps offer expressive embeddings for classical learning tasks, and augmenting sparse identification of nonlinear dynamics (SINDy) with
arXiv:2604.16396v1 Announce Type: new Abstract: Islamic inheritance law (ilm al-mawar{i}th) presents a challenging domain for evaluating large language models' structured reasoning capabilities, requi
arXiv:2602.22639v2 Announce Type: replace Abstract: In structure from motion, quadrifocal tensors capture more information than their pairwise counterparts (essential matrices), yet they have often be
arXiv:2604.16432v1 Announce Type: cross Abstract: AI in applications like screening job applicants had become widespread, and may contribute to unemployment especially among the young. Biases in the A
arXiv:2604.17842v1 Announce Type: new Abstract: LLM benchmarks are increasingly dynamic: instead of containing a fixed set of questions, they define templates and parameters that can generate an effec
arXiv:2604.17321v1 Announce Type: new Abstract: Face morphing attacks pose a substantial risk to the reliability of face recognition systems used in passport issuance, border control, and digital iden
arXiv:2506.07826v2 Announce Type: replace Abstract: Validating autonomous driving (AD) systems requires diverse and safety-critical testing, making photorealistic virtual environments essential. Tradi
arXiv:2504.07415v2 Announce Type: replace-cross Abstract: Automated radiology report generation (RRG) holds potential to reduce the workload of radiologists, and recent advances in multimodal large la
arXiv:2510.04008v5 Announce Type: replace Abstract: Softmax Attention has a quadratic time complexity in sequence length, which becomes prohibitive to run at long contexts, even with highly optimized
arXiv:2604.16310v1 Announce Type: cross Abstract: Evaluating Retrieval-Augmented Generation (RAG) systems using static multi-turn datasets fails to capture the dynamic nature of real-world dialogues.
arXiv:2512.24086v2 Announce Type: replace Abstract: In video and image generation tasks, Diffusion Transformer (DiT) models incur extremely high computational costs due to attention mechanisms, which
arXiv:2604.18450v1 Announce Type: cross Abstract: Empirical studies of trained models often report a transient regime in which signal is detectable in a finite gradient descent time window before over
arXiv:2604.16591v1 Announce Type: new Abstract: Large language models (LLMs) sometimes memorize undesirable knowledge, which must be removed after deployment. Prior work on machine unlearning has focu
arXiv:2604.18390v1 Announce Type: new Abstract: In self-supervised learning, self-distilled methods have shown impressive performance, learning representations useful for downstream tasks and even dis
arXiv:2604.17805v1 Announce Type: new Abstract: Pairwise ranking systems based on Maximum Likelihood Estimation (MLE), such as the Bradley-Terry model, are widely used to aggregate preferences from pa
arXiv:2604.18026v1 Announce Type: new Abstract: Many deployed systems expose black-box objectives whose minimizing configuration shifts with an externally observed context. When contexts revisit a sma
arXiv:2601.22002v3 Announce Type: replace Abstract: Transformers achieve superior performance on many tasks, but impose heavy compute and memory requirements during inference. This inference can be ma
arXiv:2307.08336v2 Announce Type: replace Abstract: Despite the numerous applications of convex constraints in Robotics, enforcing them within learning-based frameworks remains an open challenge. Exis
arXiv:2604.17807v1 Announce Type: new Abstract: Text-to-motion (T2M) generation aims to control the behavior of a target character via textual descriptions. Leveraging text-motion paired datasets, exi
arXiv:2604.17530v1 Announce Type: cross Abstract: Posture is a critical factor for beginning instrumental learners. Most students receive instruction only once a week, and during the intervals between
arXiv:2603.19830v2 Announce Type: replace Abstract: Efficient structural perception is essential for mapping and autonomous navigation on resource-constrained robots. Existing 3D methods are computati