A Markov Chain Approach to Preference Alignment
arXiv:2606.22652v1 Announce Type: new Abstract: We propose Markov Chain from Human Feedback (MCHF), an elementary approach for aligning generative models from pairwise human preferences. Unlike Reinfo
Knowledge catalogue
arXiv:2606.22652v1 Announce Type: new Abstract: We propose Markov Chain from Human Feedback (MCHF), an elementary approach for aligning generative models from pairwise human preferences. Unlike Reinfo
arXiv:2606.20031v2 Announce Type: replace Abstract: Dynamic environmental changes, confined workspaces, and stringent real-time constraints make pathfinding in Robotic Mobile Fulfillment Systems (RMFS
arXiv:2606.21509v1 Announce Type: new Abstract: End-to-end autonomous driving systems tightly couple perception and decision-making through latent representations. Consequently, updates to perception
arXiv:2606.22360v1 Announce Type: new Abstract: Successful conversations require speakers to align on the meaning of concepts, a challenging but crucial task for human-robot interaction. Understanding
arXiv:2606.20681v1 Announce Type: new Abstract: Slope hazards constitute a major safety threat to expressway infrastructure, and their evolution is typically manifested as slow surface deformation. Co
arXiv:2606.23574v1 Announce Type: cross Abstract: Vision-language-action (VLA) models and world-action models (WAM) are the generative models now driving general-purpose robot control, turning raw cam
arXiv:2606.23662v1 Announce Type: cross Abstract: Bayesian experimental design (BED) has traditionally been based on maximising expected uncertainty reductions from prior to posterior. A major shortfa
arXiv:2602.02451v2 Announce Type: replace Abstract: Discovering causal relationships requires controlled experiments, but experimentalists face a sequential decision problem: each intervention reveals
arXiv:2606.20880v1 Announce Type: cross Abstract: Decision-making under partial or adversarial observability requires accurate inference of the environment's latent state and its associated uncertaint
Gary Marcus argues that data security is fundamentally fragile and that mass surveillance and data harvesting practices create persistent security vulnerabilities that cannot be eliminated. The post e
arXiv:2606.21013v1 Announce Type: cross Abstract: Forecasting future events is a critical challenge for large language model (LLM) agents, spanning domains from elections and monetary policy to financ
AI Capex pushback from GS: 'If frontier intelligence can increasingly be developed in the East at a fraction of the cost incurred in the West … then the largest capital allocators are also the ones mo
AI data centres are hungry for land, water & power. I’m calling on every major AI company to publicly disclose the full environmental impact of its systems – as a matter of transparency No more hidden
arXiv:2602.12691v3 Announce Type: replace Abstract: We study how to improve large foundation vision-language-action (VLA) systems through human-in-the-loop reinforcement learning (RL) in real-world en
Although I supported Musk in his suit against OpenAI, and admire what he did for electric cars, I am afraid my considered overall view is not very different from this: Candidly I have no idea why anyo
arXiv:2606.20640v1 Announce Type: cross Abstract: Autonomous vehicles offer the potential for safer and more efficient mobility, yet public trust remains limited due to the lack of transparency in the
arXiv:2606.22278v1 Announce Type: cross Abstract: Ensuring safety of learning-enabled robotic manipulation across diverse embodiments and tasks still requires significant manual engineering. Existing
arXiv:2505.10022v4 Announce Type: replace Abstract: Learning natural, animal-like locomotion from demonstrations has become a core paradigm in legged robotics. While motion tracking can reproduce refe
arXiv:2606.20687v1 Announce Type: new Abstract: Multi-Camera Multi-Target (MCMT) tracking has emerged as a critical capability for applications ranging from autonomous driving to animal behavior monit
arXiv:2606.22480v1 Announce Type: new Abstract: Learning visuomotor policies for long-horizon manipulation remains a fundamental challenge. Recent skill-based imitation learning methods based on discr
arXiv:2606.21525v1 Announce Type: new Abstract: Model-free reinforcement learning algorithms such as Proximal Policy Optimization (PPO) treat the environment as a black box, estimating policy gradient
arXiv:2606.21172v1 Announce Type: new Abstract: Video world models are increasingly used in autonomous driving to forecast future scene evolution and provide future-aware spatio-temporal representatio
arXiv:2606.21498v1 Announce Type: cross Abstract: Autoregressive text-to-image (T2I) generation has recently advanced rapidly, yet aligning generated images with human preferences remains challenging.
arXiv:2606.20701v1 Announce Type: cross Abstract: Learned communication improves coordination in cooperative multi-agent reinforcement learning, but it also creates a trust problem: a trained policy m
arXiv:2606.21014v1 Announce Type: new Abstract: Robots must generate trajectories that remain faithful to learned expert behavior while satisfying safety constraints and task-specific objectives speci
arXiv:2606.21645v1 Announce Type: cross Abstract: Large language models (LLMs) can readily reproduce conventional expressions, yet their ability to model gradient frequency distributions remains under
arXiv:2606.20812v1 Announce Type: new Abstract: EEG foundation models can learn generalizable representations from large-scale EEG corpora to enable single-backbone transfer across diverse clinical an
arXiv:2606.23531v1 Announce Type: new Abstract: Endoscopic retrograde cholangiopancreatography (ERCP) demands precise endoscopic navigation and stable biliary cannulation within a narrow monocular fie
arXiv:2606.22456v1 Announce Type: new Abstract: Maintaining accurate navigation during GNSS outages remains a significant challenge for autonomous systems relying on low-cost inertial sensors. While c
arXiv:2601.22100v3 Announce Type: replace Abstract: Optimizing Conditional Value-at-risk (CVaR) using policy gradient (a.k.a CVaR-PG) faces significant challenges of sample inefficiency. This ineffici
arXiv:2606.23199v1 Announce Type: new Abstract: Attributed graph clustering partitions nodes by jointly exploiting node attributes and graph topology. It remains challenging due to attribute heterogen
arXiv:2606.22389v1 Announce Type: new Abstract: Singular Learning Theory leverages the Local Learning Coefficient (LLC) to quantify the geometry of neural network loss landscapes. However, mean-energy
arXiv:2606.21981v1 Announce Type: cross Abstract: While Large Language Models (LLMs) can generate fluent Arabic text, their ability to reliably control readability levels remains unclear. We propose a
arXiv:2606.21809v1 Announce Type: new Abstract: The presence of confounding bias poses a key challenge in policy evaluation, as the target causal effects of actions are not identifiable (i.e., underde
arXiv:2505.01652v2 Announce Type: replace Abstract: Fair machine learning seeks to identify and mitigate biases in predictions against unfavorable populations characterized by demographic attributes,
arXiv:2606.23206v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in multimodal reasoning. However, prevailing reinforcement learning (RL)
arXiv:2602.03389v2 Announce Type: replace Abstract: Offline goal-conditioned reinforcement learning remains challenging for long-horizon tasks. While hierarchical approaches mitigate this issue by dec
arXiv:2507.08262v2 Announce Type: replace-cross Abstract: The spatial information inherent in 3D point clouds is crucial for robotic manipulation. However, existing 3D pre-training methods face a fund
arXiv:2606.19380v2 Announce Type: replace-cross Abstract: Software engineering and deployment are increasingly delegated to AI coding agents. The scale of their adoption is surfacing rare, but highly
arXiv:2603.22169v2 Announce Type: replace Abstract: We propose a new Verbal Reinforcement Learning (VRL) framework for interpretable task-level planning in mobile robotic systems operating under execu
arXiv:2606.21885v1 Announce Type: new Abstract: Foundation models have achieved remarkable performance across medical question answering, imaging, and electronic health record (EHR) tasks, yet reliabl
arXiv:2606.22963v1 Announce Type: new Abstract: Concept segmentation models like Segment Anything Model 3 (SAM3) show strong generalization on natural images, yet their performance degrades in medical
arXiv:2602.03762v3 Announce Type: replace-cross Abstract: Visually-guided acoustic highlighting seeks to rebalance audio in alignment with the accompanying video, creating a coherent audio-visual expe
arXiv:2602.11973v2 Announce Type: replace Abstract: In critical decision support systems based on medical imaging, the reliability of AI-assisted decision-making is as relevant as predictive accuracy.
arXiv:2606.21577v1 Announce Type: new Abstract: Quadratic programs (QPs) using Control Barrier Functions (CBFs) and Control Lyapunov Functions (CLFs) are widely used for safe control in reach-and-avoi
arXiv:2606.21900v1 Announce Type: cross Abstract: Continuous authentication for mobile and zero-trust systems requires nonintrusive evidence confirming the enrolled user-device context remains valid a
arXiv:2606.21021v1 Announce Type: new Abstract: Long-horizon spacecraft trajectory forecasting suffers from error accumulation due to the absence of corrective observations in the forecast regime, mak
arXiv:2606.23680v1 Announce Type: cross Abstract: Humanoid loco-manipulation is often simplified into a stop-and-go process: walking to an object, stopping to manipulate it, and then resuming locomoti
arXiv:2606.21627v1 Announce Type: cross Abstract: As agentic systems tackle increasingly complex multi-step tasks, evaluating their trajectories presents a major bottleneck - human annotation of a sin
arXiv:2606.23015v1 Announce Type: new Abstract: Optimizing instructional policies in Intelligent Tutoring Systems (ITS) typically requires costly online experimentation or student simulators that may
arXiv:2606.21613v1 Announce Type: new Abstract: Scaling wildlife monitoring for real-world conservation deployments requires automated analysis of smart sensors that operate under severe annotation sc
arXiv:2606.22347v1 Announce Type: new Abstract: Identity-Preserving Text-to-Video Generation (IPT2V) seeks to synthesize a temporally coherent video from a reference image and a textual description, w
arXiv:2606.20725v1 Announce Type: new Abstract: Accurate, up-to-date representations of road structures are critical for the safe operation of autonomous vehicles. Existing systems rely either on cost
arXiv:2606.20622v1 Announce Type: cross Abstract: The goal of artificial intelligence is to create agents capable of general, adaptive behaviour in open-ended environments. Guided by the 'Bitter Lesso
arXiv:2511.20906v2 Announce Type: replace Abstract: Diffusion- and flow-based policies deliver state-of-the-art performance on long-horizon robotic manipulation and imitation learning tasks. However,
arXiv:2602.10155v2 Announce Type: replace-cross Abstract: Accurate compensation of brain deformation is critical for reliable image-guided neurosurgery. Surgical manipulation and tumor resection induc
arXiv:2505.09603v2 Announce Type: replace-cross Abstract: Recently, the robotics community has amassed ever larger and more diverse datasets to train generalist policies. However, while these policies
arXiv:2606.22829v1 Announce Type: new Abstract: Intraoperative Adverse Events (IAEs) detection is critical for improving surgical safety, with bleeding being among the most frequent events across many
arXiv:2010.14694v4 Announce Type: replace-cross Abstract: This paper integrates deep neural networks (DNNs) into structural models to increase flexibility and capture rich heterogeneity while preservi
arXiv:2606.22159v1 Announce Type: cross Abstract: Spacecraft operations scheduling is a highly constrained, long-horizon combinatorial optimization problem that traditionally relies on heuristics, con