World Models for Robotic Manipulation: A Survey
arXiv:2606.00113v1 Announce Type: new Abstract: Robotic manipulation depends on the ability to anticipate how actions reshape objects, contacts, and scene geometry before execution. Learned world mode
Knowledge catalogue
arXiv:2606.00113v1 Announce Type: new Abstract: Robotic manipulation depends on the ability to anticipate how actions reshape objects, contacts, and scene geometry before execution. Learned world mode
arXiv:2606.02027v1 Announce Type: cross Abstract: Robot learning must produce policies that generalize to new combinations of constraints, teammates, and environments. To achieve this, we must structu
arXiv:2605.24892v2 Announce Type: replace Abstract: Physical world knowledge resides mainly in videos. Equipping Vision-Language-Action (VLA) models with such knowledge is fundamental for safe and gen
arXiv:2602.03896v2 Announce Type: replace-cross Abstract: Poisson-distributed latent variable models are widely used in computational neuroscience, but differentiating through discrete stochastic samp
arXiv:2605.30843v1 Announce Type: new Abstract: In the forward reinforcement-learning problem, the reward is fixed and known; the learner is asked to find a good policy or value function. Here we turn
arXiv:2605.31021v1 Announce Type: new Abstract: Current alignment paradigms for generative artificial intelligence rely predominantly on monolithic benchmarking frameworks that reduce the plurality of
arXiv:2605.30989v1 Announce Type: new Abstract: Robot teleoperation enables safe, non-contact task execution in hazardous environments where direct human access is difficult, and its application has e
arXiv:2605.30452v1 Announce Type: cross Abstract: Many machine learning problems involve multiple inherent trade-offs that are best addressed by gradient-based multi-objective optimization (MOO) algor
arXiv:2605.30625v1 Announce Type: cross Abstract: Inferring continuous probability paths from sparse snapshots is a fundamental challenge in domains like single-cell biology, where high-fidelity data
arXiv:2605.31405v1 Announce Type: new Abstract: This paper addresses the challenge of simultaneously compensating for state-dependent uncertainties and enforcing time-varying state constraints in Eule
arXiv:2605.30406v1 Announce Type: cross Abstract: Recent research demonstrating AI systems exhibiting deception and shutdown resistance suggests that AI loss of control (LOC) is an urgent policy conce
arXiv:2605.31034v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) and group-based policy optimization methods such as GRPO update a stochastic policy by sampling
arXiv:2605.30742v1 Announce Type: new Abstract: This paper addresses the task of temporal sentence grounding (TSG). Although many respectable works have made decent achievements in this important topi
arXiv:2605.31490v1 Announce Type: new Abstract: On-policy distillation (OPD) provides dense teacher feedback along rollouts generated by the student and has emerged as a promising post-training paradi
arXiv:2605.30585v1 Announce Type: cross Abstract: Effective prognostics and health management of modern engines relies on accurate turbine gas temperature predictions and robust uncertainty quantifica
arXiv:2605.31153v1 Announce Type: new Abstract: Given the surge of harmful AI-generated imagery online, reliably distinguishing authentic images from generated ones has become an urgent research topic
arXiv:2602.10117v5 Announce Type: replace-cross Abstract: Large Language Models (LLMs) often provide chain-of-thought (CoT) reasoning traces that appear plausible, but may hide internal biases. We cal
arXiv:2605.30972v1 Announce Type: new Abstract: Accurate 3D medical image segmentation requires both long-range volumetric context and fine boundary preservation. CNN-based methods have limited global
arXiv:2510.11683v3 Announce Type: replace-cross Abstract: A key challenge in applying reinforcement learning (RL) to diffusion large language models (dLLMs) is the intractability of their likelihood f
arXiv:2411.13865v4 Announce Type: replace-cross Abstract: Modern recommender systems often create information cocoons, restricting users' exposure to diverse content. The central challenge is to balan
arXiv:2605.31110v1 Announce Type: new Abstract: Generalization in robotics requires prior knowledge about how the world is structured, yet this structure changes from one situation to the next. This p
arXiv:2601.22412v2 Announce Type: replace Abstract: Video-based human movement analysis holds potential for movement assessment in clinical practice and research. However, the clinical implementation
arXiv:2605.31066v1 Announce Type: new Abstract: Recent aerial vision-language-action (VLA) models show promising single-UAV capabilities, such as tracking moving objects and navigating to language-spe
arXiv:2602.02819v3 Announce Type: replace Abstract: Membership Inference Attacks (MIAs) aim to distinguish training points (members) from unseen data (non-members), and are widely used to quantify mem
arXiv:2605.30635v1 Announce Type: new Abstract: Inferring dynamics from population snapshots is a fundamental challenge in machine learning and biology. In scRNA-sequencing (scRNA-seq), destructive me
arXiv:2605.30641v1 Announce Type: cross Abstract: Large language models (LLMs) can reveal and amplify societal biases during chain-of-thought (CoT) generation. We present COFT (Chain of Fair Thought),
arXiv:2605.30838v1 Announce Type: new Abstract: LLM-powered search agents enable multi-step reasoning and tool use. However, these capabilities introduce retrieval-induced safety degradation, as harmf
arXiv:2605.30487v1 Announce Type: new Abstract: Aligning large language models (LLMs) to heterogeneous and rapidly evolving safety requirements remains a critical challenge. Existing instruction-tuned
arXiv:2601.06453v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly grounded in sensor data to perceive and reason about human physiology and the physical world. However,
arXiv:2605.31073v1 Announce Type: new Abstract: Reasoning-based LLM guardrails improve safety moderation by generating explicit rationales before issuing final decisions. However, their rationales do
arXiv:2605.31388v1 Announce Type: new Abstract: Multi-Objective Reinforcement Learning (MORL) extends standard RL by optimizing policies with respect to multiple, often conflicting, objectives. While
arXiv:2605.31291v1 Announce Type: cross Abstract: Recommender systems may operate under multiple, competing objectives. For example, audience reach, cultural values, public service mandate, and operat
arXiv:2507.12453v5 Announce Type: replace Abstract: In automated machine learning, scientific discovery, and other applications of Bayesian optimization, deciding when to stop evaluating expensive bla
arXiv:2605.30590v1 Announce Type: cross Abstract: Two clinical AI systems can score nearly identically on coverage-based rubrics yet behave radically differently when their patient inputs change: one
arXiv:2501.01926v3 Announce Type: replace-cross Abstract: Large vision-language models (LVLMs) have shown remarkable capabilities in visual-language understanding. Despite their success, LVLMs still s
Current AI is bandaids all the way down; LLMs can’t play nicely with basic tools like databases and knowledge graphs, and you never know what you are going to get. It’s long time to face facts: LLMs h
arXiv:2605.31360v1 Announce Type: cross Abstract: The Artificial Intelligence (AI) life cycle requires a thorough understanding of the underlying data dynamics for robust, safe and cost-effective AI d
arXiv:2509.20941v2 Announce Type: replace Abstract: As surgical AI transitions from pixel-level detection to complex reasoning, Scene Graphs (SGs) offer the structured, relational representations nece
arXiv:2605.31174v1 Announce Type: new Abstract: Object detection in real-world scenarios remains challenging due to diverse image degradations and heterogeneous object distributions, which significant
arXiv:2605.31354v1 Announce Type: new Abstract: Modular visual reasoning systems increasingly rely on shared working memory for multi-step collaboration, yet the failure dynamics of intermediate state
arXiv:2602.00521v2 Announce Type: replace Abstract: While LLM-as-a-Judge is widely used in automated evaluation, existing validation practices primarily operate at the level of observed outputs, offer
arXiv:2605.30808v1 Announce Type: cross Abstract: Preference alignment is a crucial post-training step for large language models (LLMs) to ensure their outputs align with human values. However, post-t
arXiv:2506.11653v3 Announce Type: replace-cross Abstract: Dataset bias often leads deep learning models to exploit spurious correlations instead of task-relevant signals. We introduce the Standard Ant
arXiv:2605.30861v1 Announce Type: new Abstract: Post-training for reasoning models typically combines supervised fine-tuning with reinforcement learning from verifiable rewards, most commonly with GRP
arXiv:2605.30915v1 Announce Type: new Abstract: Real-world images rarely suffer from a single degradation, and the order in which degradations are removed substantially affects the final restoration q
arXiv:2605.31432v1 Announce Type: cross Abstract: Simultaneous speech-to-text translation (SimulST) generates translations while speech is still unfolding, requiring a streaming policy that decides wh
arXiv:2605.31041v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated promising capability in autonomous driving, highlighting the potential of unified multimodal arc
arXiv:2605.31455v1 Announce Type: cross Abstract: Large language models are increasingly deployed in multi-turn interactive settings where users or environments can iteratively provide lightweight fee
arXiv:2509.24319v4 Announce Type: replace-cross Abstract: Large language models can express values in two main ways: (1) intrinsic expression, reflecting the model's inherent values learned during tra
arXiv:2605.31228v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards is an effective route for post-training to strengthen the reasoning capability of large language models
arXiv:2605.30776v1 Announce Type: new Abstract: Offline-to-Online Reinforcement Learning (O2O-RL) leverages an offline, pre-trained policy to minimize costly online interactions. Although data-efficie
arXiv:2605.30484v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have shown promise for robotic manipulation, yet most existing policies operate reactively by directly regressing ac
Elon was right that what OpenAI did in its for profit turn was shameful. But what he is doing with the SpaceX, at the likely expense of many people’s retirement funds, is just as shameful. SpaceX bein
arXiv:2605.30363v1 Announce Type: cross Abstract: Regime shifts in financial markets reorganise the joint dynamics of asset prices and macro variables, breaking any single-regime calibration. They are
arXiv:2605.31250v1 Announce Type: cross Abstract: We propose a unified framework for addressing three key challenges of distribution shift: (1) estimating a model's performance on an unlabeled target
arXiv:2605.31266v1 Announce Type: cross Abstract: The layout-to-image (L2I) task enables fine-grained control over image generation via object categories and spatial layouts. However, existing L2I met
arXiv:2605.30705v1 Announce Type: new Abstract: Geometry-aware generative models and novel view synthesis approaches have shown strong potential in visual fidelity and consistency. In parallel, equiva
arXiv:2605.30617v1 Announce Type: new Abstract: Robust and efficient state estimation is crucial for perception, navigation, and control in robotics. State estimation problems are conveniently modeled
arXiv:2605.31143v1 Announce Type: cross Abstract: Rising household debt and cost-of-living pressures in the United Kingdom have intensified the role of AI-driven financial technologies in mediating cr
arXiv:2602.07285v2 Announce Type: replace Abstract: Binary classification based on predicted probabilities (scores) is a fundamental task in supervised machine learning. While thresholding scores is B