Safety
Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response
arXiv:2607.27508v1 Announce Type: new Abstract: Assistance games formalize human-robot collaboration under asymmetric information: the human knows the goal, while the robot must infer it from observat
arXiv:2607.27508v1 Announce Type: new Abstract: Assistance games formalize human-robot collaboration under asymmetric information: the human knows the goal, while the robot must infer it from observation and interaction in order to assist effectively. In general, computing optimal assistance game strategies online is intractable, since exact solutions require planning in a POMDP. We identify a class of assistance games in which pragmatic-pedagogic reasoning resolves goal uncertainty in a single time step, rendering the full-horizon game exactly solvable by a tractable best-response procedure. Within this class, we show that mainstream inverse optimal control exhibits an inference ceiling that hinders alignment, while pragmatic-pedagogic reasoning overcomes this barrier by immediately disambiguating goals through actions that look equivalent under task execution alone. Finally, we validate our theoretical results and proposed method on a simple collaborative block-building example.
Related
- Learning Robot Visual Navigation in Crowds via Intention-Aware Scene Representations
- What Is My Robot Thinking? Design Considerations for Transparent and Trustworthy Shared Autonomy
- Dialogue based Interactive Explanations for Safety Decisions in Human Robot Collaboration
- OHP-RL: Online Human Preference as Guidance in Reinforcement Learning for Robot Manipulation
Source: arXiv cs.RO | 2026-07-31