Agents
Opponent Aware Reinforcement Learning
arXiv:1908.08773v3 Announce Type: replace Abstract: In certain reinforcement learning (RL) scenarios there are adversaries trying to interfere with the underlying reward process for their own benefit.
arXiv:1908.08773v3 Announce Type: replace Abstract: In certain reinforcement learning (RL) scenarios there are adversaries trying to interfere with the underlying reward process for their own benefit. We introduce Threatened Markov Decision Processes (TMDPs) as a framework to support an agent against potential opponents in an RL context as well as schemes resulting in novel learning approaches to deal with TMDPs. After introducing our framework and deriving theoretical results, empirical evidence is given via extensive experiments, showing the importance for an RL agent of acknowledging adversarial awareness.
Related
- Minimax Optimal Strategy for Delayed Observations in Online Reinforcement Learning
- Generalized Intention Modeling in Multi-Agent Reinforcement Learning
Source: arXiv cs.LG | 2026-08-26