Research
Hierarchical Behaviour Spaces
arXiv:2604.24558v1 Announce Type: new Abstract: Recent work in hierarchical reinforcement learning has shown success in scaling to billions of timesteps when learning over a set of predefined option r
arXiv:2604.24558v1 Announce Type: new Abstract: Recent work in hierarchical reinforcement learning has shown success in scaling to billions of timesteps when learning over a set of predefined option reward functions. We show that, instead of using a single reward function per option, the reward functions can be effectively used to induce a space of behaviours, by letting the controller specify linear combinations over reward functions, allowing a more expressive set of policies to be represented. We call this method Hierarchical Behaviour Spaces (HBS). We evaluate HBS on the NetHack Learning Environment, demonstrating strong performance. We conduct a series of experiments and determine that, perhaps going against conventional wisdom, the benefits of hierarchy in our method come from increased exploration rather than long term reasoning.
Related
- Synthetic POMDPs to Challenge Memory-Augmented RL: Memory Demand Structure Modeling
- Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers
- Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning
- A Systematic Review and Taxonomy of Reinforcement Learning-Model Predictive Control Integration for Linear Systems
- Leveraging Human Feedback for Semantically-Relevant Skill Discovery
Source: arXiv cs.AI | 2026-04-28