DiPRL: Learning Discrete Programmatic Policies via Architecture Entropy Regularization
DGX agentarXiv:2605.18508v1 Announce Type: cross Abstract: Programmatic reinforcement learning (PRL) offers an interpretable alternative to deep reinforcement learning by representing policies as human-readabl