ProDVI: Programmatic Dynamics Priors for Value Network Initialization
DGX agentarXiv:2608.06015v1 Announce Type: cross Abstract: Deep Reinforcement Learning (RL) is notoriously sample inefficient. One contributing factor is that RL agents are typically initialized from scratch,