Agents
Hyperparameters fine tuning for MARL comparative study [D]
hello everyone. I'm training PPO variants on different multi-agent tasks from the VMAS library (Independent PPO / Graph PPO and such, see HetGPPO by Bettini et al.). I noticed that for every architect
hello everyone. I'm training PPO variants on different multi-agent tasks from the VMAS library (Independent PPO / Graph PPO and such, see HetGPPO by Bettini et al.). I noticed that for every architecture/scenario couple, the optimal hyperparameters sometimes tend to vary (learning rate, entropy coefficient, KL coefficient, SGD batch size, etc). do I need - methodologically speaking - to unify the hyperparameters of all models in order to make a fair and correct comparison of architectures later on? note: sometimes unifying these HP leads to some non converging models. note 2 : my objective is to test these models' robustness under adversarial attack in test-time (frozen models). thank you in advance. submitted by /u/ham_bam0 [link] [comments]
Related
- Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic
- MARFT: Multi-Agent Reinforcement Fine-Tuning
- Zero Shot Coordination for Sparse Reward Tasks with Diverse Reward Shapings
Source: r/MachineLearning | 2026-08-24