RewardFlow: Topology-Aware Reward Propagation on State Graphs for Agentic RL with Large Language Models
DGX agentarXiv:2603.18859v2 Announce Type: replace Abstract: Reinforcement learning (RL) shows promise for enhancing LLM agentic reasoning, yet sparse terminal rewards hinder fine-grained optimization. Process