TRCA: Transition-wise Rubric Credit Assignment for Long-horizon LLM Agents
arXiv:2608.16156v1 Announce Type: new Abstract: Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment across multi