What do Reward Models Memorize?
DGX agentarXiv:2607.24484v1 Announce Type: cross Abstract: This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference dataset
Knowledge catalogue
arXiv:2607.24484v1 Announce Type: cross Abstract: This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference dataset
arXiv:2607.22781v1 Announce Type: cross Abstract: High task performance does not show whether a model retains prediction-relevant structural information in its internal representation. Temporal graph