Revisiting TD Target Aggregation under Uncertainty in Q-Learning
DGX agentarXiv:2608.03069v1 Announce Type: new Abstract: Deep Q-Networks (DQNs) learn value functions through bootstrapped temporal-difference updates, where future returns are approximated using a greedy maxi