Fewer Tokens, Smaller Cache: Reward-Coordinated Efficient Reasoning
DGX agentarXiv:2608.04771v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) excel on complex tasks through long chain-of-thought (CoT) reasoning, but their lengthy intermediate steps cause severe ov