LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning
DGX agentarXiv:2604.14140v1 Announce Type: new Abstract: As language models are increasingly deployed for complex autonomous tasks, their ability to reason accurately over longer horizons becomes critical. An