PreAct-Bench: Benchmarking Predictive Monitoring in LLMs
DGX agentarXiv:2606.09890v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents capable of executing multi-step action trajectories toward a given objecti