SocietyBench: Forecasting Counterfactual Social-World Evolution
DGX agentarXiv:2608.04009v1 Announce Type: new Abstract: Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task -- fix a bug, drive a b