E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios
arXiv:2607.23722v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as agents that interact with stateful environments over multiple steps: gathering hidden informat