LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
arXiv:2604.13072v1 Announce Type: new Abstract: LLM-based agents are increasingly expected to handle real-world assistant tasks, yet existing benchmarks typically evaluate them under isolated sources