I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications
arXiv:2606.00750v1 Announce Type: new Abstract: Recent advances in visual language models have enabled autonomous agents for complex reasoning, tool use, and document understanding. However, existing