Model Releases
ESF-Bench: Benchmarking Challenging Slot-Filling Scenarios for Real-World Enterprise Applications
arXiv:2607.23326v1 Announce Type: new Abstract: The rapid rise of large language models (LLMs) has driven transformative adoption across enterprises. However, deploying these models in real-world sett
arXiv:2607.23326v1 Announce Type: new Abstract: The rapid rise of large language models (LLMs) has driven transformative adoption across enterprises. However, deploying these models in real-world settings presents unique challenges due to complex system constraints and unexpected user behaviors. Among these applications, slot filling is essential for converting unstructured input into structured, actionable data. In this work, we introduce ESF-Bench, a challenging Enterprise Slot Filling benchmark consisting of 810 multi-turn samples and 6530 slots over 8 unique domains. Curated using a taxonomy of the 57 most challenging slot-filling scenarios observed during real-world enterprise deployments, ESF-Bench exposes notable limitations in current state-of-the-art LLMs, with GPT-OSS-120b low successfully extracting slots for only 20.7% of benchmark samples. To support continued research in this area, we publicly release the benchmark dataset, taxonomy, and accompanying evaluation code on GitHub.
Related
- MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment Simulation
- SupChain-Bench: Benchmarking Large Language Models for Real-World Supply Chain Management
- Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities
- SQLyzr: A Comprehensive Benchmark and Evaluation Platform for Text-to-SQL
Source: arXiv cs.AI | 2026-07-28