Research
ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling
arXiv:2510.14703v2 Announce Type: replace Abstract: Large language models (LLMs) excel at function calling, but inference scaling has been explored mainly for unstructured generation. We propose an in
arXiv:2510.14703v2 Announce Type: replace Abstract: Large language models (LLMs) excel at function calling, but inference scaling has been explored mainly for unstructured generation. We propose an inference-scaling framework for structured outputs that combines fine-grained beam search with extbf{ToolPRM}, a process reward model scoring each intra-call decision (function name and argument filling). We build the first fine-grained intra-call supervision dataset via function masking, rollout collection, and step-level annotation. ToolPRM outperforms outcome and coarse-grained reward models in predictive accuracy and yields consistent test-time gains on multiple function-calling benchmarks. We further show that structured generation follows ``extbf{explore more but retain less}'', since early JSON errors are unrecoverable.
Source: arXiv cs.AI | 2026-04-30