Model Releases
arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table Generation
arXiv:2504.10284v5 Announce Type: replace Abstract: Literature review tables are essential for summarizing and comparing collections of scientific papers. In this paper, we study the automatic generat
arXiv:2504.10284v5 Announce Type: replace Abstract: Literature review tables are essential for summarizing and comparing collections of scientific papers. In this paper, we study the automatic generation of such tables from a pool of papers to satisfy a user's information need. Building on recent work (Newman et al., 2024), we move beyond oracle settings by (i) simulating well-specified yet schema-agnostic user demands that avoid leaking gold column names or values, (ii) explicitly modeling retrieval noise via semantically related but out-of-scope distractor papers verified by human annotators, and (iii) introducing a lightweight, annotation-free, utilization-oriented evaluation that decomposes utility into schema coverage, unary cell fidelity, and pairwise relational consistency, while measuring paper selection through a two-way QA procedure (gold to system and system to gold) with recall, precision, and F1. To support reproducible evaluation, we introduce arXiv2Table, a benchmark of 1,957 tables referencing 7,158 papers, with human-verified distractors and rewritten, schema-agnostic user demands. We also develop an iterative, batch-based generation method that co-refines paper filtering and schema over multiple rounds. We validate the evaluation protocol with human audits and cross-evaluator checks. Extensive experiments show that our method consistently improves over strong baselines, while absolute scores remain modest, underscoring the task's difficulty. Our data and code is available at https://github.com/JHU-CLSP/arXiv2Table.
Related
- PeReGrINE: Evaluating Personalized Review Fidelity with User Item Graph Context
- Sell More, Play Less: Benchmarking LLM Realistic Selling Skill
- AutoReproduce: Automatic AI Experiment Reproduction with Paper Lineage
- E2Edev: Benchmarking Large Language Models in End-to-End Software Development Task
- Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
Source: arXiv cs.CL | 2026-04-10