Tutorials

Faster Results from a Smarter Schedule: Reframing Collegiate Cross Country through Analysis of the National Running Club Database

arXiv:2509.10600v5 Announce Type: replace-cross Abstract: Collegiate cross country teams often build their season schedules on intuition rather than evidence, partly because large-scale performance da

DGX agentpaper
tutorialsarxiv-cs-ai

arXiv:2509.10600v5 Announce Type: replace-cross Abstract: Collegiate cross country teams often build their season schedules on intuition rather than evidence, partly because large-scale performance datasets were not publicly accessible prior to the National Running Club Database (NRCD). We analyze the comprehensive-era Cross Country subset of NRCD, 23,355 results from 7,083 athletes (2023-2025; >97% course/weather coverage). Under leakage control and temporal validation, race-result features do not support out-of-year forecasting of individual improvement (best men's R^2=0.043; women's -0.029), capturing only a small fraction of the outcome's reliability ceiling (~0.23-0.28). Against this null, team race frequency associates with nationals placement (pooled RR =2.09; GEE OR =2.56/SD). Program-wide opportunity (roster depth; Effective Racing Opportunity) outranks a single workhorse's max race count cross-sectionally, but overall team depth for race count is controlled. Converted Only' times (not adjusted for weather and elevation) overstate mean first-to-last gains by 15-21 seconds relative to Standardized'. These results challenge coaching practices that treat schedule design as purely anecdotal and show how NRCD enables evidence-based decision-making in collegiate cross country.

Source: arXiv cs.AI | 2026-08-12

Loading related sources…