Racing to the Starting Line: Measuring Opportunity and Improvement in Cross Country

arXiv:2509.10600 · cs.CY, cs.AI, cs.LG · Submitted 2025-09-12 · Read on arXiv

cs.CY, cs.AI, cs.LG

Submitted: 2025-09-12

Updated: 2026-09-24

Code: https://github.com/National-Running-Club-Database/nrcd_xc_paper

License: http://creativecommons.org/licenses/by/4.0/

The gist: Collegiate cross country teams often build their season schedules on intuition rather than evidence, partly because large-scale performance datasets were not publicly accessible prior to the National

Terminology

Abstract

Collegiate cross country teams often build their season schedules on intuition rather than evidence, partly because large-scale performance datasets were not publicly accessible prior to the National Running Club Database (NRCD). We analyze the comprehensive-era Cross Country subset of NRCD, 23,360 results from 7,056 athletes (2023-2025; >99% course/weather coverage). Under leakage control and temporal validation, race-result features do not support out-of-year forecasting of individual improvement (best men's R squared = 0.044; women's-0.018), capturing only a small fraction of the outcome's reliability ceiling (approximately 0.23-0.28). Against this null, team race frequency associates with nationals placement (pooled RR = 2.09; GEE OR = 2.56/SD). Program-wide opportunity (roster depth; Effective Racing Opportunity) outranks a single workhorse's max race count cross-sectionally, but overall team depth for race count is controlled. Converted Only times (not adjusted for weather and elevation) overstate mean first-to-last gains by 15-21 s relative to Standardized. These results challenge coaching practices that treat schedule design as purely anecdotal and show how NRCD enables evidence-based decision-making in collegiate cross country.

Sources

Related papers