SoccerSynth-detection: a synthetic dataset for soccer player detection
Abstract
Abstract Soccer player detection is a fundamental task in soccer video analysis, supporting downstream applications such as event analysis and tactical reconstruction. However, annotated soccer player data from broadcast videos are difficult to obtain and use at scale because of copyright restrictions and high annotation costs. In contrast, synthetic data can complement real training data through scalable generation and automatic annotation, yet dedicated synthetic resources for soccer player detection remain limited. To address this gap, we build a data-generation pipeline based on the soccer simulator developed in our previous work and introduce SoccerSynth-Detection, which, to the best of our knowledge, is the first publicly released synthetic dataset providing bounding-box annotations specifically for soccer-player detection from a single main broadcast-camera viewpoint. We evaluate the proposed dataset using multiple detector architectures and two public soccer benchmarks through the Transfer Experiment, the Pretraining and Finetuning Experiment, and the Mixed Training Experiment. In the Multiple Seed Evaluation Experiment on the official benchmark partitions, Mixed Training achieves the highest mean AP50:95 across all six detector and benchmark combinations. Under cross-validation, Mixed Training yields higher mean AP50:95 than Real Only in five of the six detector–benchmark combinations and lower mean AP50:95 in one, indicating that the effect remains partition-dependent. We further test whether difficult samples generated by augmenting real images can substitute for corresponding simulator-generated samples. Both tested augmentation-based replacements yield lower mean AP50:95 than the Original Mix in five of the six detector and benchmark combinations, indicating that the tested blur and overlap augmentations did not reproduce the performance of the Original Mix in most evaluated settings. The dataset is available at https://doi.org/10.5281/zenodo.20839314 .
// Source
Authors: Calvin Yeung, Rikuhei Umemoto, Keisuke Fujii
Institutions: Nagoya University, Japan Science and Technology Agency, RIKEN Center for Advanced Intelligence Project