A model trained on camera-trap images used visual examples of body regions, including male horns, to explain its classifications.
The study used 1,833 cropped images from unconstrained camera-trap footage of endangered mountain gazelles. The researchers built on an existing deep-learning system by allowing it to learn visual examples of different sizes and shapes, rather than relying only on fixed-size image patches.
The system’s explanations often focused on the animals’ central body, legs and head. Male-related examples also frequently included horns, a known difference between the sexes, while differences in body proportions helped distinguish males from females. The researchers report a global accuracy and F1 score of 75%.
Evidence and caveats
This is a model-development study based on 1,833 cropped bounding boxes from unconstrained camera-trap images. The reported performance was 75% for both global accuracy and F1 score, but the abstract does not describe an independent field test, a comparison with other models, or how performance varied across sites or conditions. Camera-trap images can differ in pose, lighting and occlusion, so the findings do not establish how well the system will generalize to all wild mountain gazelle populations. The dataset is planned for public release upon publication.