The iWildCam 2021 Competition Dataset

Sara Beery, Arushi Agarwal, Elijah Cole, Vighnesh Birodkar

Introduction

The computer vision community has been making steady progress improving automated systems for species classification and localization in camera trap images over the past decade . Classifications of species seen in a given image or sequence are used by ecologists to generate species richness models , species occurrence models or species distribution models , which describe (stated simply) where in a region or around the world a species might live (or be able to live). However, these types of models do not typically describe the abundance (population size of a given species in an area) or density (how that population is spatially distributed ) of the species. A common method for population estimation is mark-recapture, which requires individual animals to be identified and recognized in future imagery . Though strides are being made in visual re-identification for species with strong biometric markings such as zebras , many species are not visually re-identifiable by humans, making data collection and analysis difficult. To address this, ecological models have been developed that estimate abundance from counts of individuals of a species captured in each camera across short time windows . The iWildCam 2021 competition iWildCam 2021 is hosted on Kaggle: https://www.kaggle.com/c/iwildcam2021-fgvc8 seeks to automate that counting process to enable abundance estimation to scale efficiently to large data collections, and one day to global data repositories such as Wildlife Insights .

Competitors will categorize and count species across short bursts of images in the test data. No count labels have been provided for the training set, in hopes that competitors will develop methods that can learn to count without explicit training labels, as most public camera trap data is not labeled with counts . We provide competitors with species labels along with weakly-supervised detections and instance segmentations to help them to disambiguate individuals. The competition also maintains the multi-modal aspects of the iWildcam 2020 challenge by providing citizen science images for the species of interest, remote sensing imagery for each camera location, and obfuscated geolocation for most cameras.

Data Preparation

The dataset consists of three primary components: (i) camera trap images, (ii) citizen science images, and (iii) multispectral imagery for each camera location. Each component represents one technique for monitoring an ecosystem, each of which has unique strengths and limitations. Species classification and localization performance has been shown to improve by using information beyond the image itself so we hope that participants will find creative and effective uses for this data and see similar improvements in species counting.

The camera trap data (along with expert species annotations) is provided by the Wildlife Conservation Society (WCS) . Camera trap images are taken automatically based on a motion-triggered sensor, so there is no guarantee that the animal will be centered, focused, well-lit, or at an appropriate scale (they can be either very close or very far from the camera, each causing its own problems). See Fig. 2 for examples of these challenges. Empty images pose an additional challenge, as up to 70% of the photos at any given location may be triggered by something other than an animal, such as wind in the trees.

If we wish to build systems that are trained to detect and classify animals and then deployed to new locations without further training, we must measure the ability of machine learning and computer vision to generalize to new environments . This is central to the 2018 , 2019 , 2020 and 2021 iWildCam challenges, all of which split train and test data by camera location, so no images from the test cameras are included in the training set to avoid overfitting to one set of backgrounds .

2 iWildCam 2021 Dataset

The 2021 training set contains 203,314203,314 images from 323323 locations, and the WCS test set contains 60,21460,214 images from 9191 locations. These 414414 locations are spread across 1212 countries in different parts of the world. Each image is associated with a location ID so that images from the same location can be linked. In some cases, WCS biologists placed multiple cameras at the same location. We denote this with a sub-location ID, which communicates that the background and hardware of the camera at these sub-locations is different, but the physical location is the same. As is typical for camera traps, approximately 50% of the total number of images are empty (this varies per location). The iWildCam 2021 dataset is slightly smaller than the iWildCam 2020 dataset. We removed images from iWildCam 2020 that were found to be corrupted, mislabeled, or labeled with ambiguous categories like ‘start’.

There are 206206 species represented in the camera trap images. The class distribution is long-tailed, as shown in Fig. 3. Since we have split the data by location, some classes appear only in the training set. Any images with classes that appeared only in the test set were removed.

Count labels for the test data were collected in collaboration with Centaur Labs . We showed human annotators sequences of images that they could freely scroll through. Each sequence was labeled by between 3 and 30 individual annotators, with additional annotations collected for examples where annotators did not agree. Final counts were determined by majority vote, weighted by annotator performance on an expert-labeled subset. Sequences found to have multiple species were manually annotated by experts.

2.2 Obfuscated GPS Locations

In order to allow competitors to try to use the geographic location of the cameras to improve their classification , we worked with WCS to release obfuscated GPS coordinates for most of the camera trap locations. The precise coordinates of the cameras have been obfuscated randomly to within 1 km for privacy and security reasons, and correspond to the centers of the provided remote sensing imagery. Some of the obfuscated GPS locations were not released at the request of WCS, but we can confirm that all locations without GPS are from the same country.

3 iNaturalist Data

iNaturalist is an online community where citizen scientists post photos of plants and animals and collaboratively identify the species . Similar to iWildCam 2020, we provide a mapping from our classes into the iNaturalist taxonomy.Note that for the purposes of the competition, competitors may only use iNaturalist data from the 2017-2021 iNaturalist competition datasets. We also provide the subsets of the iNaturalist 2017-2019 competition datasets that correspond to species seen in the camera trap data. This curated set provides 13,05113,051 additional images for training, covering 7575 classes.

Though small relative to the camera trap data, the iNaturalist data has some unique characteristics. First, the class distribution is completely different (though it is still long tailed). Second, iNaturalist images are typically higher quality than the corresponding camera trap images, providing valuable examples for hard classes. See for a comparison between iNaturalist images and camera trap images.

4 Remote Sensing Data

In addition to the raw remote sensing data for each camera location outlined in , this year we have provided pre-extracted ImageNet features. We use an ImageNet-pretrained ResNet-50 to extract features from the RGB channels of each multispectral image.

5 Provided Models

Competitors are free to use the Microsoft AI for Earth MegaDetector (a general and robust camera trap detection model )as they see fit. Megadetector V3 detects animal and human classes, while the MegaDetector V4 adds a vehicle class. Any version of the MegaDetector is allowed to be used in this competition. The models can be downloaded on the Microsoft Camera Traps GitHub repository . We provide the top MegaDetector V3 boxes and associated confidences along with our WCS image metadata.

5.2 DeepMAC

Along with MegaDetector box labels, we also provide a method to extract corresponding segmentation masks within each detected box. The segmentations are derived from the DeepMAC model . Although DeepMAC is designed as an instance segmentation model (i.e. detection+segmentation), for this competition we provide an instance of the model which takes boxes as input from the user. Combined with the MegaDetector box labels, or a user-provided detection model, this can be used to extract a per-detection segmentation mask. We provide the DeepMAC masks associated with MegaDetector V3 boxes on Kaggle. Examples of segmentation results paired with MegaDetector V3 boxes can be seen in Fig. 4. The DeepMAC model was originally trained on all of COCO and achieves a detection and mask mAP of 44.5 % and 39.7 % respectively.

Evaluation

Let X∈{0,1,2,…}n×mX\in\{0,1,2,\ldots\}^{n\times m} be a matrix of predictions, so each entry xijx_{ij} is the predicted count for species j∈{1,…,m}j\in\{1,\ldots,m\} in sequence i∈{1,…,n}i\in\{1,\ldots,n\}. Let Y∈{0,1,2,…}n×mY\in\{0,1,2,\ldots\}^{n\times m} be the matrix of corresponding ground truth counts. Submissions will be evaluated using mean columnwise root-mean-squared error (MCRMSE) given by

We selected this metric out of the options provided by Kaggle in order to capture both species identification mistakes and count mistakes as well as to ensure false predictions on empty sequences would contribute to the error. Because many sequences are empty in camera trap data and because many species are rare, the metric tends to be a small number even when the actual errors in counts are large. To convert the metric to something more interpretable, we can un-normalize the metric from MCRMSE to the summed columnwise root summed squared error (SCRSSE) given by

Baseline Results

We built our simple counting baselines from our iWildCam 2020 classification baseline model (see details in ), the iWildCam 2020 winning submission, and the provided MegaDetector V3 results. The results can be seen in Table 1, and the simple baselines are described below.

Max boxes:. We assume that all high-confidence animal boxes (≥0.8\geq 0.8) for an image are correct, and that the species in all boxes match our majority-vote classification prediction for that sequence. We take the maximum number of boxes from any image in the sequence and use that as our count. This will be a lower bound on the actual number of individuals across the sequence since it prevents double counting multiple images of the same individual. Example in Fig 5.

Sum boxes: We assume that all high-confidence animal boxes (≥0.8\geq 0.8) for each image are correct, and that the species in all boxes match our majority-vote classification prediction for that sequence. We take the sum of boxes across the sequence and use that as our count. This will be a upper bound on the actual number of individuals since individuals seen in multiple frames will be double counted.

One per predicted species: We add a count of one for each unique species predicted by our image-level classification model across the sequence. This will be a lower bound on the actual number of individuals across the sequence as it just assumes that one animal was seen per species, regardless of detection results.

All zeros: Just predict zero for all instances. Under our chosen metric this performs surprisingly well. This is for two reasons. First, camera trap data frequently has a small number of animals for any given species. Second, the model is double penalized if the count is correct but the species is incorrect (one penalty for missing the correct species count and one for overpredicting the incorrect species count).

Conclusion

The iWildCam 2021 dataset presents a new challenge for computer vision: counting the number of individuals across low-frame-rate sequences of images . In subsequent years, we plan to extend the iWildCam challenge by adding additional data streams and tasks, such as detection, segmentation, or distance estimation. We hope to use the knowledge we gain throughout these challenges to facilitate the development of systems that can accurately provide real-time species ID and counts in camera trap images at a global scale. Any forward progress made will have a direct impact on the scalability of biodiversity research geographically, temporally, and taxonomically.

Acknowledgements

We would like to thank Dan Morris and Siyu Yang (Microsoft AI for Earth) for their help curating the dataset, providing bounding boxes from the MegaDetector, and hosting the data on Azure. We would like to thank Jonathan Huang and the Visual Dynamics Team at Google Research for providing segmentation labels. We thank the Wildlife Conservation Society for providing the camera trap data and species annotations, and Centaur Labs for working with us to label counts on the test set. We thank Kaggle for supporting the iWildCam competition for the past four years. Thanks also to the FGVC Workshop, Visipedia, and our advisor Pietro Perona for continued support. This work was supported in part by NSF GRFP Grant No. 1745301. The views are those of the authors and do not necessarily reflect the views of the NSF.

References