Common Limitations of Image Processing Metrics: A Picture Story

Annika Reinke, Minu D. Tizabi, Carole H. Sudre, Matthias Eisenmann, Tim Rädsch, Michael Baumgartner, Laura Acion, Michela Antonelli, Tal Arbel, Spyridon Bakas, Peter Bankhead, Arriel Benis, Matthew Blaschko, Florian Buettner, M. Jorge Cardoso, Jianxu Chen, Veronika Cheplygina, Evangelia Christodoulou, Beth Cimini, Gary S. Collins, Sandy Engelhardt, Keyvan Farahani, Luciana Ferrer, Adrian Galdran, Bram van Ginneken, Ben Glocker, Patrick Godau, Robert Haase, Fred Hamprecht, Daniel A. Hashimoto, Doreen Heckmann-Nötzel, Peter Hirsch, Michael M. Hoffman, Merel Huisman, Fabian Isensee, Pierre Jannin, Charles E. Kahn, Dagmar Kainmueller, Bernhard Kainz, Alexandros Karargyris, Alan Karthikesalingam, A. Emre Kavur, Hannes Kenngott, Jens Kleesiek, Andreas Kleppe, Sven Kohler, Florian Kofler, Annette Kopp-Schneider, Thijs Kooi, Michal Kozubek, Anna Kreshuk, Tahsin Kurc, Bennett A. Landman, Geert Litjens, Amin Madani, Klaus Maier-Hein, Anne L. Martel, Peter Mattson, Erik Meijering, Bjoern Menze, David Moher, Karel G. M. Moons, Henning Müller, Brennan Nichyporuk, Felix Nickel, M. Alican Noyan, Jens Petersen, Gorkem Polat, Susanne M. Rafelski, Nasir Rajpoot, Mauricio Reyes, Nicola Rieke, Michael Riegler, Hassan Rivaz, Julio Saez-Rodriguez, Clara I. Sánchez, Julien Schroeter, Anindo Saha, M. Alper Selver, Lalith Sharan, Shravya Shetty, Maarten van Smeden, Bram Stieltjes, Ronald M. Summers, Abdel A. Taha, Aleksei Tiulpin, Sotirios A. Tsaftaris, Ben Van Calster, Gaël Varoquaux, Manuel Wiesenfarth, Ziv R. Yaniv, Paul Jäger, Lena Maier-Hein

Purpose

Validation of biological and medical image analysis algorithms is of the utmost importance for making scientific progress and for translating methodological research into practice. Validation metrics not to be confused with distance metrics in the strict mathematical sense, the measures according to which performance of algorithms is quantified, constitute a core component of validation design. While metrics can measure various quantities of interest, including speed, memory consumption or carbon footprint, most metrics applied today are reference-based metrics, which have the purpose of measuring the agreement of an algorithm prediction with a given reference. The reference, in turn, serves as an approximation of the (typically unknown) ground truth.

Knowing the properties of metrics in use and making educated choices is essential for meaningful and reliable validation in image analysis. Although several papers highlight specific strengths and weaknesses of common metrics [Kofler et al. 2021, Gooding et al. 2018, Vaassen et al. 2020, Konukoglu et al. 2012, Margolin et al. 2014], an international survey [Maier-Hein et al. 2018] revealed the choice of inappropriate metrics as one of the core problems related to performance assessment in medical image analysis. Similar problems are present in other imaging domains [Honauer et al. 2015, Correia and Pereira 2006]. Under the umbrella of Helmholtz Imaging (HI) \hyper@normalise\hyper@linkurlhttps://www.helmholtz-imaging.de/https://www.helmholtz-imaging.de/, three international initiatives have now joined forces to address these issues: the Biomedical Image Analysis Challenges (BIAS) initiative \hyper@normalise\hyper@linkurlhttps://www.dkfz.de/en/cami/research/topics/biasInitiative.htmlhttps://www.dkfz.de/en/cami/research/topics/biasInitiative.html, the Medical Image Computing and Computer Assisted Interventions (MICCAI) Society’s special interest group on challenges \hyper@normalise\hyper@linkurlhttps://miccai.org/index.php/special-interest-groups/challenges/https://miccai.org/index.php/special-interest-groups/challenges/, as well as the benchmarking working group of the MONAI framework \hyper@normalise\hyper@linkurlhttps://monai.io/https://monai.io/. A core mission is to provide researchers with guidelines and tools to choose the performance metrics in a problem- and context-aware manner. This dynamically updated document aims to illustrate important pitfalls and drawbacks of metrics commonly applied in the field of image analysis. The current version is based on a Delphi process on metrics conducted with an international consortium of medical image analysis experts. A Delphi process is a multi-stage survey process designed to pool the knowledge of several experts to arrive at a consensus decision [Brown 1968].

The Delphi consortium focused on problems reporting biomedical research that can be phrased as image-level classification, semantic segmentation, instance segmentation or object detection (Figure 1). Essentially, these can all be interpreted as a classification task at different scales and thus share many aspects in terms of validation (Figure 2). For example, an object detection task can be interpreted as an object-/instance-level classification task, while a segmentation task can be interpreted as a pixel-level classification task. We will refer to these four different task types as problem categories. Please note that we will use the term "pixel" even for three-dimensional (or n-dimensional) images for increased readability instead of referring to "pixels/voxels". Most of the examples are shown for two-dimensional images and can be translated to the n-dimensional case.

The manuscript is structured as follows. As a foundation, we first review the most commonly applied metrics for the problem categories addressed in this paper (Sec. 2). Since a common problem in the biomedical image analysis community is the selection of metrics from the wrong problem category, Sec. 3 highlights pitfalls relevant in this context. The following sections then present pitfalls for image-level classification (Sec. 4), image segmentation, including semantic and instance segmentation (Sec. 5), and object detection, including instance segmentation (Sec. 6). Finally, cross-topic pitfalls are highlighted (Sec. 7). An overview of all figures is presented in Table 1.

Fundamentals

The present work focuses on biomedical image analysis problems that can be interpreted as a classification task at image, object or pixel level. The vast majority of metrics for these problem categories is directly or indirectly based on epidemiological principles of TP (TP), FN (FN), FP (FP), TN (TN), i.e. the cardinalities of the so-called confusion matrix, depicted in Figure 2. The TP/FN/FP/TN, from now on referred to as cardinalities. In the case of more than two classes CC we also refer to the entries of the C×CC\times C confusion matrix as cardinalities. For simplicity and clarity in notation, we restrict ourselves to the binary case in most examples. Cardinalities can be computed for image (segment), object or pixel level. They are typically computed by comparing the prediction of the algorithm to a reference annotation. Modern neural network-based approaches typically require a threshold to be set in order to convert the algorithm output comprising predicted class scores (also referred to as continuous class scores) to a confusion matrix, as illustrated in Figure 2. For the purpose of metric recommendation, the available metrics can be broadly classified as follows (see also [Cao et al. 2020]):

Counting metrics operate directly on the confusion matrix and express the metric value as a function of the cardinalities (see Figs. 4, 5 and 7). In the context of segmentation, they have typically been referred to as overlap-based metrics [Taha and Hanbury 2015]. We distinguish multi-class counting metrics, which are defined for an arbitrary number of classes, from per-class counting metrics, which are computed by treating one class as foreground/positive class and all other classes as background. Popular examples for the former include MCC (MCC), Accuracy, and Cohen’s Kappa κ\kappa, while examples for the latter are Sensitivity, Specificity, PPV (PPV), DSC (DSC) and IoU (IoU).

Multi-threshold metrics operate on a dynamic confusion matrix, reflecting the conflicting properties of interest, such as high Sensitivity and high Specificity. Popular examples include the AUROC (AUROC) (see Figure 6) and AP (AP) (see Figure 15).

Distance-based metrics have been designed for semantic and instance segmentation tasks. They operate exclusively on the TP and rely on the explicit definition of object boundaries (see Figs. 8 and 10). Popular examples are the HD (HD) and the NSD (NSD) (see Figure 10).

Depending on the context (e.g. image-level classification vs. semantic segmentation task) and the community (e.g. medical imaging community vs. computer vision community), identical metrics are referred to with different terminology. For example, Sensitivity, TPR (TPR) and Recall refer to the same concept. The same holds true for the DSC and the F1 Score. The most relevant metrics for the problem categories in the scope of this paper are introduced in the following.

Most metrics are recommended to be applied per class (except the multi-class counting metrics), meaning that a potential multi-class problem is converted to multiple binary classification problems, such that each relevant class serves as the positive class once. This results in different confusion matrices depending on which class is used as the positive class.

Image-level classification refers to the process of assigning one or multiple labels, or classes, to an image. If there is only one class of interest (e.g. cancer vs. no cancer), we speak of binary classification, otherwise of categorical classification. Modern algorithms usually output predicted class probabilities (or continuous class scores) between 0 and 1 for every image and class, indicating the probability of the image belonging to a specific class. By introducing a threshold (e.g. 0.5), predictions are considered as positive (e.g. cancer = true) if they are above the threshold or negative if they are below the threshold. Afterwards, predictions are assigned to the cardinalities (e.g. a cancer patient with prediction cancer = true is considered as TP) [Davis and Goadrich 2006]. The most popular classification metrics are counting metrics, operating on a confusion matrix with fixed threshold on the class probabilities, and multi-threshold metrics, as detailed in the following.

The most common binary counting metrics used for image-level classification are presented in Figs. 4 (per-class) and 5 (multi-class). Please note that these metrics are also commonly used in segmentation and object detection tasks. For segmentation tasks, they are often referred to as overlap-based metrics. Each of the presented metrics covers specific properties.

The Sensitivity (also referred to as Recall, TPR or Hit rate) focuses on the actual positives (TP and FN) and represents the fraction of positives that were correctly detected as such. In contrast, PPV (or Precision) divides the TP by the total number of predicted positive cases, thus aiming to represent the probability of a positive prediction corresponding to an actual positive. A value of 1 would imply that all positive predicted cases are actually positives, but it might still be the case that positive cases were missed. Please note that the term Precision has multiple meanings. In the context of computer assisted interventions, for example, it typically refers to the measured variance. Hence, the usage of its synonym PPV may be preferred.

In analogy to the Sensitivity for positives, Specificity (also referred to as Selectivity or TNR) focuses on the negative cases by computing the fraction of negatives that were correctly detected as such. Similarly to the PPV, the NPV (NPV) divides the TN by the total number of predicted negative cases and measures how many of the predicted negative samples were actually negative. Specificity and NPV require the definition of TN cases, which is not always possible. In object detection tasks, for example (see Sec. 2.3), TN are typically ill-defined and not provided. Therefore, these measures can not be computed in those cases.

PPV and NPV provide quantities of direct interest to the medical practitioner, namely the probability of a certain event given the prediction of the classifier. However, unlike Sensitivity and Specificity, they are not only functions of the classifier but also of the study population. Hence, one cannot extrapolate from one population to another, e.g. from a case-control study to the general population without correcting for prevalence. In the case of the prevalence being unequal to 0.5, prevalence correction should be applied (see Figure 35):

As illustrated in Figure 25, reporting of a single metric such as Sensitivity, PPV or Specificity can be highly misleading because, for example, non-informative classifiers can achieve high values on imbalanced classes. The F1 Score (also known as DSC in the context of segmentation), overcomes this issue by representing the harmonic mean of PPV and Sensitivity and therefore penalizing extreme values of either metric [Hicks et al. 2021], while being relatively robust against imbalanced data sets [Sun et al. 2009]. The F1 Score is a specification of the Fβ\beta score, which adds a weighting between PPV and Sensitivity. For higher values of β\beta, Sensitivity is given a higher weight over PPV, which might be desired depending on the application. More specifically, the Fβ\beta Score weights between FP and FN samples. Another metric which combines the metrics above is the LR+ (LR+), which is the ratio of the Sensitivity and 1-Specificity (FPR). This metric is described as the "probability that a positive test would be expected in a patient divided by the probability that a positive test would be expected in a patient without a disease" [Shreffler and Huecker 2020]. It is defined as odds ratio and therefore invariant to the prevalence [Šimundić 2009].

All of the metrics presented so far are bounded between 0 and 1 with 1 representing a perfect value and 0 the worst possible prediction of this metric. However, all of them rely on the definition of the positive class, which may be straightforward in some cases but can be based on a rather arbitrary choice in others. Notably, metric values may be completely different depending on the choice of positive class [Powers 2020].

To overcome the need for selecting one class as the positive class, other metrics have been suggested that can be based on all entries of a multi-class confusion matrix, in which each class is assigned a row and a column of the matrix. The Accuracy is one of the most commonly used metrics and measures the ratio between all correct predictions (TP and TN) and the total number of samples. Accuracy is not robust against imbalanced data sets (see Figure 25), and is therefore often replaced by the more robust BA (BA) that averages the Sensitivity over all classes [Grandini et al. 2020]. For two classes, the metric averages Sensitivity and Specificity. An alternative is Youden’s Index J or BM (BM), similarly summing up Sensitivity per class (or for two classes, summing up Sensitivity and Specificity). Note that Youden’s Index J and BA are closely related and can directly be calculated from the other as J=2BA−1J=2BA-1 and BA=(J+1)/2BA=(J+1)/2. Both BA and Youden’s Index J are invariant to the prevalence.

The MCC, also known as Phi Coefficient, measures the correlation between the actual and predicted class. The metric is bounded between -1 and 1, with high positive values referring to a good prediction which can only be achieved when all cardinalities are good, i.e. with a low number of FP/FN and high TP/TN. Another popular metric is Cohen’s Kappa κ\kappa, which calculates the agreement between the reference and prediction while incorporating information on the agreement by chance. It is therefore a form of chance-corrected Accuracy. Similarly to MCC, it incorporates all values of the confusion matrix and is bounded between -1 and 1. In contrast to MCC, negative values do not indicate anti-correlation, but less agreement than expected by chance. Cohen’s Kappa κ\kappa can be generalized by introducing a weighting scheme for the cardinalities in the Weighted Cohen’s Kappa κ\kappa metric. For those three metrics, a value of 0 refers to a prediction which is not better than random guessing.

All of the presented binary counting metrics can be transferred to the multi-class case [Hossin and Sulaiman 2015, Grandini et al. 2020], where MCC and (Weighted) Cohen’s Kappa κ\kappa have explicit definitions, whereas the others become the implicit result of an aggregation across a rotating one-versus-the-rest binary perspective for each of the classes.

The EC (EC) is a generalization of the probability of error (which, in turn, is 1 - Accuracy) for cases in which errors cannot all be considered to have equally severe consequences. It is defined as the expectation of the cost, where the cost incurred on a certain sample depends on the sample’s class and the decision made for that sample. In practice, the expectation can be estimated as a simple average of the costs over the validation samples. EC describes the weighted sum of error rates. It can also be used to measure discrimination and calibration in one score. A variant of EC normalizes EC by the EC of a naive system.

Many of the presented metrics depend on the prevalence. Figure 3 illustrates how their values change based on the prevalence.

Multi-threshold metrics

The classical counting metrics presented above rely on fixed thresholds to be set on the predicted class probabilities (if available), resulting in them being based on the cardinalities of the confusion matrix. Multi-threshold metrics overcome this limitation by calculating metric scores based on multiple thresholds. For instance, to emphasize how well a prediction distinguishes between the positive and negative class, the AUROC can be utilized. The ROC (ROC) curve plots the FPR, which is equal to 1−Specificity1-\text{\emph{Specificity}}, against the Sensitivity for multiple thresholds of the predicted class probabilities, contrarily to just choosing one fixed threshold. For computation of the ROC curve, the class scores can be ordered in descending order and each score regarded as a potential threshold. For each threshold, the resulting Sensitivity and Specificity are computed, and the resulting tuple is added to the ROC curve as one point (cf. Figure 6); note that the lower the threshold, the higher the Sensitivity but the lower (potentially) the Specificity. This leads to a monotonic increase of the curve. To interpolate between all points, i.e. to approximate the values between the calculated Sensitivity and Specificity tuples, a simple linear interpolation can be employed by drawing a line between each pair of points [Davis and Goadrich 2006]. An optimal classifier would lead to Sensitivity and Specificity of 1 (1-Specificity of 0), therefore corresponding to a single point (0,1)(0,1) on the ROC curve. In contrast, a classifier with no skill level (random guessing) would result in a diagonal line from (0,0)(0,0) to (1,1)(1,1) (dashed line in Figure 6). The area under the ROC curve is referred to as AUROC, also called AUC ROC or simply AUC.

AUROC comes with two advantages: threshold and scale invariance. AUROC measures the quality of the predictions regardless of the threshold, as it is calculated over a number of thresholds. Furthermore, AUROC does not focus on the absolute values of predictions, but rather on how well they are ranked. However, those properties are not always desired. If a specific penalization of FP or FN is desired (cf. Figure 30), AUROC is not the best metric choice as it is invariant to the threshold. If the predicted class probabilities are intended to be well calibrated, the scale invariance feature will prevent from doing so.

Per definition, AUROC measures the complete area under the ROC curve. If only a specific range is of interest, a partial or ranged AUROC can also be computed [Ma et al. 2013]. Similarly, metrics can be assessed at a certain point of the ROC curve, for example the Sensitivity value at a specific score of the Specificity (e.g. 0.9), also referred to as Sensitivity@Specificity. This approach can similarly be used for other curve measures, e.g. the PR (PR) curve, introduced in Sec. 2.3, and is illustrated in Figure 6. Please note that we will use the synonyms Precision instead of PPV and Recall instead of Sensitivity in the case of the PR curve, given the common use of these terms.

Calibration metrics

While most research in biomedical image analysis focuses on the discrimination capabilities of classifiers, a complementary property of relevance is the calibration of predicted class scores (also known as confidence scores). Intuitively speaking, a system is well-calibrated if the predicted class scores (i.e., the output of the model) reflect the true probabilities of the outcome. In practice, this means that calibrated scores match the empirical success rate of associated predictions. For a binary classification task, calibration implies that of all the data samples assigned a predicted score of, for example, 0.80.8 for the positive class, empirically, 80%80\% belong to this class [Maier-Hein et al. 2022].

Calibration is typically assessed by either calculating calibration errors (for example as done for the ECE (ECE) or MCE (MCE) [Naeini et al. 2015]) or by proper scoring rules (also referred to as overall performance measures [Steyerberg et al. 2010]), which measure discrimination and calibration in a single score (for example the BS (BS)). A detailed overview of calibration metrics is given in [Maier-Hein et al. 2022].

2. Semantic Segmentation

Semantic segmentation is commonly defined as the process of partitioning an image into multiple segments/regions. To this end, one or multiple labels are assigned to every pixel such that pixels with the same label share certain characteristics. Semantic segmentation can therefore also be regarded as pixel-level classification. As in image-classification problems, predicted class probabilities are typically calculated for each pixel deciding on the class affiliation based on a threshold over the class scores [Asgari Taghanaki et al. 2021]. In semantic segmentation problems, the pixel-level classification is typically followed by a post-processing step, in which connected components are defined as objects, and object boundaries are created accordingly. Semantic segmentation metrics can roughly be classified into three classes: (1) counting metrics or overlap-based metrics, for measuring the overlap between the reference annotation and the prediction of the algorithm, (2) distance-based metrics, for measuring the distance between object boundaries, and (3) problem-specific metrics, measuring, for example, the volume of objects.

The most frequently used segmentation metrics are counting metrics. In the context of segmentation they are also referred to as overlap metrics, as they essentially measure the overlap between a reference mask and the algorithm prediction. According to a comprehensive analysis of biomedical image analysis challenges [Maier-Hein et al. 2018], the DSC [Dice 1945] is the by far most widely used metric in the field of medical image analysis. As illustrated in Figure 7, it yields a value between 0 (no overlap) and 1 (full overlap). The DSC is identical to the F1 Score and closely related to the IoU, which is identical to the Jaccard Index:

Distance-based metrics

Overlap-based metrics are often complemented by distance-based metrics that operate exclusively on the TP and compute one or several distances between the reference and the prediction. Apart from a few exceptions, distance-based metrics are often boundary-based metrics which focus on assessing the accuracy of object boundaries. According to [Maier-Hein et al. 2018], the HD and its 95% percentile variant ( HD95 (HD95)) [Huttenlocher et al. 1993] are the most commonly used boundary-based metrics. The HD calculates the maximum of all shortest distances for all points from one object boundary to the other, which is why it is also known as the Maximum Symmetric Surface Distance [Yeghiazaryan and Voiculescu 2018]. The HD95 calculates the 95% percentile instead of the maximum, therefore disregarding outliers (see Figure 8). Another popular metric is the ASSD (ASSD), measuring the average of all distances for every point from one object to the other and vice versa [Van Ginneken et al. 2007, Yeghiazaryan and Voiculescu 2018]. However, if one boundary is much larger than the other, it will impact the score much more. This is avoided by the MASD (MASD) [Beneš and Zitová 2015], which treats both structures equally by computing the average distance from structure AA to structure BB and the average distance from structure BB to structure AA and averaging both (see Figure 9). For the HD(95), MASD and ASSD metrics, a value of 0 refers to a perfect prediction (distance of 0 to the reference boundary), while there exists no fixed upper bound. The Boundary IoU (cf. Figure 12) is another option for measuring the boundary quality of a prediction. It measures the overlap between prediction and reference up to a certain width (which is controlled by the width parameter dd) (see Sec. 2.3 for details) and is bounded between 0 and 1.

A major problem related to boundary-based metrics are the error-prone reference annotations (see Figs. 62 and 63). In fact, domain experts often disagree on the definition and annotation of objects and their boundaries [Joskowicz et al. 2019]. While the HD(95) and ASSD are not robust with respect to uncertain reference annotations, the NSD was explicitly designed for this purpose as a hybrid metric between boundary-based and counting-based approaches. Known uncertainties in the reference as well as acceptable deviations of the predicted boundary from the reference are captured by a threshold τ\tau [Nikolov et al. 2021], as shown in Figure 10. Only boundary parts within the border regions defined by τ\tau are counted as TP. The metric is bounded between 0 (no boundary overlap) and 1 (full boundary overlap), so that it can be interpreted similarly to the classical DSC (though restricted to the boundary). Please note that τ\tau is another important hyperparameter which should be chosen wisely, based on inter-rater agreement, for example.

Problem-specific segmentation metrics

While overlap-based metrics and distance-based metrics are the standard metrics used by the general computer vision community, biomedical applications often have special domain-specific requirements. In medical imaging, for example, the actual volume of an object may be of particular interest (for example tumor volume). In this case, volume metrics such as the Absolute or Relative Volume Error and the Symmetric Relative Volume Difference can be computed [Nai et al. 2021]. However, they are less common than overlap metrics, as the location of objects is not considered at all (see Figure 58). If the structure center or center line is of particular interest (e.g. in cells or vessels), connectivity metrics come into play, which measure the agreement of the center line between two objects. This is of special interest if linear or tube-like objects are present in a data set, for example in brain vascular analysis. In these cases, over- or undersegmentation are typically not of special interest, with the focus rather being on connectivity or network topology. For this purpose, the clDice (clDice) [Shit et al. 2021] has been designed. For computation, the reference and prediction skeletons need to be extracted from the binary masks. Multiple approaches have been proposed, as described in [Shit et al. 2021]. As shown in Figure 11, the clDice is then defined as the harmonic mean of Topology Precision and Topology Sensitivity. The Topology Precision measures the ratio of the overlap of the skeleton of the prediction skeleton and the reference mask and the skeleton of the prediction. Similarly to Precision, it measures the FP samples. On the other hand, Topology Sensitivity incorporates FN samples by measuring the ratio of the overlap of the skeleton of the reference and the prediction mask and the reference skeleton.

3. Object Detection

Object detection refers to the detection of one or multiple objects (or: instances) of a particular class (e.g. lesion) in an image [Lin et al. 2014]. The following description assumes single-class problems, but translation to multi-class problems is straightforward, as validation for multiple classes on object level is performed individually per class. Notably, as multiple predictions and reference instances may be present in one image, the predictions need to include localization information, such that a matching between reference and predicted objects can be performed. Important design choices with respect to the validation of object detection methods include:

How to represent an object? Representation is typically composed of location information and a class affiliation. The former may take the form of a bounding box (i.e. a list of coordinates), a pixel mask, or the object’s center point. Additionally, modern algorithms typically assign a confidence value to each object, representing the probability of a prediction corresponding to an actual object of the respective class. Note that a confusion matrix is later computed for a fixed threshold on the predicted class probabilities. Please note that we will use the term confidence scores analogously to predicted class probabilities in the context of object detection and instance segmentation.

How to decide whether a reference instance was correctly detected? This step is achieved by applying the localization criterion. This may, for example, be based on comparing the object centers of the reference and prediction or computing their overlap (Figs. 68 and 70).

How to resolve assignment ambiguities? The above step might lead to ambiguous matchings, such as two predictions being assigned to the same reference object. Several strategies exist for resolving such cases.

The following sections provide details on (1) applying the localization criterion, (2) applying the assignment strategy and (3) computing the actual performance metrics.

As one image may contain multiple objects or no object at all, the localization criterion or hit criterion measures the (spatial) similarity between a prediction (represented by a bounding box, pixel mask, center point or similar) and a reference object. It defines whether the prediction hit/detected (TP) or missed (FP) the reference. Any reference object not detected by the algorithm is defined as FN. Please note that TN are not defined for object detection tasks, which has several implications on the applicable metrics, as detailed below.

There are multiple ways to define the localization or hit criterion (see Figs. 7, 12 and 68). Popular center-based localization criteria are (a) the center-cover criterion, for which the reference object is considered hit if the center of the reference object is inside the predicted detection, (b) the distance-based hit criterion, which considers a TP if the distance dd between the center of the reference and the detected object is smaller than a certain threshold τ\tau and (c) the center-hit criterion, which holds true if the center of the predicted object is inside the reference bounding box or mask.

The most commonly used overlap-based hit criterion is determined by computing the IoU [Jaccard 1912] (cf. Figure 7b). The prediction is considered as TP if the overlap is larger than a certain threshold (e.g. 0.3 or 0.5) and as FP otherwise. If bounding boxes are considered, the IoU is computed between the reference and predicted bounding boxes (Box IoU). For more fine-grained annotations in form of a pixel mask, the IoU may be computed for the complete mask (Mask IoU). The Mask IoU is less sensitive to structure boundary quality in larger objects (cf. Sec. 5). This is due to the fact that boundary pixels will increase linearly while pixels inside the structure will increase quadratically with an increase in structure size. The Boundary IoU measures the IoU of two structures for mask pixels within a certain distance dd from the structure boundaries [Cheng et al. 2021], as illustrated in Figure 12.

The localization criterion should be carefully chosen according to the underlying motivation and research question and depending on the available coarseness of annotations. However, it should be noted that annotations of a lower resolution will result in an information loss, as illustrated in Figure 13. For example, the Box IoU is sometimes used although pixel-mask annotations are available because algorithms are expected to output rough localization in the shape of boxes. Such a simplification might cause problems if structures are not well-approximated by a box shape, or if structures can overlap causing multi-component masks (cf. Sec. 6, Figure 73). Lastly, it should be noted that the decision for a cutoff value on the localization criterion leads to instabilities in the validation (e.g. see Figure 70). For this reason, it is common practice in the computer vision community to average metrics over multiple cutoff values (default for IoU criteria: 0.50:0.05:0.95 [Lin et al. 2014]). Generally speaking, the cutoff values should be chosen according to the driving biomedical question. For example, if particular interest lies on the exact outlines, higher thresholds should be chosen. On the other hand, for noisy reference standards, a low cutoff value is preferable.

Assignment strategy

The localization criterion alone is not sufficient to extract the final confusion matrix based on a fixed threshold for the predicted class probabilities (confidence scores), as ambiguities can occur. For example, two predictions may have been assigned to the same reference object in the localization step, or vice versa. These ambiguities need to be resolved in a further assignment step, as exemplarily shown in Figure 14.

This assignment and thus the resolving of potential assignment ambiguities can be done via different strategies. The most common strategy in the computer vision community is the Greedy by Score strategy [Everingham et al. 2015]. All predictions in an image are ranked by their predicted class probability and iteratively (starting with the highest probability) assigned to the reference object with the highest localization criterion for this prediction. The selected reference object is subsequently removed from the process since it can not be matched to any other prediction (unless double assignments are allowed). The Hungarian Matching [Kuhn 1955] is associated with a cost function, usually depending on the localization criterion, which is minimized to find the optimal assignment of predictions and reference. In the biomedical domain, more sophisticated matching strategies are often avoided by setting the localization criterion threshold to IoU > 0.5 and only allowing non-overlapping object predictions (which inherently avoids matching conflicts). In the case of a high ratio of touching reference objects and common non-split errors, meaning that one prediction overlaps with multiple reference objects, the Intersection over Reference (IoR) [Matula et al. 2015] might be considered as an alternative to IoU [Matula et al. 2015] (see Figure 72).

Metric computation

Similar to image-level classification and semantic segmentation algorithms, object detection algorithms are commonly assessed with counting metrics, assuming a fixed confusion matrix, (cf. Figs. 4 and 5). However, one of the most popular object detection metrics is the multi-threshold metric AP (AP) [Lin et al. 2014], which is the area under the PR (PR) curve for a certain interpolation scheme. The PR curve is computed similarly to the ROC curve by scanning over confidence thresholds and computing the Precision (PPV) and Recall (Sensitivity) for every threshold (cf. Figure 6). Note in this context that the popular ROC curve is not applicable in object detection tasks because TN are not available. Also, while the ROC curve is monotonically rising, this behavior may not be expected from the PR curve, which typically features a zigzag shape, as illustrated in Figure 15. Specifically "as the level of Recall varies, the Precision does not necessarily change linearly due to the fact that FP replaces FN in the denominator of the Precision metric." [Davis and Goadrich 2006]. A linear interpolation would therefore be overly optimistic, which is why more complex interpolation is needed, as detailed in [Davis and Goadrich 2006].

The area under the PR curve is typically calculated as the AP implying a conservative simplification of curve interpolation,

with RiR_{i} and PiP_{i} denoting the Recall and Precision at the iith threshold https://scikit-learn.org/stable/modules/generated/sklearn.metrics.average_precision_score.html (cf. dashed gray line in Figure 15). For the PR curve, an optimal model would lead to Recall (Sensitivity) and Precision (PPV) of 1, therefore being the point (1,1)(1,1) on the PR curve. Conversely, a model with no skill level (random guessing) would result in a horizontal line with a precision proportional to the portion of positive samples, i.e. the prevalence (dashed line in Figure 6). For computation of the metric (for a given class), the predictions are sorted in descending order of the confidence for each prediction (for that class). For each possible confidence threshold, the cardinalities are computed and the resulting tuple of Recall (Sensitivity) and Precision is added to the curve (cf. Figure 15). In the case of multiple classes, the AP is calculated per class and averaged to obtain the mAP (mAP) score.

In contrast to drawing the PR curve and computing the AP, FROC (FROC) curve is often favoured in the clinical context due to its easier interpretability. It operates at object level and plots the average number of FPPI (in contrast to the FPR) against the Sensitivity [Chakraborty and Zhai 2019, Van Ginneken et al. 2010] (cf. Figure 16). The FROC Score, measured as the area under the FROC curve, however, is not bounded between 0 and 1 and the employed FPPI scores vary across studies, such that there exists no standardized definition of an area under the respective curve. Overall, the decision between the two metrics often boils down to a decision between a standardized and technical validation versus an interpretable and application-focused validation.

Counting cardinalities at different scales

In image-level classification problems, validation is naturally performed on the entire data set, while segmentation typically relies on computing metrics for each image and then aggregating metric values. This latter approach is not applicable in object detection in a straightforward manner because of the relatively small number of samples per image (typically a few objects rather than thousands of pixels). Figure 17 illustrates the per-image and the per-data set validation of objects. In the per-image aggregation approach, special care needs to be taken in the case of an empty reference or prediction, as detailed in Sec. 6 (Figure 76).

A few comments on how to compute evaluation metrics while aggregating matched objects per image:

Counting metrics: Computing per-class metrics such as Sensitivity, PPV, or F1 Score (see Figure 4) works intuitively (as depicted in Figure 17) by computing a score per image and averaging scores over the data set while considering the NaN (NaN) conventions shown in Figure 76.

Multi-threshold metrics: While scanning over all thresholds in the data set (analogously to per data set evaluation), for every threshold the Precision (PPV) and Recall (Sensitivity) scores are computed per image/patient (while considering the NaN conventions in Figure 76) and averaged over the data set. The averaged Precision and Recall score pairs per threshold form the PR curve. The AP (see Figure 15) can be computed for this curve analogously to the per-data set evaluation. The FROC curve computation works analogously while replacing Precision per image with FPPI.

Single working point evaluation: A single working point (e.g. PPV@Sensitivity) can be retrieved based on the curves resulting from multi-threshold evaluation described above.

4. Instance Segmentation

In contrast to semantic segmentation, instance segmentation problems distinguish different instances of the same class (e.g. different lesions). Similarly to object detection problems, the task is to detect individual instances of the same class, but detection performance is measured by pixel-level correspondences (as in semantic segmentation problems). Optionally, instances can be applied to one of multiple classes. Validation metrics in instance segmentation problems often combine common detection metrics with segmentation metrics applied per instance.

If detection and segmentation performance should be assessed simultaneously in a single score, the PQ (PQ) metric can be utilized [Kirillov et al. 2019]. As shown in Figure 18, the segmentation quality is assessed by averaging the IoU scores for all TP instances. While this part alone would neglect FP and FN predictions, it is multiplied with the detection quality, which is equal to the F1 score. For perfect segmentation results, i.e. an average IoU of 1, the PQ would simply equal the F1 Score. In other words, while segmentation quality in the F1 Score is merely integrated via a hard cutoff during localization (e.g. IoU > 0.5) affecting the counts of TP versus FP and FN, PQ allows for a continuous measurement of segmentation quality by 1) employing a lose prior localization criterion (e.g. IoU > 0, i.e. any prediction overlapping with a reference instance is treated as TP) and 2) replacing the simple counting of TP by adding up the continuous IoU scores per TP. PQ was introduced for panoptic segmentation problems that are a combination of semantic and instance segmentation [Kirillov et al. 2019]. It is computed per class (including background classes) and can therefore be directly used for instance segmentation problems by only assessing the foreground instances. If needed, the background quality can also be assessed.

PQ covers segmentation and detection quality in a single score. However, this can be misleading, as shown in Fig. 19. Prediction 1 with a perfect segmentation and poor detection achieves a PQ score similar to that of Prediction 2, which detects all objects correctly (without any FP or FN), but only provides moderate segmentation results. If segmentation and detection quality should be assessed individually, two separate metrics - one for segmentation quality and one for detection quality - should be preferred.

It should be noted that instance segmentation problems are often phrased as semantic segmentation problems with an additional post-processing step, such as connected component analysis [Rosenfeld and Pfaltz 1966]. In practice, predicted class probabilities, yielded by modern segmentation algorithms, are often discarded in the post-processing step and are thus not available for subsequent validation. Figure 20 illustrates how to overcome this potential problem.

Pitfalls due to category-metric mismatch

Performance metrics are typically expected to reflect a domain-specific validation goal (e.g. clinical goal). Previous research, however, suggests, that this is often not the case [Saha et al. 2021]. Before choosing validation metrics, the correct problem category needs to be defined. In the following, we will describe pitfalls related to metrics not being applied to the appropriate problem category.

A common problem is that segmentation metrics, such as the DSC, are applied to object detection tasks [Carass et al. 2020, Jäger 2020], as illustrated in Figure 21. From a clinical perspective, for example, the algorithm producing Prediction 2 and covering all three structures of interest (e.g. tumors) would be clinically much more valuable compared to the one producing a highly accurate segmentation for one structure but missing the other two in Prediction 1. This is not reflected in the metric values, which are substantially higher for Prediction 1. In general, the DSC is strongly biased against single objects, therefore not appropriate for the detection of multiple structures [Yeghiazaryan and Voiculescu 2018, Kirillov et al. 2019].

Mismatch semantic ↔\leftrightarrow instance segmentation

In segmentation problems, the driving research question should decide whether semantic or instance segmentation should be chosen for validation. This is particularly relevant when multiple objects within one image overlap or touch, as often occurring in cell images. For semantic segmentation problems, overlapping or touching objects may end up merged into a single object without clear boundaries or distinction between the single objects. Instance segmentation problems, on the other hand, ensure that the borders of touching or overlapping structures can be accurately assigned and that objects can be differentiated. If instance segmentation is preferred, the labels need to be chosen accordingly. An example is shown in Figure 22: The desired annotation consists of two different instances, but only semantic labels are available (middle). A prediction will only be as accurate as the reference, hence detecting only one instance but yielding a perfect metric score although the desired task is not solved.

Mismatch image-level classification ↔\leftrightarrow object detection

Tasks that should be validated at image level are sometimes erroneously approached with object detection models instead of image-level classification models [Jaeger et al. 2020]. Object detection models are designed to handle different objects in an image rather than the complete image and will naturally introduce problems in a validation setting on image level. Object detection tasks are dependent on choosing a proper localization criterion, which is not needed for an image-level classification problem. For example, a ROC curve, typically used for assessing the performance of image-level classification algorithms, does not consider the localization step needed in object detection tasks, as it was designed to validate at image rather than object level. It therefore does not take into account whether a detected object is at the correct location in the image. Moreover, when validation on image level is conducted by using an object detection model, the detection with the largest class probability (confidence score) of all detections in one image is usually taken, neglecting all other predictions. This does not capture the performance of the model accurately. Figure 23 illustrates some of the resulting problems:

The image-level ROC curve does not measure the localization performance. As can be seen from Figure 23a, the validation is performed per image, not per object, therefore not considering whether an object is actually hit (see Prediction 2).

The image-level ROC curve is invariant to the number of annotated objects. As can be seen from Figure 23b, the curve can not discriminate between a model detecting all objects in an image (Prediction 1) or just detecting one object (Prediction 2), as long as the largest score is the same across predictions.

The image-level ROC curve is invariant to the number of detected objects. As can be seen from Figure 23c, the curve can not discriminate between a model detecting many FP objects in an image (Prediction 2) or only detecting one FP (Prediction 1), as long as the largest score is the same.

No matching problem category

Metrics should reflect a domain-specific validation goal. This goal may not align with commonly used technical measures like the DSC. Figure 24 shows an example with the property of interest being the accuracy of the ratio between two structure volumes, indicating, for example, the percentage of blood volume ejected in each cardiac cycle [Bamira and Picard 2018]. Both predictions will result in similar averaged DSC scores, although the ratio of the volumes vastly differs. A common segmentation metric thus does not reflect the actual research question in this case.

Pitfalls related to image-level classification

Most issues related to classification metrics are related to one of the following properties of the underlying biomedical problem:

Presence of more than two classes (Figure 28)

Unequal severity of class confusions (Figure 29)

Interdependencies between classes (Figure 31)

Importance of cost-benefit analysis (Figure 33)

Importance of confidence awareness (Figures 38 - 41)

Presence of ordinal classes (Figures 42 - 43)

Furthermore, metric-specific limitations may arise (Figures 44 - 49). Please note that all of these also apply to semantic/instance segmentation or object detection problems. The discourse focuses on the most commonly used image-level classification metrics, as presented in Figs. 4, 5 and 6. For most of the problems, it focuses on the Accuracy, Sensitivity or PPV, because those are the most common metrics for image-level classification in biomedical image analysis. Please note that we do not recommend their indiscriminate use, as they come with limitations (discussed in the following paragraphs), but rather wish to spotlight the problems and pitfalls of those most commonly used metrics.

To preserve the clarity of the illustrations, the most important of the presented metric values may be highlighted with color. Green metric values correspond to a "good" metric value (e.g. a high Sensitivity score), whereas red values correspond to a "bad" value (e.g. a low Sensitivity). Green check marks indicate desirable behavior of metrics, red crosses indicate undesirable behavior. Please note that a low metric value is not automatically a "bad" score. A metric value should always be put into perspective and compared to inter-rater variability. For simplicity, we still use the terms "good" and "bad/poor" throughout the section. Finally, our illustrations do not provide the concrete class probabilities of the presented classifiers.

Accuracy is one of the commonly applied metrics in classification problems, presumably because it is particularly straightforward to interpret. However, the metric is not designed to handle imbalanced data sets, which often occur across all domains. Figure 25 provides an example in which the positive class (orange circle) is heavily underrepresented. While Prediction 1 gives a reasonable separation of the classes, Prediction 2 results in the same Accuracy value (0.97) although the algorithm only provides the majority vote as a result. In this specific example, Sensitivity, PPV and F1 score reveal the issue, as does Matthews correlation coefficient (MCC), a metric designed to handle class imbalance which reflects that Prediction 2 is not better than a random guess (0.00) [Chicco and Jurman 2020]. As many classification measures are easily computable using the number of TP, TN, FP and FN samples, it is highly recommended to report these TP, TN, FP and FN values explicitly and then compute multiple metrics [Hicks et al. 2021].

Plotting the ROC and PR curves (Figs. 25b and c) also reveals the limitations of Prediction 2, which yields an AUROC of 0.52 and AP of 0.04, indicating that the prediction is not better than random guessing.

Some metrics inherently address the problem of performance overestimation or distortion caused by class imbalance. MCC (MCC) and Cohen’s Kappa κ\kappa, for example, are designed in such a way that they assign a (prevalence-aware) random guessing algorithm the value 00. Similarly, AUROC (AUROC) assigns the value 0.50.5 and Balanced Accuracy yields a value of 1C\frac{1}{C}, where CC is the number of classes. For all other metrics (see Figure 25), it is generally recommended to put the resulting values into perspective of with respect to a (prevalence-aware) random guessing algorithm baseline.

Given their prevalence invariance (see Figure 36), metrics such as BA are a good choice in many scenarios. However, they may be misleading in the case of underrepresented classes (see also use case 2 from [Chicco et al. 2021]). In such a scenario, BA yields high scores although the prediction is far from perfect. Here, the MCC and PPV reveal that the prediction is quite inefficient by assessing predictive values (see Figure 26).

Prevalence-dependent metrics yield different metric scores for imbalanced data sets. Figure 27 shows the metric landscapes for a balanced data set, meaning a prevalence of 50%, in comparison to those for imbalanced data (prevalence of 5%) [Brown 2018].

More than two classes available

Many binary metrics can directly be translated to the multi-class case by expanding the confusion matrix to all classes. These classes are often hierarchically structured, for example in the shape of one negative class (e.g. no pathology) and multiple positive classes (e.g. different types of pathologies). Figure 28 shows an example of a classification into triangles and circles, for which the circle class is further separated into two distinct classes (green and orange). The binary performance into triangle vs. circle, shown in the middle, is good (Accuracy of 0.88). But when considering the three classes separately, the prediction struggles to identify the color of the circles, causing their per-class accuracy scores to drop significantly.

Unequal severity of class confusions

In biomedical applications, classes are often not equally important. Consider the task of colon polyp detection in the gastrointestinal tract, for example. To provide the patient with the best care, it is crucial to detect all of these precancerous lesions. This requires a particular penalization of those samples containing a polyp which have been marked as ’no polyp’ (FN), and the metrics need to be chosen accordingly. The PPV, for example, does not include the FN in its definition, hence would not be appropriate for this research question. Sensitivity, on the other hand, would show the desired poor performance in the presence of many FN predictions, as seen in the top row of Figure 29.

For image retrieval, the task of finding images for a specific content, it is not important to find every single existing image, but the images found should be correct. In this setting, the FP (assigning an incorrect image as correct) need to be penalized. Since it includes the computation of FP, in this case, the PPV would be a good metric. In contrast, Sensitivity does not consider FP, therefore being inappropriate in this context (see bottom row of Figure 29). Penalization in both cases is especially important in cases of imbalanced data sets (see Figure 25).

Unequal handling of classes

When it comes to multi-class problems, different approaches may be chosen to compute the metric values. One possibility is to first compute the metric values per class and aggregate them subsequently. Special care has to be taken in the case of unequal importance of the different classes. For example, identifying whether a patient harbors a pathology in general might be more important than identifying the specific type of pathology. In this case, one should not just average over all class metric scores, but instead apply a sufficient weighting scheme. In the example of Figure 30, the triangle class is the most important class but also the one with the lowest per-class Accuracy. Simple averaging, so-called macro-averaging, would ignore that property and thus result in a higher aggregated Accuracy than merited. This effect can be compensated with the Weighted Accuracy.

Interdependencies between classes

If multiple classes are visible in the data set, one should carefully account for interdependencies between the classes. Interdependencies can happen in cases of multi-colinearities in which two classes are correlated, either inherently, such as for the body mass index (BMI) and the body fat percentage, or in the case of dependent data settings, for example multiple images per patient or the presence of confounders. An algorithm aiming to classify the dark blue triangle class in Figure 31 may result in a nearly perfect Accuracy of 0.94, but only because the dark blue triangle almost always appears in conjunction with the orange square. Computing the Accuracy for those images individually without the square class would lead to a much lower performance.

Stratification based on meta-information

Different kinds of meta-information may be available for a data set, including the presence and relevance of artifacts or artificial structures (e.g. metal artifacts in CT images or text overlay in endoscopic data) as well as specifics of acquisition protocols (e.g. acquisition angle or viewpoint) or grid size (cf. Figure 65). Another typical example is the gender of a patient, as shown in Figure 32. In this case, the Accuracy is computed over twelve cases, disregarding the available meta-information (gender). Stratification based on gender will reveal that the prediction performs much worse for women compared to men. These aspects are currently investigated in the fairness literature [Barocas et al. 2017].

Importance of cost-benefit analysis

Most common performance metrics fall short when it comes to determining whether an algorithm can lead to improved clinical decisions. Knowing that an algorithm performs well in terms of calibration, discrimination, and accuracy, for instance, does not necessarily imply that decisions guided by the algorithm will lead to improved clinical outcomes. Motivated by the idea of taking into account the tradeoff between the benefit resulting from the detection of a TP case and the cost of a FP, a risk threshold should be defined for a specific prediction problem. NB (NB) is a metric that incorporates the tradeoff between costs and benefits resulting from the selection of a specific risk threshold in the validation of an algorithm, and thus facilitates informed medical decision-making [Vickers et al. 2016]. In Figure 33, we present a similar example as the one presented by [Vickers et al. 2016]. In this example, nine unnecessary biopsies are deemed acceptable to detect one lesion, resulting in an exchange rate of 1:9. Two scenarios are to be considered here. On the one hand, applying the biopsy to every patient would ensure that all patients with the disease receive a biopsy (benefit) but also result in 75 patients with no disease undergoing biopsy (harm). On the other hand, a marker-based biopsy decision would result in only 20 patients with the disease undergoing biopsy, thus missing 15 of them, while reducing the number of unnecessarily biopsied healthy patients to 60. Only considering the Accuracy would rate both scenarios similarly. However, incorporating the cost-benefit analysis and the resulting exchange rate, NB would indicate that performing biopsy on all patients yields better clinical outcome. Please note that defining risk cutoffs and exchange rates is highly subjective. This is why NB is often considered over various thresholds.

Definition of class labels

Binary classification problems typically involve the step of defining one class as positive and the other as negative. The final scores of per-class counting metrics rely on the definition of the positive class. Figure 34 shows an example in which the class labels are reversed in the same data set. The scores of per-class counting metrics dramatically change when the class labels are switched. On the other hand, multi-class counting metrics remain stable as they rely on the entire confusion matrix and are invariant to changes in the class labels.

Prevalence dependency

The PPV and the NPV (NPV) are common measures to validate classification performances. In many cases, binary classification is considered, for example presence or absence of a disease. In contrast to Sensitivity and Specificity, in case-control studies, PPV and NPV should be seen as the conditional probability of a disease being present based on a test result and the prevalence in a general population [Molinaro 2015].

However, PPV and NPV are frequently used incorrectly. This is due to the fact that many practitioners assume the same prevalence in an analyzed case-control study group as in the general population. However, a study group is often heavily biased, either due to the study design or due to the observation of patient groups from specialized clinics. Thus, the assumed disease prevalence in scientific literature is often higher than that found in the general population (cf. Figure 35). This problem is amplified by default implementations (e.g. in scipy [Virtanen et al. 2020]) which disregard wider population prevalence and calculate prevalence from the study group. Without prevalence correction, this can lead to misleading results, confusion among patients and ill-informed policy-making.

Metrics depending on the prevalence cannot easily be compared across data sets with different prevalences, as shown in Figure 36. Here, two situations with the same Sensitivity and Specificity values but with different prevalences (50% vs. 90%) are presented. Only metrics that do not rely on the prevalence (such as BA and EC) can be used for a comparison across data sets.

For a prevalence unequal to 50%, BA and Youden’s Index J may lead to different rankings compared to MCC and Cohen’s Kappa κ\kappa [Chicco et al. 2021], as illustrated in Figure 37. In this example, predictions for three different prevalences (50%, 40% and 60%) are shown. Only in the case of a 50% prevalence will the rankings generated by all metrics be the same, with all preferring Prediction 2. For different prevalences, rankings may differ: While BA and Youden’s Index J prefer Prediction 2 over Prediction 1, MCC and Cohen’s Kappa κ\kappa favor the Prediction 1.

Discrimination vs. calibration

A model’s discrimination capability refers to how well it separates samples with and without the class of interest. If all predicted class scores for the samples with the class of interest are higher than for the others, the model discriminates perfectly, reflected in an AUROC score of 1. A calibrated model outputs predicted class scores that match the empirical success rate (e.g. outputs with score 0.8 for a specific class, empirically belong to this class in 80%80\% of the cases) [Cook 2007]. However, it may happen a perfectly discriminating model, such as the one presented in Figure 38, is not well calibrated. The Prediction gives a confidence score of 0.52 for all samples of the positive class (orange circle) and a score of 0.51 for all samples of the negative class (blue triangle). Although this leads to a perfect discrimination, the calibration is very poor and the predicted probabilities would not be helpful in practice at all.

Similarly to standard validation metric measuring the discrimination power of algorithms, calibration metrics come with limitations. For example, the BS [Brier et al. 1950, Gneiting and Raftery 2007] is defined as the squared difference between a predicted class score ftf_{t} and the actual outcome oto_{t} for one of nn events tt, typically defined as 1 if the event occurred and 0 otherwise.

However, the BS should be used with caution in the case of ordinal classes, in which the classes underlie an order, such as in the example in Figure 39. Here, three classes are present which denote patient survival times in years. The first class refers to zero to one year, the second class to two to three years, and the third class denotes a survival time of three or more years. While the actual survival is zero to one year (class 1; orange circle), Predictions 1 and 2 are predicting classes 2 (blue triangle) and 3 (green square), both with a probability of 100%. Although both predictions are incorrect, Prediction 1 would be preferable given the ordinal scale of variables: The predicted time frame of one to two years of survival time is closer to the actual survival time of zero to one year than that of Prediction 2. However, both predictions would yield the same BS.

Calibration errors [Guo et al. 2017a], on the other hand, compute the difference between the accuracy of a prediction and the average predicted class scores. Classifier outputs are generally continuous, which often reduces the number of available samples per prediction to one. Strategies for alleviating the sparse sampling problem include binning the continuous scale [Maier-Hein et al. 2022], as done by the ECE and MCE [Guo et al. 2017a, Naeini et al. 2015]. However, the binning itself is not strictly defined, neither in regard to whether the bins should be equidistant nor in the number of bins. Depending on the binning strategy, the scores of the ECE and MCE will differ, as shown for the example in Figure 40.

Calibration can be defined on different levels. While top-label and class-wise calibration only consider certain parts of the predicted class score vectors, solely the canonical definition of calibration computes the calibration errors based on the full probability vectors. As shown in Figure 41(a), the CE (CE) may imply a perfect calibration although the prediction is not perfectly calibrated, as indicated by the canonical CE [Gruber and Buettner 2022, Vaicenavicius et al. 2019]. Moreover, the CE depends on the sample size [Gruber and Buettner 2022]. Although a model may be perfectly calibrated, the ECE only converges towards a perfect score with an increasing number of data samples, as shown in Figure 41(b).

Classification metrics in the case of ordinal classes

In the case of ordinal classes, similarly to calibration metrics, the scores of classification metrics should be interpreted with caution, as shown in Figure 39. The classification or grading of patients according to a scale of disease severity would be a typical use case. In Figure 42, the severity of the disease is given by three classes, with class 0 denoting the lowest severity and class 2 the highest. Predictions 1 and 2 agree in the grading of the first four patients. However, Prediction 1 classifies Patient 3 into the lowest severity, and thus performs much more poorly than Prediction 2, which is only one class of severity below the reference. However, both Accuracy and MCC do not spot this difference. Underestimating the severity of disease in a patient, however, would lead to serious consequences in clinical practice and thus needs to be heavily penalized. Only metrics with pre-defined weights would highly penalize Prediction 1 (here: EC and quadratic-weighted Cohen’s Kappa).

Label shift in the case of ordinal classes

In ordinal class scenarios, one should also be aware of potential label shifts, in which the distribution or number of labels changes over time or data sets [Zhang et al. 2021]. Figure 43 shows the metric scores for two data sets with the same labels and same number of data, but a differing number of images per label. While Data set 1 is balanced, Data set 2 underwent a label shift with different per-class prevalences. While the Accuracy values do not change, the BA and MCC are slightly affected by the label shift. The most dramatic change is observed for the quadratic-weighted Cohen’s Kappa, with a drop of 0.17 after label shift. Thus, when comparing the performance of a model on data sets with varied regional or temporal origins, one should take precautions when utilizing the quadratic-weighted Cohen’s Kappa Equal Sensitivities of 0.8 per class and uninformed guessing in the remaining 20% of cases were assumed. See https://github.com/agaldran/kappa_pitfall/blob/main/kappa_pitfall.ipynb for details. .

Same metric values for different confusion matrices

Per definition, the LR+, Youden Index J and BA are calculated from Sensitivity and Specificity. However, the exact same LR+ or BA values can occur for very different specifications of the confusion matrix. Figure 44a shows three examples of predictions, yielding quite different confusion matrices. For example, the leftmost prediction shows a very high number of TN and FN, while the prediction in the middle only has very few false predictions. The prediction on the right correctly classifies most of the negative samples. Despite the different Sensitivity and Specificity values, LR+ will be the same for all three examples [Dujardin et al. 1994]. The same can occur for the BA, as shown in Figure 44b. In both cases, the substantial differences between the predictions remain hidden unless Sensitivity and Specificity are reported in addition. Similar issues hold true for other metrics that rely on multiple measures. An AUROC of 0.9, for example, may correspond to various appearances of the ROC curve, and various underlying distributions of positive and negative samples and predictions [Wald and Bestwick 2014].

Upper bound in Cohen’s κ\kappa calculation

Cohen’s κ\kappa measures the agreement between ratings while incorporating information on the Accuracy by chance. It therefore investigates how well a prediction follows the distribution of the actual class. The maximum Cohen’s κ\kappa helps interpreting the calculated κ\kappa score by symbolizing the corner case in which either the FP or FN are equal to 0 [Umesh et al. 1989]:

The maximum Cohen’s κ\kappa score will be lower as the number of the predicted positive and negative samples diverges more from the actual number of positive and negative samples. This is shown in Figure 45 with two predictions. Prediction 1 achieves lower Accuracy and Cohen’s κ\kappa scores compared to Prediction 2, as it only predicts a very low number of TP. However, the predicted distribution of samples in Prediction 1 is closer to the actual distribution of samples (13 circle predictions vs. 15 actual circles and 87 triangle predictions vs. 85 actual triangles). The distribution of samples of Prediction 2 differs more from the actual distribution, yielding a lower Cohen’s κmax\kappa_{max} value https://www.knime.com/blog/cohens-kappa-an-overview.

Similarly, the theoretical lower bound for the MCC is not always achievable, such as in the example in Figure 46.

Determination of a global threshold for all classes

In a scenario with multiple classes, a single cutoff value for a threshold needs to be chosen for all classes. Multi-threshold metrics, such as AUROC, may be overly optimistic as the optimal threshold range for one class may differ from the optimal threshold range for another. Figure 47 illustrates the case of three classes that yield three perfect AUROC scores. However, when choosing a single threshold of 0.8 based on class 1 for all classes, the respective counting metrics will yield very poor results for classes 2 and 3.

Model bias

Image analysis models might be affected by image features that human experts naturally ignore and that are not truly relevant for the prediction task. Such model biases might be revealed by analyzing data external to those used for development [Kleppe et al. 2021]. However, multi-threshold metrics such as AUROC will inherently obscure model biases that affect the predicted class scores but not their ranking, if model calibration is not considered. Such a pitfall is illustrated in Figure 48 (similar to [Kleppe 2022]). The example assumes a model that works perfectly on a subset of a larger dataset used for training. Generalizing to a different dataset when using a model with a severe bias is simulated by linear rescaling, i.e., dividing all model scores by 10, meaning that all predicted class scores will be 0.1 or below, while the ranking of the scores remains the same. Since the ranking remains the same, the AUROC value will not change. However, given the low predicted class scores, the model would be very confident that all cases actually belong to the negative class. Here, the substantial model bias is thus not captured by multi-threshold metrics such as AUROC, which would yield perfect metric values on the external dataset. In general, any metric calculation that adapts to the score distribution of the test dataset will obscure such model biases, including metrics such as Specificity at a particular level of Sensitivity (see [Kleppe 2022] for details).

Small sample sizes

Small sample sizes are a common issue in the biomedical image analysis domain. Caution should be applied when calculating, for example, the AUROC in the presence of only very few images. Figure 49 provides an example of six images (three positive, three negative samples) for two data sets and the respective predicted class scores of an algorithm. The data sets only differ in a single image. However, this apparently minor difference will result in a difference in the AUROC scores of 0.11. By calculating the 95% CI (CI), we see that the CI are very large and do not allow for proper interpretation of AUROC scores in the case of very small sample sizes.

Pitfalls related to segmentation

All pitfalls compiled for this work and relevant for semantic or instance segmentation are summarized in Table 1. This section focuses on limitations for semantic segmentation, but some of them are also transferable to other problem categories, as indicated in the table. Limitations of metrics are typically related to the following properties:

Small size of structures relative to pixel size (Figures 50 - 51)

High variability of structure sizes (Figures 52 - 53)

Complex shapes of structures (Figures 54 - 56)

Particular importance of structure volume (Figure 57)

Particular importance of structure center (Figure 58)

Particular importance of structure boundaries (Figure 59 - 60)

Possibility of multiple labels per unit (Figure 61)

Possibility of outliers in reference annotation (Figure 63)

Possibility of reference or prediction without the target structure (Figure 64)

Further pitfalls are related to technical peculiarities, such as the image resolution (Figure 65), preference for over- vs. undersegmentation (Figure 66) and the choice of global decision threshold for creating the confusion matrix (Figure 67).

The limitations are presented for the most commonly used overlap segmentation metrics, namely DSC, IoU, and the most common boundary-based metrics, namely HD, HD95, ASSD, MASD and NSD. The NSD calculation is based on a user-defined threshold (cf. Figure 10). Results differ for different thresholds. Unless stated otherwise, we set the threshold to τ=1\tau=1.

To preserve clarity of the illustrations, specific values may only be highlighted for one metric from each metric family, if the other metrics share similar properties (e.g. DSC and IoU share the same properties). Green metric values correspond to a "good" value (e.g. a high DSC or a low HD score), whereas red values correspond to a "bad" value (e.g. a low DSC or a high HD score). Green check marks indicate metric scores reflecting the research question, red crosses show those that do not. Please note that a low DSC value (or similar) is not automatically a "bad" score. A metric value should always be put into perspective and compared to inter-rater variability. We only use the terms "good" and "bad/poor" for simplicity.

Segmentation of small structures, such as brain lesions or cells imaged at low magnification, is essential for many image processing applications. In these cases, the DSC or IoU may not be appropriate metrics, as illustrated in Figure 50 (cf. [Cheng et al. 2021]). In fact, a single-pixel difference between two predictions can have a large impact on the metric values. Given that the correct outlines (e.g. of pathologies) are often unknown and taking into account the potentially high inter-observer variability related to generating reference annotations [Joskowicz et al. 2019], it is typically not desirable for few pixels to influence the metrics as much. This problem is particularly amplified in cases of large variability of structure sizes (cf. Figure 52). The same problem arises for other versions of the DSC or IoU. For example, the clDice is often used in the case of tubular structures. Similarly to the original DSC metric, a single-pixel difference will have a larger influence on the clDice for small tube structures compared to larger ones. This pitfall also applies to object detection tasks. It should be noted that once a data set exclusively contains only very tiny structures, one may consider this problem to be an object detection rather than a segmentation problem.

High variability of structure sizes

The size of target structures may vary substantially, both within an image and across images. For example, in medical instrument segmentation in laparoscopic video data, an image frame may contain full-sized instruments as well as only the tip of an instrument just entering the scene [Roß et al. 2021]. In these cases, metrics need to be chosen carefully. As shown in the example above (Figure 50), metrics such as the DSC or IoU are typically not well-suited for very small structures. Furthermore, size stratification – the aggregation of metric values for objects of similar sizes to uncover differences between them – should be employed. Figure 52 shows an exemplary data set of four images, containing three large structures and one small structure. When aggregating over all DSC values, the average DSC is 0.82. Computing the average for large and small structures separately, however, shows that the performance is much lower for the small structures compared to the large ones, demonstrating the large influence of the low metric values of small objects. This pitfall also applies to object detection tasks and other metrics.

The ASSD and MASD often result in very similar metric scores, as both compute the average over distances from the boundaries. While ASSD calculates the general average over all distances, the MASD computes the average distance for each of the structures and calculates the mean over those averages. MASD, however, substantially rewards situations in which the predicted structure is significantly smaller than the reference object and is located near the surface of the reference, as shown in the example of Figure 53. In this case, the average distance from the prediction to the reference is approximately zero. Thus, only the average distance from the reference to the prediction decides the score. For MASD, this distance is halved, which leads to a significantly better score compared to that obtained with ASSD.

Complex shapes of structures

Metrics measuring the overlap between objects are not designed to uncover differences in shapes. This is an important problem in many applications such as radiotherapy, for which identifying and treating all parts of the tumor is essential to avoid recurrence [Burnet et al. 2004]. Figure 54 illustrates that completely different object shapes may lead to the exact same DSC and IoU values. Boundary-based measures are able to detect the changes in shapes [Taha and Hanbury 2015]. Note that this pitfall also applies to object detection tasks.

Shape unawareness is especially harmful in the case of very complex shapes, for example in bronchi. These structures typically feature many small branches, as depicted in a simplified manner in Figure 55. In this example, Prediction 1 only shows the root of the structure, while Prediction 2 focuses on its center line, which would be the preferred option. However, common overlap-based metrics such as the DSC yield the same value for both predictions. The clDice on the other hand, designed to capture the center line focus, recognizes the difference and penalizes Prediction 1 with a substantially lower value.

While the clDice can typically validate tree-like structures such as vessels better than other overlap-based measures, it still faces limitations. Clinicians are sometimes particularly interested in predicting the length of a vessel, whereas its width may be of lesser importance. As shown in Figure 56, clDice captures the length of the vessel-like structure in a similar manner as another proposed metric called Length [Gegúndez-Arias et al. 2011]. However, it does not recognize that Prediction 2 is spatially shifted to the left. This shift could be captured by domain-specific metrics such as the Area [Gegúndez-Arias et al. 2011], introducing a tolerance factor to vessel width.

Particular importance of structure volume

Depending on the domain focus, a surgeon, radiologist or similar may be especially interested in the volume of a segmented structure. The most commonly used metrics may, however, result in predictions at entirely wrong locations if boundary or overlap are not considered. Figure 57 shows two predictions of a 3x3 square structure, both of them being at the wrong position. While the volume difference is correct for both predictions, the overlap is zero. Only boundary-based metrics will indicate the magnitude of mislocalization of the predicted objects.

Particular importance of structure center

The structure center point or center line may be more important than an accurate boundary or overlap of the structure, as for example in nerve segmentation [Mlynarski et al. 2020]. In these cases, the accuracy of the center point or line should be examined via an additional metric to make sure the center is correct for the prediction. Figure 58 shows two predictions yielding the same DSC values, as they have the same overlap to the reference annotation. However, only Prediction 1 is centered around the same point as the reference, while Prediction 2 is shifted slightly towards the upper left corner and thus centered incorrectly. This pitfall also applies to object detection tasks. It should be noted that once the center location is of particular importance to the task, one may consider it an object detection rather than a segmentation problem.

Particular importance of structure boundaries

While boundary-based metrics such as the HD(95), ASSD and others can help to detect shape differences between the reference and the predicted object, they do not focus on the object itself. As shown in Figure 59(top), the boundary-based metrics do not recognize a prediction with a large hole inside as poor (Prediction 2). Furthermore, in Figure 59(bottom) [Taha and Hanbury 2015], those metrics do not punish the spotted pattern within the object. It should be noted that this behavior may also be desirable. For example, it may be highly difficult to decide whether a necrotic core (hole) is present in a tumor or not. A boundary-based metric would not punish errors resulting from such annotation uncertainties.

Please note that boundary-based metrics are not appropriate under several circumstances. In scenarios in which multiple structures of the same type are present within the same image (e.g., in multiple sclerosis lesion segmentation), for example, a potential pitfall is related to comparing a predicted structure boundary to the boundary of the wrong instance in the reference, as shown in Figure 60. In this example, the prediction misses the reference object R1, while poorly segmenting R2. In the case of a semantic segmentation, the distance between the red pixel and the object boundary would be zero, which may not be desired. If phrased as an instance segmentation problem, the distance would be higher, and thus penalized.

Possibility of multiple labels per unit

In several biomedical imaging scenarios, multiple labels per pixel may be possible. A prominent example would be the tumor core inside the tumor [Menze et al. 2014]. Often, however, prior knowledge related to such scenarios (e.g. a tumor core cannot lie outside the tumor) is not reflected by common metrics, which simply calculate the agreement of the reference and prediction per class. Figure 61 shows two predictions for a multi-label example. The DSC value of Label 2, which is required to be inside of Label 1, is higher for Prediction 2 although Label 2 is also found outside the Label 1 area. For simplicity, we only show the results for the DSC metric. This pitfall also applies to object detection tasks.

Imperfect reference standard

A high quality reference annotation is crucial to determine the performance of a supervised learning algorithm. A prediction can be almost perfect, but low quality reference images will still result in a bad metric score. Especially in the medical domain, the inter-rater variability is often very high as domain knowledge is required and experts themselves often disagree [Joskowicz et al. 2019]. Figure 62 shows two masks from different annotators approximating the same structure. Although the annotations differ only slightly at the boundary, the DSC score is 0.7. With such inter-rater variability, a DSC score of 1 would not be achievable in practice. To address this issue, the NSD metric can be applied as an alternative or additional metric, as it is designed to allow a certain tolerance of outline pixels based on the threshold τ\tau. This pitfall can also be translated to object detection and image-level classification tasks.

Possibility of outliers in reference annotation

The presence of spatial outliers, such as noise or reference annotation artifacts, may severely impact performance metric values. Figure 63 demonstrates how a single erroneous pixel in the reference annotation (or the prediction) leads to a substantial decrease in the measured performance, especially in the case of the HD. Using the 95% percentile instead of the maximum (HD95) to compute the distance substantially improves the metric score as it can handle outliers. Please note that the presented example may also be seen vice versa, with a prediction including single pixel errors. It should further be noted that whether or not outliers should be considered depends on the respective research question.

Possibility of reference/prediction without target structure(s)

A given data set may contain reference annotations without the target structure(s). For example, the data set may consist of healthy and sick patients. A healthy patient will not have a tumor in the image, yielding an empty reference if the tumor is the targeted structure. An algorithm should be careful not to classify a healthy patient as tumourous as this may lead to unnecessary medical interventions. Similarly, a patient with a tumor should not be classified as healthy (empty prediction). These cases require special care to be taken in the validation, because some metrics may be undefined due to division by zero errors or similar. It is necessary to either choose appropriate metrics that consider empty references (or predictions) or account for it in the metric implementation. For example, boundary-based metrics such as the HD(95) and ASSD will be NaN if one of the structures is empty. Figure 64 shows three examples, for which several counting- and boundary-based metrics were computed. The top row depicts the case of an empty reference and a prediction of an object. Given the number of TP and FN being 0, this will result in a division by zero in the Sensitivity calculation, yielding a NaN score. A similar case is given in the second row, showing an empty prediction for a given target structure in the reference annotation, yielding an undefined Precision. When both reference and prediction are empty (bottom row), all scores will be undefined. Please note that this example is shown for a validation per image, as done for segmentation tasks. For classification and object detection tasks, the validation is typically performed over the whole data set, which would possibly preclude this problem. The presented pitfall also applies to object detection tasks.

Technical peculiarities

Several technical peculiarities also have an impact on metric behavior. For example, the image resolution and pixel sizes highly influence the reference annotation and the predicted shapes in image processing tasks. Figure 65 illustrates how the reference annotation differs between a low resolution image (top) and a high resolution image (bottom) compared to a circle. The latter is more exact. A prediction of the same size will therefore lead to different corresponding metric values, independent of the type of the metric. This pitfall also applies to object detection tasks.

In some applications such as radiotherapy, it may be highly relevant whether an algorithm tends to over- or undersegment the target structure. The DSC metric, however, does not represent over- and undersegmentation equally [Yeghiazaryan and Voiculescu 2018]. As depicted in Figure 66, a difference of a single layer of pixels in the outline yields different DSC scores (oversegmentation preferred) [Taha and Hanbury 2015]. Other boundary-based performance values such as the HD are invariant to these properties.

Another technical peculiarity is the choice of global decision threshold. Most methods in modern image analysis output continuous class scores. While it is quite common to provide those scores in image-level classification and object detection tasks, segmentation architectures often do not output class probabilities per pixel. However, fuzzy segmentation masks are getting more and more common (for instance, see [Nida et al. 2019, AlZu’bi et al. 2020, Kaftan et al. 2008]) and the choice of a global decision threshold τ\tau is very important for the algorithm’s result. Figure 67 (cf. [Nair 2018]) shows the predicted class probabilities for a reference annotation. For a binarization typically required for segmentation outputs, a threshold needs to be defined based on which a pixel is assigned to a class (here: a pixel with class probability < τ\tau corresponds to the background class, otherwise to the foreground class). The resulting segmentation masks are shown for the thresholds 0.2, 0.5 and 0.8. It can be seen that the respective masks completely differ across the thresholds. Consequently, metric values will also vastly change.

Pitfalls related to object detection

All pitfalls compiled for this work and relevant for object detection are summarized in Table 1. Note that these pitfalls equally apply to instance segmentation problems. While most issues related to the actual metric selection have already been mentioned in the previous paragraphs, this section is primarily dedicated to technical peculiarities related to the localization and assignment criteria. These include:

Mathematical implications of center-based localization criteria (Figures 68 - 69)

Mathematical implications of IoU-based localization criteria (Figure 70)

Mathematical implications of the choice of assignment strategies (Figure 71 - 73)

Effect of small structures on localization criterion (Figure 74)

Perfect Boundary IoU for imperfect prediction (Figure 75)

Possibility of reference or prediction without the target structure and NaN handling (Figure 76)

Average Precision vs. Free-response ROC score (Figure 77)

Free-response ROC score is not standardized (Figure 78)

Effect of predicted class probabilities on multi-threshold metrics (Figures 79 - 81)

Non-standardized metric definition (Figure 82)

Before calculating metrics for object detection tasks, it is necessary to define what qualifies a detection as a hit (TP) or miss (FP). There are multiple ways to define a hit, all of which come with their specific limitations. Below, the most commonly used center-based localization criteria are presented For more details, please refer to the blogpost ”Evaluation curves for object detection algorithms in medical images”: https://medium.com/lunit/evaluation-curves-for-object-detection-algorithms-in-medical-images-4b083fddce6e. The presented pitfalls are also valid for approximations or bounding circles instead of the shown bounding boxes..

For the center-cover criterion, the reference object is considered a hit if the center of the reference object is inside the predicted detection. Figure 68a shows how this criterion can be fooled by a model outputting very large boxes to maximize the chance of a correct detection. The same issue holds true for the Point inside Mask/Box criterion, which is fulfilled if a single predicted point lies inside the reference object.

In the case of the distance-based hit criterion, a prediction is considered a hit if the distance dd between the center of the reference and the detected object is smaller than a certain threshold τ\tau. In Figure 68b, both predictions have the same distance to their corresponding reference object centers. However, the prediction on the top right shows no overlap with the reference and should therefore not be counted as a hit.

The center-hit criterion holds true if the center of the predicted object is inside the reference bounding box or contour. Given this definition, large reference objects are more likely to be hit, as shown in Figure 68c. The left prediction is defined as a missed object (FP), the right detection as a hit because of its larger size.

The center point might not be a good reference for complex shapes such as tubular structures (Figure 69). In those cases (and if annotations are provided in the form of masks), a binary “Point inside Mask” criterion might be the better choice.

Mathematical implications of IoU-based localization criteria

The most commonly used hit criterion is determined by computing the IoU between the predicted and the reference mask/bounding box/boundary. Pitfalls related to IoU-based criteria are mainly related to the setting of the threshold (Figure 70). Many biomedical applications involve 3D rather than 2D images. When working with a higher dimension, it should be kept in mind that metrics may be affected. The additional dimension will lead to overlap errors being punished even more. Figure 70a shows a comparison of the IoU for two rectangles (or bounding boxes) in 2D and 3D. Being mistaken by one voxel in the zz-dimension will lead to a much lower IoU score in 3D compared to the 2D case.

As IoU-based criteria take the overlap between regions into account, it is only possible to cheat with very large boxes if the IoU threshold is set to a very small value (here: 0), as shown in Figure 70b. However, special care should be taken when applying the Box IoU in the presence of highly concave or elongated structures, as illustrated in Figure 70c. This is because bounding boxes may quickly grow for narrow and diagonally placed objects, such as medical instruments, and result in FP although visual inspection would indicate a correct prediction.

Mathematical implications of the choice of assignment strategies

Different assignment strategies may yield different proportions of TP, FP and FN. This is shown in Figure 71a, in which two different assignment strategies based on the center distance are shown for the same situation with three reference and four predicted objects. Matching one reference object at a time (Greedy by Center Distance Matching) yields two TP matches: [R1, P1] and [R2, P2]. This is due to the fact that the distance between R2 and P2 is smaller than the distance between R2 and P3 and that between R3 and P2. Thus, R2 is assigned to P2. The distance between R3 and P3 or P4 is greater than the defined tolerance, thus, P3 and P4 are seen as FP and R3 is a FN. If an Optimal (Hungarian) Matching [Kuhn 1955] is applied, matching is performed differently by optimizing a cost function. This yields a different assignment with three TP objects. Here, P3 is assigned to R2 and P2 is assigned to R3. As all reference objects are assigned, no FN appears. Only P4 is counted as a FP.

In addition, researchers should think about whether FN should be further distinguished into real FN, i.e., reference objects for which no prediction exists (see left example in Figure 71b) and predictions that are outside a distance threshold or below an overlap threshold (see right example in Figure 71b). The latter case may be more useful in the present example, as R2 is not fully overlooked although P2 lies outside of the threshold area.

Penalization of one prediction assigned to multiple reference objects requested

Object detection and instance segmentation algorithms typically involve the step of assigning predicted objects to reference objects. This may result in one prediction being assigned to multiple reference objects or vice versa. Using the IoU > 0.5 (or similar threshold) as assignment strategy may end up in a very strict penalty of two FN and one FP if the IoU for both reference objects is smaller than the threshold (as shown for two reference objects in Figure 72. Using the IoR may result in a less strict penalization for what could be interpreted as only a single error with an additional step of penalizing the non-split errors (either directly in object detection or indirectly in instance segmentation) [Matula et al. 2015].

Possibility of disconnected structures

The Box IoU is sometimes employed despite access to pixel-mask annotations. A possible explanation is that researchers want to phrase their problem as an object detection problem and then apply the most commonly used validation methods. Such simplification might cause problems if structures are not well approximated by a box shape, or if structures yield multi-component masks, appearing to be disconnected. This may occur in the case of a tubular structure shown in a 2D tomographic image or a medical instrument occluded by tissue in an endoscopic image, for example. Figure 73 provides examples of a complex diagonal (top) and a disconnected structure (bottom). Both box predictions yield a Box IoU larger than 0.3, and are thus counted as TP because of the chosen localization threshold. Nevertheless, Prediction 1 is not hitting the actual object at all. This is due to the fact that the target structures are not well approximated by the bounding box, leaving many empty pixels in the boxes.

Effect of small structures on localization criterion

Box IoU and Mask IoU are not sensitive to structure boundary quality in larger objects (cf. Section 5). This is due to the fact that boundary pixels will increase linearly (quadratically in 3D) while pixels inside the structure will increase quadratically (cubically in 3D) with an increase in structure size. In consequence, the IoU-scores tend to be higher for large objects compared to small objects. For this reason, localization criteria such as the Boundary IoU were designed.

Figure 74 shows an example of Mask IoU and Boundary IoU for a large (top) and a rather small structure (bottom). In the case of the Mask IoU, the score drops substantially for the small structure, while the scores are more consistent for the Boundary IoU when comparing small and large structures. This pitfall also applies to segmentation problems in which the (Mask) IoU and Boundary IoU are applied as overlap-based metrics.

It should further be noted that the Boundary IoU is highly dependent on the chosen distance dd, as illustrated in Figure 74 (third vs fourth column). Similarly to the example provided in Figure 59, the Boundary IoU can be fooled to result in a perfect value of 1.01.0. A prediction with a hole in the middle of the structure may result in a perfect metric score if the distance is chosen in a way that it incorporates all pixels of the predicted mask, as shown in Figure 75 [Cheng et al. 2021]. The Mask IoU, however, will be able to recognize the problem, as it completely measures the overlap between both structures. [Cheng et al. 2021] propose to use the min⁡(BoundaryIoU,MaskIoU)\min(BoundaryIoU,MaskIoU) to resolve this issue. Please note that the same limitations also affect other distance-based measures, such as the NSD or HD metrics. This pitfall also applies to segmentation tasks.

Possibility of reference/prediction without the target structure and NaN handling

When validating an object detection problem per image rather than per data set, a reference or prediction image without the target structure(s) may become problematic as some metric values will turn into NaN due to division by zero errors (cf. Figure 64). Figure 76a shows potential scenarios for a validation per image categorized by the presence and absence of TP, FP and FN. Four occurrences of NaN are presented. To proceed with the validation, namely aggregating metric values for every image over the entire data set, a NaN strategy needs to be defined for every use case.

Average Precision vs. Free-response ROC score

While the AP constitutes the standard metric for object detection and instance segmentation in the computer vision community, the FROC score is often favoured in the clinical context. In contrast to the AP, the FROC score takes into account the total number of images in the data set. As can be seen from Figure 77, both data sets D1 and D2 will yield the same AP score, although data set D1 contains two images and D2 contains four images. The FROC score, however, will reflect that the number of images is different for both data sets and that data set D2 contains two images that do not contain any FP. Thus, the FPPI will be lower in data set D2, yielding a higher FROC score.

FROC is not standardized

Although the FROC Score is often favored by clinicians over the AP given its simpler interpretation, it comes with a major drawback. While the values of the x-axis are clearly defined and bounded between 0 and 1 for the AP, there is no fixed or standardized definition for the FPPI used for the FROC curve. Depending on the defined range of the FPPI, the FROC Score will change. Figure 78 shows three examples of different bounds for the x-axes (left: , middle: , right: ) for the same prediction – all of them yielding different FROC Scores.

Multi-threshold metric-related properties

In the next paragraphs, we highlight some limitations of the multi-threshold metrics, exemplarily for the AP metric, which can be transferred to other multi-threshold metrics, such as the AUROC [Oksuz et al. 2018]. By definition, multi-threshold metrics are ranking metrics, which rank the predicted class probabilities or confidence scores (cf. Figure 15). They are not designed to reflect the calibration of confidence or class scores, as shown in the following examples. Please note that we disregard the concrete choice of the localization criterion here for simplicity.

Predicted class probabilities

The PR curve and the resulting metric score AP highly depend on the ranking of predictions, based on their predicted class probabilities or confidence scores. Small changes in the scores can therefore significantly change the metric value, as shown in Figure 79. On the other hand, as long as the ranking remains unchanged among predictions, the predicted class probabilities themselves are not important for the result, although they should be (see Figure 80).

FP with low predicted class probabilities

False positive predictions with lower predicted class probabilities than the last correctly predicted reference, corresponding to the end of the PR curve, do not affect the AP scores. Figure 81 shows two examples that are very similar, only differing in the number of wrongly predicted objects. Prediction 2, with two FP, performs worse than Prediction 1 with only one FP. Nevertheless, the AP scores are the same for both models, given the low confidence of the second FP of Prediction 2. Thus, a prediction may contain numerous FP objects with low confidence scores, but still yield a deceptively high AP score, as this practice only increases the score but never decreases it by potentially hitting more reference objects.

Non-standardized metric definition

Standard metric implementations do not always cover corner cases or handle them differently. For example, for the creation of the PR-curve and the subsequent computation of the AP metric, predictions are typically ranked by their predicted class scores. Duplicate predicted class scores can either be computed in a single step per score or can all be treated together in a common step. Based on the chosen implementation, the PR-curve will differ in appearance, causing a sometimes substantial change in the AP scores. Figure 82a provides such an example, in which the final scores differ by 0.06 solely because of two predictions having the same score.

This difference in implementation is even more critical in the extreme case of all predictions having the same score (e.g., if no class scores are available; see Figure 82b). However, it should be noted that although this practice is not uncommon (e.g., [Bai and Urtasun 2017, De Brabandere et al. 2017, Gao et al. 2019, Hirsch et al. 2020, Kulikov and Lempitsky 2020]), multi-threshold metrics such as the AP should generally be avoided if no class scores are available. Moreover, some implementations force the PR-curve to start at the point (0, 1), i.e., to always start with a PPV value of 1 instead of the actual first PPV value that was computed. This may lead to a substantial overestimation if there are no confidence scores or if the number of instances is low.

Pitfalls related to analyses and post-processing

A data set typically contains several hundreds or thousands of images. When analyzing, aggregating and combining metric values, a number of factors need to be taken into account. Pitfalls in this step are primarily related to the following aspects:

Metric aggregation for invalid algorithm output (e.g. NaN) (Figures 84 - 85)

Hierarchical data aggregation (Figure 86)

Aggregation in the presence of multiple classes (Figure 87)

Combination of related metrics (Figure 88)

Insufficient biomedical relevance of metric score differences (Fig 90)

Relying on only reporting aggregated metric scores may result in missing essential information on algorithm performance. Therefore, raw metric values (e.g. per image) should always be shown, for example in the shape of boxplots, as depicted in the top left of Figure 83. However, boxplots will only provide information on some key descriptive statistics, like median or 1st and 3rd quartiles. Another choice can be violin plots, which further visualize the raw data distribution. The top right of Figure 83 illustrates the multimodal distribution of the underlying data, invisible in the boxplot. Furthermore, using a violin plot and/or plotting the raw metric values for each data point on top (Figure 83, top right and bottom left) will reveal the complete data distribution. In the example below, many values lie below the 3rd quartile, although the box looks tight. Nevertheless, even these two visualizations may hide important information. Assume a data set with metric values of four different videos. Color- or shape-coding the metric values by the video type (Figure 83 bottom right) reveals a huge cluster of extremely low DSC values only affecting Video 4 (pink), which would have been hidden by the other two types of visualization.

Metric aggregation for invalid algorithm output (e.g. NaN)

In challenges or benchmarking experiments, metric values are often aggregated over all test cases to produce a challenge ranking [Maier-Hein et al. 2018]. Missing data plays a crucial role when aggregating metric values and occurs primarily due to two reasons: invalid output of the algorithm or metric routine output resulting in NaN, and non-submission of single cases (by accident or even for cheating [Reinke et al. 2018]). Figs. 84 and 85 illustrate why a strategy on how to handle missing values may be crucial.

In the case of metrics with fixed boundaries, such as the DSC or the IoU, missing values can easily be set to the worst possible value (here: 0). For spatial distance-based measures without lower/upper bounds, the strategy of how to treat missing values is not trivial. In the case of the HD, for example, one may choose the maximum distance of the image or normalize the metric values to $$ and use the worst possible value (here: 1). Another possibility is to employ a case-based ranking scheme [Maier-Hein et al. 2018] and assign the last rank for every missing submission. Furthermore, aggregating with the mean may not be a good choice as results are unlikely to be normally distributed. Crucially, however, every choice will produce a different aggregated value (Figure 85), thus potentially affecting the ranking. Another way of handling missing values would lie in rejecting the entire submission in a challenge.

However, metric values may also be undefined (NaN)) if either reference or prediction or both are empty. In the case of empty reference and prediction, an undefined metric value (e.g. DSC) may be a desirable outcome and should therefore not necessarily be penalized.

Hierarchical data aggregation

Nowadays, most data sets are inherently hierarchically structured, meaning that the test cases are not independent. Data may, for example, come from several centers or hospitals, and for every center or even within one, different devices may be used for image acquisition, and images may be drawn from different subjects or patients. This should be kept in mind when visualizing and aggregating data points, especially if the individual tree nodes end in a large variation in the size of images. Figure 86 shows an example of five patients with an unequal number of images associated with them. Just averaging all metric values for every image would result in a high average DSC of 0.8. Averaging metric values per patient reveals that the DSC values are much higher for Patient 1, overruling the other patients due to the high number of samples for this patient. Aggregating per patient first and averaging subsequently will resolve this issue.

Aggregation per class

Similar approaches should be chosen in the presence of multiple classes in a data set. The performance may differ significantly for the individual classes, as shown in Figure 87. The background class in particular will result in a nearly perfect averaged DSC value, whereas the average scores for classes 2 and 3 are much lower. Aggregating over all values, not considering the class, would hide this information. An alternative approach to the problem lies in the application of metrics that explicitly handle class balance, such as using the Generalized DSC [Sudre et al. 2017].

Metric combination

A single metric typically does not reflect all aspects that are essential for algorithm validation. Hence, multiple metrics with different properties are often combined. However, the selection of metrics should be well-considered as some metrics are mathematically related to each other [Taha et al. 2014, Taha and Hanbury 2015]. A prominent example is the IoU – the most popular segmentation metric in computer vision – which highly correlates with the DSC – the most popular segmentation metric in medical image analysis. In fact, the IoU and the DSC are mathematically related (see Sec. 2.2) [Taha and Hanbury 2015].

Combining metrics that are related will not provide additional information for a ranking. Figure 88 illustrates how the ranking can change when adding a metric that measures different properties.

Ranking uncertainty

Rankings themselves may be unstable. [Maier-Hein et al. 2018] and [Wiesenfarth et al. 2021] demonstrated that rankings are highly sensitive to altering the metric aggregation operators, the underlying data set, or the general ranking method. Disregarding the robustness of rankings may thus lead to the winning algorithm ranking first solely by chance rather than true superiority. Such a case is illustrated in Fig. 89. Here, the boxplots of two different benchmarking experiments are presented, one of which represents a very clear ranking, the other showing a similar performance of all five algorithms. However, both scenarios would yield the same ranking table. The common practice of solely presenting ranking tables without further visualization [Wiesenfarth et al. 2021] may thus obscure important information.

Insufficient biomedical relevance of metric score differences

Rankings may be uninformative in the domain context. For example, algorithms A1 and A2 in Fig. 90 only differ by a very small amount (0.0001). While this difference would make one algorithm numerically superior, it may not be relevant for the actual application. Therefore, the algorithms should not receive different ranks, but rather share the same rank. On the other hand, algorithms A2 and A3 differ by a substantial amount, thus, ranking them differently is clinically relevant.

Non-determinism of algorithms

AI (AI) algorithms are subject to non-determinism. When training an algorithm under identical conditions (e.g., the same libraries and architecture), results will be subject to sometimes substantial changes if the random seeds are exchanged, as shown in Fig. 12 (left). Although fixing random seeds will reduce the variability of results, there may still be differences in metric scores, for example caused by parallel processing [Pham et al. 2020] or usage of more than two GPU.

Non-standardized metric implementation

Metric implementation is typically not standardized, as can for instance be seen in Figures 78 and 82. This was further highlighted in a study presented in [Gooding et al. 2022], in which the authors analyzed the differences between how metric implementations differ and their implications. The authors found, for example, that seven out of thirteen distinct methods were used for sampling the surface for distance-based metrics such as HD or that the thirteen participants of the study used five different definitions for the symmetric average distance [Gooding et al. 2022]. It was further shown that these differences result in a high variation in the final metric score, for instance of up to 9% in the DSC and of up to 40% for the HD metric [Gooding et al. 2022].

Conclusion

Choosing the right metric for a specific image processing task is a nontrivial undertaking. With this (dynamic) paper, we wish to raise awareness about some of the common flaws of the most frequently used reference-based validation metrics in the field of image processing and provide guidance of their use, encouraging researchers to reconsider common workflows.

Acknowledgements

This work was initiated by the Helmholtz Association of German Research Centers in the scope of the Helmholtz Imaging Incubator (HI), the MICCAI Special Interest Group on biomedical image analysis challenges and the benchmarking working group of the MONAI initiative. It received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation programme (grant agreement No. , NEURAL SPICING). It was further supported in part by the Intramural Research Program of the National Institutes of Health (NIH) Clinical Center as well as by the National Cancer Institute (NCI) and the National Institute of Neurological Disorders and Stroke (NINDS) of the NIH, under award numbers NCI:U01CA242871 and NINDS:R01NS042645. The content of this publication is solely the responsibility of the authors and does not represent the official views of the NIH. T.A. acknowledges the Canada Institute for Advanced Research (CIFAR) AI Chairs program, the Natural Sciences and Engineering Research Council of Canada. F.B. was co-funded by the European Union (ERC, TAIPO, 101088594). Views and opinions expressed are however those of the authors only and do not necessarily reflect those of the European Union or the European Research Council. Neither the European Union nor the granting authority can be held responsible for them. V.C. acknowledges funding from NovoNordisk Foundation (NNF21OC0068816) and Independent Research Council Denmark (1134-00017B). B.A.C. was supported by NIH grant P41 GM135019 and grant 2020-225720 from the Chan Zuckerberg Initiative DAF, an advised fund of the Silicon Valley Community Foundation. G.S.C. was supported by Cancer Research UK (programme grant: C49297/A27294). M.M.H. is supported by the Natural Sciences and Engineering Research Council of Canada (RGPIN-2022-05134). A.Kara. is supported by French State Funds managed by the “Agence Nationale de la Recherche (ANR)” - “Investissements d’Avenir” (Investments for the Future), Grant ANR-10-IAHU-02 (IHU Strasbourg). M.K. was supported by the Ministry of Education, Youth and Sports of the Czech Republic (Project LM2018129). T.K. was supported in part by 4UH3-CA225021-03, 1U24CA180924-01A1, 3U24CA215109-02, and 1UG3-CA225-021-01 grants from the National Institutes of Health. G.L. receives research funding from the Dutch Research Council, the Dutch Cancer Association, HealthHolland, the European Research Council, the European Union, and the Innovative Medicine Initiative. S.M.R. wishes to acknowledge the Allen Institute for Cell Science founder Paul G. Allen for his vision, encouragement and support. C.H.S. is supported by an Alzheimer’s Society Junior Fellowship (AS-JF-17-011). M.R is supported by Innosuisse grant number 31274.1 and Swiss National Science Foundation Grant Number 205320_212939. R.M.S. is supported by the Intramural Research Program of the NIH Clinical Center. A.T. acknowledges support from Academy of Finland (Profi6 336449 funding program), University of Oulu strategic funding, Finnish Foundation for Cardiovascular Research, Wellbeing Services County of North Ostrobothnia (VTR project K62716), and Terttu foundation. S.A.T. acknowledges the support of Canon Medical and the Royal Academy of Engineering and the Research Chairs and Senior Research Fellowships scheme (grant RCSRF1819\8\25). B.V.C. was supported by Research Foundation Flanders (FWO grant G097322N) and Internal Funds KU Leuven (grant C24M/20/064).

We would like to thank Amine Yamlahi, Niklas Holzwarth, Marco Hübner, Amith Kamath, Dominik Michael, You Suhang, and Yannick Suter for proofreading the document and proposing pitfalls.

References

Appendix

Appendix A Acronyms