Measuring Discrimination to Boost Comparative Testing for Multiple Deep Learning Models
Linghan Meng, Yanhui Li, Lin Chen, Zhi Wang, Di Wu, Yuming Zhou, Baowen Xu
I Introduction
Deep learning (DL) supports a general-purpose learning procedure that discovers high-level representations of input samples with multiple layers of abstraction based on artificial neural networks (ANNs), which has shown significant advantages in establishing intricate structures of high dimensional data when tackling complex classification tasks . Along with increases in computation power and data size , DL technology achieves great success in constructing deeper layers of more effective abstraction to enhance classification performance, and has beaten human experts and traditional machine-learning technology in many areas , including image recognition , speech recognition , autonomous driving , playing Go , and so on. Meanwhile, concern about the reliability of DL models has been raised, which calls for novel testing techniques to deal with new DL testing scenarios and challenges.
Most current DL testing techniques try to validate the quality of DL models in two testing scenarios: debug testing and operational testing . On the one hand, debug testing considers DL testing as a technology to improve reliability by finding faultsThe faults of DL models are usually considered as the mismatching between the real labels and predicted labels of the input samples. , where various testing criteria (e.g., Neuron Activation Coverage and Neuron Boundary Coverage ) have been proposed to generate or select error-inducing inputs which trigger faults. On the other hand, operational testing aims to make reliability assessment for DL models in the objective testing contexts. Li et al. proposed an effective operational testing technique to estimate the accuracy of a single DL model by constructing probabilistic models for the distribution of testing contexts .
The boom of DL technology leads to DL models with ever-increasing functionality scale and complexity, i.e., complex DL models combine multi-function from multiple primitive DL models. Exposing code and data to build models and sharing model files (e.g., h5 files) boost the acquisition of DL models, which drive developers to build complex models by reusing available DL models achieving specific primitive functionality. One statistic in the previous study indicates that more than 13.7% of complex DL models on Github reuse at least one primitive DL model. On the positive side, this “plug-and-play” pattern has greatly facilitated the construction and application of complex DL models. On the negative side, for a given DL task, it is tough to select suitable models because of the advent of numerous DL models constructed by mass developers. These multiple models are produced by third-part developers and are trained on samples with different distributions. Therefore, their actual performance on the target application domain is not guaranteed, and they are needed to be tested.
These above points expedite the emergence of a new testing scenario “comparative testing”, where testers may encounter multiple DL models with the same functionality built by different developers, all of which are considered as candidates to accomplish a specific task, and testers are expected to rank them to choose the more suitable models in the testing contexts. Generally speaking, comparative testing is different from current DL testing, i.e., debug and operational testing, in the following two points:
The testing object is multiple DL models instead of a single DL model;
The testing aim is comparing performances among multiple DL models instead of improving/assessing performances for a single DL model.
Figure 1 shows an example of comparative testing scenarios considering multiple real world DL models available on GitHub. Hypothetically, in this scenario, the target application requires an implementation of written digit identification, and multiple candidate DL models are found to achieve this functionality. Testers are expected to compare the accuracy of written digit identification among these models and choose the more suitable ones to meet the requirements. As stated in many previous studies , sample labeling is the bottle neck of testing resources for DL models, which spends much manpower and is time-consuming. Due to the limitation of labeling effort, testers can label only a very small part from the whole testing contexts. Therefore, as shown in Figure 1, testers are asked to execute comparative testing by selecting and labeling a small but efficient subset of testing samples extracted from the testing context, and ranking multiple models based on their performance of the selected samples.
As mentioned above, comparative testing brings out a new problem of DL testing: given limited labeling effort, how to select an efficient subset of samples (label and test them) to rank multiple DL models as precise as possible? To tackle this problem, we propose a novel algorithm named Sample Discrimination based Selection (SDS) to measure the sample discrimination and select samples with higher discrimination. The main idea of our algorithm is to focus on efficient samples that could discriminate multiple models, i.e., the prediction behaviors (right/wrong) of these samples would be helpful to indicate the trend of model performance. Specifically, SDS combines two aspects of technical thoughts: majority voting in ensemble learning and item discrimination in test analysis, which are introduced to estimate the sample discrimination with the lack of actual labels (details are in Section III).
We evaluate our approach on three widely-used image datasets MNIST , Fashion-MNIST , and CIFAR-10 , each of which contains 10000 testing samples. To simulate the comparative testing scenarios where multiple DL modes are developed/submitted for the same task (e.g., digital identification with MNIST and clothing classification with Fashion-MNIST), we introduce totally 80 models from GitHub, including 28 models for MNIST, 25 for Fashion-MNIST, and 27 for CIFAR-10. To assess the performance of SDS, we introduce three sample selection methods as the baselines: one state-of-the-art method from debug testing (DeepGini at ISSTA’2020 ), one state-of-the-art method from operational testing (CES at FSE’2019 ) and the simple random selection (SRS). The experimental results indicate that our algorithm SDS is an effective and efficient sample selection method for comparative testing to solve the problem “ranking multiple DL models under limited labeling efforts”.
Our study makes the following contributions:
Dimension. This study opens a new dimension of DL testing “comparative testing” for DL models, which focuses on comparing multiple DL models instead of improving/assessing a single DL model.
Strategy. This paper proposes a novel selection method SDS to measure the discrimination of samples and select samples with higher discrimination to rank multiple DL models.
Study. This paper contains an extensive empirical study of 80 models with three datasets containing 10000 testing inputs. The experimental results indicate that compared with the baseline methods, SDS is an effective and efficient sample selection method for comparative testing.
The rest of this paper is organized as follows. In Section II, we introduce a motivation example to show the difference between comparative testing and debug/operational testing. In Section III, we present a detailed description of our algorithm SDS. In Section IV, we present our experimental settings, including studied datasets and models, baseline methods, research questions, and so on. Section V explains experimental results and discoveries. Section VI further discusses some important experimental details. Sections VII and VIII are threats to validity and related works, respectively. Section IX presents the conclusion of our paper.
II The Motivation Example
As we mentioned in Section I, the aim of comparative testing is comparing the performances of multiple models. Here we introduce an example to show the differences between comparative testing and debug/operational testing.
Figure 2 presents an example of comparative testing scenarios containing six testing samples and three DL models , with the prediction results of samples predicted by models. ✓/ indicates that the prediction results of these models running against samples are right/wrong (i.e., the predicted labels are identical/different with the actual ones). By calculating the numbers of ✓/, we can obtain the accuracies of three models, i.e., for , for , and for , respectively. As a result, the actual rank of accuracies () for these models is . As shown in Figure 2, we have the following observations:
As only six samples are considered, we can easily find that the most efficient subset to indicate the actual rank of these models is : has two ✓under , has one, and has none. We can obtain the same rank of models for the accuracies () w.r.t. : .
is not the target sampling subset in operational testing, as it assesses model performance imprecisely: the accuracy under is //, which is much different from the actual //.
is also not the target sampling subset in debug testing. Debug testing would consider with the highest priority, since it triggers the mismatching behaviors of all models.
These observations indicate that the differences of aims between comparative testing and debug/operational testing lead to the different sampling priority. In comparative testing, we focus on the samples that could discriminate multiple models, e.g., and in Figure 2. In the next section, we will introduce a novel algorithm to measure sample discrimination and select samples with higher discrimination.
III Methodology
In this section, we present the detailed description of our approach. First, we present the studied problem. Next, we show an algorithm named Sample Discrimination based Selection (SDS) to measure the sample discrimination and select samples with higher discrimination.
We first introduce some symbols and definitions, which are helpful for readers to understand the rest of our paper.
Definition 1 (DL models). A DL model is usually regarded as an implementation of complex classification task based on the layer structure of artificial neural networks, which achieves a function mapping the high dimensional samples (e.g., a gray value matrix for figures) to labels in a given label set : .
Definition 2 (Accuracy). A DL model is tested under the testing context containing samples . Let and be the predicted label generated by and the actual label of , respectively. The accuracy of w.r.t. is defined as follows:
We introduce accuracy as the main indicator to measure the performance for comparing multiple DL models, as it has been widely used in evaluating the performance of DL models . Based on above definitions and symbols, we present the studied problem “given limited labeling effort, for multiple DL models, tester aim to select an efficient subset of samples (label and test them) to rank these models as precise as possible” specifically:
Problem. are tested under the testing context . and all samples in are unlabeled. Given limited labeling effort (), the task is to select and label an efficient subset () from , and employ the results (i.e., on to estimate the rank of model performance (i.e., ) on the whole testing context , with an as small rank error as possible.
III-B Sample Discrimination based Selection
As shown in the motivation example, comparative testing need samples that could discriminate the multiple models. In this subsection, we propose a novel algorithm named Sample Discrimination based Selection (SDS) to measure the sample discrimination and select samples with higher discrimination. Generally, SDS combines two aspects of technical thoughts:
Majority voting . Majority voting is a simple weighting method in ensemble learning, which selects the class with the most votes as the final decision. As our algorithm has the precondition that all samples are unlabeled, we employ majority voting as a procedure to deal with the lack of actual labels, i.e., for a given sample, we choose the predicted label with the most models as the estimation of the actual label.
Item discrimination . Item discrimination is an indicator to describe to what extent test items can discriminate between good and poor students, which is widely used in test analysishttps://www.medsci.ox.ac.uk/divisional-services/support-services-1/learning-technologies/faqs/what-do-difficulty-correlation-discrimination-etc-in-the-question-analysis-mean. We introduce the idea of item discrimination to measure sample discrimination, i.e., estimate discrimination by calculating the difference performance between good and bad models under each sample.
Specifically, given multiple DL models , the testing context with unlabeled samples , the label set , and the labeling effort , SDS is composed of the following five steps, as shown in Algorithm 1.
Step 1: Extract prediction results. We run multiple DL models against the testing context (line 9). For model and sample , we record the predicted label in the element of the prediction matrix (line 10).
Step 2: Vote for estimated labels. For any sample , we compute the frequency of predicted labels created by multiple models (line 13). We choose the predicted label with the max frequency, i.e., majority voting, as the estimated label (line 14), which is the basic of the following steps.
Step 3: Classify top/bottom models. We employ the voted labels to score the predicted results of DL models on samples one by one, if the predicted label equals to the voted label, we add one score for the current model (line 19). After we go through all the samples, we obtain an estimated score for this model. We sort models in descending order by their estimated score (line 21). According to the classification in , we classify models into three classes (line 22): top class containing the top 27% models, bottom class containing the bottom 27% models, and other class containing other models.
Step 4: Compute sample discrimination. We employ difference performance of models in top/bottom class to calculate discrimination. Specifically, for each sample , the value of discrimination is the number of models with right prediction in top class minus the number in bottom class (line 23-31). Intuitively, if the number in top class is much larger than the number in bottom class, the result of this sample is more identical with the rank, i.e., it would be helpful to estimate the rank of model performance. Finally, we normalize and store the sample discrimination (line 32).
Step 5: We consider the samples with higher discrimination as the ones which are more helpful to rank multiple DL models. To eliminate the effects of outlier samples with higher discrimination, we introduce random selecting instead of direct selecting from higher discrimination to lower discrimination. Specifically, we choose 25% as the cutoff point to construct the subset of samples with higher discrimination since quartering is common for dataset partition in software engineering , i.e., we consider the top 25% samples as the candidates (line 34) and randomly select samples from them according to the given labeling effort (line 35).
Figure 3 shows an example of SDS running on four DL models , …, with the testing context containing four samples , which are classified into three classes , , and . Four subfigures show the running results of the first four stepsAs step 5 is easy to understand, we omit its running here. of SDS, respectively, where the entries with a gray background indicates the target information obtained in each step. Next, we describe the subfigures one by one.
Figure 3(a) shows that SDS constructs the prediction matrix, where , , and are the predicted labels.
Figure 3(b) presents that SDS employs majority voting to obtain the estimation of actual labels. For example, for , three models predict it as , and one as . Therefore, SDS adds as its estimated label.
Figure 3(c) presents that SDS estimates the scores of models based on estimated labels, e.g., since have three right and one wrong prediction, is scored 3; and SDS classify into the top class and into the bottom class.
Figure 3(d) shows SDS counts the number of models with right prediction in top class minus the number in bottom class, e.g., for , both and predict right, the discrimination of is .
IV EXPERIMENTAL SETUPS
In this section, we present the experimental setup to evaluate the performance of SDS.
We introduce three widely used datasets MNIST , Fashion-MNIST , and CIFAR-10 to conduct our experiments. MNIST is a dataset of handwritten digit images with 60000 training samples and 10000 testing samples. Samples in MNIST are pixel grayscale images to denote handwritten digits from 0 to 9. Fashion-MNIST is similar with MNIST, containing 60000 training samples and 10000 testing samples which are pixel grayscale images to describe ten types of clothing. CIFAR-10 contains 60000 pixel color images (50000 for training and 10000 for testing), which are equally distributed into 10 classes, e.g., cat, dog, ship, and truck. In summary, each dataset supports 10000 testing samples, which are considered as the testing context in the following experiment.
For these three datasets, we extract a large amount of (80) models on Github, 28 models for MNIST, 25 for Fashion-MNIST, and 27 for CIFAR-10, respectively. To simulate the different implements of the same tasks, we choose these DL models with different stars (from a few to tens of thousands) on Github, different model structures, and different accuracies. For each model, if the model files (e.g., saved as h5 file) are provided in the repository on GitHub, we reuse them directly; otherwise, we employ the code and data provided to train the studied models. Table I presents the detailed description of these studied models with their GitHub repositories, parameters of model structure, and actual accuracies in the testing context. As shown in Table I, some of studied models come from the same repository, we put them together and provide the minimum and maximum values of layers, parameters, and accuracies among them.
IV-B Experimental Settings
This section will describe some details in the experimental settings in the following aspects.
Sampling Size. As mentioned above, the labeling effort is the bottle neck, i.e., tester are limited to label only the very small percent of testing samples. Following the experiment design in , we focus on results on each sample size from 35 to 180 with intervals of 5 (i.e., 35, 40, …, 180), which are 0.35%-1.8% samples selected from the whole testing context.
Baseline Method. Given multiple DL models, our goal is to rank the performance of these models by selecting and labeling a discriminative subset. It’s worth pointing out that, as comparative testing is a new testing scenario proposed in this paper, there are not existing baseline methods. To clarify the performance of our method, we conduct comparative experiments with three baselines: two state-of-the-art sample selection methods (CES at FSE’2019 and DeepGini at ISSTA’2020 ) in current DL testing and random selection.
CES: Li et al. proposed an effective method named CES to select samples for DL testing to assess the accuracy of the single DL model. We choose CES as a baseline as it also aims to select representative subsets of sample and reduce the labeling costs. Since CES runs based on the single model, given DL models, CES may construct selections of samples for models, respectively. Here, we introduce the best of selections (i.e., choose the subset that gets the highest performance of ranking) as the result generated by CES, which is a stronger baseline to show the advance of our method.
DeepGini: Feng et al. proposed a technique called DeepGini to help prioritize testing DL models, which measures the likelihood of misclassification by calculating the set impurity of prediction probabilities for multiple classification. DeepGini supports a deterministic baseline method, i.e., it sorts the test samples according to the calculated likelihood and selects the samples according to sampling size. As SDS and CES are with randomness, we combine random selection and DeepGini to construct a new baseline, in which we perform random sampling in the first 25% (the same cutoffs in SDS) samples according to the rank of . To differentiate these two baseline methods, we call the the former deterministic DeepGini (DDG), and the latter random DeepGini (RDG).
SRS: Simple Random Sampling (SRS) is a basic method for subset selection, which is used as baseline for many studies . We randomly select a subset from the testing set and test the ranking performance of this subset.
We implement SDS and baseline methods in python 3.6.3 with the frameworks including Tensorflow 2.3.0 and Keras 2.4.3. Our experiments are performed on a Ubuntu 18.04 server with 8 GPU cores “Tesla V100 SXM2 32GB”. We provide the replication package including the detailed description of our proposed methods SDS and source code online (see Section X).
Repetition. As SDS and several baseline methods are with randomness, we conduct the experiment 50 times and report the average of calculated results.
IV-C Evaluation Indicators
To evaluate to what extend the estimated rank w.r.t. selected samples are identical with the actual rank w.r.t. the whole testing context, we introduce two indicators Spearman’s rank correlation coefficient and Jaccard similarity coefficient.
Spearman’s rank correlation coefficient is a measure of the correlation between two variables and . It can be calculated by the following formula:
The value of ranges from -1 to 1: the closer it is to 1 (-1), the two sets of variables are positively (negatively) correlated.
Besides, we introduce Jaccard similarity coefficient (denoted as ) to evaluate the similarity between the top- model sets generated by the estimated rank and actual rank. For example, the estimated rank is , and the actual rank is . The Jaccard similarity coefficient between two top-3 model sets and is calculated as:
As we encounter dozens of models in the testing context (e.g., 28 models for MNIST), we focus on to evaluate the performance of our method on different cutoff points. We take as the representative to report the evaluation under Jaccard similarity coefficient in the experimental results. The other Jaccard coefficient (when ) will be discussed in the discussion part.
IV-D Analysis Method
First, we employ Wilcoxon rank sum test to verify the difference of the rank performance between our method and the baselines. If the -value are less than 0.05, the two sets of data are considered significantly different.
Next, we introduce Cliff’s delta , which measures the effect size for comparing two ordinal data lists. We judge the difference between the two sets of data based on the range of : negligible, if ; small, if ; medium, if , and large, if .
Finally, we use “W/T/L” to compare the results of our approach and the baseline, where “W” means our approach wins, “T” means the results are tie, and “L” means our approach loses. Reaching the two standards shows that our approach wins: (a) the -value of Wilcoxon rank sum test is less than 0.05 (), which means the results between our approach and baseline are significantly different; (b) the Cliff’s delta is larger than 0.147 (), which means the difference between the two results are positive and not negligible. If and , we consider our approach loses. Otherwise, the result of comparision is tie.
IV-E Research Questions
We are committed to promoting the ranking performance of the multiple models under limited labeling effort in the comparative testing scenario. We propose the following two research questions (RQs) to organize our experiments:
RQ1 (Effectiveness): Whether our method SDS can surpass the state-of-the-art methods in ranking multiple models?
RQ2 (Efficiency): Compared with the state-of-the-art methods, is our method SDS efficient?
V EXPERIMENT RESULTS
In this section, we present the results of the experiments and answer the above two RQs.
Motivation and Approach. Our problem is to obtain effective ranking results of model performance with a very low labeling percent of testing samples for multiple models in the testing scenario. We hope to verify whether our proposed approach SDS is more effective than the baseline methods with limited labelling effort. To achieve this aim, we compare the ranking performance of SDS with the baseline methods in three testing contexts (containing the 10000 testing samples from MNIST, Fashion-MNIST, and CIFAR-10, respectively) under the sampling sizes from 35 to 180 with intervals of 5 (i.e., 35, 40, …, 180). Specifically, we employ these five studied methods to sample the subset under different sampling sizes, and use the ranking result on the subset to estimate the rank on the whole testing context. We repeat running these methods 50 times, and report the average of calculated results.
Results. Figure 4 shows the comparison results of our approach and the four baselines for ranking model performance. The three subgraphs in the first row show the comparison results of Spearman coefficient , and the subgraphs in the second row present the results of Jaccard coefficientWe take as the representative to report the evaluation under Jaccard similarity coefficient in the experimental results. The other Jaccard coefficient (when ) will be discussed in the discussion part. (). In each subfigure, the -axis indicates the number of samples, from 35 to 180, and the -axis represents the values of Spearman/Jaccard coefficient. These five studied methods are denoted by lines with different colors, i.e., for SDS, for CES, for DDG, for RDG, and for SRS. It can be seen in Figure 4 that in all sub-graphs, our method SDS is obviously better than the other baselines under all sampling sizes from 35 to 180. Besides, our approach is very stable; on the contrary, some baselines have a strong volatility, e.g., DDG has wild gyrations when measuring Jaccard coefficient for MINIST, Fashion-MINIST, and CIFAR-10.
In order to show more details of the experiment results, we choose six sampling points (35, 60, 90, 120, 150, and 180) as the representatives. Table II presents the detailed results under these six points, with the mean values of Spearman coefficient and Jaccard coefficient of ranking multiple models obtained by 50 repetitions of running the five studied methods. The best numbers are highlighted in bold. In Table II, if our approach wins the baseline method (that is to say, the value is less than 0.05 and the is greater than 0.147), then we add the gray background to the value of the baseline method. Based on the values of Spearman and Jaccard coefficients, we have added two rows: “Average” to calculate the average value of each column and “W/T/L” to record the number of times our approach win/tie/lose other baselines.
From Table II, we have the following observations. (a) From the average value of each point, our approach is higher than all baselines, under both Spearman coefficient and Jaccard coefficient. (b) The gray background indicates that our approach wins other baselines at the most of points. (c) The results of W/T/L shows that our approach is not only higher than other baselines in mean, but also significantly better.
Answer to RQ1: In ranking multiple DL models, our approach is significantly better than all other baselines in effectiveness.
V-B RQ2: Efficiency
Motivation and Approach. In RQ1, we have observed that our approach SDS is significantly better than other baselines under both Spearman coefficient and Jaccard coefficient in raking multiple DL models. The process of sample selection may be time consuming. In this RQ, we want to check the efficiency of our approach compared with other baselines.
Results. Table III shows the total time consumed when running studied methods with sampling from 35 to 180. From Table 3, we find that our approach SDS takes longer than SRS because it contains sample sorting and operations on the prediction matrix. The time SDS consumes is similar to the other three baselines CES, RDG, and DDG, which is around 10,000 seconds.
Answer to RQ2: Except for SRS, our approach is similar to other baselines in time consumption.
VI Discussion
In this section, we further discuss some parameter settings and results in the experiments. First, we analyze the parameter and indicator involved in the experiments. After that, we discuss why our algorithm can effectively help multi-model performance ranking and whether our method is effective when the number of models is reduced.
In our experiment, the random sampling interval is set to the top 25% as shown in Step 5 of Algorithm 1. We want to further discuss the performance under other selection rates by conducting experiments on five different selection rates (i.e., random sampling of the top 15%, 20%, 25%, 30%, and 35% intervals). The results are shown in Figure 5, where the performances of different rates are denoted by line with different colors, i.e., for 15%, for 20%, for 25%, for 30%, and for 35%, respectively.
It can be seen that the performances of different selection rates on different datasets vary a lot. Generally speaking, there is no obvious trend in all subfigures. In addition, the 25% sampling interval (the green line) we set in the experiment obtains the best ranking performance under the most of sampling sizes in the CIFAR-10 dataset. As quartering is common for dataset partition in software engineering and easy to implement, we still suggest applying the 25% interval in our algorithm.
VI-B The performance of Jaccard coefficient with k=1,3,5𝑘135k=1,3,5
We employ Jaccard coefficient to measure the similarity between the two top- model sets generated by the selected subset and the whole testing context, respectively. In the previous experiments, when we use the Jaccard coefficient, we calculate it with . In this section, we will discuss whether our method has advantages when the values of are different, i.e., . Due to space limitation, we cannot display all the 33 (the former 3 for the three datasets and the latter 3 for ) subgraphs, we calculate the average of the three datasets in three subgraphs for in Figure 6.
As shown in Figure 6, we compare our approach SDS (the green line) with other baselines when . Figure 6 shows the average values of the Jaccard coefficient of the three datasets when the sampling changes. It can be seen that our approach still has advantages under the most of points, which shows that our approach is still superior in ranking models when considering .
VI-C Analysis and Insight of our algorithm
In this section we will discuss why our algorithm works. In order to illustrate this point, we conduct a two-step analysis. The first is to measure the precision of the majority voting. We compare estimated labels obtained by the majority voting with true labels. Figure 7 shows the matched rate of estimated labels with true labels when the majority voting gets different numbers of votes. It can be seen that as the number of votes obtained increases, the matched rate also rises. In general, the average matched rate of majority voting results with the true labels reaches 0.9924 for MNIST, 0.9433 for Fahion-MNIST, and 0.8613 for CIFAR10, respectively, as shown by the red line in each subfigure. In other words, majority voting is close to the true label, which is the key to explain why our method is effective. This finding leads to an insight for following studies in comparative testing: the distribution of predicted labels would be helpful to deal with the lack of actual labels, which is a main difficulty in actual testing scenarios due to the limitation of labelling effort. We encourage following researchers to employ more effective methods to measure the distribution in comparative testing.
In the second step, we analyze whether the sample discrimination is positively correlated to the ranking performance, i.e., whether higher discrimination is more helpful for ranking multiple DL models. We conduct an additional experiment. After sorting the samples according to the discrimination, we randomly select samples in the top 25%, the 25%-50%, the 50%-75%, and the 75%-100% intervals to observe the results of ranking performance. We take the averages of the three datasets and show them in the Figure 8, the blue line represents the random sampling in the first 25% interval, which is the interval used in our experiment. It can be seen that the model ranking effectiveness of random sampling in the first 25% is significantly better than other intervals. That indicates higher discrimination is more helpful for ranking multiple DL models.
VI-D The performance when there are fewer models
The previous experiment content is to calculate the ranking performance of the SDS method when the number of models is large (i.e., more than 20 models for a given task). In this section, we report the performance of SDS on the model ranking when there are few models. We have selected four models in each data set to compare the ranking effect of SDS and other baselines. We measure the Spearman coefficientAs there are only four models, Jaccard coefficient () is not applicable here. We focus on Spearman coefficient. value when the sample size is from 35 to 180, the experiment was repeated 50 times, and the average results were reported. Figure 9 shows the comparison results of SDS, SRS, and CES, which are the best three methods when the number of models is large. Figure 9 presents the average ranking performance on the three data sets, where the green curve denotes the Spearman coefficient of SDS.
We observe that SDS can still show superior performance when there are fewer models, which obviously exceeds SRS and CES. To some extent, the above result shows the generalization of the SDS method.
VI-E The ranking performance when choosing majority voting as true labels
An intuitive idea is to use the labels obtained by the majority voting as the true labels to measure the accuracy of the models, and then get the ranking performance of the models (see line 21 of Algorithm 1). In this section, we compare this intuitive method with the results of SDS to verify whether the calculation after line 21 in Algorithm 1 really plays a role in model ranking.
We show the results of the comparison in Figure 10. The blue curve in the figure is the average result of the spearman coefficient on the three data sets that vary with the sample size, and the red line is the ranking result obtained by using the majority voting results as the true labels, which is also the average result of the three data sets.
It can be seen from Figure 10 that SDS overcomes majority voting when the sample size is larger than 105, which is roughly about one percent (i.e., 10000*1%=100) of the total test set. Besides, the curve after this point still shows an upward trend along with the increasing of sampling size. That is to say, the calculation content after line 21 in Algorithm 1 is useful for the model ranking.
VII THREAD TO VALIDITY
The threat to validity is discussed in the following three aspects for our study.
First, the datasets we select may be a threat. We use three well-known graph classification datasets, which are widely used in many studies, but their complexity is not high. In the future, we will explore on larger and more diverse datasets to validate the effectiveness of our algorithm.
Second, the selection of models in the experiments could become a threat. We try to choose a wide range of models on GitHub, i.e., 28 models for MNIST, 25 for Fashion-MNIST, and 27 for CIFAR-10, respectively, which include multiple DL models with different stars (from a few to tens of thousands) on Github, different model structures, and different accuracies. However, these studied 80 models may not fully cover the real situation. More models are expected in the following studies to validate our results.
Finally, it may also be a threat to the implementation of the models. As discussed earlier, if the trained model file is provided in the GitHub repository, we will use it directly, otherwise we will use the provided python code and datasets for training. Due to the difference in the training environment, it may cause the reproduced model to be different from the original one. For new trained models, we compare the accuracies announced in the GitHub repository and actual accuracies, and find that the difference between them is slight.
VIII RELATEDWORK
In this section, we introduce the related work. In the angel of traditional software testing , on the one side, testing aims to find more bugs, which is called debug testing; on the other side, testing aims to make reliability assessment of software through conditioning, which is called operational testing.
The main body of current DL testing is to focus on debug testing, i.e., the main aim is to find bugs. Pei et al. proposed a whitebox framework named DeepXplore, which uses neuron coverage as the standard for DL model testing . Tian et al. implemented a tool named DeepTest to simulate the real world to help find behaviors that may cause accidents for DNN-driven vehicles . Zhang et al. proposed unsupervised framework for DNN named DeepRoad, and utilized GANs and metamorphic testing to test the inconsistent behaviors in self-driving car . Xie et al. proposed a coverage-based framework named DeepHunter which used metamorphic mutation to help find defects for DNNs . Sun et al. presented an approach named TransRepair to help machine translation systems test and repair inconsistency bugs . Ma et al. proposed a set of testing criteria named DeepGauge for measuring the testing adequacy of DNNs . Ma et al. proposed DeepCT, which applied the idea of combinatorial testing to DL testing, and produced a series of combinatorial testing criteria for DL systems . Tian et al. developed a technique called DeepInspect, which can detect the confusion and bias errors based class for image classification . Lee et al. presented a white-box testing approach named ADAPT, which used an adaptive neuron selection strategy to find adversarial inputs .
Meanwhile, researchers have focused on the other aspects of DL testing. Li et al. proposed an effective operational testing technique to estimate the accuracy of the DL model by constructing probabilistic models for the distribution of testing contexts . To evaluate the quality of test data, Ma et al. applied the mutation framework to DL systems, and proposed a technique named DeepMutation . Zhou et al. proposed a testing approach faced the systematic physical world called DeepBillboard, which is aimed to generate adversarial test more robust . Gerasimou et al. proposed a systematic testing approach named DeepImportance, which is mixed with an Importance-Driven (IDC) test adequacy criterion to support more robust DL systems .
IX CONCLUSION
The boom of DL technology leads to the reuse of DL models, which expedites the emergence of a new testing scenario comparative testing, where testers may encounter multiple DL models with the same functionality as candidates to accomplish a specific task, and testers are expected to rank them to choose the more suitable models in the testing contexts. Due to the limitation of labeling effort, this testing scenario brings out a new problem of DL testing: ranking multiple DL models under limited labeling efforts.
To tackle this problem, we propose a novel algorithm named Sample Discrimination based Selection (SDS) to measure the sample discrimination and select samples with higher discrimination. We evaluate our approach on three widely-used image datasets and 80 DL models. Our results lead us to conclude that SDS is an effective and efficient sample selection method for comparative testing to rank multiple DL models.
Finally, we would like to emphasize that we do not seek to claim the advantage of our method SDS. Instead, the key messages are that (a) a new testing scenario comparative testing is introduced by our paper, where the testing aims are much different with the current DL testing, i.e., debug/operational testing; (b) the new testing scenario brings out the new testing challenge ranking multiple DL models under limited labeling efforts; (c) our proposed method SDS leads to the insight which would be helpful for the following researchers.
X REPEATABILITY
We provide all datasets and code used to conduct this study at https://github.com/Testing-Multiple-DL-Models/SDS.
Acknowledgements
The work is supported by National Key R&D Program of China (Grant No. 2018YFB1003901) and the National Natural Science Foundation of China (Grant No. 61872177, 61832009, 61772259, 61772263, and 61932012). We thank the anonymous referees for their helpful comments on this paper.