Learning from Noisy Labels with Deep Neural Networks: A Survey

Hwanjun Song, Minseok Kim, Dongmin Park, Yooju Shin, Jae-Gil Lee

I Introduction

With the recent emergence of large-scale datasets, deep neural networks (DNNs) have exhibited impressive performance in numerous machine learning tasks, such as computer vision , information retrieval , and language processing . Their success is dependent on the availability of massive but carefully labeled data, which are expensive and time-consuming to obtain. Some non-expert sources, such as Amazon’s Mechanical Turk and the surrounding text of collected data, have been widely used to mitigate the high labeling cost; however, the use of these source often results in unreliable labels . In addition, data labels can be extremely complex even for experienced domain experts ; they can also be adversarially manipulated by a label-flipping attack . Such unreliable labels are called noisy labels because they may be corrupted from ground-truth labels. The ratio of corrupted labels in real-world datasets is reported to range from 8.0%8.0\% to 38.5%38.5\% .

In the presence of noisy labels, training DNNs is known to be susceptible to noisy labels because of the significant number of model parameters that render DNNs overfit to even corrupted labels with the capability of learning any complex function . Zhang et al. demonstrated that DNNs can easily fit an entire training dataset with any ratio of corrupted labels, which eventually resulted in poor generalizability on a test dataset. Unfortunately, popular regularization techniques, such as data augmentation , weight decay , dropout , and batch normalization have been applied extensively, but they do not completely overcome the overfitting issue by themselves. As shown in Figure 1, the gap in test accuracy between models trained on clean and noisy data remains significant even though all of the aforementioned regularization techniques are activated. Additionally, the accuracy drop with label noise is considered to be more harmful than with other noises, such as input noise . Hence, achieving a good generalization capability in the presence of noisy labels is a key challenge.

Several studies have been conducted to investigate supervised learning under noisy labels. Beyond conventional machine learning techniques , deep learning techniques have recently gained significant attention in the machine learning community. In this survey, we present the advances in recent deep learning techniques for overcoming noisy labels. We surveyed recent studies by recursively tracking relevant bibliographies in papers published at premier research conferences, such as CVPR, ICCV, NeurIPS, ICML, and ICLR. Although we attempted to comprehensively include all recent studies at the time of submission, some of them may not be included because of the quadratic increase in deep learning papers. The studies included were grouped into five categories, as shown in Figure 2 (see Section III for details).

Frénay and Verleysen discussed the potential negative consequence of learning from noisy labels and provided a comprehensive survey on noise-robust classification methods, focusing on conventional supervised approaches such as naïve Bayes and support vector machines. Furthermore, their survey included the definitions and sources of label noise as well as the taxonomy of label noise. Zhang et al. discussed another aspect of label noise in crowdsourced data annotated by non-experts and provided a thorough review of expectation-maximization (EM) algorithms that were proposed to improve the quality of crowdsourced labels. Meanwhile, Nigam et al. provided a brief introduction to deep learning algorithms that were proposed to manage noisy labels; however, the scope of these algorithms was limited to only two categories, i.e., the loss function and sample selection in Figure 2. Recently, Han et al. summarized the essential components of robust learning with noisy labels, but their categorization is totally different from ours in philosophy; we mainly focus on systematic methodological difference, whereas they rather focused on more general views, such as input data, objective functions, and optimization policies. Furthermore, this survey is the first to present a comprehensive methodological comparison of existing robust training approaches (see Tables II and III).

I-B Survey Scope

Robust training with DNNs becomes critical to guarantee the reliability of machine learning algorithms. In addition to label noise, two types of flawed training data have been actively studied by different communities . Adversarial learning is designed for small, worst-case perturbations of the inputs, so-called adversarial examples, which are maliciously constructed to deceive an already trained model into making errors . Meanwhile, data imputation primarily deals with missing inputs in training data, where missing values are estimated from the observed ones . Adversarial learning and data imputation are closely related to robust learning, but handling feature noise is beyond the scope of this survey—i.e., learning from noisy labels.

II Preliminaries

In this section, the problem statement for supervised learning with noisy labels is provided along with the taxonomy of label noise. Managing noisy labels is a long-standing issue; therefore, we review the basic conventional approaches and theoretical foundations underlying robust deep learning. Table I summarizes the notation frequently used in this study.

where η\eta is a learning rate specified.

Here, the risk minimization process is no longer noise-tolerant because of the loss computed by the noisy labels. DNNs can easily memorize corrupted labels and correspondingly degenerate their generalizations on unseen data . Hence, mitigating the adverse effects of noisy labels is essential to enable noise-tolerant training for deep learning.

II-B Taxonomy of Label Noise

This section presents the types of label noise that have been adopted to design robust training algorithms. Even if data labels are corrupted from ground-truth labels without any prior assumption, in essence, the corruption probability is affected by the dependency between data features or class labels. A detailed analysis of the taxonomy of label noise was provided by Frénay and Verleysen . Most existing algorithms dealt with instance-independent noise, but instance-dependent noise has not yet been extensively investigated owing to its complex modeling.

II-B2 Instance-dependent Label Noise

II-C Non-deep Learning Approaches

For decades, numerous methods have been proposed to manage noisy labels using conventional machine learning techniques. These methods can be categorized into four groups , as follows:

Data Cleaning: Training data are cleaned by excluding examples whose labels are likely to be corrupted. Bagging and boosting are used to filter out false-labeled examples to remove examples with higher weights because false-labeled examples tend to exhibit much higher weights than true-labeled examples . In addition, various methods, such as kk-nearest neighbor, outlier detection, and anomaly detection, have been widely exploited to exclude false-labeled examples from noisy training data . Nevertheless, this family of methods suffers from over-cleaning issue that overly removes even the true-labeled examples.

Surrogate Loss: Motivated by the noise-tolerance of the 0-1 loss function , many researchers have attempted to resolve its inherent limitations, such as computational hardness and non-convexity that render gradient methods unusable. Hence, several convex surrogate loss functions, which approximate the 0-1 loss function, have been proposed to train a specified classifier under the binary classification setting . However, these loss functions cannot support the multi-class classification task.

Probabilistic Method: Under the assumption that the distribution of features is helpful in solving the problem of learning from noisy labels , the confidence of each label is estimated by clustering and then used for a weighted training scheme . This confidence is also used to convert hard labels into soft labels to reflect the uncertainty of labels . In addition to these clustering approaches, several Bayesian methods have been proposed for graphical models such that they can benefit from using any type of prior information in the learning process . However, this family of methods may exacerbate the overfitting issue owing to the increased number of model parameters.

Model-based Method: As conventional models, such as the SVM and decision tree, are not robust to noisy labels, significant effort has been expended to improve the robustness of them. To develop a robust SVM model, misclassified examples during learning are penalized in the objective . In addition, several decision tree models are extended using new split criteria to solve the overfitting issue when the training data are not fully reliable . However, it is infeasible to apply their design principles to deep learning.

Meanwhile, deep learning is more susceptible to label noises than traditional machine learning owing to its high expressive power, as proven by many researchers . There has been significant effort to understand why noisy labels negatively affect the performance of DNNs . This theoretical understanding has led to the algorithmic design which achieves higher robustness than non-deep learning methods. A detailed analysis of theoretical understanding for robust deep learning was provided by Han et al. .

II-D Regression with Noisy Labels

Although regression predicts continuous values, regression and classification share the same concept of learning the mapping function from the input feature xx to the output label yy. Thus, many robust approaches for classification are easily extended to the regression problem with simple modification . Thus, in this survey, we focus on the classification setting for which most robust methods are defined.

III Deep Learning Approaches

According to our comprehensive survey, the robustness of deep learning can be enhanced in numerous approaches . Figure 3 shows an overview of recent research directions conducted by the machine learning community. All of them (i.e., §III-A – §III-E) focused on making a supervised learning process more robust to label noise:

(§III-A) Robust architecture: adding a noise adaptation layer at the top of an underlying DNN to learn label transition process or developing a dedicated architecture to reliably support more diverse types of label noise;

(§III-B) Robust regularization: enforcing a DNN to overfit less to false-labeled examples explicitly or implicitly;

(§III-C) Robust loss function: improving the loss function;

(§III-D) Loss adjustment: adjusting the loss value according to the confidence of a given loss (or label) by loss correction, loss reweighting, or label refurbishment;

(§III-E) Sample selection: identifying true-labeled examples from noisy training data via multi-network or multi-round learning.

Overall, we categorize all recent deep learning methods into five groups corresponding to popular research directions, as shown in Figure 3. In §III-D, meta learning is also discussed because it finds the optimal hyperparameters for loss reweighting. In §III-E, we discuss the recent efforts for combining sample selection with other orthogonal directions or semi-supervised learning toward the state-of-the-art performance.

Figure 2 illustrates the categorization of robust training methods using these five groups.

In numerous studies, architectural changes have been made to model the noise transition matrix of a noisy dataset . These changes include adding a noise adaptation layer at the top of the softmax layer and designing a new dedicated architecture. The resulting architectures yield improved generalization through the modification of the DNN output based on the estimated label transition probability.

From the view of training data, the noise process is modeled by discovering the underlying label transition pattern (i.e., the noise transition matrix T). Given an example xx, the noisy class posterior probability for an example xx is expressed by

Technical Detail: Webly learning first trains the base DNN only for easy examples retrieved by search engines; subsequently, the confusion matrix for all training examples is used as the initial weight W\mathcal{W} of the noise adaptation layer. It fine-tunes the entire model in an end-to-end manner for hard training examples. In contrast, the noise model initializes W\mathcal{W} to an identity matrix and adds a regularizer to force W\mathcal{W} to diffuse during DNN training. The dropout noise model applies dropout regularization to the adaptation layer, whose output is normalized by the softmax function to implicitly diffuse W\mathcal{W}. The s-model is similar to the dropout noise model but dropout is not applied. The c-model is an extension of the s-model that models the instance-dependent noise, which is more realistic than the symmetric and asymmetric noises. Meanwhile, NLNN adopts the EM algorithm to iterate the E-step to estimate the noise transition matrix and the M-step to back-propagate the DNN.

Remark: A common drawback of this family is their inability to identify false-labeled examples, treating all the examples equally. Thus, the estimation error for the transition matrix is generally large when only noisy training data is used or when the noise rate is high . Meanwhile, for the EM-based method, becoming stuck in local optima is inevitable, and high computational costs are incurred .

III-A2 Dedicated Architecture

Beyond the label-dependent label noise, several studies have been conducted to support more complex noise, leading to the design of dedicated architectures . They typically aimed at increasing the reliability of estimating the label transition probability to handle more complex and realistic label noise.

Technical Detail: Probabilistic noise modeling manages two independent networks, each of which is specialized to predict the noise type and label transition probability. Because an EM-based approach with random initialization is impractical for training the entire network, both networks are trained with massive noisy labeled data after the pre-training step with a small amount of clean data. Meanwhile, masking is a human-assisted approach to convey the human cognition of invalid label transitions. Considering that noisy labels are mainly from the interaction between humans and tasks, the invalid transition investigated by humans was leveraged to constrain the noise modeling process. Owing to the difficulty in specifying the explicit constraint, a variant of generative adversarial networks (GANs) was employed in this study. Recently, the contrastive-additive noise network was proposed to adjust incorrectly estimated label transition probabilities by introducing a new concept of quality embedding, which models the trustworthiness of noisy labels. RoG builds a simple yet robust generative classifier on top of any discriminative DNN pre-trained on noisy data.

Remark: Compared with the noise adaptation layer, this family of methods significantly improves the robustness to more diverse types of label noise, but it cannot be easily extended to other architectures in general.

III-B Robust Regularization

Regularization methods have been widely studied to improve the generalizability of a learned model in the machine learning community . By avoiding overfitting in model training, the robustness to label noise improves with widely-used regularization techniques such as data augmentation , weight decay , dropout , and batch normalization . These canonical regularization methods operate well on moderately noisy data, but they alone do not sufficiently improve the test accuracy; poor generalization could be obtained when the noise is heavy . Thus, more advanced regularization techniques have been recently proposed, which further improved robustness to label noise when used along with the canonical methods. The main advantage of this family is its flexibility in collaborating with other directions because it only requires simple modifications.

The regularization can be an explicit form that modifies the expected training loss, e.g., weight decay and dropout.

Technical Detail: Bilevel learning uses a clean validation dataset to regularize the overfitting of a model by introducing a bilevel optimization approach, which differs from the conventional one in that its regularization constraint is also an optimization problem. Overfitting is controlled by adjusting the weights on each mini-batch and selecting their values such that they minimize the error on the validation dataset. Meanwhile, annotator confusion assumes the existence of multiple annotators and introduces a regularized EM-based approach to model the label transition probability; its regularizer enables the estimated transition probability to converge to the true confusion matrix of the annotators. In contrast, pre-training empirically proves that fine-tuning on a pre-trained model provides a significant improvement in robustness compared with models trained from scratch; the universal representations of pre-training prevent the model parameters from being updated in the wrong direction by noisy labels. PHuber proposes a composite loss-based gradient clipping, which is a variation of standard gradient clipping for label noise robustness. Robust early-learning classifies critical parameters and non-critical parameters for fitting clean and noise labels, respectively. Then, it penalizes only the non-critical ones with a different update rule. ODLN leverages open-set auxiliary data and prevents the overfitting to noisy labels by assigning random labels to the open-set examples, which are uniformly sampled from the label set.

Remark: The explicit regularization often introduces sensitive model-dependent hyperparameters or requires deeper architectures to compensate for the reduced capacity, yet it can lead to significant performance gain if they are optimally tuned.

III-B2 Implicit Regularization

The regularization can also be an implicit form that gives the effect of stochasticity, e.g., data augmentation and mini-batch stochastic gradient descent.

Technical Detail: Adversarial training enhances the noise tolerance by encouraging the DNN to correctly classify both original inputs and hostilely perturbed ones. Label smoothing estimates the marginalized effect of label noise during training, thereby reducing overfitting by preventing the DNN from assigning a full probability to noisy training examples. Instead of the one-hot label, the noisy label is mixed with a uniform mixture over all possible labels,

where λ∈\lambda\in is the balance parameter between two examples. Thus, mixup extends the training distribution by updating the DNN for the constructed mini-batch.

Remark: The implicit regularization improves the generalization capability of the DNN without reducing the representational capacity. It also does not introduce sensitive model-dependent hyperparameters because it is applied to the training data. However, the extended feature or label space slows down the convergence of training.

III-C Robust Loss Function

Based on this theoretical foundation, researchers have attempted to design robust loss functions such that they achieve a small risk for unseen clean data even when noisy labels exist in the training data .

Technical Detail: Initially, Manwani and Sastry theoretically proved a sufficient condition for the loss function such that risk minimization with that function becomes noise-tolerant for binary classification. Subsequently, the sufficient condition was extended for multi-class classification using deep learning . Specifically, a loss function is defined to be noise-tolerant for a cc-class classification under symmetric noise if the function satisfies the noise rate τ<c−1c\tau<\frac{c-1}{c} and

where CC is a constant. This condition guarantees that the classifier trained on noisy data has the same misclassification probability as that trained on noise-free data under the specified assumption. An extension for multi-label classification was provided by Kumar et al. . Moreover, if RD(f∗)=0\mathcal{R}_{\mathcal{D}}(f^{*})=0, then the function is also noise-tolerant under an asymmetric noise, where f∗f^{*} is a global risk minimizer of RD\mathcal{R}_{\mathcal{D}}.

For the classification task, the categorical cross entropy (CCE) loss is the most widely used loss function owing to its fast convergence and high generalization capability. However, in the presence of noisy labels, the robust MAE showed that the mean absolute error (MAE) loss achieves better generalization than the CCE loss because only the MAE loss satisfies the aforementioned condition. A limitation of the MAE loss is that its generalization performance degrades significantly when complicated data are involved. Hence, the generalized cross entropy (GCE) was proposed to achieve the advantages of both MAE and CCE losses; the GCE loss is a more general class of noise-robust loss that encompasses both of them. Amid et al. extended the GCE loss by introducing two temperatures based on the Tsallis divergence. Bi-tempered loss introduces a proper unbiased generalization of the CE loss based on the Bregman divergence. In addition, inspired by the symmetricity of the Kullback-Leibler divergence, the symmetric cross entropy (SCE) was proposed by combining a noise tolerance term, namely reverse cross entropy loss, with the standard CCE loss.

Meanwhile, the curriculum loss (CL) is a surrogate loss of the 0-1 loss function; it provides a tight upper bound and can easily be extended to multi-class classification. The active passive loss (APL) is a combination of two types of robust loss functions, an active loss that maximizes the probability of belonging to the given class and a passive loss that minimizes the probability of belonging to other classes.

Remark: The robustness of these methods is theoretically supported well. However, they perform well only in simple cases, when learning is easy or the number of classes is small . Moreover, the modification of the loss function increases the training time for convergence .

III-D Loss Adjustment

Loss adjustment is effective for reducing the negative impact of noisy labels by adjusting the loss of all training examples before updating the DNN . The methods associated with it can be categorized into three groups depending on their adjustment philosophy: 1) loss correction that estimates the noise transition matrix to correct the forward or backward loss, 2) loss reweighting that imposes different importance to each example for a weighted training scheme, 3) label refurbishment that adjusts the loss using the refurbished label obtained from a convex combination of noisy and predicted labels, and 4) meta learning that automatically infers the optimal rule for loss adjustment. Unlike the robust loss function newly designed for robustness, this family of methods aims to make the traditional optimization process robust to label noise. Hence, in the middle of training, the update rule is adjusted such that the negative impact of label noise is minimized.

In general, loss adjustment allows for a full exploration of the training data by adjusting the loss of every example. However, the error incurred by false correction is accumulated, especially when the number of classes or the number of mislabeled examples is large .

Similar to the noise adaptation layer presented in Section III-A, this approach modifies the loss of each example by multiplying the estimated label transition probability by the output of a specified DNN. The main difference is that the learning of the transition probability is decoupled from that of the model.

where T^\hat{\text{T}} is the estimated noise transition matrix.

Furthermore, gold loss correction assumes the availability of clean validation data or anchor points for loss correction. Thus, a more accurate transition matrix is obtained by using them as additional information, which further improves the robustness of the loss correction. Recently, T-Revision provides a solution that can infer the transition matrix without anchor points, and Dual T factorizes the matrix into the product of two easy-to-estimate matrices to avoid directly estimating the noisy class posterior. Beyond the instance-independent noise assumption, Zhang et al. introduced the instance-confidence embedding to model instance-dependent noise in estimating the transition matrix. On the other hand, Yang et al. proposed to use the Bayes optimal transition matrix estimated from the distilled examples for the instance-dependent noise transition matrix.

Remark: The robustness of these approaches is highly dependent on how precisely the transition matrix is estimated. To acquire such a transition matrix, they require prior knowledge in general, such as anchor points or clean validation data.

III-D2 Loss Reweighting

Inspired by the concept of importance reweighting , loss reweighting aims to assign smaller weights to the examples with false labels and greater weights to those with true labels. Accordingly, the reweighted loss on the mini-batch Bt\mathcal{B}_{t} is used to update the DNN,

Remark: These approaches need to manually pre-specify the weighting function as well as there additional hyper-parameters, which is fairly hard to be applied in practice due to the significant variation of appropriate weighting schemes that rely on the noise type and training data.

III-D3 Label Refurbishment

Technical Detail: Bootstrapping is the first method that proposes the concept of label refurbishment to update the target label of training examples. It develops a more coherent network that improves its ability to evaluate the consistency of noisy labels, with the label confidence α\alpha obtained via cross-validation. Dynamic bootstrapping dynamically adjusts the confidence α\alpha of individual training examples. The confidence α\alpha is obtained by fitting a two-component and one-dimensional beta mixture model to the loss distribution of all training examples. Self-adaptive training applies the exponential moving average to alleviate the instability issue of using instantaneous prediction of the current DNN,

D2L trains a DNN using a dimensionality-driven learning strategy to avoid overfitting to false labels. A simple measure called local intrinsic dimensionality is adopted to evaluate the confidence α\alpha in considering that the overfitting is exacerbated by dimensional expansion. Hence, refurbished labels are generated to prevent the dimensionality of the representation subspace from expanding at a later stage of training. Recently, SELFIE introduces a novel concept of refurbishable examples that can be corrected with high precision. The key idea is to consider the example with consistent label predictions as refurbishable because such consistent predictions correspond to its true label with a high probability owing to the learner’s perceptual consistency. Accordingly, the labels of only refurbishable examples are corrected to minimize the number of falsely corrected cases. Similarly, AdaCorr selectively refurbishes the label of noisy examples, but a theoretical error-bound is provided. Alternatively, SEAL averages the softmax output of a DNN on each example over the whole training process, then re-trains the DNN using the averaged soft labels.

Remark: Differently from loss correction and reweighting, all the noisy labels are explicitly replaced with other expected clean labels (or their combination). If there are not many confusing classes in data, these methods work well by refurbishing the noisy labels with high precision. In the opposite case, the DNN could overfit to wrongly refurbished labels.

III-D4 Meta Learning

In recent years, meta learning becomes an important topic in the machine learning community and is applied to improve noise robustness . The key concept is learning to learn that performs learning at a level higher than conventional learning, thus achieving data-agnostic and noise type-agnostic rules for better practical use. It is similar to loss reweighting and label refurbishment, but the adjustment is automated in a meta learning manner.

For the label refurbishment in Eq. (12), knowledge distillation adopts the technique of transferring knowledge from one expert model to a target model. The prediction from the expert DNN trained on small clean validation data is used instead of the prediction y^\hat{y} from the target DNN. MLC updates the target model with corrected labels provided by a meta model trained on clean validation data. The two models are trained concurrently via a bi-level optimization.

Remark: By learning the update rule via meta learning, the trained network easily adapts to various types of data and label noise. Nevertheless, unbiased clean validation data is essential to minimize the auxiliary objective, although it may not be available in real-world data.

III-E Sample Selection

To avoid any false corrections, many recent studies have adopted sample selection that involves selecting true-labeled examples from a noisy training dataset. In this case, the update equation in Eq. (2) is modified to render a DNN more robust for noisy labels. Let Ct⊆Bt\mathcal{C}_{t}\subseteq\mathcal{B}_{t} be the identified clean examples at time tt. Then, the DNN is updated only for the selected clean examples Ct\mathcal{C}_{t},

where the rest mini-batch examples, which are likely to be false-labeled, are excluded to pursue robust learning.

The memorization nature of DNNs has been explored theoretically and empirically to identify clean examples from noisy training data . Specifically, assuming clusterable data where the clusters are located on the unit Euclidean ball, Li et al. proved the distance from the initial weight W0{W}_{0} to the weight Wt{W}_{t} after tt iterations,

where ∥⋅∥F\left\lVert\cdot\right\rVert_{F} is the Frobenius norm, KK is the number of clusters, and C{C} is the set of cluster centers reaching all input examples within their ϵ0\epsilon_{0} neighborhood. Eq. (15) demonstrates that the weights of DNNs start to stray far from the initial weights when overfitting to corrupted labels, while they are still in the vicinity of the initial weights at an early stage of training . In the empirical studies , the memorization effect is also observed since DNNs tend to first learn simple and generalized patterns and then gradually overfit to all noisy patterns. As such, favoring small-loss training examples as the clean ones are commonly employed to design robust training methods .

Learning with sample selection is well motivated and works well in general, but this approach suffers from accumulated error caused by incorrect selection, especially when there are many ambiguous classes in training data. Hence, recent approaches often leverage multiple DNNs to cooperate with one another or run multiple training rounds . Moreover, to benefit from even false-labeled examples, loss correction or semi-supervised learning have been recently combined with the sample selection strategy .

Collaborative learning and co-training are widely used for the multi-network training. Consequently, the sample selection process is guided by the mentor network in the case of collaborative learning or the peer network in the case of co-training.

Technical Detail: Initially, Decouple proposes the decoupling of when to update from how to update. Hence, two DNNs are maintained simultaneously and updated only the examples selected based on a disagreement between the two DNNs. Next, due to the memorization effect of DNNs, many researchers have adopted another selection criterion, called a small-loss trick, which treats a certain number of small-loss training examples as true-labeled examples; many true-labeled examples tend to exhibit smaller losses than false-labeled examples, as illustrated in Figure 5(a). In MentorNet , a pre-trained mentor network guides the training of a student network in a collaborative learning manner. Based on the small-loss trick, the mentor network provides the student network with examples whose labels are likely to be correct. Co-teaching and Co-teaching+ also maintain two DNNs, but each DNN selects a certain number of small-loss examples and feeds them to its peer DNN for further training. Co-teaching+ further employs the disagreement strategy of Decouple compared with Co-teaching. In contrast, JoCoR reduces the diversity of two networks via co-regularization, making predictions of the two networks closer.

Remark: The co-training methods help reduce the confirmation bias , which is a hazard of favoring the examples selected at the beginning of training, while the increase in the number of learnable parameters makes their learning pipeline inefficient. In addition, the small-loss trick does not work well when the loss distribution of true-labeled and false-labeled examples largely overlap, as in the asymmetric noise in Figure 5(b).

III-E2 Multi-round Learning

Without maintaining additional DNNs, multi-round learning iteratively refines the selected set of clean examples by repeating the training round. Thus, the selected set keeps improved as the number of rounds increases.

Technical Detail: ITLM iteratively minimizes the trimmed loss by alternating between selecting true-labeled examples at the current moment and retraining the DNN using them. At each training round, only a fraction of small-loss examples obtained in the current round are used to retrain the DNN in the next round. INCV randomly divides noisy training data and then employs cross-validation to classify true-labeled examples while removing large-loss examples at each training round. Here, Co-teaching is adopted to train the DNN on the identified examples in the final round of training. Similarly, O2U-Net repeats the whole training process with the cyclical learning rate until enough loss statistics of every examples are gathered. Next, the DNN is re-trained from scratch only for the clean data where false-labeled examples have been detected and removed based on statistics.

A number of variations have been proposed to achieve high performance using iterative refinement only in a single training round. Beyond the small-loss trick, iterative detection detects false-labeled examples by employing the local outlier factor algorithm . With a Siamese network, it gradually pulls away false-labeled examples from true-labeled samples in the deep feature space. MORPH introduces the concept of memorized examples which is used to iteratively expand an initial safe set into a maximal safe set via self-transitional learning. TopoFilter utilizes the spatial topological pattern of learned representations to detect true-labeled examples, not relying on the prediction of the noisy classifier. NGC iteratively constructs the nearest neighbor graph using latent representations and performs geometry-based sample selection by aggregating information from neighborhoods. Soft pesudo-labels are assigned to the examples not selected.

Remark: The selected clean set keeps expanded and purified with iterative refinement, mainly through multi-round learning. As a side effect, the computational cost for training increases linearly for the number of training rounds.

III-E3 Hybrid Approach

An inherent limitation of sample selection is to discard all the unselected training examples, thus resulting in a partial exploration of training data. To exploit all the noisy examples, researchers have attempted to combine sample selection with other orthogonal ideas.

Technical Detail: The most prominent method in this direction is combining a specific sample selection strategy with a specific semi-supervised learning model. As illustrated in Figure 6, selected examples are treated as labeled clean data, whereas the remaining examples are treated as unlabeled. Subsequently, semi-supervised learning is performed using the transformed data. SELF is combined with a semi-supervised learning approach to progressively filter out false-labeled examples from noisy data. By maintaining the running average model called the mean-teacher as the backbone, it obtains the self-ensemble predictions of all training examples and then progressively removes examples whose ensemble predictions do not agree with their annotated labels. This method further leverages unsupervised loss from the examples not included in the selected clean set. DivideMix uses two-component and one-dimensional Gaussian mixture models to transform noisy data into labeled (clean) and unlabeled (noisy) sets. Then, it applies a semi-supervised technique MixMatch . Recently, RoCL employs two-phase learning strategies: supervised training on selected clean examples and then semi-supervised learning on relabeled noisy examples with self-supervision. For selection and relabeling, it computes the exponential moving average of the loss over training iterations.

Meanwhile, SELFIE is a hybrid approach of sample selection and loss correction. The loss of refurbishable examples is corrected (i.e., loss correction) and then used together with that of small-loss examples (i.e., sample selection). Consequently, more training examples are considered for updating the DNN. The curriculum loss (CL) is combined with the robust loss function approach and used to extract the true-labeled examples from noisy data.

Remark: Noise robustness is significantly improved by combining with other techniques. However, the hyperparameters introduced by these techniques render a DNN more susceptible to changes in data and noise types, and an increase in computational cost is inevitable

IV Methodological Comparison

In this section, we compare the 6262 deep learning methods for overcoming noisy labels introduced in Section III with respect to the following six properties. When selecting the properties, we refer to the properties that are typically used to compare the performance of robust deep learning methods . To the best of our knowledge, this survey is the first to provide a systematic comparison of robust training methods. This comprehensive comparison will provide useful insights that can enlighten new future directions.

(P1) Flexibility: With the rapid evolution of deep learning research, a number of new network architectures are constantly emerging and becoming available. Hence, the ability to support any type of architecture is important. “Flexibility” ensures that the proposed method can quickly adapt to the state-of-the-art architecture.

(P2) No Pre-training: A typical approach to improve noise robustness is to use a pre-trained network; however, this incurs an additional computational cost to the learning process. “No Pre-training” ensures that the proposed method can be trained from scratch without any pre-training.

(P3) Full Exploration: Excluding unreliable examples from the update is an effective method for robust deep learning; however, it eliminates hard but useful training examples as well. “Full Exploration” ensures that the proposed methods can use all training examples without severe overfitting to false-labeled examples by adjusting their training losses or applying semi-supervised learning.

(P4) No Supervision: Learning with supervision, such as a clean validation set or a known noise rate, is often impractical because they are difficult to obtain. Hence, such supervision had better be avoided to increase practicality in real-world scenarios. “No Supervision” ensures that the proposed methods can be trained without any supervision.

(P5) Heavy Noise: In real-world noisy data, the noise rate can vary from light to heavy. Hence, learning methods should achieve consistent noise robustness with respect to the noise rate. “Heavy Noise” ensures that the proposed methods can combat even the heavy noise.

(P6) Complex Noise: The type of label noise significantly affects the performance of a learning method. To manage real-world noisy data, diverse types of label noise should be considered when designing a robust training method. “Complex Noise” ensures that the proposed method can combat even the complex label noise.

Table II shows a comparison of all robust deep learning methods, which are grouped according to the most appropriate category. In the first row, the aforementioned six properties are labeled as P1–P6, and the availability of open-source implementation is added in the last column. For each property, we assign “◯\bigcirc” if it is completely supported, “✕” if it is not supported, and “△\bigtriangleup” if it is supported but not completely. More specifically, “△\bigtriangleup” is assigned to P1 if the method can be flexible but requires additional effort, to P5 if the method can combat only moderate label noise, and to P6 if the method does not make a strict assumption about the noise type but without explicitly modeling instance-dependent noise. Thus, for P6, the method marked with “✕” only deals with the instance-independent noise, while the method marked with “◯\bigcirc” deals with both instance-independent and -dependent noises. The remaining properties (i.e., P2, P3, and P4) are only assigned “◯\bigcirc” or “✕”. Regarding the implementation, we assign “N/A” if a publicly available source code is not available.

No existing method supports all the properties. Each method achieves noise robustness by supporting a different combination of the properties. The supported properties are similar among the methods of the same (sub-)category because those methods share the same methodological philosophy; however, they differ significantly depending on the (sub-)category. Therefore, we investigate the properties generally supported in each (sub-)category and summarize them in Table III. Here, the property of a (sub-)category is marked as the majority of the belonging methods. If no clear trend is observed among those methods, then the property is marked “△\bigtriangleup”.

V Noise Rate Estimation

The estimation of a noise rate is an imperative part of utilizing robust methods for better practical use, especially with the approaches belonging to the loss adjustment and sample selection. The estimated noise rate is widely used to reweight examples for a robust classifier or to determine how many examples should be selected as clean ones . However, detailed analysis has yet to be performed properly, though many robust approaches highly rely on the accuracy of noise rate estimation. The noise rate can be estimated by exploiting the inferred noise transition matrix , the Gaussian mixture model , or the cross-validation .

The noise transition matrix has been used to build a statistically consistent robust classifier because it represents the class posterior probabilities for noisy and clean data, as in Eq. (3). The first method to estimate the noise rate is exploiting this noise transition matrix, which can be inferred or trained accurately by using perfectly clean examples, i.e., anchor points ; an example xx with its label ii is defined as an anchor point if p(y=i∣x)=1p(y=i|x)=1 and p(y=k∣x)=0p(y=k|x)=0 for k≠ik\neq i. Thus, let Ai\mathcal{A}_{i} be the set of anchor points with label ii, then the element of the noise transition matrix TijT_{ij} is estimated by

However, since the anchor points are typically unknown in real-world data, they are identified from noisy training data by either theoretical derivations or heuristics . In addition, there have been recent efforts to learn the noise transition matrix without anchor points. T-Revision initializes a transition matrix by exploiting the examples with high noisy class posterior probabilities and then refines the matrix by adding a slack variable. Dual-T introduces an intermediate class that factorizes the transition matrix into two easy-to-estimate matrices for better accuracy. VolMinNet realizes an end-to-end framework and relaxes the need for anchor points under the sufficiently scattered assumption.

V-B Gaussian Mixture Model (GMM)

The second method is exploiting a one-dimensional and two-component GMM to model the loss distribution of true-labeled and false-labeled examples . As shown in Figure 7, since the loss distribution tends to be bi-modal, the two Gaussian components are fitted to the training loss by using the EM algorithm; the probability of an example being a false-labeled one is obtained through its posterior probability. Hence, the noise rate is estimated at each epoch tt by computing the expectation of the posterior probability for all training examples,

where gg is the Gaussian component with a larger loss. However, Pleiss et al. recently pointed out that the training loss becomes less separable by the GMM as the training progresses, and thus proposed the area under the loss (AUL) curve, which is the sum of the example’s training losses obtained from all previous training epochs. Even after the loss signal decays in later epochs, the distributions remain separable. Therefore, the noise rate is finally estimated by

V-C Cross Validation

The third method is estimating the noise rate by applying cross validation, which typically requires clean validation data . However, such clean validation data is hard to acquire in real-world applications. Thus, Chen et al. leveraged two randomly divided noisy training datasets for cross validation. Under the assumption that the two datasets share exactly the same noise transition matrix, the noise rate quantifies the test accuracy of DNNs that are respectively trained and tested on the two divided sets,

Therefore, the noise rate is estimated from the test accuracy obtained by cross validation.

VI Experimental Design

This section describes the typically used experimental design for comparing robust training methods in the presence of label noise. We introduce publicly available image datasets and then describe widely-used evaluation metrics.

To validate the robustness of the proposed algorithms, an image classification task was widely conducted on numerous image benchmark datasets. Table IV summarizes popularly-used public benchmark datasets, which are classified into two categories: 1) a “clean dataset” that consists of mostly true-labeled examples annotated by human experts and 2) a “real-world noisy dataset” that comprises real-world noisy examples with varying numbers of false labels.

According to the literature , seven clean datasets are widely used: MNISThttp://yann.lecun.com/exdb/mnist, classification of handwritten digits ; Fashion-MNISThttps://github.com/zalandoresearch/fashion-mnist, classification of various clothing ; CIFAR-10https://www.cs.toronto.edu/~kriz/cifar.html and CIFAR-10052, classification of a subset of 8080 million categorical images ; SVHNhttp://ufldl.stanford.edu/housenumbers, classification of house numbers in Google Street view images ; ImageNethttp://www.image-net.org and Tiny-ImageNethttps://www.kaggle.com/c/tiny-imagenet, image database organized according to the WordNet hierarchy and its small subset . Because the labels in these datasets are almost all true-labeled, their labels in the training data should be artificially corrupted for the evaluation of synthetic noises, namely symmetric noise and asymmetric noise.

VI-A2 Real-world Noisy Datasets

Unlike the clean datasets, real-world noisy datasets inherently contain many mislabeled examples annotated by non-experts. According to the literature , six real-world noisy datasets are widely used: ANIMAL-10Nhttps://dm.kaist.ac.kr/datasets/animal-10n, real-world noisy data of human-labeled online images for 10 confusing animals ; CIFAR-10Nhttp://noisylabels.com/ and CIFAR-100N57, variations of CIFAR-10 and CIFAR-100 with human-annotated real-world noisy labels collected from Amazon’s Mechanical Turk . They provide human labels with different noise rates, as shown in Table IV; Food-101Nhttps://kuanghuei.github.io/Food-101N, real-world noisy data of crawled food images annotated by their search keywords in the Food-101 taxonomy ; Clothing1Mhttps://www.floydhub.com/lukasmyth/datasets/clothing1m, real-world noisy data of large-scale crawled clothing images from several online shopping websites ; WebVisionhttps://data.vision.ee.ethz.ch/cvl/webvision/download.html, real-world noisy data of large-scale web images crawled from Flickr and Google Images search . To support sophisticated evaluation, most real-world noisy datasets contain their own clean validation set and provide the estimated noise rate of their training set.

VI-B Evaluation Metrics

A typical metric to assess the robustness of a particular method is the prediction accuracy for unbiased and clean examples that are not used in training. The prediction accuracy degrades significantly if the DNN overfits to false-labeled examples . Hence, test accuracy has generally been adopted for evaluation . For a test set T={(xi,yi)}i=1∣T∣\mathcal{T}=\{(x_{i},y_{i})\}_{i=1}^{|\mathcal{T}|}, let y^i\hat{y}_{i} be the predicted label of the ii-th example in T\mathcal{T}. Subsequently, the test accuracy is formalized by

If the test data are not available, validation accuracy can be used by replacing T\mathcal{T} in Eq. (21) with validation data V={(xi,yi)}i=1∣V∣\mathcal{V}=\{(x_{i},y_{i})\}_{i=1}^{|\mathcal{V}|} as an alternative,

Furthermore, if the specified method belongs to the “sample selection” category, label precision and label recall can be used as the metrics,

where St\mathcal{S}_{t} is the set of selected clean examples in a mini-batch Bt\mathcal{B}_{t}. The two metrics are performance indicators for the examples selected from the mini-batch as true-labeled ones .

Meanwhile, if the specified method belongs to the “label refurbishment” category, correction error can be used as an indicator of how many examples are incorrectly refurbished,

where R\mathcal{R} is the set of examples whose labels are refurbished by Eq. (12) and yirefurby_{i}^{refurb} is the refurbished label of the ii-th examples in R\mathcal{R}.

VII Future Research Directions

With recent efforts in the machine learning community, the robustness of DNNs becomes evolving in several directions. Thus, the existing approaches covered in our survey face a variety of future challenges. This section provides discussion for future research that can facilitate and envision the development of deep learning in the label noise area.

Existing theoretical and empirical studies for robust loss function and loss correction are largely built upon the instance-independent noise assumption that the label noise is independent of input features . However, this assumption may not be a good approximation of the real-world label noise. In particular, Chen et al. conducted a theoretical hypothesis testingIn Clothing1M, the result showed that the instance-independent noise happens with probability lower than 10−2125010^{-21250}, which is statistically impossible. using a popular real-world dataset, Clothing1M, and proved that its label noise is statistically different from the instance-independent noise. This testing confirms that the label noise should depend on the instance.

Conversely, most methods for the other direction (especially, sample selection) work well even under the instance-dependent label noise in general since they do not rely on the assumption. Nevertheless, Song et al. pointed out that their performance could considerably worsen in the instance-dependent (or real-world) noise compared to symmetric noise due to the confusion between true-labeled and false-labeled examples. The loss distribution of true-labeled examples heavily overlaps that of false-labeled samples in the asymmetric noise, which is similar to the real-world noise, in Figure 5(b). Thus, identifying clean examples becomes more challenging when dealing with the instance-dependent label noise.

Beyond the instance-independent label noise, there have been a few recent studies for the instance-dependent label noise. Mostly, they only focus on a binary classification task or a restricted small-scale machine learning model such as logistic regression . Therefore, learning with the instance-dependent label noise is an important topic that deserves more research attention.

VII-B Multi-label Data with Label Noise

Most of the existing methods are applicable only for a single-label multi-class classification problem, where each data example is assumed to have only one true label. However, in the case of multi-label learning, each data example can be associated with a set of multiple true class labels. In music categorization, each music can belong to multiple categories . In semantic scene classification, each scene may belong to multiple scene classes . Thus, contrary to the single-label setup, the multi-label classifier aims to predict a set of target objects simultaneously. In this setup, a multi-label dataset of millions of examples are reported to contain over 26.6%26.6\% false-positive labels as well as a significant number of omitted labels .

Even worse, the difference in occurrence between classes makes this problem more challenging; some minor class labels occur less in training data than other major class labels. Considering such aspects that can arise in multi-label classification, the simple extension of existing methods may not learn the proper correlations among multiple labels. Therefore, learning from noisy labels with multi-label data is another important topic for future research. We refer the readers to a recent study that discusses the evaluation of multi-label classifiers trained with noisy labels.

VII-C Class Imbalance Data with Label Noise

The class imbalance in training data is commonly observed, where a few classes account for most of the data. Especially when working with large data in many real-world applications, this problem becomes more severe and is often associated with the problem of noisy labels . Nevertheless, to ease the label noise problem, it is commonly assumed that training examples are equally distributed over all class labels in the training data. This assumption is quite strong when collecting large-scale data, and thus we need to consider a more realistic scenario in which the two problems coexist.

Most of the existing robust methods may not work well with the class imbalance, especially when they rely on the learning dynamics of DNNs, e.g., the small-loss trick or memorization effect. Under the existence of the class imbalance, the training model converges to major classes faster than minor classes such that most examples in the major class exhibit small losses (i.e., early memorization). That is, there is a risk of discarding most examples in the minor class. Furthermore, in terms of example importance, high-loss examples are commonly favored for the class imbalance problem , while small-loss examples are favored for the label noise problem. This conceptual contradiction hinders the applicability of the existing methods that neglect the class imbalance. Therefore, these two problems should be considered simultaneously to deal with more general situations.

VII-D Robust and Fair Training

Machine learning classifiers can perpetuate and amplify the existing systemic injustices in society . Hence, fairness is becoming another important topic. Traditionally, robust training and fair training have been studied by separate communities; robust training with noisy labels has mostly focused on combating label noise without regarding data bias , whereas fair training has focused on dealing with data bias, not necessarily noise . However, noisy labels and data bias, in fact, coexist in real-world data. Satisfying both robustness and fairness is more realistic but challenging because the bias in data is pertinent to label noise.

In general, many fairness criteria are group-based, where a target metric is equalized or enforced over subpopulations in the data, also known as protected groups such as race or gender . Accordingly, the goal of fair training is building a model that satisfies such fairness criteria for the true protected groups. However, if the noisy protection group is involved, such fairness criteria cannot be directly applied. Recently, mostly after 2020, a few pioneering studies have emerged to consider both robustness and fairness objectives at the same time under the binary classification setting . Therefore, more research attention is needed for the convergence of robust training and fair training.

VII-E Connection with Input Perturbation

There has been a lot of research on the robustness of deep learning under input perturbation, mainly in the field of adversarial training where the input feature is maliciously perturbed to distort the output of the DNN . Although learning with noisy labels and learning with noisy inputs have been regarded as separate research fields, their goals are similar in that they learn noise-robust representations from noisy data. Based on this common point of view, a few recent studies have investigated the interaction of adversarial training with noisy labels .

Interestingly, it was turned out that adversarial training makes DNNs robust to label noise . Based on this finding, Damodaran et al. proposed a new regularization term, called Wasserstein adversarial regularization, to address the problem of learning with noisy labels. Zhu et al. proposed to use the number of projected gradient descent steps as a new criterion for sample selection such that clean examples are filtered out from noisy data. These approaches are regarded as a new perspective on label noise compared to traditional work. Therefore, understanding the connection between input perturbation and label noise could be another future topic for better representation learning toward robustness.

VII-F Efficient Learning Pipeline

The efficiency of the learning pipeline is another important aspect to design deep learning approaches. However, for robust deep learning, most studies have neglected the efficiency of the algorithm because their main goal is to improve the robustness to label noise. For example, maintaining multiple DNNs or training a DNN in multiple rounds is frequently used, but these approaches significantly degrade the efficiency of the learning pipeline. On ther other hand, the need for more efficient algorithms is increasing owing to the rapid increase in the amount of available data .

According to our literature survey, most work did not even report the efficiency (or time complexity) of their approaches. However, it is evident that saving the training time is helpful under the restricted budget for computation. Therefore, enhancing the efficiency will significantly increase the usability of robust deep learning in the big data era.

VIII Conclusion

DNNs easily overfit to false labels owing to their high capacity in totally memorizing all noisy training samples. This overfitting issue still remains even with various conventional regularization techniques, such as dropout and batch normalization, thereby significantly decreasing their generalization performance. Even worse, in real-world applications, the difficulty in labeling renders the overfitting issue more severe. Therefore, learning from noisy labels has recently become one of the most active research topics.

In this survey, we presented a comprehensive understanding of modern deep learning methods to address the negative consequences of learning from noisy labels. All the methods were grouped into five categories according to their underlying strategies and described along with their methodological weaknesses. Furthermore, a systematic comparison was conducted using six popular properties used for evaluation in the recent literature. According to the comparison results, there is no ideal method that supports all the required properties; the supported properties varied depending on the category to which each method belonged. Several experimental guidelines were also discussed, including noise rate estimation, publicly available datasets, and evaluation metrics. Finally, we provided insights and directions for future research in this domain.

References

Acknowledgements

This work was supported by Institute of Information & Communications Technology Planning & Evaluation (IITP) grant funded by the Korea government (MSIT) (No. 2020-0-00862, DB4DL: High-Usability and Performance In-Memory Distributed DBMS for Deep Learning).