Long-tail learning via logit adjustment

Aditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain, Andreas Veit, Sanjiv Kumar

Introduction

Real-world classification problems typically exhibit a long-tailed label distribution, wherein most labels are associated with only a few samples (Van Horn and Perona, 2017; Buda et al., 2017; Liu et al., 2019). Owing to this paucity of samples, generalisation on such labels is challenging; moreover, naïve learning on such data is susceptible to an undesirable bias towards dominant labels. This problem has been widely studied in the literature on learning under class imbalance (Cardie and Howe, 1997; Chawla et al., 2002; Qiao and Liu, 2009; He and Garcia, 2009; Wallace et al., 2011) and the related problem of cost-sensitive learning (Elkan, 2001; Zadrozny and Elkan, 2001; Masnadi-Shirazi and Vasconcelos, 2010; Dmochowski et al., 2010).

Recently, long-tail learning has received renewed interest in the context of neural networks. Two active strands of work involve post-hoc normalisation of the classification weights (Zhang et al., 2019; Kim and Kim, 2019; Kang et al., 2020; Ye et al., 2020), and modification of the underlying loss to account for varying class penalties (Zhang et al., 2017; Cui et al., 2019; Cao et al., 2019; Tan et al., 2020). Each of these strands is intuitive, and has proven empirically successful. However, they are not without limitation: e.g., weight normalisation crucially relies on the weight norms being smaller for rare classes; however, this assumption is sensitive to the choice of optimiser (see §2). On the other hand, loss modification sacrifices the consistency that underpins the softmax cross-entropy (see §5.2). Consequently, existing techniques may result in suboptimal solutions even in simple settings (§6.1).

In this paper, we present two simple modifications of softmax cross-entropy training that unify several recent proposals, and overcome their limitations. Our techniques revisit the classic idea of logit adjustment based on label frequencies (Provost, 2000; Zhou and Liu, 2006; Collell et al., 2016), applied either post-hoc on a trained model, or as a modification of the training loss. Conceptually, logit adjustment encourages a large relative margin between a pair of rare and dominant labels. This has a firm statistical grounding: unlike recent techniques, it is consistent for minimising the balanced error (cf. (2)), a common metric in long-tail settings which averages the per-class errors. This grounding translates into strong empirical performance on real-world datasets.

In summary, our contributions are: (i) we present two realisations of logit adjustment for long-tail learning, applied either post-hoc (§4.2) or during training (§5.2) (ii) we establish that logit adjustment overcomes limitations in recent proposals (see Table 1), and in particular is Fisher consistent for minimising the balanced error (cf. (2)); (iii) we confirm the efficacy of the proposed techniques on real-world datasets (§6). In the course of our analysis, we also present a general version of the softmax cross-entropy with a pairwise label margin (11), which offers flexibility in controlling the relative contribution of labels to the overall loss.

Problem setup and related work

Broadly, extant approaches to coping with class imbalance (see also Table 2) modify:

the inputs to a model, for example by over- or under-sampling (Kubat and Matwin, 1997; Chawla et al., 2002; Wallace et al., 2011; Mikolov et al., 2013; Mahajan et al., 2018; Yin et al., 2018)

the outputs of a model, for example by post-hoc correction of the decision threshold (Fawcett and Provost, 1996; Collell et al., 2016) or weights (Kim and Kim, 2019; Kang et al., 2020)

the internals of a model, for example by modifying the loss function (Xie and Manski, 1989; Morik et al., 1999; Cui et al., 2019; Zhang et al., 2017; Cao et al., 2019; Tan et al., 2020)

One may easily combine approaches from the first stream with those from the latter two. Consequently, we focus on the latter two in this work, and describe some representative recent examples from each.

While intuitive, balancing has minimal effect in separable settings: solutions that achieve zero training loss will necessarily remain optimal even under weighting (Byrd and Lipton, 2019). Intuitively, one would like instead to shift the separator closer to a dominant class. Li et al. (2002); Wu et al. (2008); Masnadi-Shirazi and Vasconcelos (2010); Iranmehr et al. (2019); Gottlieb et al. (2020) thus proposed to add per-class margins into the hinge loss. (Cao et al., 2019) proposed to add a per-class margin into the softmax cross-entropy:

Limitations of existing approaches. Each of the above methods are intuitive, and have shown strong empirical performance. However, a closer analysis identifies some subtle limitations.

Limitations of loss modification. Enforcing a per-label margin per (5) and (6) is intuitive, as it allows for shifting the decision boundary away from rare classes. However, when doing so, it is important to ensure Fisher consistency (Lin, 2004) (or classification calibration (Bartlett et al., 2006)) of the resulting loss for the balanced error. That is, the minimiser of the expected loss (equally, the empirical risk in the infinite sample limit) should result in a minimal balanced error. Unfortunately, both (5) and (6) are not consistent in this sense, even for binary problems; see §5.2, §6.1 for details.

Logit adjustment for long-tail learning: a statistical view

The above suggests there is scope for improving performance on long-tail problems, both in terms of post-hoc correction and loss modification. We now show how a statistical perspective on the problem suggests simple procedures of each type, both of which overcome the limitations discussed above.

i.e., we translate the (unknown) distributional scores or logits based on the class priors. This simple fact immediately suggests two means of optimising for the balanced error:

Such logit adjustment techniques — which have been a classic approach to class-imbalance (Provost, 2000) — neatly align with the post-hoc and loss modification streams discussed in §2. However, unlike most previous techniques from these streams, logit adjustment is endowed with a clear statistical grounding: by construction, the optimal solution under such adjustment coincides with the Bayes-optimal solution (7) for the balanced error, i.e., it is Fisher consistent for minimising the balanced error. We shall demonstrate this translates into superior empirical performance (§6). Note also that logit adjustment may be easily extended to cover performance measures beyond the balanced error, e.g., with distinct costs for errors on dominant and rare classes; we leave a detailed study and contrast to existing cost-sensitive approaches (Iranmehr et al., 2019; Gottlieb et al., 2020) to future work.

We now study each of the techniques (i) and (ii) in turn.

Post-hoc logit adjustment

We now detail to perform post-hoc logit adjustment on a classifier trained on long-tailed data. We further show this bears similarity to recent weight normalisation schemes, but has a subtle advantage.

In post-hoc logit adjustment, we propose to instead predict, for suitable τ>0\tau>0:

One may treat τ\tau as a tuning parameter to be chosen based on some measure of holdout calibration, e.g., the expected calibration error (Murphy and Winkler, 1987; Guo et al., 2017), probabilistic sharpness (Gneiting et al., 2007; Kuleshov et al., 2018), or a proper scoring rule such as the log-loss or squared error (Gneiting and Raftery, 2007). One may alternately fix τ=1\tau=1 and aim to learn inherently calibrated probabilities, e.g., via label smoothing (Szegedy et al., 2016; Müller et al., 2019).

2 Comparison to existing post-hoc techniques

Post-hoc logit adjustment with τ=1\tau=1 is not a new idea in the class imbalance literature. Indeed, this is a standard technique when creating stratified samples (King and Zeng, 2001), and when training binary classifiers (Fawcett and Provost, 1996; Provost, 2000; Maloof, 2003). In multiclass settings, this has been explored in Zhou and Liu (2006); Collell et al. (2016). However, τ≠1\tau\neq 1 is important in practical usage of neural networks, owing to their lack of calibration. Further, we now explicate that post-hoc logit adjustment has an important advantage over recent post-hoc weight normalisation techniques.

Recall that weight normalisation involves learning a scorer fy(x)=wy⊤Φ(x)f_{y}(x)=w_{y}^{\top}\Phi(x), and then post-hoc normalising the weights via wy/νyτ{w_{y}}/{\nu_{y}^{\tau}} for τ>0\tau>0. We demonstrated in §2 that using νy=∥wy∥2\nu_{y}=\|w_{y}\|_{2} may be ineffective when using adaptive optimisers. However, even with νy=πy\nu_{y}=\pi_{y}, there is a subtle contrast to post-hoc logit adjustment: while the former performs a multiplicative update to the logits, the latter performs an additive update. The two techniques may thus yield different orderings over labels, since

The logit adjusted softmax cross-entropy

We now show how to directly bake logit adjustment into the softmax cross-entropy. We show that this approach has an intuitive relation to existing loss modification techniques.

Given a scorer that minimises the above, we now predict argmax⁡y∈[L]fy(x){\operatorname{argmax}}_{y\in[L]}f_{y}(x) as usual.

Compared to the standard softmax cross-entropy (1), the above applies a label-dependent offset to each logit. Compared to (9), we directly enforce the class prior offset while learning the logits, rather than doing this post-hoc. The two approaches have a deeper connection: observe that (10) is equivalent to using a scorer of the form gy(x)=fy(x)+τ⋅log⁡πyg_{y}(x)=f_{y}(x)+\tau\cdot\log\pi_{y}. We thus have argmax⁡y∈[L]fy(x)=argmax⁡y∈[L]gy(x)−τ⋅log⁡πy{\operatorname{argmax}}_{y\in[L]}f_{y}(x)={\operatorname{argmax}}_{y\in[L]}g_{y}(x)-\tau\cdot\log\pi_{y}. Consequently, one can equivalently view learning with this loss as learning a standard scorer g(x)g(x), and post-hoc adjusting its logits to make a prediction. For convex objectives, we thus do not expect any difference between the solutions of the two approaches. For non-convex objectives, as encountered in neural networks, the bias endowed by adding τ⋅log⁡πy\tau\cdot\log\pi_{y} to the logits is however likely to result in a different local minima.

For more insight into the loss, consider the following pairwise margin loss

for label weights αy>0\alpha_{y}>0, and pairwise label margins Δyy′\Delta_{yy^{\prime}} representing the desired gap between scores for yy and y′y^{\prime}. For τ=1\tau=1, our logit adjusted loss (10) corresponds to (11) with αy=1\alpha_{y}=1 and Δyy′=log⁡(πy′πy)\Delta_{yy^{\prime}}=\log\left(\frac{\pi_{y^{\prime}}}{\pi_{y}}\right). This demands a larger margin between rare positive (πy∼0\pi_{y}\sim 0) and dominant negative (πy′∼1\pi_{y^{\prime}}\sim 1) labels, so that scores for dominant classes do not overwhelm those for rare ones.

2 Comparison to existing loss modification techniques

A cursory inspection of (5), (6) reveals a striking similarity to our logit adjusted softmax cross-entropy (10). The balanced loss (4) also bears similarity, except that the weighting is performed outside the logarithm. Each of these losses are special cases of the pairwise margin loss (11) enforcing uniform margins that only consider the positive or negative label, unlike our approach.

For example, αy=1πy\alpha_{y}=\frac{1}{\pi_{y}} and Δyy′=0\Delta_{yy^{\prime}}=0 yields the balanced loss (4). This does not explicitly enforce a margin between the labels, which is undesirable for separable problems (Byrd and Lipton, 2019). When αy=1\alpha_{y}=1, the choice Δyy′=πy−1/4\Delta_{yy^{\prime}}=\pi_{y}^{-1/4} yields (5). Finally, Δyy′=log⁡F(πy′)\Delta_{yy^{\prime}}=\log F(\pi_{y^{\prime}}) yields (6), where F ⁣:→(0,1]F\colon\to(0,1] is some non-decreasing function, e.g., F(z)=zτF(z)=z^{\tau} for τ>0\tau>0. These losses thus either consider the frequency of the positive yy or negative y′y^{\prime}, but not both simultaneously.

The above choices of α\alpha and Δ\Delta are all intuitively plausible. However, §3 indicates that our loss in (10) has a firm statistical grounding: it ensures Fisher consistency for the balanced error.

Observe that when δy=πy\delta_{y}=\pi_{y}, we immediately deduce that the logit-adjusted loss of (10) is consistent. Similarly, δy=1\delta_{y}=1 recovers the classic result that the balanced loss is consistent. While the above is only a sufficient condition, it turns out that in the binary case, one may neatly encapsulate a necessary and sufficient condition for consistency that rules out other choices; see Appendix B.1. This suggests that existing proposals may thus underperform with respect to the balanced error in certain settings, as verified empirically in §6.1.

3 Discussion and extensions

More broadly, however, there is value in combining logit adjustment with other techniques. For example, Theorem 1 implies that it is sensible to combine logit adjustment with loss weighting; e.g., one may pick Δyy′=τ⋅log⁡(πy′/πy)\Delta_{yy^{\prime}}={\tau}\cdot\log\left({{\pi_{y^{\prime}}}/{\pi_{y}}}\right), and αy=πyτ−1\alpha_{y}=\pi_{y}^{\tau-1}. This is similar to Cao et al. (2019), who found benefits in combining weighting with their loss. One may also generalise the formulation in Theorem 1, and employ Δyy′=τ1⋅log⁡πy−τ2⋅log⁡πy′\Delta_{yy^{\prime}}=\tau_{1}\cdot\log\pi_{y}-\tau_{2}\cdot\log\pi_{y^{\prime}}, where τ1,τ2\tau_{1},\tau_{2} are constants. This interpolates between the logit adjusted loss (τ1=τ2\tau_{1}=\tau_{2}) and a version of the equalised margin loss (τ1=0\tau_{1}=0).

Cao et al. (2019, Theorem 2) provides a rigorous generalisation bound for the adaptive margin loss under the assumption of separable training data with binary labels. The inconsistency of the loss with respect to the balanced error concerns the more general scenario of non-separable multiclass data data, which may occur, e.g., owing to label noise or limitation in model capacity. We shall subsequently demonstrate that encouraging consistency can lead to gains in practical settings. We shall further see that combining the Δ\Delta implicit in this loss with our proposed Δ\Delta can lead to further gains, indicating a potentially complementary nature of the losses.

Interestingly, for τ=−1\tau=-1, a similar loss to (10) has been considered in the context of negative sampling for scalability (Yi et al., 2019): here, one samples a small subset of negatives based on the class priors π\pi, and applies logit correction to obtain an unbiased estimate of the unsampled loss function based on all the negatives (Bengio and Senecal, 2008). Losses of the general form (11) have also been explored for structured prediction (Zhang, 2004; Pletscher et al., 2010; Hazan and Urtasun, 2010).

Experimental results

We now present experiments confirming our main claims: (i) on simple binary problems, existing weight normalisation and loss modification techniques may not converge to the optimal solution (§6.1); (ii) on real-world datasets, our post-hoc logit adjustment outperforms weight normalisation, and one can obtain further gains via our logit adjusted softmax cross-entropy (§6.2).

i.e., it is a linear separator passing through the origin. We compare this separator against those found by several margin losses based on (11): standard ERM (Δyy′=0\Delta_{yy^{\prime}}=0), the adaptive loss (Cao et al., 2019) (Δyy′=πy−1/4\Delta_{yy^{\prime}}=\pi_{y}^{-1/4}), an instantiation of the equalised loss (Tan et al., 2020) (Δyy′=log⁡πy′\Delta_{yy^{\prime}}=\log\pi_{y^{\prime}}), and our logit adjusted loss (Δyy′=log⁡πy′πy\Delta_{yy^{\prime}}=\log\frac{\pi_{y^{\prime}}}{\pi_{y}}). For each loss, we train an affine classifier on a sample of 10,00010,000 instances, and evaluate the balanced error on a test set of 10,00010,000 samples over 100100 independent trials.

Figure 2 confirms that the logit adjusted margin loss attains a balanced error close to that of the Bayes-optimal, which is visually reflected by its learned separator closely matching that in (12). This is in line with our claim of the logit adjusted margin loss being consistent for the balanced error, unlike other approaches. Figure 2 also compares post-hoc weight normalisation and logit adjustment for varying scaling parameter τ\tau (c.f. (3), (9)). Logit adjustment is seen to approach the performance of the Bayes predictor; any weight normalisation is however seen to hamper performance. This verifies the consistency of logit adjustment, and inconsistency of weight normalisation (§4.2).

2 Results on real-world datasets

Baselines. We consider: (i) empirical risk minimisation (ERM) on the long-tailed data, (ii) post-hoc weight normalisation (Kang et al., 2020) per (3) (using νy=∥wy∥2\nu_{y}=\|w_{y}\|_{2} and τ=1\tau=1) applied to ERM, (iii) the adaptive margin loss (Cao et al., 2019) per (5), and (iv) the equalised loss (Tan et al., 2020) per (6), with δy′=F(πy′)\delta_{y^{\prime}}=F(\pi_{y^{\prime}}) for the threshold-based FF of Tan et al. (2020). Cao et al. (2019) demonstrated superior performance of their adaptive margin loss against several other baselines, such as the balanced loss of (4), and that of Cui et al. (2019). Where possible, we report numbers for the baselines (which use the same setup as above) from the respective papers. See also our concluding discussion about extensions to such methods that improve performance.

We compare the above methods against our proposed post-hoc logit adjustment (9), and logit adjusted loss (10). For post-hoc logit adjustment, we fix the scalar τ=1\tau=1; we analyse the effect of tuning this in Figure 3. We do not perform any further tuning of our logit adjustment techniques.

Results and analysis. Table 3 summarises our results, which demonstrate our proposed logit adjustment techniques consistently outperform existing methods. Indeed, while weight normalisation offers gains over ERM, these are improved significantly by post-hoc logit adjustment (e.g., 8% relative reduction on CIFAR-10). Similarly loss correction techniques are generally outperformed by our logit adjusted softmax cross-entropy (e.g., 6% relative reduction on iNaturalist).

Figure 3 studies the effect of tuning the scaling parameter τ>0\tau>0 afforded by post-hoc weight normalisation (using νy=∥wy∥2\nu_{y}=\|w_{y}\|_{2}) and post-hoc logit adjustment. Even without any scaling, post-hoc logit adjustment generally offers superior performance to the best result from weight normalisation (cf. Table 3); with scaling, this is further improved. See Appendix 11 for a plot on ImageNet-LT.

Figure 4 breaks down the per-class accuracies on CIFAR-10, CIFAR-100, and iNaturalist. On the latter two datasets, for ease of visualisation, we aggregate the classes into ten groups based on their frequency-sorted order (so that, e.g., group comprises the top L10\frac{L}{10} most frequent classes). As expected, dominant classes generally see a lower error rate with all methods. However, the logit adjusted loss is seen to systematically improve performance over ERM, particularly on rare classes.

While our logit adjustment techniques perform similarly, there is a slight advantage to the loss function version. Nonetheless, the strong performance of post-hoc logit adjustment corroborates the ability to decouple representation and classifier learning in long-tail settings (Zhang et al., 2019).

Discussion and extensions Table 3 shows the advantage of logit adjustment over recent post-hoc and loss modification proposals, under standard setups from the literature. We believe further improvements are possible by fusing complementary ideas, and remark on four such options.

First, one may use a more complex base architecture; our choices are standard in the literature, but, e.g., Kang et al. (2020) found gains on ImageNet-LT by employing a ResNet-152, with further gains from training it for 200200 as opposed to the customary 9090 epochs. Table 4 confirms that logit adjustment similarly benefits from this choice. For example, on iNaturalist, we obtain an improved balanced error of 31.15%{31.15}\% for the logit adjusted loss. When training for more (200200) epochs per the suggestion of Kang et al. (2020), this further improves to 30.12%{30.12}\%.

Second, one may combine together the Δ\Delta’s for various special cases of the pairwise margin loss. Indeed, we find that combining our relative margin with the adaptive margin of Cao et al. (2019) — i.e., using the pairwise margin loss with Δyy′=log⁡πy′πy+1πy1/4\Delta_{yy^{\prime}}=\log\frac{\pi_{y^{\prime}}}{\pi_{y}}+\frac{1}{\pi_{y}^{1/4}} — results in a top-11 accuracy of 31.56%{31.56}\% on iNaturalist. When using a ResNet-152, this further improves to 29.22%{29.22}\% when trained for 90 epochs, and 28.02%\mathbf{28.02}\% when trained for 200 epochs. While such a combination is nominally heuristic, we believe there is scope to formally study such schemes, e.g., in terms of induced generalisation performance.

Third, Cao et al. (2019) observed that their loss benefits from a deferred reweighting scheme (DRW), wherein the model begins training as normal, and then applies class-weighting after a fixed number of epochs. On CIFAR-10-LT and CIFAR-100-LT, this achieves 22.97%22.97\% and 57.96%57.96\% error respectively; both are outperformed by our vanilla logit adjusted loss. On iNaturalist with a ResNet-50, this achieves an error of 32.0%32.0\%, outperforming our 33.6%33.6\%. (Note that our simple combination of the relative and adaptive margins outperforms these reported numbers of DRW.) However, given the strong improvement of our loss over that in Cao et al. (2019) when both methods use SGD, we expect that employing DRW (which applies to any loss) may be similarly beneficial for our method.

Fourth, per §2, one may perform data augmentation; e.g., see Tan et al. (2020, Section 6). While further exploring such variants are of empirical interest, we hope to have illustrated the conceptual and empirical value of logit adjustment, and leave this for future work.

References

Appendix A Proofs of results in body

Consequently, under constant weights αy=1\alpha_{y}=1, the Bayes-optimal score will satisfy fy∗(x)+log⁡δy=log⁡ηy(x)f^{*}_{y}(x)+\log\delta_{y}=\log\eta_{y}(x), or fy∗(x)=log⁡ηy(x)δyf^{*}_{y}(x)=\log\frac{\eta_{y}(x)}{\delta_{y}}.

where πˉy∝πy⋅αy\bar{\pi}_{y}\propto{\pi_{y}\cdot\alpha_{y}}. Consequently, learning with the weighted loss is equivalent to learning with the original loss, on a distribution with modified base-rates πˉ\bar{\pi}. Under such a distribution, we have class-conditional distribution

Consequently, suppose αy=δyπy\alpha_{y}=\frac{\delta_{y}}{\pi_{y}}. Then, fy∗(x)=log⁡ηˉy(x)δy=log⁡ηy(x)πy+C(x)f^{*}_{y}(x)=\log\frac{\bar{\eta}_{y}(x)}{\delta_{y}}=\log\frac{\eta_{y}(x)}{\pi_{y}}+C(x), where C(x)C(x) does not depend on yy. Consequently, argmax⁡y∈[L] fy∗(x)=argmax⁡y∈[L] ηy(x)πy{\operatorname{argmax}\nolimits_{y\in[L]}}\,f^{*}_{y}(x)={\operatorname{argmax}\nolimits_{y\in[L]}}\,\frac{\eta_{y}(x)}{\pi_{y}}, which is the Bayes-optimal prediction for the balanced error.

In sum, a consistent family can be obtained by choosing any set of constants δy>0\delta_{y}>0 and setting

Appendix B On the consistency of binary margin-based losses

We study two properties of this family losses. First, under what conditions are the losses Fisher consistent for the balanced error? We shall show that in fact there is a simple condition characterising this. Second, do the losses preserve properness of the original binary logistic loss? We shall show that this is always the case, but that the losses involve fundamentally different approximations.

The losses in (13) are consistent for the balanced error iff

From the above, some admissible parameter choices include:

ω+1=1π\omega_{+1}=\frac{1}{\pi}, ω−1=11−π\omega_{-1}=\frac{1}{1-\pi}, δ±1=1\delta_{\pm 1}=1; i.e., the standard weighted loss with a constant margin

ω±1=1\omega_{\pm 1}=1, δ+1=1γ⋅log⁡1−ππ\delta_{+1}=\frac{1}{\gamma}\cdot\log\frac{1-\pi}{\pi}, δ−1=1γ⋅log⁡π1−π\delta_{-1}=\frac{1}{\gamma}\cdot\log\frac{\pi}{1-\pi}; i.e., the unweighted loss with a margin biased towards the rare class, as per our logit adjustment procedure

The second example above is unusual in that it requires scaling the margin with the temperature; consequently, the margin disappears as γ→+∞\gamma\to+\infty. Other combinations are of course possible, but note that one cannot arbitrarily choose parameters and hope for consistency in general. Indeed, some inadmissible choices are naïve applications of the margin modification or weighting, e.g.,

ω+1=1π\omega_{+1}=\frac{1}{\pi}, ω−1=11−π\omega_{-1}=\frac{1}{1-\pi}, δ+1=1γ⋅log⁡1−ππ\delta_{+1}=\frac{1}{\gamma}\cdot\log\frac{1-\pi}{\pi}, δ−1=1γ⋅log⁡π1−π\delta_{-1}=\frac{1}{\gamma}\cdot\log\frac{\pi}{1-\pi}; i.e., combining both weighting and margin modification

ω±1=1\omega_{\pm 1}=1, δ+1=1γ⋅(1−π)\delta_{+1}=\frac{1}{\gamma}\cdot(1-\pi), δ−1=1γ⋅π\delta_{-1}=\frac{1}{\gamma}\cdot{\pi}; i.e., specific margin modification

Note further that the choices of Cao et al. , Tan et al. do not meet the requirements of Lemma 2.

We make two final remarks. First, the above only considers consistency of the result of loss minimisation. For any choice of weights and margins, we may apply suitable post-hoc correction to the predictions to account for any bias in the optimal scores. Second, as γ→+∞\gamma\to+\infty, any constant margins δ±1>0\delta_{\pm 1}>0 will have no effect on the consistency condition, since σ(γ⋅δ±1)→1\sigma(\gamma\cdot\delta_{\pm 1})\to 1. The condition will be wholly determined by the weights ω±1\omega_{\pm 1}. For example, we may choose ω+1=1π\omega_{+1}=\frac{1}{\pi}, ω−1=11−π\omega_{-1}=\frac{1}{1-\pi}, δ+1=1\delta_{+1}=1, and δ−1=π1−π\delta_{-1}=\frac{\pi}{1-\pi}; the resulting loss will not be consistent for finite γ\gamma, but will become so in the limit γ→+∞\gamma\to+\infty. For more discussion on this particular loss, see Appendix C.

B.2 Properness of the pairwise margin loss

The losses in (13) are proper composite, with link function

where a=ω+1ω−1⋅eγ⋅δ+1eγ⋅δ−1a=\frac{\omega_{+1}}{\omega_{-1}}\cdot\frac{e^{\gamma\cdot\delta_{+1}}}{e^{\gamma\cdot\delta_{-1}}}, b=eγ⋅δ−1b=e^{\gamma\cdot\delta_{-1}}, c=eγ⋅δ+1c=e^{\gamma\cdot\delta_{+1}}, and q=1−ppq=\frac{1-p}{p}.

The above family of losses is proper composite iff the function

is invertible [Reid and Williamson, 2010, Corollary 12]. We have

The invertibility of Ψ−1\Psi^{-1} is immediate. To compute the link function Ψ\Psi, note that

where a=ω+1ω−1⋅eγ⋅δ+1eγ⋅δ−1a=\frac{\omega_{+1}}{\omega_{-1}}\cdot\frac{e^{\gamma\cdot\delta_{+1}}}{e^{\gamma\cdot\delta_{-1}}}, b=eγ⋅δ−1b=e^{\gamma\cdot\delta_{-1}}, c=eγ⋅δ+1c=e^{\gamma\cdot\delta_{+1}}, g=eγ⋅fg=e^{\gamma\cdot f}, and q=1−ppq=\frac{1-p}{p}. Thus,

As a sanity check, suppose a=b=c=γ=1a=b=c=\gamma=1. This corresponds to the standard logistic loss. Then,

Figure 5 and 6 compares the link functions for a few different settings:

the balanced loss, where ω+1=1π\omega_{+1}=\frac{1}{\pi}, ω−1=11−π\omega_{-1}=\frac{1}{1-\pi}, and δ±1=1\delta_{\pm 1}=1

an unequal margin loss, where ω±1=1\omega_{\pm 1}=1, δ+1=1γ⋅log⁡1−ππ\delta_{+1}=\frac{1}{\gamma}\cdot\log\frac{1-\pi}{\pi}, and δ−1=1γ⋅log⁡π1−π\delta_{-1}=\frac{1}{\gamma}\cdot\log\frac{\pi}{1-\pi}

a balanced + margin loss, where ω+1=1π\omega_{+1}=\frac{1}{\pi}, ω−1=11−π\omega_{-1}=\frac{1}{1-\pi}, δ+1=1\delta_{+1}=1, and δ−1=π1−π\delta_{-1}=\frac{\pi}{1-\pi}.

To better understand the effect of parameter choices, Figure 7 illustrates the conditional Bayes risk curves, i.e.,

We remark here that for the balanced error, this function takes the form L‾(p)=p⋅⟦p<π⟧+(1−p)⋅⟦p>π⟧\underline{L}(p)=p\cdot\llbracket p<\pi\rrbracket+(1-p)\cdot\llbracket p>\pi\rrbracket, i.e., it is a “tent shaped” concave function with a maximum at p=πp=\pi.

For ease of comparison, we normalise this curves to have a maximum of 11. Figure 7 shows that simply applying unequal margins does not affect the underlying conditional Bayes risk compared to the standard log-loss; thus, the change here is purely in terms of the link function. By contrast, either balancing the loss or applying a combination of weighting and margin modification results in a closer approximation to the conditional Bayes risk curve for the cost-sensitive loss with cost π\pi.

Appendix C Relation to cost-sensitive SVMs

We recapitulate the analysis of Masnadi-Shirazi and Vasconcelos in our notation. Consider a binary cost-sensitive learning problem with cost parameter c∈(0,1)c\in(0,1). The Bayes-optimal classifier for this task corresponds to f∗(x)=⟦η(x)>c⟧f^{*}(x)=\llbracket\eta(x)>c\rrbracket. The case c=0.5c=0.5 is the standard classification problem.

Suppose we wish to design a weighted, variable margin SVM for this task, i.e.,

where ω±1,δ±1≥0\omega_{\pm 1},\delta_{\pm 1}\geq 0. The conditional risk for this loss is

As this is a piecewise linear function, which is decreasing for f<−δ−1f<-\delta_{-1} and increasing for f>δ+1f>\delta_{+1}, the only possible minimum is at {δ+1,−δ−1}\{\delta_{+1},-\delta_{-1}\}. To ensure consistency, we seek the minimum to be δ+1\delta_{+1} iff η>c\eta>c. Observe that

Observe here that the margin terms δ±1\delta_{\pm 1} do not appear in the consistency condition: thus, as long as the weights are suitably chosen, any choice of margin terms will result in a consistent loss.

However, the margins do influence the form conditional Bayes risk: this is

For the purposes of normalisation, it is natural to require this function to attain a maximum at 11. This corresponds to choosing

In the class-imbalance setting, c=πc=\pi, and so we require

for consistency and normalisation respectively. This gives two degrees of freedom: the choice of ω+1\omega_{+1} (which determines ω−1\omega_{-1}), and then the choice of δ+1\delta_{+1} (which determines δ−1\delta_{-1}). For example, we could pick ω+1=1π\omega_{+1}=\frac{1}{\pi}, ω−1=11−π\omega_{-1}=\frac{1}{1-\pi}, δ+1=1\delta_{+1}=1, δ−1=π1−π\delta_{-1}=\frac{\pi}{1-\pi}.

To relate this to Masnadi-Shirazi and Vasconcelos , the latter considered separate costs C−1,C+1C_{-1},C_{+1} for a false positive and false negative respectively. With this, they suggested to use Masnadi-Shirazi and Vasconcelos [2010, Equation 34]

with δ+1=ed=1\delta_{+1}=\frac{e}{d}=1, d=ω+1=C+1d=\omega_{+1}=C_{+1}, a=ω−1=2C−1−1a=\omega_{-1}=2C_{-1}-1, and δ−1=ba=1a\delta_{-1}=\frac{b}{a}=\frac{1}{a}. The constraints C1≥2C−1−1C_{1}\geq 2C_{-1}-1 and C−1≥1C_{-1}\geq 1 are also enforced.

Under this setup, the cost ratio is C−1C−1+C+1\frac{C_{-1}}{C_{-1}+C_{+1}}. In the class-imbalance setting, we have C−1C−1+C+1=π\frac{C_{-1}}{C_{-1}+C_{+1}}=\pi, and so C+1=1−ππ⋅C−1C_{+1}=\frac{1-\pi}{\pi}\cdot C_{-1}. By the consistency condition, we have C+1=ω+1=1−ππ⋅ω−1=1−ππ⋅(2C−1−1)C_{+1}=\omega_{+1}=\frac{1-\pi}{\pi}\cdot\omega_{-1}=\frac{1-\pi}{\pi}\cdot(2C_{-1}-1). Thus, we must set C−1=1C_{-1}=1, and so C+1=1−ππC_{+1}=\frac{1-\pi}{\pi}. Thus, we obtain the parameters ω+1=1−ππ\omega_{+1}=\frac{1-\pi}{\pi}, ω−1=1\omega_{-1}=1, δ+1=1\delta_{+1}=1, δ−1=π1−π\delta_{-1}=\frac{\pi}{1-\pi}. By rescaling the weights, we obtain ω+1=1π\omega_{+1}=\frac{1}{\pi}, ω−1=11−π\omega_{-1}=\frac{1}{1-\pi}, δ+1=1\delta_{+1}=1, δ−1=π1−π\delta_{-1}=\frac{\pi}{1-\pi}. Observe that this is exactly one of the losses considered in Appendix B.1.

Appendix D Experimental setup

Intending a fair comparison, we use the same setup for all the methods for each dataset. All networks are trained with SGD with a momentum value of 0.9. Unless otherwise specified, linear learning rate warm-up is used in the first 5 epochs to reach the base learning rate, and a weight decay of 10−410^{-4} is used. Other dataset specific details are given below.

CIFAR-10 and CIFAR-100: We use a CIFAR ResNet-32 model trained for 200 epochs. The base learning rate is set to 0.1, which is decayed by 0.1 at the 160th epoch and again at the 180th epoch. Mini-batches of 128 images are used.

We also use the standard CIFAR data augmentation procedure used in previous works such as Cao et al. , He et al. , where 4 pixels are padded on each size and a random 32×3232\times 32 crop is taken. Images are horizontally flipped with a probability of 0.5.

ImageNet: We use a ResNet-50 model trained for 90 epochs. The base learning rate is 0.4, with cosine learning rate decay. We use a batch size of 512 and the standard data augmentation comprising of random cropping and flipping as described in Goyal et al. . Following Kang et al. , we use a weight decay of 5×10−45\times 10^{-4} on this dataset.

iNaturalist: We again use a ResNet-50 and train it for 90 epochs with a base learning rate of 0.4 and cosine learning rate decay. The data augmentation procedure is the same as the one used in ImageNet experiment above. We use a batch size of 512512.

Appendix E Additional experiments

we present results for CIFAR-10 and CIFAR-100 on the Step profile [Cao et al., 2019] with ρ=100\rho=100

we further verfiy that weight norms may not correlate with class priors under Adam

we include the results of post-hoc correction, and a breakdown of per-class errors, on ImageNet-LT

Table 5 summarises results on the Step-100 profile. Here, with τ=1\tau=1, weight normalisation slightly outperforms logit adjustment. However, with τ>1\tau>1, logit adjustment is again found to be superior (54.80); see Figure 8.

E.2 Per-class errors on ImageNet-LT

Figure 9 breaks down the per-class accuracies on ImageNet-LT. As before, the logit adjustment procedure shows significant gains on rarer classes.

E.3 Post-hoc correction on ImageNet-LT

Figure 10 compares post-hoc correction techniques as the scaling parameter τ\tau is varied on ImageNet-LT. As before, logit adjustment with suitable tuning is seen to be competitive with weight normalisation.

E.4 Per-group errors

Following Liu et al. , Kang et al. , we additionally report errors on a per-group basis, where we construct three groups of classes: “Many”, comprising those with at least 100 training examples; “Medium”, comprising those with at least 20 and at most 100 training examples; and “Few”, comprising those with at most 20 training examples. This is a coarser level of granularity than the grouping employed in the previous section, and the body. Figure 11 shows that the logit adjustment procedure shows consistent gains over all three groups.

Appendix F Does weight normalisation increase margins?

Suppose that one uses SGD with a momentum, and finds solutions where ∥wy∥2\|w_{y}\|_{2} tracks the class priors. One intuition behind normalisation of weights is that, drawing inspiration from the binary case, this ought to increase the classification margins for tail classes.

This generalises the classical binary margin, wherein by convention Y={±1}\mathscr{Y}=\{\pm 1\}, w−1=−w1w_{-1}=-w_{1}, and

which agrees with (16) upto scaling. One may also define the geometric margin in the binary case to be the distance of (x,y)(x,y) from its classifier:

Clearly, γg,b(x)=∣γf(x,y)∣∥w1∥2\gamma_{g,{\rm b}}(x)=\frac{|\gamma_{\rm f}(x,y)|}{\|w_{1}\|_{2}}, and so for fixed functional margin, one may increase the geometric margin by minimising ∥w1∥2\|w_{1}\|_{2}. However, the same is not necessarily true in the multiclass setting, since here the functional and geometric margins do not generally align [Tatsumi et al., 2011, Tatsumi and Tanino, 2014]. In particular, controlling each ∥wy∥2\|w_{y}\|_{2} does not necessarily control the geometric margin.

Appendix G Bayes-optimal classifier under Gaussian class-conditionals

for suitable μy\mu_{y} and σ\sigma. Then,

Now use the fact that in our setting, ∥μ+1∥2=∥μ−1∥2\|\mu_{+1}\|^{2}=\|\mu_{-1}\|^{2}.

We remark also that the class-probability function is