Excessive Invariance Causes Adversarial Vulnerability

Jörn-Henrik Jacobsen, Jens Behrmann, Richard Zemel, Matthias Bethge

Introduction

Adversarial vulnerability is one of the most iconic failure cases of modern machine learning models (Szegedy et al., 2013) and a prime example of their weakness in out-of-distribution generalization. It is particularly striking that under i.i.d. settings deep networks show superhuman performance on many tasks (LeCun et al., 2015), while tiny targeted shifts of the input distribution can cause them to make unintuitive mistakes. The reason for these failures and how they may be avoided or at least mitigated is an active research area (Schmidt et al., 2018; Gilmer et al., 2018b; Bubeck et al., 2018).

So far, the study of adversarial examples has mostly been concerned with the setting of small perturbation, or ϵ\epsilon-adversaries (Goodfellow et al., 2015; Madry et al., 2017; Raghunathan et al., 2018).

Perturbation-based adversarial examples are appealing because they allow to quantitatively measure notions of adversarial robustness (Brendel et al., 2018). However, recent work argued that the perturbation-based approach is unrealistically restrictive and called for the need of generalizing the concept of adversarial examples to the unrestricted case, including any input crafted to be misinterpreted by the learned model (Song et al., 2018; Brown et al., 2018). Yet, settings beyond ϵ\epsilon-robustness are hard to formalize (Gilmer et al., 2018a).

We argue here for an alternative, complementary viewpoint on the problem of adversarial examples. Instead of focusing on transformations erroneously crossing the decision-boundary of classifiers, we focus on excessive invariance as a major cause for adversarial vulnerability. To this end, we introduce the concept of invariance-based adversarial examples and show that class-specific content of almost any input can be changed arbitrarily without changing activations of the network, as illustrated in figure 1 for ImageNet. This viewpoint opens up new directions to analyze and control crucial aspects underlying vulnerability to unrestricted adversarial examples.

The invariance perspective suggests that adversarial vulnerability is a consequence of narrow learning, yielding classifiers that rely only on few highly predictive features in their decisions. This has also been supported by the observation that deep networks strongly rely on spectral statistical regularities (Jo & Bengio, 2017), or stationary statistics (Gatys et al., 2017) to make their decisions, rather than more abstract features like shape and appearance. We hypothesize that a major reason for this excessive invariance can be understood from an information-theoretic viewpoint of cross-entropy, which maximizes a bound on the mutual information between labels and representation, giving no incentive to explain all class-dependent aspects of the input. This may be desirable in some cases, but to achieve truly general understanding of a scene or an object, machine learning models have to learn to successfully separate essence from nuisance and subsequently generalize even under shifted input distributions.

We identify excessive invariance underlying striking failures in deep networks and formalize the connection to adversarial examples.

We show invariance-based adversarial examples can be observed across various tasks and types of deep network architectures.

We propose an invertible network architecture that gives explicit access to its decision space, enabling class-specific manipulations to images while leaving all dimensions of the representation seen by the final classifier invariant.

From an information-theoretic viewpoint, we identify the cross-entropy objective as a major reason for the observed failures. Leveraging invertible networks, we propose an alternative objective that provably reduces excessive invariance and works well in practice.

Two Complementary Approaches to Adversarial Examples

In this section, we define pre-images and establish a link to adversarial examples.

where (i) ⊂\subset (ii) ⊂\subset (iii) by the compositional nature of DD. Moreover, the (sub-)network is invariant to perturbations Δx\Delta x which satisfy x∗=x+Δxx^{*}=x+\Delta x.

Non-trivial pre-images (pre-images containing more elements than input xx) after the ii-th layer occur if the chain fi∘⋯∘f1f_{i}\circ\cdots\circ f_{1} is not injective, for instance due to subsampling or non-injective activation functions like ReLU (Behrmann et al., 2018a). This accumulated invariance can become problematic if not controlled properly, as we will show in the following.

We define perturbation-based adversarial examples by introducing the notion of an oracle (e.g., a human decision-maker or the unknown input-output function considered in learning theory):

Usually, such examples are constructed as ϵ\epsilon-bounded adversarial examples (Goodfellow et al., 2015). However, as our goal is to characterize general invariances of the network, we do not restrict ourselves to bounded perturbations.

Let GG denote the ii-th layer, logits or the classifier (Definition 1) and let x∗≠xx^{*}\neq x be in the GG pre-image of xx and and oo an oracle (Definition 2). Then, an invariance-based adversarial example fulfills o(x)≠o(x∗)o(x)\neq o(x^{*}), while G(x)=G(x∗)G(x)=G(x^{*}) (and hence D(x)=D(x∗)D(x)=D(x^{*})).

Intuitively, adversarial perturbations cause the output of the classifier to change while the oracle would still consider the new input x∗x^{*} as being from the original class. Hence in the context of ϵ\epsilon-bounded perturbations, the classifier is too sensitive to task-irrelevant changes. On the other hand, movements in the pre-image leave the classifier invariant. If those movements induce a change in class as judged by the oracle, we call these invariance-based adversarial examples. In this case, however, the classifier is too insensitive to task-relevant changes. In conclusion, these two modes are complementary to each other, whereas both constitute failure modes of the learned classifier.

When not restricting to ϵ\epsilon-perturbations, perturbation-based and invariance-based adversarial examples yield the same input x∗x^{*} via

with different reference points x1x_{1} and x2x_{2}, see Figure 2. Hence, the key difference is the change of reference, which allows us to approach these failure modes from different directions. To connect these failure modes with an intuitive understanding of variations in the data, we now introduce the notion of invariance to nuisance and semantic variations, see also (Achille & Soatto, 2018).

For example, such a nuisance perturbation could be a translation or occlusion in image classification. Further in Appendix A, we discuss the synthetic example called Adversarial Spheres from (Gilmer et al., 2018b), where nuisance and semantics can be explicitly formalized as rotation and norm scaling.

Bijective classifiers with simplified readout. We build deep networks that give access to their decision space by removing the final linear mapping onto the class probes in invertible RevNet-classifiers and call these networks fully invertible RevNets. The fully invertible RevNet classifier can be written as Dθ=arg max⁡k=1,…,C softmax(Fθ(x)k)D_{\theta}=\operatorname*{arg\,max}_{k=1,\ldots,C}\,softmax(F_{\theta}(x)_{k}), where FθF_{\theta} represents the bijective network. We denote z=Fθ(x)z=F_{\theta}(x), zs=z1,...,Cz_{s}=z_{1,...,C} as the logits (semantic variables) and zn=zC+1,...,dz_{n}=z_{C+1,...,d} as the nuisance variables (znz_{n} is not used for classification). In practice we choose the first C indices of the final zz tensor or apply a more sophiscticated DCT scheme (see appendix D) to set the subspace zsz_{s}, but other choices work as well. The architecture of the network is similar to iRevNets (Jacobsen et al., 2018) with some additional Glow components like actnorm (Kingma & Dhariwal, 2018), squeezing, dimension splitting and affine block structure (Dinh et al., 2017), see Figure 3 for a graphical description. As all components are common in the bijective network literature, we refer the reader to Appendix D for exact training and architecture details. Due to its simple readout structure, the resulting invertible network allows to qualitatively and quantitatively investigate the task-specific content in nuisance and logit variables. Despite this restriction, we achieve performance on par with commonly-used baselines on MNIST and ImageNet, see Table 1 and Appendix D.

Attack on adversarial spheres. First, we evaluate our analytic attack on the synthetic spheres dataset, where the task is to classify samples as belonging to one out of two spheres with different radii. We choose the sphere dimensionality to be d=100d=100 and the radii: R1=1R_{1}=1, R2=10R_{2}=10. By training a fully-connected fully invertible RevNet, we obtain 100% accuracy. After training we visualize the decision-boundaries of the original classifier DD and a posthoc trained classifier on znz_{n} (nuisance classifier), see Figure 4. We densely sample points in a 2D subspace, following Gilmer et al. (2018b), to visualize two cases: 1) the decision-boundary on a 2D plane spanned by two randomly chosen data points, 2) the decision-boundary spanned by metameric sample xmetx_{met} and reference point xx. In the metameric sample subspace we identify excessive invariance of the classifier. Here, it is possible to move any point from the inner sphere to the outer sphere without changing the classifiers predictions. However, this is not possible for the classifier trained on znz_{n}. Most notably, the visualized failure is not due to a lack of data seen during training, but rather due to excessive invariance of the original classifier DD on zsz_{s}. Thus, the nuisance classifier on znz_{n} does not exhibit the same adversarial vulnerability in its subspace.

Attack on MNIST and ImageNet. After validating its potential to uncover adversarial subspaces, we apply metameric sampling to fully invertible RevNets trained on MNIST and Imagenet, see Figure 5. The result is striking, as the nuisance variables znz_{n} are dominating the visual appearance of the logit metamers, making it possible to attach any semantic content to any logit activation pattern. Note that the entire 1000-dimensional feature vector containing probabilities over all ImageNet classses remains unchanged by any of the transformations we apply. To show our findings are not a particular property of bijective networks, we attack an ImageNet trained ResNet152 with a gradient-based version of our metameric attack, also known as feature adversaries (Sabour et al., 2016). The attack minimizes the mean squared error between a given set of logits from one image to another image (see appendix B for details). The attack shows the same failures for non-bijective models. This result highlights the general relevance of our finding and poses the question of the origin of this excessive invariance, which we will analyze in the following section.

Overcoming Insufficiency of Crossentropy-based Information-maximization

In this section we identify why the cross-entropy objective does not necessarily encourage to explain all task-dependent variations of the data and propose a way to fix this. As shown in figure 4, the nuisance classifier on znz_{n} uses task-relevant information not captured by the logit classifier DθD_{\theta} on zsz_{s} (evident by its superior performance in the adversarial subspace).

We leverage the simple readout-structure of our invertible network and turn this observation into a formal explanation framework using information theory: Let (x,y)∼D(x,y)\sim\mathcal{D} with labels y∈{0,1}Cy\in\{0,1\}^{C}. Then the goal of a classifier can be stated as maximizing the mutual information (Cover & Thomas, 2006) between semantic features zsz_{s} (logits) extracted by network FθF_{\theta} and labels yy, denoted by I(y;zs)I(y;z_{s}).

Adversarial distribution shift. As the previously discussed failures required to modify input data from distribution D\mathcal{D}, we introduce the concept of an adversarial distribution shift DAdv≠D\bm{\mathcal{D}_{Adv}\neq\mathcal{D}} to formalize these modifications. Our first assumptions for DAdv\mathcal{D}_{Adv} is IDAdv(zn;y)≤ID(zn;y)I_{\mathcal{D}_{Adv}}(z_{n};y)\leq I_{\mathcal{D}}(z_{n};y). Intuitively, the nuisance variables znz_{n} of our network do not become more informative about yy. Thus, the distribution shift may reduce the predictiveness of features encoded in zsz_{s}, but does not introduce or increase the predictive value of variations captured in znz_{n}. Second, we assume IDAdv(y;zs∣zn)≤IDAdv(y;zs)I_{\mathcal{D}_{Adv}}(y;z_{s}|z_{n})\leq I_{\mathcal{D}_{Adv}}(y;z_{s}), which corresponds to positive or zero interaction information, see e.g. (Ghassami & Kiyavash, 2017). While the information in zsz_{s} and znz_{n} can be redundant in this assumption, synergetic effects where conditioning on znz_{n} increase the mutual information between yy and zsz_{s} are excluded.

Bijective networks FθF_{\theta} capture all variations by design which translates to information preservation I(y;x)=I(y;Fθ(x))I(y;x)=I(y;F_{\theta}(x)), see (Kraskov et al., 2004). Consider the reformulation

by the chain rule of mutual information (Cover & Thomas, 2006), where I(y;zn∣zs)I(y;z_{n}|z_{s}) denotes the conditional mutual information. Most strikingly, equation 5 offers two ways forward:

Indirect increase of I(y;zs∣zn)I(y;z_{s}|z_{n}) via decreasing I(y;zn)I(y;z_{n}).

Usually in a classification task, only I(y;zs)I(y;z_{s}) is increased actively via training a classifier. While this approach is sufficient in most cases, expressed via high accuracies on training and test data, it may fail under DAdv\mathcal{D}_{Adv}. This highlights why cross-entropy training may not be sufficient to overcome excessive semantic invariance. However, by leveraging the bijection FθF_{\theta} we can minimize the unused information I(y;zn)I(y;z_{n}) using the intuition of a nuisance classifier.

The underlying principles of the nuisance classification loss LnCE\mathcal{L}_{nCE} can be understood using a variational lower bound on mutual information from Barber & Agakov (2003). In summary, the minimization is with respect to a lower bound on ID(y;zn)I_{\mathcal{D}}(y;z_{n}), while the maximization aims to tighten the bound (see Lemma 10 in Appendix C). By using these results, we now state the main result under the assumed distribution shift and successful minimization (proof in Appendix C.1):

Let DAdv\mathcal{D}_{Adv} denote the adversarial distribution and D\mathcal{D} the training distribution. Assume ID(y;zn)=0I_{\mathcal{D}}(y;z_{n})=0 by minimizing LiCE\mathcal{L}_{iCE} and the distribution shift satisfies IDAdv(zn;y)≤ID(zn;y)I_{\mathcal{D}_{Adv}}(z_{n};y)\leq I_{\mathcal{D}}(z_{n};y) and IDAdv(y;zs∣zn)≤IDAdv(y;zs)I_{\mathcal{D}_{Adv}}(y;z_{s}|z_{n})\leq I_{\mathcal{D}_{Adv}}(y;z_{s}). Then,

Thus, incorporating the nuisance classifier allows for the discussed indirect increase of IDAdv(y;zs)I_{\mathcal{D}_{Adv}}(y;z_{s}) under an adversarial distribution shift, visualized in Figure 6.

To aid stability and further encourage factorization of zsz_{s} and znz_{n} in practice, we add a maximum likelihood term to our independence cross-entropy objective as

where det(Jθx)\text{det}(J_{\theta}^{x}) denotes the determinant of the Jacobian of Fθ(x)F_{\theta}(x) and pk∼N(βk,γk)p_{k}\sim\mathcal{N}(\beta_{k},\gamma_{k}) with βk,γk\beta_{k},\gamma_{k} learned parameter. The log-determinant can be computed exactly in our model with negligible additional cost. Note, that optimizing LMLEn\mathcal{L}_{MLE_{n}} on the nuisance variables together with LsCE\mathcal{L}_{sCE} amounts to maximum-likelihood under a factorial prior (see Lemma 11 in Appendix C).

Just as in GANs the quality of the result relies on a tight bound provided by the nuisance classifier and convergence of the MLE term. Thus, it is important to analyze the success of the objective after training. We do this by applying our metameric sampling attack, but there are also other ways like evaluating a more powerful nuisance classifier after training.

Applying Independence Cross-entropy

In this section, we show that our proposed independence cross-entropy loss is effective in reducing invariance-based vulnerability in practice by comparing it to vanilla cross-entropy training in four aspects: (1) error on train and test set, (2) effect under distribution shift, perturbing nuisances via metameric sampling, (3) evaluate accuracy of a classifier on the nuisance variables to quantify the class-specific information in them and (4) on our newly introduced shiftMNIST, an augmented version of MNIST to benchmark adversarial distribution shifts according to Theorem 6.

For all experiments we use the same network architecture and settings, the only difference being the two additional loss terms as explained in Definition 5 and equation 6. In terms of test error of the logit classifier, both losses perform approximately on par, whereas the gap between train and test error vanishes for our proposed loss function, indicating less overfitting. For classification errors see Table 2 in appendix D.

Robustness under metameric sampling attack. To analyze if our proposed loss indeed leads to independence between znz_{n} and labels yy, we attack it with our metameric sampling procedure. As we are only looking on data samples and not on samples from the model (factorized gaussian on nuisances), this attack should reveal if the network learned to trick the objective. In Figure 7 we show interpolations between original images and logit metamers in CE- and iCE-trained fully invertible RevNets. In particular, we are holding the activations zsz_{s} constant, while linearly interpolating nuisances znz_{n} down the column. The CE-trained network allows us to transform any image into any class without changing the logits. However, when training with our proposed iCE, the picture changes fundamentally and interpolations in the pre-image only change the style of a digit, but not its semantic content. This shows our loss has the ability to overcome excessive task-related invariance and encourages the model to explain and separate all task-related variability of the input from the nuisances of the task.

A classifier trained on the nuisance variables of the cross-entropy trained model performs even better than the logit classifier. Yet, a classifier on the nuisances of the independence cross-entropy trained model is performing poorly (Table 2 in appendix D). This indicates little class-specific information in the nuisances znz_{n}, as intended by our objective function. Note also that this inability of the nuisance classifier to decode class-specific information is not due to it being hard to read out from znz_{n}, as this would be revealed by the metameric sampling attack (see Figure 7).

shiftMNIST: Benchmarking adversarial distribution shift. To further test the efficacy of our proposed independence cross-entropy, we introduce a simple, but challenging new dataset termed shiftMNIST to test classifiers under adversarial distribution shifts DAdv\mathcal{D}_{Adv}. The dataset is based on vanilla MNIST, augmented by introducing additional, highly predictive features at train time that are randomized or removed at test time. Randomization or removal ensures that there are no synergy effects between digits and planted features under DAdv\mathcal{D}_{Adv}. This setup allows us to reduce mutual information between category and the newly introduced feature in a targeted manner. (a) Binary shiftMNIST is vanilla MNIST augmented by coding the category for each digit into a single binary pixel scheme. The location of the binary pixel reveals the category of each image unambigiously, while only minimally altering the image’s appearance.

At test time, the binary code is not present and the network can not rely on it anymore. (b) Textured shiftMNIST introduces textured backgrounds for each digit category which are patches sampled from the describable texture dataset (Cimpoi et al., 2014).

At train time the same type of texture is underlayed each digit of the same category, while texture types across categories differ. At test time, the relationship is broken and texture backgrounds are paired with digits randomly, again minimizing the mutual information between background and label in a targeted manner. See Figure 8 for examplesLink to code and dataset: https://github.com/jhjacobsen/fully-invertible-revnet.

It turns out that this task is indeed very hard for standard classifiers and their tendency to become excessively invariant to semantically meaningful features, as predicted by our theoretical analysis. When trained with cross-entropy, ResNets and fi-RevNets make zero errors on the train set, while having error rates of up to 87% on the shifted test set. This is striking, given that e.g. in binary shiftMNIST, only one single pixel is removed under DAdv\mathcal{D}_{Adv}, leaving the whole image almost unchanged. When applying our independence cross-entropy, the picture changes again. The errors made by the network improve by up to almost 38% on binary shiftMNIST and around 28% on textured shiftMNIST. This highlights the effectiveness of our proposed loss function and its ability to minimize catastrophic failure under severe distribution shifts exploiting excessive invariance.

Related Work

Adversarial examples. Adversarial examples often include ϵ\epsilon-norm restrictions (Szegedy et al., 2013), while (Gilmer et al., 2018a) argue for a broader definition to fully capture the implications for security. The ϵ\epsilon-adversarial examples have also been extended to ϵ\epsilon-feature adversaries (Sabour et al., 2016), which are equivalent to our approximate metameric sampling attack. Some works (Song et al., 2018; Fawzi et al., 2018) consider unrestricted adversarial examples, which are closely related to invariance-based adversarial vulnerability. The difference to human perception revealed by adversarial examples fundamentally questions which statistics deep networks use to base their decisions (Jo & Bengio, 2017; Tsipras et al., 2019).

Relationship between standard and bijective networks. We leverage recent advances in reversible (Gomez et al., 2017) and bijective networks (Jacobsen et al., 2018; Ardizzone et al., 2019; Kingma & Dhariwal, 2018) for our analysis. It has been shown that ResNets and iRevNets behave similarly on various levels of their representation on challenging tasks (Jacobsen et al., 2018) and that iRevNets as well as Glow-type networks are related to ResNets by the choice of dimension splitting applied in their residual blocks (Grathwohl et al., 2019). Perhaps unsurprisingly, given so many similarities, ResNets themselves have been shown to be provably bijective under mild conditions (Behrmann et al., 2018b). Further, excessive invariance of the type we discuss here has been shown to occur in non residual-type architectures as well (Gilmer et al., 2018b; Behrmann et al., 2018a). For instance, it has been observed that up to 60% of semantically meaningful input dimensions on the adversarial spheres problem are learned to be ignored, while retaining virtually perfect performance (Gilmer et al., 2018b). In summary, there is ample evidence that RevNet-type networks are closely related to ResNets, while providing a principled framework to study widely observed issues related to excessive invariance in deep learning in general and adversarial robustness in particular.

Information theory. The information-theoretic view has gained recent interest in machine learning due to the information bottleneck (Tishby & Zaslavsky, 2015; Shwartz-Ziv & Tishby, 2017; Alemi et al., 2017) and usage in generative modelling (Chen et al., 2016; Hjelm et al., 2019). As a consequence, the estimation of mutual information (Barber & Agakov, 2003; Alemi et al., 2018; Achille & Soatto, 2018; Belghazi et al., 2018) has attracted growing attention. The concept of group-wise independence between latent variables goes back to classical independent subspace analysis (Hyvärinen & Hoyer, 2000) and received attention in learning unbiased representations, e.g. see the Fair Variational Autoencoder (Louizos et al., 2015). Furthermore, extended cross-entropy losses via entropy terms (Pereyra et al., 2017) or minimizing predictability of variables (Schmidhuber, 1991) has been introduced for other applications. Our proposed loss also shows similarity to the GAN loss (Goodfellow et al., 2014). However, in our case there is no notion of real or fake samples, but exploring similarities in the optimization are a promising avenue for future work.

Conclusion

Failures of deep networks under distribution shift and their difficulty in out-of-distribution generalization are prime examples of the limitations in current machine learning models. The field of adversarial example research aims to close this gap from a robustness point of view. While a lot of work has studied ϵ\epsilon-adversarial examples, recent trends extend the efforts towards the unrestricted case. However, adversarial examples with no restriction are hard to formalize beyond testing error. We introduce a reverse view on the problem to: (1) show that a major cause for adversarial vulnerability is excessive invariance to semantically meaningful variations, (2) demonstrate that this issue persists across tasks and architectures; and (3) make the control of invariance tractable via fully-invertible networks.

In summary, we demonstrated how a bijective network architecture enables us to identify large adversarial subspaces on multiple datasets like the adversarial spheres, MNIST and ImageNet. Afterwards, we formalized the distribution shifts causing such undesirable behavior via information theory. Using this framework, we find one of the major reasons is the insufficiency of the vanilla cross-entropy loss to learn semantic representations that capture all task-dependent variations in the input. We extend the loss function by components that explicitly encourage a split between semantically meaningful and nuisance features. Finally, we empirically show that this split can remove unwanted invariances by performing a set of targeted invariance-based distribution shift experiments.

Acknowledgements

We thank Ryota Tomioka for spotting a mistake in the proof for Theorem 6. We thank thank the anonymous reviewers, Ricky Chen, Will Grathwohl and Jesse Bettencourt for helpful comments on the manuscript. We gratefully acknowledge the financial support from the German Science Foundation for the CRC 1233 on ”Robust Vision” and RTG 2224 ”π3\pi^{3}: Parameter Identification - Analysis, Algorithms, Applications”. This work was supported by the German Federal Ministry of Education and Research (BMBF): Tübingen AI Center, FKZ: 01IS18039A.

References

Appendix A Semantic and Nuisance Variation on Adversarial Spheres

Consider classifying inputs xx from two classes given by radii R1R_{1} or R2R_{2}. Further, let (r,ϕ)(r,\phi) denote the spherical coordinates of xx. Then, any perturbation Δx\Delta x, x∗=x+Δxx^{*}=x+\Delta x with r∗≠rr^{*}\neq r is semantic. On the other hand, if r∗=rr^{*}=r the perturbation is a nuisance with respect to the task of discriminating two spheres.

In this example, the max-margin classifier D(x)=sign(∥x∥−R1+R22)D(x)=sign\left(\|x\|-\frac{R_{1}+R_{2}}{2}\right) is invariant to any nuisance perturbation, while being only sensitive to semantic perturbations. In summary, the transform to spherical coordinates allows to linearize semantic and nuisance perturbations. Using this notion, invariance-based adversarial examples can be attributed to perturbations of x∗=x+Δxx^{*}=x+\Delta x with following two properties

Perturbation Δx\Delta x is semantic, as o(x)≠o(x+Δx)o(x)\neq o(x+\Delta x).

Thus, the failure of the classifier DD can be thought of a mis-alignment between its invariance (expressed through the pre-image) and the semantics of the data and task (expressed by the oracle).

which computes the norm of xx from its first d−1d-1 cartesian-coordinates. Then, DD is invariant to a semantic perturbation with Δr=R2−R1\Delta r=R_{2}-R_{1} if only changes in the last coordinate xdx_{d} are made.

We empirically evaluate the classifier in equation 7 on the spheres problem (10M/2M samples setting (Gilmer et al., 2018b)) and validate that it can reach perfect classification accuracy. However, by construction, perturbing the invariant dimension xd∗=xd+Δxdx^{*}_{d}=x_{d}+\Delta x_{d} allows us to move all samples from the inner sphere to the outer sphere. Thus, the accuracy of the classifier drops to chance level when evaluating its performance under such a distributional shift. To conclude, this underlines how classifiers with optimal performance on finite samples can exhibit non-intuitive failure modes due to excessive invariance with respect to semantic variations.

Appendix B Approximate Gradient-based Metameric Samples

in the 1000-dimensional semantic logit space via stochastic gradient descent. We optimize with Adam in Pytorch default settings and a learning rate of 0.01 for 3000 iterations. The optimization thus takes the form of an adversarial attack targeting all logit entries and with no norm restriction on the input distance. Note that our metameric sampling attack in bijective networks is the analytic reverse equivalent of this attack. It leads to the exact solution at the cost of one inverse pass instead of an approximate solution here at the cost of thousands of gradient steps.

Appendix C Information Theory

Computing mutual information is often intractable as it requires the joint probability p(x,y)p(x,y), see (Cover & Thomas, 2006) for an extensive treatment of information theory. However, following variational lower bound can be used for approximation, see (Barber & Agakov, 2003).

Let X,YX,Y be random variables with conditional density p(y∣x)p(y|x). Further, let qθ(y∣x)q_{\theta}(y|x) be a variational density depending on parameter θ\theta. Then, the lower bound

holds with equality if p(y∣x)=qθ(y∣x)p(y|x)=q_{\theta}(y|x).

Define semantics as zs=Fθ(x)1,…,Cz_{s}=F_{\theta}(x)_{1,\dots,C} and nuisances as zn=Fθ(x)C+1,…,dz_{n}=F_{\theta}(x)_{C+1,\dots,d}, where (x,y)∼D(x,y)\sim\mathcal{D}. Then, the nuisance classification loss yields

Minimization of lower bound on ID(y;zn)\bm{I_{\mathcal{D}}(y;z_{n})}: θ∗=arg min⁡θLnCE(θ,θnc∗)\theta^{*}=\operatorname*{arg\,min}_{\theta}\mathcal{L}_{nCE}(\theta,\theta_{nc}^{*}) minimizes Iθnc∗(y;zn)I_{\theta_{nc}^{*}}(y;z_{n}), where Iθnc∗(y;zn)≤ID(y;zn)I_{\theta_{nc}^{*}}(y;z_{n})\leq I_{\mathcal{D}}(y;z_{n}) and θnc∗=arg max⁡θ2LnCE(θ,θnc)\theta_{nc}^{*}=\operatorname*{arg\,max}_{\theta_{2}}\mathcal{L}_{nCE}(\theta,\theta_{nc}).

Maximization to tighten bound on ID(y;zn)\bm{I_{\mathcal{D}}(y;z_{n})}: Under a perfect model of the conditional density, Dθnc∗(zn)=p(y∣zn)D_{\theta_{nc}^{*}}(z_{n})=p(y|z_{n}), it holds Iθnc∗(y;zn)=ID(y;zn).I_{\theta_{nc}^{*}}(y;z_{n})=I_{\mathcal{D}}(y;z_{n}).

To proof above result, we need to draw the connection to the variational lower bound on mutual information from Lemma 9. Let the nuisance classifier Dθnc(zn)D_{\theta_{nc}}(z_{n}) model the variational posterior qθnc(y∣zn)q_{\theta_{nc}}(y|z_{n}). Then we have the lower bound

From Lemma 9 follows, that if Dθnc(zn)=p(y∣zn)D_{\theta_{nc}}(z_{n})=p(y|z_{n}), it holds I(y;zn)=Iθnc(y;zn)I(y;z_{n})=I_{\theta_{nc}}(y;z_{n}). Hence, the nuisance classifier needs to model the conditional density perfectly. Estimating this bound via Monte Carlo simulation requires sampling from the conditional density p(y∣zn)p(y|z_{n}). Following (Alemi et al., 2017), we have the Markov property y↔x↔zny\leftrightarrow x\leftrightarrow z_{n} as labels yy interact with inputs xx and representation znz_{n} interacts with inputs xx. Hence,

Including above and assuming Fθ(x)=znF_{\theta}(x)=z_{n} to be a deterministic function, we have

Define semantics as zs=Fθ(x)1,…,Cz_{s}=F_{\theta}(x)_{1,\dots,C} and nuisances as zn=Fθ(x)C+1,…,dz_{n}=F_{\theta}(x)_{C+1,\dots,d}, where (x,y)∼D(x,y)\sim\mathcal{D}. Then, the MLE-term in equation 6 together with cross-entropy on the semantics

minimizes the mutual information I(zs;zn)I(z_{s};z_{n}).

From the assumptions follows IDAdv(y;zn)=0I_{\mathcal{D}_{Adv}}(y;z_{n})=0. Furthermore, we have the assumption

excluding synergetic effects in the interaction information (Ghassami & Kiyavash, 2017). By information preservation under homeomorphisms (Kraskov et al., 2004) and the chain rule of mutual information (Cover & Thomas, 2006), we have

As zs=F(x)1,…,Cz_{s}=F(x)_{1,\ldots,C} is obtained by the deterministic transform FF, by the data processing inequality (Cover & Thomas, 2006) we have the inequality IDAdv(y;x)≥IDAdv(y;zs)I_{\mathcal{D}_{Adv}}(y;x)\geq I_{\mathcal{D}_{Adv}}(y;z_{s}). Thus, the claimed equality must hold.

C.2 Mutual information bounded

Since our goal is to maximize the mutual information I(y;zs)I(y;z_{s}) while minimizing I(y;zn)I(y;z_{n}), we need to ensure that this objective is well defined as mutual information can be unbounded from above for continuous random variables. However, due to the data processing inequality (Cover & Thomas, 2006) we have I(y;zn)=I(y;Fθ(x))≤I(y;x)I(y;z_{n})=I(y;F_{\theta}(x))\leq I(y;x). Hence, we have a fixed upper bound given by our data (x,y)(x,y). Compared to (Belghazi et al., 2018) there is thus no need for gradient clipping or a switch to the bounded Jensen-Shannon divergence as in (Hjelm et al., 2019) is not necessary.

Appendix D Training and Architectural Details

All experiments were based on a fully invertible RevNet model with different hyperparameters for each dataset. For the spheres experiment we used Pytorch (Paszke et al., 2017) and for MNIST, as well as Imagenet Tensorflow (Abadi et al., 2016).

The network is a fully connected fully invertible RevNet. It has 4 RevNet-type ReLU bottleneck blocks with additive couplings and uses no batchnorm. We train it via cross-entropy and use the Adam optimizer (Kingma & Ba, 2014) with a learning rate of 0.0001 and otherwise default Pytorch settings. The nuisance classifier is a 3 layer ReLU network with 1000 hidden units per layer.

We choose the spheres to be 100-dimensional, with R1=1R_{1}=1 and R2=10R_{2}=10, train on 500k samples for 10 epochs and then validate on another 100k holdout set. We achieve 100% train and validation accuracy for logit and nuisance classifier.

D.2 MNIST experiments

We use a convolutional fully invertible RevNet with additional actnorm and invertible 1x1 convolutions between each layer as introduced in Kingma & Dhariwal (2018). The network has 3 stages, after which half of the variables are factored out and an invertible downsampling, or squeezing (Dinh et al., 2017; Jacobsen et al., 2018) is applied. The network has 16 RevNet blocks with batch norm per stage and 128 filters per layer. We also dequantize the inputs as is typically done in flow-based generative models.

The network is trained via Adamax (Kingma & Ba, 2014) with a base learning rate of 0.001 for 100 epochs and we multiply the it with a factor of 0.2 every 30 epochs and use a batch size of 64 and l2 weight decay of 1e-4. For training we compare vanilla cross-entropy training with our proposed independence cross-entropy loss. To have a more balanced loss signal, we normalize LnCE\mathcal{L}_{nCE} by the number of input dimensions it receives for the maximization step. The nuisance classifier is a fully-connected 3 layer ReLU network with 512 units. As data-augmentation we use random shifts of 3 pixels. For classification errors of the different architectures we compare, see Table 2.

D.3 ImageNet experiments

We use a convolutional fully invertible RevNet with 4 stages, 4 RevNet blocks per stage and invertible downsampling after each stage, as well as two invertible downsamplings on the input of the network. The first three stages consist of additive and the last of affine coupling layers. After the final layer we apply an orthogonal 2D DCT type-II to all feature maps and read out the classes in the low-pass components of the transformation. This effectively gives us an invertible global average pooling and makes our network even more similar to ResNets, that always apply global average pooling on their final feature maps. We train the network with momentum SGD for 128 epochs, a batch size of 480 (distributed to 6 GPUs), a base learning rate of 0.1, which is reduced by a factor of 0.1 every 32 epochs. We apply momentum of 0.9 and l2 weight decay of 1e-4.