DisCo Fever: Robust Networks Through Distance Correlation

Gregor Kasieczka, David Shih

Appendix A More details on the methods

One method to reduce correlation is to remove discriminating information carried by a variable. The approach of giving weights to training events so the distributions for different classes are identical has been long used experimentally See e.g. Reweight2011; JME15002; ATLASRun1W; ATLASRun1Top and recently was studied for understanding network decisions Planing and resonance tagging BryanOverview. Specifically a weight wi,Cw_{i,C} for event with index ii of class CC is calculated by building a histogram of the feature xx so that njn_{j} denotes the number of events in bin jj. Due to the explicit use of histogramming, it can be difficult to generalise planing to multiple variables. The weight can then be calculated as:

where AA is a per-class normalisation factor.

Planing weights are then used in the training of an e.g. neural network classifier and modify the contribution of each event to the loss function. When applying the algorithm to events of unknown class in the testing phase no weights are used (i.e. weights are set equal to one).

A.2 Designed decorrelated taggers

For decorrelating a classifier for a single selection efficiency, a transformation of the output using the expected shape of the background distribution after the training is completed is possible as well DDT. This approach is named Designed decorrelated taggers (DDT). Concretely, to decorrelate feature yy against xx, it is transformed according to:

where OO in an offset and MM is a slope parameter dydx\frac{dy}{dx} extracted for the background.

A.3 Fixed efficiency regression

It is also possible to design decorrelated variables for non-linear relations between features by subtracting the expected response for background examples ATLASDecor. This average response can also be parametrised against multiple features. Take for example the de-correlation of a feature yy against xx and x′x^{\prime}.

The decorrelated yk-NNy^{\textrm{k-NN}} can be calculated as

with the threshold y(P  %)(x,x′)y^{(P\;\%)}(x,x^{\prime}) corresponding to a true positive rate for background events PP interpoldated using a kk-nearest neighbour regression fit kNN.

A.4 uBoost

The uBoost approach is a modified training methods for boosted decision trees (BDTs). A decision tree is a series of binary selection criteria that subsequently divide the data. Boosting refers to a combination of multiple decision trees to maximise a chosen classification metric such as the Gini coefficient or cross entropy. uBoost uBoost introduces an additional weight term in the boosting procedure so that regions in mass with low efficiency receive a higher weigth and regions with large efficiency receive a lower weight.

Following ATLAS, we used the implementation provided in the hep_mlv0.6.0 package hepml. The hyperparameters were n_estimators=500\texttt{n\_estimators}=500, learning_rate=0.5\texttt{learning\_rate}=0.5, and base_estimator was the DecisionTreeClassifier from sklearn with max_depth=20\texttt{max\_depth}=20 and min_samples_leaf=0.01\texttt{min\_samples\_leaf}=0.01. For the uBoost uniforming rate (the analogue of λ\lambda for DisCo and adversary), we scanned the range 0–3. We performed 5 independent trainings per uniforming rate and observed that the results were quite stable and consistent between them. Larger values of the uniforming rate were observed to populate lower R50R_{50} but with a unreliably large variation in JSD50, so they were not included in this study.

A.5 DNN classifier

As in the ATLAS study, we use for the DNN classifier a fully connected network consisting of 3 hidden layers with 64 nodes each. Except for the final softmax layer we use ReLU activations everywhere. Unlike the ATLAS study, we chose to include a batchnorm layer BatchNorm after the first hidden layer, as we found this improved the stability of the outcome.

A.6 Training with adversary

Adversarial training follows the approach outlined in the Introduction, with the adversary attempting to learn the PDF of the mass. The training objective is given by

and is usually solved by alternating training of the two networks. The different training objectives between discriminator and adversary are implemented using gradient reversal. Here λ\lambda is a tunable hyperparameter defining the relative weight of classification and decorrelation objective. The classifier loss term LCL_{C} is the usual cross-entropy term, while for the output of the adversary AA is the probability density produced by the Gaussian mixture model and LA=−log⁡AL_{A}=-\log A is evaluated at the true value of the mass. An alternative approach to adversarial decorrelation attempts to infer the mass itself, in which case the adversarial loss LA(θC,θA)L_{A}(\theta_{C},\theta_{A}) would take the form of a regression term or cross-entropy between different mass bins Shimmin:2017mfk.

The adversary predicts a probability distribution function for the mass. The function is parametrised by a sum of 20 Gaussian distributions in a Gaussian mixture model. This means the network outputs 60 quantitites, interpreted as relative normalisation, mean and variance of 20 Gaussian distributions with a two layer fully connected network for each parameter where the first layer has 64 nodes and is shared. The output of the discriminator and pTp_{T} are used as inputs to the adversary.

Training the adversary is done in three phases: only training the discriminator for 200 epochs (45 epochs for the CNN); only training the adversary with fixed discriminator for 20 epochs (30 epochs for the CNN); and joint training of both networks for 200 epochs (25 epochs for the CNN). For the DNN, the initial learning rates for the three phases are λC=0.01\lambda_{C}=0.01, λA=0.05\lambda_{A}=0.05 and (λC,λA)=(3⋅10−6,0.0001)(\lambda_{C},\lambda_{A})=(3\cdot 10^{-6},0.0001) respectively. The initial learning rates are subject to an exponential decay of dC=0.98d_{C}=0.98, dA=0.98d_{A}=0.98, (dC,dA)=(0.97,0.97)(d_{C},d_{A})=(0.97,0.97). For the CNN the initial learning rates are λC=0.0001\lambda_{C}=0.0001, λA=0.0005\lambda_{A}=0.0005, and (λC,λA)=(0.000001,0.001)(\lambda_{C},\lambda_{A})=(0.000001,0.001) for the three phases. No exponential decay is used for pre-training the classifier, and decay rates of dA=0.98d_{A}=0.98 and (dC,dA)=(0.95,0.99)(d_{C},d_{A})=(0.95,0.99) are used for the second and third phase. We verified for some representative values of λ\lambda that greatly increasing the number of epochs (training up to 300 epochs) did not noticeably improve performance; nor did changing the model selection to the lowest loss instead of the final epoch. The batch size is 8192 (1000) for the DNN (CNN) approach.

Appendix B Top tagging

We will also compare the performance of DisCo to adversarial decorrelation in the case of top tagging. For top tagging, we use the QCD and top samples in Landscape, and we restrict our comparison to CNNs trained on jet images (with the same specifications as the WW-tagging). Fig. 5 shows the average top and QCD images. Despite the much higher possible discriminating power in top tagging, we again see that DisCo is comparable to the adversary, demonstrating that DisCo is indeed a powerful and sensitive measure of nonlinear correlation and a very effective penalty term for decorrelation.

References

References