A guide for deploying Deep Learning in LHC searches: How to achieve optimality and account for uncertainty
Benjamin Nachman
Introduction
Since the first studies of deep learning Here, ‘deep learning’ means machine learning with modern neural networks. These networks are deeper and more flexible than artificial neural networks from the past and can now readily process high-dimensional feature spaces with complex structure. In the context of classification, deep neural networks are powerful high-dimensional likelihood-ratio approximators, as described in later sections. Machine learning has a long history in HEP and there are too many references to cite here – see for instance Ref. Hocker:2007ht and the many papers that cite it and came before going back to at least Ref. Lonnblad:1990qp. in high energy physics (HEP) Baldi:2014kfa; deOliveira:2015xxd, there has been a rapid growth in the adaptation and development of techniques for all areas of data analysis (see Ref. Larkoski:2017jix; Radovic:2018dip; Guest:2018yhq and references therein). One of the most exciting prospects of deep learning is the opportunity to exploit all of the available information to significantly increase the power of Standard Model measurements and searches for new particles.
While the first analysis-specific deep learning results are only starting to become public (see e.g. Aad:2019yxi; ATLAS-CONF-2019-017; CMS-PAS-SUS-19-009), analysis-non-specific deep learning has been used for a few years starting with flavor tagging CMS-DP-2017-005; ATL-PHYS-PUB-2017-003. In addition, there are a plethora of experimental exp and phenomenological pheno studies for additional methods which will likely be realized as part of physics analyses in the near future.
The goal of this paper is to clearly and concisely describe how to achieve optimality and account for uncertainty using deep learning in LHC data analysis. One of the most common questions when an analysis wants to use deep learning is ‘…but what about the uncertainty on the network?’ Ideally the discussion here will help clarify and direct this question. The concepts of accuracy/bias and optimality/precision will be distinguished when covering uncertainties. In particular, most applications of neural networks do not result in uncertainties on the networks themselves related to the accuracy of the analysis. This statement will be qualified in detail. The exposition is based on a mixture of old and new insights, with references given to the foundational papers for further reading – although it is likely that even older references exist in the statistics literature. Section 2 introduces a simple example that will be used for illustration throughout the paper. Sections 3 and 4 discuss achieving optimality and including uncertainties, respectively. A brief discussion with future outlook is contained in Sec. 5.
Illustrative Model
One of the most widely used techniques to search for new particles in HEP is to seek out a localized feature on top of a smooth background from known phenomena, also known as the ‘bump hunt’. This methodology was used to discover the Higgs boson Aad:2012tfa; Chatrchyan:2012xdj and has a rich history, dating back at least to the discovery of the meson PhysRev.126.1858. The localized feature in these searches is often the invariant mass of two or more decay products of the hypothetical new resonance. A simplified version of this search is used as an illustrative example.
Consider a simple approximation to the bump hunt where the background distribution is uniform and the signal is a -function. Let be a random variable corresponding to the bump hunt feature, which can be thought of as the invariant mass of two objects such as high energy jets. Under this model, and . Following the analogy with jets, the random variable will be built from two other random variables representing the jet energies. As the energy of the daughter objects would be approximately half of the mass of the parent, let the decay products be . Detector distortions will perturb the . Let represent the measured versions of the . The experimental resolution will affect and independently. To model these distortions, let where . Figure 1 schematically illustrates the connection between the back-to-back decay of a massive particle produced at rest and the approximation used here.
For simplicity, all arithmetic is performed ‘mod ’ so that everything can be visualized inside a compact interval. This means that an integer is added or subtracted until the resulting value is in the range . For example, if and then . The distributions of are shown in Fig. 2.
The reconstructed mass is given by and can be used for the resonance search. By construction, contains no useful information for distinguishing the signal and background processes (see right plot of Fig. 2), but is sensitive to the value of . Two scenarios for will be investigated in later sections: (1) is constant and the same for all events In practice, may vary from event-to-event in a known way, but there is one global nuisance parameter. and (2) is different, but known for each event. The former is the usual case for a global systematic uncertainty and in this simple example is analogous to the jet energy scale resolution Khachatryan:2016kdb; Aad:2014bia. The latter is the opposite extreme, where there is a precise event-by-event constraint. The physics analog of (2) would be jet-by-jet estimates of the jet energy uncertainty or the number of additional proton-proton collisions (‘pileup’) in an event, which degrades resolutions, but can be measured for each bunch crossing.
The experimental goal is to identify if the data are consistent with the background only, or if there is evidence for a non-zero contribution of signal.
Achieving Optimality
The usual way of performing an analysis with is to define a signal region and then count the number of events in data and compare it to the number of predicted events See Ref. PhysRevD.98.030001 for a review of statistical procedures in HEP.. For a fixed and known value Section 4 will revisit the above setup when is not known. of , and a particular signal model, the probability distribution for the number of observed events follows a Poisson distribution:
where , and are the predicted signal and background event yields. Viewed as a function of for fixed , Eq. 1 is the likelihood . One can express , where is an event-selection efficiency, is the cross-section for producing events, and is the instantaneous luminosity. These yields may be derived directly from simulation or completely / partially constrained from control regions in data. The parameter distinguishes the null hypothesis This is for setting limits. For discovery, the null hypothesis is the Standard Model, . from the alternative hypothesis .
By the Neyman-Pearson lemma neyman1933ix, the most powerful This means that for a fixed probability of rejecting the null hypothesis when it is true, this test has the highest probability of rejecting the null hypothesis when the alternative is true. test for this analysis is performed with the likelihood ratio test statistic:
At the LHC, it is more common to use the profile likelihood ratio instead:
There is no uniformly most powerful test in the presence of nuisance parameter, but Eq. 3 is still the most widely-used test statistic. Two related quantities are the and , which are the -values associated with the Eq. 3 under the null and alternative hypotheses:
where is the observed test statistics. In the absence of signal, , so this is nearly the same as . A generic feature of hypothesis tests where the two hypothesis do not span the space of possibilities is that both the null and alternative hypothesis can be inconsistent with the data. The HEP solution to this phenomenon is to exclude a model if the value of is small instead of Read:2002hq; Junk:1999kv. The does not maximize statistical power and is not even a proper -value. Other proposals for regulating the -value have been discussed (see e.g. Ref. Cowan:2011an or Sec. 7.1. in Ref. Nachman:2016qyc) but are not used in practice. The results that follow will focus on and the lessons learned are likely approximately valid for the HEP solution as well The neural network approximates below naturally extend to the profiling case via parameterized neural networks Cranmer:2015bka; Baldi:2016fzo, where the dependence on is explicitly learned..
2 Cuts, bins, and beyond
Suppose momentarily that the probability density for is piece-wise constant over a finite number of patches . Each patch is independent, so the full likelihood is the product over patches:
where is the probability to observe . Equation 7 is often called the extended likelihood Barlow:1990vc. The important feature about Eq. 7 is that it does not depend on the patches and so in the continuum limit, it is also appropriate for the full phase space likelihood where is now a probability density.
The likelihood ratio statistic using Eq. 7 is then given by
The optimal use of the full phase space would then be to exclude the model if is small. Especially when is high-dimensional, , , and are not known analytically. Using a small number of bins is an approximation to Eq. 8 and can provide additional power beyond Eq. 1. Adding bins is only helpful if the events in each bin have a different likelihood ratio. The loss functions used in deep learning ensure that the network output is monotonically related to the likelihood ratio (more on this below). Therefore, bins chosen based on the output of a neural network typically enhance the statistical power of an analysis. Unless the likelihood ratio only takes on a small number of values, binning will necessarily be less powerful than an unbinned approach using the full likelihood for . Using the output of a neural network to construct bins does help to reduce the potentially high-dimensional problem to a one-dimensional one. However, a small number of bins may not be sufficient to capture all of the salient structures in the likelihood ratio, especially when is high-dimensional.
where the output range $01p(S+B|Y)$. This is not the likelihood ratio, but some symbolic manipulation shows that it is monotonically related to it:
has the property that . The loss function proposed in Ref. DAgnolo:2018cun approaches directly, which may be useful when considering the logarithm of Eq. 8 for the statistical test. One potential advantage of learning the ratio first and then taking the logarithm is that one only needs to achieve proportionality while proportionality constants for the loss designed to learn the logarithm of are suboptimal. See Appendix A for the derivations involving these loss functions.
Figure 3 illustrates the above concepts using the simple example from Sec. 2. These plots are trained only with for simplicity. The left plot of Fig. 3 presents a histogram of the ratio of neural network outputs for . The neural network is parameterized in Keras keras with the Tensorflow tensorflow backend with three fully connected hidden layers using 10, 20, 50 hidden nodes and the exponential linear unit activation function clevert2015fast with 10% dropout JMLR:v15:srivastava14a. The last hidden layer is connected to a one-node output using the sigmoid activation function and the loss was binary cross-entropy. The networks are optimized using Adam adam over three epochs with 500,000 examples and a batch size of 50. None of these parameters were optimized, as the problem is sufficiently simple that the specifics of the training are not important for the message presented with the results below.
As desired, the ratio of signal to background in Fig. 3 is monotonically increasing from left to right and this ratio is the same as the value on the horizontal axis. Since this problem is one-dimensional, it is possible to readily visualize the functional form of , shown in the right plot of Fig. 3. With a uniform background, the likelihood ratio should be simply the signal probability distribution, which is a Gaussian In fact, this (train against a uniform distribution) can be used as a general method for using a neural network to learn a generic probability density. However, it becomes inefficient when the dimensionality of is large. One may be able to overcome this by using some localized (but known) density first before applying this procedure. Thank you to David Shih for useful discussions on this point.. For comparison to the neural network, a binned version of the likelihood ratio is presented alongside the analytic result assuming so that edge effects are not relevant. The neural network output can be used to well-approximate the likelihood ratio.
Note that Eq. 9 is set up such that the classifier learns to separate from . It is more common and often more pragmatic to train a classifier to distinguish from directly. The resulting classifier will be monotonically related to the one resulting from the versus classification Metodiev:2017vrx. Writing with , a neural network trained with the binary cross entropy will produce
Therefore, one can use the predictions for the yields and to correct the versus classifier.
A complete illustration of cuts-versus-bins-versus-deep learning is presented in Fig. 4 for the simple example from Sec. 2. As the above example shows that a NN can be used to well-approximate the likelihood ratio, the analytic result is used for this comparison. The horizontal axis in Fig. 4 is the ‘level’ or type I error of the test while the vertical axis is the power or (1-type II error rate). For the Inclusive scenario, is not used and the test is based on alone. For the other cases, the bins and cuts are based on the likelihood ratio. For the Fixed cut, the value is 0.5, and for the Two bins case, the bins boundaries are at 0.5 and 2, which were optimized to capture most of the information. The Many bins case uses 20 bins evenly spaced between 0 and 3. The Optimal procedure uses Eq. 8. For this simple example, two bins are nearly sufficient to capture all of the available information and by twenty bins, the procedure has converged to the one from Eq. 8.
3 Nuisance features
Given the high-dimensionality of LHC data, it is often necessary to only consider a redacted set of features for deep learning. It is tempting to only use features for which is very different from unity. However, there are often features that are directly related to the resolution or uncertainty of other observables. Such features may have on their own, but can enhance the potential of other observables when combined Nachman:2013bia. Examples of this type were mentioned in Sec. 2. A neural network approximation to the likelihood will naturally make the optimal use of these ‘nuisance features’. Removing these features from consideration or even purposefully reducing their impact on the directly discriminative features Louppe:2016ylz will necessarily reduce the analysis optimality. Nuisance features are not the same as nuisance parameters, where the value is unknown and a direct source of uncertainty. The interplay of nuisance parameters and neural network uncertainty is discussed in Sec. 4.3.
The simplest way to incorporate nuisance features is to simply treat them in the same way as directly discriminative features. Figure 5 illustrates this difference with the toy model from Sec. 2, where inference is performed with a classifier using only and one using and , where is uniformly distributed between 0 and 0.29. As expected, when trying to determine , the example that used in addition to is able to achieve a superior statistical precision. This case is nearly the same to the one where there is a global nuisance parameter and it is well constrained by some auxiliary data, i.e. constraining with . In that case, one may use the techniques of parameterized classifiers Cranmer:2015bka; Baldi:2016fzo to construct the NN, which is the same as treating as a discriminating feature, only that a small number of values may be available for training. This is discussed in more detail in Sec. 4.3.
Sources of Uncertainty
Uncertainty quantification is an essential part of incorporating deep learning into a HEP analysis framework. One of the most often-expressed phrases when someone proposes to use a deep neural network in an analysis is ‘what is the uncertainty on that procedure?’ The goal of this section is to be explicit about sources of uncertainty, how they impact the scientific result, and what can be done to reduce them.
There are generically two sources of uncertainty. One source of uncertainty decreases with more events (statistical or aleatoric uncertainty) and one represents potential sources of model bias that are independent of the number of events (systematic or epistemic uncertainty). These uncertainties are relevant for data as well as the models used to interpret the data, and in general there can be sources of uncertainty that have components due to both types. For most searches, the analysis strategy is designed prior This is not the case for a recent anomaly detection proposal where the event selection depends on the data Collins:2018epr; Collins:2019jip; Hajer:2018kqm; Farina:2018fyg; Heimel:2018mkt; Cerri:2018anq; Roy:2019jae; Blance:2019ibf; Dillon:2019cqt; DAgnolo:2018cun; DeSimone:2018efk. to any statistical tests on data (‘unblinding’). In the deep learning context, this means that the neural network training is separate from the statistical analysis. As such, it is useful to further divide uncertainty sources into two more types: uncertainty on the precision/optimality of the procedure and uncertainty on the accuracy/bias of the procedure. These will be described in more detail below.
Consider the neural network setup from Sec. 3. If the network architecture is not flexible enough, there were not enough training examples, or the network was not trained for long enough, it may be that the likelihood ratio is not well-approximated. This means that the procedure will be suboptimal and will not achieve the best possible precision. However, if the classifier is well-modeled by the simulation, then -values computed from the classifier may be accurate, which means that the results are unbiased. Conversely, a well-trained network may result in a biased result if the simulation used to estimate the -value is not accurate. From the point of view of accuracy, the neural network is just a fixed non-linear high-dimensional function whose probability distribution must be modeled to compute -values. In other words, the NN itself has no uncertainty in its accuracy - its evaluation is only uncertain through its inputs. A useful analogy is to consider common high-dimensional non-linear functions like the jet mass, which clearly have no uncertainty on their definition.
Figure 6 summarizes the various sources of uncertainty related to neural networks, broken down into the four categories described above. A machine learning model is trained on (usually) simulation following the distribution . Given the trained model, the probability density of is determined with another simulation following the distribution . It is often the case that . Systematic uncertainties affecting the accuracy of the result originate from differences between and the true density while systematic uncertainties related to the optimality of the procedure originate from differences between and .
The precision/optimality uncertainty is practically important for analysis optimization. If this uncertainty is large, one may want to modify some aspect of the analysis design (more on this in Sec. 4.3). The precision/optimality uncertainty is often estimated by rerunning the training with different random initializations of the network parameters. This procedure is sensitive to both the finite size of the training dataset as well as the flexibility of the optimization procedure. One can also bootstrap the training data for fixed weight initialization to uniquely probe the statistical uncertainty from the training set size. An automated approach to estimate these uncertainties that does not require retraining multiple networks is Bayesian Neural Networks mackay; oldBNN; Xu:2007ad; Bhat:2005hq; Bollweg:2019skg. Estimating the uncertainty from the input feature accuracy can be performed by varying the inputs within their systematic uncertainty (see Sec. 4.4). This can be incorporated into network training via parameterized networks Cranmer:2015bka; Baldi:2016fzo with profiling (see Sec. 4.3). Determining the uncertainty from the model flexibility is challenging and there is currently no automated way for including this in the training. One (likely insufficient) possibility is to probe the sensitivity of the network performance to small perturbations in the network architecture.
Before turning to more specific details in the following sections, it is useful to consider efforts by the non-HEP machine learning community for uncertainties related to deep learning. An often-cited discussion of model uncertainty (not necessarily for deep learning) is Ref. 2cfc44fb6aa842f5ac03e51409667171, which lists seven sources of uncertainty. Many of these align well with those presented in Fig. 6. However, one key difference between HEP and industrial (and other scientific) applications of deep learning is the high-quality of HEP simulation. For instance, consider a charged particle with momentum that is measured with momentum . Industrial applications may treat as a source of uncertainty while in HEP, if is well-modeled by the simulation, there is no uncertainty at all. Therefore, the tools and strategies for uncertainty in the machine learning literature are not always directly applicable to HEP. See e.g. Ref. Tagasovska2019 (and the many references therein) for a recent discussion of uncertainties related to deep learning models.
2 Asymptotic formulae with classifiers
3 Learning to profile: reducing the optimality systematic uncertainty
If the classification is particularly sensitive to a source of systematic uncertainty, one may want to reduce the dependence of the neural network on the corresponding nuisance parameter This is in contrast to ‘nuisance features’, where it is argued in Sec. 3.3 should never be removed from the neural network training except to reduce model complexity. - see e.g. Ref. Cranmer:2015bka for an automated method for achieving this goal. While removing the dependence on such features may reduce model complexity, it will generally not improve the overall analysis sensitivity. By construction, if the classification is sensitive to a given nuisance parameter, removing the dependence on that parameter will reduce the nominal model performance. The significance will only degrade if the uncertainty is sufficiently large. To see this, suppose that there are two independent features to be used in training and one of them has an uncertainty for the background. In the asymptotic limit (including , Cowan:2010js), the question of deriving additional benefit from the uncertain feature is given symbolically by
where is the efficiency of classifier for class and is the uncertainty on the efficiency. Equation 15 is equivalent to
Independent of the uncertainty, it only makes sense to use the classifier if the first term . For the second term, if each of the efficiencies and relative uncertainties are , then would need to be in order for the additional uncertain feature to detract from the analysis sensitivity. Even if the uncertainty is large, there are methods which construct classifiers using the information about how they will be used (‘inference-aware’) and therefore should never do worse than the case where the uncertain features are removed from the start deCastro:2018mgh.
Especially for deep learning, one should proceed with caution when removing the dependence on single nuisance parameters that represent many sources of uncertainty. For example, it is common to use a single nuisance parameter to encode all of the fragmentation uncertainty. This is already tenuous when using high-dimensional, low-level inputs, as the uncertainty covariance is highly constrained. If the sensitivity to such a nuisance parameter is removed, it does not mean that the network is insensitive to fragmentation - it only means that it is not sensitive to the fragmentation variations encoded by the single nuisance parameter. This may also apply to other sources of theory uncertainty such as scale variations for estimating uncertainties from higher-order effects. In some cases, higher-order terms may be known to be small and can justify reducing the sensitivity to scale variations Englert:2018cfo, but these terms are typically not known.
The above arguments can be complicated when the two features are not independent and the background is estimated entirely from data via the ABCD method For this method, there are two independent features which are used to form four regions called based on threshold cuts on the two features. One of these regions is enriched in signal and the other three are used to estimate the background. or a sideband fit. In that case, the strength of the feature dependence can increase the background uncertainty. When the dependence is strong enough, it may no longer be possible to estimate the background. There are a variety of neural network Shimmin:2017mfk; Bradshaw:2019ipy and other Dolen:2016kst; Moult:2017okx; Stevens:2013dya; ATL-PHYS-PUB-2018-014 approaches to achieve this decorrelation.
Instead of removing the dependence on uncertain features, a potentially more powerful way to reduce precision systematic uncertainties is to do exactly the opposite - depend explicitly on the nuisance parameters Cranmer:2015bka; Baldi:2016fzo. By parameterizing a neural network as a function of the nuisance parameters , one can achieve the best performance for each value of (such as the variations). Furthermore, this can be combined with profiling so that when the data are fit to determine and constrain its uncertainty, the neural network is accordingly modified. The left plot of Fig 7 shows that the idea of parameterized classifiers Cranmer:2015bka; Baldi:2016fzo works well for the toy example from Sec. 2. The training was performed with values . The right plot of Fig. 7 shows the uncertainty on when performing a statistical test with a neural network trained with a sample generated with nuisance parameter . As expected, the uncertainty is smallest when so that is the optimal classifier for that value of the nuisance parameter (clearly, the uncertainty is worse when is large). Therefore, if the fitted value of is the true value (as is hopefully true when it is profiled), the statistical procedure will make the best use of the data.
In practice, it may be challenging to generate multiple training datasets with different values of . Neural networks are excellent at interpolating between parameter values, but there must be enough values to ensure an accurate interpolation. This can be especially challenging if is multi-dimensional. In practice, learning to profile will likely work well for nuisance parameters that only require a variation in the final analysis inputs (such as the jet energy scale variation) and not for parameters that require rerunning an entire detector simulation (such varying fragmentation model parameters). For the latter case, one may be able to use high-dimensional reweighting to emulate parameter variations without expensive detector simulations Andreassen:2019nnm.
4 High-dimensional bias uncertainties
The single biggest challenge to using high-dimensional features for neural networks is estimating high-dimensional uncertainties. Many sources of experimental uncertainty factorize into independent terms for each object. However, physics modeling uncertainties are often grouped into two-point variations that cover many physical effects all at once. These uncertainties may no longer be appropriate when the input features are high-dimensional (see also Sec. 4.3). There are additional complications when computing uncertainties beyond and even for the uncertainties if the NN is a non-monotonic transformation of the input as quantiles are not preserved.
The fact that this section is short is an indication that new ideas are needed in this area.
Conclusions and Outlook
This paper has reviewed how deep learning can be used to make the best use of data for new particle searches at the LHC. Deep learning-based classifiers can serve as surrogates to the likelihood ratio in order to achieve an optimal test statistic. Nuisance features can improve the performance of such classifiers even if they are not individually useful for distinguishing signal and background. The ways in which uncertainties affects deep learning-based inference were discussed and categorized into precision uncertainties related to the optimality of the procedure and accuracy uncertainties related to the bias of the method. While both sources of uncertainty are useful to quantify, the latter is much more important for the utility of the results. Precision uncertainties can be reduced by letting the deep learning models depend explicitly on the nuisance parameters and then profiling them during the statistical analysis.
As deep learning-based search strategies become more common, it will be important to discuss all of these topics in more detail and develop strategies to ensure that the precious data from the LHC are used in the best way possible to learn the most about the fundamental properties of nature.
References
Appendix A Loss functions and asymptotic behavior
The results presented here Heavily borrowed from the appendix in Ref. Andreassen:2019nnm. can be found (as exercises) in textbooks, but are repeated here for completeness. Let be some discriminating features and is another random variable representing class membership (signal versus background). Consider the general problem of minimizing an average loss for the function :
As a first concrete example, consider the mean-squared error loss: . One can compute
where again, the optimal value is . The same analysis can be applied to the loss in Eq. 11:
Equation 10 in the text shows that the above is proportional to the likelihood ratio.