Neural Networks for Full Phase-space Reweighting and Parameter Tuning
Anders Andreassen, Benjamin Nachman
References
Appendix A Optimal Functions
The results presented here can be found (as exercises) in textbooks, but are repeated here for easy access. Let be some discriminating features and is another random variable representing class membership. Consider the general problem of minimizing some average loss for the function :
Now, consider the case where the loss is cross-entropy:
Therefore, the output is proportional to the likelihood ratio. The proportionality constant is the ratio of fractions of the two classes used during the training. In the paper, the two classes always have the same number of examples and thus this factor is unity.
Appendix B Alternative Fitting Method
In the main body, it was shown how a continuously parameterized NN used for reweighting:
This works well when the reweighting and fitting happen on the same ‘level’. However, if the reweighting happens at truth level (before detector simulation) while the fit happens in data (after the effects of the detector), this procedure will not work. It works only if the reweighting and fitting both happen at detector-level or both happen at truth-level. The following is an alternative method:
where is a trained Dctr using binary cross entropy as in the main body. The intuition of the above equation is that the classifier is trying to distinguish the two samples and we try to find a that makes ’s task maximally hard. If cannot tell apart the two samples, then the reweighting has worked. This is similar to the minimax graining of a GAN, only now the analog of the generator network is the reweighting network which is fixed and thus the only trainable parameters are the . The advantage of this second approach is that it readily generalizes to the case where the reweighting happens on a different level:
where is the truth value and is the detector-level value. In simulation (the second sum), these come in pairs and so one can apply the reweighting on one level and the classification on the other.
Asymptotically, both this method and the one in the body of the DCTR paper learn the same result: . To see this for the second method, consider the same logic as in Appendix A. Conditioning on and , the optimal is given by