Collapse of Deep and Narrow Neural Nets

Lu Lu, Yanhui Su, George Em Karniadakis

Introduction

With regards to optimum activation function employed in the NN approximation, before 2010 the two commonly used non-linear activation functions were the logistic sigmoid 1/(1+e−x)1/(1+e^{-x}) and the hyperbolic tangent (tanh⁡\tanh); they are essentially the same function by simple re-scaling, i.e., tanh⁡(x)=2 sigmoid(2x)−1\tanh(x)=2\text{ sigmoid}(2x)-1. The deep neural networks with these two activations are difficult to train (Glorot & Bengio, 2010). The non-zero mean of the sigmoid induces important singular values in the Hessian (LeCun et al., 1998), and they both suffer from the vanishing gradient problem, especially through neurons near saturation (Glorot & Bengio, 2010). In 2011, ReLU was proposed, which avoids the vanishing gradient problem because of its linearity, and also results in highly sparse NNs (Glorot et al., 2011). Since then, ReLU and its variants including leaky ReLU (LReLU) (Maas et al., 2013), parametric ReLU (PReLU) (He et al., 2015) and ELU (Clevert et al., 2015) are favored in almost all deep learning models. Thus, in this study, we focus on the ReLU activation.

While the aforementioned theoretical results are very powerful, they do not necessarily coincide with the results of training of NNs in practice which is NP-hard (Šíma, 2002). For example, while the theory may suggest that the approximation of a multi-dimensional smooth function is accurate for NN with 10 layers and 5 neurons per layer, it may not be possible to realize this NN approximation in practice. Fukumizu & Amari (2000) first proved that existence of local minima poses a serious problem in learning of NNs. After that, more work has been done to understand bad local minima under different assumptions (Zhou & Liang, 2017; Du et al., 2017; Safran & Shamir, 2017; Wu et al., 2018; Yun et al., 2018). Besides local minima, singularity (Amari et al., 2006) and bad saddle points (Kawaguchi, 2016) also affect training of NNs. Our paper focuses on a particular kind of bad local minima, i.e., those encountered in deep and narrow neural networks collapse with high probability. This is the topic of our work presented in this paper. Our results are summarized in Fig. 6, which shows a diagram of the safe region of training to achieve the theoretically expected accuracy. As we show in the next section through numerical simulations as well as in subsequent sections through theoretical results, there is very high probability that for deep and narrow ReLU NNs will converge to an erroneous state, which may be the mean value of the function or its partial mean value. However, if the NN is trained with proper normalization techniques, such as batch normalization (Ioffe & Szegedy, 2015), the collapse can be avoided. Not every normalization technique is effective, for example, weight normalization (Salimans & Kingma, 2016) leads to the collapse of the NN.

Collapse of deep and narrow neural networks

In this section, we will present several numerical tests for one- and two-dimensional functions of different regularity to demonstrate that deep and narrow NNs collapse to the mean value or partial mean value of the function.

It is well known that it is hard to train deep neural networks. Here we show through numerical simulations that the situation gets even worse if the neural networks is narrow. First, we use a 10-layer ReLU network with width 2 to approximate y(x)=∣x∣y(x)=|x|, and choose the mean squared error (MSE) as the loss. In fact, y(x)y(x) can be represented exactly by a 2-layer ReLU NN with width 2, ∣x∣=ReLU(x)+ReLU(−x)=[11]ReLU([1−1]x)|x|=\text{ReLU}(x)+\text{ReLU}(-x)=\begin{bmatrix}1&1\end{bmatrix}\text{ReLU}(\begin{bmatrix}1\\ -1\end{bmatrix}x). However, our numerical tests show that there is a high probability (∼90%\sim 90\%) for the NN to collapse to the mean value of y(x)y(x) (Fig. 1), no matter what kernel initializers (He normal (He et al., 2015), LeCun normal (LeCun et al., 1998; Klambauer et al., 2017), Glorot uniform (Glorot & Bengio, 2010)) or optimizers (first order or second order including SGD, SGDNesterov (Sutskever et al., 2013), AdaGrad (Duchi et al., 2011), AdaDelta (Zeiler, 2012), RMSProp (Hinton, 2014), Adam (Kingma & Ba, 2015), BFGS (Nocedal & Wright, 2006), L-BFGS (Byrd et al., 1995)) are employed. The training data were sampled from a uniform distribution on [−3,3][-\sqrt{3},\sqrt{3}], and the minibatch size was chosen as 128 during training. We find that when this happens, in most cases the bias in the last layer is the mean value of the function y(x)y(x), and the composition of all the previous layers is equivalent to a zero function. It can be proved that under these conditions, the gradient vanishes, i.e., the optimization stops (Corollary 5). For functions of different regularity, we observed the same collapse problem, see Fig. 2 for the C∞C^{\infty} function y(x)=xsin⁡(5x)y(x)=x\sin(5x) and Fig. 3 for the L2L^{2} function y(x)=1{x>0}+0.2sin⁡(5x)y(x)=1_{\{x>0\}}+0.2\sin(5x).

For multi-dimensional inputs and outputs, this collapse phenomenon is also observed in our simulations. Here, we test the target function y(x)\mathbf{y}(\mathbf{x}) with din=2d_{in}=2 and dout=2d_{out}=2, which can be represented by a 2-layer neural network with width 4, y(x)=[∣x1+x2∣∣x1−x2∣]=[1111]ReLU([11−1−11−1−11]x)\mathbf{y}(\mathbf{x})=\begin{bmatrix}|\mathbf{x}_{1}+\mathbf{x}_{2}|\\ |\mathbf{x}_{1}-\mathbf{x}_{2}|\end{bmatrix}=\begin{bmatrix}1&1&&\\ &&1&1\end{bmatrix}\text{ReLU}(\begin{bmatrix}1&1\\ -1&-1\\ 1&-1\\ -1&1\end{bmatrix}\mathbf{x}). When training a 10-layer ReLU network with width 4, there is a very high probability for the NN to collapse to the mean value or with low probability to the partial mean value of y(x)\mathbf{y}(\mathbf{x}) (Fig. 4).

We also observed the same collapse problem for other losses, such as the mean absolute error (MAE); the results are summarized in Fig. 5 for three different functions with varying regularity. Furthermore, we find that for MSE loss, the constant is the mean value of the target function, while for MAE it is the median value.

Initialization of ReLU nets

As we demonstrated above, when the weights of the ReLU NN are randomly initialized from a symmetric distribution, the deep and narrow NN will collapse with high probability. This type of initialization is widely used in real applications. Here, we demonstrate that this initialization avoids the problem of exploding/vanishing mean activation length, therefore this is beneficial for training neural networks.

where ϕ\phi is a component-wise activation function.

Following the work in Poole et al. (2016), we investigate how the length of the input propagates through neural networks. The normalized squared length of the vector before activation at each layer is defined as

where hil\mathbf{h}^{l}_{i} denotes the entry ii of the vector hl\mathbf{h}^{l}. If the weights and biases are drawn i.i.d. from a zero mean Gaussian with variance σw2/Nl−1\sigma_{w}^{2}/N^{l-1} and σb2\sigma_{b}^{2} respectively, then the length at layer ll can be obtained from its previous layer (Poole et al., 2016; Long & Sedghi, 2019)

Theoretical analysis of the collapse problem

Remark: We point out here that the connectedness in assumption A1 is a very weak requirement for the input space. The weights in a neural network are usually sampled independently from continuous distributions in real applications, and thus the assumption A2 is satisfied at the NN initialization stage; during training, the assumption A2 is usually maintained due to stochastic gradients of minibatch.

With assumptions A1 and A2, if N(x0)\mathcal{N}(\mathbf{x}^{0}) is a constant function, then there exists a layer l∈{1,…,L−1}l\in\{1,\dots,L-1\} such that hl≤0\mathbf{h}^{l}\leq\mathbf{0}a≤b\mathbf{a}\leq\mathbf{b} denotes ai≤bi\mathbf{a}_{i}\leq\mathbf{b}_{i} for any index ii, i.e., component-wise. Similarly for <<, >> and ≥\geq. and xl=0 ∀x0∈Ω\mathbf{x}^{l}=\mathbf{0}~{}\forall\mathbf{x}^{0}\in\Omega, with probability 1 (wp1).

With assumptions A1 and A2, if N(x0)\mathcal{N}(\mathbf{x}^{0}) is bias-free and a constant function, then there exists a layer l∈{1,…,L−1}l\in\{1,\dots,L-1\} such that for any n≥ln\geq l, it holds hn≤0\mathbf{h}^{n}\leq\mathbf{0} and xn=0\mathbf{x}^{n}=\mathbf{0} wp1.

With assumptions A1 and A2, if N(x0)\mathcal{N}(\mathbf{x}^{0}) is a constant function, then any order gradients of the loss function with respect to the weights and biases in layers 1,…,l1,\dots,l vanish, where ll is the layer obtained in Lemma 1.

Remark: See Appendices A, B, C and D for the proofs of Lemma 1, Corollary 2, Lemma 3 and Theorem 4, respectively. MAE and MSE loss used in practice are discrete versions of L1L^{1} and L2L^{2} loss, respectively, if the size of minibatch is large.

Corollary 5 can be generalized to the following corollary including more general converged mean states.

With assumptions A1 and A2, for a ReLU feed-forward neural network N(x0)\mathcal{N}(\mathbf{x}^{0}) and any function y(x0)\mathbf{y}(\mathbf{x}^{0}), x0∈Ω\mathbf{x}^{0}\in\Omega, if ∃K1,…,Kn⊂Ω\exists K_{1},\dots,K_{n}\subset\Omega and each KiK_{i} is a connected domain with at least two points, such that

then the gradients of the loss function with respect to any weight or bias vanish when using the L2L^{2} loss. Here xKi0\mathbf{x}^{0}_{K_{i}} is the random variable of x0\mathbf{x}^{0} restricted to KiK_{i}.

See Appendices E and F for the proofs of Corollaries 5 and 6. We can see that Corollary 5 is a special case of Corollary 6 with ∪i=1nKi=Ω\cup_{i=1}^{n}K_{i}=\Omega.

Let us assume that a one-layer ReLU feed-forward neural network N1\mathcal{N}_{1} is initialized independently by symmetric nonzero distributions, i.e., any weight or bias of N1\mathcal{N}_{1} is initialized by a symmetric nonzero distribution, which can be different for different parameters. Then, for any fixed input the corresponding output is zero with probability (1/2)dout(1/2)^{d_{out}}, except the special case where all biases and the input are zero yielding that the output is always zero.

If a ReLU feed-forward neural network N\mathcal{N} with LL layers assembled width N1,…,NLN^{1},\dots,N^{L} is initialized randomly by symmetric nonzero distributions for weights and zero biases, then for any fixed nonzero input, the corresponding output is zero with probability 1−Πl=1L(1−(1/2)Nl)1-\Pi_{l=1}^{L}(1-(1/2)^{N^{l}}) if the last layer also employs ReLU activation, otherwise with the probability 1−Πl=1L−1(1−(1/2)Nl)1-\Pi_{l=1}^{L-1}(1-(1/2)^{N^{l}}).

See Appendices G and H for the proofs of Lemma 7 and Theorem 8. Although biases are initialized to 0 in most applications, for the sake of completeness, we also consider the case where biases are not initialized to 0.

If a ReLU feed-forward neural network N\mathcal{N} with LL layers assembled width N1,…,NLN^{1},\dots,N^{L} is initialized randomly by symmetric nonzero distributions (weights and biases), then for any fixed nonzero input, the corresponding output is zero with probability (1/2)NL(1/2)^{N^{L}} if the last layer also employs ReLU activation, otherwise the output is equal to the last bias bL\mathbf{b}^{L} with probability (1/2)NL−1(1/2)^{N^{L-1}}.

See Appendix I for the proof of Proposition 9. We note that Theorem 8 provides the probability for any given input, but in Theorem 4 it requires that the entire neural network is a zero function. Hence, the probability in Theorem 8 is an upper bound. In the following theorem, we give a theoretical formula of the probability for the NN with width 2.

Suppose the origin is an interior point of Ω\Omega. Consider a bias-free ReLU neural network with din=1d_{in}=1, width 2 and LL layers, and weights are initialized randomly by symmetric nonzero distributions. Then for this neural network, the probability of being initialized to a constant function is the last component of πL\pi^{L}, where

with π1\pi^{1} and PP being the probability distribution after the first layer and the probability transition matrix when one more layer is added, respectively. Here every layer employs the ReLU activation.

See Appendix J for the derivation of π1\pi^{1} and PP. For general cases, we found that it is hard to obtain an explicit expression for the probability, so we used numerical simulations instead, where 1 million samples of random initialization are used to calculate each probability estimation. We show both theoretically (Theorem 8, Propositions 9 and 10) and numerically that NN has the same probability to collapse no matter what symmetric distributions are used, even if different distributions are used for different weights. On the other hand, to keep the collapse probability less than pp, because the probability obtained in Theorem 8 is an upper bound, which corresponds to a safer maximum number of layers, we have that 1−Πl=1L(1−(1/2)N)≤p1-\Pi_{l=1}^{L}(1-(1/2)^{N})\leq p, which implies the upper bound of the depth of NN

Theorem 8 shows that when the NN gets deeper and narrower, the probability of the NN initialized to a zero function is higher (Fig. 6A). Hence, we have higher probability of vanishing gradient in almost all the layers, rather than just some neurons. In our experiments, we also found that there is very high probability that the gradient is 0 for all parameters except in the last layer, because ReLU is not used in the last layer. During the optimization, the neural network thus can only optimize the parameters in the last layer (Theorem 4). When we design a neural network, we should keep the probability less than 1% or 10%. As a practical guide, we constructed a diagram shown in Fig. 6B that includes both theoretical predictions and our numerical tests. We see that as the number of layers increases, the numerical tests match closer the theoretical results. It is clear from the diagram that a 10-layer NN of width 10 has a probability of only 1% to collapse whereas a 10-layer NN of width 5 has a probability greater than 10% to collapse; for width of three the probability is greater than 60%.

Training deep and narrow neural networks

In this section, we present some training techniques and examine which ones do not suffer from the collapse problem.

Our analysis applies for any symmetric initialization, so it is straightforward to consider asymmetric initializations. The asymmetric initializations proposed in the literature include orthogonal initialization (Saxe et al., 2014) and layer-sequential unit-variance (LSUV) initialization (Mishkin & Matas, 2016). LSUV is the orthogonal initialization combined with rescaling of weights such that the output of each layer has unit variance. Because weight rescaling cannot make the output escape from the negative part of ReLU, it is sufficient to consider the orthogonal initialization. The probability of collapse when using orthogonal initialization is very close to and a little lower than that when using symmetric distributions (Fig. 7). Therefore, orthogonal initialization cannot treat the collapse problem.

2 Normalization and dropout

As we have shown in the previous section, deep and narrow neural networks cannot be trained well directly with gradient-based optimizers. Here, we employ several widely used normalization techniques to train this kind of networks. We do not consider some methods, such as Highway (Srivastava et al., 2015) and ResNet (He et al., 2016), because in these architectures the neural nets are no longer the standard feed-forward neural networks. Current normalization methods mainly include batch normalization (BN) (Ioffe & Szegedy, 2015), layer normalization (LN) (Ba et al., 2016), weight normalization (WN) (Salimans & Kingma, 2016), instance normalization (IN) (Ulyanov et al., 2016), group normalization (GN) (Wu & He, 2018), and scaled exponential linear units (SELU) (Klambauer et al., 2017). BN, LN, IN and GN are similar techniques and follow the same formulation, see Wu & He (2018) for the comparison.

Because we focus on the performance of these normalization methods on narrow nets and the width of the neural network must be larger than the dimension of the input to achieve a good approximation, we only test the normalization methods on low dimensional inputs. However, LN, IN and GN perform normalization on each training data individually, and hence they cannot be used in our low-dimensional situations. Hence, we only examine BN, WN and SELU. BN is applied before activations while for SELU LeCun normal initialization is used (Klambauer et al., 2017). Our simulations show that the neural network can successfully escape from the collapsed areas and approximate the target function with a small error, when BN or SELU are employed. BN changes the weights and biases not only depending on the gradients, and different from ReLU the negative values do not vanish in SELU. However, WN failed because it is only a simple re-parameterization of the weight vectors.

Moreover, our simulations show that the issue of collapse cannot be solved by dropout, which induces sparsity and more zero activations (Srivastava et al., 2014).

Conclusion

We consider here ReLU neural networks for approximating multi-dimensional functions of different regularity, and in particular we focus on deep and narrow NNs due to their reportedly good approximation properties. However, we found that training such NNs is problematic because they converge to erroneous means or partial means or medians of the target function. We demonstrated this collapse problem numerically using one- and two-dimensional functions with C0C^{0}, C∞C^{\infty} and L2L^{2} regularity. These numerical results are independent of the optimizers we used; the converged state depends on the loss but changing the loss function does not lead to correct answers. In particular, we have observed that the NN with MSE loss converges to the mean or partial mean values while the NN with MAE loss converges to the median values. This collapse phenomenon is induced by the symmetric random initialization, which is popular in practice because it maintains the length of the outputs of each layer as we show theoretically in Section 3.

We analyze theoretically the collapse phenomenon by first proving that if a NN is a constant function then there must exist a layer with output 0 and the gradients of weights and biases in all the previous layers vanish (Lemma 1, Corollary 2, and Lemma 3). Subsequently, we prove that if such conditions are met, then the NN will converge to a constant value depending on the loss function (Theorem 4). Furthermore, if the output of NN is equal to the mean value of the target function, the gradients of weights and biases vanish (Corollaries 5 and 6). In Lemma 7 and Theorem 8 and Proposition 9, we derive estimates of the probability of collapse for general cases, and in Proposition 10, we derive a more precise estimate for deep NNs with width 2. These theoretical estimates are verified numerically by tests using NNs with different layers and widths. Based on these results, we construct a diagram which can be used as a practical guideline in designing deep and narrow NNs that do not suffer from the collapse phenomenon.

Finally, we examine different methods of preventing deep and narrow NNs from converging to erroneous states. In particular, we find that asymmetric initializations including orthogonal initialization and LSUV cannot be used to avoid this collapse. However, some normalization techniques such as batch normalization and SELU can be used successfully to prevent the collapse of deep and narrow NNs; on the other hand, weight normalization fails. Similarly, we examine the effect of dropout which, however, also fails.

This work received support by the DARPA EQUiPS grant N66001-15-2-4055, the NSF grant DMS-1736088, the AFOSR grant FA9550-17-1-0013. The research of the second author was partially supported by the NSF of China 11771083 and the NSF of Fujian 2017J01556, 2016J01013.

References

Appendix A Proof of Lemma 1

Now let us go back to the proof of Lemma 1.

By assumption A2 and Lemma 11, N(x0)=WLxL−1(x0)+bL\mathcal{N}(\mathbf{x}^{0})=\mathbf{W}^{L}\mathbf{x}^{L-1}(\mathbf{x}^{0})+\mathbf{b}^{L} is a constant function, iff xL−1(x0)\mathbf{x}^{L-1}(\mathbf{x}^{0}) is a constant function with respect to x0\mathbf{x}^{0}. So we can assume that there is ReLU in the last layer, and prove that there exists a layer l∈{1,…,L}l\in\{1,\dots,L\}, s.t., hl≤0\mathbf{h}^{l}\leq\mathbf{0} and xl=0\mathbf{x}^{l}=\mathbf{0} wp1 for every x0∈Ω\mathbf{x}^{0}\in\Omega. We proceed in two steps.

ii) Assume the theorem is true for LL. Then for L+1L+1, if x1=0\mathbf{x}^{1}=0, choose l=1l=1 and we are done; otherwise, consider the NN without the first layer with x1∈Ω1\mathbf{x}^{1}\in\Omega^{1} as the input, denoted N1\mathcal{N}_{1}. By i, Ω1\Omega^{1} is a connected space with at least two points. Because N1\mathcal{N}_{1} is a constant function of x1\mathbf{x}^{1} and has LL layers, by induction, there exists a layer whose output is zero. Therefore, for the original neural network N\mathcal{N}, the output of such layer is also zero.

By i and ii, the statement is true for any LL. ∎

Appendix B Proof of Corollary 2

By Lemma 1, there exists a layer l∈{1,…,L−1}l\in\{1,\dots,L-1\}, s.t. hl≤0\mathbf{h}^{l}\leq\mathbf{0} and xl=0\mathbf{x}^{l}=\mathbf{0} wp1. Because N\mathcal{N} is bias-free, hl+1=Wl+1xl=0\mathbf{h}^{l+1}=\mathbf{W}^{l+1}\mathbf{x}^{l}=\mathbf{0} and xl+1=ReLU(hl+1)=0\mathbf{x}^{l+1}=\text{ReLU}(\mathbf{h}^{l+1})=\mathbf{0} wp1. By induction, for any n≥ln\geq l, hn≤0\mathbf{h}^{n}\leq\mathbf{0} and xn=0\mathbf{x}^{n}=\mathbf{0} wp1. ∎

Appendix C Proof of Lemma 3

Because xl≡0\mathbf{x}^{l}\equiv\mathbf{0}, it is then obvious by backpropagation. ∎

Appendix D Proof of Theorem 4

Appendix E Proof of Corollary 5

Appendix F Proof of Corollary 6

It suffices to show that gradients vanish for x0∈Ki\mathbf{x}^{0}\in K_{i}, i=1,…,ni=1,\dots,n and x0∈Ω∖∪i=1nKi\mathbf{x}^{0}\in\Omega\setminus\cup_{i=1}^{n}K_{i}.

ii) For x0∈Ω∖∪i=1nKi\mathbf{x}^{0}\in\Omega\setminus\cup_{i=1}^{n}K_{i}, the loss at x0\mathbf{x}^{0} is 0, so gradients vanish.

By i and ii, gradients vanish when using the L2L^{2} (MSE) loss. ∎

Appendix G Proof of Lemma 7

Let x=(x1,x2,…,xdin)\mathbf{x}=(x_{1},x_{2},\dots,x_{d_{in}}) be any input, and y=(y1,y2,…,ydout)\mathbf{y}=(y_{1},y_{2},\dots,y_{d_{out}}) be the corresponding output. For i=1,…,douti=1,\dots,d_{out},

Because (wi1,…,widin,bi)(w_{i1},\dots,w_{id_{in}},b_{i}) is a (din+1)(d_{in}+1)-dim vector initialized by a symmetric distribution, then

Appendix H Proof of Theorem 8

Appendix I Proof of Proposition 9

If the last layer also has ReLU activation, by Lemma 7,

If the last layer does not have ReLU activation, and L≥2L\geq 2, then

For L=1L=1, N\mathcal{N} is a single layer perceptron, which is a trivial case. ∎

Appendix J Proof of Proposition 10

where w1w_{1}, w2w_{2}, w1∗w_{1}^{*}, w2∗w_{2}^{*} are some coefficients.

Each case in the llth hidden layer may also induce all 16 cases in the (l+1)(l+1)th layer. For any given case in the llth hidden layer, we will compute the probabilities of these 16 cases for the (l+1)(l+1)th layer as follows.

Note that ω=(ω1,ω2)\boldsymbol{\omega}=(\omega_{1},\omega_{2}) lies in the first quadrant, and ω∗=(ω1∗,ω2∗)\boldsymbol{\omega}^{*}=(\omega_{1}^{*},\omega_{2}^{*}) lies in the third quadrant. Then the output of the next layer is

Since the matrix \left(\begin{array}[]{cc}A_{11}&A_{12}\\ A_{21}&A_{22}\end{array}\right) is random, for fixed ω\boldsymbol{\omega} and ω∗\boldsymbol{\omega}^{*}, the probability of case (1) is (∠(ω,ω∗)2π)2\left(\frac{\angle(\boldsymbol{\omega},\boldsymbol{\omega}^{*})}{2\pi}\right)^{2}. Without loss of generality, we can assume that ∥ω∥=∥ω∗∥=1\|\boldsymbol{\omega}\|=\|\boldsymbol{\omega}^{*}\|=1, and hence we can assume that ω=(cos⁡θ,sin⁡θ)\boldsymbol{\omega}=(\cos\theta,\sin\theta), θ∈(0,π2)\theta\in(0,\frac{\pi}{2}) and ω∗=(cos⁡ψ,sin⁡ψ)\boldsymbol{\omega}^{*}=(\cos\psi,\sin\psi), ψ∈(π,3π2)\psi\in(\pi,\frac{3\pi}{2}). It is easy to see that

Since ω,ω∗\boldsymbol{\omega},\boldsymbol{\omega}^{*} are random, the probability of case (1) is

Similarly, the probability of cases (6), (11) and (16) in the (l+1)(l+1)th layer are also 1796\frac{17}{96}. For cases (2), (3), (5), (8), (9), (12), (14) and (15), the probability is

For cases (4), (7), (10) and (13), the probability is

ii) Case (2) (the same method can be applied for cases (3), (5) and (9))

Note that in this case we can assume that ω=(cos⁡θ,sin⁡θ)\boldsymbol{\omega}=(\cos\theta,\sin\theta), θ∈(0,π2)\theta\in(0,\frac{\pi}{2}) and ω∗=(−1,0)\boldsymbol{\omega}^{*}=(-1,0) is a constant vector. It is easy to see that ∠(ω,ω∗)=π−θ\angle(\boldsymbol{\omega},\boldsymbol{\omega}^{*})=\pi-\theta, and hence the probabilities of cases (1), (6), (11) and (16) are

Similarly, the probabilities of cases (2), (3), (5), (8), (9), (12), (14) and (15) are

and the probabilities of cases (4), (7), (10) and (13) are

iii) Case (4) (the same method can be applied for cases (8) and (12))

It is easy to see that the probabilities of cases (4), (8), (12) and (16) are 14\frac{1}{4}, and the probabilities of all other cases are 0.

iv) Case (6) (the same method can be applied for case (11))

Note that in this case, ω1>0\omega_{1}>0 and ω1∗<0\omega_{1}^{*}<0, and thus it is not hard to see that the probabilities of cases (1), (6), (11) and (16) are 14\frac{1}{4}, and the probabilities of all the other cases are 0.

v) Case (7) (the same method can be applied for case (10))

Therefore, the probabilities of all the 16 cases are 116\frac{1}{16}.

vi) Case (13) (the same method can be applied for cases (14) and (15))

Similar to the argument of the case (4)(4), it is easily to see that the probabilities for cases (13), (14), (15) and (16) are 14\frac{1}{4}, and the probabilities for all other cases are .

The output of the next layer is the case (16) with probability 11.

By i, ii, iii, iv, v, vi and vii, we can get the probability transition matrix

where PjiP_{ji} is the probability of that the (l+1)(l+1)th layer is case jj when the iith layer is case ii.

Furthermore, direct computations show that the probability distribution vector of the first hidden layer π1\pi^{1} is

Therefore, the probability distribution of the llth hidden layer is