A powerful tool from fixed point theory to analyze and solve optimization and inclusion problems in a real Hilbert space H is the class of averaged nonexpansive operators, which was introduced in . Recall that an operator T:H→H is nonexpansive if it is 1-Lipschitzian, and α-averaged for some α∈]0,1] if there exists a nonexpansive operator Q:H→H such that T=(1−α)Id+αQ; if α=1/2, T is firmly nonexpansive. The importance of firmly nonexpansive operators in convex optimization and variational methods has long been recognized . More generally, averaged operators were shown in to play a prominent role in the analysis of convex feasibility problems. In this context the underlying problem is to find a common fixed point of averaged operators. In , it was shown that many convex minimization and monotone inclusion problems reduce to the more general problem of finding a fixed point of compositions of averaged operators, which provided a unified analysis of various proximal splitting algorithms. Along these lines, several fixed point methods based on various combinations of averaged operators have since been devised, see for recent work. Motivated by deep neural network structures with thus far elusive asymptotic properties, we investigate in the present paper a novel averaged operator model involving a mix of nonlinear and linear operators.
Artificial neural networks have attracted considerable attention as a tool to better understand, model, and imitate the human brain . In a Hilbertian setting , an (n+1)-layer feed-forward neural network architecture acting on real Hilbert spaces (Hi)0⩽i⩽n is defined as the composition of operators Rn∘(Wn⋅+bn)∘⋯∘R1∘(W1⋅+b1) where, for every i∈{1,…,n}, Ri:Hi→Hi is a nonlinear operator known as an activation operator, Wi:Hi−1→Hi is a linear operator, known as a weight operator, and bi∈Hi is a so-called bias parameter. Deep neural networks feature a (possibly large) number n of layers. In recent years, they have been found to be quite successful in a wide array of classification, recognition, and prediction tasks; see and its bibliography. Despite their success, the operational structure and properties of deep neural networks are not yet well understood from a mathematical viewpoint. In the present paper, we propose to analyze them within the following iterative model. We emphasize that our purpose is not to study the training of the network, which consists of optimally setting the weight operators and bias parameters from data samples, but to analyze mathematically such a structure once it is trained. Our model is also of general interest in constructive fixed point theory.
Our contributions are articulated around the following findings.
We show that a wide range of activation operators used in neural networks are actually proximity operators, which paves the way to the analysis of such networks via fixed point theory.
We provide a new analysis of compositions of proximity and affine operators, establishing mild conditions that guarantee that the resulting operator is averaged.
We show that, under suitable assumptions, the asymptotic output of the network converges to a point defined via a variational inequality. Furthermore, in general, this variational inequality does not derive from a minimization problem.
The remainder of the paper is organized as follows. In Section 2, we bring to light strong connections between the activation functions employed in neural networks and the theory of proximity operators in convex analysis. In Section 3 we derive new results on the averagedness properties of compositions of proximity and affine operators acting on different spaces. In Section 4, we investigate the asymptotic behavior of a class of deep neural networks structures and show that their fixed points solve a variational inequality. The main assumption on this subclass of Model 1.1 is that the structure of the network is periodic in the sense that a group of layers is repeated. Finally, in Section 5, the same properties are established for a broader class of networks.
Proximal activation in neural networks
Let φ∈Γ0(H). Then the following hold:
[8, Proposition 12.29] Fixproxφ=Argminφ.
[8, Corollary 24.5] Let g∈Γ0(H) be such that φ=g−∥⋅∥2/2. Then proxφ=∇g∗.
which was initially proposed in . As this discontinuous activation model may lead to unstable neural networks, various continuous approximations have been proposed. Our key observation is that a vast array of activation functions used in neural networks belong to the following class.
Remarkably, we can precisely characterize this class of activation functions as that of proximity operators.
The most basic activation function is ϱ=Id=prox0. It is in particular useful in dictionary learning approaches, which correspond to the linear special case of Model 1.1 .
can be written as ϱ=proxϕ, where ϕ is the indicator function of $$.
The rectified linear unit (ReLU) activation function
can be written as ϱ=proxϕ, where ϕ is the indicator function of [0,+∞[.
Let α∈]0,1]. The parametric rectified linear unit activation function is
We have ϱ=proxϕ, where
Proof. This follows from [23, Lemma 2.6 and Example 2.18].
Proof. Let ξ∈]−1,1[=dom∇ϕ=dom∂ϕ=ranproxϕ. Then ξ+ϕ′(ξ)=ξ/1−ξ2 and therefore proxϕ=(Id+ϕ′)−1:μ↦μ/1+μ2.
The inverse square root linear unit activation function
can be written as ϱ=proxϕ, where
Proof. Let ξ∈]−1,+∞[=dom∇ϕ=ranproxϕ. Then ξ+ϕ′(ξ)=ξ if ξ⩾0, and ξ+ϕ′(ξ)=ξ/1−ξ2 if ξ<0. Hence, ϱ=(Id+ϕ′)−1 is given by (2.8).
The arctangent activation function (2/π)arctan is the proximity operator of
Proof. Let ξ∈]−1,1[=dom∇ϕ=ranproxϕ. Then ξ+ϕ′(ξ)=tan(πξ/2) and therefore ϱ=(Id+ϕ′)−1=(2/π)arctan.
The hyperbolic tangent activation function tanh is the proximity operator of
Proof. Let ξ∈]−1,1[=dom∇ϕ=ranproxϕ. Then ξ+ϕ′(ξ)=arctanh(ξ) and therefore ϱ=(Id+ϕ′)−1=tanh.
Proof. Let ξ∈]−1/2,1/2[=dom∇ϕ=ranproxϕ. Then ξ+ϕ′(ξ)=ln((1+2ξ)/(1−2ξ)) and therefore proxϕ=(Id+ϕ′)−1:μ↦(1/2)(eμ−1)/(eμ+1)=1/(1+e−μ)−1/2.
Examples 2.12 and 2.13 are closely related in the sense that the function of (2.12) can be written as ϱ=(1/2)tanh(⋅/2).
Proof. Let ξ∈]−1,1[=dom∇ϕ=ranproxϕ. Then ξ+ϕ′(ξ)=ξ/(1−∣ξ∣) and therefore proxϕ=(Id+ϕ′)−1:μ↦μ/(1+∣μ∣).
The inverse hyperbolic sine activation function arcsinh is the proximity operator of ϕ=cosh−∣⋅∣2/2.
Proof. We have ϕ′:ξ↦sign(ξ)(e∣ξ∣−1)−ξ. Hence (Id+ϕ′):ξ↦sign(ξ)(e∣ξ∣−1) and, in turn, proxϕ=(Id+ϕ′)−1:ξ↦sign(ξ)ln(1+∣ξ∣).
Proof. (i)–(iii): This follows at once from Definition 2.2.
(iv)–(v): The fact that the resulting operators are proximity operators is established in [21, Section 3.3]. The fact that they are proximity operators of a function ϕ∈Γ0(H) that is minimal at is equivalent to the fact that proxϕ0=0 Lemma 2.1(i). This identity is easily seen to hold in each instance.
(vi): Set ϱ=ϱ1∘(2ϱ2−Id)+Id−ϱ2. Then ϱ is firmly nonexpansive [8, Proposition 4.31(ii)]. It is therefore increasing and nonexpansive. Finally, ϱ(0)=0.
Using Proposition 2.18, the above examples can be combined to obtain additional activation functions. For instance, it follows from Example 2.5 and Proposition 2.18(iv) that the soft thresholder
2 Activation operators
In Section 2.1, we have described activation functions which model neuronal activity in terms of a scalar function. In this section, we extend this notion to more general activation operators.
Let H be a real Hilbert space and let R:H→H. Then R belongs to the class A(H) if there exists a function φ∈Γ0(H) which is minimal at the zero vector and such that R=proxφ.
Property (ii) below shows that activation operators in A(H) have strong stability properties. On the other hand, the boundedness property (iv) is important in neural network-based functional approximation .
Let H be a real Hilbert space and let R∈A(H). Then the following hold:
Let x and y be in H. Then ∥Rx−Ry∥2⩽∥x−y∥2−∥x−y−Rx+Ry∥2.
Let x∈H. Then ∥Rx∥⩽∥x∥.
Let φ∈Γ0(H) be such that R=proxφ. Then ranR is bounded if and only if domφ is bounded.
Proof. (i): This follows from Lemma 2.1(i).
(ii): This follows from the firm nonexpansiveness of proximity operators [8, Proposition 12.28].
(iv): We have ranR=ran(Id+∂φ)−1=dom(Id+∂φ)=dom∂φ. On the other hand, dom∂φ is a dense subset of domφ [8, Corollary 16.39].
Let H and G be real Hilbert spaces. Then the following hold:
Let L∈B(H,G) be such that ∥L∥⩽1 and let R∈A(H). Then L∗∘R∘L∈A(H).
Let (Ri)i∈I be a finite family in A(H) and let (ωi)i∈I be real numbers in ]0,1] such that ∑i∈Iωi=1. Then ∑i∈IωiRi∈A(H).
Let R∈A(H). Then Id−R∈A(H).
Let R1∈A(H) and R2∈A(H). Then (R1−R2+Id)/2∈A(H).
Proof. The fact that the resulting operators are proximity operators is established in [21, Section 3.3]. In addition, is clearly a fixed point of the resulting operators. In view of Lemma 2.1(i), the proof is complete.
Then ψ=g−∥⋅∥2/2 and [40, Section 16] asserts that
Since ∇g∗=R+u, according to Lemma 2.1(ii), R=proxψ−u. We complete the proof by invoking the shift properties of proximity operators [8, Proposition 24.8(iii)].
Separable activation operators supply another important instance of activation operators.
Compositions of firmly nonexpansive and affine operators
Our analysis will revolve around the following property for a family of linear operators (Wi)1⩽i⩽m+1.
Let m⩾0 be an integer, let (Hi)0⩽i⩽m be real Hilbert spaces, set Hm+1=H0, and let α∈[1/2,1]. For every i∈{1,…,m+1}, let Wi∈B(Hi−1,Hi) and set
It is required that, for every x=(xi)0⩽i⩽m∈H0×⋯×Hm such that
In Condition 3.1, we take α⩾1/2 because, if x=(xi)0⩽i⩽m∈(H0∖{0})×H1×⋯×Hm satisfies (3.3), then 2m+1(1−α)∥x0∥⩽∥Lm+1x−2m+1(1−α)x0∥+∥Lm+1x∥⩽2m+1α∥x0∥.
We establish some preliminary results before providing properties that imply Condition 3.1.
Let m⩾1 be an integer, let (Hi)0⩽i⩽m be real Hilbert spaces, and set θ0=1. For every i∈{1,…,m}, let Wi∈B(Hi−1,Hi) and set
Let (xi)0⩽i⩽m∈H0×⋯×Hm be such that (3.2) is satisfied. Then the following hold:
(∀i∈{1,…,m})θi=∑k=0i−1θk∥Wi∘⋯∘Wk+1∥.
(∀i∈{1,…,m})∥xi∥⩽θi∥x0∥.
Proof. (i): This follows recursively from (3.4).
(ii): For every i∈{1,…,m}, let Li be as in (3.1). We proceed by induction on m. We first observe that the inequality is satisfied if m=1 since ∥x1∥⩽∥L1x0∥=∥W1x0∥⩽∥W1∥∥x0∥=θ1∥x0∥. Now assume that m⩾2 and that the inequalities hold for (x1,…,xm−1). Then, since (i) yields
Let H be a real Hilbert space, and let x and y be in H. Then
Proof. Since ∥x+y∥2−2∥x+y∥(∥x∥+∥y∥)+(∥x∥+∥y∥)2⩾0, we have
Let m⩾0 be an integer, and let (Hi)0⩽i⩽m be real Hilbert spaces. Let X be the standard vector space H0×⋯×Hm equipped with the norm ∥⋅∥X:x=(xi)0⩽i⩽m↦max0⩽i⩽m∥xi∥ and let Y be the standard vector space H0×H0 equipped with the norm ∥⋅∥Y:y=(y1,y2)↦∥y1∥+∥y2∥. Henceforth, the norm of M∈B(X,Y) is denoted by ∥M∥X,Y.
Let m⩾0 be an integer, let (Hi)0⩽i⩽m be nonzero real Hilbert spaces, set Hm+1=H0, and use Notation 3.5. For every i∈{1,…,m+1}, let Wi∈B(Hi−1,Hi). Further, let α∈[1/2,1], let θ0=1, let (θi)1⩽i⩽m+1 be as in (3.4), and set
There exists i∈{1,…,m+1} such that Wi=0.
∥M∥X,Y⩽1.
∥W−2m+1(1−α)Id∥−∥W∥+2θm+1⩽2m+1α.
α=1, for every i∈{1,…,m+1}Wi=0, and there exists η∈[0,α/((1−α)θm+1)] such that
Then (Wi)1⩽i⩽m+1 satisfies Condition 3.1.
Proof. We use the operators (Li)1⩽i⩽m+1 introduced in Condition 3.1. Per Notation 3.5 and (3.9d),
Now let x∈X be such that
(i): We assume that m⩾1. For every k∈{i,…,m}, it follows from (3.4) that θk=0 and in turn from Lemma 3.3(ii) and (3.13) that xk=0. Therefore,
(ii): In view of (i), we assume that, if m⩾1, (∀i∈{1,…,m})Wi=0. We then derive from (3.4) that (∀i∈{1,…,m})θi⩾∏k=1i∥Wk∥>0. If x0=0, (3.3) trivially follows from Lemma 3.3(ii), we therefore assume otherwise. Now set
According to Lemma 3.3(ii), (∀i∈{0,…,m})∥yi∥⩽1. On the other hand, it follows from (3.9c), (3.15), and (3.1) that My=Lm+1x/∥x0∥. Altogether, we deduce from (3.12) that (3.3) holds.
(iii)⇒(ii): Take y∈X such that ∥y∥X⩽1. Then it follows from (3.9c) and Lemma 3.3(i) that
In turn, (3.11) yields ∥M∥X,Y⩽1.
(iv)⇒(ii): Let y=(y0,…,ym)∈X be such that ∥y0∥=⋯=∥ym∥=1, and set
Therefore, since (3.21) implies that α−η(1−α)∥u∥⩾0, it results from (3) that
However, since (3.20) implies that α−η(1−α)∥W∥⩾0, while (3.17) implies that ∥u∥⩽θm+1−∥W∥, we derive from (3) that
Hence, using (3), (3.27), (3.9c), (3.9a), and (3.9d) we obtain
Now set \boldsymbol{C}=\big{\{}{\boldsymbol{y}\in{\boldsymbol{\mathcal{X}}}}~{}\big{|}~{}{\|y_{0}\|=\cdots=\|y_{m}\|=1}\big{\}}. Then, in view of (3.11), (3), and [8, Proposition 11.1(ii)], we conclude that ∥M∥X,Y=supy∈convC∥My∥Y=supy∈C∥My∥Y⩽1.
The next result establishes a link between deep neural network structures and the operators introduced in (3.1).
Let m⩾1 be an integer and let (Hi)0⩽i⩽m+1 be nonzero real Hilbert spaces. For every i∈{1,…,m+1}, let Wi∈B(Hi−1,Hi) and let Li be as in (3.1). Further, for every i∈{1,…,m}, let Pi:Hi→Hi be firmly nonexpansive. Set
let x and y be distinct points in H0, and set v0=(x−y)/∥x−y∥. Then there exists (v1,…,vm)∈H1×⋯×Hm such that
Proof. For every i∈{1,…,m}, since Pi is firmly nonexpansive, there exists a nonexpansive operator Qi:Hi→Hi such that
We proceed by induction on m. Suppose that m=1 and set
which implies that ∥v1∥⩽∥W1(x−y)∥/∥x−y∥=∥L1v0∥. Then
Thus, (3.30) holds for m=1. Next, we assume that m>1 and that there exists (v1,…,vm−1)∈H1×⋯×Hm−1 such that
In addition, it follows from (3.34) and (3.35) that
We now establish connections between Condition 3.1 for linear operators and the concept of averagedness for composite nonlinear operators.
Let m⩾1 be an integer, let (Hi)0⩽i⩽m−1 be nonzero real Hilbert spaces, set Hm=H0, and let α∈[1/2,1]. For every i∈{1,…,m}, let Wi∈B(Hi−1,Hi) and let Pi:Hi→Hi be firmly nonexpansive. Suppose that (Wi)1⩽i⩽m satisfies Condition 3.1. Then Pm∘Wm∘⋯∘P1∘W1 is α-averaged.
Proof. Set T=Pm∘Wm∘⋯∘P1∘W1. We must show that
is nonexpansive. By assumption, for every i∈{1,…,m}, there exists a nonexpansive operator Qi:Hi→Hi such that (3.31) holds. Let (Li)1⩽i⩽m be as in (3.1) and let x and y be distinct points in H0. According to Lemma 3.7, there exists v=(v0,…,vm−1)∈H0×⋯×Hm−1 such that
In turn, we derive from (3.38) and (3.31) that
which establishes the nonexpansiveness of Q.
Consider Theorem 3.8 with m=2. In view of Proposition 3.6(iii), P2∘W2∘P1∘W1 is α-averaged if ∥W2∘W1−4(1−α)Id∥+∥W2∘W1∥+2∥W2∥∥W1∥⩽4α. In particular, if α=1, this condition is obviously less restrictive than requiring that W1 and W2 be nonexpansive.
A variational inequality model
In this section, we first investigate an autonomous version of Model 1.1.
where x=(x1,…,xm) denotes a generic element in H.
We start with a property of the compositions of the operators (Ti)1⩽i⩽m of (4.1).
Consider the setting of Model 4.1, let i and j be integers such that 1⩽j⩽i⩽m, and let x∈Hj−1. Then
Proof. In view of (4.1), the property is satisfied when i=j. We now assume that i>j. Since Ri∈A(Hi), Proposition 2.21(i) yields
Next, we establish a connection between Model 4.1 and a variational inequality.
In the setting of Model 4.1, consider the variational inequality problem
Suppose that (Wi)1⩽i⩽m satisfies Condition 3.1 for some α∈[1/2,1]. Then F is closed and convex.
Suppose that (Wi)1⩽i⩽m satisfies Condition 3.1 for some α∈[1/2,1] and that one of the following holds:
ran(Tm∘⋯∘T1) is bounded.
There exists j∈{1,…,m} such that domφj is bounded.
Then F and F are nonempty.
Suppose that Id−W∘S is monotone. Then F is closed and convex. In addition, F and F are nonempty if any of the following holds:
Id−W∘S+∂φ is surjective.
∂φ−W∘S is maximally monotone.
max1⩽i⩽m∥Wi∥⩽1, S∗−W has closed range, and ker(S−W∗)={0}.
max1⩽i⩽m∥Wi∥⩽1 and, for every i∈{1,…,m}, domφi∗=Hi.
For every i∈{1,…,m}, domφi=H and domφi∗=Hi.
S∗−W has closed range, ker(S−W∗)={0}, and, for every i∈{1,…,m}, domφi=Hi.
For every i∈{1,…,m}, domφi is bounded.
Proof. We first observe that S∈B(H,H→), W∈B(H→,H), φ∈Γ0(H), and ψ∈Γ0(H).
(i): Let x∈H. Then
(ii): Let x∈H. Using (4.2), we obtain
(iii): Clear from the definitions of F and F.
(iv): Define m firmly nonexpansive operators by (∀i∈{1,…,m})Pi:Hi→Hi:y↦Ri(y+bi). Then it follows from (4.1) and Theorem 3.8 applied to (Pi)1⩽i⩽m that Tm∘⋯∘T1 is nonexpansive. In turn, we derive from [8, Corollary 4.24] that its fixed point set F is closed and convex.
(v): Thanks to (iii), it is enough to show that F=∅. Set T=Tm∘⋯∘T1 and recall that it is nonexpansive by virtue of Theorem 3.8.
(v)(a): Let C be a closed ball such that ranT⊂C and set S=T∣C. Then S:C→C is nonexpansive and therefore [8, Proposition 4.29] asserts that FixT=FixS=∅.
(v)(b)⇒(v)(a): We have ranTj⊂ranRj=ranproxφj=dom(Id+∂φj)=dom∂φj⊂domφj. Hence ranTj is bounded and Proposition 4.2 (with i=m) implies that
(vi): Set A=Id−W∘S+∂ψ. Since Id−W∘S is monotone and continuous, it is maximally monotone [8, Corollary 20.28], with H as its domain. Since ∂ψ is also maximally monotone [8, Theorem 20.25], A is likewise [8, Corollary 25.5(i)] and hence F=zerA is closed and convex [8, Proposition 23.39]. Next, we note that, in view of (iii), F=∅⇔F=∅.
(vi)(a): The hypothesis implies that (bi)1⩽i⩽m∈ran(Id−W∘S+∂φ) and therefore that (4.5) has a solution, i.e., F=∅.
(vi)(b)⇒(vi)(a): The claim follows from Minty’s theorem [8, Theorem 21.1].
(vi)(c)⇒(vi)(a): We have ∥W∘S∥=∥W∥=max1⩽i⩽m∥Wi∥⩽1. Therefore, −W∘S is nonexpansive, which implies that (Id−W∘S)/2 is firmly nonexpansive [8, Corollary 4.5], that is (∀x∈H)⟨x−W(Sx)∣x⟩⩾∥x−W(Sx)∥2/2. Consequently, Id−W∘S is 3∗ monotone [8, Proposition 25.16], while ∂φ is also 3∗ monotone [8, Example 25.13]. Finally, since S is unitary,
which shows that Id−W∘S is surjective. Altogether, since [8, Corollary 25.5(i)] implies that Id−W∘S+∂φ is maximally monotone, it follows from [8, Corollary 25.27(i)] that Id−W∘S+∂φ is surjective.
(vi)(d)⇒(vi)(a): We have domφ∗=H. Hence since intdomφ∗⊂dom∂φ∗ [8, Proposition 16.27], we have ran∂φ=dom(∂φ)−1=dom∂φ∗=H. Hence, ∂φ is surjective. We conclude using the same arguments as in (vi)(c): ∂φ and Id−W∘S are both 3∗ monotone and their sum is maximally monotone, which allows us to invoke [8, Corollary 25.27(i)].
(vi)(e)⇒(vi)(a): As seen in (vi)(d), ∂φ is surjective. We have H=intdomφ⊂dom∂φ [8, Proposition 16.27]. Consequently, H=dom(Id−W∘S)⊂dom∂φ. Altogether, since ∂φ is 3∗ monotone, it follows from [8, Corollary 25.27(ii)] that Id−W∘S+∂φ is surjective.
(vi)(f)⇒(vi)(a): As seen in (vi)(c), Id−W∘S is surjective and ∂φ is 3∗ monotone. In addition, dom(Id−W∘S)⊂dom∂φ since H=intdomφ⊂dom∂φ [8, Proposition 16.27]. Altogether, it follows from [8, Corollary 25.27(ii)] that Id−W∘S+∂φ is surjective.
(vi)(g): Here \text{\rm dom}\,\boldsymbol{A}=\text{\rm dom}\,\partial\boldsymbol{\varphi}\subset\text{\rm dom}\,\boldsymbol{\varphi}=\raisebox{-1.42262pt}{\mbox{\LARGE{\times}}}_{\!i=1}^{\!m}\text{\rm dom}\,\varphi_{i} is bounded. Hence, F=zerA=∅ [8, Proposition 23.36(iii)].
In Proposition 4.3(vi), it is required that Id−W∘S be monotone, or equivalently, that its self-adjoint part Id−(W∘S+S∗∘W∗)/2 be positive. In a finite-dimensional setting, this just means that the eigenvalues of the matrix WS+S∗W∗ are in ]−∞,2].
where, given x∈Hi−1, [Wix]k is the kth component of Wix and
Altogether, we conclude that F is a closed convex polyhedron.
2 Asymptotic analysis
Next, we investigate the asymptotic behavior of (1.2) in the context of Model 4.1.
In the setting of Model 4.1, set T=Tm∘⋯∘T1, let α∈[1/2,1], and suppose that the following hold:
(Wi)1⩽i⩽m satisfies Condition 3.1 with parameter α.
λn≡1/α=1 and Txn−xn→0.
For every i∈{1,…,m−1}, Ri is weakly sequentially continuous.
For every i∈{1,…,m−1}, Ri is a separable activation operator in the sense of Proposition 2.24.
For every i∈{1,…,m−1}, Hi is finite-dimensional.
Proof. We first derive from (1.2) and Model 4.1 that
Now set (∀i∈{1,…,m})Pi:Hi→Hi:y↦Ri(y+bi). Then (4.1) yields T=Pm∘Wm∘⋯∘P1∘W1 and, since the operators (Ri)1⩽i⩽m are firmly nonexpansive, the operators (Pi)1⩽i⩽m are likewise. Hence, it follows from (b), Theorem 3.8, and (4.2) that
We now prove the convergence of the individual sequences under each assumption.
(iii): We have already established that xn⇀xm. Since W1 is weakly continuous as a bounded linear operator, so is T1 in (4.1). Hence, (1.2) implies that x1,n=T1xn⇀T1xm=x1. Likewise, we obtain successively x2,n=T2x1,n⇀T2x1=x2, x3,n=T3x2,n⇀T3x2=x3,…, xm,n=Tmxm−1,n⇀Tmxm−1=xm.
(iv)⇒(iii): See [8, Proposition 24.12(iii)].
(v)⇒(iii): A proximity operator is nonexpansive and therefore continuous, hence weakly continuous in a finite-dimensional setting.
(vi): As shown above, xn⇀xm∈F. It follows from Proposition 3.6(iii) and Theorem 3.8 (applied with m=1) that, for every i∈{1,…,m}, Ti is βi-averaged. Hence, upon applying [24, Theorem 3.5(ii)] with α as an averaging constant of T, we infer that
Thus, x1,n−xn=T1xn−xn→T1xm−xm, which implies that x1,n=(x1,n−xn)+xn⇀(T1xm−xm)+xm=T1xm. However, since x2,n−x1,n=(T2∘T1)xn−T1xn→(T2∘T1)xm−T1xm, we obtain x2,n⇀(T2∘T1)xm. Continuing this telescoping process yields the claim.
The next result covers the case when the variational inequality problem (4.5) has no solution.
To model closely existing deep neural networks, we have chosen the activation operators in Definition 2.20 and Model 4.1 to be proximity operators. However, as is clear from the results of Section 3 and in particular the central Theorem 3.8, an activation operator Ri:Hi→Hi could more generally be a firmly nonexpansive operator that admits as a fixed point. By [8, Corollary 23.9], this means that Ri is the resolvent of some maximally monotone operator such Ai:Hi→2Hi (i.e., Ri=(Id+Ai)−1) such that 0∈Ai0. In this context, the variational inequality (4.5) assumes the more general form of a system of monotone inclusions, namely,
Analysis of nonperiodic networks
We analyze the deep neural network described in Model 1.1 in the following scenario.
In the setting of Model 1.1, suppose that Assumption 5.1 is satisfied, let i∈{1,…,m}, and set
We can now present the main result of this section on the asymptotic behavior of Model 1.1. The proof of this result relies on Theorem 4.7, which it extends.
Consider the setting of Model 1.1 and let α∈[1/2,1]. Suppose that Assumption 5.1 is satisfied as well as the following:
F=FixT=∅, where T=Tm∘⋯∘T1.
(Wi)1⩽i⩽m satisfies Condition 3.1 with parameter α.
λn≡α=1 and Txn−xn→0.
For every i∈{1,…,m−1}, Ri is weakly sequentially continuous.
For every i∈{1,…,m−1}, Ri is a separable activation function in the sense of Proposition 2.24.
For every i∈{1,…,m−1}, Hi is finite-dimensional.
In (c)(i)–(c)(ii) above, Proposition 4.3(iii) ensures that (T1xm,(T2∘T1)xm,…,(Tm−1∘⋯∘T1)xm,xm) solves (4.5).
(vi): For every i∈{1,…,m}, set
Thus, since Proposition 3.6(iii) and Theorem 3.8 imply that the operators (Ti)1⩽i⩽m are averaged, the proof can be completed as that of Theorem 4.7(vi) since [24, Theorem 3.5(ii)] asserts that (4.15) remains valid under (5.21).