Deep Neural Network Structures Solving Variational Inequalities

Patrick L. Combettes, Jean-Christophe Pesquet

Introduction

A powerful tool from fixed point theory to analyze and solve optimization and inclusion problems in a real Hilbert space H{\mathcal{H}} is the class of averaged nonexpansive operators, which was introduced in . Recall that an operator T ⁣:H→HT\colon{\mathcal{H}}\to{\mathcal{H}} is nonexpansive if it is 11-Lipschitzian, and α\alpha-averaged for some α∈]0,1]\alpha\in\left]0,1\right] if there exists a nonexpansive operator Q ⁣:H→HQ\colon{\mathcal{H}}\to{\mathcal{H}} such that T=(1−α)Id⁡+αQT=(1-\alpha)\operatorname{Id}+\alpha Q; if α=1/2\alpha=1/2, TT is firmly nonexpansive. The importance of firmly nonexpansive operators in convex optimization and variational methods has long been recognized . More generally, averaged operators were shown in to play a prominent role in the analysis of convex feasibility problems. In this context the underlying problem is to find a common fixed point of averaged operators. In , it was shown that many convex minimization and monotone inclusion problems reduce to the more general problem of finding a fixed point of compositions of averaged operators, which provided a unified analysis of various proximal splitting algorithms. Along these lines, several fixed point methods based on various combinations of averaged operators have since been devised, see for recent work. Motivated by deep neural network structures with thus far elusive asymptotic properties, we investigate in the present paper a novel averaged operator model involving a mix of nonlinear and linear operators.

Artificial neural networks have attracted considerable attention as a tool to better understand, model, and imitate the human brain . In a Hilbertian setting , an (n+1)(n+1)-layer feed-forward neural network architecture acting on real Hilbert spaces (Hi)0⩽i⩽n({\mathcal{H}}_{i})_{0\leqslant i\leqslant n} is defined as the composition of operators Rn∘(Wn⋅+bn)∘⋯∘R1∘(W1⋅+b1)R_{n}\circ(W_{n}\cdot+b_{n})\circ\cdots\circ R_{1}\circ(W_{1}\cdot+b_{1}) where, for every i∈{1,…,n}i\in\{1,\ldots,n\}, Ri ⁣:Hi→HiR_{i}\colon{\mathcal{H}}_{i}\to{\mathcal{H}}_{i} is a nonlinear operator known as an activation operator, Wi ⁣:Hi−1→HiW_{i}\colon{\mathcal{H}}_{i-1}\to{\mathcal{H}}_{i} is a linear operator, known as a weight operator, and bi∈Hib_{i}\in{\mathcal{H}}_{i} is a so-called bias parameter. Deep neural networks feature a (possibly large) number nn of layers. In recent years, they have been found to be quite successful in a wide array of classification, recognition, and prediction tasks; see and its bibliography. Despite their success, the operational structure and properties of deep neural networks are not yet well understood from a mathematical viewpoint. In the present paper, we propose to analyze them within the following iterative model. We emphasize that our purpose is not to study the training of the network, which consists of optimally setting the weight operators and bias parameters from data samples, but to analyze mathematically such a structure once it is trained. Our model is also of general interest in constructive fixed point theory.

Our contributions are articulated around the following findings.

We show that a wide range of activation operators used in neural networks are actually proximity operators, which paves the way to the analysis of such networks via fixed point theory.

We provide a new analysis of compositions of proximity and affine operators, establishing mild conditions that guarantee that the resulting operator is averaged.

We show that, under suitable assumptions, the asymptotic output of the network converges to a point defined via a variational inequality. Furthermore, in general, this variational inequality does not derive from a minimization problem.

The remainder of the paper is organized as follows. In Section 2, we bring to light strong connections between the activation functions employed in neural networks and the theory of proximity operators in convex analysis. In Section 3 we derive new results on the averagedness properties of compositions of proximity and affine operators acting on different spaces. In Section 4, we investigate the asymptotic behavior of a class of deep neural networks structures and show that their fixed points solve a variational inequality. The main assumption on this subclass of Model 1.1 is that the structure of the network is periodic in the sense that a group of layers is repeated. Finally, in Section 5, the same properties are established for a broader class of networks.

Proximal activation in neural networks

Let φ∈Γ0(H)\varphi\in\Gamma_{0}({\mathcal{H}}). Then the following hold:

[8, Proposition 12.29] Fix proxφ=Argmin φ\text{\rm Fix}\,\text{\rm prox}_{\varphi}=\text{\rm Argmin}\,\varphi.

[8, Corollary 24.5] Let g∈Γ0(H)g\in\Gamma_{0}({\mathcal{H}}) be such that φ=g−∥⋅∥2/2\varphi=g-\|\cdot\|^{2}/2. Then proxφ=∇g∗\text{\rm prox}_{\varphi}=\nabla g^{*}.

which was initially proposed in . As this discontinuous activation model may lead to unstable neural networks, various continuous approximations have been proposed. Our key observation is that a vast array of activation functions used in neural networks belong to the following class.

Remarkably, we can precisely characterize this class of activation functions as that of proximity operators.

The most basic activation function is ϱ=Id⁡=prox0\varrho=\operatorname{Id}=\text{\rm prox}_{0}. It is in particular useful in dictionary learning approaches, which correspond to the linear special case of Model 1.1 .

can be written as ϱ=proxϕ\varrho=\text{\rm prox}_{\phi}, where ϕ\phi is the indicator function of $$.

The rectified linear unit (ReLU) activation function

can be written as ϱ=proxϕ\varrho=\text{\rm prox}_{\phi}, where ϕ\phi is the indicator function of [0,+∞[\left[0,+\infty\right[.

Let α∈]0,1]\alpha\in\left]0,1\right]. The parametric rectified linear unit activation function is

We have ϱ=proxϕ\varrho=\text{\rm prox}_{\phi}, where

Proof. This follows from [23, Lemma 2.6 and Example 2.18].

Proof. Let ξ∈]−1,1[=dom ∇ϕ=dom ∂ϕ=ran proxϕ\xi\in\left]-1,1\right[=\text{\rm dom}\,\nabla\phi=\text{\rm dom}\,\partial\phi=\text{\rm ran}\,\text{\rm prox}_{\phi}. Then ξ+ϕ′(ξ)=ξ/1−ξ2\xi+\phi^{\prime}(\xi)=\xi/\sqrt{1-\xi^{2}} and therefore proxϕ=(Id⁡+ϕ′)−1 ⁣:μ↦μ/1+μ2\text{\rm prox}_{\phi}=(\operatorname{Id}+\phi^{\prime})^{-1}\colon\mu\mapsto\mu/\sqrt{1+\mu^{2}}.

The inverse square root linear unit activation function

can be written as ϱ=proxϕ\varrho=\text{\rm prox}_{\phi}, where

Proof. Let ξ∈]−1,+∞[=dom ∇ϕ=ran proxϕ\xi\in\left]-1,{+\infty}\right[=\text{\rm dom}\,\nabla\phi=\text{\rm ran}\,\text{\rm prox}_{\phi}. Then ξ+ϕ′(ξ)=ξ\xi+\phi^{\prime}(\xi)=\xi if ξ⩾0\xi\geqslant 0, and ξ+ϕ′(ξ)=ξ/1−ξ2\xi+\phi^{\prime}(\xi)=\xi/\sqrt{1-\xi^{2}} if ξ<0\xi<0. Hence, ϱ=(Id⁡+ϕ′)−1\varrho=(\operatorname{Id}+\phi^{\prime})^{-1} is given by (2.8).

The arctangent activation function (2/π)arctan(2/\pi)\text{arctan} is the proximity operator of

Proof. Let ξ∈]−1,1[=dom ∇ϕ=ran proxϕ\xi\in\left]-1,1\right[=\text{\rm dom}\,\nabla\phi=\text{\rm ran}\,\text{\rm prox}_{\phi}. Then ξ+ϕ′(ξ)=tan(πξ/2)\xi+\phi^{\prime}(\xi)=\text{tan}(\pi\xi/2) and therefore ϱ=(Id⁡+ϕ′)−1=(2/π)arctan\varrho=(\operatorname{Id}+\phi^{\prime})^{-1}=(2/\pi)\text{arctan}.

The hyperbolic tangent activation function tanh is the proximity operator of

Proof. Let ξ∈]−1,1[=dom ∇ϕ=ran proxϕ\xi\in\left]-1,1\right[=\text{\rm dom}\,\nabla\phi=\text{\rm ran}\,\text{\rm prox}_{\phi}. Then ξ+ϕ′(ξ)=arctanh(ξ)\xi+\phi^{\prime}(\xi)=\text{arctanh}(\xi) and therefore ϱ=(Id⁡+ϕ′)−1=tanh\varrho=(\operatorname{Id}+\phi^{\prime})^{-1}=\text{tanh}.

Proof. Let ξ∈]−1/2,1/2[=dom ∇ϕ=ran proxϕ\xi\in\left]-1/2,1/2\right[=\text{\rm dom}\,\nabla\phi=\text{\rm ran}\,\text{\rm prox}_{\phi}. Then ξ+ϕ′(ξ)=ln⁡((1+2ξ)/(1−2ξ))\xi+\phi^{\prime}(\xi)=\ln((1+2\xi)/(1-2\xi)) and therefore proxϕ=(Id⁡+ϕ′)−1 ⁣:μ↦(1/2)(eμ−1)/(eμ+1)=1/(1+e−μ)−1/2\text{\rm prox}_{\phi}=(\operatorname{Id}+\phi^{\prime})^{-1}\colon\mu\mapsto(1/2)(e^{\mu}-1)/(e^{\mu}+1)=1/(1+e^{-\mu})-1/2.

Examples 2.12 and 2.13 are closely related in the sense that the function of (2.12) can be written as ϱ=(1/2)tanh(⋅/2)\varrho=(1/2)\text{tanh}(\cdot/2).

Proof. Let ξ∈]−1,1[=dom ∇ϕ=ran proxϕ\xi\in\left]-1,1\right[=\text{\rm dom}\,\nabla\phi=\text{\rm ran}\,\text{\rm prox}_{\phi}. Then ξ+ϕ′(ξ)=ξ/(1−∣ξ∣)\xi+\phi^{\prime}(\xi)=\xi/(1-|\xi|) and therefore proxϕ=(Id⁡+ϕ′)−1 ⁣:μ↦μ/(1+∣μ∣)\text{\rm prox}_{\phi}=(\operatorname{Id}+\phi^{\prime})^{-1}\colon\mu\mapsto\mu/(1+|\mu|).

The inverse hyperbolic sine activation function arcsinh is the proximity operator of ϕ=cosh−∣⋅∣2/2\phi=\text{cosh}-|\cdot|^{2}/2.

Proof. We have ϕ′ ⁣:ξ↦sign(ξ)(e∣ξ∣−1)−ξ\phi^{\prime}\colon\xi\mapsto\text{\rm sign}(\xi)(e^{|\xi|}-1)-\xi. Hence (Id⁡+ϕ′) ⁣:ξ↦sign(ξ)(e∣ξ∣−1)(\operatorname{Id}+\phi^{\prime})\colon\xi\mapsto\text{\rm sign}(\xi)(e^{|\xi|}-1) and, in turn, proxϕ=(Id⁡+ϕ′)−1 ⁣:ξ↦sign(ξ)ln⁡(1+∣ξ∣)\text{\rm prox}_{\phi}=(\operatorname{Id}+\phi^{\prime})^{-1}\colon\xi\mapsto\text{\rm sign}(\xi)\ln(1+|\xi|).

Proof. (i)–(iii): This follows at once from Definition 2.2.

(iv)–(v): The fact that the resulting operators are proximity operators is established in [21, Section 3.3]. The fact that they are proximity operators of a function ϕ∈Γ0(H)\phi\in\Gamma_{0}({\mathcal{H}}) that is minimal at is equivalent to the fact that proxϕ0=0\text{\rm prox}_{\phi}0=0 Lemma 2.1(i). This identity is easily seen to hold in each instance.

(vi): Set ϱ=ϱ1∘(2ϱ2−Id⁡)+Id⁡−ϱ2\varrho=\varrho_{1}\circ(2\varrho_{2}-\operatorname{Id})+\operatorname{Id}-\varrho_{2}. Then ϱ\varrho is firmly nonexpansive [8, Proposition 4.31(ii)]. It is therefore increasing and nonexpansive. Finally, ϱ(0)=0\varrho(0)=0.

Using Proposition 2.18, the above examples can be combined to obtain additional activation functions. For instance, it follows from Example 2.5 and Proposition 2.18(iv) that the soft thresholder

2 Activation operators

In Section 2.1, we have described activation functions which model neuronal activity in terms of a scalar function. In this section, we extend this notion to more general activation operators.

Let H{\mathcal{H}} be a real Hilbert space and let R ⁣:H→HR\colon{\mathcal{H}}\to{\mathcal{H}}. Then RR belongs to the class A(H)\mathcal{A}({\mathcal{H}}) if there exists a function φ∈Γ0(H)\varphi\in\Gamma_{0}({\mathcal{H}}) which is minimal at the zero vector and such that R=proxφR=\text{\rm prox}_{\varphi}.

Property (ii) below shows that activation operators in A(H)\mathcal{A}({\mathcal{H}}) have strong stability properties. On the other hand, the boundedness property (iv) is important in neural network-based functional approximation .

Let H{\mathcal{H}} be a real Hilbert space and let R∈A(H)R\in\mathcal{A}({\mathcal{H}}). Then the following hold:

Let xx and yy be in H{\mathcal{H}}. Then ∥Rx−Ry∥2⩽∥x−y∥2−∥x−y−Rx+Ry∥2\|Rx-Ry\|^{2}\leqslant\|x-y\|^{2}-\|x-y-Rx+Ry\|^{2}.

Let x∈Hx\in{\mathcal{H}}. Then ∥Rx∥⩽∥x∥\|Rx\|\leqslant\|x\|.

Let φ∈Γ0(H)\varphi\in\Gamma_{0}({\mathcal{H}}) be such that R=proxφR=\text{\rm prox}_{\varphi}. Then ran R\text{\rm ran}\,R is bounded if and only if dom φ\text{\rm dom}\,\varphi is bounded.

Proof. (i): This follows from Lemma 2.1(i).

(ii): This follows from the firm nonexpansiveness of proximity operators [8, Proposition 12.28].

(iv): We have ran R=ran (Id⁡+∂φ)−1=dom (Id⁡+∂φ)=dom ∂φ\text{\rm ran}\,R=\text{\rm ran}\,(\operatorname{Id}+\partial\varphi)^{-1}=\text{\rm dom}\,(\operatorname{Id}+\partial\varphi)=\text{\rm dom}\,\partial\varphi. On the other hand, dom ∂φ\text{\rm dom}\,\partial\varphi is a dense subset of dom φ\text{\rm dom}\,\varphi [8, Corollary 16.39].

Let H{\mathcal{H}} and G{\mathcal{G}} be real Hilbert spaces. Then the following hold:

Let L∈B (H,G)L\in\mathcal{B}\,({\mathcal{H}},{\mathcal{G}}) be such that ∥L∥⩽1\|L\|\leqslant 1 and let R∈A(H)R\in\mathcal{A}({\mathcal{H}}). Then L∗∘R∘L∈A(H)L^{*}\circ R\circ L\in\mathcal{A}({\mathcal{H}}).

Let (Ri)i∈I(R_{i})_{i\in I} be a finite family in A(H)\mathcal{A}({\mathcal{H}}) and let (ωi)i∈I(\omega_{i})_{i\in I} be real numbers in ]0,1]\left]0,1\right] such that ∑i∈Iωi=1\sum_{i\in I}\omega_{i}=1. Then ∑i∈IωiRi∈A(H)\sum_{i\in I}\omega_{i}R_{i}\in\mathcal{A}({\mathcal{H}}).

Let R∈A(H)R\in\mathcal{A}({\mathcal{H}}). Then Id⁡−R∈A(H)\operatorname{Id}-R\in\mathcal{A}({\mathcal{H}}).

Let R1∈A(H)R_{1}\in\mathcal{A}({\mathcal{H}}) and R2∈A(H)R_{2}\in\mathcal{A}({\mathcal{H}}). Then (R1−R2+Id⁡)/2∈A(H)(R_{1}-R_{2}+\operatorname{Id})/2\in\mathcal{A}({\mathcal{H}}).

Proof. The fact that the resulting operators are proximity operators is established in [21, Section 3.3]. In addition, is clearly a fixed point of the resulting operators. In view of Lemma 2.1(i), the proof is complete.

Then ψ=g−∥⋅∥2/2\psi=g-\|\cdot\|^{2}/2 and [40, Section 16] asserts that

Since ∇g∗=R+u\nabla g^{*}=R+u, according to Lemma 2.1(ii), R=proxψ−uR=\text{\rm prox}_{\psi}-u. We complete the proof by invoking the shift properties of proximity operators [8, Proposition 24.8(iii)].

Separable activation operators supply another important instance of activation operators.

Compositions of firmly nonexpansive and affine operators

Our analysis will revolve around the following property for a family of linear operators (Wi)1⩽i⩽m+1(W_{i})_{1\leqslant i\leqslant m+1}.

Let m⩾0m\geqslant 0 be an integer, let (Hi)0⩽i⩽m({\mathcal{H}}_{i})_{0\leqslant i\leqslant m} be real Hilbert spaces, set Hm+1=H0{\mathcal{H}}_{m+1}={\mathcal{H}}_{0}, and let α∈[1/2,1]\alpha\in[1/2,1]. For every i∈{1,…,m+1}i\in\{1,\ldots,m+1\}, let Wi∈B (Hi−1,Hi)W_{i}\in\mathcal{B}\,({\mathcal{H}}_{i-1},{\mathcal{H}}_{i}) and set

It is required that, for every x=(xi)0⩽i⩽m∈H0×⋯×Hm\boldsymbol{x}=(x_{i})_{0\leqslant i\leqslant m}\in{\mathcal{H}}_{0}\times\cdots\times{\mathcal{H}}_{m} such that

In Condition 3.1, we take α⩾1/2\alpha\geqslant 1/2 because, if x=(xi)0⩽i⩽m∈(H0∖{0})×H1×⋯×Hm\boldsymbol{x}=(x_{i})_{0\leqslant i\leqslant m}\in({\mathcal{H}}_{0}\smallsetminus\{0\})\times{\mathcal{H}}_{1}\times\cdots\times{\mathcal{H}}_{m} satisfies (3.3), then 2m+1(1−α)∥x0∥⩽∥Lm+1x−2m+1(1−α)x0∥+∥Lm+1x∥⩽2m+1α∥x0∥2^{m+1}(1-\alpha)\|x_{0}\|\leqslant\|L_{m+1}\boldsymbol{x}-2^{m+1}(1-\alpha)x_{0}\|+\|L_{m+1}\boldsymbol{x}\|\leqslant 2^{m+1}\alpha\|x_{0}\|.

We establish some preliminary results before providing properties that imply Condition 3.1.

Let m⩾1m\geqslant 1 be an integer, let (Hi)0⩽i⩽m({\mathcal{H}}_{i})_{0\leqslant i\leqslant m} be real Hilbert spaces, and set θ0=1\theta_{0}=1. For every i∈{1,…,m}i\in\{1,\ldots,m\}, let Wi∈B (Hi−1,Hi)W_{i}\in\mathcal{B}\,({\mathcal{H}}_{i-1},{\mathcal{H}}_{i}) and set

Let (xi)0⩽i⩽m∈H0×⋯×Hm(x_{i})_{0\leqslant i\leqslant m}\in{\mathcal{H}}_{0}\times\cdots\times{\mathcal{H}}_{m} be such that (3.2) is satisfied. Then the following hold:

(∀i∈{1,…,m})(\forall i\in\{1,\ldots,m\}) θi=∑k=0i−1θk∥Wi∘⋯∘Wk+1∥\theta_{i}=\sum_{k=0}^{i-1}\theta_{k}\|W_{i}\circ\cdots\circ W_{k+1}\|.

(∀i∈{1,…,m})(\forall i\in\{1,\ldots,m\}) ∥xi∥⩽θi∥x0∥\|x_{i}\|\leqslant\theta_{i}\|x_{0}\|.

Proof. (i): This follows recursively from (3.4).

(ii): For every i∈{1,…,m}i\in\{1,\ldots,m\}, let LiL_{i} be as in (3.1). We proceed by induction on mm. We first observe that the inequality is satisfied if m=1m=1 since ∥x1∥⩽∥L1x0∥=∥W1x0∥⩽∥W1∥ ∥x0∥=θ1∥x0∥\|x_{1}\|\leqslant\|L_{1}x_{0}\|=\|W_{1}x_{0}\|\leqslant\|W_{1}\|\,\|x_{0}\|=\theta_{1}\|x_{0}\|. Now assume that m⩾2m\geqslant 2 and that the inequalities hold for (x1,…,xm−1)(x_{1},\ldots,x_{m-1}). Then, since (i) yields

Let H{\mathcal{H}} be a real Hilbert space, and let xx and yy be in H{\mathcal{H}}. Then

Proof. Since ∥x+y∥2−2∥x+y∥(∥x∥+∥y∥)+(∥x∥+∥y∥)2⩾0\|x+y\|^{2}-2\|x+y\|(\|x\|+\|y\|)+(\|x\|+\|y\|)^{2}\geqslant 0, we have

Let m⩾0m\geqslant 0 be an integer, and let (Hi)0⩽i⩽m({\mathcal{H}}_{i})_{0\leqslant i\leqslant m} be real Hilbert spaces. Let X{\boldsymbol{\mathcal{X}}} be the standard vector space H0×⋯×Hm{\mathcal{H}}_{0}\times\cdots\times{\mathcal{H}}_{m} equipped with the norm ∥⋅∥X ⁣:x=(xi)0⩽i⩽m↦max⁡0⩽i⩽m∥xi∥\|\cdot\|_{{\boldsymbol{\mathcal{X}}}}\colon\boldsymbol{x}=(x_{i})_{0\leqslant i\leqslant m}\mapsto\max_{0\leqslant i\leqslant m}\|x_{i}\| and let Y{\boldsymbol{\mathcal{Y}}} be the standard vector space H0×H0{\mathcal{H}}_{0}\times{\mathcal{H}}_{0} equipped with the norm ∥⋅∥Y ⁣:y=(y1,y2)↦∥y1∥+∥y2∥\|\cdot\|_{{\boldsymbol{\mathcal{Y}}}}\colon\boldsymbol{y}=(y_{1},y_{2})\mapsto\|y_{1}\|+\|y_{2}\|. Henceforth, the norm of M∈B (X,Y)\boldsymbol{M}\in\mathcal{B}\,({\boldsymbol{\mathcal{X}}},{\boldsymbol{\mathcal{Y}}}) is denoted by ∥M∥X,Y\|\boldsymbol{M}\|_{{\boldsymbol{\mathcal{X}}},{\boldsymbol{\mathcal{Y}}}}.

Let m⩾0m\geqslant 0 be an integer, let (Hi)0⩽i⩽m({\mathcal{H}}_{i})_{0\leqslant i\leqslant m} be nonzero real Hilbert spaces, set Hm+1=H0{\mathcal{H}}_{m+1}={\mathcal{H}}_{0}, and use Notation 3.5. For every i∈{1,…,m+1}i\in\{1,\ldots,m+1\}, let Wi∈B (Hi−1,Hi)W_{i}\in\mathcal{B}\,({\mathcal{H}}_{i-1},{\mathcal{H}}_{i}). Further, let α∈[1/2,1]\alpha\in[1/2,1], let θ0=1\theta_{0}=1, let (θi)1⩽i⩽m+1(\theta_{i})_{1\leqslant i\leqslant m+1} be as in (3.4), and set

There exists i∈{1,…,m+1}i\in\{1,\ldots,m+1\} such that Wi=0W_{i}=0.

∥M∥X,Y⩽1\|\boldsymbol{M}\|_{{\boldsymbol{\mathcal{X}}},{\boldsymbol{\mathcal{Y}}}}\leqslant 1.

∥W−2m+1(1−α)Id⁡∥−∥W∥+2θm+1⩽2m+1α\|W-2^{m+1}(1-\alpha)\operatorname{Id}\|-\|W\|+2\theta_{m+1}\leqslant 2^{m+1}\alpha.

α≠1\alpha\neq 1, for every i∈{1,…,m+1}i\in\{1,\ldots,m+1\} Wi≠0W_{i}\neq 0, and there exists η∈[0,α/((1−α)θm+1)]\eta\in[0,\alpha/((1-\alpha)\theta_{m+1})] such that

Then (Wi)1⩽i⩽m+1(W_{i})_{1\leqslant i\leqslant m+1} satisfies Condition 3.1.

Proof. We use the operators (Li)1⩽i⩽m+1(L_{i})_{1\leqslant i\leqslant m+1} introduced in Condition 3.1. Per Notation 3.5 and (3.9d),

Now let x∈X\boldsymbol{x}\in{\boldsymbol{\mathcal{X}}} be such that

(i): We assume that m⩾1m\geqslant 1. For every k∈{i,…,m}k\in\{i,\ldots,m\}, it follows from (3.4) that θk=0\theta_{k}=0 and in turn from Lemma 3.3(ii) and (3.13) that xk=0x_{k}=0. Therefore,

(ii): In view of (i), we assume that, if m⩾1m\geqslant 1, (∀i∈{1,…,m})(\forall i\in\{1,\ldots,m\}) Wi≠0W_{i}\neq 0. We then derive from (3.4) that (∀i∈{1,…,m})(\forall i\in\{1,\ldots,m\}) θi⩾∏k=1i∥Wk∥>0\theta_{i}\geqslant\prod_{k=1}^{i}\|W_{k}\|>0. If x0=0x_{0}=0, (3.3) trivially follows from Lemma 3.3(ii), we therefore assume otherwise. Now set

According to Lemma 3.3(ii), (∀i∈{0,…,m})(\forall i\in\{0,\ldots,m\}) ∥yi∥⩽1\|y_{i}\|\leqslant 1. On the other hand, it follows from (3.9c), (3.15), and (3.1) that My=Lm+1x/∥x0∥M\boldsymbol{y}=L_{m+1}\boldsymbol{x}/\|x_{0}\|. Altogether, we deduce from (3.12) that (3.3) holds.

(iii)⇒\Rightarrow(ii): Take y∈X\boldsymbol{y}\in{\boldsymbol{\mathcal{X}}} such that ∥y∥X⩽1\|\boldsymbol{y}\|_{{\boldsymbol{\mathcal{X}}}}\leqslant 1. Then it follows from (3.9c) and Lemma 3.3(i) that

In turn, (3.11) yields ∥M∥X,Y⩽1\|\boldsymbol{M}\|_{{\boldsymbol{\mathcal{X}}},{\boldsymbol{\mathcal{Y}}}}\leqslant 1.

(iv)⇒\Rightarrow(ii): Let y=(y0,…,ym)∈X\boldsymbol{y}=(y_{0},\ldots,y_{m})\in{\boldsymbol{\mathcal{X}}} be such that ∥y0∥=⋯=∥ym∥=1\|y_{0}\|=\cdots=\|y_{m}\|=1, and set

Therefore, since (3.21) implies that α−η(1−α)∥u∥⩾0\alpha-\eta(1-\alpha)\|u\|\geqslant 0, it results from (3) that

However, since (3.20) implies that α−η(1−α)∥W∥⩾0\alpha-\eta(1-\alpha)\|W\|\geqslant 0, while (3.17) implies that ∥u∥⩽θm+1−∥W∥\|u\|\leqslant\theta_{m+1}-\|W\|, we derive from (3) that

Hence, using (3), (3.27), (3.9c), (3.9a), and (3.9d) we obtain

Now set \boldsymbol{C}=\big{\{}{\boldsymbol{y}\in{\boldsymbol{\mathcal{X}}}}~{}\big{|}~{}{\|y_{0}\|=\cdots=\|y_{m}\|=1}\big{\}}. Then, in view of (3.11), (3), and [8, Proposition 11.1(ii)], we conclude that ∥M∥X,Y=sup⁡y∈conv C∥My∥Y=sup⁡y∈C∥My∥Y⩽1\|\boldsymbol{M}\|_{{\boldsymbol{\mathcal{X}}},{\boldsymbol{\mathcal{Y}}}}=\sup_{\boldsymbol{y}\in\text{\rm conv}\,\boldsymbol{C}}\|\boldsymbol{M}\boldsymbol{y}\|_{{\boldsymbol{\mathcal{Y}}}}=\sup_{\boldsymbol{y}\in\boldsymbol{C}}\|\boldsymbol{M}\boldsymbol{y}\|_{{\boldsymbol{\mathcal{Y}}}}\leqslant 1.

The next result establishes a link between deep neural network structures and the operators introduced in (3.1).

Let m⩾1m\geqslant 1 be an integer and let (Hi)0⩽i⩽m+1({\mathcal{H}}_{i})_{0\leqslant i\leqslant m+1} be nonzero real Hilbert spaces. For every i∈{1,…,m+1}i\in\{1,\ldots,m+1\}, let Wi∈B (Hi−1,Hi)W_{i}\in\mathcal{B}\,({\mathcal{H}}_{i-1},{\mathcal{H}}_{i}) and let LiL_{i} be as in (3.1). Further, for every i∈{1,…,m}i\in\{1,\ldots,m\}, let Pi ⁣:Hi→HiP_{i}\colon{\mathcal{H}}_{i}\to{\mathcal{H}}_{i} be firmly nonexpansive. Set

let xx and yy be distinct points in H0{\mathcal{H}}_{0}, and set v0=(x−y)/∥x−y∥v_{0}=(x-y)/\|x-y\|. Then there exists (v1,…,vm)∈H1×⋯×Hm(v_{1},\ldots,v_{m})\in{\mathcal{H}}_{1}\times\cdots\times{\mathcal{H}}_{m} such that

Proof. For every i∈{1,…,m}i\in\{1,\ldots,m\}, since PiP_{i} is firmly nonexpansive, there exists a nonexpansive operator Qi ⁣:Hi→HiQ_{i}\colon{\mathcal{H}}_{i}\to{\mathcal{H}}_{i} such that

We proceed by induction on mm. Suppose that m=1m=1 and set

which implies that ∥v1∥⩽∥W1(x−y)∥/∥x−y∥=∥L1v0∥\|v_{1}\|\leqslant{\|W_{1}(x-y)\|}/{\|x-y\|}=\|L_{1}v_{0}\|. Then

Thus, (3.30) holds for m=1m=1. Next, we assume that m>1m>1 and that there exists (v1,…,vm−1)∈H1×⋯×Hm−1(v_{1},\ldots,v_{m-1})\in{\mathcal{H}}_{1}\times\cdots\times{\mathcal{H}}_{m-1} such that

In addition, it follows from (3.34) and (3.35) that

We now establish connections between Condition 3.1 for linear operators and the concept of averagedness for composite nonlinear operators.

Let m⩾1m\geqslant 1 be an integer, let (Hi)0⩽i⩽m−1({\mathcal{H}}_{i})_{0\leqslant i\leqslant m-1} be nonzero real Hilbert spaces, set Hm=H0{\mathcal{H}}_{m}={\mathcal{H}}_{0}, and let α∈[1/2,1]\alpha\in[1/2,1]. For every i∈{1,…,m}i\in\{1,\ldots,m\}, let Wi∈B (Hi−1,Hi)W_{i}\in\mathcal{B}\,({\mathcal{H}}_{i-1},{\mathcal{H}}_{i}) and let Pi ⁣:Hi→HiP_{i}\colon{\mathcal{H}}_{i}\to{\mathcal{H}}_{i} be firmly nonexpansive. Suppose that (Wi)1⩽i⩽m(W_{i})_{1\leqslant i\leqslant m} satisfies Condition 3.1. Then Pm∘Wm∘⋯∘P1∘W1P_{m}\circ W_{m}\circ\cdots\circ P_{1}\circ W_{1} is α\alpha-averaged.

Proof. Set T=Pm∘Wm∘⋯∘P1∘W1T=P_{m}\circ W_{m}\circ\cdots\circ P_{1}\circ W_{1}. We must show that

is nonexpansive. By assumption, for every i∈{1,…,m}i\in\{1,\ldots,m\}, there exists a nonexpansive operator Qi ⁣:Hi→HiQ_{i}\colon{\mathcal{H}}_{i}\to{\mathcal{H}}_{i} such that (3.31) holds. Let (Li)1⩽i⩽m(L_{i})_{1\leqslant i\leqslant m} be as in (3.1) and let xx and yy be distinct points in H0{\mathcal{H}}_{0}. According to Lemma 3.7, there exists v=(v0,…,vm−1)∈H0×⋯×Hm−1\boldsymbol{v}=(v_{0},\ldots,v_{m-1})\in{\mathcal{H}}_{0}\times\cdots\times{\mathcal{H}}_{m-1} such that

In turn, we derive from (3.38) and (3.31) that

which establishes the nonexpansiveness of QQ.

Consider Theorem 3.8 with m=2m=2. In view of Proposition 3.6(iii), P2∘W2∘P1∘W1P_{2}\circ W_{2}\circ P_{1}\circ W_{1} is α\alpha-averaged if ∥W2∘W1−4(1−α)Id⁡∥+∥W2∘W1∥+2∥W2∥ ∥W1∥⩽4α\|W_{2}\circ W_{1}-4(1-\alpha)\operatorname{Id}\|+\|W_{2}\circ W_{1}\|+2\|W_{2}\|\,\|W_{1}\|\leqslant 4\alpha. In particular, if α=1\alpha=1, this condition is obviously less restrictive than requiring that W1W_{1} and W2W_{2} be nonexpansive.

A variational inequality model

In this section, we first investigate an autonomous version of Model 1.1.

where x=(x1,…,xm)\boldsymbol{x}=(x_{1},\ldots,x_{m}) denotes a generic element in H{\boldsymbol{\mathcal{H}}}.

We start with a property of the compositions of the operators (Ti)1⩽i⩽m(T_{i})_{1\leqslant i\leqslant m} of (4.1).

Consider the setting of Model 4.1, let ii and jj be integers such that 1⩽j⩽i⩽m1\leqslant j\leqslant i\leqslant m, and let x∈Hj−1x\in{\mathcal{H}}_{j-1}. Then

Proof. In view of (4.1), the property is satisfied when i=ji=j. We now assume that i>ji>j. Since Ri∈A(Hi)R_{i}\in\mathcal{A}({\mathcal{H}}_{i}), Proposition 2.21(i) yields

Next, we establish a connection between Model 4.1 and a variational inequality.

In the setting of Model 4.1, consider the variational inequality problem

The set of solutions to (4.5) is F\boldsymbol{F}.

F=zer (Id⁡ −W∘S+∂ψ)=Fix (proxψ∘W∘S)\boldsymbol{F}=\text{\rm zer}\,({\boldsymbol{\operatorname{Id}}\,}-\boldsymbol{W}\circ\boldsymbol{S}+\partial\boldsymbol{\psi})=\text{\rm Fix}\,(\text{\rm prox}_{\boldsymbol{\psi}}\circ\boldsymbol{W}\circ\boldsymbol{S}).

\boldsymbol{F}=\big{\{}{(T_{1}\overline{x}_{m},(T_{2}\circ T_{1})\overline{x}_{m},\ldots,(T_{m-1}\circ\cdots\circ T_{1})\overline{x}_{m},\overline{x}_{m})}~{}\big{|}~{}{\overline{x}_{m}\in F}\big{\}}.

Suppose that (Wi)1⩽i⩽m(W_{i})_{1\leqslant i\leqslant m} satisfies Condition 3.1 for some α∈[1/2,1]\alpha\in[1/2,1]. Then FF is closed and convex.

Suppose that (Wi)1⩽i⩽m(W_{i})_{1\leqslant i\leqslant m} satisfies Condition 3.1 for some α∈[1/2,1]\alpha\in[1/2,1] and that one of the following holds:

ran (Tm∘⋯∘T1)\text{\rm ran}\,(T_{m}\circ\cdots\circ T_{1}) is bounded.

There exists j∈{1,…,m}j\in\{1,\ldots,m\} such that dom φj\text{\rm dom}\,\varphi_{j} is bounded.

Then FF and F\boldsymbol{F} are nonempty.

Suppose that Id⁡ −W∘S{\boldsymbol{\operatorname{Id}}\,}-\boldsymbol{W}\circ\boldsymbol{S} is monotone. Then F\boldsymbol{F} is closed and convex. In addition, FF and F\boldsymbol{F} are nonempty if any of the following holds:

Id⁡ −W∘S+∂φ{\boldsymbol{\operatorname{Id}}\,}-\boldsymbol{W}\circ\boldsymbol{S}+\partial\boldsymbol{\varphi} is surjective.

∂φ−W∘S\partial\boldsymbol{\varphi}-\boldsymbol{W}\circ\boldsymbol{S} is maximally monotone.

max⁡1⩽i⩽m∥Wi∥⩽1\max_{1\leqslant i\leqslant m}\|W_{i}\|\leqslant 1, S∗−W\boldsymbol{S}^{*}-\boldsymbol{W} has closed range, and ker⁡(S−W∗)={0}\ker(\boldsymbol{S}-\boldsymbol{W}^{*})=\{\boldsymbol{0}\}.

max⁡1⩽i⩽m∥Wi∥⩽1\max_{1\leqslant i\leqslant m}\|W_{i}\|\leqslant 1 and, for every i∈{1,…,m}i\in\{1,\ldots,m\}, dom φi∗=Hi\text{\rm dom}\,\varphi_{i}^{*}={\mathcal{H}}_{i}.

For every i∈{1,…,m}i\in\{1,\ldots,m\}, dom φi=H\text{\rm dom}\,\varphi_{i}={\mathcal{H}} and dom φi∗=Hi\text{\rm dom}\,\varphi_{i}^{*}={\mathcal{H}}_{i}.

S∗−W\boldsymbol{S}^{*}-\boldsymbol{W} has closed range, ker⁡(S−W∗)={0}\ker(\boldsymbol{S}-\boldsymbol{W}^{*})=\{\boldsymbol{0}\}, and, for every i∈{1,…,m}i\in\{1,\ldots,m\}, dom φi=Hi\text{\rm dom}\,\varphi_{i}={\mathcal{H}}_{i}.

For every i∈{1,…,m}i\in\{1,\ldots,m\}, dom φi\text{\rm dom}\,\varphi_{i} is bounded.

Proof. We first observe that S∈B (H,H→)\boldsymbol{S}\in\mathcal{B}\,({\boldsymbol{\mathcal{H}}},\overset{\rightarrow}{{\boldsymbol{\mathcal{H}}}}), W∈B (H→,H)\boldsymbol{W}\in\mathcal{B}\,(\overset{\rightarrow}{{\boldsymbol{\mathcal{H}}}},{\boldsymbol{\mathcal{H}}}), φ∈Γ0(H)\boldsymbol{\varphi}\in\Gamma_{0}({\boldsymbol{\mathcal{H}}}), and ψ∈Γ0(H)\boldsymbol{\psi}\in\Gamma_{0}({\boldsymbol{\mathcal{H}}}).

(i): Let x∈H\boldsymbol{x}\in{\boldsymbol{\mathcal{H}}}. Then

(ii): Let x∈H\boldsymbol{x}\in{\boldsymbol{\mathcal{H}}}. Using (4.2), we obtain

(iii): Clear from the definitions of FF and F\boldsymbol{F}.

(iv): Define mm firmly nonexpansive operators by (∀i∈{1,…,m})(\forall i\in\{1,\ldots,m\}) Pi ⁣:Hi→Hi ⁣:y↦Ri(y+bi)P_{i}\colon{\mathcal{H}}_{i}\to{\mathcal{H}}_{i}\colon y\mapsto R_{i}(y+b_{i}). Then it follows from (4.1) and Theorem 3.8 applied to (Pi)1⩽i⩽m(P_{i})_{1\leqslant i\leqslant m} that Tm∘⋯∘T1T_{m}\circ\cdots\circ T_{1} is nonexpansive. In turn, we derive from [8, Corollary 4.24] that its fixed point set FF is closed and convex.

(v): Thanks to (iii), it is enough to show that F≠∅F\neq{\varnothing}. Set T=Tm∘⋯∘T1T=T_{m}\circ\cdots\circ T_{1} and recall that it is nonexpansive by virtue of Theorem 3.8.

(v)(a): Let CC be a closed ball such that ran T⊂C\text{\rm ran}\,T\subset C and set S=T∣CS=T|_{C}. Then S ⁣:C→CS\colon C\to C is nonexpansive and therefore [8, Proposition 4.29] asserts that Fix T=Fix S≠∅\text{\rm Fix}\,T=\text{\rm Fix}\,S\neq{\varnothing}.

(v)(b)⇒\Rightarrow(v)(a): We have ran Tj⊂ran Rj=ran proxφj=dom (Id⁡+∂φj)=dom ∂φj⊂dom φj\text{\rm ran}\,T_{j}\subset\text{\rm ran}\,R_{j}=\text{\rm ran}\,\text{\rm prox}_{\varphi_{j}}=\text{\rm dom}\,(\operatorname{Id}+\partial\varphi_{j})=\text{\rm dom}\,\partial\varphi_{j}\subset\text{\rm dom}\,\varphi_{j}. Hence ran Tj\text{\rm ran}\,T_{j} is bounded and Proposition 4.2 (with i=mi=m) implies that

(vi): Set A=Id⁡ −W∘S+∂ψ\boldsymbol{A}={\boldsymbol{\operatorname{Id}}\,}-\boldsymbol{W}\circ\boldsymbol{S}+\partial\boldsymbol{\psi}. Since Id⁡ −W∘S{\boldsymbol{\operatorname{Id}}\,}-\boldsymbol{W}\circ\boldsymbol{S} is monotone and continuous, it is maximally monotone [8, Corollary 20.28], with H{\boldsymbol{\mathcal{H}}} as its domain. Since ∂ψ\partial\boldsymbol{\psi} is also maximally monotone [8, Theorem 20.25], A\boldsymbol{A} is likewise [8, Corollary 25.5(i)] and hence F=zer A\boldsymbol{F}=\text{\rm zer}\,\boldsymbol{A} is closed and convex [8, Proposition 23.39]. Next, we note that, in view of (iii), F≠∅F\neq{\varnothing} ⇔\Leftrightarrow F≠∅\boldsymbol{F}\neq{\varnothing}.

(vi)(a): The hypothesis implies that (bi)1⩽i⩽m∈ran (Id⁡ −W∘S+∂φ)(b_{i})_{1\leqslant i\leqslant m}\in\text{\rm ran}\,({\boldsymbol{\operatorname{Id}}\,}-\boldsymbol{W}\circ\boldsymbol{S}+\partial\boldsymbol{\varphi}) and therefore that (4.5) has a solution, i.e., F≠∅\boldsymbol{F}\neq{\varnothing}.

(vi)(b)⇒\Rightarrow(vi)(a): The claim follows from Minty’s theorem [8, Theorem 21.1].

(vi)(c)⇒\Rightarrow(vi)(a): We have ∥W∘S∥=∥W∥=max⁡1⩽i⩽m∥Wi∥⩽1\|\boldsymbol{W}\circ\boldsymbol{S}\|=\|\boldsymbol{W}\|=\max_{1\leqslant i\leqslant m}\|W_{i}\|\leqslant 1. Therefore, −W∘S-\boldsymbol{W}\circ\boldsymbol{S} is nonexpansive, which implies that (Id⁡ −W∘S)/2({\boldsymbol{\operatorname{Id}}\,}-\boldsymbol{W}\circ\boldsymbol{S})/2 is firmly nonexpansive [8, Corollary 4.5], that is (∀x∈H)(\forall\boldsymbol{x}\in{\boldsymbol{\mathcal{H}}}) ⟨x−W(Sx)∣x⟩⩾∥x−W(Sx)∥2/2{\left\langle{{\boldsymbol{x}-\boldsymbol{W}(\boldsymbol{S}\boldsymbol{x})}\mid{\boldsymbol{x}}}\right\rangle}\geqslant\|\boldsymbol{x}-\boldsymbol{W}(\boldsymbol{S}\boldsymbol{x})\|^{2}/2. Consequently, Id⁡ −W∘S{\boldsymbol{\operatorname{Id}}\,}-\boldsymbol{W}\circ\boldsymbol{S} is 3∗3^{*} monotone [8, Proposition 25.16], while ∂φ\partial\boldsymbol{\varphi} is also 3∗3^{*} monotone [8, Example 25.13]. Finally, since S\boldsymbol{S} is unitary,

which shows that Id⁡ −W∘S{\boldsymbol{\operatorname{Id}}\,}-\boldsymbol{W}\circ\boldsymbol{S} is surjective. Altogether, since [8, Corollary 25.5(i)] implies that Id⁡ −W∘S+∂φ{\boldsymbol{\operatorname{Id}}\,}-\boldsymbol{W}\circ\boldsymbol{S}+\partial\boldsymbol{\varphi} is maximally monotone, it follows from [8, Corollary 25.27(i)] that Id⁡ −W∘S+∂φ{\boldsymbol{\operatorname{Id}}\,}-\boldsymbol{W}\circ\boldsymbol{S}+\partial\boldsymbol{\varphi} is surjective.

(vi)(d)⇒\Rightarrow(vi)(a): We have dom φ∗=H\text{\rm dom}\,\boldsymbol{\varphi}^{*}={\boldsymbol{\mathcal{H}}}. Hence since int dom φ∗⊂dom ∂φ∗\text{int\,dom}\,\boldsymbol{\varphi}^{*}\subset\text{\rm dom}\,\partial\boldsymbol{\varphi}^{*} [8, Proposition 16.27], we have ran ∂φ=dom (∂φ)−1=dom ∂φ∗=H\text{\rm ran}\,\partial\boldsymbol{\varphi}=\text{\rm dom}\,(\partial\boldsymbol{\varphi})^{-1}=\text{\rm dom}\,\partial\boldsymbol{\varphi}^{*}={\boldsymbol{\mathcal{H}}}. Hence, ∂φ\partial\boldsymbol{\varphi} is surjective. We conclude using the same arguments as in (vi)(c): ∂φ\partial\boldsymbol{\varphi} and Id⁡ −W∘S{\boldsymbol{\operatorname{Id}}\,}-\boldsymbol{W}\circ\boldsymbol{S} are both 3∗3^{*} monotone and their sum is maximally monotone, which allows us to invoke [8, Corollary 25.27(i)].

(vi)(e)⇒\Rightarrow(vi)(a): As seen in (vi)(d), ∂φ\partial\boldsymbol{\varphi} is surjective. We have H=int dom φ⊂dom ∂φ{\boldsymbol{\mathcal{H}}}=\text{int\,dom}\,\boldsymbol{\varphi}\subset\text{\rm dom}\,\partial\boldsymbol{\varphi} [8, Proposition 16.27]. Consequently, H=dom (Id⁡ −W∘S)⊂dom ∂φ{\boldsymbol{\mathcal{H}}}=\text{\rm dom}\,({\boldsymbol{\operatorname{Id}}\,}-\boldsymbol{W}\circ\boldsymbol{S})\subset\text{\rm dom}\,\partial\boldsymbol{\varphi}. Altogether, since ∂φ\partial\boldsymbol{\varphi} is 3∗3^{*} monotone, it follows from [8, Corollary 25.27(ii)] that Id⁡ −W∘S+∂φ{\boldsymbol{\operatorname{Id}}\,}-\boldsymbol{W}\circ\boldsymbol{S}+\partial\boldsymbol{\varphi} is surjective.

(vi)(f)⇒\Rightarrow(vi)(a): As seen in (vi)(c), Id⁡ −W∘S{\boldsymbol{\operatorname{Id}}\,}-\boldsymbol{W}\circ\boldsymbol{S} is surjective and ∂φ\partial\boldsymbol{\varphi} is 3∗3^{*} monotone. In addition, dom (Id⁡ −W∘S)⊂dom ∂φ\text{\rm dom}\,({\boldsymbol{\operatorname{Id}}\,}-\boldsymbol{W}\circ\boldsymbol{S})\subset\text{\rm dom}\,\partial\boldsymbol{\varphi} since H=int dom φ⊂dom ∂φ{\boldsymbol{\mathcal{H}}}=\text{int\,dom}\,\boldsymbol{\varphi}\subset\text{\rm dom}\,\partial\boldsymbol{\varphi} [8, Proposition 16.27]. Altogether, it follows from [8, Corollary 25.27(ii)] that Id⁡ −W∘S+∂φ{\boldsymbol{\operatorname{Id}}\,}-\boldsymbol{W}\circ\boldsymbol{S}+\partial\boldsymbol{\varphi} is surjective.

(vi)(g): Here \text{\rm dom}\,\boldsymbol{A}=\text{\rm dom}\,\partial\boldsymbol{\varphi}\subset\text{\rm dom}\,\boldsymbol{\varphi}=\raisebox{-1.42262pt}{\mbox{\LARGE{\times}}}_{\!i=1}^{\!m}\text{\rm dom}\,\varphi_{i} is bounded. Hence, F=zer A≠∅\boldsymbol{F}=\text{\rm zer}\,\boldsymbol{A}\neq{\varnothing} [8, Proposition 23.36(iii)].

In Proposition 4.3(vi), it is required that Id⁡ −W∘S{\boldsymbol{\operatorname{Id}}\,}-\boldsymbol{W}\circ\boldsymbol{S} be monotone, or equivalently, that its self-adjoint part Id⁡ −(W∘S+S∗∘W∗)/2{\boldsymbol{\operatorname{Id}}\,}-(\boldsymbol{W}\circ\boldsymbol{S}+\boldsymbol{S}^{*}\circ\boldsymbol{W}^{*})/2 be positive. In a finite-dimensional setting, this just means that the eigenvalues of the matrix WS+S∗W∗\boldsymbol{W}\boldsymbol{S}+\boldsymbol{S}^{*}\boldsymbol{W}^{*} are in ]−∞,2]\left]{-\infty},2\right].

where, given x∈Hi−1x\in{\mathcal{H}}_{i-1}, [Wix]k[W_{i}x]_{k} is the kkth component of WixW_{i}x and

Altogether, we conclude that F\boldsymbol{F} is a closed convex polyhedron.

2 Asymptotic analysis

Next, we investigate the asymptotic behavior of (1.2) in the context of Model 4.1.

In the setting of Model 4.1, set T=Tm∘⋯∘T1T=T_{m}\circ\cdots\circ T_{1}, let α∈[1/2,1]\alpha\in[1/2,1], and suppose that the following hold:

(Wi)1⩽i⩽m(W_{i})_{1\leqslant i\leqslant m} satisfies Condition 3.1 with parameter α\alpha.

λn≡1/α=1\lambda_{n}\equiv 1/\alpha=1 and Txn−xn→0Tx_{n}-x_{n}\to 0.

For every i∈{1,…,m−1}i\in\{1,\ldots,m-1\}, RiR_{i} is weakly sequentially continuous.

For every i∈{1,…,m−1}i\in\{1,\ldots,m-1\}, RiR_{i} is a separable activation operator in the sense of Proposition 2.24.

For every i∈{1,…,m−1}i\in\{1,\ldots,m-1\}, Hi{\mathcal{H}}_{i} is finite-dimensional.

Proof. We first derive from (1.2) and Model 4.1 that

Now set (∀i∈{1,…,m})(\forall i\in\{1,\ldots,m\}) Pi ⁣:Hi→Hi ⁣:y↦Ri(y+bi)P_{i}\colon{\mathcal{H}}_{i}\to{\mathcal{H}}_{i}\colon y\mapsto R_{i}(y+b_{i}). Then (4.1) yields T=Pm∘Wm∘⋯∘P1∘W1T=P_{m}\circ W_{m}\circ\cdots\circ P_{1}\circ W_{1} and, since the operators (Ri)1⩽i⩽m(R_{i})_{1\leqslant i\leqslant m} are firmly nonexpansive, the operators (Pi)1⩽i⩽m(P_{i})_{1\leqslant i\leqslant m} are likewise. Hence, it follows from (b), Theorem 3.8, and (4.2) that

We now prove the convergence of the individual sequences under each assumption.

(iii): We have already established that xn ⇀ x‾mx_{n}\>\rightharpoonup\>\overline{x}_{m}. Since W1W_{1} is weakly continuous as a bounded linear operator, so is T1T_{1} in (4.1). Hence, (1.2) implies that x1,n=T1xn ⇀ T1x‾m=x‾1x_{1,n}=T_{1}x_{n}\>\rightharpoonup\>T_{1}\overline{x}_{m}=\overline{x}_{1}. Likewise, we obtain successively x2,n=T2x1,n ⇀ T2x‾1=x‾2x_{2,n}=T_{2}x_{1,n}\>\rightharpoonup\>T_{2}\overline{x}_{1}=\overline{x}_{2}, x3,n=T3x2,n ⇀ T3x‾2=x‾3x_{3,n}=T_{3}x_{2,n}\>\rightharpoonup\>T_{3}\overline{x}_{2}=\overline{x}_{3},…, xm,n=Tmxm−1,n ⇀ Tmx‾m−1=x‾mx_{m,n}=T_{m}x_{m-1,n}\>\rightharpoonup\>T_{m}\overline{x}_{m-1}=\overline{x}_{m}.

(iv)⇒\Rightarrow(iii): See [8, Proposition 24.12(iii)].

(v)⇒\Rightarrow(iii): A proximity operator is nonexpansive and therefore continuous, hence weakly continuous in a finite-dimensional setting.

(vi): As shown above, xn ⇀ x‾m∈Fx_{n}\>\rightharpoonup\>\overline{x}_{m}\in F. It follows from Proposition 3.6(iii) and Theorem 3.8 (applied with m=1m=1) that, for every i∈{1,…,m}i\in\{1,\ldots,m\}, TiT_{i} is βi\beta_{i}-averaged. Hence, upon applying [24, Theorem 3.5(ii)] with α\alpha as an averaging constant of TT, we infer that

Thus, x1,n−xn=T1xn−xn→T1x‾m−x‾mx_{1,n}-x_{n}=T_{1}x_{n}-x_{n}\to T_{1}\overline{x}_{m}-\overline{x}_{m}, which implies that x1,n=(x1,n−xn)+xn ⇀ (T1x‾m−x‾m)+x‾m=T1x‾mx_{1,n}=(x_{1,n}-x_{n})+x_{n}\>\rightharpoonup\>(T_{1}\overline{x}_{m}-\overline{x}_{m})+\overline{x}_{m}=T_{1}\overline{x}_{m}. However, since x2,n−x1,n=(T2∘T1)xn−T1xn→(T2∘T1)x‾m−T1x‾mx_{2,n}-x_{1,n}=(T_{2}\circ T_{1})x_{n}-T_{1}x_{n}\to(T_{2}\circ T_{1})\overline{x}_{m}-T_{1}\overline{x}_{m}, we obtain x2,n ⇀ (T2∘T1)x‾mx_{2,n}\>\rightharpoonup\>(T_{2}\circ T_{1})\overline{x}_{m}. Continuing this telescoping process yields the claim.

The next result covers the case when the variational inequality problem (4.5) has no solution.

To model closely existing deep neural networks, we have chosen the activation operators in Definition 2.20 and Model 4.1 to be proximity operators. However, as is clear from the results of Section 3 and in particular the central Theorem 3.8, an activation operator Ri ⁣:Hi→HiR_{i}\colon{\mathcal{H}}_{i}\to{\mathcal{H}}_{i} could more generally be a firmly nonexpansive operator that admits as a fixed point. By [8, Corollary 23.9], this means that RiR_{i} is the resolvent of some maximally monotone operator such Ai ⁣:Hi→2HiA_{i}\colon{\mathcal{H}}_{i}\to 2^{{\mathcal{H}}_{i}} (i.e., Ri=(Id⁡+Ai)−1R_{i}=(\operatorname{Id}+A_{i})^{-1}) such that 0∈Ai00\in A_{i}0. In this context, the variational inequality (4.5) assumes the more general form of a system of monotone inclusions, namely,

Analysis of nonperiodic networks

We analyze the deep neural network described in Model 1.1 in the following scenario.

In the setting of Model 1.1, suppose that Assumption 5.1 is satisfied, let i∈{1,…,m}i\in\{1,\ldots,m\}, and set

We can now present the main result of this section on the asymptotic behavior of Model 1.1. The proof of this result relies on Theorem 4.7, which it extends.

Consider the setting of Model 1.1 and let α∈[1/2,1]\alpha\in[1/2,1]. Suppose that Assumption 5.1 is satisfied as well as the following:

F=Fix T≠∅F=\text{\rm Fix}\,T\neq{\varnothing}, where T=Tm∘⋯∘T1T=T_{m}\circ\cdots\circ T_{1}.

(Wi)1⩽i⩽m(W_{i})_{1\leqslant i\leqslant m} satisfies Condition 3.1 with parameter α\alpha.

λn≡α=1\lambda_{n}\equiv\alpha=1 and Txn−xn→0Tx_{n}-x_{n}\to 0.

For every i∈{1,…,m−1}i\in\{1,\ldots,m-1\}, RiR_{i} is weakly sequentially continuous.

For every i∈{1,…,m−1}i\in\{1,\ldots,m-1\}, RiR_{i} is a separable activation function in the sense of Proposition 2.24.

For every i∈{1,…,m−1}i\in\{1,\ldots,m-1\}, Hi{\mathcal{H}}_{i} is finite-dimensional.

In (c)(i)–(c)(ii) above, Proposition 4.3(iii) ensures that (T1x‾m,(T2∘T1)x‾m,…,(Tm−1∘⋯∘T1)x‾m,x‾m)(T_{1}\overline{x}_{m},(T_{2}\circ T_{1})\overline{x}_{m},\ldots,(T_{m-1}\circ\cdots\circ T_{1})\overline{x}_{m},\overline{x}_{m}) solves (4.5).

(vi): For every i∈{1,…,m}i\in\{1,\ldots,m\}, set

Thus, since Proposition 3.6(iii) and Theorem 3.8 imply that the operators (Ti)1⩽i⩽m(T_{i})_{1\leqslant i\leqslant m} are averaged, the proof can be completed as that of Theorem 4.7(vi) since [24, Theorem 3.5(ii)] asserts that (4.15) remains valid under (5.21).

References