Convergence Analysis of Alternating Direction Method of Multipliers for a Family of Nonconvex Problems
Mingyi Hong, Zhi-Quan Luo, Meisam Razaviyayn
Introduction
Consider the following linearly constrained (possibly nonsmooth or/and nonconvex) problem with blocks of variables :
where is a constant representing the primal penalty parameter.
To solve problem (1), let us consider a popular algorithm called the alternating direction method of multipliers (ADMM), whose steps are given below:
Algorithm 0. ADMM for Problem (1) At each iteration , update the primal variables: (3) Update the dual variable: (4)
The ADMM algorithm was originally introduced in early 1970s , and has since been studied extensively . Recently it has become widely popular in modern big data related problems arising in machine learning, computer vision, signal processing, networking and so on; see and the references therein. In practice, the algorithm often exhibits faster convergence than traditional primal-dual type algorithms such as the dual ascent algorithm or the method of multipliers . It is also particularly suitable for parallel implementation .
Unlike the convex case, for which the behavior of ADMM has been investigated quite extensively, when the objective becomes nonconvex, the convergence issue of ADMM remains largely open. Nevertheless, it has been observed by many researchers that the ADMM works extremely well for various applications involving nonconvex objectives, such as the nonnegative matrix factorization , phase retrieval , distributed matrix factorization , distributed clustering , sparse zero variance discriminant analysis , polynomial optimization, tensor decomposition , matrix separation , matrix completion , asset allocation , sparse feedback control and so on. However, to the best of our knowledge, existing convergence analysis of ADMM for nonconvex problems is very limited — all known global convergence analysis needs to impose uncheckable conditions on the sequence generated by the algorithm. For example, references show global convergence of the ADMM to the set of stationary solutions for their respective nonconvex problems, by making the key assumptions that the limit points do exist, and that the successive differences of the iterates (both primal and dual) converge to zero. However such assumption is nonstandard and overly restrictive. It is not clear whether the same convergence result can be claimed without making assumptions on the iterates. Reference analyzes a family of splitting algorithms (which includes the ADMM as a special case) for certain nonconvex quadratic optimization problem, and shows that they converge to the stationary solution when certain condition on the dual stepsize is met. We note that there has been many recent works proposing new algorithms to solve nonconvex and nonsmooth problems, for example . However, these works do not deal with nonconvex problems with linearly coupling constraints, and their analysis does not directly apply to the ADMM-type methods.
The Nonconvex Consensus Problem
Consider the following nonconvex global consensus problem with regularization
where ’s are a set of smooth, possibly nonconvex functions, while is a convex nonsmooth regularization term. This problem is related to the convex global consensus problem discussed heavily in [8, Section 7], but with the important difference that ’s can be nonconvex.
In many practical applications, ’s need to be handled by a single agent, such as a thread or a processor. This motivates the following consensus formulation. Let us introduce a set of new variables , and transform problem (5) equivalently to the following linearly constrained problem
We note that after reformulation, the problem dimension is increased by due to the introduction of auxiliary variables . Consequently, solving the reformulated problem (6) distributedly may not be as efficient (in terms of total number of iterations required) as applying the centralized algorithms directly to the original problem (5). Nonetheless, a major benefit of solving the reformulated problem (6) is the flexibility of allowing each distributed agent to handle a single local variable and a local function .
The augmented Lagrangian function is given by
Note that this augmented Lagrangian is slightly different from the one expressed in (2), as we have used a set of different penalization parameters , one for each equality constraint . We note that there can be many other variants of the basic consensus problem, such as the general form consensus optimization, the sharing problem and so on. We will discuss some of those variants in the later sections.
2 The ADMM Algorithm for Nonconvex Consensus
The problem (6) can be solved distributedly by applying the classical ADMM. The details are given in the table below.
Algorithm 1. The Classical ADMM for Problem (6) At each iteration , compute: (8) Each node computes by solving: (9) Each node updates the dual variable: (10)
In the update step, if the nonsmooth penalization does not appear in the objective, then this step can be written as
Note that the above algorithm has the exact form as the classical ADMM described in , where the variable is taken as the first block of primal variable, and the collection as the second block. The two primal blocks are updated in a sequential (i.e., Gauss-Seidel) manner, followed by an inexact dual ascent step.
In what follows, we consider a more general version of ADMM which includes Algorithm 1 as a special case. In particular, we propose a flexible ADMM algorithm in which there is a greater flexibility in choosing the order of the update of both the primal and the dual variables. Specifically, we consider the following two types of variable block update order rules: let be the indices for the primal variable blocks , and let {\mbox{\mathcal{C}}}^{t}\subseteq\{0,1,\cdots,K\} denote the set of variables updated in iteration , then
Randomized update rule: At each iteration , a variable block is chosen randomly with probability ,
Essentially cyclic update rule: There exists a given period during which each index is updated at least once. More specifically, at iteration , update all the variables in an index set {\mbox{\mathcal{C}}}^{t} whereby
We call this update rule a period- essentially cyclic update rule.
Algorithm 2. The Flexible ADMM for Problem (6) Let {\mbox{\mathcal{C}}}^{1}=\{0,\cdots,K\}, . At each iteration , do: If , pick an index set {\mbox{\mathcal{C}}}^{t+1}\subseteq\{0,\cdots,K\}. If 0\in{\mbox{\mathcal{C}}}^{t+1}, compute: (14) Else . If and k\in{\mbox{\mathcal{C}}}^{t+1}, node computes by solving: (15) Update the dual variable: (16) Else , .
We note that the randomized version of Algorithm 2 is similar to that of the convex consensus algorithms studied in and . It is also related to the randomized BSUM-M algorithm studied in . The difference with the latter is that in the randomized BSUM-M, the dual variable is viewed as an additional block that can be randomly picked (independent of the way that the primal blocks are picked), whereas in Algorithm 2, the dual variable is always updated whenever the corresponding primal variable is updated. To the best of our knowledge, the period- essentially cyclic update rule is a new variant of the ADMM.
Notice that Algorithm 1 is simply the period-1 essentially cyclic rule, which is a special case of Algorithm 2. Therefore we will focus on analyzing Algorithm 2. To this end, we make the following assumption.
There exists a positive constant such that
Moreover, is convex (possible nonsmooth); is a closed convex set.
For all , the penalty parameter is chosen large enough such that:
For all , the subproblem (15) is strongly convex with modulus ;
For all , and .
is bounded from below over , that is,
We have the following remarks regarding to the assumptions made above.
As inceases, the subproblem (15) will be eventually strongly convex with respect to . The corresponding strong convexity modulus is a monotonic increasing function of .
Whenever is nonconvex (therefore ), the condition implies .
By construction, is also strongly convex with respect to , with a modulus .
Assumption A makes no assumption on the iterates generated by the algorithm. This is in contrast to the existing analysis of the nonconvex ADMM algorithms .
Now we begin to analyze Algorithm 2. We first make several definitions. Let (resp. ) denote the latest iteration index that (resp. ) is updated before iteration , i.e.,
This definition implies that for all .
We first show that the size of the successive difference of the dual variables can be bounded above by that of the primal variables.
Suppose Assumption A holds. Then for Algorithm 2 with either randomized or essentially cyclic update rule, the following are true
Proof. We will show the first inequality. The second inequality follows a similar line of argument.
To prove (19a), first note that the case for k\notin{\mbox{\mathcal{C}}}^{t+1} is trivial, as both sides of (19a) evaluate to zero. Suppose k\in{\mbox{\mathcal{C}}}^{t+1}. From the update step (15) we have the following optimality condition
Combined with the dual variable update step (16) we obtain
Combining this with Assumption A1, and noting that for any given , and are always updated in the same iteration, we obtain for all k\in{\mbox{\mathcal{C}}}^{t+1}/\{0\}
Next, we use (19a) to bound the difference of the augmented Lagrangian.
For Algorithm 2 with either randomized or period-T essentially cyclic update rule, we have the following
Proof. We first split the successive difference of the augmented Lagrangian by
The first term in (2.2) can be bounded by
where in we have use (16), and the fact that for all variable block that has not been updated (i.e., k\neq 0,k\notin{\mbox{\mathcal{C}}}^{t+1}). The second term in (2.2) can be bounded by
where in we have used the fact that is strongly convex w.r.t. each and , with modulus and , respectively, and that
is some subgradient vector; in we have used the fact that when k\notin{\mbox{\mathcal{C}}}^{t+1} (resp. 0\notin{\mbox{\mathcal{C}}}^{t+1}), (resp. ), and we have defined \iota\{0\in{\mbox{\mathcal{C}}}^{t+1}\} as the indicator function that takes the value if 0\in{\mbox{\mathcal{C}}}^{t+1} is true, and takes value otherwise; in we have used the optimality of each subproblem (15) and (where is specialized to the subgradient vector that satisfies the optimality condition for problem (14)).
Combining the above two inequalities (24) and (25), we obtain
where the last inequality is due to (19a). The desired result is obtained by noticing the fact that when 0\notin{\mbox{\mathcal{C}}}^{t+1}, we have .
The above result implies that if the following condition is satisfied:
then the value of the augmented Lagrangian function will always decrease. Note that as long as , one can always find a large enough such that the above condition is satisfied, as the left hand side (lhs) of is monotonically increasing w.r.t. , while the right hand side (rhs) is a constant.
Next we show that is in fact convergent.
Suppose Assumption A is true. Let be generated by Algorithm 2 with either the essentially cyclic rule or the randomized rule. Then the following limit exists and is lower bounded by defined in Assumption A3:
Proof. Notice that the augmented Lagrangian function can be expressed as
where comes from the Lipschitz continuity of the gradient of ’s (Assumption A1), and the fact that for all (Assumption A2). To see why is true, we first observe that due to (21), we have for all and k\in{\mbox{\mathcal{C}}}^{t+1}
For all and k\notin{\mbox{\mathcal{C}}}^{t+1}, it follows from and that
Combining these two cases shows that is true.
Clearly, (2.2) and Assumption A3 together imply that is lower bounded. This combined with (2) says that whenever the penalty parameter ’s are chosen sufficiently large (as per Assumption A2), is monotonically decreasing and is convergent. This completes the proof.
We are now ready to prove our first main result, which asserts that the sequence of iterates generated by Algorithm 2 converges to the set of stationary solution of problem (6).
Assume that Assumption A is satisfied. Then we have the following
We have , deterministically for the essentially cyclic update rule and almost surely for the randomized update rule.
Let denote any limit point of the sequence generated by Algorithm 2. Then the following statement is true (deterministically for the essentially cyclic update rule and almost surely for the randomized update rule)
That is, any limit point of Algorithm 2 is a stationary solution of problem (6).
If is a compact set, then the sequence of iterates generated by Algorithm 2 converges to the set of stationary solutions of problem (6). That is,
where is the set of primal-dual stationary solutions of problem (6); denotes the distance between a vector and the set , i.e.,
Proof. We first show part (1) of the theorem. For the essentially cyclic update rule, Lemma 2 implies that
where the last equality follows from the fact if k\not\in{\mbox{\mathcal{C}}}^{t+i} and . Using the fact that each index in will be updated at least once during , as well as Lemma 3 and the bounds for ’s in Assumption A2, we have
By Lemma 1, we further obtain for all . In light of the dual update step of Algorithm 2, the fact that implies that .
For the randomized update rule, we can take the conditional expectation (over the choice of the blocks) on both sides of (2) and obtain
where in the last two inequalities, we have used the fact that ’s satisfy Assumption A2, hence for all ; the last inequality follows from the fact that for all . Note that by Lemma 3, for all , where is defined in Assumption A3. Then let us substract both sides of the above inequality by , and invoke the Supermartigale Convergence Theorem [57, Proposition 4.2]. We conclude that is convergent almost surely (a.s.), and that
By Lemma 1, we further obtain and for all . Finally, from the definition of , we see that a.s. implies that a.s. for all .
Next we show part (2) of the theorem. For simplicity, we consider only the essentially cyclic rule as the proof for the randomized rule is similar. We begin by examining the optimality condition for the and subproblems at iteration . Suppose k\neq 0,\;k\in{\mbox{\mathcal{C}}}^{t+1}, then we have
Similarly, suppose 0\in{\mbox{\mathcal{C}}}^{t+1}, then there exists an such that
Using the definition of the essentially cyclic update rule, we have that for all
Note that is finite, and that , and , we have
Using this result, taking limit for (34), and using the fact that , , , for all , we have
Due to the fact that for all , we have that the primal feasibility is achieved in the limit, i.e.,
This set of equalities together with (2.2) imply
To prove part 3, we first show that there exists a limit point for each of the sequences , and . Let us consider only the essentially cyclic rule. Due to the compactness assumption of , it is obvious that must have a limit point. Also by a similar argument leading to (30), we see that , thus for each , must also lie in a compact set thus have a limit point. Note that the Lipschitz continuity of combined with the compactness of the set implies that the set is bounded, therefore is a bounded sequence. Using (21), we conclude that that is also a bounded sequence, therefore must have at least one limit point.
We prove part 3 by contradiction. Because the feasible set is compact, then lies in a compact set. From the argument in the previous part it is easy to see that , also lie in some compact sets. Then every subsequence will have a limit point. Suppose that there exists a subsequence , and such that
where is some limit point, and by part 2, we have . By further restricting to a subsequence if necessary, we can assume that is the unique limit point.
Suppose that this sequence does not converge to the set of stationary solutions, i.e.,
Then it follows that there exists some such that
By the definition of the distance function we have
Combining the above two inequalities we must have
This contradicts to (40). The desired result is proven.
The analysis presented above is different from the conventional analysis of the ADMM algorithm where the main effort is to bound the distance between the current iterate and the optimal solution set. The above analysis is partly motivated by our previous analysis of the convergence of ADMM for multi-block convex problems, where the progress of the algorithm is measured by the combined decrease of certain primal and dual gaps; see [27, Theorem 3.1]. Nevertheless, the nonconvexity of the problem makes it difficult to estimate either the primal or the dual optimality gaps. Therefore we choose to use the decrease of the augmented Lagrangian as a measure of the progress of the algorithm.
Next we analyze the iteration complexity of the vanilla ADMM (i.e., Algorithm 1). To state our result, let us define the proximal gradient of the augmented Lagrangian function as
where is the proximity operator. We will use the following quantity to measure the progress of the algorithm
It can be verified that if , then a stationary solution of the problem (6) is obtained. We have the following iteration complexity result.
Suppose Assumption A is satisfied. Let denote an iteration index in which the following inequality is achieved
for some . Then there exists some constant such that
where is defined in Assumption A3.
Proof. We first show that there exists a constant such that
This proof follows similar steps of [27, Lemma 2.5]. From the optimality condition of the update step (14) we have
where in the last inequality we have used the nonexpansiveness of the proximity operator.
Similarly, the optimality condition of the subproblem is given by
Therefore, combining (48) and (49), we have
By taking , (47) is proved.
The inequalities (50) – (51) implies that for some
According to Lemma 2, there exists a constant such that
Summing both sides of the above inequality over , we have
where in the last inequality we have used the fact that is decreasing and lower bounded by (cf. Lemmas 2–3).
By utilizing the definition of and , the above inequality becomes
Dividing both sides by , and by setting , the desired result is obtained.
3 The Proximal ADMM
One potential limitation of Algorithms 1 and 2 is the requirement that each subproblem (15) needs to be solved exactly, while in certain practical applications cheap iterations are preferred. In this section, we consider an important extension of Algorithm 1–2 in which the above restriction is removed. The main idea is to take a proximal step instead of minimizing the augmented Lagrangian function exactly with respect to each variable block. Like in the previous section, we will analyze a generalized version, termed the flexible proximal ADMM, where there is more freedom in choosing the update schedules.
Algorithm 3. A Flexible Proximal ADMM for Problem (6) At each iteration , compute: (55) Pick a set {\mbox{\mathcal{C}}}^{t+1}\subseteq\{1,\cdots,K\}. If k\in{\mbox{\mathcal{C}}}^{t+1}, update by solving: (56) Update the dual variable: (57) Else let , .
Notice that the update step is different from the conventional proximal update (e.g., ). In particular, the linearization is done with respect to instead of computed in the previous iteration. This modification is instrumental in the convergence analysis of Algorithm 3.
Here we use the period- essentially cyclic rule to decide the set {\mbox{\mathcal{C}}}^{t+1} at each iteration. We note that there is a slight difference of the update schedule used in Algorithm 3 and Algorithm 2. In Algorithm 3, the block variable is updated in every iteration while in Algorithm 2 the update of is also governed by block selection rules.
Now we begin analyzing Algorithm 3. We make the following assumptions in this section (in addition to Assumption A1 and A3).
Assumption B. For all , the penalty parameter is chosen large enough such that:
Again let denote the last iteration that is updated before , i.e.,
Note that we do not need anymore since is updated in every iteration. Clearly, we have and as a result, . We have the following result.
Suppose Assumption B and Assumptions A1, A3 are satisfied. Then for Algorithm 3, the following is true for the essentially cyclic block selection rule
Proof. Suppose k\notin{\mbox{\mathcal{C}}}^{t+1}, then the inequality is trivially true, as .
For any k\in{\mbox{\mathcal{C}}}^{t+1}, we observe from the update of step (56) that the following is true
Therefore we have, for all k\in{\mbox{\mathcal{C}}}^{t+1}
where the last step follows from triangular inequality and the fact (cf. the definition of ). The above result further implies that
Next, we upper bound the successive difference of the augmented Lagrangian. To this end, let us define the following functions
Using these short-hand definitions, we have
Suppose Assumption A1 is satisfied. Let be generated by Algorithm 3 with essential cyclic block update rule. Then we have the following
Observe that when k\in{\mbox{\mathcal{C}}}^{t+1}, is generated according to (67). Due to the strong convexity of with respect to , we have
Further, we have the following series of inequalities
where the first two inequalities follow from Assumption A1. Combining (69) – (2.3) we obtain
Next, we bound the difference of the augmented Lagrangian function values.
Assume the same set up as in Lemma 7. Then we have
where we and are the positive constants defined in (58) and (59).
Proof. We first bound the successive difference . Again we decompose it as in (2.2), and bound the resulting two differences separately.
The first term in (2.2) can be again expressed as
To bound the second term in (2.2), we use Lemma 7. We use an argument similar to the proof of (25) to obtain
where the last inequality follows from Lemma 7 and the strong convexity of with respect to the variable (with modulus ) at .
Combining the above two inequalities, we obtain
where in we have used (62); in we have used the fact that ; in the last inequality we have used the definition of the period- essentially cyclic update rule which implies that
Then for any given , the difference is obtained by summing (2.3) over all iterations. Specifically, we obtain
We conclude that to make the rhs of (8) negative at each iteration, it is sufficient to require that and for all , or more specifically:
Note that one can always find a set of ’s large enough such that the above condition is satisfied.
Next we show that is convergent.
Suppose Assumption A1, A3 and Assumption B are satisfied. Then Algorithm 3 with period- essentially cyclic update rule generates a sequence of augmented Lagrangian, whose limit exists and is bounded below by .
Proof. Observe that the augmented Lagrangian can be expressed as
where is from (64); is due to the following inequalities
Clearly, combining the inequality (2.3) with Assumptions B and A3 yields that is lower bounded. It follows from Lemma 8 that whenever the penalty parameter ’s are chosen sufficiently large (as per Assumption B), will monotonically decrease and is convergent. This completes the proof.
Using Lemmas 6–9, we arrive at the following convergence result. The proof is similar to Theorem 4, and is thus omitted.
Suppose that Assumptions A1, A3 and B hold. Then the following is true for Algorithm 3.
We have , .
Let denote any limit point of the sequence generated by Algorithm 3 with period- essentially cyclic block update rule. Then is a stationary solution of problem (6).
If is a compact set, then Algorithm 3 with period-T essentially cyclic block update rule converges to the set of stationary solutions of problem (6). That is, the following is true
where is the set of primal-dual stationary solutions of problem (6).
The Nonconvex Sharing Problem
Consider the following well-known sharing problem (see, e.g., [8, Section 7.3] for motivation)
The augmented Lagrangian for this problem is given by
Note that we have chosen a special reformulation in (80): a single variable is introduced which leads to a problem with a single linear constraint. Applying the classical ADMM to this reformulation leads to a multi-block ADMM algorithm in which block variables are updated sequentially. As mentioned in the introduction, even in the case where the objective is convex, it is not known whether the multi-block ADMM converges in this case. Variants of the multi-block ADMM has been proposed in the literature to solve this type of multi-block problems; see recent developments in and the references therein.
The analysis of Algorithm 4 follows similar argument as that of Algorithm 3. Therefore we will only provide an outline for it.
First, we make the following assumptions in this section.
There exists a positive constant such that
Moreover, ’s are closed convex sets; each is full column rank so that , where denotes the minimum eigenvalue of a matrix.
The penalty parameter is chosen large enough such that:
Each subproblem (82) as well as the subproblem (83) is strongly convex, with modulus and , respectively.
, and that .
is lower bounded over .
is either smooth nonconvex or convex (possibly nonsmooth). For the former case, there exists such that , .
Note that compared with Assumptions A and B, in this case we no longer require that each to be smooth. Define an index set {\mbox{\mathcal{K}}}\subseteq\{1,\cdots,K\}, such that is convex if k\in{\mbox{\mathcal{K}}}, and nonconvex smooth otherwise. Further, the requirement that is full column rank is needed to make the subproblem (82) strongly convex.
Our convergence analysis consists of a series of lemmas whose proofs, for the most part, are omitted since they are similar to that of Lemma 1–Lemma 3.
Suppose Assumption C is satisfied. Then for Algorithm 4 with either essentially cyclic rule or the randomized rule, the following is true
Suppose Assumption C is satisfied. Then for Algorithm 4 with either essentially cyclic rule or the randomized rule, the following is true
Assume the same set up as in Lemma 12. Then the following limit exists and is bounded from below
Proof. We have the following series of inequalities
The last inequality comes from the fact that
Using assumptions C2.– C3. leads to the desired result.
We have the following main result for the nonconvex consensus problem.
Suppose that Assumption C holds. Then the following is true for Algorithm 4, either deterministically for the essentially cyclic update rule or almost surely for the randomized update rule.
We have , .
Let denote any limit point of the sequence generated by Algorithm 4. Then is a stationary solution of problem (80) in the sense that
If is a compact set for all , then Algorithm 4 converges to the set of stationary solutions of problem (80), i.e.,
where is the set of primal-dual stationary solution for problem (80).
The penalty parameter is chosen large enough such that .
Then the flexible ADMM algorithm (i.e., Algorithm 4), converges to the set of primal dual optimal solution of problem (6), either deterministically for the essentially cyclic update rule or almost surely for the randomized update rule.
Similar to the consensus problem, one can extend Algorithm 4 to its proximal version. Here the benefit offered by the proximal-type algorithms is twofold: i) one can remove the strong convexity requirement posed in Assumption C2-(1) ; ii) one can allow inexact and simple update for each block variable. However, the analysis is a bit more involved, as the penalty parameter as well as the proximal coefficient for each subproblem needs to be carefully bounded. Due to the fact that the analysis follows almost identical steps as those in Section 2.3, we will not present them here.
Extensions
In this paper, we analyze the behavior of the ADMM method in the absence of convexity. We show that when the penalty parameter is chosen sufficiently large, the ADMM and several of its variants converge to the set of stationary solutions for certain consensus and sharing problems.
Our analysis is based on using the augmented Lagrangian as a potential function to guide the iterate convergence. This approach may be extended to other nonconvex problems. In particular, if the following set of sufficient conditions (see Assumption D below) are satisfied, then the convergence of the ADMM is guaranteed for the nonconvex problem (1). It is important to note that in practice these conditions should be verified case by case for different applications, just like what we have done for the consensus and sharing problems.
The iterations are well defined, meaning the function is uniformly lower bounded for all .
There exists a constant such that , for all .
The penalty parameter is chosen large enough such that each subproblem is strongly convex with modulus , which is a nondecreasing function of . Further, for all .
Following a similar argument leading to Theorem 4, we can show that as long as Assumption D is satisfied, then the primal feasibility gap goes to zero in the limit, and that every limit point of the sequence is a stationary solution of problem (1). A few remarks on Assumption D are in order:
Assumption D1 is necessary for showing convergence. Without D1, even if one is able to show that the augmented Lagrangian is decreasing, one cannot claim the convergence to stationary solutions. The reason is that the augmented Lagrangian may go to In fact, it is very easy to modify the algorithm so that the augmented Lagrangian reduces at each iteration – just change the “+” in the dual update (16) to “-”. However, it is obvious that by doing this the dual variables will become unbounded, and the primal feasibility will never be satisfied. , therefore there is no way to guarantee that the successive difference of the iterates goes to , or the primal feasibility is satisfied in the limit.
The main drawback of Assumption D is that it is made on the iterates rather than on the problem. For different linearly constrained optimization problems, one still needs to verify that these conditions are indeed valid, as we have done for the consensus and the sharing problem considered in this paper.
Here we mention one more family of problems for which Assumption D can be verified. Consider
where is a convex possibly nonsmooth function; is a possibly nonconvex function, and has Lipschitzian gradient with modulus ; ; is an invertible matrix; and are lower bounded over the set . Consider the following ADMM method, where the iterate generated at iteration is given by
By using steps in Lemma 2.1-Lemma 2.3, one can verify that if , then Assumptions D1 holds true. By having large enough and by using the invertibility of , we can make the subproblem strongly convex, then Assumption D4 holds true. Other assumptions can be verified along similar lines. Note that in this case the convergence can be obtained with a slightly weaker condition in which the subproblem is convex but not necessarily strongly convex.