A forward-backward view of some primal-dual optimization methods in image recovery
Patrick L. Combettes, Laurent Condat, Jean-Christophe Pesquet, Bang Cong Vu
Introduction
Many image recovery problems can be formulated in Hilbert spaces and as structured optimization problems of the form
where, for every , is a proper lower semicontinuous convex function from to and is a bounded linear operator from to . For example, the functions may model data fidelity terms, smooth or nonsmooth measures of regularity, or hard constraints on the solution. In recent years, many algorithms have been developed to solve such a problem by taking advantage of recent advances in convex optimization, especially in the development of proximal tools (see and the references therein). In image processing, however, solving such a problem still poses a number of conceptual and numerical challenges. First of all, one often looks for methods which have the ability to split the problem by activating each of the functions through elementary processing steps which can be computed in parallel. This makes it possible to reduce the complexity of the original problem and to benefit from existing parallel computing architectures. Secondly, it is often useful to design algorithms which can exploit, in a flexible manner, the structure of the problem. In particular, some of the functions may be Lipschitz differentiable in which case they should be exploited through their gradient rather than through their proximity operator, which is usually harder to implement (examples of proximity operators with closed-form expression can be found in ). In some problems, the functions can be expressed as the infimal convolution of simpler functions (see and the references therein). Last but not least, in image recovery, the operators may be of very large size so that their inversions are costly (e.g., in reconstruction problems). Finding algorithms which do not require to perform inversions of these operators is thus of paramount importance.
Note that all the existing convex optimization algorithms do not have these desirable properties. For example, the Alternating Direction Method of Multipliers (ADMM) requires a stringent assumption of invertibility of the involved linear operator. Parallel versions of ADMM and related Parallel Proximal Algorithm (PPXA) usually necessitate a linear inversion to be performed at each iteration. Also, early primal-dual algorithms did not make it possible to handle smooth functions through their gradients. Only recently, have primal-dual methods been proposed with this feature. Such work was initiated in in the line of and subsequent developments can be found in . As will be seen in the present paper, another advantage of these approaches is that they can be coupled with variable metric strategies which can potentially accelerate their convergence.
In Section 2, we provide some background on convex analysis and monotone operator theory. In Section 3, we introduce a general form of the forward-backward algorithm which uses a variable metric. This algorithm is employed in Section 4 to develop a versatile family of primal-dual proximal methods. Several particular instances of this framework are discussed. Finally, we provide illustrating numerical results in Section 5.
Notation and background
Monotone operator theory provides a both insightful and elegant framework for dealing with convex optimization problems and developing new solution algorithms that could not be devised using purely variational tools. We summarize a number of related concepts that will be needed.
Throughout, , , and are real Hilbert spaces. We denote the scalar product of a Hilbert space by and the associated norm by . The symbol denotes weak convergence, In a finite dimensional space, weak convergence is equivalent to strong convergence. and denotes the identity operator. We denote by the space of bounded linear operators from to , we set , where denotes the adjoint of . The Loewner partial ordering on is denoted by . For every , we set and we denote by the square root of . Moreover, for every and , we define the norm .
We denote by the Hilbert direct sum of the Hilbert spaces , i.e., their product space equipped with the scalar product where and denote generic elements in .
Let be a set-valued operator. We denote by the graph of , by the set of zeros of , and by its range. The inverse of is , and the resolvent of is . Moreover, is monotone if
and maximally monotone if it is monotone and there exists no monotone operator such that and . An operator is -cocoercive for some if
The conjugate of a function is
and the infimal convolution of with is
The class of lower semicontinuous convex functions such that is denoted by . If , then and the subdifferential of is the maximally monotone operator
Let for some . The proximity operator of relative to the metric induced by is [22, Section XV.4]
When , we retrieve the standard definition of the proximity operator . Let be a nonempty subset of . The indicator function of is defined on as
A general form of Forward-Backward algorithm
Optimization problems can often be reduced to finding a zero of a sum of two maximally monotone operators and acting on . When is cocoercive (see (3)), a useful algorithm to solve this problem is the forward-backward algorithm, which can be formulated in a general form involving a variable metric as shown in the next result.
Then for some .
A variable metric primal-dual method
A wide array of optimization problems encountered in image processing are instances of the following one, which was first investigated in and can be viewed as a more structured version of the minimization problem in (1):
Let us now introduce the product space and the operators
The operator can be shown to be maximally monotone,whereas is cocoercive. A key observation in this context is that, if there exists such that , then is a pair of primal-dual solutions to Problem 4.1 . This connection with the construction for a zero of makes it possible to apply a forward-backward algorithm as discussed in Section 3, by using a linear operator to change the metric at each iteration . Depending on the form of this operator various algorithms can be obtained.
2 A first class of primal-dual algorithms
The following result constitutes a direct extension of [14, Example 6.4]:
3 A second class of primal-dual algorithms
where is given by (18).
The following result can then be deduced from Theorem 3.1. Its proof is skipped due to the lack of space.
Application to image restoration
An estimate of is computed as a solution to (12) where , , , ,
The restored image is displayed in Fig. 1(d). Fig. 2 shows the convergence profile of the algorithm. We plot the evolution of the normalized Euclidean distance (in log scale) between the iterates and in terms of computational time (Matlab R2011b codes running on a single-core Intel i7-2620M CPU@2.7 GHz with 8 GB of RAM). An approximation of obtained after 5000 iterations is used. This result illustrates the fact that an appropriate choice of the metric may be beneficial in terms of speed of convergence.