Gradient Descent-Ascent Provably Converges to Strict Local Minmax Equilibria with a Finite Timescale Separation
Tanner Fiez, Lillian Ratliff
Introduction
In this paper we study learning in zero-sum games of the form
As a result of this perspective, there has been significant interest in the study of gradient descent-ascent owing to the fact that the learning rule is computationally efficient and a natural analogue to gradient descent from function optimization. Formally, the learning dynamics are given by each player myopically updating a strategy with an individual gradient as follows:
The analysis of gradient descent-ascent is complicated by the intricate optimization landscape in non-convex, non-concave zero-sum games. To begin, there is the fundamental question of what type of solution concept is desired. Given the class of games under consideration, local solution concepts have been proposed and are often taken to be the goal of a learning algorithm. The primary notions of equilibrium that have been adopted are the local Nash and local minmax/Stackelberg concepts with a focus on the set of strict local equilibrium that can be characterized by gradient-based sufficient conditions. Following several past works, from here on we refer to strict local Nash equilibrium and strict local minmax/Stackelberg equilibrium as differential Nash equilibrium and differential Stackelberg equilibrium, respectively.
Regardless of the equilibrium notion under consideration, a number of past works highlight failures of standard gradient descent-ascent in non-convex, non-concave zero-sum games. Indeed, it has been shown gradient descent-ascent with a shared learning rate () is prone to reaching critical points that are neither differential Nash equilibrium nor differential Stackelberg equilibrium (Daskalakis and Panageas, 2018; Mazumdar et al., 2020; Jin et al., 2020). While an important negative result, it does not rule out the prospect that gradient descent-ascent may be able to guarantee equilibrium convergence as it fails to account for a key structural parameter of the learning dynamics, namely the ratio of learning rates between the players.
Motivated by the observation that the order of play between players is fundamental to the definition of the game, the role of timescale separation in gradient descent-ascent has been explored theoretically in recent years (Heusel et al., 2017; Chasnov et al., 2019; Jin et al., 2020). On the empirical side of past work, it has been widely demonstrated and prescribed that timescale separation in gradient descent-ascent between the generator and discriminator, either by heterogeneous learning rates or unrolled updates, is crucial to improving the solution quality when training generative adversarial networks (Goodfellow et al., 2014; Arjovsky et al., 2017; Heusel et al., 2017). Denoting as the learning rate of the player 1, the learning rate of player 2 can be redefined as where is the ratio of learning rates or timescale separation parameter. The work of Jin et al. (2020) took a meaningful step toward understanding the effect of timescale separation in gradient descent-ascent by showing that as the stable critical points of the learning dynamics coincide with the set of differential Stackelberg equilibrium. In simple terms, the aforementioned result implies that all ‘bad critical points’ (that is, critical points lacking game-theoretic meaning) become unstable as the timescale separation approaches infinity and that all ‘good critical points’ (that is, game-theoretically meaningful equilibria) remain or become stable as the timescale separation approaches infinity. While a promising theoretical development on the local stability of the underlying dynamics, it does not lead to a practical, implementable learning rule or necessarily provide an explanation for the satisfying performance in applications of gradient descent-ascent with a finite timescale separation. It remains an open question to fully understand gradient descent-ascent as a function of the timescale separation and to determine whether the desirable behavior with an infinite timescale separation is achievable for a range of finite learning rate ratios.
This paper continues the theoretical study of gradient descent-ascent with timescale separation in non-convex, non-concave zero-sum games. We focus our attention on answering the remaining open questions regarding the behavior of the learning dynamics with finite learning rate ratios and provide a number of conclusive results. Notably, we develop necessary and sufficient conditions for a critical point to be stable for a range of finite learning rate ratios. The results imply that differential Stackelberg equilibria are stable for a range of finite learning rate ratios and that non-equilibria critical points are unstable for a range of finite learning rate ratios. Together, this means gradient descent-ascent only converges to differential Stackelberg equilibrium for a range of finite learning rate ratios. To our knowledge, this is the first provable guarantee of its kind for an implementable first-order method. Moreover, the technical results in this work rely on tools that have not appeared in the machine learning and optimization communities analyzing games and expose a number interesting directions of future research. Explicitly, the notion of a guard map, which is arguably even an obscure tool in modern control and dynamical systems theory, is ‘rediscovered’ in this work as a technique for analyzing the stability of game dynamics.
To motivate our primary theoretical results, we present a self-contained description of what is known about the local stability of gradient descent-ascent around critical points in Section 3.1. The existing results primarily concern gradient descent-ascent without timescale separation and with a ratio of learning rates approaching infinity (see Figure 1 for a graphical depiction of known results in each regime). In contrast, this paper is focused on characterizing the stability and convergence of gradient descent-ascent across a range of finite learning rate ratios. To hint at what is achievable in this realm, we present simple examples for which gradient descent-ascent converges to non-equilibrium critical points and games with differential Stackelberg equilibrium that are unstable with respect to gradient descent-ascent without timescale separation (see Examples 1 and 2, Section 3). While the existence of such examples is known (Daskalakis and Panageas, 2018; Mazumdar et al., 2020; Jin et al., 2020), we demonstrate in them that a finite timescale separation is sufficient to remedy the undesirable stability properties of gradient descent-ascent without timescale separation.
Toward characterizing this phenomenon in its full generality, we provide intermediate results which are known, but we prove using technical tools not yet broadly seen and exploited by this community. To begin, it is known that the set of differential Nash equilibrium are stable with respect to gradient descent-ascent (Mazumdar and Ratliff, 2019; Daskalakis and Panageas, 2018), and that they remain stable for any timescale separation parameter (Jin et al., 2020). We provide a proof for this result (Proposition 3.1) using the concept of quadratic numerical range (Tretter, 2008). Furthermore, Jin et al. (2020) recently showed that as the timescale separation , the stable critical points of gradient descent-ascent coincide with the set of differential Stackelberg equilibrium. We reveal that this result has long existed in the literature on singularly perturbed systems (Kokotovic et al., 1986, Chapter 2 and the citations within) and provide a proof (see Proposition 3.2) using analysis methods from the aforementioned line of work that are novel to the literature on learning in games from the machine learning and optimization communities in recent years.
A relevant line of study on singularly perturbed systems is that of characterizing the range of perturbation parameters for which a system is stable (Kokotovic et al., 1986; Saydy et al., 1990; Saydy, 1996). Debatably introduced by Saydy et al. (1990), guardian or guard maps act as a certificate that the roots of a polynomial lie in a particular guarded domain for a range of parameter values. Historically, guard maps serve as a tool for studying the stability of parameterized families of dynamical systems. We bring this tool to learning in games and construct a map that guards a class of Hurwitz stable matrices parameterized by the timescale separation parameter in order to analyze the range of learning rate ratios for which a critical point is stable with respect to gradient descent-ascent. This technique leads to the following result.
Consider a sufficiently regular critical point of gradient descent-ascent. There exists a such that is stable for all if and only if is a differential Stackelberg equilibrium.
Theorem 3.1 confirms that there does indeed exist a range of finite learning ratios such that a differential Stackelberg equilibrium is stable with respect to gradient descent-ascent. Moreover, such a range of learning rate ratios only exists if a critical point is a differential Stackelberg equilibrium. As we show in Corollary 4.1, the former implication of Theorem 3.1 nearly immediately implies there exists a such that gradient descent-ascent converges locally asymptotically for all if and only if is a differential Stackelberg equilibrium given a suitably chosen learning rate and deterministic gradient feedback. We give an explicit asymptotic rate of convergence in Theorem 4.1 and characterize the iteration complexity in Corollary 4.2. Moreover, we extend the convergence guarantees to stochastic gradient feedback in Theorem 4.2.
The latter implication of Theorem 3.1 says that there exists a finite learning rate ratio such that a non-equilibrium critical point of gradient descent-ascent is unstable. Building off of this, we complement the stability result of Theorem 3.1 with the following analagous instability result.
Consider any stable critical point of gradient descent-ascent which is not a differential Stackelberg equilibrium. There exists a finite learning rate ratio such that is unstable for all .
Theorem 3.2 establishes that there exists a range of finite learning ratios non-equilibrium critical points are unstable with respect to gradient descent-ascent. This implies that for a suitably chosen finite timescale separation, gradient descent-ascent avoids critical points lacking game-theoretic meaning. Together, Theorem 3.1 and Theorem 3.2 answer affirmatively that gradient descent-ascent with timescale separation can guarantee equilibrium convergence, which answers a standing open question. Moreover, we provide explicit constructions for computing and given a critical point. In fact our construction of in Theorem 3.1 is tight, and this is confirmed by our numerical experiments.
We finish the theoretical analysis of gradient descent-ascent in this paper by connecting to the literature on generative adversarial networks. We show under common assumptions on generative adversarial networks (Nagarajan and Kolter, 2017; Mescheder et al., 2018) that the introduction of gradient penalty based regularization to the discriminator does not change the set of critical points for the dynamics and, further, there exists a finite learning rate ratio such that for any learning rate ratio and any non-negative, finite regularization parameter , the continuous time limiting regularized learning dynamics remain stable, and hence, there is a range of learning rates for which the discrete time update locally converges asymptotically.
The theoretical results we provide are complemented by extensive experiments. In simulation, we explore a number of interesting behaviors of gradient descent-ascent with timescale separation analyzed theoretically including differential Stackelberg equilibria shifting from being unstable to stable and non-equilibrium critical points moving from being stable to unstable. Furthermore, we examine how the vector field and the spectrum of the game Jacobian evolve as a function of the timescale separation and explore the relationship with the rate of convergence. We experiment with gradient descent-ascent on the Dirac-GAN proposed by Mescheder et al. (2018) and illustrate the interplay between timescale separation, regularization, and rate of convergence. Building on this, we train generative adversarial networks on the CIFAR-10 and CelebA datasets with regularization and demonstrate that timescale separation can benefit performance and stability. In the experiments we observe that regularization and timescale separation are intimately connected and there is an inherent tradeoff between them. This indicates that insights made on simple generative adversarial network formulations may carry over to the complex problems where players are parameterized by neural networks.
Collectively, the primary contribution of this paper is the near-complete characterization of the behavior of gradient descent-ascent with finite timescale separation. Moreover, by introducing a novel set of analysis tools to this literature, our work opens a number of future research questions. As an aside, we believe these technical tools open up novel avenues for not only proving results about learning dynamics in games, but also for synthesizing algorithms.
2 Organization
The organization of this paper is as follows. Preliminaries on game theoretic notions of equilibria, gradient-based learning algorithms, and dynamical systems theory are reviewed in Section 2.
Convergence analysis proceeds in two phases. In Section 3, we study the stability properties of the continuous time limiting dynamical system given a timescale separation between the minimizing and maximizing players. Specifically, we show the first result on necessary and sufficient conditions for convergence of the continuous time limiting system corresponding to gradient descent-ascent with time scale separation to game theoretically meaningful equilibria (i.e., local minmax equilibria in zero-sum games). Following this, in Section 4, we provide convergence guarantees for the original discrete time dynamical system of interest (namely, gradient descent ascent). Using the results in the proceeding section, we show that gradient descent-ascent converges to a critical point if and only if it is a differential Stackelberg equilibrium (i.e., a sufficiently regular local minmax). In addition, we characterize the iteration complexity of gradient descent-ascent dynamics and provide finite-time bounds on local convergence to approximate local Stackelberg equilibria.
We apply the main results in the preceding sections to generative adversarial networks in Section 5, and in Section 6 we present several illustrative examples including generative adversarial networks where we show that tuning the learning rate ratio along with regularization and the exponential moving average hyperparameter significantly improves the Fréchet Inception Distance (FID) metric for generative adversarial networks.
Given its length, prior to concluding in Section 7, we review related work drawing connections to solution concepts, gradient descent-ascent learning dynamics, applications to adversarial learning where the success of heuristics provide strong motivation for the theoretical work in this paper, and historical connections to dynamical systems theory. Throughout the sections proceeding Section 7, we draw connections to related works and results in an effort to place our results in the context of the literature. We conclude in Section 8 with a discussion on the significance of the results and open questions. The appendix includes the majority of the detailed proofs as well as additional experiments and commentary.
Preliminaries
In this section, we review game theoretic and dynamical systems preliminaries. Additionally, we formulate the class of learning rules analyzed in this paper.
There are two natural equilibrium concepts for such games depending on the order of play—i.e., the Nash equilibrium concept in the case of simultaneous play and the Stackelberg equilibrium concept in the case of hierarchical play. Each notion of equilibria can be characterized as the intersection points of the reaction curves of the players (Başar and Olsder, 1998).
The joint strategy is a local Nash equilibrium on , where , if , for all and for all . Furthermore, if the inequalities are strict, we say is a strict local Nash equilibrium.
Consider for where, without loss of generality, player 1 is the leader (minimizing player) and player 2 is the follower (maximizing player). The strategy is a local Stackelberg solution for the leader if, ,
where is the reaction curve. Moreover, for any , the joint strategy profile is a local Stackelberg equilibrium on .
Predicated on existence,Characterizing existence of equilibria is outside the scope of this work. However, we remark that Nash equilibria exist for convex costs on compact and convex strategy spaces and Stackelberg equilibria exist on compact strategy spaces (Başar and Olsder, 1998, Thm. 4.3, Thm. 4.8, & Sec. 4.9). equilibria can be characterized in terms of sufficient conditions on player costs. Indeed, in continuous games, first and second order conditions on player cost functions leads to a differential characterization (i.e., necessary and sufficient conditions) of local Nash equilibria reminiscent of optimality conditions in nonlinear programming (Ratliff et al., 2016).The differential characterization of local Nash equilibria in continuous games was first reported in (Ratliff et al., 2013). Genericity and structural stability we studied in general-sum settings in (Ratliff et al., 2014) and in zero-sum settings in (Mazumdar and Ratliff, 2019).
We denote as the derivative of with respect to , as the partial derivative of with respect to , as the partial derivative of with respect to , and as the total derivative.Example: given , .
If is a local Nash equilibrium of the zero-sum game , then , , and . On the other hand, if , , and and , then is a local Nash equilibrium.
The following definition, characterized by sufficient conditions for a local Nash equilibrium as defined in Definition 2.1, was first introduced in (Ratliff et al., 2013).
The joint strategy is a differential Nash equilibrium if , , and .
The joint strategy is a differential Stackelberg equilibrium if , , , and .
Observe that in a general sum setting the first order conditions for player are equivalent the total derivative of being zero at the candidate critical point where is implicitly defined as a function of via the implicit mapping theorem applied to . Since in this paper and in Definition 2.4, the class of games is zero sum, and (along with the condition that which is implied by the second order conditions) are sufficient to imply that the total derivative is zero.
The Jacobian of the first order necessary and sufficient condition—i.e., conditions that define potential candidate differential Nash and/or Stackelberg equilibria—is a useful mathematical object for understanding convergence properties of gradient based learning rules as we will see in subsequent sections. Consider the vector of individual gradients which define first order conditions for a differential Nash equilibrium. Let denote the Jacobian of which is defined by
We recall from Fiez et al. (2020) an alternative (to Definition 2.4, but equivalent set of sufficient conditions for a differential Stackelberg in terms of . Let denote the Schur complement of with respect to the block-row matrix in .
2 Gradient-based learning algorithms
As noted above, in this paper we focus on settings in which agents or players in this game are seeking equilibria of the game via a learning algorithm. We study arguably the most natural learning rule in zero-sum continuous games: gradient descent-ascent (GDA). This gradient-based learning rule is a simultaneous gradient play algorithm in that agents update their actions at each iteration simultaneously.
Gradient descent-ascent is defined as follows. At iteration , each agent updates their choice variable by the process
where is agent ’s learning rate, and is agent ’s gradient-based update mechanism. For simultaneous gradient play,
is the vector of individual gradients and in a zero-sum setting, GDA is defined using where the first player is the minimizing player and the second player is the maximizing player.
We analyze the iteration complexity or local asymptotic rate of convergence of learning rules of the form (2) in the neighborhood of an equilibrium. Given two real valued functions and , we write if there exists a positive constant such that . For example, consider iterates generated by (2) with initial condition and critical point . Suppose that we show . Then, we write where .
3 Dynamical Systems Primer
In this paper, we study learning rules employed by agents seeking game-theoretically meaningful equilibria in continuous games. Dynamical systems tools for both continuous and discrete time play a crucial role in this analysis.
Before we proceed, we recall and remark on some facts from dynamical systems theory concerning stability of equilibria in the continuous-time dynamics
relevant to convergence analysis for the discrete-time learning dynamics in (2). Observe that equilibria are shared between (2) and (5). Our focus is on the subset of equilibria that satisfy Definition 2.4, and the subset thereof defined in Definition 2.3. Recall the following equivalent characterizations of stability for an equilibrium of (5) in terms of the Jacobian matrix .
The continuous time dynamical system takes the form due to the timescale separation . Such a system is known as a singularly perturbed system or a multi-timescale system in the dynamical systems theory literature (Kokotovic et al., 1986), particularly where is small. Singularly perturbed systems are classically expressed as
where is most often a physically meaningful quantity inherent to some dynamical system that describes the evolution of some physical phenomena; e.g., in circuits it may be a constant related to device material properties, and in communication networks, it is often the speed at which data flows through a physical medium such as cable.
Given a dynamical system , the state or solution of the system at time starting from at time is called the flow and is denoted .
Consider the -dimensional dynamical system with equilibrium point . If has no zero or purely imaginary eigenvalues, there is a homeomorphism defined on a neighborhood of taking orbits of the flow to those of the linear flow of —that is, the flows are topologically conjugate. The homeomorphism preserves the sense of the orbits and is chosen to preserve parameterization by time.
The above theorem says that the qualitative properties of the nonlinear system in the vicinity (which is determined by the neighborhood ) of an isolated equilibrium are determined by its linearization if the linearization has no eigenvalues on the imaginary axes in the complex plane. We also remark that Hartman-Grobman can also be applied to discrete time maps (cf. Sastry (1999, Thm. 2.18)) with the same qualitative outcome.
In proving results for stochastic gradient descent-ascent, we leverage what is known as the ordinary differential equation method in which the flow of the limiting continuous time system starting at sample points from the stochastic updates of the players actions is compared to asymptotic psuedo-trajectories—i.e., linear interpolations between sample points. To understand stability in the stochastic case, we need the notion of internally chain transitive sets. For more detail, the reader is referred to (Alongi and Nelson, 2007, Chap. 2–3).
Stability of Continuous Time GDA with Timescale Separation
To characterize the convergence of -GDA, we begin by studying its continuous time limiting system
By analyzing the stability of the continuous time system as a function of the timescale separation using the Jacobian from (8) in this section, we can then draw conclusions about the stability and convergence of the discrete time system -GDA in Section 4.
The organization of this section is as follows. To begin, we present a collection of preliminary observations in Section 3.1 regarding the stability of continuous time gradient descent-ascent with timescale separation to motivate the results in the subsequent subsections by establishing known results and introducing alternative analysis methods that the technical results in this paper build on. Then, in Sections 3.2 and 3.3 respectively, we present necessary and sufficient conditions for stability of the continuous time system around critical points in terms of the learning rate ratio along with sufficient conditions to guarantee the instability of the continuous time system around non-equilibrium critical points in terms of the timescale separation.
In Figure 1 we present a graphical representation of known results on the stability of gradient descent-ascent with timescale separation in continuous time, where we remark that such results nearly directly imply equivalent conclusions regarding the discrete time system -GDA with a suitable choice of learning rate . The primary focus of past work has been on the edge cases of and . For , the set of differential Nash equilibrium are stable, but differential Stackelberg equilibrium may be stable or unstable, and non-equilibrium critical points can be stable. As , the set of differential Nash equilibrium remain stable, each differential Stackelberg equilibrium is guaranteed to become stable, and each non-equilibrium critical point must be unstable. We fill the gap between the known results by providing results as a function of finite . With an eye toward this goal, we now provide examples and preliminary results that illustrate the type of guarantees that may be achievable for a range of finite learning rate ratios.
To start off, we consider the set of differential Nash equilibrium. It is nearly immediate from the structure of the Jacobian that each differential Nash equilibrium is stable for (Mazumdar et al., 2020; Daskalakis and Panageas, 2018). Moreover, Jin et al. (2020) showed that regardless of the value of , the set of differential Nash equilibrium remain stable. In other words, the desirable stability characteristics of differential Nash equilibrium are retained for any choice of timescale separation. We state this result as a proposition for later reference and since our proof technique relies on the concept of quadratic numerical range (Tretter, 2008), which has not appeared previously in this context. The proof of Proposition 3.1 is provided in Appendix B.
Fiez et al. (2020) show that the set of differential Nash equilibrium is a subset of the set of differential Stackelberg equilibrium. In other words, any differential Nash equilibrium is a differential Stackelberg equilibrium, but a differential Stackelberg equilibrium need not be a differential Nash equilibrium. Moreover, Jin et al. (2020) show that the result of Proposition 3.1 fails to extend from differential Nash equilibria to the broader class of differential Stackelberg equilibrium. Indeed, not all differential Stackelberg equilibrium are stable with respect to the continuous time limiting dynamics of gradient descent-ascent without timescale separation. However, as the following example demonstrates, differential Stackelberg equilibrium that are unstable without timescale separation can become stable for a range of finite timescale learning rate ratios.
Within the class of zero-sum games, there exists differential Stackelberg equilibrium that are unstable with respect to and stable with respect to for all where is finite. Indeed, consider the quadratic zero-sum game defined by the cost
We explore Example 1 further via simulations in Section 6.1. The key takeaway from Example 1 is that it is clearly not always necessary for the timescale separation to approach infinity in order to guarantee the stability of a differential Stackelberg equilibrium and instead there exists a sufficient finite learning rate ratio. Put simply, the undesirable property of differential Stackelberg equilibria not being stable with respect to gradient descent-ascent without timescale separation can potentially be remedied with only a finite timescale separation.
It is well-documented that some stable critical points of the continuous time gradient descent-ascent limiting dynamics without timescale separation can lack game-theoretic meaning, as they may be neither a differential Nash equilibria nor differential Stackelberg equilibria (Mazumdar et al., 2020; Daskalakis and Panageas, 2018; Jin et al., 2020). The following example demonstrates that such undesirable critical points that are stable without timescale separation can become unstable for a range of finite learning ratios.
Within the class of zero-sum games, there exists non-equilibrium critical points that are stable with respect to and unstable with respect to for all where is finite. Indeed, consider a zero sum game defined by the cost
The game construction from (9) is quadratic and as a result has a unique critical point. Games can be constructed in which critical points lacking game-theoretic meaning that are stable without timescale separation become unstable for all even in the presence of multiple equilibria. Indeed, consider a zero-sum game defined by the cost
We investigate the game defined in (10) from Example 2 with simulations in Section 6.2. In an analogous manner to Example 1, Example 2 demonstrates that it is not always necessary for the timescale separation to approach infinity in order to guarantee non-equilibrium critical points become unstable as there can exist a sufficient finite learning rate ratio. This is to say that the unwanted property of non-equilibrium critical points being stable without timescale separation can also potentially be remedied with only a finite timescale separation.
The examples of this section have provided evidence that there exists a range of finite learning rate ratios for which differential Stackelberg equilibrium are stable and a range of learning rate ratios for which non-equilibrium critical points are unstable. Yet, no result has appeared in the literature on gradient descent-ascent with timescale separation confirming this behavior in general. We focus on doing precisely that in the subsection that follows. Before doing so, we remark on the closest existing result. As mentioned previously Jin et al. (2020) show that as , the set of stable critical points with respect to the dynamics coincide with the set of differential Stackelberg equilibrium. However, an equivalent result in the context of general singularly perturbed systems has been known in the literature (cf. Kokotovic et al. 1986, Chap. 2). We give a proof based on this type of analysis because it reveals a new set of analysis tools to the study of game-theoretic formulations of machine learning and optimization problems; a proof sketch is given below while the full proof is given in Appendix F.
The basic idea in showing this result is that there is a (local) transformation of coordinates from the linearized dynamics of , which we write as
in a neighborhood of a critical point to an upper triangular system that depends parametrically on and hence, the asymptotic behavior is readily obtainable from the block diagonal components of the system in the new coordinates. Indeed, consider the change of variables for the second player so that
A transformation of coordinates such that always exists (cf. Lemma F.1, Appendix F). Hence, the characteristic equation of (11) can be expressed as
where and with . As , . Consequently, of the eigenvalues of , denoted by , are the roots of the slow characteristic equation and the rest of the eigenvalues are denoted by for and where are the roots of the fast characteristic equation . The roots of are precisely those of the (first) Schur complement of while the roots of are precisely those of . ∎
This simple transformation of coordinates to an upper triangular dynamical system shown in (11) leads immediately to the asymptotic result in Proposition 3.2. It also shows that if the eigenavlues of are distinctDistinct eigenvalues is a generic property in the space of real matrices. and similarly, so are those of (although, and are allowed to have eigenvalues in common), then the asymptotic results from Proposition 3.2 imply the following approximations for the elements of :
This follows simply by observing that when the eigenvalues are distinct, the derivatives and are well-defined by the implicit mapping theorem and the total derivative of and , respectively.
2 Necessary and Sufficient Conditions for Stability
The proof of Proposition 3.2 provides some intuition for the next result, which is one of our main contributions. Indeed, as shown in Kokotovic et al. (1986, Chap. 2), as the first eigenvalues of tend to fixed positions in the complex plane defined by the eigenvalues of , while the remaining eigenvalues tend to infinity, with the linear rate , along as asymptotes defined by the eigenvalues of . The asymptotic splitting of the spectrum provides some intuition for the following result.
Before getting into the proof sketch, we provide some intuition for the construction of and along the way revive an old analysis tool from dynamical systems theory which turns out to be quite powerful in analyzing stability properties of parameterized systems.
There is still the question of how to construct such a and do so in a way that is as tight as possible. Recall Theorem 2.1 which states that a matrix is exponentially stable if and only if there exists a symmetric positive definite such that . The operator is known as the Lyapunov operator. Given a positive definite , is stable if and only if there exists a unique solution to
so that if is a non-degenerate (a condition implied by the hyperbolicity of for and ), does not change the properties of the guard map. In particular, the values of where does not depend on . Hence, we can use the reduced guard map
Reflecting back to (12), we see that this guard map in is closely related to the vectorization of the Lyapunov operator and of course, this is not a coincidence. For any symmetric positive definite , there will be a symmetric positive definite solution of the Lyapunov equation
The ‘necessary’ direction follows directly from the above observation, while the ‘sufficiency’ direction follows by construction.
In short, selecting the maximum value of over the finite set of equilibria guarantees that the local linearization of around any differential Stackelberg equilibria is stable, and hence, the nonlinear system is locally stable around each of these critical points.
Before moving on, we remark on the utility of the algebraic tools we use in the proof for Theorem 3.1. Indeed, the guard map concept is extremely powerful for understanding stability of parameterized families of dynamical systems, and it is not limited to single parameter families. Hence, there is potential to extend the above results to games with more than two players or additional parameters. In fact, we do exactly this in Section 5 where we present results for GANs trained with gradient-penalty type regularizers for the discriminator. Moreover, it is fairly easy to construct analogous guard maps for non-zero sum games. Many of the tools and constructions readily extend. We leave these results to a different paper so as to not create too much clutter in the present work.
3 Sufficient Conditions for Instability
Unlike Theorem 3.1, in Theorem 3.2 is not tight in the sense that may become unstable for . The reason for this is that there are potentially many matrices and that satisfy such that and have the same inertia; an analogous statement holds for , and . The choice of these matrices impact the value of . Hence, the question of finding the exact value of beyond which a spurious stable critical point for -GDA is unstable remains open.
Provable Convergence of GDA with Timescale Separation
In this section, derive convergence guarantees for -GDA to differential Stackelberg equilibria in both the deterministic (i.e., where agents have oracle access to their individual gradients) and the stochastic (i.e., where agents have an unbiased estimator of their individual gradient) settings.
As a corollary to Theorem 3.1, we first show that the discrete time -GDA update is locally asymptotically stable for a range of learning rates .
We need the following lemma to prove asymptotic convergence as well as the subsequent results on convergence rates.
In addition to showing asymptotic convergence, we also provide an asymptotic convergence rate. To prove the main theorems on convergence rates for both differential Stackelberg and differential Nash equilibria, we use a common argument which is summarized in the lemma below.
The full proof of the above lemma is provided in Appendix C, and a short proof sketch with the main ideas is summarized below.
Consider a zero-sum game as articulated in the lemma statement where is either a differential Stackelberg or differential Nash equilibrium and fix any such that .
For the discrete time dynamical system , it is well known that if is chosen such that , then locally asymptotically converges to (cf. Proposition A.1, Appendix A). With this in mind, we formulate an optimization problem to find the upper bound on the learning rate such that for all , the spectral radius of the local linearization of the discrete time map is a contraction which is precisely . The optimization problem is given by
The above lemma provides a convergence rate given a differential Stackelberg equilibrium and a learning rate ratio such that is stable with respect to the dynamics .
The following theorem—which uses Lemma 4.2 in its proof—characterizes the iteration complexity for -GDA. Specifically, the result leverages Theorem 3.1 to construct a finite such that is (Hurwitz) stable, and then for any , Lemma 4.2 implies a local asymptotic convergence rate.
is locally asymptotically (in fact, exponentially) stable by the Hartman-Grobman theorem (cf. Theorem 2.2). Therefore, for any , by Lemma 4.2, -GDA converges with a rate of . ∎
Theorem 4.1 directly implies a finite time convergence guarantee for obtaining an -differential Stackelberg equilibrium—i.e., an point with an -ball around a differential Stackelberg .
Given , under the assumptions of Theorem 4.1, -GDA obtains an –differential Stackleberg equilibrium in iterations for any with where is the local Lipschitz constant of .
We note that we have essentially given a proof that there exists a neighborhood on which -GDA converges. Of course, due to the non-convexity of the problem in general, this neighborhood could be arbitrarily small. We provide an estimate of the neighborhood size using the local Lipschitz constant of the local linearization . One way to better understand the size of this neighborhood is to use Lyapunov analysis, a tool which is well explored in the singular perturbation theory (Kokotovic et al., 1986). In particular, Lyapunov methods can be applied directly to the nonlinear system if one can construct Lyapunov functions for the fast and slow subsystems individually—also known as the boundary layer model and reduced order model. With these Lyapunov functions in hand, one can “stitch” the two together (via convex combination) and show under some reasonable assumptions that this combined function is a Lyapunov function for the overall singularly perturbed system. The benefit of this analysis is that the Lyapunov function gives one an estimate of the region of attraction (via, e.g., the level sets); however, it is not easy to construct a Lyapunov function for a nonlinear system in general. We leave expanding to such methods to future work.
Before turning to the stochastic setting, we comment on saddle point avoidance in the deterministic setting. It was shown by Mazumdar et al. (2020) that gradient-based learning in continuous games with heterogeneous learning rates avoids saddles on all but a set of measure zero initializations. Hence, -GDA avoids saddles for almost every initialization. We also know that all differential Nash equilibria are locally asymptotically stable for zero-sum settings. Hence, there are no differential Nash equilibria that are saddle points of the dynamics . On the other hand, as Example 1 shows, there are differential Stackelberg equilibria which correspond to saddle points of the dynamics for some choices of —in particular, in that example. Theorem 3.1 and Corollary 3.1, however, implies that for a given zero-sum game (or minmax problem), there exists a finite such that all locally asymptotically stable equilibria are differential Stackelberg equilibria. Hence, an ‘almost sure’ saddle point avoidance result together with the local convergence guarantee provided by Theorem 4.1 provides a strong characterization of long-run learning behavior.
Avoidance of saddles nor the if and only if convergence guarantee of Theorem 3.1 are, however, enough to ensure avoidance of limit cycles. In fact, it is known that limit cycles can exist in zero sum games (Daskalakis et al., 2018; Mazumdar et al., 2020). Understanding when such complex phenomena exist in games and determining how to ascribe meaning the behavior is an active area of study (see, e.g., the work of Papadimitriou and Piliouras (2019)).
2 Convergence of Stochastic GDA with Timescale Separation
In this section, we analyze convergence when players do not have oracle access to their gradients but instead have an unbiased estimator in the presence of zero mean, finite variance noise. Specifically, we show that the agents will converge locally asymptotically almost surely to a differential Stackelberg equilibrium.
The stochastic form of the update is given by
where is a zero mean, finite variance random variable and is the learning rate sequence.
The stochastic process is a martingale difference sequence with respect to the increasing family of -fields defined by
We note that this assumption has been relaxed in the literature (cf. Thoppe and Borkar (2019)), however simplicity, we state the theorem with the most accessible criteria. We remark below in the paragraph on extensions to concentration bounds on the nature of the relaxed assumptions.
The convergence of to a, possibly sample path dependent, compact connected internally chain transitive invariant set of follows from classical results in stochastic approximation theory (cf. Borkar (2008, Chap. 2); Benaim (1996)).
While (local) almost sure convergence in gradient descent-ascent (Chasnov et al., 2019) to a critical pointTo date it has not been shown that for a sufficient separation in timescale the only critical point attractors are local minmax. in the stochastic setting, the result requires time varying learning rates with a sufficient separation in timescale. Specifically, the players need to be using learning rate sequences for each such that (without loss of generality) not only is it assumed that , but also and for each . The challenge with these assumptions on the learning rate sequences is that empirically the sequences that satisfy them result in poor behavior along the learning path such as getting stuck at saddle points or making no progress. This is, in essence, due to the fact that the faster player—i.e., player 2 if —equilibriates too quickly causing progress to stall. This can result in undesirable behavior such as vanishing gradients (so that the discriminator does not provide enough information for the generator to make progress), mode collapse, or failure to converge in practical applications such as generative adversarial networks.
On the other hand, our convergence result gives a similar guarantee with less restrictive requirements on the stepsize sequence. In particular, only a single stepsize sequence is required (so that the algorithm can be viewed as a single timescale stochastic approximation update) as long as the fast player (who, without loss of generality, is player 2 in this paper) scales their estimated gradient by where is as in Theorem 3.1.
Furthermore, we note that in applications such as generative adversarial networks, while it has been observed that timescale separation heuristics such as unrolling or annealing the stepsize of the discriminator work well, in the stochastic case, summmable/square-summable assumptions on stepsizes are generally too restrictive in practice since they lead to a rapid decay in the stepsize which, in turn, can stall progress. On the other hand, stepsize sequences such as for —a sequence which satisfies the assumptions posed in (Thoppe and Borkar, 2019)—tend not to have this issue of decaying too rapidly for appropriately chosen , while also maintaining the guarantees of the theoretical results. We state a convergence guarantee under these relaxed assumptions in Proposition J.1 which is contained in Appendix J.
Regularization with Applications to Adversarial Learning
In this section, we focus on generative adversarial networks with regularization. Specifically, using the theory developed so far, we extend the results in Mescheder et al. (2018) to provide a convergence guarantee for a range of regularization parameters and learning rate ratios.
As has been repeatedly observed in recent theoretical works on generative adversarial networks, the training dynamics of generative adversarial networks are not well understood even though we have seen impressive practical advances over the last few years. In an attempt to address this, recent works—e.g., (Nagarajan and Kolter, 2017; Mescheder et al., 2017; Fiez et al., 2020; Berard et al., 2020) amongst others—study the optimization landscape of generative adversarial networks through the lens of dynamical systems theory which provides analysis tools for convergence based on the eigen-structure of the local linearization of the learning dynamics. Nagarajan and Kolter (2017) show, under suitable assumptions, that gradient-based methods for training generative adversarial networks are locally convergent assuming the data distributions are absolutely continuous. However, as observed by Mescheder et al. (2018), such assumptions not only may not be satisfied by many practical generative adversarial network training scenarios such as natural images, but it can often be the case that the data distribution is concentrated on a lower dimensional manifold. The latter characteristic leads to nearly purely imaginary eigenvalues and highly ill-condition problems.
Mescheder et al. (2018) provide an explanation for observed instabilities consequent of the true data distribution being concentrated on a lower dimensional manifold using discriminator gradients orthogonal to the tangent space of the data manifold. Further, the authors introduce regularization via gradient penalties that leads to convergence guarantees under less restrictive assumptions than were previously known. Here, we further extend these results to show that convergence to differential Stackelberg equilibria is guaranteed under a wide array of hyperparameter configurations.
Consider the training objective of the form
It is observed in Lemma 3.1 Mescheder et al. (2018) that GDA will not generally converge even with unrolling of the discriminator, a proxy for timescale separation. We observe, however, that this is because the equilibrium is not hyperbolic, and hence it is not structurally stable (Broer and Takens, 2010). In fact, introducing regularization remedies these issues. Under reasonable generative adversarial network assumptions, we show that the introduction of a gradient penalty based regularization to the discriminator does not change the set of critical points for the dynamics and, further, for any learning rate ratio and any positive, finite regularization parameter , the continuous time limiting regularized learning dynamics remain stable, and hence, there is a range of learning rates for which the discrete time update locally converges asymptotically.
Gradient penalties ensure that the discriminator cannot create a non-zero gradient which is orthogonal to the data manifold without suffering a loss. Introduced by Roth et al. (2017) and refined in Mescheder et al. (2018), we consider training generative adversarial networks with one of two fairly natural gradient-penalties used to regularize the discriminator:
Also following Mescheder et al. (2018), we use relaxed assumptions—as compared to the work by Nagarajan and Kolter (2017)—which allow us to consider generative adversarial networks with data distributions that do not (locally) have the same support and hence, are concentrated on lower dimensional manifolds, a commonly observed phenomena in practice (Arjovsky et al., 2017).
The Jacobian of the regularized dynamics, for either or , is of the form
It is straightforward to compute the block components of the Jacobian. Observe that Assumption 2.a implies that is identically zero, and hence is never a differential Nash equilibrium. However, we show that is not only a differential Stackelberg equilibrium, but also characterize the learning rate ratio and regularization parameter range for which is (locally) stable with respect to -GDA and give a convergence rate.
To proof that is a differential Stackelberg equilibrium follows analogous arguments to those in the proof of Theorem 4.1 in Mescheder et al. (2018). Given any positive regularization parameter , to prove the stability of for any fixed , we leverage the concept of the quadratic numerical range of a block operator which is a superset of the spectrum of the operator (cf. Appendix A). The key for both arguments is that Assumption 2 implies that the Jacobian of the regularized game has a specific structure. Indeed, observe that the structural form of is
The proof of the above corollary follows a similar line of reasoning as the proof of Corollary 4.1.
Theorem A.7 of Mescheder et al. (2018) shows that matrices of the form
are stable if is full rank and . The following proposition provides necessary conditions on the sizes of the network architectures for the discriminator and generator network for stability.
Experiments
We now present extensive numerical experiments examining gradient descent-ascent with timescale separation. As we explored theoretically so far, the stability of gradient descent-ascent critical points has an intricate relationship with timescale separation. We begin to investigate this behavior empirically by simulating the gradient descent-ascent dynamics for the games from Examples 1 and 2 and examining how the spectrum of the game Jacobian evolves as a function of the timescale separation. Then, on a polynomial game, we demonstrate how timescale separation warps the vector field of gradient descent-ascent and consequently shapes the region of attraction around critical points in the optimization landscape. There are a number of both qualitative and quantitative theoretical questions that remain open related to characterizing the region of attraction and how it depends parameterically on .
After exploring the optimization landscape, we focus in on gradient descent-ascent in the Dirac-GAN game and illustrate the interplay between timescale separation, regularization, and rate of convergence. Finally, we train generative adversarial networks on the CIFAR-10 and CelebA datasets with regularization and show timescale separation can significantly improve stability and performance. Moreover, we find that several of the insights we draw from the Dirac-GAN game carry over to this complex setting. Appendix L contains several more experimental results including a generative adversarial network formulation to learn a covariance matrix and a torus game. The code for our experiments is available at github.com/fiezt/Finite-Learning-Ratio.
We now revisit the game from Example 1 that demonstrated there exists differential Stackelberg equilibrium that are unstable for choices of the timescale separation . To be clear, we repeat the game construction and some characteristics of the game. Let us consider the quadratic zero-sum game defined by the cost
For this experiment, we select and simulate -GDA from the initial condition with and . In Figures 4a and 4b, we show the trajectories of the players coordinate pairs and , respectively. We observe that -GDA cycles around the equilibrium with since it is marginally stable with respect to the dynamics. For , the equilibrium is stable and -GDA ends up converging to it at a rate that depends on the choice of . We demonstrate how the convergence rate depends on the choice of in Figure 4c by showing the distance from the equilibrium along the learning path for each of the trajectories. The primary observation is that the cyclic behavior of -GDA dissipates as grows and as a result the dynamics then rapidly converge to the equilibrium.
The behavior of the learning dynamics as a function of the timescale separation can be further explained by evaluating the eigenvalues of the game Jacobian at the equilibrium. We show the eigenvalues of the Jacobian at the equilibrium in several forms in Figures 4e, 4f, and 4g. Analyzing the spectrum, we are able to verify that for all the equilibrium is indeed stable. Moreover, we see that the imaginary parts of the conjugate pairs of eigenvalues decay after and , and then the eigenvalues of the conjugate pairs eventually become purely real at and , respectively. After the eigenvalues of a conjugate pair become purely real, they split so that one of the eigenvalues asymptotically converges to an eigenvalue of by moving back along the real line, while the other eigenvalue tends toward an eigenvalue of . This occurrence is exactly what was described in Section 3 as an immediate implication of Proposition 3.2 when the eigenvalues of and are distinct. The convergence rate is in fact limited by the eigenvalues splitting since as grows, the spectrum of the Jacobian is limited by the eigenvalues of the Schur complement which remain constant. A related open question centers on finding the worst case convergence rate as a function of the spectral properties of and . Finally, the evolution of the eigenvalues as a function of the timescale separation demonstrates that the rotational dynamics in -GDA vanish as the ratio between the magnitude of the real and imaginary parts of the eigenvalues grows.
2 Polynomial Game: Timescale Separation and Non-Equilibrium Stability
We now return to the game from Example 2 that showed a non-equilibrium critical point which is stable without timescale separation and becomes unstable for a range of finite learning ratios with multiple equilibria in the vicinity. Again, we repeat the game construction along with some of the key characteristics that were previously presented in Example 2. Consider a zero-sum game defined by the cost
This game has critical points at , , and . Among the critical points, only and are game-theoretically meaningful equilibrium. In fact, they are each differential Nash equilibrium and are locally stable for any choice of as a result of Proposition 3.1. On the other hand, the critical point is neither a differential Nash equilibrium nor a differential Stackelberg equilibrium. However, is stable for and it is marginally stable for . In general, convergence to the non-equilibrium critical point in the presence of multiple game-theoretically meaningful equilibrium would be viewed as undesirable. In fact, this is precisely the type of critical point that sophisticated schemes for converging to only differential Nash equilibria or only differential Stackelberg equilibria seek to avoid (Adolphs et al., 2019; Mazumdar et al., 2019; Wang et al., 2020; Fiez et al., 2020). We show in this example that the simple inclusion of timescale separation in gradient descent-ascent is sufficient to avoid and instead converge to a differential Nash equilibrium.
3 Polynomial Game: Vector Field Warping and Region of Attraction
Consider a zero-sum game defined by the cost
The cost structure of this game is visualized in Figure 6a, where we present a three dimensional view of along with the cost contours and the locations of critical points. This game has eleven critical points including one differential Nash equilibrium and two differential Stackelberg equilibria that are not a differential Nash equilibrium. The critical points that are neither a differential Nash equilibrium nor a differential Stackelberg equilibrium are unstable for any choice of timescale separation . The differential Nash equilibrium is at and it is stable for all by Proposition 3.1. The differential Stackelberg equilibria are at and ; each is stable for all . We computed for the pair of differential Stackelberg equilibrium using the theoretical construction from Theorem 3.1 and observed that it properly recovered for each equilibrium as the timescale separation such that the continuous time system is stable for all . Finally, we note that while the set of equilibrium follow a linear translation, this game is generic and the equilibria are in fact isolated.
In Figure 6b, we show the trajectories of -GDA with and given the initialization near the differential Stackelberg equilibrium at . Moreover, in Figure 7a, we overlay the trajectories on the vector field generated by the respective timescale separation parameters. As expected, the choice of results in a trajectory that cycles around the equilibrium in a closed curve since it is marginally stable and has purely imaginary eigenvalues. Notably, as grows, the cyclic behavior dissipates as the timescale separation reshapes the vector field until the trajectory moves near directly to the zero derivative line of the maximizing player and then follows a path along that line toward the equilibrium and converges rapidly. The eigenvalues of as a function of are presented in Figures 6c and 6d. As was the case for the previous experiments, we observe that after the eigenvalues become purely real as grows, they then split and asymptotically converge toward the eigenvalues of and . It is worth noting that much of the rotational behavior in the dynamics and vector field disappears as a result of timescale separation well before the eigenvalues become purely real; this seems to occur after the timescale separation is such that the magnitude of the real part of the eigenvalues is greater than that of the imaginary part.
Finally, in Figure 7b, we demonstrate how the choice of timescale separation not only warps the vector field but also shapes the regions of attraction around critical points. The vector field is again shown for each , but now zoomed out to include each of the equilibria. The colors overlayed on the vector field indicate the equilibria that the dynamics converge to given an initialization at that position. Positions in the strategy space without color did not converge to an equilibrium in the fixed horizon of 75000 iterations with . This is explained by the fact that the dynamics are not guaranteed to be globally convergent and may get stuck in limit cycles or may simply move slowly for a long time in flat regions of the optimization landscape. We produced this experiment by running -GDA for a dense set of initial conditions chosen uniformly over the space of interest. It is clear from the experiment that the choice of timescale separation determines not only the stability of equilibria, but also has a fundamental impact on the equilibria the dynamics converge to from a given initial condition as a result of the warping of the vector field. As a concrete example, given an initialization of , the dynamics with converge to the differential Nash equilibria at . However, for any , the dynamics instead converge to the differential Stackelberg equilibrium at that is significantly closer to the initial condition. This example motivates future work on methods for obtaining accurate estimates of the regions of attraction around critical points and techniques to design in order to explicitly shape the region of attraction around an equilibrium of interest. We refer to the end of Section 4.1 for further discussion on potentially relevant analysis methods in this direction.
4 Dirac-GAN: Regularization, Timescale Separation, and Convergence Rate
In Section 5, we studied gradient descent-ascent with regularization in generative adversarial networks and showed that the general theory we provide can be extended to such a formulation. Recall that the training objective for generative adversarial networks is of the form
The unique critical point of gradient descent-ascent is a local Nash equilibrium given by . However, the structure of the game is such that
Mescheder et al. (2018) proposed to remedy the degeneracy issues of generative adversarial networks by using the following gradient penalties to regularize the discriminator with :
The zero-sum game corresponding to the Dirac-GAN with regularization can be defined by the cost
The unique critical point of the game remains at , but we now see that
From Figures 8a and 8d, we observe that the impact of timescale separation with regularization is that the trajectory is not as oscillatory since it moves faster to the zero line of and then follows along that line until reaching the equilibrium. We further see from Figure 8b that with regularization , -GDA with converges faster to the equilibrium than -GDA with , despite the fact that the former exhibits some cyclic behavior in the dynamics while the latter does not. The eigenvalues of the Jacobian with regularization presented in Figure 8e explains this behavior since the imaginary parts are non-zero with and zero with , but the eigenvalue with the minimum real part is greater at than at . This example highlights that a degree of oscillatory behavior in the dynamics is not always harmful for convergence and it can even speed up the rate of convergence. Building off of this, for regularization and timescale separation , Figures 8a and 8b show that even though -GDA follows a direct path toward the equilibrium and does not cycle since the eigenvalues of the Jacobian are purely real, the trajectory converges slowly to the equilibrium. While not presented, we ran experiments with with as well and timescale separation only made the convergence rate worse. The eigenvalues of the Jacobian with each regularization parameter presented in Figures 8e and 8f are able to explain this phenomenon. Indeed, for each regularization parameter, the eigenvalues split after becoming purely real and then converge toward the eigenvalues of and . Since and , there is a trade-off between the choice of regularization and the timescale separation on the conditioning of the Jacobian matrix. As shown in Figures 8e and 8f, the minimum real part of the eigenvalues with is significantly larger than with after sufficient timescale separation, which makes the convergence rate faster. Together, this example demonstrates that there may often be a delicate relationship between timescale separation, regularization, and convergence rate, where after a certain threshold each parameter choice may inhibit the rate of convergence.
5 Generative Adversarial Networks: Image Datasets
We now investigate the role timescale separation has on training generative adversarial networks parameterized by deep neural networks. The empirical benefits of training with a timescale separation have been documented previously. For example, Heusel et al. (2017) showed on a number of image datasets that a timescale separation between the generator and discriminator improves generation performance as measured by the Frechet Inception Distance (FID). Since then a significant number of papers have presented results training generative adversarial networks with timescale separation. Moreover, it is common in the literature for the discriminator to be updated multiple times between each update of the generator (Arjovsky et al., 2017). Indeed, it has been widely demonstrated that this heuristic improves the stability and convergence of the training process and locally it has a similar effect as including a timescale separation between the generator and discriminator. The disadvantage of this approach is that the number of gradient calls per generator update increases and consequently the convergence is then slower in terms of wall clock time when a similar effect could potentially be achieved by a learning rate separation between the generator and discriminator. We remark that it appears to be reasonably common for practitioners to fix a shared learning rate for the generator and discriminator along with a pre-selected number of discriminator updates per generator update and not thoroughly investigate the impact timescale separation has on the training process.
The goal of our generative adversarial network experiments is to reinforce the importance of the timescale separation between the generator and the discriminator as a hyperparameter in the training process, demonstrate how it changes the behavior along the learning path, and show that it is compatible with a number of common training heuristics. This is to say that our goal is not necessarily to show state-of-the art performance, but rather to perform experiments that allow us to make insights relevant to the theory in this paper. We remark that our empirical work on training generative adversarial networks is distinct from and complimentary to that of Heusel et al. (2017) in several ways. The theory given by Heusel et al. (2017) only applies to stochastic stepsizes, however in the experiments they implemented constant step sizes. We train with mini-batches and decaying stepsizes, which does satisfy the theory we provide. Moreover, by and large, the experiments by Heusel et al. (2017) compare a fixed learning rate ratio between the generator and discriminator to multiple fixed shared learning rates for the generator and discriminator. In contrast, we fix a learning rate for the generator and explore the behavior of the training process as the timescale parameter is swept over a given range.
We build our experiments based on the methods and implementations of Mescheder et al. (2018) and explore both the CIFAR-10 and CelebA image datasets. We train the generative adversarial networks with the non-saturating objective function and the gradient penalty proposed by Mescheder et al. (2018) with regularization parameters . We note that the non-saturating objective results in a game that is not zero-sum, however it is commonly used in practice and under the realizable assumptions is it locally equivalent to the zero-sum objective as discussed at the end of Section 6.4. The network architectures for the generator and discriminator are both based on the ResNet architecture. The initial learning rate for the generator in all of our experiments is fixed to be and we decay the stepsizes so that at update the learning rate of the generator is given by where and the learning rate of the discriminator is . For each experiment the batch size is , the latent data is drawn from a standard normal distribution of dimension , and the resolution of the images is . Finally, as an optimizer, we run RMSprop with parameter . Again, the theory we provide does not strictly apply to using RMSprop, but it is ubiquitous in practice for training generative adversarial networks and if the timescale separation is sufficiently large so that the eigenvalues are purely real in the Jacobian then the theory we provide is applicable as remarked previously. We provide further details on the network and hyperparameters in Appendix L.4. A final heuristic and hyperparameter that we explore in conjunction with the timescale separation is that of using an exponential moving average to produce the model that is evaluated. This means that at each update , given that the parameters of the generator are given by , the moving average is kept where . Experimental studies have shown that this heuristic can yield a significant improvement in terms of both the inception score and the FID (Yazici et al., 2019; Gidel et al., 2019a). The success of this method is thought to be a result of dampening both rotational dynamics and the noise from the randomness in the mini-batches of data.
We run the training algorithm with the learning rate ratio belonging to the set and the regularization parameter belonging to the set . For each choice of and , we retain exponential moving averages of the generator parameters for . The training process is repeated times for each hyperparameter configuration to rule out noise from random seeds and the performance is evaluated along the learning path at every 10,000 updates in terms of the FID score. We report the mean scores and the standard error of the mean over the repeated experiments. We run the experiments with for 150k mini-batch updates and the experiments with for 300k mini-batch updates.
The results for each dataset across the hyperparameter configurations are presented in numeric form in Figure 12. Figure 11 shows some generated samples selected at random for each dataset with the hyperparameter configuration that performed best in terms of the FID score at the end of the training process. Figure 17 in Appendix L.4 has several more generated samples for each dataset selected at random. We now describe the key observations from the experiments for each dataset.
The FID scores along the learning path for CIFAR-10 with and are presented in Figures 9a and 9b, respectively. The corresponding scores in numeric form are given in Figures 12a, 12c, and 12e for at 150k iterations and at 150k and 300k iterations, respectively. To begin, we observe that the exponential moving average significantly improves performance, and of the parameters considered, performed best. Relevant to this work, we observe that the performance gain from using an exponential moving average appears to be maximized when the ratio of learning rates is smallest. This may indicate that some of the dynamics in -GDA are dampened by timescale separation in this generative adversarial network experiment, similarly to as observed for the simpler experiments presented previously. Moreover, we that timescale separation also has a significant impact on the FID score of the training process. Indeed, even selecting versus can yield an impressive performance gain. In this experiment for each regularization parameter, converges fastest, followed by , then , and finally . Finally, observe that the performance with regularization is far superior to that with regularization . Interestingly, the last pair of conclusions are in line with the insights drawn from the simple Dirac-GAN experiment in Section 6.4. In particular, timescale separation only speeds up to convergence until hitting a limiting value and there is a fundamental interplay between timescale separation, regularization, and convergence rate. This indicates that it may be possible to transfer some of the insights made on simple generative adversarial network formulations to the much more complex problem where players are parameterized by neural networks.
The FID scores along the learning path for CIFAR-10 with and are presented in Figures 10a and 10b, respectively. The corresponding scores in numeric form are given in Figures 12b, 12d, and 12f for at 150k iterations and at 150k and 300k iterations, respectively. In this experiment we that while the exponential moving average helps performance, the gain is not as drastic as it was for CIFAR-10. However, timescale separation in combination with the regularization does have a major effect on the the FID score of the training process in this experiment. For regularization , the timescale parameters of and outperform and by a wide margin, again highlighting that timescale separation can speed up convergence until a certain point where it can potentially slow it down owing to the effect on the conditioning of the problem locally. A similar trend can be observed with regularization , but with performing closer to and . We again observe in this experiment that for all timescale separation parameters, the performance is significantly improved with regularization as compared with . This once again highlights the importance of considering how this the hyperparameters of regularization and timescale interact and dictate the local convergence rates.
In summary, we took a well-performing method and implementation for training generative adversarial networks and demonstrated that timescale separation is an extremely important, and easy to implement, hyperparameter that is worth careful consideration since it can have a major impact on the convergence speed and final performance of the training process.
Related Work
In this section, we provide a review of related work at the intersection machine learning and game theory, as well as connections to dynamical systems theory and control.
Given the extensive work on the topic of learning in games in machine learning that has gone on over the last several years, we cannot cover all of it and instead focus our attention on only the most relevant to this paper. We begin by reviewing solution concepts developed for the class of games under consideration and then discuss some learning dynamics studied in the literature beyond gradient descent-ascent. Following this, we delineate the related work studying gradient descent-ascent in non-convex, non-concave zero-sum games and finish by making note of the literature on non-convex, concave optimization.
Owing to the numerous applications in machine learning, a significant portion of the modern work on learning in games has focused on the zero-sum formulation with non-convex, non-concave cost functions. Most recently, Daskalakis et al. (2020) tout the importance and significance of this class of games in a paper on the complexity of finding equilibria (in particular, in the constrained setting) in such games. Consequently, local solution concepts have been broadly adopted. Compared to the standard game-theoretic notions of equilibrium that characterize player’s incentive to deviate given the game and information structure, local equilibrium concepts restrict the deviation search space to a suitable local neighborhood. Following the standard game-theoretic viewpoint, a vast number of works in machine learning study the local Nash equilibrium concept and critical points satisfying gradient-based sufficient conditions for the equilibrium, which are often referred to as differential Nash equilibria (Ratliff et al., 2013, 2014, 2016). Based on the observation that in non-convex, non-concave zero-sum games the order of play is fundamental in the definition of the game, there has been a push toward considering local notions of the Stackelberg equilibrium concept, which is the usual game-theoretic equilibrium when there is an explicit order of play between players. In the zero-sum formulation, Stackelberg equilibrium are often referred to as minmax equilibria. Similar to as for the Nash equilibrium, gradient-based sufficient conditions for local minmax/Stackelberg equilibrium have been given (Fiez et al., 2020; Jin et al., 2020; Daskalakis and Panageas, 2018) and such critical points have been referred to as differential Stackelberg equilibria (Fiez et al., 2020). We remark that it has been shown that local/differential Nash equilibria are a subset of local/differential Stackelberg equilibria (Jin et al., 2020; Fiez et al., 2020). Following past works, we adopt the terminology of differential Nash equilibrium and differential Stackelberg equilibrium in this paper as the meaning of strict local Nash equilibrium and strict local minmax/Stackelberg equilibrium, respectively. Finally we mention the proximal equilibria proposed by Farnia and Ozdaglar (2020), which we do not consider in this work, that depending on a regularization parameter can interpolate between the local Nash and local Stackelberg equilibrium notions.
Given that the focus of this work is on gradient descent-ascent, we center our coverage of related work on papers analyzing its behavior. Nonetheless, we mention that a significant number of learning dynamics for zero-sum games have been developed in the past few years, in some cases motivated by the shortcomings of gradient descent-ascent without timescale separation. The methods include optimistic and extra-gradient algorithms (Daskalakis et al., 2018; Gidel et al., 2019a; Mertikopoulos et al., 2019), negative momentum (Gidel et al., 2019b), gradient adjustments (Mescheder et al., 2017; Balduzzi et al., 2018; Letcher et al., 2019a), and opponent modeling methods (Zhang and Lesser, 2010; Metz et al., 2017; Foerster et al., 2018; Letcher et al., 2019b; Schäfer and Anandkumar, 2019), among others. While the aforementioned learning dynamics possess some desirable characteristics, they cannot guarantee that the set of stable critical points coincide with a set of local equilibria for the class of games under consideration. However, there have been a select few learning dynamics proposed that can guarantee the stable critical points coincide with either the set of differential Nash equilibria (Adolphs et al., 2019; Mazumdar et al., 2019) or the set of differential Stackelberg equilibria (Fiez et al., 2020; Wang et al., 2020)—effectively solving the problem of guaranteeing local convergence to only a class of local equilibria. However, since each of the algorithms achieving the equilibrium stability guarantee require solving a linear equation in each update step, they are not efficient and can potentially suffer from degeneracies along the learning path in applications such as generative adversarial networks. These practical shortcomings motivate either proving that existing learning dynamics using only first-order gradient feedback achieve analogous theoretical guarantees or developing novel computationally efficient learning dynamics that can match the theoretical guarantee of interest.
Gradient descent-ascent has been studied extensively in non-convex, non-concave zero-sum games since it is a natural analogue to gradient descent from optimization, is computationally efficient, and has been shown to be effective in practice for applications of interest when combined with common heuristics. A prevailing approach toward gaining understanding of the convergence characteristics of gradient descent-ascent has been to analyze the local stability around critical points of the continuous time limiting dynamical system. The majority of this work has not considered the impact of timescale separation. Numerous papers have pointed out that the stable critical points of gradient descent-ascent without timescale separation may not be game-theoretically meaningful. In particular, it has been shown that there can exist stable critical points that are not differential Nash equilibrium (Daskalakis and Panageas, 2018; Mazumdar et al., 2020). Furthermore, it is known that there can exist stable critical points that are not differential Stackelberg equilibria (Jin et al., 2020). The aforementioned results rule out the possibility that gradient descent-ascent without timescale separation can guarantee equilibrium convergence. In terms of the stability of equilibria, it is known that differential Nash equilibrium are stable for gradient descent-ascent without timescale separation (Daskalakis and Panageas, 2018; Mazumdar et al., 2020), but that there can exist differential Stackelberg equilibria which are not stable with respect to gradient descent-ascent without timescale separation.
The work of Jin et al. (2020) is the most relevant exploring how the aforementioned stability properties of gradient descent-ascent change with timescale separation. In particular, Jin et al. (2020) investigate whether the desirable stability characteristics (stability of differential Nash equilibria) and undesirable stability characteristics (stability of non-equilibrium critical points and instability of differential Stackelberg equilibria) of gradient descent without timescale separation are maintained and remedied, respectively with timescale separation. In terms of the former query, extending the examples shown in Mazumdar et al. (2020) and Daskalakis and Panageas (2018), Jin et al. (2020) show that differential Nash equilibrium are stable for gradient descent-ascent with any amount of timescale separation.
On the other hand, for the latter query, Jin et al. (2020) shows (in Proposition 27) two interesting examples: (a) for an a priori fixed , there exists a game with a differential Stackelberg equilibrium that is not stable and (b) for an a priori fixed , there exists a game with a stable critical point that is not a differential Stackelberg equilibrium. However, (a) does not imply that for the constructed game, there does not exist another (finite) —independent of the game parameters—such the differential Stackelebrg equilibrium is stable for all larger . In simple language, the result summarized in (a) says the following: if a bad timescale separation is chosen, then convergence may not be guaranteed. Similarly, (b) does not imply that there is no such that for all larger for the constructed game instance, the critical point becomes unstable. Again, in simple language, the result summarized in (b) says the following: if a bad timescale separation is chosen, then non-game theoretically meaningful equilibria may persist. While at first glance this set of results may appear to indicate that the undesirable stability characteristics of gradient descent without timescale separation cannot be averted by any finite timescale separation, it is important to emphasize that these results do not answer the questions of whether there (a) exists a game with a critical point that is not a differential Stackelberg equilibrium which is stable with respect to gradient descent-ascent without timescale separation and remains stable for all finite timescale separation ratios or (b) exists a game with a differential Stackelberg equilibrium that is not stable for all finite timescale separation ratios. The preceding questions are left open from previous work and are exactly the focus of this paper. In Appendix K, we go into greater detail on the comparison between Proposition 27 of Jin et al. (2020) as we believe this to be an important point of distinction between Theorem 3.1 and 3.2 in this paper.
Finally, Jin et al. (2020) study the an infinite timescale separation ratio and show that the stable critical points of gradient descent-ascent coincide with the set of differential Stackelberg equilibria in this regime. This result effectively shows that gradient descent-ascent can guarantee only equilibrium convergence with timescale separation, albeit infinite. We remark that an equivalent result in the context of general singularly perturbed systems has been known in the literature (Kokotovic et al., 1986, Chap. 2) as we discuss further in Section 3.1. Finally, we point out that since an infinite timescale separation does not result in an implementable algorithm, fully understanding the behavior with a finite timescale separation is of fundamental importance and the motivation for our work.
Beyond the work of Jin et al. (2020) considering timescale separation in gradient descent-ascent, it is worth mentioning the work of Chasnov et al. (2019) and Heusel et al. (2017). Chasnov et al. (2019) study the impact of timescale separation on gradient descent-ascent, but focus on the convergence rate as a function of it given an initialize around a differential Nash equilibrium and do not consider the stability questions examined in this paper. Heusel et al. (2017) study stochastic gradient descent-ascent with timescale separation and invoke the results of Borkar (2008) for analysis. The stochastic approximation results the claims rely on guarantee the convergence of the system locally to a stable critical point. Consequently, the claim of convergence to differential Nash equilibria of stochastic gradient descent-ascent given by Heusel et al. (2017) only holds given an initialization in a local neighborhood around a differential Nash equilibrium. In this regard, the issue of the local stability of the types of critical point is effectively assumed away and not considered. In contrast, we are able to combine our stability results for gradient descent-ascent with timescale separation together with the stochastic approximation theory of Borkar (2008) to guarantee local convergence to a differential Stackelberg equilibrium in Section 4.2. We remark that Heusel et al. (2017) empirically demonstrate that timescale separation can significantly improve the performance of gradient descent-ascent when training generative adversarial networks.
The final relevant line of work studying gradient descent-ascent is specific to generative adversarial networks. The results from this literature develop assumptions relevant to generative adversarial networks and then analyze the stability and convergence properties of gradient descent-ascent under them (see, e.g., works by Metz et al. (2017); Goodfellow et al. (2014); Daskalakis et al. (2018); Nagarajan and Kolter (2017); Mescheder et al. (2018)). Within this body of work, there has been a significant amount of effort focusing on how the stability (and, hence, convergence properties) of gradient descent-ascent in generative adversarial networks can be enhanced with regularization methods. Nagarajan and Kolter (2017) show, under suitable assumptions, that gradient-based methods for training generative adversarial networks are locally convergent assuming the data distributions are absolutely continuous. However, as observed by Mescheder et al. (2018), such assumptions not only may not be satisfied by many practical generative adversarial network training scenarios such as natural images, but it can often be the case that the data distribution is concentrated on a lower dimensional manifold. The latter characteristic leads to nearly purely imaginary eigenvalues and highly ill-condition problems. Mescheder et al. (2018) provide an explanation for observed instabilities consequent of the true data distribution being concentrated on a lower dimensional manifold using discriminator gradients orthogonal to the tangent space of the data manifold. Further, the authors introduce regularization via gradient penalties that leads to convergence guarantees under less restrictive assumptions than were previously known. In this paper, we further extend these results to show that convergence to differential Stackelberg equilibria is guaranteed under a wide array of hyperparameter configurations (i.e., learning rate ratio and regularization).
2 Historical Perspective: Dynamical Systems and Control
The study of gradient descent-ascent dynamics with timescale separation between the minimizing and maximizing players is closely related to that of singularly perturbed dynamical systems (Kokotovic et al., 1986). Such systems arise in classical control and dynamical systems in the context of physical systems that either have multiple states which evolve on different timescales due to some underlying immutable physical process or property, or a single dynamical system which evolves on a sub-manifold of the larger state-space. For example, robot manipulators or end effectors often have have slower mechanical dynamics than electrical dynamics. On the other hand, in electrical circuits or mechanical systems, certain resistor-capacitor circuits or spring-mass systems have a state which evolves subject to a constraint equation (Sastry and Desoer, 1981; Lagerstrom and Casten, 1972). Due to their prevalence, singularly perturbed systems have been studied extensively with one of the outcomes being a number of works on determining the range of perturbation parameters for which the overall system is stable (Kokotovic et al., 1986; Saydy et al., 1990; Saydy, 1996). We exploit these results and analysis techniques to develop novel results for learning in games. One of contributions of this work is the introduction of the algebraic analysis techniques to the machine learning and game theory communities. These tools open up new avenues for algorithm synthesis; we comment on potential directions in the concluding discussion section.
This being said, there are a couple key difference between the present setting and that of the classical literature including the following:
The perturbation parameter is no longer an immutable characteristic of the physical system, but rather a hyperparameter subject to design. Indeed, in singular perturbation theory, the typical dynamical system studied takes the form
where the –player seeks to minimize with respect to and the –player seeks to maximize with respect to , and is the ratio of learning rates (without loss of generality) of the maximizing to the minimizing player. These learning rates—and hence the value of —are hyperparameters subject to design in most machine learning and optimization applications. Another feature of (27) as compared to (26), is that the dynamics are partial derivatives of a function , which leads to the second key difference.
There is structure in the dynamical system that arises from gradient-play which reflects the underlying game theoretic interactions between players. This structure can be exploited in obtaining convergence guarantees in machine learning and optimization applications of game theory. For instance, minmax optimization is analogous to a zero sum game for which the local linearization of gradient descent-ascent dynamics has the structure
where and and is the learning rate ratio or timescale separation parameter. Such block matrices have very interesting properties. In particular, second order optimality conditions for a minmax equilibrium correspond to positive definiteness of the first Schur complement , and of (Fiez et al., 2020). This turns out to be keenly important for understanding convergence of gradient descent-ascent. Furthermore, due to the structure of , tools from the theory of block operators (see, e.g., works by Tretter (2008); Magnus (1988); Lancaster and Tismenetsky (1985)) such as the quadratic numerical range can be exploited (and combined with singular perturbation theory) to understand the effects of hyperparameters such as (the learning rate ratio) and regularization (which is common in applications such as generative adversarial networks) on convergence.
Discussion
In this paper, we prove a necessary and sufficient condition for the convergence of gradient descent-ascent with timescale separation to differential Stackelberg equilibria in zero sum games. This answers a long standing open question about provable convergence of first order methods for zero-sum games to local minimax equilibria. Specifically, we provide necessary and sufficient conditions for the convergence of -GDA to differential Stackelberg equilibria. A key component of the proof is the construction of a (tight) finite lower bound on the learning rate ratio for which stability of the game Jacobian is guaranteed, and hence local asymptotic convergence of -GDA. In addition, we provide results on iteration complexity and convergence rate and apply the results to generative adversarial networks under mild assumptions on the data distribution. For both differential Nash equilibira and the superset of differential Stackelberg equilibria, we provide estimates on the neighborhood on which convergence is guaranteed.
This being said, the question of the size of the region of attraction remains open. As commented on earlier in the paper, an alternative but related technique tackles the nonlinear system directly. The downside of this technique is that one needs to have in hand (or be able to construct) Lyapunov functions for both the boundary layer model (i.e., the system that arises from treating the choice variable of the slow player as being ‘static’) and the reduced order model (i.e., the system that arises from plugging in the implicit mapping from the fast player’s action to the slow player’s action into the slow player’s dynamics). A convex combination of these functions provides a Lyapunov function for the original system . The level sets of this combined Lyapunov function then give a better sense of the region of attraction and, in fact, one can optimize over the weighting in the convex combination in order to obtain better estimates of the region of attraction. This is an interesting avenue to explore in the context of learning in games with lots of intrinsic structure that can potentially be exploited to improve both the rate of convergence and the region on which convergence is guaranteed.
Another significant contribution of this work is the fact that we introduce tools that are arguably new to the machine learning and optimization communities and expose interesting new directions of research. In particular, the notion of a guard map, which is arguably even an obscure tool in modern control and dynamical systems theory, is ‘rediscovered’ in this paper. The is potential to leverage this concept in not only providing certificates for performance (e.g., beyond stability to robustness) but also in synthesizing algorithms with performance guarantees. For instance, one observation from our empirical analysis is that convergence rate is not only limited by the eigenvalues of the Schur complement of the Jacobian, but the fastest convergence appears to occur when there are complex components of the eigenvalues. In short, some cycling is beneficial. Better understanding this fact from a theoretical perspective is an open question, as is optimizing the rate of convergence by exploiting these observations.
Finally, another set of related open questions center on practical considerations for the efficient use of first order methods. For instance, with respect to generative adversarial networks, the exponential moving average is known to empirically reduce the negative effects of cycling. Additionally, increasing the learning rate ratio does lead to predominantly real eigenvalues which in turn reduces cycling. Understanding the trade offs between not only these two hyperparameters but also regularization is very important for practical implementations. Empirically, we study the tradeoffs between the learning rate ratio, regularization parameter, and the parameter controlling the degree of “smoothness” in the exponential moving average, another common heuristic that performs well in practice. There is an open line of research related to analytically characterizing the tradeoffs between these three hyperparameters. However, in the absence of theoretical tools for exploring these issues, what are reasonable and principled heuristics?
To conclude, while we arguably definitively address a standing open question for first order methods for learning in zero-sum games/minmax optimization problems, there a many open directions exposed by the tools introduced and empirical observations discovered in this work.
Acknowledgements
This work is funded by the Office of Naval Research (YIP Award) and National Science Foundation CAREER Award (CNS-1844729). Tanner Fiez is also funded by a National Defense Science and Engineering Graduate Fellowship. We thank Daniel Calderone for the helpful discussions, in particular on linear algebra results as they pertain to the results in this paper. Finally, we thank Mescheder et al. (2018) for providing a high quality open source implementation of the generative adversarial network experiments they performed, which facilitated and expedited the experiments we performed.
References
Appendix A Helper Lemmas and Additional Mathematical Preliminaries
In this appendix, we present a handful of technical lemmas and review some additional mathematical preliminaries excluded from the main body but which are important in proving the results in the paper.
The following technical lemma is used in proving an upper bound on the spectral radius of the linearization of the discrete time update -GDA a requirement for obtaining the convergence rate results.
The function satisfies for all .
Since and , we simply need to show that on to get that is a decreasing function on . Indeed, since for all . ∎
The following proposition is a well-known result in numerical analysis and can be found in a number of books and papers on the subject. Essentially, it provides an asymptotic convergence guarantee for a discrete time update process or dynamical system.
Let be a fixed point for the discrete dynamical system . If the spectral radius of the Jacobian satisfies , then is a contraction at and hence, is asymptotically stable.
The following technical lemma, due to Mustafa and Davidson (1994), is used in constructing the finite learning rate ratio.
For completeness (and because there is a typo in the original manuscript), we provide the proof here.
Suppose that and are non-singular so that the partial Schur decomposition
Applying the determinant operator, we have that
Combining (28) with (31) in (29) gives exactly the claimed result. ∎
The following lemma is Theorem 2 Lancaster and Tismenetsky (1985, Chap. 13.1). We use this lemma several times in the proofs of Theorem 3.1 and 3.2 so we include it here for ease of reference. For a given matrix , , , and are the number of eigenvalues of the argument that have positive, negative and zero real parts, respectively.
If is a symmetric matrix such that where , then is nonsingular and and have the same inertia—i.e.,
On the other hand, if , then there exists a matrix and a matrix such that and and have the same inertia (i.e., (32) holds).
The quadratic numerical range of is defined by
where denotes the spectrum of its argument.
The quadratic numerical range can be described as the set of solutions of the characteristic polynomial
for and . We use the notation to denote the inner product. Note that is a (potentially non-convex) subset of and contains .
Appendix B Proof of Proposition 3.1
Before proving this result, we note that the result has already been shown in the literature by Jin et al. (2020). We included the result primarily because the proof approach is different and the tools we use (in particular, the quadratic numerical range) have not been utilized before in this type of analysis. Hence, we view the proof technique itself as a contribution.
Then, the elements of are of the form
where , and for vectors and .
We claim that for any , for all ,, and where and since is a differential Nash equilibrium.
Indeed, we argue this by considering the two possible cases: (1) or (2) .
Case 1: Suppose is such that . Then, trivially since .
Case 2: Suppose is such that . In this case, we want to ensure that
The last inequality is equivalent to . Indeed,
Moreover, holds for any pair of vectors such that and since and .
Appendix C Proof of Lemma 4.1 and Lemma 4.2
In this appendix section, we prove Lemma 4.1 and Lemma 4.2 from Section 4. We note that the proof of Lemma 4.2 starts where the proof of Lemma 4.1 leaves off.
as claimed. From this argument, it is clear that for any , for all .
Hence, the so that an application of Proposition A.1 gives us the desired result.
C.2 Proof of Lemma 4.2
To prove this lemma, we build directly on the conclusion of the proof of Lemma 4.1. Indeed, since
given there exists a norm (cf. Lemma 5.6.10 in Horn and Johnson (1985))The norm that exists can easily be constructed as essentially a weighted induced -norm. Note that the norm construction is not unique. The proof in Horn and Johnson (1985) is by construction and the construction of this norm can be found there. such that
where the last inequality holds by Lemma A.1. Taking the Taylor expansion of around , we have
where is the remainder term satisfying as .The notation as means . This implies that there is a such that whenever . Hence,
where the last inequality holds again by Lemma A.1. Hence,
whenever which verifies the claimed convergence rate.
Appendix D Proof of Corollary 4.2
Let be the norm that exists (via construction a la Horn and Johnson (1985, Lem. 5.6.10)) in the proof of Lemma 4.2 which is given in Appendix C. Following standard arguments, (35) in the proof of Lemma 4.2 implies a finite time convergence guarantee. Indeed, let be given. Since we have that . Hence,
In turn, this implies that , meaning that is a -differential Stackelberg equilibrium for all whenever .
we have that the such that is .
Appendix E Proof of Theorem 3.1 and Corollary 4.1
Towards this end, we need to introduced some notation as well as formal definitions for important concepts such as the guard map.
Given a square matrix , let be the largest positive real eigenvalue of and if does not have a positive real eigenvalue then it is zero.
The use of guardian maps for studying stability of parameterized families of dynamical systems was arguably introduced by Saydy et al. (1990). Guardian or guard maps act as a certificate for a performance criteria such as stability.
Formally, let be the set of all real matrices or the set of all polynomials of degree with real coefficients. Consider an open subset of with closure and boundary .
The following result gives a necessary and sufficient condition for stability of parameterized families of matrices relative to some open subset of the complex plane.
from which it can be shown that the eigenvalues of are for where for are the eigenvalues of .
Indeed, let be a non-singular matrix such that where is upper triangular with on its diagonal. Observe that for any matrix , and . Hence, using properties of the Kronecker product (namely, that ), we have that
so that the spectrum of and coincide. Now, since is upper triangular, is upper triangular with diagonal elements () which can be verified by direct computation and using the definition of . This implies that () are exactly the eigenvalues of . ∎
We note that there are several other guard maps for the space of Hurwtiz stable matrices including . To give some intuition for this map, it is fairly straightforward to see that the Kronecker sum has spectrum where . The operator is simply a more computationally efficient expression of , and as such the eigenvalues of are those of removing redundancies. We use specifically because of its computational advantages in computing .
E.2 Proof of Theorem 3.1
Then we prove the other direction. That is, if there exists a finite such that for all , is exponentially stable for , then is a differential Stackelberg equilibrium. We prove this by contradiction.
Towards this end, for a critical point , let
Note that this is equivalent to the first Schur complement of (i.e., when ) since the and cancel, and by assumption the first Schur complement of is positive definite. Suppose that is a differential Stackelberg equilibrium so that and .
In particular, if we envision as the input to and simply vary (holding all the entries of otherwise fixed), then can be thought of simply as a function of which guards the set of Hurwitz stable matrices via the reasoning describe above. Indeed, by slightly overloading the notation for ,
Hence, for intuition, observe that as decreases (towards zero) stability is first lost when at least one eigenvalue of reaches the imaginary axis, at which point .
Recall that we have assumed that is a differential Stackelberg equilibrium (i.e., and ). We will show next (by way of explicit construction of ) that we are always in case 2.
We note that there are more elegant, simpler constructions, but to our knowledge this construction gives the tightest bound on the range of for which is guarnateed to be Hurwitz stable. Recall that
Let denote the identity matrix.
The finite learning rate ratio is where
with and .
We apply basic properties of the Kronecker product and sum as well as Schur’s determinant formula to obtain a reduced form of the guard map. To this end, we have that
Now, we apply Schur’s determinant formula to get that
From here, we apply Lemma A.2 to further reduce the guard map. First, note that
Let , , , , and . Using the two properties of the Kronecker product and , we have that
where (40) holds since . Now, define where
The assumptions that and together imply that and . Hence, if and only if since . The determinant expression is exactly an eigenvalue problem.
Since by assumption the Schur complement of and the individual Hessian are positive definite (i.e., is a differential Stackelberg equilibrium), Thus, the largest positive real root of is
where is the largest positive real eigenvalue of its argument if one exists and otherwise its zero. Using properties of the Kronecker product and duplication matrices, it can easily be seen that . ∎
The proof of this direction is argued by contradiction. Consider a critical point (i.e., where such that and have no zero eigenvalues—that is, and .
Since and , by Lemma A.3.b, there exists non-singular Hermitian matrices and positive definite Hermitian matrices such that and . Further, and have the same inertia, meaning
where for a given matrix , , , and are the number of eigenvalues of the argument that have positive, negative and zero real parts, respectively. Similarly, and have the same inertia:
Since has at least one strictly positive eigenvalue, .
which can be verified by straightforward calculations.
Observe that is equivalent to and both matrices are symmetric so that if and only if and where
Now, is also a real symmetric matrix, and hence, it is positive definite if and only if all its eigenvalues are positive. To determine the range of such that is positive definite, we can formulate an eigenvalue problem to determine the value of such that the matrix becomes singular. This is analogous to the guard map approach used in the proof in the previous subsection for the other direction of the proof, and in this case, we are varying from zero to infinity and finding the point such that for all larger , is positive definite. Intuitively, such an argument works since scales the positive definite matrix . Towards this end, consider the eigenvalue problem in given by
Let be the maximum positive eigenvalue, and zero otherwise. Then, since eigenvalues vary continuously, for all , so that by Lemma A.3.a we conclude that and have the same inertia, but this contradicts the stability of for all since .
E.3 Proof of Corollary 4.1
Appendix F Proof of Proposition 3.2
The structure of this proof is as follows. We begin by introducing general background for analyzing general singularly perturbed systems. Following this, we consider the linearization of the singularly perturbed system that approximates the simultaneous gradient dynamics and describe how insights made about this system translate to the corresponding nonlinear system. Finally, we analyze the stability of the linear system around a critical point to arrive at the stated result. The analysis is primarily from Kokotovic et al. (1986).
where and are assumed to be sufficiently many times continuously differential functions of the arguments , , , and . Observe that when , the dimension of the system given in (44) drops from to since degenerates into the equation
where the notation of indicates that the variables belong to the system with . We further require the assumption that (45) has isolated roots, which for each are given by
We now define an -dimensional manifold for any characterized by the expression
where is sufficiently many times continuously differentiable function of and . For to be an invariant manifold of the system given in (44), the expression in (46) must hold for all if it holds for . Formally, if
then is an invariant manifold for (44). Differentiating the expression in (47) with respect to , we obtain
Now, multiplying the expression in (48) by and substituting in the forms of , , and from (44) and (46), the manifold condition becomes
which must satisfy for all of interest and all , where is a positive constant.
Then, in terms of and , the system becomes
One interesting observation is that the above system is exactly the continuous time limiting system for the -Stackelberg learning update in Fiez et al. (2020) under a simple transformation of coordinates.
Observe that the invariant manifold is characterized by the fact that implies for all for which the manifold condition from (49) holds. This implies that if , it is sufficient to solve the system
This system is often referred to as the exact slow model and is valid for all and known as the slow manifold of (52).
We now consider the singularly perturbed system for simulataneous gradient descent given by
Let us linearize the system around a point . Then,Here, the means, e.g., , and similarly for .
Defining and and considering a point such that and , then linearized singularly perturbed system is given by
To simplify notation, let us define as follows
Then, an equivalent form of (52) is given by
The manifold condition from (49) for the system in (53) is given by
We claim that (54) can be satisfied by a function that is linear in . Indeed, defining
and then substituting back into (49), we get the simplified manifold condition of
Before we prove that an always exists to satisfy (55), consider the change of variables
The change of variables transforms the system from (53) into the equivalent representation
Consider that . Then, the system from (56) has the upper block-triangular form
which has the effect of generating a replacement fast subsystem given by
We now proceed to show that an such that always exists.
If is such that , there is an such that for all , there exists a solution to the matrix quadratic equation
To begin, observe that for , the unique solution to (59) is given by . Now, differentiating from (59) with respect to , we find
The unique solution of this equation at is
Accordingly, (60) represents the first two terms of the MacLaurin series for . ∎
We remark that as defined in (60) is unique in the sense that even though as given in (59) may have several real solutions, only one is approximated by (60).
The characteristic equation of (58) is equivalent to that for the system from (53) owing to the similarity transform between the systems. The block-triangular form of (53) admits a characteristic equation given by
is the characteristic polynomial of the slow subsystem, and
is the characteristic polynomial of the fast subsystem in the timescale . Consequently, of the eigenvalues of (53) denoted by are the roots of the slow characteristic equation and the rest of the eigenvalues are denoted by for and where are the roots of the fast characteristic equation .
The roots of at , given by the solution to
are the eigenvalues of the matrix defined in (61) since as shown in Lemma F.1. The roots of the fast characteristic equation at , given by the solution to
are the eigenvalues of the matrix . We now proceed by characterizing how closely the eigenvalues of the system at approximate the eigenvalues of the system from (53) as .
If , then as , eigenvalues of the system given in (53) tend toward the eigenvalues of the matrix while the remaining eigenvalues of the system from (53) tend to infinity with the rate along asymptotes defined by the eigenvalues of given as as a result of the continuity of coefficients of the polynomials from (63) and (64) with respect to .
Now, consider the special (but generic) case in which the eigenvalues of are distinct and the eigenvalues of are distinct, but and may have common eigenvalues. Then, taking the total derivative of (62) with respect to we have that
Now, observe that since the eigenvalues of are distinct.Recall that having distinct eigenvalues is a generic condition for a matrix an matrix, though not explicitly required for the asymptotic results; its only a condition for the big-O approximation for and where for . For each , this gives us a well-defined derivative (by the implicit mapping theorem) and hence, with , the approximation of follows directly. That is,
Similarly, taking the total derivative of and again applying the implicit function theorem, we have
where we have used the fact that .
Appendix G Proof of Theorem 3.2
Let be a stable critical point of -GDA which is not a differential Stackelberg equilibrium. Without loss of generality, suppose that has at least one eigenvalue with strictly positive real part.
Since both and have no zero valued eigenvalues, by Lemma A.3.b, there exists non-singular Hermitian matrices and positive definite Hermitian matrices such that and . Further, and have the same inertia, meaning
where for a given matrix , , , and are the number of eigenvalues of the argument that have positive, negative and zero real parts, respectively. Similarly, and have the same inertia:
Recall that we assumed has at least one eigenvalue with strictly positive real part. Hence, .
which can be verified by straightforward calculations.
Now, is also a real symmetric matrix, and hence, it is positive definite if and only if all its eigenvalues are positive. To determine the range of for which , we simply need to solve the eigenvalue problem
and extract the maximum eigenvalue, namely,
To provide some context for the proof approach, it follows the same idea as the proof of Theorem 3.1 in Appendix E.2.2. Indeed, to determine the range of such that is positive definite, we can formulate an eigenvalue problem to determine the value of such that the matrix becomes singular. We vary from zero to infinity in order to find the point such that for all larger , is positive definite. Intuitively, such an argument works since scales the positive definite matrix .
Appendix H Proof of Theorem 5.1
To prove the first part of this result, we following similar arguments to Theorem 4.1 of (Mescheder et al., 2018). To prove the second part, we leverage the concept of the quadratic numerical range. For both components of the proof, we will use the following form of the Jacobian of the regularized game. Indeed, first observe that the structural form of is
Examining (67), it is straightforward to see that the quadratic numerical range has eigenvalues of the form
Case 1: Suppose that . Then, trivially since .
Case 2: Suppose that . In this case, we want to ensure that
Appendix I Proof of Proposition 5.1
This proposition follows immediately from observing the structure of the Jacobian: for any matrix of the form
Appendix J Extensions in the Stochastic Setting
where and .
The stochastic process is a martingale difference sequence with respect to the increasing family of -fields defined by
Suppose that Assumption 3 holds and that is a differential Stackelberg equilibrium. Let where . There exists a and an such that for any fixed , there exists functions and so that when and where is such that for all , the stochastic iterates of -GDA with stepsize sequence and timescale separation satisfy
The utility of this result is that it provides a guarantee in the stochastic setting for a more reasonable and practically useful stepsize sequence. However, constructing the constants such as , and is highly non-trivial as can be seen in the work of Thoppe and Borkar (2019) and similar works in the area of stochastic approximation (Borkar, 2008). One direction of future work is examining the Lyapunov approach for directly analyzing the nonlinear singularly perturbed system; it is known, however, that the stochastic singularly perturbed systems have much weaker guarantees in terms of stability (Kokotovic et al., 1986, Chap. 4).
Appendix K Further Details on Related Work
In this section, we provide further details on the discussion from Section 7 regarding the results presented by Jin et al. (2020) on the local stability of gradient descent-ascent with a finite timescale separation. The purpose of this discussion is to make clear that Proposition 27 from the work of Jin et al. (2020) does not disagree with the results we provide in Theorem 3.1 and Theorem 3.2 and is instead complementary. In what follows, we recall Proposition 27 of Jin et al. (2020) in separate pieces in the terminology of this paper and delineate its meaning from our results on the stability of gradient descent-ascent with a finite timescale separation.
To begin, we consider the component of Proposition 27 from Jin et al. (2020) which says that given any fixed and finite timescale separation , a zero-sum game can be constructed with a differential Stackelberg equilibrium that is not stable with respect to the continuous time limiting system of -GDA given by the dynamics .
We now explain the proof. Let us consider any and the game
At the unique critical point , the Jacobian of the dynamics is given by
Moreover, observe that is a differential Stackelberg equilibrium and not a differential Nash equilibrium since , and . Finally, the spectrum of the Jacobian is
Let us now fix as any arbitrary positive value. Then, consider the game construction from (68) with . For the fixed choice of and subsequent game construction, we get that
This in turn means the differential Stackelberg equilibrium is not stable with respect to the dynamics for the given choice of . Since the choice of was arbitrary, this is a valid procedure to generate a game with a differential Stackelberg equilibrium that is not stable with respect to given a choice of beforehand.
We now move on to examining the portion of Proposition 27 from Jin et al. (2020) which says that given any fixed and finite timescale separation , a zero-sum game can be constructed with a critical point that is not a differential Stackelberg equilibrium which is stable with respect to the continuous time limiting system of -GDA given by .
In a similar manner as following Proposition K.1, we now explain the proof of Proposition K.2 and then contrast the result with Theorem 3.2. Again, consider any , along with the game construction
At the unique critical point , the Jacobian of the dynamics is given by
Observe that is neither a differential Nash equilibrium nor a differential Stackelberg equilibrium since and are both indefinite. The spectrum of the Jacobian is
Now, fix as any arbitrary positive value, then consider the game construction from (69) with . For the fixed choice of and resulting game construction given the choice of , we have that
This indicates that the non-equilibrium critical point is stable with respect to the dynamics where for the given choice of . Similar to the proof of Proposition K.1, since the choice of was arbitrary, the procedure to generate a game with a non-equilibrium critical point that is stable with respect to is valid given a choice of beforehand.
As a result, given the unique critical point of the game there is a finite such that the non-equilibrium critical point is not stable with respect to for all . In summary, Proposition K.2 is showing that there is exists a continuum of games for which a non-equilibrium critical point is stable given an unsuitable choice of finite learning rate ratio . In contrast, Theorem 3.2 is showing that given a game with a non-equilibrium critical point, there exists a range of finite learning rate ratios such that it is not stable.
To recap, the discussion in this section is meant to explicitly contrast Proposition 27 from the work of Jin et al. (2020) with Theorem 3.1 and Theorem 3.2 since they may potentially appear contradictory to each other without close inspection. The result of Jin et al. (2020) shows that (i) given a fixed finite learning ratio, there exists a game for with a differential Stackelberg equilibria that is not stable and (ii) given a fixed finite learning ratio, there exists a game with a non-equilibrium critical point that is stable. From a different perspective, we show that (i) given a fixed game and differential Stackelberg equilibrium, there exists a range of finite learning rate ratios for which the equilibrium is stable (Theorem 3.1) and (ii) given a fixed game and a non-equilibrium critical point, there exists a range of finite learning rate ratios for which the critical point is not stable (Theorem 3.2).
Appendix L Experiments Supplement
In this section we present several experiments not included in the body of the paper along with supplemental simulation results and details for the experiments presented in Section 6. We study a torus game in Section L.1 and examine the connection between timescale separation and the region of attraction. Then, in Section L.2, we return to the Dirac-GAN game and consider the non-saturating objective function. In Section L.3, we explore a generative adversarial network formulation using the Wasserstein cost function with a linear generator and quadratic discriminator for the problem of learning a covariance matrix. We finish in Section L.4 by presenting further results and details on our experiments training generative adversarial networks on image datasets.
We use the example in this section to further study the role of timescale separation on the regions of attraction around critical points. Consider the zero-sum game defined by the cost
This game can be interpreted as a location game on the torus. Specifically, the first player seeks to be far from the second player but near zero, while the second player seeks to be near the first player. This is a non-convex game on a non-convex strategy space. The critical points are given by the setNote that because the joint strategy space is a torus, , , and .:
The critical points and are the only differential Stackelberg equilibrium and neither is a differential Nash equilibrium. The differential Stackelberg equilibrium at is stable for all where and the differential Stackelberg equilibrium is stable for all where . The rest of the critical points are unstable for any choice of . We remark that we computed for each differential Stackelberg equilibrium using the construction from Theorem 3.1 in Section 3 and it again gave the exact value of such that the system is stable for all .
In Figure 13a, we show the trajectories of -GDA with and given the initializations and overlayed on the vector field generated by the respective timescale separation parameters. We observe that as the timescale separation grows, the rotational dynamics in the vector field dissipate and the directions of movement become sharp. As we mentioned in previous examples, -GDA moves directly to the zero line of and then along that line to an equilibrium given sufficient timescale separation. The warping of the vector field that occurs as a result of timescale separation impacts the equilibrium that the dynamics converge to from a fixed initial condition and the neighborhood on which -GDA converges to an equilibrium. In other words, the region of attraction around critical points depends heavily on the timescale separation .
To illustrate this fact, in Figure 13b we show the regions of attraction for each choice of timescale separation. The vector fields are again shown for each , but now with colors overlayed indicating the equilibria that the dynamics converge to given an initialization at that position. This experiment was generated by running -GDA with a dense set of initial conditions chosen uniformly over the strategy space. Positions in the strategy space without color did not converge to an equilibrium in the fixed horizon of 20000 iterations with . This happens when -GDA is not initialized in the local neighborhood of attraction around a stable equilibrium. For the choice of , is the only stable equilibrium. However, as demonstrated in Figure 13a, -GDA fails to converge to the equilibrium from the initial conditions and . This behavior is further demonstrated over the strategy space in Figure 13b and highlights the local nature of the guarantees since convergence is only assured given an initialization in a suitable local neighborhood around a stable critical point. On the other hand, -GDA converges to an equilibrium from any initial condition for as can be seen by Figure 13b. Notably, the equilibrium to which the learning dynamics converge depends on the timescale separation and initial condition. To give a concrete example, consider the initial conditions shown in Figure 13a of and . For the initial condition , -GDA converges to the equilibrium at for each . Yet, for the initial condition , -GDA converges to the equilibrium at for the respective choices of . In other words, the region of attraction around the critical points changes so that from a fixed initial condition -GDA may converge to distinct equilibrium depending on the initial condition. From Figure 13b, we see that the region of attraction around grows from to and , but then shrinks at . This example highlights that timescale separation has a fundamental impact on the region of attraction around critical points and as grows it is possible for the region of attraction around an equilibrium to shrink. Collectively, this motivates explicit methods for trying to shape the region of attraction around desirable equilibria.
L.2 Dirac-GAN and Regularization: Non-Saturating Formulation
In Section 6.4, we presented experiments for the Dirac-GAN game studied by Mescheder et al. (2017) using the original generative adversarial network formulation of Goodfellow et al. (2014). In this section, we revisit the Dirac-GAN game using the non-saturating generative adversarial network formulation also proposed by Goodfellow et al. (2014). While we refer the reader back to Section 6.4 for complete details on the Dirac-GAN, we do recall some key components of the formulation. Recall that the zero-sum game which arises from the original objective with regularization is defined by the cost
As discussed in Section 6.4, the unique critical point of the game is and it corresponds to the local Nash equilibrium of the unregularized game and a differential Stackelberg equilibrium of the regularized game. Moreover, the equilibrium is stable with respect to the continous time dynamics for all and so that the discrete time update -GDA converges with a suitable learning rate .
L.3 Generative Adversarial Network: Learning a Covariance Matrix
We now consider a generative adversarial network formulation presented by Daskalakis et al. (2018) for learning a covariance matrix. This is a simple example with degeneracies much like the Dirac-GAN game, but it can be generalized to arbitrary dimensional strategy spaces and has served as a benchmark for comparing convergence rates in a number of recent papers on learning in games. Often, the example is used to show that gradient descent-ascent cycles and converges slowly. However, by and large, timescale separation is not considered. We show that gradient descent-ascent converges fast in this game with suitable timescale separation and further explore the interplay between timescale separation, regularization, and rate of convergence. We primarily follow the notation of Daskalakis et al. (2018) when describing the problem.
As shown by Daskalakis et al. (2018), the cost function can be simplified to be expressed as
With this cost, the individual gradients for gradient descent-ascent are given by
From the individual gradients, it is clear that the critical points of the game are given by such that and . Moreover, given the form of , the game Jacobian at any critical point is of the form
Consequently, the eigenvalues of the game Jacobian are purely imaginary and the critical points are not stable. To fix this problem, Daskalakis et al. (2018) regularized both the generator and discriminator. We only regularize the discriminator in this example. The cost function of the zero-sum game with regularization is given by
The individual gradients for gradient descent-ascent in this regularized game are then
We begin by considering the simplest form of this problem, which is that . The critical points with this restriction are and and the game Jacobian evaluated at them is
From the eigenvalue trajectories, we see that as grows, the eigenvalues become purely real at a smaller value of . Moreover, as increases, the magnitude of the real and imaginary parts of the eigenvalues decreases. We observe the effect of this on the convergence, where the dynamics do not cycles as much for larger . Again, we see the trade-off between timescale separation, regularization, and convergence. For example, despite the eigenvalues being purely real with and so that there is no rotational dynamics, the convergence is slower than for where there is some non-zero imaginary piece of the eigenvalues.
Figures 15g, 15h, and 15i show the distance from a critical point along the learning path of -GDA with given a fixed initial condition with learning rate , regularization , and the dimension of the problem among the set , respectively. The primary purpose of showing this set of results is simply to be clear that the behavior for , which is easier to explain and visualize, transfers over to higher dimensional formulations of this problem. This is to be expected since the problem dimension is not necessarily fundamental to the convergence rate, but rather it depends on the conditioning of and each was chosen so that the behavior was comparable for each choice of dimension.
L.4 Generative Adversarial Networks: Image Data
In this section we provide additional results and details from the experiments we ran training generative adversarial networks on the CIFAR-10 and CelebA datasets. In Figure 17 we show more generated samples on each of the datasets. We ran our simulations based on the work of Mescheder et al. (2018) and used the publicly available code from the link https://github.com/LMescheder/GAN_stability. We refer the readers to (Mescheder et al., 2018) for details on the implementation and architectures, as we primarily only changed the learning rates used to run the experiments. For the networks, we ran the experiments using the architecture provided in the gan_training/models/resnet.py file of the repository. In Figure 18 we include the hyperparameters we used for the experiments. To be clear, we used the same exact setup for both training CIFAR-10 and CelebA datasets. We computed the Frechet Inception Distance using 10k samples from the real and generated data. For both experiments and across the set of hyperparameters we did the evaluation using a fixed random noise vector to make for an equal comparison and a fixed set of real images which were randomly selected. The evaluation was done using the training data. We used the FID score implementation in pytorch available at https://github.com/mseitzer/pytorch-fid.