Differentiable Game Mechanics

Alistair Letcher, David Balduzzi, Sebastien Racaniere, James Martens, Jakob Foerster, Karl Tuyls, Thore Graepel

Introduction

A significant fraction of recent progress in machine learning has been based on applying gradient descent to optimize the parameters of neural networks with respect to an objective function. The objective functions are carefully designed to encode particular tasks such as supervised learning. A basic result is that gradient descent converges to a local minimum of the objective function under a broad range of conditions (Lee et al., 2017). However, there is a growing set of algorithms that do not optimize a single objective function, including: generative adversarial networks (Goodfellow et al., 2014; Zhu et al., 2017), proximal gradient TD learning (Liu et al., 2016), multi-level optimization (Pfau and Vinyals, 2016), synthetic gradients (Jaderberg et al., 2017), hierarchical reinforcement learning (Wayne and Abbott, 2014; Vezhnevets et al., 2017), intrinsic curiosity (Pathak et al., 2017; Burda et al., 2019), and imaginative agents (Racanière et al., 2017). In effect, the models are trained via games played by cooperating and competing modules.

The time-average of iterates of gradient descent, and other more general no-regret algorithms, are guaranteed to converge to coarse correlated equilibria in games (Stoltz and Lugosi, 2007). However, the dynamics do not converge to Nash equilibria – and do not even stabilize in general (Mertikopoulos et al., 2018; Papadimitriou and Piliouras, 2018). Concretely, cyclic behaviors emerge even in simple cases, see example 1.

This paper presents an analysis of the second-order structure of game dynamics that allows to identify two classes of games, potential and Hamiltonian, that are easy to solve separately. We then derive symplectic gradient adjustmentSource code is available at https://github.com/deepmind/symplectic-gradient-adjustment. (SGA), a method for finding stable fixed points in games. SGA’s performance is evaluated in basic experiments.

Tractable algorithms that converge to Nash equilibria have been found for restricted classes of games: potential games, two-player zero-sum games, and a few others (Hart and Mas-Colell, 2013). Finding Nash equilibria can be reformulated as a nonlinear complementarity problem, but these are ‘hopelessly impractical to solve’ in general (Shoham and Leyton-Brown, 2008) because the problem is PPAD hard (Daskalakis et al., 2009).

Players are primarily neural nets in our setting. For computational reasons we restrict to gradient-based methods, even though game-theorists have considered a much broader range of techniques. Losses are not necessarily convex in any of their parameters, so Nash equilibria are not guaranteed to exist. Even leaving existence aside, finding Nash equilibria in nonconvex games is analogous to, but much harder than, finding global minima in neural nets – which is not realistic with gradient-based methods.

There are at least three problems with gradient-based methods in games. Firstly, the potential existence of cycles (recurrent dynamics) implies there are no convergence guarantees, see example 1 below and Mertikopoulos et al. (2018). Secondly, even when gradient descent converges, the rate of convergence may be too slow in practice because ‘rotational forces’ necessitate extremely small learning rates, see figure 4. Finally, since there is no single objective, there is no way to measure progress. Concretely, the losses obtained by the generator and the discriminator in GANs are not useful guides to the quality of the images generated. Application-specific proxies have been proposed, for example the inception score for GANs (Salimans et al., 2016), but these are of little help during training. The inception score is domain specific and is no substitute for looking at samples. This paper tackles the first two problems.

2 Outline and summary of main contributions

We start with the basic case of a zero-sum bimatrix game: example 1. It turns out that the dynamics under simultaneous gradient descent can be reformulated in terms of Hamilton’s equations. The cyclic behavior arises because the dynamics live on the level sets of the Hamiltonian. More directly useful, gradient descent on the Hamiltonian converges to a Nash equilibrium.

Lemma 1 shows that the Jacobian of any game decomposes into symmetric and antisymmetric components. There are thus two ‘pure’ cases corresponding to when the Jacobian is symmetric and anti-symmetric. The first case, known as potential games (Monderer and Shapley, 1996), have been intensively studied in the game-theory literature because they are exactly the games where gradient descent does converge.

Algorithms.

The general case, neither potential nor Hamiltonian, is more difficult and is therefore the focus of the remainder of the paper. Section 3 proposes symplectic gradient adjustment (SGA), a gradient-based method for finding stable fixed points in general games. Appendix A contains TensorFlow code to compute the adjustment. The algorithm computes two Jacobian-vector products, at a cost of two iterations of backprop. SGA satisfies a few natural desiderata explained in section 3.1: (D1D1) it is compatible with the original dynamics; and it is guaranteed to find stable equilibria in (D2D2) potential and (D3D3) Hamiltonian games.

For general games, correctly picking the sign of the adjustment (whether to add or subtract) is critical since it determines the behavior near stable and unstable equilibria. Section 2.4 defines stable equilibria and contrasts them with local Nash equilibria. Theorem 10 proves that SGA converges locally to stable fixed points for sufficiently small parameters (which we quantify via the notion of an additive condition number). While strong, this may be impractical or slow down convergence significantly. Accordingly, lemma 11 shows how to set the sign so as to be attracted towards stable equilibria and repelled from unstable ones. Correctly aligning SGA allows higher learning rates and faster, more robust convergence, see theorem 15. Finally, theorem 17 tackles the remaining class of saddle fixed points by proving that SGA locally avoids strict saddles for appropriate parameters.

Experiments.

We investigate the empirical performance of SGA in four basic experiments. The first experiment shows how increasing alignment allows higher learning rates and faster convergence, figure 4. The second set of experiments compares SGA with optimistic mirror descent on two-player and four-player games. We find that SGA converges over a much wider range of learning rates.

The last two sets of experiments investigate mode collapse, mode hopping and the related, less well-known problem of boundary distortion identified in Santurkar et al. (2018). Mode collapse and mode hopping are investigated in a setup involving a two-dimensional mixture of 16 Gaussians that is somewhat more challenging than the original problem introduced in Metz et al. (2017). Whereas simultaneous gradient descent completely fails, our symplectic adjustment leads to rapid convergence – slightly improved by correctly choosing the sign of the adjustment.

Finally, boundary distortion is studied using a 75-dimensional spherical Gaussian. Mode collapse is not an issue since there the data distribution is unimodal. However, as shown in figure 10, a vanilla GAN with RMSProp learns only one of the eigenvalues in the spectrum of the covariance matrix, whereas SGA approximately learns all of them.

The appendix provides some background information on differential and symplectic geometry, which motivated the developments in the paper. The appendix also explores what happens when the analogy with classical mechanics is pushed further than perhaps seems reasonable. We experiment with assigning units (in the sense of masses and velocities) to quantities in games, and find that type-consistency yields unexpected benefits.

3 Related work

Nash (1950) was only concerned with existence of equilibria. Convergence in two-player games was studied in Singh et al. (2000). WoLF (Win or Learn Fast) converges to Nash equilibria in two-player two-action games (Bowling and Veloso, 2002). Extensions include weighted policy learning (Abdallah and Lesser, 2008) and GIGA-WoLF (Bowling, 2004). Infinitesimal Gradient Ascent (IGA) is a gradient-based approach that is shown to converge to pure Nash equilibria in two-player two-action games. Cyclic behaviour may occur in case of mixed equilibria. Zinkevich (2003) generalised the algorithm to nn-action games called GIGA. Optimistic mirror descent approximately converges in two-player bilinear zero-sum games (Daskalakis et al., 2018), a special case of Hamiltonian games. In more general settings it converges to coarse correlated equilibria.

Convergence has also been studied in various nn-player settings, see Rosen (1965); Scutari et al. (2010); Facchinei and Kanzow (2010); Mertikopoulos and Zhou (2016). However, the recent success of GANs, where the players are neural networks, has focused attention on a much larger class of nonconvex games where comparatively little is known, especially in the nn-player case. Heusel et al. (2017) propose a two-time scale methods to find Nash equilibria. However, it likely scales badly with the number of players. Nagarajan and Kolter (2017) prove convergence for some algorithms, but under very strong assumptions (Mescheder et al., 2018). Consensus optimization (Mescheder et al., 2017) is closely related to our proposad algorithm, and is extensively discussed in section 3. A variety of game-theoretically or minimax motivated modifications to vanilla gradient descent have been investigated in the context of GANs, see Mertikopoulos et al. (2019); Gidel et al. (2018).

Learning with opponent-learning awareness (LOLA) infinitesimally modifies the objectives of players to take into account their opponents’ goals (Foerster et al., 2018). However, Letcher et al. (2019) recently showed that LOLA modifies fixed points and thus fails to find stable equilibria in general games.

Symplectic gradient adjustment was independently discovered by Gemp and Mahadevan (2018), who refer to it as “crossing-the-curl”. Their analysis draws on powerful techniques from variational inequalities and monotone optimization that are complementary to those developed here – see for example Gemp and Mahadevan (2016, 2017); Gidel et al. (2019). Using techniques from monotone optimization, Gemp and Mahadevan (2018) obtained more detailed and stronger results than ours, in the more particular case of Wasserstein LQ-GANs, where the generator is linear and the discriminator is quadratic (Feizi et al., 2017; Nagarajan and Kolter, 2017).

Network zero-sum games are shown to be Hamiltonian systems in Bailey and Piliouras (2019). The implications of the existence of invariant functions for games is just beginning to be understood and explored.

Dot products are written as v⊺w{\mathbf{v}}^{\intercal}{\mathbf{w}} or ⟨v,w⟩\langle{\mathbf{v}},{\mathbf{w}}\rangle. The angle between two vectors is θ(v,w)\theta({\mathbf{v}},{\mathbf{w}}). Positive definiteness is denoted S≻0{\mathbf{S}}\succ 0.

The Infinitesimal Structure of Games

In contrast to the classical formulation of games, we do not constrain the parameter sets to the probability simplex or require losses to be convex in the corresponding players’ parameters. Our motivation is that we are primarily interested in use cases where players are interacting neural nets such as GANs (Goodfellow et al., 2014), a situation in which results from classical game theory do not straightforwardly apply.

It is sometimes convenient to write w=(wi,w−i){\mathbf{w}}=({\mathbf{w}}_{i},{\mathbf{w}}_{-i}) where w−i{\mathbf{w}}_{-i} concatenates the parameters of all the players other than the ithi^{\text{th}}, which is placed out of order by abuse of notation.

The simultaneous gradient is the gradient of the losses with respect to the parameters of the respective players:

By the dynamics of the game, we mean following the negative of the vector field, −ξ-\boldsymbol{\xi}, with infinitesimal steps. There is no reason to expect ξ\boldsymbol{\xi} to be the gradient of a single function in general, and therefore no reason to expect the dynamics to converge to a fixed point.

Hamiltonian mechanics is a formalism for describing the dynamics in classical physical systems, see Arnold (1989); Guillemin and Sternberg (1990). The system is described via canonical coordinates (q,p)({\mathbf{q}},{\mathbf{p}}). For example, q{\mathbf{q}} often refers to position and p{\mathbf{p}} to momentum of a particle or particles.

The Hamiltonian of the system H(q,p){\mathcal{H}}({\mathbf{q}},{\mathbf{p}}) is a function that specifies the total energy as a function of the generalized coordinates. For example, in a closed system the Hamiltonian is given by the sum of the potential and kinetic energies of the particles. The time evolution of the system is given by Hamilton’s equations:

An importance consequence of the Hamiltonian formalism is that the dynamics of the physical system – that is, the trajectories followed by the particles in phase space – live on the level sets of the Hamiltonian. In other words, the total energy is conserved.

2 Hamiltonian Mechanics in Games

The next example illustrates the essential problem with gradients in games and the key insight motivating our approach.

has a Nash equilibrium at (x,y)=(0,0)({\mathbf{x}},{\mathbf{y}})=({\mathbf{0}},{\mathbf{0}}). The simultaneous gradient ξ(x,y)=(Ay,−A⊺x)\boldsymbol{\xi}({\mathbf{x}},{\mathbf{y}})=({\mathbf{A}}{\mathbf{y}},-{\mathbf{A}}^{\intercal}{\mathbf{x}}) rotates around the Nash, see figure 1.

Remarkably, the dynamics can be reformulated via Hamilton’s equations in the coordinates given by the SVD of A{\mathbf{A}}:

The vector field ξ\boldsymbol{\xi} cycles around the equilibrium because ξ\boldsymbol{\xi} conserves the Hamiltonian’s level sets (i.e. ⟨ξ,∇H⟩=0\langle\boldsymbol{\xi},\nabla{\mathcal{H}}\rangle=0). However, gradient descent on the Hamiltonian converges to the Nash equilibrium. The remainder of the paper explores the implications and limitations of this insight.

Papadimitriou and Piliouras (2016) recently analyzed the dynamics of Matching Pennies (essentially, the above example) and showed that the cyclic behavior covers the entire parameter space. The Hamiltonian reformulation directly explains the cyclic behavior via a conservation law.

3 The Generalized Helmholtz Decomposition

The Jacobian of a game with dynamics ξ\boldsymbol{\xi} is the (d×d)(d\times d)-matrix of second-derivatives {\mathbf{J}}({\mathbf{w}}):=\nabla_{\mathbf{w}}\cdot\boldsymbol{\xi}({\mathbf{w}})^{\intercal}=\Big{(}\frac{\partial\xi_{\alpha}({\mathbf{w}})}{\partial w_{\beta}}\Big{)}_{\alpha,\beta=1}^{d}, where ξα(w)\boldsymbol{\xi}_{\alpha}({\mathbf{w}}) is the αth\alpha^{\text{th}} entry of the dd-dimensional vector ξ(w)\boldsymbol{\xi}({\mathbf{w}}). Concretely, the Jacobian can be written as

The Jacobian of any vector field decomposes uniquely into two components J(w)=S(w)+A(w){\mathbf{J}}({\mathbf{w}})={\mathbf{S}}({\mathbf{w}})+{\mathbf{A}}({\mathbf{w}}) where S≡S⊺{\mathbf{S}}\equiv{\mathbf{S}}^{\intercal} is symmetric and A+A⊺≡0{\mathbf{A}}+{\mathbf{A}}^{\intercal}\equiv 0 is antisymmetric.

Proof Any matrix decomposes uniquely as M=S+A{\mathbf{M}}={\mathbf{S}}+{\mathbf{A}} where S=12(M+M⊺){\mathbf{S}}=\frac{1}{2}({\mathbf{M}}+{\mathbf{M}}^{\intercal}) and A=12(M−M⊺){\mathbf{A}}=\frac{1}{2}({\mathbf{M}}-{\mathbf{M}}^{\intercal}) are symmetric and antisymmetric. The decomposition is preserved by orthogonal change-of-coordinates: given orthogonal matrix P{\mathbf{P}}, we have P⊺MP=P⊺SP+P⊺AP{\mathbf{P}}^{\intercal}{\mathbf{M}}{\mathbf{P}}={\mathbf{P}}^{\intercal}{\mathbf{S}}{\mathbf{P}}+{\mathbf{P}}^{\intercal}{\mathbf{A}}{\mathbf{P}} since the terms remain symmetric and antisymmetric. Applying the decomposition to the Jacobian yields the result.

The connection to the classical Helmholtz decomposition in calculus is sketched in appendix B. Two natural classes of games arise from the decomposition:

A game is a potential game if the Jacobian is symmetric, i.e. if A(w)≡0{\mathbf{A}}({\mathbf{w}})\equiv 0. It is a Hamiltonian game if the Jacobian is antisymmetric, i.e. if S(w)≡0{\mathbf{S}}({\mathbf{w}})\equiv 0.

Potential games are well-studied and easy to solve. Hamiltonian games are a new class of games that are also easy to solve. The general case is more difficult, see section 3.

4 Stable Fixed Points (SFPs) vs Local Nash Equilibria (LNEs)

There are (at least) two possible solution concepts in general differentiable games: stable fixed points and local Nash equilibria.

We introduce local Nash equilibria because finding global Nash equilibria is unrealistic in games involving neural nets. Gradient-based methods can reliably find local – but not global – optima of nonconvex objective functions (Lee et al., 2016, 2017). Similarly, gradient-based methods cannot be expected to find global Nash equilibria in nonconvex games.

A fixed point w∗{\mathbf{w}}^{*} with ξ(w∗)=0\boldsymbol{\xi}({\mathbf{w}}^{*})=0 is stable if J(w∗)⪰0{\mathbf{J}}({\mathbf{w}}^{*})\succeq 0 and J(w∗){\mathbf{J}}({\mathbf{w}}^{*}) is invertible, unstable if J(w∗)≺0{\mathbf{J}}({\mathbf{w}}^{*})\prec 0 and a strict saddle if J(w∗){\mathbf{J}}({\mathbf{w}}^{*}) has an eigenvalue with negative real part. Strict saddles are a subset of unstable fixed points.

The definition is adapted from Letcher et al. (2019), where conditions on the Jacobian hold at the fixed point; in contrast, Balduzzi et al. (2018a) imposed conditions on the Jacobian in a neighborhood of the fixed point. We motivate this concept as follows.

Another viewpoint is that invertibility and positive semidefiniteness of the Hessian together imply positive definiteness, and the notion of stable fixed point specializes, in a one-player game, to local minima that are detected by the second partial derivative test. These minima are precisely those which gradient-like methods provably converge to. Stable fixed points are defined by analogy, though note that invertibility and semidefiniteness do not imply positive definiteness in nn-player games since J{\mathbf{J}} may not be symmetric.

The conditions J(w∗)⪰0{\mathbf{J}}({\mathbf{w}}^{*})\succeq 0 and J(w∗)≺0{\mathbf{J}}({\mathbf{w}}^{*})\prec 0 are equivalent to the conditions on the symmetric component S(w∗)⪰0{\mathbf{S}}({\mathbf{w}}^{*})\succeq 0 and S(w∗)≺0{\mathbf{S}}({\mathbf{w}}^{*})\prec 0 respectively, since

for all u{\mathbf{u}}, by antisymmetry of A{\mathbf{A}}. This equivalence will be used throughout.

Stable fixed points and local Nash equilibria are both appealing solution concepts, one from the viewpoint of optimisation by analogy with single objectives, and the other from game theory. Unfortunately, neither is a subset of the other:

5 Potential Games

Potential games were introduced by Monderer and Shapley (1996). It turns out that our definition of potential game above coincides with a special case of the potential games of Monderer and Shapley (1996), which they refer to as exact potential games.

for all ii and all wi′,wi′′,w−i{\mathbf{w}}_{i}^{\prime},{\mathbf{w}}_{i}^{\prime\prime},{\mathbf{w}}_{-i}, see Monderer and Shapley (1996).

If αi=1\alpha_{i}=1 for all ii then equation (11) is equivalent to requiring that the Jacobian of the game is symmetric.

Proof In an exact potential game, the Jacobian coincides with the Hessian of the potential function ϕ\phi, which is necessarily symmetric.

Monderer and Shapley (1996) refer to the special case where αi=1\alpha_{i}=1 for all ii as an exact potential game. We use the shorthand ‘potential game’ to refer to exact potential games in what follows.

Potential games have been extensively studied since they are one of the few classes of games for which Nash equilibria can be computed (Rosenthal, 1973). For our purposes, they are games where simultaneous gradient descent on the losses corresponds to gradient descent on a single function. It follows that descent on ξ\boldsymbol{\xi} converges to a fixed point that is a local minimum of ϕ\phi or a saddle.

6 Hamiltonian Games

Hamiltonian games, where the Jacobian is antisymmetric, are a new class games. They are related to the harmonic games introduced in Candogan et al. (2011), see section B.4. An example from Balduzzi et al. (2018b) may help develop intuition for antisymmetric matrices:

Suppose nn competitors play one-on-one and that the probability of player ii beating player jj is pijp_{ij}. Then, assuming there are no draws, the probabilities satisfy pij+pji=1p_{ij}+p_{ji}=1 and pii=12p_{ii}=\frac{1}{2}. The matrix A=(log⁡pij1−pij)i,j=1n{\mathbf{A}}=\left(\log\frac{p_{ij}}{1-p_{ij}}\right)_{i,j=1}^{n} of logits is then antisymmetric. Intuitively, antisymmetry reflects a hyperadversarial setting where all pairwise interactions between players are zero-sum.

Hamiltonian games are closely related to zero-sum games.

However, in general there are Hamiltonian games that are not zero-sum and vice versa.

Fix constants aa and bb and suppose players 11 and 22 minimize losses

with respect to xx and yy respectively.

The game actually has potential function ϕ(x,y)=x2−y2\phi(x,y)=x^{2}-y^{2}.

Hamiltonian games are quite different from potential games. In a Hamiltonian game there is a Hamiltonian function H{\mathcal{H}} that specifies a conserved quantity. In potential games the dynamics equal ∇ϕ\nabla\phi; in Hamiltonian games the dynamics are orthogonal to ∇H\nabla{\mathcal{H}}. The orthogonality implies the conservation law that underlies the cyclic behavior in example 1.

Let H(w):=12∥ξ(w)∥22{\mathcal{H}}({\mathbf{w}}):=\frac{1}{2}\|\boldsymbol{\xi}({\mathbf{w}})\|^{2}_{2}. If the game is Hamiltonian then

∇H=A⊺ξ\nabla{\mathcal{H}}={\mathbf{A}}^{\intercal}\boldsymbol{\xi} and

ξ\boldsymbol{\xi} preserves the level sets of H{\mathcal{H}} since ⟨ξ,∇H⟩=0\langle\boldsymbol{\xi},\nabla{\mathcal{H}}\rangle=0.

If the Jacobian is invertible and lim⁡∥w∥→∞H(w)=∞\lim_{\|{\mathbf{w}}\|\rightarrow\infty}{\mathcal{H}}({\mathbf{w}})=\infty then gradient descent on H{\mathcal{H}} converges to a stable fixed point.

Proof Direct computation shows ∇H=J⊺ξ\nabla{\mathcal{H}}={\mathbf{J}}^{\intercal}\boldsymbol{\xi} for any game. The first statement follows since J=A{\mathbf{J}}={\mathbf{A}} in Hamiltonian games.

For the second statement, the directional derivative is DξH=⟨ξ,∇H⟩=ξ⊺A⊺ξD_{\boldsymbol{\xi}}{\mathcal{H}}=\langle\boldsymbol{\xi},\nabla{\mathcal{H}}\rangle=\boldsymbol{\xi}^{\intercal}{\mathbf{A}}^{\intercal}\boldsymbol{\xi} where ξ⊺A⊺ξ=(ξ⊺A⊺ξ)⊺=ξ⊺Aξ=−(ξ⊺A⊺ξ)\boldsymbol{\xi}^{\intercal}{\mathbf{A}}^{\intercal}\boldsymbol{\xi}=(\boldsymbol{\xi}^{\intercal}{\mathbf{A}}^{\intercal}\boldsymbol{\xi})^{\intercal}=\boldsymbol{\xi}^{\intercal}{\mathbf{A}}\boldsymbol{\xi}=-(\boldsymbol{\xi}^{\intercal}{\mathbf{A}}^{\intercal}\boldsymbol{\xi}) since A=−A⊺{\mathbf{A}}=-{\mathbf{A}}^{\intercal} by anti-symmetry. It follows that ξ⊺A⊺ξ=0\boldsymbol{\xi}^{\intercal}{\mathbf{A}}^{\intercal}\boldsymbol{\xi}=0.

For the third statement, gradient descent on H{\mathcal{H}} will converge to a point where ∇H=J⊺ξ(w)=0\nabla{\mathcal{H}}={\mathbf{J}}^{\intercal}\boldsymbol{\xi}({\mathbf{w}})=0. If the Jacobian is invertible then clearly ξ(w)=0\boldsymbol{\xi}({\mathbf{w}})=0. The fixed-point is stable since 0≡S⪰00\equiv{\mathbf{S}}\succeq 0 in a Hamiltonian game, recall remark 1.

In fact, H{\mathcal{H}} is a Hamiltonian function for the game dynamics, see appendix B for a concise explanation. We use the notation H(w)=12∥ξ(w)∥2{\mathcal{H}}({\mathbf{w}})=\frac{1}{2}\|\boldsymbol{\xi}({\mathbf{w}})\|^{2} throughout the paper. However, H{\mathcal{H}} can only be interpreted as a Hamiltonian function for ξ\boldsymbol{\xi} when the game is Hamiltonian.

There is a precise mapping from Hamiltonian games to symplectic geometry, see appendix B. Symplectic geometry is the modern formulation of classical mechanics (Arnold, 1989; Guillemin and Sternberg, 1990). Recall that periodic behaviors (e.g. orbits) often arise in classical mechanics. The orbits lie on the level sets of the Hamiltonian, which expresses the total energy of the system.

Algorithms

We have seen that fixed points of potential and Hamiltonian games can be found by descent on ξ\boldsymbol{\xi} and ∇H\nabla{\mathcal{H}} respectively. This section tackles finding stable fixed points in general games.

There are two classes of games where we know how to find stable fixed points: potential games where ξ\boldsymbol{\xi} converges to a local minimum and Hamiltonian games where ∇H\nabla{\mathcal{H}}, which is orthogonal to ξ\boldsymbol{\xi}, finds stable fixed points.

In the general case, the following desiderata provide a set of reasonable properties for an adjustment ξλ\boldsymbol{\xi}_{\lambda} of the game dynamics. Recall that θ(u,v)\theta({\mathbf{u}},{\mathbf{v}}) is the angle between the vectors u{\mathbf{u}} and v{\mathbf{v}}.

To find stable fixed points, an adjustment ξλ\boldsymbol{\xi}_{\lambda} to the game dynamics should satisfy

compatibleTwo nonzero vectors are compatible if they have positive inner product. with game dynamics: ⟨ξλ,ξ⟩=α1⋅∥ξ∥2\langle\boldsymbol{\xi}_{\lambda},\boldsymbol{\xi}\rangle=\alpha_{1}\cdot\|\boldsymbol{\xi}\|^{2};

compatible with potential dynamics: if the game is a potential game then ⟨ξλ,∇ϕ⟩=α2⋅∥∇ϕ∥2\langle\boldsymbol{\xi}_{\lambda},\nabla\phi\rangle=\alpha_{2}\cdot\|\nabla\phi\|^{2};

compatible with Hamiltonian dynamics: If the game is Hamiltonian then ⟨ξλ,∇H⟩=α3⋅∥∇H∥2\langle\boldsymbol{\xi}_{\lambda},\nabla{\mathcal{H}}\rangle=\alpha_{3}\cdot\|\nabla{\mathcal{H}}\|^{2};

attracted to stable equilibria: in neighborhoods where S≻0{\mathbf{S}}\succ 0, require θ(ξλ,∇H)≤θ(ξ,∇H)\theta(\boldsymbol{\xi}_{\lambda},\nabla{\mathcal{H}})\leq\theta(\boldsymbol{\xi},\nabla{\mathcal{H}});

repelled by unstable equilibria: in neighborhoods where S≺0{\mathbf{S}}\prec 0, require θ(ξλ,∇H)≥θ(ξ,∇H)\theta(\boldsymbol{\xi}_{\lambda},\nabla{\mathcal{H}})\geq\theta(\boldsymbol{\xi},\nabla{\mathcal{H}}).

for some α1,α2,α3>0\alpha_{1},\alpha_{2},\alpha_{3}>0.

Desideratum D1D1 does not guarantee that players act in their own self-interest – this requires a stronger positivity condition on dot-products with subvectors of ξ\boldsymbol{\xi}, see Balduzzi (2017). Desiderata D2D2 and D3D3 imply that the adjustment behaves correctly in potential and Hamiltonian games respectively.

To understand desiderata D4D4 and D5D5, observe that gradient descent on H=12∥ξ∥2{\mathcal{H}}=\frac{1}{2}\|\boldsymbol{\xi}\|^{2} will find local minima that are fixed points of the dynamics. However, we specifically wish to converge to stable fixed points. Desideratum D4D4 and D5D5 require that the adjustment improves the rate of convergence to stable fixed points (by finding a steeper angle of descent), and avoids unstable fixed points.

More concretely, desiderata D4D4 can be interpreted as follows. If ξ\boldsymbol{\xi} points at a stable equilibrium then we require that ξλ\boldsymbol{\xi}_{\lambda} points more towards the equilibrium (i.e. has smaller angle). Conversely, desiderata D5D5 requires that if ξ\boldsymbol{\xi} points away then the adjustment should point further away.

The unadjusted dynamics ξ\boldsymbol{\xi} satisfies all the desiderata except D3D3.

2 Consensus Optimization

Since gradient descent on the function H(w)=12∥ξ∥2{\mathcal{H}}({\mathbf{w}})=\frac{1}{2}\|\boldsymbol{\xi}\|^{2} finds stable fixed points in Hamiltonian games, it is natural to ask how it performs in general games. If the Jacobian J(w){\mathbf{J}}({\mathbf{w}}) is invertible, then ∇H=J⊺ξ=0\nabla{\mathcal{H}}={\mathbf{J}}^{\intercal}\boldsymbol{\xi}=0 iff ξ=0\boldsymbol{\xi}=0. Thus, gradient descent on H{\mathcal{H}} converges to fixed points of ξ\boldsymbol{\xi}.

However, there is no guarantee that descent on H{\mathcal{H}} will find a stable fixed point. Mescheder et al. (2017) propose consensus optimization, a gradient adjustment of the form

Unfortunately, consensus optimization can converge to unstable fixed points even in simple cases where the ‘game’ is to minimize a single function:

Note that ∥ξ∥2=κ2(x2+y2)\|\boldsymbol{\xi}\|^{2}=\kappa^{2}(x^{2}+y^{2}) and

Descent on ξ+λ⋅J⊺ξ\boldsymbol{\xi}+\lambda\cdot{\mathbf{J}}^{\intercal}\boldsymbol{\xi} converges to the global maximum (x,y)=(0,0)(x,y)=(0,0) unless λ<1κ\lambda<\frac{1}{\kappa}.

Although consensus optimization works well in two-player zero-sum, it cannot be considered a candidate algorithm for finding stable fixed points in general games since it fails in the basic case of potential games. Consensus optimization only satisfies desiderata D3D3 and D4D4.

3 Symplectic Gradient Adjustment

The problem with consensus optimization is that it can perform worse than gradient descent on potential games. Intuitively, it makes bad use of the symmetric component of the Jacobian. Motivated by the analysis in section 2, we propose symplectic gradient adjustment, which takes care to only use the antisymmetric component of the Jacobian when adjusting the dynamics.

satisfies D1D1–D3D3 for λ>0\lambda>0, with α1=1=α2\alpha_{1}=1=\alpha_{2} and α3=λ\alpha_{3}=\lambda.

Proof First claim: λ⋅ξ⊺A⊺ξ=0\lambda\cdot\boldsymbol{\xi}^{\intercal}{\mathbf{A}}^{\intercal}\boldsymbol{\xi}=0 by anti-symmetry of A{\mathbf{A}}. Second claim: A≡0{\mathbf{A}}\equiv 0 in a potential game, so ξλ=ξ=∇ϕ\boldsymbol{\xi}_{\lambda}=\boldsymbol{\xi}=\nabla\phi. Third claim: ⟨ξλ,∇H⟩=⟨ξλ,J⊺ξ⟩=⟨ξλ,A⊺ξ⟩=λ⋅ξ⊺AA⊺ξ=λ⋅∥∇H∥2\langle\boldsymbol{\xi}_{\lambda},\nabla{\mathcal{H}}\rangle=\langle\boldsymbol{\xi}_{\lambda},{\mathbf{J}}^{\intercal}\boldsymbol{\xi}\rangle=\langle\boldsymbol{\xi}_{\lambda},{\mathbf{A}}^{\intercal}\boldsymbol{\xi}\rangle=\lambda\cdot\boldsymbol{\xi}^{\intercal}{\mathbf{A}}{\mathbf{A}}^{\intercal}\boldsymbol{\xi}=\lambda\cdot\|\nabla{\mathcal{H}}\|^{2} since J=A{\mathbf{J}}={\mathbf{A}} by assumption. Note that desiderata D1D1 and D2D2 are true even when λ<0\lambda<0. This will prove useful, since example 9 shows that it may be necessary to pick negative λ\lambda near S≺0{\mathbf{S}}\prec 0. Section 3.5 shows how to also satisfy desiderata D4D4 and D5D5.

4 Convergence

We begin by analysing convergence of SGA near stable equilibria. The following lemma highlights that the interaction between the symmetric and antisymmetric components is important for convergence. Recall that two matrices A{\mathbf{A}} and S{\mathbf{S}} commute iff [A,S]:=AS−SA=0[{\mathbf{A}},{\mathbf{S}}]:={\mathbf{A}}{\mathbf{S}}-{\mathbf{S}}{\mathbf{A}}={\mathbf{0}}. That is, A{\mathbf{A}} and S{\mathbf{S}} commute iff AS=SA{\mathbf{A}}{\mathbf{S}}={\mathbf{S}}{\mathbf{A}}. Intuitively, two matrices commute if they have the same preferred coordinate system.

If S⪰0{\mathbf{S}}\succeq 0 is symmetric positive semidefinite and S{\mathbf{S}} commutes with A{\mathbf{A}} then ξλ\boldsymbol{\xi}_{\lambda} points towards stable fixed points for non-negative λ\lambda:

Proof First observe that ξ⊺ASξ=ξ⊺S⊺A⊺ξ=−ξ⊺SAξ\boldsymbol{\xi}^{\intercal}{\mathbf{A}}{\mathbf{S}}\boldsymbol{\xi}=\boldsymbol{\xi}^{\intercal}{\mathbf{S}}^{\intercal}{\mathbf{A}}^{\intercal}\boldsymbol{\xi}=-\boldsymbol{\xi}^{\intercal}{\mathbf{S}}{\mathbf{A}}\boldsymbol{\xi}, where the first equality holds since the expression is a scalar, and the second holds since S=S⊺{\mathbf{S}}={\mathbf{S}}^{\intercal} and A=−A⊺{\mathbf{A}}=-{\mathbf{A}}^{\intercal}. It follows that ξ⊺ASξ=0\boldsymbol{\xi}^{\intercal}{\mathbf{A}}{\mathbf{S}}\boldsymbol{\xi}=0 if SA=AS{\mathbf{S}}{\mathbf{A}}={\mathbf{A}}{\mathbf{S}}. Finally rewrite the inequality as

since ξ⊺ASξ=0\boldsymbol{\xi}^{\intercal}{\mathbf{A}}{\mathbf{S}}\boldsymbol{\xi}=0 and by positivity of S{\mathbf{S}}, λ\lambda and AA⊺{\mathbf{A}}{\mathbf{A}}^{\intercal}.

The lemma suggests that in general the failure of A{\mathbf{A}} and S{\mathbf{S}} to commute should be important for understanding the dynamics of ξλ\boldsymbol{\xi}_{\lambda}. We therefore introduce the additive condition number κ\kappa to upper-bound the worst-case noncommutativity of S{\mathbf{S}}, which allows to quantify the relationship between ξλ\boldsymbol{\xi}_{\lambda} and ∇H\nabla{\mathcal{H}}. If κ=0\kappa=0, then S=σ⋅I{\mathbf{S}}=\sigma\cdot{\mathbf{I}} commutes with all matrices. The larger the additive condition number κ\kappa, the larger the potential failure of S{\mathbf{S}} to commute with other matrices.

Let S{\mathbf{S}} be a symmetric matrix with eigenvalues σmax≥⋯≥σmin\sigma_{\text{max}}\geq\cdots\geq\sigma_{\text{min}}. The additive condition numberThe condition number of a positive definite matrix is σmaxσmin\frac{\sigma_{\text{max}}}{\sigma_{\text{min}}}. of S{\mathbf{S}} is κ:=σmax−σmin\kappa:=\sigma_{\text{max}}-\sigma_{\text{min}}. If S⪰0{\mathbf{S}}\succeq 0 is positive semidefinite with additive condition number κ\kappa then λ∈(0,4κ)\lambda\in(0,\frac{4}{\kappa}) implies

If S{\mathbf{S}} is negative semidefinite, then λ∈(0,4κ)\lambda\in(0,\frac{4}{\kappa}) implies

The inequalities are strict if J{\mathbf{J}} is invertible.

Proof We prove the case S⪰0{\mathbf{S}}\succeq 0; the case S⪯0{\mathbf{S}}\preceq 0 is similar. Rewrite the inequality as

Since S{\mathbf{S}} is positive semidefinite, there exists an upper-triangular square-root matrix TT such that T⊺T=S{\mathbf{T}}^{\intercal}{\mathbf{T}}={\mathbf{S}} and so ξ⊺Sξ=∥Tξ∥2\boldsymbol{\xi}^{\intercal}{\mathbf{S}}\boldsymbol{\xi}=\|{\mathbf{T}}\boldsymbol{\xi}\|^{2}. Further,

since ∥T∥2=σmax\|{\mathbf{T}}\|_{2}=\sqrt{\sigma_{\text{max}}}. Putting the observations together obtains

Set α=λ\alpha=\sqrt{\lambda} and η=σmax\eta=\sqrt{\sigma_{max}}. We can continue the above computation

Finally, 2α−α2η>02\alpha-\alpha^{2}\eta>0 for any α\alpha in the range (0,2η)(0,\frac{2}{\eta}), which is to say, for any 0<λ<4σmax0<\lambda<\frac{4}{\sigma_{max}}. The kernel of S{\mathbf{S}} and the kernel of T{\mathbf{T}} coincide. If ξ\boldsymbol{\xi} is in the kernel of A{\mathbf{A}}, resp. T{\mathbf{T}}, it cannot be in the kernel of T{\mathbf{T}}, resp. A{\mathbf{A}} and the term (∥Tξ∥−α∥Aξ∥)2(\|{\mathbf{T}}\boldsymbol{\xi}\|-\alpha\|{\mathbf{A}}\boldsymbol{\xi}\|)^{2} is positive. Otherwise, the term ∥Aξ∥∥Tξ∥\|{\mathbf{A}}\boldsymbol{\xi}\|\|{\mathbf{T}}\boldsymbol{\xi}\| is positive.

Proof This is a standard result on fixed-point iterations, adapted from Ortega and Rheinboldt (2000, 10.1.3).

A matrix M{\mathbf{M}} is called positive stable if all its eigenvalues have positive real part. Assume w∗{\mathbf{w}}^{*} is a fixed point of a differentiable game such that (I+λA⊺)J(w∗)({\mathbf{I}}+\lambda{\mathbf{A}}^{\intercal}){\mathbf{J}}({\mathbf{w}}^{*}) is positive stable for λ\lambda in some set Λ\Lambda. Then SGA converges locally to w∗{\mathbf{w}}^{*} for λ∈Λ\lambda\in\Lambda and α>0\alpha>0 sufficiently small.

Proof Let X=(I+λA⊺)X=({\mathbf{I}}+\lambda{\mathbf{A}}^{\intercal}). By definition of fixed points, ξ(w∗)=0\boldsymbol{\xi}({\mathbf{w}}^{*})=0 and so

is positive stable by assumption, namely has eigenvalues ak+ibka_{k}+ib_{k} with ak>0a_{k}>0. Writing F(w)=w−αXξ(w)F({\mathbf{w}})={\mathbf{w}}-\alpha X\boldsymbol{\xi}({\mathbf{w}}) for the iterative procedure given by SGA, it follows that

has eigenvalues 1−αak−iαbk1-\alpha a_{k}-i\alpha b_{k}, which are in the unit circle for small α\alpha. More precisely,

which is always possible for ak>0a_{k}>0. Hence ∇F(w∗)\nabla F({\mathbf{w}}^{*}) has eigenvalues in the unit circle for 0<α<min⁡k2ak/(ak2+bk2)0<\alpha<\min_{k}2a_{k}/(a_{k}^{2}+b_{k}^{2}), and we are done by Ostrowski’s Theorem since w∗{\mathbf{w}}^{*} is a fixed point of FF.

Let w∗{\mathbf{w}}^{*} be a stable fixed point and κ\kappa the additive condition number of S(w∗){\mathbf{S}}({\mathbf{w}}^{*}). Then SGA converges locally to w∗{\mathbf{w}}^{*} for all λ∈(0,4κ)\lambda\in(0,\frac{4}{\kappa}) and α>0\alpha>0 sufficiently small.

Proof By Theorem 5 and the assumption that w∗{\mathbf{w}}^{*} is a stable fixed point with invertible Jacobian, we know that

for λ∈(0,4κ)\lambda\in(0,\frac{4}{\kappa}). The proof does not rely on any particular property of ξ\boldsymbol{\xi}, and can trivially be extended to the claim that

for all non-zero vectors u{\mathbf{u}}. In particular this can be rewritten as

which implies positive definiteness of J(I+λA⊺){\mathbf{J}}({\mathbf{I}}+\lambda{\mathbf{A}}^{\intercal}). A positive definite matrix is positive stable, and any matrices ABAB and BABA have identical spectrum. This implies also that (I+λA⊺)J({\mathbf{I}}+\lambda{\mathbf{A}}^{\intercal}){\mathbf{J}} is positive stable, and we are done by the corollary above.

We conclude that SGA converges to an SFP if λ\lambda is small enough, where ‘small enough’ depends on the additive condition number.

5 Picking sign(λ)sign𝜆\operatorname*{sign}(\lambda)

This section explains desiderata D4D4–D5D5 and shows how to pick sign⁡(λ)\operatorname*{sign}(\lambda) to speed up convergence towards stable and away from unstable fixed points. In the example below, almost any choice of positive λ\lambda results in convergence to an unstable equilibrium. The problem arises from the combination of a weak repellor with a strong rotational force.

with an unstable equilibrium at (0,0)(0,0). The dynamics are

which converges to the unstable equilibrium if λ>ϵ\lambda>\epsilon.

We now show how to pick the sign of λ\lambda to avoid unstable equilibria. First, observe that ⟨ξ,∇H⟩=ξ⊺(S+A)⊺ξ=ξ⊺Sξ\langle\boldsymbol{\xi},\nabla{\mathcal{H}}\rangle=\boldsymbol{\xi}^{\intercal}({\mathbf{S}}+{\mathbf{A}})^{\intercal}\boldsymbol{\xi}=\boldsymbol{\xi}^{\intercal}{\mathbf{S}}\boldsymbol{\xi}. It follows that for ξ≠0\boldsymbol{\xi}\neq 0:

A criterion to probe the positive/negative definiteness of S{\mathbf{S}} is thus to check the sign of ⟨ξ,∇H⟩\langle\boldsymbol{\xi},\nabla{\mathcal{H}}\rangle. The dot product can take any value if S{\mathbf{S}} is neither positive nor negative (semi-)definite. The behavior near saddle points will be explored in Section 3.7.

Recall that desiderata D4D4 requires that, if ξ\boldsymbol{\xi} points at a stable equilibrium then we require that ξλ\boldsymbol{\xi}_{\lambda} points more towards the equilibrium (i.e. has smaller angle). Conversely, desiderata D5D5 requires that, if ξ\boldsymbol{\xi} points away then the adjustment should point further away. More formally,

Let u{\mathbf{u}} and v{\mathbf{v}} be two vectors. The infinitesimal alignment of ξλ:=u+λ⋅v\boldsymbol{\xi}_{\lambda}:={\mathbf{u}}+\lambda\cdot{\mathbf{v}} with a third vector w{\mathbf{w}} is

If u{\mathbf{u}} and w{\mathbf{w}} point the same way, u⊺w>0{\mathbf{u}}^{\intercal}{\mathbf{w}}>0, then align⁡>0\operatorname*{align}>0 when v{\mathbf{v}} bends u{\mathbf{u}} further toward w{\mathbf{w}}, see figure 2A. Otherwise align⁡>0\operatorname*{align}>0 when v{\mathbf{v}} bends u{\mathbf{u}} away from w{\mathbf{w}}, see figure 2B.

The following lemma allows us to rewrite the infinitesimal alignment in terms of known (computable) quantities, from which we can deduce the correct choice of λ\lambda.

When ξλ\boldsymbol{\xi}_{\lambda} is the symplectic gradient adjustment,

where the denominator has no linear term in λ\lambda because ξ⊥A⊺ξ\boldsymbol{\xi}\perp{\mathbf{A}}^{\intercal}\boldsymbol{\xi}. It follows that the sign of the infinitesimal alignment is

Intuitively, computing the sign of ⟨ξ,∇H⟩\langle\boldsymbol{\xi},\nabla{\mathcal{H}}\rangle provides a check for stable and unstable fixed points. Computing the sign of ⟨A⊺ξ,∇H⟩\langle{\mathbf{A}}^{\intercal}\boldsymbol{\xi},\nabla{\mathcal{H}}\rangle checks whether the adjustment term points towards or away from the nearby fixed point. Putting the two checks together yields a prescription for the sign of λ\lambda, as follows.

Desiderata D4D4–D5D5 are satisfied for λ\lambda such that λ⋅⟨ξ,∇H⟩⋅⟨A⊺ξ,∇H⟩≥0\lambda\cdot\langle\boldsymbol{\xi},\nabla{\mathcal{H}}\rangle\cdot\langle{\mathbf{A}}^{\intercal}\boldsymbol{\xi},\nabla{\mathcal{H}}\rangle\geq 0.

Proof If we are in a neighborhood of a stable fixed point then ⟨ξ,∇H⟩≥0\langle\boldsymbol{\xi},\nabla{\mathcal{H}}\rangle\geq 0. It follows by lemma 11 that \operatorname*{sign}\Big{(}\operatorname*{align}(\boldsymbol{\xi}_{\lambda}),\nabla{\mathcal{H}})\Big{)}=\operatorname*{sign}\Big{(}\langle{\mathbf{A}}^{\intercal}\boldsymbol{\xi},\nabla{\mathcal{H}}\rangle\Big{)} and so choosing \operatorname*{sign}(\lambda)=\operatorname*{sign}\Big{(}\langle{\mathbf{A}}^{\intercal}\boldsymbol{\xi},\nabla{\mathcal{H}}\rangle\Big{)} leads to the angle between ξλ\boldsymbol{\xi}_{\lambda} and ∇H\nabla{\mathcal{H}} being smaller than the angle between ξ\boldsymbol{\xi} and ∇H\nabla{\mathcal{H}}, satisfying desideratum D4D4. The proof for the unstable case is similar.

Gradient descent is also known as the method of steepest descent. In general games, however, ξ\boldsymbol{\xi} does not follow the steepest path to fixed points due to the ‘rotational force’, which forces lower learning rates and slows down convergence.

The following lemma provides some intuition about alignment. The idea is that, the smaller the cosine between the ‘correct direction’ w{\mathbf{w}} and the ‘update direction’ ξ\boldsymbol{\xi}, the smaller the learning rate needs to be for the update to stay in a unit ball, see figure 3.

If w{\mathbf{w}} and ξ\boldsymbol{\xi} are unit vectors with 0<w⊺ξ0<{\mathbf{w}}^{\intercal}\boldsymbol{\xi} then ∥w−η⋅ξ∥≤1\|{\mathbf{w}}-\eta\cdot\boldsymbol{\xi}\|\leq 1 for 0≤η≤2w⊺ξ=2cos⁡θ(w,ξ)0\leq\eta\leq 2{\mathbf{w}}^{\intercal}\boldsymbol{\xi}=2\cos\theta({\mathbf{w}},\boldsymbol{\xi}). In other words, ensuring that w−ηξ{\mathbf{w}}-\eta\boldsymbol{\xi} is closer to the origin than w{\mathbf{w}} requires smaller learning rates η\eta as the angle between w{\mathbf{w}} and ξ\boldsymbol{\xi} gets larger.

Proof Check ∥w−η⋅ξ∥2=1+η2−2η⋅w⊺ξ≤1\|{\mathbf{w}}-\eta\cdot\boldsymbol{\xi}\|^{2}=1+\eta^{2}-2\eta\cdot{\mathbf{w}}^{\intercal}\boldsymbol{\xi}\leq 1 iff η2≤2η⋅w⊺ξ\eta^{2}\leq 2\eta\cdot{\mathbf{w}}^{\intercal}\boldsymbol{\xi}. The result follows.

The next lemma is a standard technical result from the convex optimization literature.

Finally, we show that increasing alignment helps speed convergence:

Suppose ff is convex and Lipschitz smooth with ∥∇f(x)−∇f(y)∥≤L⋅∥x−y∥\|\nabla f({\mathbf{x}})-\nabla f({\mathbf{y}})\|\leq L\cdot\|{\mathbf{x}}-{\mathbf{y}}\|. Let wt+1=wt−η⋅v{\mathbf{w}}_{t+1}={\mathbf{w}}_{t}-\eta\cdot{\mathbf{v}} where ∥v∥=∥∇f(wt)∥\|{\mathbf{v}}\|=\|\nabla f({\mathbf{w}}_{t})\|. Then the optimal step size is η∗=cos⁡θL\eta^{*}=\frac{\cos\theta}{L} where θ:=θ(∇f(wt),v)\theta:=\theta(\nabla f({\mathbf{w}}_{t}),{\mathbf{v}}), with

The proof of Theorem 15 adapts lemma 14 to handle the angle arising from the ‘rotational force’.

to obtain η∗=αL\eta^{*}=\frac{\alpha}{L} and Δ(η∗)=−α22L\Delta(\eta^{*})=-\frac{\alpha^{2}}{2}L as required.

Increasing the cosine with the steepest direction improves convergence. The alignment computation in algorithm 1 chooses λ\lambda to be positive or negative such that ξλ\boldsymbol{\xi}_{\lambda} is bent towards stable (increasing the cosine) and away from unstable fixed points. Adding a small ϵ>0\epsilon>0 to the computation introduces a weak bias towards stable fixed points.

6 Aligned Consensus Optimization

The stability criterion in (36) also provides a simple way to prevent consensus optimization from converging to unstable equilibria. Aligned consensus optimization is

where in practice we set λ=1\lambda=1. Aligned consensus optimization satisfies desiderata D3–D5. However, it behaves strangely in potential games. Multiplying by the Jacobian is the ‘inverse’ of Newton’s method since for potential games the Jacobian of ξ\boldsymbol{\xi} is the Hessian of the potential function. Multiplying by the Hessian increases the gap between small and large eigenvalues, increasing the (usual, multiplicative) condition number and slows down convergence. Nevertheless, consensus optimization works well in GANs (Mescheder et al., 2017), and aligned consensus may improve performance, see experiments below.

Dropping the first term ξ\boldsymbol{\xi} from (48) yields a simpler update that also satisfies D3D3–D5D5. However, the resulting algorithm performs poorly in experiments (not shown), perhaps because it is attracted to saddles.

7 Avoiding Strict Saddles

How does SGA behave near saddles? We show that Symplectic Gradient Adjustment locally avoids strict saddles, provided that λ\lambda and α\alpha are small and parameters are initialized with (arbitrarily small) noise. More precisely, let F(w)=w−αξλ(w){\mathbf{F}}({\mathbf{w}})={\mathbf{w}}-\alpha\boldsymbol{\xi}_{\lambda}({\mathbf{w}}) be the iterative optimization procedure given by SGA. Then every strict saddle w∗{\mathbf{w}}^{*} has a neighbourhood UU such that {w∈U∣Fn(w)→w∗ as n→∞}\{{\mathbf{w}}\in U\mid{\mathbf{F}}^{n}({\mathbf{w}})\to{\mathbf{w}}^{*}\text{ as }n\to\infty\} has measure zero for small α>0\alpha>0 and λ\lambda.

Intuitively, the Taylor expansion around a strict saddle w∗{\mathbf{w}}^{*} is locally dominated by the Jacobian at w∗{\mathbf{w}}^{*}, which has a negative eigenvalue. This prevents convergence to w∗{\mathbf{w}}^{*} for random initializations of w{\mathbf{w}} near w∗{\mathbf{w}}^{*}. The argument is made rigorous using the Stable Manifold Theorem following Lee et al. (2017).

It follows that if ∇F(w∗)\nabla F({\mathbf{w}}^{*}) has at least one eigenvalue ∣σ∣>1\left|\sigma\right|>1 then EuE^{u} has dimension at least 11. Since WW has tangent space EsE^{s} at w∗{\mathbf{w}}^{*} with codimension at least one, we conclude that WW has measure zero. This is central to proving that the set of nearby initial points which converge to a given strict saddle w∗{\mathbf{w}}^{*} has measure zero. Since w{\mathbf{w}} is initialized randomly, the following theorem is obtained.

SGA locally avoids strict saddles almost surely, for α>0\alpha>0 and λ\lambda small.

Proof Let w∗{\mathbf{w}}^{*} a strict saddle and recall that SGA is given by

All terms involved are continuously differentiable and we have

by assumption that ξ(w∗)=0\boldsymbol{\xi}({\mathbf{w}}^{*})=0. Since all terms except I{\mathbf{I}} are of order at least α\alpha, ∇F(w∗)\nabla F({\mathbf{w}}^{*}) is invertible for all α\alpha sufficiently small. By the inverse function theorem, there exists a neighbourhood UU of w∗{\mathbf{w}}^{*} such that FF is has a continuously differentiable inverse on UU. Hence FF restricted to UU is a C1C^{1} diffeomorphism with fixed point w∗{\mathbf{w}}^{*}.

By definition of strict saddles, J(w∗){\mathbf{J}}({\mathbf{w}}^{*}) has an eigenvalue with negative real part. It follows by continuity that (I−αJ)J(w∗)({\mathbf{I}}-\alpha{\mathbf{J}}){\mathbf{J}}({\mathbf{w}}^{*}) also has an eigenvalue a+iba+ib with a<0a<0 for α\alpha sufficiently small. Finally,

has an eigenvalue σ=1−αa−iαb\sigma=1-\alpha a-i\alpha b with

It follows that EsE^{s} has codimension at least one, implying in turn that the local stable set WW has measure zero. We can now prove that

Now F−1F^{-1} is C1C^{1}, hence locally Lipschitz and thus preserves sets of measure zero, so that F−n(W)F^{-n}(W) has measure zero for each nn. Countable unions of measure zero sets are still measure zero, so we conclude that ZZ also has measure zero. In other words, SGA converges to w∗{\mathbf{w}}^{*} with zero probability upon random initialization of w{\mathbf{w}} in UU.

Unlike stable and unstable fixed points, it is unclear how to avoid strict saddles using only alignment, that is, independently from the size of λ\lambda.

Experiments

We compare SGA with simultaneous gradient descent, optimistic mirror descent (Daskalakis et al., 2018) and consensus optimization (Mescheder et al., 2017) in basic settings.

We investigate the effect of SGA when a weak attractor is coupled to a strong rotational force:

Gradient descent is extremely sensitive to the choice of learning rate η\eta, top row of figure 4. As η\eta increases through {0.01,0.032,0.1}\{0.01,0.032,0.1\} gradient descent goes from converging extremely slowly, to diverging slowly, to diverging rapidly. SGA yields faster, more robust convergence. SGA converges faster with learning rates η=0.01\eta=0.01 and η=0.032\eta=0.032, and only starts overshooting the fixed point for η=0.1\eta=0.1.

2 Basic adversarial games

Optimistic mirror descent is a family of algorithms that has nice convergence properties in games (Rakhlin and Sridharan, 2013; Syrgkanis et al., 2015). In the special case of optimistic gradient descent the updates are

The left panel shows the number of steps to convergence (when convergence occurs) over a range of learning rates. OMD’s peak performance is better than SGA, where the red curve dips below the blue. Howwever, we find that SGA converges – and does so faster – for a much wider range of learning rates. OMD diverges for learning rates not in the range [0.3, 1.2]. Simultaneous gradient descent oscillates without converging (not shown). The right panel shows the average performance of OMD and SGA on the last 10 steps. Once again, here SGA consistently performs better over a wider range of learning rates. Individual runs are shown in figure 6.

Figure 7 shows time to convergence (using the same convergence criterion as above) for optimistic mirror descent and SGA. The games are constructed with four players, each of which controls one parameter. The losses are

where ϵ=1100\epsilon=\frac{1}{100} in the left panel and ϵ=0\epsilon=0 in the right panel. The antisymmetric component of the game Jacobian is

OMD converges considerably slower than SGA across the full range of learning rates. It also diverges for learning rates >0.22>0.22. In contrast, SGA converges more quickly and robustly.

3 Learning a two-dimensional mixture of Gaussians

We apply SGA to a basic Generative Adversarial Network setup adapted from Metz et al. (2017). Data is sampled from a highly multimodal distribution designed to probe the tendency of GANs to collapse onto a subset of modes during training. The distribution is a mixture of 16 Gaussians arranged in a 4×44\times 4 grid. Figure 8 shows the probability distribution that is sampled to train the generator and discriminator. The generator and discriminator networks both have 6 ReLU layers of 384 neurons. The generator has two output neurons; the discriminator has one.

Figure 9 shows results after {2000,4000,6000,8000}\{2000,4000,6000,8000\} iterations. The networks are trained under RMSProp. Learning rates were chosen by visual inspection of grid search results at iteration 8000. More precisely, grid search was over learning rates {\{1e-5, 2e-5,5e-5, 8e-5, 1e-4, 2e-4, 5e-4}\} and then a more refined linear search over [[8e-5, 2e-4]]. Simultaneous gradient descent and SGA are shown in the figure.

The last two rows of figure 9 show the performance of consensus optimization without and with alignment. Introducing alignment slightly improves speed of convergence (second column) and final result (fourth column), although intermediate results in third column are ambiguous.

Simultaneous gradient descent exhibits mode collapse followed by mode hopping in later iterations (not shown). Mode hopping is analogous to the cycles in example 1. Unaligned SGA converges to the correct distribution; alignment speeds up convergence slightly. Consensus optimization performs similarly in this GAN example. However, consensus optimization can converge to local maxima even in potential games, recall example 8.

4 Learning a high-dimensional unimodal Gaussian

Mode collapse is a well-known phenomenon in GANs. A more subtle phenomenon, termed boundary distortion, was identified in Santurkar et al. (2018). Boundary distortion is a form of covariate shift where the generator fails to model the true data distribution.

Santurkar et al demonstrate boundary distortion using data sampled from a 75-dimensional unimodal Gaussian with spherical covariate matrix. Mode collapse is not a problem in this setting because the data distribution is unimodal. Nevertheless, they show that vanilla GANs fail to learn most of the spectrum of the covariate matrix.

Figure 10 reproduces their result. Panel A shows the ground truth: all 75 eigenvalues are equal to 1.0. Panel B shows the spectrum of the covariance matrix of the data generated by a GAN trained with RMSProp. The GAN concentrates on a single eigenvalue and essentially ignores the remaining 74 eigenvalues. This is similar to, but more extreme than, the empirical results obtained in Santurkar et al. (2018). We emphasize that the problem is not mode collapse, since the data is unimodal (although, it’s worth noting that most of the mass of a high-dimensional Gaussian lies on the “shell”).

Finally, panel C shows the spectrum of the covariance matrix of the data sampled from a GAN trained via SGA. The GAN approximately learns all the eigenvalues, with values ranging between 0.6 and 1.5.

Discussion

Modern deep learning treats differentiable modules like plug-and-play lego blocks. For this to work, at the very least, we need to know that gradient descent will find local minima. Unfortunately, gradient descent does not necessarily find local minima when optimizing multiple interacting objectives. With the recent proliferation of algorithms that optimize more than one loss, it is becoming increasingly urgent to understand and control the dynamics of interacting losses. Although there is interesting recent work on two-player adversarial games such as GANs, there is essentially no work on finding stable fixed points in more general games played by interacting neural nets.

The generalized Helmholtz decomposition provides a powerful new perspective on game dynamics. A key feature is that the analysis is indifferent to the number of players. Instead, it is the interplay between the simultaneous gradient ξ\boldsymbol{\xi} on the losses and the symmetric and antisymmetric matrices of second-order terms that guides algorithm design and governs the dynamics under gradient adjustments.

Symplectic gradient adjustment is a straightforward application of the generalized Helmholtz decomposition. It is unlikely that SGA is the best approach to finding stable fixed points. A deeper understanding of the interaction between the potential and Hamiltonian components will lead to more effective algorithms. Reinforcement learning algorithms that optimize multiple objectives are increasingly common, and second-order terms are difficult to estimate in practice. Thus, first-order methods that do not use Jacobian-vector products are of particular interest.

Finally, it is worth raising a philosophical point. In this paper we are concerned with finding stable fixed points (because, for example, they yield pleasing samples in GANs). We are not concerned with the losses of the players per se. The gradient adjustments may lead to a player acting against its own self-interest by increasing its loss. We consider this acceptable insofar as it encourages convergence to a stable fixed point. The players are but a means to an end.

We have argued that stable fixed points are a more useful solution concept than local Nash equilibria for our purposes. However, neither is entirely satisfactory, and the question “What is the right solution concept for neural games?” remains open. In fact, it likely has many answers. The intrinsic curiosity module introduced by Pathak et al. (2017) plays two objectives against one another to drive agents to search for novel experiences. In this case, converging to a fixed point is precisely what is to be avoided.

It is remarkable – to give a few examples sampled from many – that curiosity, generating photorealistic images, and image-to-image translation (Zhu et al., 2017) can be formulated as games. What else can games do?

Acknowledgements.

We thank Guillaume Desjardins and Csaba Szepesvari for useful comments.

References

A TensorFlow code to compute SGA

Source code is available at https://github.com/deepmind/symplectic-gradient-adjustment. Since computing the symplectic adjustment is quite simple, we include an explicit description here for completeness.

The code requires a list of nn losses, Ls\mathtt{Ls}, and a list of variables for the nn players, xs\mathtt{xs}. The function fwd_gradients which implements forward mode auto-differentiation is in the module tf.contrib.kfac.utils.

% compute Jacobian-vector product Jv{\mathbf{J}}{\mathbf{v}}

def jac_vec(ys,xs,vs):\mathtt{\texttt{def jac\_vec}(ys,xs,vs):} return fwd_gradients(ys,xs,\texttt{}\hskip 11.38092pt\quad\mathtt{\texttt{return fwd\_gradients}(ys,xs,} grad_xs=vs,stop_gradients=xs)\texttt{}\mathtt{\texttt{grad\_xs}=vs,\texttt{stop\_gradients}=xs)}

% compute Jacobian⊺-vector product J⊺v{\mathbf{J}}^{\intercal}{\mathbf{v}}

def jac_tran_vec(ys,xs,vs):\mathtt{\texttt{def jac\_tran\_vec}(ys,xs,vs):} dydxs=tf.gradients(ys,xs,grad_ys=vs,\texttt{}\hskip 11.38092pt\mathtt{dydxs=\texttt{tf.gradients}(ys,xs,\texttt{grad\_ys}=vs,} stop_gradients=xs)\mathtt{\texttt{stop\_gradients}=xs)} return [tf.zeros_like(x) if dydx is None\texttt{}\hskip 11.38092pt\mathtt{\texttt{return }[\texttt{tf.zeros\_like}(x)\texttt{ if }dydx\texttt{ is }None} else dydx\texttt{}\hskip 54.06006pt\mathtt{\texttt{else }dydx} for (x,dydx) in zip(xs,dydxs)]\texttt{}\mathtt{\texttt{for }(x,dydx)\texttt{ in zip}(xs,dydxs)]}

% compute Symplectic Gradient Adjustment A⊺ξ{\mathbf{A}}^{\intercal}\xi

B Helmholtz, Hamilton, Hodge, and Harmonic games

This section explains the mathematical connections with the Helmholtz decomposition, symplectic geometry and the Hodge decomposition. The discussion is not necessary to understand the main text. It is also not self-contained. The details can be found in textbooks covering differential and symplectic geometry (Arnold, 1989; Guillemin and Sternberg, 1990; Bott and Tu, 1995).

The classical Helmholtz decomposition states that any vector field ξ\boldsymbol{\xi} in 3-dimensions is the sum of curl-free (gradient) and divergence-free (infinitesimal rotation) components:

We explain the link between curl and the antisymmetric component of the game Jacobian. Recall that gradients of functions are actually differential 1-forms, not vector fields. Differential 1-forms and vector fields on a manifold are canonically isomorphic once a Riemannian metric has been chosen. In our case, we are implicitly using the Euclidean metric. The antisymmetric matrix A{\mathbf{A}} is the differential 2-form obtained by applying the exterior derivative dd to the 1-form ξ\boldsymbol{\xi}.

In 3-dimensions, the Hodge star operator is an isormorphism from differential 2-forms to vector fields, and the curl can be reformulated as curl(∙)=∗d(∙)\text{curl}(\bullet)=*d(\bullet). In claiming A{\mathbf{A}} is analogous to curl, we are simply dropping the Hodge-star operator.

Finally, recall that the Lie algebra of infinitesimal rotations in dd-dimensions is given by antisymmetric matrices. When d=3d=3, the Lie algebra can be represented as vectors (three numbers specify a 3×33\times 3 antisymmetric matrix) with the ×\times-product as Lie bracket. In general, the antisymmetric matrix A{\mathbf{A}} captures the infinitesimal tendency of ξ\boldsymbol{\xi} to rotate at each point in the parameter space.

B.2 Hamiltonian Mechanics

The function is then referred to as the Hamiltonian function of the vector field. In our case, the antisymmetric matrix A{\mathbf{A}} is a closed 2-form because A=dξ{\mathbf{A}}=d\boldsymbol{\xi} and d∘d=0d\circ d=0. It may however be degenerate. It is therefore a presymplectic form (Bottacin, 2005).

Setting ω=A\omega={\mathbf{A}}, equation (58) can be rewritten in our notation as

justifying the terminology ‘Hamiltonian’.

B.3 The Hodge Decomposition

The exterior derivative dk:Ωk(M)→Ωk+1(M)d_{k}:\Omega^{k}(M)\rightarrow\Omega^{k+1}(M) is a linear operator that takes differential kk-forms on a manifold MM, Ωk(M)\Omega^{k}(M), to differential k+1k+1-forms, Ωk+1(M)\Omega^{k+1}(M). In the case k=0k=0, the exterior derivative is the gradient, which takes 0-forms (that is, functions) to 1-forms. Given a Riemannian metric, the adjoint of the exterior derivative δ\delta goes in the opposite direction. Hodge’s theorem states that kk-forms on a compact manifold decompose into a direct sum over three types:

Setting k=1k=1, we recover a decomposition that closely resembles the generalized Helmholtz decomposition:

B.4 Harmonic and Potential Games

Candogan et al. (2011) derive a Hodge decomposition for games that is closely related in spirit to our generalized Helmholtz decomposition – although the details are quite different. Candogan et al. (2011) work with classical games (probability distributions on finite strategy sets). Their losses are multilinear, which is easier than our setting, but they have constrained solution sets, which is harder in many ways. Their approach is based on combinatorial Hodge theory (Jiang et al., 2011) rather than differential and symplectic geometry. Finding a best-of-both-worlds approach that encompasses both settings is an open problem.

C Type Consistency

The next two sections carefully work through the units in classical mechanics and two-player games respectively. The third section briefly describes a use-case for type consistency.

where qq is position, p=μ⋅q˙p=\mu\cdot\dot{q} is momentum, μ\mu is mass, κ\kappa is surface tension and H{\mathcal{H}} measures energy. The units (denoted by τ\tau) are

where mm is meters, kgkg is kilograms and ss is seconds. Energy is measured in joules, and indeed it is easy to check that τ(H)=kg⋅m2s2\tau({\mathcal{H}})=\frac{kg\cdot m^{2}}{s^{2}}.

Note that the units for differentation by xx are τ(∂∂x)=1τ(x)\tau(\frac{\partial}{\partial x})=\frac{1}{\tau(x)}. For example, differentiating by time has units 1s\frac{1}{s}. Hamilton’s equations state that q˙=∂H∂p=1μ⋅p\dot{q}=\frac{\partial H}{\partial p}=\frac{1}{\mu}\cdot p and p˙=−∂H∂q=−κ⋅q\dot{p}=-\frac{\partial{\mathcal{H}}}{\partial q}=-\kappa\cdot q where

The resulting flow describing the dynamics of the system is

with units τ(ξ)=1s\tau(\boldsymbol{\xi})=\frac{1}{s}. Hamilton’s equations can be reformulated more abstractly via symplectic geometry. Introduce the symplectic form

Observe that contracting the flow with the Hamiltonian obtains

with units τ(dH)=τ(H)=kg⋅m2s2\tau(d{\mathcal{H}})=\tau({\mathcal{H}})=\frac{kg\cdot m^{2}}{s^{2}}.

Although there is no notion of “loss” in classical mechanics, it is useful (for the next section) to keep pushing the formal analogy. Define the “losses”

The duality between vector fields and differential forms.

Finally recall that the symplectic form in games was not “pulled out of thin air” as ω=dq∧dp\omega=dq\wedge dp, but rather derived as ω=dξ♭\omega=d\boldsymbol{\xi}^{\flat}, where ξ♭\boldsymbol{\xi}^{\flat} is the differential form corresponding to the vector field ξ\boldsymbol{\xi} under the musical isomorphism ♭:TM→T∗M\flat:TM\rightarrow T^{*}M.

It is instructive to compute ξ♭\boldsymbol{\xi}^{\flat} in the case of a classical mechanical system and see what happens. Naively, we would guess that the musical isomorphism is (∂∂q)♭=dq\left(\frac{\partial}{\partial q}\right)^{\flat}=dq and (∂∂p)♭=dp\left(\frac{\partial}{\partial p}\right)^{\flat}=dp. However, applying the naive musical isomorphism to ξ\boldsymbol{\xi} to get

and we cannot add objects with different types.

To correct the type inconsistency, define the musical isomorphism as

The correction terms in the direction ♭:TM→T∗M\flat:TM\rightarrow T^{*}M invert the coupling terms κ\kappa and 1μ\frac{1}{\mu} that were originally introduced into the Hamiltonian for physical reasons. Applying the corrected musical isomorphism to ξ\boldsymbol{\xi} yields

The two terms of ξ♭\boldsymbol{\xi}^{\flat} then have coherent types

which recovers the symplectic form (up to sign) with units τ(ω)=kg⋅m2s\tau(\omega)=\frac{kg\cdot m^{2}}{s} as required. Finally, observe that

C.2 Units in Two-Player Games

Without loss of generality let w=(x;y){\mathbf{w}}=({\mathbf{x}};{\mathbf{y}}) where we refer to x{\mathbf{x}} as position and y{\mathbf{y}} as momentum so that τ(x)=m\tau({\mathbf{x}})=m and τ(y)=kg⋅ms\tau({\mathbf{y}})=\frac{kg\cdot m}{s}. The aim of this section is to check type-consistency under these, rather arbitrarily assigned, units. Since we are considering a game, we do not require that x{\mathbf{x}} and y{\mathbf{y}} have the same dimension – even though this would necessarily be the case for a physical system. The goal is to verify that units can be consistently assigned to games.

Consider a quadratic two player game of the form

The presymplectic form ω=dξ♭\omega=d\boldsymbol{\xi}^{\flat} makes use of the musical isomorphism ♭:TM→T∗M\flat:T^{M}\rightarrow T^{*}M. As in section C.1, if we naively define (∂∂x)♭=dx(\frac{\partial}{\partial{\mathbf{x}}})^{\flat}=d{\mathbf{x}} and (∂∂y)♭=dy(\frac{\partial}{\partial{\mathbf{y}}})^{\flat}=d{\mathbf{y}} then

It is necessary, as in section C.1, to correct the naive musical isomorphism by taking into account the coupling constants for the mixed position-momentum terms. In the classical setup the coupling constants were the scalars 1μ\frac{1}{\mu} and κ\kappa, whereas in a game they are the off-diagonal blocks A12{\mathbf{A}}_{12} and C12{\mathbf{C}}_{12}.

Apply singular value decomposition to factorize

where the entries of the diagonal matrices have types τ(DA)=1kg\tau({\mathbf{D}}_{\mathbf{A}})=\frac{1}{kg} and τ(DC)=kgs2\tau({\mathbf{D}}_{\mathbf{C}})=\frac{kg}{s^{2}}, and the types of the orthogonal matrices U{\mathbf{U}} and V{\mathbf{V}} are pure scalars. The diagonal matrices DA{\mathbf{D}}_{\mathbf{A}} and DC{\mathbf{D}}_{\mathbf{C}} have the same types as 1μ\frac{1}{\mu} and κ\kappa in the classical system since they play the same coupling role.

Extending the procedure adopted in the section C.1, fix the type-inconsistency by defining the musical isomorphisms as

Alternatively, the isomorphisms can be computed by noting that UA⊺DA−1UA=(A12A21)−1{\mathbf{U}}_{\mathbf{A}}^{\intercal}{\mathbf{D}}_{\mathbf{A}}^{-1}{\mathbf{U}}_{\mathbf{A}}=(\sqrt{{\mathbf{A}}_{12}{\mathbf{A}}_{21}})^{-1} and VC⊺DC−1VC=(C21C12)−1{\mathbf{V}}_{\mathbf{C}}^{\intercal}{\mathbf{D}}_{\mathbf{C}}^{-1}{\mathbf{V}}_{\mathbf{C}}=(\sqrt{{\mathbf{C}}_{21}{\mathbf{C}}_{12}})^{-1}.

The dual isomorphism ♯:T∗M→TM\sharp:T^{*}M\rightarrow TM is then

where the notation ωτ\omega_{\tau} emphasizes that the two-form is type-consistent.

C.3 What Does Type-Consistency Buy?

Although ξ\boldsymbol{\xi} is not a potential field, there is a family of functions on which ξ\boldsymbol{\xi} performs gradient descent – albeit with coordinate-wise learning rates that may not be optimal. The vector field ξ\boldsymbol{\xi} arguably does not require adjustment. This kind of situation often arises when the learning rates of different parameters are set adaptively during training of neural nets, by rescaling them by positive numbers.

The vanilla and type-consistent 1-forms corresponding to ξ\boldsymbol{\xi} are, respectively,

It follows that the type-consistent symplectic gradient adjustment is zero. Type-consistency ‘detects’ that no gradient adjustment is needed in example 10.