Learning Stable Deep Dynamics Models
Gaurav Manek, J. Zico Kolter
Introduction
In this paper, we propose an approach to learning neural network dynamics that are provably stable over the entirety of the state space. To do so, we jointly learn the system dynamics and a Lyapunov function. This stability is a hard constraint imposed upon the model: unlike recent approaches, we do not enforce stability via an imposed loss function but build it directly into the dynamics of the model (i.e. Even a randomly initialized model in our proposed model class will be provably stable everywhere in state space). The key to this is the design of a proper Lyapunov function, based on input convex neural networks , which ensures global exponential stability to an equilibrium point while still allowing for expressive dynamics.
Using these methods, we demonstrate learning dynamics of physical models such as -link pendulums, and show a substantial improvement over generic networks. We also show how such dynamics models can be integrated into larger network systems to learn dynamics over complex output spaces. In particular, we show how to combine the model with a variational auto-encoder (VAE) to learn dynamic “video textures” .
Background and related work
Showing that these conditions imply the various forms of stability is relatively straightforward, but showing the converse (that any stable system must obey this property for some ) is relatively more complex. In this paper, however, we are largely concerned with the “simpler” of these two directions, as our goal is to enforce conditions that ensure stability.
Stability of linear systems.
it is well-established that the system is stable if and only if the real components of the the eigenvalues of are all strictly negative (). Equivalently, the same same property can be shown via a positive definite quadratic Lyapunov function
for . In this case, by Equation 4, the following ensures stability:
i.e., if we can find a positive definite matrix with that negative semidefinite. Such bounds (and much more complex extensions) for the basis for using linear matrix inequalities (LMIs), as a method to ensure stability of linear dynamical systems. The methods also have applicability to non-linear systems, and several authors have used LMI analysis to learn non-linear dynamical systems by constraining the linearization of the systems to have global Lyapunov functions ,
The point we want to emphasize from the above discussion, though, is that the task of learning even a stable linear dynamical system is not a convex problem. Although the constraints
are convex in and separately, they are not convex in and jointly. Thus, the problem of jointly learning a stable linear dynamical system and its corresponding Lyapunov function, even for the simple linear-quadratic setting, is not a convex optimization problem, and alternative techniques such as alternating minimization need to be employed instead. Alternatively, past work has also looked at different heuristics, such as approximately projecting a dynamics function onto the (non-convex) stable set of matrices with eigenvalues .
Stability of non-linear systems
For general non-linear systems, establishing stability via Lyapunov techniques is typically even more challenging. For the typical task here, which is that of establishing stability of some known dynamics , finding a suitable Lyapunov function is often more an art than a science. Although some general techniques such as sum-of-squares certification provide general methods for certifying stability of e.g., polynomial systems, these are often expensive and don’t easily scale to high dimensional systems.
Notably, our proposed approach here is able to learn provably stable systems without solving this (generally hard) problem. Specifically, while it is difficult to find a Lyapunov function that certifies the stability of some known system, we exploit the fact that it is relatively much easier to enforce some function to behave in a stable manner according to a Lyapunov function.
Lyapunov functions in deep learning
Finally, there has been a small set of recent work exploring the intersection of deep learning and Lyapunov analysis . Although related to our work here, the approach in this past work is quite different. As is more common in the control setting, these papers try to learn neural-network-based Lyapunov functions for control policies, but in way that enforces stability via a loss penalty. For instance Richards et al., optimize a loss function that encourages for in some training set. In contrast, our work guarantees absolute stability everywhere in the state space, not just at a small set of points; but only for a simpler setting where the entire dynamics are to be learned (and hence can be “forced” to be stable) rather than a stabilizing controller for known dynamics.
Joint learning of dynamics and Lyapunov functions
The intuition of the approach we propose in this paper is straightforward: instead of learning a dynamics function and attempting to separately verify its stability via a Lyapunov function, we propose to jointly learn a dynamics model and Lyapunov function, where the dynamics is inherently constrained to be stable (everywhere in the state space) according to the Lyapunov function.
where denotes the orthogonal projection of onto the point , and where the second equation follows from the analytical projection of a point onto a halfspace. As long as is defined using automatic differentiation tools, it is straightforward to include the gradient terms into the definition of , and our final network can be trained just like any other function. The general approach here is illustrated in Figure 1.
Although the treatment above seems to make the problem of learning stable systems quite straightforward, the sublety of the approach lies in the choice of the function . Specifically, as mentioned previously, needs to be positive definite, but additionally needs to have no local optima except . This is due to Lyapunov decrease condition: recall that we are attempting to guarantee stability to the equilibrium point , yet the decrease condition imposed upon the dynamics means that is decreasing along trajectories of . If has a local optimum away from the origin, the dynamics can in theory get stuck in this location; this manifests itself by the term going to zero, which results in the dynamics becoming undefined at the optima.
To enforce these conditions, we make the following design decisions regarding :
We represent via an input-convex neural network (ICNN) function , which enforces the condition that be convex in its inputs . A fairly generic form of such networks consists is given by the recurrence
where are real-valued weights mapping from inputs to the layer activations; are positive weights mapping previously layer activations to the next layer; are real-valued biases;and are convex, monotonically non-decreasing non-linear activations, such as the ReLU or smooth variants. It is straightforward to show that with this formulation, is convex in , and indeed any convex function can be approximated by such networks .
Positive definite.
While the ICNN property can enforce that have only a single global optima, it does not necessarily enforce that this optima be at . While one could fix this by e.g., removing the biases term (but this imposes substantial limitations on the representable functions, which can no longer be arbitrary convex functions) or by shifting whatever global minima exists to the origin (but this requires finding the global minimum during training, which itself is computationally expensive), we take an alternative approach and simply shift the function such that , and add a small quadratic regularization term to ensure strict positive definiteness.
where is a positive convex non-decreasing function with , is the ICNN defined previously, and is a small constant. These terms together still enforce (strong) convexity and positive definiteness of .
Continuously differentiable.
Although not always required, several of the conditions for Lyapunov stability are simplified is is continuously differentiable. To achieve this, rather than use ReLU activationsNote that the typical softplus smoothed approximation of the ReLU will not work for all purposes above, since we require an activation with , we use a smoothed version that replaces the purely linear ReLU with a quadratic region in
An illustration of this activation is shown in Figure 2.
(Optional) Warped input space.
as the Lyapunov function. Invertibility ensures that the sublevel sets of (which are convex sets, by definition) map to contiguous regions of the composite function , thus ensuring that no local optima exist in this composed function.
With these conditions in place, we have the following result.
defined by from (10) and from (12) or (14) are globally exponentially stable to the equilibrium point , for any (bounded weight) networks defining the and functions.
The proof is straightforward, and relies on the properties of the networks created above. First, note that by our definitions we have, for some ,
where the lower bound follows by definition and the fact that is positive. The upper bound follows from the fact that the activation as defined is linear for large and quadratic around 0. This fact in turn implies that behaves linearly as , and is quadratic around the origin, so can be upper bounded by some quadratic .
The fact the is continuously differentiable means that (in ) is defined everywhere, bounds on for all follows from the the Lipschitz property of , the fact that , and the term
where denotes the operator norm when applied to a matrix. This implies that the dynamics are defined and bounded everywhere owing to the choice of function .
Now, consider some initial state . The definition of implies that
Integrating this equation gives the bound
and applying the lower and upper bounds gives
as required for global exponential convergence. ∎
Empirical results
We illustrate our technique on several example problems, first highlighting the (inherent) stability of the method for random networks, demonstrating learning on simple -link pendulum dynamics, and finally learning high-dimensional stable latent space dynamics for dynamic video textures via a VAE model.
Although we mention this only briefly, it is interesting to visualize the dynamics created by random networks according to our process, i.e., before any training at all. Because the dynamics models are inherently stable, these random networks lead to stable dynamics with interesting behaviors, illusrated in Figure 3. Specifically, we let be defined by a 2-100-100-2 fully connected network, and be a 2-100-100-1 ICNN, with both networks initialized via the default weights of PyTorch (the Kaiming uniform initialization ) and with the ICNN having it’s weights further put through a softplus unit to make them positive.
2 n𝑛n-link pendulum
Next we look at the ability of our approach to model a physically-based dynamical system, specifically the -link pendulum. A damped, rigid -link pendulum’s state can be described by the angular position and angular velocity of each link . As before is a -100-100- network, and the Lyapunov function is a -60-60-1 ICNN with properties described in Section 3.1. Models are trained with pairs of data produced by the symbolic algebra solver sympy, using simulation code adapted from .
In Figure 4, we compare the simulated dynamics with the learned dynamics in the case of a simple damped pendulum (i.e. with ), showing both the streamplot of the vector field and a single simulated trajectory, and draw a contour plot of the learned Lyapunov function. As seen, the system is able to learn dynamics that can accurately predict motion of the system even over long time periods.
We also evaluate the learned dynamics quantitatively varying and the time horizon of simulation. Figure 5 presents the total error over time for the 8-link pendulum, and the average cumulative error over 1000 time steps for different values of . While both the simple and our stable models show increasing mean error at the start of the trajectory, our model is able to capture the contraction in the physical system (implied by conservation of energy) and in fact exhibits decreasing error towards the end of the simulation (the true and simulated dynamics are both stable). In comparison, the error in the simple model increases.
3 Video Texture Generation
We train the model on pairs of successive frames sampled from videos. To generate video textures, we seed the dynamics model with the encoding of a single frame and numerically integrate the dynamics model to obtain a trajectory. The VAE decoder converts each step of the trajectory into a frame. In Figure 7, we present sample stable trajectories and frames produced by our network. For comparison, we also include an example trajectory and resulting frames when the dynamics are modelled without the stability constraint (i.e. letting in the above loss be a generic neural network). For the naive model, the dynamics quickly diverge and produce a static image, whereas for our approach, we are able to generate different (stable) trajectories that keep generating realistic images over long time horizons.
Conclusion
In this paper we proposed a method for learning stable non-linear dynamical systems defined by neural network architectures. The approach jointly learns a convex positive definite Lyapunov function along with dynamics constrained to be stable according to these dynamics everywhere in the state space. We show that these models can be integrated into other deep architectures such as VAEs, and learn complex latent space dynamics is a fully end-to-end manner. Although we have focused here on the autonomous (i.e., uncontrolled) setting, the method opens several directions for future work, such as integration into dynamical systems for control or reinforcement learning settings. Have stable systems as a “primitive” can be useful in a large number of contexts, and combining these stable systems with the representational power of deep networks offers a powerful tool in modeling and controlling dynamical systems.