Euclideanizing Flows: Diffeomorphic Reduction for Learning Stable Dynamical Systems

Muhammad Asif Rana, Anqi Li, Dieter Fox, Byron Boots, Fabio Ramos, Nathan Ratliff

Introduction

In many applications, robots are required to execute complex motions in potentially dynamic and unstructured environments. Since hand-coding such motions can be cumbersome or even infeasible, learning from demonstration (LfD) enables robots to, instead, acquire new motion skills by observing humans. In this work, we focus on learning goal-directed motions from human demonstrations, i.e. motions that stop at a given target location. This is without loss of generality since complex tasks can often be achieved by an ordered execution of goal-directed motions .

Human motions naturally preserve regularity properties, e.g. continuity, smoothness and boundedness . Since we assume goal-directed motions, human demonstrations should also naturally be stable. Exploiting these regularity properties can significantly improve the sample efficiency of the learning algorithms, which is critical since the number of human demonstrations are often limited. Thus, we model human motions as rollouts from a stable dynamical system. Motions governed by stable dynamical systems exhibit some additional desirable properties. First, such motions can react to temporal and spatial perturbations, which is necessary in dynamic task settings. Second, the stability property formally guarantees that the motions converge to the goal region.

The challenge, however, is to encode stable human motions into dynamical systems that are stable by construction. A number of approaches have been proposed, which address this problem by explicitly parameterizing the class of stable dynamical systems and the notions of stability [4; 5; 6]. In this work, we take a fundamentally different approach than the aforementioned methods. Instead of explicitly learning a stable dynamical system, we view demonstrations as motions on a Riemannian manifold which is linked, under a smooth bijective map, i.e. a diffeomorphism, to a latent Euclidean space. This diffeomorphism implicitly gives rise to an inherently stable dynamical system. The learning problem thus involves finding a diffeormorphism which explains the observed demonstrations. Compared to existing diffeomorphism learning approaches to encoding motions [7; 8], our formulation, Stable Dynamical System learning using Euclideanizing Flows (SDS-EF\xspace), is based on a more expressive formulation of diffeomorphisms.

We present an approach for learning a time-invariant continuous-time dynamical system (or reactive motion policy), which is globally asymptotically stable. Our dynamics formulation allows encoding severely curved goal-directed motions, and can be learned from a few demonstrations with minimal parameter tuning. Our specific contributions include: (i) a formulation of stable dynamics through warping curves into simple motions on a latent space using diffeomorphisms, and (ii) an expressive class of diffeomorphisms suitable for learning stable and smooth dynamical systems. We demonstrate the effectiveness of our approach on a standard handwriting dataset , and data collected on a robot manipulator .

Related Work

Over the past decade, a number of approaches have been proposed towards learning goal-directed time-invariant dynamical systems from human demonstrations . An early approach is SEDS . SEDS assumes the demonstrations comply with a Lyapunov (or Energy) function given by the squared-distance to the goal, and thus is restricted to motions which monotonically converge to the goal over time. A relaxation to the stability criterion was proposed in CLF-DM , which instead assumes Lyapunov functions in the form of a weighted sum of asymmetric quadratic functions (WSAQF). The parameters of the Lyapunov function are learned independently from the (unstable) dynamics, and then used in an online fashion to generate stabilizing controls. The online correction scheme may interfere significantly with the learned dynamics . A more recent approach, CDSP instead enforces incremental stability, a notion concerned with relative displacement of motions. Instead of learning a Lyapunov function, CDSP proposed learning a positive-definite contraction metric. However, CDSP restricts the class of contraction metrics to (potentially limiting) sum of squared polynomials.

There are a few approaches that learn stable dynamical systems via learned diffeomorphisms. However, by definition, a diffeomorphism must be invertible, making the learning problem non-trivial. One approach , realizes a diffeomorphism as a composition of locally weighted translations. The authors only apply this approach to the problem of learning a single (or average) demonstration. Another strategy, τ\tau-SEDS , learns diffeomorphisms from multiple demonstrations. However, τ\tau-SEDS defines a diffeomorphism by the square-root of a WSAQF, thus restricting the hypothesis class. In contrast to this prior work, we learn diffeomorphisms by using flexible function approximators including kernel methods and neural networks .

Our diffeomorphism learning approach builds on normalizing flows [15; 16], which have recently been successfully used for density estimation. The goal of normalizing flows is to map a simple base distribution into a complicated probability distribution over observed data by applying the change of variable theorem sequentially. Our problem is similar to the normalizing flows problem as we seek to map simple straight-line motions to complicated motions captured from human demonstrations. However, our problem is fundamentally different in two ways. First, normalizing flows only require the mapping to be bijective, a property less strict than diffeomorphism (see Section 3). Second, normalizing flows are concerned primarily with mapping scalar functions to scalar functions (i.e. probability densities). In our method, we seek to map vector fields to vector fields.

Background: Stability, Diffeomorphism, and Riemannian Manifolds

We briefly summarize the theoretical background for this paper. First, we introduce the concept of global asymptotic stability and how it can be shown through an auxiliary function, i.e. a Lyapunov function. Next, we discuss how one dynamical system can be described in different ways through a change of coordinates using a diffeomorphism. Finally, we shed light into the geometric interpretation of diffeomorphisms, especially when applied to gradient descent dynamical systems.

2 Change of Coordinates for Dynamical Systems

where Jψ(x)=∂ψ∂x\mathbf{J}_{\psi}(\mathbf{x})=\frac{\partial\psi}{\partial\mathbf{x}} is the Jacobian matrix of ψ\psi. In the special case where both the system dynamics ff and the diffeomorphism ψ\psi are linear maps, this change of coordinates reduces to change of basis for linear dynamical systems .

3 Riemannian Manifolds and Natural Gradient Descent

Figure 4(a)–(b) shows the iso-contours of a potential function Φ\Phi which generate straight-line motions on the Euclidean space, and the corresponding potential function Φ∘ψ\Phi\circ\psi on the Riemannian manifold. The isocontours show equidistant points to the goal in the Riemannian manifold, thus revealing the underlying geometry. Figure 4(c)–(d) shows the trajectories generated by the same underlying system observed in the corresponding coordinate systems.

Learning Stable Dynamics Using Diffeomorphisms

We view human demonstrations as goal-directed motions on an nn-dimensional Riemannian manifold, governed by a stable dynamical system of the form (4). We view this dynamical system to be equivalent, under a change of coordinates, to another system defined on a latent space. Our key insight here is: a diffeomorphism can warp a simple potential function into a more complicated one, and hence transform straight-line trajectories into severely curved motions. As a result, the problem of learning stable dynamical systems reduces to a diffeomorphism learning problem.

Although both the diffeomorphism ψ\psi and the potential function Φ\Phi can shape the dynamical system (4), we view them as playing fundamentally different roles: the potential function Φ\Phi dictates the theoretical property of the dynamics, e.g. stability guarantees, while the diffeomorphism ψ\psi provides expressivity to our hypothesis class. As a result, we specify a simple potential function Φ(y)=∥y−y∗∥\Phi(\mathbf{y})=\|\mathbf{y}-\mathbf{y}^{*}\|, with y∗=ψ(x∗)\mathbf{y}^{*}=\psi(\mathbf{x}^{*}). This potential function generates unit-velocity straight-line motions to the globally asymptotically stable equilibrium point y∗\mathbf{y}^{*}. The diffeomorphism acts to deform these straight lines to arbitrarily curved motions converging to x∗\mathbf{x}^{*}, where the demonstrations converge. With a parameterized diffeomorphism ψθ\psi_{\theta}, the learning problem reduces to solving,

2 A Class of Expressive Diffeormorphisms

In contrast to , for the scaling and translation functions, we use single-layer neural networks with the layer resembling an approximated kernel machine . This special network structure leverage the advantages of both kernel machines and neural networks: (i) the kernel functions act as regularizers to enforce desired properties on the learned function, e.g. smoothness, which can further improve the sample efficiency of the learning algorithm, and (ii) as a parameterized model, the composed diffeomorphism can be efficiently trained using learning techniques for neural networks.

We use a matrix-valued Gaussian separable kernel with length-scale ll, defined as K(z,z′)=exp⁡(−∥z−z′∥22l2)IK(\mathbf{z},\mathbf{z}^{\prime})=\exp(-\frac{\|\mathbf{z}-\mathbf{z}^{\prime}\|^{2}}{2l^{2}})\mathbf{I}. The Gaussian kernel restricts the hypothesis class to the class of C∞C^{\infty} vector-valued functions, imposing a stricter smoothness constraint than continuous differentiability, i.e. C1C^{1}. This is desirable since human motions are known to be maximizing smoothness (or minimizing jerk) . We approximate the aforementioned kernel by mm randomly sampled Fourier features , such that the scaling and translation functions are given by linear combinations of these features,

3 Practical Considerations

To learn (5) through back-propogation, the Jacobian of the diffeomorphism can be calculated analytically from (6) and (7), or through auto-differentiation packages, e.g. pytorch .

Experimental Results

The LASA dataset consists of a library of 30 two-dimensional handwritten letters, each with 7 demonstrations. The scale of the dataset is 100mm×100mm100mm\times 100mm. For each letter, we find a diffeomorphism using (5). Fig. 8 shows the vector fields as governed by (4) on a subset of letters.

In all the plots, the rollouts (in red) closely match the demonstrations (in white), coming to rest at the goal. Furthermore, due to the structure imposed in learning, the dynamical system generalizes smooth and stable motions throughout the state-space. Also shown in Fig. 7 are isocontours of the potential function Φ(ψ(x))\Phi({\psi}(x)). For quantitative evaluations, we employ three error metrics: root mean squared error (RMSE), dynamic time warping distance (DTWD) , and Frechet disance (FD) . These metrics evaluate performance of our approach in terms of its capability to reproduce the demonstrated motions. Fig. 6 reports these metrics, in millimeters, evaluated over 210210 demonstrations (77 letters ×\times 3030 demonstrations). Each aforementioned metric focuses on different aspects of the motions. RMSE penalizes both spatial and temporal misalignment between demonstrated and reproduced motions. On the other hand, DTWD and FD disregard time misalignment, and instead focus solely on the spatial misalignment between motions. Since DTWD between any two trajectories is a time-aggregated measure of error, we report average DTWD, found by dividing the DTWD by the number of points TiT_{i} in a trajectory. In Fig. 6, the median and mean errors are observed to be small relative to the scale of the data, signifying that the learned dynamical systems are able to accurately reproduce most motions. However, higher errors are occasionally observed, accounting for outliers. This is mostly due to intersecting demonstrations which can not be modeled by a first-order dynamical system.

For demonstrations collected on a Franka Emika robot, we evaluate on two tasks: door reaching, and drawer closing . Each task dataset consists of 6 three-dimensional end-effector motions collected by physically guiding the robot. The door reaching task required the robot to start from inside a cabinet and reach the door handle, while the drawer closing task required reaching a drawer handle and pushing the drawer close. Fig. 9 shows the tasks, demonstrated motions (blue), as well as the reproduced motions (red). Our learning approach is observed to accurately reproduce three-dimensional motions collected on a real robot.

Conclusion

We have presented SDS-EF\xspace, an approach for learning motion skills from a few human demonstrations. SDS-EF\xspaceencodes complex human motions as generated from a dynamical system, linked under a learnable diffeomorphism, to a simple gradient-descent dynamical system on a latent space. A class of parameterized diffeomorphisms is proposed to learn a wide range of motions and generalize to different tasks with minimal parameter tuning. Experimental validation on a handwriting dataset and data collected on a real robot is provided to show the efficacy of the proposed approach.

This work was supported in part by NVIDIA Research.

References