Path Integral Networks: End-to-End Differentiable Optimal Control
Masashi Okada, Luca Rigazio, Takenobu Aoshima
Introduction
Recently, deep architectures such as convolutional neural networks have been successfully applied to difficult control tasks such as autonomous driving , robotic manipulation and playing games . In these settings, a deep neural network is typically trained with reinforcement learning or imitation learning to represent a control policy which maps input states to control sequences. However, as already discussed in , the resulting networks and encoded policies are inherently reactive, thus unable to execute planning to decide following actions, which may explain poor generalization to new or unseen environments. Conversely, optimal control algorithms utilize specified models of system dynamics and a cost function to predict future states and future cost values. This allows to compute control sequences that minimize expected cost. Stated differently, optimal control executes planning for decision making to provide better generalization.
The main practical challenge of optimal control is specifying system dynamics and cost models. Model-based reinforcement learning can be used to estimate system dynamics by interacting with the environment. However in many robotic applications, accurate system identification is difficult. Furthermore, predefined cost models accurately describing controller goals are required. Inverse optimal control or inverse reinforcement learning estimates cost models from human demonstrations , but require perfect knowledge of system dynamics. Other inverse reinforcement learning methods such as do not require system dynamics perfect knowledge, however, they limit the policy or cost model to the class of time-varying linear functions.
In this paper, we propose a new approach to deal with these limitations. The key observation is that control sequences resulting from a specific optimal control algorithm, the path integral control algorithm , are differentiable with respect to all of the controller internal parameters. The controller itself can thus be represented by a special kind recurrent network, which we call path integral network (PI-Net). The entire network, which includes dynamics and cost models, can then be trained end-to-end using standard back-propagation and stochastic gradient descent with fully specified or approximated cost models and system dynamics. After training, the network will then execute planning by effectively running path integral control utilizing the learned system dynamics and cost model. Furthermore, the effect of modeling errors in learned dynamics can be mitigated by end-to-end training because cost model could be trained to compensate the errors.
We demonstrate the effectiveness of PI-Net by training the network to imitate optimal controllers of two control tasks: linear system control and pendulum swing-up task. We also demonstrate that dynamics and cost models, latent in demonstrations, can be adequately extracted through imitation learning.
Path Integral Optimal Control
Path integral control provides a framework for stochastic optimal control based on Monte-Carlo simulation of multiple trajectories . This framework has generally been applied to policy improvements for parameterized policies such as dynamic movement primitives . Meanwhile in this paper, we focus on a state-of-the-art path-integral optimal control algorithm developed for model predictive control (MPC; a.k.a. receding horizon control). In the rest of this section, we briefly review this path integral optimal algorithm.
Eq. (3) can be implemented on digital computers by approximating the expectation value with the Monte Carlo method as shown in Alg. 1.
Different from other general optimal control algorithms, such as iterative linear quadratic regulator (iLQR) , path integral optimal control does not require first or second-order approximation of the dynamics and a quadratic approximation of the cost model, naturally allowing for non-linear system dynamics and cost models. This flexibility allows us to use general function approximators, such as neural networks, to represent dynamics and cost models in the most general possible form.
Path Integral Networks
We illustrate the architecture of PI-Net in Figs. 1 (a)-(d). The architecture encodes Alg. 1 as a fully differentiable recurrent network representation. Namely, the forward pass of this network completely imitates the iterative execution of Alg. 1.
PI-Net Kernel in Fig. 1 (b) contains three modules: Noise Generator, Monte-Carlo Simulator and Control Sequence Updater. First, the Noise Generator procures Gaussian noise vectors sampled from . Then the noise vectors are input to Monte-Carlo Simulator along with and , which estimates running- and terminal-cost values (denoted as ) of different trajectories. Finally, the estimated cost values are fed into the Control Sequence Updater to improve the initial control sequence.
2 Learning schemes
We remark that all the nodes in the computational graph of PI-Net are differentiable. We can therefore employ the chain rule to differentiate the network end-to-end, concluding that PI-Net is fully differentiable. If an objective function with respect to the network control output, denoted as , is defined, then we can differentiate the function with the internal parameters (). Therefore, we can tune the parameters by optimizing the objective function with gradient descent methods. In other words, we can train internal dynamics and/or cost models end-to-end through the optimization. For the optimization, we can re-use all the standard Deep Learning machinery, including back-propagation and stochastic gradient descent, and a variety of Deep Learning frameworks. We implemented PI-Net with TensorFlow . Interestingly, all elemental operations of PI-Net can be described as TensorFlow nodes, allowing to utilize automatic differentiation.
A general use case of PI-Net is imitation learning to learn dynamics and cost models latent in experts’ demonstrations. Let us consider an open loop control setting and suppose that a dataset is available; is a state observation and is a corresponding control sequence generated by an expert. In this case, we can supervisedly train the network by optimizing , i.e., the errors between the expert demonstration and the network output . For closed loop infinite time horizon control setting, the network can be trained as an MPC controller. If we have a trajectory by an expert , we can construct a dataset and then optimize the estimation errors between the expert control and the first value of output control sequence output. If sparse reward function is available, reinforcement learning could be introduced to train PI-Net. The objective function here is expected return which can be optimized by policy gradient methods such as REINFORCE .
Loss functions In addition to , we can append other loss functions to make training faster and more stable. In an MPC scenario, we can construct a dataset in another form . In this case, a loss function with respect to internal dynamics output can be introduced; i.e., state prediction errors between and . Furthermore, we can employ loss functions regarding cost models. In many cases on control domains, we know goal states in prior and we can assume cost models have optimum points at . Therefore, loss functions, which penalize conditions of , can be employed to help the cost models have such property. This is a useful approach when we utilize highly expressive approximators (e.g., neural networks) to cost models. In the later experiments, mean squared error (MSE) was used for . was defined as , where is the ramp function. The sum of these losses can be jointly optimized in a single optimization loop. Of course, dynamics model can be pre-trained independently by optimizing .
3 Discussion of computational complexity
The complexity can be alleviated by data parallel approach, in which a mini-batch is divided and processed in parallel with distributed computers. Therefore, we can reduce the batch size processed on a single computer. Another possible approach is to reduce ; the recurrence number of the PI-Net Kernel module. In the experiment, initial control sequence is filled with a constant value (i.e., zero) and is set to be large enough (e.g., ). In our preliminary experiment, we found that inputting desired output (i.e., demonstrations) as initial sequences and training PI-Net with small did not work; the trained PI-Net just passed through the initial sequence, resulting in poor generalization performance. In the future, a scheme to determine good initial sequences, which reduces while achieving good generalization, must be established.
Note that the memory problem is valid only during training phase because the network does not need to store input values during control phase. In addition, the mini-batch size is obviously in that phase. Further in MPC scenarios, we can employ warm start settings to reduce , under which output control sequences are re-used as initial sequences at next timestep. For instance in , real-time path integral control has been realized by utilizing GPU parallelization.
Related Work
Conceptually PI-Net is inspired by the value iteration network (VIN) , a differentiable network representation of the value iteration algorithm designed to train internal state-transition and reward models end-to-end. The main difference between VIN and PI-Net lies in the underlying algorithms: the value iteration, generally used for discrete Markov Decision Process (MDP), or path integral optimal control, which allows for continuous control. In , VIN was applied to 2D navigation task, where 2D space was discretized to grid map and a reward function was defined on the discretized map. In addition, action space was defined as eight directions to forward. The experiment showed that adequately estimated reward map can be utilized to navigate an agent to goal states by not only discrete control but also continuous control. Let us consider a more complex 2D navigation task on continuous control, in which velocity must be taken into reward functionSuch as a task to control a mass point to trace a fixed track while forwarding it as fast as possible.. In order to design such the reward function with VIN, 4D state space (position and velocity) and 2D control space (vertical and horizontal accelerations) must be discretized. This kind of discretization could cause combinatorial explosion especially for higher dimensional tasks.
Generally used optimal controller, linear quadratic regulator (LQR), is also differentiable and Ref. employs this insight to re-shape original cost models to improve short-term MPC performance. The main advantage of the path integral control over (iterative-)LQR is that we do not require a linear and quadratic approximation of non-linear dynamics and cost model. In order to differentiate iLQR with non-linear models by back-propagation, we must iteratively differentiate the functions during preceding forward pass, making the backward pass very complicated.
Policy Improvement with Path Integrals and Inverse Path Integral Inverse Reinforcement Learning are policy search approaches based on the path integral control framework, which train a parameterized control policy via reinforcement learning and imitation learning, respectively. These methods have been succeeded to train policies for complex robotics tasks, however, they assume trajectory-centric policy representation such as dynamic movement primitives ; such the policy is less generalizable for unseen settings (e.g., different initial states).
Since PI-Net is a policy representation of optimal control, trainable end-to-end by standard back-propagation, a wide variety of learning to control approaches may be applied, including:
Reinforcement learning Deep Deterministic Policy Gradient , A3C , Trust Region Optimization , Guided Policy Search and Path Integral Guided Policy Search .
Experiments
We conducted experiments to validate the viability of PI-Net. These experiments are meant to test if the network can effectively learn policies in an imitation learning setting. We did this by supplying demonstrations generated by general optimal control algorithms, with known dynamics and cost models (termed teacher models in the rest of this paper). Ultimately, in real application scenarios, demonstrations may be provided by human experts.
The objective of this experiment is to validate that PI-Net is trainable and it can jointly learn dynamics and cost models latent in demonstrations.
PI-Net settings Internal dynamics and cost models were also linear and quadratic form whose initial parameters were different from the teacher models’. PI-Net was supervisedly trained by optimizing . We did not use and in this experiment.
Results Fig. 2 shows the results of this experiments. Fig. 2 (a) illustrates loss during training epochs, showing good convergence to a lower fixed point. This validates that PI-Net was indeed training well and the trained policy generalized well to test samples. Fig. 2 (b, c) exemplifies state and cost trajectories predicted by trained dynamics and cost models, which were generated by feeding certain initial state and corresponding optimal control sequence into the models. Fig. 2 (c) are trajectories by the teacher models. State trajectories in Fig. 2(b, c) approximate each other, indicating the internal dynamics model learned to approximate the teacher dynamics. It is well-known that different cost functions could result in same controls , and indeed cost trajectories in Fig. 2(b, c) seem not similar. However, this would not be a problem as long as a learned controller is well generalized to unseen state inputs.
2 Pendulum swing up
Next, we tried to imitate demonstrations generated from non-linear dynamics and non-quadratic cost models while validating the PI-Net applicability to MPC tasks. We also compared PI-Net with VIN.
Demonstrations The experiment focuses on the classical inverted pendulum swing up task . Teacher cost models were , where is pendulum angle and is angular velocity. Under this model assumptions, we firstly conducted 40s MPC simulations with iLQR to generate trajectories. We totally generated trajectories and then produced , where control is torque to actuate a pendulum. We also generated trajectories for .
PI-Net settings We used a more general modeling scheme where internal dynamics and cost models were represented by neural networks, both of which had one hidden layer. The number of hidden nodes was 12 for the dynamics model, and 24 for the cost model. First, we pre-trained the internal dynamics independently by optimizing and then the entire network was trained by optimizing . In this final optimization, the dynamics model was freezed to focus on cost learning. Goal states used to define were . We prepared a model variant, termed freezed PI-Net, whose internal dynamics was the above-mentioned pre-trained one and cost model was teacher model as is. The freezed PI-Net was not trained end-to-end.
Results The results of training and MPC simulations with trained controllers are summarized in Table 1. In the simulations, we observed success rates of task completion (defined as keeping the pendulum standing up more than 5s) and trajectory cost calculated by the teacher cost. For each network, ten 60-second simulations were conducted starting from different initial states. In the table, freezed PI-Net showed less generalization performance although it was equipped with the teacher cost. This degradation might result from the modeling errors of the learned dynamics. On the other hand, trained PI-Net achieved the best performance both on generalization and control, suggesting that adequate cost model was trained to imitate demonstrations while compensating the dynamics errors. Fig. 3 illustrates visualized cost models where the cost map of the trained PI-Net resembles the teacher model well. The reason of VIN failures must result from modeling difficulty on continuous control tasks. Fine discretization of state and action space would be necessary for good approximation; however, this results in the explosion of parameters to be trained, making optimizations difficult. CNN-representation of transition kernels would not work because this is very rough approximation for the most of control systems. Therefore, one can conclude that the use of PI-Net would be more reasonable on continuous control because of the VIN modeling difficulty.
Conclusion
In this paper, we introduced path integral networks, a fully differentiable end-to-end trainable representation of the path integral optimal control algorithm, which allows for optimal continuous control. Because PI-Net is fully differentiable, it can rely on powerful Deep Learning machinery to efficiently estimate in high dimensional parameters spaces from large data-sets. To the best of our knowledge, PI-Net is the first end-to-end trainable differentiable optimal controller directly applicable to continuous control domains.
PI-Net architecture is highly flexible, allowing to specify system dynamics and cost models in an explicit way, by analytic models, or in an approximate way by deep neural networks. Parameters are then jointly estimated end-to-end, in a variety of settings, including imitation learning and reinforcement learning. This may be very useful for non-linear continuous control scenarios, such as the “pixel to torques” scenario, and in situations where it’s difficult to fully specify system dynamics or cost models. We postulate this architecture may allow to train approximate system dynamics and cost models in such complex scenarios while still carrying over the advantages of optimal control from the underlying path integral optimal control. We show promising initial results in an imitation learning setting, comparing against optimal control algorithms with linear and non-linear system dynamics. Future works include a detailed comparison to other baselines, including other learn-to-control methods as well as experiments in high-dimensional settings. To tackle high-dimensional problems and to accelerate convergence we plan to combine PI-Net with other powerful methods, such as guided cost learning and trust region policy optimization .
References
Appendix A Supplements of Experiments
We used RMSProp for optimization. Initial learning rate was set to be and the rate was decayed by a factor of 2 when no improvement of loss function was observed for five epochs. We set the hyper-parameters appeared in the path integral algorithm as , .
A.2 Linear System
Dynamics parameters , were randomly determined by following equations;
where indicates the matrix exponential and the time step size is 0.01. Cost parameters , were deterministically given by;
Elements of a state vector were sampled from and then the vector was input to LQR to generate a control sequence whose length was .
PI-Net settings
The internal dynamics parameters were initialized in the same manner as Eqs. (6, 7). According to the internal cost parameters, all values were sampled from . The number of trajectories was and the number of PI-Net Kernel recurrence was . We used the standard deviation to generate Gaussian noise .
Results
Fig. 4 exemplifies control sequences for an unseen state input estimated by trained PI-Net and LQR.
A.3 Pendulum Swing-Up
We supposed that the system is dominated by the time-evolution: , where was set to be 0.5. We discretized this continuous dynamics by using the forth-order Runge-Kutta method, where the time step size was .
MPC simulations were conducted under the following conditions. Initial pendulum-angles in trajectories were uniformly sampled from at random. Initial angular velocities were uniformly sampled from $N=30$, and only the first value in the sequence was utilized for actuation at each time step.
PI-Net settings
A neural dynamics model with one hidden layer (12 hidden nodes) and one output node was used to approximate . and were estimated by the Euler method utilizing their time derivatives. The cost model is represented as , where is a neural network with one hidden layer (12 hidden nodes) and 12 output nodes. In the both neural networks, hyperbolic tangent is used as activation functions for hidden nodes. Other PI-Net parameters were: , , and .