Optimal Control Via Neural Networks: A Convex Approach
Yize Chen, Yuanyuan Shi, Baosen Zhang
I Introduction
Decisions on how to best operate and control complex physical systems such as the power grid, commercial and industrial buildings, transportation networks and robotic systems are of critical societal importance. These systems are often challenging to control because they tend to have complicated and poorly understood dynamics, sometimes with legacy components are built over a long period of time . Therefore detailed models for these systems may not be available or may be intractable to construct. For instance, since buildings account for 40% of the global energy consumption , many approaches have been proposed to operate buildings more efficiently by controlling their heating, ventilation, and air conditioning (HVAC) systems . Most of these methods, however, suffer from two drawbacks. On one hand, a detailed physics model of a building can be used to accurately describe its behavior, but this model can take years to develop. On the other hand, simple control algorithms have been developed by using linear (RC circuit) models to represent buildings, but the performance of these models may be poor since the building dynamics can be far from linear .
In this paper, we leverage the availability of data to strike a balance between requiring painstaking manual construction of physics based models and the risk of not capturing rich and complex system dynamics through models that are too simplistic. In recent years—with the growing deployment of sensors in physical and robotics systems—large amount of operational data have been collected, such as in smart buildings , legged robotics and manipulators . Using these data, the system dynamics can be learned directly and then automatically updated at periodic intervals. One popular method is to parameterize these complex system dynamics using deep neural networks to capturing complex relationships , yet few research investigated how to integrate deep learning models into real-time closed-loop control of physical systems.
A key reason that deep neural networks have not been directly applied in control is that even though they provide good performances in learning system behaviors, optimization on top of these networks is challenging . Neural networks, because of their structures, are generally not convex from input to output. Therefore, many control applications (e.g., where real-time decisions need to be made) choose to favor the computational tractability offered by linear models despite their poor fitting performances.
In this paper we tackle the modeling accuracy and control tractability tradeoff by building on the input convex neural networks (ICNN) in to both represent system dynamics and to find optimal control policies. By making the neural network convex from input to output, we are able to obtain both good predictive accuracies and tractable computational optimization problems. The overall methodology is shown in Fig. 1. Our proposed method (shown in Fig. 1 (b)) firstly utilizes an input convex network model to learn the system dynamics and then computes the best control decisions via solving a convex model predictive control (MPC) problem, which is tractable and has optimality guarantees. This is different from existing methods that uses model-free end-to-end controller which directly maps input to output (shown in Fig. 1 (a)). Another major contribution of our work is that we explicitly prove that ICNN can represent all convex functions and systems dynamics, and is exponentially more efficient than widely used convex piecewise linear approximations .
The work in was an impetus for this paper. The key differences are that the goal in is to show that ICNN can achieve similar classification performances as conventional neural networks and how the former can be used in inference and prediction problems. Our goal is to use these networks for optimization and closed-loop control, and in a sense that we are more interested in the overall system performances and not directly the performance of the networks. We also extend the class of networks to include RNNs to capture dynamical systems.
Control and decision-making have used deep learning mainly in model-free end-to-end controller settings (shown in Fig. 1 (a)), such as sequential decision making in game , robotics manipulation , and control of cyber-physical systems . However, much of the success relies heavily on a reinforcement learning setup where the optimal state-action relationship can be learned via a large number of samples. However, many physical systems do not fit into the reinforcement learning process, where both the sample collection is limited by real-time operations, and there are physical model constraints hard to represent efficiently.
To address the above sample efficiency, safety and model constraints incompatibility concerns faced by model-free reinforcement learning algorithms in physical system control, we consider a model-based control approach in this work. Model-based control algorithms often involve two stages – system identification and controller design. For the system identification stage, the goal is to learn a fixed form of system model to minimize some prediction error . Most efficient model-based control algorithms have used a relatively simple function estimator for the system dynamics identification , such as linear model and Gaussian processes . These simplified models are sample-efficient to learn, and can be nicely incorporated in the sub-sequent optimal control problems. However, such simple models may not have enough representation capacity in modeling large-scale or high-dimension systems with nonlinear dynamics. Deep neural networks (DNNs) feature powerful representation capability, while the main challenge of using DNNs for system identification is that such models are typically highly non-linear and non-convex , which causes great difficulty for following decision making. A recent work from is close in spirit as our proposed method. Similarly, the authors use a model-based approach for robotics control, where they first fit a neural network for the system dynamics and then use the fitted network in an MPC loop. However, since use conventional NN for system identification, they cannot solve the MPC problem to global optimality. Our work shows how the proposed ICNN control algorithm achieves the benefits from both sides of the world. The optimization with respect to inputs can be implemented using off-the-shelf deep learning optimizers, while we are able to obtain good identification accuracies and tractable computational optimization problems by using proposed method at the same time.
II Closed-loop control with input convex neural networks
In this paper, we consider the settings where a neural network is used in a closed-loop system. The fundamental goal is to optimize system performance which is beyond the learning performance of network on its own. In this section we describe how input convex neural networks (ICNN) can be extremely useful in these systems by considering two related problems. First, we show how ICNN perform in single-shot optimization problems. Then we extend the results to an input convex recurrent neural networks (ICRNN), which allows us to both capture systems’ complex dynamics and make time-series decisions.
The following proposition states a simple sufficient condition for a neural network to be input convex:
The feedforward neural network in Fig. 2(a) is convex from input to output given that all weights between layers and weights in the “passthrough” layers are non-negative, and all of the activation functions are convex and nondecreasing (e.g. ReLU).
An simple example that demonstrates how the proposed ICNN can be used to fit a convex function comes form fitting the function. This function is convex and both decreasing and increasing. Let the activation function be . We can write . However, in this representation, we need a negative weight, the in front of , and this would be troublesome if we compose several networks together. In our proposed ICNN structure with all positive weights and input negation duplicates, we can write , where we impose a constraint . Such doubline on the number of input variables may potentially make the network harder to train. Yet during control, having all of the weights positive maintains the convexity between inputs and outputs even if multiple steps are considered which will be discussed in Section II-B. The constraint is linear and can be easily included in any convex optimization.
This proposition follows directly from composition of convex functions . Although it allows for any increasing convex activation functions, in this paper we work with the popular ReLU activation function. Two notable additions in ICNN compared with conventional feedforward neural networks are: 1) Addition of the direct “passthrough” layers connecting inputs to hidden layers and conventional feedforward layers connecting hidden layers for better representation power. 2) the expanded inputs that include both and . The proposed ICNN structure is shown in Fig. 2(a). Note that such construction guarantees that the network is convex and non-decreasing with respect to the expanded inputs , while the output can achieve either decreasing or non-decreasing functions over .
Fundamentally, ICNN allows us to use neural networks in decision making processes by guaranteeing the solution is unique and globally optimal. Since many complex input and output relationships can be learned through deep neural networks, it is natural to consider using the learned network in an optimization problem in the form of
where is a convex feasible space. Then if is an ICNN, optimizing over is a convex problem, which can be solved efficiently to global optimality. Note that we will always duplicate the variables by introducing , but again this does not change the convexity of the problem. Of course, since the weights of the network are restricted to be nonnegative, the performance of the network (e.g., classification) may be worse. A common thread we observe in this paper is that trading off classification performance with tractability can be preferable.
II-B Closed-loop control and recurrent neural networks
In addition to the single-shot optimization problem in (1), we are interested in optimally controlling a dynamical system. To model the temporal dependency of the system dynamics, we propose to use recurrent neural networks (instead of feed-forward neural networks). Recurrent networks carry an internal state of the system, which introduces coupling with previous inputs to the system. Fig. 2(b) shows the proposed input convex recurrent neural networks (ICRNN) structure. This network maps from input to output with memory unit according to the following Eq. (2),
where , and are added direct “passthrough” layers for augmenting representation power. If we unroll the dynamics with respect to time, we have where are network parameters, and denote the nonlinear activation functions. The next proposition states a sufficient condition for the network to be input convex.
The network shown in Fig. 2(b) is a convex function from inputs to output if all weights are non-negative, and all activation functions are convex and nondecreasing (e.g. ReLU).
The proof of this proposition again follows directly from the composition rule of convex functions. Similarly to the ICNN case, by expanding the inputs vector to include both and and restricting all weights to be non-negative, the resulted ICRNN structure is a convex and non-decreasing mapping from inputs to output.
The proposed ICRNN structure can be leveraged to represent system dynamics for close-loop control. Consider a physical system with discrete-time dynamics, at time step , let’s define as the system states, as the control actions, and as the system output. For example, for the real-time control of a building system, includes the room temperature, humidity, etc; denotes the building appliance scheduling, room temperature set-points, etc; and output is the building energy consumption. In addition, there maybe exogenous variables that impact the output of the system, for example, outside temperature will impact the energy consumption of the building. However, since the exogenous variables are not impacted by any of the control actions we take, we suppress them in the formulation below. The time evolution of a system is described by
where (4b) describes the coupling between the current inputs to the future system states. Physical systems described by (4) may have significant inertia in the sense that the outcome of any control actions is delayed in time and there are significant couplings across time periods.
Since we use ICRNNs to represent both the system dynamics and the output , the control variable expands as . The optimal receding horizon control problem at time can be written as,
where a new variable is introduced for notational simplicity, which called system inputs. It is the collection of system states and duplicated control actions and , therefore ensuring the mapping from to any future states and outputs remains convex. is the control system cost incurs at time , that is a function of both the system inputs and output . The functions and in Eq. (5b)-(5c) are parameterized as ICRNNs, which represent the system dynamics from sequence of inputs to the system output , and the dynamics from control actions to system states, respectively. is the memory window length of the recurrent neural network. The equations (5d) and (5e) duplicate the input variables and enforce the consistency condition between and its negation . Lastly, (5f) and (5g) are the constraints on feasible system states and control actions respectively. Note that as a general formulation, we do not include the duplication tricks on state variables, so the dynamics fitted by (5b) and (5c) are non-decreasing over state space, which are not equivalent to those dynamics represented by linear systems. However, since we are not restricting the control space, and we have explicitly included multiple previous states in the system transition dynamics, so the non-decreasing constraint over state space should not restrict the representation capacity by much. In Section.III we theoretically prove the representability of proposed networks.
Optimization problem in (5) is a convex optimization with respect to (w.r.t.) inputs , provided the cost function is convex w.r.t. , and convex, nondecreasing w.r.t. and . A problem is convex if and only if both the objective function and constraints are convex. In the above problem, is convex and nondecreasing w.r.t. and ; and are parameterized as ICRNNs, i.e., (5a) and (5b), such that they are convex w.r.t. . Therefore following the composition rule of convex functions, the objective function is convex w.r.t. inputs . Besides, all the equality constraints (5d) and (5e) are affine. Suppose both the state feasibile set (5f) and action feasibile set (5g) are convex, the overall optimization is convex.
The convexity of the problem in (5) guarantees that it can be solved efficiently and optimally using gradient descend method. Since both the objective function (5a) and the constraints (5b)-(5c) are parameterized as neural networks, and their gradients can be calculated via back-propagation with the modification where cost is propagated to the input rather than the weights of the network. For implementation, the gradients can be convinently calculated via existing modules such as Tensorflow viaback-propagation. Let be the optimal solution of the optimization problem at time . Then the first element of is implemented to the real-time system control, that is . The optimization problem is repeated at time , based on the updated state prediction using , yielding a model predictive control strategy.
III Efficiency and representation power of ICNN
Besides the computational traceability of the input convex networks, as an system identification model, we are also interested its predictive accuracies and capacity. This section provides theoretical analysis on the representation ability and efficiency of input convex neural networks.
[Representation power of ICNN] For any Lipschitz convex function over a compact domain, there exists a neural network with nonnegative weights and ReLU activation functions that approximates it within .
Supposing Lemma 1 is true, the proof of Theorem 1 boils down to showing that neural network with nonnegative weights and ReLU activation functions can exactly represent a maximum of affine functions. The proof is constructive. We first construct a neural network with ReLU activation functions and both positive and negative weights, then we show that the weights between different layers of the network can be restricted to be nonnegative by a simple duplication trick. Specifically, since the weights in the input layer and passthrough layers in the ICNN can be negative, we simply add a negation of each input variable (e.g. both and are given as inputs) to the network. These variables need satisfy a consistency constraint since one is the negation of the other. Since this constraint is linear, it preserves the convexity of optimization problems. The details of the proofs are given in the Appendix B.
This proof is similar in spirit to theorems in . The key new result is a simpler construction than the one used in and the restriction to nonnegative weights between the layers. ∎
Similar to Theorem 1, an analogous result about the representation power of ICRNN can be shown for systems with convex dynamics. Given a dynamical system described by rolled out system dynamics is convex, then there exists a recurrent neural network with nonnegative weights and ReLU activation functions that approximates it within . A broad range of systems can be captured by this model. For example, the linear quadratic (Gaussian) regulator problem can be described using a ICRNN if we identify as the cost of the regulator .It’s important to note that is usually used as the system output of a linear system, but in our context, we are using it to refer to the quadratic cost with respect to the system states and the control input. An example of a nonlinear system is the control of electrochemical batteries. It can be shown from first principles that the degradation of these types of batteries is convex in their charge and discharge actions and our framework offers a powerful data-driven way to control batteries found in electric vehicles, cell phones, and power systems.
III-B ICNN vs. convex piecewise linear fitting
In the proof of Theorem 1, we first approximate a convex function by a maximum of affine functions then construct a neural network according to this maximum. Then a natural question is why learn a neural network and not directly the affine functions in the maximum? This approach was taken in , where a convex piecewise-linear function (max of affine functions) are directly learned from data through a regression problem.
A key reason that we propose to use ICNN (or ICRNN) to fit a function rather than directly finding a maximum of affine functions is that the former is a much more efficient parameterization than the latter. As stated in Theorem 2, a maximum of affine functions can be represented by an ICNN with layers, where each layer only requires a single ReLU activation function. However, given a single layer ICNN with ReLU activation functions, it may take a maximum of affine functions to represent it exactly. Therefore in practice, it would be much easier to train a good ICNN than finding a good set of affine functions.
The proof of this theorem is given in Appendix C.
IV Experiments
In this section, we verify the effectiveness of ICNN and ICRNN by presenting experimental results on two decision-making problems: continuous control benchmarks on MuJoco locomotion tasks and energy management of reference large-scale commercial building , respectively. The proposed method can be used as a flexible building block in decision making problems, where we use ICNN to represent system dynamics for MuJoco simulators, and we use ICRNN in an end-to-end fashion to find the optimal control inputs. Both examples demonstrate that proposed method: 1) discovers the connection between controllable variables and the system dynamics or cost objectives; 2) is lightweight and sample-efficient; 3) achieves generalizable and more stable control performances compared with previous model-based reinforcement learning and simplified linear control approaches.
Experimental Setup We consider four simulated robotic locomotion tasks: swimmer, half-cheetah, hopper, ant implemented in MuJoCo under the OpenAI rllab framework . We train and represent the locomotion state transition dynamics Note that for notation convenience, in this example and the following building example, we use to represent the expanded control vector including its negation. For system state , if , convexity means that each dimension of is convex w.r.t. the function inputs. using a 2-layer ICNN with ReLU activations, which could be integrated into the following finite-horizon control problem to find the optimal action sequence t+T for fixed looking ahead horizon :
where the objective (6a) is convex because is a concave reward function related to system states such as velocity and control actions (the detailed forms of for different locomotion tasks are listed in Appendix D). To achieve better model generalization on locomotion dynamics, we also followed , and applied DAGGER to iteratively collect labeled robotic rollouts and train the supervised dyamics model (6b) using on-policy locomotion samples. See Appendix D for furthur simulation hyperparameters and experimental details. For each aggregated iterations of collecting rollouts data and training ICNN model, we validate the controller performance on standalone validation rollouts by optimally solving (6).
Baselines We compare our system modeling and continuous control method with state-of-the-art model-based RL algorithm , where the authors used a normal multi-layer perceptrons (MLP) model to parameterize the system dynamics (6b). We refer to their method as random-shooting algorithm, since they can not solve (6) to optimality, and they used pre-defined number of random-shooting control sequences (denoted as ) to query the trained MLP and find a best sequence as the rollout policy. Such a method is able to find good control policies in the degree of timesteps, which are much more sample-efficient than model-free RL methods . To make fair comparisons with baseline method, we keep the same setup on the rollouts number and initial random action training. Our framework makes the neural networks convex w.r.t input by adding passthrough links to the 2-layer model and keeping all the layer weights nonnegative. We evaluate the performance of both algorithms on three randomly selected fixed random seeds for four tasks. Similar to the fine tuning steps in , control policies found by ICNN can also be plugged in as initialized policies for subsequent model-free reinforcement learning algorithms.
Continuous Control Performance During training, we found both ICNN and MLP are able to predict robotic states quite accurately based on (6b). This provides a good system dynamics model which is beneficial to solve control policies. The control performances are shown in Fig. 3, where we compare the average reward of proposed method and random-shooting method with over validation rollouts during each aggregated iteration (see Fig. 8 in Appendix D.4 for random shooting performance with varying ). The policy found by ICNN outperforms the random-shooting method in all settings with varying horizon for all of the four locomotion tasks.
Intuitively, ICNN should perform better when the action space is larger, since random-shooting method can not search through the action space efficiently with a fixed . This is illustrated in the example of ant, where with more training samples aggregated and MLP model representing more accurate dynamics, random-shooting gets stuck to find better control policies and there is little improvement reflected in the control performance. Moreover, since we are skipping the expensive process on calculating rewards of each random shooting trajectory and finding the best one, our method only implements ICNN inference step based on (6) and is much faster than random shooting methods in most settings, especially when is large (see Table. II for wall-clock time in Appendix D.3). For instance, in the case of Swimmer, our proposed method only uses of time compared to . This also indicates that our method is even much more sample-efficient than off-the-shelf model-free RL methods, where we use two orders of magnitude less training data to reach similar validation rewards (see Fig. 9 in Appendix D.4).
IV-B Building Energy Management
Experimental Setup We now move on to optimally control a dynamical system with significant inertia. We consider the real-time control problem of building’s HVAC (heating, ventilation, and air conditioning) system to reduce its energy consumption. Building energy management remains to be a hard problem in control area. The exact system dynamics are unknown and hard to model due to the complex heating transfer dynamics, time-varying environments and the scale of the system in terms of states and actions . At time , we assume the building’s running profile is available, where denotes building system states, including outside temperature, room temperature measurements, zone occupancies and etc. denotes a collection of control actions such as room temperature set points and appliance schedule. Output is the electricity consumption .
This is a model predictive control problem in the sense that we want to find the best control inputs that minimize the overall energy consumption of building by looking ahead several time steps. To achieve this goal, we firstly learn an ICRNN model of the building dynamics, which is trained to minimize the error between and , while denotes the memory window of recurrent neural networks. Then we solve:
where the objective (7a) is minimizing the total energy consumption in future steps ( is the model predictive control horizon), and (7b) is used for modeling building states, in which are parameterized as ICRNNs. Note that the formulation (7) is also flexible with different loss functions. For instance, in practice, we could reuse trained dynamics model (7b), and integrate electricity prices into the overall objective so that we could directly learn real-time actions to minimize electricity bills (please refer to Appendix E for more results). The constraints on control actions and system states t are given in (7c) and (7d). For instance, the temperature set points as well as real measurements should not exceed user-defined comfort regions.
To test the performance of the proposed method, we set up a 12-story large office building, which is a reference EnergyPlus commercial building model from US Department of Energy (DoE) Energyplus is an open-source whole-building energy modeling software, which is developed by US DoE for standard building energy simulation, with a total floor area of square feet which is divided into separate zones. By using the whole year’s weather profile, we simulate the building running through the year and record (, ) with a resolution of minutes. We use months’ data to train the ICRNN and subsequent months’ data for testing. We use building system state variables (uncontrollable), along with control variables . Output is a single value of building energy consumption at each time step. We set the model predictive control horizon (six hours). We employ an ICRNN with recurrent layer of dimension to fit the building input-output dynamics . The model is trained to minimize the MSE between its predictions and the actual building energy consumption using stochastic gradient descent. We use the same network structure and training scheme to fit state transition dynamics .
Baseline We set the model-based forecasting and optimization benchmark using an linear resistor-circuit (RC) circuit model to represent the heat transfer in building systems, and solve for the optimal control actions via MPC . At each step, MPC algorithm takes into account the forecasted states of the building based on the fitted RC model and implements the current step control actions. We also compare the performance of ICRNN against the conventionally trained RNN in terms of building dynamics fitting performance and control performance. To solve the MPC problem with conventional RNN models, we also use gradient-based method with respect to controls. However, since conventional RNN models are generally not convex from input to output, there is no guarantee to reach a global optimum (or even a local one).
Results In terms of the fitting performance, ICRNN provides a competitive result compared to conventional RNN model. The overall test root mean square error (RMSE) is for ICRNN and for conventional RNN, both of which are much smaller than the error made by RC model (0.240). Fig. 4(a) shows the fitting performance on 5 working days in test data. This illustrates the good performance of ICRNN in modeling building HVAC system dynamics. Then by using the learned ICRNN model of building dynamics, we obtain the suggested room control actions by solving the optimal building control problem (7). As shown in Fig. 4(b), with the same constraints on building temperature interval of , the building energy consumption is reduced by after implementing the new temperature set points calculated by ICRNN. On the contrary, since there is no guarantee for finding optimal control actions by optimizing over conventional RNN’s input, the control solutions given by conventional RNN could only reduce of electricity. Solutions given by RC model only saves of electricity. More importantly, in Fig. 4(c) we demonstrate the control actions outputted by our method against MPC with conventional RNN in two randomly selected building zones, the building basement and top floor central area. It shows that our proposed approach is able to find a group of stable control actions for the building system control. While in the conventional RNN case, it generates control set points which have undesirable, drastic variations.
V Summary and Discussion
In this work we proposed a novel optimal control framework that uses deep neural networks engineered to be convex from the input to the output. This framework bridges machine learning and control by representing system dynamics using input convex (recurrent) neural networks. We show that many interesting data-driven control problems can be cast as convex optimization problems using the proposed network architecture. Experiments on both benchmark MuJoCo locomotion tasks and building energy management demonstrate our methodology’s potential in a variety of control and optimization problems.
References
Appendix
Figure 5 shows the decision boundaries for and , respectively. These networks are composed of 2 hidden layers, with neurons in each layer, and are trained using the same random seed, same number of samples (100) until loss convergence. The decision boundaries of a conventional network have many “zigzags”, which makes solving (1) challenging, especially if is constrained. In contrast, the ICNN has convex level sets (by construction) as decision boundaries, which leads to a convex optimization problem.
Appendix B. Proof of Theorem 1
Lemma 1 follows from well established facts in function analysis stating that piecewise linear functions are dense in the space of all continuous functions over compact sets and convex piecewise linear functions are dense in the space of all convex continuous functions . Using the fact that convex piecewise linear functions can be represented as a maximum of affine functions gives the desired result in the lemma.
As a starting example, consider a maximum of two affine functions
To obtain the exact same function using a neural network, we first rewrite it as
Now define a two-layer neural network with layers and as shown in Fig. 6:
where is the ReLU activation function and the second layer is linear. By construction, this neural network is the same function as given in (8).
The above argument extends directly to a maximum of linear functions. Suppose
Again the trick is to rewrite as a nested maximum of affine functions. For notational convenience, let , . Then
The last equation describes a layer neural network, where the layers are:
Each layer of of this neural network uses only a single activation function.
Then any inner product of the form can be written as
where all coefficients are nonnegative in the above sum.
Therefore any inner product between a coefficient vector and the input can be written as an inner product between a nonnegative coefficient vector and the expanded input . Therefore, without loss of generality, we can limit all of the weights between layers to be nonnegative, and thus the neural network to be input convex. Note that in optimization problems, we need to enforce consistency in be including (12) as a constraint. However, this is a linear equality constraint, which maintains the convexity of the optimization problem.
Appendix C. Proof of Theorem 2
The second statement of Theorem 2 directly follows the construction in the proof of Theorem 1, which shows that a maximum of affine functions can be represent by a -layer ICNN (with a single ReLU function in each layer). So it remains to show the first statement of Theorem 2.
To show that a maximum of affine functions can require exponential number of pieces to approximate a function specified by an ICNN with activation functions, consider a network with 1 hidden layer of K nodes and the weights of direct “passthrough” layers are set to 0:
In order to represent the same function by a maximum of affine functions, we need to assess the value of every activation unit . If , ; otherwise, . In total, we have potential combinations of piecewise-linear function, including
So the following maximum over pieces is required to represent the single linear ICNN:
Appendix D. Experimental Details on MuJoCo Tasks
Rollout Samples To train the neural network dynamics model (both ICNN and MLP), we first collect initial rollout data using fully random action sequences with a random chosen initial state. During the data collection process in aggregated iterations, to improve model generalization and explore larger state spaces, we add Gaussian noise to the optimal control policies .
Neural Networks Training We represent the MuJoCo dynamics with a 2-hidden-layer neural networks with hidden sizes . The passthrough links of ICNN are of same size of corresponding added layers. We train both models using Adam optimizer with a learning rate 0.001 and a mini-batch size of 512. Due to the different complexity of MuJoCo tasks, we vary training epochs and summarize the training details in Table. I.
D.2 Environment Details
In all of the MuJoCo locomotion tasks, includes state variables such as robot positions, velocity along each axis; includes action efforts for the agent. We use standard reward functions for moving tasks, which could be also promptly calculated in (6a) as the control objective. For the ease of neural network training and action sampling, we normalize all the action and states in the range of . We use DAGGER for aggregated iterations for all cases, and during aggregated iteration, we use a split of random rollouts collected as described in D.1 Data Collection, and other coming from past iterations’ control policies (on-policy rollouts). Note that we use random control sequences in our method to initialize the policy finding approach and avoid the long computation time for taking gradients on finding optimal . Other environment parameters are described in Table. I.
D.3 Wall-Clock Time
In Table.II, we show the average run time for the total of aggregation iterations over runs. Finding control policies via ICNN is using less or equal training time compared to random-shooting method with , while achieving better task rewards than for different control horizons. All the experiments are running on a computer with 8 cores Intel I7 6700 CPU. Note that we do not use GPU for accelerating ICNN optimization step (6), which could furthur improve our method’s efficiency.
D.4 Details of Simulation Results
MuJoCo Dynamics Modeling In Fig. 7, we compare the ICNN and normal MLP fitting performance of the MuJoCo dynamics modeling (6b), which illustrates that both MLP and ICNN are able to find a data-driven dynamics model for ant MuJoCo agent, which is of the most complex dynamics we considered for locomotion tasks. The multi-step prediction errors of ICNN is comparable to normal MLP used in for different length of rollout steps.
More Simulation Results In Fig. 8, we compare our control method with random-shooting approach with varying settings on shooting number , which shows that our approach is more efficient in finding control policies.
In Fig. 9, we compare our control method with the rllab implementation of trust region policy optimization (TRPO) , an end-to-end deep reinforcement learning approach for mujoco locomotion tasks. More specifically, we compare the algorithms’ performances with relatively few available rollout samples. While our approach quickly learns the dynamics and then find control actions via optimization steps, TRPO is hard to learn the actions directly with few provided rollouts. Similarly to the model-based and model-free (Mb-Mf) approach described in , our control method could provide good initialization samples for the model-free algorithms, which could greatly accelerate the training process of model-free algorithms.
Appendix E. Details on Building Energy Management
To further demonstrate the potential of our proposed control framework in dealing with different real world tasks, we modify the setting of the building control example in Section 4.2 to a more complicated case. Instead of directly minimize the total energy consumption of building, we aim to minimize the total energy cost of building which subject to a varying time-of-use electrical price .
The optimization problem in (7) should be re-written as,
where the objective (14a) is minimizing the total energy cost of building in future steps ( is the model predictive control horizon) subject to time-of-use electricity price , and (14b) is used for modeling building states, in which are parameterized as ICRNNs. Same as the previous building control case, we have constraints on both control actions and system states t are given in (14c) and (14d). For instance, the temperature set points as well as real measurements should not exceed user-defined comfort regions. In Fig. 10 we visualize our model flexibility by using Seattle’s Time-of-Use (TOU) price from Seattle City Light http://www.seattle.gov/light/, and minimizing one week’s electricity bills. We could see ICRNN capture the long term relationships between control variables and final costs, and raise the energy consumption during off-peak price a little, but reduce the energy consumption during peak hours.
E.2 Control Constraints Effects
In Fig. 11 we add one more comparison on the control constraints effects on the final control performance by using ICRNN. Interestingly, with different set point constraints, the ICRNN finds similar solutions for off-peak electricity usage, which may correspond to necessary energy consumptions, such as lightning and ventilation. Moreover, when we set no constraints on the system, it would cut down more than of total energy during peak hours.