A Model-based Approach for Sample-efficient Multi-task Reinforcement Learning
Nicholas C. Landolfi, Garrett Thomas, Tengyu Ma
Introduction
Reinforcement learning has achieved significant success in domains such as game-playing (Mnih et al. 2015), recommender systems (Zhang et al. 2019), and robotic control (Levine et al. 2016). When developing systems that must solve several tasks rather than a single task, it is natural to try to leverage shared structure across tasks to make the learning process more efficient (Caruana 1997). Moreover, exploiting shared structure enables faster adaptation to unseen tasks in the future.
The Model Agnostic Meta Learning (MAML) algorithm has proven to be a successful algorithm for multi-task learning in both supervised and reinforcement learning settings (Finn et al. 2017). In the reinforcement learning (RL) setting, MAML learns a shared policy initialization across tasks. This policy initialization is then adapted to new tasks at test time by policy gradient updates.
Transferring a policy between tasks faces two challenges. First, transferring a policy proves difficult on task distributions in which the optimal policies vary dramatically across tasks. We must assume the existence of a policy initialization from which many task-specific policies may be found via a few local updates. This assumption may fail if policies vary widely across tasks. While Finn and Levine 2017 show that a single step or multiple steps of gradient-based adaptation is powerful and universal in theory, this analysis may require very complex neural networks, which in turn need many samples to generalize.
Second, transferring policies also affects the time and sample efficiency of the adaptation phase. Even if an algorithm exhaustively trains against all tasks in the family, the adaptation phase still requires samples.
We aim to address these limitations by revisiting the model-based approach to multi-task reinforcement learning. First, we learn and adapt with a shared dynamical model, rather than a policy. Second, we propose a “warm-up" phase of adaptation in which we train a policy on our learned dynamical model. We assume that the tasks encountered have the same (or similar) associated dynamical models, but we allow the tasks, and the policies required to solve them, to vary arbitrarily. Adaptation with our algorithm requires a modest amount of time (because we train a separate policy for each task), but the adaptation requires few, if any, samples. Consider the extreme case in which we acquire a perfect dynamical model of the environment during training: adapting a policy to a new task may be computationally challenging, but will require no new samples.
Conventional wisdom may suggest the learned dynamical model will fail to transfer; i.e., a dynamical model learned for one task will not work for other tasks. Perhaps surprisingly, we find that the dynamical model can transfer well; especially well if we also collect a small amount of new samples on a task during adaptation. We find our model-based approach provides significant gains in sample complexity over prior state-of-the-art: policies produced by our algorithm obtain better performance on test tasks, despite using less than 1% as much environment interaction at train time (e.g., 0.4 vs. 80 million samples).
A model-based multi-task RL algorithm, Sequential Multi-task Learning (Algorithm 1).
Numerical experiments comparing our approach to MAML on several continuous control MuJoCo reinforcement learning benchmarks of varying difficulty (Figure 1).
Numerical experiments suggesting our algorithm can handle (a) out-of-distribution tasks, (b) a shift in dynamical model (both Figure 2) and (c) active task selection (Figure 3).
Preliminaries
In multi-task learning we train on a variety of tasks to acquire some shared parameters that help us learn new, “similar," tasks with few samples. In this section we clarify the multi-task RL formalism addressed by our work, review the Model Agnostic Meta Learning (MAML) algorithm in this setting, and recall SLBO, a recent state-of-the-art model-based reinforcement learning algorithm.
In practice, we truncate the infinite sum to a finite horizon . The expected return gives a one-number summary of a given policy’s performance on an environment with dynamics . We seek a policy which maximizes .
2 Multi-task Reinforcement Learning
We aim to find a shared structure which enables good post-adaptation performance, in expectation, over a given task distribution :
If is non-deterministic, the expectation above is also taken over the randomness in .
3 Model Agnostic Meta Learning
The celebrated Model Agnostic Meta Learning (MAML) algorithm, as applied to multi-task reinforcement learning, is a model-free policy-search initialization method (Finn et al. 2017). In MAML, the shared structure learned at train time is a set of policy parameters, i.e., . The adaptation algorithm updates the parameters of the policy for a new task so that they will perform well on the new task after a single step of gradient ascent (multiple steps may be used in practice):
4 Stochastic Lower Bound Optimization
The recently proposed Stochastic Lower Bound Optimization (SLBO) method is a model-based reinforcement learning algorithm (Luo et al. 2019). SLBO achieves state-of-the-art sample efficiency by interleaving dynamical model fitting, policy training, and sample collection.
For a learned dynamical model , define the -step prediction on state and action sequence by and = for . Define the -step prediction loss:
Consider a family of policies and dynamics models (for our purposes, fully connected neural networks). SLBO approximately solves:
Related Work
A number of previous works have considered multi-task reinforcement learning (Oh et al. 2017; Ammar et al. 2014; Wilson et al. 2007; Lazaric and Ghavamzadeh 2010), and Taylor and Stone 2009 survey the area. One line of work studies the problem of producing a single policy which can effectively perform a variety of tasks. A common strategy here, often referred to as distillation, is to transfer knowledge from one policy to another. The basic approach trains a single policy to imitate many task-specific expert policies (Parisotto et al. 2016; Rusu et al. 2015). One can additionally regularize the task-specific policies so that they remain somewhat close to the distilled policy (Teh et al. 2017).
Multi-task learning is closely related to meta-learning, where a learning algorithm attempts to “learn how to learn”, somehow leveraging prior experience to learn future tasks more quickly (Schmidhuber 1987; Thrun and Pratt 1998). Beyond reinforcement learning, meta-learning has been successfully applied to the problem of few-shot classification, where the goal is to learn (at test time) to accurately classify instances of new classes not encountered during training, given just a few examples from these classes (Santoro et al. 2016; Vinyals et al. 2016; Ravi and Larochelle 2017). The MAML algorithm (Finn et al. 2017) approaches meta-learning by learning an initialization that can be quickly adapted to new tasks via gradient-based optimization. Another approach is training a recurrent neural network that is a function of the training history, implicitly encoding a learning algorithm in its weights (Hochreiter et al. 2001). This strategy has been applied in reinforcement learning by Duan et al. 2016 and Wang et al. 2016. Clavera et al. 2019 develop both MAML-like and RNN-based meta-learning approaches to learning and adapting dynamical models online in changing environments.
Model-based approaches have long been recognized as a promising avenue for reducing the sample complexity of RL algorithms (Sutton and Barto 1998; Deisenroth and Rasmussen 2011). We use an existing model-based RL algorithm (SLBO), noting that in principle our approach can be used with any model-based RL algorithm, and that SLBO is related to other recently proposed algorithms. The special case where (i.e. the model is re-trained once each time new data is observed) has been described several times under different names: it is referred to as MB-TRPO in (Luo et al. 2019), as “Vanilla Model-Based Deep Reinforcement Learning” in (Kurutach et al. 2018), and as SimPLe in (Kaiser et al. 2019). However, it is usually challenging to obtain a perfect dynamics model (Abbeel et al. 2006), and a policy trained on a single estimated model is prone to overfit to particular inaccuracies in that model. Kurutach et al. 2018 propose to use an ensemble of models to mitigate this issue. Using multiple inner iterations in SLBO is similar to using an ensemble; due to stochasticity in the optimizer, the intermediate models obtained after varied numbers of optimization steps are likely to agree on the training data, but may differ outside of the training distribution.
Training a policy in a virtual environment is not the only way to use a learned dynamical model (Sutton and Barto 1998). For example, one can produce policies based on model predictive control (MPC), where at each time step the model is used to perform planning over a short horizon to select the next action (Chua et al. 2018). Nagabandi et al. 2018 use MPC to initalize policies and then fine-tune them using model-free RL. Alternatively, one can use the model to generate “imagined” trajectories and use these as additional inputs to a policy (Racanière et al. 2017).
Our Approach
Recall that two areas of interest addressed by our approach are (a) sample complexity in the training time and adaptation time and (b) zero-shot adaptation to similar tasks. To achieve this we leverage two insights:
We transfer the parameters of the learned dynamical model, rather than the policy.
We warm-up a task’s policy by training on learned “virtual" dynamics prior to interaction.
Our algorithm proceeds in three iterated stages: (1) task sampling, (2) policy warm-up training, and (3) one or more interleaved data collection and policy training phases. In essence, we perform model-based reinforcement learning on a sequence of tasks sampled from the task distribution . We present details on each step below and pseudo-code in Algorithm 1.
(1) Task Sampling. We sample tasks from the distribution . We could augment this form of sampling by choosing tasks actively, or in some adversarial manner. Later we suggest some heuristics for choosing tasks actively (Section 4.4), and we leave adversarial multi-task setting to later work.
(2) Policy Warm-Up. We initialize a random policy and warm it up on the new task by training on learned dynamical models; for the first task, we skip the warm-up. We call this VirtualTraining because it requires no real samples, and trains the policy well against only learned dynamical models of the environment. We intend, of course, that this warm-up procedure on the learned dynamical model will produce a reasonable policy without interacting with the new environment (validated empirically in Section 5).
(3) Data Collection. Here we alternately collect new data, fit a dynamics model and then perform VirtualTraining for iterations. This stage is properly viewed as running SLBO, a model-based RL algorithm, on the new task with a warmed-up policy and previously selected data. Only here (line 1) do we collect samples from the real environment.
2 Key Design Choices
Inner Iterations of Virtual Training. In steps (2) and (3), we train the policy against several learned dynamical models. By intentional over-parameterization of the neural network representing the dynamical model, we cause the policy to see different dynamical models (all, however consistent with the data) at each iteration of the virtual training process. Intuitively, this process prevents the policy from over-fitting to a particular dynamical model. Empirically, (Luo et al. 2019) found this improved performance in the single-task traditional RL setting.
Warm-up. We warm-up a policy on the trained model. Firstly, this allows us to adapt the policy without any new samples. If we can collect new samples from the environment, this process ensures that we obtain informative data even on the first roll-out. Ultimately when we test our algorithm, our adaptation will be exactly this warm-up phase (followed, perhaps, by a few iterations of data collection and training phases).
Sequential vs. Joint Training. We have chosen to perform training sequentially, meaning that we select a task and then train and collect data on it in isolation, before moving to the next task. In contrast, we could instead sample a batch of tasks, then iteratively collect data and train policies for each task jointly. The primary advantage of sequential training is that for later tasks, we start with a reasonable policy, and so data collected is very informative. Whereas in joint training, we waste samples early on by collecting data on all tasks when each policy is poor.
3 Adaptation / Test Phase
where vt denotes the VirtualTraining routine (lines 1-1 in Algorithm 1) which returns updated policy parameters, and slbo denotes some number of steps of the SLBO algorithm (lines 1-1 in Algorithm 1).
In Section 5.3 we experimentally analyze our algorithm’s performance when adapting to (a) different task distributions at test time and (b) different dynamical models at test time.
4 Active Task Selection
As mentioned earlier, our tasks need not be sampled uniformly. In particular, Line 1 of Algorithm 1 can be extended to select tasks actively. This enables us to select a particularly difficult or diverse sequence of tasks, which may be desirable if it speeds training.
Rating Tasks. One sensible rating function is the difference between a virtual model’s predicted policy reward and the real reward:
where are parameters of the policy after warm-up. Computing such a rating requires samples, but we can estimate it using several pairs of virtual models. We experimentally explore active task sampling in Section 5.4.
Experiments
We evaluate our approach on a variety of continuous control tasks based on environments from the rllab benchmark (Duan et al. 2016), which uses the MuJoCo physics simulator (Todorov et al. 2012). We first describe the tasks considered, explain the setup of the experiments, present results comparing our algorithm and MAML. Next, we describe and give results for modified settings where the test tasks are not drawn from the same distribution as the training tasks. Finally, we include preliminary results for methods in which we sample tasks adaptively, rather than independently, at train time.
We implemented our sequential multi-task training algorithm by extending the code that Luo et al. 2019 provide. https://github.com/roosephu/slbo We use the MAML implementation that Finn et al. 2017 provide. https://github.com/cbfinn/maml_rl In all cases, we use a horizon of time steps. Each task involves one of three standard MuJoCo models: half-cheetah, ant, or simple humanoid.
Goal velocity tasks. In goal velocity tasks, a velocity is given, and the agent’s goal is to match its own velocity to . In our experiments, is sampled uniformly at random from an interval . We experimented with the following variants: half-cheetah with , half-cheetah with , ant with , ant with , humanoid with , humanoid with , and ant with and velocities . and forward/backward tasks first appeared in Finn et al. 2017. We remove contact cost, a negligible part of total reward, as we do not include this in our learned dynamical model.
Forward/backward tasks. In forward/backward tasks, a direction (where indicates forward and indicates backward) is given, and the agent’s goal is to maximize its speed in the given direction. In our experiments, is sampled uniformly at random from .
We compare the performance of our method to MAML in adapting to new tasks sampled from the task distribution. The output of MAML’s meta-training process is a set of initial policy parameters , while the output of our method is a set of initial dynamical model parameters and a dataset .
Training. The training process for our algorithm is outlined in Algorithm 1. The training process for MAML is outlined in Finn et al. 2017, Algorithm 1. For both algorithms, and all environments, each policy is implemented as a fully connected feedforward network with two hidden layers of 100 units each, using the ReLU activation . The sizes of the input and output layers are matched with the dimensions of the state and action spaces, respectively. We list all the hyperparameters used for both algorithms in the Appendix.
We note here that the MAML training process (which we did not modify from Finn et al. 2017) requires 80 million samples. We ran our algorithm for iterations, requiring 0.4 million samples, less than 1% required by MAML Ours: 100 tasks by 4000 samples/task. MAML: 500 meta-steps by 40 tasks/meta-step by 4000 samples/tasks..
Evaluation. To evaluate our algorithm against MAML we sample tasks from and train a separate policy for each. For our algorithm we warm-up a policy on the virtual environment using data sampled from , and then train on stages of SLBO. For MAML, we initialize to the policy and then train policy gradient steps. In both cases, the amount of data collected at each iteration is the same, and is plotted on the horizontal axis of all plots.
2 Comparison to MAML
For the eight task/environment pairs described in Section 5, we compare our algorithm against (a) MAML and (b) an oracle policy which is trained jointly across the task distribution but receives the task parameter as an additional input (as in (Finn et al. 2017)). We produce the curves in Figure 1 by (1) sampling 40 independently drawn tasks, (2) training a policy by collecting roll-outs (described in Section 5.1), and (3) estimating the policy’s average return by sampling new roll-outs. The plots show the return averaged across tasks, with 95% confidence intervals computed by a bootstrap estimator. Although different tasks have different reward ranges (i.e., the maximum achievable reward depends on the task parameter ), the expectation over tasks is comparable. The oracle policy acts roughly as an upper bound of the few-step performance of MAML, but we note that our approach can and often does exceed the performance of the oracle policy, because we train a separate, independent policy for each test task, whereas the oracle policy takes in the task parameters as inputs. MAML also trains a policy for each test task, but these are all constrained to have the same initialization, effectively limiting their ability to cover the task distribution.
Analysis: Firstly, we out-perform MAML across all tasks; even having trained with fewer samples. As we mentioned earlier, if the policy must vary substantially on different tasks within the same task family, there is no guarantee that there exists an initialization from which near-optimal policies for all tasks can be reached within one or a few gradient steps. This is a fundamental limitation of MAML which is not shared by our approach, and allows us to substantially out-perform MAML. Secondly, training policies in a virtual environment provided by the learned dynamics model enables zero-shot adaptation to new tasks if the model is sufficiently accurate. That is, the policy produced by the warm-up stage (before any samples are collected from the test environment) may already be high-performing. Indeed, we observe our algorithm out-performing the oracle without any samples.
3 Task Distribution and Dynamical Model Shift
We also examine how our approach performs when the distribution of tasks at test time differs from that at train time. This scenario is referred to as domain adaptation in the context of supervised learning (Ben-David et al. 2010). We experimented with four such tasks, results in Figure 2.
Changing Reward Distribution. Our first pair of set-ups (one for half-cheetah, and one for ant) evaluate transfer to different reward distributions, without changing the dynamics. In these positive to negative tasks, we train on positive velocities ( for ant) but we test on velocities from the negated interval (i.e., , respectively).
Changing True Dynamical Model. Our second two set-ups maintain the same reward distribution, but evaluate transfer to somewhat different underlying dynamical models. In the ant low friction scenario, the setup we train on the usual ant forward/backward task family, but test in a MuJoCo environment with all friction coefficients halved. In the ant crippled scenario, we train on the usual ant forward/backward task family, but test in a MuJoCo environment with one ant leg disabled.
Analysis: We see that in the changing reward distributions (Figures 2(a), 2(b)), our algorithm is able to still achieve some gains over MAML as a result of the warm-up phase (compare with Figures 1(e) and 1(f)). As expected, if the underlying dynamical model changes (i.e., the friction in Figure 2(c)) our warm-up is not as effective. Especially so in the case when we cripple the ant, compare the mean reward in Figure 2(d) (reward ~) with that in Figure 1(d) (reward ~).
4 Active Task Selection
We validate two instances of the active task selection algorithm described in (Section 4.4): one using the real reward difference rating (requiring extra samples) and one estimating the reward difference (requiring no samples). Figure 3 shows the advantage over the non-active Algorithm 1 on the two-dimensional ant velocity task in three regimes: 5, 20 and 30 tasks.
Analysis: We see that active sampling can not help much with only 5 tasks, helps a modest amount with 20 tasks, and by 30 tasks, our baseline is already doing so well that the active sampling no longer helps. These results suggest the promise of active sampling, but also suggest that the tasks considered in this work are too easy to see any significant gains; our non-active algorithm works well already. We leave it to future work to explore harder tasks and different rating functions.
Conclusion
Limitations. Of course, our ability to outperform MAML is enabled by the underlying assumption of similar dynamical models across tasks. In many applications, e.g. robotics, this assumption is reasonable. Still, we do not pretend to solve the general problem when the transition dynamics vary significantly among the tasks to be learned. In that case, zero-shot adaptation becomes impossible, but a partial dynamical model may accelerate adaptation. Second, virtual training, although it requires no new samples, is compute-intensive and the warm-up phase proves the longest part of our algorithm. Though the trade-off with samples is favorable: according to our rough measurements, our less sample-intensive training process still requires less wall-time than MAML.
Future directions. Three clear directions lie ahead. (1) Apply Sequential Multi-task Learning to more difficult settings. In fact, we perform well on most setups in Section 5 within 30 tasks. Also, harder settings would allow us to evaluate active task selection algorithms. (2) Sample tasks actively from the distribution. We sketched one possibility in Section 5.4, but upon reaching a sample budget of 30 we observe vanishing gains from active sampling. The questions of rating and comparing tasks, especially with no new samples, remain open. (3) Analyze the online setting and adversarial tasks. Our algorithm handles tasks sequentially so the tasks need not come from any distribution. Analysing worst case adaptation performance would be interesting.
Summary. Our model-based approach to multi-task reinforcement learning (a) delivers sample efficiency and (b) enables zero-shot adaptation. We transfer the shared structure of a learned dynamical model and adapt to a new task by training on this learned dynamical model. Our empirical results support the effectiveness of model-based multi-task reinforcement learning.
References
Appendix
For the purpose of reproducibility, we list here all the hyperparameters used in our experiments.
We first describe the settings used for our algorithm. We parameterized the dynamical model by a simple feedforward neural network with two hidden layers of width 500. We used Adam to optimize the reconstruction loss, with a learning rate of 0.001 and a training bath size of 128. We parameterized the policy by a simple feed-forward neural network with two hidden layers of width 32. We used TRPO to optimize the policy with standard parameters from Luo et al. 2019. As mentioned earlier, most experiments (except notably the active task selection ones) trained against 100 tasks (i.e., ). Additionally, , , and (i.e., we only collected samples once for each task after warm-up). For the adaptation phase, of course, . On half cheetah environment we collected 4000 samples from the real environment per iteration (with a horizon of 200, this corresponds to 20 trajectories on average. For ant and humanoid we collected 8000 samples from the real environment per iteration (with a horizon of 200, this corresponds to 40 trajectories on average). We always collected the same number of samples when performing TRPO on the virtual dynamics (i.e., or respectively). We collect the same number of samples at adaptation time.
We now describe the settings used when evaluating MAML. They were mostly taken directly from Finn et al. 2017 or their supplied implementation. The adaptation steps are computed by standard policy gradients [Williams 1992], while the meta-updates are computed by TRPO [Schulman et al. 2015]. We performed iterations of MAML training. The learning rate (referred to as in Finn et al. 2017) is set to during the meta-train phase; during the meta-test phase, for the first step and for subsequent steps. The meta-learning rate (referred to as in Finn et al. 2017) is set to . The batch size (the number of rollouts used to compute the policy gradient updates) is , except for the ant forward/backward and all humanoid tasks, where we use . The meta-batch size (the number of tasks sampled at each iteration of MAML) is in all cases.