What Matters for Adversarial Imitation Learning?

Manu Orsini, Anton Raichuk, Léonard Hussenot, Damien Vincent, Robert Dadashi, Sertan Girgin, Matthieu Geist, Olivier Bachem, Olivier Pietquin, Marcin Andrychowicz

Introduction

Reinforcement Learning (RL) has shown its ability to perform complex tasks in contexts where clear reward functions can be set-up (e.g. +1 for winning a chess game) but for many real-world applications, designing a correct reward function is either tedious or impossible , while demonstrating a correct behavior is often easy and cheap. Therefore, imitation learning (IL, ) might be the key to unlock the resolution of more complex tasks, such as autonomous driving, for which reward functions are much harder to design.

The simplest approach to IL is Behavioral Cloning (BC, ) which uses supervised learning to predict the expert’s action for any given state. However, BC is often unreliable as prediction errors compound in the course of an episode. Adversarial Imitation Learning (AIL, ) aims to remedy this using inspiration from Generative Adversarial Networks (GANs, ) and Inverse RL : the policy is trained to generate trajectories that are indistinguishable from the expert’s ones. As in GANs, this is formalized as a two-player game where a discriminator is co-trained to distinguish between the policy and expert trajectories (or states). See App. B for a brief introduction to AIL.

A myriad of improvements over the original AIL algorithm were proposed over the years , from changing the discriminator’s loss function to switching from on-policy to off-policy agents . However, their relative performance is rarely studied in a controlled setting, and never these changes were all compared together. The performance of these high-level choices may also depend on the low-level implementation details, which might be not even mentioned in the publications , as well as the hyperparameters (HPs) used. Thus, assessing whether the proposed changes are the reason for the presented improvements becomes extremely difficult. This lack of proper comparisons slows down the overall research in imitation learning and the industrial applicability of these methods.

We investigate such high- and low-level choices in depth and study their impact on the algorithm performance. Hence, as our key contributions, we (1) implement a highly-configurable generic AIL algorithm, with various axes of variation (>50 HPs), including 4 different RL algorithms and 7 regularization schemes for the discriminator, (2) conduct a large-scale (>500k trained agents) experimental study on 10 continuous-control tasks A task is defined by an environment and the demonstrator type (either human or RL agent). and (3) analyze the experimental results to provide practical insights and recommendations for designing novel and using existing AIL algorithms.

While many of our findings confirm common practices in AIL research, some of them are surprising or even contradict prior work. In particular, we find that standard regularizers from Supervised Learning — dropout and weight decay often perform similarly to the regularizers designed specifically for adversarial learning like gradient penalty . Moreover, for easier environments (which were often the only ones used in prior work), we find that it is possible to achieve excellent results without using any explicit discriminator regularization.

Most surprising finding #2: human demonstrations.

Not only does the performance of AIL heavily depend on whether the demonstrations were collected from a human operator or generated by an RL algorithm, but the relative performance of algorithmic choices also depends on the demonstration source. Our results suggest that artificial demonstrations are not a good proxy for human data and that the very common practice of evaluating IL algorithms only with synthetic demonstrations may lead to algorithms which perform poorly in the more realistic scenarios with human demonstrations.

Paper outline.

In Sec. 2, we describe our experimental setting and the performance metrics used. We then present and analyze the results related to the agent (Sec. 3) and discriminator (Sec. 4) training. Afterwards, we compare RL-generated and human-collected demonstrations (Sec. 5) and analyze the choices influencing the computational cost of running the algorithm (Sec. 6). The appendices contain the details of the different algorithmic choices in AIL (App. C) as well as the raw results of the experiments (App. F–H).

Experimental design

We focus on continuous-control tasks as robotics appears as one of the main potential applications of IL and a vast majority of the IL literature thus focuses on it. In particular, we run experiments with five widely used environments from OpenAI Gym : HalfCheetah-v2, Hopper-v2, Walker2d-v2, Ant-v2, and Humanoid-v2 and three manipulation environments from Adroit : pen-v0, door-v0, and hammer-v0. All the environments are shown in Fig. 1. The Adroit tasks consist in aligning a pen with a target orientation, opening a door and hammering a nail with a 5-fingered hand. These two benchmarks bring orthogonal contributions. The former focuses on locomotion but has 5 environments with different state/action dimensionality. The latter, more varied in term of tasks, has an almost constant state-action space.

Demonstrations.

For the Gym tasks, we generate demonstrations with a SAC agent trained on the environment reward. For the Adroit environments, we use the “expert” and “human” datasets from D4RL , For pen, we only use the “expert” dataset, the “human” one consisting of a single (yet very long) trajectory. which are, respectively, generated by an RL agent and collected from a human operator. As far as we know, our work is the first to solve these tasks with human datasets in the imitation setup (most of the prior work concentrated on Offline RL). For all environments, we use 11 demonstration trajectories. Following prior work , we subsample expert demonstrations by only using every 20th20^{\text{th}} state-action pair to make the tasks harder.

Adversarial Imitation Learning algorithms.

We researched prior work on AIL algorithms and made a list of commonly used design decisions like policy objectives or discriminator regularization techniques. We also included a number of natural options which we have not encountered in literature (e.g. dropout in the discriminator or clipping rewards bigger than a threshold). All choices are listed and explained in App. C. Then, we implemented a single highly-configurable AIL agent which exposes all these choices as configuration options in the Acme framework using JAX for automatic differentiation and Flax for neural networks computation. The configuration space is so wide that it covers the whole family of AIL algorithms, in particular, it mostly covers the setups from AIRLSeminal AIRL uses TRPO to train the policy, not supported in our implementation (PPO used here). and DAC . We plan to open source the agent implementation.

Experimental design.

We created a large HP sweep (57 HPs swept, >120k agents trained) in which each HP is sampled uniformly at random from a discrete set and independently from the other HPs. We manually ensured that the sampling ranges of all HPs are appropriate and cover the optimal values. Then, we analyzed the results of this initial experiment (called wide, detailed description and results in App. F), removed clearly suboptimal options and ran another experiment with the pruned sampling ranges (called main, 43 HPs swept, >250k agents trained, detailed description and results in App. G). The latter experiment serves as the basis for most of the conclusions drawn in this paper but we also run a few additional experiments to investigate some additional questions (App. H and App. I).

This pruning of the HP space guarantees that we draw conclusions based on training configurations which are highly competitive (training curves can be found in Fig. 27) while using a large HP sweep (including, for example, multiple different RL algorithms) ensures that our conclusions are robust and valid not only for a single RL algorithm and specific values of HPs, but are more generally applicable. Moreover, many choices may have strong interactions with other related choices, for example we find a surprisingly strong interaction between the discriminator regularization scheme and the discriminator learning rate (Sec. 4). This means that such choices need to be tuned together (as it is the case in our study) and experiments where only a single choice is varied but the interacting choices are kept fixed may lead to misleading conclusions.

Performance measure.

For each HP configuration and each of the 10 environment-dataset pairs we train a policy and evaluate it 10 times through the training by running it for 50 episodes and computing the average undiscounted return using the environment reward. We then average these scores to obtain a single performance score which approximates the area under the learning curve. This ensures we assign higher scores to HP configurations that learn quickly.

Analysis.

We consider two different analyses for each choice This analysis is based on a similar type of study focused on on-policy RL algorithms .:

Conditional 95th percentile: For each potential value of that choice (e.g., RL Algorithm = PPO), we look at the performance distribution of sampled configurations with that value. We report the 95th percentile of the performance as well as error bars based on bootstrapping. We compute each metric 2020 times based on a randomly selected half of all training runs, and then report the mean of these 20 measurements while the error bars show mean-std and mean+std. This corresponds to an estimate of the performance one can expect if all other choices were tuned with random search and a limited budget of roughly 13 HP configurationsThe probability that all 13 configurations score worse than the 95th percentile is equal 0.9513≈50%0.95^{13}\approx 50\%.. All scores are normalized so that corresponds to a random policy and 11 to the expert performance (expert scores can be found in App. E).

Distribution of choice within top 5% configurations. We further consider for each choice the distribution of values among the top 5% HP configurations. In particular, we measure the ratio of the frequency of the given value in the top 5% of HP configurations with the best performance to the frequency of this value among all HP configurations. If certain values are over-represented in the top models (ratio higher than 1), this indicates that the specific choice is important for good performance.

What matters for the agent training?

The AIRL reward function perform best for synthetic demonstrations while −ln⁡(1−D)-\ln(1-D) is better for human demonstrations. Using explicit absorbing state is crucial in environments with variable length episodes. Observation normalization strongly affects the performance. Using an off-policy RL algorithm is necessary for good sample complexity while replaying expert data and pretraining with BC improves the performance only slightly.

Implicit reward function.

In this section, we investigate choices related to agent training with AIL, the most salient of which is probably the choice of the implicit reward function. Let D(s,a)D(s,a) be the probability of classifying the given state-action pair as expert by the discriminator Some prior works, including GAIL, use the opposite notation, with D(s,a)D(s,a) the non-expert probability.. In particular, we run experiments with the following reward functions: r(s,a)=−log⁡(1−D(s,a))r(s,a)=-\log(1-D(s,a)) (used in the original GAIL paper ), r(s,a)=log⁡D(s,a)−log⁡(1−D(s,a))r(s,a)=\log D(s,a)-\log(1-D(s,a)) (called the AIRL reward ), r(s,a)=log⁡D(s,a)r(s,a)=\log D(s,a) (a natural choice we have not encountered in literature), and the FAIRL reward function r(s,a)=−h(s,a)⋅eh(s,a)r(s,a)=-h(s,a)\cdot e^{h(s,a)}, where h(s,a)h(s,a) is the discriminator logitIt can also be expressed as h(s,a)=log⁡D(s,a)−log⁡(1−D(s,a))h(s,a)=\log D(s,a)-\log(1-D(s,a)).. It can be shown that, under the assumption that all episodes have the same length, maximizing these reward functions corresponds to the minimization of different divergences between the marginal state-action distribution of the expert and the policy. See for an in-depth discussion on this topic. We also consider clipping the rewards with absolute values bigger than a threshold which is a HP.

The FAIRL reward performed much worse than all others in the initial wide experiment (Fig. 2a) and therefore was not included in our main experiment. This is mostly caused by its inferior performance with off-policy RL algorithms (Fig. 25). Moreover, reward clipping significantly helps the FAIRL reward (Fig. 26) while it does not help the other reward functions apart from some small gains for −ln⁡(1−D)-\ln(1-D) (Fig. 85). Therefore, we suspect that the poor performance of the FAIRL reward function may be caused by its exponential term which may have very high magnitudes. Moreover, the FAIRL paper mentions that the FAIRL reward is more sensitive to HPs than other reward functions which could also explain its poor performance in our experiments.

Fig. 28 shows that the ln⁡(D)\ln(D) reward functions performs a bit worse than the other two reward functions in the main experiment. Five out of the ten tasks used in our experiments have variable length episodes with longer episodes correlated with better behaviourThe episodes are terminated earlier if the simulated robot falls over or if the pen is dropped. (Hopper, Walker2d, Ant, Humanoid, pen) — on these tasks we can notice that r(s,a)=−ln⁡(1−D(s,a))r(s,a)=-\ln(1-D(s,a)) often performs best and r(s,a)=ln⁡D(s,a)r(s,a)=\ln D(s,a) worst. This can be explained by the fact that −ln⁡(1−D(s,a))>0-\ln(1-D(s,a))>0 and ln⁡D(s,a)<0\ln D(s,a)<0 which means that the former reward encourages longer episodes and the latter one shorter ones . Absorbing state (described in App. C.2) is a technique introduced in the DAC paper to mitigate the mentioned bias and encourage the policy to generate episodes of similar length to demonstrations. In Fig. 2b-c we show how the performance of different reward functions compares in the environments with variable length episodes depending on whether the absorbing state is used. We can notice that without the absorbing state r(s,a)=−ln⁡(1−D(s,a))>0r(s,a)=-\ln(1-D(s,a))>0 performs much better in the environments with variable episode length which suggests that the learning is driven to a large extent by the reward bias and not actual imitation of the expert behaviour . This effect disappears when the absorbing state is enabled (Fig. 2c).

Fig. 75 shows the performance of different reward functions in all environments conditioned on whether the absorbing state is used. If the absorbing state is used, the AIRL reward function performs best in all the environments with RL-generated demonstrations, and ln⁡(D)\ln(D) performs only marginally worse. The −ln⁡(1−D)-\ln(1-D) reward function underperforms on the Humanoid and pen tasks while performing best with human datasets. We provide some hypothesis for this behaviour in Sec. 5, where we discuss human demonstrations in more details.

Observation normalization.

We consider observation normalization which is applied to the inputs of all neural networks involved in AIL (policy, critic and discriminator). The normalization aims to transform the observations so that that each observation coordinate has mean and standard deviation 11. In particular, we consider computing the normalization statistics either using only the expert demonstrations so that the normalization is fixed throughout the training, or using data from the policy being trained (called online). See App. C.6 for more details. Fig. 3 shows that input normalization significantly influences the performance with the effects on performance being often much larger than those of algorithmic choices like the reward function or RL algorithm used. Surprisingly, normalizing observations can either significantly improve or diminish performance and whether the fixed or online normalization performs better is also environment dependent.

Replaying expert data.

When demonstrations as well as external rewards are available, it is common for RL algorithms to sample batches for off-policy updates from the demonstrations in addition to the replay buffer . We varied the ratio of the policy to expert data being replayed but found only very minor gains (Fig. 86). Moreover, in the cases when we see some benefits, it is usually best to replay 16–64 times more policy than expert data. On some tasks (Humanoid) replaying even a single expert transitions every 256 agent ones significantly hurts performance. We suspect that, in contrast to RL with demonstrations, we see little benefit from replaying expert data in the setup with learned rewards because (1) replaying expert data mostly helps when the reward signal is sparse (not the case for discriminator-based rewards), and (2) discriminator may overfit to the expert demonstrations which could result in incorrectly high rewards being assigned to expert transitions.

Pretraining with BC.

We also experiment with pretraining a policy with Behavioral Cloning (BC, ) at the beginning of training. Despite starting from a much better policy than a random one, we usually observe that the policy quality deteriorates quickly at the beginning of training (see the pen task in Fig. 6) due to being updated using randomly initialized critic and discriminator networks, and the overall gain from pretraining is very small in most environments (Fig. 30).

RL algorithms.

We run experiments with four different RL algorithms, three of which are off-policy algorithms (SAC , TD3 and D4PG ), as well as PPO which is nearly on-policy. Fig. 4 shows that the sample complexity of PPO is significantly worse than that of the off-policy algorithms while all off-policy algorithms perform overall similarly.

RL algorithms HPs.

Fig. 11 shows that the discount factor is one of the most important HPs with the values of 0.97−0.990.97-0.99 performing well on all tasks. Fig. 32 shows that in most environments it is better not to erase any data from the RL replay buffer and always sample from all the experience encountered so far. It is common in RL to use a noise-free version of the policy during evaluation and we observe that it indeed improves the performance (Fig. 33). The policy MLP size does not matter much (Figs. 34-35) while bigger critic networks perform significantly better We thus only include critics with at least two hidden layers with the size at least 128 in the main experiment. (Figs. 12-13). Regarding activation functions We use the same activation function in the policy and critic networks., relu performs on par or better than tanh in all environments apart from door in which tanh is significantly better (Fig. 36). Our implementation of TD3 optionally applies gradient clippingThe reason for that is that the DAC paper uses TD3 with gradient clipping. but it does not affect the performance much (Fig. 37). D4PG can use n-step returns, this improves the performance on the Adroit tasks but hurts on the Gym suite (Fig. 38).

What matters for the discriminator training?

MLP discriminators perform on par or better than AIL-specific architectures. Explicit discriminator regularization is only important in more complicated environments (Humanoid and harder ones). Spectral norm is overall the best regularizer but standard regularizers from supervised learning often perform on par. Optimal learning rate for the discriminator may be 2–2.5 orders of magnitude lower than the one for the RL agent.

Discriminator input.

In this section we look at the choices related to the discriminator training. Fig. 52 shows how the performance depends on the discriminator input. We can observe that while it is beneficial to feed actions as well as states to the discriminator, the state-only demonstrations perform almost as well. Interestingly, on the door task with human data, it is better to ignore the expert actions. We explore the results with human demonstrations in more depth in Sec. 5.

Discriminator architecture.

Regarding the discriminator network, our basic architecture is an MLP but we also consider two modifications introduced in AIRL : a reward shaping term and a log⁡π(a∣s)\log\pi(a|s) logit shift which introduces a dependence on the current policy (only applicable to RL algorithms with stochastic policies, which in our case are PPO and SAC). See App. C.3 for a detailed description of these techniques. Fig. 16 shows that the logit shift significantly hurts the performance. This is mainly due to the fact that it does not work well with SAC which is off-policy (Fig. 24). Fig. 53 shows that the shaping term does not affect the performance much. While the modifications from AIRL does not improve the sample complexity in our experiments, it is worth mentioning that they were introduced for another purpose, namely the recovery of transferable reward functions.

Regarding the size of the discriminator MLP(s), the best results on all tasks are obtained with a single hidden layer (Fig. 54), while the size of the hidden layer is of secondary importance (if it is not very small) with the exception of the tasks with human data where fewer hidden units perform significantly better (Fig. 55). All tested discriminator activation functions perform overall similarly while sigmoid performs best with human demonstrations (Fig. 56).

Discriminator training.

Fig. 57 shows that it is best to use as large as possible replay buffers for sampling negative examples (i.e. agent transitions). As noticed in prior work, the initialization of the last policy layer can significantly influence the performance in RL , thus we tried initializing the last discriminator layer with smaller weights but it does not make much difference (Fig 58).

Discriminator regularization.

An overfitting or too accurate discriminator can make agent’s training challenging, and therefore it is common to use additional regularization techniques when training the AIL discriminator (or GANs in general). We run experiments with a number of regularizers commonly used with AIL, namely Gradient Penalty (GP, used e.g. in ), spectral norm (e.g. in ), Mixup (e.g. in ), as well as using the PUGAIL loss instead of the standard cross entropy loss to train the discriminator. Apart from the above regularizers, we also run experiments with regularizers commonly used in Supervised Learning, namely dropout , the weight decay variant from AdamW as well as the entropy bonus of the discriminator output treated as a Bernoulli distribution. The detailed description of all these regularization techniques can be found in App. C.5.

Fig. 5 shows how the performance depends on the regularizer. Spectral normalization performs overall best, while GP, dropout and weight decay all perform on par with each other and only a bit worse than spectral normalization. We find this conclusion to be quite surprising given that we have not seen dropout or weight decay being used with AIL in literature. We also notice that the regularization is generally more important on harder tasks like Humanoid or the tasks in the Adroit suite (Fig. 59).

Most of the regularizers investigated in this section have their own HPs and therefore the comparison of different regularizers depends on how these HPs are sampled. As we randomly sample the regularizer-specific HPs in this analysis, our approach favours regularizers that are not too sensitive to their HPs. At the same time, there might be regularizers that are sensitive to their HPs but for which good settings may be easily found. Fig. 70 shows that even if we condition on choosing the optimal HPs for each regularizer, the relative ranking of regularizers does not change.

Moreover, there might be correlations between the regularizer and other HPs, therefore their relative performance may depend on the distribution of all other HPs. In fact, we have found two such surprising correlations. Fig. 76 shows the performance conditioned on the regularizer used as well as the discriminator learning rate. We notice that for PUGAIL, entropy and no regularization, the performance significantly increases for lower discriminator learning rates and the best performing discriminator learning rate (10−610^{-6}) is in fact 2-2.5 orders of magnitude lower than the best learning rate for the RL algorithm (0.00010.0001–0.00030.0003, Figs. 17, 41, 42, 44, 50). The optimal learning rate for those regularizers was the smallest one included in the main experiment. We also run an additional sweep with smaller rates but found that even lower ones do not perform better (Fig. 84). On the other hand, the remaining regularizers are not too sensitive to the discriminator learning rate. This means that the performance gap between PUGAIL, entropy and no regularization and the other regularizers is to some degree caused by the fact that the former ones are more sensitive to the learning rate and may be smaller than suggested by Fig. 5 if we adjust for the appropriate choice of the discriminator learning rate. We can notice that PUGAIL and entropy are the only regularizers which only change the discriminator loss but do not affect the internals of the discriminator neural network. Given that they are the only two regularizers benefiting from very low discriminator learning rate, we suspect that it means that a very low learning rate can play a regularizing role in the absence of an explicit regularization inside the network.

Another surprising correlation is that in some environments, the regularizer interacts strongly with observation normalization (described App. C.6) employed on discriminator inputs (see Fig. 77 for an example on Ant). These two correlations highlight the difficulty of comparing regularizers, and algorithmic choices more broadly, as their performance significantly depends on the distribution of other HPs.

We also supplement our analysis by comparing the performance of different regularizers for the best found HPs. More precisely, we choose the best value for each HP in the main experiment (listed in App. D) and run them with different regularizers. To account for the mentioned correlations with the discriminator learning rate and observation normalization, we also include these two choices in the HP sweep and choose the best performing variant (as measure by the area under the learning curve) for each regularizer and each environment.

While it is not guaranteed that the performance is going to be good at all because we greedily choose the best performing value for each HP and there might be some unaccounted HP correlations, we find that the performance is very competitive (Fig. 6). Notice that we use the same HPs in all environments Apart from the discriminator learning rate and observation normalization used. and the performance can be probably improved by varying some HPs between the environments, or at least between the two environment suites.

We notice that on the four easiest tasks (HalfCheetah, Hopper, Walker2d, Ant), investigated discriminator regularizers provide no, or only minor performance improvements and excellent results can be achieved without them. On the tasks where regularization is beneficial, we usually see that there are multiple regularizers performing similarly well, with spectral normalization being one of the best regularizers in all tasks apart from the two tasks with human data where PUGAIL performs better.

Regularizers-specific HPs.

For GP, the target gradient norm of 11 is slightly better in most environments but the value of is significantly better in hammer-human (Fig. 60), while the penalty strength of 11 performs best overall (Fig. 61). For dropout, it is important to apply it not only to hidden layers but also to inputs (Fig. 62) and the best results are obtained for 50% input dropout and 75% hidden activations dropout (Figs. 62, 63 and 70). For weight decay, the optimal decay coefficient in the AIL setup is much larger than the values typically used for Supervised Learning, the value λ=10\lambda=10 performs best in our experiments (Fig. 64). For Mixup, α=1\alpha=1 outperforms the other values on almost all tested environments (Fig. 65). For PUGAIL, the unbounded version performs much better on the Adroit suite, while the bounded version is better on the gym tasks (Fig. 66), and positive class prior of η=0.7\eta=0.7 performs well on most tasks (Fig. 67). For the discriminator entropy bonus, the values around 0.030.03 performed best overall (Fig. 68). All experiments with spectral normalization enforce the Lipschitz constant of 11 for each weight matrix.There may be different Lipschitz constants for different networks depending on those of related activations.

Are synthetic demonstrations a good proxy for human data?

Human demonstrations significantly differ from synthetic ones. Learning from human demonstrations benefits more from discriminator regularization and may work better with different discriminator inputs and reward functions than RL-generated demonstrations.

Using a dataset of human demonstrations comes with a number of additional challenges. Compared to synthetic demonstrations, the human policy can be multi-modal in that for a given state different decisions might be chosen. A typical example occurs when the human demonstrator remains idle for some time (for example to think about the next action) before taking the actual relevant action: we have two modes in that state, the relevant action has a low probability while the idle action has a very high probability. The human policy might not be exactly markovian either. Those differences are significant enough that the conclusions on synthetic datasets might not hold anymore.

In this section, we focus on the Adroit door and hammer environments for which we run experiments with human as well as synthetic demonstrations. For pen, we only use the “expert” dataset, the “human” one consists of a single (yet very long) trajectory. Note that on top of the aforementioned challenges, the setup with the Adroit environments using human demonstrations exhibits a few additional specifics. The demonstrations were collected letting the human decide when the task is completed: said in a different way, the demonstrator is offered an additional action to jump directly to a terminal state and this action is not available to the agent imitating the expert. The end result is a dataset of demonstrations of variable length while the agent can only generate episodes consisting of exactly 200 transitions. Note that there was no time limit imposed on the demonstrator and some of the demonstrations have a length greater than 200 transitions. Getting to the exact same state distribution as the human expert may be impossible, and imitation learning algorithms may have to make some trade-offs. The additional specificity of that setup is that the reward of the environment is not exactly what the human demonstrator optimized. In the door environment, the reward provided by the environment is the highest when the door is fully opened while the human might abort the task slightly before getting the highest reward. However, overall, we consider the reward provided by the environment as a reasonable metric to assess the quality of the trained policies. Moreover, in the hammer environment, some demonstrations have a low return and we suspect those are not successful demonstrations.D4RL datasets contain only the policy observations and not the simulator states and therefore it is not straightforward to visualize the demonstrations.

Discriminator regularization.

When comparing the results for RL-generated (adroit-expert We do not include pen in the adroit-expert plots so that both adroit-expert and adroit-human show the results averages across the door and hammer tasks and differ only in the demonstrations used.) and human demonstrations (adroit-human) we can notice differences on a number of HPs related to the discriminator training. Human demonstrations benefit more from using discriminator regularizers (Fig. 59) and they also work better with smaller discriminator networks (Fig. 55) trained with lower learning rates (Fig. 69). The increased need for regularization suggest that it is easier to overfit to the idiosyncrasies of human demonstrations than to those of RL policies.

Discriminator input.

Fig. 7a shows the performance given the discriminator input depending on the demonstration source. For most tasks with RL-generated demonstrations, feeding actions as well as states improves the performance (Fig. 52). Yet, the opposite holds when human demonstrations are used. We suspect that it might be caused by the mentioned issue with demonstrations lengths which forces the policy to repeat a similar movement but with a different speed than the demonstrator.

Reward functions.

Finally, we look at how the relative performance of different reward functions depends on the demonstration source. Fig. 7b shows that for RL-generated demonstrations the best reward function is AIRL while −ln⁡(1−D)-\ln(1-D) performs better with human demonstrations. Under the assumption that the discriminator is optimal, these two reward functions correspond to the minimization of different divergences between the state (or state-action depending on the discriminator input) occupancy measures of the policy and the expert — See Table 1 for the details.

While it is hard to draw any general conclusions only from the two investigated environments for which we had access to human demonstrations, our analysis shows that the differences between synthetic and human-generated demonstrations can influence the relative performance of different algorithmic choices. This suggests that RL-generated data are not a good proxy for human demonstrations and that the very common practice of evaluating IL only with synthetic demonstrations may lead to algorithms which perform poorly in the more realistic scenarios with human demonstrations.

How to train efficiently?

So far we have analysed how HPs affect the performance of AIL algorithms measured after fixed numbers of environment steps. Here we look at the HPs which influence sample complexity as well as the computational cost of running an algorithm. Raw experiment report can be found in App. H.

One of the main factors influencing the throughput of a particular imitation algorithm is the number of times each transition is replayed on average and the batch size used.We use the same batch size for the policy and actor networks while the discriminator batch size is effectively two times larger because its batches contain always batch size (CC.1) demonstration transitions and batch size (CC.1) policy transitions. The replay ratio is the same for all networks with the exception of the discriminator which can have its replay ratio doubled depending on the value of discriminator to RL updates ratio (CC.4). See App. C.4 for details. See App. C.1 for the detailed description of the HPs involved. Fig. 79 shows that smaller batches perform overall better (given a fixed replay ratio) and increasing the replay ratio improves the performance, at least up to some threshold depending on the environment (Fig. 80). There is a very strong correlation between the two HPs — Fig. 83 shows that for most batch sizes, the optimal replay ratio is equal to the batch size, which corresponds to replaying exactly one batch of data per environment step. If we compare different batch sizes under the ratio of batches to environment steps fixed to one, the performance is mostly independent of the batch size (Fig. 83).

While in most of our experiment the discriminator and the RL agent are trained with exactly the same number of batches, we also tried doubling the number of discriminator batches. Fig. 81 shows that it improves the performance slightly on the Adroit suite.

Combining multiple batches.

We also consider processing multiple batches at once for improved accelerator (GPU or TPU) utilization. In particular, we sample an NN-times larger batch from a replay buffer, split it back into NN smaller/proper batches on an accelerator, and process them sequentially. In order to keep the replay ratio unaffected, we decrease the frequency of updates accordingly, e.g. instead of performing one gradient update for every environment step, we perform NN gradients updates every NN environment steps. We apply this technique to the discriminator as well as the RL agent training. The effect on the sample complexity of the algorithm can be seen in Fig. 82. There is a small negative effect for values larger or equal to 16. The effect of this parameter on the throughput of our system could be observed in Fig. 8. The value of 8 provides a good compromise: almost no noticeable sample complexity regression while decreasing the training time by 2–3 times.

Related work

The most similar work to ours is probably which compares the performance of different discriminator regularizers and concludes that gradient penalty is necessary for achieving good performance with off-policy AIL algorithms. In contrast to , which uses a single HP configuration, we run large-scale experiments with very wide HP sweeps which allows us to reach more robust conclusions. In particular, we are able to achieve excellent sample complexity on all the environments used in evaluated gradient penalty in the off-policy setup in the following environments: HalfCheetah, Hopper, Walker2d and Ant, as well as InvertedPendulum which we did not use due to its simplicity. without using any explicit discriminator regularizer (Fig. 6).

Another empirical study of IL algorithms is , which investigates the problem of HP selection in IL under the assumption that the reward function is not available for the HP selection.

The methodology of our study is mostly based on which analyzed the importance of different choices for on-policy actor-critic methods. Our work is also similar to other large-scale studies done in other fields of Deep Learning, e.g. model-based RL , GANs , NLP , disentangled representations and convolution network architectures .

Conclusions

In this empirical study, we investigate in depth many aspects of the AIL framework including discriminator architecture, training and regularization as well as many choices related to the agent training. Our key findings can be divided into three categories: (1) Corroborating prior work, e.g. for the underlying RL problem, off-policy algorithms are more sample efficient than on-policy ones; (2) Adding nuances to previous studies, e.g. while the regularization schemes encouraging Lipschitzness improve the performance, more classical regularizers like dropout or weight decay often perform on par; (3) Raising concerns: we observe a high discrepancy between the results for RL-generated and human data. We hope this study will be helpful to anyone using or designing AIL algorithms.

Acknowledgments

We thank Kamyar Ghasemipour for the discussions related to the FAIRL reward function and Lucas Beyer for the feedback on an earlier version of the manuscript.

References

Appendix A Reinforcement Learning Background

Appendix B Adversarial Imitation Learning Background

See App. A for a very brief introduction to RL and the notation used in this section.

Drawing inspiration from Inverse Reinforcement Learning and Generative Adversarial Networks (GANs, ), adversarial imitation learning aims at learning a behavior similar to that of the expert given a set of expert demonstrations Dexpert\mathcal{D}_{expert} and the ability to interact with the environment.

To do so, the agent with policy π\pi is initialized randomly and interacts with the environment. A discriminator network DD is trained to distinguish between samples coming from the agent (st,at,st+1)∼Dπ(s_{t},a_{t},s_{t+1})\sim\mathcal{D}_{\pi} and samples coming from the expert dataset (st,at,st+1)∼Dexpert(s_{t},a_{t},s_{t+1})\sim\mathcal{D}_{expert} with a cross-entropy loss. A reward function for the policy is then defined based on the discriminator prediction, e.g. r(s,a)=−ln⁡(1−D(s,a))r(s,a)=-\ln(1-D(s,a)), where D(s,a)D(s,a) denotes the probability of classifying the state-action pair as expert by the discriminator. The agent is then trained with an RL algorithm to maximize this reward and thus fool the discriminator. As in GANs, the training of the discriminator and that of the agent (here playing the role of the generator) are interleaved. Therefore, at the high level, the algorithm repeats the following steps in a loop: (1) interact with the environment using the current policy and store the experience in a replay buffer, (2) update the discriminator, (3) perform an RL update accordingly to the RL algorithm used.

Appendix C List of Investigated Choices

In this section we list all algorithmic choices which we consider in our experiments. See App. B for an introduction to adversarial imitation and the notation used in this section. For convenience, we mark each of the choices with a number (e.g., CC.1) and a fixed name (e.g. RL Algorithm (CC.1)) that can be easily used to find a description of the choice in this section.

In all experiments we use MLPs for the policy and critic/value networks and sample the following HPs controlling the networks architectures: policy MLP depth (CC.1) (the number of hidden layers), policy MLP width (CC.1), critic MLP depth (CC.1), critic MLP width (CC.1), RL activation (CC.1), as well as discount γ\gamma (CC.1) and batch size (CC.1). All networks are optimized with the Adam optimizer.

We sample RL Algorithm (CC.1) from the following options:

For PPO, batch size (CC.1) denotes the number of experience fragments, each of consisting PPO unroll length (CC.1) transitions, collected in each policy update step. In each policy update step, we perform PPO number of epochs (CC.1) passes over the gathered data when in each pass the data is split into PPO number of minibatches (CC.1) minibatches. We use the PPO loss with the clipping threshold set by PPO clipping ϵ\epsilon (CC.1) and add an entropy loss with the coefficient specified by PPO entropy cost (CC.1). We also sample PPO learning rate (CC.1), and the GAE returns mixing coefficient GAE λ\lambda (CC.1).

Soft Actor Critic (SAC, [29])

We use a version of SAC with a policy entropy constraint . In particular, we choose SAC entropy per dimension (CC.1) and that set the entropy constraint so that the policy entropy is not lower than the number of action dimensions times this value. We also sweep SAC learning rate (CC.1) and the target network polyak averaging coefficient SAC polyak τ\tau (CC.1) (the target network is updates after each minibatch).

Twin Delayed Deep Deterministic Policy Gradient (TD3, [27])

For TD3, we sweep TD3 policy learning rate (CC.1) and TD3 critic learning rate (CC.1) separately, as well as sample behavioral policy noise (CC.1). Following the original publication, we update the actor only using every other minibatch while the critic networks uses all minibatches. The target network is updated after every minibatch with the polyak coefficient fixed to 0.0050.005. Following DAC , we clip actor gradients with magnitudes bigger than TD3 gradient clipping (CC.1).

Distributed Distributional Deterministic Policy Gradients (D4PG, [25])

This algorithm is similar to TD3 but uses a distributional C51-style critic outputting distributions over number of atoms (CC.1) atoms spaced equally between -VMax (CC.1) and VMax (CC.1) as well as N-step returns (CC.1) returns. In contrast to the original D4PG , we use a single actor and do not use prioritized replay. The target network is fully updated every 100100 training batches. As usual, we also sweep D4PG learning rate (CC.1).

Moreover, for off-policy algorithm (SAC, TD3 and D4PG) we sample replay ratio (CC.1) which denotes the average number of times each transition is replayed. This is achieved in the following way — if replay ratio (CC.1) ≥\geq batch size (CC.1) than we replay replay ratio (CC.1) // batch size (CC.1) batches (each with batch size (CC.1) transitions) after every environment step. If batch size (CC.1) >> replay ratio (CC.1), we replay a single batch every batch size (CC.1) // replay ratio (CC.1) transitions. The transitions for replay are sampled uniformly from a FIFO replay buffer of size RL replay buffer size (CC.1) and we start training whenever we have at least 10k transition in the buffer.

stochastic: sample from the distribution (same as behavioral policy used during training),

mode: use the mode of the Gaussian instead of sampling,

average: sample five action from the distribution and take the average of them.

C.2 Imitation-specific changes to RL

Let DD denote the probability that a state-action pair (s,a)(s,a) is classified as expert by the discriminator while hh is the discriminator logit, i.e. D=σ(h)D=\sigma(h) where σ\sigma denotes the sigmoid function. Depending on the value of reward function (CC.2) we use one of the following reward functions (for completeness we write the formulas as a function of DD as well as hh):

r(s,a)=ln⁡D−ln⁡(1−D)=hr(s,a)=\ln D-\ln(1-D)=h (introduced in AIRL ).

We also clip rewards with the absolute values higher than max reward magnitude (CC.2).

Absorbing state

We optionally (if absorbing state (CC.2)=True) apply the absorbing state technique from DAC . This technique encourages the agent to generate episodes of similar length to the ones of the expert. In particular, the demonstration and agent episodes are processed in the following way: for each terminal transition, we replace it with a non-terminal transition to a special absorbing stateIn practice, this is done by adding a special bit to every observation which is set to zero for normal observations and one for the absorbing state. The remaining bits of the absorbing state are all zeros. and also add a transition from the absorbing state to itself with a zero action.

Replaying demonstrations

For off-policy RL algorithms, we optionally (if policy-to-expert replay ratio (CC.2)≠∞\not=\infty) sample batches for RL training not only from the replay buffer, but also from the demonstrations. In particular, the ratio of policy to expert data in each minibatch is equal to policy-to-expert replay ratio (CC.2).

Initialization with behavior cloning

We optionally (if BC pretraining (CC.2)=True) pre-train the policy network offline at the beginning of training using Behavior Cloning . In particular, we perform 100k gradient steps with Adam on the MSE loss, using learning rate 10−410^{-4} and batch size 256256.

C.3 Discriminator parameterization

Depending on the value of discriminator input (CC.3), the discriminator is fed single states, state-action pairs, state-state pairs or state-action-state tuples.

Our basic discriminator architecture is an MLP with discriminator MLP depth (CC.3) hidden layers, each of size discriminator MLP width (CC.3) with the activation function specified by discriminator activation (CC.3). Its output is interpreted as the logit of the probability of being classified as expert, i.e. for a state-action-state tuple (s,a,s′)(s,a,s^{\prime}) we have D(s,a,s′)=σ(f(s,a,s′))D(s,a,s^{\prime})=\sigma(f(s,a,s^{\prime})), where DD is the probability of classifying the tuple (s,a,s′)(s,a,s^{\prime}) as expert, σ\sigma denotes the sigmoid function, and ff is a learnable function represented as an MLP.

We also consider two modifications introduced in the AIRL paper. The first one (enabled if reward shaping (CC.3)=True) adds a reward shaping term where the ff function is parameterized in the following way: f(s,a,s′)=g(s,a,s′)+γh(s′)−h(s)f(s,a,s^{\prime})=g(s,a,s^{\prime})+\gamma h(s^{\prime})-h(s) where gg and hh are MLPs parameterized as described above, and γ\gamma is the RL discount factor.The inputs fed to gg are specified by discriminator input (CC.3). The second modification (enabled if subtract log-pi (CC.3)=True) parameterizes the discriminator as D(s,a,s′)=exp⁡(f(s,a,s′))exp⁡(f(s,a,s′))+π(a∣s)D(s,a,s^{\prime})=\frac{\exp(f(s,a,s^{\prime}))}{\exp(f(s,a,s^{\prime}))+\pi(a|s)}, where π\pi is the current agent policy. It can be easy shown that it is equivalent to D(s,a,s′)=σ((f(s,a,s′)−log⁡π(a∣s))D(s,a,s^{\prime})=\sigma((f(s,a,s^{\prime})-\log\pi(a|s)) so this just shifts the logits by log⁡π(a∣s)\log\pi(a|s).

C.4 Discriminator training

All discriminator weight matrices use the lecun_uniform initializer from JAX . The last discriminator layer initialization is additionally multiplied by discriminator last layer init scale (CC.4).

The discriminator is trained with the Adam optimizer, the learning rate specified by discriminator learning rate (CC.4) and the cross-entropy loss. Each data batch contains exactly batch size (CC.1) expert transitions and batch size (CC.1) policy transitions. The policy transitions are sampled uniformly from a FIFO replay buffer of size discriminator replay buffer size (CC.4).

We perform discriminator to RL updates ratio (CC.4) discriminator gradient steps for each RL gradient step. More precisely, after each environment step, we compute the number of RL gradient steps as described in App. C.1, and perform discriminator to RL updates ratio (CC.4) that many discriminator gradient steps before performing the RL update.

C.5 Discriminator regularization

Depending on the value of discriminator regularizer (CC.5), we optionally apply one of the following regularizers to the discriminator:

Gradient penalty is parameterized with gradient penalty k (CC.5) and gradient penalty λ\lambda (CC.5). This regularizer adds an extra term in the discriminator loss that encourages the discriminator gradient to be close to kk on a convex combination of positive (expert) and negative (policy) data. In particular, for an expert data x∼Dexpertx\sim\mathcal{D}_{expert} and policy data x~∼Dπ\widetilde{x}\sim\mathcal{D}_{\pi}, the gradient penalty is defined as λ(∣∣∇x^D(x^)∣∣2−k)2\lambda(||\nabla_{\hat{x}}D(\hat{x})||_{2}-k)^{2}, where x^\hat{x} is a convex combination of xx and x~\widetilde{x}, i.e. x^:=ϵx+(1−ϵ)x~\hat{x}:=\epsilon x+(1-\epsilon)\widetilde{x} and ϵ\epsilon follows a uniform distribution: ϵ∼U\epsilon\sim U. In practice, kk is usually chosen to be (penalty for high gradients) or 11 (penalty for gradients with norms far from 11). Our gradient penalty implementation uses the gradient of the discriminator logit instead of the classification probability.

Spectral normalization [36]

Spectral normalization guarantees that the discriminator is 1-Lipschitz: ∣D(x2)−D(x1)∣≤∣∣x2−x1∣∣|D(x_{2})-D(x_{1})|\leq||x_{2}-x_{1}||. It does so by dividing each dense layer matrix by its highest eigenvalue which can be efficiently computed with the power iteration method. See for details.

Mixup [24]

Mixup is parameterized with mixup α\alpha (CC.5) and relies on training the discriminator on a convex combination of positive (expert) and negative (policy) data. With expert data x∼Dexpertx\sim\mathcal{D}_{expert} and policy data x~∼Dπ\widetilde{x}\sim\mathcal{D}_{\pi}, let ϵ\epsilon follow a Beta distribution: ϵ∼Beta(α,α)\epsilon\sim Beta(\alpha,\alpha). Instead of training the discriminator on xx and x~\widetilde{x} separately, we only train it on the convex combination of them x^:=ϵx+(1−ϵ)x~\hat{x}:=\epsilon x+(1-\epsilon)\widetilde{x} with the label being the convex combinations of the labels, i.e. expert with probability ϵ\epsilon and non-expert with probability 1−ϵ1-\epsilon, so that the loss is −ϵln⁡D(x^)−(1−ϵ)ln⁡(1−D(x^))-\epsilon\ln D(\hat{x})-(1-\epsilon)\ln(1-D(\hat{x})).

Positive Unlabeled GAIL (PUGAIL, [42])

Normally the discriminator is trained under the assumption that expert trajectories are positive examples and policy trajectories are negative examples. The PUGAIL loss assumes instead that policy trajectories are a mix of positive and negative examples.

With PUGAIL η\eta (CC.5) denoting the assumed proportion of positive samples in the policy data and PUGAIL β\beta (CC.5) being a clipping threshold, the discriminator is trained with the following loss:

Dropout [10]

We apply dropout to the hidden layers (dropout hidden rate (CC.5)) as well as inputs (dropout input rate (CC.5)). See for the description of dropout.

Weight decay [1, 34]

Weight decay is parameterized with a parameter controlling its strength weight decay λ\lambda (CC.5). Normally, weight decay is applied by adding a sum of the squares of the network parameters to the loss. However, this may interact negatively with an adaptive gradient optimizer like Adam unless the optimizer is modified appropriately . In our experiments, we use a version of Adam with weight decay called AdamW from the Optax library . See for the details.

Entropy bonus

Similarly to entropy bonus in RL, we also experiment with adding to the discriminator loss a term proportional to the entropy of the discriminator output treated as a Bernoulli distribution: λ(Dln⁡D+(1−D)ln⁡(1−D))\lambda\left(D\ln D+(1-D)\ln(1-D)\right) where entropy λ\lambda (CC.5) is a HP.

C.6 Observation normalization

We optionally apply input normalization (choice observation normalization (CC.6)) which transforms linearly the observations to all neural networks (in the RL algorithm as well as the discriminator) so that each coordinate has approximately mean equal zero and standard deviation equal one. This is done by subtracting from each observation μ\mu and dividing by max⁡(ρ,0.001)\max(\rho,0.001), where μ\mu and ρ\rho are the empirical mean and standard deviation of either all demonstrations (we call it fixed normalization because it does not change during training) or the empirical mean and standard deviation of all the observations encountered by the policy being trained so far (called online because it changes during training).

C.7 Combining multiple batches

We consider processing multiple batches at once for improved accelerator (GPU or TPU) utilization (choice number of combined batches (CC.7)). In particular, we sample an NN-times larger batch from a replay buffer, split it back into NN smaller/proper batches on an accelerator, and process them sequentially. In order to keep the replay ratio unaffected, we decrease the frequency of updates accordingly, e.g. instead of performing one gradient update for every environment step, we perform NN gradients updates every NN environment steps. We apply this technique to the discriminator as well as the RL agent training.

Appendix D Best hyperparameter values

Table 2 shows the best value found for each HP in the main experiment. See App. G for the full experimental report. The sample complexity can be slightly improved by decreasing number of combined batches (CC.7) and increasing discriminator to RL updates ratio (CC.4). We used the suboptimal values from Table 2 because they give a good trade-off between sample complexity and runtime. discriminator learning rate (CC.4) equal 10−610^{-6} is better when PUGAIL, entropy or no discriminator regularizer is used, and 3⋅10−53\cdot 10^{-5} is better otherwise. The performance of observation normalization schemes depends heavily on the environment and discriminator regularization used. For completeness, we present the best HPs for all discriminator regularizers.

Appendix E Expert and random policy scores

Appendix F Experiment wide

For each of the 10 tasks, we sampled 12083 choice configurations where we sampled the following choices independently and uniformly from the following ranges:

RL Algorithm (CC.1): {d4pg, ppo, sac, td3}

For the case “RL Algorithm (CC.1) = sac”, we further sampled the sub-choices:

SAC learning rate (CC.1): {0.0001, 0.0003, 0.001}

SAC entropy per dimension (CC.1): {-2.0, -1.0, -0.5, 0.0}

SAC polyak τ\tau (CC.1): {0.001, 0.003, 0.01, 0.03}

For the case “RL Algorithm (CC.1) = d4pg”, we further sampled the sub-choices:

D4PG learning rate (CC.1): {3e-05, 0.0001, 0.0003}

behavioral policy noise (CC.1): {0.1, 0.2, 0.3, 0.5}

number of atoms (CC.1): {51.0, 101.0, 201.0, 401.0}

For the case “RL Algorithm (CC.1) = td3”, we further sampled the sub-choices:

TD3 policy learning rate (CC.1): {0.0001, 0.0003, 0.001}

TD3 critic learning rate (CC.1): {0.0001, 0.0003, 0.001}

TD3 gradient clipping (CC.1): {40.0, ∞\infty}

behavioral policy noise (CC.1): {0.1, 0.2, 0.3, 0.5}

For the case “RL Algorithm (CC.1) = ppo”, we further sampled the sub-choices:

PPO learning rate (CC.1): {3e-05, 0.0001, 0.0003}

PPO number of epochs (CC.1): {2.0, 5.0, 10.0, 20.0}

PPO entropy cost (CC.1): {0.0, 0.001, 0.003, 0.01, 0.03, 0.1}

PPO number of minibatches (CC.1): {8.0, 16.0, 32.0, 64.0}

PPO unroll length (CC.1): {4.0, 8.0, 16.0, 32.0}

PPO clipping ϵ\epsilon (CC.1): {0.1, 0.2, 0.3}

GAE λ\lambda (CC.1): {0.8, 0.9, 0.95, 0.99}

RL replay buffer size (CC.1): {300000.0, 1000000.0, 3000000.0}

policy MLP width (CC.1): {64, 128, 256, 512}

critic MLP width (CC.1): {64, 128, 256, 512}

discount γ\gamma (CC.1): {0.9, 0.97, 0.99, 0.997}

discriminator replay buffer size (CC.4): {300000, 1000000, 3000000}

discriminator input (CC.3): {s, sa, sas, ss}

discriminator MLP depth (CC.3): {1, 2, 3}

discriminator MLP width (CC.3): {16, 32, 64, 128, 256, 512}

discriminator activation (CC.3): {elu, leaky_relu, relu, sigmoid, swish, tanh}

discriminator last layer init scale (CC.4): {0.001, 1.0}

discriminator regularizer (CC.5): {GP, Mixup, No regularizer, PUGAIL, dropout, entropy, spectral norm, weight decay}

For the case “discriminator regularizer (CC.5) = GP”, we further sampled the sub-choices:

gradient penalty λ\lambda (CC.5): {0.1, 1.0, 10.0}

For the case “discriminator regularizer (CC.5) = Mixup”, we further sampled the sub-choices:

For the case “discriminator regularizer (CC.5) = PUGAIL”, we further sampled the sub-choices:

PUGAIL β\beta (CC.5): {0.0, 0.7, ∞\infty}

For the case “discriminator regularizer (CC.5) = entropy”, we further sampled the sub-choices:

entropy λ\lambda (CC.5): {0.0003, 0.001, 0.003, 0.01, 0.03, 0.1, 0.3}

For the case “discriminator regularizer (CC.5) = weight decay”, we further sampled the sub-choices:

weight decay λ\lambda (CC.5): {0.3, 1.0, 3.0, 10.0, 30.0}

For the case “discriminator regularizer (CC.5) = dropout”, we further sampled the sub-choices:

dropout input rate (CC.5): {0.0, 0.25, 0.5, 0.75}

dropout hidden rate (CC.5): {0.25, 0.5, 0.75}

observation normalization (CC.6): {fixed, none}

evaluation behavior policy type (CC.1): {average, mode, stochastic}

discriminator learning rate (CC.4): {1e-06, 3e-06, 1e-05, 3e-05, 0.0001, 0.0003}

max reward magnitude (CC.2): {0.5, 1.0, 2.0, 5.0, 10.0, 50.0, ∞\infty}

reward function (CC.2): {-ln(1-D), AIRL, FAIRL, ln(D)}

discriminator to RL updates ratio (CC.4): {1}

F.2 Results

For each of the sampled choice configurations we compute the performance metric as described in Section 2. We report aggregate statistics of the experiment in Tables 4–7 as well as training curves in Figure 9. We further provide per-choice analyses in Figures 10-23.

Appendix G Experiment main

For each of the 10 tasks, we sampled 25334 choice configurations where we sampled the following choices independently and uniformly from the following ranges:

For the case “RL Algorithm (CC.1) = sac”, we further sampled the sub-choices:

SAC learning rate (CC.1): {0.0001, 0.0003, 0.001}

SAC entropy per dimension (CC.1): {-2.0, -1.0, -0.5, 0.0}

SAC polyak τ\tau (CC.1): {0.001, 0.003, 0.01, 0.03}

For the case “RL Algorithm (CC.1) = d4pg”, we further sampled the sub-choices:

D4PG learning rate (CC.1): {3e-05, 0.0001, 0.0003}

behavioral policy noise (CC.1): {0.1, 0.2, 0.3, 0.5}

number of atoms (CC.1): {51.0, 101.0, 201.0, 401.0}

For the case “RL Algorithm (CC.1) = td3”, we further sampled the sub-choices:

TD3 policy learning rate (CC.1): {0.0001, 0.0003, 0.001}

TD3 critic learning rate (CC.1): {0.0001, 0.0003, 0.001}

TD3 gradient clipping (CC.1): {40.0, ∞\infty}

behavioral policy noise (CC.1): {0.1, 0.2, 0.3, 0.5}

RL replay buffer size (CC.1): {300000, 1000000, 3000000}

policy MLP width (CC.1): {64, 128, 256, 512}

discriminator replay buffer size (CC.4): {300000, 1000000, 3000000}

discriminator input (CC.3): {s, sa, sas, ss}

discriminator MLP depth (CC.3): {1, 2, 3}

discriminator MLP width (CC.3): {16, 32, 64, 128, 256, 512}

discriminator activation (CC.3): {elu, leaky_relu, relu, sigmoid, swish, tanh}

discriminator last layer init scale (CC.4): {0.001, 1.0}

discriminator regularizer (CC.5): {GP, Mixup, No regularizer, PUGAIL, dropout, entropy, spectral norm, weight decay}

For the case “discriminator regularizer (CC.5) = GP”, we further sampled the sub-choices:

gradient penalty λ\lambda (CC.5): {0.1, 1.0, 10.0}

For the case “discriminator regularizer (CC.5) = Mixup”, we further sampled the sub-choices:

For the case “discriminator regularizer (CC.5) = PUGAIL”, we further sampled the sub-choices:

PUGAIL β\beta (CC.5): {0.0, 0.7, ∞\infty}

For the case “discriminator regularizer (CC.5) = entropy”, we further sampled the sub-choices:

entropy λ\lambda (CC.5): {0.0003, 0.001, 0.003, 0.01, 0.03, 0.1, 0.3}

For the case “discriminator regularizer (CC.5) = weight decay”, we further sampled the sub-choices:

weight decay λ\lambda (CC.5): {0.3, 1.0, 3.0, 10.0, 30.0}

For the case “discriminator regularizer (CC.5) = dropout”, we further sampled the sub-choices:

dropout input rate (CC.5): {0.0, 0.25, 0.5, 0.75}

dropout hidden rate (CC.5): {0.25, 0.5, 0.75}

observation normalization (CC.6): {fixed, none, online}

evaluation behavior policy type (CC.1): {average, mode, stochastic}

discriminator learning rate (CC.4): {1e-06, 3e-06, 1e-05, 3e-05, 0.0001, 0.0003}

reward function (CC.2): {-ln(1-D), AIRL, ln(D)}

discriminator to RL updates ratio (CC.4): {1}

G.2 Results

For each of the sampled choice configurations we compute the performance metric as described in Section 2. We report aggregate statistics of the experiment in Tables 8–11 as well as training curves in Figure 27. We further provide per-choice analyses in Figures 40-74.

Appendix H Experiment trade-offs

For each of the 10 tasks, we sampled 7991 choice configurations where we sampled the following choices independently and uniformly from the following ranges:

For the case “RL Algorithm (CC.1) = sac”, we further sampled the sub-choices:

SAC learning rate (CC.1): {0.0001, 0.0003, 0.001}

SAC entropy per dimension (CC.1): {-2.0, -1.0, -0.5, 0.0}

SAC polyak τ\tau (CC.1): {0.001, 0.003, 0.01, 0.03}

For the case “RL Algorithm (CC.1) = d4pg”, we further sampled the sub-choices:

D4PG learning rate (CC.1): {3e-05, 0.0001, 0.0003}

behavioral policy noise (CC.1): {0.1, 0.2, 0.3, 0.5}

number of atoms (CC.1): {51.0, 101.0, 201.0, 401.0}

For the case “RL Algorithm (CC.1) = td3”, we further sampled the sub-choices:

TD3 policy learning rate (CC.1): {0.0001, 0.0003, 0.001}

TD3 critic learning rate (CC.1): {0.0001, 0.0003, 0.001}

TD3 gradient clipping (CC.1): {40.0, ∞\infty}

behavioral policy noise (CC.1): {0.1, 0.2, 0.3, 0.5}

RL replay buffer size (CC.1): {300000, 1000000, 3000000}

policy MLP width (CC.1): {64, 128, 256, 512}

discriminator replay buffer size (CC.4): {300000, 1000000, 3000000}

discriminator input (CC.3): {s, sa, sas, ss}

discriminator MLP depth (CC.3): {1, 2, 3}

discriminator MLP width (CC.3): {16, 32, 64, 128, 256, 512}

discriminator activation (CC.3): {elu, leaky_relu, relu, sigmoid, swish, tanh}

discriminator last layer init scale (CC.4): {0.001, 1.0}

discriminator regularizer (CC.5): {GP, Mixup, No regularizer, PUGAIL, dropout, entropy, spectral norm, weight decay}

For the case “discriminator regularizer (CC.5) = GP”, we further sampled the sub-choices:

gradient penalty λ\lambda (CC.5): {0.1, 1.0, 10.0}

For the case “discriminator regularizer (CC.5) = Mixup”, we further sampled the sub-choices:

For the case “discriminator regularizer (CC.5) = PUGAIL”, we further sampled the sub-choices:

PUGAIL β\beta (CC.5): {0.0, 0.7, ∞\infty}

For the case “discriminator regularizer (CC.5) = entropy”, we further sampled the sub-choices:

entropy λ\lambda (CC.5): {0.0003, 0.001, 0.003, 0.01, 0.03, 0.1, 0.3}

For the case “discriminator regularizer (CC.5) = weight decay”, we further sampled the sub-choices:

weight decay λ\lambda (CC.5): {0.3, 1.0, 3.0, 10.0, 30.0}

For the case “discriminator regularizer (CC.5) = dropout”, we further sampled the sub-choices:

dropout input rate (CC.5): {0.0, 0.25, 0.5, 0.75}

dropout hidden rate (CC.5): {0.25, 0.5, 0.75}

observation normalization (CC.6): {fixed, none}

evaluation behavior policy type (CC.1): {average, mode, stochastic}

discriminator learning rate (CC.4): {1e-06, 3e-06, 1e-05, 3e-05, 0.0001, 0.0003}

replay ratio (CC.1): {64, 128, 256, 512, 1024}

batch size (CC.1): {64, 128, 256, 512, 1024}

discriminator to RL updates ratio (CC.4): {1, 2}

number of combined batches (CC.7): {1, 2, 4, 8, 16, 32, 64}

reward function (CC.2): {-ln(1-D), AIRL, ln(D)}

H.2 Results

For each of the sampled choice configurations we compute the performance metric as described in Section 2. We report aggregate statistics of the experiment in Tables 12–15 as well as training curves in Figure 78. We further provide per-choice analyses in Figures 79-82.

Appendix I Additional experiments