Generalization in Reinforcement Learning with Selective Noise Injection and Information Bottleneck
Maximilian Igl, Kamil Ciosek, Yingzhen Li, Sebastian Tschiatschek, Cheng Zhang, Sam Devlin, Katja Hofmann
Introduction
Deep Reinforcement Learning (rl) has been used to successfully train policies with impressive performance on a range of challenging tasks, including Atari , continuous control and tasks with long-ranged temporal dependencies . In those settings, the challenge is to be able to successfully explore and learn policies complex enough to solve the training tasks. Consequently, the focus of these works was to improve the learning performance of agents in the training environment and less attention was being paid to generalization to testing environments.
However, being able to generalize is a key requirement for the broad application of autonomous agents. Spurred by several recent works showing that most rl agents overfit to the training environment , multiple benchmarks to evaluate the generalization capabilities of agents were proposed, typically by procedurally generating or modifying levels in video games . How to learn generalizable policies in these environments remains an open question, but early results have shown the use of regularization techniques (like weight decay, dropout and batch normalization) established in the supervised learning paradigm can also be useful for rl agents . Our work builds on these results, but highlights two important differences between supervised learning and rl which need to be taken into account when regularizing agents.
First, because in rl the training data depends on the model and, consequently, the regularization method, stochastic regularization techniques like Dropout or BatchNorm can have adverse effects. For example, injecting stochasticity into the policy can lead to prematurely ending episodes, preventing the agent from observing future rewards. Furthermore, stochastic regularization can destabilize training through the learned critic and off-policy importance weights. To mitigate those adverse effects and effectively apply stochastic regularization techniques to rl, we propose Selective Noise Injection (sni). It selectively applies stochasticity only when it serves regularization and otherwise computes the output of the regularized networks deterministically. We focus our evaluation on Dropout and the Variational Information Bottleneck (vib), but the proposed method is applicable to most forms of stochastic regularization.
A second difference between rl and supervised learning is the non-stationarity of the data-distribution in rl. Despite many rl algorithms utilizing millions or even billions of observations, the diversity of states encountered early on in training can be small, making it difficult to learn general features. While it remains an open question as to why deep neural networks generalize despite being able to perfectly memorize the training data , it has been shown that the optimal point on the worst-case generalization bound requires the model to rely on a more compressed set of features the fewer data-points we have . Therefore, to bias our agent towards more general features even early on in training, we adapt the Information Bottleneck (ib) principle to an actor-critic agent, which we call Information Bottleneck Actor Critic (ibac). In contrast to other regularization techniques, ibac directly incentivizes the compression of input features, resulting in features that are more robust under a shifting data-distribution and that enable better generalization to held-out test environments.
We evaluate our proposed techniques using Proximal Policy Optimization (ppo), an off-policy actor-critic algorithm, on two challenging generalization tasks, Multiroom and Coinrun . We show the benefits of both ibac and sni individually as well as in combination, with the resulting ibac-sni significantly outperforming the previous state of the art results.
Background
We consider having a distribution of Markov decision processes (mdps) , with being a tuple consisting of state-space , action-space , transition distribution , reward function and initial state distribution . For training, we either assume unlimited access to (like in section 5.2, Multiroom) or restrict ourselves to a fixed set of training environments , (like in section 5.3, Coinrun).
with for epochs. The advantage is computed as in A2C . This is an efficient approximate trust region method , optimizing a pessimistic lower bound of the objective function on the collected data. It corresponds to estimating the gradient w.r.t the policy conservatively, since moving further away from , such that moves outside a chosen range , is only taken into account if it decreases performance. Similarly, the value function loss minimizes an upper bound on the squared error:
with a bootstrapped value function target and previous value function . The overall minimization objective is then:
where denotes an entropy bonus to encourage exploration and prevent the policy to collapse prematurely. In the following, we discuss regularization techniques that can be used to mitigate overfitting to the states and mdps so far seen during training.
In supervised learning, classifiers are often regularized using a variety of techniques to prevent overfitting. Here, we briefly present several major approaches which we either utilize as baseline or extend to rl in section 4.
Weight decay, also called L2 regularization, reduces the magnitude of the weights by adding an additional loss term . With a gradient update of the form , this decays the weights in addition to optimizing , i.e. we have .
Data augmentation refers to changing or distorting the available input data to improve generalization. In this work, we use a modified version of cutout , proposed by , in which a random number of rectangular areas in the input image is filled by random colors.
Batch Normalization normalizes activations of specified layers by estimating their mean and variance using the current mini-batch. Estimating the batch statistics introduces noise which has been shown to help improve generalization in supervised learning.
Another widely used regularization technique for deep neural networks is Dropout . Here, during training, individual activations are randomly zeroed out with a fixed probability . This serves to prevent co-adaptation of neurons and can be applied to any layer inside the network. One common choice, which we are following in our architecture, is to apply it to the last hidden layer.
Lastly, we will briefly describe the Variational Information Bottleneck (vib), a deep variational approximation to the Information Bottleneck (ib). While not typically used for regularization in deep supervised learning, we demonstrate in section 5 that our adaptation ibac shows strong performance in rl. Given a data distribution , the learned model is regularized by inserting a stochastic latent variable and minimizing the mutual information between the input and , , while maximizing the predictive power of the latent variable, i.e. . The vib objective function is:
where is the encoder, the decoder, the approximated latent marginal often fixed to a normal distribution and is a hyperparameter. For a normal distributed , eq. 4 can be optimized by gradient decent using the reparameterization trick .
The Problem of Using Stochastic Regularization in RL
We now take a closer look at a prototypical objective for training actor-critic methods and highlight important differences to supervised learning. Based on those observations, we propose an explanation for the finding that some stochastic optimization methods are less effective or can even be detrimental to performance when combined with other regularization techniques (see appendix D).
where we utilize a rollout policy {\color[rgb]{1,0.23,0.13}\definecolor[named]{pgfstrokecolor}{rgb}{1,0.23,0.13}\pgfsys@color@cmyk@stroke{0}{0.77}{0.87}{0}\pgfsys@color@cmyk@fill{0}{0.77}{0.87}{0}\pi^{r}_{\theta}} to collect trajectories. It can deviate from \color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}\pi_{\theta} but should be similar to keep the off-policy correction term {\nicefrac{{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}\pi_{\theta}}}{{\color[rgb]{1,0.23,0.13}\definecolor[named]{pgfstrokecolor}{rgb}{1,0.23,0.13}\pgfsys@color@cmyk@stroke{0}{0.77}{0.87}{0}\pgfsys@color@cmyk@fill{0}{0.77}{0.87}{0}\pi^{r}_{\theta}}}} low variance. In eq. 5, only the term {\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}\pi_{\theta}(a_{t}|s_{t})} is being updated and we highlight in orange all the additional influences of the learned policy and critic on the gradient.
From eqs. 5 and 6 we can see that the injection of noise into the computation of {\color[rgb]{1,0.23,0.13}\definecolor[named]{pgfstrokecolor}{rgb}{1,0.23,0.13}\pgfsys@color@cmyk@stroke{0}{0.77}{0.87}{0}\pgfsys@color@cmyk@fill{0}{0.77}{0.87}{0}\pi_{\theta}^{r}} and {\color[rgb]{1,0.23,0.13}\definecolor[named]{pgfstrokecolor}{rgb}{1,0.23,0.13}\pgfsys@color@cmyk@stroke{0}{0.77}{0.87}{0}\pgfsys@color@cmyk@fill{0}{0.77}{0.87}{0}V_{\theta}} can degrade performance in several ways: i) During rollouts using the rollout policy {\color[rgb]{1,0.23,0.13}\definecolor[named]{pgfstrokecolor}{rgb}{1,0.23,0.13}\pgfsys@color@cmyk@stroke{0}{0.77}{0.87}{0}\pgfsys@color@cmyk@fill{0}{0.77}{0.87}{0}\pi_{\theta}^{r}}, it can lead to undesirable actions, potentially ending episodes prematurely, and thereby deteriorating the quality of the observed data; ii) It leads to a higher variance of the off-policy correction term \nicefrac{{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}\pi_{\theta}}}{{\color[rgb]{1,0.23,0.13}\definecolor[named]{pgfstrokecolor}{rgb}{1,0.23,0.13}\pgfsys@color@cmyk@stroke{0}{0.77}{0.87}{0}\pgfsys@color@cmyk@fill{0}{0.77}{0.87}{0}\pi^{r}_{\theta}}} because the injected noise can be different for \color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}\pi_{\theta} and \color[rgb]{1,0.23,0.13}\definecolor[named]{pgfstrokecolor}{rgb}{1,0.23,0.13}\pgfsys@color@cmyk@stroke{0}{0.77}{0.87}{0}\pgfsys@color@cmyk@fill{0}{0.77}{0.87}{0}\pi_{\theta}^{r}, increasing gradient variance; iii) It increases variance in the gradient updates of both the policy and the critic through variance in the computation of {\color[rgb]{1,0.23,0.13}\definecolor[named]{pgfstrokecolor}{rgb}{1,0.23,0.13}\pgfsys@color@cmyk@stroke{0}{0.77}{0.87}{0}\pgfsys@color@cmyk@fill{0}{0.77}{0.87}{0}V_{\theta}}.
Method
To utilize the strength of noise-injecting regularization techniques in rl, we introduce Selective Noise Injection (sni)in the following section. Its goal is to allow us to make use of such techniques while mitigating the adverse effects the added stochasticity can have on the rl gradient computation. Then, in section 4.2, we propose Information Bottleneck Actor Critic (ibac)as a new regularization method and detail how sni applies to ibac, resulting in our state-of-the art method ibac-sni.
To apply sni to a regularization technique relying on noise-injection, we need to be able to temporarily suspend the noise and compute the output of the model deterministically. This is possible for most techniquesIn this work, we will focus on vib and Dropout as those show the most promising results without sni (see section 5) and will leave its application to other regularization techniques for future work.: For example, in Dropout, we can freeze one particular dropout mask, in vib we can pass in the mode instead of sampling from the posterior distribution and in Batch Normalization we can either utilize the moving average instead of the batch statistics or freeze and re-use one statistic multiple times. Formally, we denote by the version of a component , with the injected regularization noise suspended. Note that this does not mean that is deterministic, for example when the network approximates the parameters of a distribution.
Then, for sni we modify the policy gradient loss as follows: i) We use \color[rgb]{1,0.23,0.13}\definecolor[named]{pgfstrokecolor}{rgb}{1,0.23,0.13}\pgfsys@color@cmyk@stroke{0}{0.77}{0.87}{0}\pgfsys@color@cmyk@fill{0}{0.77}{0.87}{0}\bar{V}_{\theta} as critic instead of \color[rgb]{1,0.23,0.13}\definecolor[named]{pgfstrokecolor}{rgb}{1,0.23,0.13}\pgfsys@color@cmyk@stroke{0}{0.77}{0.87}{0}\pgfsys@color@cmyk@fill{0}{0.77}{0.87}{0}V_{\theta} in both eqs. 5 and 6, eliminating unnecessary noise through the critic; ii) We use \color[rgb]{1,0.23,0.13}\definecolor[named]{pgfstrokecolor}{rgb}{1,0.23,0.13}\pgfsys@color@cmyk@stroke{0}{0.77}{0.87}{0}\pgfsys@color@cmyk@fill{0}{0.77}{0.87}{0}\bar{\pi}^{r} as rollout policy instead of \color[rgb]{1,0.23,0.13}\definecolor[named]{pgfstrokecolor}{rgb}{1,0.23,0.13}\pgfsys@color@cmyk@stroke{0}{0.77}{0.87}{0}\pgfsys@color@cmyk@fill{0}{0.77}{0.87}{0}\pi^{r}. For some regularization techniques this will reduce the probability of undesirable actions; iii) We compute the policy gradient as a mixture between gradients for \color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}\pi_{\theta} and \color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}\bar{\pi}_{\theta} as follows:
The first term guarantees a lower variance of the off-policy importance weight, which is especially important early on in training when the network has not yet learned to compensate for the injected noise. The second term uses the noise-injected policy for updates, thereby taking advantage of its regularizing effects while still reducing unnecessary variance through the use of \color[rgb]{1,0.23,0.13}\definecolor[named]{pgfstrokecolor}{rgb}{1,0.23,0.13}\pgfsys@color@cmyk@stroke{0}{0.77}{0.87}{0}\pgfsys@color@cmyk@fill{0}{0.77}{0.87}{0}\bar{\pi}^{r} and \color[rgb]{1,0.23,0.13}\definecolor[named]{pgfstrokecolor}{rgb}{1,0.23,0.13}\pgfsys@color@cmyk@stroke{0}{0.77}{0.87}{0}\pgfsys@color@cmyk@fill{0}{0.77}{0.87}{0}\bar{V}_{\theta}. Note that sharing the rollout policy \color[rgb]{1,0.23,0.13}\definecolor[named]{pgfstrokecolor}{rgb}{1,0.23,0.13}\pgfsys@color@cmyk@stroke{0}{0.77}{0.87}{0}\pgfsys@color@cmyk@fill{0}{0.77}{0.87}{0}\bar{\pi}^{r} between both terms allows us to use the same collected data. Furthermore most computations are shared between both terms or can be parallelized.
2 Information Bottleneck Actor Critic
Early on in training an rl agent, we are often faced with little variation in the training data. Observed states are distributed only around the initial states , making spurious correlations in the low amount of data more likely. Furthermore, because neither the policy nor the critic have sufficiently converged yet, we have a high variance in the target values of our loss function.
This combination makes it harder and less likely for the network to learn desirable features that are robust under a shifting data-distribution during training and generalize well to held-out test mdps. To counteract this reduced signal-to-noise ratio, our goal is to explicitly bias the learning towards finding more compressed features which are shown to have a tighter worst-case generalization bound . While a higher compression does not guarantee robustness under a shifting data-distribution, we believe this to be a reasonable assumption in the majority of mdps, for example because they rely on a consistent underlying transition mechanism like physical laws.
To incentivize more compressed features, we use an approach similar to the vib , which minimizes the mutual information between the state and its latent representation while maximizing , the predictive power of on actions . To do so, we re-interpret the policy gradient update as maximization of the log-marginal likelihood of under the data distribution with discounted state distribution , advantage function and normalization constant . Taking the semi-gradient of this objective, i.e. assuming to be fixed, recovers the policy gradient:
Now, following the same steps as , we introduce a stochastic latent variable and minimize while maximizing under , resulting in the new objective:
We take the gradient and use the reparameterization trick to write the encoder as deterministic function with :
resulting in the overall loss function of the proposed Information Bottleneck Actor Critic (ibac)
with the hyperparameters , and balancing the loss terms.
Experiments
In the following, we present a series of experiments to show that the ib finds more general features in the low-data regime and that this translates to improved generalization in rl for ibac agents, especially when combined with sni. We evaluate our proposed regularization techniques on two environments, one grid-world with challenging generalization requirements in which most previous approaches are unable to find the solution and on the recently proposed Coinrun benchmark . We show that ibac-sni outperforms previous state of the art on both environments by a large margin. Details about the used hyperparameters and network architectures can be found in the Appendix, code to reproduce the results can be found at https://github.com/microsoft/IBAC-SNI/.
First we start in the supervised setting and show on a synthetic dataset that the vib is particularly strong at finding more general features in the low-data regime and in the presence of multiple signals with varying degrees of generality. Our motivation is that the low-data regime is commonly encountered in rl early on in training and many environments allow the agent to base its decision on a variety of features in the state, of which we would like to find the most general ones.
In fig. 1 we measure how the test performance of fully trained classification models varies for different regularization techniques when we i) vary the generality of and ii) vary the number of data-points in the training set. We find that most techniques perform comparably with the exception of the vib which is able to find more general features both in the low-data regime and in the presence of multiple features with only small differences in generality. In the next section, we show that this translates to faster training and performance gains in rl for our proposed algorithm ibac.
2 Multiroom
In this section, we show how ibac can help learning in rl tasks which require generalization. For this task, we do not distinguish between training and testing, but for each episode, we draw randomly from the full distribution over mdps . As the number of mdps is very large, learning can only be successful if the agent learns general features that are transferrable between episodes.
This experiment is based on . The aim of the agent is to traverse a sequence of rooms to reach the goal (green square in fig. 2) as quickly as possible. It takes discrete actions to rotate in either direction, move forward and toggle doors to be open or closed. The observation received by the agent includes the full grid, one pixel per square, with object type and object status (like direction) encoded in the 3 color channels. Crucially, for each episode, the layout is generated randomly by placing a random number of rooms in a sequence connected by one door each.
The results in fig. 2 show that ibac agents are much better at successfully learning to solve this task, especially for layouts with more rooms. While all other fully trained agents can solve less than 3% of the layouts with two rooms and none of the ones with three, ibac-sni still succeeds in an impressive 43% and 21% of those layouts. The difficulty of this seemingly simple task arises from its generalization requirements: Since the layout is randomly generated in each episode, each state is observed very rarely, especially for multi-room layouts, requiring generalization to allow learning. While in the 1 room layout the reduced policy stochasticity of the sni agent slightly reduces performance, it improves performance for more complex layouts in which higher noise becomes detrimental. In the next section we will see that this also holds for the much more complex Coinrun environment in which sni significantly improves the ibac performance.
3 Coinrun
On the previous environment, we were able to show that ibac and sni help agents to find more general features and to do so faster. Next, we show that this can lead to a higher final performance on previously unseen test environments. We evaluate our proposed regularization techniques on Coinrun , a recently proposed generalization benchmark with high-dimensional observations and a large variety in levels. Several regularization techniques were previously evaluated there, making it an ideal evaluation environment for ibac and sni. We follow the setting proposed in , using the same 500 levels for training and evaluate on randomly drawn, new levels of only the highest difficulty.
As have shown, combining multiple regularization techniques can improve performance, with their best- performing agent utilizing data augmentation, weight decay and batch normalization. As our goal is to push the state of the art on this environment and to accurately compare against their results, fig. 3 uses weight decay and data-augmentation on all experiments. Consequently, ‘Baseline’ in fig. 3 refers to only using weight decay and data-augmentation whereas the other experiments use Dropout, Batch Normalization or ibac in addition to weight decay and data-augmentation. Results without those baseline techniques can be found in appendix D.
First, we find that almost all previously proposed regularization techniques decrease performance compared to the baseline, see fig. 3 (left), with batch normalization performing worst, possibly due to its unusual interaction with weight decay . Note that this combination with batch normalization was the highest performing agent in . We conjecture that regularization techniques relying on stochasticity can introduce additional instability into the training update, possibly deteriorating performance, especially if their regularizing effect is not sufficiently different from what weight decay and data-augmentation already achieve. This result applies to both batch normalization and Dropout, with and without sni, although sni mitigates the adverse effects. Consequently, we can already improve on the state of the art by only relying on those two non-stochastic techniques. Furthermore, we find that ibac in combination with sni is able to significantly outperform our new state of the art baseline. We also find that for ibac, achieves better performance than , justifying using both terms in eq. 7.
Related Work
Generalization in RL can take a variety of forms, each necessitating different types of regularization. To position this work, we distinguished two types that, whilst not mutually exclusive, we believe to be conceptually distinct and found useful to isolate when studying approaches to improve generalization.
The first type, robustness to uncertainty refers to settings in which the unobserved mdp influences the transition dynamics or reward structure. Consequently the current state might not contain enough information to act optimally in the current mdp and we need to find the action which is optimal under the uncertainty about . This setting often arises in robotics and control where exact physical characteristics are unknown and domain shifts can occur . Consequently, domain randomization, the injection of randomness into the environment, is often purposefully applied during training to allow for sim-to-real transfer . Noise can be injected into the states of the environment or the parameters of the transition distribution like friction coefficients or mass values . The noise injected into the dynamics can also be manipulated adversarially . As the goal is to prevent overfitting to specific mdps, it also has been found that using smaller or simpler networks can help. We can also aim to learn an adaptive policy by treating the environment as partially observable Markov decision process (pomdp) (similar to viewing the learning problem in the framework of Bayesian RL ) or as a meta-learning problem .
On the other hand, we distinguish feature robustness, which applies to environments with high-dimensional observations (like images) in which generalization to previously unseen states can be improved by learning to extract better features, as the focus for this paper. Recently, a range of benchmarks, typically utilizing procedurally generated levels, have been proposed to evaluate this type of generalization .
Improving generalization in those settings can rely on generating more diverse observation data , or strong, often relational, inductive biases applied to the architecture . Contrary to the results in continuous control domains, here deeper networks have been found to be more successful . Furthermore, this setting is more similar to that of supervised learning, so established regularization techniques like weight decay, dropout or batch-normalization have also successfully been applied, especially in settings with a limited number of training environments . This is the work most closely related to ours. We build on those results and improve upon them by taking into account the specific ways in which rl is different from the supervised setting. They also do not consider the vib as a regularization technique.
Combining rl and vib has been recently explored for learning goal-conditioned policies and meta-RL . Both of these previous works also differ from the ibac architecture we propose by conditioning action selection on both the encoded and raw state observation. These studies complement the contribution made here by providing evidence that the vib can be used with a wider range of rl algorithms including demonstrated benefits when used with Soft Actor-Critic for continuous control in MuJoCo and on-policy A2C in MiniGrid and MiniPacMan .
Conclusion
In this work we highlight two important differences between supervised learning and rl: First, the training data is generated using the learned model. Consequently, using stochastic regularization methods can induce adverse effects and reduce the quality of the data. We conjecture that this explains the observed lower performance of Batch Normalization and Dropout. Second, in rl, we often encounter a noisy, low-data regime early on in training, complicating the extraction of general features.
We argue that these differences should inform the choice of regularization techniques used in RL. To mitigate the adverse effects of stochastic regularization, we propose Selective Noise Injection (sni)which only selectively injects noise into the model, preventing reduced data quality and higher gradient variance through a noisy critic. On the other hand, to learn more compressed and general features in the noisy low-data regime, we propose Information Bottleneck Actor Critic (ibac), which utilizes an variational information bottleneck as part of the agent.
We experimentally demonstrate that the vib is able to extract better features in the low-data regime and that this translates to better generalization of ibac in rl. Furthermore, on complex environments, sni is key to good performance, allowing the combined algorithm, ibac-sni, to achieve state of the art on challenging generalization benchmarks. We believe the results presented here can inform a range of future works, both to improve existing algorithms and to find new regularization techniques adapted to rl.
Acknowledgments
We would like to thank Shimon Whiteson for his helpful feedback, Sebastian Lee, Luke Harris, Hiske Overweg and Patrick Fernandes for help with experimental evaluations and Adrian O’Grady, Jaroslaw Rzepecki and Andre Kramer for help with the computing infrastructure. M. Igl is supported by the UK EPSRC CDT on Autonomous Intelligent Machines and Systems.
References
Appendix A Dropout with SNI
In order to apply sni to Dropout, we need to decide how to ‘suspend’ the noise to compute . While one could apply no dropout mask and scale the activations accordingly, we empirically found it to be better to instead sample one dropout mask and keep it fixed for all gradient updates using the thus collected data. This follows the implementation used in .
Appendix B Supervised Classification Task
The network consist of a 1D-convolutional layer with 10 filters and a kernel size of 11 followed by two hidden, fully connected layers of size 1024 and 256 and the last layer which outputs logits. When the vib or Dropout are used, they are applied to the last hidden layer. We use a learning rate of . The relative weight for weight decay was , which performed best out of . For the vib we used , which performed best out of . Lastly, For dropout we tested the dropout rates , out of which performed best. Our results were stable across a range of hyperparameters, see fig. 5.
Data generation
For each data-point, after drawing the class label , we want to encode the information about in two ways, using the encoding functions and which use one of and different patterns to encode the information. The larger is, the less general the encoding is as it applies to fewer data-points. Note that there are different patterns per class.
We generate the patterns by first generating a set of random functions and by randomly drawing Fourier coefficients from $d_{x}f^{c}d_{x}g^{c}d_{g} To encode the information about we first choose one pattern from (slightly overloading notation between functions and pattern-vectors) and add some noise: Next, to also encode the information about using , we choose one of the patterns and replace a part of the vector , which is possible because the patterns are shorter: . The location of replacement is randomly drawn for each data-point, but restricted to a a set of possible locations which are also random, but kept fixed for the experiment and the same between training and testing set. The process is pictured in fig. 4. By changing the number of possible locations and the strength of the noise added to , ,we can tune the relative difficulty of learning to recognize patterns and , allowing us to find a regime where both can be found. Within this regime, our qualitative results were stable. We use and . Furthermore, we have for the dimension of of the observations , and for the size of the patterns we have . We use different classes. The observation space measures where the channels are used to encode object type and object features like orientation or ‘open/closed’ and ‘color’ for doors on each of the spatial locations (see fig. 2 for a typical layout for ). The agent uses a 3-layer CNN with and filters respectively. All layers use a kernel of size . After the CNN, it uses one hidden layer of size 64 to which ibac or Dropout are applied if they are used. Dropout uses and was tested for . Both weight decay and ibac were tried with a weighting factor of , with performing best for weight decay and performing best for ibac. The output of the hidden layer is fed into a value function head and the policy head. We use a discount factor , a learning rate of , generalized value estimation with , an entropy coefficient of , value loss coefficient , gradient clipping at , and ppo with the Adam optimizer . We use the same architecture (’Impala’) and default policy gradient hyperparameters as well as the codebase (https://github.com/openai/coinrun) from the authors of to ensure staying as closely as possible to their proposed benchmark. Dropout and ibac where applied to the last hidden layer and both, as well as weight decay, were tried with the same set of hyperparameters as in Multiroom. The best performance was achieved with for Dropout and for ibac and weight decay. Batch normalization was applied between the layers of the convolutional part of the network. Note that the original architecture in uses Dropout also on earlier layers, however, we achieve higher performance with our implementation. In fig. 6 (left) we show results for Dropout with and without sni and for and . We find that learns fastest, possible due to the high importance weight variance in the stochastic term in sni for (see fig. 3 (right)). However, all Dropout implementations converge to roughly the same value, significantly below the ‘baseline’ agent, indicating that Dropout is not suitable for combination with weight decay and data augmentation. In fig. 6 (right) we show the test performance for ibac and Dropout with and without sni, without using weight decay and data-augmentation. Again, we can see that sni helps the performance. Interestingly, we can see that ibac does not prevent overfitting by itself (one can see the performance decreasing for longer training) but does lead to faster learning. Our conjecture is that it finds more general features early on in training, but ultimately overfits to the test-set of environments without additional regularization. This further indicates that it’s regularization is different to techniques such as weight decay, explaining why their combination synergizes well. In fig. 7 we show the training set performance of our experiments.Appendix C Multiroom
Appendix D Coinrun