Generalization and Regularization in DQN

Jesse Farebrother, Marlos C. Machado, Michael Bowling

Introduction

Recently, reinforcement learning (RL) algorithms have proven very successful on complex high-dimensional problems, in large part due to the use of deep neural networks for function approximation (e.g., Mnih et al., 2015; Silver et al., 2016). Despite the generality of the proposed solutions, applying these algorithms to slightly different environments often requires agents to learn the new task from scratch. The learned policies rarely generalize to other domains and the learned representations are seldom reusable. On the other hand, deep neural networks are lauded for their generalization capabilities (e.g., Lecun et al., 1998), with some communities heavily relying on reusing learned representations in different problems. In light of the successes of supervised learning methods, the lack of generalization or reusable knowledge (i.e., policies, representation) acquired by current deep RL algorithms is somewhat surprising.

In this paper we investigate whether the representations learned by deep RL methods can be generalized, or at the very least reused and refined on small variations to the task at hand. We evaluate the generalization capabilities of DQN (Mnih et al., 2015), one of the most representative algorithms in the family of value-based deep RL methods; and we further explore whether the experience gained by the supervised learning community to improve generalization and to avoid overfitting can be used in deep RL. We employ conventional supervised learning techniques such as regularization and fine-tuning (i.e., reusing and refining the representation) to DQN and we show that a learned representation trained with regularization allows us to learn more general features that can be reused and fine-tuned.

We are interested in agents that generalize across tasks that have similar underlying dynamics but that have different observation spaces. In this context, we see generalization as the agent’s ability to abstract aspects of the environment that do not matter. The main contributions of this work are:

We propose the use of the new modes and difficulties of Atari 2600 games as a platform for evaluating generalization in RL and we provide the first baseline results in this platform. These game modes allow agents to be trained in one environment and evaluated in a slightly different environment that still captures key concepts of the original environment (e.g., game sprites, dynamics).

Under this new notion of generalization in RL, we thoroughly evaluate the generalization capabilities of DQN and we provide evidence that it exhibits an overfitting trend.

Inspired by the current literature in regularizing deep neural networks to improve robustness and adaptability, we apply regularization techniques to DQN and show they vastly improve its sample efficiency when faced with new tasks. We do so by analyzing the impact of regularization on the policy’s ability to not only perform zero-shot generalization, but to also learn a more general representation amenable to fine-tuning on different problems.

Background

We begin our exposition with an introduction of basic terms and concepts for supervised learning and reinforcement learning. We then discuss the related work, focusing on generalization in reinforcement learning.

Another popular regularization technique is dropout (Srivastava et al., 2014). When using dropout, during forward propagation each neural unit is set to zero according to a Bernoulli distribution with probability p∈p\in, referred to as the dropout rate. Dropout discourages the network from relying on a small number of neurons to make a prediction, making memorization of the dataset harder.

Prior to training, the network parameters are usually initialized through a stochastic process such as Xavier initialization (Glorot and Bengio, 2010). We can also initialize the network using pre-trained weights from a different task. If we reuse one or more pre-trained layers we say the weights encoded by those layers will be fine-tuned during training (e.g., Razavian et al., 2014; Long et al., 2015), a topic we explore in Section 6.

2 Reinforcement Learning

where tt denotes the current timestep and α\alpha the step size. Generally, due to the exploding size of the state space in many real-world problems, it is intractable to learn a state-action pairing for the entire MDP. Instead we learn an approximation to the true function qπq_{\pi}.

DQN approximates the state-action value function such that Q(s,a; θ)≈qπ(s,a)Q(s,a;\,\theta)\approx q_{\pi}(s,a), where θ\theta denotes the weights of a neural network. The network takes as input some encoding of the current state StS_{t} and outputs ∣A∣|\mathcal{A}| scalars corresponding to the state-action values for StS_{t}. DQN is trained to minimize

where τ=(St,At,Rt+1,St+1)\tau=(S_{t},A_{t},R_{t+1},S_{t+1}) are uniformly sampled from U(⋅)U(\cdot), the experience replay buffer filled with experience collected by the agent. The weights θ−\theta^{-} of a duplicate network are updated less frequently for stability purposes.

3 Related Work

In reinforcement learning, regularization is rarely applied to value-based methods. The few existing studies often focus on single-task settings with linear function approximation (e.g., Farahmand et al., 2008; Kolter and Ng, 2009). Here we look at the reusability, in different tasks, of learned representations. The closest work to ours is Cobbe et al.’s (2019), which also looks at regularization techniques applied to deep RL. However, different from Cobbe et al., here we also evaluate the impact of regularization when fine-tuning value functions. Moreover, in this paper we propose a different platform for evaluating generalization in RL, which we discuss below.

There are several recent papers that support our results with respect to the limited generalization capabilities of deep RL agents. Nevertheless, they often investigate generalization in light of different aspects of an environment such as noise (e.g., Zhang et al., 2018a) and start state distribution (e.g., Rajeswaran et al., 2017; Zhang et al., 2018a, b). There are also some proposals for evaluating generalization in RL through procedurally generated or parametrized environments (e.g., Finn et al., 2017; Juliani et al., 2019; Justesen et al., 2018; Whiteson et al., 2011; Witty et al., 2018). These papers do not investigate generalization in deep RL the same way we do. Moreover, as aforementioned, here we also propose using a different testbed, the modes and difficulties of Atari 2600 games. With respect to that, Witty et al.’s (2018) work is directly related to ours, as they propose parameterizing a single Atari 2600 game, Amidar, as a way to evaluate generalization in RL. The use of modes and difficulties is much more comprehensive and it is free of experimenters’ bias.

In summary, our work adds to the growing literature on generalization in reinforcement learning. To the best of our knowledge, our paper is the first to discuss overfitting in Atari 2600 games, to present results using the Atari 2600 modes as testbed, and to demonstrate the impact regularization can have in value function fine-tuning in reinforcement learning.

The ALE as a Platform for Evaluating Generalization in Reinforcement Learning

The Arcade Learning Environment (ALE) is a platform used to evaluate agents across dozens of Atari 2600 games (Bellemare et al., 2013). It is one of the standard evaluation platforms in the field and has led to several exciting algorithmic advances (e.g., Mnih et al., 2015). The ALE poses the problem of general competency by having agents use the same learning algorithm to perform well in as many games as possible, without using any game specific knowledge. Learning to play multiple games with the same agent, or learning to play a game by leveraging knowledge acquired in a different game is harder, with fewer successes being known (Rusu et al., 2016; Kirkpatrick et al., 2016; Parisotto et al., 2016; Schwarz et al., 2018; Espeholt et al., 2018).

Throughout this paper we evaluate the generalization capabilities of our agents using hold out test environments. We do so with different modes and difficulties of Atari 2600 games, features the ALE recently started to support (Machado et al., 2018). Game modes, which were originally native to the Atari 2600 console, generally give us modifications of each Atari 2600 game by modifying sprites, velocities, and the observability of objects. These modes offer an excellent framework for evaluating generalization in RL. They were designed several decades ago and remain free from experimenter’s bias as they were not designed with the goal of being a testbed for AI agents, but with the goal of being variedThere are 48 Atari 2600 games with more than one flavour in the ALE. These games have 414 different flavours (Machado et al., 2018). Notice that, on average, each game has less than 10 flavours though. This is another challenge since other settings often assume access to many more environment variations (e.g., via procedural content generation). and entertaining to humans. Figure 1 depicts some of the different modes and difficulties available in the ALE. As Machado et al. (2018), hereinafter we call each mode/difficult pair a flavour.

Besides having the properties that made the ALE successful in the RL community, the different game flavours allow us to look at the problem of generalization in RL from a different perspective. Because of hardware limitations, the different flavours of an Atari 2600 game could not be too different from each other.The Atari 2600 console has only 2KB of RAM. Therefore, different flavours can be seen as small variations of the default game, with few latent variables being changed. In this context, we pose the problem of generalization in RL as the ability to identify invariances across tasks with high-dimensional observation spaces. Such an objective is based on the assumption that the underlying dynamics of the world does not vary much. Instead of requiring an agent to play multiple games that are visually very different or even non-analogous, the notion of generalization we propose requires agents to play games that are visually very similar and that can be played with policies that are conceptually similar, at least from a human perspective. In a sense, the notion of generalization we propose requires agents to be invariant to changes in the observation space.

Introducing flavours to the ALE is not one of our contributions, this was done by Machado et al. (2018). Nevertheless, here we provide a first concrete suggestion on how to use these flavours in reinforcement learning. Our paper also provides the first baseline results for different flavours of Atari 2600 games since Machado et al. (2018) incorporated them to the ALE but did not report any results on them. The baseline results for the traditional deep RL setting are available in Table 5 while the full baseline results for regularization are available in Table 6. Because these baseline results are quite broad, encompassing multiple games and flavours, and because we wanted to first discuss other experiments and analyses, Tables 5 and 6 are at the end of the paper. They follow Machado et al.’s (2018) suggestions on how to report Atari 2600 games results.

We believe our proposal is a more realistic and tractable way of defining generalization in decision-making problems. Instead of focusing on the samples (s,a,s′,r)(s,a,s^{\prime},r), simply requiring them be drawn from the same distribution, we look at a more general notion of generalization where we consider multiple tasks, with the assumption that tasks are sampled from the same distribution, similar to the meta-RL setting. Nevertheless, we concretely constrain the distribution of tasks with the notion that only few latent variables describing the environment can vary. This also allows us to have a new perspective towards an agents’ inability to succeed in slightly different tasks from those they are trained on. At the same time, this is more challenging than using, for example, different parametrizations of an environment, as often done when evaluating meta-RL algorithms. In fact, we could not obtain any positive results in these Atari 2600 games with traditional meta-RL algorithms (e.g., Finn et al., 2017; Nichol et al., 2018a) and to the best of our knowledge, there are no reports of meta-RL algorithms succeeding in Atari 2600 games. Because of that, we do not further discuss these approaches.

In this paper we focus on a subset of Atari 2600 games with multiple flavours. Because we wanted to provide exhaustive results averaging over multiple trials, here we use 1313 flavours obtained from 44 games: Freeway, HERO, Breakout, and Space Invaders. In Freeway, the different modes vary the speed and number of vehicles, while different difficulties change how the player is penalized for running into a vehicle. In HERO, subsequent modes start the player off at increasingly harder levels of the game. The mode we use in Breakout makes the bricks partially observable. Modes of Space Invaders allow for oscillating shield barriers, increasing the width of the player sprite, and partially observable aliens. Figure 1 depicts some of these flavours and Figure 2 further explains the difference between the ALE flavours we used.Videos of the different modes are available in the following link: https://goo.gl/pCvPiD.

Generalization of the Policies Learned by DQN

In order to test the generalization capabilities of DQN, we first evaluate whether a policy learned in one flavour can perform well in a different flavour. As aforementioned, different modes and difficulties of a single game look very similar. If the representation encodes a robust policy we might expect it to be able to generalize to slight variations of the underlying reward signal, game dynamics, or observations. Evaluating the learned policy in a similar but different flavour can be seen as evaluating generalization in RL, similar to cross-validation in supervised learning.

To evaluate DQN’s ability to generalize across flavours, we evaluate the learned ϵ\epsilon-greedy policy on a new flavour after training for 50M frames in the default flavour, m0d0 (mode 0, difficulty 0). We measure the cumulative reward averaged over 100 episodes in the new flavour, adhering to the evaluation protocol suggested by Machado et al. (2018). The results are summarized in Table 1. Baseline results where the agent is trained from scratch for 50M frames in the target flavour used for evaluation are reported in the baseline column Learn Scratch. Theoretically, this baseline can be seen as an upper bound on the performance DQN can achieve in that flavour, as it represents the agent’s performance when evaluated in the same flavour it was trained on. Full baseline results with the agent’s performance after different number of frames can be found in Tables 5 and 6.

We can see in the results that the policies learned by DQN do not generalize well to different flavours, even when the flavours are remarkably similar. For example, in Freeway, a high-level policy applicable to all flavours is to go up while avoiding cars. This does not seem to be what DQN learns. For example, the default flavour m0d0 and m4d0 comprise of exactly the same sprites, the only difference is that in m4d0 some cars accelerate and decelerate over time. The close to optimal policy learned in m0d0 is only able to score 15.8 points when evaluated on m4d0, which is approximately half of what the policy learned from scratch in that flavour achieves (29.9 points). The learned policy when evaluated on flavours that differ more from m0d0 perform even worse (for example, when a new sprite is introduced, or when there are more cars in each lane).

As aforementioned, the different modes of HERO can be seen as giving the agent a curriculum or a natural progression. Interestingly, the agent trained in the default mode for 50M frames can progress to at least level 3 and sometimes level 4. Mode 1 starts the agent off at level 5 and performance in this mode suffers greatly during evaluation. There are very few game mechanics added to level 5, indicating that perhaps the agent is memorizing trajectories instead of learning a robust policy capable of solving each level.

Results in some flavours suggest that the agent is overfitting to the flavour it is trained on. We tested this hypothesis by periodically evaluating the learned policy in each other flavour of that game. This process involved taking checkpoints of the network every 500,000500{,}000 frames and evaluating the ϵ\epsilon-greedy policy in the prescribed flavour for 100100 episodes, further averaged over five runs. The results obtained in Freeway, the most pronounced game in which we observe overfitting, are depicted in Figure 3. Learning curves for all flavours can be found in the Appendix.

In Freeway, while we see the policy’s performance flattening out in m4d0, we do see the traditional bell-shaped curve associated to overfitting in the other modes. At first, improvements in the original policy do correspond to improvements in the performance of that policy in other flavours. With time, it seems that the agent starts to refine its policy for the specific flavour it is being trained on, overfitting to that flavour. With other game flavours being significantly more complex in their dynamics and gameplay, we do not observe this prominent bell-shaped curve.

In conclusion, when looking at Table 1, it seems that the policies learned by DQN struggle to generalize to even small variations encountered in game flavours. The results in Freeway even exhibit a troubling notion of overfitting. Nevertheless, being able to generalize across small variations of the task the agent was trained on is a desirable property for truly autonomous agents. Based on these results we evaluate whether deep RL can benefit from established methods from supervised learning promoting generalization.

Regularization in DQN

We follow the same evaluation scheme described when evaluating the non-regularized policy to different flavours. We evaluate the policy learned after 50M frames of the default mode of each game. We contrast these results with the results presented in the previous section. This evaluation protocol allows us to directly evaluate the effect of regularization on the learned policy’s ability to generalize. The results are presented in Table 2, on the next page, and the evaluation curves are available in the Appendix.

When using regularization during training we sometimes observe a performance hit in the default flavour. Dropout generally requires increased training iterations to reach the same level of performance one would reach when not using dropout. However, maximal performance in one flavour is not our goal. We are interested in the setting where one may be willing to take lower performance on one task in order to obtain higher performance, or adaptability, on future tasks. Full baseline results using regularization can also be found in Table 6.

In most flavours, when looking at Table 2, we see that evaluating the policy trained with regularization does not negatively impact performance when compared to the performance of the policy trained without regularization. In some flavours we even see an increase in performance. When using regularization the agent’s performance in Freeway improves for all flavours and the agent even learns a policy capable of outperforming the baseline learned from scratch in two of the three flavours. Moreover, in Freeway we now observe increasing performance during evaluation throughout most of the learning procedure as depicted in Figure 4, on the next page. These results seem to confirm the notion of overfitting observed in Figure 3.

Despite slight improvements from these techniques, regularization by itself does not seem sufficient to enable policies to generalize across flavours. Learning from scratch in these new flavours is still more beneficial than re-using a policy learned with regularization. As shown in the next section, the real benefit of regularization in deep RL seems to come from the ability to learn more general features. These features lead to a more adaptable representation which can be reused and subsequently fine-tuned on other flavours.

Value function fine-tuning

We hypothesize that the benefit of regularizing deep RL algorithms may not come from improvements during evaluation, but instead in having a good parameter initialization that can be adapted to new tasks that are similar. We evaluate this hypothesis using two common practices in machine learning. First, we use the weights trained with regularization as the initialization for the entire network. We subsequently fine-tune all weights in the network. This is similar to what classification methods do in computer vision problems (e.g., Razavian et al., 2014). Secondly, we evaluate reusing and fine-tuning only early layers of the network. This has been shown to improve generalization in some settings (e.g., Yosinski et al., 2014), and is sometimes used in natural language processing problems (e.g., Mou et al., 2016; Howard and Ruder, 2018).

In this setting we take the weights of the network trained in the default flavour for 50M frames and use them to initialize the network commencing training in the new flavour for 50M frames. We perform this set of experiments twice (for the weights trained with and without regularization, as described in the previous section). Each run is averaged over five seeds. For comparison, we provide a baseline trained from scratch for 50M and 100M frames in each flavour. Directly comparing the performance obtained after fine-tuning to the performance after 50M frames (Scratch) shows the benefit of re-using a representation learned in a different task instead of randomly initializing the network. Comparing the performance obtained after fine-tuning to the performance of 100M frames (Scratch) lets us take into consideration the sample efficiency of the whole learning process. The results are presented on the next page, in Table 3.

Fine-tuning from a non-regularized representation yields conflicting conclusions. Although in Freeway we obtained positive fine-tuning results, we note that rewards are so sparse in mode 1 that this initialization is likely to be acting as a form of optimistic initialization, biasing the agent to go up. The agent observes rewards more often, therefore, it learns quicker about the new flavour. However, the agent is still unable to reach the maximum score in these flavours.

The results of fine-tuning the regularized representation are more exciting. In Freeway we observe the highest scores on m1d0 and m1d1 throughout the whole paper. In HERO we vastly outperform fine-tuning from a non-regularized representation. In Space Invaders we obtain higher scores across the board when comparing to the same amount of experience. These results suggest that reusing a regularized representation in deep RL might allow us to learn more general features which can be more successfully fine-tuned.

Initializing the network with a regularized representation also seems to be better than initializing the network randomly, that is, when learning from scratch. These results are impressive when we consider the potential regularization has in reducing the sample complexity of deep RL algorithms. Initializing the network with a regularized representation seems even better than learning from scratch when we take the total number of frames seen between two flavours into consideration. When we look at the rows Regularized Fine-tuning and Scratch in Table 3 we are comparing two algorithms that observed 100M frames. However, to generate the results in the column Scratch for two flavours we used 200M frames while we only used used 150M frames to generate the results in the column Regularized Fine-tuning (50M frames are used to learn in the default flavour and then 50M frames are used in each flavour you actually care about). Obviously, this distinction becomes larger as more tasks are taken into consideration.

2 Fine-Tuning Early Layers to Learn Co-Adaptations

We also investigated which layers may encode general features able to be fine-tuned. We were inspired by other studies showing that neural networks can re-learn co-adaptations when their final layers are randomly initialized, sometimes improving generalization (Yosinski et al., 2014). We conjectured DQN may benefit from re-learning the co-adaptations between early layers comprising general features and the randomly initialized layers which ultimately assign state-action values. We hypothesized that it might be beneficial to re-learn the final layers from scratch since state-action values are ultimately conditioned on the flavour at hand. Therefore, we also evaluated whether fine-tuning only the convolutional layers, or the convolutional layers and the first fully connected layer, was more effective than fine-tuning the whole network. This does not seem to be the case. The performance when we fine-tune the whole network is consistently better than when we re-learn co-adaptations, as shown in Table 4.

Discussion and conclusion

Many studies have tried to explain generalization of deep neural networks in supervised learning settings (e.g., Zhang et al., 2018b; Dinh et al., 2017). Analyzing generalization and overfitting in deep RL has its own issues on top of the challenges posed in the supervised learning case. Actually, generalization in RL can be seen in different ways. We can talk about generalization in RL in terms of conditioned sub-goals within an environment (e.g., Andrychowicz et al., 2017; Sutton, 1995), learning multiple tasks at once (e.g., Teh et al., 2017; Parisotto et al., 2016), or sequential task learning as in a continual learning setting (e.g., Schwarz et al., 2018; Kirkpatrick et al., 2016). In this paper we evaluated generalization in terms of small variations of high-dimensional control tasks. This provides a candid evaluation method to study how well features and policies learned by deep neural networks in RL problems can generalize. The approach of studying generalization with respect to the representation learning problem intersects nicely with the aforementioned problems in RL where generalization is key.

Some of the results here can also be seen under the light of curriculum learning. The regularization techniques we have evaluated here seem to be effective in leveraging situations where an easier task is presented first, sometimes leading to unseen performance levels (e.g., Freeway).

Finally, it is obvious that we want algorithms that can generalize across tasks. Ultimately we want agents that can keep learning as they interact with the world in a continual learning fashion. We believe the flavours of Atari 2600 games can be a stepping stone towards this goal. Our results suggested that regularizing and fine-tuning representations in deep RL might be a viable approach towards improving sample efficiency and generalization on multiple tasks. It is particularly interesting that fine-tuning a regularized network was the most successful approach because this might also be applicable in the continual learning settings where the environment changes without the agent being told so, and re-initializing layers of a network is obviously not an option.

Acknowledgments

The authors would like to thank Matthew E. Taylor, Tom van de Wiele, and Marc G. Bellemare for useful discussions, as well as Vlad Mnih for feedback on a preliminary draft of the manuscript. This work was supported by funding from NSERC and Alberta Innovates Technology Futures through the Alberta Machine Intelligence Institute (Amii). Computing resources were provided by Compute Canada through CalculQuébec. Marlos C. Machado performed part of this work while at the University of Alberta.

References

Appendix

We provide a brief description of each game flavour used in the paper.

In Freeway a chicken must cross a road containing multiple lanes of moving traffic within a prespecified time limit. In all modes of Freeway the agent is rewarded for reaching the top of the screen and is subsequently teleported to the bottom of the screen. If the chicken collides with a vehicle in difficulty 0 it gets bumped down one lane of traffic, alternatively, in difficulty 1 the chicken gets teleported to its starting position at the bottom of the screen. Mode 1 changes some vehicle sprites to include buses, adds more vehicles to some lanes, and increases the velocity of all vehicles. Mode 4 is almost identical to Mode 1; the only difference being vehicles can oscillate between two speeds.

Hero

In Hero you control a character who must navigate a maze in order to save a trapped miner within a cave system. The agent scores points for any forward progression such as clearing an obstacle or killing an enemy. Once the miner is rescued, the level is terminated and you continue to the next level with a different maze. Some levels have partially observable rooms, more enemies, and more difficult obstacles to traverse. Past the default mode, each subsequent mode starts off at increasingly harder levels denoted by a level number increasing by multiples of 55. The default mode starts you off at level 11, mode 1 starts at level 55, and so on.

Breakout

In Breakout you control a paddle which can move horizontally along the bottom of the screen. At the beginning of the game, or on a loss of life the ball is set into motion and can bounce off the paddle and collide with bricks at the top of the screen. The objective of the game is to break all the bricks without having the ball fall below your paddles horizontal plane. Subsequently, mode 12 of Breakout hides the bricks from the player until the ball collides with the bricks in which case the bricks flash for a brief moment before disappearing again.

Space Invaders

When playing Space Invaders you control a spaceship which can move horizontally along the bottom of the screen. There is a grid of aliens above you and the objective of the game is to eliminate all the aliens. You are afforded some protection from the alien bullets with three barriers just above your spaceship. Difficulty 1 of Space Invaders widens your spaceships sprite making it harder to dodge enemy bullets. Mode 1 of Space Invaders causes the shields above you to oscillate horizontally. Mode 9 of Space Invaders is similar to Mode 12 of Breakout where the aliens are partially observable until struck with the player’s bullet.

Experimental Details

All experiments performed in this paper utilized the neural network architecture proposed by Mnih et al. (2015). That is, a convolutional neural network with three convolutional layers and two fully connected layers. A visualization of this network can be found in Figure 5. Unless otherwise specified, hyperparametes are kept consistent with the ALE baselines discussed by Machado et al. (2018). A summary of the parameters, which were consistent across all experiments, can be found in in Table Architecture and hyperparameters.

Evaluation

We adhere to the evaluation methodologies set out by Machado et al. (2018). This includes the use of all 18 primitive actions in the ALE, not utilizing loss of life as episode termination, and the use of sticky actions to inject stochasticity. Each result outlined in this paper averages the agents performance over 100100 episodes further averaged over five runs. We do not take the maximum over runs nor the maximum over the learning curve.

When comparing results in this paper and with other evaluation methodologies it is worth noting the following terminology and time scales. We use a frame skip of 55 frames, i.e., following every action executed by the agent the simulator advances 55 frames into the future. The agent will take \nicefrac# frames5\nicefrac{{\text{\# frames}}}{{5}} actions within the environment over the duration of each experiment. One step of stochastic gradient descent to update the network parameters is performed every 44 actions. The training routine will perform \nicefrac# frames5⋅4\nicefrac{{\text{\# frames}}}{{5\cdot 4}} gradient updates over the duration of each experiment. Therefore, when we discuss experiments with a duration of 50M frames this is in actuality 50M simulator frames, 10M agent steps, and 2.5M gradient updates.

Regularization Ablation Study

We begin by analyzing the training performance for DQN in Freeway m0d0 for different values of λ\lambda. We also provide evaluation curves for m1d0, and m4d0 of Freeway. Both sets of experiments are presented in Figure 6.

Dropout

We now examine the evaluation performance for each parameter configuration in both Freeway m1d0, and Freeway m4d0. These results are presented in Figure 9 for m1d0, and Figure 10 for m4d0.

Policy Evaluation Learning Curves