Efficient Exploration through Bayesian Deep Q-Networks
Kamyar Azizzadenesheli, Animashree Anandkumar
Introduction
One of the central challenges in reinforcement learning (RL) is to design algorithms with efficient exploration-exploitation trade-off that scale to high-dimensional state and action spaces. Recently, deep RL has shown significant promise in tackling high-dimensional (and continuous) environments. These successes are mainly demonstrated in simulated domains where exploration is inexpensive and simple exploration-exploitation approaches such as -greedy or Boltzmann strategies are deployed. -greedy chooses the greedy action with probability and randomizes over all the actions, and does not consider the estimated Q-values or its uncertainties. The Boltzmann strategy considers the estimated Q-values to guide the decision making but still does not exploit their uncertainties in estimation. For complex environments, more statistically efficient strategies are required. One such strategy is optimism in the face of uncertainty (OFU), where we follow the decision suggested by the optimistic estimation of the environment and guarantee efficient exploration/exploitation strategies. Despite compelling theoretical results, these methods are mainly model based and limited to tabular settings .
An alternative to OFU is posterior sampling (PS), or more general randomized approach, is Thompson Sampling which, under the Bayesian framework, maintains a posterior distribution over the environment model, see Table 1. Thompson sampling has shown strong performance in many low dimensional settings such as multi-arm bandits and small tabular MDPs . Thompson sampling requires sequentially sampling of the models from the (approximate) posterior or uncertainty and to act according to the sampled models to trade-off exploration and exploitation. However, the computational costs in posterior computation and planning become intractable as the problem dimension grows.
To mitigate the computational bottleneck, Osband et al., consider episodic and tabular MDPs where the optimal Q-function is linear in the state-action representation. They deploy Bayesian linear regression (BLR) to construct an approximated posterior distributing over the Q-function and employ Thompson sampling for exploration/exploitation. The authors guarantee an order optimal regret upper bound on the tabular MDPs in the presence of a Dirichlet prior on the model parameters. Our paper is a high dimensional and general extension of .
While the study of RL in general MDPs is challenging, recent advances in the understanding of linear bandits, as an episodic MDPs with episode length of one, allows tackling high dimensional environment. This class of RL problems is known as LinReL. In linear bandits, both OFU and Thompson sampling guarantee promising results for high dimensional problems. In this paper, we extend LinReL to MDPs.
Contribution 1 – Bayesian and frequentist regret analysis: We study RL in episodic MDPs where the optimal Q-function is a linear function of a -dimensional feature representation of state-action pairs. We propose two algorithms, LinPSRL, a Bayesian method using PS, and LinUCB, a frequentist method using OFU. LinPSRL constructs a posterior distribution over the linear parameters of the Q-function. At the beginning of each episode, LinPSRL draws a sample from the posterior then acts optimally according to that model. LinUCB constructs the upper confidence bound on the linear parameters and in each episode acts optimally with respect to the optimistic model. We provide theoretical performance guarantees and show that after episodes, the Bayesian regret of LinPSRL and the frequentist regret of LinUCB are both upper bounded by The dependency in the episode length is more involved and details are in Section 2..
Contribution 2 – From theory to practice: While both LinUCB and LinPSRL are statistically designed for high dimensional RL, their computational complexity can make them practically infeasible, e.g., maintaining the posterior can become intractable. To mitigate this shortcoming, we propose a unified method based on the BLR approximation of these two methods. This unification is inspired by the analyses in Abeille and Lazaric, , Abbasi-Yadkori et al., for linear bandits. 1) For LinPSRL: we deploy BLR to approximate the posterior distribution over the Q-function using conjugate Gaussian prior and likelihood. In tabular MDP, this approach turns out to be similar to Osband et al., We refer the readers to this work for an empirical study of BLR on tabular environment.. 2) For LinUCB: we deploy BLR to fit a Gaussian distribution to the frequentist upper confidence bound constructed in OFU (Fig 2 in Abeille and Lazaric, ). These two approximation procedures result in the same Gaussian distribution, and therefore, the same algorithm. Finally, we deploy Thompson sampling on this approximated distribution over the Q-functions. While it is clear that this approach is an approximation to PS, Abeille and Lazaric, show that this approach is also an approximation to OFU An extra expansion of the Gaussian approximation is required for the theoretical analysis. . For practical use, we extend this unified algorithm to deep RL, as described below.
Contribution 3 – Design of BDQN: We introduce Bayesian Deep Q-Network (BDQN), a Thompson sampling based deep RL algorithm, as an extension of our theoretical development to deep neural networks. We follow the DDQN architecture and train the Q-network in the same way except for the last layer (the linear model) where we use BLR instead of linear regression. We deploy Thompson sampling on the approximated posterior of the Q-function to balance between exploration and exploitation. Thus, BDQN requires a simple and a minimal modification to the standard DDQN implementation.
We empirically study the behavior of BDQN on a wide range of Atari games . Since BDQN follows an efficient exploration-exploitation strategy, it reaches much higher cumulative rewards in fewer interactions, compared to its -greedy predecessor DDQN. We empirically observed that BDQN achieves DDQN performance in less than 5M1M interactions for almost half of the games while the cumulative reward improves by a median of with a maximum of on all games. Also, BDQN has (mean and standard deviation) improvement over these games on the area under the performance measure. Thus, BDQN achieves better sample complexity due to a better exploration/exploitation trade-off.
Comparison: Recently, many works have studied efficient exploration/exploitation in high dimensional environments. proposes a variational inference-based approach to help the exploration. Bellemare et al., proposes a surrogate for optimism. Osband et al., proposes an ensemble of many DQN models. These approaches are significantly more expensive than DDQN while also require a massive hyperparameter tuning effort .
In contrast, our approach has the following desirable properties: 1) Computation: BDQN nearly has a same computational complexity as DDQN since there is no backpropagation in the last layer of BDQN (therefore faster), but instead, there is a BLR update which requires inverting a small matrix (order of less than a second), once in a while. 2) Hyperparameters: no exhaustive hyper-parameter tuning. We spent less than two days of academic level GPU time on hyperparameter tuning of BLR in BDQN which is another evidence on its significance. 3) Reproducibility: All the codes, with detailed comments and explanations, are publicly available.
Linear Q-function
Consider an episodic MDP , with horizon length , state space , closed action set , transition kernel , initial state distribution , reward distribution , discount factor . For any natural number , . The time step within the episode, , is encoded in the state, i.e., and . We drop in state-action definition for brevity. denotes the spectral norm and for any positive definite matrix , denotes the matrix-weighted spectral norm. At state of time step , we define Q-function as agent’s expected return after taking action and following policy , a mapping from state to action.
Following the Bellman optimality in MDPs, we have that for the optimal Q-function
2 LinReL
At each time step and state action pair , define a mean zero random variable that captures stochastic reward and transition at time step :
where is the reward at time step . Definition of plays an important role since knowing and reduces the learning of to the standard Martingale based linear regression problem. Of course we neither have nor have .
LinPSRL(Algorithm 1): In this Bayesian approach, the agent maintains the prior over the vectors and given the collected experiences, updates their posterior at the beginning of each episode. At the beginning of each episode , the agent draws from the posterior, and follows their induced policy , i.e., .
LinUCB(Algorithm 2): In this frequentist approach, at the beginning of ’th episode, the agent exploits the so-far collected experiences and estimates up to a high probability confidence intervals i.e., . At each time step , given a state , the agent follows the optimistic policy: .
Following the standard assumption in the self normalized analysis of linear regression and linear bandit , we have:
: the noise vector is a -sub-Gaussian vector. (refer to Assumption 1)
we have , , a.s.
Then, define such that:
similar to the ridge linear regression analysis in Hsu et al., , we require . This requirement is automatically satisfied if the optimal Q-function is bounded away from zero (all features have large component at least in one direction). Let denote the following combination of :
For any prior and likelihood satisfying these assumptions, we have:
Proof is given in the Appendix E.3. These regret upper bounds are similar to those in linear bandits and linear quadratic control , i.e. . Since linear bandits are special cases of episodic continuous MDPs, when horizon is equal to , we observe that our Bayesian regret upper bound recovers and our frequentist regret upper bound recovers the bound in . While our regret upper bounds are order optimal in , and , they have bad dependency in the horizon length . In our future work, we plan to extensively study this problem and provide tight lower and upper bound in terms of and .
Bayesian Deep Q-Networks
We propose Bayesian deep Q-networks (BDQN) an efficient Thompson sampling based method in high dimensional RL problems. In value based RL, the core of most prominent approaches is to learn the Q-function through minimizing a surrogate to Bellman residual using temporal difference (TD) update . Van Hasselt et al., carries this idea, and propose DDQN (similar to its predecessor DQN ) where the Q-function is parameterized by a deep network. DDQN employ a target network , target value , where the tuple are consecutive experiences, . DDQN learns the Q function by approaching the empirical estimates of the following regression problem:
The DDQN agent, once in a while, updates the network by setting it to the Q network, and follows the regression in Eq.1 with the new target value. Since we aim to empirically study the effect of Thompson sampling, we directly mimic the DDQN to design BDQN.
In DDQN, we match to using the regression in Eq. 1. This regression problem results in a linear regression in the last layer, ’s. BDQN follows all DDQN steps except for the learning of the last layer ’s. BDQN deploys Gaussian BLR instead of the plain linear regression, resulting in an approximated posterior on the ’s and consequently on the Q-function. As discussed before, BLR with Gaussian prior and likelihood is an approximation to LinPSRL and LinUCB . Through BLR, we efficiently approximate the distribution over the Q-values, capture the uncertainty over the Q estimates, and design an efficient exploration-exploitation strategy using Thompson Sampling.
which is the derivation of well-known BLR, with mean zero prior as well as and , variance of prior and likelihood respectively. Fig. 1 demonstrate the mean and covariance of the over for each action . A BDQN agent deploys Thompson sampling on the approximated posteriors every to balance exploration and exploitation while updating the posterior every .
Experiments
We empirically study BDQN behaviour on a variety of Atari games in the Arcade Learning Environment using OpenAI Gym , on the measures of sample complexity and score against DDQN, Fig 2, Table 2. All the codeshttps://github.com/kazizzad/BDQN-MxNet-Gluon, Double-DQN-MxNet-Gluon, DQN-MxNet-Gluon, with detailed comments and explanations are publicly available and programmed in MxNet .
We implemented DDQN and BDQN following Van Hasselt et al., . We also attempted to implement a few other deep RL methods that employ strategic exploration (with advice from their authors), e.g., . Unfortunately we encountered several implementation challenges that we could not address since neither the codes nor the implementation details are publicly available (we were not able to reproduce their results beyond the performance of random policy). Along with BDQN and DDQN codes, we also made our implementation of Osband et al., publicly available. In order to illustrate the BDQN performance we report its scores along with a number of state-of-the-art deep RL methods 2. For some games, e.g., Pong, we ran the experiment for a longer period but just plotted the beginning of it in order to observe the difference. Due to huge cost of deep RL methods, for some games, we run the experiment until a plateau is reached. The BDQN and DDQN columns are scores after running them for number steps reported in Step column without Since the regret is considered, no evaluation phase designed for them. is the reported scores of DDQN in Van Hasselt et al., at evaluation time where the . We also report scores of Bootstrap DQN , NoisyNet , CTS, Pixel, Reactor . For NoisyNet, the scores of NoisyDQN are reported. To illustrate the sample complexity behavior of BDQN we report SC: the number of interactions BDQN requires to beat the human score ( means BDQN could not beat human score), and : the number of interactions the BDQN requires to beat the score of . Note that Table 2 does not aim to compare different methods. Additionally, there are many additional details that are not included in the mentioned papers which can significantly change the algorithms behaviors ), e.g., the reported scores of DDQN in Osband et al., are significantly higher than the reported scores in the original DDQN paper, indicating many existing non-addressed advancements (Appendix A.2).
We also implemented DDQN drop-out a Thomson Sampling based algorithm motivated by Gal and Ghahramani, . We observed that it is not capable of capturing the statistical uncertainty in the function and falls short in outperforming a random(uniform) policy. Osband et al., investigates the sufficiency of the estimated uncertainty and hardness in driving suitable exploitation out of it. It has been observed that drop-out results in the ensemble of infinitely many models but all models almost the same Appendix A.1.
As mentioned before, due to an efficient exploration-exploitation strategy, not only BDQN improves the regret and enhance the sample complexity, but also reaches significantly higher scores. In contrast to naive exploration, BDQN assigns less priority to explore actions that are already observed to be not worthy, resulting in better sample complexity. Moreover, since BDQN does not commit to adverse actions, it does not waste the model capacity to estimate the value of unnecessary actions in unnecessary states as good as the important ones, resulting in saving the model capacity and better policies.
For the game Atlantis, reaches score of during the evaluation phase, while BDQN reaches score of after time steps. After multiple run of BDQN, we constantly observed that its performance suddenly improves to around in the vicinity of time steps. We closely investigate this behaviour and realized that BDQN saturates the Atlantis game and reaches the internal OpenAIGym limit of . After removing this limit, BDQN reaches score after . Please refer to Appendix A for the extensive empirical study.
Related Work
The complexity of the exploration-exploitation trade-off has been deeply investigated in RL literature for both continuous and discrete MDPs . Jaksch et al., investigate the regret analysis of MDPs with finite state and action and deploy OFU to guarantee a regret upper bound, while Ortner and Ryabko, relaxes it to a continuous state space and propose a sub-linear regret bound. Azizzadenesheli et al., 2016b deploys OFU and propose a regret upper bound for Partially Observable MDPs (POMDPs) using spectral methods . Furthermore, Bartók et al., tackles a general case of partial monitoring games and provides minimax regret guarantee. For linear quadratic models OFU is deployed to provide an optimal regret bound . In multi-arm bandit, Thompson sampling have been studied both from empirical and theoretical point of views . A natural adaptation of this algorithm to RL, posterior sampling RL (PSRL) Strens, also shown to have good frequentist and Bayesian performance guarantees . Inevitably for PSRL, these methods also have hard time to become scalable to high dimensional problems, .
Exploration-exploitation trade-offs has been theoretically studied in RL but a prominent problem in high dimensional environments . Recent success of Deep RL on Atari games , the board game Go , robotics , self-driving cars , and safety in RL propose promises on deploying deep RL in high dimensional problem.
To extend the exploration-exploitation efficient methods to high dimensional RL problems, Osband et al., suggest bootstrapped-ensemble approach that trains several models in parallel to approximate the posterior distribution. Bellemare et al., propose a way to come up with a surrogate to optimism in high dimensional RL. Other works suggest using a variational approximation to the Q-networks or a concurrent work on noisy network suggest to randomize the Q-network. However, most of these approaches significantly increase the computational cost of DQN, e.g., the bootstrapped-ensemble incurs a computation overhead that is linear in the number of bootstrap models.
Concurrently, Levine et al., proposes least-squares temporal difference which learns a linear model on the feature representation in order to estimate the Q-function. They use -greedy approach and provide results on five Atari games. Out of these five games, one is common with our set of 15 games which BDQN outperforms it by a factor of (w.r.t. the score reported in their paper). As also suggested by our theoretical derivation, our empirical study illustrates that performing Bayesian regression instead, and sampling from the result yields a substantial benefit. This indicates that it is not just the higher data efficiency at the last layer, but that leveraging an explicit uncertainty representation over the value function is of substantial benefit.
Conclusion
In this work, we proposed LinPSRL and LinUCB, two LinReL algorithms for continuous MDPs. We then proposed BDQN, a deep RL extension of these methods to high dimensional environments. BDQN deploys Thompson sampling and provides an efficient exploration/exploitation in a computationally efficient manner. It involved making simple modifications to the DDQN architecture by replacing the linear regression learning of the last layer with Bayesian linear regression. We demonstrated significantly improvement training, convergence, and regret along with much better performance in many games.
While our current regret upper bounds seem to be sub-optimal in terms of (we are not aware of any tight lower bound), in the future, we plan to deploy the analysis in and develop a tighter regret upper bounds as well as an information theoretic lower bound. We also plan to extend the analysis in Abeille and Lazaric, and develop Thompson sampling methods with a performance guarantee and finally go beyond the linear models . While finding optimal continuous action given a Q function can be computationally intractable, we aim to study the relaxation of these approaches in continuous control tasks in the future.
Acknowledgments
The authors would like to thank Emma Brunskill, Zachary C. Lipton, Marlos C. Machado, Ian Osband, Gergely Neu, Kristy Choi, particularly Akshay Krishnamurthy, and Nan Jiang during ICLR2019 openreview, for their feedbacks, suggestions, and helps. K. Azizzadenesheli is supported in part by NSF Career Award CCF-1254106 and AFOSR YIP FA9550-15-1-0221. This research has been conducted when the first author was a visiting researcher at Stanford University and Caltech. A. Anandkumar is supported in part by Bren endowed chair, Darpa PAI, and Microsoft, Google, Adobe faculty fellowships, NSF Career Award CCF-1254106, and AFOSR YIP FA9550-15-1-0221. All the experimental study have been done using Caltech AWS credits grant.
References
Appendix A Empirical Study
We empirically study the behavior of BDQN along with DDQN as reported in Fig. 2 and Table 2. We run each of these algorithms for the number steps mentioned in the Step column of Table 2. We observe that BDQN significantly improves the sample complexity over DDQN and reaches the highest scores of DDQN in a much fewer number of interactions required by DDQN. Due to BDQN’s better exploration-exploitation strategy, we expected BDQN to improve the convergence and enhance the sample complexity, but we further observed a significant improvement in the scores as well. It is worth noting that since the purpose of this study is sample complexity and regret analysis, design of evaluation phase (as it is common for -greedy methods) is not relevant. In our experiments, e.g., in the game Pong, DDQN reaches the score of 18.82 during the learning phase. When we set to a quantity close to zero, DDQN reaches the score of while BDQN just converges to 21 as uncertainty decays.
In addition to the Table 2, we also provided the score ratio as well as the area under the performance plot ratio comparisons in Table 3.
We observe that BDQN learns significantly better policies due to its efficient explore/exploit in a much shorter period of time. Since BDQN on game Atlantis promise a big jump around time step , we ran it five more times in order to make sure it was not just a coincidence Fig. 4. For the game Pong, we ran the experiment for a longer period but just plotted the beginning of it in order to observe the difference. Due to cost of deep RL methods, for some games, we run the experiment until a plateau is reached.
For the game Atlantis, reaches the score of during the evaluation phase, while BDQN reaches score of after interactions. As it is been shown in Fig. 2, BDQN saturates for Atlantis after interactions. We realized that BDQN reaches the internal OpenAIGym limit of , where relaxing it improves score after steps to .
After removing the maximum episode length limit for the game Atlantis, BDQN gets the score of 62M. This episode is long enough to fill half of the replay buffer and make the model perfect for the later part of the game but losing the crafted skill for the beginning of the game. We observe in Fig. 3 that after losing the game in a long episode, the agent forgets a bit of its skill and loses few games but wraps up immediately and gets to score of . To overcome this issue, one can expand the replay buffer size, stochastically store samples in the reply buffer where the later samples get stored with lowest chance, or train new models for the later parts of the episode. There are many possible cures for this interesting observation and while we are comparing against DDQN, we do not want to advance BDQN structure-wise.
We would like to acknowledge that we provide this study based on request by our reviewer during the last submission.
Dropout, as another randomized exploration method, is proposed by Gal and Ghahramani, , but Osband et al., argue about the deficiency of the estimated uncertainty and hardness in driving a suitable exploration and exploitation trade-off from it (Appendix A in ). They argue that Gal and Ghahramani, does not address the fundamental issue that for large networks trained to convergence all dropout samples may converge to every single datapoint. As also observed by , dropout might results in a ensemble of many models, but all almost the same (converge to the very same model behavior). We also implemented the dropout version of DDQN, Dropout-DDQN, and ran it on four randomly chosen Atari games (among those we ran for less than time steps). We observed that the randomization in Dropout-DDQN is deficient and results in performances worse than DDQN on these four Atari games, Fig. 5. In Table 4 we compare the performance of BDQN, DDQN, , and Dropout-DDQN, as well as the performance of the random policy, borrowed from Mnih et al., . We observe that the Dropout-DDQN not only does not outperform the plain -greedy DDQN, it also sometimes underperforms the random policy. For the game Pong, we also ran Dropout-DDQN for time steps but its average performance did not get any better than -17. For the experimental study we used the default dropout rate of to mitigate its collapsing issue.
A.2 Further discussion on Reproducibility
In Table 2, we provide the scores of bootstrap DQN and NoisyNet This work does not have scores of Noisy-net with DDQN objective function but it has Noisy-net with DQN objective which are the scores reported in Table 2, and along with BDQN. These score are directly copied from their original papers and we did not make any change to them. We also report the scores of count-based method despite its challenges. Ostrovski et al., does not provide the scores table.
Table 2 shows, despite the simplicity of BDQN, it provides a significant improvement over these baselines. It is important to note that we can not scientifically claim that BDQN outperforms these baselines by just looking at the scores in Table2 since we are not aware of their detailed implementation as well as environment details. For example, in this work, we directly implemented DDQN by following the implementation details mentioned in the original DDQN paper and the scores of our DDQN implementation during the evaluation time almost matches the scores of DDQN reported in the original paper. But the reported scores of implemented DDQN in Osband et al., are significantly different from the reported score in the original DDQN paper.
Appendix B Why Thompson Sampling and not ε𝜀\varepsilon-greedy or Boltzmann exploration
In value approximation RL algorithms, there are different ways to manage the exploration-exploitation trade-off. DQN uses a naive -greedy for exploration, where with probability it chooses a random action and with probability it chooses the greedy action based on the estimated function. Note that there are only point estimates of the function in DQN. In contrast, our proposed Bayesian approach BDQN maintains uncertainties over the estimated function, and employs it to carry out Thompson Sampling based exploration-exploitation. Here, we demonstrate the fundamental benefits of Thompson Sampling over -greedy and Boltzmann exploration strategies using simplified examples. In Table 1, we list the three strategies and their properties.
-greedy is among the simplest exploration-exploitation strategies and it is uniformly random over all the non-greedy actions. Boltzmann exploration is an intermediate strategy since it uses the estimated function to sample from action space. However, it does not maintain uncertainties over the function estimation. In contrast, Thompson sampling incorporates the estimate as well as the uncertainties in the estimation and utilizes the most information for exploration-exploitation strategy.
Consider the example in Figure 6(a) with our current estimates and uncertainties of the function over different actions. -greedy is not compelling since it assigns uniform probability to explore over and , which are sub-optimal when the uncertainty estimates are available. In this setting, a possible remedy is Boltzmann exploration since it assigns lower probability to actions and but randomizes with almost the same probabilities over the remaining actions.
However, Boltzmann exploration is sub-optimal in settings where there is high uncertainty. For example if the current estimate is according to Figure 6(b), then Boltzmann exploration assigns almost equal probability to actions and , even though action has much higher uncertainty and needs to be explored more.
Thus, both -greedy and Boltzmann exploration strategies are sub-optimal since they do not maintain an uncertainty estimate over the estimation. In contrast, Thompson sampling uses both estimated function and its uncertainty estimates to carry out a more efficient exploration.
Appendix C BDQN Implementation
In this section, we provide the details of BDQN algorithm, mentioned in Algorithm 3.
As mentioned, BDQN agent deploys Thompson sampling on approximated posteriors over the weights of the last layer ’s. At every time step, BDQN samples a new model to balance exploration and exploitation while updating the posterior every time steps using experience in the replay buffer, chosen uniformly at random.
BDQN agent samples a new Q function, (i.e., ), every time steps and act accordingly . As suggested by derivations in Section 2, is chosen to be in order of episode length for the Atari games. As mentioned before, the feature network is trained similar to DDQN using .
Similar to DDQN, we update the target network every steps and set to . We update the posterior distribution using a minibatch of randomly chosen experiences in the replay buffer every , and set the , the mean of the posterior distribution Alg. 3.
The input to the network part of BDQN is tensor with a rescaled and averaged over channels of the last four observations. The first convolution layer has filters of size with a stride of . The second convolution layer has filters of size with stride . The last convolution layer has filters of size followed by a fully connected layer with size . We add a BLR layer on top of this.
C.2 Choice of hyper-parameters:
For BDQN, we set the values of to the mean of the posterior distribution over the weights of BLR with covariances and draw ’s from this posterior. For the fixed set of ’s and , we randomly initialize the parameters of network part of BDQN, , and train it using RMSProp, with learning rate of , and a momentum of , inspired by where the discount factor is , the number of steps between target updates steps, and weights are re-sampled from their posterior distribution every steps. We update the network part of BDQN every steps by uniformly at random sampling a mini-batch of size samples from the replay buffer. We update the posterior distribution of the weight set every using mini-batch of size , with entries sampled uniformly form replay buffer. The experience replay contains the most recent transitions. Further hyper-parameters are equivalent to ones in DQN setting.
For the BLR, we have noise variance , variance of prior over weights , sample size , posterior update period , and the posterior sampling period . To optimize for this set of hyper-parameters we set up a very simple, fast, and cheap hyper-parameter tuning procedure which proves the robustness of BDQN. To find the first three, we set up a simple hyper-parameter search. We used a pretrained DQN model for the game of Assault, and removed the last fully connected layer in order to have access to its already trained feature representation. Then we tried combination of , , and and test for episode of the game. We set these parameters to their best .
The above hyper-parameter tuning is cheap and fast since it requires only a few times the number of forwarding passes. Moreover, we did not try all the combination, and rather each separately. We observed that perform similarly with slightly better. But both perform better than . For . For the remaining parameters, we ran BDQN ( with weights randomly initialized) on the same game, Assault, for time steps, with a set of and , where BDQN performed better with choice of . For both choices of , it performs almost equal and we choose the higher one to reduce the computation cost. We started off with the learning rate of and did not tune for that. Thanks to the efficient Thompson sampling exploration and closed form BLR, BDQN can learn a better policy in an even shorter period of time. In contrast, it is well known for DQN based methods that changing the learning rate causes a major degradation in the performance (Fig. 7). The proposed hyper-parameter search is very simple and an exhaustive hyper-parameter search is likely to provide even better performance.
C.3 Learning rate:
It is well known that DQN and DDQN are sensitive to the learning rate and change of learning rate can degrade the performance to even worse than random policy. We tried the same learning rate as BDQN, 0.0025, for DDQN and observed that its performance drops. Fig. 7 shows that the DDQN with higher learning rates learns as good as BDQN at the very beginning but it can not maintain the rate of improvement and degrade even worse than the original DDQN with learning rate of 0.00025.
C.4 Computational and sample cost comparison:
For a given period of game time, the number of the backward pass in both BDQN and DQN are the same where for BDQN it is cheaper since it has one layer (the last layer) less than DQN. In the sense of fairness in sample usage, for example in duration of , all the layers of both BDQN and DQN, except the last layer, sees the same number of samples, but the last layer of BDQN sees times fewer samples compared to the last layer of DQN. The last layer of DQN for a duration of , observes (4 is back prob period) mini batches of size , which is , where the last layer of BDQN just observes samples size of . As it is mentioned in Alg. 3, to update the posterior distribution, BDQN draws samples from the replay buffer and needs to compute the feature vector of them. Therefore, during the interactions for the learning procedure, DDQN does of forward passes and of backward passes, while BDQN does same number of backward passes (cheaper since there is no backward pass for the final layer) and of forward passes. One can easily relax it by parallelizing this step along the main body of BDQN or deploying on-line posterior update methods.
C.5 Thompson sampling frequency:
The choice of Thompson sampling update frequency can be crucial from domain to domain. Theoretically, we show that for episodic learning, the choice of sampling at the beginning of each episode, or a bounded number of episodes is desired. If one chooses too short, then computed gradient for backpropagation of the feature representation is not going to be useful since the gradient get noisier and the loss function is changing too frequently. On the other hand, the network tries to find a feature representation which is suitable for a wide range of different weights of the last layer, results in improper waste of model capacity. If the Thompson sampling update frequency is too low, then it is far from being Thompson sampling and losses the randomized exploration property. We are interested in a choice of which is in the order of upper bound on the average length of each episode of the Atari games. The current choice of is suitable for a variety of Atari games since the length of each episode is in range of and is infrequent enough to make the feature representation robust to big changes.
Appendix D A short discussion on safety
In BDQN, as mentioned in Eq. 2, the prior and likelihood are conjugate of each others. Therefore, we have a closed form posterior distribution of the discounted return, , approximated as
One can use this distribution and come up with a safe RL criterion for the agent . Consider the following example; for two actions with the same mean, if the estimated variance over the return increases, then the action becomes more unsafe Fig. 8. By just looking at the low and high probability events of returns under different actions we can approximate whether an action is safe to take.
Appendix E Bayesian and frequentist regrets, Proof of Theorems 1 and 2
where and denote the state and action random variables. Since we assume the environment is an MDP, following the Bellman optimality we have for
where the optimal policy denotes a deterministic mapping from states to actions. Following the Bellman optimality, conditioned on , one can rewrite the reward at time step , , as follows;
Where is the noise in the reward, and it is a mean zero random variable due to bellman optimality. In other word, condition on , the distribution of one step reward can be described as:
where the equality is in distribution. We also could extend the randomness in the noise of the return one step further and include the randomness in the transition. It means, condition on and following at a time step , for the distribution of one step return we have:
where here the encodes the randomness in the reward as well as the transition kernel, and it is a mean zero random variable. The equality is in distribution. If instead of following after time step , we follow policies other than , e.g., then condition on and following at a time step , and afterward, for the distribution of one step return we have:
where then noise process is not mean zero anymore (it is biased), except for final time step . We can quantify this bias as follows:
where is now an unbiased zero mean random variable due to equality of two reward distributions in Eq. 3 and Eq. 4.
We use this intuition and definition of noises in order to later turn the learning of model parameters to unbiased Martingale linear regression.
E.2 Linear Q-function
In the following, we consider the case when the optimal Q function is a linear transformation of a feature representation vector for any pair of state and actions at any time step . In other words, at any time step we have
To keep the notation simple, while all the policies, Q functions, and feature presentations are functions of time step , we encode the into the state .
As mentioned in the section 2, we consider the feature represent and the weight vectors of satisfy the following conditions;
In the following, we denote the agent policy at ’th time step of ’th episode as . We show how to estimate using data collected under . We denote as the estimation of for any at an episode . We show over time for any our estimations of’s concentrate around their true parameters . We further show how to deploy this concentrations and construct two algorithms, one based on PSRL, LinPSRL, and another based on Optimism OFU , LinUCB, to guaranteed Bayesian and frequentist regret upper bounds respectively. The main body of the following analyses on concentration of measure are based on the contextual linear bandit literature and self normalized processes .
For the following we consider the case where and then show how to extend the results to any discounted case.
We define the modified target value, condition on as follows:
where we restate the mean zero noise as:
At time step of episode, conditioned on the event up to time step of the episode is a sub-Gaussian random variable, i.e. there exists a parameter such that
where is -measurable, and is -measurable.
A similar assumption on the noise model is considered in the prior works in linear bandits .
Expected Reward and Return:
The expected reward and expected return are in $$.
Regret Definition
Let denote the value of policy under a model parameter . The regret definition (pseudo regret) for the frequentist regret is as follows;
Where is the agent policy during episode .
When there is a prior over the the expected Bayesian regret might be the target of the study. The Bayesian regret is as follows;
E.3 Optimism: Regret bound of LinUCB, Algorithm. 2
In optimism we approximate the desired model parameters up to their high probability confidence set where with probability at least . In optimism we choose the most optimistic models from these plausible sets. Let denote the chosen most optimistic parameters during an episode . The most optimistic models for each is;
and also the most pessimistic models ;
Since in RL, we do not have access to the true model parameters, we do not observe the modified target value. Therefore for eath and we define as the observable target value:
where is the policy corresponding to . For we have .
The modified target values and the target value have the following relationship
These definitions of , , and and their relationship allows us to efficiently construct a self normalized process to estimate the true model parameters.
where is a ridge regularization matrix and is usually set to .
For the optimal actions and time step , denotes the spectral bound on;
Let denote the following combination of ;
It is worth noting that a lose upper bound on when is .
For all , let denote the estimation of given and ;
then, with probability at least
for all . For , we set and have
Let denote the event that the confidence bounds in Lemma 1 holds at least until episode.
Lemma 1 states that under event , . Furthermore, we define state and policy dependent optimistic parameter as follows;
Following OFU, Algorithm 2, we set , the optimistic policies. By the definition we have
We use this inequality to derive an upper bound for the regret;
Let us defined as the value function at , following policies in for the first time steps then switching to the optimal policies.
Given the linear model of the function we have;
For we deploy the similar decomposition and upper bound as
Since for both of and we follow the same policy on the same model for 1 time step, the reward at the first time step has the same distribution, therefore we have;
Similarly we can defined . Therefore;
Since the maximum expected cumulative reward, condition on states of a episode is at most , we have;
where we exploit the fact that is an increasing function of and consider the harder case where .
Moreover, at time , we can use Jensen’s inequality
Now, using the fact that for any scalar such that , then we can rewrite the latter part of Eq. E.3
By applying the Lemma 2 and substituting the RHS of Eq. 5 into Eq. E.3, we get
with probability at least . If we set then the probability that the event holds is and we get regret of at most the RHS of Eq. 9, otherwise with probability at most we get maximum regret of , therefore
For the case of discounted reward, substituting instead of results in the theorem statement.
E.4 Bayesian Regret of Alg. 1
The analysis developed in the previous section, up to some minor modification, e.g., change of strategy to PSRL, directly applies to Bayesian regret bound, with a farther expectation over models.
When there is a prior over the the expected Bayesian regret might be the target of the study.
Here is a multivariate random sequence which indicates history at the beginning of episode and are the policies following PSRL. For the remaining denotes the PSRL policy at each time step . As it mentioned in the Alg. 1, at the beginning of an episode, we draw from the posterior and the corresponding policies are;
Condition on the history , i.e., the experiences by following the agent policies for each episode , we estimate the as follows;
Lemma 1 states that under event , . Conditioned on , the and are equally distributed, then we have
Similar to optimism and defining we have;
For we deploy the similar decomposition and upper bound as
Since for both of and we follow the same distribution over policies and models for 1 time step, the reward at the first time step has the same distribution, therefore we have;
Similarly we can defined . The condition on was required to come up with the mentioned decomposition through and it is not needed anymore, therefore;
Again similar to optimism we have the maximum expected cumulative reward condition on states of a episode is at most , we have;
Moreover, at time , we can use Jensen’s inequality, exploit the fact that is an increasing function of and have
Under event which holds with probability at least . If we set then the probability that the event holds is and we get regret of at most the RHS of Eq. E.4, otherwise with probability at most we get maximum regret of , therefore
For the case of discounted reward, substituting instead of results in the theorem statement.
E.5 Proof of Lemmas
Let be any vector and for any define
Then, for a stopping time under the filtration , .
We first show that is a supermartingale sequence. Let
[Extension to Self-normalized bound in ] For a stopping time and filtration , with probability at least
of Lemma 4. Given the definition of the parameters of self-normalized process, we can rewrite as follows;
Since and are positive semi definite and positive definite respectively, we have
Where the Markov inequality is deployed for the final step. The stopping is considered to be the time step as the first time in the sequence when the concentration in the Lemma 4 does not hold. ∎
Consider the following estimator :
As a results, applying Cauchy-Schwarz inequality and inequalities we get
where applying self normalized Lemma 4, for all with probability at least
hold for any . By plugging in we get the following;
For the last term in the above equation is zero therefore we have;
For we need to account for the bias introduced by . Due to the pesimism we have
also due to the optimal action of the pesimistic model we have
The can be written as therefore, . We also know that;
A matrix extension to Lemma 10 and 11 in results in while we have ,
therefore the main statement of Lemma 1 goes through. ∎
Lemma 2 We have the following for the determinant of through first matrix determinant lemma and second Sylvester’s determinant identity;
Using the fact that we have