Worst Cases Policy Gradients
Yichuan Charlie Tang, Jian Zhang, Ruslan Salakhutdinov
Introduction
One of the key challenges for building intelligent systems is developing the capability to make robust and safe sequential decisions in complex environments. Towards this goal, the recent breakthroughs in deep reinforcement learning (RL) are very encouraging and have led to super human performance in various video games and board games .
As we move towards real-world applications and learning from increasingly diverse environments, we may encounter both parametric and inherent uncertainties due to the stochastic nature of the model and the environments . One cause of stochasticity is the inability to fully observe the state (e.g. intentions, beliefs) of other agents in a multi-agent environment. Robust handling of uncertainties and risks are a must before we can fully leverage the power of deep RL in real world safety-critical applications such as self-driving . However, standard RL optimizes the average expected return and is not risk sensitive .
In this paper, we take a step towards this goal by developing a novel deep RL architecture that optimizes a risk-sensitive criterion. In contrast, the vast majority of existing deep RL techniques maximize the expected value over possible future returns . Maximizing expectation, however, is not risk-sensitive as it does not explicitly penalize rare occurrences of catastrophic events. Working under the assumption that the future return is inherently stochastic, we first start by modeling its distribution. The risk of various actions can then be computed from this distribution. Specifically, we use the conditional Value-at-Risk as the criterion to maximize. Our architecture is based on the actor-critic , but we modify our critic to predict the full distribution over future returns instead of simply the expectation. The proposed framework (Fig. 1), which we call worst cases policy gradients (WCPG), learns a continuous family of policies, each optimizing for their respective level of risk: . Depending on the risk appetite, the learned policy can choose to act differently even from the same state.
After introducing our framework and the learning algorithm in Section 3, we demonstrate the effectiveness of WCPG on two traffic simulation scenarios, where an agent must learn to safely interact with other agents and to achieve their own goals. In Section 4, we describe the details of our environments, training procedure, and quantitative results, where we find WCPG learns risk-averse policies which are significantly more robust in test scenarios.
Preliminaries
Policy gradient methods directly optimize the policy parameters to maximize the expected total return. Popular in continuous control, they compute the gradients with respect to the objective :
Using the Policy Gradient Theorem and the log-derivative trick , the gradient of the objective can be written as:
where the gradient does not dependent on the state distribution . have shown that it may be beneficial to substitute with other terms, where it can be one of several expressions: , , or the advantage function , leading to a lower variance of the gradients.
While Eq. 2 holds for any stochastic policy, introduced the Deterministic Policy Gradient (DPG) theorem for deterministic policies, whose gradients can be estimated more efficiently. DPG states (under certain regularity conditions) that for a deterministic policy :
The deep deterministic policy gradient (DDPG) framework extended DPG to large continuous state action space environments. DDPG works well for continuous control and is a popular variant of the actor-critic policy gradients method. A critic function learn to estimate using temporal difference bootstrapping and provides the training signal for the actor (policy network). Variants of DDPG have also been explored, including actors with parametrized actions .
Worst Cases Policy Gradients
Standard reinforcement learning maximizes for the expected (possibly discounted) future return . However, maximizing for average return is not sensitive to the possible risks when the future return is stochastic, due to the inherent randomness of the environment not captured by the observable state. For example, when the return distribution has high variance or is heavy-tailed, finding a policy which maximizes the expectation of the distribution might not be ideal: a high variance policy (and therefore higher risk) that has higher return in expectation is preferred over low variance policies with lower expected returns. Instead, we want to learn more robust policies by minimizing long-tail risks, reducing the likelihoods of bad outcomes.
Formally, the criterion that we care about is the -percentile of the distribution over returns. This is captured by the conditional Value-at-Risk:
where is the -percentile of . As , the policy will focus on performing well in the “worst cases” scenarios. Computing CVaR metric directly for long time horizons (e.g. a brute-forced approach like sampling ) would be prohibitively expensive. As an alternative, we extend DDPG (an actor-critic architecture) so that the critic learns to model the distribution over expected total discount reward. Provided a good distributional critic can be learned, CVaR can be computed from the distribution and we can train the actor by backpropagating the gradient back through the actor or policy network.
At this point, the reader might wonder if it is possible to first train (offline) a critic for and then obtain a risk-sensitive policy (online) simply by selecting actions to maximize the -percentile of at every state. However, this approach would fail as is conditioned on executing policy for all subsequent steps. Selecting the action (which modifies the original policy) that gave the best percentile at time is not meaningful if we revert to policy for time and beyond. In other words, multiple actions at a given state would have similar distributions over , as a single action (especially continuous) at time can be nullified by future actions.
To learn the critic, we must define the equivalent Bellman operator for distributions. Let us first define as the distribution over total future (possibly discounted) return generated by executing policy until reaching a terminal state. We define as the transition operator:
There are two sets of projection equations for the distributional Bellman operator. The projection for (mean) is the standard Bellman equation, similar as in Q-learning: . The projection for (variance) is:
The following equation for the variance of the future return exists.
A straight-forward proof is given in the Appendix. Eq. 1 is the parametric distributional Bellman operator for the distributional state-action value function. In practice, due to a relatively large continuous state and action spaces, we use function approximation for representing the critic. Specifically, a neural network parameterized by is used for estimating the critic’s mean and variance: .
We learn the critic by using the Temporal Difference (TD) learning, where the TD error is obtained by assuming Gaussianity (maximum entropy distribution) and using the Wasserstein metric as the loss. While other losses such as KL-divergence could also be used, the distributional Bellman operator using the -th Wasserstein metric has been shown to be a contraction operator for policy evaluations . The -th Wasserstein distance between two probability distributions and is defined as :
where is the inverse cumulative distribution function (CDF). Assuming and , the 2-Wasserstein distance simplifies to:
The critic will try to minimize the Wasserstein distance by backpropagating the gradient .
2 CVaR Actor
For every state action , we can use the critic’s estimated (mean and variance of a Gaussian) to compute the measure in closed-form:
where is the standard normal distribution and is its CDF: \Phi(x)=\frac{1}{2}\big{(}1+\operatorname{erf}(x/\sqrt{2})\big{)}. We use to explicitly denote the reliance of on both the state, action, as well as a specific policy . We stress that is a scalar term computable at every state-action pair . In plain language, is the expected future CVaR when in state and executing action , and following policy hereafter.
We now replace the standard objective of Eq. 1 by the risk-averse objective:
Suppose that the MDP satisfied conditions A.1 and A.2 (see Appendix), then:
Proof. The proof mainly follows the original Policy Gradient Theorem (PGT) , see Appendix.
Given Eq. 11, we follow DPG from Eq. 3 to derive the equivalent deterministic policy gradient, where we backpropagate the CVaR loss gradients through the critic and down to the actor/policy networks. We obtain the deterministic gradient for the actor network by following a similar derivation as in :
Note that our new objective is dependent on “risk level” , and the question now arises as to how to choose the percentile , which ranges from to . One strategy is to discretize into discrete values and train separate policy networks optimizing for different settings. Instead of this naive approach with times more parameters, we learn a single network which is conditioned on as an additional input (see Fig. 1). The advantage of this approach is that we can learn a continuous family of conditional policies with varying risk-sensitivity.
During training, a different input will have a different loss function . To train for all s, we uniformly sample during the start of an episode and fix for the entirety of that episode. During inference, can output different actions given the same exact state , conditioned on the setting of . Intuitively, a small leads to conservative actions while a larger leads to more aggressive actions.
We now have the learning equations for both the actor and the critic. We employ an off-policy training algorithm by using an experience replay buffer. Algorithm 1 in the Appendix outlines the entire WCPG training procedure. Fig. 1 provides an overall illustration of the entire WCPG architecture and gradient flow during training.
Experimental Results
We tested our algorithm on two continuous-action 2D driving environments, focusing on two of the more critical scenarios: unprotected turns and merges (Figs 3, 4). In our environments, each agent is a vehicle, with the goal of getting from point A to point B while staying on the road and avoiding collisions. We simulate the vehicle dynamics using a discrete time kinematics bicycle model . The simulation timestep is 100 milliseconds. We restrict our action space to be the acceleration of the ego vehicleEgo refers to the vehicle for which we are learning a policy. We bound the acceleration and deceleration of all vehicles to , similar to that of a typical real vehicle.. The steering is determined by the Stanley controller, a non-linear closed loop feedback steering controller . The physical dimension of our simulation environments is approximately 200 meters by 200 meters, where agents are randomly spawned on specific “birth” lanes with a given probability. Ego’s initial velocity is randomly chosen between 5 to 20 m/s.
As a part of our multi-agent environments, non-ego agents are controlled by rule-based behaviors. Their velocity is randomly chosen and they have the ability to perform adaptive cruise control: slowing down and speeding up according to the vehicle in front of them. They can also perform safe lane changes by dynamically planning a smooth trajectory in order to merge from one lane to another. To introduce more realistic behaviors, the non-ego agents, during spawning, each samples randomly from one of three behaviors (when close to ego): yield, ignore, or accelerate, mimicking human drivers who might be conservative, distracted, or aggressive.
The reward for catastrophic failure (e.g. collision) is , while the reward for successful completion of a maneuver depends on time-to-completion: . A failure to complete the maneuver results in a reward of . We initially attempted to learn safer policies by simply scaling up the collision penalties (e.g. ), however we found that the resulting policies learned to be overly conservative and often did not even attempt the maneuvers. See Appendix for more details.
2 WCPG Network
Our network is an actor-critic based on DDPG with two modifications. The first is that an additional CVaR input is added to both the actor and the critic. The second is that instead of producing a scalar value for approximating , the critic outputs parameters which govern the distribution of future returns. We use the softplus function to guarantee that the critic’s estimation of variance will always be positive. See Fig. 1 for the diagram of our network model. The input to our network is dimensional and consists of encoding the closest vehicles around ego (in sorted order) and their states: position, heading, and velocity. The actor consists of 3 hidden ReLU layers with 32 hidden units each. The output of the actor is ego’s acceleration. The critic network consists of 4 hidden layers with 64 hidden units each. We did not find the performance to be particularly sensitive to the choice of number of layers and layer size. See Appendix for more details.
Off-policy training is performed with an experience replay buffer of up to tuples. The learning rate for both the actor and the critic is , and the minibatch size is . The same action is executed times in a row for a total of ms. Exploration noise is Gaussian with a standard deviation of 2.0. Input mean and variance for normalization is calculated in an online moving average fashion. is randomly sampled from . Training is run until convergence for 5000 episodes, where each episode can last up to 30 seconds in simulation (300 timesteps).
3 Testing Performance
We now look at WCPG performs on test environmentsThe random seeds for the testing environments are different compared to training environments.. Results are shown for the two environments in Fig. 5. In the left panel, we can see that as decreases, the policy becomes risk-averse and more robust, reducing the probability of collisions. In the middle, we can see that smaller s leads to more conservative behavior (e.g. waiting for a wider gap), increasing the time-to-completion for both environments. Finally, the right panel shows that the critic’s own estimate of uncertainty grows with an increasing . This also means that the actor also has a higher risk-tolerance with an increasing .
Params WCPG (varying ) State-of-the-art baselines vel. spwn. 0.02 0.1 0.3 0.6 1.0 DDPG PPO C51 D4PG (train) 1% 0% (100%) 0% (100%) 0% (100%) 2% (98%) 15% (85%) 0% (100%) 4% (96%) 2% (98%) 2% (98%) +5 m/s 5% 0% (100%) 0% (100%) 0% (100%) 0% (100%) 15% (85%) 13% (87%) 24% (76%) 10% (90%) 4% (96%) +10 m/s 5% 0% (100%) 0% (100%) 0% (100%) 0% (100%) 15% (85%) 18% (82%) 40% (60%) 11% (89%) 4% (96%) +15 m/s 5% 0% (100%) 0% (100%) 0% (100%) 3% (97%) 21% (79%) 20% (80%) 49% (51%) 14% (86%) 2% (98%) +0 m/s 2% 0% (100%) 0% (100%) 0% (100%) 0% (100%) 19% (81%) 7% (93%) 15% (85%) 4% (96%) 5% (95%) +0 m/s 8% 0% (100%) 0% (100%) 0% (100%) 3% (97%) 26% (74%) 6% (94%) 15% (85%) 5% (95%) 1% (99%) +10 m/s 8% 0% (100%) 0% (100%) 0% (100%) 2% (98%) 23% (77%) 19% (81%) 40% (60%) 11% (89%) 2% (98%)
Params WCPG (varying ) State-of-the-art baselines vel. spwn. 0.02 0.1 0.3 0.6 1.0 DDPG PPO C51 D4PG (train) 1% 0% (100%) 0% (100%) 0% (100%) 0% (100%) 8% (92%) 0% (100%) 0% (96%) 0% (100%) 6% (94%) +5 m/s 1% 0% (4%) 1% (5%) 0% (54%) 2% (90%) 10% (88%) 4% (96%) 31% (68%) 9% (91%) 9% (91%) +10 m/s 1% 0% (2%) 0% (5%) 0% (50%) 0% (93%) 8% (92%) 3% (97%) 26% (73%) 5% (95%) 8% (92%) +15 m/s 1% 0% (1%) 0% (5%) 0% (57%) 1% (94%) 9% (91%) 1% (99%) 24% (76%) 6% (94%) 9% (91%) +0 m/s 2% 1% (7%) 0% (11%) 1% (50%) 2% (78%) 18% (76%) 5% (95%) 40% (59%) 8% (92%) 14% (86%) +0 m/s 3% 0% (6%) 0% (10%) 0% (43%) 2% (70%) 17% (79%) 5% (95%) 39% (59%) 9% (91%) 12% (88%) +10 m/s 3% 0% (4%) 0% (8%) 1% (43%) 2% (76%) 11% (79%) 4% (96%) 33% (67%) 5% (95%) 10% (90%)
4 Extrapolation Performance
A challenging test is in the ability for policies to extrapolate, where the parameters of testing environments are “out-of-distribution” or extrapolated from the training distribution. Specifically, we significantly increase the upper-bound of non-ego agents’ velocity by an additional m/s and the spawn rate upwards of . The total number of spawned agents is also increased (up to 4 additional agents) to make the scene denser.
Tables 1, 2 show the performance of WCPG and baselines on the extrapolation environments. We report collision rates for 100 random trials. Our results are highlighted in green. The training results are reported in the first row of each environment. We compare with 4 state-of-the-art algorithms in DDPG, proximal policy optimization (PPO) , C51/Rainbow , and D4PG . We used open source implementations from OpenAI and Dopaminehttps://github.com/openai/baselines, https://github.com/google/dopamine.. For DDPG training, we used default parameters, but ensured that the model architecture, batch size, exploration noise, number of updates and other parameters matched that of the WCPG training parameters. PPO training was performed with minibatch size of 32, , , , and learning rate of . Training was performed until convergence.
For all of the methods, training performance is quite good with almost all of them achieving crash rates. However, if we use out-of-distribution environment parameters, PPO shows severe overfitting. DDPG is better but also starts to experience collisions. Note that C51/Rainbow and D4PG are distributional RL techniques (highlighted in blue), but they do not optimize for any risk-sensitive criteria. For all baseline methods, the performance significantly degrades in the presences of changing environmental parameters. In contrast, WCPG policies are more robust as we reduce , achieving crash rate for the left turn environment and between and for the merge environment, while keeping the success rate high. See Appendix for full results.
5 Uncertainty Modeling
It is also interesting to examine how the critic predicts uncertainty by looking at the critic’s own estimate of the standard deviation of future return. Fig. 6 shows six intermediate steps from a test episode along with critic’s estimations. As ego approaches the intersection and begins to make the left turn, the variance gradually increases and crescendos when ego is in the path of oncoming traffic. As ego completes the turn, uncertainty quickly reduces.
6 Carla Simulation
We tested how well our policy can transfer to similar scenarios in the CARLA simulator . It currently contains seven different towns, dozens of different vehicles, and simple traffic law abiding “auto-pilot” CPU agents. For our left turn and merge environments, we found locations within CARLA that had similar geometry (See Fig. 7). We extract the position, heading, and velocity of vehicles from the CARLA simulator and fed into our previously trained WCPG policy networksDue to a much slower simulation speed, we leave training directly on CARLA as future work.. We report collision and success rates for 100 random trials in Tab. 3. We can see that even in a different simulation, lower values improved robustness dramatically.
Unprotected Left Turn: (Town05) 0% (100%) 24% (76%) 42% (58%) Merge: (Town04) 2% (83%) 4% (89%) 24% (76%)
Related Work
Risk-sensitive, safe RL and the related robust MDPs have been extensively studied in literature . A recent comprehensive survey on the topic is provided by , where safe RL is categorized into two main types. The first type is based on the modification of the exploration process to avoid unsafe exploratory actions. Strategies consist of incorporating external knowledge or using risk-directed explorations . The second type modifies the optimality criterion used during training. Our work falls under this latter category, where we try to optimize our policy to strike a balance between pay-off and avoiding catastrophic events.
Different types of “safe” criteria have been proposed: exponential utility functions , linear combination of return and variance , percentile performance , Sharpe ratio and other variance-related criteria . The worst-case minimax criterion can also be directly maximized . In , -Learning was introduced, where the function is the lower bound of the function. However, optimizing the minimax criterion can result in overly pessimistic policies . Optimizing a constrained criterion (return) subjected to a bounded variance was proposed in .
Conditional Value-at-Risk (CVaR) is another criterion that is gaining popularity in various fields such as engineering and finance . It has various desirable properties, including coherence and the easy of computation for certain underlying distributions. Combing CVaR with reinforcement learning, policy gradients for a bounded CVaR was proposed by , while a sampling based algorithm for CVaR was discussed in . However, these approaches directly estimate the gradients to CVaR and are often computationally expensive or require extensive trajectory roll-outs. In contrast, WCPG optimizes for CVaR indirectly by first using distributional RL techniques to estimate the distribution of return and then compute CVaR from this distribution. Recently, there have been a resurgence of interest in distributional RL . However, the estimated distributions have not been used to minimize any risk-sensitive criteria. While implicit quantile networks do optimize for risk-sensitive measures, a disadvantage to their approach is the approximation of risk and its gradients via discrete samples, which greatly increases computation complexity.
Our proposed WCPG builds on DDPG, a popular deep RL method for continuous control. We retain DDPG’s advantage of a powerful and sample efficient off-policy method that works well for large continuous state and action space. Instead of directly optimizing for CVaR, which can be difficult, we compute CVaR in closed-form from our distributional critic’s estimation of future return, without resorting to sampling. In addition, both our actor and critic take risk-tolerance as input during training, which allows the learned policy to operate with varying levels of risk after training.
Discussions
We have proposed a novel actor-critic framework to learn risk-sensitive policies by maximizing CVaR. Our policies can be adjusted dynamically after deployment to select risk-sensitive actions. In simulated driving environments, we perform significantly better with a smaller , when compared to other top RL algorithms. Modeling uncertainty with heavy-tailed or mixture distributions could be explored in the future. It would also be interesting to see if the automatic selection of would be possible for completing a maneuver while minimizing risk exposure.
Acknowledgements We thank Barry Theobald, Hanlin Goh, Nitish Srivastava, Johannes Heinrich, and the anonymous reviewers for making this a better manuscript.
References
APPENDIX
Appendix A Proof of Proposition 1
For convenience, we first restate the definitions here:
We now take out the first term out of the summation:
Now, we can reuse the definition of and to arrive at the recursion:
Appendix B Proof of Theorem 2
The policy function is deterministic: , where and is the Dirac delta function. This is a mild condition as our proposed WCPG is based on the deterministic policy gradient (DDPG), which also assumes deterministic policy functions.
Condition A.2
Given a particular state and action , the environment transition is deterministic, that is: , where is the Dirac delta function. We note here that environment determinism for continuous control is often made by widely popular RL environments. For example, both Mujoco simulator and the DeepMind Control Suite are deterministic. The original Atari Learning Environment (ALE) also has deterministic dynamics.
For clarity and without loss of generality, we define without conditioning on risk parameter :
where is an dependent constant.
Let us define to be the expected conditional Value-at-Risk for policy starting from state , and given Condition A.1 is satisfied,
Proof. It is trivial to see that as approach the Dirac delta function, probability mass concentrate on the chosen deterministic action .
Provided that Condition A.2 is satisfied,
From Eq. 4 in proof of Proposition 1, we substitute :
Given Condition 2, we simplify expectations involving summations over , and rearranging:
We mainly follow the start state formulation of the Policy Gradient Theorem of (Sutton et al. 2000), but also making use of Propositions 3 and 4.
where we have used to denote the probability of going to state from state under policy in steps. This can be seen after unrolling (Eq. 22) for a few steps. It can then be seen that:
Using the log-derivative trick we can further write the gradient as:
And for conditional objective conditioned on , we have:
Appendix C Algorithm
Appendix D Distributional Critic
Our distributional critic uses the Gaussian approximation as it allows us to have closed form CVaR computation, which is critical as we must compute CVaR for every update step, for every data tuple. In addition, ease of computation is important when we want to learn a policy which is conditioned on the risk , where we have to compute CVaR for multiple s.
While a Gaussian is unimodal and not heavy-tailed, our environmental reward is lower-bounded by the negative reward caused by a collision. In this case, Gaussian distribution’s light-tail is not terribly restrictively, since we do not have unbounded loss. In addition, it is important to note that the unimodal Gaussian approximation is for a particular state-action pair. The value distribution for a particular state (e.g. distributional ), which marginalizes over the actions, could still be multi-modal with even potentially different variances for different modes.
Appendix E Experiments
We focused our experiments on self-driving environments as safety and risk-sensitive decision making is paramount in this application. They are also multi-agent environments, where the latent behaviors of other agents are not always observable. Specifically, as noted in the paper, the other on-coming vehicles have a random probability of 3 behaviors: yielding, ignoring, or aggressively accelerate. This stochasticity creates inherent uncertainty in the environment and leads to the need for a distributional RL.
See Table 4 for the details of the training environment parameters.
E.1.2 Network Parameters
We show the exact network parameters for the WCPG actor and critic network used for the driving environment simulations in Table 5.
E.2 Additional Proof-of-concept Experiment
We demonstrate the concept of WCPG in a simple discrete setting where we must learn a policy to decide to drive in the fast lane or the slow lane on a highway. The problem is a highly simplified proof-of-concept MDP where we can leverage brute-force sampling to perform policy improvement and find the optimal policy. This allows us to test the core concepts behind WCPG independently from the various approximations needed in learning deep RL policies for more complex environments.
Imagine a vehicle is traveling on a two lane freeway and needs to learn a policy for performing lane selection. The left lane is the fast lane while the right lane is the slower lane. At every timestep, the rewards are stochastic and the fast lane reward has both higher mean and higher variance than the slow lane. The stochastic policy we wish to learn is the probability of lane change at each timestep.
We formulate this as a discrete finite-horizon MDP with two states, , for the vehicle being in the left and right lanes, respectively. The MDP is shown in Fig. 8, where the total number of time steps is 4. The initial state at is the right lane and at each timestep the policy outputs the probability of lane change: . When the vehicle is in the left lane, it receives a reward randomly sampled from . When the vehicle in the right lane, its reward is sampled from . This roughly models the assumption of higher risk and rewards for the fast lane, while the slow lane might be safer but would take longer to get to the destination.
Using this environment, we are interested in learning an conditioned optimal policy and to verify our hypothesis that different s will lead to different policies. The MDP is simple enough such that sampling based policy improvement schemes will arrive at the optimal solution.
The Fast-Slow Lanes MDP can be solved via brute force sampling based policy improvement. For a given , we grid the entire policy space. Specifically, we enumerate from to with 32 intervals. Iterating through the policies, we numerically evaluate each policy using 1000 random trials, where the CVaR objective is computed from the statistics of the trials. The optimal policy is then found by selecting the policy that achieves the highest CVaR for any given . We plot the optimal policies in Fig. 9.
The optimal policy is a scalar specifying the probability for left/right lane selection. When is small, the optimal policy is to stay in the right lane as the rewards have a much smaller variance. Conversely, when approaches , the optimal policy is to choose the left lane as we obtain higher rewards in expectation. This results verifies our hypothesis and is expected given the specified distribution of rewards. While this is a simple proof-of-concept problem, it demonstrates that when a globally optimal solution can be found, the proposed CVaR objective does indeed lead to conditional policies that varies depending on user specified risk appetite.
E.3 Extrapolation Experiments Full Results
We report the full extrapolation results for the unprotected left and merge environments in Tables E.3, 7, 8, 9. The numbers are all in percentages and we also report the standard errors of the mean, over 100 trials with random environment seeds.
Params WCPG (varying ) State-of-the-art baselines vel. spwn. 0.02 0.1 0.3 0.6 1.0 DDPG PPO C51 D4PG (train) 1% +5 m/s 5% +10 m/s 5% +15 m/s 5% +0 m/s 2% +0 m/s 8% +10 m/s 8%
Params WCPG (varying ) State-of-the-art baselines vel. spwn. 0.02 0.1 0.3 0.6 1.0 DDPG PPO C51 D4PG (train) 1% +5 m/s 5% +10 m/s 5% +15 m/s 5% +0 m/s 2% +0 m/s 8% +10 m/s 8%
Params WCPG (varying ) State-of-the-art baselines vel. spwn. 0.02 0.1 0.3 0.6 1.0 DDPG PPO C51 D4PG (train) 1% +5 m/s 1% +10 m/s 1% +15 m/s 1% +0 m/s 2% +0 m/s 3% +10 m/s 3%
Params WCPG (varying ) State-of-the-art baselines vel. spwn. 0.02 0.1 0.3 0.6 1.0 DDPG PPO C51 D4PG (train) 1% +5 m/s 1% +10 m/s 1% +15 m/s 1% +0 m/s 2% +0 m/s 3% +10 m/s 3%