The Distracting Control Suite -- A Challenging Benchmark for Reinforcement Learning from Pixels

Austin Stone, Oscar Ramirez, Kurt Konolige, Rico Jonschkowski

I Introduction

The DeepMind Control Suite (DM Control) is one of the main benchmarks for continuous control in the reinforcement learning (RL) community. By providing a challenging set of tasks with a fixed implementation and a simple interface, it has enabled a number of advances in RL – most recently a set of methods that solve the benchmark as well and efficiently from pixels as from states . Simulation-based benchmarks like DM Control have many advantages: they are easy to distribute, they are hermetic and repeatable, and they are fast to train and iterate on. However, the DeepMind Control Suite is a poor proxy for real robot learning from visual input, which remains inefficient despite the advances we have seen in DM Control. To enable research gains in simulated benchmarks to better translate to gains in real world vision-based control, we need a new simulated benchmark that more closely mirrors perceptual challenges of real environments, most importantly visual distractions – variations in the input that are irrelevant for the task.

A major challenge in perception is to extract only the task-relevant information from sensory input and remove distractions which might otherwise lead to spurious correlations in downstream tasks . DM Control does not contain such distractions, as the agent is shown from a constant camera view under constant lighting against a singular, static background. Since every change in the observation is tied to the change of a task-relevant state variable, DM Control does not allow measuring or making progress on the ability of filtering out irrelevant variations through perception.

To address this problem, we present the Distracting Control Suite, an extension of DM Control created with real-world robot learning in mind. Our extension adds three distinct types of distractions: random color changes of all objects in the scene, random video backgrounds, and random continuous changes of the camera pose (see Fig. 1). Each of these distractions can be applied in a static setting where changes only occur at episode transitions or in a dynamic setting where distractions change smoothly between frames. All distractions can be scaled in their difficulty from barely perceptible to severely distracting. All three distraction types can be arbitrarily combined with each other.

We implement these distractions on top of DM Control to retain the same simple interface. Our suite works by accessing and modifying scene properties (color, camera position, and background textures) at run time before visual observations are rendered. The underlying physics and control properties of the tasks are kept exactly the same to facilitate comparisons to work performed on the original DM Control.

Using the Distracting Control Suite, we perform an empirical analysis of state of the art methods in reinforcement learning from pixels, comparing different combinations of SAC and QT-Opt with RAD and DrQ . We analyze a) the sensitivity to distractions during inference when no distractions are present during training, b) the effect of each individual distraction on RL performance at different difficulties, and c) a combination of all three distractions which proves to be challenging for existing methods.

We have three main contributions: 1) The design and implementation of the Distracting Control Suite, which is available at https://github.com/google-research/google-research/tree/master/distracting_control and which we hope will facilitate future advances in vision-based control. 2) The definition of a benchmark with results for the current state of the art that future work can compare against. 3) A set of empirical observations about RL from pixels when faced with distractions, such as i) methods are relatively robust to especially color distractions without training on them but struggle to improve substantially from seeing distractions during training, ii) distractions interact in a way that makes combinations of them especially difficult, iii) the relative performance of different methods changes significantly between the DM Control and our Distracting Control benchmarks. We think that these observations are especially relevant to real world, robot RL where task-irrelevant visual input is very common. We hope that our benchmark can be a useful proxy for learning visual control in the real world and therefore facilitate advances in robot learning.

II Related Work

Learning successful policies from pixels in the Atari environment was a major breakthrough in reinforcement learning that produced a surge of interest and advances in pixel-based RL. The work in simulation first focused on Atari, but later also included DM Control from pixels . Recently, CURL , DrQ , and RAD have established that different versions of applying image cropping augmentation can greatly improve results up to a point where DM Control can be solved similarly from pixels as from states. Reinforcement learning has also been successfully applied to robotics training in the real world .

An alternative approach to training on real robots is to use domain randomization to train in a very diverse set of simulated environments that enables transfer to the real world . Domain randomization is the extension of data augmentation, which has been used in computer vision since the inception of convolutional networks , from data sets to simulators. Randomizing many aspects of the simulation that do not match the real world forces the learned model to be robust to these variations.

Distractions, which this paper focuses on, can look technically similar to domain randomization but distractions as we define them here are part of the problem that the agent has to solve rather than part of the solution. As a result, the agent does not have control over distractions, i.e. cannot affect these distractions, cannot arbitrarily sample more of them, and has to handle them during evaluation. The importance of visual distractions for studying perception and control was first demonstrated in simple environments and has recently been applied to more complex ones , including different modifications to the DeepMind Control Suite .

The goal of our work is to provide a unifying benchmark with visual distractions to enable comparability between approaches for pixel-based RL that currently rely on different sets of distractions. Compared to distractions that were added to DM Control in previous or concurrent work, our benchmark combines camera, color, and background distractions, and presents an in-depth study of state of the art methods in this new setting. We hope that our empirical observations and our Distracting Control Suite with clearly defined benchmarks will facilitate future research in this direction.

III The Distracting Control Suite

This work extends the DeepMind Control Suite to make its perception aspect more challenging by adding visual distractions. The resulting Distracting Control Suite applies random changes to camera pose, object colors, and background. The magnitude of each distraction type can be controlled by a “difficulty magnitude” scalar between 0 and 1. Distractions can be set to either change during episodes or change only between episodes, which we will refer to as dynamic and static settings, respectively.

For the viewing camera, the difficulty magnitude scales both the span of camera poses and the camera velocity. For the color change augmentations, the difficulty magnitude scales the maximum allowable color change and the speed of color changes, and for the background distractors it scales the number of unique videos used or (for one of our experiments) the weight for blending between the background videos and the original skybox background.

We parameterize the camera pose by c=(ϕ,θ,r,θroll)c=(\phi,\theta,r,\theta_{roll}), corresponding to the spherical angles ϕ\phi and θ\theta and radius rr, which define the camera position, and an additional angle θroll\theta_{roll} that specifies the roll. The camera’s pitch and yaw are not randomly varied. Depending on whether the task uses a tracking camera, e.g. for cheetah and walker, or a “fixed” camera, e.g. for cartpole, pitch and yaw are calculated to focus on the agent’s current or starting position, respectively. The difficulty scale defines a viewing range of the camera as a subset of the upper frontal hemisphere for azimuth and elevation that scales the maximum distance (see Fig. 2). Based on the difficulty scale βcam∈\beta_{\text{cam}}\in, we set ϕmax=θmax=θroll max=πβcam2\phi_{max}=\theta_{max}=\theta_{roll\,max}=\frac{\pi\beta_{\text{cam}}}{2}, rmin=roriginal(1−0.5βcam)r_{min}=r_{\text{original}}(1-0.5\beta_{\text{cam}}), and rmax=roriginal(1+1.5βcam)r_{max}=r_{\text{original}}(1+1.5\beta_{\text{cam}}). Therefore, 0≤ϕ,θ,θroll≤π20\leq\phi,\theta,\theta_{roll}\leq\frac{\pi}{2}, and 0.5roriginal≤r≤2.5roriginal0.5r_{\text{original}}\leq r\leq 2.5r_{\text{original}}. In the static setting, we uniformly sample the camera pose from this range at the start of each episode and keep it constant during the episode. In the dynamic setting, we sample the camera’s starting pose in the same way, but additionally maintain a camera velocity vtv_{t} that is updated via a random walk at each time step.

Velocity is stored as both an (x˙,y˙,z˙)(\dot{x},\dot{y},\dot{z}) spatial vector and a θ˙\dot{\theta} roll velocity. The random walk’s standard deviation and maximum velocity are scaled relative to the viewing range, vmax=2βcam5v_{max}=\frac{2\beta_{\text{cam}}}{5}, σ=βcam10\sigma=\frac{\beta_{\text{cam}}}{10}, vroll max=πβcam50v_{roll\,max}=\frac{\pi\beta_{\text{cam}}}{50}, σroll=πβcam300\sigma_{roll}=\frac{\pi\beta_{\text{cam}}}{300}. The random walk is clipped to within the maximum velocity and camera pose parameters.

III-B Object Colors

For this distraction type, we change the colors of all bodies in the simulation, where the difficulty scalar βrgb∈\beta_{\text{rgb}}\in defines the maximum distance per color channel. At the start of each episode, all colors are sampled uniformly per channel x0∼U(x−βrgb,x+βrgb)x_{0}\thicksim\mathcal{U}(x-\beta_{\text{rgb}},x+\beta_{\text{rgb}}), where x0x_{0} is a sampled color value and xx is the original color in DM Control. In the static setting, the colors remain constant throughout the episode. In the dynamic setting, they change randomly xn=xn−1+N(0,0.03⋅βrgb)x_{n}=x_{n-1}+\mathcal{N}(0,0.03\cdot\beta_{\text{rgb}}), but are clipped to never exceed the maximum distance βrgb\beta_{\text{rgb}} from the original color.

III-C Background

Here, we project random backgrounds from videos of the DAVIS 2017 dataset onto the skybox of the scene. To make these backgrounds visible for all tasks and views, we make the floor plane transparent except for walking tasks where it is small and task relevant (we set the ground plane opacity α=1.0\alpha=1.0 for cheetah and walker, α=0\alpha=0 for reacher, α=0.3\alpha=0.3 for all other tasks). Depending on the experiment, we use a different number of background videos b∈b\in – the task is more difficult when more scenes are used. We take the bb first videos in the DAVIS 2017 training set and randomly sample a video and a frame from it at the start of every episode. In the static setting, that frame stays constant. In the dynamic setting, the video plays forwards or backwards until the last or first frame is reached at which point the playing direction is reversed. This way, the background motion is always smooth and without “cuts”. In one experiment, we smoothly blend between the distraction background and the original skybox background with weights βbg\beta_{\text{bg}} and 1−βbg1-\beta_{\text{bg}} respectively (see Fig. 3).

IV Methods for RL from Pixels

Our experiments compare SAC and QT-Opt with and without random cropping following the RAD approach with a single random cropping per sample or averaging over two crops as detailed in DrQ . This can also be viewed as using DrQ with K=M∈{0,1,2}K=M\in\{0,1,2\}. While QT-Opt in fact already includes random cropping in its description, for consistency, we will refer to that approach as QT-Opt+RAD and have QT-Opt denote the method without cropping.

To implement QT-Opt+DrQ, we modify the Bellman error minimization. Originally QT-Opt proposes

where the cross-entropy function is used as the divergence metric DD and QtQ_{t} is the target value defined by Qt(s,a,s′)=r(s,a)+γV(s′)Q_{t}(\mathbf{s},\mathbf{a},\mathbf{s^{\prime}})=r(\mathbf{s},\mathbf{a})+\gamma V(\mathbf{s^{\prime}}). To compute VV, QT-Opt estimates a QQ-function and uses CEM to select the best action according to the current Q^\hat{Q} estimate. Adding DrQ augmentations requires two changes to the algorithm. First, we need to average the target value over K=2K=2 random image crops,

where ff is the image transformation function and υk∼U\upsilon_{k}\sim\mathcal{U} a random sample of image augmentation parameters. Second, we need to average the QQ estimates in the loss,

All methods use faithful replications of their published hyperparameters without special tuning to the Distracting Control Suite. Note that while SAC+DrQ was originally tuned for DM Control from pixels, QT-Opt was only tuned for DM Control with state input.

V Experiments

In this section, we analyze state of the art reinforcement learning methods in our Distracting Control Suite, which yields a number of interesting results: 1) Methods trained without any distractions are fairly robust to color distractions an somewhat robust to camera distractions during inference. 2) Training with distractions does not substantially improve this robustness, except for background distractions where performance improves but only up to a point. For random color changes in particular, the improvement from training with these distractions is minor compared to not training with them. 3) Training with random video backgrounds performs better than training with random static backgrounds. Generalization to new backgrounds is limited and does not improve when training on additional background scenes. 4) The degrading effects of distractions on task performance are more than multiplicative. As a result, current methods perform rather poorly in our benchmarks that combine three kinds of distractions, even in the easiest settings. 5) The ranking of methods changes from the standard DM Control benchmark – where SAC-based and QT-Opt-based methods perform comparably – to our distracting benchmarks, where QT-Opt with RAD or DrQ augmentations performs best. Generally, we found that RAD variants worked equally well or better than DrQ across our experiments.

All methods use the same model architecture from DRQ . A shared image encoder applies four convolutional layers using 3×33\times 3 kernels and 32 filters with an stride of 2 for the first layer and 1 for others. ReLU activations are applied after each convolution. A final 50 dimensional output dense layer normalized by LayerNorm is applied with a tanh activation. Both critic and actor networks (in the case of SAC) are parametrized with a 3-layer MLP using ReLU activations up until the last layer. The output dimension of these layers is 1024. In the critic this reduces to a single Q-Value prediction, and in the case of the actor it predicts a mean and covariance for each action. The image encoder weights are shared when using SAC across the critic and the actor, and gradients are only computed through the critic optimizer.

Tasks and Experiment Parameters

Training is performed with batch size 512, and alternates one learning step with each sample collection step. Tasks and action repeats are adopted from the Planet benchmark (see Table II). All experiments report results after 500K environment steps, evaluated for 100 episodes. Unless otherwise noted all experiments are performed with five random seeds per task used to compute means and standard errors of their evaluations. In tables, results are boldfaced if they have the highest mean or if they do not have a statistically significant difference (p<0.1p<0.1) from the result with the highest mean.

V-A Robustness to Distractions During Inference

In this experiment, we analyze how well methods trained on the standard DM Control benchmark generalize to unseen distractions during inference. After training each method without any distractions, we then test them separately for each type of distraction with different amounts of distracting variation βrgb\beta_{\text{rgb}}, βcam\beta_{\text{cam}}, βbg\beta_{\text{bg}} from 0 to 1. The number of background scenes b=60b=60.

Table II shows the results in DM Control without distractions and verifies that methods are learning to solve the tasks. We see that using one or two cropping augmentations (RAD or DrQ) is necessary for reaching high performance and that SAC-based methods and QT-Opt based methods perform comparably. Figure 4 evaluates these trained models with camera, color, and background distractions of different intensities. As expected, all methods lose performance with increasing distraction intensity, but the robustness to these distractions varies with the distraction type and method. All methods cope best with color distractions (b), less well with camera pose distractions (a), and are highly sensitive to unseen backgrounds even when blended with the skybox background (c, visualized in Fig. 3). The points where the top methods lose half of their score are at camera scale βcam=0.2\beta_{\text{cam}}=0.2 (corresponding to camera views in column 3 of Fig. 1), at color scale βrgb=0.6\beta_{\text{rgb}}=0.6 (corresponding to color changes in column 7 of Fig. 1), and at a background weight βbg<0.1\beta_{\text{bg}}<0.1, which corresponds to column 2 in Fig. 3. It seems to be irrelevant if the distractions are dynamic or static over an episode (dashed vs. solid lines). Interestingly, SAC-based methods appear more robust to color distractions than QT-Opt-based methods (b).

V-B Training with Distractions

In this section, we apply distractions during both training and evaluation. For the background, we vary the number of background videos during training, using the fully opaque distracting background (βbg=1\beta_{\text{bg}}=1). Here we also look at generalization to unseen backgrounds during evaluation using the 30 videos from the test split of DAVIS 2017.

The results are shown in Figure 5. As before, performance drops with increasing distraction scale, which indicates how challenging it is for the agent to learn effectively in the presence of distractions. Training with distractions improves performance compared to the previous experiments for camera distractions and especially for background distractions, but not for color distractions (compare Fig. 4a,b,c and Fig. 5a,b,c and note that for backgrounds the mixture weight is 1 in Fig. 5c,d). For background distractions, we can see that with more different training videos, the performance with these same videos decreases (Fig. 5c), while the performance on unseen videos increases and then levels off (see Fig. 5d).

Compared to the previous experiment, the static / dynamic setting appears to make a difference when training with distractions to camera pose and background. The dynamic setting (i.e. with moving cameras and video backgrounds, dashed lines) produces higher scores than the static setting (solid lines). This might result from allowing the agent to see a larger variety of distractions during training, i.e. a different distraction instance per frame instead of per episode.

And contrary to the previous experiment, DrQ-based approaches are consistently outperforming SAC-based ones in all settings when training with distractions (compare blue/green to orange/pink lines in Fig. 5).

V-C A New Benchmark for Control from Pixels

Here we combine all three distraction types. We envision this combined setting as a new benchmark for pixel-based RL that measures the ability to extract task-relevant information from visual input in the presence of visual distractions. To provide a set of competitive baselines for this benchmark, we evaluate the different combinations of SAC and QT-Opt with RAD and DrQ on this benchmark.

To decide on the right values for the severity of distractions in the benchmark, we conducted experiments to generate an “easy” and “medium” difficulty for the tested methods. We also added a “blind” baseline to estimate lower bound of the performance in these tasks without seeing the relevant objects. In the easy setting, we use βcam=βrgb=0.1\beta_{\text{cam}}=\beta_{\text{rgb}}=0.1, and b=4b=4 background videos. In the medium setting we use βcam=βrgb=0.2\beta_{\text{cam}}=\beta_{\text{rgb}}=0.2 and b=8b=8. In the blind setting, we use the same parameters as the medium setting, but turn the camera backwards so that it cannot see any task-relevant information. All experiments are run with static as well as with dynamic distractions.

Figures VI & VI show the average results across all tasks with no, easy, and medium distractions and for the blind benchmark. Detailed benchmark results can be found in Tables II, VI, and VI. The observations from these results are: 1) Sensitivity to distractions is task-dependent. The cheetah and walker tasks receive lower scores than ball in cup, cartpole, finger spin or reacher tasks. In the easy benchmark, the finger spin task works better in the static than in the dynamic setting, but for the reacher task it is flipped. 2) The performance degradation in these benchmarks is larger than the product of the individual performance reductions with the same parameters shown in Figure 5. Table VII shows relative performance per distraction and reveals that their product is generally above the actual benchmark performance, which is also visualized in Figures VI & VI. We find that the distractors have a compounding effect: combined, the distractors degrade performance more than individually. This outcome is stronger in the dynamic than in the static setting. 3) In the medium benchmark, the static setting appears to be easier than the dynamic setting, where current methods only barely outperform the blind baseline experiment. Combined the easy and medium benchmarks should be a good metric for future research as they provide a lot of room for improvement, but still allow current methods to learn some meaningful behaviors. 4) The ranking of methods changes in the easy and medium benchmarks vs. no distractions, as QT-Opt methods now significantly outperform SAC-based methods. Random cropping is still essential to improve performance but does not “solve” these settings.

VI Conclusion

We have presented the Distracting Control Suite, a new benchmark for pixel-based control in the presence of different types of visual distractions. We found that these distractions are challenging for current methods, especially when multiple distractions are applied at the same time. Between the methods that we compared, we found that random cropping was essential for good performance but DrQ did not outperform the simpler RAD approach. We also found that while SAC-based and QT-Opt-based methods perform similarly on the original DM Control benchmark, QT-Opt-based methods perform better in the presence of distractions, indicating that prior work on simpler environments might not transfer to more realistic settings. We hope that our benchmarkCode is available at https://github.com/google-research/google-research/tree/master/distracting_control and analysis will facilitate progress towards algorithms that can efficiently handle the visual complexities of the real world.

References