Learning by Cheating
Dian Chen, Brady Zhou, Vladlen Koltun, Philipp Krähenbühl
Introduction
How should we teach autonomous systems to drive based on visual input? One family of approaches that has demonstrated promising results is imitation learning . The agent is given trajectories generated by an expert driver, along with the expert’s sensory input. The goal of learning is to produce a policy that will mimic the expert’s actions given corresponding input .
Despite impressive progress, learning vision-based urban driving by imitation remains hard. The agent is tasked with organizing the “blooming, buzzing confusion” of the visual world by correlating it with a set of actions shown in the demonstration. A recent study argues that even with tens of millions of examples, direct imitation learning does not yield satisfactory driving policies .
In this paper, we show that imitation learning for vision-based urban driving can be made much more effective by decomposing the learning process into two stages. First, we train an agent that has access to privileged information: it can directly observe the layout of the environment and the positions of other traffic participants. This privileged agent is trained to imitate the expert trajectories. In the second stage, a sensorimotor agent that has no privileged information is trained to imitate the privileged agent. The privileged agent cheats by accessing the ground-truth state of the environment for both training and deployment. The final sensorimotor agent doesn’t: it only uses visual input from legitimate sensors (a single forward-facing camera in our experiments) and does not use any privileged information.
The effectiveness of this decomposition is counter-intuitive. If direct imitation learning – from expert trajectories to vision-based driving – is hard, why is the decomposition of the learning process into two stages, both of which perform imitation, any better?
Conceptually, direct sensorimotor learning conflates two difficult tasks: learning to see and learning to act. Our procedure tackles these in turn. For the privileged agent, in the first stage, perception is solved by providing direct access to the environment’s state, and the agent can thus focus on learning to act. In the second stage, the privileged agent acts as a teacher, and provides abundant supervision to the sensorimotor student, whose primary responsibility is learning to see. This is illustrated in Figure 1.
Concretely, the decomposition provides three advantages. First, the privileged agent operates on a compact intermediate representation of the environment, and can thus learn faster and generalize better . In particular, the representation we use (a bird’s-eye view) enables simple and effective data augmentation that facilitates generalization.
Second, the trained privileged agent can provide much stronger supervision than the original expert trajectories. It can be queried from any state of the environment, not only states that were visited in the original trajectories. This enables automatic DAgger-like training in which supervision from the privileged agent is gathered adaptively via online rollouts of the sensorimotor agent . It turns passive expert trajectories into an online agent that can provide adaptive on-policy supervision.
The third advantage is that the privileged agent produced in the first stage is a “white box”, in the sense that its internal state can be examined at will. In particular, if the privileged agent is trained via conditional imitation learning , it can provide an action for each possible command (e.g., “turn left”, “turn right”) in the second stage, all at once, in any state of the environment. Thus all conditional branches of the privileged agent can train all branches of the sensorimotor agent in parallel. In every state visited during training, the sensorimotor student can in effect ask the privileged teacher “What would you do if you had to turn left here?”, “What would you do if you had to turn right here?”, etc. This is both a powerful form of data augmentation and a high-capacity learning signal.
While our training is conducted in simulation – and indeed relies on simulation in order to access privileged information during the first stage – the final sensorimotor policy does not rely on any privileged information and is not restricted to simulation. It can be transferred to the physical world using any approach from sim-to-real transfer .
We validate the presented approach via extensive experiments on the CARLA benchmark and the recent NoCrash benchmark . Our approach achieves, for the first time, 100% success rate on all tasks in the original CARLA benchmark. We also set a new record on the NoCrash benchmark, advancing the state of the art by 18 percentage points (absolute) in the hardest, dense-traffic condition. Compared to the recent state-of-the-art CILRS architecture , our approach reduces the frequency of infractions by at least an order of magnitude in most conditions.
Background. Imitation is one of the earliest approaches to learning to drive. It was pioneered by Pomerleau and developed further in subsequent work . While these earlier investigations focus on lane following and obstacle avoidance, more recent work pushes into urban driving, with nontrivial road layouts and traffic . Our work fits into this line and advances it. In particular, we substantially improve upon the state-of-the-art in recent urban driving benchmarks .
Our work builds on online (or “on-policy”) imitation learning . Sun et al. summarize this line of work and argue that the availability of an optimal oracle can substantially accelerate training. These ideas were applied to off-road racing by Pan et al. , who trained a neural network policy to imitate a classic MPC expert that had access to expensive sensors. Our work develops related ideas in the context of urban driving.
Method
The goal of our sensorimotor agent is to control an autonomous vehicle by generating steering , throttle , and braking signals in each time step. The agent’s input is a monocular RGB image from a forward-facing camera, the speed of the vehicle, and a high-level command (“follow-lane”, “turn left”, “turn right”, “go straight”) . Command input guides the agent along reproducible routes and reduces ambiguity at intersections .
Our approach first trains a privileged agent. This privileged agent gets access to a map that contains ground-truth lane information, location and status of traffic lights, and vehicles and pedestrians in its vicinity. This map is not available to the final sensorimotor agent, and is only accessible by the privileged agent. Both agents predict a series of waypoints for the vehicle to steer towards. A low-level PID controller then translates these waypoints into control commands (steering , throttle , brake ). The privileged and sensorimotor agents are illustrated schematically in Figure 2.
Training proceeds in two stages. In the first stage, we train the privileged agent from a set of expert demonstrations. Next, we train the sensorimotor agent off-policy by imitating the privileged agent on the same set of states as in the first stage, using offline behavior cloning on all command-conditioned branches. Finally, we train the sensorimotor agent on-policy, using the privileged agent as an oracle that provides adaptive on-demand supervision in any state reached by the sensorimotor student .
In the following subsections, we explain each aspect of the approach in detail.
The privileged agent sees the world through a ground-truth map , anchored at the agent’s current position. This map contains binary indicators for objects, road features, and traffic lights. See Figure 3(a) for an illustration. The task of the privileged agent is to predict waypoints that the vehicle should travel to. The agent additionally observes the speed of the vehicle and the high-level command . We parameterize this agent as a convolutional network that outputs a series of heatmaps , one for each waypoint and high-level command . We convert the heatmaps to waypoints using a soft-argmax . The convolutional network and soft-argmax are end-to-end differentiable. This representation has the advantage that the input and the intermediate output are in perfect alignment, exploiting the spatial structure of the CNN.
The privileged agent is trained using behavior cloning from a set of expert driving trajectories . For each trajectory , we store the ground-truth road map , high-level navigation command , and the agent’s velocity , position , and orientation in world coordinates. We generate the ground-truth waypoints from future locations of the agent’s vehicle , where the inverse rotation matrix rotates the relative offset into the agent’s reference frame. Given a set of ground-truth trajectories and waypoints, our training objective is to imitate the training trajectories as well as possible, by minimizing the distance between the future waypoints and the agent’s predictions:
The training data , , , and are sampled from an offline dataset of expert driving trajectories .
In prior imitation learning approaches, data augmentation, namely multiple camera angles or trajectory noise injection , was a crucial ingredient. In our setup, the driving trajectories are noise-free. We simulate trajectory noise by shifting and rotating the ground-truth map , and propagating the same geometric transformations to the waypoints . This is illustrated in Figure 3(b). The agent is thus placed in a variety of perturbed configurations (e.g., facing the sidewalk or the opposite lane) and learns to find its way back onto the road by predicting waypoints that lie in the correct lane. Random rotation and shifting of the map mimic both the multi-camera augmentations and the trajectory perturbations used in other imitation learning setups. The augmentation is completely offline and does not require any modifications to the data collection procedure or the expert trajectory.
2 Sensorimotor agent
The sensorimotor agent is trained to imitate the privileged agent using an loss:
where is a dataset of corresponding road maps , images , and velocities .
This stage has two major advantages. First, sampling is no longer restricted to the offline trajectories provided by the original expert. In particular, the learning algorithm can sample states adaptively by rolling out the sensorimotor agent during training . The second advantage is that the sensorimotor agent can be supervised on all its waypoints and across all commands at once. Both of these capabilities provide a significant boost to the driving performance of the resulting sensorimotor agent.
3 Low-level controller
The privileged and sensorimotor agents both rely on a low-level controller to translate waypoints into driving commands. Given a set of waypoints predicted or projected into the vehicle’s coordinate frame, the goal of this controller is to produce steering, throttle, and braking commands (, , and , respectively). We use two independent PID controllers for this purpose.
A longitudinal PID controller tries to match a target velocity as closely as possible. This target velocity is the average velocity that the vehicle needs to pass through all waypoints:
where is the temporal spacing between waypoints and . The longitudinal PID controller then computes the throttle to minimize the error , where is the current speed of the vehicle. We ignore negative throttle commands, and only brake if the predicted velocity is below some threshold . We use .
A lateral PID controller tries to match a target steering angle . Instead of directly steering toward one of the predicted points, we first fit an arc to all waypoints and steer towards a point on the arc, as shown in Figure 4. This averages out prediction error in individual waypoints. Specifically, we fit a parametrized circular arc to all waypoints using least-squares fitting. We then steer towards a point on the arc. The target steering angle is . The point is a projection of one of the predicted waypoints onto the arc. We use for the straight and follow-the-road commands, for right turn, and for left turn. Later waypoints allow for a larger turning radius. These hyperparameters and all parameters of the PID controllers were tuned using a subset of the training routes.
Implementation details
Network architecture. We use a lightweight keypoint detection architecture to predict waypoints. The privileged agent uses a randomly initialized ResNet-18 backbone, while the sensorimotor agent uses a ResNet-34 backbone pretrained on ImageNet . Both architectures use three up-convolutional layers to produce an output feature map. Each up-convolutional layer additionally sees the vehicle velocity as an input. The network predicts each waypoint heatmap in a separate output channel using a convolution and a linear classifier from a shared feature map. Following prior work , the network branches into four heads, where each head produces a -channel heatmap. The branches represent one of the four high-level commands (“follow-lane”, “turn left”, “turn right”, “go straight”). A differentiable soft-argmax then converts the heatmaps into spatial coordinates.
The input resolution of the privileged agent is and the resolution of its output heatmap is . The input image is cropped such that the center of the agent’s vehicle is at the bottom of the map. During training we apply random rotation and shift augmentation to the map. We first rotate the input image by an angle of pixels uniformly at random. This shift corresponds to a m offset in the simulated world.
The sensorimotor agent sees a RGB image as input, and produces heatmaps. We use the same image augmentations as CIL , including pixel dropout, blurring, Gaussian noise, and color perturbations. The sensorimotor agent predicts waypoints in camera coordinates, which are then projected into the vehicle’s coordinate frame.
Training. We first train the sensorimotor agent on the same trajectories used to train the privileged agent. Next, we train the sensorimotor agent online via DAgger , using the privileged agent as an oracle. The second stage alone – without pre-training on the original trajectories – works equally well in the final accuracy. However, the online training with DAgger is slower than training on pre-existing trajectories, so we use the first stage to accelerate the overall training process.
We resample the data following Bastani et al. 2018. Critical states with higher loss are sampled more frequently.
Results
We perform all experiments in the open-source CARLA simulator . We train the privileged agent from trajectories of a handcrafted expert autopilot that leverages the internal state of the simulator to navigate through fine-grained hand-designed waypoints. We collect 100 training trajectories at 10 fps, which amount to about 157K frames (174K in our CARLA 0.9.6 implementation) and 4 hours of driving. Each frame contains world position , rotation , velocity , monocular RGB camera image , high-level command , and privileged information in the form of a bird’s-eye view map . High-level commands are computed using a topological graph and simulate a simple navigation system.
Training and validation frames are collected in Town1, under the four training weathers specified by the CARLA benchmark . We use Town2 for the test town evaluation, and do not train or tune any parameters on it.
Experimental setup. We evaluate the presented approach on the original CARLA benchmark (subsequently referred to as CoRL2017) and on the recent NoCrash benchmark . At each frame, agents receive a monocular RGB image , velocity , and a high-level command to compute steering , throttle , and brake , in order to navigate to the specified goals. Agents are evaluated in an urban driving setting, with intersections and traffic lights. The CoRL2017 benchmark consists of four driving conditions, each with predefined navigation routes. The four driving conditions are: driving straight, driving with one turn, full navigation with multiple turns, and the same full navigation routes but with traffic. A trial on a given route is considered successful if the agent reaches the goal within a certain time limit. The time limit corresponds to the amount of time needed to drive the route at a cruising speed of 10 km/h. In the CoRL2017 benchmark, collisions and red light violations do not count as failures. We thus conduct a separate infraction analysis to examine the behavior of the agents in more detail.
The NoCrash benchmark consists of three driving conditions, each on a shared set of predefined routes with comparable difficulty to the full navigation condition in the CoRL2017 benchmark. The three conditions differ in the presence of traffic: no traffic, regular traffic, and dense traffic, respectively. As in CoRL2017, a trial is considered successful if the agent reaches the goal within a given time limit. The time limit corresponds to the amount of time needed to drive the route at a cruising speed of 5 km/h. In addition, in NoCrash, a trial is considered a failure if a collision above a preset threshold occurs.
Both benchmarks are evaluated under six weather conditions, four of which were seen during training and the other two only used at test time. The training weathers are “Clear noon”, “Clear noon after rain”, “Heavy raining noon”, and “Clear sunset”. For CoRL2017, the test weathers are “Cloudy noon after rain” and “Soft raining sunset” . For NoCrash, the test weathers are “After rain sunset” and “Soft raining sunset”. We train a single agent for all conditions.
The CARLA simulator underwent a significant revision in version 0.9.6, including an update of the rendering engine and pedestrian logic. This makes CARLA 0.9.5 and prior versions not comparable to the current CARLA versions. We thus compare to all prior work on the older CARLA 0.9.5, but also provide numbers on the newer 0.9.6 version for future reference. For a fair comparison, we reran the current state-of-the-art CILRS model on CARLA 0.9.5, but were unable to run other methods due to the lack of open implementations or support for newer versions of CARLA. See the supplement for more detail on the effect of the simulator version on the benchmark.
Ablation study. Table 1 compares our full learning-by-cheating (LBC) approach to simpler baselines on the CoRL2017 benchmark (“navigation” condition).
“Direct” one-stage training of the sensorimotor policy by imitation of the autopilot expert does not perform well in our experiments. This is in part due to the lack of trajectory augmentations in our training data. Prior direct imitation approaches heavily relied on trajectory noise injection during training.
Vanilla two-stage training suffers from the same issues and does not perform better than a direct supervision. However, simply supervising all conditional branches during training (“white-box”) significantly increases the performance of the model. White-box supervision appears to function as powerful data augmentation that is extremely effective even in the off-policy setting. Consider an intersection with left and right turns. In traditional imitation learning, the student only receives gradients for the branch taken by the supervising agent, and only on the supervising agent’s near-perfect trajectory. With white-box supervision, the sensorimotor student receives supervision on what it would need to do even if it suddenly had to turn right in the middle of a left turn. Multi-branch training also helps the sensorimotor model decorrelate its outputs across branches, as it gets to see multiple different signals for each training example.
“On-policy” refers to the sensorimotor agent rolling out its own policy during training. Training with both white-box multi-branch supervision and student rollouts (bottom row in Table 1) yields the best results and achieves 100% success rate in all conditions. We use this setting in all experiments that follow.
Comparison to the state of the art. Tables 3 and 6 compare the performance of our final sensorimotor agent to the state of the art on the CoRL2017 and NoCrash benchmarks, respectively. We substantially outperform the prior state of the art on both benchmarks. On CoRL2017, we achieve 100% success rate on all routes in the full-generalization setting (new town, new weather). On NoCrash, we outperform the recent CILRS model by significant factors, achieving 100% success rate without traffic and reaching 85% success rate or higher in all conditions.
Infraction analysis. To examine the driving behavior of the agents in further detail, we conduct an infraction analysis on all routes from the NoCrash benchmark in CARLA 0.9.5. We compare the presented approach (LBC) with the previous state of the art (CILRS ). We measure the average number of traffic light violations (i.e., running a red light) and collisions per 10 km. The results are summarized in Figure 5. Our approach cuts the frequency of infractions by at least an order of magnitude in most conditions.
Conclusion
We showed that imitation learning for vision-based urban driving can be made much more effective by decomposing the learning process into two stages: first training a privileged (“cheating”) agent and then using this privileged agent as a teacher to train a purely vision-based system. This decomposition partially decouples learning to act from learning to see and has a number of advantages. We have validated these advantages experimentally and have used the presented approach to train a vision-based urban driving system that substantially outperforms the state of the art on standard benchmarks.
Our training procedure leverages simulation, and indeed highlights certain benefits of simulation. (“Cheating” by accessing the ground-truth state of the environment is difficult in the physical world.) However, the procedure yields genuine vision-based driving systems that are not tied to simulation in any way. They can be transferred to the physical world using any procedure for sim-to-real transfer . We leave such demonstration to future work and hope that the presented ideas will serve as a powerful shortcut en route to safe and robust autonomous driving systems.
Another exciting opportunity for future work is to combine the presented ideas with reinforcement learning and train systems that exceed the capabilities of the expert that provides the initial demonstrations .
Our implementation and benchmark results are available at https://github.com/dianchen96/LearningByCheating.
We acknowledge the Texas Advanced Computing Center (TACC) for providing computing resources. This work has been supported in part by the National Science Foundation under grants IIS-1845485 and CNS-1414082. We thank Xingyi Zhou for constructive feedback on the network architecture, and Felipe Codevilla for image augmentation code.
References
Appendix A Additional details
Data collection. To train the privileged agent, we use 157K training frames and 39K validation frames collected at 10 fps by a hand-crafted autopilot. We use 174K training frames in our 0.9.6 implementation. For both our offline dataset collection and privileged rollouts, we collect the frames using four training weather conditions uniformly sampled in the training town. We add other vehicles to share the traffic with the ego-vehicle. We add pedestrians in our 0.9.6 implementation.
Hyperparameters. We use the Adam optimizer with initial learning rate and no weight decay to train all our models. We use batch size to train all of our models in 0.9.5, batch size to train the privileged model in 0.9.6, and batch size to train the sensorimotor model in 0.9.6. We used the batch augmentation trick with for our image model training in 0.9.6. For the spatial argmax layers in both privileged and sensorimotor agent, we fix temperature instead of a learnable parameter. We use PyTorch 1.0 to train and evaluate our models.
Image model warm-up. Since a randomly initialized network returns the canvas center at the end of the spatial argmax layer in the image coordinate, it corresponds to infinite distance when projected. This causes gradients to explode in the backward pass. To address this issue, we warm up our image model by first supervising it with loss in the projected image coordinate space for 1K iterations before the two-stage training.
Appendix B Additional experiments
If an agent can perform near perfectly using a map representation, why not simply try to predict the map representation from raw pixels, then act on that? This approach resembles that of Müller et al. 2018, where perception and control are explicitly decoupled and trained separately.
We train two networks. The first network is used for perception and directly predicts the privileged representation from an RGB image. We resize the RGB image to and feed this into a ResNet34 backbone , followed by five layers of bilinear-upsampling + convolution + ReLU to produce a map of the original resolution of . The perception network is trained using an L1 loss between the network’s output, and the ground truth privileged representation. The second network is used for action and predicts waypoints from the output of the first network. We use the same architecture and training procedure as the privileged agent in our main experiments, and we freeze the weights of the perception network during training.
We use a dataset of 150K frames collected from an expert in training conditions, with no trajectory noise. The offline map predictions on the training and validation sets are quite good, but we notice that during evaluation, even slightly out-of-distribution observations produce erroneous map predictions, causing the waypoint network to fail. To address this, we collect another dataset of the same size and employ trajectory noise in 20% of the frames, to broaden the states seen by the perception network. Table 4 shows the results. Both map prediction agents performs significantly worse than our two-stage agent.
Appendix C Benchmark results
For completeness, Table 5 and Table 6 show the training town performance in the CARLA CoRL 2017 and NoCrash benchmarks, respectively. We again compare to MP , CIL , CAL , CIRL , and CILRS .
Table 7 compares different CARLA versions, 0.8 and 0.9.5. We compare the performance of CILRS on the Navigation Dynamic task with and without a working pedestrian autopilot (versions 0.8 and 0.9.5, respectively). The CILRS performance in 0.9.5 matches the older CARLA version in test weathers and is slightly lower in the training weathers. This indicates that CARLA 0.9.5 does not make the task easier. We report the higher numbers from the CILRS paper .