Autonomous Braking System via Deep Reinforcement Learning

Hyunmin Chae, Chang Mook Kang, ByeoungDo Kim, Jaekyum Kim, Chung Choo Chung, Jun Won Choi

I INTRODUCTION

Safety is one of top priorities that should be pursued in realizing fully autonomous driving vehicles. For safe autonomous driving, autonomous vehicles should perceive the environments using the sensors and control the vehicle to travel to the destination without any accidents. Since it is inevitable for an autonomous vehicle to encounter with unexpected and risky situations, it is critical to develop the reliable autonomous control systems that can cope well with such uncertainties. Recently, several safety systems including collision avoidance, pedestrian detection, and front collision warning (FCW) have been proposed to enhance the safety of the autonomous vehicle .

One critical component for enabling safe autonomous driving is the autonomous braking systems which can reduce the velocity of the vehicle automatically when a threatening obstacle is detected. The autonomous braking should offer safe and comfortable brake control without exhibiting too early or too late braking. Most conventional autonomous braking systems are rule-based, which designate the specific brake control protocol for each different situation. Unfortunately, this approach is limited in handling all scenarios that can happen in real roads. Hence, the intelligent braking system should be developed to avoid the accidents in a principled and goal-oriented manner.

Recently, interest in machine learning has explosively grown up with the rise of parallel computing technology and a large amount of training data. In particular, the success of deep neural network (DNN) technique led the researchers to investigate the application of machine learning for autonomous driving. The DNN has been applied to autonomous driving from camera-based perception to end-to-end approach which learns mapping from the sensing to the control . Reinforcement learning (RL) technique has also been improved significantly as DNN was adopted. The technique, called deep reinforcement learning (DRL), has shown to perform reasonably well for various challenging robotics and control problems. In , the DRL technique called Deep Q-network (DQN) was proposed, which approximates Q-value function using DNN. It was shown that the DQN can outperform human experts in various Atari video games. Recently, the DRL is applied to control systems for autonomous driving vehicle in .

In this paper, we propose a new autonomous braking system based DRL, which can intelligently control the velocity of the vehicle in situations where collision is expected if no action is taken. The proposed autonomous braking system is described in Fig. 1. The agent (vehicle) interacts with the uncertain environment where the position of the obstacle could change in time and thus the risk of collision at each time step varies as well. The agent receives the information of the obstacle’s position using the sensors and adapts the brake control to the state change such that the chance of accident is minimized.

In our work, we design the autonomous braking system for the urban road scenario where a vehicle faces a pedestrian who crosses the street at a random timing. In order to find the desirable brake action for the given pedestrian’s location and vehicle’s speed, we need to allocate appropriate reward function for each state-action pair. In our work, we focus on finding the desirable reward function which strikes the balance between the penalty imposed to the agent when accident happens and the reward obtained when the vehicle quickly gets out of risk. Using the reward function we carefully designed, we train DQN to learn the policy that decides the timing of brake based on the given pedestrian’s state. We also provide a new DQN design which can rapidly learn the policy to avoid rare accidents.

Via computer simulations, we evaluate the performance of the proposed autonomous braking system. In simulations, we consider the uncertainty of the vehicle’s initial velocity, pedestrian’s initial position, and whether the pedestrian will cross or not. The experimental results show that the proposed braking system exhibits desirable control behavior for various test scenarios including autonomous emergency braking (AEB) test administrated by Euro NCAP.

The rest of this paper is organized as follows. In Section II, we describe the basic scenarios and the framework of the proposed system. In Section III, we provide the details of the DQN design for autonomous braking. The experimental results are provided in Section IV and the paper is concluded in Section V.

II System Description

In this section, we describe the overall structure of the autonomous braking system. We first define the possible scenarios for autonomous braking and explain the detailed operation of the proposed system.

One of the factors that hinders safe driving in autonomous driving is the threat from nearby objects, e.g. pedestrians. Many accidents could happen when the vehicle fails to stop ahead of it when a pedestrian crosses the road. Hence, in order to avoid accidents, the vehicle should detect the threat that can potentially cause accidents in advance and perform appropriate brake actions to stop vehicle in front of the obstacle. However, there exist various degrees of uncertainty which make the design of autonomous braking challenging such as

Even if a pedestrian is detected accurately, it is hard to know when it can become a threat to the vehicle. Hence we need appropriate braking strategy for different situations. (see Fig. 2.) That is, for the given state of the pedestrian (i.e. position, velocity), the autonomous braking system should decide what brake action to apply.

In our system, we consider the scenario where behavior of the pedestrian follows the discrete-state Markov process described in Fig. 3. The state SnobodyS_{nobody} implies that the sensors have not detected any obstacle. Once a pedestrian is detected, the state SnobodyS_{nobody} can change to the state SstayS_{stay} or the state ScrossS_{cross}, where SstayS_{stay} is the state that the pedestrian stays at sidewalk and ScrossS_{cross} is the state that the pedestrian crosses the road. The pedestrian’s initial position can be either from far-side and near-side of the vehicle and the pedestrian walking speed can vary between vpedmin m/s{v_{ped}}^{min}\ m/s and vpedmax m/s{v_{ped}}^{max}\ m/s. Note that the vehicle’s initial velocity is distributed between vvehmin m/s{v_{veh}}^{min}\ m/s and vvehmax m/s{v_{veh}}^{max}\ m/s. In practical scenarios, it is difficult to know the transition probabilities of the Markov process and the distribution of the pedestrian’s states. Therefore, reinforcement learning approach can be applied to learn the brake control policy through the interaction with environment.

II-B Autonomous Braking System

The detailed operation of the proposed autonomous braking system is depicted in Fig. 4. The vehicle is moving at speed vvehv_{veh} from the position (vehposx,vehposy)(vehpos_{x},vehpos_{y}). As soon as a pedestrian is detected, the autonomous braking system receives the relative position of the pedestrian, i.e., (pedposx−vehposx,pedposy−vehposy)(pedpos_{x}-vehpos_{x},pedpos_{y}-vehpos_{y}) from the sensor measurements where (pedposx,pedposy)(pedpos_{x},pedpos_{y}) is the location of the pedestrian. Using the vehicle’s velocity vvehv_{veh} and the relative position (pedposx−vehposx,pedposy−vehposy)(pedpos_{x}-vehpos_{x},pedpos_{y}-vehpos_{y}), the vehicle decides whether it will step brake at each time step. The interval between consecutive time steps is given by ΔT\Delta T. We consider four brake actions; no braking anothing{a_{nothing}} and braking ahigh,amida_{high},a_{mid} and alowa_{low} with different intensities. We can include more brake actions with more refined steps or continuous brake action which are not considered in this work.

III Deep Reinforcement Learning for Autonomous Braking System

In this section, we present the details of the proposed DRL-based autonomous braking system. We first introduce the structure of the DQN and explain the reward function used to train the DQN in details.

Our system follows the basic RL structure. The agent performs an action At\it{A_{t}} given state St\it{S_{t}} under policy π\pi. The agent receives the state as feedback from the environment and gets the reward rtr_{t} for the action taken. The state feedback that the agent takes from sensors consists of the velocity of the vehicle vveh{v_{veh}} and the relative position to the pedestrian, (pedposx−vehposx,pedposy−vehposy)(pedpos_{x}-vehpos_{x},pedpos_{y}-vehpos_{y}) for the past nn time steps. Possible action that agent can choose is among deceleration ahigh,amid,alow{a_{high},a_{mid},a_{low}} and keeping the current speed anothing{a_{nothing}}. The goal of our proposed autonomous braking system is to maximize the expected accumulated reward called “value function” that will be received in the future within an episode. Using the simulations, the agent learns from interaction with environment episode-by-episode. One episode starts when a pedestrian is detected. Note that the initial position of the pedestrian and the initial velocity of the vehicle are random. The vehicle drives on a straight way based on the brake policy π{\pi}. If the distance between the vehicle and the pedestrian is less than the safety distance ll, it is considered as a collision event. (see Fig. 4.) The episode ends if at least one of the following events occurs

StopStop : the vehicle completely stops, i.e., vveh=0v_{veh}=0.

BumpBump : the vehicle passes the safety line ll when the pedestrian is crossing road.

PassPass : the vehicle passes the pedestrian without accident.

CrossCross : the pedestrian completely crosses the road and reaches the opposite side.

Once one episode ends, the next episode starts with the state of environment and the value function reset.

III-B Deep Q-Network

Q-learning is one of the popular RL methods which searches for the optimal policy in an iterative fashion . Basically, the Q-value function qπ(s,a)q_{\pi}(s,a) is defined as

for the given state ss and action aa, where rtr_{t} is the reward received at the time step t. The Q-value function is the expected sum of the future rewards which indicates how good the action aa is given the state ss under the policy of the agent π\pi. The contribution to the Q-value function decays exponentially with the discounting factor γ\gamma for the rewards with far-off future. For the given Q-value function, the greedy policy is obtained as

One can show that for the policy in (2), the following Bellman equation should hold ;

In practice, since it is hard to obtain the exact value of qπ(s,a)q_{\pi}(s,a) satisfying the Bellman equation, the Q-learning method uses the following update rule for the given one step backups StS_{t}, AtA_{t}, rt+1r_{t+1}, St+1S_{t+1};

However, when the state space is continuous, it is impossible to find the optimal value of the state-action pair q∗(s,a){q_{*}(s,a)} for all possible states. To deal with this problem, the DQN method was proposed, which approximates the state-action value function q(s,a){q(s,a)} using the DNN, i.e., q(s,a)≈qθ(s,a){q(s,a)\approx{q}_{\theta}(s,a)} where θ\theta is the parameter of the DNN . The parameter θ\theta of the DNN is then optimized to minimize the squared value of the temporal difference error δt\delta_{t}

For better convergence of the DQN, instead of estimating both q(St,At)q(S_{t},A_{t}) and q(St+1,a′)q(S_{t+1},a^{\prime}) in (5), we approximate q(St,At)q(S_{t},A_{t}) and q(St+1,a′)q(S_{t+1},a^{\prime}) using the Q-network and the target network parameterized by θ\theta and θ−\theta^{-}, respectively . The update of the target network parameter θ−\theta^{-} is done by cloning Q-network parameter θ\theta, periodically. Thus, (5) becomes

To speed up convergence further, replay memory is adopted to store a bunch of one step backups and use a part of them chosen randomly from the memory by batch size . The backups in the batch is used to calculate the loss function LL which is given by

where BreplayB_{replay} is the backups in the batch selected from replay memory. Note that the optimization of parameter θ\theta for minimizing the loss LL is done through the stochastic gradient decent method.

III-C Reward Function

Unlike video games, the reward should be appropriately defined by a system designer in autonomous braking system. As mentioned, the reward function determines the behavior of the brake control. Hence, in order to ensure the reliability of the brake control, it is crucial to use the properly defined reward function. In our model, there is conflict between two intuitive objectives for brake control; 1) collision should be avoided no matter what happens and 2) the vehicle should get out of the risky situation quickly. If it is unbalanced, the agent becomes either too conservative or reckless. Therefore, we should use the reward function which balances two conflicting objectives. Taking this into consideration, we propose the following reward function

where vtv_{t} is the velocity of the vehicle at the time step tt, deceldecel is difference between vtv_{t} and vt−1v_{t-1} and 1(x=y)\textbf{1}(x=y) has a value of 11 if the statement inside is true and otherwise. The first term −(α(pedposx−vehposx)2+β)decel-(\alpha(pedpos_{x}-vehpos_{x})^{2}+\beta)decel in the reward function prevents the agent from braking too early by giving penalty proportional to squared distance between the vehicle and pedestrian. It guides the vehicle to drive without deceleration if the pedestrian is far from the vehicle. On the other hand, the term −(ηvt2+λ)1(St=bump)-(\eta{v_{t}}^{2}+\lambda)\textbf{1}(S_{t}=bump) indicates the penalty that the agent receives when the accident occurs. Note that this penalty is a function of the vehicle’s velocity, which reflects the severe damage to the pedestrian in case of high velocity at collision. Without such dependency on the velocity, the agent would not reduce the speed in situation when the accident is not avoidable. The constants α\alpha, β\beta, η\eta and λ\lambda are the weight parameters that controls the trade-off between two objectives.

III-D Trauma Memory

As mentioned in the previous section, autonomous braking systems should learn both of the conflicting objectives. However, when we train the DQN with the reward function in (III-C), we find that the learning performance is not stable since collision events rarely happen and thus there remains only a few one-step backups associated with the collisions in the replay memory. As a result, the probability of picking such one-step backups is small and the DQN does not have enough chance to learn to avoid accidents in practical learning stage. To solve this issue, we propose so called ”trauma” memory which is used to store only the one-step backups for the rare events (e.g., collision events in our scenario). While the one step backups are randomly picked from the replay memory, some fixed number of backups associated with the collision events are randomly selected from the trauma memory and used for training together. In other words, with the trauma memory, the loss function LL is modified to

where BtraumaB_{trauma} is the backups randomly picked from trauma memory. Trauma memory persistently reminds the agent of the memory on the accidents regardless of the current policy, thus allowing the agent to learn to maintain speed and avoid collisions reliably.

IV Experiments

In this section, we evaluate the performance of the proposed autonomous braking system via computer simulations.

In simulations, we used the commercial software PreScan which models vehicle dynamics in real time . We generated the environment in order to train the DQN by simulating the random behavior of the pedestrian. In the simulations, we assume that the relative location of the pedestrian is provided to the agent. To make the system practical, we add slight measurement noise to it. In each episode, the initial position of vehicle is set to (0,0)(0,0). Time-to-collision TTCTTC is chosen according to the uniform distrubution between 1.5 s1.5\ s and 4 s4\ s. The initial velocity of the vehicle is uniformly distributed between vinitmin=2.78 m/s (10 km/s){v_{init}}^{min}=2.78\ m/s\ (10\ km/s) and vinitmax=16.67 m/s (60 km/h){v_{init}}^{max}=16.67\ m/s\ (60\ km/h). At the beginning of the episodes, the position of the pedestrian is fixed to 5∗vinit5*v_{init} meters away from the position of the vehicle. The pedestrian stands either at the far-side or at near-side of the vehicle with equal probability. The behavior of the pedestrian follows one of two scenarios below;

During training, either of two scenarios is selected with equal probability. In Scenario 1, the pedestrian starts to move when the vehicle is crossing at the “pedestrian crossing point” ptrig=(5−TTC)∗vinitp_{trig}=(5-TTC)*v_{init}. (see Fig. 4.) The safety distance ll for the pedestrian is set to 33 m. The agent chooses the brake control among ahigh=−9.8 m/s2a_{high}=-9.8\ m/s^{2}, amid=−5.9 m/s2a_{mid}=-5.9\ m/s^{2}, ahigh=−2.9 m/s2a_{high}=-2.9\ m/s^{2} and anothing=0 m/s2a_{nothing}=0\ m/s^{2} every ΔT=0.1\Delta T=0.1 second. The detailed simulation setup is summarized below.

Initial velocity of vehicle vinit∼U(2.78,16.67) m/sv_{init}\sim U(2.78,16.67)\ m/s

Velocity of pedestrian vped∼U(2,4) m/sv_{ped}\sim U(2,4)\ m/s

Initial pedestrian position pedposx=5∗vinit mpedpos_{x}=5*v_{init}\ m

Trigger point ptrig=(5−TTC)∗vinit mp_{trig}=(5-TTC)*v_{init}\ m

ahigh,amid,alow,anothing={−9.8,−5.9,−2.9,0} m/s2a_{high},a_{mid},a_{low},a_{nothing}=\{-9.8,-5.9,-2.9,0\}\ m/s^{2}

IV-B Training of DQN

The neural network used for the DQN consists of the fully-connected layers with five hidden layers. RMSProp algorithm is used to minimize the loss with learning rate μ=0.0005\mu=0.0005. The number of position data samples used as a state is set to n=5n=5. We set the size of the replay memory to 10,000 and that of the trauma memory to 1,000. We set the replay batch size to 32 and trauma batch size to 10. The summary of the DQN configurations used for our experiments is provided below;

Network architecture: fully-connected feedforwared network

Number of nodes for each layers : [15(Input layer), 100, 70, 50, 70, 100, 4(Output layer)]

RMSProp optimizer with learning rate 0.0005

Reward function: α=0.001, β=0.1, η=0.01, λ=100\alpha=0.001,\ \beta=0.1,\ \eta=0.01,\ \lambda=100

Fig. 5 provides the plot of the total accumulated rewards i.e., value function achieved for each episode when training is conducted with and without trauma memory. We observe that with trauma memory the value function converges after 2,000 episodes and high total reward is steadily attained after convergence while without trauma memory the policy does not converge and keeps fluctuating.

IV-C Test Results

Safety test was conducted for several different TTCTTC values. Collision rate is measured for 10,000 trials for each TTC value.

Table I provides the collision rate for each TTCTTC value for the test performed for Scenario 1. The agent avoids collision successfully for TTCTTC values above 1.5s. For the cases with TTCTTC values less than 1.5s, we observe some collisions. According to our analysis on the trajectory of braking actions, these are the cases where collision was not avoidable due to the high initial velocity of the vehicle even though full braking actions were applied. The agent passed the pedestrian without unnecessary stop for all cases in the Scenario 2. The detailed trajectory of the brake actions for one example case is shown in Fig. 6. Fig. 6 (a) shows the trajectory of the position of the vehicle and the pedestrian recorded every 0.10.1 s. The velocity of the agent and the brake actions applied are shown in Fig. 6 (b) and (c), respectively. The vehicle starts to decelerate about 20 m away from the pedestrian and completely stops about 5 m ahead, thereby accomplishing collision avoidance. We observe that weak braking actions are applied in the beginning part of deceleration and then strong braking actions come as the agent gets close to the pedestrian.

Fig. 7 shows how the initial position of the pedestrian and the relative distance between the pedestrian and vehicle are distributed for 1,000 trials in the scenario 1. We see that the vehicle stops around 5 m5\ m in front of the pedestrian for most of cases. This seems to be reasonable safe braking operation considering the safety distance of l=3l=3 m. Note that this distance can be adjusted by changing the reward parameters. Overall, the experimental results show that the proposed autonomous braking system exhibits consistent brake control performance for all cases considered.

IV-D Test Results for Euro NCAP AEB Pedestrian Test

Additional autonomous emergency braking (AEB) pedestrian tests are conducted. We follow the test procedure specified by Euro NCAP test protocol , . Tests are conducted for both farside (CVFA test) and nearside (CVNA test) under the velocity range between 20 to 60 km/hkm/h with 5 km/hkm/h interval. TTCTTC is set to 4 ss and the pedestrian crosses the road at 8 km/hkm/h for CVFA and 5 km/hkm/h for CVNA. Tests are scored according to the rating parameters and the metric suggested in . The proposed system passed all tests without collision and the rating scores acquired by the proposed method are shown in Table II.

V CONCLUSIONS

We have presented the new autonomous braking system based on the deep reinforcement learning. The proposed system learns an intelligent way of brake control from the experiences obtained under the simulated environment. We designed the autonomous braking systems using the DQN method with carefully designed reward function and enhanced stability of learning process by modifying the structure of the DQN. We showed through computer simulations that the proposed autonomous braking system exhibits desirable and consistent brake control behavior for various scenarios where behavior of the pedestrian is uncertain.

References