Traffic Control Gesture Recognition for Autonomous Vehicles

Julian Wiederer, Arij Bouazizi, Ulrich Kressel, Vasileios Belagiannis

I INTRODUCTION

Part of autonomous driving incorporates the vehicle interaction with humans. In urban traffic situations, the interaction engages pedestrians, school traffic patrols and traffic officers among others. The latter two examples are particularly interesting for the road traffic control. While a driver has learnt to recognise the traffic hand signals, it is not the same for the autonomous vehicle. Traffic control signals, i.e. hand gestures, need to be “taught” to the autonomous vehicle by means of learning databases. Understanding those gestures is essential for achieving proactive and safe autonomous driving.

On one hand, recent perception databases for autonomous driving, e.g. Cityscapes , ApolloScape or Eurocity Persons Dataset , contain thousands of pedestrians, road users and cyclists, however, due to the rareness of gestures they lack scenarios with human-vehicle interaction. Road traffic controllers do not exist in this kind of databases. On the other hand, gesture recognition databases include body-language, as well as human-human and human-machine interactions, but they lack of road traffic control gestures, such as stop or go. It becomes, thus, a necessity to create a public database for road traffic control gesture recognition.

In this work, we introduce the TCG dataset for road traffic control gesture recognition, targeted on autonomous vehicles. We define the gestures as a set of landmarks that belong to the general body pose, represented by a three-dimensional skeleton. The aim of the dataset is to classify the traffic control gestures in every time step from the sequential skeleton-based input. With the progress in human pose estimation the body skeleton representation has become a standard input for gesture and activity recognition . Moreover, it allows generalization to any kind of road traffic controller since it does not depend on the individual’s appearance. Capturing outdoors skeleton-based traffic control gestures is not trivial though. Motion capture on public roads is forbidden due to road obstruction. To address this limitation, we work on a closed environment where we portray road intersections with multiple vehicles and the road traffic controller involved. Our recordings include all possible traffic control scenarios for road intersections with a large amount of human motion variance. Finally, our quantitative evaluations on real-world sequences show that our studio-based recordings capture the variance of the real-world. This is the first public dataset for traffic control gesture recognition to the best of our knowledge.

Alongside with the dataset, we examine a plethora of neural network approaches for gesture recognition from sequential data. In traffic control gesture recognition, we have a sequence to sequence problem where the gesture classification happens for each input of a 3D body skeleton. This mapping is modeled with recurrent neural networks (RNNs), including attention models, temporal convolutional networks (TCNs) and graph-based networks (GCNs). In total, we examine eight different neural networks architectures, demonstrating the advantages and limitations for each model. For that reason, we provide an extensive evaluation on our dataset and real-world sequences for cross-subject and cross-view settings, using multiple metric scores. On the real-world evaluation (see Fig. 1), we demonstrate that our dataset generalizes well outdoors, although it has been captured on a closed environment.

To sum up, our work makes the following contributions: 1. The first public traffic control gesture recognition dataset for autonomous vehicles. 2. An extensive evaluation of eight sequence modelling approaches, including recurrent networks, attention mechanism, TCN and GCN models. 3. An quantitative evaluation on real-world sequences to show generalization.

II RELATED WORK

Gesture recognition for human-machine and human-human interaction is a long studied problem . Below, we discuss the related datasets and approaches to gesture recognition, where our focus is on human-vehicle interaction.

Human-vehicle interaction. Autonomous vehicles need to interact with humans inside the vehicle , e.g. driver, cyclist and passengers, as well as outside the vehicle, e.g. pedestrians and police . According to these studies, comprehensive understanding of the body language is important in order to react according to the human intentions. In particular, hand gestures are a common mean of interaction between the vehicle and human . Fortunately, the state-of-the-art on gesture recognition allows to make easily accurate predictions. However, modeling the traffic control gestures can be challenging due to the intercultural differences . For example, the traffic control hand gestures differ from country to country. In addition, road traffic control gestures are unique and they are not included in general gesture recognition datasets. In this work, we focus on the German traffic control gestures, which are also common in Europe.

Traffic control gesture recognition. Although traffic control gesture recognition becomes increasingly important in autonomous driving, the prior work on the problem is rather limited. Recently, Ma et al. have developed a spatiotemporal convolutional neural network (CNN) to spot Chinese traffic command gestures. Similarly, a long short-term memory (LSTM) network is employed in for classifying also Chinese traffic police gestures. Both approaches rely on human body skeleton input to perform the recognition. As the human body pose is in general a strong feature for activity recognition , we also build our baselines with skeleton-based input. Compared to these prior approaches, we do not only study the problem by providing a number of algorithmic solutions, motivated by general gesture recognition, but we additionally release a public database for traffic control gesture recognition.

Existing gesture recognition databases. A reason for the limited research on traffic control gesture recognition is due to the lack of public data. While there are several hand gesture databases for indoor scenarios , general gestures and for specific applications such as sign language recognition or egocentric gesture recognition ; the publicly available databases for traffic control hand gesture recognition are inexcitable. Consequently, our new public database on traffic control hand gesture supports the further research on the problem. Next, we introduce our dataset and then present the baseline algorithms for evaluation.

III TRAFFIC CONTROL GESTURE DATASET

We introduce TCG, a dataset for traffic control gesture recognition, that covers all possible road traffic control variations for European road intersections. By modeling road intersections, we automatically include the non-intersection situations as well. We consider road traffic control gesture recognition as a classification task from 3D body pose skeleton input data over time. As a result, our dataset consists of 3D human body skeleton sequences represented by joint sets and the respective label per skeleton. Below, we discuss the data collection, labelling and properties, as well as the experimental setup.

We asked from 5 individuals of different body types to regulate the traffic on road intersections. We chose a T- and X-junction where the individual makes uses of the hands for regulation, without additional control devices like whistle or traffic paddle. We also defined 5 different scenarios for each junction, with variable number of involved vehicles. Fig. 3 shows all scenarios in bird’s-eye view, while Fig. 4 presents the data distribution for all scenarios and individuals. The vehicles are specified based on their driving intention, i.e. straight, left turn or right turn, and driving order.

Since staging in real traffic situations is not permitted, we simulate the above scenarios in a closed environment, including intersection layouts, vehicles and the traffic controller. For that reason, we used colored discs to mark the streets and stopping lines. Additional colored markers were placed at the positions of the waiting vehicles to simulate the interaction partners. This helps the actors to adapt their sight according to the marker they are interacting with. In this way the setting facilitates realistic head and body orientations.

To capture the body motion of the traffic controller, each actor has been centered in the road intersection and wore an IMUInertial Measurement Unit (IMU).-based motion capture suit above the clothes. In total the suit is composed of 17 high-quality MEMSMicro Electro Mechanical Systems (MEMS). inertial sensors (accelerometer, magnetometer and gyroscope) and two pressure insoles to record smooth orientation measurements in high resolution. All sensors were synchronously sampled on 100 Hz and streamed to a computer via integrated Wi-Fi transmitter. Since the computing resources are valuable and limited in an autonomous vehicle, the sampling frequencies can not be very high. For this reason, we sub-sample to 20 Hz as a reasonable frequency for autonomous vehicles. An implemented kinematic body model computes exact 3D locations and orientations of the body joints. In total, we have a skeleton model with 17 3D body joints as it is depicted in Fig. 2. Of course, the recordings would not be easily feasible outdoors. This is the advantage of the closed environment data collection.

During recordings, a lightweight script helped the actors to keep the correct order of commands, i.e. which car needs to be stopped next and which one should proceed, while they are completely free in the duration of the commands. The script is intended as a high-level guidance rather than a detailed story line, since strong restrictions could lead to insecure and unrealistic behavior. Each scenario is repeated 5 times. In early repetitions we request the actors to perform road traffic control gestures, but after increasing the repetitions, the actors are allowed to use their own, spontaneous gestures in order to control the situation. As a starting point, all actors learn the standard European traffic control gestures, i.e. stop, go and clear. With this loose recording procedure, we achieve granularity in motion complexity, while the different actors contributed to high motion diversity.

III-B Label Definition

In autonomous driving, the perception provides the environmental state, e.g. object locations or lane markings in each time step. The next action is then planned based on the history and the current state. As a result, traffic control gesture recognition should also happen continuously. To follow this principle, we build our dataset with gesture labels per time step. We reach high annotation quality with trained annotators and consequent quality-checks.

According to and the German regulations, we differentiate three active gesture classes, go, clear and stop, as well as an inactive class. To increase the diversity of the inactive class, we actively enrich motions with daily activities, like rubbing hands, taking sunglasses on or looking at watch. Fig. 2a to 2d show examples for each class of our dataset. Additionally, the dataset provides annotation for the evaluation of transition phases, e.g. from Stop to Go. This can give insights for the decision boundaries of the gesture classifier, e.g. a gesture classifier that detects a stop gesture early in the transition phase might be a solution for autonomous driving compared to another one with larger detection latency. For the main classes, go, clear and stop, we sub-categorize the motion of the active hand, e.g. left, right or both, into static and dynamic. Fig. 5 compares the class distribution for the 5 subjects with overall 2,886 unique time intervals annotated with a major class label. The inactive class dominates over the classes as expected for real traffic situations. Table I provides quantitative insights of the label distribution. For most of the time, the go commands are indicated in a dynamic way, while road traffic controllers signal stop and clear in a more static way. Dynamic stop gestures with both hands are very rare, while dynamic go with a right waving or pointing is highly present.

III-C Dataset Properties

The dataset includes 250 unique 3D human body pose sequences, ranging from 16 to 90 seconds per sequence. We consider the directional property of gestures. This means that the gesture interpretation strongly depends on the viewpoint. For instance, a static stop gesture from one viewpoint will be a go gesture from another orthogonal viewpoint; or a dynamic go to the right does not mean any signal to the other participants. Therefore, the 3D body poses are transformed in the corresponding coordinate systems of the involved vehicles, i.e. the autonomous vehicle. On average, every sequence is transformed in 2.2 viewpoints, which results in 550 perspectives.

All sequences are recorded in high temporal resolution of 100 Hz and comprise 140 minutes of realistic human body motion, in total 839,350 frames. As shown in Fig. 4a, the amount of frames are evenly distributed on the 5 subjects. Apparently, with the complexity of the scenes, i.e. from T1 to T5 and X1 to X5, sequences become longer (Fig. 4b). The pie chart over viewpoints, Fig. 4c, shows an under-representation of vehicles coming from the lower street, since it does not appear in the T-junction layout. Based on the design of the scenarios, most of the vehicles approach from the left and right. The proposed TCG dataset can serve the community as a considerable learning base for continuous gesture recognition in the context of self-driving cars.

IV GESTURE RECOGNITION MODELS

We define hand gesture recognition as sequence modeling, where the input sequence is the track of 3D body pose skeletons x0,…,xT\mathbf{x}_{0},\dots,\mathbf{x}_{T} and the output sequence is the gesture category y0,…,yT\mathbf{y}_{0},\dots,\mathbf{y}_{T}. At each time step t∈Tt\in T, the body skeleton xt∈R3×N\mathbf{x}_{t}\in\mathcal{R}^{3\times N} is composed of NN body joints, represented as a vector. The ground-truth gesture category yt∈NK\mathbf{y}_{t}\in\mathcal{N}^{K} is an one-hot vector of KK classes. Our goal is to learn the mapping from the input skeleton to the class category from a set of training data. Without loss of generality, we represent that mapping as:

where f:R3×N×T→NK×Tf:\mathcal{R}^{3\times N\times T}\rightarrow\mathcal{N}^{K\times T} is the mapping function. We propose to approximate the mapping function based on deep neural networks. We consider recurrent, temporal convolutional and graph convolution neural networks as three different ways to approach the problem. For all network architectures, the learning goal is to minimize the difference between the predictions and ground-truth. This can be formalized by the loss function that is given by:

that is cross-entropy for problem. Finally, the training is accomplished with back-propagation and stochastic gradient descent. Note, that we do not assume access to future time steps, i.e. T+1T+1. Next, we comment on the neural network models for each architecture type.

Skeleton-based action recognition approaches traditionally make use of RNNs to model the temporal dynamics. GRU-, LSTM-cells or more complex structures, such as bidirectional networks are the common network architectures since vanilla RNNs do not capture long dependencies. In our evaluation, we consider all these types of RNNs for gesture recognition.

IV-B Attention Mechanism

Modeling long sequences can be accomplished with an attention mechanism as well. Song et al. have shown an end-to-end spatial and temporal attention model for human action recognition. The model is trained to pay more attention on discriminative joints of the skeleton within each frame and to estimate the importance of frames in the sequence. Attention has also been used for spatiotemporal attention networks to model the evolution of dynamic hand gestures . To retrieve a better semantic information, a novel model with self-attention network (SAN) was proposed by . We also examine the potential of self-attention in combination with the LSTM cells.

IV-C Temporal Convolutional Networks

Recently, it has been shown that convolutional network architectures are on par with recurrent networks on sequence modeling . At the same time, the idea of temporal convolutions has been established for visual tasks , audio generation and signal processing . We study the effect of temporal convolutions in our problems as well. The temporal convolutions process the 3D body joints, independently, over-time.

IV-D Graph Convolutional Architectures

Graph neural networks are well-suited to non-structured data such as the human body, represented by a skeleton model . Yan et al. proposed a spatiotemporal graph convolutional network to perform activity recognition from skeletal data. The skeletons are composed of 2D or 3D joint positions. We rely on the same idea to perform gesture recognition. We present a graph convolutional network that processes 3D body joints to classify traffic control activities.

V EXPERIMENTS

We evaluate the presented dataset for different sequence modelling strategies, as they have been presented in Sec. IV. The experiments include six recurrent network models, one temporal convolution network (TCN) and a spatio-temporal graph convolution network (GCN). Similar to gesture recognition approaches , the evaluation metrics are accuracy, as well as Jaccard index , F1-score and the confusion matrix. At last, we present an image-based evaluation on real-world sequences with the traffic officer and the autonomous vehicle.

We provide the implementation details for each neural network model individually. In general, all models have been trained from scratch with grid hyper-parameter search. Moreover, the activation function is non-linear, dropout is applied everywhere with rate 0.5 and the training takes place until convergence. Class confidences are computed with a dense layer and the softmax function on top of the high-level feature representations provided by the temporal models. The optimizer is the adaptive learning rate optimization algorithm (Adam) , with initial learning rate 0.001, unless it is differently reported. Below, the specific configuration for each temporal model is reported.

We consider six types of recurrent neural networks, combined with a fully connected layer and the softmax activation function to perform gesture classification. In detail, the encoder is modeled as vanilla-RNN, GRU, LSTM, bidirectional-GRU or bidirectional-LSTM. Since the sequence length varies, we adopt a masking mechanism for the input 3D body skeletons as in to overcome the zero-padding problem. For the bidirectional-LSTM, we adopt the architecture of . For the other models, our architecture is presented in Fig. 6. In all cases, we rely on 100 cells and a single hidden layer.

Attention Model

We add an attention layer on top of the LSTM encoder. In particular, we transform the LSTM to Attention-LSTM and make use of same architecture as before, however, empirically select 50 cells for the hidden layer and 50 attention units.

Temporal Convolutional Networks

We adopt to implement our TCN. We build it though simpler, because it does not deal with image data. It consists of 1D convolution kernels of size 2 and 64 filters where the dilation rate goes from 2 and reaches 32, by doubling it in each layer. The parameter optimization is accomplished with Adam, with learning rate 0.001 and back-propagation. An illustration of the model is shown in Fig. 6.

Graph Convolutional Networks

We rely on the GCN of for our problem. The body is represented as an undirected spatiotemporal graph with 17 joints and T time steps, where T=20 for making a single prediction. In the training phase we randomly sample sequences of 20 3D body skeletons of each class. During testing continuous predictions are required. Therefore we predict with a sliding window of stride 1 to obtain continuous predictions that equally compare with the other sequence modelling approaches. The initial learning rate for this model is 0.1.

V-B Cross-Subject & Cross-View Protocol

We define the cross-subject and cross-view evaluation protocol, similar to the gesture recognition approaches . Note we explicitly aim to distinguish gestures dependent on a specific viewpoint, i.e. view variant recognition. In the cross-view evaluation, this means that the labels differ depending on the vehicle’s viewpoint. As a result, the model is trained on all sequences of 3 viewpoints, e.g left, top and right, and evaluated on the omitted set of sequences, e.g. bottom. In the cross-subject evaluation, which is considered to be more challenging, the model is trained on 4 actors and tested on the remaining actor. The process is repeated for all combinations.

V-C Real-World Experiment Description

Our ultimate goal is to make use of our dataset for real-world traffic control gesture recognition. For that reason, we captured 5 image-based sequences of real traffic control scenarios. They consist of the traffic regulator on a T-junction road intersection and the autonomous vehicle. After labeling the sequences, we perform an evaluation of the presented gesture recognition approaches.

To obtain the 3D body skeleton from the image input, first, the traffic regulator is detected and the body 2D pose is extracted with a pre-trained Mask-RNN model. Since the 3D pose is necessary, we rely on the approach of Pavllo et al. to lift the 2D body poses to 3D body pose skeletons based on a sequence of 2D poses. Second, the estimated 3D body pose skeletons are provided to the gesture recognition approach for classification. We have performed this experiment off-line for being able to follow our evaluation protocol. Our approach is outlined in Fig. 1.

V-D Quantitative Evaluation

The dataset evaluation is performed for the 4-class problem, i.e. go, clear, stop and inactive. Both training and test sets come from our dataset according to the cross-subject and cross-view protocol. For the real-world evaluation, the test set is the image-based real-world sequences, as described in Sec. V-C. Since one actor of the dataset also appears in the real-world image sequences, we exclude the actor from the dataset and re-train all models. The dataset results for cross-subject, cross-view, as well as the real-world evaluation are presented in Table II. Especially in unbalanced recognition tasks, a fair metric is required to take the distribution of classes into account. For that reason we consider the Jaccard index as the most representative metric .

The three evaluation metrics have similar behaviour for the cross-subject and cross-view. The best performing approach is the LSTM in the bidirectional formulation for both cases as shown in Table II. Only, the accuracy of bidirectional-GRU is slightly higher than bidirectional-LSTM for the cross-view evaluation. The recurrent networks have in overall comparable performance other than the vanilla RNN. The temporal convolutional network has consistent results both for cross-subject and cross-view, but it is behind the recurrent models. At last, the graph convolutional network has much lower performance compared to all other models. In addition, it had difficulties to converge. We additionally provide the confusion matrices for the cross-subject (Fig. 7) and cross-view (Fig. 8) evaluation. All classifiers are able to distinct the active classes from the inactive class. Notable is the performance on go compared to stop. In most of the cases, the recognition performance is higher on the latter, which we explain with the larger amount of dynamic gestures in the go class.

Real-World

For the real-world evaluation, all metrics agree on the best performing approach as well. The LSTM model delivers the best results on the real-world sequences (see Table II), while here the bidirectional formulation does not further improve the final outcome. Next, the behavior of the temporal convolutional network is similar to the cross-subject and cross-view evaluation. In total, the real-world evaluation delivers considerable worse performance than the cross-view and cross-subject evaluation. This is expected given that the 3D body pose skeletons are algorithmically computed and thus include some sort of error.

V-E Ablation Study

We consider another classification scheme of 15-class15-Classes: inactive; stop: both-static, both-dynamic, left-static, left-dynamic, right-static, right-dynamic; clear: left-static, right-static; go: both-static, both-dynamic, left-static, left-dynamic, right-static, right-dynamic. problem. By moving from 4 to 15 gesture categories, our aim is to study how the static and dynamic gestures affect the classification performance. All experimental settings are the same with Sec. V-D except the loss function that is optimized for 15 classes. The results are reported in Table III. We report the results of all methods except the graph convolutional network because it has shown unstable convergence during training and thus reached poor performance.

The accuracy for cross-subject and cross-view is comparable to the 4-class problem. However, the jaccard index and F1-score show that the 15-class problem results in a descent performance reduction. This observation holds for all models. The best performing model is again the bidirectional-LSTM for both evaluations. The bidirectional-GRU accuracy is the best for the cross-subject evaluation, but the bidirectional-LSTM is in general on par with it. The other model have similar behavior by comparing the 4-class results of Table II with Table III.

Real-World

Unlike the cross-subject and cross-view results, the real-world performance is similar to the 4-class problem (see Table III). Considering the standard deviation, all models show great variation between runs as a result of the domain difference between the motion capture data for training and the estimated 3D poses for testing. The clear observation is that the LSTM model delivers promising performance.

VI CONCLUSION

We introduced a road traffic control gesture recognition dataset in the context of autonomous driving. Our dataset consists of 3D body skeleton data and gesture category for every time step. To perform gesture classification, we presented eight sequential processing models based on deep neural networks, such as recurrent networks, temporal convolutional networks and graph convolutional networks. Finally, we demonstrated promising performance on real-world sequences, which indicates the representativity for our dataset.

ACKNOWLEDGMENT

Part of the research was conducted within @CITY-AF (Research project No. 19 A 18003 A.), funded by BMWi (Federal Ministry for Economic Affairs and Energy).

References