Interpretable 3D Human Action Analysis with Temporal Convolutional Networks

Tae Soo Kim, Austin Reiter

Introduction

Human activity analysis is a crucial yet challenging research area of computer vision. Applications of human activity recognition ranges from video surveillance, human-computer interaction, robotics and skill evaluation . At the core of successful systems for human activity recognition lies an effective representation that can model both the spatial and temporal dynamics of human motion.

Traditionally, the community has focused on activity recognition in the domain of RGB videos . For a RGB video, complex human motion in 3D euclidean space is projected on to a series of 2D images and in the process, loss of valuable 3D spatio-temporal information is inevitable. In recent years, we have witnessed a drastic improvement of cost-effective depth sensors in the form of Microsoft Kinect . Naturally, computer vision methods leveraging on the 3D structure provided by such 3D sensors, namely RGB+D methods, have been an active area of research . Applied to human activity recognition, 3D information of how a human body articulates comes in the form of time series sequence of 3D skeletons. Such representations describe human motion as a collection of trajectories in 3D euclidean space of key human joints. Even without the context information and visual cues, early work in biological perception and more recent methods provide strong evidence that encoding humans as a 3D skeleton yields both a discriminative and a robust representation for activity analysis. Given the recent progress of powerful human pose estimators from RGB or RGBD data , human activity recognition model that builds on top of 3D skeletons is a promising direction.

Despite this significant progress, the inner workings of such complex temporal models still remain mostly black-boxes. Without the capability to interpret learning based models, we inevitably lack the power to fully support a model’s decision regardless of its correctness . Such short-comings may hinder practical deployment of even the strongest models. The ability to understand and explain precisely how a model came to a wrong prediction is a fundamental first step towards improving the potential of our current methods.

In this light, we propose Temporal Convolutional Neural Networks (TCN) applied to 3D Human Action Recognition. Through the lens of TCN, we wish to uncover what exactly learning-based temporal models leverage on especially when trained on interpretable data such as a sequence of 3D skeletons. We re-design the original TCN by factoring out the deeper layers into additive residual terms which yields both interpretable hidden representations and model parameters. Using the resulting architecture, Res-TCN, we validate our approach on currently the largest 3D human activity recognition dataset, NTU-RGBD and obtain state-of-the-art results.

Related Work

In this section, we first provide a literature review on recent developments in learning based 3D human action recognition models. We focus our narrative on models that employ LSTM-based Recurrent Neural Networks. We also extend our review to works focusing on model interpret-ability and visualization of deep learning models.

Traditional recurrent neural network models suffer from vanishing/exploding gradient problem during optimization and are difficult to train correctly . By formulation, LSTM neurons begin to address such optimization problems and are capable of modeling long-term dependencies . Given the temporal recurrent nature of human action analysis, most leading methods in 3D human action recognition adopt LSTM-based RNNs.

Hierarchical recurrent neural network of combines the features of different body parts hierarchically. At the initial layer, each sub-network extracts features over a single joint and these representations are fused hierarchically in the deeper layers. A final prediction is made when all joint information is combined. In part-aware LSTM model introduced in , individual body joints are grouped together in five groups based on their spatial context. The memory units of the LSTM are learned independently per group and the information from different parts is aggregated to produce a final prediction. The work of leverages on similar intuition that co-occurrence of joints is a strong discriminative feature for human action recognition. A group sparsity constraint on the connection matrix pushes the network to learn the mappings between co-occurring joints and the human activity. Deep spatio-temporal LSTM with Trust Gates is introduced in to learn features both in the temporal and spatial domains. Similarly, the authors of propose a spatio-temporal attention model for LSTM-based RNNs. The method comprises of three LSTM networks: a spatial attention sub-LSTM, a temporal attention sub-LSTM and a main LSTM. Both temporal and spatial attention modules are pre-trained separately initially and the entire network is trained end-to-end.

In the above methods, the key intuition is that a certain subset of joints are more important for recognizing human activities. However, it is difficult to interpret what the model parameters of each LSTM layer represent. In our proposed version of TCN, we show that our model also learns both spatial and temporal attention without the need for initial pre-training stage as in . Moreover, by model design based on temporal convolutions and residual connections , we can begin to directly interpret what our model parameters and features represent.

2 Model Interpretability and Visualization

Here, we focus our discussion on interpretability of supervised machine learning models. Post-hoc approaches are often considered to provide interpretation of models. This means that once a model is learned, post-operative experiments are conducted to gain insight into what the model has learned. proposes a method to find the optimal stimulus for each unit in a deep neural network by performing back propagation with respect to image space to maximally stimulate a neuron. Obtained images give us an insight into the appearance of the input that a neuron is most likely to activate. The work of sheds light on what spatial context of the image the convolutional neural network (CNN) is leveraging on for image classification through saliency maps. Similarly, the authors of use a deconvnet to map the activities of intermediate layers back to the input pixel space so that inputs that maximally activates an intermediate layer can be directly visualized as an image. The methods mentioned above uncover that CNNs learn to decompose the image space into hierarchical modular patterns. However, not all visualized patterns are necessarily interpretable or understandable. In such post-hoc approaches, there is no control over how the model is optimized in the first place. Though it is valuable and interesting to expose what the model has learned after-the-fact, we wish to take a more active approach to the problem. We focus our investigation on how to improve model interpretability by design.

Another popular direction in post-hoc approaches is the use of examples and prototypes. Example-based explanations and classifiers have shown to offer a condensed view of a dataset, potentially offering a reason why a classifier came to a certain conclusion through other data points in the dataset . The work of pushes this idea further by forcing a model to produce both exemplars and criticisms. Even with such examples, the causality between model parameters and the final prediction of the model is still unclear. In our work, we strive to take a more direct approach on model interpretability. We focus on two key questions: 1. How do we interpret the representations learned using TCN and 2. How can we design a deep learning architecture that provides readily interpretable hidden representations and model parameters in the context of 3D human action analysis?

Overview of Temporal Convolutional Neural Networks

In this section, we provide a brief overview of the structure of a TCN as provided in the original paper . Note that the original TCN is designed for temporal action segmentation in video and it follows a convolutional encoder-decoder design. We adapt the encoder portion of the net for action recognition. The properties of a TCN follow those of a modern spatial Convolutional Neural Network (CNN) for recognition and segmentation tasks . The network is built from stacked units of 1-dimensional convolution followed by a non-linear activation function. The 1-dimensional convolution is across the temporal domain.

where ff is a non-linear activation function such as ReLU. The whole network is trained with back-propagation. In an attempt to further improve the interpretability of a TCN, we adopt residual connections of . In the following sections, we discuss how such skip connections and the resulting TCN architecture, namely Res-TCN, leads to improved interpretability of 3D human action recognition models.

Interpretability of TCNs with Residual Connections

The biggest road-block in interpreting current spatio-temporal models such as LSTM-based RNNs for 3D human action analysis stems from the lack of clear connection between the learned model parameters and their hidden representations. However, for TCNs, the formulation of hidden representations from its model parameters is straightforward to comprehend: activation maps are computed by convolving a learnable temporal filter across time and passing the output through a ReLU unit. In a ReLU network, after an iteration of a forward-backward pass, the network parameters are optimized such that convolution of a filter across the characteristic regions of the input more likely produces a positive value in the next iteration. We can exploit such behavior of the model to improve the model interpretability by re-formulating the TCN with residual connections .

As introduced in , skip connection with identity mapping introduces beneficial properties for network convergence even for very deep networks. We observe that such design for CNNs improves model interpretability as well given input with semantic meaning. Our Res-TCN model architecture is shown in Figure 1.

Res-TCN stacks building blocks called Residual Units as introduced in and adapts the pre-activation scheme of . Each unit in layer ll performs the following computation:

For prediction, we apply global average pooling after the last merge layer across the entire temporal sequence and attach a softmax layer with number of neurons equal to number of classes.

2 A Closer Look at Model Parameters

In a Res-TCN architecture, Equation 4 suggests that the representational power of the entire model depends heavily on producing discriminative X1X_{1} through filters in W1W_{1}. In this section, we analyze what each filter in W1W_{1} represents.

3 A Deeper Look at Model Parameters

Let us now extend our analysis to deeper layers in the model. In a Res-TCN formulation, deeper layers are factored out into residual units and an output from a residual unit is simply merged by adding to the input of the residual unit. For example, consider the hidden representation after two convolution layers:

Experiments

We validate that our approach not only leads to an interpretable representation but also to a discriminative one. We evaluate Res-TCN on 3D skeleton based human activity recognition dataset of NTU . We also provide interpretations on our model predictions based on the concepts discussed in the previous sections.

NTU RGB+D dataset is currently the largest human activity recognition dataset with full 3D skeleton annotations. It contains 56880 training videos ranging over 60 action classes. The dataset provides two train/test split paradigms: Cross-Subject (CS) and Cross-View (CV) settings. The dataset covers 40 distinct subjects with varying physical traits. In terms of camera viewpoints, three cameras are placed in three different angles:−45∘-45^{\circ}, 0∘0^{\circ} and +45∘+45^{\circ}.

Implementation Details: We follow the skeleton feature construction procedure as adapted in . However, in contrast to their feature extraction stage, we do not perform view normalization prior to feeding the features into our Res-TCN. We take the raw (X,Y,Z) values of each skeleton joint and concatenate all values to form a skeleton feature per frame. Given that there are at most two actors in the scene and there are 25 joints per skeleton, a skeleton feature per frame is a 2∗25∗3=1502*25*3=150 dimensional vector. We use the Keras deep learning framework with a TensorFlow backend . We use an initial learning rate of 0.01 and decrease the learning rate by a factor of 10 when the testing loss plateaus for more than 10 epochs. We use stochastic gradient descent with nesterov acceleration with a momentum of 0.9. L-1 regularizer with a weight of 1e−41e^{-4} is applied to all convolution layers. We use a batch size of 128. Dropout with rate 0.5 is applied after all activation layers to prevent overfitting. We perform all our experiments on a Nvidia K80 GPU. The implementation and converged model weights will be made publicly available 111https://github.com/TaeSoo-Kim/TCNActionRecognition/.

2 Why Did My Model Predict This?

Leveraging on the explainable structure of Res-TCN, we wish to provide an answer to the question: ”How/why did my model come to this conclusion?”, using only the model parameters and hidden representations as the basis for providing such an explanation.

Let us choose an arbitrary video clip from NTURGB+D. The particular sequence of skeletons that we visualize in Figure 4 is of class kicking something and is approximately 70 frames long. The output of the first block (Block-A in Figure 1) is displayed above. As discussed in section 4.3, we can trace which of the W1W_{1} filters had the largest influence on any given deeper hidden representation XnX_{n} where n>1n>1. For clarity in visualization, for each time step, we only plot the activation values that are within the top 20 percentile. Each column denotes the activation values from all filters in W4W_{4} and each row denotes the corresponding filter’s response over time. Consider the dimension in the activation map that is color coded with green in Figure 4. By following the logic described in section 4.2, we found that this particular filter produces a high positive response for translational movement of the left ankle and left hip. The yellow filter has high magnitude parameters associated with the right knee joint. And finally, the blue filter picks up signals from the right ankle and the left wrist joint. The activation map of X4X_{4} and the corresponding W1W_{1} filters tell a rather detailed and precisely timed story about the input skeleton sequence: the left ankle and hip joints first translate followed by a sudden movement of the right knee, all the while the left wrist and the right ankle undergo a swinging motion.

The bit about the swinging motion can be inferred from the relative change in the magnitude of the activation in the dimension corresponding to the blue filter. Figure 5 zooms into this particular set of filters and shows their activation magnitudes over the entire video sequence. What is very interesting here is that the activation of the filter corresponding to left ankle and hip joints is close to zero at the peak of the kicking motion. At approximately the same time step, the activation magnitude of the dimension corresponding to the right knee joint peaks. The story that the filters are explaining makes sense. The sequence description that we can interpret from the filters and their activations provides insight into why the model arrived to a certain prediction. During a kicking activity, we first step towards the target with our pivot foot, firmly plant the pivot foot (in this case, the left foot), swing the kicking foot around and step back to return to original position. Note that we focused our analysis on selected interesting dimensions of the hidden representation with significant weight magnitudes. It is important to note that all other positive dimensions also factor into the final decision of the classifier but our discussion was focused on the significant and interesting ones.

3 Comparison to Other State-of-the-Art

We focused most of our narrative on how a Res-TCN formulation yields explainable spatio-temporal representation compared to state-of-the-art LSTM-based RNN counterparts. We also validate the effectiveness of our model on producing discriminative spatio-temporal features for 3D human action analysis. We compare the performances of published methods on NTURGB+D dataset and show that we improve on the current state-of-the-art on both Cross-Subject and Cross-View settings.

Conclusion

We present a new approach to performing 3D human action analysis with a Res-TCN. We discuss how such an architecture enhances the interpretability of model parameters and features compared to other popular RNN based approaches. Given an interpretable input such as sequence of human skeletons positions, we can begin to explain what each of the learned filters in a Res-TCN are leveraging on to make a prediction. We show that the model learns to pay different levels of attention both spatially and temporally. Experimentally, we validate that our model is explainable and produces a discriminative representation for human activity analysis, improving upon the state-of-the-art.

References