When will you do what? - Anticipating Temporal Occurrences of Activities

Yazan Abu Farha, Alexander Richard, Juergen Gall

Introduction

In the last years, we have seen a tremendous progress in the capabilities of computer systems to classify and segment activities in videos, e.g. . These systems, however, analyze the past or in the case of real-time systems the present with a delay of a few milliseconds. For applications, where a moving system has to react or interact with humans, this is insufficient. For instance, collaborative robots that work closely with humans have to anticipate the activities of a human in the future. In contrast to humans that are very good in anticipating activities, developing methods that anticipate future activities from video data is very challenging and has just recently received an increase of interest.

Current works anticipate activities only for a very short time horizon of a few seconds. While early activity detection addresses the problem of inferring the class label of an action at the point when the activity starts or shortly thereafter , other works predict the class label of the action that will happen next . In the recent work , the starting time of the future activity is estimated as well.

In this work, we go beyond the recognition of an ongoing activity or the anticipation of the next activity. We address the problem of anticipating all activities that will be happening within a time horizon of up to 5 minutes. This includes the classes and order of the activities that will occur as well as when each activity will start and end. Figure 6 shows a few example predictions.

To address this problem, we propose two novel approaches for this task. In both cases, we first infer the activities from the observed part of the video using an RNN-HMM . The first approach builds on a recurrent neural network (RNN) that predicts for a given sequence of inferred activities the remaining duration of the ongoing activity as well as the duration and class of the next activity. The anticipated activities are then fed back to the RNN in order to anticipate activities for a longer time horizon. The second approach builds on a convolutional neural network (CNN). To this end, we convert the temporal sequence of inferred activities in a matrix that encodes both the length and the action label information. The CNN then predicts a matrix that encodes the length and the action labels of the anticipated activities. In contrast to the RNN approach, the CNN approach anticipates all activities in one pass.

We have evaluated the two approaches on two challenging datasets that contain long sequences and large variations. Both approaches outperform by a large margin two baselines, a grammar based baseline and a nearest neighbor baseline. While the RNN and CNN perform similarly for a long time horizon of more than 40 seconds, the RNN performs better for shorter time horizons less than 20 seconds. Both approaches also outperform the method of that does not anticipate activities directly but visual representations of future frames, which can then be used to classify the activities.

Related Work

Predicting future frames, poses or image segmentations in videos has been studied in several works . Approaches that predict future frames at a pixel-level, however, are limited to a few frames. Instead of predicting frames, a deep network is trained in to predict visual representations of future frames. The predicted visual representations can then be used to classify actions or objects using standard classifiers. This approach, however, is also limited to a very short time horizon of only 5 seconds. This work has been extended in , where they use an encoder-decoder network to predict a sequence of future representations based on a history of previous frames. To get the predicted actions, the output of the encoder-decoder network is passed through another network that generates the action labels.

The task of early activity detection is also related, but it assumes that a partial observation of the ongoing activity is available, and the goal is to recognize this activity with the least possible amount of observations . Recent approaches for this task use Long Short-Term Memory (LSTM) networks with special loss functions that encourage early activity detection .

A slightly longer time horizon is considered for approaches that predict the next action that will happen. predict future actions from hierarchical representations of short clips or still images. They encode the observed frames in a multi-granular hierarchical representation of movements, and train different classifiers at each level in the hierarchy in a max-margin framework. Recently, train a deep network to predict the future activity and its starting time from a sequence of preceding activities. Their model relies on appearance based and motion based features extracted from the observed activities to predict what will happen next in the future. combine on-wrist motion sensing and visual observations to anticipate daily intentions. In , observed activities are modeled with spatio-temporal graphs which are used for anticipating object affordances, trajectories, and sub-activities. Besides of activities, anticipating goal states from first person videos has also been addressed. model the human behavior with a Markov Decision Process (MDP) and use inverse reinforcement learning to infer all elements of the MDP and the reward function from first person vision videos. Then, they use the learned MDP to anticipate goal states and the length of trajectories towards them. Inverse reinforcement learning is also used in for visual sequence forecasting.

Other works investigate future prediction in the sports domain. For example, use augmented hidden conditional random field to predict the location of future events in sports, like predicting shot location in tennis. Recently, a framework has been introduced in to predict the next move of players or the location of the ball in the immediate future from sports videos.

In contrast to previous work that anticipate the next activity only within a short time horizon of a couple of seconds, we address the problem of anticipating a sequence of activities including the start and end points within time horizons of up to 5 minutes.

Anticipating Activities

Our goal is to anticipate from an observed video what will happen next in the video for a given time horizon, which can take up to 5 minutes. As shown in Figure 6, we aim to predict for each frame in the future the label of the activity that will happen. More formally, let x1T=(x1,…,xT)\mathbf{x}_{1}^{T}=(x_{1},\dots,x_{T}) be a video with TT frames. Given the first tt frames x1t\mathbf{x}_{1}^{t}, the task is to predict the actions happening from frame t+1t+1 to TT. That is, we aim to assign action labels ct+1T=(ct+1,…,cT)\mathbf{c}_{t+1}^{T}=(c_{t+1},\dots,c_{T}) to each of the unobserved frames.

Instead of predicting the future actions ct+1T\mathbf{c}_{t+1}^{T} directly from the video frames x1t\mathbf{x}_{1}^{t}, we first infer the actions c1,…,ctc_{1},\dots,c_{t} for the given frames x1,…,xtx_{1},\dots,x_{t} and then predict the future actions ct+1,…,cTc_{t+1},\dots,c_{T} from the inferred actions c1t\mathbf{c}_{1}^{t} as it is illustrated in Figure 1. This has the advantage that we can separately study the impact of the network, which infers activities from observed video sequences, and the network that anticipates the future activities. In our experiments, we will also show that the predictor network performs worse if activities are directly anticipated from the observed video frames.

For inferring the activities c1t\mathbf{c}_{1}^{t} from x1t\mathbf{x}_{1}^{t}, we use a hybrid RNN-HMM approach . In contrast to , which train the method in a weakly supervised setting, we train the model fully supervised since in our training set each video frame xtx_{t} is labeled with a class ctc_{t}.

For inferring the activities ct+1T\mathbf{c}_{t+1}^{T} from c1t\mathbf{c}_{1}^{t}, we investigate two architectures. The first architecture is based on a recurrent neural network (RNN) and will be described in Section 3.2. The second architecture is based on a convolutional neural network (CNN) and will be described in Section 3.3.

The source code for both the RNN and CNN models is publicly available at https://github.com/yabufarha/anticipating-activities.

2 RNN-based Anticipation

We can interpret future action prediction as a recursive sequence prediction: As input sequence, the RNN obtains all observed segments and predicts the remainder of the last segment as well as the next segment. This is repeated until the desired amount of future frames is reached.

More precisely, for each observed segment, the RNN gets its class label in form of a 11-hot encoding and the corresponding segment length, which is normalized by the video length, as input. Sequentially forwarding all those segments, three output predictions are made: the remaining length of the last observed segment as well as a label and a length for the next segment. This prediction is concatenated with the observed segments to form a new input for the network. The new input is again forwarded through the network to produce the next prediction. The final result is obtained by repeatedly forwarding the previously generated prediction until the desired amount of frames is predicted, see Figure 2.

As RNN architecture, we use two stacked layers of 256256 gated recurrent units and fully connected layers at the input and output. As output layer for both length predictions, remaining length of current action and length of next action, we use a rectified linear unit to ensure positive length outputs. The label prediction is done via a softmax layer as usual for classification tasks.

The training data generation for the RNN is illustrated in Figure 3. Given a ground-truth labeling of a training sequence with nn action segments, n−1n-1 training examples are generated out of it. For a segment i<ni<n, a random split point is defined. Everything before this point is encoded as a sequence of ii tuples containing the length of the observed segment and its label as 11-hot encoding. Each such sequence is an input training example for the network. For segment i+1i+1, another random split point is defined. The values between the first and second split point define the target the network should predict: A triplet consisting of the remaining length of segment ii (lrl_{r}), the length of the next action i+1i+1 from its start up to the split point (lnl_{n}), and the label of the next action (cc). By processing each training sequence like this, a large amount of input tuple sequences and target triplets is generated.

As loss for a single training example, we use

where l^r\hat{l}_{r} denotes the predicted remaining length of the current action, l^n\hat{l}_{n} denotes the predicted length of the next action, and p^c\hat{p}_{c} the predicted class probability of the next action. For training, we minimize the loss, which is summed over all training examples, by backpropagation through time.

3 CNN-based Anticipation

The CNN-based anticipation approach aims at predicting all actions directly in one single step, rather than relying on a recursive strategy such as the RNN. The given framewise labels c1t\mathbf{c}_{1}^{t} are encoded in a matrix XX with CC columns and SS rows. While the columns correspond to the CC action classes, the rows correspond to action segments. The number of rows for an action segment of length ll is given by ⌊ltS⌋\lfloor\frac{l}{t}S\rfloor and, for each row ss, Xsc=1X_{sc}=1 for the label cc of the corresponding action segment and zero otherwise. The matrix is filled in the order the actions occur as illustrated in Figure 4. We set SS large enough such that each action segment covers at least one row.

and concatenate c^s\hat{c}_{s} where each row corresponds to ⌊T−tS⌋\lfloor\frac{T-t}{S}\rfloor frames.

Since the CNN approach predicts all actions directly while the RNN uses a recursive strategy, we also have to prepare the training data slightly differently. For each video in the training set, we generate 44 training examples by using the first 10%10\%, 20%20\%, 30%30\%, and 50%50\% of the video, respectively, as observation and the following 50%50\% of the video as ground-truth for the prediction. For each training example, we convert the sequence of action labels c1t\mathbf{c}_{1}^{t} into the matrix XX and the labels of the following 50%50\% of the video frames into the ground-truth matrix YY. To train the network, we use the squared error criterion over all output elements

Experiments

We conduct our experiments on two benchmark datasets for action recognition. The Breakfast dataset contains 1,7121,712 videos of 5252 different actors making breakfast. Overall, there are 4848 fine-grained action classes and about 66 action instances for each video. The average duration of the videos is 2.32.3 minutes. We use the four splits as proposed in .

The 50Salads dataset contains 5050 videos with 1717 fine-grained action classes. With an average length of 6.46.4 minutes, the videos are comparably long and contain 2020 action instances per video on average. Following , we use five-fold cross-validation.

The longest video in both datasets is 10 minutes. As evaluation metric, we report the accuracy of the predicted frames as mean over classes (MoC).

Parameters.

For the CNN approach, we set the number of rows SS of the matrix XX to 128 for Breakfast. This ensures that each action segment covers at least one row. Since the average video length for 50Salads is about four times larger than for Breakfast, we use 512 for 50Salads. Since increasing SS and therefore sampling at a finer temporal resolution did not significantly change the results, we stick to these values for the remainder of the paper. For the post-processing, we use Gaussian smoothing with σ=3\sigma=3 for Breakfast and σ=13\sigma=13 for 50Salads. For the RNN approach, the normalized length input is scaled by the average number of actions in the videos to ensure numerical stability. In all experiments, the Adam optimizer is used with a learning rate of 0.0010.001.

Baselines.

We define two baselines for future action prediction to compare against our two proposed systems. The first baseline uses a grammar and the mean length of each action class. The mean length of each action class is estimated from the training data. The finite grammar generates all action sequences that have been observed during training. This is the same grammar as proposed in . Given the observed actions c1t\mathbf{c}_{1}^{t}, either based on the ground truth or on the action decoder , we randomly select an action sequence from the grammar that has c1t\mathbf{c}_{1}^{t} as prefix. We then predict action labels ct+1T\mathbf{c}_{t+1}^{T} such that the labels are consistent with the chosen action sequence of the grammar. Each action class from the chosen grammar path that has not been observed, i.e. that is to appear in the future, is added with its mean class length to the prediction until all required frames from t+1t+1 to TT are predicted.

As a second baseline, we use a nearest neighbor approach. We search the nearest neighbor in the training set using frame-wise error of the observed part as distance, and use the remaining part as future prediction.

2 Prediction with Ground-Truth Observations

In order to provide a fair evaluation, we first assume that the observed segmentation c1t\mathbf{c}_{1}^{t} of the video is perfect, i.e. that the observed labels do not contain any errors. While this is not the case in most realistic settings, it allows to get a clean evaluation how well the systems can predict the future. With noisy observations, the results are more delusive as errors in the observed part are propagated to the future.

In Table 1 and Figure 5, the results on Breakfast and 50Salads are shown. Both the RNN model and the CNN model show good performance compared to the baseline. Independent of the fraction of the video that has been observed, i.e. 20%20\%, or 30%30\%, however, the RNN outperforms the CNN in most cases. In general, the RNN is better for short term prediction, i.e. 10%10\%, or 20%20\%, while the CNN performs similarly or even sometimes better than the RNN for longer prediction. The reason lies in the recursive structure of the RNN predictions: once a segment is predicted, it is appended to the observed part and used as input to predict the next sequence. Consequently, if the RNN outputs an erroneous prediction at some point, this error is likely to propagate through time. The CNN, on the contrary, uses the observed part of the video to predict all future actions directly, so errors are less likely to propagate from one segment to another. However, this slightly better performance of the CNN comes with the drawback of favouring long action segments over short ones. The behaviour can also be observed in the qualitative results in Figure 6 (a) and (c). For example in the case of (c), the CNN tends to miss small action segments, whereas the RNN seems to be more reliable.

3 Prediction without Ground-Truth Observations

In this section, we evaluate the performance of our proposed systems given noisy annotations. In contrast to the clean ground-truth observations from the previous section, we now assume that the observed part of the video has been decoded using the system of and, thus, c1t\mathbf{c}_{1}^{t} is not perfect anymore but is likely to contain errors. The mean over frames accuracy of when observing 20%20\% of the video, for instance, is only 37%37\% on Breakfast and 67%67\% on 50Salads. While the accuracy when observing 30%30\% of the video is 43%43\% on Breakfast and 68%68\% on 50Salads.

Compared to the amount of noise in the observed video labeling, the prediction results are still surprisingly stable, see Table 2 and Figure 7. While the drop in performance on the Breakfast dataset (Table 1 vs. Table 2) is comparably large, both systems still achieve a good performance compared to the baseline. On 50Salads, the overall loss of accuracy compared to the system with perfect observations is surprisingly small. This can on the one hand be attributed to the better performance of the decoder . On the other hand, the inter-class dependencies on 50Salads are very strong, making it easier for both RNN and CNN to learn valid action sequences.

We would also like to put emphasis on the qualitative results for noisy observations, see Figure 6 (b) and (d). Particularly, the case of (d) is interesting: the decoder mistakenly predicted incorrect label for the last observed segment. For both models, RNN and CNN, this error propagates further to future segment predictions.

As the videos have different lengths, the proposed models might behave differently depending on the length of the videos. An evaluation on three categories of videos from Breakfast based on the prediction length is shown in Figure 8 for the case of observing 30%30\% of the video and predicting the 50%50\% that follows. While the models perform well on both short and long videos, they achieve a better performance for shorter videos. This is mainly because the duration of the predicted future is less for shorter videos, which makes the prediction task much easier compared to long video sequences. A similar behaviour can also be observed when considering the number of actions that are predicted in the future. As shown in Table 3, the accuracy drops as we predict more in the future.

4 Analysis of the CNN Model

Compared to commonly used CNN architectures such as VGG-16 or ResNet, our model is comparably small. We stick to this simple architecture since we have very limited amount of training data, and complex models would easily overfit on the training set. A comparison of our architecture to a VGG-16 on Breakfast with and without ground truth annotation is provided in Figure 5 and Figure 7, respectively. The performance of our architecture compared to VGG-16 is clearly better with an improvement up to 9%9\% when using the ground truth annotation as observations. Note that our architecture already has around 66m parameters due to the fully connected layers at the end.

Loss Function

5 Future Prediction Directly from Features

So far, we considered observations that are either frame-wise ground truth action labels, or those labels that are obtained by decoding the observed frame-wise features. In this section, we evaluate the performance of our models when applied directly on the observed frame-wise features and compare it to the two-step approach. We only use the CNN model for this evaluation. Originally, the input of the model is a matrix with CC columns that correspond to the number of classes, and SS rows that represent the temporal resolution. Each row is a 1-hot encoding that represents the action label of the corresponding frames in the observed sequence. When applying this model to features directly, CC is equal to the dimensionality of the features, which is 64 in our case for the Fisher vectors features. SS is kept at 128128 as before. The observed video features are down- or upsampled to have exactly 128 frames. We use the same training protocol that is used for the previous experiments, by considering the set {10%, 20%, 30%, 50%}\{10\%,\ 20\%,\ 30\%,\ 50\%\} as observation percentages, and the target is always the 50%50\% that follows immediately after the observations. Table 5 shows the results of the CNN model when applied on features directly compared to the two-step approach. As shown in the table, using the two-step approach outperforms the direct prediction from features by a large margin, i.e. up to +5%+5\%. Predicting the future directly from features is a harder problem since the model has to recognize the observed actions and capture the relevant information to anticipate the future, while for the two-step approach these two tasks are decoupled. This allows to use a strong decoding model to recognize the observed actions, and restricting the future predictor to capture the context over action classes only instead of frame-wise features. A similar conclusion was reached by where semantic labels of the unseen parts are predicted in the spatial domain of an image. It has been shown that segmenting the observed part of the image first and then using the segmented image for prediction achieves better results than using the RGB image directly.

6 Comparison with the State-of-the-Art

To the best of our knowledge, long term future action prediction has not been addressed before. Most works focus on predicting the immediate future. Vondrick et al. , for instance, train a model to predict AlexNet features of a frame one second in the future. Based on these predictions, an SVM is trained to determine the action label of the future frame.

Since train their future prediction model on 600h600h of videos from the web, which are not made publicly available, we use the recent Kinetics network to generate deep CNN features that have shown to generalize extremely well on several action recognition datasets and are the current state-of-the-art . We run the approach of on both, Breakfast and 50Salads, and compare their results to our model. To provide a fair comparison, our model is trained in a way that the input sequences always end one second before the next action segment starts. The results are shown in Table 6. Our approach outperforms the system of by a large margin. Note that for both datasets, predicting an action label only based on a single frame is particularly hard and the framewise action classifier achieves less than 10%10\% accuracy. In contrast to , our approaches make use of temporal context of the previously observed video content, which is crucial for reliable predictions.

Conclusion

We have introduced two efficient methods to predict future actions in videos, which is a task that has not been addressed before. While most existing prediction approaches focus on early anticipation of an already ongoing action or predict at most one action in the future, our methods are the first to predict video content of up to several minutes length. Proposing two models to address the task, an RNN and a CNN, we obtain accurate predictions that scale well along different datasets and videos with varying lengths, varying quality of observed data, and huge variations in the possible future actions.

The work has been financially supported by the DFG project GA 1927/4-1 (DFG Research Unit FOR 2535 Anticipating Human Behavior), the ERC Starting Grant ARCA (677650) and the German Academic Exchange Service (DAAD).

References