Deep Learning for Detecting Multiple Space-Time Action Tubes in Videos
Suman Saha, Gurkirt Singh, Michael Sapienza, Philip H. S. Torr, Fabio Cuzzolin
Introduction
Recent advances in object detection via convolutional neural networks (CNNs) [Girshick et al.(2014)Girshick, Donahue, Darrel, and Malik] have triggered a significant performance improvement in the state-of-the-art action detectors [Gkioxari and Malik(2015), Weinzaepfel et al.(2015)Weinzaepfel, Harchaoui, and Schmid]. However, the accuracy of these approaches is limited by their relying on unsupervised region proposal algorithms such as Selective Search [Gkioxari and Malik(2015)] or EdgeBoxes [Weinzaepfel et al.(2015)Weinzaepfel, Harchaoui, and Schmid] which, besides being resource-demanding, cannot be trained for a specific detection task and are disconnected from the overall classification objective. Moreover, these approaches are computationally expensive as they follow a multi-stage classification strategy which requires CNN fine-tuning and intensive feature extraction (at both training and test time), the caching of these features onto disk, and finally the training of a battery of one-vs-all SVMs for action classification. On large datasets such as UCF-101 [Soomro et al.(2012)Soomro, Zamir, and Shah], overall training and feature extraction takes a week using 7 Nvidia Titan X GPUs, plus one extra day for SVM training. At test time, detection is slow as features need to be extracted for each region proposal via a CNN forward pass.
To overcome these issues we propose a novel action detection framework which, instead of adopting an expensive multi-stage pipeline, takes advantage of the most recent single-stage deep learning architectures for object detection [Ren et al.(2015)Ren, He, Girshick, and Sun], in which a single CNN is trained for both detecting and classifying frame-level region proposals in an end-to-end fashion. Detected frame-level proposals are subsequently linked in time to form space-time ‘action tubes’[Gkioxari and Malik(2015)] by solving two optimisation problems via dynamic programming. We demonstrate that the proposed action detection pipeline is at least faster in training and faster in test time detection speeds as compared to [Gkioxari and Malik(2015), Weinzaepfel et al.(2015)Weinzaepfel, Harchaoui, and Schmid]. In the supplementary material, we present a comparative analysis of the training and testing time requirements of our approach with respect to [Gkioxari and Malik(2015), Weinzaepfel et al.(2015)Weinzaepfel, Harchaoui, and Schmid] on the UCF-101 [Soomro et al.(2012)Soomro, Zamir, and Shah] and J-HMDB-21 [Jhuang et al.(2013)Jhuang, Gall, Zuffi, Schmid, and Black] datasets. Moreover, our pipeline consistently outperforms previous state-of-the-art results (§ 4).
Overview of the approach. Our approach is summarised in Fig. 1. We train two pairs of Region Proposal Networks (RPN) [Ren et al.(2015)Ren, He, Girshick, and Sun] and Fast R-CNN [Girshick(2015)] detection networks - one on RGB and another on optical-flow images [Gkioxari and Malik(2015)]. For each pipeline, the RPN (b), takes as input a video frame (a), and generates a set of region proposals (c), and their associated ‘actionness’ [Chen et al.(2014)Chen, Xiong, Xu, and Corso] scoresA softmax score for a region proposal containing an action or not. . Next, a Fast R-CNN [Ren et al.(2015)Ren, He, Girshick, and Sun] detection network (d) takes as input the original video frame and a subset of the region proposals generated by the RPN, and outputs a ‘regressed’ detection box and a softmax classification score for each input proposal, indicating the probability of an action class being present within the box. To merge appearance and motion cues, we fuse (f) the softmax scores from the appearance- and motion-based detection boxes (e) (§ 3.3). We found that this strategy significantly boosts detection accuracy.
After fusing the set of detections over the entire video, we identify sequences of frame regions most likely to be associated with a single action tube. Detection boxes in a tube need to display a high score for the considered action class, as well as a significant spatial overlap for consecutive detections. Class-specific action paths (g) spanning the whole video duration are generated via a Viterbi forward-backward pass (as in [Gkioxari and Malik(2015)]). An additional second pass of dynamic programming is introduced to take care of temporal detection (h). As a result, our action tubes are not constrained to span the entire video duration, as in [Gkioxari and Malik(2015)]. Furthermore, extracting multiple paths allows our algorithm to account for multiple co-occurring instances of the same action class (see Fig. 2).
Although it makes use of existing RPN [Ren et al.(2015)Ren, He, Girshick, and Sun] and Fast R-CNN [Girshick(2015)] architectures, this work proposes a radically new approach to spatiotemporal action detection which brings them together with a novel late fusion approach and an original action tube generation mechanism to dramatically improve accuracy and detection speed. Unlike [Gkioxari and Malik(2015), Weinzaepfel et al.(2015)Weinzaepfel, Harchaoui, and Schmid], in which appearance and motion information are fused by combining fc7 features, we follow a late fusion approach [Simonyan and Zisserman(2014a)]. Our novel fusion strategy boosts the confidence scores of the detection boxes based on their spatial overlaps and their class-specific softmax scores obtained from appearance and motion based networks (§ 3.3). The 2 pass of dynamic programming, we introduce for action tube temporal trimming, contributes to a great extent to significantly improve the detection performance (§ 4).
Contributions. In summary, this work’s main contribution is a novel action detection pipeline which:
incorporates recent deep Convolutional Neural Network architectures for simultaneously predicting frame-level detection boxes and the associated action class scores (§ 3.1-3.2);
uses an original fusion strategy for merging appearance and motion cues based on the softmax probability scores and spatial overlaps of the detection bounding boxes (§ 3.3);
brings forward a two-pass dynamic programming (DP) approach for constructing space time action tubes (§ 3.4).
An extensive evaluation on the main action detection datasets demonstrates that our approach significantly outperforms the current state-of-the-art, and is 5 to 10 times faster than the main competitors at detecting actions at test time (§ 4). Thanks to our two-pass action tube generation algorithm, in contrast to most existing action classification [Wang et al.(2011)Wang, Kläser, Schmid, and Liu, Wang and Schmid(2013), Ji et al.(2013)Ji, Xu, Yang, and Yu, Karpathy et al.(2014)Karpathy, Toderici, Shetty, Leung, Sukthankar, and Fei-Fei, Simonyan and Zisserman(2014a)] and localisation [Gkioxari and Malik(2015), Weinzaepfel et al.(2015)Weinzaepfel, Harchaoui, and Schmid] approaches, our method is capable of detecting and localising multiple co-occurring action instances in temporally untrimmed videos (see Fig. 2).
Related work
Recently, inspired by the record-breaking performance of CNNs in image classification [Krizhevsky et al.(2012)Krizhevsky, Sutskever, and Hinton] and object detection from images [Girshick et al.(2014)Girshick, Donahue, Darrel, and Malik], deep learning architectures have been increasingly applied to action classification [Ji et al.(2013)Ji, Xu, Yang, and Yu, Karpathy et al.(2014)Karpathy, Toderici, Shetty, Leung, Sukthankar, and Fei-Fei, Simonyan and Zisserman(2014a)], spatial [Gkioxari and Malik(2015)] or spatio-temporal [Weinzaepfel et al.(2015)Weinzaepfel, Harchaoui, and Schmid] action localisation, and event detection [Xu et al.(2014)Xu, Yang, and Hauptmann].
The action localisation problem, in particular, can be addressed by leveraging video segmentation methods. An example is the unsupervised greedy agglomerative clustering approach of [Jain et al.(2014)Jain, Van Gemert, Jégou, Bouthemy, and Snoek], which resembles Selective Search space-time video blocks. Since [Jain et al.(2014)Jain, Van Gemert, Jégou, Bouthemy, and Snoek] does not exploit the representative power of CNN features, they fail to achieve state-of-the-art results. Soomro et al\bmvaOneDot [Soomro et al.(2015)Soomro, Idrees, and Shah] learn the contextual relations between different space-time video segments. Such ‘supervoxels’, however, may end up spanning very long time intervals, failing to localise each action instance individually. Similarly, [van Gemert et al.(2015)van Gemert, Jain, Gati, and Snoek] uses unsupervised clustering to generate a small set of bounding box-like spatio-temporal action proposals. However, since the approach in [van Gemert et al.(2015)van Gemert, Jain, Gati, and Snoek] employs dense-trajectory features [Wang et al.(2011)Wang, Kläser, Schmid, and Liu], it does not work on actions characterised by small motions [van Gemert et al.(2015)van Gemert, Jain, Gati, and Snoek].
The temporal detection of actions [Jiang et al.(2014)Jiang, Liu, Roshan Zamir, Toderici, Laptev, Shah, and Sukthankar, Gorban et al.(2015)Gorban, Idrees, Jiang, Zamir, Laptev, Shah, and Sukthankar] and gestures [Escalera et al.(2014)Escalera, Baró, Gonzalez, Bautista, Madadi, Reyes, Ponce-López, Escalante, Shotton, and Guyon] in temporally untrimmed videos has also recently attracted much interest [Yeung et al.(2015)Yeung, Russakovsky, Jin, Andriluka, Mori, and Fei-Fei, Evangelidis et al.(2014)Evangelidis, Singh, and Horaud]. Sliding window approaches have been extensively used [Laptev and Pérez(2007), Gaidon et al.(2013)Gaidon, Harchaoui, and Schmid, Tian et al.(2013)Tian, Sukthankar, and Shah, Wang et al.(2014)Wang, Qiao, and Tang]. Unlike our approach, these methods [Tian et al.(2013)Tian, Sukthankar, and Shah, Wang et al.(2014)Wang, Qiao, and Tang, Yeung et al.(2015)Yeung, Russakovsky, Jin, Andriluka, Mori, and Fei-Fei] only address temporal detection, and suffer from the inefficient nature of temporal sliding windows. Our framework is based on incrementally linking frame-level region proposals and temporal smoothing (in a similar fashion to [Evangelidis et al.(2014)Evangelidis, Singh, and Horaud]), an approach which is computationally more efficient and can handle long untrimmed videos.
Indeed methods which connect frame-level region proposals for joint spatial and temporal localisation have risen to the forefront of current research. Gkioxari and Malik [Gkioxari and Malik(2015)] have extended [Girshick et al.(2014)Girshick, Donahue, Darrel, and Malik] and [Simonyan and Zisserman(2014a)] to tackle action detection using unsupervised Selective-Search region proposals and separately trained SVMs. However, as the videos used to evaluate their work only contain one action and were already temporally trimmed (J-HMDB-21 [Jhuang et al.(2013)Jhuang, Gall, Zuffi, Schmid, and Black]), it is not possible to assess their temporal localisation performance. Weinzaepfel et al.’s approach [Weinzaepfel et al.(2015)Weinzaepfel, Harchaoui, and Schmid], instead, first generates region proposals using EdgeBoxes [Zitnick and Dollár(2014)] at frame level to later use a tracking-by-detection approach based on a novel track-level descriptor called a Spatio-Temporal Motion Histogram. Moreover, [Weinzaepfel et al.(2015)Weinzaepfel, Harchaoui, and Schmid] achieves temporal trimming using a multi-scale sliding window over each track, making it inefficient for longer video sequences. Our approach improves on both [Gkioxari and Malik(2015), Weinzaepfel et al.(2015)Weinzaepfel, Harchaoui, and Schmid] by using an efficient two-stage single network for detection of region proposals and two passes of dynamic programming for tube construction.
Some of the reviewed approaches [Weinzaepfel et al.(2015)Weinzaepfel, Harchaoui, and Schmid, van Gemert et al.(2015)van Gemert, Jain, Gati, and Snoek] could potentially be able to detect co-occurring actions. However, [Weinzaepfel et al.(2015)Weinzaepfel, Harchaoui, and Schmid] limit their method to produce maximum of two detections per class, while [van Gemert et al.(2015)van Gemert, Jain, Gati, and Snoek] does so on the MSRII dataset [Cao et al.(2010)Cao, Liu, and Huang] which only contains three action classes of repetitive nature (clapping, boxing and waving). Klaser at al. [Kläser et al.(2010)Kläser, Marszałek, Schmid, and Zisserman] use a space-time descriptor and a sliding window classifier to detect the location of only two actions (phoning and standing up). In contrast, in our LIRIS-HARL tests (§ 4) we consider 10 diverse action categories.
Methodology
As outlined in Figure 1, our approach combines a region-proposal network (§ 3.1-Fig. 1b) with a detection network (§ 3.2-Fig. 1d), and fuses the outputs (§ 3.3-Fig. 1f) to generate action tubes (§ 3.4-Fig. 1g-h). All components are described in detail below.
To generate rectangular action region hypotheses in a video frame we adopt the Region Proposal Network (RPN) approach of [Ren et al.(2015)Ren, He, Girshick, and Sun], which is built on top of the last convolutional layer of the VGG-16 architecture by Simonyan and Zisserman [Simonyan and Zisserman(2014b)]. To generate region proposals, this mini-network slides over the convolutional feature map outputted by the last layer, processing at each location an spatial window and mapping it to a lower dimensional feature vector (512-d for VGG-16). The feature vector is then passed to two fully connected layers: a box-regression layer and a box-classification layer.
During training, for each image location, region proposals (also called ‘anchors’) [Ren et al.(2015)Ren, He, Girshick, and Sun] are generated. We consider those anchors with a high Intersection-over-Union () with the ground-truth boxes () as positive examples, whilst those with as negatives. Based on these training examples, the network’s objective function is minimised using stochastic gradient descent (SGD), encouraging the prediction of both the probability of an anchor belonging to action or no-action category (a binary classification), and the 4 coordinates of the bounding box.
2 Detection network
For the detection network we use a Fast R-CNN net [Girshick(2015)] with a VGG-16 architecture [Simonyan and Zisserman(2014b)]. This takes the RPN-based region proposals (§ 3.1) and regresses a new set of bounding boxes for each action class and associates classification scores. Each RPN-generated region proposal leads to (number of classes) regressed bounding boxes with corresponding class scores.
Analogously to the RPN component, the detection network is also built upon the convolutional feature map outputted by the last layer of the VGG-16 network. It generates a feature vector for each proposal generated by RPN, which is again fed to two sibling fully-connected layers: a box-regression layer and a box-classification layer. Unlike what happens in RPNs, these layers produce multi-class softmax scores and refined boxes (one for each action category) for each input region proposal.
We employ a variation on the training strategy of [Ren et al.(2015)Ren, He, Girshick, and Sun] to train both the RPN and Fast R-CNN networks. Shaoqing et al\bmvaOneDot [Ren et al.(2015)Ren, He, Girshick, and Sun] suggested a 4-steps ‘alternating training’ algorithm in which in the first 2 steps, a RPN and a Fast R-CNN nets are trained independently, while in the 3 and 4 steps the two networks are fine-tuned with shared convolutional layers. In practice, we found empirically that the detection accuracy on UCF101 slightly decreases when using shared convolutional features, i.e., when fine tuning the RPN and Fast-RCNN trained models obtained after the first two steps. As a result, we train the RPN and the Fast R-CNN networks independently following only the 1 and 2 steps of [Ren et al.(2015)Ren, He, Girshick, and Sun], while neglecting the 3 and 4 steps suggested by [Ren et al.(2015)Ren, He, Girshick, and Sun].
3 Fusion of appearance and motion cues
In a work by Redmon et al\bmvaOneDot [Redmon et al.(2015)Redmon, Divvala, Girshick, and Farhadi], the authors combine the outputs from Fast R-CNN and YOLO (You Only Look Once) object detection networks to reduce background detections and improve the overall detection quality. Inspired by their work, we use our motion-based detection network to improve the scores of the appearance-based detection net (c.f. Fig. 1f).
The second term adds to the existing score of the appearance-based detection box a proportion, equal to the amount of overlap, of the motion-based detection score. In our tests we set .
4 Action tube generation
The output of our fusion stage (§ 3.3) is, for each video frame, a collection of detection boxes for each action category, together with their associated augmented classification scores (1). Detection boxes can then be linked up in time to identify video regions most likely to be associated with a single action instance, or action tube. Action tubes are connected sequences of detection boxes in time, without interruptions, and unlike those in [Gkioxari and Malik(2015)] they are not constrained to span the entire video duration.
Once an optimal path has been found, we remove all the detection boxes associated with it and recursively seek the next best action path. Extracting multiple paths allows our algorithm to account for multiple co-occurring instances of the same action class.
Smooth path labelling and temporal trimming.
As the resulting action-specific paths span the entire video duration, while human actions typically only occupy a fraction of it, temporal trimming becomes necessary. The first pass of dynamic programming (2) aims at extracting connected paths by penalising regions which do not overlap in time. As a result, however, not all detection boxes within a path exhibit strong action-class scores.
where is a scalar parameter weighting the relative importance of the pairwise term. The pairwise potential is defined to be:
where is a class-specific constant parameter which we set by cross validation. In the supplementary material, we show the impact of the class-specific on the detection accuracy. Equation (4) is the standard Potts model which penalises labellings that are not smooth, thus enforcing a piecewise constant solution. Again, we solve (3) using the Viterbi algorithm.
Experimental validation and discussion
In order to evaluate our spatio-temporal action detection pipeline we selected what are currently considered among the most challenging action detection datasets: UCF-101 [Soomro et al.(2012)Soomro, Zamir, and Shah], LIRIS HARL D2 [Wolf et al.(2012)Wolf, Mille, Lombardi, Celiktutan, Jiu, Baccouche, Dellandréa, Bichot, Garcia, and Sankur], and J-HMDB-21 [Jhuang et al.(2013)Jhuang, Gall, Zuffi, Schmid, and Black]. UCF-101 is the largest, most diverse and challenging dataset to date, and contains realistic sequences with a large variation in camera motion, appearance, human pose, scale, viewpoint, clutter and illumination conditions. Although each video only contains a single action category, it may contain multiple action instances of the same action class. To achieve a broader comparison with the state-of-the-art, we also ran tests on the J-HMDB-21 [Jhuang et al.(2013)Jhuang, Gall, Zuffi, Schmid, and Black] dataset. The latter is a subset of HMDB-51 [Kuehne et al.(2011)Kuehne, Jhuang, Garrote, Poggio, and Serre] with 21 action categories and 928 videos, each containing a single action instance and trimmed to the action’s duration. The reported results were averaged over the three splits of J-HMDB-21. Finally we conducted experiments on the more challenging LIRIS-HARL dataset, which contains 10 action categories, including human-human interactions and human-object interactions (e.g., ‘discussion of two or several people’, and ‘a person types on a keyboard’http://liris.cnrs.fr/voir/activities-dataset). In addition to containing multiple space-time actions, some of which occurring concurrently, the dataset contains scenes where relevant human actions take place amidst other irrelevant human motion.
For all datasets we used the exact same evaluation metrics and data splits as in the original papers. In the supplementary material, we further discuss all implementation details, and propose an interesting quantitative comparison between Selective Search- and RPN-generated region proposals.
Table 1 presents the results we obtained on UCF-101, and compares them to the previous state-of-the-art [Weinzaepfel et al.(2015)Weinzaepfel, Harchaoui, and Schmid, Yu and Yuan(2015)]. We achieve an mAP of compared to reported by [Weinzaepfel et al.(2015)Weinzaepfel, Harchaoui, and Schmid] (a gain), at the standard threshold of . At a threshold of we still get a high score of , (comparable to [Weinzaepfel et al.(2015)Weinzaepfel, Harchaoui, and Schmid] at ). Note that we are the first to report results on UCF-101 up to , attesting to the robustness of our approach to more accurate localisation requirements. Although our separate appearance- and motion-based detection pipelines already outperform the state-of-the-art (Table 1), their combination (§ 3.3) delivers a significant performance increase.
Some representative example results from UCF-101 are shown in Fig. 3. Our method can detect several (more than ) action instances concurrently, as shown in Fig. 2, in which three concurrent instances and in total six action instances are detected correctly. Quantitatively, we report class-specific video AP (average precision in ) of , and on the UCF-101 action categories ‘Fencing’, ‘SalsaSpin’ and ‘IceDancing’, respectively, which all concern multiple inherently co-occurring action instances. Class-specific video APs on UCF-101 are reported in the supplementary material.
Performance comparison on J-HMDB-21.
The results we obtained on J-HMDB-21 are presented in Table 2. Our method again outperforms the state-of-the-art, with an mAP increase of and at as compared to [Gkioxari and Malik(2015)] and [Weinzaepfel et al.(2015)Weinzaepfel, Harchaoui, and Schmid], respectively. Note that our motion-based detection pipeline alone exhibits superior results, and when combined with appearance-based detections leads to a further improvement of at . These results attest to the high precision of the detections - a large portion of the detection boxes have high IoU overlap with the ground-truth boxes, a feature due to the superior quality of RPN-based region proposals as opposed to Selective Search’s (a direct comparison is provided in the supplementary material). Sample detections on J-HMDB-21 are shown in Figure 4. Also, we list our classification accuracy results on J-HMDB-21 in Table 3, where it can be seen that our method achieves an gain compared to [Gkioxari and Malik(2015)].
Performance comparison on LIRIS-HARL.
LIRIS HARL allows us to demonstrate the efficacy of our approach on temporally un-trimmed videos with co-occurring actions. For this purpose we use LIRIS-HARL’s specific evaluation tool - the results are shown in Table 4. Our results are compared with those of i) VPULABUAM-13 [SanMiguel and Suja(2012)] and ii) IACAS-51 [He et al.(2012)He, Liu, Sui, Xiang, and Pan] from the original LIRIS HARL detection challenge. In this case, our method outperforms the competitors by an even larger margin. We report space-time detection results by fixing the threshold quality level to 10% for the four thresholds [Wolf et al.(2014)Wolf, Mille, Lombardi, Celiktutan, Jiu, Dogan, Eren, Baccouche, Dellandrea, Bichot, Garcia, and Sankur] and measuring temporal precision and recall along with spatial precision and recall, to produce an integrated score. We refer the readers to [Wolf et al.(2014)Wolf, Mille, Lombardi, Celiktutan, Jiu, Dogan, Eren, Baccouche, Dellandrea, Bichot, Garcia, and Sankur] for more details on LIRIS HARL’s evaluation metrics.
We also report in Table 5 the mAP scores obtained by the appearance, motion and the fusion detection models, respectively (note that there is no prior state of the art to report in this case). Again, we can observe an improvement of mAP at due to our fusion strategy. To demonstrate the advantage of our 2nd pass of DP (§ 3.4), we also generate results (mAP) using only the first DP pass (§ 3.4). Without the 2 pass performance decreases by , highlighting the importance of temporal trimming in the construction of action tubes.
Test-time detection speed comparison.
Finally, we compared detection speed at test time of the combined region proposal generation and CNN feature extraction approach used in ([Gkioxari and Malik(2015), Weinzaepfel et al.(2015)Weinzaepfel, Harchaoui, and Schmid]) to our neural-net based, single stage action proposal and classification pipeline on the J-HMDB-21 dataset.We found our method to be faster than [Gkioxari and Malik(2015)] and faster than [Weinzaepfel et al.(2015)Weinzaepfel, Harchaoui, and Schmid], with a mean of 113.52 [Gkioxari and Malik(2015)], 52.23 [Weinzaepfel et al.(2015)Weinzaepfel, Harchaoui, and Schmid] and 10.89 (ours) seconds per video, averaged over all the videos in J-HMDB-21 split1. More timing comparison details and qualitative results (images and video clips) can be found in the supplementary material.
Discussion.
The superior performance of the proposed method is due to a number of reasons. 1) Instead of using unsupervised region proposal algorithms as in [Uijlings et al.(2013)Uijlings, van de Sande, Gevers, and Smeulders, Zitnick and Dollár(2014)], our pipeline takes advantage of a supervised RPN-based region proposal approach which exhibits better recall values than [Uijlings et al.(2013)Uijlings, van de Sande, Gevers, and Smeulders] (supplementary-material). 2) Our fusion technique improves the mAPs (over the individual appearance or motion models) by , and on the UCF-101, J-HMDB-21 and LIRIS HARL datasets respectively. We are the first to report an ablation study (supplementary-material) where it is shown that the proposed fusion strategy (§ 3.3) improves the class-specific video APs of UCF-101 action classes. 3) Our original 2nd pass of DP is responsible for significant improvements in mAP by on LIRIS HARL and on UCF-101 (supplementary-material). Additional qualitative results are provided in the supplementary video https://www.youtube.com/embed/vBZsTgjhWaQ, and on the project web page http://sahasuman.bitbucket.org/bmvc2016, where the code has also been made available.
Conclusions and future work
In this paper, we presented a novel human action recognition approach which addresses in a coherent framework the challenges involved in concurrent multiple human action recognition, spatial localisation and temporal detection, thanks to a novel deep learning strategy for simultaneous detection and classification of region proposals and an improved action tube generation approach. Our method significantly outperforms the previous state-of-the-art on the most challenging benchmark datasets, for it is capable of handling multiple concurrent action instances and temporally untrimmed videos.
Its combination of high accuracy and fast detection speed at test time is very promising for real-time applications, for instance smart car navigation. As the next step we plan to make our tube generation and labelling algorithm fully incremental and online, by only using region proposals from independent frames at test time and updating the dynamic programming optimisation step at every incoming frame.
This work was partly supported by ERC grant ERC-2012-AdG 321162-HELIOS, EPSRC grant Seebibyte EP/M013774/1 and EPSRC/MURI grant EP/N019474/1.