Weakly Supervised Action Labeling in Videos Under Ordering Constraints

Piotr Bojanowski, Rémi Lajugie, Francis Bach, Ivan Laptev, Jean Ponce, Cordelia Schmid, Josef Sivic

Introduction

Significant progress towards action recognition in realistic video settings has been achieved in the past few years . However action recognition is often cast as a classification or detection problem using fully annotated data, where the temporal boundaries of individual actions, e.g. in the form of pre-segmented video clips, are given during training. The goal of this paper is to exploit the supervisory power of the temporal ordering of actions in a video stream, as illustrated in figure 1.

Gathering fully annotated videos with accurately time-stamped action labels is quite time consuming in practice. This limits the utility of fully supervised machine learning techniques on large-scale data. Using data redundancy, weakly and semi-supervised methods are a promising alternative in this case. On the other hand, it is easy to gather videos with some level of textual annotation but poor temporal localization, from movie scripts for example. This type of weak supervisory signal has been used before in classification and temporal localization tasks. However, the crucial information on the ordering of actions has, to the best of our knowledge, been ignored so far in the weakly supervised setting. Following recent work on discriminative clustering , image and video cosegmentation, we propose to exploit this information in a discriminative framework where both the action model and the optimal assignments under temporal constraints are learned together.

The temporal ordering of actions, e.g. in the form of Markov models or action grammars, have been used to constrain action prediction in videos . These kinds of spatial and temporal constraints have been also used in the context of group activity recognition . Similar to us, these papers exploit the temporal structure of videos, but focus on inferring action sequences from noisy but pre-defined action detectors, often in constrained surveillance and laboratory settings with a limited number of actions and static cameras. In contrast, in this work we explore the temporal structure of actions for learning action classifiers in a weakly supervised set-up and show results on challenging videos from feature length movies.

Related is also work on recognition of composite activities , where atomic action models (“cut”, “open”) are learned given full supervision on a cooking video dataset. Composite activity models (“prepare pizza”) are learned on top of the atomic actions, using the prediction scores for the atomic actions as features. Annotations are, however, used without taking into account the ordering of actions.

Temporal models for recognition of individual actions have been explored in e.g. . Implicit models in the form of temporal pyramids have been used with bag-of-features representations . Others have used more explicit temporal models in the form of, e.g. latent action parts or hidden Markov models . Contrary to these methods, we do not use an a priori model of the temporal structure of individual actions, but instead exploit the given ordering constraints between actions to learn better individual actions models.

Weak supervision for learning actions has been explored in . These methods use uncertain temporal annotations of actions provided by movie scripts. Contrary to these works our method learns multiple actions simultaneously and incorporates temporal ordering constraints on action labels obtained, e.g. from the movie scripts.

Dynamic time warping algorithms (DTW) can be used to match temporal sequences, and are extensively used in speech recognition, e.g. . In computer vision, the temporal order of events has been exploited in , where a DTW-like algorithm is used at test time to improve the performance of non-maximum suppression on the output of pre-trained action detectors.

Discriminative clustering is an unsupervised method that partitions data by minimizing a discriminative objective, optimizing over both classifiers and labels . Convex formulations of discriminative clustering have been explored in . In computer vision these methods have been successfully applied to co-segmentation . The approach presented in this paper is inspired by this framework, but adds to it the use of ordering constraints.

In this work, we make use of the Frank-Wolfe algorithm (a.k.a conditional gradient) to minimize our cost function. The Frank-Wolfe algorithm is a classical convex optimization procedure that permits optimizing a continuously differentiable convex function over a convex compact domain only by optimizing linear functions over the domain. In particular, it does not require any projection steps. It has recently received increased attention in the context of large-scale optimization .

2 Problem Statement and Contributions

The temporal assignment problem addressed in the rest of this paper and illustrated by Fig. 1 can be stated as follows: We are given a set of NN video clips (or clips for short in what follows). A clip is defined as a contiguous video segment consisting of FF frames, and may correspond, for example, to a scene (as defined in a movie script) or a collection of subsequent shots. Each clip is divided into TT small time intervals (chunks of videos consisting of F/T=10F/T=10 frames in our case), and annotated by an ordered list of KK elements taken from some action set A{\cal A} of size A=∣A∣A=|{\cal A}| (that may consist of labels such as “open door”, “stand up”, “answer phone”, etc., as in Fig. 1 for example). Note that clips are not of the same length but for the sake of simplicity, we assume they are. We address the problem of assigning to each time interval of each clip one action in A{\cal A}, respecting the order in which the actions appear in the original annotation list (Fig. 2).

We make the following contributions: (i) we propose a discriminative clustering model (section 2) that handles weak supervision in the form of temporal ordering constraints and recovers a classifier for each action together with the temporal localization of each action in each video clip; (ii) we design a convex relaxation of the proposed model and show it can be efficiently solved using the conditional gradient (Frank-Wolfe) algorithm (section 3); and finally (iii) we demonstrate improved performance of our model on a new action dataset for the tasks of temporal localization (section 6) and action classification (section 7). All the data and code are publicly available at http://www.di.ens.fr/willow/research/ordering.

Discriminative Clustering with Ordering Constraints

In this section we describe the proposed discriminative clustering model that incorporates label ordering constraints. The input is a set of video clips, each annotated with an ordered list of action labels specifying the sequence of actions present in the clip. The output is the temporal assignment of actions to individual time intervals in each clip respecting the ordering constraint provided by the annotations together with a learnt classifier for each action, common for all clips. In the following, we first formulate the temporal assignment of actions to individual frames as discriminative clustering (section 2.1), then introduce a parametrization of temporal assignments using indicator variables (section 2.2), and finally we describe the choice of a loss function for the discriminative clustering that leads to a convex cost (section 2.3).

Let us denote by M\mathcal{M} the set of admissible assignments on {1,…,T}\{1,\dots,T\}, that is, the set of sequences m=(m1,…,mT)m=(m_{1},\ldots,m_{T}) with elements in {1,…,K}\{1,\dots,K\} such that m1=1m_{1}=1, mT=Km_{T}=K, and mt+1=mtm_{t+1}=m_{t} or mt+1=mt+1m_{t+1}=m_{t}+1 for all tt in {1,…,T−1}\{1,\dots,T-1\} Such an assignment is illustrated in Fig. 2.

with respect to assignment mm in M\mathcal{M}. The regularizer Ω\Omega prevents overfitting and we therefore define a scalar parameter λ\lambda to control this effect. Jointly learning the classifiers and solving the assignment problem corresponds to the following optimization problem:

2 Parameterization Using an Assignment Matrix

Because of temporal constraints, we want the assignment matrices ZZ to correspond to valid assignments mm. This amounts to imposing some constraints on ZZ. Let us therefore define Z\mathcal{Z}, the set of all valid assignment matrices as:

There is a bijection between the sets Z\mathcal{Z} and M\mathcal{M}. For each mm in M\mathcal{M} there exists a unique corresponding ZZ in Z\mathcal{Z} and vice versa. Figure 3 gives an intuitive illustration of this bijection.

The set Z\mathcal{Z} is a subset of the set of stochastic matrices (positive matrices whose rows sum up to 1), formed by the matrices whose columns consist of exactly KK blocks of contiguous ones occurring in a predefined order (K=6K=6 in Fig. 3). There are as many elements in Z\mathcal{Z} as ways of choosing (K−1)(K-1) transitions among (T−1)(T-1) possibilities, thus ∣Z∣=(T−1K−1)\left|\mathcal{Z}\right|=\binom{T-1}{K-1}, which can be extremely large in our setting (in our setting T≈100T\approx 100 and K≈10K\approx 10). Furthermore, it is very difficult to describe explicitly the algebraic constraints on stochastic matrices that define Z\mathcal{Z}. This point will prove important in Sec. 3 when we propose an optimization algorithm for learning our model. Using these notations, Eq. (2) is equivalent to:

3 Quadratic Cost Functions

This is exactly a ridge regression cost. Minimizing this cost with respect to WW and bb for fixed ZZ can be done in closed form . Setting the partial derivatives with respect to WW and bb to zero and plugging the solution back yields the following equivalent problem:

and the matrix Πp\Pi_{p} is the p×pp\times p centering matrix Ip−1p1p1pTI_{p}-\frac{1}{p}\mathbf{1}_{p}\mathbf{1}_{p}^{T}. This corresponds to implicitly learning the classifier while finding the optimal Z by solving a quadratic optimisation problem in ZZ. The implicit classifier parameters WW and bb are shared among all video clips and can be recovered in closed-form as:

Convex Relaxation and the Frank-Wolfe Algorithm

In Sec. 2, we have seen that our model can be interpreted as the minimization of a convex quadratic function (BB is positive semidefinite) over a very large but discrete domain. As is usual for this type of hard combinatorial optimization problem, we replace the discrete set Z\mathcal{Z} by its convex hull Z‾\overline{\mathcal{Z}}. This allows us to find a continuous solution of the relaxed problem using an appropriate and efficient algorithm for convex optimization.

We want to carry out the minimization of a convex function over a complex polytope Z‾\overline{\mathcal{Z}}, defined as the convex hull of a large but finite set of integer points defined by the constraints associated with admissible assignments. When it is possible to optimize a linear function over a constraint set of this kind, but other usual operations (like projections) are not tractable, a good way to optimize a convex objective function is to use the iterative Frank-Wolfe algorithm (a.k.a. conditional gradient method) . We show in Sec. 3.2 that we can minimize linear functions over Z‾\overline{\mathcal{Z}}, so this is an appropriate choice in our case.

The idea behind the Frank-Wolfe algorithm is rather simple. An affine approximation of the objective function is minimized yielding a point Z∗Z^{*} on the edge of Z‾\overline{\mathcal{Z}}. Then a convex combination of Z∗Z^{*} and the current point ZZ is computed. This is repeated until convergence. The interpolation parameter γ\gamma can be chosen either by using the universal step size 2p+1\frac{2}{p+1}, where pp is the iteration counter (see and references therein) or, in the case of quadratic functions, by solving a univariate quadratic equation. In our implementation, we use the latter. A good feature of the Frank-Wolfe algorithm is that it provides for free a duality gap (referred to as the linearization duality gap ) that can be used as a certificate of sub-optimality and stopping criterion. The procedure is described in the special case of our relaxed problem in Algorithm 1. Figure 4 illustrates one step of the optimization.

2 Linear Function Minimization over 𝒵¯¯𝒵\overline{\mathcal{Z}}

The optimal value PT∗(K)P_{T}^{*}(K) can be computed in O(TK)O(TK) using dynamic programming, by precomputing the matrix DD, incrementally computing the corresponding Pt∗(k)P_{t}^{*}(k) values, and maintaining at each node (t,k)(t,k) back pointers to the appropriate neighbors.

3 Rounding

At convergence, the Frank-Wolfe algorithm finds the (non-integer) global optimum Z∗Z^{*} of Eq. (6) over Z‾\overline{\mathcal{Z}}. Given Z∗Z^{*}, we want to find an appropriate nearby point ZZ in Z\mathcal{Z}. The simplest geometric rounding scheme consists in finding the closest point of Z\mathcal{Z} according to the Frobenius distance : min⁡Z∈Z∥Z∗−Z∥F2\min_{Z\in\mathcal{Z}}\|Z^{*}-Z\|_{F}^{2}. Expanding the norm yields: ∥Z∗−Z∥F2=Tr(Z∗TZ∗)+Tr(ZTZ)−2Tr(Z∗TZ)\|Z^{*}-Z\|_{F}^{2}=\text{Tr}({Z^{*}}^{T}Z^{*})+\text{Tr}(Z^{T}Z)-2\text{Tr}({Z^{*}}^{T}Z).

Since Z∗Z^{*} is fixed, its norm is a constant. Moreover, since ZZ is an element of Z\mathcal{Z}, its squared norm is constant and equal to TT. The rounding problem is therefore equivalent to: min⁡Z∈Z−2Tr(Z∗TZ)\min_{Z\in\mathcal{Z}}-2\text{Tr}({Z^{*}}^{T}Z), that is to the minimization of a linear function over Z‾\overline{\mathcal{Z}}. This can be done, as in Sec. 3.2, using dynamic programming.

Practical Concerns

In this section, we detail some refinements of our model. First we show how to tackle a semi-supervised setting where some time-stamped annotations are available. Secondly, we discuss how to avoid the trivial solutions, a common issue in discriminative clustering methods .

This supervised model does not change the optimization procedure, which remains valid.

2 Minimum size constraints

There are two inherent problems with discriminative clustering First, the constant assignment matrix is typically a trivial optimum. As explained in this occurs when the optimization domain is symmetric over permutations of the labels of the assignment matrices. Due to our temporal constraints, the set Z\mathcal{Z} is not symmetric and thus we are not subject to this effect.

The second difficulty is linked to the use of the centering matrix ΠT\Pi_{T} in the expression of the quadratic cost matrix BB. Indeed, we notice that the constant vector of length TT is an eigen vector of ΠT\Pi_{T}. Therefore, the column-wise constant matrices are trivial solutions to our problem. These piecewise-constant solutions are not admissible for our problem due to the temporal constraints. In practice however, we have have observed that the algorithm returned an assignment with almost all points being affected to the background label ∅\varnothing. We consider two ways to get rid of the trivial solutions.

3 Linear penalty.

To avoid solutions with dominant classes we add constraints over the fraction of clip intervals affected to each class. Ideally, we would like to incorporate a hard constraint over the proportions of each class as in , that is, to add to the problem formulated in Eq. (6), a constraint of the type:

Note that, with this simple modification, we can still use Alg. 1.

4 Balancing the loss.

Dataset and Features

Dataset. Our input data consists of challenging video clips annotated with sequences of actions. One possible source for such data is movies with their associated scripts . The annotations provided by this kind of data are noisy and do not provide ground-truth time-stamps for evaluation. To address this issue, we have constructed a new action dataset, containing clips annotated by sequences of actions. We have taken the 69 movies from which the clips of the Hollywood2 dataset were extracted , and manually added full time-stamped annotation for 16 classes (12 of these classes are already present in Hollywood2). To build clips that form our input data, we search in the annotations for action chains containing at least two elements. To do so, we pad the temporal action annotations by 250 frames and search for overlapping intervals. A chain of such overlapping annotations forms one video clip with associated action sequence in our dataset. In the end we obtain 937 clips, with number of actions ranging from 2 to 11. We subdivide each clip into temporal intervals of length 10 frames. Clips contain on average 84 intervals, the shortest containing 11, the longest 289.

Action Labeling Experiments

Experimental Setup. To carry out the action labeling experiments, we split 90% of the dataset into three parts (Fig. 5) that we denote Sup (for supervised), Eval (for evaluation) and Val (for validation). Sup is the part of data that has time-stamped annotations, and it is used only in the semi-supervised setting described in Sec. 4.1. Val is the set of examples on which we automatically adjust the hyper-parameters for our method (λ,κ,D\lambda,\kappa,D). In practice we fix the Val set to contain 5% of the dataset. This set is provided with fully time-stamped annotations, but these are not used during the cost optimization. None of the reported results are computed on this set. We evaluate the quality of the assignment on the Eval set.

Note that we carry out the Frank-Wolfe optimization on the union of all three sets. The annotations from the Sup set are used to constrain ZZ in the semi-supervised setup while those from the Val set are only used for choosing our hyper parameters. The supervisory information used over the rest of the data are the ordered annotations without time stamps. Please also keep in mind that there are no “training” and “testing” phases per se in this primary assignment task. All our experiments are conducted over five random splits of the data. This allows us to present results with error bars.

Performance Measure. Several measures may be used to evaluate the performance of discriminative clustering algorithms. Some authors propose to use the output classifier to perform a classification task or use the output partition of the data as a solution of the segmentation task . Yet another way to evaluate is to use a loss between partitions as in . Note that because of temporal constraints, for every clip we have a set of corresponding (prediction, ground-truth) pairs. We have thus chosen to measure the assignment quality for every ground-truth action interval I*I^{\text{*}} and prediction II as ∣I∩I∗∣/∣I∣|I\cap I^{*}|/|I|. This measure is similar to the standard Jaccard measure used for comparing ensembles . Therefore, with a slight abuse of notation, we refer to this measure as the Jaccard measure. This performance measure is well suited for our problem since it respects the following properties: (1) it is high if the action predicted is included in the ground-truth annotation, (2) it is low if the prediction is bigger than the annotation, (3) it is lowest if the prediction is out of the annotation, (4) it does not take into account the prediction of the background class. The score is averaged across all ground-truth intervals. The perfect score of 1 is achieved when all actions are aligned to the correct annotations, but accurate temporal segmentation is not required as long as the predicted labels are within the ground truth interval.

Baselines. We compare our method to the three following baselines. All these are trained using the same features as the ones used for our method. For all baselines, we round the obtained solution ZZ using the scheme described in Sec. 3.3.

Normalized Cuts (NCUT). We compare our method to normalized cuts (or spectral clustering) . Let us define BB as the symmetric Laplacian of the matrix EE: B=I−D−12ED−12B=I-D^{-\frac{1}{2}}ED^{-\frac{1}{2}}, where D=Diag(E1)D=\text{Diag}\left(E\mathbf{1}\right). EE measures both the proximity and appearance similarity of intervals. For all (i,j)(i,j) in {1,…,T}2\{1,\dots,T\}^{2}, we compute: Eij=e−α∣i−j∣−βdχ2(Xi,Xj) \dsrom1∣i−j∣<dminE_{ij}=e^{-\alpha|i-j|-\beta d_{\chi^{2}}(X_{i},X_{j})}\ \textrm{\dsrom{1}}_{|i-j|<d_{\text{min}}}, where dχ2d_{\chi^{2}} is the Chi-squared distance. More precisely, we minimize over all cuts ZZ the cost g(Z)=Tr(ZZTB)g(Z)=\text{Tr}\left(ZZ^{T}B\right). gg is convex (BB is positive semidefinite) and we can use the Frank-Wolfe optimization scheme developed for our model. Intuitively, this baseline is searching for a partition of the video such that time intervals falling into the same segments have close-by features according to the Chi-squared distance.

Bojanowski et al. . We also consider our own implementation of the weakly-supervised approach proposed in . We replace our ordering constraints by the corresponding “at least one” constraints. When an action is mentioned in the sequence, we require it appears at least once in the clip. This corresponds to a set of linear constraints on ZZ. We adapt this technique in order to work on our dataset. Indeed, the available implementation requires storing a square matrix of the size of the problem. Instead, we choose to minimize the convex objective of using the Frank-Wolfe algorithm which is more scalable.

Supervised Square Loss (SL). For completeness, we also compare our method to a fully supervised approach. We train a classifier using the square loss over the annotated Sup set and score all time intervals in Eval. We use the square loss since it is used in our method and all other baselines.

Weakly Supervised Setup. In this setup, all baselines except (SL) have only access to weak supervision in the form of ordering constraints. Figure 6 (left) illustrates the quality of the predicted asignmentss and compares our method to baselines. Our method performs better than all other weakly-supervised methods. Both the Bojanowski et al. and NCUT baselines have low scores in the weakly-supervised setting. This shows the advantage of exploiting temporal constraints as weak supervisory signal. The fully supervised baseline (blue) eventually recovers a better alignment than our method as the fraction of fully annotated data increases. This occurs (when the red line crosses the blue line) at the 25% mark, as the supervised data makes up for the lack of ordering constraints. Fully time-stamped annotated data are expensive to produce whereas movies scripts are often easy to get. It appears thus that manually annotated videos are not always necessary since good performance is reached simply by using weak supervision. Figure 7 shows the results for all weakly-supervised methods for all classes. We notice that we outperform the baselines on the most frequent classes (such as “Open Door”, “Sit Down” and “Stand Up”).

Semi-supervised Setup. Figure 6 (right) illustrates the performance of our model when some supervised data is available. The fraction of the supervised data is given on the x-axis. First, note that our semi-supervised method (red) is always and consistently (Cf error bars) above the square loss baseline (blue). Of course, during the optimization, our method has access to weak annotations over the whole dataset, and to full annotations on the Sup set whereas the SL baseline has access only to the latter. This demonstrates the benefits of exploiting temporal constraints during learning. The semi-supervised Bojanowski et al. baseline (orange) has low performance, but it improves with the amount of full supervision provided.

Classification Experiments

The experiments in the previous section evaluate the quality of the recovered assignment matrix ZZ. Here we evaluate instead the quality of the recovered classifiers on a held-out test set of data for an action classification task. We recover these classifiers as explained later in this section. We can treat them as KK independent, one-versus-rest classifiers and use them to score the samples from the test set. We evaluate this performance by computing per-class precision and recall and report the corresponding average precision for each class.

Experimental setup. The models are trained following the procedure described in the previous section. To test the performance of our classifiers, we use the held out set of clips. This set is made of 10% of the clips from the original data. The clips from this set are identical in nature to the ones used to train the models. We also perform multiple random splits to report results with error bars.

Recovering the classifiers. One of the nice features of our method is that we can estimate the implicit classifiers corresponding to our solution Z∗Z^{*}. We do so using the expression from Eq. 7.

Baselines. We compare the classifiers obtained by our method to those obtained by the Bojanowski et al. baseline . We also compare them to the classifiers learned using the (SL) baseline.

Weakly Supervised Setup. Classification results are presented in Fig. 8 (left). We observe a behavior similar to the action labeling experiment. But the supervised classifier (SL) trained on the Sup set using the square loss (blue) always performs worse than our model (red). This can be explained by the fact that the proposed model makes use of mode data. Even though our model has only access to weak annotation, it can prove sufficient to train good classifiers. The weakly-supervised method from Bojanowski et al. (orange) is performing worst, exactly as in the previous task. This can be explained by the fact that this method does not have access to full supervision or ordering constraints.

Semi-supervised Setup. In the semi-supervised setting (Fig. 8 (left)), our method (red) performs better than the supervised SL baseline (blue). The action model we recover is consistently better than the one obtained using only fully supervised data. Thus, our method is able to perform well semi-supervised learning. The Bojanowski et al. baseline (orange) improves when the fraction of annotated examples increases. Nonetheless, we see that making use of ordering constraints as used by our method signigicantly improves over simple linear inequalities (“at least one” constraints as formulated in ).

Acknowledgements. This work was supported by the European integrated project AXES, the MSR-INRIA laboratory, EIT-ICT labs, a Google Research Award, a PhD fellowship from the EADS Foundation, the Institut Universitaire de France and ERC grants ALLEGRO, VideoWorld, Activia and Sierra.

References