Near-Online Multi-target Tracking with Aggregated Local Flow Descriptor

Wongun Choi

Introduction

The goal of multiple target tracking is to automatically identify objects of interest and reliably estimate the motion of targets over the time. Thanks to the recent advancement in image-based object detection methods , tracking-by-detection has become a popular framework to tackle the multiple target tracking problem. The advantages of the framework are that it naturally identifies new objects of interest entering the scene, that it can handle video sequences recorded using mobile platforms, and that it is robust to a target drift. The key challenge in this framework is to accurately group the detections into individual targets with high accuracy (data association), so one target could be fully represented by a single estimated trajectory. Mistakes made in the identity maintenance could result in a catastrophic failure in many high level reasoning tasks, such as future motion prediction, target behavior analysis, etc.

To implement a highly accurate multiple target tracking algorithm, it is important to have a robust data association model and an accurate measure to compare two detections across time (pairwise affinity measure). Recently, much work is done in the design of the data association algorithm using global (batch) tracking framework . Compared to the online counterparts , these methods have a benefit of considering all the detections over entire time frames. With a help of clever optimization algorithms, they achieve higher data association accuracy than traditional online tracking frameworks. However, the application of these methods is fundamentally limited to post-analysis of video sequences. On the other hand, the pairwise affinity measure is relatively less investigated in the recent literature despite its importance. Most methods adopt weak affinity measures (see Fig. 1) to compare two detections across time, such as spatial affinity (e.g. bounding box overlap or euclidean distance ) or simple appearance similarity (e.g. intersection kernel with color histogram ). In this paper, we address the two key challenging questions of the multiple target tracking problem: 1) how to accurately measure the pairwise affinity between two detections (i.e. likelihood to link the two) and 2) how to efficiently apply the ideas in global tracking algorithms into an online application.

As the first contribution, we present a novel Aggregated Local Flow Descriptor (ALFD) that encodes the relative motion pattern between two detection boxes in different time frames (Sec. 3). By aggregating multiple local interest point trajectories (IPTs), the descriptor encodes how the IPTs in a detection moves with respect to another detection box, and vice versa. The main intuition is that although each individual IPT may have an error, collectively they provide a strong information for comparing two detections. With a learned model, we observe that ALFD provides strong affinity measure, thereby providing strong cues for the association algorithm.

As the second contribution, we propose an efficient Near-Online Multi-target Tracking (NOMT) algorithm. Incorporating the robust ALFD descriptor as well as long-term motion/appearance models motivated by the success of modern batch tracking methods, the algorithm produces highly accurate trajectories, while preserving the causality property and running in real-time (∼10\sim 10 FPS). In every time frame tt, the algorithm solves the global data association problem between targets and all the detections in a temporal window [t\scalebox{0.5}[1.0]{-}\tau,t] of size τ\tau (see Fig. 2). The key property is that the algorithm is able to fix any association error made in the past when more detections are provided. In order to achieve both accuracy and efficiency, the algorithm generates candidate hypothetical trajectories using ALFD driven tracklets and solve the association problem with a parallelized junction tree algorithm (Sec. 4).

We perform a comprehensive experimental evaluation on two challenging datasets: KITTI and MOT Challenge datasets. The proposed algorithm achieves the best accuracy with a large margin over the state-of-the-arts (including batch algorithms) in both datasets, demonstrating the superiority of our algorithm. The rest of the paper is organized as follows. Sec. 2 discusses the background and related work in multiple target tracking literature. Sec. 3 describes our newly proposed ALFD. Sec. 4 presents overview of NOMT data association model and the algorithm. Sec. 5 discusses the details of model design. We show the analysis and experimental evaluation in Sec. 6, and finally conclude with Sec. 7.

Background

Most of multiple target tracking algorithms/systems can be classified into two categories: online method and global (batch) method.

2 Affinity Measures in Visual Tracking

The importance of a robust pairwise affinity measure (i.e. likelihood of did_{i} and djd_{j} being the same target) is relatively less investigated in the multi-target tracking literature. Most of the recent literature employs a spatial distance and/or an appearance similarity with simple features (such as color histograms). In order to learn a discriminative affinity metric, Kuo et al. introduces an online appearance learning with boosting algorithm using various feature inputs such as HoG , texture feature, and RGB color histogram. Milan et al. and Zamir et al. proposed to use a global appearance consistency measure to ensure a target has a similar (or smoothly varying) appearance over a long term. Although there have been many works exploiting appearance information or spatial smoothness, we are not aware of any work employing optical flow trajectories to define a likelihood of matching detections. Recently, Fragkiadaki et al. introduced a method to track multiple targets while jointly clustering optical flow trajectories. The work presents a promising result, but the model is complicated due to the joint inference on both target and flow level association. In contrast, our ALFD provides a strong pairwise affinity measure that is generally applicable in any tracking model.

Aggregated Local Flow Descriptor

The Aggregated Local Flow Descriptor (ALFD) encodes the relative motion pattern between two bounding boxes in a temporal distance (Δt=∣ti−tj∣\Delta t=|t_{i}-t_{j}|) given interest point trajectories . The main intuition in ALFD is that if the two boxes belong to the same target, we shall observe many supporting IPTs in the same relative location with respect to the boxes. In order to make it robust against small localization errors in detections, targets’ orientation change, and outliers/errors in the IPTs, we build the ALFD using spatial histograms. Once the ALFD is obtained, we measure the affinity between two detections using the linear product of a learned model parameter wΔtw_{\Delta t} and ALFD, i.e. aA(di,dj)=wΔt⋅ρ(di,dj)a_{A}(d_{i},d_{j})=w_{\Delta t}\cdot\rho(d_{i},d_{j}). In the following subsections, we discuss the details of the design.

We obtain Interest Point Trajectories using a local interest point detector and optical flow algorithm . The algorithm is designed to produce a set of long and accurate point trajectories, combining various well-known computer vision techniques. Given an image ItI_{t}, we run the FAST interest point detector to identify “good points” to track. In order to avoid having redundant points, we compute the distance between the newly detected interest points and the existing IPTs and keep the new points sufficiently far from the existing IPTs (>4>4 px). The new points are assigned unique IDs. For all the IPTs in tt, we compute the forward (t→t+1t\rightarrow t+1) and backward (t+1→tt+1\rightarrow t) optical flow using . The starting points of backward flows are given by the forward flows’ end point. Any IPT having a large disagreement between the two (>10>10 px) is terminated.

2 ALFD Design

Let us define the necessary notations to discuss ALFD. κid∈K\kappa_{id}\in\mathcal{K} represents an IPT with a unique idid that is parameterized by pixel locations (κid(t)[x],κid(t)[y])(\kappa_{id}(t)[x],\kappa_{id}(t)[y]) during the time of presence. κid(t)\kappa_{id}(t) denotes the pixel location at the frame tt. If κid\kappa_{id} does not exist at tt (terminated or not initiated), \o\o is returned.

We first define a unidirectional ALFD ρ′(di,dj)\rho^{\prime}(d_{i},d_{j}), i.e. from did_{i} to djd_{j}, by aggregating the information from all the IPTs that are located inside of did_{i} box and existing at tjt_{j}. Formally, we define the IPT set as K(di,dj)={κid∣κid(ti)∈di & κid(tj)≠\o}\mathcal{K}(d_{i},d_{j})=\{\kappa_{id}|\kappa_{id}(t_{i})\in d_{i}\ \&\ \kappa_{id}(t_{j})\neq\o\}. For each κid∈K(di,dj)\kappa_{id}\in\mathcal{K}(d_{i},d_{j}), we compute the relative location ri(κid)=(x,y)r_{i}(\kappa_{id})=(x,y) of each κid\kappa_{id} at tit_{i} by r_{i}(\kappa_{id})[x]=(\kappa_{id}(t_{i})[x]\scalebox{0.5}[1.0]{-}d_{i}[x])/d_{i}[w] and r_{i}(\kappa_{id})[y]=(\kappa_{id}(t_{i})[y]\scalebox{0.5}[1.0]{-}d_{i}[y])/d_{i}[h]. We compute rj(κid)r_{j}(\kappa_{id}) similarly. Notice that ri(κid)r_{i}(\kappa_{id}) are bounded between $,but, butr_{j}(\kappa_{id})arenotboundedsinceare not bounded since\kappa_{id}canbeoutsideofcan be outside ofd_{j}.Giventhe. Given ther_{i}(\kappa_{id})andandr_{j}(\kappa_{id}),wecomputethecorrespondingspatialgridbinindicesasshownintheFig.3andaccumulatethecounttobuildthedescriptor.Wedefine, we compute the corresponding spatial grid bin indices as shown in the Fig. 3 and accumulate the count to build the descriptor. We define4\times 4gridsforgrids forr_{i}(\kappa_{id})andand4\times 4+2gridsforgrids forr_{j}(\kappa_{id})wherethelastwhere the last2binsareaccountingfortheoutsideregionofthedetection.Thefirstoutsidebindefinestheneighborhoodofthedetection(bins are accounting for the outside region of the detection. The first outside bin defines the neighborhood of the detection (

Using a pair of unidirectional ALFDs, we define the ALFD as ρ(di,dj)=(ρ′(di,dj)+ρ′(dj,di)) / n(di,dj)\rho(d_{i},d_{j})=(\rho^{\prime}(d_{i},d_{j})+\rho^{\prime}(d_{j},d_{i}))\ /\ n(d_{i},d_{j}), where n(di,dj)n(d_{i},d_{j}) is a normalizer. The normalizer nn is defined as n(di,dj)=∣K(di,dj)∣+∣K(dj,di)∣+λn(d_{i},d_{j})=|\mathcal{K}(d_{i},d_{j})|+|\mathcal{K}(d_{j},d_{i})|+\lambda, where ∣K(⋅)∣|\mathcal{K}(\cdot)| is the count of IPTs and λ\lambda is a constant. λ\lambda ensures that the L1 norm of the ALFD increases as we have more supporting K(di,dj)\mathcal{K}(d_{i},d_{j}) and converges to 11. We use λ=20\lambda=20 in practice.

3 Learning the Model Weights

The algorithm computes a weighted average with a sign over all the ALFD patterns, where the weights are determined by the overlap between targets and detections. Intuitively, the ALFD pattern between detections that matches well with GT contributes more on the model parameters. The advantage of the weighted voting method is that each element in wΔtw_{\Delta t} are bounded in $,thustheALFDmetric,, thus the ALFD metric,a_{A}(d_{i},d_{j}),isalsoboundedby, is also bounded bysincesince||\rho(d_{i},d_{j})||_{1}\leq 1$. Fig. 4 shows two learned model using our method. One can adopt alternative learning algorithms like SVM .

4 Properties

In this section, we discuss the properties of ALFD affinity metric aA(di,dj)a_{A}(d_{i},d_{j}). Firstly, unlike appearance or spatial metrics, ALFD implicitly exploit the information in all the images between tit_{i} and tjt_{j} through IPTs. Secondly, thanks to the collective nature of ALFD design, it provides strong affinity metric over arbitrary length of time. We observe a significant benefit over the appearance or spatial metric especially over a long temporal distance (see Sec. 6.1 for the analysis). Thirdly, it is generally applicable to any scenarios (either static or moving camera) and for any object types (person or car). A disadvantage of the ALFD is that it may become unreliable when there is an occlusion. When an occlusion happens to a target, the IPTs initiated from the target tend to adhere to the occluder. It motivates us to combine target dynamics information discussed in Sec. 5.1.

Near Online Multi-target Tracking (NOMT)

where Ψ(⋅)\Psi(\cdot) encodes individual target’s motion, appearance, and ALFD metric consistency, and Φ(⋅)\Phi(\cdot) represent an exclusive relationship between different targets (e.g. no two targets share the same detection). If there are hypotheses for newly entering targets, we define the corresponding target as an empty set, Am∗t−1=\oA_{m}^{*t-1}=\o.

The potential measures the compatibility of a hypothesis Hm,xmtH_{m,x_{m}}^{t} to a target A_{m}^{*t\scalebox{0.5}[1.0]{-}1}. Mathematically, this can be decomposed into unary, pairwise and high order terms as follows:

ψu\psi_{u} encodes the compatibility of each detection did_{i} in the target hypothesis Hm,xmtH_{m,x_{m}}^{t} using the ALFD affinity metric and Target Dynamics feature (Sec. 5.1). ψp\psi_{p} measures the pairwise compatibility (self-consistency of the hypothesis) between detections within Hm,xmtH_{m,x_{m}}^{t} (Sec. 5.2) using the ALFD metric. Finally, ψh\psi_{h} implements a long-term smoothness constraint and appearance consistency (Sec. 5.3).

This potential penalizes choosing two targets with large overlap in the image plane (repulsive force) as well as duplicate assignments of a detection. Instead of using “hard” exclusion constraints as in the Hungarian Algorithm , we use “soft” cost function for flexibility and computational simplicity. If the single target consistency is strong enough, soft penalization cost could be overcome. Also, this formulation makes it possible to reuse popular graph inference algorithms discussed in Sec. 4.3. The potential can be written as follows:

2 Hypothesis Generation

3 Inference with Dynamic Graphical Model

Once we have all the hypotheses for all the new and existing targets, the problem (eq. 2) can be formulated as an inference problem with an undirected graphical model, where one node represents a target and the states are hypothesis indices as shown in Fig. 5 (c). The main challenges in this problem are: 1) there may exist loops in the graphical model representation and 2) the structure of graph is different depending on the hypotheses at each circumstance. In order to obtain the exact solution efficiently, we first analyze the structure of the graph on the fly and apply appropriate inference algorithms based on the structure analysis.

Given the graphical model, we find independent subgraphs (shown as dashed boxes in Fig. 5 (c)) using connected component analysis and perform individual inference algorithm per each subgraph in parallel. If a subgraph is composed of more than one node, we use junction-tree algorithm to obtain the solution for corresponding subgraph. Otherwise, we choose the best hypothesis for the target.

Model Details

In this section, we discuss the details of the potentials described in the Eq. 3.

As discussed in the previous sections, we utilize the ALFD metric as the main affinity metric to compare detections. The unary potential for each detection in the hypothesis is measured by:

where N\mathcal{N} is a predefined set of neighbor frame distances and d(A_{m}^{*t\scalebox{0.5}[1.0]{-}1},t_{i}) gives the associated detection of A_{m}^{*t\scalebox{0.5}[1.0]{-}1} at tit_{i}. Although we can define an arbitrarily large set of N\mathcal{N}, we choose N={1,2,5,10,20}\mathcal{N}=\{1,2,5,10,20\} for computational efficiency while modeling long term affinity measures.

Although ALFD metric provides very strong information in most of the cases, there are few failure cases including occlusions, erroneous IPTs, etc. To complement such cases, we design an additional Target Dynamics (TD) feature \mu_{T}(A_{m}^{*t\scalebox{0.5}[1.0]{-}1},d_{i}). Using the same polynomial least square predictor discussed in Sec. 4.2, we define the feature as follows:

where η\eta is a decay factor (0.980.98) that discounts long term prediction, f(A_{m}^{*t\scalebox{0.5}[1.0]{-}1}) denotes the last associated frame of A_{m}^{*t\scalebox{0.5}[1.0]{-}1}, o2o^{2} represents IoU2IoU^{2} discussed in the Sec. 4.1, and pp is the polynomial least square predictor described in Sec. 4.2.

Using the two measures, we define the unary potential \psi_{u}(A_{m}^{*t\scalebox{0.5}[1.0]{-}1},d_{i}) as:

where sis_{i} represents the detection score of did_{i}. The minmin operator enables us to utilize the ALFD metric in most cases, but activate the TD metric only when it is very confident (more than 0.50.5 overlap between the prediction and the detection). If A_{m}^{*t\scalebox{0.5}[1.0]{-}1} is empty, the potential becomes −si-s_{i}.

2 Pairwise potential

The pairwise potential ψp(⋅)\psi_{p}(\cdot) is solely defined by the ALFD metric. Similarly to the unary potential, we define the pairwise relationship between detections in Hm,xmtH_{m,x_{m}}^{t},

It measures the self-consistency of a hypothesis Hm,xmtH_{m,x_{m}}^{t}.

3 High-order potential

We incorporate a high-order potential to regularize the target association process with a physical feasibility and appearance similarity. Firstly, inspired by , we implement the physical feasibility by penalizing the hypotheses that present an abrupt motion. Secondly, we encodes long term appearance similarity between all the detections in A_{m}^{*t\scalebox{0.5}[1.0]{-}1} and Hm,xmtH_{m,x_{m}}^{t} similarly to . The intuition is encoded by the following potential:

where γ,ϵ,θ\gamma,\epsilon,\theta are scalar parameters, ξ(a,b)\xi(a,b) measures the sum of squared distances in (x,y,height)(x,y,height) of the two boxes, that is normalized by the mean height of pp in [t\scalebox{0.5}[1.0]{-}\tau,t], and K(di,dj)K(d_{i},d_{j}) represents the intersection kernel for color histograms associated with the detections. We use a pyramid of LAB color histogram where the first layer is the full box and the second layer is 3×33\times 3 grids. Only the A and B channels are used for the histogram with 44 bins per each channel (resulting in 4×4×(1+9)4\times 4\times(1+9) bins). We use (γ,ϵ,θ)=(20,0.4,0.8)(\gamma,\epsilon,\theta)=(20,0.4,0.8) in practice.

Experimental Evaluation

In order to evaluate the proposed algorithm, we use the KITTI object tracking benchmark and MOT challenge dataset . KITTI tracking benchmark is composed of about 19,00019,000 frames (∼32\sim 32 minutes). The dataset is composed of 2121 training and 2929 testing video sequences that are recorded using cameras mounted on top of a moving vehicle. Each video sequence has a variable number of frames from 7878 to 11761176 frames having a variable number of target objects (Car, Pedestrian, and Cyclist). The videos are recorded at 1010 FPS. The dataset is very challenging since 1) the scenes are crowded (occlusion and clutter), 2) the camera is not stationary, and 3) target objects appears in arbitrary location with variable sizes. Many conventional assumptions/techniques adopted in multiple target tracking with a surveillance camera is not applicable in this case (e.g. fixed entering/exiting location, background subtraction, etc). MOT challenge is composed of 11,28611,286 frames (∼16.5\sim 16.5 minutes) with varying FPS. The dataset is composed of 1111 training and 1111 testing video sequences. Some of the videos are recorded using mobile platform and the others are from surveillance videos. All the sequences contain only Pedestrians. As it is composed of videos with various configuration, tracking algorithms that are particularly tuned for a specific scenario would not work well in general. For the evaluation, we adopt the widely used CLEAR MOT tracking metrics . For a fair comparison to the other methods, we use the reference object detections provided by the both datasets.

We first run an ablative analysis on our ALFD affinity metric. We choose two sequences, KITTI’s 0001 and MOT’s PETS09-S2L1 both from the training sets, for the analysis. Given all the detections and the ground truth annotations, we first find the label association between detections and annotations. For each detection, we assign ground truth id if there is larger than 0.50.5 overlap. We collect all possible pairs of detections in 1,2,5,10,201,2,5,10,20 frame distance (Δt\Delta t), to obtain the positive and negative pairs. As the baseline affinity measures, we use the L2 distance between bottom center of the detections that is normalized by the mean height of the two (NDist2) and the intersection kernel between the color histograms of the two (HistIK). Fig. 6 and Table. 1 show the ROC curve and AUC of each affinity metric. We observe that ALFD affinity metric performs the best in all temporal distance regardless of the camera configuration and object type. As the temporal distance increases, the other metrics become quickly unreliable as expected, whereas our ALFD metric still provides strong cue to compare different detections.

2 KITTI Testing Benchmark Evaluation

As shown in the table, we observe that our algorithm (NOMT) outperforms the other state-of-the-art methods in most of the metrics with significant margins. Our method produces much larger numbers of mostly tracked targets (MT) in both Car and Pedestrian experiments with smaller numbers of mostly lost targets (ML). This is thanks to the highly accurate identity maintenance capability of our algorithm demonstrated in the low number of identity switch (IDS) and fragmentation (FRAG). In turn, our method achieves highest MOTA compared to other state-of-the-arts (>10%>10\% for Car and >8%>8\% for Pedestrian), which summarize all aspects of tracking evaluation. Notice that the higher tracking accuracy results in the higher detection accuracy as shown in Recall, Precision, and F1 metrics. Our own HM baseline also performs better than the other state-of-the-art methods, which demonstrates the robustness of ALFD metric. However, due to the nature of pure online association and lack of high order potential, it ends up missing more targets as shown in the MT and ML measures.

3 MOT Challenge Evaluation

Table. 3 summarizes the evaluation accuracy of our method (NOMT) and the other state-of-the-art algorithms on the MOT test video sequencesThe comparison is also available at http://nyx.ethz.ch/view_results.php?chl=2.. The website provides a set of reference detections obtained using .

Similarly to the KITTI experiment, we observe that our algorithm outperforms the other state-of-the-art methods with significant margins. Our method achieves the lowest identity switch and fragmentation while achieving the highest detection accuracy (lowest False Positives (FP) and False Negatives (FN)). In turn, our method records the highest MOTA compared to the other state-of-the-arts with a significant margin (>14%>14\%). The two experiments demonstrate that our ALFD metric and NOMT algorithm is generally applicable to any application scenario. Fig. 7 shows some qualitative examples of our results.

4 Timing Analysis

Our algorithm is not only highly accurate, but also very efficient. Leveraging on the parallel computation, we achieve a real-time efficiency (∼10FPS\sim 10FPS) using a 2.5GHz CPU with 16 cores. Table. 4 summarizes the time spent in each computational module.

Conclusion

In this paper, we propose a novel Aggregated Local Flow Descriptor that enables us to accurately measure the affinity between a pair of detections and a Near Online Muti-target Tracking that takes the advantages of both the pure online and global tracking algorithms. Our controlled experiment demonstrates that ALFD based affinity metric is significantly better than other conventional affinity metrics. Equipped with ALFD, our NOMT algorithm generates significantly better tracking results on two challenging large-scaler datasets. In addition, our method runs in real-time that enables us to apply the method in a variety of applications including autonomous driving, real-time surveillance, etc.

References