Kinematic 3D Object Detection in Monocular Video
Garrick Brazil, Gerard Pons-Moll, Xiaoming Liu, Bernt Schiele
Introduction
The detection of foreground objects is among the most critical requirements to facilitate self-driving applications . Recently, 3D object detection has made significant progress , even while using only a monocular camera . Such works primarily look at the problem from the perspective of single frames, ignoring useful temporal cues and constraints.
Computer vision cherishes inverse problems, e.g., recovering the 3D physical motion of objects from monocular videos. Motion information such as object velocity in the metric space is highly desirable for the path planning of self-driving. However, single image-based 3D object detection can not directly estimate physical motion, without relying on additional tracking modules. Therefore, video-based 3D object detection would be a sensible choice to recover such motion information. Furthermore, without modeling the physical motion, image-based 3D object detectors are naturally more likely to suffer from erratic and unnatural changes through time in orientation and localization (as exemplified in Fig. 1(a)). Therefore, we aim to build a novel video-based 3D object detector which is able to provide accurate and smooth 3D object detection with per-object velocity, while also prioritizing a compact and efficient model overall.
Yet, designing an effective video-based 3D object detector has challenges. Firstly, motion which occurs in real-world scenes can come from a variety of sources such as the camera atop of an autonomous vehicle or robot, and/or from the scene objects themselves — for which most of the safety-critical objects (car, pedestrian, cyclist ) are typically dynamic. Moreover, using video inherently involves an increase in data consumption which introduces practical challenges for training and/or inference including with memory or redundant processing.
To address such challenges, we propose a novel framework to integrate a 3D Kalman filter into a 3D detection system. We find Kalman is an ideal candidate for three critical reasons: (1) it allows for use of real-world motion models to serve as a strong prior on object dynamics, (2) it is inherently efficient due to its recursive nature and general absence of parameters, (3) the resultant behavior is explainable and provides useful by-products such as the object velocity.
Furthermore, we observe that objects predominantly move in the direction indicated by their orientation. Fortunately, the benefit of Kalman allows us to integrate this real-world constraint into the motion model as a compact scalar velocity. Such a constraint helps maintain the consistency of velocity over time and enables the Kalman motion forecasting and fusion to perform accurately.
However, a model restricted to only move in the direction of its orientation has an obvious flaw — what if the orientation itself is inaccurate? We therefore propose a novel reformulation of orientation in favor of accuracy and stability. We find that our orientation improves the 3D localization accuracy by a margin of and reduces the orientation error by , which collectively help enable the proposed Kalman to function more effectively.
A notorious challenge of using Kalman comes in the form of uncertainty, which is conventionally assumed to be known and static, e.g., from a sensor. However, 3D objects in video are intuitively dependent on more complex factors of image features and cannot necessarily be treated like a sensor measurement. For a better understanding of 3D uncertainty, we propose a 3D self-balancing confidence loss. We show that our proposed confidence has higher correlation with the 3D localization performance compared to the typical classification probability, which is commonly used in detection .
To complete the full understanding of the scene motion, we elect to estimate the ego-motion of the capturing camera itself. Hence, we further narrow the work of Kalman to account for only the object’s motion. Collectively, our proposed framework is able to model important scene dynamics, both ego-motion and per-object velocity, and more precisely detect 3D objects in videos using a stabilized orientation and 3D confidence estimation. We demonstrate that our method achieves state-of-the-art (SOTA) performance on monocular 3D Object Detection and Bird’s Eye View (BEV) tasks in the KITTI dataset .
In summary, our contributions are as follows:
We propose a monocular video-based 3D object detector, leveraging realistic motion constraints with an integrated ego-motion and a 3D Kalman filter.
We propose to reformulate orientation into axis, heading and offset along with a self-balancing 3D localization loss to facilitate the stability necessary for the proposed Kalman filter to perform more effectively.
Overall, using only a single model our framework develops a comprehensive 3D scene understanding including object cuboids, orientation, velocity, object motion, uncertainty, and ego-motion, as detailed in Fig. 1 and 2.
We achieve a new SOTA performance on monocular 3D object detection and BEV tasks using comprehensive metrics within the KITTI dataset.
Related Work
We first provide context of our novelties from the perspective of monocular 3D object detection (Sec. 2.1) with attention to orientation and uncertainty estimation. We next discuss and contrast with video-based object detection (Sec. 2.2).
Monocular 3D object detection has made significant progress . Early methods such as began by generating 3D object proposals along a ground plane using object priors and estimated point clouds, culminating in an energy minimization approach. utilize additional domains of semantic segmentation, object priors, and estimated depth to improve the localization. Similarly, create a pseudo-LiDAR map using SOTA depth estimator , which is respectfully passed into detection subnetworks or LiDAR-based 3D object detection works . In strong 2D detection systems are extended to add cues such as object orientation, then the remaining 3D box parameters are solved via 3D box geometry. extends the region proposal network (RPN) of Faster R-CNN with 3D box parameters.
Prior monocular 3D object detectors estimate orientation via two main techniques. The first method is to classify orientation via a series of discrete bins then regress a relative offset . The bin technique requires a trade-off between the quantity/coverage of the discretized angles and an increase in the number of estimated parameters (bin ). Other methods directly regress the orientation angle using quaternion or Euler angles. Direct regression is comparatively efficient, but may lead to degraded performance and periodicity challenges , as exemplified in Fig. 1.
In contrast, we propose a novel orientation decomposition which serves as an intuitive compromise between the bin and direct approaches. We decompose the orientation estimation into three components: axis and heading classification, followed by an angle offset. Thus, our technique increases the parameters by a static factor of compared to a bin hyperparameter, while drastically reducing the offset search space for each orientation estimation (discussed in Sec. 3.1).
1.2 Uncertainty Estimation:
Although it is common to utilize the classification score to rate boxes in 2D object detection or explicitly model uncertainty as parametric estimation , prior works in monocular 3D object detection realize the need for 3D box uncertainty/confidence . defines confidence using the 3D IoU of a box and ground truth after center alignment, thus capturing the confidence primarily of the 3D object dimensions. predicts a confidence by re-mapping the 3D box loss into a probability range, which intuitively represents the confidence of the overall 3D box accuracy.
In contrast, our self-balancing confidence loss is generic and self-supervised, with two benefits. (1) It enables estimation of a 3D localization confidence using only the loss values, thus being more general than 3D IoU. (2) It enables the network to naturally re-balance extremely hard 3D boxes and focus on relatively achievable samples. Our ablation (Sec. 4.4) shows the importance of both effects.
2 Video-based Object Detection
Video-based object detection is generally less studied than single-frame object detection . A common trend in video-based detection is to benefit the accuracy-efficiency trade-off via reducing the frame redundancy . Such works are applied primarily on domains of ImageNet VID 2015 , which contain less ego-motion from the capturing camera than self-driving scenarios . As such, the methods are designed to use 2D transformations, which lack the consistency and realism of 3D motion modeling.
In comparison, to our knowledge this is the first work that utilizes video cues to improve the accuracy and robustness of monocular 3D object detection. In the domain of 2D/3D object tracking, experiments using Kalman Filters, Particle Filters, and Gaussian Mixture Models, and observe Kalman to be the most effective aggregation method for tracking. An LSTM with depth ordering and IMU camera ego-motion is utilized in to improve the tracking accuracy. In contrast, we explore how to naturally and effectively leverage a 3D Kalman filter to improve the accuracy and robustness of monocular 3D object detection. We propose novel enhancements including estimating ego-motion, orientation, and a 3D confidence, while efficiently using only a single model.
Methodology
Our proposed kinematic framework is composed of three primary components: a 3D region proposal network (RPN), ego-motion estimation, and a novel kinematic model to take advantage of temporal motion in videos. We first overview the foundations of a 3D RPN. Then we detail our contributions of orientation decomposition and self-balancing 3D confidence, which are integral to the kinematic method. Next we detail ego-motion estimation. Lastly, we present the complete kinematic framework (Fig. 2) which carefully employs a 3D Kalman to model realistic motion using the aforementioned components, ultimately producing a more accurate and comprehensive 3D scene understanding.
Our measurement model is founded on the 3D RPN , enhanced using novel orientation and confidence estimations. The RPN itself acts as a sliding window detector following the typical practices outlined in Faster R-CNN and . Specifically, the RPN consists of a backbone network and a detection head which predicts 3D box outputs relative to a set of predefined anchors.
1.2 3D Box Outputs:
Lastly, the regression targets for 3D dimensions GTs are defined as:
The remaining targets for our novel orientation estimation , , and 3D self-balancing confidence are defined in subsequent sections.
1.3 Orientation Estimation:
We propose a novel object orientation formulation, with a decomposition of three components: axis, heading, and offset (Fig. 3). Intuitively, the axis estimation represents the probability an object is oriented towards the vertical axis () or the horizontal axis (), with its label formally defined as: , where is the ground truth object orientation in radians from a bird’s eyes view (BEV) with bounded range.We then compute an orientation with a restricted range relative to its axis, e.g., when , and when . We start with then add or subtract from until the desired range is satisfied.
Intuitively, loses its heading since the true rotation may be . We therefore preserve the heading using a separate , which represents the probability of being rotated by with its GT target defined as:
Lastly, we encode the orientation offset transformation which is relative to the corresponding anchor, axis, and restricted orientation as: . The reverse decomposition is where denotes round.
In designing our orientation, we first observed that the visual difference between objects at opposite headings of is low, especially for far objects. In contrast, classifying the axis of an object is intuitively more clear since the visual features correlate with the aspect ratio. Note that disentangle these two objectives. Hence, while the axis is being determined, the heading classifier can focus on subtle clues such as windshields, headlights and shape.
We further note that our binary classifications have the same representational power as bins following . Specifically, bins of . However, it is more common to use considerably more bins (such as in ). An important distinction is that bin-based approaches require the network decide axis and heading simultaneously, whereas our method disentangles the orientation into the two distinct and explainable objectives. We provide ablations to compare our decomposition and the bin method using $$ bins in Sec. 4.4.
1.4 Self-Balancing Loss:
The novel 3D localization confidence follows a self-balancing formulation closely coupled to the network loss. We first define the 2D and 3D loss terms which comprise the general RPN loss. We unroll and match all box estimations to their respective ground truths. A box is matched as foreground when sufficient () 2D intersection over union (IoU) is met, otherwise it is considered background () and all loss terms except for classification are ignored. The 2D box loss is thus defined as:
where CE denotes a softmax activation followed by logistic cross-entropy loss over the ground truth class , and IoU uses predicted and ground truth . Similarly, the 3D localization loss for only foreground () is defined as:
where BCE denotes a sigmoid activation followed by binary cross-entropy loss. Next we define the final self-balancing confidence loss with the estimation as:
where is the rolling mean of the most recent losses per mini-batch. Since is predicted per-box via a sigmoid, the network can intuitively balance whether to use the loss of or incur a proportional penalty of . Hence, when the confidence is high () we infer that the network is confident in its 3D loss . Conversely, when the confidence is low (), the network is uncertain in , thus incurring a flat penalty is preferred. At inference, we fuse the self-balancing confidence with the classification score as .
The proposed self-balancing loss has two key benefits. Firstly, it produces a useful 3D localization confidence with inherent correlation to 3D IoU (Sec. 4.4). Secondly, it enables the network to re-balance samples which are exceedingly challenging and re-focus on the more reasonable targets. Such a characteristic can be seen as the inverse of hard-negative mining, which is important while monocular 3D object detection remains highly difficult and unsaturated (Sec. 4.1).
2 Ego-motion
A challenge with the dynamics of urban scenes is that not only are most foreground objects in motion, but the capturing camera itself is dynamic. Therefore, for a full understanding of the scene dynamics, we design our model to additionally predict the self-movement of the capturing camera, e.g., ego-motion.
3 Kinematics
In order to leverage temporal motion in video, we elect to integrate our RPN and ego-motion into a novel kinematic model. We adopt a 3D Kalman due to its notable efficiency, effectiveness, and interpretability. We next detail our proposed motion model, the procedure for forecasting, association, and update.
3.2 Forecasting:
The forecasting step aims to utilize the tracked state variables and covariances of time to estimate the state of a future time . The equation to forecast a state variable into is: , where is the state transition model at . Note that both objects and the capturing camera may have independent motion between consecutive frames. Therefore, we lastly apply the estimated ego-motion to all available tracks’ 3D center by:
where denotes the average self-balancing confidence of a track’s life. Hence, the resultant track states and track covariances represent the Kalman filter’s best forecasted estimation with respect to frame .
3.3 Association:
3.4 Update:
After making associations between tracks and measurements , the next step is to utilize the track covariance and measured confidence to update each track to its final state and covariance . Firstly, we formally define the equation for computing the Kalman gain as:
where represents the incoming measurement covariance matrix, and the forecasted covariance of the track. Next, given the Kalman gain , forecasted state , forecasted covariance , and measured box , the final track state and covariance are defined as:
We lastly aggregate each track’s overall confidence over time as a running average of , where is the measured confidence.
4 Implementation Details
Our framework is implemented in PyTorch , with the 3D RPN settings of .We release source code at http://cvlab.cse.msu.edu/project-kinematic.html. We use a batch size of and learning rate of . We set , , , , , , , and .To ease training, we implement three phases. We first train the 2D-3D RPN with , then the self-balancing loss of Eq. 7, for and iterations. We freeze the RPN to train ego-motion using for . Our backbone is DenseNet121 where . Inference uses frames as provided by .
Experiments
We benchmark our kinematic framework on the KITTI dataset. We comprehensively evaluate on 3D Object Detection and Bird’s Eye View (BEV) tasks.We then provide ablation experiments to better understand the effects and justification of our core methodology. We show qualitative examples in Fig. 6.
The KITTI dataset is a popular benchmark for self-driving tasks. The official dataset consists of training and testing images including annotations for 2D/3D objects, ego-motion, and temporally adjacent frames. We evaluate on the most widely used validation split as proposed in , which consists of training and validation images. We focus primarily on the car class.
Average precision (AP) is utilized for object detection in KITTI. Following , the KITTI metric has updated to include ( ) recall points while skipping the first. The AP40 metric is more stable and fair overall . Due to the official adoption of AP40, it is not possible to compute AP11 on test. Hence, we elect to use the AP40 metric for all reported experiments.
2 3D Object Detection
We evaluate our proposed framework on the task of 3D object detection, which requires objects be localized in 3D camera coordinates as well as supplying the 3D dimensions and BEV orientation relative to the XZ plane. Due to the strict requirements of IoU in three dimensions, the task demands precise localization of an object to be considered a match (3D IoU ). We evaluate our performance on the official test dataset in Tab. 1 and the validation split in Tab. 2.
We emphasize that our method improves the SOTA on KITTI test by a significant margin of compared to on the moderate configuration with IoU , which is the most common metric used to compare. Further, we note that require multiple encoder-decoder networks which add overhead compared to our single network approach. Hence, their runtime is (Tab. 1) compared to ours, self-reported on similar but not identical GPU hardware. Moreover, is the most comparable method to ours as both utilize a single network and an RPN archetype. We note that our method significantly outperforms and many other recent works by .
We further evaluate our approach on the KITTI validation split using the AP40 for available approaches and observe similar overall trends as in Tab. 2. For instance, compared to competitive approaches our method improves the performance by for the challenging IoU criteria of . Similarly, our performance on the more relaxed criteria of IoU increases by . We additionally visualize detailed performance characteristics on AP at discrete depth meters and IoU matching criterias in Fig. 5.
3 Bird’s Eye View
The Bird’s Eye View (BEV) task is similar to 3D object detection, differing primarily in that the 3D boxes are firstly projected into the XZ plane then 2D object detection is calculated. The projection collapses the Y-axis degree of freedom and intuitively results in a less precise but reasonable localization.
We note that our method achieves SOTA performance on the BEV task regarding the moderate setting of the KITTI test dataset as detailed in Tab. 1. Our method performs favorably compared with SOTA works (e.g., ), and similarly to at a notably lower runtime cost. We suspect that our method, especially the self-balancing confidence (Eq. 7), prioritizes precise localization which warrants more benefit in full 3D Object Detection task compared to the Bird’s Eye View task.
Our method performs similarly on the validation split of KITTI (Tab. 2). Specifically, compared to our proposed method outperforms by a range of , which is consistent to the same methods on test .
4 Ablation Study
To better understand the characteristics of our proposed kinematic framework, we perform a series of ablation experiments and analysis, summarized in Tab. 3. We adopt without hill-climbing or depth-aware layers as our baseline method. Unless otherwise specified we use the experimental settings outlined in Sec. 3.4.
The orientation of objects is intuitively a critical component when modeling motion. When the orientation is decomposed into axis, heading, and offset the overall performance significantly improves, e.g., by in AP and in AP, as detailed within Tab. 3. We compute the mean angle error of our baseline, orientation decomposition, and kinematics method which respectively achieve , , and (), suggesting our proposed methodology is significantly more stable.
We compare our orientation decomposition to bin-based methods following general idea of . We specifically change our orientation definition into which includes a bin classification and an offset. We experiment with the number of bins set to which are uniformly spread from . Note that bins have the same representational power as using binary . We observe that the ablated bin-based methods achieve in AP. In comparison, our decomposed orientation achieves in AP. We provide additional detailed experiments in our supplemental material.
Further, we find that our proposed kinematic motion model (as in Sec. 3.3) degrades in performance when a comparatively erratic baseline (Row 1. Tab. 3) orientation is utilized instead ( on AP), reaffirming the importance of having a consistent/stable orientation when aggregating through time.
4.2 Self-balancing Confidence:
We observe that the self-balancing confidence is important from two key respects. Firstly, its integration in Eq. 7 enables the network to re-weight box samples to focus more on reasonable samples and incur a flat penalty (e.g., of Eq. 7) on the difficult samples. In a sense, the self-balancing confidence loss is the inverse of hard-negative mining, allowing the network to focus on reasonable estimations. Hence, the loss on its own improves performance for AP by and AP by .
The second benefit of self-balancing confidence is that by design has an inherent correlation with the 3D object detection performance. Recall that we fuse with the classification score to produce a final box rating of , which results in an additional gain of in AP and in AP. We further analyze the correlation of with 3D IoU, as is summarized in Fig. 5. The correlation coefficient with the classification score is significantly lower than the correlation using instead ( vs. ). In summary, the use of the Eq. 7 and account for a gain of in AP and in AP.
4.3 Temporal Modeling:
The use of video and kinematics is a significant motivating factor for this work. We find that the use of kinematics (detailed in Sec. 3.3) results in a gain of in AP and in AP, as shown in Tab. 3. We emphasize that although the improvement is less dramatic versus orientation and self-confidence, the former are important to facilitate temporal modeling. We find that if orientation decomposition and uncertainty are removed, by using the baseline orientation and setting to be a static constant, then the kinematic performance drastically reduces from in AP.
We emphasize that kinematic framework not only helps 3D object detection, but also naturally produces useful by-products such as velocity and ego-motion. Thus, we evaluate the respective average errors of each motion after applying the camera capture rate to convert the motion into miles per hour (MPH). We find that the per-object velocity and ego-motion speed errors perform reasonably at MPH and MPH respectively. We depict visual examples of all dynamic by-products in Fig. 6 and additionally in supplemental video.
Conclusions
We present a novel kinematic 3D object detection framework which is able to efficiently leverage temporal cues and constraints to improve 3D object detection. Our method naturally provides useful by-products regarding scene dynamics, e.g., reasonably accurate ego-motion and per-object velocity. We further propose novel designs of orientation estimation and a self-balancing 3D confidence loss in order to enable the proposed kinematic model to work effectively. We emphasize that our framework efficiently uses only a single network to comprehensively understand a highly dynamic 3D scene for urban autonomous driving. Moreover, we demonstrate our method’s effectiveness through detailed experiments on the KITTI dataset across the 3D object detection and BEV tasks.
Acknowledgments: Research was partially sponsored by the Army Research Office under Grant Number W911NF-18-1-0330. The views and conclusions contained in this document are those of the authors and should not be interpreted as representing the official policies, either expressed or implied, of the Army Research Office or the U.S. Government. The U.S. Government is authorized to reproduce and distribute reprints for Government purposes notwithstanding any copyright notation herein. This work is further partly funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) - 409792180 (Emmy Noether Programme, project: Real Virtual Humans).
Supplementary Material: Kinematic 3D Object Detection in Monocular Video
Garrick Brazil Gerard Pons-Moll Xiaoming Liu Bernt Schiele
Orientation Ablations
We provide detailed experiments on 3D object detection and Bird’s Eye View tasks to compare our orientation decomposition performance with bin-based approaches such as within Tab. 1. Recall that bin-based orientation first classifies the best bin for orientation then predicts an offset with respect to the bin. In contrast, our method disentangles the bin classification into a distinct explainable objectives such as an axis classification and a heading classification. For such experiments we change our formulation to use bins of , where bins has a similar representational power as two binary classifications . The bins are spread uniformly from and an offset is predicted afterwards. We use the settings in Sec. in main paper. We emphasize that our method outperforms the bin-based approaches between on AP and on AP using the standard moderate setting and IoU.
Kalman Forecasting
Since our method uses ego-motion and a 3D Kalman filter to aggregate temporal information, the approach can be modified to act as a box forecaster. Although our method was not strictly designed for the tracking and forecasting task, we evaluate the 3D object detection and Bird’s Eye View performance after forecasting frames into the future. We assume a static ego-motion for unknown frames and otherwise use the Kalman equations described in the main paper Sec. 3.3 to forecast the tracked boxes.
For all forecasting experiments we process temporally adjacent frames before forecasting. Since KITTI only provides a current frame and proceeding frames, we carefully map images back to the raw dataset in order to forecast. For instance, when we infer using frames $n_{f}{}_{\text{3D}}{}_{\text{BEV}}1-211210.64\%5.10\%{}_{\text{3D}}$ respectively, which are competitive to methods on test.
Qualitative Video
We further provide a qualitative demonstration video at http://cvlab.cse.msu.edu/project-kinematic.html. The video demonstrates our framework’s ability to determine a full scene understanding including 3D object cuboids, per-object velocity and ego-motion. We compare to a related monocular work of M3D-RPN , plot ground truths, image view, Bird’s Eye View, and the track history.