Orthographic Feature Transform for Monocular 3D Object Detection
Thomas Roddick, Alex Kendall, Roberto Cipolla
Introduction
The success of any autonomous agent is contingent on its ability to detect and localize the objects in its surrounding environment. Prediction, avoidance and path planning all depend on robust estimates of the 3D positions and dimensions of other entities in the scene. This has led to 3D bounding box detection emerging as an important problem in computer vision and robotics, particularly in the context of autonomous driving. To date the 3D object detection literature has been dominated by approaches which make use of rich LiDAR point clouds , while the performance of image-only methods, which lack the absolute depth information of LiDAR, lags significantly behind. Given the high cost of existing LiDAR units, the sparsity of LiDAR point clouds at long ranges, and the need for sensor redundancy, accurate 3D object detection from monocular images remains an important research objective. To this end, we present a novel 3D object detection algorithm which takes a single monocular RGB image as input and produces high quality 3D bounding boxes, achieving state-of-the-art performance among monocular methods on the challenging KITTI benchmark .
Images are, in many senses, an extremely challenging modality. Perspective projection implies that the scale of a single object varies considerably with distance from the camera; its appearance can change drastically depending on the viewpoint; and distances in the 3D world cannot be inferred directly. These factors present enormous challenges to a monocular 3D object detection system. A far more innocuous representation is the orthographic birds-eye-view map commonly employed in many LiDAR-based methods . Under this representation, scale is homogeneous; appearance is largely viewpoint-independent; and distances between objects are meaningful. Our key insight therefore is that as much reasoning as possible should be performed in this orthographic space rather than directly on the pixel-based image domain. This insight proves essential to the success of our proposed system.
It is unclear, however, how such a representation could be constructed from a monocular image alone. We therefore introduce the orthographic feature transform (OFT): a differentiable transformation which maps a set of features extracted from a perspective RGB image to an orthographic birds-eye-view feature map. Crucially, we do not rely on any explicit notion of depth: rather our system builds up an internal representation which is able to determine which features from the image are relevant to each location on the birds-eye-view. We apply a deep convolutional neural network, the topdown network, in order to reason locally about the 3D configuration of the scene.
The main contributions of our work are as follows:
We introduce the orthographic feature transform (OFT) which maps perspective image-based features into an orthographic birds-eye-view, implemented efficiently using integral images for fast average pooling.
We describe a deep learning architecture for predicting 3D bounding boxes from monocular RGB images.
We highlight the importance of reasoning in 3D for the object detection task.
The system is evaluated on the challenging KITTI 3D object benchmark and achieves state-of-the-art results among monocular approaches.
Related Work
Detecting 2D bounding boxes in images is a widely studied problem and recent approaches are able to excel even on the most formidable datasets . Existing methods may broadly be divided into two main categories: single stage detectors such as YOLO , SSD and RetinaNet which predict object bounding boxes directly and two-stage detectors such as Faster R-CNN and FPN which add an intermediate region proposal stage. To date the vast majority of 3D object detection methods have adopted the latter philosophy, in part due to the difficulty in mapping from fixed-sized regions in 3D space to variable-sized regions in the image space. We overcome this limitation via our OFT transform, allowing us to take advantage of the purported speed and accuracy benefits of a single-stage architecture.
3D object detection is of considerable importance to autonomous driving, and a large number of LiDAR-based methods have been proposed which have enjoyed considerable success. Most variation arises from how the LiDAR point clouds are encoded. The Frustrum-PointNet of Qi et al. and the work of Du et al. operate directly on the point clouds themselves, considering a subset of points which lie within a frustrum defined by a 2D bounding box on the image. Minemura et al. and Li et al. instead project the point cloud onto the image plane and apply Faster-RCNN-style architectures to the resulting RGB-D images. Other methods, such as TopNet , BirdNet and Yu et al. , discretize the point cloud into some birds-eye-view (BEV) representation which encodes features such as returned intensity or average height of points above the ground plane. This representation turns out to be extremely attractive since it does not exhibit any of the perspective artifacts introduced in RGB-D images for example, and a major focus of our work is therefore to develop an implicit image-only analogue to these birds-eye-view maps. A further interesting line of research is sensor fusion methods such as AVOD and MV3D which make use of 3D object proposals on the ground plane to aggregate both image-based and birds-eye-view features: an operation which is closely related to our orthographic feature transform.
Obtaining 3D bounding boxes from images, meanwhile, is a much more challenging problem on account of the absence of absolute depth information. Many approaches start from 2D bounding boxes extracted using standard detectors described above, upon which they either directly regress 3D pose parameters for each region or fit 3D templates to the image . Perhaps most closely related to our work is Mono3D which densely spans the 3D space with 3D bounding box proposals and then scores each using a variety of image-based features. Other works which explore the idea of dense 3D proposals in the world space are 3DOP and Pham and Jeon , which rely on explicit estimates of depth using stereo geometry. A major limitation of all the above works is that each region proposal or bounding box is treated independently, precluding any joint reasoning about the 3D configuration of the scene. Our method performs a similar feature aggregation step to , but applies a secondary convolutional network to the resulting proposals whilst retaining their spatial configuration.
Integral images have been fundamentally associated with object detection ever since their introduction in the seminal work of Viola and Jones . They have formed an important component in many contemporary 3D object detection approaches including AVOD , MV3D , Mono3D and 3DOP . In all of these cases however, integral images do not backpropagate gradients or form part of a fully end-to-end deep learning architecture. To our knowledge, the only prior work to do so is that of Kasagi et al. , which combines a convolutional layer and an average pooling layer to reduce computational cost.
3D Object Detection Architecture
In this section we describe our full approach for extracting 3D bounding boxes from monocular images. An overview of the system is illustrated in Figure 3. The algorithm comprises five main components:
A front-end ResNet feature extractor which extracts multi-scale feature maps from the input image.
A orthographic feature transform which transforms the image-based feature maps at each scale into an orthographic birds-eye-view representation.
A topdown network, consisting of a series of ResNet residual units, which processes the birds-eye-view feature maps in a manner which is invariant to the perspective effects observed in the image.
A set of output heads which generate, for each object class and each location on the ground plane, a confidence score, position offset, dimension offset and a orientation vector.
A non-maximum suppression and decoding stage, which identifies peaks in the confidence maps and generates discrete bounding box predictions.
The remainder of this section will describe each of these components in detail.
The first element of our architecture is a convolutional feature extractor which generates a hierarchy of multi-scale 2D feature maps from the raw input image. These features encode information about low-level structures in the image, which form the basic components used by the topdown network to construct an implicit 3D representation of the scene. The front-end network is also responsible for inferring depth information based on the size of image features since subsequent stages of the architecture aim to eliminate variance to scale.
2 Orthographic feature transform
In order to reason about the 3D world in the absence of perspective effects, we must first apply a mapping from feature maps extracted in the image space to orthographic feature maps in the world space, which we term the Orthographic Feature Transform (OFT).
where is the camera focal length and the principle point.
We can then assign a feature to the appropriate location in the voxel feature map by average pooling over the projected voxel’s bounding box in the image feature map :
Transforming to an intermediate voxel representation before collapsing to the final orthographic feature map has the advantage that the information about the vertical configuration of the scene is retained. This turns out to be essential for downstream tasks such as estimating the height and vertical position of object bounding boxes.
A major challenge with the above approach is the need to aggregate features over a very large number of regions. A typical voxel grid setting generates around 150k bounding boxes, which far exceeds the 2k regions of interest used by the Faster R-CNN architecture, for example. To facilitate pooling over such a large number of regions, we make use of a fast average pooling operation based on integral images . An integral image, or in this case integral feature map, , is constructed from an input feature map using the recursive relation
Given the integral feature map , the output feature corresponding to the region defined by bounding box coordinates and (see Equation 1), is given by
The complexity of this pooling operation is independent of the size of the individual regions, which makes it highly appropriate for our application where the size and shape of the regions varies considerably depending on whether the voxel is close to or far from the camera. It is also fully differentiable in terms of the original feature map and so can be used as part of an end-to-end deep learning framework.
3 Topdown network
A core contribution of this work is to emphasize the importance of reasoning in 3D for object recognition and detection in complex 3D scenes. In our architecture, this reasoning component is performed by a sub-network which we term the topdown network. This is a simple convolutional network with ResNet-style skip connections which operates on the the 2D feature maps generated by the previously described OFT stage. Since the filters of the topdown network are applied convolutionally, all processing is invariant to the location of the feature on the ground plane. This means that feature maps which are distant from the camera receive exactly the same treatment as those that are close, despite corresponding a much smaller region of the image. The ambition is that the final feature representation will therefore capture information purely about the underlying 3D structure of the scene and not its 2D projection.
4 Confidence map prediction
Among both 2D and 3D approaches, detection is conventionally treated as a classification problem, with a cross entropy loss used to identify regions of the image which contain objects. In our application however we found it to be more effective to adopt the confidence map regression approach of Huang et al. . The confidence map is a smooth function which indicates the probability that there exists an object with a bounding box centred on location , where is the distance of the ground plane below the camera. Given a set of ground truth objects with bounding box centres , we compute the ground truth confidence map as a smooth Gaussian region of width around the center of each object. The confidence at location is given by
5 Localization and bounding box estimation
The confidence map encodes a coarse approximation of the location of each object as a peak in the confidence score, which gives a position estimate accurate up to the resolution of the feature maps. In order to localize each object more precisely, we append an additional network output head which predicts the relative offset from grid cell locations on the ground plane to the center of the corresponding ground truth object :
We use the same scale factor as described in Section 3.4 to normalize the position offsets within a sensible range. A ground truth object instance is assigned to a grid location if any part of the object’s bounding box intersects the given grid cell. Cells which do not intersect any ground truth objects are ignored during training.
In addition to localizing each object, we must also determine the size and orientation of each bounding box. We therefore introduce two further network outputs. The first, the dimension head, predicts the logarithmic scale offset between the assigned ground truth object with dimensions and the mean dimensions over all objects of the given class.
The second, the orientation head, predicts the sine and cosine of the objects orientation about the y-axis:
6 Non-maximum suppression
Similarly to other object detection algorithms, we apply a non-maximum suppression (NMS) stage to obtain a final discrete set of object predictions. In a conventional object detection setting this step can be expensive since it requires bounding box overlap computations. This is compounded by the fact that pairs of 3D boxes are not necessarily axis aligned, which makes the overlap computation more difficult compared to the 2D case. Fortunately, an additional benefit of the use of confidence maps in place of anchor box classification is that we can apply NMS in the more conventional image processing sense, i.e. searching for local maxima on the 2D confidence maps . Here, the orthographic birds-eye-view again proves invaluable: the fact that two objects cannot occupy the same volume in the 3D world means that peaks on the confidence maps are naturally separated.
To alleviate the effects of noise in the predictions, we first smooth the confidence maps by applying a Gaussian kernel with width . A location on the smoothed confidence map is deemed to be a maximum if
Of the produced peak locations, any with a confidence smaller than a given threshold are eliminated. This results in the final set of predicted object instances, whose bounding box center , dimensions , and orientation , are given by inverting the relationships in Equations 7, 8 and 9 respectively.
Experiments
For our front-end feature extractor we make use of a ResNet-18 network without bottleneck layers. We intentionally choose the front-end network to be relatively shallow, since we wish to put as much emphasis as possible on the 3D reasoning component of the model. We extract features immediately before the final three downsampling layers, resulting in a set of feature maps at scales of 1/8, 1/16 and 1/32 of the original input resolution. Convolutional layers with 11 kernels are used to map these feature maps to a common feature size of 256, before processing them via the orthographic feature transform to yield orthographic feature maps . We use a voxel grid with dimensions 80m4m80m, which is sufficient to include all annotated instances in KITTI, and set the grid resolution to be 0.5m. For the topdown network, we use a simple 16-layer ResNet without any downsampling or bottleneck units. The output heads each consist of a single 11 convolution layer. Throughout the model we replace all batch normalization layers with group normalization which has been found to perform better for training with small batch sizes.
We train and evaluate our method using the KITTI 3D object detection benchmark dataset . For all experiments we follow the train-val split of Chen et al. which divides the KITTI training set into 3712 training images and 3769 validation images.
Since our method relies on a fixed mapping from the image plane to the ground plane, we found that extensive data augmentation was essential for the network to learn robustly. We adopt three types of widely-used augmentations: random cropping, scaling and horizontal flipping, adjusting the camera calibration parameters and accordingly to reflect these perturbations.
The model is trained using SGD for 600 epochs with a batch size of 8, momentum of 0.9 and learning rate of . Following , losses are summed rather than averaged, which avoids biasing the gradients towards examples with few object instances. The loss functions from the various output heads are combined using a simple equal weighting strategy.
2 Comparison to state-of-the-art
We evaluate our approach on two tasks from the KITTI 3D object detection benchmark. The 3D bounding box detection task requires that each predicted 3D bounding box should intersect a corresponding ground truth box by at least 70% in the case of cars and 50% for pedestrians and cyclists. The birds-eye-view detection task meanwhile is slightly more lenient, requiring the same amount of overlap between a 2D birds-eye-view projection of the predicted and ground truth bounding boxes on the ground plane. At the time of writing, the KITTI benchmark included only one published approach operating on monocular RGB images alone (), which we compare our method against in Table 1. We therefore perform additional evaluation on the KITTI validation split set out by Chen et al. (2016) ; the results of which are presented in Table 2. For monocular methods, performance on the pedestrian and cyclist classes is typically insufficient to obtain meaningful results and we therefore follow other works and focus our evaluation on the car class only.
It can be seen from Tables 1 and 2 that our method is able to outperform all comparable (i.e. monocular only) methods by a considerable margin across both tasks and all difficulty criteria. The improvement is particularly marked on the hard evaluation category, which includes instances which are heavily occluded, truncated or far from the camera. We also show in Table 2 that our method performs competitively with the stereo approach of Chen et al. (2015) , achieving close to or in one case better performance than their 3DOP system. This is in spite of the fact that unlike , our method does not have access to any explicit knowledge of the depth of the scene.
3 Qualitative results
We provide a qualitative comparison of predictions generated by our approach and Mono3D in Figure LABEL:fig:mono3d. A notable observation is that our system is able to reliably detect objects at a considerable distance from the camera. This is a common failure case among both 2D and 3D object detectors, and indeed many of the cases which are correctly identified by our system are overlooked by Mono3D. We argue that this ability to recognise objects at distance is a major strength of our system, and we explore this capacity further in Section 5.1. Further qualitative results are included in supplementary material.
A unique feature of our approach is that we operate largely in the orthographic birds-eye-view feature space. To illustrate this, Figure LABEL:fig:heatmaps shows examples of predicted confidence maps both in the topdown view and projected into the image on the ground plane. It can be seen that the predicted confidence maps are well localized around each object center.
4 Ablation study
A central claim of our approach is that reasoning in the orthographic birds-eye-view space significantly improves performance. To validate this claim, we perform an ablation study where we progressively remove layers from the topdown network. In the extreme case, when the depth of the topdown network is zero, the architecture is effectively reduced to RoI pooling over projected bounding boxes, rendering it similar to R-CNN-based architectures. Figure 4 shows a plot of average precision against the total number of parameters for two different architectures.
The trend is clear: removing layers from the topdown network significantly reduces performance. Some of this decline in performance may be explained by the fact that reducing the size of the topdown network reduces the overall depth of the network, and therefore its representational power. However, as can be seen from Figure 4, adopting a shallow front-end (ResNet-18) with a large topdown network achieves significantly better performance than a deeper network (ResNet-34) without any topdown layers, despite the two architectures having roughly the same number of parameters. This strongly suggests that a significant part of the success of our architecture comes from its ability to reason in 3D, as afforded by the 2D convolution layers operating on the orthographic feature maps.
Discussion
Motivated by the qualitative results in Section 4.2, we wished to further quantify the ability of our system to detect and localize distant objects. Figure 5 plots performance of each system when evaluated only on objects which are at least the given distance away from the camera. Whilst we outperform Mono3D over all depths, it is also apparent that the performance of our system degrades much more slowly as we consider objects further from the camera. We believe that this is a key strength of our approach.
2 Evolution of confidence maps during training
While the confidence maps predicted by our network are not necessarily calibrated estimates of model certainty, observing their evolution over the course of training does give valuable insights into the learned representation. Figure LABEL:fig:evolution shows an example of a confidence map predicted by the network at various points during training. During the early stages of training, the network very quickly learns to identify regions of the image which contain objects, which can be seen by the fact that high confidence regions correspond to projection lines from the optical center at which intersect a ground truth object. However, there exists significant uncertainty about the depth of each object, leading to the predicted confidences being blurred out in the depth direction. This fits well with our intuition that for a monocular system depth estimation is significantly more challenging than recognition. As training progresses, the network is increasingly able to resolve the depth of the objects, producing sharper confidence regions clustered about the ground truth centers. It can be observed that even in the latter stages of training, there is considerably greater uncertainty in the depth of distant objects than that of nearby ones, evoking the well-known result from stereo that depth estimation error increases quadratically with distance.
Conclusions
In this work we have presented a novel approach to monocular 3D object detection, based on the intuition that operating in the birds-eye-view domain alleviates many undesirable properties of images which make it difficult to infer the 3D configuration of the world. We have proposed a simple orthographic feature transform which transforms image-based features into this birds-eye-view representation, and described how to implement it efficiently using integral images. This was then incorporated into part of a deep learning pipeline, in which we particularly emphasized the importance of spatial reasoning in the form of a deep 2D convolutional network applied to the extracted birds-eye-view features. Finally, we experimentally validated our hypothesis that reasoning in the topdown space does achieve significantly better results, and demonstrated state-of-the-art performance on the KITTI 3D object benchmark.