Cross-view Transformers for real-time Map-view Semantic Segmentation

Brady Zhou, Philipp Krähenbühl

Introduction

Autonomous vehicles depend on robust scene understanding and online mapping to navigate the world. To drive safely, these systems not only reason about the semantics of their surroundings, but also a spatial understanding due to the geometric nature of navigation. Many prior approaches directly model geometry and relationships between different view and a canonical map representation . They require an explicit or probabilistic estimate of depth in either image or map-view. However, this explicit modeling can be hard. First, image-based depth estimates are error-prone, as monocular depth estimates scale poorly with the distance to the observer. Second, depth-based projections are a fairly inflexible and rigid bottleneck to map between views. In this work, we take a different approach.

We learn to map from camera-view to a canonical map-view representation using a cross-view transformer architecture. The transformer does not perform any explicit geometric reasoning but instead learns to map between views through a geometry-aware positional embedding. Multi-head attention then learns to map features from camera-view into a canonical map-view representation using a learned map-view positional embedding. We learn a single map-view positional embedding for all cameras and perform attention across all views. The model thus learns to link up different map locations to both cameras and locations within each camera. Our cross-view transformer refines the map-view embedding through multiple attention and MLP blocks. The cross-view transformer allows the network to learn any geometric transformation implicitly and directly from data. It learns an implicit estimate of depth through the camera-dependent map-view positional embedding by performing the downstream task as accurately as possible.

The simplicity of our model is a key strength. The model performs at state-of-the-art on the nuScenes dataset for vehicle and road segmentation in the map-view and comfortably runs in real-time (35 FPS) on a single RTX 2080 Ti GPU. The model is easy to implement and trains within 32 GPU hours. The learned attention mechanism learns accurate correspondences between camera and map-view directly from data.

Related Works

Map-view semantic segmentation lies at the intersection of 3D recognition, depth estimation, and mapping.

Monocular detection aims to find objects in a scene, estimate their real-world size, orientation, and placement in the 3D scene. Most common approaches reduce the problem to 2D object detection and infer monocular depth . CenterNet directly predicts depth for each image coordinate. ROI-10D lifts 2D detections into 3D using depth estimates then regresses 3D bounding boxes. Psuedo-lidar based approaches project to 3D points using a depth estimate and leverage 3D point based architectures (e.g. ) with 2D labels. This family of algorithms directly benefits from advances in monocular depth estimation and 3D vision.

Monocular 3D object detection is both easier and harder than mapping from multiple cameras. The overall problem setup deals with just a single camera and does not need to merge multiple sources of inputs. However, it strongly relies on a good explicit monocular depth estimate, which may be harder to obtain.

Depth estimation.

Depth is a core ingredient in many multi-view mapping approaches. Classic structure-from-motion approaches leverage epipolar geometry and triangulation to explicitly compute camera extrinsics and depth. Stereo matching finds corresponding pixels, from which depth can be explicitly computed . Recent deep learning approaches directly regress depth from images .

While convenient, explicit depth is challenging to utilize for downstream tasks. It is camera-dependent and requires an accurate calibration and fusion of multiple noisy estimates. Our approach side-steps explicit depth estimation and instead allows an attention mechanism with positional embedding to take its place. Our cross-view transformer learns to reproject camera views into a common map representation as part of training.

Semantic mapping in the map-view.

Driven by ever larger 3D recognition datasets , a number of works have focused on perception in the map-view. This problem is particularly challenging as the inputs and outputs lie in different coordinate frames. Inputs are recorded in calibrated camera views, outputs are rasterized onto a map. Most prior works differ in the way the transformation is modeled. One common technique is to assume the scene is mostly planar and represent the image to map-view transformation as a simple homography . A second family of methods directly produces map-view predictions from input images, with no explicit geometric modeling. VED uses a Variational Auto Encoder to produce a semantic occupancy grid from a single monocular camera-view.

Closely related in spirit to our method, VPN learns a common feature representation across multiple views with their proposed view relation module - an MLP that outputs map-view features from inputs across all views. Both VED and VPN show carefully-designed networks trained with sufficient training data can jointly learn the map-view transformation and perform prediction. However, these methods do suffer certain drawbacks as they do not model the geometric structure of the scene. They forgo the inherit inductive biases contained in a calibrated camera setup and instead need to learn an implicit model of camera calibration baked into the network weights. Our cross-view transformer instead uses positional embeddings derived from calibrated camera intrinsics and extrinsics. The transformer can learn a camera-calibration-dependent mapping akin to raw geometric transformations.

Most recently, top-performing methods returned back to explicit geometric reasoning . Orthographic Feature Transform (OFT) creates a map-view intermediate representation from a monocular image by average pooling image features from the 2D projection that corresponds with the pillar in map-view. This pooling operation foregoes an explicit depth estimate and instead averages all possible image locations a map-view object could take. Lift-Splat-Shoot (LSS) constructs an intermediate map-view representation in a similar fashion. However, they allow the model to learn a soft depth estimate and average across different bins using a learned depth-estimate-dependent weight. Their downstream decoder can account for uncertainty in depth. This weighted averaging operation closely mimics the attention used in a transformer. However, their “attention weights” are derived from geometric principles and not learned from data. The original Lift-Splat-Shoot approach considers multiple views within a single timestep. Recent methods have extended this further to take aggregate features from previous timesteps , and use multi-view, multi-timestep observations to do motion forecasting .

In this work, we show that implicit geometric reasoning performs as well as explicit geometric models. The added benefit of our implicit handling of geometry is an improvement in inference speed compared to explicit models. We simply learn a set of positional embeddings, and attention will reproject the camera to map-view.

Cross-view transformers

We design a simple, yet effective encoder-decoder architecture for map-view semantic segmentation. An image-encoder produces a multi-scale feature representation of each input image. A cross-view cross-attention mechanism then aggregates multi-scale features into a shared map-view representation. The cross-view attention relies on a positional embedding that is aware of the geometric structure of the scene and learns to match up camera-view and map-view locations. All cameras share the same image-encoder, but use a positional embedding dependent on their individual camera calibration. Finally, a lightweight convolutional decoder upsamples the refined map-view embedding and produces the final segmentation output. The entire network is end-to-end differentiable and learned jointly. Figure 2 shows an overview of the full architecture.

In Section 3.1, we first present the core cross-view attention mechanism and positional embedding that underlies our entire architecture. Section 3.2 then combines multiple cross-view attention layers into the final map-view segmentation model.

Here, ≃\simeq describes equality up to a scale factor, and x(I)=(⋅,⋅,1)x^{(I)}=(\cdot,\cdot,1) uses homogeneous coordinates. However, without an accurate depth estimate in camera view or height-above-ground estimate in map-view, the world coordinate x(W)x^{(W)} is ambiguous. We do not learn an explicit estimate of depth but encode any depth ambiguity in the positional embeddings and let a transformer learn a proxy for depth.

We start by rephrasing the geometric relationship between world and image coordinates in Equation 1 as a cosine similarity for use in an attention mechanism.

This similarity still relies on the exact world coordinate w(W)w^{(W)}. Next, we replace all geometric components of this similarity with positional encodings that can learn both geometric and appearance features.

The camera-aware positional encoding starts from the unprojected image coordinate dk,i=Rk−1Kk−1xi(I)d_{k,i}=R_{k}^{-1}K_{k}^{-1}x^{(I)}_{i} for each image coordinate xi(I)x^{(I)}_{i}. The unprojected image coordinate dk,id_{k,i} describes a direction vector from the origin tkt_{k} of camera kk to the image plane at depth 11. The direction vector uses world coordinates.

Next, we show how to build an equivalent representation for the map-view queries. This embedding can no longer rely on exact geometric inputs and instead needs to learn geometric reasoning in consecutive layers of the transformer.

Map-view latent embedding.

Cross-view attention.

Our cross-view transformer combines both positional encodings through a cross-view attention mechanism. We allow each map-view coordinate to attend one or more image locations. Crucially, not every map-view location has a corresponding image patch in each view. Front-facing cameras do not see the back, rear-facing cameras do not see the front. We allow the attention mechanism to select both camera and location within each camera when corresponding map-view and camera-view perspectives. To this end, we first combine all camera-aware positional embeddings δ1,δ2,…\delta_{1},\delta_{2},\ldots from all views into a single key vector δ=[δ1,δ2,…]\delta=\left[\delta_{1},\delta_{2},\ldots\right]. At the same time, we combine all image features ϕ1,ϕ2,…\phi_{1},\phi_{2},\ldots into a single value vector ϕ=[ϕ1,ϕ2,…]\phi=\left[\phi_{1},\phi_{2},\ldots\right]. We combine camera-aware positional embeddings δ\delta and image features ϕ\phi to compute attention keys. Finally, we perform softmax-cross-attention between keys [δ,ϕ][\delta,\phi], values ϕ\phi, and map-view queries c−τkc-\tau_{k}.

The softmax attention uses a cosine similarity between keys and queries as a basic building block

This cosine similarity follows the geometric interpretation in Equation 2. This cross-view attention forms the basic building block of our cross-view transformer architecture.

2 A cross-view transformer architecture

The first stage of the network builds up a camera-view representation for each input image. We feed each image IiI_{i} into feature extractor (EfficientNet-B4 ) and get a multi-resolution patch embedding {ϕ11,ϕ12,…,ϕnR}\{\phi_{1}^{1},\phi_{1}^{2},\ldots,\phi_{n}^{R}\}, where RR is the number of resolutions we consider. We found R=2R=2 resolutions to produce sufficiently accurate results. We process each resolution separately. We start from the lowest resolution and project all image features into map-view using cross-view attention. We then refine the map-view embedding and repeat the process for higher resolutions. Finally, we use three up-convolutional layers to produce the full resolution output.

A detailed overview of this architecture is shown in Figure 2. The final network is end-to-end trainable. We train all layers using ground truth semantic map-view annotations and a focal loss .

Implementation Details

We use (and fine-tune) a pre-trained EfficientNet-B4 to compute image features at two different scales - (28, 60) and (14, 30), which correspond to a 8x and 16x downscaling, respectively. The initial map-view positional embedding is a tensor of learned parameters w×h×Dw\times h\times D, where D=128D=128. For computational efficiency, we choose w=h=25w=h=25 as the cross-attention function grows quadratically with grid size. The encoder consists of two cross-attention blocks: one for each scale of patch features. We use multi-head attention with 44 heads and an embedding size dhead=64d_{head}=64. The decoder consists of three (bilinear upsample + conv) layers to upsample the latent representation to the final output size. Each upsampling layer increases the resolution by a factor of 2 up to a final output resolution of 200×200200\times 200. This corresponds to a 100×100100\times 100 meter area centered around the ego-vehicle.

Training.

We train all models using a focal loss , with a batch size of 4 (per GPU) for 30 epochs. We optimize using the AdamW optimizer with the one-cycle learning rate scheduler . Training converges within 8 hours on a 4 GPU machine.

Results

We evaluate our cross-view transformer on vehicle and road map-view semantic segmentation on the nuScenes and Argoverse datasets.

The nuScenes dataset is a collection of 1000 diverse scenes collected over a variety of weathers, time-of-day, and traffic conditions. Each scene lasts 20 seconds and contains 40 frames for a total of 40k total samples in the dataset. The recorded data captures a full 360° view around the ego-vehicle and is composed of 6 camera views. Each camera view has calibrated intrinsics KK and extrinsics (R,t)(R,t) at every timestep. We resize every image to 224×448224\times 448 unless specified otherwise. The Argoverse dataset contains 10k total frames.

Vehicles and other objects in the scene are tracked across frames and annotated with 3D bounding boxes using LiDAR data. Using the pose of the ego-vehicle, we generate the ground-truth labels yy, a binary vehicle occupancy mask rendered at a resolution of (200, 200) by orthographically projecting 3D box annotations onto the ground plane, following standard practice .

Evaluation.

There are two commonly used evaluation settings for map-view vehicle segmentation. Setting 1 uses a 100m×\times50m area around the vehicle and samples a map at a 25cm resolution. The setting, popularized by Roddick et al. , serves as the main comparison to prior work. Setting 2 uses a 100m×\times100m area around the vehicle, with a 50cm sampling resolution. This setting was popularized by Philion and Fidler and serves as a comparison to Lift-Splat-Shoot and FIERY . We use Setting 2 for all ablations. In both settings, we use the Intersection-over-Union (IoU) score between the model predictions and the ground truth map-view labels as the main performance measure. We additionally report inference speeds measured on an RTX 2080 Ti GPU.

1 Comparison to prior work

We compare our model to the five most competitive prior approaches on online mapping. For a fair comparison, we use single-timestep models only and do not consider temporal models. We compare to Pyramid Occupancy Networks (PON) , Orthographic Feature Transform (OFT) , View Parsing Network (VPN) , Spatio-temporal Aggregation (STA) , Lift-Splat-Shoot , and FIERY . PON, VPN, STA only report numbers in Setting 1, while Lift-Splat-Shoot only uses Setting 2.

In both settings, our cross-view transformer and FIERY outperform all alternative approaches by a significant margin. Our cross-view transformer and FIERY perform comparably. We have a slight edge in Setting 2, FIERY in Setting 1. The main advantage of our model is simplicity and inference speed, along with the accompanying edge in model size. Our model trains significantly faster (32 GPU hours vs 96 GPU hours) and performs 4×4\times faster inference. We intentionally use the same image feature extractor (EfficientNet-B4) and similar decoder architecture as FIERY. This suggests our cross-view transformer is capable of combining features from multiple views in a more efficient manner.

2 Ablations of cross-view attention

The core ingredient of our approach is the cross-view attention mechanism. It combines camera-aware embeddings and image features as keys and learned map-view positional embeddings as queries. The map-view embeddings are allowed to update across multiple iterations, while the camera-aware embeddings contain some geometric information. Table 3 compares the impact of each of the components of the attention mechanism on the resulting map-view segmentation system. For each ablation, we trained a model from scratch using equivalent experimental settings, changing a single component at the time.

The most important component of our system is the camera-aware positional embedding. It bestows the attention mechanism with the ability to reason about the geometric layout of the scene. Without it, attention has to rely on the image feature to reveal its own location. It is possible for the network to learn this localization due to the size of the receptive field and zero padding around the boundary of the image. However, an image feature alone struggles to properly link up map-view and camera-view perspectives. It also needs to explicitly infer the direction each image is facing to disambiguate different views. On the other hand, a purely geometric camera-aware positional embedding alone is also insufficient. The network likely uses both semantic and geometric cues to align map-view and camera-view, especially after the refinement of the map-view embedding. Finally, using a single fixed map-view embedding also degrades the performance of the model. The final model performs best with all its attention components.

3 Camera-aware positional embeddings

As we have previously seen, the camera-aware positional embedding plays a major role in the success of the cross-view transformer. Table 4 compares different choices for this embedding. We ablate just the positional embedding and keep all other model and training parameters fixed.

Not using any positional embedding performs poorly. The attention mechanism has a hard time localizing features and identifying cameras. A learned embedding per camera performs surprisingly well. This is likely because the camera calibration stays mostly static and a learned embedding simply bakes in all geometric information. A camera-aware embedding with either a linear or Random Fourier projection performs best. This should not come as a surprise as both can learn a compact embedding that directly captures the geometry of the scene.

4 Accuracy vs distance

Next, we evaluate how well our model performs as the distance to the ego-vehicle increases. For this experiment, we measure the intersection-over-union accuracy, but ignore all predictions that are closer than a certain distance to the ego-vehicle. Figure 3 compares to our closest competitor FIERY.

Both models have close to identical error modes. As the distance to the camera increases, the models get less accurate. This is easiest explained through actual qualitative results in Figure 5. Farther away vehicles are often (partially) occluded and thus much harder to detect and segment. Our approach degrades slower for close-by distances, but slightly under-performs FIERY at longer ranges.

Partially occluded far-away samples have fewer corresponding image features, thus learning a mapping from map-view to camera-view directly is harder: There is less training data and fewer geometric priors to rely upon for our model. We anticipate more data to make up for this difference.

5 Robustness to sensor dropout

We take a model trained on all six inputs and evaluate the intersection over union (IoU) metric by randomly dropping mm cameras for each sample in the validation set. Figure 4 shows how the performance decreases linearly with the number of cameras dropped. This is quite intuitive as different cameras only overlap marginally. Thus each removed camera reduces the visible area linearly and drops the performance in the unobserved area. Note that the transformer-based model is generally quite robust to this camera dropout and the overall performance does not degrade beyond unobserved parts of the scene.

6 Qualitative Results

Figure 5 shows qualitative results on a variety of scenes. For each row, we show the six input camera views and the predicted map-view segmentation along with the ground truth segmentation. Our presented method can accurately segments nearby vehicles, but does not sense far away or occluded vehicles well.

7 Geometric reasoning in cross-view attention

Our quantitative experiments indicate that cross-view attention can learn some geometric reasoning. In Figure 6, we visualize the image-view attention for several points in the map-view. Each point corresponds to part of a vehicle. From qualitative evidence, the attention mechanism can highlight closely corresponding map-view and camera-view locations.

Conclusion

We present a novel map-view segmentation approach based on a cross-view transformer architecture built on top of camera-aware positional embeddings. The proposed approach achieves state-of-the-art performance, is simple to implement, and runs in real time.

This material is based upon work supported by the National Science Foundation under Grant No. IIS-1845485 and IIS-2006820.

References