Translating Images into Maps

Avishkar Saha, Oscar Mendez Maldonado, Chris Russell, Richard Bowden

I Introduction

Many tasks in autonomous driving are substantially easier from a top-down, map or bird’s-eye view (BEV). As many autonomous agents are restricted to the ground-plane, an overhead map is a convenient low-dimensional representation, ideal for navigation, that captures relevant obstacles and hazards. For scenarios such as autonomous driving, semantically segmented BEV maps must be generated on the fly as an instantaneous estimate, to cope with freely moving objects and scenes that are visited only once.

Inferring BEV maps from images requires determining the correspondence between image elements and their location in the world. Multiple works guide their transformation with dense depth and image segmentation maps , while others have developed approaches which resolve depth and semantics implicitly. Although some exploit the camera’s geometric priors , they do not explicitly learn the interaction between image elements and the BEV-plane.

Unlike previous approaches, we treat the transformation to BEV as an image-to-world translation problem, where the objective is to learn an alignment between vertical scan lines in the image and polar rays in BEV. The projective geometry therefore becomes implicit to the network. For our alignment model, we adopt transformers , an attention-based architecture for sequence prediction. With its attention mechanisms, we explicitly model pairwise interactions between vertical scanlines in the image and their polar BEV projections. Transformers are well-suited to the image-to-BEV transformation problem, as they can reason about interdependence between objects, depths and the lighting of the scene to achieve a globally consistent representation.

We embed our transformer-based alignment model within an end-to-end learning formulation which takes as input a monocular image with its intrinsic matrix, and predicts semantic BEV maps for static and dynamic classes.

The contributions of our paper are (1) We formulate generating a BEV map from an image as a set of 1D sequence-to-sequence translations. (2) By physically grounding our formulation we construct a restricted data-efficient transformer network that is convolutional with respect to the horizontal x-axis, yet spatially-aware. (3) By combining our formulation with monotonic attention from the language domain, we show that knowledge of what is below a point in an image is more important than knowledge of what is above it for accurate mapping; although using both leads to best performance. (4) We show how axial attention improves performance by providing temporal awareness and demonstrate state-of-the-art results across three large-scale datasets.

II Related Work

BEV object detection: Early approaches detected objects in the image and then regressed 3D pose parameters . The Mono3D model instead generated 3D bounding box proposals on the ground plane and scored each one by projecting into the image. However, all these works lacked global scene reasoning in 3D as each proposal was generated independently. OFTNet overcame this by generating 3D features from projecting a 3D voxel grid into the image, and performing 3D object detection over those features. While it reasons directly in BEV, the context available to each voxel depends upon its distance from the camera, in contrast, we decouple this relationship to allow each BEV position access to the entire vertical axis of the image.

Inferring semantic BEV maps: BEV object detection has been extended to building semantic maps from images for both static and dynamic objects. Early work in road layout estimation performed semantic segmentation in the image-plane and assumed a flat world mapping to the ground plane via a homography. However, as the flat world assumption leads to artifacts for dynamic objects such as cars and pedestrians, others exploit depth and semantic segmentation maps to lift objects into BEV. While such intermediate representations provide strong priors, they require image depth and segmentation maps as additional input.

Several works instead reason about semantics and depth implicitly. Some use camera geometry to transform the image into BEV while others learn this transformation implicitly . Current state-of-the-art approaches can be categorised as taking a ‘compression’ or ‘lift’ approach to the transformation. ‘Compression’ approaches vertically condense image features into a bottleneck representation and then expand out into BEV, thus creating an implicit relationship between an object’s depth and the context available to it. This increases its susceptibility to ignore small, distant objects. ‘Lift’ approaches instead expand each image into a frustum of features to learn a depth distribution for each pixel. However, each pixel is given the entire image as context, potentially increasing overfitting due to redundancies in the image. Furthermore, neither approaches have spatial awareness, meaning they are unable to leverage the structured environments of urban scenes. We overcome issues with both these approaches by (1) maintaining the spatial structure of the image to explicitly model its alignment with the BEV-plane and (2) adding spatial awareness which allows the network to assign image context across the ray space based on both content and position.

Encoder-decoder transformers: Attention mechanisms were first proposed by Bahdanau et al. for machine translation to learn an alignment between source and target sequences using recurrent neural networks (RNNs). Transformers, introduced by Vaswani et al. , instead implemented attention within an entirely feed-forward network, leading to state-of-the-art performance in many tasks .

Like us, the 2D detector DETR performs decoding in a spatial domain through attention. However, their predicted output sequences are sets of object detections, which have no intrinsic order to them, and permits the use of attention’s permutation invariant nature without any spatial awareness. In contrast, the order of our predicted BEV ray sequences is inherently spatial and so we need spatial awareness and therefore permutation equivariance in our decoding.

III Method

where Φ\Phi is a neural network trained to resolve both semantic and positional uncertainties.

We propose learning the alignment between input scanlines and output polar rays through an attention mechanism . We employ attention in two ways: (1) inter-plane attention as shown in Fig.1b, which initially assigns features from a scanline to a ray and (2) polar ray self-attention that globally reasons about its positional assignments across the ray. We motivate both uses below, starting with inter-plane attention.

Inter-plane attention: Consider a semantically segmented image column and its corresponding polar BEV ground truth. Here, alignment between the column and the ground truth ray is ‘hard’, i.e. each pixel in the polar ray corresponds to a single semantic category from the image column. Thus, the only uncertainty that must be resolved to make this a hard-assignment is the depth of each pixel. However, when making this assignment, we need to assign features that aid in resolving semantics and depth. Hence, a hard assignment would be detrimental. Instead, we want a soft-alignment, where every pixel in the polar ray is assigned a combination of elements in the image column, i.e. a context vector. Concretely, when generating each radial element Siϕ(BEV)S^{\phi(BEV)}_{i}, we want to give it a context cic_{i} based on a convex combination of elements in the image column SIS^{I} and the radial position rir_{i} of the element Siϕ(BEV)S^{\phi(BEV)}_{i} along the polar ray. This need for context assignment motivates our use of soft-attention between the image column and its polar ray, as illustrated in Fig. 1.

Following common terminology, we refer to QQ and KK as ‘queries’ and ‘keys’ respectively. After projection, an unnormalized alignment score ei,je_{i,j} is produced between each memory-query combination using the scaled-dot product :

Finally, the context vector is computed as a weighted sum of KK:

Generating the context this way allows each radial slot rir_{i} to independently gather relevant information from the image column; and represents an initial assignment of components from the image to their BEV locations. Such an initial assignment is analogous to lifting a pixel based on its depth. However, it is lifted to a distribution of depths and thus should be able to overcome common pitfalls of sparsity and elongated object frustums. This means that the image-context available to each radial slot is decoupled from its distance to the camera. Finally, to generate BEV feature Siϕ(BEV)S^{\phi(BEV)}_{i} at radial position rir_{i}, we globally operate on the assigned contexts for all radial positions c={c1,...,cr}\mathbf{c}=\{c_{1},...,c_{r}\}:

where g(.)g(.) is a nonlinear function reasoning across the entire polar ray. We describe its role below.

Polar ray self-attention: The need for the non-linear function g(.)g(.) as a global operator arises out of the limitations brought about by generating each context vector cic_{i} independently. Given the absence of global reasoning for each context cic_{i}, the spatial distribution of features across the ray is unlikely to be congruent with object shape, locally or globally. Rather, this distribution may only represent scattered suggestions of object-part positions. Therefore, we need to operate globally across the ray to allow the assigned scanline features to reason about their placement within the context of the entire ray, and thus aggregate information in a manner that generates coherent object shapes.

Global computation across the polar ray is computed much like soft-attention outlined in Eq. (2) - (5), except that the self-attention is applied to the ray only. Eq. (2) is recalculated with a new set of weight matrices with inputs to both equations replaced with the context vector cic_{i}.

Extension to transformers: Our inter-plane attention can be extended to attention between the encoder-decoder of transformers by replacing the key K(hj)K(\mathbf{h}_{j}) in Eq. (5) with another projection of the memory h\mathbf{h}, the ‘value’. Similarly, polar-ray self-attention can be placed within a transformer-decoder by replacing the key in Eq. (5) with a projection of the context cic_{i} to represent the value.

III-B Infinite lookback monotonic attention

Although soft-attention is sufficient for learning an alignment between any arbitrary pair of source-target sequences, our sequences exist in the physical world where the alignment exhibits physical properties based on their spatial ordering. Typically, in urban environments, depth monotonically increases with height i.e., as you move up the image, you move further away from the camera. We enforce this through monotonic attention with infinite lookback . This constrains radial depth intervals to observe elements of the image column that are monotonically increasing in height but also allows context from the bottom of the column (or equivalently, previous memory entries).

Monotonic attention (MA) was originally proposed for computing alignments for simultaneous machine translation . However, the ‘hard’ assignment between source and target sequence means important context is neglected. This led to the development of MA with infinite lookback (MAIL) , which combined hard MA with soft-attention that extends from the hard assignment to the beginning of the source sequence. We adopt MAIL as a way of constraining our attention mechanism to potentially prevent overfitting by ignoring the redundant context in the vertical scan line of an image. The primary objective of our adoption of MAIL is to understand whether context below a point in an image is more helpful than what is above.

We employ MAIL by first calculating a hard-alignment using monotonic attention. This makes a hard assignment of context cic_{i} to an element of the memory hj\mathbf{h}_{j}, after which a soft-attention mechanism over previous memory entries h1,...,hj−1\mathbf{h}_{1},...,\mathbf{h}_{j-1} is applied. Formally, for each radial position yi∈y\mathbf{y}_{i}\in\mathbf{y} along the polar ray, the decoder begins scanning memory entries from index j=ti−1j=t_{i-1}, where tit_{i} is the index of the memory entry chosen for position yi\mathbf{y}_{i}. For each memory entry, it produces a selection probability pi,jp_{i,j}, which corresponds to the probability of either stopping and setting ti=jt_{i}=j and ci=htic_{i}=\mathbf{h}_{t_{i}}, or moving onto the next memory entry j+1j+1. As hard assignment is not differentiable, training is instead carried out with respect to the expected value of cic_{i}, with the monotonic alignments αi,j\alpha_{i,j} calculated as follows:

This effectively represents a distribution over image-elements which lie below a point in the image; to calculate a distribution over only what lies above a point in the image, the image column can be flipped. The context vector is calculated similar to inter-plane attention, where ci=∑j=1Hβi,jK(hj)c_{i}=\sum_{j=1}^{H}\beta_{i,j}K(\mathbf{h}_{j}).

III-C Model architecture

We build an architecture that facilitates our goal of predicting a semantic BEV map from a monocular image around this alignment model. As shown in Fig. 1, it contains three main components: a standard CNN backbone which extracts spatial features in the image-plane, encoder-decoder transformers to translate features from the image-plane to BEV and finally a segmentation network which decodes BEV features into semantic maps.

Our transformer encoder and decoder use the same set of projection matrices for every sequence-to-sequence translation, giving it a structure that is convolutional along the xx-axis and allowing us to make efficient use of data when training. We constrain our translations to 1D sequences as opposed to using the entire image to make learning easier, a decision we analyze in section IV-A.

Polar-adaptive context assignment: The positional encodings applied to the transformer so far have all been 1D. While this allows our convolutional transformer to leverage spatial relationships between height in the image and depth, it remains agnostic to polar angle. However, the angular domain plays an important role in urban environments. For instance, images display a broadly structured distribution of object classes across their width (e.g. pedestrians are typically only seen on sidewalks, which lie towards the edges of the image). Furthermore, object appearance is also structured along the width of the image as they are typically orientated along orthogonal axes and viewing angle changes its appearance. To account for such variations in appearance and distribution across the image, we add additional positional information by encoding polar angle in our 1D scanline-to-ray translations.

where yiky_{i}^{k} is the ground truth binary variable grid cell, y^ik\hat{y}_{i}^{k} the predicted sigmoid output of the network, and ϵ\epsilon is a constant used to prevent division by zero.

IV Experiments and Results

We evaluate the effectiveness of treating the image-to-BEV transformation as a translation problem on the nuScenes dataset ; with ablations on lookback direction in monotonic attention, the utility of long-range horizontal context and the effect of polar positional information. Finally, we compare our approach to current state-of-the-art approaches on the nuScenes , Argoverse and Lyft datasets.

Dataset: The nuScenes dataset consists of 1000 20-second clips captured across Boston and Singapore, annotated with 3D bounding boxes and vectorized road maps. We follow ’s data generation process, object classes and training/validation splits to allow fair comparison. We use nuScenes for our ablation studies as it is considerably larger and contains more object categories.

Implementation: Our frontend uses a pretrained ResNet-50 with a feature pyramid on top. BEV feature maps built by the transformer decoder have a resolution of 100×\times100 pixels, with each pixel representing 0.5m2 in the world. Our spatiotemporal model takes a 6Hz sequence of 4 images, where the final frame is the time step we make the prediction for. Our largest scale output is 100×\times100 pixels, which we upsample to 200×\times200 for fair evaluation with the literature. We train our network end-to-end with an Adam optimizer, batch size 8 and initial learning rate of 5e−55e{-5}, which we decay by 0.99 every epoch for 40 epochs.

Which way to look? In Table II (top) we compare soft-attention (looking both ways), monotonic attention with lookback towards the bottom of the image (looking down) and monotonic attention with lookback towards the top of the image (looking up). The results indicate looking downwards from a point in the image is better than looking upwards. This is consistent with how humans try to determine the distance of an object in an urban environment — along with local textural clues of scale, we make use of where the object intersects the ground plane. The results also show that looking in both directions further increases accuracy, making it more discriminative for depth reasoning.

Long-range horizontal dependencies: As our image-to-BEV transformation is carried out as a set of 1D sequence-to-sequence translations, a natural question is what happens when the entire image is translated to BEV (similar to ‘lift’ approaches ). Given the quadratic computation time and memory required to produce attention maps, this is prohibitively expensive. However, we can approximate the contextual benefits of using the entire image by applying horizontal axial-attention on the image-plane features before the transformation. With axial-attention across the rows of the image, the pixels in the vertical scanline now have long-range horizontal context, after which we provide long-range vertical context as before by translating between 1D sequences.

Table II (middle) shows that incorporating long-range horizontal context does not benefit the model and its impact is slightly detrimental. This suggests two things. Firstly, every transformed ray does not need information from the entire width of the input image, or rather, the long-range context does not provide any additional benefit over the context that has already been aggregated through the convolutions of our frontend. This indicates that performing the translation using the entire image would not increase model accuracy over the constrained formulation of our baseline. Finally, the decrease in performance from the introduction of horizontal axial-attention is possibly a sign of the difficulty in training using attention for sequences which are the width of the image; we should expect that using the entire image as the input sequence would be much harder to train.

Polar-agnostic vs polar-adaptive transformers: Table II (bottom) compares a polar-agnostic (Po-Ag) transformer to its polar-adaptive (Po-Ad) variants. A Po-Ag model has no polar-positional information, Po-Ad in the image-plane involves polar encodings added to the transformer encoder while for the BEV-plane this information is added to the decoder. Adding polar encodings to any one plane provides similar benefit over an agnostic model, with dynamic classes increasing the most. Adding it to both planes increases this further, but has the largest impact on static classes.

IV-B Comparison to state-of-the-art

Baselines: We compare against a number of prior state-of-the-art methods. We begin our comparison against ‘compression’ approaches on nuScenes and Argoverse using the train/val splits of . We then compare against the ‘lift’ approach of on nuScenes and Lyft.

In Table I, our spatial model outperforms the current state-of-the-art compression approach of STA-S with a mean relative improvement 15%. It is the smaller dynamic classes in particular on which we show significant improvement, with buses, trucks, trailers and barriers all increasing by a relative 35-45%. This is supported by our qualitative results in Fig. 2, where our models show greater structural similarity to the ground truths and a better sense of shape. This difference can be partly attributed to the fully-connected layer (FCL) used in compression: when detecting small, distant objects, a large portion of the image is redundant context. Expecting the weights of the FCL to ignore redundancies to maintain only the small objects in the bottleneck is a challenge. Furthermore, objects such as pedestrians are often partially occluded by vehicles. In such cases, the FCL would be inclined to ignore the pedestrian and instead maintain the vehicle’s semantics. Here the attention method shows its advantages as each radial depth can independently attend to the image — so the further depths can look at the pedestrian’s visible body, while depths before can attend to the vehicle. Our results on the Argoverse dataset in Table III demonstrate similar patterns, where we improve upon PON by a relative 30%.

In Table IV we outperform LSS and FIERY on nuScenes and Lyft (FIERY uses the ‘lift’ approach of ). A true comparison on Lyft is not possible as it doesn’t have a canonical train/val split and we were unable to acquire those used by . While we used splits of similar sizes to , the exact scenes are unknown. As a ‘lift’ approach bears some similarity to our translation approach in that the network is able to select how to distribute image context across its polar ray, the difference in performance here can likely be attributed to our constrained, spatially-aware translations between scanlines and rays. One of the avenues for future work is improving localisation accuracy for distant objects, and their false negatives. Finally, our approach is easily transferrable to indoor mobile robotics applications once ground truth has been collected to train the models.

V Conclusion

We proposed a novel use of transformer networks to map from images and video sequences to an overhead map or bird’s-eye-view of the world. We combine our physical-grounded and constrained formulation, with ablation studies that make use of progress in monotonic attention to confirm our intuitions whether context above or below a point is more important for this form of map generation. Our novel formulation obtains state-of-the-art results for instantaneous mapping of three well-established datasets.

Acknowledgements

This project was supported by the EPSRC project ROSSINI (EP/S016317/1) and studentship 2327211 (EP/T517616/1).

References