ONCE-3DLanes: Building Monocular 3D Lane Detection

Fan Yan, Ming Nie, Xinyue Cai, Jianhua Han, Hang Xu, Zhen Yang, Chaoqiang Ye, Yanwei Fu, Michael Bi Mi, Li Zhang

Introduction

The perception of lane structure is one of the most fundamental and safety-critical tasks in autonomous driving system. It is developed with the desired purpose of preventing accidents, reducing emissions and improving the traffic efficiency . It serves as a key roll of many applications, such as lane keeping, high-definition (HD) map modeling , trajectory planning etc. In light of its importance, there has been a recent surge of interest in monocular 3D lane detection . However, existing 3D lane detection datasets are either unpublished, or synthesized in a simulated environment due to the difficulty of data acquisition and high labor costs of annotation. With the only synthesized data, the model inevitably lacks generalization ability in real-word scenarios. Although benefiting from the development of domain adaptation method , it still cannot completely alleviate the domain gap.

Most existing image-based lane detection methods have exclusively focused on formulating the lane detection problem as a 2D task , in which a typical pipeline is to firstly detect lanes in the image plane based on semantic segmentation or coordinate regression and then project the detected lanes in top view by assuming the ground is flat . With well-calibrated camera extrinsics, the inverse perspective mapping (IPM) is able to obtain an acceptable approximation for 3D lane in the flat ground plane. However, in real-world driving environment, roads are not always flat and camera extrinsics are sensitive to vehicle body motion due to speed change or bumpy road, which will lead to the incorrect perception of the 3D road structure and thus unexpected behavior may happen to the autonomous driving vehicle.

To overcome above shortcomings associated with flat-ground assumption, 3D-LaneNet , directly predicts the 3D lane coordinates in an end-to-end manner, in which the camera extrinsics are predicted in a supervised manner to facilitate the projection from image view to top view. In addition, an anchor-based lane prediction head is proposed to produce the final 3D lane coordinates from the virtual top view. Despite the promising result exhibits the feasibility of this task, the virtual IPM projection is difficult to learn without the hard-to-get extrinsics and the model is trained under the assumption that zero degree of camera roll towards the ground plane. Once the assumption is challenged or the need of extrinsic parameters is not satisfied, this method can barely work.

In this work, we take steps towards addressing above issues. For the first time, we present a real-world 3D lane detection dataset ONCE-3DLanes, consisting of 211K images with labeled 3D lane points. Compared with previous 3D lane datasets, our dataset is the largest real-world lane detection dataset published up to now, containing more complex road scenarios with various weather conditions, different lighting conditions as well as a variety of geographical locations. An automatic data annotation pipeline is designed to minimize the manual labeling effort. Comparing to the method of using multi-sensor and expensive HD maps, ours is simpler and easier to be implemented. In addition, we introduce a spatial-aware lane detection method, dubbed SALAD, in an extrinsic-free and end-to-end manner. Given a monocular input image, SALAD directly predicts the 2D lane segmentation results and the spatial contextual information to reconstruct the 3D lanes without explicit or implicit IPM projection.

The contributions of this work are summarized as follows: (i) For the first time, we present a largest 3D lane detection dataset ONCE-3DLanes, alongside a more generalized evaluation metric to revive the interest of such task in a real-world scenario; (ii) We propose a method, SALAD that directly produce 3D lane layout from a monocular image without explicit or implicit IPM projection.

Related work

There are various methods proposed to tackle the problem of 2D lane detection. Segmentation-based methods predict pixel-wise segmentation labels and then cluster the pixels belonging to the same label together to predict lane instances. Proposal-based methods first generate lane proposals either from the vanishing points or from the edges of image , and then optimize the lane shape by regressing the lane offset. There are some other methods trying to project the image into the top view and using the properties that lanes are almost parallel and can be fitted by lower order polynomial in the top view to fit lanes. However, most methods are limited in the image view, lack the image-to-world step or suffer from the untenable flat-ground assumption. As a result, formulating lane detection as a 2D task may cause inappropriate behaviors for autonomous vehicle when encountering hilly or slope roads .

2 3D lane detection

LiDAR-based lane detection. Several methods have been proposed using LiDAR to detect 3D lanes, use the characteristic that the intensity values of different material are different to filter out the point clouds of the lanes by a certain intensity threshold and then cluster them to obtain 3D lanes. However, it’s hard to determine the specific intensity threshold since the material used in different countries or regions is distinct, and the intensity value varies much in various weather conditions, e.g., rainy or snowy.

Multi-sensor lane detection. Other methods try to aggregate information from both camera and LiDAR sensors to tackle lane detection task. Specifically, predicts the ground height from LiDAR points to project the image to the dense ground. It combines the image information with LiDAR information to produce lane boundary detection results. Nevertheless, it’s difficult to guarantee that the image and the point clouds appear in pairs in the real scenes, e.g., CULane dataset only contains images.

Monocular lane detection. Recently there are a few methods trying to address this problem by directly predicting from a single monocular image. The pioneering work 3D LaneNet predicts camera extrinsics in a supervised manner to learn the inverse perspective mapping(IPM) projection, by combining the image-view features with the top-view features. Gen-LaneNet proposed a new geometry-guided lane anchor in the virtual top view. By decoupling the learning of image segmentation and 3D lane prediction, it achieves higher performance and are more generalizable to unobserved scenes. Instead of associating each lane with a predefined anchor, 3D-LaneNet+ proposes an anchor-free, semi-local representation method to represent lanes. Although the ability to detect more lane topology structures shows the anchor-free method’s power. However, all the above methods need to learn a projection matrix in a supervised way to align the image-view features with top-view features, which may cause height information loss. While our proposed method directly regresses the 3D coordinates in the image view without considering camera extrinsics.

3 Lane datasets

Existing 3D lane detection datasets are either unpublished or synthesized in simulated environment. Gen-Lanenet uses Unity game engine to build 3D worlds and releases a synthetic 3D lane dataset, Apollo-Sim-3D, containing 10.5K images. 3D-LaneNet adopts a graphics engine to model terrains using a Mixture of Gaussians distribution. Lanes modeled by a 4th4^{th} degree polynomial in top view are placed on the terrains to generate the synthetic 3D lanes dataset synthetic-3D-lanes, containing 306K images with a resolution of 360×480360\times 480. A real 3D lane dataset Real-3D-lanes with 85K images is also created using multiple sensors including camera, LiDAR scanner and IMU as well as the expensive HD maps in .

In this paper, we publish the first real-world 3D lane dataset ONCE-3DLanes which contains 211K images and covers abundant scenes with various weather conditions, different lighting conditions as well as a variety of geographical locations. Comprehensive comparisons of 3D lane detection datasets are shown in Table 1.

ONCE-3DLanes

Raw data. We construct our ONCE-3DLanes dataset based on the most recent large-scale autonomous driving dataset ONCE (one million scenes) considering its superior data quality and diversity. ONCE contains 1 million scenes and 7 million corresponding images, the 3D scenes are recorded with 144 driving hours covering different time periods including morning, noon, afternoon and night, various weather conditions including sunny, cloudy and rainy days, as well as a variety of regions including downtown, suburbs, highway, bridges and tunnels. Since the camera data is captured at a speed of two frames per second and most adjacent frames are very similar, we take one frame every five frames to build our dataset to reduce data redundancy. Also, the distortions are removed to enhance the image quality and improve the projection accuracy from LiDAR to camera. Thus, by downsampling the ONCE five times, our dataset contains 211k images taken by a front-facing camera.

Lane representation. A lane LkL^{k} in 3D space is represented by a series of points {(xik,yik,zik)}i=1n\left\{(x_{i}^{k},y_{i}^{k},z_{i}^{k})\right\}_{i=1}^{n}, which are recorded in the 3D camera coordinate system with unit meter. The camera coordinate system is placed at the optical center of the camera, with X-axis positive to the right, Y-axis downward and Z-axis forward.

The projection error from front-view to top-view mainly occurs in the situation of slope ground, so we focus on analyzing the slope statistics on ONCE-3DLanes. The mean slope of lanes in each scene is utilized to represent the slope of this scene. The slope of a specific lane in forward direction which is considered to be the most important is calculated as follows:

where (x1,y1,z1)(x_{1},y_{1},z_{1}) and (x2,y2,z2)(x_{2},y_{2},z_{2}) are the start point and the end point of the lane respectively. The distribution of the slope conditions and the histogram of the number of lanes per image are shown in Figure 2. It shows that our dataset is full of complexity and contains enough various slope scenes with different illumination conditions.

Dataset splits. Follow the ONCE dataset, our benchmark contains the same 3K scenes for validation and 8K scenes for testing. To fully make use of the raw data, the training dataset not only contains the original 5K scenes, but also the unlabeled 200K scenes.

2 Annotation pipeline

Lanes are a series of points on the ground, which are hard to be identified in point clouds. Hence the high-quality annotations of 3D lanes are expensive to obtain, whilst it is much cheaper to annotate the lanes in 2D images. The paired LiDAR point clouds and image pixels are thoroughly investigated and used to construct our 3D Lane dataset. An overview of dataset construction pipeline is shown in Figure 3. The pipeline consists of five steps: Ground segmentation, Point cloud projection, Human labeling/Auto labeling, Adaptive lanes blending and Point cloud recovery. These steps are described in detail below.

Ground segmentation. Lanes are painted on the ground, which is a strong prior to locate the precise coordinate in the 3D space. To make full use of human prior and avoid the reflection of point clouds aliasing between lanes and other objects, the ground segmentation algorithm is utilized to get the ground LiDAR points at first.

The ground segmentation is performed in a coarse-to-fine manner. In the coarse way, since the height of the LiDAR points reflected by the ground always settles in certain intervals, a pre-defined threshold is adopted to filter out those points lying on the ground coarsely based on the height statistics of the LiDAR points among the whole dataset as seen in Figure 2(b). In the fine way, several points in front of the vehicle are sampled randomly as seeds and then the classic region growth method is applied to get the fine segmentation result.

Point cloud projection. In this step, the previous extracted ground LiDAR points are projected to the image plane with the help of calibrated LiDAR-to-camera extrinsics and camera intrinsics based on the classic homogeneous transformation, which reveals the explicit corresponding relationship between the 3D ground LiDAR points and the 2D ground pixels in the image.

Human labeling / Auto labeling. To obtain 2D lane labels in the images and alleviate the taggers’ burden, a robust 2D lane detector which is trained in million-level scenes is firstly used to automatically pre-annotate pseudo lane labels. At the same time, professional taggers are required to verify and correct the pseudo labels to ensure the annotation accuracy and quality.

Adaptative lanes blending. After getting the accurate 2D lane labels and ground points, in order to judge whether a ground point belongs to the lane marker or not, we broaden 2D lane labels with an appropriate and adaptive width to get lane regions. Due to the perspective principle, a lane is broadened with different widths according to the distance from camera. After the point clouds projection procedure, we consider the ground points which are contained in the lane regions as the lane point clouds.

Point cloud recovery. Finally, we select these lane point clouds out. And for a specific lane, the lane point clouds in the same beam are clustered to get the lane center points to represent this lane.

To ensure the accuracy of annotations, we do not interpolate between center points in data collection stage. While in training stage, we use cubic spline interpolation to generate dense supervised labels. we also compared our annotations with manually labeling results on a small portion of data, which shows the high quality of our annotations. The interpolation code will be made public along with our dataset.

SALAD

In this section, we introduce SALAD, a spatial-aware monocular lane detection method to perform 3D lane detection directly on monocular images. In contrast to previous 3D lane detection algorithms , which project the image to top view and adopt a set of predefined anchors to regress 3D coordinates, our method does not require human-crafting anchors and the supervision of extrinsic parameters. Inspired by SMOKE , SALAD consists of two branches: semantic awareness branch and spatial contextual branch. The overall structure of our model is illustrated in Figure 4. In addition, we also adopt a revised joint 3D lane augmentation strategy to improve the generalization ability. The details of our network architecture and augmentation methods are discussed in the following parts.

2 Semantic awareness branch

Traditional 3D lane detection methods directly projecting feature map from front view to top view, are not reasonable and harmful to the performance of prediction as the feature map may not be organized following the perspective principle.

According to , inverse projection from 2D image to 3D space is an underdetermined problem. Based on the segmentation map generated by semantic awareness branch, spatial information of each pixel is also required to transfer this segmentation map from 2D image plane to 3D space.

3 Spatial contextual branch

Due to the downsampling and lacking of global information, the locations of the predicted lane points are not accurate enough. Our spatial contextual branch accepts feature FF and outputs a pixel-level offset map, which predicts the spatial location shift δu\delta_{u} and δv\delta_{v} of lane points along uu and vv axis on the image plane. With the predictions of pixel location offsets δu\delta_{u} and δv\delta_{v}, the rough estimation of lane point locations is modified by global spatial context:

In order to recover 3D lane information, the spatial contextual branch also generates a dense depth map to regress on depth offset δz\delta_{z} for each pixel of the lane markers. Considering the depth of the ground on image plane increases along rows, we assign each row of the depth map a predefined shift αr\alpha_{r} and scale βr\beta_{r}, and perform regression in a residual way. The standard depth value zz is recovered as following:

The ground-truth depth map is generated by projecting 3D lane points {(xik,yik,zik)}i=1n\left\{(x_{i}^{k},y_{i}^{k},z_{i}^{k})\right\}_{i=1}^{n} on the image plane to get pixel coordinates {(uik,vik,zik)}i=1n\left\{(u_{i}^{k},v_{i}^{k},z_{i}^{k})\right\}_{i=1}^{n}. Then at each pixel (uik,vik)(u_{i}^{k},v_{i}^{k}), its corresponding depth value is assigned to zikz_{i}^{k}. Following , we apply depth completion on the sparse depth map to get the dense depth map DgtD_{gt} to provide sufficient training signals for our spatial contextual branch.

4 Spatial reconstruction

The spatial information predicted by our spatial contextual branch of our model plays a virtual role in 3D lane reconstruction. To map 2D lane coordinates back to 3D spatial location in camera coordinate system, the depth information is an indispensable element. To be concrete, given the camera intrinsic matrix K3×3K_{3\times 3}, a 3D point (x,y,z)(x,y,z) in camera coordinate system can be projected to a 2D image pixel (u,v)(u,v) as:

where fxf_{x} and fyf_{y} represent the focal length of the camera, (cx,cy)(c_{x},c_{y}) is the principal point and ss is the axis skew. Thus, given a 2D lane point in the image with pixel coordinates (u,v)(u,v) along with its depth information dd, noted that the depth denotes the distance to the camera plane, so the depth dd is the same as the zz in the camera coordinate system. Thus 3D lane point in camera coordinate system (x,y,z)(x,y,z) can be restored as follows:

Utilizing the fixed parameters of the camera intrinsic, we can project the 2D lane proposal points back to 3D locations to reconstruct our 3D lanes.

5 Loss function

Given an image and its corresponding ground-truth 3D lanes, the loss function between predicted lanes and ground-truth lanes are formulated as:

Lseg\mathcal{L}_{seg} is for the binary segmentation branch, with a cross-entropy loss in a pixel-wise manner on the segmentation map. For a specific pixel in segmentation map SS, yiy_{i} is the label and pip_{i} is the probability of being foreground pixels:

Lreg\mathcal{L}_{reg} is for the spatial contextual branch, which predicts the spatial offsets O=[δu,δv,δz]TO=[\delta_{u},\delta_{v},\delta_{z}]^{T}. we choose smooth L1 loss to regress these spatial contextual information OO:

λ\lambda denotes the penalty term for regression loss and is set to 11 in our experiments.

6 Data augmentation

Randomly horizontal flip and image scaling are common data augmentation methods to improve the generalization ability of 2D lane detection models. However, it is worth noting that the image shift and scale augmentation methods will cause 3D information inconsistent with the data augmentation . We revise it by proposing a joint scale strategy.

To ensure we can restore the same size of the original image from the scaled image, we first crop the top of the image with size c. Then we scale the cropped image with the proportion s. According to the similar triangles theorems, it is proved that the relationship of 3D information of a specific pixel before scaling (x,y,z)(x,y,z) and after scaling (x^,y^,z^)(\hat{x},\hat{y},\hat{z}) is:

Take scaling factor ss which is less than one for example. If the image is scaled with factor ss, in the camera coordinate system, it is like the camera is moving forward in the ZZ direction. As for a specific point, the xx and yy keep the same while the zz become smaller and the new zz equals z⋅sz\cdot s. Using this strategy, we can ensure that the 3D ground truth remain consistent during 2D image data augmentation.

Experiment

In this section, our experiments are presented as follows. First we introduce our experimental setups, including evaluation metrics and implementation details. Then we evaluate our baseline method on our ONCE-3DLanes dataset and investigate the evaluation performance of different hyper-parameter settings. Next we compare our proposed method with the prior state-of-the-art to prove the superiority of our proposed method. Finally, we conduct several ablations studies to show the significance of modules in our network.

Evaluation metric is set to measure the similarity between the predicted lanes and the ground-truth lanes. Previous evaluation metric , which set predefined yy-position to regress the xx and zz coordinates of lane points, is not sophisticated. Due to the fixed anchor design, this metric essentially performs badly when the lanes are horizontal.

To tackle this problem, we propose a two-stage evaluation metric, which regards 3D lane evaluation problem as a point cloud matching problem combined with top-view constraints. Intuitively, two 3D lane pairs are well matched when achieving little difference in the zz-xx plane (top view) alongside a close height distribution. Matching in the top view constrains the predicted lanes in the correct forward direction, and close point clouds distance in 3D space ensures the accuracy of the predicted lanes in spatial height.

Our proposed metric first calculate the matching degree of two lanes on the zz-xx plane. To be concrete, lane is represented as Lk={(xik,yik,zik)}i=1nL^{k}=\left\{(x_{i}^{k},y_{i}^{k},z_{i}^{k})\right\}_{i=1}^{n}. To judge whether predicted lane LpL^{p} matches ground-truth lane LgL^{g}, the first matching process is done in the z-x plane, namely top-view, we use the traditional IoU method to judge whether LpL^{p} matches LgL^{g}. If the IoU is bigger than the IoU threshold, further we use a unilateral Chamfer Distance (CDCD) to calculate the curves matching error in the camera coordinates. The curve matching error CDp,gCD_{p,g} between LpL^{p} and LgL^{g} is calculated as follows:

where Ppj=(xpj,ypj,zpj)P_{p_{j}}=(x_{p_{j}},y_{p_{j}},z_{p_{j}}) and Pgi=(xgi,ygi,zgi)P_{g_{i}}=(x_{g_{i}},y_{g_{i}},z_{g_{i}}) are point of LpL^{p} and LgL^{g} respectively, and P^pj\hat{P}_{p_{j}} is the nearest point to the specific point PgiP_{g_{i}}. mm represents the number of points token at an equal distance from the ground-truth lane. If the unilateral chamfer distance is less than the chamfer distance threshold, written as τCD\tau_{CD}. We consider LpL^{p} matches LgL^{g} and accept LpL^{p} as a true positive. The legend to calculate chamfer distance error is shown in Figure 5.

This evaluation metric is intuitive and strict, and more importantly, it applies to more lane topologies such as vertical lanes, so it is more generalized. At last, since we know how to judge a predicted lane is true positive or not, we use the precision, recall and F-score as the evaluation metric.

2 Implementation details

Our experiments are carried out on our proposed ONCE-3DLanes benchmark. We use the Segformer as our backbone with two branches. The Segformer encoder is pre-trained with Imagenet . The input resolution is set to 320 ×\times 800 with our augmentation strategy during training. The augmentation is turned off during testing. The Adamw optimizer is adopted to train 20 epochs, with an initial learning rate at 3e-3 and applied with a poly scheduler by default. For evaluation, We set the IoU threshold as 0.3, the chamfer distance thresh τCD\tau_{CD} as 0.3m. We also test our model using the Mindspore .

3 Benchmark performance

We train our model with the whole training set of 200k images and report the detection performance on the test set of ONCE-3DLanes dataset. To verify the rationality of the hyper-parameters setting in our evaluation metric, we evaluate our model under different Chamfer Distance threshold τCD\tau_{CD} and report the testing results in Table 2.

We report our performance in different settings of τCD\tau_{CD}, in order to fully investigate the impact of tightening or loosening the criteria on model performance. The illustrations of criteria adopting various τCD\tau_{CD} are also proposed in Figure 6. It can be seen that under the threshold of 0.5, some predicted lanes relatively far from the ground-truth are judged as true positives. While under the τCD\tau_{CD} of 0.15, the criteria seems too harsh. The τCD\tau_{CD} of 0.3 is more reasonable and our SALAD achieves 64.07%64.07\% F1-score based on this threshold. Besides, since the distance is calculated based on real scene so it is highly adaptable to real-world datasets. In the remaining parts, the experimental results are reported under the threshold of 0.3.

3.2 Results of 3D lane detection methods

In order to further verify the authenticity of our dataset and the superiority of our method, which is extrinsic-free, we also conduct some experiments of other 3D lane detection algorithms on our dataset.

It is worth noting that all existing 3D lane detection algorithms need to provide camera poses as supervision information and have strict assumptions that the camera is installed at zero degrees roll relative to the ground plane . However, our method requires no external parameter information. To make comparison, we used the camera-pose parameters provided by ONCE to provide supervision signals for counterpart methods and ultimately evaluated them on our test set. The performance of 3D lane detection algorithms are presented in Table 3.

The comparison shows our method outperforms other 3D lane detection method in ONCE-3DLanes dataset. The comparison result indicates that extrinsic-required methods under the assumption of fixed camera pose and zero-degree camera roll may suffer in real 3D scenarios.

3.3 Results of extended 2D lane detection methods

ONCE-3DLanes dataset is a newly released 3D lane detection dataset and no previous work addresses the problem of extrinsic-free 3D lane detection. In order to verify the validity of our dataset and the efficiency of our method, we extend the existing 2D lane detection model and evaluate their performance on our dataset.

Different from 3D lane detection methods, 2D lane detection algorithms can only detect the pixel coordinates of lanes in the image plane, but cannot recover the spatial information of lanes. To obtain 3D lane detection results, we use a pre-trained depth estimation model MonoDepth2 (finetuned to the depth scale of our dataset) to estimate the pixel-level depth of the image. It is worth mentioning that the depth model is finetuned on full ONCE dataset in order to avoid under-fitting caused by sparse supervision provided by lane points, which also indicates that this pipeline is difficult to be extended on other 3D lane benchmarks. Combined with the detection results of extended 2D lane detection model, the spatial position of 3D lanes are reconstructed and our evaluation metric is used for performance evaluation. The results are showed in Table 4.

Experimental results show that the extended 2D models are effective to perform 3D lane detection task on our ONCE-3DLanes dataset. It can also be found that our proposed method can reach 64.07%64.07\% F1-score, outperforms the best of other methods 56.57%56.57\% by 7.5%7.5\%, which shows the superiority of our method.

4 Ablation study

To verify the efficiency of revised data augmentation methods, we conduct ablation experiments by gradually turning off data augmentation strategy. As shown in Table 5, the performance of our method is constantly improved with the gradual introduction of our data augmentation strategy. The flip method brings a 1.06%1.06\% improvement to our model and the 3D scale provides a further 1.83%1.83\% improvement, which proves the validity of our augmentation strategy.

Conclusion and limitations

In this paper, we have presented a largest real-world 3D lane detection benchmark ONCE-3DLanes. To revive the interest of 3D lane detection, we benchmark the dataset with a novel evaluation metric and propose an extrinsic-free and anchor-free method, dubbed SALAD, directly predicting the 3D lane from a single image in an end-to-end manner. We believe our work can lead to the expected and unexpected innovations in communities of both academia and industry.

As our dataset construction requires LiDAR to provide 3D information, occlusions would lead to short interruption and we have used interpolation to fix it. Missing points problem in the distance still exists due to the low resolution of LiDAR. Future work will focus on the ground point clouds completion to generate full information for 3D lanes.

This work was supported in part by National Natural Science Foundation of China (Grant No. 6210020439), Lingang Laboratory (Grant No. LG-QS-202202-07), Natural Science Foundation of Shanghai (Grant No. 22ZR1407500), Shanghai Municipal Science and Technology Major Project (Grant No. 2018SHZDZX01 and 2021SHZDZX0103), Science and Technology Innovation 2030 - Brain Science and Brain-Inspired Intelligence Project (Grant No. 2021ZD0200204), MindSpore and CAAI-Huawei MindSpore Open Fund.

References

Appendix A Appendix

We conduct ablation studies to show the rationality of our experiment settings including loss function, backbone network and regression method.

As shown in Table 6, for the spatial contextual branch, we study different loss functions. Results show the smooth L1 loss outperforms L1 and L2 loss functions at all metrics.

We compare the SegFormer with Unet for the backbone network. Moreover, different attention mechanisms are added to Unet to help learn the global information of lane structures. Table 7 shows that model with SegFormer beat the variants of Unet by a clear margin.

We also conduct an ablation study to evaluate the way to predict the depth information in Table 8. Our method regress in a residual manner is referred as relative method, and the method directly regress the depth information without pre-defined shift and scale is called absolute method. Table 8 shows the relative outperforms absolute by a large margin.

A.2 More qualitative results

We present the qualitative results of SALAD lane prediction in Figure 7. 2D projections are shown in the left and 3D visualizations are presented in the right.