CurveFormer: 3D Lane Detection by Curve Propagation with Curve Queries and Attention

Yifeng Bai, Zhirong Chen, Zhangjie Fu, Lang Peng, Pengpeng Liang, Erkang Cheng

I INTRODUCTION

Lane detection is a critical component of an autonomous driving system, and it plays an important role in lane keeping assist, lane departure warning, etc. Most of the current lane detection approaches are developed with 2D images using semantic segmentation or line regression . However, downstream tasks like planning and control prefer lanes that are represented by the curve parameters in 3D space. Subject to the lack of depth information and accurate real-time camera extrinsic parameters, the projection from the image plane to the BEV perspective is prone to the error propagation problem (as shown in Fig. 1 (a)). Additionally, these methods suffer from complex and time-consuming post-processing steps, such as cluster and curve fitting.

In order to mitigate the drawbacks of post-processing in two-stage methods, CNN-based approaches have been proposed for end-to-end 3D lane detection task . As shown in Fig. 1 (b), 3D-LaneNet proposes an anchor-based 3D lane representation and predicts camera pose to project 2D features with Inverse Projective Mapping (IPM). 3D-LaneNet+ reformulates 3D lane as an anchor-free representation to consider the restriction of the lane direction, and learns lane curve clustering in the network. Gen-LaneNet proposes a virtual top view to align the BEV features projected by IPM and lanes in the real-world. Although these methods make end-to-end 3D lane detection possible, the loss of lane height and the accuracy of camera pose estimation would affect the robustness of these methods. On the contrary, ONCE performs 2D lane semantic segmentation and depth estimation, and integrates these information to obtain 3D lanes. A problem of ONCE is that depth estimation might bring about errors at the far end of the lane.

Recently, inspired by the successes of Transformer in various vision and robotic tasks , several Transformer-based lane detection algorithms have been proposed. LSTR introduces Transformer to the lane detection task and predicts 2D lane parameters directly. But it encounters difficulties in representing sharp curves or lanes with complex topology. STSU follows the sparse query-based framework to produce topologically accurate lane graphs. Similar to the above CNN-based 3D methods, CLGo applies Transformer to enhance images features and predicts 3D lanes from distance-invariant top-view image in the second stage. PersFormer builds a dense BEV query and uses Transformer to interact queries from BEV with image features (as shown in Fig. 1 (c)). Although these methods try to utilize Transformer to the 3D lane detection task, the lack of image depth or BEV map height restricts their performance as they can not obtain the features that exactly correspond to the query.

To address the above challenges, we propose CurveFormer, a Transformer-based method for 3D lane detection (Fig. 1 (d)). Lanes are defined as sparse curve queries consisting of lane confidence, two polynomials and start and end points (Fig. 2 (a)). Inspired by DAB-DETR , we propose a set of 3D dynamic anchor points to interact curve queries with image features. Since the 3D anchor point (x,y,z)(x,y,z) has height information, we can use camera extrinsic parameters to obtain accurate image features corresponding to the point. Dynamic anchor point set is iteratively refined within the sequence of Transformer decoders. We introduce a novel curve cross-attention module in the decoder part to investigate the effect of curve queries and dynamic anchor point set. Different from standard Deformable-DETR that directly predicts sampling offsets from the query, we introduce a context sampling unit to predict offsets from the combination of reference features and queries to guide sampling offsets learning. In addition, an auxiliary segmentation branch is adopted to enhance the shared CNN backbone. In this way, our design of CurveFormer lends itself to 3D lane detection.

To verify the performance of the proposed algorithm, we evaluate our CurveFormer on the Apollo Synthetic dataset and OpenLane dataset . Our proposed CurveFormer sets a new state-of-the-art performances for 3D lane detection on the Apollo Synthetic test set. It also achieves promising performance on the OpenLane dataset compared with recently proposed Transformer-based 3D lane detection approaches. The effectiveness of each component is validated as well.

In general, our main contributions are three-fold:

We propose CurveFormer, a novel Transformer-based 3D lane detection algorithm, by formulating queries in decoder layers as dynamic anchor point set, and a curve cross-attention module is applied to compute the query-to-image similarity.

We introduce a context sampling unit to predict offsets from the combination of reference features and queries to guide sampling offsets learning.

Experimental results show that our method achieves promising performance compared with both CNN-based and Transformer-based state-of-the-art approaches.

II RELATED WORK

2D Lane Detection. It detects lanes in the image plane and projects them to 3D space with camera pose. In general, advanced monocular lane detectors can be categorized into segmentation approaches and regression approaches .

SCNN is proposed to propagate context by slice-by-slice convolutions within feature maps. LaneNet introduces an instance segmentation approach for lane detection which combines a binary segmentation branch and an embedding branch. SAD allows a lane detection network to reinforce representation learning of itself without the need of additional labels and external supervisions. RESA aggregates information in vertical and horizontal directions by shifting sliced feature maps recurrently.

Lane regression algorithms can be grouped into key points estimation , anchor-based regression and row-wise regression . PINet combines key points estimation and instance segmentation, and GANet represents lanes as a set of key points which are only related to the start point. PointLaneNet and CurveLane-NAS separate images into non-overlapping grids and regress lanes based on vertical anchors. Line-CNN and LaneATT regress lanes on the pre-defined ray-anchors, while CLRNet dynamically refines the start point and angle of ray-anchors through pyramidal features. Ultra-Fast introduces a novel row-wise classification method with remarkable speed. Laneformer applies row-column self-attentions to accommodate the conventional Transformer to capture the shape characteristics and semantic contexts of lanes.

Except for point regression, polynomial regression is also a method for 2D lane detection task. PolyLaneNet uses a fully connected layer to directly predict the polynomial coefficients of lanes in the image plane. PRNet decomposes lane detection into three parts: polynomial regression, initial classification and height regression. Method in applies IPM and least square fitting to predict parabolic equations in BEV perspective. LSTR introduces a Transformer-based network to predict lane parameters which reflect road structures and the camera pose.

3D Lane Detection. 3D lane detection has attracted more attention than its 2D counterpart recently, and a reason is that the results of the latter lack depth information and spatial transformations have the error propagation problem. 3D-LaneNet is a dual-pathway architecture based on intra-network inverse-perspective mapping and anchor-based lane representation. 3D-LaneNet+ divides the BEV features into non-overlapping cells and detects lanes by regressing the lateral offset distance relative to the cell center, line angle and height offset. Method in introduces uncertainty estimation in order to enhance the capabilities of the network. Gen-LaneNet first introduces a new geometry-guided lane anchor representation in virtual top-view coordinate frame rather than the ego-vehicle coordinate frame, and applies a specific geometric transformation to calculate 3D lane points from the network output directly. CLGo replaces the CNN backbone with Transformer to predict camera pose and polynomial parameters. PersFormer builds a dense BEV query with known camera pose, and unifies 2D and 3D lane detection under one framework.

III METHOD

Fig. 3 shows the overview of our CurveFormer. It consists of three major components: (1) a Shared CNN Backbone takes a single front-view image as input and outputs multi-scale feature maps; (2) a Transformer Encoder to enhance the multi-scale feature maps subsequently and (3) a curve Transformer Decoder to propagates curve queries by curve cross-attention and iteratively refine anchor point sets. Finally, a prediction head is applied to output 3D lane parameters. The ii-th output can be represented as Predi=(pi,yistart,yiend,{ai,bi}r=0R)\text{Pred}_{i}=(p_{i},y_{i}^{start},y_{i}^{end},\{a_{i},b_{i}\}_{r=0}^{R}), where pip_{i} is the foreground confidence, yistarty_{i}^{start} and yiendy_{i}^{end} are start and end point in the YY direction. Two polynomials of 3D lane are denoted by aia_{i} and bib_{i} with order RR to model a traffic lane in X-O-Y and Y-O-Z plane, respectively.

III-B Shared Backbone and Transformer Encoder

The backbone takes an input image and outputs multi-scale feature maps. We add an auxiliary segmentation branch in the training stage to enhance the shared CNN backbone.

Similar to , in the decoder part, we apply multi-scale deformable self-attention module for each scale feature map to exchange information among different scales. The multi-scale feature maps are written as X={xl}l=1L\mathbf{X}=\left\{\mathbf{x}^{l}\right\}_{l=1}^{L}.

III-C Representing Sparse Curve Query with Dynamic Anchor Point Set

where positional encoding (PE) generates embeddings from floating numbers, and the parameters of the MLP are shared among all layers.

By representing a curve query as an ordered anchor point set {p1…pN}\{p_{1}\dots p_{N}\}, we can refine the curve query layer-by-layer in the Transformer decoder. Specifically, each Transformer decoder estimates relative positions ({Δx}1N,{Δz}1N)(\{\Delta x\}_{1}^{N},\{\Delta z\}_{1}^{N}) by a shared parameters linear layer. In this way, the curve query representation is suitable for 3D lane detection and is able to accelerate the learning convergence via layer-by-layer refinement scheme. Fig. 2 (b) shows the iterative refine process in the image plane.

III-D Curve Transformer Decoder

Our curve Transformer decoder contains a multi-head self-attention module, a context sampling module and a curve cross-attention module. We apply deformable attention in the self-attention module which focuses on a small set of key sampling points around the reference point, regardless of the spatial size of the feature map.

Context Sampling Module In deformable DETR , a learnable linear layer is used to predict offsets of the sampling locations corresponding to the reference points by queries, which are irrelevant to the image features. Different from it, we introduce a context sampling module to predict sampling offsets by incorporating more relative image features. Fig. 4 illustrates the difference between the standard sample offset module (a) and our context sampling module (b).

First, a dynamic anchor point set CiC_{i} is projected to the image view with camera parameters. We apply bilinear interpolation to extract features from these projected points Ci2D={p12D=(u1i,v1i),⋯ ,pN2D=(uNi,vNi)}C_{i}^{2D}=\{p_{1}^{2D}=(u_{1}^{i},v_{1}^{i}),\cdots,p_{N}^{2D}=(u_{N}^{i},v_{N}^{i})\} on multi-scale feature maps X{\mathbf{X}}. The final feature fCif_{C_{i}} is computed by

where σln\sigma_{ln} is used to determine whether a projected point pn2Dp_{n}^{2D} is outside ll-th feature map. And ϵ\epsilon is a small number to avoid division by zero.

We then use a learnable linear layer to predict KK sampling offsets. Typically, for a curve query Zq\mathbf{Z}_{q} with anchor point set CiC_{i}, the context sampling module denotes as:

Curve Cross Attention. We adapt deformable attention module in Deformable DETR to our curve cross-attention module. Mathematically, let qq be a query element in Zq\mathbf{Z}_{q}, and its anchor point set CiC_{i}, our curve cross-attention is calculated as:

where (m,l,n)(m,l,n) index the attention head, feature level and the sampling point. Δpmln\Delta\mathbf{p}_{mln} and AmlnA_{mln} denote sampling offsets and attention weights of the nn-th sampling point in the ll-th feature level and the mm-th attention head. The scalar attention weight AmlnA_{mln} is normalized to sum as 1. ϕc(⋅)\phi_{c}(\cdot) re-scales the normalized coordinates to input feature maps.

III-E Curve Training Supervision

In addition to the refined anchor point set P={pn}n=1N\mathbf{P}=\{p_{n}\}_{n=1}^{N}, the prediction head of our CurveFormer outputs curve parameters of LL 3D lanes, where LL is larger than the maximum number of labeled lanes across the training set. Similar to , we first associate the predicted curves Predi=(pi,yistart,yiend,{ai,bi}r=0R)\text{Pred}_{i}=(p_{i},y_{i}^{start},y_{i}^{end},\{a_{i},b_{i}\}_{r=0}^{R}) and ground truth lanes GTi=(p^i,y^istart,y^iend,L^i={p^n}1N)\text{GT}_{i}=(\hat{p}_{i},\hat{y}_{i}^{start},\hat{y}_{i}^{end},\hat{\mathbf{L}}_{i}=\{\hat{p}_{n}\}_{1}^{N}) by solving a bipartite matching problem, where c∈{0,1}c\in\{0,1\} (0: background, 1: lane). We sample a set of 3D point Li={pn}1N)\mathbf{L}_{i}=\{p_{n}\}_{1}^{N}) using the predicted curve parameters to compute the matching and training loss. The lane boundary (starting and ending points) is denoted by Lib={yistart,yiend}\mathbf{L}_{i}^{b}=\{y_{i}^{start},y_{i}^{end}\}.

Let Ω={wl=Predl}l=1L\Omega=\left\{w_{l}=\text{Pred}_{l}\right\}_{l=1}^{L} be the set of predicted 3D lanes and Π={π^l=GTl}l=1L\Pi=\left\{\hat{\pi}_{l}=\text{GT}_{l}\right\}_{l=1}^{L} be the set of groundtruth. Note that Π\Pi is padded with non-lanes to fill enough the number of ground truth lanes to LL. The matching problem is formulated as a cost minimization problem by searching an optimal injective function z:Π→Ωz:\Pi\rightarrow\Omega, where z(l)z(l) is the index of a 3D lane prediction ωz(l)\omega_{z(l)} which is assigned to ll-th ground truth 3D lane π^l\hat{\pi}_{l}:

where α1\alpha_{1}, α2\alpha_{2}, and α3\alpha_{3} are coefficients which adjust the loss effects of classification, polynomial fitting and boundary regression, and \mathds1\mathds{1} is an indicator function.

After solving Eq. 6 by Hungarian algorithms , the final training loss can be written as Ltotal=Lcurve+Lquery+LsegL_{total}=L_{curve}+L_{query}+L_{seg}, where LcurveL_{curve} is the curve prediction loss, LqueryL_{query} is the deep supervision of refined anchor point set for each curve, and LsegL_{seg} is an auxiliary segmentation loss. The curve prediction loss is defined as:

where α1\alpha_{1}, α2\alpha_{2}, and α3\alpha_{3} are the same coefficients with Eq. 7, and deep supervision of refined anchor point set is:

IV EXPERIMENTS

Apollo 3D Lane Synthetic Dataset. Apollo Synthetic dataset consists of over 10k 1080 × 1920 images which are built using unity 3D engine, including highway, urban, residential and downtown environments. The dataset is split into three different scenes: balanced scenes, rarely observed scenes and scenes with visual variations for evaluating algorithms from different perspectives.

OpenLane Dataset. OpenLane Dataset is the first real world 3D lane dataset which consists of over 200K frames at a frequency of 10 FPS based on Waymo Open dataset. In total, it has a training set with 157k images and a validation set of 39k images. The dataset provides camera intrisics and extrinsics following the same data format as Waymo Open Dataset.

IV-B Experiment Settings

Implementation Details. We use EfficientNet as backbone which gives 4 scale feature maps. The input image is resized to size of 360×480360\times 480. The 3D-space range is set to [−30m,30m]×[3m,103m]×[−10m,10m][-30m,30m]\times[3m,103m]\times[-10m,10m] along x,yx,y and zz axis respectively. For curve representation, we use fixed y-positions {5,10,15,20,30,40,50,60,80,100}\{5,10,15,20,30,40,50,60,80,100\}. We set coefficients to α1=2\alpha_{1}=2, α2=5\alpha_{2}=5, α3=2\alpha_{3}=2, and α4=2\alpha_{4}=2. All experiments are performed with known camera poses and intrinsic parameters provided by two datasets. Our network uses Adam optimizer , with a base learning rate of 2×10−42\times 10^{-4} and weight decay of 10−410^{-4}. All models are trained from scratch with 100 epochs and the per-GPU batch size is set to 4.

IV-C Evaluation Metrics and Results

Evaluation metrics. We follow the evaluation metrics designed by Gen-LaneNet . Point-wise Euclidean distance is calculated when a yy-position is covered by both prediction and the ground-truth. For each predicted lane, we consider it matched when 75%75\% of its covered yy-positions have point-wise euclidean distance less than the max-allowed distance (1.5 meters). We report Average Precision (AP) , F-score, and errors (near range and far range) to investigate the performance of our model.

Results on Apollo 3D Lane Synthetic Dataset. As shown in Table. I, we compare our CurveFormer with CNN-based 3D lane detection and Transformer-based 3D lane detection. Experimental results verify that our method outperforms the previous state-of-the-art approaches on Apollo 3D Lane Synthetic dataset. CurveFormer achieves the best F-Score and AP on every scene. Compared to PersFormer on three different scenes, CurveFormer significantly improved F-Score by 3.1%3.1\%, 9.0%9.0\% and 1.3%1.3\%, respectively.

Results on OpenLane Dataset. For OpenLane dataset, we evaluate CurveFormer on entire validation set and different scenario sets. In Table. II, our CurveFormer also gets comparable results compared to previous methods on the entire validation set, and achieves the highest F-Score on five scenario sets. We present detailed comparison with previous 3D lane detection SOTAs in Table. III.

IV-D Ablation Study

In this section, we analyze the effects of the proposed key components via the ablation study conducted on Apollo 3D Lane Synthetic dataset .

Network Output: Curve Parameters vs Anchor Point Set. We compare two different network outputs for 3D lane detection, curve parameters estimation and anchor point set prediction. The latter one is further interpolated as 3D curve for evaluation. Table. V lists the performance comparison on Apollo 3D Lane Synthetic dataset. It shows that curve parameter prediction largely surpasses using anchor point set as network output, due to that lane parameter prediction can preserve the geometric property of 3D lane compared to predicting separated points.

Context Sampling. We study the impact of the different ways to produce sampling offsets corresponding to 3D lane reference points. As shown in Table. VI, using context sampling offset (CSO) achieves best F-Score and AP compared to using standard sampling offset (SO). The results demonstrate the significance of image-query correlation for feature aggregation in our cross-attention module.

Number of Decoder Layer. We vary the number of decoder layers and the performance of the model is shown in Table. IV. It shows that using 4 decoder layers in our network achieves best performance. We can use few decoder layers due to curve propagation scheme of our CurveFormer method. Therefore, we set the number of decoder layers in our CurveFormer to 4 by default in the experiments.

Auxiliary Segmentation. Lastly, we study the effect of the auxiliary segmentation branch. Experimental results show that the auxiliary segmentation branch can slightly improve F-Score by 0.13, AP by 0.06 on the Balanced Scenes test set.

V CONCLUSIONS

In this paper, we introduce CurveFormer, a Transformer-based 3D lane detection method. It uses dynamic anchor point set to construct queries, and refines it layer-by-layer in Transformer decoders. In addition, to attend to more relevant image features, we present a curve cross-attention module and a context sampling module to compute the key-to-image similarity. In the experiments, we show that CurveFormer achieves promising results compared with both CNN-based and Transformer-based approaches. In future work, we would like to explore video-based 3D lane detection for autonomous driving.

References