LiDAR-based Online 3D Video Object Detection with Graph-based Message Passing and Spatiotemporal Transformer Attention

Junbo Yin, Jianbing Shen, Chenye Guan, Dingfu Zhou, Ruigang Yang

Introduction

LiDAR-based 3D object detection plays a critical role in a wide range of applications, such as autonomous driving, robot navigation and virtual/augmented reality geiger2012we; song2019apollocar3d. The majority of current 3D object detection approaches shi2019pointrcnn; yang2019std; chen2019fast; zhou2018voxelnet; lang2019pointpillars follow the single-frame detection paradigm, while few of them perform detection in the point cloud video. A point cloud video is defined as a temporal sequence of point cloud frames. For instance, in the nuScenes dataset caesar2019nuScenes, 2020 point cloud frames can be captured per second with a modern 32-beam LiDAR sensor. Detection in single frame may suffer from several limitations due to the sparse nature of point cloud. In particular, occlusions, long-distance and non-uniform sampling inevitably occur on a certain frame, where a single-frame object detector is incapable of handling these situations, leading to a deteriorated performance, as shown in Fig 1. However, a point cloud video contains rich spatiotemporal information of the foreground objects, which can be explored to improve the detection performance. The major concern of constructing a 3D video object detector is how to model the spatial and temporal feature representation for the consecutive point cloud frames. In this work, we propose to integrate a graph-based spatial feature encoding component with an attention-aware spatiotemporal feature aggregation component, to capture the video coherence in consecutive point cloud frames, which yields an end-to-end online solution for the LiDAR-based 3D video object detection.

Popular single-frame 3D object detectors tend to first discretize the point cloud into voxel or pillar girds zhou2018voxelnet; yan2018second; lang2019pointpillars, and then extract the point cloud features using stacks of convolutional neural networks (CNNs). Such approaches incorporate the success of existing 2D or 3D CNNs and usually gain better computational efficiency compared with the point-based methods shi2019pointrcnn; qi2019deep. Therefore, in our spatial feature encoding component, we also follow this paradigm to extract features for each input frame. However, a potential problem with these approaches lies in that they only focus on a locally aggregated feature, i.e., employing a PointNet qi2017pointnet to extract features for separate voxels or pillars as in zhou2018voxelnet and lang2019pointpillars. To further enlarge the receptive fields, they have to apply the stride or pooling operations repeatedly, which will cause the loss of the spatial information. To alleviate this issue, we propose a novel graph-based network, named Pillar Message Passing Network (PMPNet), which treats a non-empty pillar as a graph node and adaptively enlarges the receptive field for a node by aggregating messages from its neighbors. PMPNet can mine the rich geometric relations among different pillar grids in a discretized point cloud frame by iteratively reasoning on a kk-NN graph. This effectively encourages information exchanges among different spatial regions within a frame.

After obtaining the spatial features of each input frame, we assemble these features in our spatiotemporal feature aggregation component. Since ConvGRU ballas2016delving has shown promising performance in the 2D video understanding field, we suggest an Attentive Spatiotemporal Transformer GRU (AST-GRU) to extend ConvGRU to the 3D field through capturing dependencies of consecutive point cloud frames with an attentive memory gating mechanism. Specifically, there exist two potential limitations when considering the LiDAR-based 3D video object detection in autonomous driving scenarios. First, in the bird’s eye view, most foreground objects (e.g., cars and pedestrians) occupy small regions, and the background noise is inevitably accumulated as computing the new memory in a recurrent unit. Thus, we propose to exploit the Spatial Transformer Attention (STA) module, an intra-attention derived from vaswani2017attention; wang2018non, to suppress the background noise and emphasize the foreground objects by attending each pixel with the context information. Second, when updating the memory in the recurrent unit, the spatial features of the two inputs (i.e., the old memory and the new input) are not well aligned. In particular, though we can accurately align the static objects across frames using the ego-pose information, the dynamic objects with large motion are not aligned, which will impair the quality of the new memory. To address this, we propose a Temporal Transformer Attention (TTA) module that adaptively captures the object motions in consecutive frames with a temporal inter-attention mechanism. This will better utilize the modified deformable convolutional layers zhu2019deformable; zhu2019empirical. Our AST-GRU can better handle the spatiotemporal features and produce a more reliable new memory, compared with the vanilla ConvGRU. To summarize, we propose a new LiDAR-based online 3D video object detector that leverages the previous long-term information to improve the detection performance. In our model, a novel PMPNet is introduced to adaptively enlarge the receptive field of the pillar nodes in a discretized point clod frame by iterative graph-based message passing. The output sequential features are then aggregated in the proposed AST-GRU to mine the rich coherence in the point cloud video by using an attentive memory gating mechanism. Extensive evaluations demonstrate that our 3D video object detector achieves better performance against the single-frame detectors on the large-scale nuScenes benchmark.

Related Work

Both the VFE layers and the PFN only take into account separate voxels or pillars when generating the grid-level representation, which ignores the information exchange in larger spatial regions. In contrast, our PMPNet encodes the pillar feature from a global perspective by graph-based message passing, and thus promotes the representation with the non-local property. Besides, all these single-frame 3D object detectors can only process the point cloud data frame-by-frame, lacking the exploration of the temporal information. Though luo2018fast applies temporal 3D ConvNet on point cloud sequences, it encounters the feature collapse issue when downsampling the features in the temporal domain. Moreover, it cannot deal with long-term sequences with multi-frame labels. Our AST-GRU instead captures the long-term temporal information with an attentive memory gating mechanism, which can fully mine the spatiotemporal coherence in the point cloud video.

Graph Neural Networks. Graph Neural Networks (GNNs) are first introduced by Gori et al. gori2005new to model the intrinsic relationships of the graph-structured data. Then Scarselli et al. scarselli2008graph extend it to different types of graphs. Afterward, GNNs are explored in two directions in terms of different message propagation strategies. The first group li2016gated; kearnes2016molecular; zayats2018conversation; peng2017cross; qi2018learning uses the gating mechanism to enable the information to propagate across the graph. For instance, Li et al. li2016gated leverage the recurrent neural networks to describe the state of each graph node. Then, Gilmer et al. gilmer2017neural generalizes a framework to formulate the graph reasoning as a parameterized message passing network. Another group bruna2014spectral; henaff2015deep; defferrard2016convolutional; hu2020infinitely; li2020self integrates convolutional networks to the graph domain, named as Graph Convolutional Neural Networks (GCNNs), which update node features via stacks of graph convolutional layers. GNNs have achieved promising results in many areas defferrard2016convolutional; fan2019understanding; wang2018attentive; battaglia2016interaction; wang2020hierarchical due to the great expressive power of graphs. Our PMPNet belongs to the first group by capturing the pillar features with a gated message passing strategy, which is used to construct the spatial representation for each point cloud frame.

Model Architecture

In this section, we elaborate on our online 3D video object detection framework. As shown in Fig. 2, it consists of a spatial feature encoding component and a spatiotemporal feature aggregation component. Given the input sequences {It}t=1T\{\bm{I}_{t}\}_{t=1}^{T} with TT frames, we first convert the point cloud coordinates from the previous frames {It}t=1T−1\{\bm{I}_{t}\}_{t=1}^{T-1} to the current frame IT\bm{I}_{T} using the GPS data, so as to eliminate the influence of the ego-motion and align the static objects across frames. Then, in the spatial feature encoding component, we extract features for each frame with the Pillar Message Passing Network (PMPNet) (§3.1) and a 2D backbone, producing sequential features {Xt}t=1T\{\bm{X}_{t}\}_{t=1}^{T}. After that, these features are fed into the Attentive Spatiotemporal Transformer Gated Recurrent Unit (AST-GRU) (§3.2) in the spatiotemporal feature aggregation component, to generate the new memory features {Ht}t=1T\{\bm{H}_{t}\}_{t=1}^{T}. Finally, a RPN head is applied on {Ht}t=1T\{\bm{H}_{t}\}_{t=1}^{T} to give the final detection results {Yt}t=1T\{\bm{Y}_{t}\}_{t=1}^{T}. Some network architecture details are provided in §3.3.

Previous point cloud encoding layers (e.g., the VFE layers in zhou2018voxelnet and the PFN in lang2019pointpillars) for voxel-based 3D object detection typically encode each voxel or pillar separately, which limits the expressive power of the grid-level representation due to the small receptive field of each local grid region. Our PMPNet instead seeks to explore the rich spatial relations among different gird regions by treating the non-empty pillar grids as graph nodes. Such design effectively reserves the non-Euclidean geometric characteristics of the original point clouds and enhance the output pillar features with a non-locality property.

Given an input point cloud frame It\bm{I}_{t}, we first uniformly discretize it into a set of pillars P\mathcal{P}, with each pillar uniquely associated with a spatial coordinate in the x-y plane as in lang2019pointpillars. Then, PMPNet maps the resultant pillars to a directed graph G=(V,E)\mathcal{G}=(\mathcal{V},\mathcal{E}), where node vi∈Vv_{i}\in\mathcal{V} represents a non-empty pillar Pi∈PP_{i}\in\mathcal{P} and edge ei,j∈E{e}_{i,j}\in\mathcal{E} indicates the message passed from node viv_{i} to vjv_{j}. For reducing the computational overhead, we define G\mathcal{G} as a kk-nearest neighbor (kk-NN) graph, which is built from the geometric space by comparing the centroid distance among different pillars.

To explicitly mine the rich relations among different pillar nodes, PMPNet performs iterative message passing on G\mathcal{G} and updates the nodes state at each iteration step. Concretely, given a node viv_{i}, we first utilize a pillar feature network (PFN) lang2019pointpillars to describe its initial state hi0\bm{h}_{i}^{0} at iteration step s=0s=0:

Next, we elaborate on the message passing process. One iteration step of message propagation is illustrated in Fig. 3. At step ss, a node viv_{i} aggregates information from all the neighbor nodes vj∈Ωviv_{j}\in\bm{\Omega}_{v_{i}} in the kk-NN graph. We define the incoming edge feature from node vjv_{j} as ej,is\bm{e}_{j,i}^{s}, indicating the relation between node viv_{i} and vjv_{j}. Inspired by wang2019dynamic, the incoming edge feature ej,is\bm{e}_{j,i}^{s} is given by:

which is an asymmetric function encoding the local neighbor information. Accordingly, we have the message passed from vjv_{j} to viv_{i}, which is denoted as:

where ϕθ\phi_{\theta} is parameterized by a fully connected layer, which takes as input the concatenation of his\bm{h}_{i}^{s} and ej,is\bm{e}_{j,i}^{s}, and yields a L′L^{\prime}-dim feature.

After computing all the pair-wise relations between viv_{i} and the neighbors vj∈Ωvi{v_{j}}\in{\bm{\Omega}_{v_{i}}} of , we summarize the received kk messages with a maximum operation:

Then, we update the node state his\bm{h}_{i}^{s} with his+1\bm{h}_{i}^{s+1} for node viv_{i}. The update process should consider both the newly collected message mis+1\bm{m}_{i}^{s+1} and the previous state his\bm{h}_{i}^{s}. Recurrent neural network and its variants hochreiter1997long; sutskever2014sequence can adaptively capture dependencies in different time steps. Hence, we utilize Gated Recurrent Unit (GRU) cho2014learning as the update function for its better convergence characteristic. The update process is then formulated as follows:

In this way, the new node state his+1\bm{h}_{i}^{s+1} contains the information from all the neighbor nodes of viv_{i}. Moreover, a neighbor node vjv_{j} also collects information from its own neighbors Ωvj\bm{\Omega}_{v_{j}}. Consequently, after the totally SS iteration steps, node viv_{i} is able to aggregate information from the high-order neighbors. This effectively enlarges the perceptual range for each pillar grid and enables our model to better recognize objects from a global view.

where FBF_{\text{B}} denotes the backbone network and Xt\bm{X}_{t} is the spatial features of It\bm{{I}}_{t}. Details of the PMPNet and the backbone network can be found in §3.3.

2 Attentive Spatiotemporal Transformer GRU

Since the sequential features {Xt}t=1T\{\bm{X}_{t}\}_{t=1}^{T} produced by the spatial feature encoding component are regular tensors, we can employ the ConvGRU ballas2016delving to fuse these features in our spatiotemporal feature aggregation component. However, it may suffer from two limitations when directly applying the ConvGRU. On the one hand, the interest objects are relatively small in the bird’s eye view compared with those in the 2D images (e.g., an average of 18×818\times 8 pixels for cars with the pillar size of 0.252 m20.25^{2}~m^{2}). This may cause the background noise to dominate the results when computing the memory. On the other hand, though the static objects can be well aligned across frames using the GPS data, the dynamic objects with large motion still lead to an inaccurate new memory. To address the above issues, we propose the AST-GRU to equip the vanilla ConvGRU ballas2016delving with a spatial transformer attention (STA) module and a temporal transformer attention (TTA) module. As illustrated in Fig. 4, the STA module stresses the foreground objects in {Xt}t=1T\{\bm{X}_{t}\}_{t=1}^{T} and produces the attentive new input {Xt′}t=1T\{\bm{X}_{t}^{{}^{\prime}}\}_{t=1}^{T}, while the TTA module aligns the dynamic objects in {Ht−1}t=1T\{\bm{H}_{t-1}\}_{t=1}^{T} and {Xt′}t=1T\{\bm{X}_{t}^{{}^{\prime}}\}_{t=1}^{T}, and outputs the attentive old memory {Ht−1′}t=1T\{\bm{H}_{t-1}^{{}^{\prime}}\}_{t=1}^{T}. Then, {Xt′}t=1T\{\bm{X}_{t}^{{}^{\prime}}\}_{t=1}^{T} and {Ht−1′}t=1T\{\bm{H}_{t-1}^{{}^{\prime}}\}_{t=1}^{T} are used to generate the new memory {Ht}t=1T\{\bm{H}_{t}\}_{t=1}^{T}, and further produce the final detections {Yt}t=1T\{\bm{Y}_{t}\}_{t=1}^{T}. Before giving the details of the STA and TTA modules, we first review the vanilla ConvGRU.

Spatial Transformer Attention. The core idea of the STA module is to attend each pixel-level feature x∈Xt\bm{x}\in\bm{X}_{t} with a rich spatial context, to better distinguish a foreground object from the background noise. Basically, a transformer attention receives a query xq∈Xt\bm{x}_{q}\in\bm{X}_{t} and a set of keys xk∈Ωxq\bm{x}_{k}\in\bm{\Omega}_{\bm{x}_{q}} (e.g., the neighbors of xq\bm{x}_{q}), to calculate an attentive output yq\bm{y}_{q}. The STA is designed as an intra-attention, which means both the query and key are from the same input feature Xt\bm{X}_{t}.

Formally, given a query xq∈Xt\bm{x}_{q}\in\bm{X}_{t} at location q∈w×hq\in{w\times{h}}, the attentive output yq\bm{y}_{q} is computed by:

where A(⋅,⋅)A(\cdot,\cdot) is the attention weight. ϕK\phi_{K}, ϕQ\phi_{Q} and ϕV\phi_{V} are the linear layers that map the inputs xq,xk∈Xt\bm{x}_{q},\bm{x}_{k}\in\bm{X}_{t} into different embedding subspaces. The attention weight A(⋅,⋅)A(\cdot,\cdot) is computed from the embedded query-key pair (ϕQ(xq),ϕK(xk))(\phi_{Q}(\bm{x}_{q}),\phi_{K}(\bm{x}_{k})), and is then applied to the neighbor values ϕV(xk)\phi_{V}(\bm{x}_{k}).

Temporal Transformer Attention. To adaptively align the features of dynamic objects from Ht−1\bm{H}_{t-1} to Xt′\bm{X}_{t}^{{}^{\prime}}, we apply the modified deformable convolutional layers zhu2019deformable; zhu2019empirical as a special instantiation of the transformer attention. The core is to attend the queries in Ht−1\bm{H}_{t-1} with adaptive supporting key regions computed by integrating the motion information.

Specifically, given a vanilla deformable convolutional layer with kernel size 3×33\times 3, let wm\bm{w}_{m} denotes the learnable weights, and pm∈{(−1,−1),(−1,0),...,(1,1)}p_{m}\in\{(-1,-1),(-1,0),...,(1,1)\} indicates the predetermined offset in total M=9M=9 grids. The output hq′\bm{h}_{q}^{{}^{\prime}} for input hq∈Ht−1\bm{h}_{q}\in{\bm{H}_{t-1}} at location q∈w×hq\in{w\times{h}} can be expressed as:

where ϕV\phi_{\text{V}} is an identity function, and wm\bm{w}_{m} acts as the weights in different attention heads zhu2019empirical, with each head corresponding to a sampled key position k∈Ωqk\in{\bm{\Omega}_{q}}. G(⋅,⋅)G(\cdot,\cdot) is the attention weight defined by a bilinear interpolation function, such that G(a,b)=max(0,1−∣a−b∣)G(a,b)=max(0,1-|a-b|).

The supporting key regions Ωq{\bm{\Omega}_{q}} play an important role in attending hq\bm{h}_{q}, which are determined by the deformation offset Δpm∈ΔPt−1\Delta{p_{m}}\in\Delta{\bm{P}_{t-1}}. In our TTA module, we compute ΔPt−1\Delta{\bm{P}_{t-1}} not only through Ht−1\bm{H}_{t-1}, but also through a motion map, which is defined as the difference of Ht−1\bm{H}_{t-1} and Xt′\bm{X}_{t}^{{}^{\prime}}:

where ΦR\Phi_{R} is a regular convolutional layer with the same kernel size as that in the deformable convolutional layer, and [⋅,⋅][\cdot,\cdot] is the concatenation operation. The intuition is that, in the motion map, the features response of the static objects is very low since they have been spatially aligned in Ht−1\bm{H}_{t-1} and Xt′\bm{X}_{t}^{{}^{\prime}}, while the features response of the dynamic objects remains high. Therefore, we integrate Ht−1\bm{H}_{t-1} with the motion map, to further capture the motions of dynamic objects. Then, ΔPt−1\Delta{\bm{P}_{t-1}} is used to select the supporting key regions and further attend Ht−1\bm{H}_{t-1} for all the query regions q∈w×hq\in{w\times{h}} in terms of Eq. 15, yielding a temporally attentive memory Ht−1′\bm{H}_{t-1}^{{}^{\prime}}. Since the supporting key regions are computed from both Ht−1\bm{H}_{t-1} and Xt′\bm{X}_{t}^{{}^{\prime}}, our TTA module can be deemed as an inter-attention.

Additionally, we can stack multiple modified deformable convolutional layers to get a more accurate Ht−1′\bm{H}_{t-1}^{{}^{\prime}}. In our implementation, we adopt two layers. The latter layer takes as input [Ht−1′,Ht−1′−Xt′][\bm{H}_{t-1}^{{}^{\prime}},\bm{H}_{t-1}^{{}^{\prime}}-\bm{X}_{t}^{{}^{\prime}}] to predict the deformation offset according to Eq. 16, and the offset is then used to attend Ht−1′\bm{H}_{t-1}^{{}^{\prime}} via Eq. 15. Accordingly, we can now utilize the temporally attentive memory Ht−1′\bm{H}_{t-1}^{{}^{\prime}} and the spatially attentive input Xt′\bm{X}_{t}^{{}^{\prime}} to compute the new memory Ht\bm{H}_{t} in the recurrent unit (see Fig. 4). Finally, a RPN detection head is applied on Ht\bm{H}_{t} to produce the final detection results Yt\bm{Y}_{t}.

3 Network Details

AST-GRU Module. In our STA module, all the linear functions in Eq. 11 and Wout\bm{W}_{\text{out}} in Eq. 13 are 1×11\times{1} convolution layers. In our TTA module, the regular convolutional layers, the deformable convolutional layers and the ConvGRU all have learable kernels of size 3×33\times{3}.

Detection Head. The detection head in zhou2018voxelnet is applied on the attentive memory features. In particular, the smooth L1 loss and the focal loss lin2017focal count for the object bounding box regression and classification, respectively. A corss-entropy loss is used for the orientation classification. For the velocity regression required by the nuScenes benchmark, a simple L1 loss is adopted and shows substantial results.

Experimental Results

Implementation Details. For each keyframe, we consider the point clouds within range of ××\times\times meters along the X, Y and Z axes. The pillar resolution on the X-Y plane is 0.252 m20.25^{2}~m^{2}. The pillar number PP used in PMPNet is 16,384, sampled from the total 25,000 pillars, with each pillar containing most N=60N=60 points. The input point cloud is a D=5D=5 dimensions representation (x,y,z,r,Δt)(x,y,z,r,\Delta t), which are then embedded into L=L′=64L=L^{{}^{\prime}}=64 dimensions feature space after the total S=3S=3 graph iteration steps. The convolutional kernels in the 2D backbone are of size Z=3Z=3 and the output channel number CC in each block is (64,128,256)(64,128,256). The upsampling layer has kernel size 3 and channel number 128. Thus, the final features map produced by the 2D backbone has a resolution of 100×100×384100\times 100\times 384. We calculate anchors for different classes using the mean sizes and set the matching threshold according to the class instance number. The coefficients of the loss functions for classification, localization and velocity prediction are set to 1, 2 and 0.1, respectively. NMS with IOU threshold 0.5 is utilized when generating the final detections. In both training and testing phases, we feed most 3 consecutive keyframes to the model due to the memory limitation. The training procedure has two stages. In the first stage, we pre-train the spatial features encoding component using the one-cycle policy smith2017cyclical with a maximum learning rate of 0.003. Then, we fix the learning rate to 0.0002 in the second stage to train the full model. We train 50 epochs for both stages with batch size 3. Adam optimizer kingma2015adam is used to optimize the loss functions.

We present the performance comparison of our algorithm and other state-of-the-art approaches on the nuScenes benchmark in Table 1. PointPillars lang2019pointpillars, SARPNET ye2020sarpnet, WYSIWYG hu2019you and Tolist leaderboard are all voxel-based single-frame 3D object detectors. In particular, PointPillars is used as the baseline of our model. WYSIWYG is a recent algorithm that extends the PointPillars with a voxelized visibility map. Tolist uses a multi-head network that contains multiple prediction heads for different classes. Our 3D video object detector outperforms these approaches by a large margin. In particular, we improve the official PointPillars algorithm by 15%. Please note that there is a severe class imbalance issue in the nuScenes dataset. The approach in zhu2019class designs a class data augmentation algorithm. Further integrating with these techniques can promote the performance of our model. But we focus on exploring the spatiotemporal coherence in the point cloud video, and handling the class imbalance issue is not the purpose in this work. In addition, we further show some qualitative results in Fig. 5. Besides the occlusion situation in Fig. 1, we present another case of detecting the distant car (the car on the top right), whose point clouds are especially sparse, which is very challenging for the single-frame detectors. Again, our 3D video object detector effectively detects the distant car using the attentive temporal information.

2 Ablation Study

In this section, we investigate the effectiveness of each module in our algorithm. Since the training samples in the nuScenes is 7×\times as many as those in the KITTI (28,130 vs 3,712), it is non-trivial to train multiple models on the whole dataset. Hence, we use a mini train set for validation purposes. It contains around 3,500 samples uniformly sampled from the original train set. Besides, PointPillars lang2019pointpillars is used as the baseline detector in our model.

Conclusion

This paper proposed a new 3D video object detector for exploring the spatiotemporal information in point cloud video. It has developed two new components: spatial feature encoding component and spatiotemporal feature aggregation component. We first introduce a novel PMPNet that considers the spatial features of each point cloud frame. PMPNet can effectively enlarge the receptive field of each pillar grid through iteratively aggregating messages on a kk-NN graph. Then, an AST-GRU module composed of STA and TTA is presented to mine the spatiotemporal coherence in consecutive frames by using an attentive memory gating mechanism. The STA focuses on detecting the foreground objects, while the TTA aims to align the dynamic objects. Extensive experiments on the nuScenes benchmark have proved the better performance of our model.

References