Panoptic NeRF: 3D-to-2D Label Transfer for Panoptic Urban Scene Segmentation

Xiao Fu, Shangzhan Zhang, Tianrun Chen, Yichong Lu, Lanyun Zhu, Xiaowei Zhou, Andreas Geiger, Yiyi Liao

Introduction

Semantic instance segmentation is an important perception task for autonomous driving. It is widely acknowledged that large-scale training data with high-quality annotations is critical to propel the performance of segmentation models. However, manual annotation of pixel-accurate segmentation masks is highly expensive and time-consuming. For example, annotating all instances in a single street scene image requires up to 1.5 hours . Recently, a few urban datasets propose to annotate in 3D space using coarse bounding primitives (e.g., cuboids and ellipsoids) and transfer 3D labels to 2D , significantly reducing the annotation time to 0.75 minutes per image . Besides, it is often easier to separate instances in 3D rather than in 2D image space (e.g., pedestrian in front of building). Thus, there is an increasing demand for accurately transferring coarse 3D annotations to 2D semantic and instance labels.

There are only a few existing attempts in this direction. Huang et al. perform manual post-processing to obtain accurate 2D labels based on 3D coarse annotations. Another line of work additionally leverages 2D image cues (e.g., noisy 2D semantic predictions) to avoid human intervention, achieved by combining 3D annotations and 2D image cues based on conditional random fields (CRF) . This CRF-based approach relies on intermediate 3D reconstructions for projecting non-occluded 3D points to 2D, and then performs inference in 2D image space. The 3D reconstruction cannot be jointly optimized in the CRF model and thus erroneous reconstruction leads to inaccurate label transfer results. To alleviate this problem, we propose Panoptic NeRF, a novel label transfer method built on NeRF that infers geometry and semantic jointly in 3D space to render dense 2D semantic and instance labels, i.e., panoptic segmentation labels (see Fig. 1).

A naïve solution is to first train a vanilla NeRF model, and then render segmentation maps based on semantic/instance labels determined by the 3D bounding primitives. However, as illustrated in Fig. 1, inaccurate geometric reconstruction leads to wrong semantic/instance maps, yet it is hard to obtain accurate geometry using the vanilla NeRF in the driving scenario where input views are sparse. Furthermore, label ambiguity at overlapping regions of the 3D bounding primitives also yields inaccurate 2D labels.

In this work, we aim to tackle both challenges. Inspired by , we combine 3D annotations and noisy 2D semantic predictions transferred from existing datasets to fully automate the label transfer process. Specifically, our method consists of a radiance field and dual semantic fields, supervised by 3D and 2D weak semantic information as well as posed 2D RGB images. To improve the underlying geometry, we propose a semantically-guided geometry optimization strategy based on a fixed semantic field determined by the 3D bounding primitives. With the semantic field fixed, we demonstrate that the geometry can be improved guided by noisy 2D semantic predictions. The semantic rendering is then further refined by joint geometry and semantic optimization, where a learned semantic field is adopted to fuse information of the 3D bounding primitives and the 2D noisy predictions. As evidenced by our experiments, this fusion procedure is able to resolve the label ambiguity of the 3D bounding primitives and largely eliminate noise in the 2D predictions. Furthermore, Panoptic NeRF enables rendering globally consistent 2D instance maps across multiple frames, where each object has a unique instance index determined by the 3D bounding primitives. Utilizing 3D bounding primitives of the recently released KITTI-360 dataset, Panoptic NeRF outperforms existing 3D-to-2D and 2D-to-2D label transfer methods, providing a promising approach to efficiently develop large-scale and densely labeled datasets for autonomous driving.

We summarize our contributions as follows: 1) We propose to perform 3D-to-2D label transfer by inferring in the 3D space. This allows us to unify easy-to-obtain 3D bounding primitives and noisy 2D semantic predictions in a single model, yielding high-quality panoptic labels. 2) By leveraging a novel dual formulation of the semantic fields, Panoptic NeRF effectively improves the geometric reconstruction given sparse views, yielding accurate object boundaries. Moreover, it is able to resolve label ambiguities and eliminates label noise based on the improved geometry. 3) Our Panoptic NeRF achieves superior performance compared to existing label transfer methods in terms of both semantic and instance predictions. Furthermore, our 2D semantic and instance labels are multi-view and spatio-temporally consistent by design. Finally, our method enables rendering RGB images and semantic/instance labels at novel viewpoints.

Related Work

Urban Scene Segmentation: Semantic instance segmentation is a critical task for autonomous vehicles . Learning-based algorithms have achieved compelling performance , but rely on large-scale training data. Unfortunately, annotating images at pixel level is extremely time-consuming and labor-intensive, especially for instance-level annotation. While most urban datasets provides labels in 2D image space , autonomous vehicles are usually equipped with 3D sensors . KITTI-360 demonstrates that annotating the scene in 3D can significantly reduce the annotation time. However, transferring coarse 3D labels to 2D remains challenging. In this work, we focus on developing a novel 3D-to-2D label transfer method, exploiting recent advances in neural scene representations.

Label Transfer: There have been several attempts at improving label efficiency for individual frames . In this paper we focus on efficient labeling of video sequences. Existing works in this area can be divided into two categories: 2D-to-2D and 3D-to-2D. 2D-to-2D label transfer approaches reduce the workload by propagating labels across 2D images , whereas 3D-to-2D methods exploit additional information in 3D for efficient labeling . To obtain dense labels in 2D image space, some works perform per-frame inference jointly over the 3D point clouds and 2D pixels using a non-local multi-field CRF model. However, these methods require reconstructing a 3D mesh to project 3D point clouds to 2D. As it is treated as a pre-processing step, the mesh reconstruction is not jointly optimized in the CRF model. Thus, inaccurate reconstruction hinders label transfer performance. In contrast, Panoptic NeRF provides a novel end-to-end method for 3D-to-2D label transfer where geometry and semantic estimations are jointly optimized.

Coordinate-based Neural Representations: Recently, coordinate-based neural representations has received wide attention. in many areas, including 3D reconstruction , novel view synthesis , and 3D generative modeling . In this paper, we focus on utilizing coordinate-based representations to estimate the semantics of the scene. Towards this goal, NeSF focuses on generalizable semantic field learning from density grids supervised by 2D GT labels, but we concentrate on the 3D-to-2D label transfer task without access to 2D GT. A closely related work Semantic NeRF explores NeRF for semantic fusion. However, Semantic NeRF takes as input ground truth 2D labels or synthetic noise labels, which struggles to produce correct labels given real-world predictions from pre-trained 2D models. Moreover, Semantic NeRF operates in indoor scenes with dense RGB inputs and degenerates in challenging outdoor driving scenarios with sparse input views. Finally, Semantic NeRF is limited to rendering semantic labels, whereas our method can render panoptic labels. A concurrent work PNF allows for rendering papotic labels. While PNF focuses on parsing the scene using known classes of the Cityscapes dataset based on pre-trained segmentation and detection models, we aim to transfer 3D annotations to 2D image space for arbitrary classes to enable the development of new datasets, e.g., providing instance labels for buildings that are not available in Cityscapes.

Background

NeRF: NeRF models a 3D scene as a continuous neural radiance field fθf_{\theta}. Specifically, it maps a 3D coordinate x\mathbf{x} and a viewing direction d\mathbf{d} to a volume density σ\sigma and an RGB color value c\mathbf{c}:

Let r(t)=o+td\mathbf{r}(t)=\mathbf{o}+t\mathbf{d} denote a camera ray. The color at the corresponding pixel can be obtained by volume rendering

where σi\sigma_{i} and ci\mathbf{c}_{i} denote the density and color value at a point ii sampled along the ray, TiT_{i} denotes the transmittance at the sample point, and δk=ti+1−ti\delta_{k}=t_{i+1}-t_{i} is the distance between adjacent samples. Let π\pi denote the volume rendering process of one ray. Enabled by volume rendering, NeRF learns fθf_{\theta} from a set of 2D RGB images with known camera poses.

Problem Formulation: As shown in Fig. 1, Panoptic NeRF aims to transfer coarse 3D bounding primitives to dense 2D semantic and instance labels. In addition to a sparse set of posed RGB images, we assume a set of 3D bounding primitives β={Bk}k=1K\beta=\left\{B_{k}\right\}_{k=1}^{K} to be available. These 3D bounding primitives cover the full scene in the form of cuboids, ellipsoids and extruded polygons. Each 3D bounding primitive BkB_{k} has a semantic label, belonging to either “stuff” or “thing”. For “thing” classes, BkB_{k} is additionally associated with a unique instance ID. We further apply a pre-trained semantic segmentation model to the RGB images to obtain a 2D semantic prediction for each image. With this input information, our primary goal is to generate multi-view consistent, semantic and instance labels at the input frames, whereas rendering RGB images and panoptic labels at novel viewpoints is also enabled by inferring in the 3D space.

Methodology

Panoptic NeRF provides a novel method for label transfer from 3D to 2D. Fig. 2 gives an overview of our method. We first map a 3D point x\mathbf{x} to a density σ\sigma and a color value c\mathbf{c} using a radiance field, as well as two semantic categorical distributions s^\hat{\mathbf{s}} and s\mathbf{s} based on our dual semantic fields (Section 4.1). Correspondingly, for each camera ray, two semantic categorical distributions S^\hat{\mathbf{S}} and S\mathbf{S} in 2D image space via volume rendering π\pi are obtained. Based on semantic losses in both 3D and 2D space (Section 4.2), the fixed semantic field sβs_{\beta} serves to improve geometry, while the learned semantic field sϕs_{\phi} results in improved semantics. With the 3D bounding primitives, we further define a fixed instance field tβt_{\beta} that allows for rendering panoptic label T\mathbf{T} when combined with the learned semantic field (Section 4.3).

To jointly improve geometry and semantics, we define dual semantic fields, one is determined by the 3D bounding primitives β\beta and the other is learned by a semantic head ϕ\phi

where MsM_{s} denotes the number of semantic classes. In combination with the volume density of the radiance field fθf_{\theta}, two semantic distributions S^(r)\hat{\mathbf{S}}({\mathbf{r}}) and S(r){\mathbf{S}}({\mathbf{r}}) can be obtained at each camera ray r\mathbf{r} via the volume rendering operation π\pi:

Note that S^(r)\hat{\mathbf{S}}(\mathbf{r}) and S(r){\mathbf{S}}(\mathbf{r}) are both normalized distributions when ∑i=1NTi(1−exp⁡(−σiδi))=1\sum_{i=1}^{N}T_{i}(1-\exp(-\sigma_{i}\delta_{i}))=1. We set the background class to sky if ∑i=1NTi(1−exp⁡(−σiδi))<1\sum_{i=1}^{N}T_{i}(1-\exp(-\sigma_{i}\delta_{i}))<1. We apply losses to both S^(r)\hat{\mathbf{S}}({\mathbf{r}}) and S(r){\mathbf{S}}({\mathbf{r}}) for training. During inference, the semantic label is determined as the class of the maximum probability in S(r){\mathbf{S}}(\mathbf{r}).

Fixed Semantic Field: If x\mathbf{x} is uniquely enclosed by a 3D bounding primitive BkB_{k}, s^\hat{\mathbf{s}} is a fixed one-hot categorical distribution of the category of BkB_{k}. For a point x\mathbf{x} enclosed by multiple 3D bounding boxes of different semantic categories, we assign equal probability to these plausible categories and to the others. As explained in Section 4.2, the semantic field sβs_{\beta} is able to improve the geometry but cannot resolve the label ambiguity at the overlapping region.

Learned Semantic Field: We add a semantic head parameterized by ϕ\phi to NeRF to learn the semantic distribution s{\mathbf{s}}. We apply a softmax operation at each 3D point to ensure that s{\mathbf{s}} is a categorical distribution. The detailed network structure can be found in the supplementary material.

2 Loss Functions

Semantically-Guided Geometry Optimization: In the driving scenario considered in our setting, the RGB images are sparse and the depth range is infinite. We observe that the vanilla NeRF fails to recover reliable geometry in this setting. However, we find that leveraging noisy 2D semantic predictions as pseudo ground truth is able to boost density prediction when applied to the fixed semantic fields sβs_{\beta}

where S^k(r)\hat{\mathbf{S}}_{k}(\mathbf{r}) denotes the probability of the camera ray r\mathbf{r} belonging to the class kk, and Sk∗(r)\mathbf{S}^{*}_{k}(\mathbf{r}) denotes the corresponding pseudo-2D ground truth. As illustrated in Fig. 3, the key to improve density is to directly apply the semantic loss to the fixed semantic field sβs_{\beta}, where LS^2D\mathcal{L}^{\text{2D}}_{\hat{\mathbf{S}}} can only be minimized by updating the density σ\sigma. Fig. 3 shows that a correct S∗\mathbf{S}^{*} increases the volume density of 3D points inside the correct bounding primitive and suppresses the density of others. When S∗\mathbf{S}^{*} is wrong, the negative impact can be mitigated: 1) If S∗\mathbf{S}^{*} does not match any bounding primitive along the ray, it has no impact on the radiance field fθf_{\theta}. 2) If S∗\mathbf{S}^{*} exists in one of the bounding primitives along the ray, it means S∗\mathbf{S}^{*} corresponds to an occluding/occluded bounding primitive with wrong depth. To compensate, we introduce a weak depth supervision Ld\mathcal{L}_{d} based on stereo matching to alleviate the misguidance of LS^2D\mathcal{L}^{\text{2D}}_{\hat{\mathbf{S}}}. Although Ld\mathcal{L}_{d} improves the overall geometry as shown in our ablation study, it fails to produce accurate object boundaries when used alone (see supplementary). Adding our semantically-guided geometry optimization yields more accurate density estimation as pre-trained segmentation models usually perform well on frequently occurring classes, e.g., cars and roads.

Joint Geometry and Semantic Optimization: While enabling improved geometry, the 3D label of the overlapping regions remains ambiguous in the fixed semantic field. We leverage sϕs_{\phi} to address this problem by jointly learning the semantic and the radiance fields. We apply a cross-entropy loss LS2D\mathcal{L}^{\text{2D}}_{\mathbf{S}} to each camera ray based on the filtered 2D pseudo ground truth, where u(r)\mathbf{u}(\mathbf{r}) is set to 1 if S∗(r)\mathbf{S}^{*}(\mathbf{r}) matches the semantic class of any bounding primitive along the ray and otherwise . To further suppress noise in the 2D predictions, we add a per-point semantic loss Ls3D\mathcal{L}^{\text{3D}}_{\mathbf{s}} based on the 3D bounding primitives

where ui\mathbf{u}_{i} is a per-point binary mask. ui\mathbf{u}_{i} is set to 11 if (1) xi\mathbf{x}_{i} has a unique 3D semantic label and (2) the density σ\sigma is above a threshold σth\sigma_{th} to focus on the object surface. As illustrated in Fig. 3, LS2D(θ,ϕ)\mathcal{L}^{\text{2D}}_{\mathbf{S}}(\theta,\phi) does not necessarily improve the underlying geometry as the network can simply adjust the semantic head sϕs_{\phi} to satisfy the loss. This behavior is also observed in novel view synthesis where NeRF does not necessarily recover good geometry when optimized for image reconstruction alone, specifically given sparse input views .

Total Loss: Together, the total loss takes the form as

where Lp=∑r∈R∥C∗(r)−C(r)∥22\mathcal{L}_{p}=\sum_{\mathbf{r}\in\mathcal{R}}\left\|\mathbf{C}^{*}(\mathbf{r})-\mathbf{C}(\mathbf{r})\right\|_{2}^{2} and Ld=∑r∈R∥D∗(r)−D(r)∥22\mathcal{L}_{d}=\sum_{\mathbf{r}\in\mathcal{R}}\left\|\mathbf{D}^{*}(\mathbf{r})-\mathbf{D}(\mathbf{r})\right\|_{2}^{2} denote the photometric loss and the depth loss, respectively. λS^\lambda_{\hat{\mathbf{S}}}, λS\lambda_{{\mathbf{S}}}, λs\lambda_{{\mathbf{s}}}, λC\lambda_{\mathbf{C}}, and λd\lambda_{d} are constant weighting parameters. C∗(r)\mathbf{C}^{*}(\mathbf{r}) and C(r)\mathbf{C}(\mathbf{r}) are the ground truth and rendered RGB colors for ray r\mathbf{r}. D∗(r)\mathbf{D}^{*}(\mathbf{r}) and D(r)\mathbf{D}(\mathbf{r}) are pseudo ground truth depth generated by stereo matching and rendered depth, respectively. Please refer to the supplementary for more details of D∗\mathbf{D}^{*}.

3 Rendering of Panoptic Labels

Based on our learned semantic field sϕs_{\phi} and the 3D bounding primitives β\beta with instance IDs, we can easily render a panoptic segmentation map. Specifically, for a camera ray r\mathbf{r}, the panoptic label directly takes the class with maximum probability in S(r){\mathbf{S}}(\mathbf{r}) if it is a “stuff” class. For “thing” classes, we render an instance distribution T(r)\mathbf{T}(\mathbf{r}) based on the bounding primitives β\beta to replace S{\mathbf{S}} with T\mathbf{T}. Our instance field is defined as follow

where Mt{M_{t}} is the number of the things in the scene and t\mathbf{t} denotes a categorical distribution indicating which thing it belongs to. Note that t\mathbf{t} is determined by the bounding primitives and is a one-hot vector if x\mathbf{x} is uniquely enclosed by a bounding primitive of a thing. As overlap often occurs at the intersection of stuff and thing region, the bounding primitives of things rarely overlap with each other. Thus, this deterministic instance field leads to reliable performance in practice. To ensure that the instance label of this ray is consistent with the semantic class defined by S{\mathbf{S}}, we mask out instances belonging to other semantic classes by setting their probabilities to in T\mathbf{T}.

4 Implementation Details

Sampling Strategy and Sky Modeling: With the 3D bounding primitives covering the full scene, we sample points inside the bounding primitives to skip empty space. For each ray, we optionally sample a set of points to model the sky after the furthest bounding primitive. More details regarding the sampling strategy can be found in the supplementary. Our sampling strategy allows the network to focus on the non-empty region. As evidenced by our experiments, this is particularly beneficial in unbounded outdoor environments.

Training: We optimize one Panoptic NeRF model per scene, using a single NVIDIA 3090. For each scene, we set the origin to the center of the scene. We use Adam with a learning rate of 5e-4 to train our models. We set loss weights to λS^=1,λS=1,λs=1,λC=1,λd=0.1\lambda_{\hat{\mathbf{S}}}=1,\lambda_{{\mathbf{S}}}=1,\lambda_{{\mathbf{s}}}=1,\lambda_{\mathbf{C}}=1,\lambda_{d}=0.1, and the density threshold to σth=0.1\sigma_{th}=0.1. We optimize the total loss L\mathcal{L} for 80,000 iterations.

Experiments

Dataset: We conduct experiments on the recently released KITTI-360 dataset. KITTI-360 is collected in suburban areas and provides 3D bounding primitives covering the full scene. Following , we evaluate Panoptic NeRF on manually annotated frames from 5 static suburbs. We split these 5 suburbs into 10 scenes, comprising 128 consecutive frames each with an average travel distance of 0.8m between frames. We leverage all 128 pairs of posed stereo images for training. KITTI-360 provides a set of manually labeled frames sampled in equidistant steps of 5 frames. We use half of the manually labeled frames for evaluation and provide the other half as input to 2D-to-2D label transfer baselines. We improve the quality of the manually labeled ground truth which is inaccurate at ambiguous regions, see supplementary for details.

Baselines: We compare against top-performing baselines in two categories: (1) 2D-to-2D label transfer baselines, including Fully Connected CRF (FC CRF) and Semantic NeRF . For both baselines, we provide manually annotated 2D frames as input, sparsely sampled at equidistant steps of 10 frames. Note that labeling these 2D frames takes similar or longer compared to annotating 3D bounding primitives . As these 2D annotations are extremely sparse, we further provide the same pseudo-2D labels used in our method to Semantic NeRF. (2) 3D-to-2D label transfer baselines, including PSPNet*, 3D Primitives/Meshs/Points+GC , and 3D-2D CRF . All these baselines leverage the same 3D bounding primitives to transfer labels to 2D. Here, PSPNet* is considered 3D-to-2D as it is pre-trained on Cityscapes and fine-tuned on KITTI-360 based on the 3D sparse label projections. The second set of baselines first project 3D primitives/meshes/points to 2D and then apply Graph Cut to densify the label. The 3D-2D CRF densely connects 2D image pixels and 3D LiDAR points, performing inference jointly on these two fields with a set of consistency constraints.

Pseudo 2D GT: We use PSPNet* to provide pseudo ground truth in our main experiment to supervise our dual semantic fields. This ensures fair comparison to the 3D-2D CRF, which takes the predictions of PSPNet* as unary terms. Note that PSPNet* is fine-tuned on KITTI-360. To further simplify the entire process, in the ablation study we investigate the performance of our method using pre-trained models on Cityscape without any fine-tuning, including PSPNet , Deeplab and Tao et al. .

Metrics: We evaluate semantic labels by the mean Intersection over Union (mIoU) and the average pixel accuracy (Acc) metrics. To quantitatively evaluate multi-view consistency (MC), we utilize LiDAR points to retrieve corresponding pixel pairs between two consecutive evaluation frames. The MC metric is then calculated as the ratio of pixels pairs with consistent semantic labels over all pairs. For evaluating panoptic segmentation, we report Panoptic Quality (PQ) , which can be decomposed into Segmentation Quality (SQ) and Recognition Quality (RQ). We additionally adopt PQ† as PQ over-penalizes errors of stuff classes. To verify that Panoptic NeRF is able to improve the underlying geometry, we further evaluate the rendered depth compared to sparse depth maps obtained from LiDAR using Root Mean Squared Error (RMSE) and the ratio of accurate predictions (δ1.25\delta_{1.25}) .

We evaluate our model and compare it to our baselines on KITTI-360. As most baselines are not designed for panoptic label transfer, we first compare the semantic predictions of all methods and then compare our panoptic predictions to the 3D-2D CRF.

Semantic Label Transfer: As evidenced by Table 1, our method achieves the highest mIoU and Acc over a line of baselines. Specifically, compared to 3D-2D CRF, we obtain an absolute improvement of 1.6% (79.5%→81.1%)1.6\%~{}(79.5\%\rightarrow 81.1\%) on mIoU. Despite PSPNet* being finetuned on KITTI-360 which reduces the performance gap, our method outperforms PSPNet* by a significant margin. Supervised by the extremely sparse manually annotated GT, Semantic NeRF struggles to produce reliable performance. Using pseudo labels of PSPNet*, Semantic NeRF is capable of denoising and thus improving performance (67.2%→68.6%)(67.2\%\rightarrow 68.6\%). However, both variants of Semantic NeRF are inferior in the urban scenario when the input views are sparse. Moreover, in terms of MC, our method slightly outperforms the 3D-2D CRF and significantly surpasses 2D-to-2D label transfer methods. While our method slightly lags behind 3D Point + GC in terms of MC, it is reasonable as the label consistency is evaluated on the sparsely projected 3D points which GC takes as input to generate a dense label map.

Panoptic Label Transfer: We split the instance labels into things and stuff classes in KITTI-360. As the class “building” is classified as thing in KITTI-360 but stuff in Cityscapes, our 2D baselines are not suitable for testing performance on KITTI-360. Therefore, we ignore pre-trained 2D SOTA baselines. As shown in Table 2, our proposed method outperforms the 3D-2D CRF in both things and stuff classes. A visual comparison is shown in Fig. 5. As can be seen, we can deal well with overexposure at buildings, which is a challenge for the 3D-2D CRF, as it projects LiDAR points to reconstruct the intermediate meshes whose quality suffers on building class.

Novel View Label Synthesis: Panoptic NeRF can render RGB images and panoptic labels at novel viewpoints, whereas 3D-2D CRF is not capable of doing so. Please refer to the supplementary material for details.

2 Ablation Study

We validate our pipeline’s design modules with an extensive ablation in Table 3 by removing one component at a time. As there is a positive correlation between semantic labels and panoptic labels, we focus on semantic segmentation for this experiment on one scene.

Geometric Reconstruction: We now verify that our method effectively improves the underlying geometry leveraging semantic information. We first remove all the other losses except for Lp\mathcal{L}_{p}, leading to a baseline similar to NeRF but uses our proposed sampling strategy (NeRF*). In this case we render a semantic map based on the fixed semantic field sβs_{\beta}. As can be seen from Table 3 and Fig. 6, the underlying geometry of NeRF* drops significantly with only Lp\mathcal{L}_{p}. More importantly, the depth prediction also degrades considerably when removing LS^2D\mathcal{L}^{\text{2D}}_{\hat{\mathbf{S}}} (w/o LS^2D\mathcal{L}^{\text{2D}}_{\hat{\mathbf{S}}}), indicating the importance of the fixed semantic field in improving the underlying geometry. Fig. 6 shows that the full model has sharper edges while removing the fixed semantic fields leads to over-smooth object boundaries. We further show that eliminating Ld\mathcal{L}_{d} (w/o Ld\mathcal{L}_{d}) or replacing the sampling strategy with standard uniform sampling (Uniform S.) both impair the geometric reconstruction, and consequently the semantic estimation as well.

Semantic Segmentation: When removing LS2D\mathcal{L}^{\text{2D}}_{\mathbf{S}} (w/o LS2D\mathcal{L}^{\text{2D}}_{\mathbf{S}}), the performance also drops as the learned semantic field sϕs_{\phi} is only supervised by weak 3D supervision. Interestingly, this baseline still outperforms the semantic map rendered by the fixed semantic field despite that they share the same geometry. This observation suggests that the weak 3D supervision provided by Ls3D\mathcal{L}^{\text{3D}}_{\mathbf{s}} also allows to address the label ambiguity of the overlapping region to a certain extent. Therefore, it is not surprising that removing Ls3D\mathcal{L}^{\text{3D}}_{\mathbf{s}} (w/o Ls3D\mathcal{L}^{\text{3D}}_{\mathbf{s}}) worsens the semantic prediction compared to the full model. We observe that the performance is also deteriorated without ray masking (w/o u(r)\mathbf{u}(\mathbf{r})).

2D Pseudo GT: Finally, we evaluate how the quality of the pseudo 2D ground truth affects our method in Table 4. As some classes are not considered during training in Cityscapes, we additionally report mIoUsub{}_{\text{sub}} over the remaining classes. It is worth noting that using models pre-trained on Cityscapes without any fine-tuning leads to promising results, where Ours w/ Deeplab and Ours w/ Tao et al. are very close to Ours + PSPNet* in terms of mIoUsub{}_{\text{sub}}. More importantly, our method consistently outperforms the corresponding pseudo GT leveraging the 3D bounding primitives.

3 Limitations

Our method performs per-scene optimization and training takes 4 hours on one scene. Training time needs to be reduced to scale our method to large-scale scenes, e.g., by adopting improvements in speeding up NeRF training . In addition, we consider label transfer on static scenes in this work. We plan to extend our method to dynamic scenes leveraging recent advances in dynamic radiance field estimation .

Conclusion

We present Panoptic NeRF that infers in 3D space and renders per-pixel semantic and instance labels for 3D-to-2D label transfer. By combining coarse 3D bounding primitives and noisy 2D predictions using our dual semantic fields, Panoptic NeRF is capable of improving the underlying geometry given sparse input views and resolving label noise. Moreover, it enables label synthesis at novel view points. We believe that our method is a step towards more efficient data annotation, while simultaneously providing a 3D consistent continuous panoptic representation of the scene.

References

Appendix A Implementation Details

Fig. 7 shows the trainable part of our Panoptic NeRF model. We adopt the same network architecture in all experiments. The network takes as input the 3D location x\mathbf{x} (each element normalized to $)andtheviewingdirection) and the viewing direction\mathbf{d}.FollowingNeRF,both. Following NeRF , both\mathbf{x}andand\mathbf{d}$ are mapped to a higher dimensional space using a positional encoding (PE):

To learn high frequency components in unbounded outdoor environments, we set L=15L=15 for γ(x)\gamma(\mathbf{x}) and L=4L=4 for γ(d)\gamma(\mathbf{d}). Our learned semantic field is conditioned only on the 3D location x\mathbf{x} rather than the viewing direction d\mathbf{d} in order to predict view-independent semantic distributions.

A.2 Sampling Strategy

We sample points within the bounding primitives to skip empty space. As our bounding primitives are convexThe cuboids and ellipsoid are both convex. The extruded 3D plane is convex in a local region., each ray intersects a bounding primitive exactly twice which determines the sampling interval. For each camera ray, we sort all bounding primitives that the ray hits from near to far and save the intersections offline. To save storage and to speed up training, we keep the first 10 sorted bounding primitives as the rest are highly likely to be occluded. If a camera ray intersects less than 10 bounding primitives, we additionally sample a set of points to model the sky in [tmax,tmax+tint][t_{max},t_{max}+t_{int}], where tmaxt_{max} denotes the distance from the origin to the furthest bounding primitive in the scene and tintt_{int} is a constant distance interval.

A.3 Evaluation Metric

We evaluate mIoU and pixel accuracy following standard practice . Here, we provide more details of the multi-view consistency and panoptic quality metrics.

Multi-view Consistency: To evaluate multi-view consistency, we use depth maps obtained from LiDAR points to retrieve matching pixels across two consecutive frames. A similar multi-view consistency metric is considered in where optical flow is used to find corresponding pixel pairs. We instead use LiDAR depth maps as they are more accurate compared to optical flow estimations. The details of generating the LiDAR depth maps will be introduced in Section B.2. Given LiDAR depth maps at two consecutive test frames, we first unproject them into 3D space and find matching points. Two LiDAR points are considered matched if their distance in 3D is smaller than 0.1 meters. For each pair of matched points, we retrieve the corresponding 2D semantic labels and evaluate their consistency. The MC metric is evaluated as the number of consistent pairs over all matched pairs. Despite being not 100%\% accurate as the 3D points may not match exactly in 3D space, we find this metric meaningful in reflecting multi-view consistency.

Panoptic Quality: Following , we use the PQ metric to evaluate the performance of panoptic segmentation. As the ground truth panoptic labels are not precise in distant areas and have a lot of small noises of things, we set ground truth labels of areas less than 100 pixels to “void”. Correspondingly, segment matching will not be performed in void regions. In addition, Panoptic maps of the 3D-2D CRF and our method are obtained by 3D primitives, thus containing very far objects. In fact, these far objects may only occupy very small areas, usually less than 100 pixels, on 2d images. To avoid being biased by those extremely far objects in the segment matching, we omit them by setting the predicted labels of the areas less than 100 pixels to the “sky” class. To ensure a fair comparison across all methods, we adopt the same evaluation protocol for all baselines and our method.

A.4 Training and Inference

As mentioned in Section 4.2 of the main paper, our total loss function comprises five terms, including three semantic losses LS^2D\mathcal{L}^{2D}_{\hat{\mathbf{S}}}, LS2D\mathcal{L}^{2D}_{\mathbf{S}}, Ls3D\mathcal{L}^{3D}_{\mathbf{s}}, the photometric loss Lc\mathcal{L}_{\mathbf{c}} and the depth loss Ld\mathcal{L}_{d}. During per-scene optimization, the photometric loss Lc\mathcal{L}_{\mathbf{c}} is defined on the posed stereo images. The 2D semantic losses LS^2D\mathcal{L}^{2D}_{\hat{\mathbf{S}}}, LS2D\mathcal{L}^{2D}_{\mathbf{S}} are applied to the left images only. While our method allows for using noisy 2D semantic predictions on the right images, this ensures fair comparison to the 3D-2D CRF which is not capable of using predictions on other viewpoints for inference. We apply the 3D semantic loss Ls3D\mathcal{L}^{3D}_{\mathbf{s}} directly on 3D points sampled along the camera rays of the left images. The depth loss Ld\mathcal{L}_{d} is also defined on the left images as the information gain is marginal on the right views.

For inference, we compare our method to the baselines on the left views of which the manually labeled 2D Ground Truth is defined. Note that our method is not constrained to the left views during inference. We show label transfer results on the right camera views in Section C.3 and novel view label synthesis in Section C.4.

Appendix B Data Preparation

To provide weak depth supervision to Panoptic NeRF, we use Semi-Global Matching (SGM) to estimate depth given a stereo image pair. We perform a left-right consistency check and a multi-frame consistency check in a window of 55 consecutive frames to filter inconsistent predictions. We further omit depth predictions further than 15 meters for each frame as disparity is better estimated in nearby regions, see Fig. 8.

B.2 LiDAR Depth for Evaluation

We evaluate the rendered depth maps against the LiDAR measurements. We refrain from using LiDAR as input as 1) this allows us to evaluate our depth prediction against LiDAR and 2) it makes our method more flexible to work with settings without any LiDAR observations. As LiDAR observations at each frame are sparse, we accumulate multiple frames of LiDAR observations and project the visible points to each frame similar to .

B.3 Manually Annotated 2D GT

The manually annotated 2D ground truth of KITTI-360 is inferior at some regions. For a fair comparison, we improve the label quality by manually relabeling ambiguous classes, see Fig. 9 for illustrations.

Appendix C Additional Experimental Results

We show that using the depth loss Ld\mathcal{L}_{d} alone is not able to recover accurate object boundaries in Fig. 10. In contrast, adding the semantic loss LS^2D\mathcal{L}^{2D}_{\hat{\mathbf{S}}} to the fixed semantic field further improves the object boundary. These improvements can be explained as follows: Firstly, the weak stereo depth supervision is not fully accurate, especially at far regions. Furthermore, even with perfect depth supervision, the model receives very small penalty if the predicted depth is close to the GT depth. In contrast, the cross entropy loss LS^2D\mathcal{L}^{2D}_{\hat{\mathbf{S}}} defined on the fixed semantic field provides a strong penalty as small errors in depth lead to wrong semantics.

C.2 Qualitative Comparison of Label Transfer

Fig. 11 shows additional qualitative comparisons corresponding to the Table 1 of the main paper. Consistent with the quantitative results, our method outperforms all baselines qualitatively. We further show qualitative comparisons to 3D-2D CRF on a set of unlabeled 2D frames, including semantic label transfer in Fig. 12 and panoptic label transfer in Fig. 13.

C.3 Stereo Label Transfer

In Fig. 14, we illustrate our stereo label transfer results. Despite that we only utilize pseudo ground truth on the left views for supervision, our model achieves consistent results on both left and right views.

C.4 Novel View Label Synthesis

Here, we evaluate our performance on novel view label synthesis by applying the photometric loss Lc\mathcal{L}_{\mathbf{c}} to the left images only. This allows us to evaluate novel view appearance and label synthesis on the right view images. As shown in Fig. 15, our method achieves promising results on novel view appearance and label synthesis. More results of appearance and label synthesis on unseen viewpoints can be found in the supplementary video.

C.5 Analysis of 3D-2D CRF

The 3D-2D CRF performs inference based on a multi-field CRF which reasons jointly about the labels of the 3D points and all pixels in the image. To obtain dense 3D points, it accumulates LiDAR observations over multiple frames and project visible 3D points to the image based on a reconstructed mesh. Fig. 16 shows depth maps of the reconstructed mesh corresponding to Fig. 5 of the main paper. As can be seen, the side of the building can hardly be scanned by the LiDAR, leading to incomplete mesh reconstruction. Consequently, 3D-2D CRF lacks 3D information in these regions and needs to distinguish building instances mainly based on 2D image cues. It is not surprising that the 3D-2D CRF fails at overexposed image regions in this case.

C.6 Failure Cases

Our method leverages a deterministic instance field defined by the 3D bounding primitives to render instance labels. Thus, our method struggles to recover accurate instance boundaries where two instance bounding primitives overlap in 3D space. This sometimes occurs on the building class where two buildings are spatially connected to each other as shown in Fig. 17.