PolyMaX: General Dense Prediction with Mask Transformer

Xuan Yang, Liangzhe Yuan, Kimberly Wilber, Astuti Sharma, Xiuye Gu, Siyuan Qiao, Stephanie Debats, Huisheng Wang, Hartwig Adam, Mikhail Sirotenko, Liang-Chieh Chen

Introduction

Entering the deep learning era , enormous efforts have been made to tackle dense prediction problems, including but not limited to, image segmentation , depth estimation , surface normal prediction , and others . Early attempts formulate these problems as per-pixel prediction (i.e., assigning a predicted value to every pixel) via fully convolutional networks . Specifically, when the desired prediction of each pixel is discrete, such as image segmentation, the task is constructed as per-pixel classification, while other tasks whose target outputs are continuous, such as depth and surface normal, are instead cast as per-pixel regression problems.

Recently, a new paradigm for segmentation tasks is drawing attention because of its superior performance compared to previous per-pixel classification approaches. Inspired by the object detection network DETR , MaX-DeepLab and MaskFormer propose to classify each segmentation mask as a whole instead of pixel-wise, by extending the concept of object queries in DETR to represent the clusters of pixels in segmentation masks. Specifically, with the help of pixel clustering via conditional convolutions , these Mask-Transformer-based works employ transformer decoders to convert object queries to a set of (mask embedding vector, class prediction) pairs, which finally yield a set of binary masks by multiplying the mask embedding vectors with the pixel features. Such methods are effective when the target domain is discretely quantized (e.g., one semantic label is encoded by one integer scalar, as in image segmentation). However, it is unclear how this framework can be generalized to other dense prediction tasks, whose outputs are continuous or even multi-dimensional, such as depth estimation and surface normal prediction.

On the contrary, in the field of depth estimation, DORN demonstrates the potential of discretizing the continuous depth range into a set of fixed bins, and AdaBins further adaptively estimates the bin centers, dependent on the input image. The continuous depth values are then estimated by linearly combining the bin centers. Promising results are achieved by jointly learning the bin centers and performing per-pixel classification on those bins. This insight – of using classification to perform dense prediction in a continuous domain – opens up the possibility of performing many other continuous dense prediction tasks within the Mask-Transformer-based framework, a powerful tool for discrete value predictions.

Consequently, a few natural questions emerge: Can we extend these Mask-Transformer-based frameworks even further, to solve more continuous dense prediction tasks? Can the resulting framework generalize to other multi-dimensional continuous domain, e.g., surface normal estimation? In this work, we provide affirmative answers to those questions by proposing a new general architecture for dense prediction tasks. Specifically, building on top of the insight from , we generalize the mask transformer framework to multiple dense prediction tasks by using the cluster centers (i.e., object queries) as the intermediate representation. We evaluate the resulting model, called PolyMaX, on the challenging NYUD-v2 and Taskonomy datasets. Remarkably, our simple yet effective approach demonstrates new state-of-the-art performance on semantic segmentation, depth estimation, and surface normal prediction, without using any extra modality as inputs (e.g., multi-modal RGB-D inputs as in CMNeXt ), heavily pretrained backbones (e.g., Stable-Diffusion as in VPD ) or complex pretraining schemes (e.g., a mix of 12 datasets as in ZoeDepth ).

Our contributions are summarized as follows:

We propose PolyMaX, which generalizes dense prediction tasks with a unified Mask-Transformer-based framework. We take surface normal as a concrete example to demonstrate how general dense prediction tasks can be solved by the proposed framework.

We evaluate PolyMaX on NYUD-v2 and Taskonomy datasets, and it sets new state-of-the-arts on multiple benchmarks on NYUD-v2, achieving 58.08 mIoU, 0.250 root-mean-square (RMS) error and 13.09 mean error on semantic segmentation, depth estimation and surface normal prediction, respectively.

We further perform the scalability study, which demonstrates that PolyMaX scales significantly better than the conventional per-pixel regression based methods as the pretraining data increases.

Lastly, we provide the high-quality pseudo-labels of semantic segmentation for Taskonomy dataset, aiming to compensate for the scarcity of existing large-scale multi-task datasets and facilitate the future research.

Related Work

Dense Prediction in the Discrete Domain Image segmentation partitions images into multiple segments by their semantic classes. For example, fully convolutional networks train semantic segmentation in an end-to-end manner mapping pixels into their classes . Atrous convolution increases the network receptive field for semantic segmentation without additional learnable parameters . Many state-of-the-art methods use an encoder-decoder meta architecture for stronger global and local information integration .

Dense Prediction in the Continuous Domain Unlike image segmentation that produces discrete predictions, depth and surface normal estimation expect the models to predict continuous values . Most early attempts solve it as a standard regression problem. To alleviate training instability, DORN proposes treating the problem as classification by discretizing the continuous depth into a set of pre-defined intervals (bins). AdaBins adaptively learns the depth bins, conditioned on the input samples. LocalBins further learns the depth range partitions from local regions instead of global distribution of depth ranges. BinsFormer views adaptive bins generation as a direct set prediction problem . The current state-of-art on depth estimation is set by VPD , which employs the Stable-Diffusion pretrained with the large-scale LAION-5B dataset. Among the non-diffusion-based models, ZoeDepth shows the strongest capability by pretraining on 12 datasets with relative depth, and then finetuning on two datasets with metric depth.

Surface Normal Prediction Despite having a continuous output domain like depth estimation, the surface normal problem remains under-explored for discretization approaches. Most prior works on surface normal estimation focus on improving the loss , reducing the distribution bias and shift , and resolving the artifacts in ground-truth by leveraging other modalities . Until recently, iDisc , a method based on discretizing internal representations, shows promise for depth estimation and surface normal prediction. However, it differs from depth estimation discretization methods by enforcing discretization in the feature space rather than the target output space. We are more interested in the latter approach, as it is more generalizable to other dense prediction tasks.

Other Efforts to Unify Dense Prediction Tasks Realizing segmentation, depth and surface normal are all pixel-wise mapping problem, previous works have deployed them to the same framework. Some works improve performance by exploring the relations among dense prediction tasks. UViM and Painter unify segmentation and depth estimation, but neither of them is based on discretizing the continuous output space, and neither considers surface normal estimation. By contrast, we focus on a complementary perspective: a unified architecture for dense prediction that models both discrete and continuous tasks, covering image segmentation, depth estimation, and surface normal prediction.

Mask Transformer Witnessing the success of transformers in NLP, the vision community has begun exploring them for more computer vision tasks . Inspired by DETR and conditional convolutions , MaX-DeepLab and MaskFormer switch the segmentation paradigm from per-pixel classification to mask classification, by introducing the mask transformer, where each input query learns to correspond to a mask prediction together with a class prediction. Several works also introduce transformer into semantic segmentation. Further improvements are made on attention mechanism to enhance the performance of mask transformers . CMT-DeepLab , kMaX-DeepLab , and ClustSeg reformulate the cross-attention learning in transformer as a clustering process. Our proposed PolyMaX builds on top of the cluster-based mask transformer architecture . Rather than focusing on segmentation tasks, PolyMaX unifies dense prediction tasks (i.e., image segmentation, depth estimation, and surface normal estimation) by extending the cluster-based mask transformer to support both discrete and continuous outputs.

Method

In this section, we first describe how segmentation and depth estimation can be transformed from per-pixel prediction problems to cluster-prediction problems (Sec. 3.1). We then introduce our proposed method, PolyMaX, a new mask transformer framework for general dense predictions (Sec. 3.2). We take surface normal prediction as a concrete example to explain how general dense prediction problems can be reformulated in a similar fashion, allowing us to unify them into the same mask transformer framework.

Cluster-Prediction Paradigm for Segmentation Two earlier works, MaX-DeepLab and MaskFormer , demonstrate how to shift image segmentation from per-pixel classification to the cluster-prediction paradigm. Along the same direction, we follow the recently proposed clustering perspective , which casts object queries to cluster centers. This paradigm is realized by the following two steps:

pixel-clustering: group pixels into KK clusters, represented by segmentation masks {mi∣mi∈H×W}i=1K\{\mathbf{m_{i}}|\mathbf{m_{i}}\in^{H\times W}\}^{K}_{i=1}, where HH and WW are height and width. Note that mi\mathbf{m_{i}} denotes soft segmentation masks.

cluster-classification: assign semantic label to each cluster with a probability distribution over CC classes. The iith cluster’s probability distribution is a 1D vector, denoted as pi\mathbf{p_{i}}, where {pi∣pi∈C,∑c=1Cpi,c=1}i=1K\{\mathbf{p_{i}}|\mathbf{p_{i}}\in^{C},\sum_{c=1}^{C}\mathbf{p}_{i,c}=1\}^{K}_{i=1}.

These clusters and probability distributions are jointly learned to predict the output SS, a set of KK cluster-probability pairs:

This cluster-prediction paradigm is general for semantic , instance , and panoptic segmentation. When only handling semantic segmentation, the framework can be further simplified by setting K=CK=C (i.e., number of clusters is equal to number of classes) which has a fixed matching between cluster centers (i.e., object queries) and semantic classes. We adopt this simplification, since the datasets we experimented with only support semantic segmentation.

Cluster-Prediction Paradigm for Depth Estimation At first glance, depth estimation seems incompatible with this cluster-prediction paradigm because of its continuous nature. However, recent works propose promising solutions by dividing the continuous depth range into KK learnable bins and regarding the task as a classification problem. By this means, the depth estimation task fits neatly into the above cluster-classification paradigm. Specifically, the pixel-clustering step outputs the range attention map {ri∣ri∈H×W}i=1K\{\mathbf{r_{i}}|\mathbf{r_{i}}\in^{H\times W}\}_{i=1}^{K}, which (after softmax) represents the predicted probability distribution over the KK bins for each pixel. While the cluster-prediction step estimates bin centers {bi}i=1K\{\mathbf{b_{i}}\}_{i=1}^{K}, adaptively discretizing the continuous depth range for each image. Here, KK controls the granularity of the depth range partition. Therefore, the output of depth estimation can be expressed as {(ri,bi)}i=1K\{(\mathbf{r_{i}},\mathbf{b_{i}})\}^{K}_{i=1}, sharing the same formulation as segmentation (Eq 1). The final depth prediction is then generated by a linear combination between the depth values of bin centers and the range attention map.

Mask Transformer Framework After describing the cluster-prediction paradigm for both segmentation and depth estimation, we now explain how to integrate them to the mask transformer framework. In the framework, the KK cluster centers are learned from the KK input queries through the transformer-decoder branch (blue block in Fig. 2). It associates the clusters, represented as query embedding vectors, with the pixel features extracted from the pixel encoder-decoder path (pink block in Fig. 2) through self-attention and cross-attention blocks. The aggregated information is gradually refined as the pixel features go from lower-resolution to higher-resolution. The probability distribution map is then generated from the following equation:

2 Mask-Transformer-Based General Dense Prediction

General dense prediction tasks can be expressed as follows:

Cluster-Prediction for Surface Normal From the geometric perspective, a surface normal lies on the surface of a unit 3D ball. Therefore, the problem essentially becomes learning the pixel mapping T\mathcal{T} from the 2D image space to the 3D point on the unit ball surface. The geometric meaning of the surface normal prediction allows us to naturally adopt the clustering-prediction approach.

Similar to , we discretize the surface normal output space into KK pieces of partitioned 3D spheres, converting the problem to classification among the KK pieces. We refer the partitioned sphere as sphere segments for simplicity Note that here the sphere segment does not rigorously match the mathematical definition of spherical segment, which refers to the solid produced by cutting a sphere with a pair of parallel planes., which can be viewed as clusters of 3D points on the unit ball surface, and thus are likely to have irregular shapes. Fortunately, with such a simplification, we can adopt the same paradigm as segmentation and depth estimation to tackle surface normals, where pi\mathbf{p_{i}} can be regarded as the probability over the KK sphere segments. We jointly learn the center coordinates of sphere segments with the 3D-point-cluster and probability distribution. The final surface normal is predicted as a linear combination of the center coordinates of KK sphere segments and the range attention map.

Model Instantiation We build the proposed general framework on top of kMaX-DeepLab with the official code-base . We call the resulting model PolyMaX, a Polymath with Mask Xformer for general dense prediction tasks. The model architecture is illustrated in Fig. 2.

Experimental Results

In this section, we first provide the details of our experimental setup. We then report our main results on NYUD-v2 and Taskonomy . We also present visualizations to obtain deeper insights of the proposed cluster-prediction paradigm, followed by ablation studies.

Dataset We focus on semantic segmentation, monocular depth estimation and surface normal tasks during experiments, as they comprehensively represent the diverse target space of dense prediction (one-dimensional to multi-dimensional, and discrete to continuous). Specifically, we use NYUD-v2 dataset. The official release provides 795 training and 654 testing images with real (instead of pseudo labels) ground-truths of those three tasks.

To complement the small scale of NYUD-v2 dataset, we also conduct experiments on the Taskonomy dataset, which is composed of 4.6 million images (train: 3.4M, val: 538K: test: 629K images) from 537 different buildings with indoor scenes. This dataset is originally developed to facilitate the study of task transfer learning, thereby containing ground-truths of various tasks, including semantic segmentation, depth estimation, surface normal, and so on. The depth Z-buffer and surface normal ground-truth are programmatically computed from image registration and mesh alignment, resulting in high quality annotations. However, the provided semantic segmentation annotations are only pseudo-labels generated by Li et al. , an out-dated model trained on COCO . After carefully examining these pseudo-labels, we notice that their quality does not satisfy the requirement for evaluating or improving state-of-the-art models. Therefore, we adopt the kMaX-DeepLab model with ConvNeXt-L backbone pretrained on COCO dataset to regenerate the pseudo-labels for semantic segmentation. We visually compare both labels in Fig. 5. The Taskonomy dataset enhanced by the new pseudo labels can serve as a meaningful complementary to this research field, particularly given the scarcity of the available large-scale high-quality multi-task dense prediction datasets. We will release the high-quality pseudo labels to facilitate future researchWill release at https://github.com/google-research/deeplab2

Evaluation Metrics Semantic segmentation is evaluated with mean Intersection-over-Union (mIoU) . For depth estimation, the metrics are root mean square error (RMS), Absolute mean relative error (A.Rel), absolute error in log-scale (Log10) and pixel inlier ratio (δi\delta_{i}) with error threshold at 1.25i1.25^{i} . For surface normal metrics, following we use mean (Mean) and median (Median) absolute error, RMS angular error, and pixel inlier ratio (δ1\delta_{1}, δ2\delta_{2}, δ3\delta_{3}) with thresholds at 11.5∘, 22.5∘ and 30∘, respectively.

Implementation Details We adopt the pixel encoder/decoder and the transformer decoder modules from the kMaX-DeepLab . Additional L2 normalization is needed at the end of surface normal head to yield 3D unit vectors. During training, we adopt the same losses from kMaX-DeepLab for semantic segmentation. For depth estimation, when training on NYUD-v2 dataset, we use scale invariant logarithmic error , relative squared error , following ViP-DeepLab . We also include the multi-scale gradient loss proposed by MegaDepth to improve the visual sharpness of the predicted depth map. While training depth estimation on Taskonomy, we switch to robust Charbonnier loss , since this dataset has maximum depth value 128m with more outliers. The loss function we apply to surface normal is simply L2 loss, since it has better training stability than truncated angular loss . The experiments on NYUD-v2 are conducted by first pretraining PolyMaX on Taskonomy dataset, and then finetuning on NYUD-v2 train split. The finetuning step is unnecessary when evaluating on Taskonomy test split. The learning rate at the pretraining and finetuning stages are 5e-4 and 5e-5, respectively. To ensure fair comparisons with previous works, we closely follow their experiment setup. However, due to the use of different training data splits for each of the three tasks in previous studies, we have to train each task independently.

2 Main Results

NYUD-v2 In Tab. 2, we compare PolyMaX with state-of-the-art models for dense prediction tasks on the NYUD-v2 dataset. We group the existing models based on the number and type of the tasks they support on this dataset. We choose the best numbers reported in the prior works. As shown in the table, in such a competitive comparison, PolyMaX still significantly outperforms all the existing models.

Specifically, in the semantic segmentation task, most of the recent models use additional modalities such as depth, but PolyMaX surpasses the existing best two models CMX and CMNeXt by 1.2%1.2\% mIoU, without using any additional modalities. In the depth estimation task, the current best model VPD is built upon Stable-Diffusion and pretrained with LAION-5B dataset . Despite of only using a pretraining dataset (i.e., Taskonomy) with 0.1%0.1\% of that scale, PolyMaX achieves better performance on all the depth metrics than VPD. Furthermore, PolyMaX breaks the close competition among non-diffusion-based models by a meaningful improvement from above 0.27 to 0.25 RMS (all the recent non-diffusion-based models achieve around 0.27 RMS), similarly for all the other metrics. This is non-trivial, particularly given that the best non-diffusion-based depth model ZoeDepth uses the pretrained BeiT384-L backbone and 2-stage training with a mixture of 12 datasets. On the surface normal benchmark, PolyMaX continues to outperform all the existing models by a substantial margin, with the mean error being reduced from above 14.60 to 13.09. Finally, when comparing with models that support all three tasks on this dataset, the improvement of PolyMaX is further amplified for all metrics. Overall, PolyMaX is not only among the few models that support all the three tasks, but also sets a new state-of-the-art on all three dense prediction benchmarks.

Taskonomy In Tab. 3, we report our results on Taskonomy, along with a solid baselineWe notice recent works also report numbers on this benchmark, but with the unreleased code and different experimental setup, we can not consider them here as fair baselines. DeepLabv3+ . The performance of the same baseline model on NYUD-v2 benchmarks in Tab. 2, compared with state-of-the-art models, can justify our choice of it as baseline on Taskonomy benchmarks. On all three tasks, PolyMaX significantly outperforms the baseline model. Specifically, the performance on semantic segmentation, depth estimation and surface normal is improved by 8.3 mIoU, 0.13 RMS and 1.0 mean error, respectively. These results further validate the effectiveness of the proposed framework.

Visualizations and Limitations We visualize the predictions of PolyMaX on all the three dense preiction tasks with their corresponding ground-truth in Fig. 6. As shown in the top row, PolyMaX successfully resolves the fine-grained details and irregular object shapes, and predicts high quality results. The bottom row shows a challenging case where PolyMaX can be further improved. Similar to many other depth models, it is difficult to correctly infer the depth map with glass, mirror and other reflective surfaces, due to the artifacts in the available ground-truth. Lastly, another direction for future improvement is the visual sharpness of the predicted surface normal. We find that the multi-scale gradient loss proposed by MegaDepth can effectively improve the edge sharpness for depth estimation, and slightly improves depth metrics (e.g., around 0.002 RMS improvement). It is promising to adapt the loss to surface normal prediction to improve visual quality.

3 Ablation Studies

Scalability of Mask-Transformer-Based Framework We investigate the scalability of PolyMaX, since this property has become increasingly crucial in the large scale model era. Fig. 7 demonstrates that PolyMaX’s scalability with the percentage of pretraining dataset. For a fair comparison, we build the baseline models upon DeepLabv3+ , using pixel-classification prediction head for semantic segmentation, and pixel-regression prediction head for depth estimation and surface normal prediction. The training configurations are the same within each task, except that the DeepLabv3+ based depth model requires the robust Charbonnier loss to overcome the typical training instability issue of regression with outliers.

As shown in Fig. 7(a), PolyMaX scales significantly better than the baseline model on semantic segmentation. When all models are pretrained with only ImageNet (i.e., 0% Taskonomy data is introduced in pretraining), PolyMaX performs worse than the baseline, especially the ConvNeXt-L model variant fails disappointingly. This is because the larger capacity of PolyMaX can not be fully exploited with only 795 training samples that are available for semantic segmentation from NYUD-v2 dataset. Thus, after including only 25%25\% of the Taskonomy data in pretraining, PolyMaX with ResNet-50 and ConvNeXt-L immediately outperform the baseline model with more than 2%2\% mIoU and 9%9\% mIoU, respectively. This performance gaps remain when the percentage of pretraining data increases, and stay at more than 1%1\% and 7%7\% mIoU after using the entire Taskonomy dataset for pretraining.

The scalability of depth estimation is shown in Fig. 7(b). When no Taskonomy pretraining data is used, the improvement of PolyMaX over the baseline (both with ResNet-50) is only 0.01 RMS, implying negligible positive impact of cluster-prediction paradigm. The more than 0.1 RMS performance improvement of PolyMaX with ConvNeXt-L mainly comes from the more powerful backbone. However, PolyMaX has better scalability with the increase of pretraining data. When 25%25\% Taskonomy data is involved in pretraining, the model performance improvements of PolyMaX with ResNet-50 and ConvNeXt-L increase from 0.01 and 0.1 RMS to 0.05 and 0.12 RMS, respectively. With all Taskonomy dataset being used in pretraining, the performance gaps of the two PolyMaX come to 0.6 and 0.12 RMS.

The scalability of PolyMaX on surface normal, showing a similar trend as depth estimation, is illustrated in Fig. 7(c). Even though PolyMaX with ResNet-50 does not perform better than the baseline model when no Tasknomy pretraining data is used, it significantly surpasses the baseline model by 0.7 and 0.9 mean error at 25%25\% and 100%100\% of pretraining data. While PolyMaX with ConvNeXt-L substantially outperforms the baseline model by 1.2 and 1.8 mean error at 25%25\% and 100%100\% of pretraining data.

To obtain deeper understanding of the model behaviour, we further look into the visualization of the learned cluster centers and the probability distribution across those clusters for depth estimation and surface normal, shown in Fig. 8. Interestingly, 12 out of the 16 probability maps from the depth estimation model look similar, implying redundancy in the learned 16 cluster centers. Similar findings are observed for surface normal. This can also explain the small impact of cluster count KK on model performance in tables of supplementary.

Despite the redundancy, the rest unique probability maps clearly depict the cluster-prediction paradigm. For instance, the last 4 probability maps of the depth model correspond to the attention on close, medium, futher and furthest regions in the scene. Similarly, the unique probability maps of surface normal illustrate the attention on surfaces with different angles, including left, right, upwards, downwards.

Conclusion

We have generalized dense prediction tasks with the same mask transformer framework, realized by casting semantic segmentation, depth estimation and surface normal to cluster-prediction paradigm. The proposed PolyMaX has demonstrated state-of-the-art results on the three benchmarks of NYUD-v2 dataset. We hope the superior performance of PolyMaX can enable many impactful applications, including but not limited to scene understanding, image generation and editing, and augmented reality.

Acknowledgement We thank Yukun Zhu, Jun Xie and Shuyang Sun for their support on the code-base.

References

Supplementary Materials

In the supplementary materials, we provide additional information as listed below:

Sec. A provides detailed training protocol used in the experiments.

Sec. B provides additional ablations studies.

Sec. C provides more visualizations of (1) model predictions, (2) failure modes, (3) learned probability distribution maps, and (4) our generated high-quality pseudo-labels for Taskonomy semantic segmentation.

Appendix A Training Protocol

The training configurations of PolyMaX closely follow kMaX-DeepLab, including the regularization, drop path , color jitting , AdamW optimizer with weight decay 0.05, and learning rate multiplier 0.1 for backbone. Additionally, for depth estimation and surface normal, we follow the data preprocessing in , except that we disable random scaling and rotation for surface normal.

Appendix B Additional Ablation Studies

Impact of Cluster Granularity We analyze the impact of cluster granularity (i.e., KK cluster centers) for depth estimation and surface normal, which are presented in Tab. 4 and Tab. 5. Note that, we skip this analysis for semantic segmentation, as we can simply assign the number of clusters as the number of classes. In both Tab. 4 and Tab. 5, we observe that the cluster granularity does not have a significant impact on the model performance on either benchmarks. Among the different cluster settings, 16 clusters and 8 clusters perform the best for depth estimation and for surface normal, respectively.

Appendix C Additional Visualization

Model Predictions In Fig. 9, we show more model predictions of semantic segmentation, depth estimation, and surface normal prediction. As shown in the figure, our proposed PolyMaX can capture fine details on scenes with complex structures.

Failure Modes To better understand the limitations of the proposed model, we also look into the failure modes. As shown in Fig. 10, PolyMaX struggles to predict the depth and surface normal for transparent and reflective objects, which are the most challenging issues in the tasks of depth and surface normal estimation. The difficulties can also be reflected by the unreliable ground-truth annotations for those cases. In Fig. 11, our model sometimes predicts over-smoothed depth and surface normal results. The findings of (e.g., a better loss function) may alleviate this issue, which is left for future exploration.

Probability Distribution Maps We provide additional visualizations of the learned probability distribution maps for depth estimation and surface normal prediction in Fig. 12 and Fig. 13, respectively. As shown in the figures, the learned probability distribution maps effectively cluster pixels for different distances (for depth task) or angles (for surface normal task).

Taskonomy Pseudo-Labels In Fig. 14, we show additional visualization of the generated high-quality pseudo-labels for Taskonomy semantic segmentation.