GMMSeg: Gaussian Mixture based Generative Semantic Segmentation Models

Chen Liang, Wenguan Wang, Jiaxu Miao, Yi Yang

Introduction

Semantic segmentation aims to explain visual semantics at the pixel level. It is typically considered as a problem of pixel-wise classification, i.e., assigning a class label c ⁣∈ ⁣{1, ⁣⋯ ⁣,C}c\!\in\!\{1,_{\!}\cdots_{\!},C\} to each pixel data x{x}. Under this regime, deep-neural solutions are naturally built as a combination of two parts (Fig. ⁣{}_{\!} 1(a)): an encoder-decoder, dense feature extractor that maps xx to a high-dimensional feature representation x\bm{x}, and a dense classifier that conducts CC-way classification given input pixel feature x\bm{x}. Starting from the first end-to-end segmentation solution – fully convolutional networks (FCN) ⁣{}_{\!} , researchers leave the classifier as parametric softmax, and fully devote to improving the dense feature extractor for learning better representation. As a result, a huge amount of FCN-based solutions ⁣{}_{\!} emerged and their state-of-the-art was further pushed forward by recent Transformer ⁣{}_{\!} -style algorithms ⁣{}_{\!} .

From a probabilistic perspective, the softmax classifier, supervised by the cross-entropy loss together with the feature extractor, directly models the class probability given an input, i.e., posterior p(c∣x)p(c|\bm{x}). ⁣{}_{\!} This ⁣{}_{\!} is ⁣{}_{\!} known ⁣{}_{\!} as ⁣{}_{\!} a ⁣{}_{\!} discriminative ⁣{}_{\!} classifier, ⁣{}_{\!} as ⁣{}_{\!} the ⁣{}_{\!} conditional ⁣{}_{\!} probability ⁣{}_{\!} distribution ⁣{}_{\!} discriminates directly between the different values of c ⁣c_{\!} . As discriminative classifiers directly find the classifica- tion rule with the smallest error rate, they often give excellent performance in downstream tasks, and hence become the de facto paradigm in segmentation. Yet, due to the discriminative nature, softmax- based ⁣{}_{\!} segmentation ⁣{}_{\!} models ⁣{}_{\!} suffer ⁣{}_{\!} from ⁣{}_{\!} several ⁣{}_{\!} limitations: ⁣{}_{\!} First, ⁣{}_{\!} they ⁣{}_{\!} only ⁣{}_{\!} learn ⁣{}_{\!} the ⁣{}_{\!} decision ⁣{}_{\!} boundary between classes, without modeling the underlying data distribution ⁣{}_{\!} . Second, as only one weight vector is learned per class, they assume unimodality for each class ⁣{}_{\!} , bearing no within-class variation. Third, they learn a prediction space where the model accuracy deteriorates rapidly away from the decision boundaries ⁣{}_{\!} and thus yield poorly calibrated predictions ⁣{}_{\!} , struggling to recognize out-of-distribution data ⁣{}_{\!} . The first two limitations may hinder the expressive power of segmentation models, and the last one challenges the adoption of segmentation models in decision-critical tasks (e.g., autonomous driving) and motivates the development of anomaly segmentation methods (which, however, rely on pre-trained discriminative segmentation models).

As an alternative of discriminative classifiers, generative classifiers first find the joint probability p(x,c)p(\bm{x},c), and use p(x,c)p(\bm{x},c) to evaluate the class-conditional densities p(x∣c)p(\bm{x}|c). Then classification is con- ducted using Bayes rule. Numerous theoretical and empirical comparisons ⁣{}_{\!} between these two approaches have been initiated even before the deep learning revolution. They reach the agreement that generative classifiers have potential to overcome shortcomings of their discriminative counterparts, as they are able to model the input data itself. This stimulates the recent investigation of generative (and discriminative-generative hybrid ⁣{}_{\!} ) classifiers in trustworthy AI ⁣{}_{\!} and semi-supervised learning ⁣{}_{\!} , while the discriminative classifiers are still dominant in most downstream tasks.

In light of this background, we propose a GMM based segmentation framework – GMMSeg – that addresses the limitations of current discriminative solutions from a generative perspective (Fig. ⁣{}_{\!} 1(b)). Our work not only represents a novel effort to advocate generative classifiers for end-to-end segmentation, but also evidences the merits of generative approaches in a challenging, dense classification task setting. In particular, we adopt a separate mixture of Gaussians for modeling the data distribution of each class in the feature space, i.e., class-conditional feature densities p(x∣c)p(\bm{x}|c). During training, GMM classifier is online optimized by a momentum version of (Sinkhorn) EM on large-scale, so as to ensure its generative nature and synchronization with the evolving feature space. Meanwhile, the feature extractor is end-to-end trained with the discriminative (cross-entropy) loss, i.e., maximizing the conditional likelihood p(c∣x)p(c|\bm{x}) derived with the generative GMM, so as to enable expressive representation learning. In this way, GMMSeg smartly learns generative classification with end-to-end discriminative representation in a compact and collaborative manner, exploiting the benefit of both generative and discriminative approaches. This also greatly distinguishes GMMSeg from most existing GMM based neural classifiers, which are either discriminatively trained ⁣{}_{\!} or trivially estimate a GMM in the feature space of a pre-trained discriminative classifier .

GMMSeg has several appealing facets: First, with the hybrid training strategy – online EM based classifier optimization and end-to-end discriminative representation learning, GMMSeg can precisely approximate the data distribution over a robust feature space. Second, the mixture components make GMMSeg a structured model that well adapts to multimodal data densities. Third, the distribution-preserving property allows GMMSeg to naturally reject abnormal inputs, without neither architectural change (like ⁣{}_{\!} ) nor re-training (like ⁣{}_{\!} ) nor post-calibration (like ⁣{}_{\!} ). Fourth, GMMSeg ⁣{}_{\!} is ⁣{}_{\!} a ⁣{}_{\!} principled ⁣{}_{\!} framework, ⁣{}_{\!} fully ⁣{}_{\!} compatible ⁣{}_{\!} with ⁣{}_{\!} modern ⁣{}_{\!} segmentation ⁣{}_{\!} network ⁣{}_{\!} architectures.

For thorough examination, in ⁣{}_{\!} §4.1, we approach GMMSeg on several representative segmentation architectures ⁣{}_{\!} (i.e., ⁣{}_{\!} DeepLabV3+ ⁣ ⁣{}_{\text{V3+}\!\!} , ⁣{}_{\!} OCRNet ⁣{}_{\!} , ⁣{}_{\!} UperNet ⁣{}_{\!} , ⁣{}_{\!} SegFormer ⁣{}_{\!} ), ⁣{}_{\!} with ⁣{}_{\!} diverse backbones (i.e., ResNet ⁣{}_{\!} , HRNet ⁣{}_{\!} , Swin ⁣{}_{\!} , MiT ⁣{}_{\!} ). Experimental results demonstrate GMMSeg even outperforms the softmax-based discriminative counterparts, e.g., 0.6% – 1.5%, 0.5% – 0.8%, and 0.7% – 1.7% mIoU gains over ADE20K ⁣{}_{\text{20K}\!} , Cityscapes ⁣{}_{\!} , and COCO-Stuff ⁣{}_{\!} , respectively. Furthermore, in §4.2, we validate our approach on anomaly segmentation. Without any modification, our Cityscapes-trained GMMSeg model is directly tested on Fishyscapes Lost&Found ⁣{}_{\!} and Road Anomaly ⁣{}_{\!} datasets, and outperforms all hand-tailored discriminative competitors.

To our best knowledge, GMMSeg is the first semantic segmentation method that reports promising results on both closed-set and open-world scenarios by using a single model instance. More notably, our impressive results manifest the advantages of generative classifiers in a large-scale real-world setting. We feel this work opens a new avenue for research in this field.

Related Work

Semantic Segmentation. Since the seminal work of FCN , deep-net segmentation solutions are typically built in a dense classification fashion, i.e., learning dense representation andcategorization end-to-end. By directly adopting discriminative softmax forcategorization,FCN-style solutions put focus on learning expressive dense representation; they modify the FCN architecture from various aspects, such as enlargingthe receptive field , modeling multi-scale context , investigating non-local operations , and exploring hierarchical information . With a similar goal of sharpening representation, later Transformer-style solutions empower attentive networks with, for instance, local contiguity and multi-level feature aggregation . Two very recent attentive models formulate the task in an alternative form of mask classification, however, still relying on discriminative softmax.

From the discussion above, we can find that current prevalent segmentation solutions are in essence a pixel-wise, discriminative classifier, which only learns decision boundaries between classes in the pixel feature space , without modeling the underlying data distribution. In contrast, our GMMSeg tackles the task from a generative viewpoint. GMMSeg deeply embeds generative optimization of GMMs into end-to-end dense representation learning, so as to comprehensively describe the class-aware knowledge in a discriminative feature space. GMMSeg is partly inspired by , that also probe data structures via intra-class clustering. However, the dense classification in the two works are achieved via non-parametric, nearest centroid retrieving – still a discriminative model. In , though data density is estimated (as a mixture of vMF distributions ), it is only used as a supervisory signal for dense embedding learning, and the final prediction is still made by a discriminative classifier – kk-NN. Our work represents the first step towards formulating (closed-set) semantic segmentation within a generative neural classification framework.

Discriminative vs Generative Classifiers. Generative classifiers and discriminative classifiers represent two contrasting ways of solving classification tasks . Basically, the generative classifiers (such as Linear Discriminant Analysis and naive Bayes) learn the class densities p(x∣c)p(\bm{x}|c), while the discriminative classifiers (such as softmax) learn the class boundaries p(c∣x)p(c|\bm{x}) without regard to the underlying class densities. In practical classification tasks, softmax discriminative classifier is used exclusively ⁣{}_{\!} , due to its simplicity and excellent discriminative performance. Nonetheless, genera- tive classifiers are widely agreed to have several advantages over their discriminative counterparts ⁣{}_{\!} , e.g., accurately modeling the input distribution, and explicitly identifying unlikely inputs in a natural way. Driven by this common belief, a surge of deep learning literature ⁣{}_{\!} investi- gated the potential (and the limitation) of generative classifiers in adversarial defense ⁣{}_{\!} , explainable AI ⁣{}_{\!} , out-of-distribution detection ⁣{}_{\!} , and semi-supervised learning ⁣{}_{\!} .

As GMMs can express (almost) arbitrary continuous distributions, it has been adopted in many neural classifiers ⁣{}_{\!} . However, most of these GMM classifiers are discriminative models ⁣{}_{\!} that ⁣{}_{\!} are ⁣{}_{\!} trained ⁣{}_{\!} ‘discriminatively’ ⁣ ⁣{}_{\!\!} (i.e., ⁣{}_{\!} maximizing ⁣{}_{\!} posteriors ⁣{}_{\!} p(c∣x)p(c|\bm{x})). ⁣{}_{\!} In ⁣{}_{\!} GMMSeg, ⁣{}_{\!} the ⁣{}_{\!} GMM ⁣{}_{\!} is ⁣{}_{\!} purely optimized via EM (i.e., estimating class densities p(x∣c)p(\bm{x}|c)) while the deep representation is trained via gradient backpropagation of the discriminative loss. Thus the whole GMMSeg is a hybrid of genera- tive GMM and discriminative representation, getting the best of two worlds. Although bearing the general idea of trading-off between generative and discriminative classifiers ⁣{}_{\!} , none of the previous hybrid algorithms demonstrate their utility in challenging segmentation tasks.

Anomaly Segmentation. Anomaly segmentation strives to identify unknown object regions, typically in road-driving scenarios . Existing solutions can be generally categorized into three classes: i) Uncertainty estimation based algorithms usually approximate the uncertainty from simple statistics of the classification probability or logits of pre-trained segmentation models , or adopt Bayesian neural networks with Monte-Carlo dropout to capture pixel uncertainty . ii) Outlier exposure based algorithms make use of auxiliary datasets as training samples of unexpected objects . Therefore, this type of algorithms requires re-training the segmentation network, resulting in performance degradation. iii) Image resynthesis based algorithms reconstruct the input image and discriminate the anomaly instances according to the reconstruction error .

With a generative classifier, our GMMSeg handles anomaly segmentation naturally, without neither external datasets of outliers, nor additional image resynthesis models. It also greatly differs from most uncertainty estimation-based methods that are post-processing techniques adjusting the prediction scores ⁣{}_{\!} of ⁣{}_{\!} softmax-based ⁣{}_{\!} segmentation ⁣{}_{\!} networks ⁣{}_{\!} . ⁣{}_{\!} The ⁣{}_{\!} most ⁣{}_{\!} relevant ⁣{}_{\!} ones ⁣{}_{\!} are ⁣{}_{\!} maybe ⁣{}_{\!} a ⁣{}_{\!} few density estimation-based models ⁣{}_{\!} , which directly measure the likelihood of samples w.r.t. the data distribution. However, they are either limited to pre-trained representation ⁣{}_{\!} or specialized for anomaly detection with simple data ⁣{}_{\!} . To our best knowledge, this is the first time to report promising results on both closed-set and open-world large-scale settings, through a single model instance without any change of network architecture as well as training and inference protocols.

Methodology

In this section, we first formalize modern semantic segmentation models within a dense discriminative classification framework and discuss defects of such discriminative regime from a probabilistic view- point (§3.1). Then we describe our new segmentation framework – GMMSeg – that brings a paradigm shift from the discriminative to generative (§3.2). Finally, in §3.3, we provide implementation details.

Recent mainstream solutions employ a deep neural network for pixel representation learning and softmax for semantic label prediction. Hence they are usually built as a composition of f ⁣∘ ⁣gf\!\circ\!g:

The feature extractor ff and softmax-based classifier gg are jointly trained end-to-end. Their corresponding parameters {θ,ω}\{\bm{\theta},\bm{\omega}\} are optimized by minimizing the so-called cross-entropy loss on D\mathcal{D}:

which is equivalent to maximizing conditional likelihood, i.e., Π(x,c)∈Dp(c∣x)\Pi_{(x,c)\in\mathcal{D}}p(c|\bm{x}). In some literature ⁣{}_{\!} , such learning strategy is called discriminative training. As softmax directly models the conditional probability distribution p(c∣x)p(c|\bm{x}) with no concern for modeling the input distribution p(x,c)p(\bm{x},c), existing softmax-based segmentation models are in essence a dense discriminative classifier.

Discriminative softmax typically gives good predictive performance, as the pixel classification rule depends only on the conditional distribution p(c∣x)p(c|\bm{x}) in the sense of minimum error rate and softmax optimizes the quantity of interest in a concise manner, i.e., learning a direct map from inputs xx to the class labels cc. In spite of its prevalence and effectiveness, this dense discriminative regime has some drawbacks that are still poorly understood: First, it attends only to learning the decision boundaries between the CC classes on the pixel embedding space, i.e., splitting the DD-dimensional feature space using CC different (D ⁣− ⁣1D\!-\!1)-dimensional hyperplanes. It achieves a simplified approach that eliminates extra parameters for modeling the data (representation) distribution ⁣{}_{\!} . However, from another perspective, it fails to capture the intrinsic class characteristics and is hard to achieve good generalization on unseen data. Second, in softmax, each class cc corresponds to only a single weight (wc,bc\bm{w}_{c},b_{c}). That means existing segmentation models rely on an implicit assumption of unimodality of data of each class in the feature space ⁣{}_{\!} . However, this unimodality assumption is rarely the case in real-world scenarios and makes the model less tolerant of intra-class variances ⁣{}_{\!} , especially when the multimodality remains in the feature space ⁣{}_{\!} . Third, softmax is not capable of inferring the data ⁣{}_{\!} distribution ⁣{}_{\!} – ⁣{}_{\!} it ⁣{}_{\!} is ⁣{}_{\!} notorious ⁣{}_{\!} with ⁣{}_{\!} inflating ⁣{}_{\!} the ⁣{}_{\!} probability ⁣{}_{\!} of ⁣{}_{\!} the ⁣{}_{\!} predicted ⁣{}_{\!} class ⁣{}_{\!} as ⁣{}_{\!} a ⁣{}_{\!} result ⁣{}_{\!} of ⁣{}_{\!} the ⁣{}_{\!} exponent ⁣{}_{\!} employed ⁣{}_{\!} on ⁣{}_{\!} the ⁣{}_{\!} network ⁣{}_{\!} outputs ⁣{}_{\!} . ⁣{}_{\!} Thus ⁣{}_{\!} the ⁣{}_{\!} prediction ⁣{}_{\!} score ⁣{}_{\!} of ⁣{}_{\!} a ⁣{}_{\!} class ⁣{}_{\!} is ⁣{}_{\!} useless ⁣{}_{\!} besides ⁣{}_{\!} its ⁣{}_{\!} comparative ⁣{}_{\!} value against other classes. This is the root cause of why existing segmentation models

are hard to identify pixel samples x′x^{\prime} of an unseen class (out-of-distribution data), i.e., c′ ⁣\centernot∈ ⁣{1, ⁣⋯ ⁣,C}c^{\prime}\!\centernot\in\!\{1,\!\cdots\!,C\}.

Accordingly, we argue that the time might be right to rethink the current de facto, discriminative segmentation regime, where the softmax classifier may actually cause more harm than good.

2 GMMSeg: Dense GMM Generative Classification

Our GMMSeg reformulates the task from a dense generative classification point of view. Instead of building posterior p(c∣x)p(c|\bm{x}) directly, generative classifiers predict labels using Bayes rule. Specifically, generative classifiers model the joint distribution p(x,c)p(\bm{x},c), by estimating the class-conditional distribu- tion p(x∣c)p(\bm{x}|c) along with the class prior p(c)p(c). Then, following Bayes rule, the posterior is derived as:

Since the class probabilities p(c)p(c) are typically set as a uniform prior (also in our case), estimating the class-conditional distributions (i.e., data densities) p(x∣c)p(\bm{x}|c) is the core and most difficult part of building a generative classifier. It is also worth noting that generative classifiers are optimized by approximating the data distribution Π(x,c)∈Dp(x∣c)\Pi_{(x,c)\in\mathcal{D}}p(\bm{x}|c), which is called generative training ⁣{}_{\!} .

Although discriminative classifiers demonstrate impressive performance in many application tasks, there are several crucial reasons for using generative rather than discriminative classifiers, which can be succinctly articulated by Feynman’s mantra “What I cannot create, I do not understand.” Surprisingly, generative classifiers have been rarely investigated in modern segmentation models.

Driven by the belief that generative classifiers are the right way to remove the shortcomings of discri- minative approaches, we revisit GMM – one of the most classic generative probabilistic classifiers. We couple the generative EM optimization of GMMs with the discriminative learning of the dense feature extractor ff – the most successful part of modern segmentation models, leading to a powerful, principled, and dense generative classification based segmentation framework – GMMSeg (Fig. ⁣{}_{\!} 2).

Specifically, GMMSeg adopts a weighted mixture of MM multivariate Gaussians for modeling the pixel data distribution of each class cc in the DD-dimensional embedding space:

To ⁣{}_{\!} find ⁣{}_{\!} the ⁣{}_{\!} optimal ⁣{}_{\!} parameters ⁣{}_{\!} of ⁣{}_{\!} the ⁣{}_{\!} GMM ⁣{}_{\!} classifier, ⁣{}_{\!} i.e., ⁣{}_{\!} {ϕc∗}c=1C\{\bm{\phi}^{\ast}_{c}\}_{c=1}^{C}, a ⁣{}_{\!} standard ⁣{}_{\!} approach ⁣{}_{\!} is ⁣{}_{\!} EM ⁣{}_{\!} , ⁣{}_{\!} i.e., ⁣{}_{\!} maximizing ⁣{}_{\!} the ⁣{}_{\!} log ⁣{}_{\!} likelihood ⁣{}_{\!} over ⁣{}_{\!} the ⁣{}_{\!} feature-label ⁣{}_{\!} pairs ⁣{}_{\!} {(xn,cn)}n=1 ⁣N\{(\bm{x}_{n},c_{n})\}_{n=1\!}^{N} in ⁣{}_{\!} the ⁣{}_{\!} training ⁣{}_{\!} dataset ⁣{}_{\!} D\mathcal{D}:

EM starts with some initial guess at the maximum likelihood parameters ϕc(0)\bm{\phi}_{c}^{(0)}, and then proceeds to iteratively create successive estimates ϕc(t) ⁣\bm{\phi}_{c}^{(t)\!} for t ⁣= ⁣1,2, ⁣⋯t\!=\!1,2,\!\cdots, by repeatedly optimizing a FF function ⁣{}_{\!} :

qc[m] ⁣= ⁣p(m∣x,c;ϕc)q_{c}[m]\!=\!p(m|\bm{x},c;\bm{\phi}_{c}) gives the probability that data x\bm{x} is assigned to component mm. FF is defined as:

where Nc ⁣N_{c\!} is the number of training samples labeled as cc and Ncm ⁣ ⁣= ⁣ ⁣∑n:cn=c ⁣ ⁣ qcn[m]N_{cm\!}\!=\!\!\sum_{n:c_{n}=c\!\!}~{}q_{cn}[m]. In E-step, we re- compute ⁣{}_{\!} the ⁣{}_{\!} posterior ⁣{}_{\!} qc(t) ⁣q^{(t)\!}_{c} over ⁣{}_{\!} the ⁣{}_{\!} M ⁣M_{\!} components ⁣{}_{\!} given ⁣{}_{\!} the ⁣{}_{\!} old ⁣{}_{\!} parameters ⁣{}_{\!} ϕ(t−1)\bm{\phi}^{(t-1)}. ⁣{}_{\!} In ⁣{}_{\!} M-step, ⁣{}_{\!} with ⁣{}_{\!} the ⁣{}_{\!} soft cluster assignment qc(t)q^{(t)}_{c}, the parameters are updated as ϕc(t) ⁣\bm{\phi}_{c}^{(t)\!} such that the FF function is maximized.

In practice, we find standard EM suffers from slow convergence and delivers unsatisfactory results (cf. ⁣{}_{\!} §4.3). A potential reason is the parameter sensitivity of EM – convergent parameters may change vastly even with slightly different initialization ⁣{}_{\!} . Drawing inspiration from recent optimal transport (OT) based clustering algorithms ⁣{}_{\!} , we introduce a uniform prior on the mixture weights πc\bm{\pi}_{c}, i.e., ∀c,m ⁣: ⁣πcm ⁣= ⁣1M\forall c,m\!:\!\pi_{cm}\!=\!\frac{1}{M}. Recalling qc[m] ⁣= ⁣p(m∣x,c)q_{c}[m]\!=\!p(m|\bm{x},c), we can derive a constraint Qc ⁣ ⁣= ⁣ ⁣{qc ⁣ ⁣: ⁣ ⁣1Nc ⁣ ⁣∑xn:cn=cp(m∣xn,c) ⁣= ⁣ ⁣1M}\mathcal{Q}_{c\!}\!=_{\!}\!\{q_{c\!\!}:\!\!\frac{1}{N_{c}\!}\!\sum_{\bm{x}_{n}:c_{n}=c}p(m|\bm{x}_{n},c)\!=_{\!}\!\frac{1}{M}\}. Then E-step in Eq. 6 is performed by restricting the optimi- zation of qcq_{c} over the set Qc\mathcal{Q}_{c}:

This can be intuitively viewed as an equipartition constraint guided clustering process: inside each class cc, we expect the NcN_{c} pixel samples to be evenly assigned to MM components. As indicated by ⁣{}_{\!} , Eq. 9 is analogous to entropy-regularized OT:

Our GMMSeg adopts a hybrid training strategy that is partly generative and partly discriminative:

In GMMSeg, GMM classifier (has C ⁣× ⁣MC\!\times\!M components in total) is purely optimized in a generative fashion, i.e., applying Sinkhorn EM to model the data densities p(x∣c)p(\bm{x}|c) within each class cc in the fea- ture space fθf_{\bm{\theta}}. The feature extractor/space fθf_{\bm{\theta}}, in contrast, is end-to-end trained in a discriminative manner, i.e., minimizing the cross-entropy loss over the posteriors output by the GMM. During each training iteration, the extractor’s parameters θ\bm{\theta} are only updated by the gradient backpropagated from the discriminative loss, while the GMM’s parameters {ϕc}c ⁣\{\bm{\phi}_{c}\}_{c\!} are only optimized by EM. To accurately estimate the GMM distributions, an external memory is adopted to store a large set of pixel representations, sampled from several preceding training batches, enabling large-scale EM. Moreover, since the feature space fθf_{\bm{\theta}} gradually evolves during training, we opt for a momentum EM: we directly use the GMM’s parameters {ϕ^c}c ⁣\{\hat{\bm{\phi}}_{c}\}_{c\!} estimated in the latest iteration as the initial guess in the current iteration {ϕc(0)}c\{{\bm{\phi}}^{(0)}_{c}\}_{c}, and adopt momentum update in the M-Step, i.e., {ϕc ⁣(t) ⁣ ⁣← ⁣(1 ⁣− ⁣τ)ϕc ⁣(t) ⁣ ⁣+ ⁣τϕ^c}c\{{\bm{\phi}}^{(t)\!}_{c\!}\!\leftarrow\!(1\!-\!\tau){\bm{\phi}}^{(t)\!}_{c\!}\!+\!\tau\hat{\bm{\phi}}_{c}\}_{c}, where the momentum coefficient is set as τ ⁣= ⁣0.999\tau\!=\!0.999. This makes our training more stable and accelerates the convergence of EM – we empirically find even one EM loop per training iteration is good enough.

This hybrid training scheme brings several advantages: First, GMMSeg achieves the merits of both generative and discriminative learning. The online EM based generative optimization enables the GMM to best fit the data distribution even on the evolving feature space. On the other hand, the feature space is discriminatively end-to-end trained under the guidance of the GMM classifier, so as to maximize the pixel-wise predictive performance. Second, as the generative EM optimization and discriminative stochastic training work in an independent yet closely collaborative manner, GMMSeg is fully compatible with modern segmentation network architectures and existing discriminative training objectives. It can be further advanced with the development of network architectures of the discriminative counterparts. Third, as GMMSeg explicitly models class-conditional data distribution p(x∣c)p(\bm{x}|c), it can naturally handle off-manifold examples, i.e., directly giving meaningful likelihood of the example fitting each class GMM distribution (see §4.2 for experiments on anomaly segmentation).

3 Implementation Details

Training In each training iteration, we conduct one loop of momentum (Sinkhorn) EM (i.e., t ⁣ ⁣= ⁣ ⁣1t_{\!}\!=_{\!}\!1) on current training batch as well as the external memory for the generative optimization of GMM, and backpropagate the gradient of the cross-entropy loss on current batch for the discriminative training of the feature extractor. The external memory maintains a queue for each component in each class; each queue gathers 32K pixel features from previous training batches in a first in, first out manner. To improve the diversity of the stored pixel features, we sample a sparse set of 100 pixels per class from each image, instead of directly storing the whole images into the memory. Note that the memory is discarded after training, and does not introduce extra overheads in inference.

Experiments

We respectively examine the efficacy and robustness of GMMSeg on semantic segmentation (§4.1) and anomaly segmentation (§4.2). In §4.3, we provide diagnostic analysis on our core model design.

Datasets. We conduct experiments on three widely used semantic segmentation datasets:

ADE20K ⁣{}_{\text{20K}\!}  ⁣{}_{\!} has ⁣{}_{\!} 20K/2K/3K ⁣{}_{\!} images ⁣{}_{\!} in ⁣{}_{\!} train/val/test ⁣{}_{\!} set, ⁣{}_{\!} with ⁣{}_{\!} 150 ⁣{}_{\!} stuff/object ⁣{}_{\!} categories ⁣{}_{\!} in ⁣{}_{\!} total.

Cityscapes ⁣{}_{\!} has 2,9752,975/500500/1,5241,524 fine-labeled images for train/val/test set with 1919 classes.

COCO-Stuff ⁣{}_{\!} has 1010K images (99K/11K for train/test), pixel-wise labeled with 171171 classes.

Base ⁣{}_{\!} Segmentation ⁣{}_{\!} Architectures ⁣{}_{\!} and ⁣{}_{\!} Backbones. ⁣{}_{\!} For ⁣{}_{\!} thorough ⁣{}_{\!} evaluation, ⁣{}_{\!} we ⁣{}_{\!} apply ⁣{}_{\!} GMMSeg ⁣{}_{\!} to four ⁣{}_{\!} famous ⁣{}_{\!} segmentation ⁣{}_{\!} architectures ⁣{}_{\!} (i.e., ⁣{}_{\!} DeepLabV3+ ⁣ ⁣{}_{\text{V3+}\!\!} , ⁣{}_{\!} OCRNet ⁣{}_{\!} , ⁣{}_{\!} UPerNet ⁣{}_{\!} , ⁣{}_{\!} Segfor- mer ⁣{}_{\!} ), with various backbones (i.e., ResNet ⁣{}_{\!} , HRNet ⁣{}_{\!} , Swin ⁣{}_{\!} , MiT ⁣{}_{\!} ). For fairness, we re-implement ⁣{}_{\!} these ⁣{}_{\!} models ⁣{}_{\!} using ⁣{}_{\!} the ⁣{}_{\!} standardized ⁣{}_{\!} hyper-parameter ⁣{}_{\!} setting ⁣{}_{\!} in ⁣{}_{\!} MMSegmentation ⁣{}_{\!} .

Training Details. GMMSeg is implemented on MMSegmentation ⁣{}_{\!} and follows the standard training setting for each dataset. All models are initialized with ImageNet-1K ⁣{}_{\!} pretrained back- bones and trained with commonly used data augmentations including resizing, flipping, color jittering and cropping. For ADE20K{}_{\text{20K}}/COCO-Stuff/Cityscapes, images are cropped to 512 ⁣× ⁣512512\!\times\!512/512 ⁣× ⁣512512\!\times\!512/768 ⁣× ⁣768768\!\times\!768 and models are trained for 160160K/8080K/8080K iterations with 1616/1616/88 batch size, using 8/16 NVIDIA Tesla A100 GPUs. Other training hyper-parameters (i.e., optimizers, learning rates, weight decays, schedulers) are set as the default in MMSegmentation and can be found in the supplementary.

Inference ⁣{}_{\!} Details. ⁣{}_{\!} For ⁣{}_{\!} ADE20K ⁣{}_{\text{20K}\!} and ⁣{}_{\!} COCO-Stuff, ⁣{}_{\!} we ⁣{}_{\!} keep ⁣{}_{\!} the ⁣{}_{\!} aspect ⁣{}_{\!} ratio ⁣{}_{\!} of ⁣{}_{\!} test ⁣{}_{\!} images ⁣{}_{\!} and ⁣{}_{\!} rescale the ⁣{}_{\!} short ⁣{}_{\!} side ⁣{}_{\!} to ⁣{}_{\!} 512. ⁣{}_{\!} For ⁣{}_{\!} Cityscapes, ⁣{}_{\!} sliding ⁣{}_{\!} window ⁣{}_{\!} inference ⁣{}_{\!} is ⁣{}_{\!} used ⁣{}_{\!} with 768 ⁣ ⁣× ⁣ ⁣768 ⁣768_{\!}\!\times_{\!}\!768_{\!} window ⁣{}_{\!} size. ⁣{}_{\!} Note that for fairness, all our results are reported without any test-time data augmentation.

Quantitative ⁣{}_{\!} Results. ⁣{}_{\!} Table ⁣{}_{\!} 1 demonstrates our quantitative results. Although mainly focusing on the comparison ⁣{}_{\!} with ⁣{}_{\!} the ⁣{}_{\!} four ⁣{}_{\!} base ⁣{}_{\!} segmentation models ⁣{}_{\!} , we further include five widely recognized ⁣{}_{\!} methods ⁣{}_{\!} for ⁣{}_{\!} completeness. ⁣{}_{\!} As ⁣{}_{\!} can ⁣{}_{\!} be ⁣{}_{\!} seen, ⁣{}_{\!} our ⁣{}_{\!} GMMSeg ⁣{}_{\!} outperforms ⁣{}_{\!} all ⁣{}_{\!} its ⁣{}_{\!} discriminative ⁣{}_{\!} counterparts ⁣{}_{\!} across ⁣{}_{\!} various ⁣{}_{\!} datasets, backbones, and network ⁣{}_{\!} architectures ⁣{}_{\!} (FCN-style ⁣{}_{\!}

ADE20K ⁣{}_{\text{20K}\!} ⁣{}_{\!}  ⁣{}_{\!} val. ⁣{}_{\!} With ⁣{}_{\!} FCN- style ⁣{}_{\!} segmentation ⁣{}_{\!} neural ⁣{}_{\!} ar- chitectures, ⁣{}_{\!} i.e., ⁣{}_{\!} DeepLabV3+ ⁣{}_{\text{V3+\!}} and ⁣{}_{\!} OCR, ⁣{}_{\!} GMMSeg ⁣{}_{\!} provides 1.2%\%/1.5%\% mIoU gains over corresponding discriminative models. ⁣{}_{\!} Similar ⁣{}_{\!} performance improvements, i.e., 1.0%\% and 0.6%\%, are also obtained with attentive neural architectures, i.e., ⁣{}_{\!} Swin-UperNet ⁣{}_{\!} and ⁣{}_{\!} SegFor- mer, manifesting the universality and efficacy of GMMSeg.

Cityscapes ⁣ ⁣{}_{\!\!}  ⁣ ⁣{}_{\!\!} val. ⁣ ⁣{}_{\!\!} Again ⁣{}_{\!} our ⁣ ⁣{}_{\!\!} GMMSeg surpasses all its discriminative counterparts by large margins, e.g., 0.5%\% over DeepLabV3+{}_{\text{V3+}}, 0.8%\% over OCRNet, 0.7%\% over Swin-UperNet, and 0.6%\% over SegFormer, suggesting its wide utility in this field.

COCO-Stuff ⁣{}_{\!}  ⁣{}_{\!} test. Our GMMSeg also demonstrates promising results. This is particularly impressive considering these results are achieved by a dense generative classifier, while the semantic segmentation task is commonly considered as a battlefield for discriminative approaches.

Qualitative Results. In Fig. 3, we illustrate the qualitative comparisons of our GMMSeg against SegFormer . It is evident that, among the representative samples in the three datasets, our method yields more accurate predictions when facing challenging scenarios, e.g., unconspicuous objects.

2 Experiments on Anomaly Segmentation

Datasets. ⁣{}_{\!} To ⁣{}_{\!} fully ⁣{}_{\!} reveal ⁣{}_{\!} the ⁣{}_{\!} merits ⁣{}_{\!} of ⁣{}_{\!} our ⁣{}_{\!} generative ⁣{}_{\!} method, ⁣{}_{\!} we ⁣{}_{\!} next ⁣{}_{\!} test ⁣{}_{\!} its ⁣{}_{\!} robustness ⁣{}_{\!} for ⁣{}_{\!} abnormal data, ⁣{}_{\!} i.e., ⁣{}_{\!} identifying ⁣{}_{\!} test ⁣{}_{\!} samples ⁣{}_{\!} of ⁣{}_{\!} unseen ⁣{}_{\!} classes, ⁣{}_{\!} using ⁣{}_{\!} two ⁣{}_{\!} popular ⁣{}_{\!} anomaly ⁣{}_{\!} segmentation ⁣{}_{\!} datasets:

Fishyscapes ⁣{}_{\!} Lost&Found ⁣{}_{\!} , ⁣{}_{\!} built ⁣{}_{\!} upon ⁣{}_{\!} , ⁣{}_{\!} has ⁣{}_{\!} 100100/275275 ⁣{}_{\!} val/test ⁣{}_{\!} images. ⁣{}_{\!} It ⁣{}_{\!} is ⁣{}_{\!} collected ⁣{}_{\!} under ⁣{}_{\!} the ⁣{}_{\!} same ⁣{}_{\!} setup ⁣{}_{\!} as ⁣{}_{\!} Cityscapes ⁣{}_{\!}  ⁣{}_{\!} but ⁣{}_{\!} with ⁣{}_{\!} real ⁣{}_{\!} obstacles ⁣{}_{\!} on ⁣{}_{\!} the ⁣{}_{\!} road. ⁣{}_{\!} Pixels ⁣{}_{\!} are ⁣{}_{\!} labeled ⁣{}_{\!} as ⁣{}_{\!} either ⁣{}_{\!} back- ground (i.e., pre-defined Cityscapes classes) or anomaly (i.e., other unexpected classes like crate).

Road Anomaly has 60 images containing anomalous objects in unusual road conditions.

Evaluation Metrics. The area under receiver operating characteristics (AUROC), average precision (AP), and false positive rate (FPR95) at a true positive rate of 95%, are adopted following ⁣{}_{\!} .

Experiment Protocol. As in ⁣{}_{\!} , we adopt ResNet101{}_{\text{101}}-DeepLabV3+{}_{\text{V3+}} architecture. For com- pleteness, we also report the results of our GMMSeg based on ResNet101{}_{\text{101}}-FCN and MiTB5{}_{\text{B5}}-SegFormer. All our models are the same ones in Table ⁣{}_{\!} 1, i.e., trained on Cityscapes train only. As GMMSeg estimates class densities p(x∣c)p(\bm{x}|c), it can naturally reject unlikely inputs (cf. ⁣{}_{\!} §3.3), i.e., directly thresholding  ⁣{}_{\!} − ⁣max⁡cp(x∣c)-\!\max_{c}p(\bm{x}|c) for computing the anomaly segmentation metrics, without any post-processing.

Quantitative Results. As shown in Table ⁣{}_{\!} 2, based on DeepLabV3+{}_{\text{V3+}} architecture, GMMSeg outper- forms ⁣{}_{\!} all ⁣{}_{\!} the ⁣{}_{\!} competitors ⁣{}_{\!} under ⁣{}_{\!} the ⁣{}_{\!} same ⁣{}_{\!} setting, ⁣{}_{\!} i.e., ⁣{}_{\!} neither ⁣{}_{\!} using ⁣{}_{\!} external ⁣{}_{\!} out-of-distribution ⁣{}_{\!} data ⁣{}_{\!} nor extra resynthesis module. Note that, rely on pre-trained discriminative segmentation models and thus have to make post-calibration. However, GMMSeg ⁣{}_{\!} directly ⁣{}_{\!} derives ⁣{}_{\!} meaningful ⁣{}_{\!} confidence ⁣{}_{\!} scores ⁣{}_{\!} from ⁣{}_{\!} likelihood ⁣{}_{\!} p(x∣c)p(\bm{x}|c). ⁣{}_{\!} Mahalanobis ⁣{}_{\!}  ⁣{}_{\!} also ⁣{}_{\!} models data density, yet, merely on pre-trained feature space with a single Gaussian per class. In contrast, GMMSeg performs much better, proving the superiority of mixture modeling and hybrid training. Even with a weaker architecture, i.e., FCN, GMMSeg still performs robustly. When adopting SegFormer, better performance is achieved.

Qualitative Results. In Fig. ⁣{}_{\!} 4, we visualize the anomaly score heatmaps generated by MSP ⁣{}_{\!} -DeepLabV3+{}_{\text{V3+}} ⁣{}_{\!} and GMMSeg-DeepLabV3+{}_{\text{V3+}}. The softmax based counterpart ignores the anomalies with overconfident predictions; in contrast, GMMSeg naturally rejects them (red colored regions).

3 Diagnostic Experiments

For in-depth analysis, we conduct ablative studies using DeepLabV3+ ⁣{}_{\text{V3+}}\! -ResNet101 ⁣{}_{\text{101}}\! segmentation architecture. Due to limited space, we put some diagnostic experiments in our supplementary material.

Online Hybrid Training. We first investigate our hybrid training strategy (cf. Eq. LABEL:eq:loss), where the discriminative feature extractor and generative GMM classifier are online optimized iteratively. Owe to this ingenious design, both components are gradually updated, aligned with and adaptive to each other, making GMMSeg a compact model. To fully demonstrate the effectiveness, we study a variant, DeepLabV3+ ⁣{}_{\text{V3+\!}} + ⁣{}_{\!} GMM, where a GMM classifier is directly fitted onto the feature space trained with the softmax classifier beforehand. As shown in Table 3, a clear performance drop is observed, i.e., mIoU: 46.0%46.0\%→\rightarrow 31.6%31.6\%, revealing the appealing efficacy of our end-to-end hybrid training strategy.

Discriminative GMMSeg vs. ⁣{}_{\!} Generative ⁣{}_{\!} GMMSeg. Our GMMSeg learns generative GMM via EM, i.e., max⁡p(x∣c;ϕ)\max p(\bm{x}|c;\bm{\phi}), with discri- minative representation learning, i.e., max⁡p(c∣x;θ)\max p(c|\bm{x};\bm{\theta}). A discriminative counterpart can be achieved by end-to-end learning all the parameters, i.e., {ϕ,θ}\{\bm{\phi},\bm{\theta}\}, with cross-entropy loss, i.e., max⁡p(c∣x;ϕ,θ)\max p(c|\bm{x};\bm{\phi},\bm{\theta}). Discriminative GMMSeg sacrifices data characterization for more flexiblility in discrimination, and yields poor performance in open-world setting. While inapparent effect on closed-set Cityscapes is observed, which in turn verifies the accurate specification of data distribution in generative GMMSeg.

Standard EM vs. Sinkhorn EM. In our GMMSeg, we leverage the entropic OT based Sinkhorn EM ⁣{}_{\!} (cf. ⁣{}_{\!} Eq. ⁣{}_{\!} 10) instead of the classic one (cf. ⁣{}_{\!} Eq. ⁣{}_{\!} 8) for the generative optimization of the GMM. In Table ⁣{}_{\!} 5a, we investigate the impacts of these two different EM algorithms and show that Sinkhorn EM is more favored. More specifically, during the E-step, rather than the vanilla EM assigning data samples to Gaussian components independently, Sinkhorn EM restricts the assignment with an equipartition constraint. As pointed out in ⁣{}_{\!} , incorporating such prior information about the mixing weights of GMM components leads to higher curvature around the global optimum. Our empirical results confirm this theoretical finding.

Number of EM Loop per Training Iteration. EM algorithm alternates between E-step and M-step for maximum-likelihood inference (cf. ⁣{}_{\!} Eq. ⁣{}_{\!} 6). In GMMSeg, in order to blend EM with stochastic gradient descent, we adopt an online version of (Sinkhorn) EM based on momentum update. In Table ⁣{}_{\!} 5a, we also study the influence of looping EM different times per training iteration. We can find that one loop per iteration is enough to catch the drift of the gradually updated feature space.

Number ⁣{}_{\!} of ⁣{}_{\!} Gaussian ⁣{}_{\!} Components ⁣{}_{\!} per ⁣{}_{\!} Class. ⁣{}_{\!} In ⁣{}_{\!} GMMSeg, ⁣{}_{\!} data ⁣{}_{\!} distribution ⁣{}_{\!} of ⁣{}_{\!} each ⁣{}_{\!} class ⁣{}_{\!} is ⁣{}_{\!} modeled by ⁣{}_{\!} a ⁣{}_{\!} mixture ⁣{}_{\!} of ⁣{}_{\!} M ⁣M_{\!} Gaussian ⁣{}_{\!} components ⁣{}_{\!} (cf. ⁣{}_{\!} Eq. ⁣{}_{\!} 4). ⁣{}_{\!} Table ⁣{}_{\!} 5b ⁣{}_{\!} shows ⁣{}_{\!} the ⁣{}_{\!} results ⁣{}_{\!} with ⁣{}_{\!} different ⁣{}_{\!} values ⁣{}_{\!} of MM. When M ⁣= ⁣1M\!=\!1, each class corresponds to a single Gaussian, which is directly estimated via Gaussian Discriminant Analysis, without EM. This baseline achieves 44.2%44.2\% mIoU. After adopting the mixture model, i.e., M ⁣ ⁣: ⁣ ⁣1 ⁣ ⁣→ ⁣ ⁣3 ⁣ ⁣→ ⁣ ⁣5M_{\!}\!:_{\!}\!1\!_{\!}\rightarrow_{\!}\!3_{\!}\!\rightarrow_{\!}\!5, the performance is greatly improved, i.e., mIoU: 44.2%44.2\%→\rightarrow 45.3%<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>→</mo></mrow><annotationencoding="application/x−tex">→</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.3669em;"></span><spanclass="mrel">→</span></span></span></span></span>46.0%45.3\%<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>→</mo></mrow><annotation encoding="application/x-tex">\rightarrow</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.3669em;"></span><span class="mrel">→</span></span></span></span></span>46.0\%. This verifies our hypothesis of class multimodality. Yet, further increasing component number (i.e., M ⁣: ⁣ ⁣5 ⁣→ ⁣15M\!:_{\!}\!5\!\rightarrow\!15) only brings marginal even negative gains, due to overparameterization.

Confidence Calibration. We further study the model calibration of GMMSeg and the discriminative counterpart, i.e., DeepLabV3+{}_{\text{V3+}} with the softmax classifier. In Fig. 5, we illustrate the Expected Calibration Error (ECE) along with reliability diagrams, which plot the expected pixel accuracy as a function of confidence . As seen, GMMSeg yields better calibrated prefictions, i.e., smaller gaps between the expected accuracy and confidence. On the other hand, the discriminative softmax produces confidences that deviate more from the true probabilities, and suffers higher calibration error accordingly, which again verifies the better reliability and interpretability of GMMSeg compared to its discriminative counterparts.

Runtime Analysis. The inference speed of GMMSeg is 13.3713.37 fps, which only yields negligible overhead w.r.t. its discriminative softmax counterpart, i.e., 13.3713.37 vs. 14.1614.16 fps. We measure the fps with a single NVIDIA GeForce RTX 3090 GPU with a batch size of one.

Conclusion

We ⁣{}_{\!} presented ⁣{}_{\!} GMMSeg, ⁣{}_{\!} the ⁣{}_{\!} first ⁣{}_{\!} generative ⁣{}_{\!} neural ⁣{}_{\!} framework ⁣{}_{\!} for ⁣{}_{\!} semantic ⁣{}_{\!} segmentation. ⁣{}_{\!} By ⁣{}_{\!} explicitly ⁣{}_{\!} modeling data distribution as GMMs, GMMSeg shows promise to solve the intrinsic limitations of current softmax based discriminative regime. It successfully optimizes generative GMM with end-to-end discriminative representation learning in a compact and collaborative manner. This makes GMMSeg principled and well applicable in both closed-set and open-world settings. We believe this work provides fundamental insights and can benefit a broad range of application tasks. As a part of our ⁣{}_{\!} future ⁣{}_{\!} work, ⁣{}_{\!} we ⁣{}_{\!} will ⁣{}_{\!} explore ⁣{}_{\!} our ⁣{}_{\!} algorithm ⁣{}_{\!} in ⁣{}_{\!} image ⁣{}_{\!} classification ⁣{}_{\!} and ⁣{}_{\!} trustworthy ⁣{}_{\!} AI ⁣{}_{\!} related ⁣{}_{\!} tasks.

References

Appendix A Detailed Training Parameters

We evaluate our GMMSeg on six base segmentation architectures. Four of them, i.e., DeepLabv3+{}_{\text{v3+}} , OCRNet, Swin-UperNet , SegFormer , are presented in our main paper. And the two additional base architectures, i.e., FCN and Mask2Former , are provided in this supplemental material (cf. §B). We follow the default training settings in the official Mask2Former codebase and MMSegmentation for Mask2Former and other base architectures respectively. In particular, we train FCN, DeepLabv3+{}_{\text{v3+}} and OCRNet using SGD optimizer with initial learning rate 0.10.1, weight decay 4e-4 with polynomial learning rate annealing; we train Swin-UperNet and SegFormer using AdamW optimizer with initial learning rate 6e-5, weight decay 1e-2 with polynomial learning rate annealing; we train Mask2Former using AdamW optimizer with initial learning rate 1e-4, weight decay 5e-2 and the learning rate is decayed by a factor of 10 at 0.9 and 0.95 fractions of the total training steps.

Appendix B More Experimental Results

More Base Segmentation Architectures. We first demonstrate the efficacy of our GMMSeg on two additional base segmentation architectures, i.e., FCN and Mask2Former , with quantitative results summarized in Table B. We train FCN based models with the according training hyperparameter

Impact of Memory Capacity. In Table 8, we further explore the influence of the memory capacity, i.e., the amount of pixel representations stored for class-wise EM estimation, with DeepLabV3+{}_{\text{V3+}}-ResNet101{}_{\text{101}} on ADE20K{}_{\text{20K}} val trained for 80K iterations. For the first row, where the memory size is set to , the EM is only performed within mini-batches. Not surprisingly, data distribution estimated at such a local scale is far from accurate, leading to inferior results. With enlarged memory capacity, the performance is increased. When the performance reaches saturation, the stored pixel samples are sufficient enough to represent the true data distribution of the whole training set.

Semantic ⁣{}_{\!} Segmentation. We ⁣{}_{\!} illustrate ⁣{}_{\!} the ⁣{}_{\!} qualitative ⁣{}_{\!} comparisons ⁣{}_{\!} of ⁣{}_{\!} GMMSeg ⁣{}_{\!} equipped ⁣{}_{\!} SegFormer -MiTB5{}_{\text{B5}} ⁣{}_{\!} against ⁣{}_{\!} the ⁣{}_{\!} original ⁣{}_{\!} model ⁣{}_{\!} on ⁣{}_{\!} ADE20K{}_{\text{20K}} ⁣{}_{\!}  ⁣{}_{\!} (Fig. ⁣{}_{\!} 6), ⁣{}_{\!} Cityscapes ⁣{}_{\!}  ⁣{}_{\!} (Fig. ⁣{}_{\!} 7) ⁣{}_{\!} and ⁣{}_{\!} COCO-Stuff ⁣{}_{\!}  ⁣{}_{\!} (Fig. ⁣{}_{\!} 8). It is evident that, benefiting from the accurate data characterization modeling, GMMSeg is less confused by object categories and gives preciser predictions than SegFormer.

Anomaly ⁣{}_{\!} Segmentation. We ⁣{}_{\!} then ⁣{}_{\!} show ⁣{}_{\!} more ⁣{}_{\!} qualitative ⁣{}_{\!} results ⁣{}_{\!} of ⁣{}_{\!} MSP -DeepLabV3+{}_{\text{V3+}}  ⁣{}_{\!} and ⁣{}_{\!} GMMSeg-DeepLabV3+{}_{\text{V3+}} ⁣{}_{\!} on ⁣{}_{\!} Fishyscapes ⁣{}_{\!} Lost&Found ⁣{}_{\!} val. As observed, different from MSP, GMMSeg gets rid of being overwhelmed by overconfident predictions and successfully identifies the anomalies.