MaskGIT: Masked Generative Image Transformer
Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, William T. Freeman
Introduction
Deep image synthesis as a field has seen a lot of progress in recent years. Currently holding state-of-the-art results are Generative Adversarial Networks (GANs), which are capable of synthesizing high-fidelity images at blazing speeds. They suffer from, however, well known issues include training instability and mode collapse, which lead to a lack of sample diversity. Addressing these issues still remains open research problems.
Inspired by the success of Transformer and GPT in NLP, generative transformer models have received growing interests in image synthesis . Generally, these approaches aim at modeling an image like a sequence and leveraging the existing autoregressive models to generate image. Images are generated in two stages; the first stage is to quantize an image to a sequence of discrete tokens (or visual words). In the second stage, an autoregressive model (e.g., transformer) is learned to generate image tokens sequentially based on the previously generated result (i.e. autoregressive decoding). Unlike the subtle min-max optimization used in GANs, these models are learned by maximum likelihood estimation. Because of the design differences, existing works have demonstrated their advantages over GANs in offering stabilized training and improved distribution coverage or diversity.
Existing works on generative transformers mostly focus on the first stage, i.e. how to quantize images such that information loss is minimized, and share the same second stage borrowed from NLP. Consequently, even the state-of-the-art generative transformers still treat an image naively as a sequence, where an image is flattened into a 1D sequence of tokens following a raster scan ordering, i.e. from left to right line-by-line (cf. Figure 2). We find this representation neither optimal nor efficient for images. Unlike text, images are not sequential. Imagine how an artwork is created. A painter starts with a sketch and then progressively refines it by filling or tweaking the details, which is in clear contrast to the line-by-line printing used in previous work . Additionally, treating image as a flat sequence means that the autoregressive sequence length grows quadratically, easily forming an extremely long sequence–longer than any natural language sentence. This poses challenges for not only modeling long-term correlation but also renders the decoding intractable. For example, it takes a considerable 30 seconds to generate a single image on a GPU autoregressively with 32x32 tokens.
This paper introduces a new bidirectional transformer for image synthesis called Masked Generative Image Transformer (MaskGIT). During training, MaskGIT is trained on a similar proxy task to the mask prediction in BERT . At inference time, MaskGIT adopts a novel non-autoregressive decoding method to synthesize an image in constant number of steps. Specifically, at each iteration, the model predicts all tokens simultaneously in parallel but only keeps the most confident ones. The remaining tokens are masked out and will be re-predicted in the next iteration. The mask ratio is decreased until all tokens are generated with a few iterations of refinement. As illustrated in Figure 2, MaskGIT’s decoding is an order-of-magnitude faster than the autoregresive decoding as it only takes 8 steps, instead of 256 steps, to generate an image and the predictions within each step are parallelizable. Moreover, instead of conditioning only on previous tokens in the order of raster scan, bidirectional self-attention allows the model to generate new tokens from generated tokens in all directions. We find that the mask scheduling (i.e. fraction of the image masked each iteration) significantly affects generation quality. We propose to use the cosine schedule and substantiate its efficacy in the ablation study.
On the ImageNet benchmark, we empirically demonstrate that MaskGIT is both significantly faster (by up to 64x) and capable of generating higher quality samples than the state-of-the-art autoregressive transformer, i.e. VQGAN, on class-conditional generation with 256256 and 512512 resolution. Even compared with the leading GAN model, i.e. BigGAN, and diffusion model, i.e. ADM , MaskGIT offers comparable sample quality while yielding more favourable diversity. Notably, our model establishes new state-of-the-arts on classification accuracy score (CAS) and on FID for synthesizing 512512 images. To our knowledge, this paper provides the first evidence demonstrating the efficacy of the masked modeling for image generation on the common ImageNet benchmark.
Furthermore, MaskGIT’s multidirectional nature makes it readily extendable to image manipulation tasks that are otherwise difficult for autoregressive models. Fig. 1 shows a new application of class-conditional image editing in which MaskGIT re-generates content inside the bounding box based on the given class while keeping the context (outside of the box) unchanged. This task, which is either infeasible for autoregressive model or difficult for GAN models, is trivial for our model. Quantitatively, we demonstrate this flexibility by applying MaskGIT to image inpainting, and image extrapolation in arbitrary directions. Even though our model is not designed for such tasks, it obtains comparable performance to the dedicated models on each task.
Related Work
Deep generative models have achieved lots of successes in image synthesis tasks. GAN based methods demonstrate amazing capability in yielding high-fidelity samples . In contrast, likelihood-based methods, such as Variational Autoencoders (VAEs) , Diffusion Models and Autoregressive Models , offer distribution coverage and hence can generate more diverse samples .
However, maximizing likelihood directly in pixel space can be challenging. So instead, VQVAE proposes to generate images in latent space in two stages. In the first stage, which is known as tokenization, it tries to compress images into discrete latent space, and primarily consists of three components:
a decoder which predicts the reconstructed image from the visual tokens .
In the second stage, it first predicts the latent priors of the visual tokens using deep autoregressive models, and then uses the decoder from the first stage to map the token sequences into image pixels. Several approaches have followed this paradigm due to the efficacy of the two-stage approach. DALL-E uses Transformers to improve token prediction in the second stage. VQGAN adds adversarial loss and perceptual loss in the first stage to improve the image fidelity. A contemporary work to ours, VIM , proposes to use a VIT backbone to further improve the tokenization stage. Since these approaches still employ an auto-regressive model, the decoding time in the second stage scales with the token sequence length.
2 Masked Modeling with Bi-directional Transformers
The transformer architecture , was first proposed in NLP, and has recently extended its reach to computer vision . Transformer consists of multiple self-attention layers, allowing interactions between all pairs of elements in the sequence to be captured. In particular, BERT introduces the masked language modeling (MLM) task for language representation learning. The bi-directional self-attention used in BERT allows the masked tokens in MLM to be predicted utilizing context from both directions. In vision, the masked modeling in BERT has been extended to image representation learning with images quantized to discrete tokens. However, few works have successfully applied the same masked modeling to image generation because of the difficulty in performing autoregressive decoding using bi-directional attentions. To our knowledge, this paper provides the first evidence demonstrating the efficacy of masked modeling for image generation on the common ImageNet benchmark. Our work is inspired by bi-directional machine translation in NLP, and our novelty lies in the proposed new masking strategy and decoding algorithm which, as substantiated by our experiments, are essential for image generation.
Method
Our goal is to design a new image synthesis paradigm utilizing parallel decoding and bi-directional generation.
We follow the two-stage recipe discussed in 2.1, as illustrated in Figure 3. Since our goal is to improve the second stage, we employ the same setup for the first stage as in the VQGAN model , and leave potential improvements to the tokenization step to future work.
For the second stage, we propose to learn a bidirectional transformer by Masked Visual Token Modeling (MVTM). We introduce MVTM training in 3.1 and the sampling procedure in 3.2. We then discuss the key technique of masking design in 3.3.
Let denote the latent tokens obtained by inputting the image to the VQ-encoder, where is the length of the reshaped token matrix, and the corresponding binary mask. During training, we sample a subset of tokens and replace them with a special [MASK] token. The token is replaced with [MASK] if , otherwise, when , will be left intact.
The sampling procedure is parameterized by a mask scheduling function , and executes as follows: we first sample a ratio from to , then uniformly select tokens in to place masks, where is the length. The mask scheduling significantly affects the quality of image generation and will be discussed in 3.3.
Denote the result after applying mask to . The training objective is to minimize the negative log-likelihood of the masked tokens:
Concretely, we feed the masked into a multi-layer bidirectional transformer to predict the probabilities for each masked token, where the negative log-likelihood is computed as the cross-entropy between the ground-truth one-hot token and predicted token. Notice the key difference to autoregressive modeling: the conditional dependency in MVTM has two directions, which allows image generation to utilize richer contexts by attending to all tokens in the image.
2 Iterative Decoding
In autoregressive decoding, tokens are generated sequentially based on previously generated output. This process is not parallelizable and thus very slow for image because the image token length, e.g. 256 or 1024, is typically much larger than that of language. We introduce a novel decoding method where all tokens in the image are generated simultaneously in parallel. This is feasible due to the bi-directional self-attention of MTVM.
In theory, our model is able to infer all tokens and generate the entire image in a single pass. We find this challenging due to inconsistency with the training task. Below, the proposed iterative decoding is introduced. To generate an image at inference time, we start from a blank canvas with all the tokens masked out, i.e. . For iteration , our algorithm runs as follows:
Mask Schedule. We compute the number of tokens to mask according to the mask scheduling function by , where is the input length and is the total number of iterations.
Mask. We obtain by masking tokens in . The mask for iteration is calculated from:
where is the confidence score for the -th token.
The decoding algorithm synthesizes an image in steps. At each iteration, the model predicts all tokens simultaneously but only keeps the most confident ones. The remaining tokens are masked out and re-predicted in the next iteration. The mask ratio is made decreasing until all tokens are generated within iterations. In practice, the masking tokens are randomly sampled with temperature annealing to encourage more diversity, and we will discuss its effect in 3. Figure 2 illustrates an example of our decoding process. It generates an image in iterations, where the unmasked tokens at each iteration are highlighted in the grid, e.g. when we only keep 1 token and mask out the rest.
3 Masking Design
We find that the quality of image generation is significantly affected by the masking design. We model the masking procedure by a mask scheduling function that computes the mask ratio for the given latent tokens. As discussed, the function is used in both training and inference. During inference time, it takes the input of indicating the progress in decoding. In training, we randomly sample a ratio in to simulate the various decoding scenarios.
BERT uses a fixed mask ratio of 15% , i.e., it always masks 15% of the tokens, which is unsuitable for our task since our decoder needs to generate images from scratch. New masking scheduling is thus needed. Before discussing specific schemes, we first examine the property of the mask scheduling function. First, needs to be a continuous function bounded between and for . Second, should be (monotonically) decreasing with respect to , and it holds that and . The second property ensures the convergence of our decoding algorithm.
This paper considers common functions and makes simple transformations so that they satisfy the properties. Figure 8 visualizes these functions which are divided into three groups:
Linear function is a straightforward solution, which masks an equal amount of tokens each time.
Concave function captures the intuition that image generation follows a less-to-more information flow. In the beginning, most tokens are masked so the model only needs to make a few correct predictions for which the model feel confident. Towards the end, the mask ratio sharply drops, forcing the model to make a lot more correct predictions. The effective information is increasing in this process. The concave family includes cosine, square, cubic, and exponential.
Convex function, conversely, implements a more-to-less process. The model needs to finalize a vast majority of tokens within the first couple of iterations. This family includes square root and logarithmic.
We empirically compare the above mask scheduling functions in 3 and find the cosine function works the best in all of our experiments.
Experiments
In this section, we empirically evaluate MaskGIT on image generation in terms of quality, efficiency and flexibility. In 4.2, we evaluate MaskGIT on the standard class-conditional image generation tasks on ImageNet 256256 and 512512. In 4.3, we show MaskGIT’s versatility by demonstrating its performance on three image editing tasks, image inpainting, outpainting, and editing. In 3, we verify the necessity of our design of mask scheduling. We will release the code and model for reproducible research.
For each dataset, we only train a single autoencoder, decoder, and codebook with 1024 tokens on cropped 256x256 images for all the experiments. The image is always compressed by a fixed factor of 16, i.e. from to a grid of tokens in the size of , where = and =. We find that this autoencoder, together with the codebook, can be reused to synthesize 512512 images.
All models in this work have the same configuration: 24 layers, 8 attention heads, 768 embedding dimensions and 3072 hidden dimensions. Our models use learnable positional embedding, LayerNorm, and truncated normal initialization (stddev=). We employ the following training hyperparameters: label smoothing=, dropout rate=, Adam optimizer with = and =. We use RandomResizeAndCrop for data augmentation. All models are trained on 4x4 TPU devices with a batch size of 256. ImageNet models are trained for 300 epochs while the Places2 model is trained for 200 epochs.
2 Class-conditional Image Synthesis
We evaluate the performance of our model on class-conditional image synthesis on ImageNet 256256 and 512512. Our main results are summarized in Table 1.
Quality. On ImageNet 256256, without any special sampling strategies such as beam-search, top-k or nucleus sampling heuristics or classifier guidance , we significantly outperform VQGAN in both Fréchet Inception Distance (FID) ( vs ) and Inception Score (IS) ( vs ). We also report the results with classifier-based rejection sampling in the appendix B.
We also train a VQGAN baseline with the same tokenizer and hyperparameters as MaskGIT’s in order to further highlight the difference between bi-directional and uni-directional transformers, and find that on both resolutions, MaskGIT still outperforms our implemented baseline by a significant margin.
Furthermore, MaskGIT improves BigGAN’s FIDs on both resolutions, achieving a new state-of-the-art on 512512 with an FID of .
Speed. We evaluate model speed by assessing the number of steps, i.e. forward passes, each model requires to generate a sample. As shown in Table 1, MaskGIT requires the fewest steps among all non-GAN-based models on both resolutions.
To further substantiate the speed difference between MaskGIT and autoregressive models, we perform a runtime comparison between MaskGIT and VQGAN’s decoding processes. As illustrated in Figure 4, MaskGIT significantly accelerates VQGAN by -x, with a speedup that gets more pronounced as the image resolution (and thus the input token length) grows.
Diversity. We consider Classification Accuracy Score (CAS) and Precision/Recall as two metrics for evaluating sample diversity, in addition to sample quality.
CAS involves first training a ResNet-50 classifier solely on the samples generated by the candidate model, and then measuring the classifier’s classification accuracy on the ImageNet validation set. The last two columns in Table 1 present the CAS results, where the scores of the classifier trained on real ImageNet training data are included for reference (76.6% and 93.1% for the top-1 and top-5 accuracy). For image resolution 256256, we follow the common practice of using data augmentation RandAugment, and report the scores trained without augmentation in the appendix B. We find that MaskGIT significantly outperforms prior work VQVAE-2 and VQGAN, establishing a new state-of-the-art of CAS on the ImageNet benchmark on both resolutions.
The Precision/Recall results in Table 1 show that MaskGIT achieves better coverage (Recall) compared to BigGAN, and better sample quality (Precision) compared to likelihood-based models such as VQVAE-2 and diffusion models. Compared to our baseline VQGAN, we improve the diversity as measured by recall while slightly boosting its precision.
In contrast to BigGAN’s samples, MaskGIT’s samples are more diverse with more varied lighting, poses, scales and context as shown in Figure 5. More comparisons are available in the appendix B.
3 Image Editing Applications
In this subsection, we present direct applications of MaskGIT on three image editing tasks: class-conditional image editing, image inpainting, and outpainting. All three tasks can be almost trivially translated to ones that MaskGIT can handle if we consider the task as just a constraint on the initial binary mask MaskGIT uses in its iterative decoding, as discussed in 3.2. We show that without modifications to the architecture or any task-specific training, MaskGIT is capable of generating very compelling results on all three applications. Furthermore, MaskGIT obtains comparable performance to dedicated models on both inpainting and outpainiting, even though it is not designed specifically for either task.
Class-conditional Image Editing. We define a new class-conditional image editing task to showcase MaskGIT’s flexibility. In this task, the model regenerates content specified inside a bounding box on the given class while preserving the context, i.e. content outside of the box. It is infeasible for autoregressive methods due to the violation to their prediction orders.
For MaskGIT, however, it is a trivial task if we consider the bounding box region as the input of initial mask to the iterative decoding algorithm. Figure 6 shows a few example results. More can be found in the appendix C.
In these examples, we observe that MaskGIT can reasonably replace the selected object while preserving, or to some extend even completing, the context in the background. Furthermore, we find that MaskGIT seems to be capable of synthesizing unnatural yet plausible combinations unseen in the ImageNet training set, e.g. a flying cat, cat in a soup bowl, and cat in a flower. This suggests that MaskGIT has incidentally learned useful representations for composition, which may be further exploited in related tasks in future works.
Image Inpainting. Image inpainting or image completion is a fundamental image editing task to synthesize contents in missing regions so that the completion looks visually realistic. Traditional patch-based methods work well on texture regions, while deep learning based methods have been demonstrated to synthesize images requiring better semantic coherence. Both approaches have been are extensively studied in computer vision.
We extend MaskGIT to this problem by tokenizing the masked image and interpreting the inpainting mask as the initial mask in our iterative decoding. We then composite the output image by linearly blending it with the input based on the masking boundary following . To match the training of our baselines, we train MaskGIT on the 512512 center-cropped images from the Places2 dataset. All hyperparameters are kept the same as the MaskGIT model trained on ImageNet.
We compare MaskGIT against common GAN-based baselines, including DeepFillv2 and HiFill, on inpainting with a central 50% 50% mask, which are evaluated on the Places2 validation set. Table 2 summarizes the quantitative comparisons. MaskGIT beats both DeepFill and HiFill in FID and IS by a significant margin, while achieving scores close to the state-of-the-art inpainting approach CoModGAN . We show more qualitative comparisons with CoModGAN in the appendix E.
Image Outpainting. Outpainting, or image extrapolation, is an image editing task that has received increased attention recently. It is seen as a more challenging task than inpainting due to the fewer constraints from surrounding pixels and thus more uncertainty in the predicted regions. Our adaptation of the problem and the model used in the following evaluation is the same as in inpainting.
We compare against common GAN-based baselines, including Boundless , In&Out , InfinityGAN, and CoModGAN on extrapolating rightward with a 50% ratio. We evaluate on the image set generously provided by the authors of InfinityGAN and In&Out.
Table 2 summarizes the quantitative comparisons. MaskGIT beats all baselines and achieves state-of-the-art FID and IS. As the examples in Figure 7 illustrate, MaskGIT is also capable of synthesizing diverse results given the same input with different seeds. We observe that MaskGIT completes objects and global structures particularly well, and hypothesize that this is thanks to the model learning useful representations with the global attentions in the transformer.
4 Ablation Studies
We conduct ablation experiments using the default setting on ImageNet 256256.
Mask scheduling. A key design of MaskGIT is the mask scheduling function used in both training and iterative decoding. We compare the functions discussed in 3.3, visualize them in Figure 8, and summarize the results in Table 3.
We observe that concave functions generally obtain better FID and IS than linear, followed by the convex functions. While cosine and square perform similarly relative to other functions, cosine slightly edges out square in all scores, making cosine the default in our model.
We hypothesize that concave functions perform favorably because they 1) challenge training with more difficult cases (i.e. encouraging larger mask ratios), and 2) appropriately prioritize the less-to-more prediction throughout the decoding. That said, over-prioritization seems to be costly as well, as shown by the cubic function being worse than square, and exponential being much worse than all other concave functions.
Iteration number. We study the effect of the number of iterations () on our model by running all candidate masking functions with different s. As shown in Figure 8, under the same setting, more iterations are not necessarily better: as increases, aside from the logarithmic function which performs poorly throughout, all other functions hit a “sweet spot” where the model’s performance peaks before it worsens again. The sweet spot also gets “delayed” as functions get less concave. As shown, among functions that achieve strong FIDs (i.e. cosine, square, and linear), cosine not only has the strongest overall score, but also the earliest sweet spot at a total of to iterations. We hypothesize that such sweet spots exist because too many iterations may discourage the model from keeping less confident predictions, which worsens the token diversity. We think further study on the masking design would be interesting for future work.
Conclusion
In this paper, we propose MaskGIT, a novel image synthesis paradigm using a bidirectional transformer decoder. Trained on Masked Visual Token Modeling, MaskGIT learns to generate samples using an iterative decoding process within a constant number of iterations. Experimental results show that MaskGIT significantly outperforms the state-of-the-art transformer model on conditional image generation, and our model is readily extendable to various image manipulation tasks. As MaskGIT achieves competitive performance with state-of-the-art GANs, applying our approach to other synthesis tasks is a promising direction for future work. Please see the appendix F for the limitations and future work.
Acknowledgement The authors would like to thank Xiang Kong for inspiring related works and anonymous reviewers for helpful comments.
References
Appendix A Discussion on Image Reconstruction
In 4.2, we primarily evaluate MaskGIT on class-conditional image generation tasks. Here we offer more discussion on its performance on image reconstruction. We set up by first randomly sampling input mask with a mask ratio of the visual tokens masked out, and then running MaskGIT’s iterative decoding algorithm to reconstruct images. Figure 10 shows the PSNR and LPIPS of the reconstructed samples as functions of , whereas Figure 9 visualizes two examples of this process with ranging from to .
We observe that MaskGIT reconstructs holistic information (e.g. pose and shape of the foreground objects) even with a very high percentage (e.g. 95%) of tokens masked out. More importantly, there seems to exist an inflection point around 90%: while both reconstruction quality and consistency improve drastically as the mask ratio decreases until 90%, after 90% further improvements are slowed down. This observation is corroborated by the large jump in the visual similarity between reconstruction samples and the original images from 95% to 90% in Figure 9, e.g. the fence in front of the tiger and the car’s color are consistently captured once the mask ratio is below 90%, but not at 95%.
In other words, we find that visual tokens are highly redundant. For a holistic reconstruction, only a very small portion (e.g. 10%) of the tokens are essential; the remaining ones merely improve the recovery of finer appearance or details. This echos our intuition behind the masking design laid out in 3.3 that the prediction of the first few tokens is key to image generation. Similar observations on the spatial redundancy of images are discussed in a concurrent paper MAE . In their work, they find that masking a high proportion of the input image yields a nontrivial and meaningful self-supervisory task for image representation learning.
Appendix B Additional Class-conditional Image Generation Results
In this section, we report additional results on class-conditional image generation.
We follow prior transformer-based methods to employ the classifier-based rejection sampling to improve the sample quality scores. Specifically, we use a pre-trained ResNet classifier to score output samples based on the predicted probability and keep samples with an acceptance rate of , as in VQGAN . As shown in Table 4, MaskGIT demonstrates consistent improvement over VQGAN, and is comparable with ADM with classifier guidance . More importantly, by adding the rejection sampling, MaskGIT achieves state-of-the-art Inception Scores ( on 256256 and on 512512).
In Table 5, we report Precision and Recall scores calculated using Inception features . In contrast to the VGG feature-based scores, which we report in Table 1 for a more direct comparison with prior work , we find that the Inception feature-based scores are more consistent with our qualitative observations that VQGAN’s samples are more diverse than BigGAN’s. Under both measures, MaskGIT ’s recall scores outperform those of BigGAN and VQGAN. We also report CAS evaluated on classifiers trained without augmentation from RandAugment. Consistent with our main results, MaskGIT outperforms BigGAN and our baseline VQGAN by a large margin.
Finally, we show a few comparisons of the class-conditional samples generated by MaskGIT with the samples generated by BigGAN-deep and VQVAE-2 in Figure 11, 12, and 13.
Appendix C Additional Examples of Class-conditional Image Editing Applications
We show more examples of class-conditional image editing in Figure 14, and examples of image-conditional panorama synthesis in Figure 15.
Appendix D Image Outpainting Comparisons with SOTA Transformer-based Approaches
In Figure 16 and 17, we show a few outpainting comparisons among MaskGIT, ImageGPT, and VQGAN. In each set of images, we show the groundtruth (left), extrapolated samples using only the top half of the groundtruth (middle), and extrapolated samples using only the bottom half of the groundtruth (right).
MaskGIT and VQGAN can both perform on higher resolutions by taking advantage of tokenization and thus achieve higher sample fidelity than ImageGPT, which runs on a maximum resolution of . At the same time, MaskGIT demonstrates stronger flexibility than ImageGPT and VQGAN in that it can outpaint in arbitrary directions (e.g. both upward and downward), while ImageGPT and VQGAN can only handle outpainting in one direction with a single model due to their autoregressive natures.
Appendix E Image Inpainting and Outpainting Comparisons with SOTA GAN-based Approaches
In this section, we show more qualitative comparisons with state-of-the-art GAN-based image completion methods in Figure 18 and Figure 19. Quantitative results have been discussed in 4.3.
We find that compared to prior GAN-based methods, MaskGIT demonstrates a stronger capability of completing structures coherently, and its samples contain fewer artifacts. In Figure 19, MaskGIT completes the bridge in row two and the building in the second to last row, which all GAN methods struggle to do in comparison.
In addition, we compare with CoModGAN on image completion tasks with large masking ratios, i.e. conditioning on the center 5050% and the center 31.25%31.25% respectively, which are challenging cases for traditional GANs. Examples are shown in Figure 20.
Appendix F Limitations and Failure Cases
In Figure 21, we show several limitations and failure cases of our approach. (A) and (B) are examples of semantic and color shifts in MaskGIT’s outpainting results. Due to its limited attention size, MaskGIT may ”forget” the synthesized semantics or color from one end when it’s outpainting the other end. (C) and (D) show cases where our approach may sometimes ignore or modify objects on the boundary when applied to outpainting and inpainting. (E) showcases MaskGIT’s failure mode in which it causes oversmoothing or creates undesired artifacts on complex structures such as human faces, text and symmetric objects. The improvement for these circumstances remains future work.