Learning Generative Structure Prior for Blind Text Image Super-resolution
Xiaoming Li, Wangmeng Zuo, Chen Change Loy
Introduction
Blind text super-resolution (SR) aims at restoring a high-resolution (HR) image from a low-resolution (LR) text image that is corrupted by unknown degradation. The problem is relatively less explored than SR of natural and face images.
This seemingly easy task actually possesses many unique challenges. In particular, each character owns its unique structure. A suboptimal restoration that destroys the structure with, e.g., distorted, missing or additional strokes, is easily perceptible and may alter the meaning of the original word. The task becomes more difficult when dealing with writing systems such as Chinese, Kanji and Hanja characters, in which the patterns of thousands of characters vary in complex ways and are mostly composed of different subunits (or compound ideographs). A slight deviation in strokes, say ‘太’ and ‘大’ carries entirely different meanings. The problem becomes more intricate given the different typefaces. In addition, complex and unknown degradation make the data-driven convolutional neural networks (CNNs) incapable of generalizing well to real-world degraded observations.
To ameliorate the difficulties in blind text SR, especially in retaining the unique structure of each character, recent studies design new loss functions to constrain SR results semantically close to the high-resolution (HR) ground-truth , or incorporate the recognition prior into the intermediate features . Figure 1 shows several representative results of these two types of methods, i.e., TBSRN that applies content-aware constraint and TATT that embeds the recognition prior into the SR process. The results are not satisfactory despite the models were retrained with data synthesized with sophisticated degradation models, i.e., BSRGAN and Real-ESRGAN . TBSRN and TATT perform well when the LR character has fewer strokes. However, they both fail to generate the desired structures when the character is composed of complex strokes, e.g., the 4, 6, 8, 11-th characters in (a). The results suggest that high-level recognition prior cannot well benefit the SR task in restoring faithful and accurate structures.
When dealing with complex characters, we can identify some commonly used strokes as constituents of characters or their subparts. These obvious structural regularities have not been studied thoroughly in the literature. In particular, most existing text SR studies focus on scene text images that are presented in different view perspectives or arranged in irregular patterns. In contrast, our goal is to explore structural regularities specific to complex characters and investigate an effective way to exploit the prior. The investigation has practical values in many applications, including the restoration of old printed media (e.g., newspapers or books) for digitalization and preservation, and the enhancement of captions in digital media.
The key idea of our solution is to mine structure prior from a generative network, and use the prior to facilitate blind SR of text images. To capture the structure prior of characters, we train a StyleGAN to generate characters of different styles. As our intent is not to generate infinite and non-repetitive results, we restrict the generative space of StyleGAN to obey the structure of characters by constraining the generation with a codebook. Specifically, the codebook stores the discrete code of each character, and each code serves as a constant to StyleGAN for generating a specific high-resolution character. The font style is governed by the latent space of StyleGAN. Once the character StyleGAN is trained, we can extract its intermediate features serving as the generative structure prior, which encapsulates the intrinsic structure and style of the LR character.
Using the prior for text SR can be conveniently achieved through an additional Transformer-based encoder and a text SR network that accepts prior. In particular, given a text image that constitutes multiple characters (e.g., Figure 1), we use a Transformer-based encoder to predict the font style, character bounding boxes, and their respective indexes in the codebook jointly. The codebook indexes drive the structure prior generation for each character. In the text SR network, each LR character is super-resolved, guided by the respective priors that are aligned using their bounding boxes.
Experiments demonstrate that our design can generate robust and consistent structure priors for diverse and severely degraded text input. With the embedded structure prior, our approach performs superior against state-of-the-art methods in generating photo-realistic text results and shows excellent generalization to real-world LR text images. While our study mainly employs Chinese characters as examples, the proposed method can be readily extended to characters of other languages. We call our approach as MARCONet. The main contributions are summarized as follows:
We show that blind SR task, especially for characters with complex structures, can be restored by using their structure prior encapsulated in a generative network.
We present a viable way of learning such generative structure prior through reformulating a StyleGAN by replacing its single constant with discrete codes that represent different characters.
To retrieve the prior accurately, we propose a Transformer-based encoder to jointly predict the font styles, character bounding boxes and their indexes in codebook from the LR input.
Related Work
Blind Image SR. Blind image SR is challenging due to the complex mixture of unknown degradations. Recent studies address the problem from two aspects, i.e., degradation estimation and establishing more realistic training data . The former paradigm mainly focuses on estimating degradation model parameters and then applies non-blind SR methods, e.g., ZSSR . The latter builds training pairs either through capturing real-world LR and HR pairs or designing elaborate degradation models that imitate real-world degradation . Since characters in text images have specific and semantic structures, we show that a good restoration performance cannot be achieved by merely using elaborately designed degradation models.
Text Image SR. Text image SR has been studied for many years. In traditional methods, maximum a posterior (MAP) and Bayesian framework are exploited for super-resolving text images. These earlier approaches are incapable of generating high-quality results. Dong et al. adopt CNNs (i.e., SRCNN ) for text image SR and achieve promising results in the ICDAR 2015 competition . Xu et al. adopt a Generative Adversarial Network (GAN) to learn category-specific prior for face and text images SR, along with the supervision from a multi-class GAN loss. Mou et al. propose to plug a SR unit into the recognition process of degraded text images. Wang et al. introduce the first real-world text SR pairs (i.e., TextZoom), which are cropped from RealSR and SRRAW . They also present a sequential residual block by incorporating Bi-directional LSTM to capture the sequential information for low-level reconstruction. Both RealSR and SRRAW are captured by different digital cameras with different focal lengths, aiming to collect natural LR/HR pairs in real-world scenarios. The setting is not designed specifically for text images. We observe that most HR text images in TextZoom are limited in clarity, which could limit the SR performance when taking them as ground-truth.
Text SR benefits from prior and auxiliary constraints. Quan et al. recover text images in a cascade model by predicting the high-frequency information. Chen et al. propose Transformer-based position-aware and content-aware modules to emphasize the position and the content of each character. Analogously, Nakaune et al. and Qin et al. respectively present the structure-aware loss and content-perceptual constraint for learning detailed structural skeletons. Zhao et al. propose a parallel contextual attention network to learn sequence-dependent features for bringing more high-frequency information into text reconstruction. Zhao et al. exploit linguistic, recognition and visual clues to jointly boost the SR performance. Chen et al. propose a stroke-aware framework by concentrating on stroke-level internal structures. Ma et al. introduce text recognition prior to text reconstruction with a Transformer-based module, leveraging the global attention mechanism. They also embed categorical text priors in the encoder and employ multi-stage refinement to progressively enhance low-resolution images .
Many methods above employ recognition information either as a loss function on the SR results or as intermediate SR features for providing high-level guidance . Albeit the high-level recognition prior helps improve the text recognition ability, it is limited in providing accurate structure and style guidance, especially for certain texts with complex structures. In this study, we show that generative structure prior benefits high-quality guidance to restore the faithful structure of LR characters.
Generative Structure Prior in Image SR. Image structure prior is proven effective in many low-level vision tasks, e.g., depth image enhancement , image inpainting , and image restoration . Most recently, by using generative structure priors obtained from pre-trained StyleGANs or codebooks , blind face restoration has achieved tremendous improvement , suggesting the apparent advantage of such prior over other methods in generating photo-realistic textures. Our study is inspired by the success of these approaches. Nevertheless, crafting a suitable structural prior for text images are harder than face images. This is because each character has its unique strokes, but may have a vast variety of font styles. Any distorted, missing or additional strokes easily change their semantic layout that is easily perceptible, and worsens their actual meaning (see Figure 1). All of these challenges aggravate the difficulties for learning a generative structure prior for the text SR task. In this study, we show an effective way of learning such prior through replacing the constant input of StyleGAN with a discrete codebook, while controlling the font style via the space.
Methodology
A GAN model can well capture the intrinsic structure prior for a category by training the model on abundant images of the same category. Many image restoration tasks have shown the benefits of using such generative priors for restoring photo-realistic details from LR inputs despite severe and complex degradation. Previous studies mainly use the prior for face restoration, while its application to text SR is under-explored. In this paper, we propose MARCONet, the first attempt to eMbed generAtive stRuCture priOr for blind text SR. The proposed pipeline is depicted in Figure 3. It mainly contains three parts, 1) the prediction of font style, character bounding boxes and their indexes in codebook based on a given LR input, 2) the generation of the structure prior for each character, 3) the SR framework that takes the generative structure prior as guidance for restoration. Next, we first describe the approach to obtain the generative structure prior for each character, followed by the way to employ such prior for the text SR.
The original StyleGANs take a learnable constant as input, and control the style of output image through the space, which is a projection from the initial latent space . Layer-wise noise is also introduced to support stochastic variation on fine-grained details. To better capture the structure of text images, we remove the layer-wise noise and replace the single constant with discrete codes that represent different characters (see details in Figure 2).
We denote each character code as , where is the codebook that stores all the character features. Each code is learnable, with a size of . Following CRNN , the cardinality of codebook is set to 6736, a size covering simplified Chinese characters, English letters, and numbers. The retrofitted StyleGAN is defined as:
where represents the network that maps to , and is the model parameters of StyleGAN. For generalizing to different scenarios, we simplify the structure image with pixel values (see the output in Figure 2). The structure prior used in this work is the intermediate features from in Eqn. (1), and can be formulated as:
where represents the output features from -th layer of .
The training of the character StyleGAN can be done on synthetic images yet with satisfying generalization ability to real-world texts. In particular, the PIL packagehttps://python-pillow.org/ is adopted to synthesize high-resolution character images with hundreds of font styles (see examples in the suppl.). We also augment the diversity with random translations, font sizes and slight affine transformation. Different from the original StyleGANs that adopt only adversarial loss , we introduce an additional recognition loss derived from a pre-trained Transformer-based recognizer as regularization. Once learned, each learnable code well captures the distinctive features of each character and can control the font styles of the output image. In the following, we will describe the approach to extract the structure prior for each character in a given LR text image.
2 Transformer Encoder
A LR text image usually is composed of several characters. To derive the generative structure prior for each character, first, we need to obtain the font style of the LR characters. Second, the indexes of each character in the codebook and their bounding boxes are also necessary for aligning the structure prior with each LR character. To this end, we adopt a Transformer-based encoder to jointly predict the needed information. Transformer is chosen so that we can better capture the dependency across different characters in the input image.
Learning to predict the font style is similar to learning-based GAN inversion , where the Transformer network serves as an encoder. It is observed that the Transformer-based encoder achieves satisfactory inversion, achieving reconstruction with low distortion and of high quality . Note that characters in one text image generally have the same font style. So we use one linear layer to map the features of all characters to a single prediction (), which is shared with all the characters in the same LR image. We also use a similar structure as prediction to carry out the regression and classification of the bounding boxes and code indexes. These three sub-tasks share the same Transformer encoder (i.e., ViT ) and a CNN backbone (i.e., ResNet45 ). Positional encoding is incorporated to inject positions of features in the sequence.
The font style prediction branch is optimized with the gradient from StyleGAN. Since the input text image has unconstrained numbers of characters, we adopt CTC loss for recognition. The loss is proposed to solve the alignment between predicted and target labels and is widely used in scene text recognition tasks. As for the learning of character bounding boxes regression, we adopt a linear combination of the Smooth loss and the generalized IoU loss . The ground-truth boxes of each character can be obtained when synthesizing the training images with PIL package. The gradient from the latter text SR network can further benefit the learning process of the whole Transformer encoder.
3 Text Super-resolution Network
Network Architecture. With the Transformer encoder, we can obtain the font style , classification label (i.e., index in codebook), and the bounding box for each character in the LR input. Then, the corresponding high-quality generative structure prior for each character can be generated with and the selected code through the pre-trained StyleGAN in Eqn. (2) (see the character structure prior in Figure 3). In this section, we describe how to use them for the text SR task. First, a simple UNet is stacked upon the ResNet45 to extract the LR features. Then, the generative structure prior of each character is embedded into their LR characters through a structure prior transform module. The process is shown in Figure 4 (more details can be found in the suppl.). For the LR features, we adopt the detected bounding boxes to crop and align each LR character by RoIAlign operation . For each character, AdaIN is adopted to normalize the prior’s distribution, and spatial feature transformation is subsequently employed to predict the affine parameters, which are applied on the LR character features. The reverse RoIAlign is finally used to paste the enhanced features back to their original locations. The structure prior transform module is adopted at two scales , allowing our MARCONet to remain high fidelity with different degradation. A Conv-ReLU-ResBlock based CNN module is stacked to generate the final SR result.
Learning Objective. We minimize the differences between the SR results and HR ground-truth on both the pixel and perceptual domains :
where , and are the feature dimensions from the -th convolution layer of the pre-trained VGG-19 model . is the trade-off parameter and is set to 0.05. The loss is applied to the whole text image.
Adversarial loss is also added to improve visual quality. Instead of constraining on the whole image as in Eqn. (3), is performed on each cropped character image together with its corresponding structure image as additional conditions . To be specific, the concatenation of HR character image and its structure image is expected to be classified as Real, while the concatenation of SR character image and its structure image is recognized as Fake. Note that is the binarized version of , which represents the ground-truth structure, and is obtained from Eqn. (1) with and predicted from LR input. Such design allows us to better constrain the embedding of generative structure prior into the SR results. The hinge version of adversarial loss for each character is defined as:
Finally, we use the loss on structure image in the text SR training stage to further fine-tune and code in codebook. For each character, this constraint is defined as:
With these learning objectives, the whole framework of MARCONet is optimized in an end-to-end manner.
Experiments
Training Data. In this work, we take Chinese characters as our initial exploration for super-resolving complex characters but with regularity. The same method can be extended to other languages by retraining the framework with specific data. Since a text image is usually composed of several characters, here we use the Chinese corpus from Xu et al. , which contains tens of millions of common items. We also collect 182 font families to introduce diverse structures. PIL toolbox is adopted to synthesize these text images, rendered with random RGB values, font sizes and locations. The text background is obtained from DIV2K and Flick2K datasets, in which each image is randomly cropped and upsampled to to synthesize the complex background. With this toolbox, we can obtain the HR text images , together with the classification label, bounding box and the ground-truth structure image of each character. The degradation pipeline presented in BSRGAN and Real-ESRGAN are applied online to degrade the HR image to LR input (see input in Figure 3). Since each text image usually contains varying numbers of characters, we pad zeros to the width dimension to keep the structure unchanged.
Implementation Details. The training of the whole MARCONet and the pre-training for generative structure prior are conducted on a server with four Tesla V100 GPUs. The batch size is set to 2 and 16, respectively. We employ Adam as the optimizer. The initial learning rate is set to and decreased by 0.5 when reaches a stable range on the validation set. The for StyleGAN in the text SR training stage is set to for fine-tuning. The height of HR images is set to 128. Color jittering is used to increase image diversity. It takes two days for pre-training StyleGAN and nearly four days for training the whole MARCONet.
Baselines. Since there are few works studying the problem of these types of text images on SR task, we compare with two types of methods following TATT , i.e., general image SR (i.e., SRCNN , ESRGAN ) and recent text image SR (i.e., TSRN , TBSRN and TATT ). For a fair comparison, we modify their codes to handle and SR tasks, and carefully retrain them with the same training data as Ours. In particular, the retrained TBSRN adopts the ground-truth text and bounding boxes for its pre-trained content-aware and position-aware modules. As for TATT, we also use the ground-truth text to pre-train its TPG module. We report the quantitative and qualitative results on our synthetic dataset, which contains 1,000 LR/HR pairs for and tasks, respectively. These LR inputs are injected with random noise, blurring and JPEG compression, to simulate real-world degradation. We also provide the results on real-world text images collected from different sources. More results can be found in the suppl.
Table 1 summarizes the PSNR, SSIM and LPIPS of different methods on the synthetic dataset. An additional pre-trained CRNN is used to compute the recognition accuracy of SR results (Acc.), serving as an auxiliary indicator for the restoration performance. Note that the synthetic data is non-trivial given its low quality. Without specifically designed components for text images, general SR methods perform poorly as expected. By incorporating the recognition information, both TBSRN and TATT show obvious improvement, especially in retaining the semantic layout of character for recognition (see Acc. columns). Our method achieves the best quantitative results and outperforms others with a large margin (e.g., 2.5 dB higher than the second-best method on average). Even with severe degradation (i.e., ), our method still attains favorable performance, which can be ascribed to the effectiveness of the learned generative structure prior for each character.
2 Qualitative Comparison
Figure 5 shows the visual comparison on our synthetic LR images for and tasks. With careful retraining, competing methods perform favorably when the LR input has slight degradation, but they are not able to recover the unique structure of characters with complex strokes. When the degradation turns more severe, all of them fail to restore satisfactory results (see the last column). The results yielded by TBSRN and TATT suggest that high-level recognition constraints bring limited benefit to the SR task, especially in coping with diverse and complex font structures and styles. With the guidance of characters’ generative structure prior, our method shows compelling results that exhibit more consistent structures with ground-truth (GT). It is interesting to observe the strong performance of MARCONet in restoring heavily degraded LR input that is challenging for humans to recognize (last column of Figure 5).
Apart from the evaluation on synthetic LR input, we also evaluate our method on real-world LR text images. We collect real-world LR text images from different sources, e.g., invoices, license plates, scanned documents, road signs, plaques, video captions, and scene texts from TextZoom . Each text region is cropped by CnSTDhttps://github.com/breezedeus/CnSTD. We compare with two competing methods, i.e., TBSRN and TATT , which are specifically designed for text images and achieve the top quantitative results in Table 1. The results are shown in Figure 6. Benefiting from our synthetic training data, the competing methods achieve comparable performance on these characters with simple structures, even for LR inputs corrupted by unknown and complex degradation. The positive results suggest that our synthetic HR text images are of high quality and well-suited as ground-truth for the SR task. The usage of degradation models from BSRGAN and Real-ESRGAN further benefits these methods in synthesizing realistic LR inputs. However, both of these two methods fail to generate plausible results when the LR character is degraded severely or contains complex structures. In contrast, with the proposed generative structure prior, our method achieves the best performance, not only in visual quality, but also in the generation of accurate structures.
3 Ablation Study
Analyses of Space. The space controls the style of a character’s structure prior. Figure 7 shows the structure image from our StyleGAN with and character indexes predicted from the LR input. One can observe that can well capture the manifold of font styles, e.g., the thickness of different font families (1-st column) and different font sizes and locations (2-nd column). We also interpolate two from different LR inputs and demonstrate their results in the rows. One can see that the generated structures change smoothly, indicating that in our framework also has the editable ability like the original StyleGAN. Finally, we use the predicted from a real-world LR video frame (The Dream of the Red Chamber) and explore font style transfer. Even though the LR input contains a traditional Chinese character and the font family does not exist in our training data, our predicted can well capture its style and successfully transfer it to other characters. Besides, the latent and character code are disentangled, judging from the good separation of roles between the and in style control and storing the unique structure for each character.
Analyses of Variants. Here we consider the following variants to evaluate each part of our MARCONet, 1) Ours (UNet) & Ours (UNet†): only taking ResNet45 and UNet to perform the SR task and increasing its model parameters, respectively, 2) Ours (CRNN) & Ours (PSP): replacing our Transformer-based character classification and inversion branches with CRNN and PSP encoder , respectively, 3) Ours (D): only taking the SR and HR images as input of discriminator without concatenating their structure images, 4) Ours (w/o S): removing StyleGAN and directly incorporating the code on each LR character features through spatial affine transformation, 5) Ours (w/o C): removing codebook and directly taking each LR character features into a pre-trained StyleGAN that produces colored output like GFPGAN , 6) Ours (#32) & Ours (#64): only using the structure prior transform module on feature sizes of either or . Their performance on our synthetic test set is shown in Table 2 and Figure 8.
We can observe that a) Ours (UNet) and Ours (UNet†) perform on par with general SR methods, and indicate that a mere increase in model capacity would not bring significant performance improvement; b) the Transformer-based encoder is more effective than CRNN and PSP encoder , which may be caused by the global attention and thus contribute better to the classification and GAN inversion tasks. It is observed that Ours (CRNN) sometimes makes wrong predictions on character. While the SR leads to a clear image, the character loses its actual semantic layout (see red box in Figure 8); c) the concatenation of structure image on discriminator is beneficial for emphasizing the structure prior more on LR input and thus boosting the final SR result; d) by removing StyleGAN and codebook, Ours (w/o S) and Ours (w/o C) show obvious inferior results, indicating that codebook can benefit the recognition accuracy while SyleGAN contributes more on the visual quality (e.g., LPIPS); e) Ours (#64) has slightly better results than Ours (#32) but both of them are inferior to Ours (Full) that deploys a multi-scale structure prior transform module. From these analyses, we conclude that the Transformer-based encoder, codebook, StyleGAN and multi-scale structure prior transform module are all crucial for achieving good performance in Ours (Full).
Conclusion
In this work, we made the first attempt to embed the generative structure prior for blind SR of text images. The combination of a codebook for storing distinctive character-specific codes and a retrofitted StyleGAN for controlling font style cope well with complicated structures, high similarity between characters, and a great variety of font styles. We have shown that such structure prior is beneficial for super-resolving LR text images, even for those with severe degradation. We can potentially extend it to other text-related tasks, e.g., text image completion for ancient documents, font style transformation, and few-shot font generation.
Acknowledgement. This study is supported under the RIE2020 Industry Alignment Fund – Industry Collaboration Projects (IAF-ICP) Funding Initiative, as well as cash and in-kind contribution from the industry partner(s).