Text Prior Guided Scene Text Image Super-resolution

Jianqi Ma, Shi Guo, Lei Zhang

Introduction

Scene text image recognition aims to recognize the text characters from the input image, which is an important computer vision task that involves text information processing. It has been widely used in text retrieval , sign recognition , license plate recognition and other scene-text-based image understanding tasks . However, due to the various issues such as low sensor resolution, blurring, poor illumination, etc., the quality of captured scene text images may not be good enough, which brings many difficulties to scene text recognition in practice. In particular, scene text recognition from low-resolution (LR) images remains a challenging problem.

In recent years, single image super-resolution (SISR) techniques have achieved significant progress owing to the rapid development of deep neural networks . Inspired by the success of SISR, researchers have started to investigate scene text image super-resolution (STISR) to improve the quality of LR text images and hence improve the text recognition accuracy. Tran et al. adapted LapSRN to STISR and significantly improved the text content details. To obtain more realistic STISR results, Bílková et al. and Wang et al. employed GAN based networks with CTC loss and text perceptual losses. In these methods, the LR images are synthesized (e.g., bicubic down-sampling) from high resolution (HR) images for SR model learning, while the image degradation process in real-world LR images can be much more complex. Recently, Wang et al. collected a real-world STISR dataset, namely TextZoom, where LR-HR image pairs captured by zooming lens are provided. Wang et al. also proposed a TSRN model for STISR, achieving state-of-the-art performance .

The existing STISR methods , however, mostly treat scene text images as natural scene images to perform super-resolution, ignoring the important categorical information brought by the text content in the image. As shown in Fig. 1(b), the result by TSRN is much better than the simple bicubic model (Fig. 1(a)), but it is still hard to tell the characters therein. Based on the observation that semantic information can help to recover the shape and texture of objects , in this paper we propose a new STISR method, namely text prior guided super-resolution (TPGSR), by introducing the categorical text prior information into the model learning process. Unlike the face segmentation prior used in and the semantic segmentation used in , the text character segmentation is hard to obtain since there are few datasets containing annotations of fine character segmentation masks. We instead employ a text recognition model (e.g., CRNN ) to extract the probability sequence of the text as the categorical prior of the given LR scene text image, and embed it into the super-resolution network learning process to guide the reconstruction of HR image. As can be seen in Fig. 1(c), the text prior information can indeed improve much the STISR results, making the text characters much more readable. On the other hand, the reconstructed HR text image can be used to refine the text prior, and consequently a multi-stage TPGSR framework can be built for effective STISR. Fig. 1(d) shows the super-resolved text image by using the refined text prior, where the text can be clearly and correctly recognized. The major contributions of our work are as follows:

We introduce the text recognition categorical probability sequence as the prior into the STISR task, and validate its effectiveness to improve the visual quality of text content and recognition accuracy of scene text images.

We propose the refinements of the text recognition categorical prior without using extra supervision of text labels except for the real HR image: refining recurrently by the estimated HR image and by fine-tuning the text prior generator with our proposed TP Loss. With such refinement, the text prior and the super-resolved text image can be jointly enhanced under our TPGSR framework.

By improving the image quality of LR text images, the proposed TPGSR improves the text recognition performance on TextZoom for different text recognition models by a large margin and demonstrates good generalization performance to other recognition datasets.

Related Works

Single Image Super Resolution (SISR). Aiming at estimating a high-resolution (HR) image from its low-resolution (LR) counterpart, SISR is a highly ill-posed problem. In the past, handcrafted image priors are commonly used to regularize the SISR model to improve the image quality. In recent years, the training of deep neural networks (DNNs) has dominated the research of SISR. The pioneer work SRCNN learns a three-layer convolutional neural network (CNN) for the SISR task. Later on, many deeper CNN models have been proposed to improve the SISR performance, e.g., deep residual block , Laplacian pyramid structure , densely connected network and channel attention mechanism . The PSNR and SSIM losses are widely used in those works to train the SISR model. In order to produce perceptually-realistic SISR results, SRGAN employs a generative adversarial network (GAN) to synthesize image details. SFT-GAN and FSRNet utilizes the GAN loss and semantic segmentation to generate visually pleasing HR image.

Scene Text Image Super Resolution (STISR). Different from the general purpose SISR that works on natural scene images, STISR focuses on text images, aiming to improve the readability of texts by improving their visual quality. Intuitively, those methods for SISR can be directly adopted for STISR. In , Dong et al. extended SRCNN to text images, and obtained the best performance in ICDAR 2015 competition . PlugNet employs a light-weight pluggable super-resolution unit to deal with LR images in feature domain. TextSR utilizes the text recognition loss and text perceptual loss to generate the desired HR images for text recognition. To improve the performance of STISR on real-world scene text images, Wang et al. built a real-world STISR image dataset, namely TextZoom, where the LR and HR text image pairs were cropped from real-world SISR datasets . They also proposed a TSRN method to use the central alignment module and sequential residual block to exploit the semantic information in internal features. SCGAN employs a multi-class GAN loss as supervision to equip the model with ability to generate more distinguishable face and text images. Further by progressively adopting the high-frequency information derived from the text image, Quan et al. proposed a multi-stage model for recovering the blurry text images. Different from the above methods, we propose employ the text recognition prior to guide the STISR model recovering better quality text images.

Scene Text Recognition. In the early stage of deep learning based scene text recognition, researchers intended to solve the problem in a bottom-up manner , i.e., extracting the text from characters into words. Some other approaches recognize the text in a top-down fashion , i.e., regarding the text image as a whole and performing a word-level classification. Taking text recognition as an image-to-sequence problem, CRNN employs the CNN to extract image feature and uses the recurrent neural networks to model the semantic information of image features. Trained with the CTC loss, the predicted sequence can be more accurately aligned with the target sequence . Recently, attention-based methods thrive due to the improvement in text recognition benchmarks and the robustness to various shapes of text images . In our method, we adopt CRNN as the text prior generator to generate categorical text priors for STISR model training. It is shown that such text priors can significantly improve the perceptual quality of super-resolved text images and consequently boost the text recognition performance.

Methodology

In this section, we will first explain what the text prior (TP) is, and then introduce the text prior guided super-resolution (TPGSR) network in detail, followed by the design of loss function.

In this paper, the TP is defined as the deep categorical representation of a scene text image generated by some text recognition models. The TP is then used as guidance information to encourage our TPGSR network to produce high quality scene text images, which are favorable to both visual perception and scene text recognition.

Specifically, we choose the classic text recognition model CRNN to be the TP Generator. CRNN uses several convolution layers to extract the text features and five max pooling layers to down-sample the features into a feature sequence. The TP is then defined as the categorical probability prediction by CRNN, which is a sequence of ∣A∣|A|-dimensional probability vectors where ∣A∣|A| denotes the number of characters learned by CRNN. Fig. 2 visualizes the TP of some scene text images, where the horizontal axis represents the sequence in left-to-right order and the vertical axis represents the categories in reverse alphabet order (e.g., ’Z’ to ’A’). In the visualization, the lighter the spot is, the higher the probability of this category will be. By using the TP as guidance, our TPGSR model can recover visually more pleasing HR images with higher text recognition accuracy, as we illustrated in Fig. 1.

2 The Architecture of TPGSR

By introducing TP into the STISR process, the main architecture of our TPGSR is illustrated in Fig. 3. Our TPGSR network has two branches: a TP generation branch and a super-resolution (SR) branch. First, the TP branch intakes the LR image to generate the TP feature. Then, the SR branch intakes the LR image and the TP feature to estimate the HR image. In the following, we introduce the two branches in detail.

TP generation branch. This branch uses the input LR image to generate the TP feature and passes the TP feature to the SR branch as guidance for more accurate STISR results. The key component of this branch is the TP Module, which consists of a TP Generator and a TP Transformer. As mentioned in Section 3.1, the TP generated by TP Generator is a probability sequence, whose size may not match the image feature map in the SR branch. To solve this problem, we employ a TP Transformer to transform the TP sequence into a feature map.

Specifically, the input LR image is first resized to match the input of TP Generator by bicubic interpolator, and then passed to the TP Generator to generate a TP matrix whose width is LL. Each position of the TP is a vector of size ∣A∣|A|, which is the number of categories of alphabet AA adopted in CRNN. To align the size of TP feature with the size of image feature, we pass the TP feature to the TP Transformer. The TP Transformer consists of 4 Deconv blocks, each of which consists of a deconvolution layer, a BN layer and a ReLU layer. For an input TP matrix with width LL and height ∣A∣|A|, the output of TP Transformer will be a feature map with recovered spatial dimension and channel 3232 after three deconvolution layers with stride (2,2)(2,2) and one deconvolution layer with stride (2,1)(2,1). The kernel size of the all deconvolution layers is 3×33\times 3.

SR branch The SR branch aims to reproduce an HR text image from the input LR image and TP guidance feature. It mainly contains an SR Module. Please note that most of the SR blocks in existing SISR and STISR methods can be adopted as our SR Module in couple with our TP guidance features. Considering that these SR blocks, such as the residual block in SRResNet and the SRB in TSRN , only take the image features as input, we need to modify them in order to embed the TP features. We call our modified SR blocks as TP-Guided SR Blocks.

The difference between previous SR Blocks and our TP-Guided SR Block is illustrated in Fig. 4. To embed the TP features into the SR Block, we concatenate them to the image features along the channel dimension. Before the concatenation, we align the spatial size of TP features to that of the image features by bicubic interpolation. Suppose that the channel number of image features is CC, then the concatenated features of C+32C+32 channels will go through a projection layer to reduce the channel number back to CC. We simply use a 1×11\times 1 kernel convolution to perform this projection. The output of projection layer is fused with the input image feature by addition. With several such TP-Guided SR Blocks, the SR branch will output the estimated HR image, as in those previous super-resolution models.

3 Multi-stage Refinement

With the TPGSR framework described in Section 3.2, we can super-resolve an LR image to a better quality HR image with the help of TP features extracted from the LR input. One intuitive question is, if we extract the TP features from the super-resolved HR image, can we use those better quality TP features to further improve the super-resolution results? Actually, multi-stage refinement has been widely adopted in many computer vision tasks such as object detection and instance segmentation to improve the prediction quality progressively. Therefore, we extend our one-stage TPGSR model to a multi-stage learning framework by passing the estimated HR text image in one stage to the TP Generator in next stage. The multi-stage TPGSR framework is illustrated in Fig. 5. In the 1st1^{\text{st}} stage, the TP Module accepts the bicubically interpolated LR image as input, while in the following stages, the TP Module accepts the HR image output from the SR Module in previous stage as input for refinement. As we will show in the ablation study in Section 4.3, both the quality of estimated HR text image and the text recognition accuracy can be progressively improved by this multi-stage refinement.

4 Training Loss

There are two types of loss functions in our TPGSR framework, one for the SR branch and another one for the TP generation branch. For the SR branch, the loss is similar to that in many previous SISR methods . Denote by I^H\hat{I}_{H} the estimate HR image from the LR input and by IHI_{H} the ground-truth HR image, the loss for the SR branch, denoted by LSL_{S}, can be commonly defined as the L1L_{1} norm distance between I^H\hat{I}_{H} and IHI_{H}, i.e., LS=∣I^H−IH∣L_{S}=|\hat{I}_{H}-I_{H}|.

Different from the many SISR works as well as the STISR works , in TPGSR we have loss functions specifically designed for the TP generation branch, which is crucial to improve the text image quality and text recognition. The TP sequence generated by the TP Generator has significant impact on the final SR results. If the TP is correct (i.e., identical to the TP of ground-truth image), it will bring positive impact on the estimated HR image, as we shown in Fig. 1. Otherwise, if the categorical information in the extracted TP is incorrect, the SR result can be much damaged (please refer to Section 4.5 for such failure cases). Therefore, we use the TP extracted from the ground-truth HR image to supervise the learning of our TPGSR network. Denote by tLt_{L} and tHt_{H} the TP extracted from LR image ILI_{L} and HR ground-truth image IHI_{H}, respectively, we use the L1L_{1} norm distance ∣tH−tL∣|t_{H}-t_{L}| and the KL divergence DKL(tL∣∣tH)D_{KL}(t_{L}||t_{H}) to measure the similarity between tLt_{L} and tHt_{H}. With the text priors tL,tH∈RL×∣A∣t_{L},t_{H}\in R^{L\times|A|} of the pair of LR and HR images, the DKL(tL∣∣tH)D_{KL}(t_{L}||t_{H}) can be calculated as follows:

where tLijt_{L}^{ij} and tHijt_{H}^{ij} denote the element in iith position and jjth dimension in tLt_{L} and tHt_{H}. ϵ\epsilon is a small positive number to avoid numeric error in division and logarithm. Together with LSL_{S}, the overall loss function for a single-stage TPGSR can is follows:

where α\alpha and β\beta are the balancing parameters.

For the multi-stage TPGSR learning, the loss for each stage, denote by LiL_{i}, can be similarly defined as in Eq. 2. Suppose there are NN stages in total, the overall loss is defined as follows:

where λi\lambda_{i} balances the loss of each stage and ∑i=1Nλi=1\sum_{i=1}^{N}\lambda_{i}=1.

Experiments

We implement our TPGSR method in Pytorch. Adam is selected as our optimizer with momentum 0.90.9. The batch size is set to 4848 and the model is trained for 500500 epochs with one NVIDIA RTX 2080Ti GPU. The TP Generator is selected as CRNN pre-trained on SynthText and MJSynth . In Eq. 2, the weights α\alpha and β\beta are both simply set to 11, while the ϵ\epsilon in Eq. 1 is set to 10−610^{-6}. The alphabet set AA includes mainly alphanumeric ( to 99 and ’a’ to ’z’) case-insensitive characters. Together with a blank label, ∣A∣|A| (i.e., the size of AA), has 3737 categories in total. For dealing with the out-of-category cases, we assign all out-of-category characters with blank label, and the reconstruction of these characters mainly depends on the SR Module.

For multi-stage TPGSR training, we adopt a well-trained single-stage model to initialize all stages and cut the gradient across stages to speed up the training process to converge. The TP Generator are non-shared while the SR Module are shared cross stages. As in previous multi-stage learning methods , higher weight is assigned to the loss on the last stage, and the other stages are assigned with smaller weights on loss. In particular, we use a 3-stage TPGSR. The parameters λi\lambda_{i} in Eq. 3 are set as λ1=14\lambda_{1}=\frac{1}{4}, λ2=14\lambda_{2}=\frac{1}{4} and λ3=12\lambda_{3}=\frac{1}{2}.

2 Datasets and Experiment Settings

Datasets. The TextZoom , ICDAR2015 and SVT datasets are used to validate the effectiveness of our proposed TPGSR method. TextZoom consists of 21,74021,740 LR-HR text image pairs collected by lens zooming of the camera in real-world scenarios. The training set has 17,36717,367 pairs, while the test set is divided into three subsets based on the camera focal length, namely easy (1,6191,619 samples), medium (1,4111,411 samples) and hard (1,3431,343 samples). The dataset also provides the text label for each pair.

ICDAR2015 is a well-known scene text recognition dataset, which contains 2,0772,077 cropped text images from street view photos for testing. SVT is also a scene text recognition dataset, which contains 647647 testing text images. Each image has a 5050-word lexicon with it.

Experiment settings. Since there are real-world LR-HR image pairs in the TextZoom dataset, we first use it to train and evaluate the proposed TPGSR model. We then apply the trained model to ICDAR2015/SVT to test its generalization performance to other datasets. Considering the fact that most of the images in ICDAR2015 and SVT have good resolution and quality, while the TextZoom training data focus on LR images, we perform the generalization test only to the low quality images in ICDAR2015/SVT whose height is less than 1616 or the recognition score is less than 0.90.9.

3 Ablation Studies

To better understand the proposed TPGSR model, in this section we conduct a series of ablation experiments on the selection of parameters in loss function, the selection of number of stages and whether the TP Generator should be fine-tuned in training. We also perform experiments to validate the effectiveness of SR Module in our TPGSR framework. We adopt TSRN as the SR Module in the experiments, and name our model as TPGSR-TSRN. All ablation experiments are performed on TextZoom and the recognition accuracies are evaluated with CRNN .

Impact of tuning the TP Generator. The loss terms for the TP branch aim to fine-tune the TP Generator. To prove the significance of TP Generator tuning, we conduct experiments by fixing and tuning the TP Generator in a one-stage TPGSR model. The text recognition accuracies are shown in Table 1(a). By fixing the TP Generator, we can enhance the SR image recognition by 3.1%3.1\% compared to the TSRN baseline . By tuning the TP Generator during the training process, the recognition accuracy can be further improved from 44.5%44.5\% to 49.8%49.8\%, achieving a performance gain of 5.3%5.3\%. This clearly demonstrates the benefits of tuning the TP Generator to the SR text recognition task.

Impact of multiple stages in TPGSR. In addition to refining the TP Generator, recurrently inputting the estimated HR image into the TPGSR can also enhance the quality of TP since the SR Module can improve the estimated HR text image in each recurrence. To find out how well the multi-stage refinement can reach, we set the stage number N=1,2,…,5N=1,2,…,5 and report the text recognition accuracy in Table 1(b). We can see that the recognition accuracy increases with the increase of NN; however, the margin of improvement decreases with N. When N=5N=5, the accuracy of ’Hard’ split begins to fall. Considering the balance between the model size and the performance gain, we set N to 3 in our following experiments.

Parameter sharing strategy. To determine the best sharing strategies, we conduct experiments to test on both the TP Module and the SR Module. As shown in Table 1(c), we find that under different settings of stage number, the setting of non-shared TP Module shows significant performance improvement. However, when we use non-shared SR Module, little performance improvement in SR image recognition is achieved. Thus we use the settings of shared SR Module and non-shared TP Module in our multi-stage model.

The effectiveness of SR in TPGSR. Since one of the goals of STISR is to improve the text recognition performance by HR image recovery, it is necessary to check if the estimated SR images truly help the final text recognition task. To this end, we evaluate the TPGSR models with both fixed and tuned TP Generator by using LR and SR images as inputs. For multi-stage version, we test all the TP Generators and pick the best LR and SR results from all TP Generators. Note that models with tuned TP Generator and LR image as input is similar to directly fine-tuning the text recognition model on the LR images. The results are shown in Table 1(d). It can be seen that by tuning TP Generator on the LR images, the text recognition accuracy can be increased. However, the recognition accuracy can be improved more by using the SR text image. For example, at stage one, the recognition accuracy of LR images by using tuned TP Generator is 45.3%45.3\%, while the accuracy of SR images even without fine-tuning the TP Generator will be 49.8%49.8\%. If the tuned TP Generator is used to generate the SR text image, the text recognition performance can be further improved compared to the fixed TP Generator. The experiments and comparisons demonstrate the effectiveness of our SR Module in improving the final SR text recognition.

4 Comparison with State-of-the-Arts

As described in Section 3.2 and illustrated in Fig. 4, the SR block of most existing representative SISR and STISR models can be adopted in the SR Module of our TPGSR framework, resulting in a new TPGSR model. To verify the superiority of our TPGSR framework, we select several popular SISR models, including SRCNN , SRResNet , RDN , and the latest state-of-the-art STISR model TSRN , and embed their SR blocks into our TPGSR framework. The corresponding STISR models are called TPGSR-SRCNN, TPGSR-SRResNet, TPGSR-RDN, and TPGSR-TSRN, respectively. The TextZoom , ICDAR2015 and SVT datasets are used to compare these models as well as their prototypes. For fair comparison, all the models are trained trained in TextZoom dataset with the same settings.

Results on TextZoom. The experimental results on TextZoom are shown in Table 2. Here we present the text recognition accuracies on STISR results by using the official ASTER , MORAN and CRNN text recognition models. In Fig. 6, we visualize the SR images by the competing models with the ground-truth text labels. From Table 2 and Fig. 6, we can have the following findings.

First, from Table 2 we see that our TPGSR framework significantly improves the text recognition accuracy of all original SISR/STISR methods under all settings. This clearly validates the effectiveness of TP in guiding text image enhancement for recognition. Second, from Fig. 6 we can see that with TPGSR, all SR models show clear improvement in text image recovery with more readable character stroke, resulting in correct text recognition. This also explains why our TPGSR can improve significantly the text recognition accuracy, as shown in Table 2.

Generalization to other datasets. As mentioned in Section 4.2, to verify the generalization performance of our model trained on TextZoom to other datasets, we apply it to the low quality images (height ≤16\leq 16 or recognition score ≤0.9\leq 0.9) in ICDAR2015 and SVT. Overall, 563563 low quality images were selected from the 2,0772,077 testing images in ICDAR2015, and 104104 images were selected from the 647647 testing images in SVT. The STISR and text image recognition experiments are then performed on the 667667 low-quality images. Since TSRN is specially designed for text image SR and it performs much better than other SISR models, we only employ TSRN and TPGSR-TSRN in this experiment. The ASTER and CRNN text recognizers as well as stronger baseline SEED are used.

The results are shown in Table 3. We can have the following findings. First, compared with the text recognition results using original images without SR, TSRN improves the performance when CRNN is used the text recognizer, but decreases the performance when ASTER or SEED is used as the recognizer. This implies that TSRN does not have stable cross-dataset generalization capability. Second, TPGSR-TSRN can consistently improve the performance over the original images in all three recognizer. This demonstrates that it has good generalization performance on cross-dataset test. Third, TPGSR-TSRN consistently outperforms TSRN under all settings.

5 Discussions

Cost vs. performance. To further examine the value of our TPGSR, we compare the computational cost of our single-stage TPGSR with the TSRN . In Table 5, the experiments and results show that straightly increasing the number of SRB blocks is not an effective way of gaining performance (1.3%1.3\% accuracy drop). However, under our designed TPGSR network, the performance shows an improvement of 8.4%8.4\% compared to TSRN with 5-SRB. It is humble to conclude that generating text prior under our TPGSR framework is more valued than the additional cost it introduces.

The PSNR/SSIM indices for STISR. In terms of STISR, better results of some objective metrics, e.g., PSNR and SSIM, do not always guarantee more accurate scene text estimation, and vice versa. Similar conclusion can be also found in that PSNR and SSIM are not stable metrics for STISR. For better interpretation of this point, we adopt some popular SR models, including SRCNN , SRResNet , RDN and TSRN , and compare the results of these models before and after they are integrated into our TPGSR framework for joint optimization. The results are demonstrated in Table 4. We can see that the accuracy of all models boosts after they are embedded to our TPGSR framework. In terms of the objective metrics, our best SR model under TPGSR consistently outperforms the competing methods for samples at the “hard” difficulty level, since TP strengthens the power of the SR models in the challenging cases. In contrast, for samples at both “easy” and “medium” difficulty levels, joint optimization using our TPGSR does not always improve PSNR/SSIM results. This is because SR models trained without TP may suffer from over-fitting in easier cases. We could alleviate this issue by introducing TP loss between LR images and the predicted HR images into our objective function as a regularization term for joint optimization. In this way, as shown in Table 6, we can obtain better perceptual quality and more accurate scene text image recognition results, which we believe are more valuable in real world applications than subtle rise in metrics such as PSNR and SSIM.

Out-of-category analysis. As mentioned in Section 4.1, in our implementation, we assign the out-of-category characters with blank label. For such characters, the STISR results will mainly depend on the SR Module in our TPGSR network. To test the SR performance of our TPGSR model on images with out-of-category characters, we applied it to some text images in Korean, Chinese and Bangla picked from the ICDAR-MLT dataset. The results are shown in Fig. 7. We see that the reconstructed HR text images by our model show clearer appearance and contour than their LR counterparts. Compared with TSRN, TPGSR-TSRN demonstrates slightly better perceptual quality. The reason may be that categorical text prior serves as a regularization term in training SR Module to avoid over-fitting. Hence, the SR Module can still produce well-recovered scene text image with null guidance in inference stage.

Failure case. Though TPGSR can improve the visual quality of SR text images and boost the performance of text recognition, it still has some limitations, as shown in Fig. 8. First, the TP Generators tuned with LR samples are robust for most cases, while in some case it may produce false prior on the input LR image, and the output SR image may show incorrect character stroke, and causing false text recognition. Fig. 8(a) illustrates such a failure case. Second, if the text instance encounters certain rotation, the benefit brought by TPGSR will be weakened. Fig. 8(b) shows such an example. Though the recognition result is correct, the improvement on image quality is not significant. Third, for the cases with extremely long text instance, as shown in Fig. 8(c), the TP Generator may fail with significantly compressed outputs. In such case, the final output of TPGSR will suffer from text distortion with wrong text recognition.

To address the above issues, in the future we could consider to adopt more powerful TP Generators to provide more robust guidance, and design new TP guidance strategies for recovering multi-oriented scene text and curve text. In addition, the failures on long text can be alleviated by lengthening the width of input to the TP Generator.

Recovering hieroglyphs (e.g. Chinese). In this work, we focused on the real-world SR of English text images, for which there is a well-prepared benchmark dataset TextZoom. It is interesting to know whether our proposed method can be adopted for hieroglyphs such as Chinese. Here we perform some preliminary experiments to validate the feasibility. We train a multilingue recognition model using CRNN on the ICPR2018-MTWI Chinese and English dataset as our TP Generator. The overall alphabet contains 3,9653,965 characters, including the English and Chinese frequently-used set. Since there is no real-world benchmark dataset with LR-HR image pairs of Chinese characters, we synthesize LR-HR text image pairs by blurring and down-sampling the MTWI text images. We inherit the splits of MTWI as our training (59,88659,886 samples) and testing (4,8384,838 samples) set. The model training and testing are conducted following the settings described in Section 4.1.

The SR text recognition results are 27.7%27.7\% (Bicubic), 41.1%41.1\% (TSRN ), 42.7%42.7\% (TPGSR-TSRN) and 56.1%56.1\% (HR). From the results, we can observe that our TPGSR framework can still achieve 1.6%1.6\% accuracy gain over TSRN. Visualization on some Chinese characters can be seen in Fig. 9. One can see that our TPGSR can improve much the visual quality of SR results. Compared to TSRN , our TPGSR can better recover the text stroke on the samples. This preliminary experiment verifies that our TPGSR framework can be extended to hieroglyphs. More investigations and real-world dataset construction will be made in our future work.

Conclusion

In this paper, we presented a novel scene text image super-resolution framework, namely TPGSR, by introducing text prior (TP) to guide the text image super-resolution (SR) process. Considering the fact that text images have distinct text categorical information compared with those natural scene images, we integrated the TP features and image features to more effectively reconstruct the text characters. The enhanced text image can produce better TP in return, and therefore multi-stage TPGSR was employed to progressively improve the SR recovery of text images. Experiments on TextZoom benchmark and other datasets showed that TPGSR can clearly improve the visual quality and readability of low-resolution text images, especially for those hard cases, and consequently improve significantly the text recognition performance on them.

References