A Text Attention Network for Spatial Deformation Robust Scene Text Image Super-resolution

Jianqi Ma, Zhetong Liang, Lei Zhang

Introduction

The text in an image is an important source of information in our daily life, which can be extracted and interpreted for different purposes. However, scene text images often encounter various quality degradation during the imaging process, resulting in low resolution and blurry structures. This problem significantly impairs the performance of the downstream high-level recognition tasks, including scene text detection , optical character recognition (OCR) and scene text recognition . Thus, it is necessary to increase the resolution as well as enhance the visual quality of scene text images.

In the past few years, many scene text image super-resolution (STISR) methods have been developed to improve the image quality of text images, with notable progress obtained by deep-learning-based methods . By using a dataset of degraded and original text image pairs, a deep convolutional neural network (CNN) can be trained to super-resolve the text image. With strong expressive capability, CNNs can learn various priors from data and demonstrate much strong performance. A recent advance is the TPGSR model , where the semantics of the text are firstly recognized as prior information and then used to guide the text reconstruction process. With the high-level prior information, TPGSR can restore the semantically correct text image with compelling visual quality.

Despite the great progress, many CNN-based methods still have difficulty in dealing with spatially-deformed text images, including those with rotation and curved shape. Two examples are shown in Fig. 1, where the text in the left image has rotation and the right one has a curved shape. One can see that the current representative methods, including TSRN and TPGSR , produce blurry texts with semantically incorrect characters. This is because the architectures in current works mainly employ locality-based operations like convolution, which are not effective in capturing the large position variation caused by the deformations. In particular, the TPGSR model adopts a simplistic approach to utilize the text prior: it merely merges text prior with image feature by convolutions. This arrangement can only let the text prior interact with the image feature within a small local range, which limits the effect of text prior on the text reconstruction process. Based on the this observation, some globality-based operations (e.g., attention) should be employed to capture long range correlation in the text image for better STISR performance.

In this paper, we propose a novel architecture, termed Text ATTention network (TATT), for spatial deformation robust text super resolution. Similar to TPGSR, we first employ a text recognition module to recognize the character semantics as text prior (TP). Then we design a transformer-based module termed TP Interpreter to enforce global interaction between the text prior and the image feature. Specifically, the TP Interpreter operates cross attention between the text prior and the image feature to capture long-range correlation between them. The image feature can then receive rich semantic guidance in spite of the spatial deformation, leading to improved text reconstruction. To further refine the text appearance under spatial deformation, we design a text structure consistency loss, which measures the structural distance between the regular and deformed texts. As can be seen in Fig. 1, the characters recovered by our method show better visual quality with correct semantics.

Overall, our contributions can be summarized as follows:

We propose a novel method to align the text prior with the spatially-deformed text image for better SR recovery by using CNN and Transformer.

We propose a text structure consistency loss to enhance the robustness of text structure recovery from spatially-deformed low-resolution text images.

Our proposed model not only achieves state-of-the-art performance on the TextZoom dataset in various evaluation metrics, but also exhibits outstanding generalization performance in recovering orientation-distorted and curve-shaped low-resolution text images.

Related Works

Single image super resolution (SISR) aims at recovering a high-resolution (HR) image from a given low-resolution (LR) input image. The traditional methods design hand-crafted image priors for this task, including statistical prior , self-similarity prior and sparsity prior . The recent deep-learning-based methods train convolutional neural networks (CNNs) to address the SISR task and achieve leading performance. The seminal work SRCNN adopts a three-layer CNN to learn the SR recovery. Later on, more complex CNN architectures have been developed to upgrade the SISR performance, e.g., residual block , Laplacian pyramid , dense connections and channel attention mechanism . Recently, generative adversarial networks have been employed in SISR to achieve photo-realistic results .

2 Scene Text Image Super Resolution (STISR)

Different from the general purposed SISR that works on natural scene images, STISR focuses on scene text images. It aims to not only increase the resolution of text image, but also reconstruct semantically correct texts that can benefit the down-stream recognition task. The early methods directly adopt the CNN architectures from SISR for the task of STISR. In , Dong et al. extended SRCNN to text images, and obtained the best performance in ICDAR 2015 competition . PlugNet adopts a pluggable super-resolution unit to deal with LR images in feature domain. TextSR utilizes the text perceptual loss to generate the desired HR images to benefit the text recognition.

To address the problem of STISR on real-world scenes, Wang et al. built a real-world STISR image dataset, namely the TextZoom, where the LR and HR text image pairs were cropped from real-world SISR datasets . They also proposed TSRN to use the sequential residual block to exploit the semantic information in internal features. SCGAN employs a multi-class GAN loss to supervise the STISR model for more perceptual-friendly face and text images. Further, Quan et al. proposed a cascading model for recovering blurry text images in high-frequency domain and image domain collaboratively. Chen et al. and Zhao et al. enhanced the network block structures to improve the STISR performance by self-attending the image features and attending channels.

3 Scene Text Recognition

Scene text recognition aims to extract text content from the input images. Some early approaches tend to recognize each character first and then interpret the whole word, while some others regard the text image as a whole and performing word-level classification . Considering text recognition as an image-to-sequence problem, CRNN extracts image features and uses the recurrent neural networks to model the semantic information. It is trained with CTC loss to align the predicted sequence and the target sequence. Recently, attention-based methods achieve a great progress due to the robustness in extracting text against shape variance of text images . Despite the great performance achieved by the recent methods, it is still difficult to recognize the text in low-resolution images. Therefore, we aim to solve the problem of high-resolution text image restoration for better recognition in this paper.

Methodology

Finally, the TP map fTMf_{TM} and the image feature fIf_{I} are passed into a reconstruction module. This module includes 55 Text-Prior Guided Blocks (TPGBs) that progressively fuse fTMf_{TM} and fIf_{I}, and a final Pixel-Shuffle layer to increase the resolution. Each of the 5 TPGBs firstly merges fTMf_{TM} and fIf_{I} by element-wise addition, followed by a Sequential-Recurrent Block (SRB) to reconstruct the high-resolution image feature. The output of this module is the super-resolved (SR) text image.

2 TP Interpreter

In the proposed architecture, the crucial part lies in the design of TP Interpreter (TPI). The TP Interpreter aims to interpret the text prior fPf_{P} to the image feature fIf_{I} so that the influence of the semantics guidance can be exerted to the correlated spatial position in the image feature domain. One intuitive idea is to enlarge fPf_{P} to the shape of fIf_{I} and then merge them by convolution. Since the convolution operation has a small effective range, the semantics of fPf_{P} cannot be assigned to the distant spatial location in fIf_{I}, especially in the case of spatially-deformed text. Thus, we turn to design a Transformer-based TP Interpreter with attention mechanism to enforce global correlation between text prior fPf_{P} and the image feature fIf_{I}.

As shown in Fig. 3, the proposed TP Interpreter consists of an Encoder part and a Decoder part. The Encoder encodes the text prior fPf_{P} by performing correlation between the semantics of each character in fPf_{P} and outputs the context-enhanced feature fEf_{E}. The decoder performs cross attention between fEf_{E} and fIf_{I} to interpret the semantic information to the image feature.

Decoder. The decoder module accepts the output from the encoder module fEf_{E} and image feature fIf_{I} to perform global cross attention. Similar to the setting in encoder, we firstly add a position encoding to fIf_{I} to incorporate position information. We design a recurrent positional encoding (RPE) to better encode the bias contained in sequential dependency of image feature in horizontal direction, and better help the model look up the text semantic features in the subsequent cross attention . In RPE, we maintain the learnable parameter with the same shape as image feature and encode the sequential dependency in horizontal direction to help the model better learn the neighboring context. See supplementary file for more details.

The position-encoded image feature, denoted by fI′f_{I}^{{}^{\prime}}, and the encoder output fEf_{E} are then delivered to the decoder module for correlation computation. We process the two inputs with a Multi-head Cross Attention (MCA) layer, which performs cross attention operation between fEf_{E} and fI′f_{I}^{{}^{\prime}}. Firstly, the features of fEf_{E} and fI′f_{I}^{{}^{\prime}} are divided into nn subgroups in the channel dimension. Then a cross attention operation CAi\textit{CA}_{i} is performed on the ii-th group of fEf_{E} and fI′f_{I}^{{}^{\prime}}:

The output of MCA is passed to a FFN for feature refinement, and then reshaped to obtain the TP map fTMf_{TM}.

By using the MCA operation, the text prior fEf_{E} can effectively interact with the image feature fI′f_{I}^{{}^{\prime}} by correlating every element in semantic domain to the position in spatial domain. Thus, the semantically meaningful regions in the spatial domain are strengthened in the TP map fTMf_{TM}, which can be used to modulate the image feature for semantic-specific text reconstruction.

3 Text Structure Consistency Loss

While the proposed TATT network can attain a good performance, the reconstructed text image still needs some refinement to improve the visual appearance. This is because it is a bit difficult for a CNN model to represent the deformed text features as it does for regular text features, and the reconstructed text image has weaker character structures with relatively low contrast. As a remedy, we simulate deformed text images and design a text structure consistency (TSC) loss to train the proposed TATT network.

We consider minimizing the distance of three images, i.e., the deformed version of the SR text image DF(Y)\mathbf{D}\mathcal{F}(Y), the SR version of the deformed LR text image F(DY)\mathcal{F}(\mathbf{D}Y) and the deformed ground truth D(X)\mathbf{D}(X), where D\mathbf{D} denotes the random deformationWe consider rotation, shearing and resizing in this paper.. By increasing the similarity among the three items, we can encourage the CNN model to reduce the performance drop when encountering spatial deformations. The proposed TSC loss firstly measures the structural similarity between the above triplet. For this purpose, we extend the Structure-Similarity Index Measure (SSIM) to a triplex SSIM (TSSIM), described as

where μx\mu_{x}, μy\mu_{y}, μz\mu_{z} and σx\sigma_{x}, σy\sigma_{y}, σz\sigma_{z} represent the mean and standard deviation of the triplet xx, yy and zz, respectively. σxy\sigma_{xy}, σyz\sigma_{yz} and σxz\sigma_{xz} denote the correlation coefficients between (x,y)(x,y), (y,z)(y,z) and (x,z)(x,z), respectively. C1C_{1} and C2C_{2} are small constants to avoid instability for dividing values close to zero. The derivation is in the supplementary file.

Lastly, TSC loss LTSCL_{TSC} is designed to measure the mutual structure difference among DF(Y)\mathbf{D}\mathcal{F}(Y), F(DY)\mathcal{F}(\mathbf{D}Y) and DX\mathbf{D}X:

4 Overall Loss Function

In the training, the overall loss function includes a super resolution loss LSRL_{SR}, a text prior loss LTPL_{TP} and the proposed TSC loss LTSCL_{TSC}. The SR loss LSRL_{SR} measures the difference between our SR output F(Y)\mathcal{F}(Y) and the ground-truth HR image XX. We adopt L2L_{2} norm for this computation. The TP loss measures the L1L_{1} norm and KL Divergence between the text prior extracted from the LR image and those from the ground truth. Together with TSC loss LTSCL_{TSC}, the overall loss function is described as follows:

where the α\alpha and β\beta are the balancing parameters.

Experiments

TATT is trained and tested on a single RTX 3090 GPU. We adopt Adam optimizer to train the model with batch size 6464. The training lasts for 500500 epochs with learning rate 10−310^{-3}. The input image of our model is of width 6464 and height 1616, while the output is the 2×2\times SR result. We set the α\alpha and β\beta in (5) to 11 and 0.10.1, respectively (see supplementary file for ablations). The deformation operation D\mathbf{D} in LTSCL_{TSC} is implemented by applying random rotation in a range of $degree,shearingandaspectratioinarangeofdegree, shearing and aspect ratio in a range of[0.5,2.0].TheheadnumbersofMSAandMCAlayersarebothsetto. The head numbers of MSA and MCA layers are both set to4(followingthebestsettingsin).Thenumberofimagefeaturechannels(following the best settings in ). The number of image feature channelsc,,d_{k}inMSA,MCAandFFNcalculationareallsettoin MSA, MCA and FFN calculation are all set to64.ThemodelsizeofTATTis. The model size of TATT is14.9Mintotal.Whentraining,theTPGisinitializedwithpretrainedweightsderivedfrom,whileotherpartsarerandomlyinitialized.Whentesting,TATTwilloccupy6.5GBofGPUmemorywithbatchsizeM in total. When training, the TPG is initialized with pretrained weights derived from , while other parts are randomly initialized. When testing, TATT will occupy 6.5GB of GPU memory with batch size50$.

2 Datasets

TextZoom. TextZoom has 21,74021,740 LR-HR text image pairs collected by changing the focal length of the camera in real-world scenarios, in which 17,36717,367 samples are used for training. The rest samples are divided into three subsets, based on the camera focal length, for testing , namely easy (1,6191,619 samples), medium (1,4111,411 samples) and hard (1,3431,343 samples). Text label is provided in TextZoom.

Scene Text Recognition Datasets. Besides experiments conducted in TextZoom, we also adopt ICDAR2015 , CUTE80 and SVTP to evaluate the robustness of our model in recovering spatially-deformed LR text images. ICDAR2015 has 2,0772,077 scene text images for testing. Most text images suffer from both low quality and perspective-distortion, making the recognition extremely challenging. CUTE80 is also collected in the wild. The test set has 288288 samples in total. Samples in SVTP are mostly curve-shaped text. The total size of the test set is 649649. Besides evaluating our model on the original samples, we further degrade the image quality to test the model generalization against unpredicted bad conditions.

3 Ablation Studies

In this section, we investigate the impact of TP Interpreter, the TSC loss function and the effectiveness of positional encoding. All evaluations in this section are performed on the real-world STISR dataset TextZoom . The text recognition is peformed by CRNN .

Impact of TP Interpreter in SR recovery. Since our TP Interpreter aims at providing better alignment between TP and the image feature and use text semantics to guide SR recovery, we compare it with other guiding strategies, e.g., first upsampling the TP to match the image feature with deconvolution layers or pixel shuffle to align text prior to image feature, and then fusing them to perform guidance with element-wise addition or SFT layers The SFT layer merges the semantics of image feature with channel-wise affine transformation.. The results are shown in Tab. 1. One can see that the proposed TP interpreter obtains that highest PSNR/SSIM, which also indicates the best SR performance.

Referring to the SR text image recognition, one can see that using Pixel-Shuffle and deconvolution strategies provides inferior guidance (46.2%46.2\% and 49.8%49.8\%). There is no stable improvement by combining them with the SFT layers (47.9%47.9\% and 48.6%48.6\%). This is because none of the competing strategies performs global correlation between the text semantics and the image feature, resulting in inferior semantic guidance for SR recovery. In contrast, our TP Interpreter can obtain a good semantics context and accurate alignment to the text region. It thus strengthens the guidance in image feature and improves the text recognition result to 52.6%52.6\%. This validates that using TP Interpreter is an effective way to utilize TP semantics for SR recovery. Some visual comparisons are shown in Fig. 4. One can see that the setting with TP interpreter can lead to the highest quality SR text image with correct semantics.

To demonstrate how the TP Interpreter provides global context, we visualize the attention heatmap provided by our MCA (the outputs from the SMSM layer in (1)) in Fig. 5. One can see that the region of the corresponding foreground character has the highest weight (highlighted). It thus proves that the ability of TP Interpreter in finding semantics in image features. Some other highlighted regions in the neighborhood also demonstrate that the TP Interpreter can be aware of the neighboring context, which can provide better guidance for final SR recovery.

Impact of training with TSC loss. To validate the effectiveness of the TSC loss in refining text structure, we compare the results of 44 models trained with and without the TSC loss, including non-TP based TSRN , TBSRN , TP based TPGSR and TATT. From the results in Tab. 2, one can see that all models lead to a performance gain (4.3%4.3\% for TSRN, 1.3%1.3\% for TBSRN, 0.8%0.8\% for TPGSR, and 1.0%1.0\% for TATT) in SR text recognition when adopting our TSC loss. Notably, though TBSRN is claimed to be robust for multi-oriented text, it can still be improved with our TSC loss, indicating that training with the TSC loss can improve the robustness of reconstructing the character structure against various spatial deformations.

Effectiveness of the RPE. We evaluate the impact of recurrent positional encoding in learning text prior guidance. We deploy different combinations of fixed positional encoding (FPE), learnable positional encoding and the proposed recurrent positional encoding (RPE) in the encoder and decoder modules, and compare the corresponding text recognition results on the SR text images. From Tab. 3, we observe that using LPE or FPE in decoder shows limited performance because they are weak in learning the sequential information. By adopting RPE in the decoder, the SR recognition is increased by 1.8%1.8\%, indicating that RPE is beneficial to text sequential semantics learning.

4 Comparison with State-of-the-Arts

Results on TextZoom. We conduct experiments on the real-world STISR dataset TextZoom to compare the proposed TATT network with state-of-the-art SISR models, including SRCNN and SRResNet and HAN , and STISR models, including TSRN , TPGSR , PCAN and TBSRN . For TPGSR, we compare two models of it, i.e., 1-stage and 3-stage (TPGSR-3). The evaluation metrics are SSIM/PSNR and text recognition accuracy. The comparison results are shown in Tab. 4 and Tab. 5.

One can see that our model trained with LTSCL_{TSC} achieves the best PSNR (21.5221.52) and SSIM (0.79300.7930) overall performance. This verifies the superiority of our method in improving the image quality. As for the SR text recognition, our method achieves new state-of-the-art accuracy under all settings by using the text recognition models of ASTER and CRNN . It even surpasses the 3-stage model TPGSR-3 by using only a single stage.

We also test the inference speed of the three most competitive STISR methods, i.e., TBSRN (982982 fps), TPGSR (1,0851,085 fps) and our TATT model (960960 fps). TATT has comparable speed with TPGSR and TBSRN, while surpasses them by 2.7%2.7\% and 3.6%3.6\% in SR image text recognition by using ASTER as the recognizer.

To further investigate the performance on spatially deformed text images, we manually pick 804804 rotated and curve-shaped samples from TextZoom test set to evaluate the compared models. Results in Tab. 6 indicate that our TATT model obtains the best performance, and the average gap over models like TPGSR and TBSRN becomes larger when encountering spatially deformed text.

We also visualize the recovery results of both regular samples and spatially-deformed samples of TextZoom in Fig. 6. Without TP guidance, TSRN and TBSRN perform far from readable and they are visually unacceptable. With the TP guidance, TPGSR is still unstable in recovering spatially-deformed images. In contrast, our TATT network performs much better in recovering text semantics in samples of all cases compared to all the competitors. With TSC loss, our model further upgrades the visual quality of the samples with better-refined character structure.

Generalization to recognition dataset. We evaluate the generalization performance of our TATT network to other real-world text image datasets, including ICDAR15 , CUTE80 and SVTP . These datasets are built for text recognition purpose and contain spatially deformed text image in natural scenes. Since some of the images in these datasets have good quality, we only pick the low-resolution images (i.e., lower than 16×6416\times 64) to form our test set with 533533 samples (391391 from ICDAR15, 33 from CUTE80 and 139139 from SVTP). Since the degradation is relatively small, we manually add some degradation on them, including contrast variation, Gaussian noise and Gaussian blurring (see details in supplementary file). We compare with TSRN , TBSRN and TPGSR in this test and evaluate the recognition accuracy on the SR results. All models are trained on TextZoom and tested on the picked low-quality images.

The results are illustrated in Tab. 7, we can see that the proposed TATT network achieves the highest recognition accuracy across all types of degradations. This indicates that our TATT network, though trained on TextZoom, can be well generalized to images in other datasets. The reconstructed high-quality text images by TATT can benefit the downstream tasks such as text recognition.

Conclusion and Discussions

In this paper, we proposed a Text ATTention network for single text image super-resolution. We leveraged a text prior, which is the semantic information extracted from the text image, to guide the text image reconstruction process. To tackle with the spatially-deformed text recovery, we developed a transformer-based module, called TP Interpreter, to globally correlate the text prior in the semantic domain to the character region in image feature domain. Moreover, we proposed a text structure consistency loss to refine the text structure by imposing structural consistency between the recovered regular and deformed texts. Our model achieved state-of-the-art performance in not only the text super resolution task but the downstream text recognition task.

Though recording state-of-the-art results, the proposed TATT network has limitation on recovering extremely blurry texts, as shown in Fig. 7. In such cases, the strokes of the characters in the text are mixed together, which are difficult to separate. In addition, the computational complexity of our TATT network grows exponentially with the length of the text in the image due to the global attention adopted in our model. It is expected to reduce the computational complexity and improve run-time efficiency of TATT, which will be our future work.

Acknowledgements

This work is supported by the Hong Kong RGC RIF grant (R5001-18). We thank Dr. Lida Li for the useful discussion on this project.

References