Robust Scene Text Recognition with Automatic Rectification

Baoguang Shi, Xinggang Wang, Pengyuan Lyu, Cong Yao, Xiang Bai

Introduction

In natural scenes, text appears on various kinds of objects, e.g. road signs, billboards, and product packaging. It carries rich and high-level semantic information that is important for image understanding. Recognizing text in images facilitates many real-world applications, such as geo-location, driverless car, and image-based machine translation. For these reasons, scene text recognition has attracted great interest from the community . Despite the maturity of the research on Optical Character Recognition (OCR) , recognizing text in natural images, rather than scanned documents, is still challenging. Scene text images exhibit large variations in the aspects of illumination, motion blur, text font, color, etc. Moreover, text in the wild may have irregular shape. For example, some scene text is perspective text , which is caused by side-view camera angles; some has curved shapes, meaning that its characters are placed along curves rather than straight lines. We call such text irregular text, in contrast to regular text which is horizontal and frontal.

Usually, a text recognizer works best when its input images contain tightly-bounded regular text. This motivates us to apply a spatial transformation prior to recognition, in order to rectify input images into ones that are more “readable” by recognizers. In this paper, we propose a recognition method that is robust to irregular text. Specifically, we construct a deep neural network that combines a Spatial Transformer Network (STN) and a Sequence Recognition Network (SRN). An overview of the model is given in Fig. 1.

In the STN, an input image is spatially transformed into a rectified image. Ideally, the STN produces an image that contains regular text, which is a more appropriate input for the SRN than the original one. The transformation is a thin-plate-spline (TPS) transformation, whose nonlinearity allows us to rectify various types of irregular text, including perspective and curved text. The TPS transformation is configured by a set of fiducial points, whose coordinates are regressed by a convolutional neural network.

In an image that contains regular text, characters are arranged along a horizontal line. It bares some resemblance to a sequential signal. Motivated by this, for the SRN we construct an attention-based model that recognizes text in a sequence recognition approach. The SRN consists of an encoder and a decoder. Given an input image, the encoder generates a sequential feature representation, which is a sequence of feature vectors. The decoder recurrently generates a character sequence conditioning on the input sequence, by decoding the relevant contents which are determined by its attention mechanism at each step.

We show that, with proper initialization, the whole model can be trained end-to-end. Consequently, for the STN, we do not need to label any geometric ground truth, i.e. the positions of the TPS fiducial points, but let its training be supervised by the error differentials back-propagated by the SRN. In practice, the training eventually makes the STN tend to produce images that contain regular text, which are desirable inputs for the SRN.

The contributions of this paper are three-fold: First, we propose a novel scene text recognition method that is robust to irregular text. Second, our model extends the STN framework with an attention-based model. The original STN is only tested on plain convolutional neural networks. Third, our model adopts a convolutional-recurrent structure in the encoder of the SRN, thus is a novel variant of the attention-based model .

Related Work

In recent years, a rich body of literature concerning scene text recognition has been published. Comprehensive surveys have been given in . Among the traditional methods, many adopt bottom-up approaches, where individual characters are firstly detected using sliding window , connected components , or Hough voting . Following that, the detected characters are integrated into words by means of dynamic programming, lexicon search , etc.. Other work adopts top-down approaches, where text is directly recognized from entire input images, rather than detecting and recognizing individual characters. For example, Almázan et al. propose to predict label embedding vectors from input images. Jaderberg et al. address text recognition with a 90k-class convolutional neural network, where each class corresponds to an English word. In , a CNN with a structured output layer is constructed for unconstrained text recognition. Some recent work models the problem as a sequence recognition problem, where text is represented by character sequence. Su and Lu extract sequential image representation, which is a sequence of HOG descriptors, and predict the corresponding character sequence with a recurrent neural network (RNN). Shi et al. propose an end-to-end sequence recognition network which combines CNN and RNN. Our method also adopts the sequence prediction scheme, but we further take the problem of irregular text into account.

Although being common in the tasks of scene text detection and recognition, the issue of irregular text is relatively less addressed in explicit ways. Yao et al. firstly propose the multi-oriented text detection problem, and deal with it by carefully designing rotation-invariant region descriptors. Zhang et al. propose a character rectification method that leverages the low-rank structures of text. Phan et al. propose to explicitly rectify perspective distortions via SIFT descriptor matching. The above-mentioned work brings insightful ideas into this issue. However, most methods deal with only one type of irregular text with specifically designed schemes. Our method rectifies several types of irregular text in a unified way. Moreover, it does not require extra annotations for the rectification process, since the STN is supervised by the SRN during training.

Proposed Model

In this section we formulate our model. Overall, the model takes an input image II and outputs a sequence l=(l1,…,lT)\mathbf{l}=(l_{1},\dots,l_{T}), where ltl_{t} is the tt-th character, TT is the variable string length.

The STN transforms an input image II to a rectified image I′I^{\prime} with a predicted TPS transformation. It follows the framework proposed in . As illustrated in Fig. 2, it first predicts a set of fiducial points via its localization network. Then, inside the grid generator, it calculates the TPS transformation parameters from the fiducial points, and generates a sampling grid on II. The sampler takes both the grid and the input image, it produces a rectified image I′I^{\prime} by sampling on the grid points.

A distinctive property of STN is that its sampler is differentiable. Therefore, once we have a differentiable localization network and a differentiable grid generator, the STN can back-propagate error differentials and gets trained.

The localization network localizes KK fiducial points by directly regressing their x,yx,y-coordinates. Here, constant KK is an even number. The coordinates are denoted by C=[c1,…,cK]∈ℜ2×K\mathbf{C}=\left[\mathbf{c}_{1},\dots,\mathbf{c}_{K}\right]\in\Re^{2\times K}, whose kk-th column ck=[xk,yk]⊺\mathbf{c}_{k}=\left[x_{k},y_{k}\right]^{\intercal} contains the coordinates of the kk-th fiducial point. We use a normalized coordinate system whose origin is the image center, so that xk,ykx_{k},y_{k} are within the range of $$.

We use a convolutional neural network (CNN) for the regression. Similar to the conventional structures , the CNN contains convolutional layers, pooling layers and fully-connected layers. However, we use it for regression instead of classification. For the output layer, which is the last fully-connected layer, we set the number of output nodes to 2K2K and the activation function to tanh⁡(⋅)\tanh(\cdot), so that its output vectors have values that are within the range of (−1,1)\left(-1,1\right). Last, the output vector is reshaped into C\mathbf{C}.

The network localizes fiducial points based on global image contexts. It is expected to capture the overall text shape of an input image, and localizes fiducial points accordingly. It should be emphasized that we do not annotate coordinates of fiducial points for any sample. Instead, the training of the localization network is completely supervised by the gradients propagated by the other parts of the STN, following the back-propagation algorithm .

1.2 Grid Generator

The grid generator estimates the TPS transformation parameters, and generates a sampling grid. We first define another set of fiducial points, called the base fiducial points, denoted by C′=[c1′,…,cK′]∈ℜ2×K\mathbf{C}^{\prime}=\left[\mathbf{c}_{1}^{\prime},\dots,\mathbf{c}_{K}^{\prime}\right]\in\Re^{2\times K}. As illustrated in Fig. 3, the base fiducial points are evenly distributed along the top and bottom edge of a rectified image I′I^{\prime}. Since KK is a constant and the coordinate system is normalized, C′\mathbf{C}^{\prime} is always a constant.

The parameters of the TPS transformation is represented by a matrix T∈ℜ2×(K+3)\mathbf{T}\in\Re^{2\times(K+3)}, which is computed by

where ΔC′∈ℜ(K+3)×(K+3)\boldsymbol{\Delta}_{\mathbf{C}^{\prime}}\in\Re^{(K+3)\times(K+3)} is a matrix determined only by C′\mathbf{C}^{\prime}, thus also a constant:

where the element on the ii-th row and jj-th column of R\mathbf{R} is ri,j=di,j2ln⁡di,j2r_{i,j}=d_{i,j}^{2}\ln d_{i,j}^{2}, di,jd_{i,j} is the euclidean distance between ci′\mathbf{c}^{\prime}_{i} and cj′\mathbf{c}^{\prime}_{j}.

The grid of pixels on a rectified image I′I^{\prime} is denoted by P′={pi′}i=1,…,N{\cal P}^{\prime}=\left\{\mathbf{p}^{\prime}_{i}\right\}_{i=1,\dots,N}, where pi′=[xi′,yi′]⊺\mathbf{p}_{i}^{\prime}=\left[x_{i}^{\prime},y_{i}^{\prime}\right]^{\intercal} is the x,y-coordinates of the ii-th pixel, NN is the number of pixels. As illustrated in Fig. 3, for every point pi′\mathbf{p}^{\prime}_{i} on I′I^{\prime}, we find the corresponding point pi=[xi,yi]⊺\mathbf{p}_{i}=\left[x_{i},y_{i}\right]^{\intercal} on II, by applying the transformation:

where di,kd_{i,k} is the euclidean distance between pi′\mathbf{p}_{i}^{\prime} and the kk-th base fiducial point ck′\mathbf{c}^{\prime}_{k}.

By iterating over all points in P′{\cal P}^{\prime}, we generate a grid P={pi}i=1,…,N{\cal P}=\left\{\mathbf{p}{}_{i}\right\}_{i=1,\dots,N} on the input image II. The grid generator can back-propagate gradients, since its two matrix multiplications, Eq. 1 and Eq. 4, are both differentiable.

1.3 Sampler

Lastly, in the sampler, the pixel value of pi′\mathbf{p}_{i}^{\prime} is bilinearly interpolated from the pixels near pi\mathbf{p}_{i} on the input image. By setting all pixel values, we get the rectified image I′I^{\prime}:

where VV represents the bilinear sampler , which is also a differentiable module.

The flexibility of the TPS transformation allows us to transform irregular text images into rectified images that contain regular text. In Fig. 4, we show some common types of irregular text, including a) loosely-bounded text, which resulted by imperfect text detection; b) multi-oriented text, caused by non-horizontal camera views; c) perspective text, caused by side-view camera angles; d) curved text, a commonly seen artistic style. The STN is able to rectify images that contain these types of irregular text, making them more readable for the following recognizer.

2 Sequence Recognition Network

Since target words are inherently sequences of characters, we model the recognition problem as a sequence recognition problem, and address it with a sequence recognition network. The input to the SRN is a rectified image I′I^{\prime}, which ideally contains a word that is written horizontally from left to right. We extract a sequential representation from I′I^{\prime}, and recognize a word from it.

In our model, the SRN is an attention-based model , which directly recognizes a sequence from an input image. The SRN consists of an encoder and a decoder. The encoder extracts a sequential representation from the input image I′I^{\prime}. The decoder recurrently generates a sequence conditioned on the sequential representation, by decoding the relevant contents it attends to at each step.

A naïve approach for extracting a sequential representation for I′I^{\prime} is to take local image patches from left to right, and describe each of them with a CNN. However, this approach does not share the computation among overlapping patches, thus inefficient. Besides, the spatial dependencies between the patches are not exploited and leveraged. Instead, following , we build a network that combines convolutional layers and recurrent networks. The network extracts a sequence of feature vectors, given an input image of arbitrary size.

2.2 Decoder: Recurrent Character Generator

The decoder recurrently generates a sequence of characters, conditioned on the sequence produced by the encoder. It is a recurrent neural network with the attention structure proposed in . In the recurrency part, we adopt the Gated Recurrent Unit (GRU) as the cell.

The generation is a TT-step process, at step tt, the decoder computes a vector of attention weights αt∈ℜL\boldsymbol{\alpha}_{t}\in\Re^{L} via the attention process described in :

where st−1\mathbf{s}_{t-1} is the state variable of the GRU cell at the last step. For t=1t=1, both s0\mathbf{s}_{0} and α0\boldsymbol{\alpha}_{0} are zero vectors. Then, a glimpse gt\mathbf{g}_{t} is computed by linearly combining the vectors in h\mathbf{h}: gt=∑i=1Lαtihi\mathbf{g}_{t}=\sum_{i=1}^{L}\alpha_{ti}\mathbf{h}_{i}. Since αt\boldsymbol{\alpha}_{t} has non-negative values that sum to one, it effectively controls where the decoder focuses on.

The state st−1\mathbf{s}_{t-1} is updated via the recurrent process of GRU :

where lt−1l_{t-1} is the (t−1)(t-1)-th ground-truth label in training, while in testing, it is the label predicted in the previous step, i.e. l^t−1\hat{l}_{t-1}.

The probability distribution over the label space is estimated by:

Following that, a character l^t\hat{l}_{t} is predicted by taking the class with the highest probability. The label space includes all English alphanumeric characters, plus a special “end-of-sequence” (EOS) token, which ends the generation process.

The SRN directly maps a input sequence to another sequence. Both input and output sequences may have arbitrary lengths. It can be trained with only word images and associated text.

3 Model Training

We denote the training set by X={(I(i),l(i))}i=1…N{\cal X}=\{(I^{(i)},\mathbf{l}^{(i)})\}_{i=1\dots N}. To train the model, we minimize the negative log-likelihood over X{\cal X}:

where the probability p(⋅)p(\cdot) is computed by Eq. 8, θ\boldsymbol{\theta} is the parameters of both STN and SRN. The optimization algorithm is the ADADELTA , which we find fast in convergence speed.

The model parameters are randomly initialized, except the localization network, whose output fully-connected layer is initialized by setting weights to zero. The initial biases are set to such values that yield the fiducial points pattern displayed in Fig. 6.a. Empirically, we also find that the patterns displayed Fig. 6.b and Fig. 6.c yield relatively poorer performance. Randomly initializing the localization network results in failure of convergence during training.

4 Recognizing With a Lexicon

When a test image is associated with a lexicon, i.e. a set of words for selection, the recognition process is to pick the word with the highest posterior conditional probability:

However, on very large lexicons, e.g. the Hunspell which contains more than 50k words, computing Eq. 10 is time consuming, as it requires iterating over all lexicon words. We adopt an efficient approximate search scheme on large lexicons. The motivation is that computation can be shared among words that share the same prefix.

We first construct a prefix tree over a given lexicon. As illustrated in Fig. 7, each node of the tree is a character label. Nodes on a path from the root to a leaf forms a word (including the EOS). In testing, we start from the root node, every time the model outputs a distribution y^t\hat{\mathbf{y}}_{t}, the child node with the highest posterior probability is selected as the next node to move to. The process repeats until a leaf node is reached, and a word is found on the path from the root to that leaf. Since the tree depth is at most the length of the longest word in the lexicon, this search process takes much less computation than the precise search.

Recognition performance could be further improved by incorporating beam search. A list of nodes is maintained, and the above search process is repeated on each of them. After each step, the list is updated to store the nodes with top-BB accumulated log-likelihoods, where BB is the beam width. Larger beam width usually results in better performance, but lower search speed.

Experiments

In this section we evaluate our model on a number of standard scene text recognition benchmarks, paying special attention to recognition performance on irregular text. First we evaluate our model on some general recognition benchmarks, which mainly consist of regular text, but irregular text also exists. Next, we perform evaluations on benchmarks that are specially designed for irregular text recognition. For all benchmarks, performance is measured by word accuracy.

Spatial Transformer Network The localization network of STN has 4 convolution layers, each followed by a 2×22\times 2 max-pooling layer. The filter size, padding size and stride are 3, 1, 1 respectively, for all convolutional layers. The number of filters are respectively 64, 128, 256 and 512. Following the convolutional and the max-pooling layers is two fully-connected layers with 1024 hidden units. We set the number of fiducial points to K=20K=20, meaning that the localization network outputs a 40-dimensional vector. Activation functions for all layers are the ReLU , except the output layer which uses tanh⁡(⋅)\tanh(\cdot).

Sequence Recognition Network In the SRN, the encoder has 7 convolutional layers, whose {filter size, number of filters, stride, padding size} are respectively {3,64,1,1}, {3,128,1,1}, {3,256,1,1}, {3,256,1,1,}, {3,512,1,1}, {3,512,1,1}, and {2,512,1,0}. The 1st, 2nd, 4th, 6th convolutional layers are each followed by a 2×22\times 2 max-pooling layer. On the top of the convolutional layers is a two-layer BLSTM network, each LSTM has 256 hidden units. For the decoder, we use a GRU cell that has 256 memory blocks and 37 output units (26 letters, 10 digits, and 1 EOS token).

Model Training Our model is trained on the 8-million synthetic samples released by Jaderberg et al. . No extra data is used. The batch size is set to 64 in training. Following , images are resized to 100×32100\times 32 in both training and testing. The output size of the STN is also 100×32100\times 32. Our model processes ∼\sim160 samples per second during training, and converges in 2 days after ∼\sim3 epochs over the training dataset.

Implementation We implement our model under the Torch7 framework . Most parts of the model are GPU-accelerated. All our experiments are carried out on a workstation which has one Intel Xeon(R) E5-2620 2.40GHz CPU, an NVIDIA GTX-Titan GPU, and 64GB RAM.

Without a lexicon, the model takes less than 2ms recognizing an image. With a lexicon, recognition speed depends on the lexicon size. We adopt the precise search (Sec. 3.4) when lexicon size ≤\leq 1k. On larger lexicons, we adopt the approximate beam search (Sec. 3.4) with a beam width of 7. With a 50k-word lexicon, the search takes ∼\sim200ms per image.

2 Results on General Benchmarks

Our model is firstly evaluated on benchmarks that are designed for general scene text recognition tasks. Samples in these benchmarks mostly contain regular text, but irregular text also exists. The benchmark datasets are:

IIIT 5K-Words (IIIT5K) contains 3000 cropped word images for testing. The images are collected from the Internet. For each image, there is a 50-word lexicon and a 1000-word lexicon. All lexicons consist of a ground truth word and some randomly picked words.

Street View Text (SVT) is collected from Google Street View. Its test dataset consists of 647 word images. Many images in SVT are severely corrupted by noise and blur, or have very low resolutions. Each sample is associated with a 50-word lexicon.

ICDAR 2003 (IC03) contains 860 cropped word images, each associated with a 50-word lexicon defined by Wang et al. . Following , we discard images that contain non-alphanumeric characters or have less than three characters. Besides, there is a “full lexicon” which contains all lexicon words, and the Hunspell lexicon which has 50k words.

ICDAR 2013 (IC13) inherits most of its samples from IC03. After filtering samples as done in IC03, the dataset contains 857 samples.

In Tab. 1 we report our results, and compare them with other methods. On unconstrained recognition tasks (recognizing without a lexicon), our model outperforms all the other methods in comparison. On IIIT5K, RARE outperforms prior art CRNN by nearly 4 percentages, indicating a clear improvement in performance. We observe that IIIT5K contains a lot of irregular text, especially curved text, while RARE has an advantage in dealing with irregular text. Note that, although our model falls behind on some datasets, our model differs from in that it is able recognize random strings such as telephone numbers, while only recognizes words that are in its 90k-dictionary. On constrained recognition tasks (recognizing with a lexicon), RARE achieves state-of-the-art or highly competitive accuracies. On IIIT5K, SVT and IC03, constrained recognition accuracies are on par with , and slightly lower than .

We also train and test a model that contains only the SRN. As reported in the last row of Tab. 1, we see that the SRN-only model is also a very competitive recognizer, achieving higher or competitive performance on most of the benchmarks.

3 Recognizing Perspective Text

To validate the effectiveness of the rectification scheme, we evaluate RARE on the task of perspective text recognition. SVT-Perspective is specifically designed for evaluating performance of perspective text recognition algorithms. Text samples in SVT-Perspective are picked from side view angles in Google Street View, thus most of them are heavily deformed by perspective distortion. Some examples are shown in Fig. 8.a. SVT-Perspective consists of 639 cropped images for testing. Each image is associated with a 50-word lexicon, which is inherited from the SVT dataset. In addition, there is a “Full” lexicon which contains all the per-image lexicon words.

We use the same model trained on the synthetic dataset without fine-tuning. For comparison, we test the CRNN model on SVT-Perspective. We also compare RARE with , whose recognition accuracies are reported in .

Tab. 2 summarizes the results. In the second and third columns, we compare the accuracies of recognition with the 50-word lexicon and the full lexicon. Our method outperforms , which is a perspective text recognition method, by a large margin on both lexicons. However, this gap is partially due to that we use a much larger training set than . In the comparisons with , which uses the same training set as RARE, we still observe significant improvements in both the Full lexicon and the lexicon-free settings. Furthermore, recall the results in Tab. 1, on SVT-Perspective RARE outperforms by a even larger margin. The reason is that the SVT-perspective dataset mainly consists of perspective text, which is inappropriate for direct recognition. Our rectification scheme can significantly alleviate this problem.

In Fig. 9 we present some qualitative analysis. Fiducial points predicted by the STN are plotted on input images in green crosses. We see that the STN tends to place fiducial points along upper and lower edges of scene text, and hence produces rectified images that are more readable for the SRN. However, the STN fails sometimes in the case of heavy perspective distortion.

4 Recognizing Curved Text

Curved text is a commonly seen artistic-style text in natural scenes. Due to its irregular character placement, recognizing curved text is very challenging. CUTE80 focuses on the recognition of curved text. The dataset contains 80 high-resolution images taken in natural scenes. Originally, the dataset is proposed for detection tasks. We crop the words, resulting in 288 word images for testing. For comparisons, we evaluate the trained models of and . All models are evaluated without a lexicon.

From the results summarized in Tab. 3, we see that RARE outperforms the other two methods by a large margin. is a constrained recognition model, it cannot recognize words that are not in its dictionary. is able to recognize arbitrary words, but it does not have a specific mechanism for handling curved text. Our model rectifies images that contain curved text before recognizing them. Therefore, it is advantageous on this task.

In Fig. 9, we demonstrate the effect of rectification through some examples. Generally, the rectification made by the STN is not perfect, but it alleviates the recognition difficulty to some extent. RARE tends to fail when curve angles are too large, as shown in the last two rows of Fig. 9.

Conclusion

We study a common but difficult problem in scene text recognition, called the irregular text problem. Traditional solutions typically use a separate text rectification component. We address this problem in a more feasible and elegant way by adopting a differentiable spatial transformer network module. In addition, the spatial transformer network is connected to an attention-based sequence recognizer, allowing us to train the whole model end-to-end. The extensive experimental results show that 1) without geometric supervision, the learned model can automatically generate more “readable” images for both human and the sequence recognition network; 2) the proposed text rectification method can significantly improve recognition accuracies on irregular scene text; 3) the proposed scene text recognition system is competitive compared with the state-of-the-arts. In the future, we plan to address the end-to-end scene text reading problem through the combination of RARE with a scene text detection method, e.g. .

Acknowledgments

This work was primarily supported by National Natural Science Foundation of China (NSFC) (No. 61222308, No. 61573160 and No. 61503145), and Open Project Program of the State Key Laboratory of Digital Publishing Technology (No. F2016001).

References