Supervised mid-level features for word image representation

Albert Gordo

Introduction

In recent years there has been an increasing interest in tasks related to text understanding in natural scenes, and, amongst them, in word recognition: given a cropped image of a word, one is interested in obtaining its transcription. The most popular approaches to address this task involve detecting and localizing individual characters in the word image and using that information to infer the contents of the word, using for example conditional random fields and language priors . As shown by Bissacco et al. , such approaches that learn directly from the annotated individual characters can obtain impressive accuracy if large volumes of training data are available. However, these approaches are not exempt from problems. First, to obtain a high accuracy, one needs to annotate very large amounts of words (in the order of millions) with character bounding boxes for training purposes, as done by Bissaco et al. . When limited training data is available, the recognition accuracy of these approaches is much lower . Then, at testing time, one needs to localize the individual characters of the word image, which is slow and error prone. Also, these approaches do not lead to a final signature of the word image that can be used for other tasks such as word image matching and retrieval.

Rather than localizing and classifying the individual characters in a word, a new trend in word image recognition and retrieval has been to describe word images with global representations using standard computer vision features (e.g. HOG , or SIFT features aggregated with bags of words or Fisher vector encodings) and apply different frameworks and machine learning techniques (such as using attribute representations, metric learning, or exemplar SVMs) on top of these global representations to learn models to perform tasks such as recognition, retrieval, or spotting . The global approaches have important advantages: they do not require that words be annotated with character bounding boxes for training and they do not require that the characters forming a word be explicitly localized at testing time. They can also produce compact signatures which are faster to compute, store and index, or compare, while still obtaining very competitive results in many tasks. The use of off-the-shelf computer vision features and machine learning techniques also makes them very attractive. Yet, one may argue that not using character bounding box annotations during training, although very convenient, may be a limiting factor for their accuracy.

The main contribution of this paper is an approach to construct a global word image representation that unites the best properties of both main trends by leveraging character bounding boxes information at training time. This is achieved by learning mid-level local features that are correlated with the characters in which they tend to appear. We use a small external dataset annotated at the character level to learn how to transform small groups of locally aggregated low-level features into mid-level semantic features suitable for text, and then we encode and aggregate these mid-level features into a global representation. This unites advantages of both paradigms: a compact signature that does not require localizing characters explicitly at test time, but still exploits information about annotated characters. Although some works have already used supervised information (in the form of text transcriptions) to project a global image representation into a more semantic space , to the best of our knowledge, no other approach that constructs global image representations has leveraged character bounding box information at training time to do so.

We test our approach on two public benchmarks of scene text showing that constructing representations using mid-level features yields large improvements over constructing them using SIFT features directly. When pairing these mid-level features with the recent attributes framework of , we significantly outperform the state-of-the-art in word image retrieval (using both images or strings as queries), and obtain results comparable to or better than Google’s PhotoOCR in recognition tasks using a tiny fraction of training data: we use less than 5,0005,000 training words annotated at the character level, while uses several millions.

The rest of the paper is organized as follows. Section 2 reviews the related work. Section 3 describes our method. Section 4 deals with the experimental evaluation. Finally, Section 5 concludes the paper.

Related Work

We now review those works which are most related to our approach.

Most works focusing on scene-text target only the problem of recognition, i.e., given the image of a word, the goal is to produce its transcription. In such a case, a priori, there are no clear advantages with producing a global image representation. Instead, most methods aim at localizing and classifying characters or character regions inside the image and using this information to infer the transcription. For example, Mishra et al. propose to detect characters using a sliding window model and produce a transcription using a conditional random field model with language priors. Neumann and Matas use Extreme Regions or Strokes to localize and describe characters. Words are then recognized using a commercial OCR or by maximizing the characters’ probability. The recent uses a mid-level representation of strokes to produce more semantic descriptions of characters, that are then classified using random forests. PhotoOCR learns a character classifier using a deep architecture with millions of annotated training characters, and produces very accurate transcriptions at the cost of requiring vast amounts of annotated data. In a different line, Jaderberg et al. propose to learn the classification task directly from the image without localizing the characters using deep convolutional neural networks. Although the results on some tasks are impressive, this approach also requires millions of annotated training samples to perform well.

Global representations for word images.

More directly related to our work are global representations. Producing image signatures opens the door to other tasks such as word retrieval, as well as easing tasks such as storing and indexing word images. Rusiñol et al. construct a bag of words over SIFT descriptors to encode word images, and use it to perform segmentation-free spotting on handwritten documents. In , this framework is enriched using textual information. In , Exemplar SVMs and HOG descriptors are also used to perform segmentation-free spotting on documents. Goel et al. propose to recognize scene-text images by synthesizing a dataset of annotated images and finding the nearest neighbor in that dataset. Rodriguez and Perronnin propose in a label embedding approach that puts word images (represented with Fisher vectors) and text strings in the same vectorial space. Similarly, Almazán et al. propose a word attributes framework that can perform both retrieval and recognition in a low dimensional space. Although they perform well in retrieval tasks, they are outperformed in recognition by methods that exploit annotated character information, such as PhotoOCR .

Learning mid-level features.

Our work is also related to the use of mid-level features (e.g. ), where “blocks” that contain some basic semantic information are discovered, learned, and/or defined. The use of mid-level features has been shown to produce large improvements in different tasks. Of those works, the most related to ours is the work of Yao et al. , which learns Strokelets, a mid-level representation that can be understood as “parts” of characters. These are then used to represent characters in a more semantic way. The main distinctions between the Strokelets and our work are that i), we exploit supervised information to learn a more semantic representation, and ii), we do not explicitly classify the character blocks, and instead use this semantic representation to construct a high-level word image signature. We show that our approach leads to significantly better results than the Strokelets . The embedding approaches of could also be understood as producing supervised mid-level features, but do not use character bounding box information to do so. To learn the semantic space, we perform supervised dimensionality reduction of local Fisher vectors that are then encoded and aggregated into a global Fisher vector. This could be seen as a deep Fisher network for image recognition , also used very recently for action recognition . The main difference is that in the supervised dimensionality reduction step is learned using the image labels, the same labels that will be used for the final classification step. In our case, the goal is to transfer knowledge from the individual character bounding boxes, which are only annotated in the training set, to produce features that are correlated with characters, and exploit this information in the target datasets, where these bounding boxes are not available. This would be similar to learning the extra layer of the deep Fisher network using the labels of the bounding boxes of the objects, instead of using the whole image label as does. Although the resulting architectures are similar, the motivation behind them is very different. In that sense, our work can also be related to works on learning with privileged information , where the information available at training time to construct the representations (in our case, character bounding boxes) is not available at test time.

Mid-Level Features for Word Images

A standard approach to construct a global word image representation is to i) extract low-level descriptors, ii) encode the descriptors, and iii) aggregate them into a global representation, potentially with spatial pyramids to add some weak geometry. This representation can then be used as input for different learning approaches, as done e.g. in . Figure 1 (top) illustrates this process.

In our proposed method, we aim at constructing global representations based on semantic mid-level features instead of using the low-level descriptors directly. The goal is to produce features that might not be good enough to predict the individual characters directly, but are more correlated with the individual characters than SIFT or other local descriptors.

This is achieved not by finding and classifying characteristic blocks as in , but by projecting all possible image blocks into a lower-dimensional space where our mid-level visual features and the characters are more correlated. By projecting all possible blocks in an image, one obtains a set of mid-level descriptors that can then be aggregated into a global image representation. The process is illustrated in Figure 1 (bottom).

In what follows, we first describe the learning process and describe how to project the image blocks into the semantic space correlated with the characters (Section 3.1). We then describe how to extract these mid-level features from a new image, and how to combine them with the word attributes framework to obtain very compact, discriminative word representations (Section 3.2).

Let us assume that, at training time, one has access to a set of NN word images and their respective annotations. Let I\mathcal{I} be one word image containing ∣I∣|\mathcal{I}| characters, and let its annotation be a list of character bounding boxes char_bbchar\_bb and character labels char_ychar\_y, CI={(char_bbi,char_yi),i=1…∣I∣}\mathcal{C}_{\mathcal{I}}=\{(char\_bb_{i},char\_y_{i}),i=1\ldots|\mathcal{I}|\}. Each character label char_yichar\_y_{i} belongs to one of the 6262 characters in the following alphabet: Σ={A,…,Z,a,…,z,0,…,9}\Sigma=\{\verb|A|,\ldots,\verb|Z|,\verb|a|,\ldots,\verb|z|,\verb|0|,\ldots,\verb|9|\}.

Let us denote with block_bbblock\_bb a square block in the image represented as a bounding box. For training purposes, let us randomly sample from every image SS blocks of different sizes (e.g. 32×3232\times 32 pixels, 48×4848\times 48, etc) at different positions. Some of these blocks may contain only background, but most of them will contain parts of a character or parts of two consecutive characters.

We then describe these blocks using two modalities. The first one is based on visual features. The second one is based on character annotations. This second modality is more discriminative but is only available during training.

Then, one can learn a mapping between the visual features and the character annotation space. Once this mapping has been learned, at testing time one can extract all possible blocks in an image, represent them first with low-level visual features, and then map them into this semantic space to obtain mid-level features, as seen in Figure 1 bottom.

Finally, stacking the descriptors of all the sampled blocks of all the training images leads to a matrix X\mathbf{X} of size NS×DvNS\times D_{v}, where we denote with DvD_{v} the dimensionality of these visual representations. In our experiments we will use visual descriptors of Dv=4,096D_{v}=4,096 dimensions.

Block annotations.

The second view is based on the annotation and contains information about the overlaps between the blocks and the characters in the word image. The goal is to encode with which characters the sampled blocks tend to overlap. In particular, we are interested in encoding what percentage of the character regions are covered by the blocks. As a first approach, we construct a DaD_{a}-dimensional label yy for each sampled block, with Da=∣Σ∣=62D_{a}=|\Sigma|=62. Given a block, its label yy is encoded as follows: for each character Σd\Sigma_{d} in the alphabet Σ\Sigma, we assign at dimension dd of the label vector the normalized overlap between the bounding box of the block and the bounding boxes of the characters of the word whose label char_yichar\_y_{i} equals Σd\Sigma_{d}. Since the word may have repeated characters overlapping with the same block, we take the maximum overlap:

where ∣⋅∣|\cdot| represents the area of the region, and δi,d\delta_{i,d} equals 11 if char_yi=Σdchar\_y_{i}=\Sigma_{d} and otherwise. Figure 2 illustrates this with an annotated image and a sampled block with its corresponding computed label yy.

Learning the mid-level space.

Discussion.

In the described system, the block labels encode which percentage of the characters or the character regions are covered by the blocks. However, this is only one possible way to encode the labels. Other options could include e.g. encoding which percentage of the block is covered, the intersection over union, or working at the pixel level instead of the region level. Any representation that relates the visual block with the character annotation could be considered. We found that the proposed approach worked well in practice and deemed the search for the optimum representation out of the scope of this work.

2 Representing Word Images with Mid-Level Features

Once the matrix U\mathbf{U} has been learned, one can use it to compute the set of mid-level features of a new word image. A naive approach would involve extracting all possible blocks, encoding them with Fisher vectors, and projecting them with U\mathbf{U} into the mid-level space. This is shown in Algorithm 2.

Building word representations. The mid-level features can then be encoded and aggregated into a global image representation using e.g. Fisher vectors. These global image representations can then be used by themselves, but can also be used as building blocks to more advanced methods that use global image signatures as input (e.g., ). We focus on the recent word attributes work of Almazán et al. , with available source code and state-of-the-art results in word image matching. This work uses Fisher vectors on SIFT descriptors as a building block to predict character attributes. The character attributes represent the presence or absence of a given character at a given relative position of the word (e.g., “word contains an a in the second half of the word” or “word contains a d in the first third of the word”). Word images are then described by the predicted attribute scores. This attribute representation is then projected into an embedded space correlated with embedded text strings using CCA, which improves its discriminative power while reducing its dimensionality. In our experiments, we will use this to produce global image signatures of only 9696 dimensions.

A great advantage of this framework is that representations of images and strings can then be compared using a cosine similarity, providing a unified framework to perform query-by-example (QBE) matching (i.e., retrieve images of a dataset given a query image), query-by-string (QBS) matching (i.e., retrieve images of a dataset given a query text string), and recognition (i.e., retrieve text strings given a query image). Since the approach is already based on Fisher vectors, it is easy to replace the SIFT descriptors in the pipeline of with our mid-level features and measure exactly their contribution.

Experiments

We start by describing the data used for learning the mid-level features and for evaluation purposes. We then describe the evaluation protocols and report the experimental results.

We evaluate our approach on two public benchmarks: IIIT5K and Street-View Text (SVT) . IIIT5K is the largest public annotated scene-text dataset to date, with 5,0005,000 cropped word images: 2,0002,000 training words and 3,0003,000 testing words. Each testing word is associated with a small text lexicon (SL) of 5050 words and a medium text lexicon (ML) of 1,0001,000 words used for recognition tasks. Note that each word has a different lexicon associated with it. SVT is another popular dataset with about 350 images harvested from Google Street View. These images contain annotations of 904904 cropped words: 257257 for training purposes and 647647 for testing purposes. Each testing image has an associated lexicon of 5050 words. We also construct a combined lexicon (CL), that contains every possible word that appears in the lexicons of each dataset (1,7871,787 unique words on IIIT5K and 4,2824,282 on SVT).

Learning dataset.

Learning the proposed transformations requires gathering training words annotated at the character level. Fortunately, pixel level annotations and character bounding boxes exist for several standard datasets . We gathered annotations for ICDAR 2003 , Sign Recognition 2009 , ICDAR 2011 , and IIIT5K – for IIIT5K, only the training set annotations were used. In total, we gathered 3,8293,829 words annotated with approximately 22,50022,500 character bounding boxes that are used to learn the mid-level features transformation.

2 Implementation details

To construct our block FVs, we extract SIFTs at 6 different scales, project them with PCA down to 6464 dimensions, and aggregate them using FVs (gradients w.r.t. means and variances) with 88 Gaussians and a spatial pyramid of 2×22\times 2, leading to a dimensionality DvD_{v} of 2×64×8×4=4,0962\times 64\times 8\times 4=4,096. During training, we sample 150150 blocks per training image. In total, we sampled approximately 600,000600,000 blocks. To learn the projection matrix UU with CCA, we use a regularization of η=1e−4\eta=1e^{-4} in all our experiments.

To construct the baseline global representation based only on SIFT (with no mid-level features) we follow a very similar approach, but instead of computing the mid-level blocks and reducing their dimensionality down to 6262 dimensions with CCA, we use the SIFT features directly, reducing their dimensionality down to 6262 dimensions with PCA and appending the normalized xx and yy coordinates of the center of the patch. This mimics the setup of the word attributes framework of . We also experiment with reducing the dimensionality of our mid-level features in an unsupervised manner with PCA instead of CCA, to separate the influence of the extra layer of the architecture from the supervised learning.

When using word attributes, one has control of the final dimensionality of the representations. We set this output dimensionality to 9696 dimensions. We observed that, in general, accuracy reached a plateau around that point: after that, increasing the number of dimensions does not significantly affect the performance. We found only one exception, where increasing the number of output dimensions beyond 9696 in one of the SVT experiments significantly improved the results.

3 Evaluation

We evaluate our approach with two different setups. In the first one, we are interested in observing the effect of using supervised and unsupervised mid-level features instead of SIFT descriptors directly when computing global word image representations, without applying any further supervised learning. We compute Fisher vectors using a) SIFT features, b) unsupervised mid-level features (i.e., dimensionality reduction of the block FVs with PCA), and c) supervised mid-level features (i.e., dimensionality reduction with CCA). In this last case, we also explore the effect of the number of character regions (CR) used to compute the labels during training (cf. Equation (2)), from a 1×11\times 1 to a 4×44\times 4 region split. These FVs can be compared using the dot-product as a similarity measure.

We measure the accuracy in a query-by-example retrieval framework, where one uses a word image as a query, and the goal is to retrieve all the images of the same word in the dataset. We use each image of the test set in a leave-one-out fashion to retrieve all the other images in the test set. Images that do not have any relevant item in the dataset are not considered as queries.

We report results in Table 1 using both mean average precision and precision at one as metrics. We highlight three aspects of the results. First, using supervised mid-level features significantly improves over using SIFT features directly, showing that the representation encodes more semantic information. Second, the improvement in the mid-level features comes from the supervised information and not from the extra layer of the architecture: the unsupervised mid-level features perform worse than the SIFT baseline. And, third, encoding which part of the characters the blocks overlap with is more informative than only encoding which characters they overlap with. We will use the 4×44\times 4 split for the rest of our experiments.

In the second set of experiments, we are interested in measuring how state-of-the-art global word image representations can benefit from these supervised features. In particular, we focus on the recent word attributes work of Almazán et al. , described in Section 3.2. We also explore the option of combining both FV representations (one based on SIFT and the other based on mid-level features), since their information may be complementary. To do so, we concatenate the global representations based on SIFT and mid-level features before learning the word attributes. To ensure the fairness of the comparisons, we learned the word attributes of using a dataset comprised of the learning dataset described in Section 4.1 plus IIIT5K and SVT (excluding the test set of the target dataset). In this manner, word attributes based on SIFT and word attributes based on mid-level features have been trained using exactly the same images, and only the type of annotations (text transcriptions only vs text transcriptions and character bounding boxes) differs.

As in , we evaluate on three tasks: query-by-example (QBE) – i.e., image-to-image retrieval –, query-by-string (QBS) – i.e., text-to-image –, and recognition –i.e., image-to-text. The accuracy of the QBE and QBS tasks is measured in terms of mean average precision. As is standard practice, we evaluate recognition using the small lexicon (SL) in IIIT5K and SVT, and the medium lexicon (ML) in IIIT5K. We also evaluate on our more challenging combined lexicon (CL). The recognition task is measured in terms of precision at 11.

Retrieval results. In Table 2 we report the retrieval results on IIIT5K and SVT. On both datasets using supervised mid-level features significantly improves over using SIFT features directly both on the query-by-example and query-by-string tasks. Combining the mid-level features and SIFT yields even further improvements on SVT. To the best of our knowledge, the best reported results on query-by-example and query-by-string on both datasets were those of , and using mid-level features significantly improves those results.

Recognition results. Table 3 shows the recognition results of our approach compared to the state-of-the-art. On IIIT5K, held the best results, and these results are improved thanks to the mid-level features. In the more difficult medium and combined lexicons, these improvements are very noticeable: from 82.0782.07 to 85.9385.93 and from 77.7777.77 to 83.0383.03. A similar trend can be observed on SVT, where the improvement on the combined lexicon is very significant. Increasing the output dimensionality from 9696 to 192192 dimensions also significantly improves the accuracy on the SVT small lexicon recognition task up to an 91.81%91.81\%. We only observed this behavior in this particular case; other datasets and tasks do not require larger representations.

The best reported results on SVT are those of and , both of which use millions of training samples. Using a fraction of the training data we obtain results better than PhotoOCR , but are outperformed by Jaderberg et al. . However, as seen on Figure 4, the performance of our method has not saturated: using more data as does will likely increase the accuracy on both datasets. Our method also has other advantages such as leading to tiny (9696-192192 dimensions) signatures that can be used for image-to-image and text-to-image matching. Compared to methods that do not use millions of training samples , our results are significantly better.

Qualitative results. Figure 3 shows qualitative results of the query-by-example task for some IIIT5K queries. In the more difficult queries (e.g. “Before” or “HOUSE”) the mid-level features are clearly superior. In general, even when the retrieved results are not correct, they are closer to the query than when using SIFT features directly.

Conclusions

In this paper we have introduced supervised mid-level features for the task of word image representation. These features are learned by leveraging character bounding box annotations at training time, and correlate visual blocks with the characters (and, most importantly, the character regions) in which such blocks tend to appear. Despite using character information at training time, one key advantage of our approach is that it does not require localizing characters explicitly at testing time. Instead, our mid-level features can be densely extracted in an efficient manner. We used these mid-level features as a building block of the word attributes framework of . The proposed mid-level features outperform equivalent representations based on SIFT on two standard benchmarks using tiny signatures of only 9696 dimensions, and obtain state-of-the-art results on retrieval and recognition tasks. We finally note that the proposed approach can be seen as a way to learn mid-level features from annotated parts at training time (characters and character regions in our case) without explicitly localizing them at testing time. We believe that the key ideas behind these mid-level features are not limited to text, and could be exploited well beyond the scene-text domain, e.g., for generic object categorization.

Appendix A Efficient Word Representation

This appendix describes an approach to compute all mid-level features of a new image exactly in an efficient manner.

Unfortunately, this process is not feasible in practice, as it would take a long time to compute and project all the FVs for all the blocks independently in a naive way. Instead, we propose an approach to compute these descriptors exactly in an efficient manner. The key idea is to isolate the contribution of each individual SIFT feature towards the mid-level feature of a block that included such a SIFT feature. If that contribution is linear, one can compute the individual contribution of each SIFT feature in the image and accumulate them in an integral representation. Then, the mid-level representation of an arbitrary block could be computed with just 22 additions and 22 subtractions over vectors, and computing all possible descriptors of all possible blocks becomes feasible and efficient.

In our experiments we use images of 120120 pixels in height, a step size of p=4p=4, a block spatial pyramid of 2×22\times 2, and K=62K=62 projections. We also extract blocks at 55 different block sizes: 16×1616\times 16, 24×2424\times 24, …, 48×4848\times 48. With this setup, we can extract and describe all blocks in an image in less than a second using a MATLAB implementation and a single core.

References