Joint Visual Semantic Reasoning: Multi-Stage Decoder for Text Recognition

Ayan Kumar Bhunia, Aneeshan Sain, Amandeep Kumar, Shuvozit Ghose, Pinaki Nath Chowdhury, Yi-Zhe Song

Introduction

Text recognition has been a popular area of research for decades thanks to its wide range of commercial applications , from translation apps in mixed reality, street signs recognition in autonomous driving to assistive technology for the visually impaired , to name a few. Significant progress in fundamental deep learning components alongside sequence-to-sequence learning frameworks , have boosted unconstrained word recognition accuracy (WRA) in recent times. Despite such developments, state-of-the-art text recognition frameworks still struggle in wild scenarios due to complex backgrounds, varying fonts, uncontrolled illuminations, distortions and other artifacts. While machines struggle with a combination of these challenges, humans recognise them easily via joint visual-semantic reasoning. Therefore, the question in focus is – how to develop a visual-semantic reasoning skill for text recognition?

State-of-the-art text recognition systems mostly rely on extracted visual features to recognize a word image as a machine readable character-sequence. Follow-up efforts have been made towards improving reasoning ability by increasing the depth of convolutional feature extractor having larger receptive fields, or introducing pyramidal pooling and stacking multiple Bi-LSTM layers . Despite all these attempts that merely lead towards a better context modeling , a semantic reasoning potential is largely missing beyond enriching the visual feature. In wild scenarios, a word image might be blurred, distorted, partly noisy or have artifacts, making recognition extremely difficult using visual feature alone. In such cases, we humans first try to interpret the easily recognizable characters using visual cues alone. A semantic reasoning skill is then applied to decode the final text by jointly processing the visual and semantic information from previously recognized character sequence. Motivated by this intuition, we propose a novel multi-stage prediction paradigm for text recognition. Here the first stage predicts using visual cues, while subsequent stages refine on top of it using joint visual-semantic information, by iteratively building up the estimates.

Designing this joint visual-semantic reasoning framework for text recognition is non-trivial. One might argue that attentional decoder being a sequence-to-sequence model, encapsulates the character dependency and caters for semantic reasoning. However, due to its auto-regressive nature , only those characters predicted previously, could provide semantic context at a given step, thus making the semantic context flow unidirectional during inference. While semantic context becomes negligible towards the initial steps, one wrong prediction here would deal a cumulative adverse impact on the later steps (which stays unrefined due to single stage prediction). Therefore, this single stage attentional decoder fails to model the global semantic context, leaving joint visual-semantic reasoning unaccomplished. To explore the entire global semantic context, we need the completely unrolled prediction from first stage, upon which we can build up the global semantic information. Hence as our first contribution we propose a multi-stage attentional decoder (Figure 1), where we build up global semantic reasoning on the initial estimate of first stage, which is further refined by subsequent stages.

Let us consider the word ‘aeroplane’. For a single stage attentional decoder, if the model predicts ‘n’ instead of ‘r’, ‘aen’ would adversely affect rest of the prediction, without any chance of refinement (being single stage). Also, it holds almost negligible semantic context while predicting the first few characters. Considering we unroll the prediction stage-wise, if a character is predicted wrongly, like ‘aenoplane’, rest of the characters provide significant context as semantic information. This helps in refining ‘n’ to ‘r’ during the later stages coupled with visual information.

Moreover, obtaining the prediction from earlier stages, needs a non-differentiable argmax operation as characters are discrete tokens. This leads to an inefficient modelling of influence of a prior stage on the next predictions. An apparent approach here might be to adapt teacher forcing for the later stages during training. The later stages intend to learn how to refine the initial (might be incorrect) hypothesis towards a correct prediction. This motivation however is defeated on feeding exact ground-truth labels as teacher forcing for subsequent stages. Consequently, we make use of Gumbel-Softmax operation bypassing non-differentiability, and making the network end-to-end trainable even across stages.

In summary our contributions are: First and foremost, we propose a multi-stage character decoding paradigm with stage-wise unrolling. While the first stage predicts using visual features, subsequent stages refine on-the-top of them using joint visual-semantic information. Secondly, we employ a Gumbel-softmax layer to make visual-to-semantic embedding layer differentiable. The model thus learns its refining strategy from initial to final prediction in an end-to-end manner. Thirdly, from the architectural design, we introduce multi-scale 2D attention to deal with varying scales of character size, and empirically found adding dense and residual connection between different stages stabilize training for better performance leading to outperforming other state-of-the-arts significantly on benchmark datasets.

Related Works

Text Recognition: While connectionist temporal classification (CTC) layer does not model dependency in the output character space , an attention based decoder encases language modeling, weakly supervised character detection and character recognition in a single paradigm. Following some seminal works , attention based decoder became state-of-the-art pipeline for text recognition which includes four successive modules: i) a rectification network to simplify irregular text image, ii) convolutional encoder for feature extraction, iii) Bi-LSTM layer for context modeling, and iv) an attentional decoder predicting the characters autoregressively.

Furthermore, the motivation of recent followed-up works can broadly be classified into following directions: (i) improve rectification network by introducing iterative pipeline and modelling geometrical attributes of text image; (ii) four directional feature encoder for better convolutional feature extraction; (iii) improving attention mechanism by extending to 2-D attention and hard character localized annotation , to better guide the attention based character alignment process. (iv) Recently, stacking multiple Bi-LSTM layers and pyramidal pooling on convolutional feature were employed towards the goal of better context modeling. These approaches however mainly focus on exploiting visual features, via different architectural modifications on top of Shi et al. , but mostly lack in any semantic reasoning capabilities.

Although some works claim to model semantic reasoning by stacking additional Bi-LSTM layers , it only helps in modelling better contextual information without having actual reasoning potential. In this context, word-embeddings from pre-trained language model were used to initialize the hidden state of attentional decoder, however we are skeptical towards this. For e.g. two related words “Chair” and “Table” may lie close in word-embedding space, but their character combination is way apart, thus questioning usage of word-embedding for text recognition. Yu et al.’s architectural design in this direction, gets severely limited on using argmax operation in visual-to-semantic embedding layer which invokes non-differentiability, restricting gradient flow from final prediction layer through this block; making learning deficient (Section 4.1). To our belief, ours is the first work employing a fully-differentiable semantic reasoning block that caters multi-stage refining objective for discrete character sequence prediction task.

Multi-Scale Learning: This learning paradigm is widely prevalent in object detection , recognition and semantic segmentation . Instead of solely relying on low resolution, semantically strong features, multi-scale framework sike MSCNN , DAG-CNNs , and FPN combine them with high-resolution, semantically weak features for object detection across a diverse range of shape and sizes. We couple multi-scale feature to generate multi-scale attention vectors for text recognition.

Multi-Stage Frameworks: In spite of computational overhead, multi-stage framework has gained popularity in computer vision task like pose estimation , object detection and action recognition for significantly improved performance. Specifically, Convolutional Pose Machine is one of the most successful and widely accepted multi-stage deep frameworks for pose-estimation.

Joint Visual-Semantic Learning: Recently, Graph Convolution Networks achieved success in object detection , image-text matching , image captioning by generating enhanced visual features with local and global semantic relationship. In our work, we use transformer network for joint visual semantic reasoning.

Methodology

Overview: Given an input word image II, we intend to predict the character sequence Y={y1,y2,...,yT}Y=\{y_{1},y_{2},...,y_{T}\}, where TT denotes the variable length of text. Our framework is two-fold: (i) a visual feature extractor extracts context-rich holistic feature and multi-scale feature maps. (ii) Following that, a multi-stage attentional decoder builds up the character sequence estimates, in a stage-wise successive manner. While dealing with irregular/curved word images , image rectification based approaches often fall short . To do away with the burden of adding a separate sophisticated rectification network entirely, we follow a 2D attention mechanism that helps to localize individual character in a weakly-supervised manner during decoding.

2 Joint Visual-Semantic Reasoning Decoder

Here, “⊛\circledast” and “⊗\otimes” denote convolution and matrix multiplication respectively. WBW_{B}, WHW_{H}, WattnW_{attn} are the learnable weights. Usually, Qt=Ht−1\mathbf{Q_{t}=H_{t-1}} containing history of prediction information is used as a query to locate yty_{t}. Moreover, query vector enriched in global semantic information (e.g. as in s≥1s\geq 1) could also be used instead, for better performance. While calculating the attention weight αi,j\alpha_{i,j} at every spatial position (i,j)(i,j), we employ a convolution operation with 3×33\times 3 kernel WBW_{\mathcal{B}} to consider the neighborhood information in 2D attention mechanism.

2.2 Decoder Stage 𝐬=𝟎𝐬0\mathbf{s=0}

2.3 Decoder Stage 𝐬≥𝟏𝐬1\mathbf{s\geq 1}

Fundamentally, there are three differences compared to basic attentional decoder (Eqn. 1):

(ii) For gtsg_{t}^{s}, we additionally use joint visual semantic information μts\mathbf{\mu_{t}^{s}} for query; thus Qts=[μts,Ht−1s]\mathbf{Q_{t}^{s}=[\mu_{t}^{s},H_{t-1}^{s}]}, and higher resolution feature-map is used as B=BL−s\mathcal{B}=B_{L-s} (e.g., BL−1B_{L-1} for s=1s=1) to couple multi-scale feature learning in a multi-stage decoder. Thus glimpse vector is gts=ψs(BL−s,[μts,Ht−1s])\mathbf{g_{t}^{s}=\psi^{s}(B_{L-s},[\mu_{t}^{s},H_{t-1}^{s}])}.

(iii) While s=0s=0 acts following baseline attentional decoder (Eqn. 1), the role for s≥1s\geq 1 is to learn refining strategy over previous predictions. Thus instead of feeding previous time-step prediction yt−1sy_{t-1}^{s}, we feed prediction from previous stage corresponding to the same time-step as yts−1y_{t}^{s-1}.

Visual-Semantic Reasoning: The visual and semantic reasoning functions ϕ(⋅)\phi(\cdot) and ω(⋅)\omega(\cdot) are employed by Transformer module that uses multi-headed self-attention mechanism to gather global context information. In brief, given key (K), query (Q) and value (V), attention is calculated as: Attention(K,Q,V)=softmax(QK⊺dim)V\textit{Attention}(K,Q,V)=\textit{softmax}(\frac{QK^{\intercal}}{\sqrt{dim}})V. At each time step output ϕt(⋅)\phi_{t}(\cdot) and ωt(⋅)\omega_{t}(\cdot), feature representation is enriched by information from remaining time-steps and thus long-range dependencies are modelled carefully. Semantic reasoning module ω(⋅)\omega(\cdot) is pre-trained separately following BERT language model training topology. We mask out (also purposefully replace by erroneous instances) certain input time steps and force to predict masked token by a linear layer. This helps the model to learn better refining potential using text-only data in advance.

3 Learning Objective

We accumulate cross-entropy loss from all stages of attentional decoder to train our text-recognition model.

Experiments

Datasets: Following the similar approach described in , we train our model on synthetic datasets (without any further fine-tuning) such as SynthText and Synth90k , which holds 6 and 8 million images respectively. The evaluation is performed without fine-tuning on datasets containing real images like: Street View Text (SVT), ICDAR 2013 (IC13), ICDAR 2015 (IC15), CUTE80, SVT-Perspective (SVT-P), IIIT5K-Words. Street View Text dataset consists of 647 images, most of which are blurred, noisy or have low resolution. While ICDAR 2013 has 1015 words, ICDAR 2015 contains a total of 2077 images of which 200 images are irregular. CUTE80 offers 288 cropped high quality curved text images. SVT-Perspective presents 645 samples from side-view angle snapshots containing perspective distortion. IIIT5K-Words distinguishes itself by presenting randomly picked 3000 cropped word images.

Implementation Details: We use ResNet architecture from with FPN heads having 256 channels in each multi-scale feature-maps. The kernel size of intermediate pooling layers is so adjusted that BL,BL−1,BL−2B_{L},B_{L-1},B_{L-2} have spatial size of 4×254\times 25, 8×258\times 25, and 16×5016\times 50 respectively. The hidden state size of two-layer encoder BLSTM and each decoder LSTM is kept at 256. Semantic (ω)(\omega) and visual (ϕ)(\phi) reasoning blocks consist of 2 stacked transformer units with 4 heads and hidden state size 256. The hidden units in attention block is of size 128. A total of 37 classes are taken including alphatbets, numbers and end-tokens; with the maximum sequence length (N) set to 25. We use ADADELTA optimizer with learning rate 1.0 and batch size 32. We resize the image to 32x100 and train our model in a 11 GB NVIDIA RTX-2080-Ti GPU using PyTorch. We first warm-up using single stage attentional decoder for 50K iterations, and then train our proposed three-stage (S={0,1,2}S=\{0,1,2\}) attentional decoder (ablation on optimal stages in Sec. 4.2) framework end-to-end, for 600K iterations with λ1,λ2,λ3\lambda_{1},\lambda_{2},\lambda_{3} set to 1, 0.1, 0.1 respectively. Please note that the first stage is fed with one-time step shifted ground truth label to accommodate teacher forcing in sequence modeling, however, later stages are fed with model’s prediction from previous stage in order to learn the data driven refining strategy.

Table 1 shows our proposed method to surpass SOTA methods by a reasonable margin. Every method’s salient contributions are briefly mentioned there as well. In this section, we first describe the limitations of the existing or alternative (naive) designs and then illustrate (using IC15) how and why all our design components/choices contribute towards superiority over others.

[i] Limitation of previous attentional decoders: Existing methods relying on unidirectional auto-regressive attentional decoders exhibit a bottleneck, and its drawback becomes evident from the following scenario : An easily recognizable character present towards the end of a word would fail to provide any contextual semantic information towards recognizing some noisy character present earlier. We on the contrary let the first stage completely unroll itself. Thereafter the prediction of previous stage (even if certain time-step’s character is incorrect) could be rectified in the subsequent stages using joint visual-semantic information. Although SCATTER stacks multiple BLSTM layers on the top of baseline design from ASTER , both methods lack semantic reasoning as they barely enrich visual feature encoding. Examples from our stage-wise decoder are shown in Figure 3.

[ii] Significance of Differentiable Semantic Space: Improving semantic reasoning for better text recognition was only considered by and among all SOTA methods. Although Qiao et al. proposed to use word embedding, such technique relies on semantic meaning of a word instead of the required character sequence. For example, the word “table” and “chair”, although semantically related have character combinations that are way-apart. Therefore, we emphasise on modelling character sequences instead, to help recognize a noisy character based on two-way information passing. Even though Yu et al. took this direction to some extent, their non-differentiable semantic-reasoning block imposes a significant limitation. We alleviate that with the help of gumbel-softmax to develop a differentiable semantic space and allow learning of multi-stage semantic reasoning. While the use of teacher forcing for later stages by feeding ground-truth label for training multi-stage decoder might seem an alternative, empirical evidence suggests otherwise. The third stage decoder obtains 74.4%74.4\% accuracy as compared to 74.5%74.5\% accuracy (on IC15) in first stage – no practical gains. Another straight-forward way is to use straight-through estimator , which simply copies gradients from argmax output to the next input. However, this results in significant instability where later stage performance drops by 3.9%3.9\% to 80.1%80.1\% due to discrepancies between forward and backward passes resulting in much higher variance than gumbel-softmax .

[iv] Why use top-down attentional decoder: While low resolution and semantically strong features are good for classification, tasks requiring focus in local regions, such as object detection and semantic segmentation, benefit even further when combined with high-resolution semantically weak features found in shallower regions of a feature extractor . Although our first stage is similar to a basic attentional decoder focusing on feature map of the last layer to benefit from rich semantic information, that is more invariant to distortion, later stages (refining stages) combine higher resolution feature-map from preceding layers. This not only handles varying character size, but also verifies prior prediction by exploiting joint information between high resolution feature and previous predictions to guide the refining process. This hypothesis is verified by contradiction, using high-resolution semantically weak feature BL−2B_{L-2} in s=0s=0 and lower resolution semantically strong features in later stages s>1s>1. We observe performance collapses to 72.1%72.1\% in IC15 dataset due to inability of high resolution semantically weak features to output the initial estimates.

[v] Significance of self-attention based Joint Visual-Semantic Reasoning: To emulate human-like inference, self-attention based reasoning functions allow two way information passing across visual and semantic spaces to obtain a joint visual-semantic context. Its significance could be empirically understood by removing the visual reasoning block and modifying the architecture accordingly, which drops result by 2.9%2.9\%. A similar drop of 4.8%4.8\% was observed when the semantic reasoning block was removed. On removing both we observe 77.1%77.1\% accuracy – a significant drop of 6.9%6.9\% from our method (Table 2).

[vi] Do multi-scale (resolution) feature maps help? We empirically validate this by excluding multi-scale feature maps and use BLB_{L}, instead of BL−sB_{L-s}, to calculate gtsg^{s}_{t} at every stage ss. Such modification drops performance by 2.7%2.7\% (against ours), to 81.3%81.3\%, which highlights the contribution of multi-scale feature maps in our method.

[vii] Comparison with alternative multi-scale attentional decoder designs: In text recognition, the only other work realising importance of multi-scale information is by Wan et al. , where pyramid pooling was used. Here visual feature maps from different spatial resolutions were concatenated, which eventually harmed downstream tasks owing to the large semantic gaps between such feature maps. Consequently, we introduce lateral connections following Feature Pyramid Networks , semantically strengthening high-resolution levels for superior performance. Simply employing pyramid pooling for all stages s={0,1,2}s=\{0,1,2\} however, drops performance by 2.1%2.1\% (against ours) to 81.9%81.9\% .

[viii] Significance of Dense and Residual Connections: Beside improving visual information flow in the forward pass, the residual connection between initial Ht0H_{t}^{0} and final HtSH_{t}^{S} ensures efficient gradient flow in visual feature networks, accelerating convergence of the whole network. Furthermore, the dense connection is used to adaptively learn a more discriminative glimpse vector by combining its features from preceding stages with the current one, thus stabilising the training of multi-stage multi-scale attentional decoder. Removing dense connection (gtg_{t} calculation) decreases the performance by 1.6%1.6\%, and removing residual connection decreases it by 1.3%1.3\%. On removing both we get an even larger drop of 1.9%1.9\%. Faster training is observed while using both dense and residual connections.

[ix] Significance of Multiple Constraints: We design experimental setups (Table 2) that reveal the following observations: (a) imposing loss LCL_{C} only in the last stage harms the model, resulting in 73.1%73.1\% accuracy. We attribute this to the poor gradient flow across stages. (b) Adding multi-stage LCL_{C} loss results in 77.1%77.1\% accuracy, performing closer to the proposed method. (c) Adding visual-semantic constraints LVL_{V} and LSL_{S} finally gives the best performance of 84.0%84.0\%. This shows multi-stage constraint is vital for training and convergence. The intuition behind multiple constraints sources from multi-task learning, which ensures better convergence, thus enriching individual character aligned feature, with better visual-semantic information.

[ix] Varying training data size: Following , we also vary the training size and evaluate our proposed framework compared to single stage baseline and Yu et al. in Table 2. Significant overhead at low data regime brings the superiority to our proposed method over others.

2 Further Analysis and Insights

[i] Design of Visual-Semantic Reasoning Module: One can capture two-way visual semantic information using (a) Bi-LSTM (b) Transformer with multi-headed self-attention mechanism. Table 3 shows Transformer to outperform LSTM by 1.3%1.3\%. Furthermore, pre-training global semantic reasoning module ω(⋅)\omega(\cdot) using BERT like training topology, scores 0.9%0.9\% higher accuracy than without it.

[iii] Computational Analysis: Each stage needs to unroll itself completely, before the next starts processing. Hence, the performance gain comes at a cost of extra computational expenses (analysis in Table 4), which is reasonable given the superior performance over strong baselines. Even so, we experimented with ResNet-101 as a backbone feature extractor, having similar number of parameters and flops to ours. This naive stacking of multiple-layers lags by 8.9%, which accredits our gain to our novel design choice.

[iv] Comparison with SOTA Language Model: We compare our framework with state-of-the-art Language Modeling (LM) based post-processing techniques based on librispeech text-corpus. Based on we adopt two techniques: (a) Shallow Fusion that results in 74.3%74.3\% and (b) Deep Fusion giving 75.9%75.9\% accuracy on IC15 (Table 3).

[v] Optimum Stages: The optimal value for the number of stages ss is found empirically on IC15. For s=1s=1 we have 80.3%80.3\% accuracy that improves at s=2s=2 to give 84.0%84.0\%, but saturates at s=3s=3 giving 83.6%83.6\%. Hence we consider s=2s=2 to be optimal. This performance saturation could be attributed to vanishing gradient problem which is addressed via residual/dense connection, but still persists to some extent. Also, for s>2s>2, the joint visual-semantic information might reach its optimum, where the result saturates. Please refer to supplementary material as well.

Conclusion

We propose a novel joint visual-semantic reasoning based multi-stage multi-scale attentional decoding paradigm. The first stage predicts from visual features, followed by refinement using joint visual-semantic information. We further exploit Gumbel-softmax operation to make visual-to-semantic embedding layer differentiable. This enables backpropagation across stages to learn the refining strategy using joint visual-semantic information. Experimental results indicate the superior efficiency of our model.

References