Learning Robust Visual-Semantic Embeddings

Yao-Hung Hubert Tsai, Liang-Kang Huang, Ruslan Salakhutdinov

Introduction

Over the past few years, due to the availability of large amount of data and the advancement of the training techniques, learning effective and robust representations directly from images or text becomes feasible . These learned representations have facilitated a number of high-level tasks, such as image recognition , sentence generation , and object detection . Despite useful representations being developed for specific domains, learning more comprehensive representations across different data modalities remains challenging. In practice, more complex tasks, such as image captioning and image tagging often involve data from different modalities (i.e., images and text). Additionally, the learning process would be faster, requiring fewer labeled examples, and hence more scalable to handling a large number of categories if we could transfer cross-domain knowledge more effectively . This motivates learning multi-modal embeddings. In this paper, we consider learning robust joint embeddings across visual and textual modalities in an end-to-end fashion under zero and few-shot setting.

Zero-shot learning aims at performing specific tasks, such as recognition and retrieval of novel classes, when no label information is available during training . On the other hand, few-shot learning enables us to have few labeled examples in our of-interest categories . In order to compensate the missing information under the zero and few-shot setting, the model should learn to associate novel concepts in image examples with textual attributes and transfer knowledge from training to test classes. A common strategy for deriving the visual-semantic embeddings is to make use of images and textual attributes in a supervised way . Specifically, one can learn transformations of images and textual attributes under the objective that the transformed visual and semantic vectors of the same class should be similar in the joint embeddings space. Despite good performance, this common strategy basically boils down to a supervised learning setting, learning from labeled or paired data only. In this paper, we show that to learn better joint embeddings across different data modalities, it is beneficial to combine supervised and unsupervised learning from both labeled and unlabeled data.

Our contributions in this work are as follows. First, to extract meaningful feature representations from both labeled and unlabeled data, one possible option is to train an auto-encoder . In this way, instead of learning representations directly to align the visual and textual inputs, we choose to learn representations in an auto-encoder using reconstruction objective. Second, we impose a cross-modality distribution matching constraint to require the embeddings learned by the visual and textual auto-decoders to have similar distributions. By minimizing the distributional mismatch between visual and textual domain, we show improved performance on recognition and retrieval tasks. Finally, to achieve better adaptation on the unlabeled data, we perform a novel unsupervised-data adaptation inference technique. We show that by adopting this technique, the accuracy increases significantly not only for our method but also for many of the existing other models. Fig. 1 illustrates our overall end-to-end differentiable model.

To summarize, our proposed method successfully combines supervised and unsupervised learning objectives, and learns from both labeled and unlabeled data to construct joint embeddings of visual and textual data. We demonstrate improved performance on Animals with Attributes (AwA\mathsf{AwA}) and Caltech-UCSD Birds 200-2011 datasets on both image recognition and image retrieval tasks under zero and few-shot setting.

Related Work

In this section, we provide an overview of learning multi-modal embeddings across visual and textual domain.

Zero-shot and few-shot learning are related problems, but somewhat different in the setting of the training data. While few-shot learning aims to learn specific classes through one or few examples, zero-shot learning aims to learn even when no examples of the classes are presented. In this setting, zero-shot learning should rely on the side information provided by other domains. In the case of image recognition, this often comes in the form of textual descriptions. Thus, the focus of zero-shot image recognition is to derive joint embeddings of visual and textual data, so that the missing information of specific classes could be transferred from the textual domain.

Since the relation between raw pixels and text descriptions is non-trivial, most of the previous work relied on learning the embeddings through a large amount of data. Witnessing the success of deep learning in extracting useful representations, much of the existing work mostly applies deep neural networks to first transform raw pixels and text into more informative representations, followed by using various techniques to further identify the relation between them. For example, Socher et al. used deep architectures to learn representations for both images and text, and then used a Bayesian framework to perform classification. Norouzi et al. introduced a simple idea that treated classification scores output by the deep network as weights in convex combination of word vectors. Fu et al. proposed a method that learns projections from low-level visual and textual features to form a hypergraph in the embedding space and performed label propagation for recognition. A number of similar methods learn transformations from input image representations to the semantic space for the recognition or retrieval purposes .

A number of recent approaches also attempt to learn the entire task with deep models in an end-to-end fashion. Frome et al. constructed a deep model that took visual embeddings extracted by CNN and word embeddings as input, and trained the model with the objective that the visual and word embeddings of the same class should be well aligned under linear transformations. Ba et al. predicted the output weights of both the convolutional and fully connected layers in a deep convolutional neural network. Instead of using textual attributes or word embeddings model, Reed et al. proposed to train a neural language model directly from raw text with the goal of encoding only the relevant visual concepts for various categories.

Liu et al. developed multi-task deep visual-semantic embeddings model for selecting video thumbnails based on side semantic information (i.e., title, description, and query). By incorporating knowledge about objects similarities between visual and semantic domains, Tang et al. improved object detection in a semi-supervised fashion. Kottur et al. proposed to learn visually grounded word embeddings (vis-w2v) and showed improvements over text only word embeddings (word2vec) on various challenging tasks. Reed et al. designed a text-conditional convolutional GAN architecture to synthesize an image from text. Recently, Wang et al. introduced structure-preserving constraints in learning joint embeddings of images and text for image-to-sentence and sentence-to-image retrieval tasks.

Unsupervised Multi-modal Representations Learning

One of our key contributions is to effectively combine supervised and unsupervised learning tasks for learning multi-modal embeddings. This is inspired and supported by several previous works that provided evidence of how unsupervised learning tasks could benefit cross-modal feature learning.

Ngiam et al. proposed various models based on Restricted Boltzmann Machine, Deep Belief Network, and Deep Auto-encoder to perform feature learning over multiple modalities. The derived multi-modal features demonstrated an improved performance over single-modal features on the audio-visual speech classification tasks. Srivastava and Salakhutdinov developed a Multimodal Deep Boltzmann Machine for fusing together multiple diverse modalities even when some of them are absent. Providing inputs of images and text, their generative model manifested noticeable performance improvement on classification and retrieval tasks.

Proposed Method

First, we define the problem setting and the corresponding notation. Let Vtr={vi(tr)}i=1Ntr\mathbf{V_{tr}}=\{v_{i}^{(tr)}\}_{i=1}^{N_{tr}} denote labeled training images from CtrC_{tr} classes, Vut={vi(ut)}i=1Nut\mathbf{V_{ut}}=\{v_{i}^{(ut)}\}_{i=1}^{N_{ut}} denote unlabeled training images from CutC_{ut} possibly different classes, and Vte={vi(te)}i=1Nte\mathbf{V_{te}}=\{v_{i}^{(te)}\}_{i=1}^{N_{te}} denote test images from CteC_{te} novel classes. For each class, following , its textual attributes are either provided from human annotated attributes or learned from unsupervised text corpora (Wikipedia) . We denote these class-specific textual attributes as Ttr={tc(tr)}c=1Ctr\mathbf{T_{tr}}=\{t_{c}^{(tr)}\}_{c=1}^{C_{tr}}, Tut={tc(ut)}c=1Cut\mathbf{T_{ut}}=\{t_{c}^{(ut)}\}_{c=1}^{C_{ut}}, and Tte={tc(te)}c=1Cte\mathbf{T_{te}}=\{t_{c}^{(te)}\}_{c=1}^{C_{te}} for labeled training, unlabeled training, and test classes, respectively.

Under zero-shot setting, our goal is to predict labels of the test images coming from novel, previously unseen, classes given textual attributes. That is, for a given test image vi(te)v_{i}^{(te)}, its label is determined by

where θ\theta denotes model parameters. We will also consider a few-shot learning, where a few labeled training images are available in each of the test classes. In the following, we omit the model parameters θ\theta for brevity.

The goal of learning multi-modal embeddings can be formulated as learning transformation functions fvf_{v} and ftf_{t}, such that given an image vv and a textual attribute tt, fv(v)f_{v}(v) should be similar to ft(t)f_{t}(t) if vv and tt are of the same class. Much of the previous work for learning multi-modal embeddings can be generalized to this formulation. For instance, in Cross-Modal Transfer (CMT) , fv(⋅)f_{v}(\cdot) can be viewed as a pre-defined feature extraction model followed by a two-layer neural network, while ft(⋅)f_{t}(\cdot) is set to an identity matrix. To be more specific, aim at learning a non-linear projection directly from visual features to semantic word vectors.

Over the past few years, deep architectures have been shown to learn useful representations that could embed high-level semantics for both visual and textual data. This gives rise to the attempts of applying successful deep architectures to learn fv(⋅)f_{v}(\cdot) and ft(⋅)f_{t}(\cdot). For example, DeViSE designed fv(⋅)f_{v}(\cdot) as a CNN model followed by a linear transformation matrix. On the other hand, they adopted the well known skip-gram text modeling architecture for learning ft(⋅)f_{t}(\cdot) from raw text on Wikipedia. It is worth noting that, to further take advantage of previous success, these deep models are often pre-trained on large datasets where they have shown to learn effective representations.

Figure 2 shows the basic formulation of the visual-semantic embeddings model. Our method is built on top of this basic architecture by adding additional components as well as modifying existing ones.

2 Reconstructing Features from Auto-Encoder

where J(⋅)J(\cdot) is the Jacobian matrix .

On the other hand, for a given semantic feature vector or textual attribute tt, we use a vanilla auto-encoder to first encode and then reconstruct from its hidden representation tht_{h}. We hence minimize the reconstruction error

Combining (2) and (3) gives us the reconstruction loss

3 Cross-Modality Distributions Matching

We can now minimize the MMD criterion between visual and textual embeddings by minimize eq. (5). This can be further regarded as shrinking the gap between information across two data modalities. In our experiments, we find that the MMD loss helps improve model performance on both recognition and retrieval tasks in zero and few-shot setting.

4 Learning

where fv′(⋅)f^{\prime}_{v}(\cdot) and ft′(⋅)f^{\prime}_{t}(\cdot) are the mapping functions from the hidden representations to the visual and textual output.

To leverage the supervised information from labeled training images Vtr\mathbf{V_{tr}} and the corresponding textual attributes Ttr\mathbf{T_{tr}}, we minimize the binary prediction loss:

where Ii,cI_{i,c} indicates a {0,1}\{0,1\} encoding of positive and negative classes and ⟨⋅⟩\left\langle\cdot\right\rangle denotes a dot-product. It is worth noting that we can adopt different loss functions, including binary cross-entropy loss or multi-class hinge loss. However, empirically, using the simple binary prediction loss results in the best performance of our model.

Similar to eq. (8), we adopt the binary prediction loss for unlabeled training images Vut\mathbf{V_{ut}} and the attributes Tut\mathbf{T_{ut}}:

We refer to eq. (9) as unsupervised-data adaptation inference, which acts as a self-reinforcing strategy using the unsupervised data with unknown labels. The intuition is that by minimizing eq. (9), we can further adapt our unlabeled data into the learning of fv′(⋅)f^{\prime}_{v}(\cdot) and ft′(⋅)f^{\prime}_{t}(\cdot) based on the empirical predictions. The choice of λ\lambda does influence its effectiveness. However, we find that setting λ=1.0\lambda=1.0 works quite well for many methods we considered in this work.

In sum, our model learns by minimizing the total loss from both supervised and unsupervised objectives:

with α\alpha, λ\lambda, and β\beta representing the trade-off parameters for different components. Note that we can also view the unsupervised objective here as a regularizer for learning more robust visual and textual representations (see Figure 1 for our overall model architecture).

Experiments

In the experiments, we denote our proposed method as ReViSE (Robust sEmi-supervised Visual-Semantic Embeddings). Extensive experiments on zero and few-shot image recognition and retrieval tasks are conducted using two benchmark datasets: Animals with Attributes (AwA\mathsf{AwA}) and Caltech-UCSD Birds 200-2011 (CUB\mathsf{CUB}) . CUB\mathsf{CUB} is a fine-grained dataset in which the objects are both visually and semantically very similar, while AwA\mathsf{AwA} is a more general concept dataset. We use the same training (+validation)/ test splits as in . Table 1 lists the statistics of the datasets.

To verify the performance of our method, we consider two state-of-the-art deep-embeddings methods: CMT and DeViSE . CMT and DeViSE can be viewed as a special case of our proposed method with α=0\alpha=0 (without using unsupervised objective in eq. (11)). The difference between them is that DeViSE learns a nonlinear transformation on raw visual images and textual attributes for the alignment purpose, while CMT only learns the nonlinear transformation from visual to semantic embeddings.

We choose GoogLeNet as the CNN model in DeViSE, CMT, and our architecture. For the textual attributes of classes, we consider three alternatives: human annotated attributes (att\mathit{att}) , Word2Vec attributes (w2v\mathit{w2v}) , and Glove attributes (glo\mathit{glo}) . att\mathit{att} are continuous attributes judged by humans: CUB\mathsf{CUB} contains 312 attributes and AwA\mathsf{AwA} contains 85 attributes. w2v\mathit{w2v} and glo\mathit{glo} are unsupervised methods for obtaining distributed text representations of words. We use the pre-extracted Word2Vec and Glove vectors from Wikipedia provided by . Both w2v\mathit{w2v} and glo\mathit{glo} are 400-dim. features.

Please see Supplementary for the design details of ReViSE and its parameters. Note that we report results averaged over 1010 random trials.

2 Zero-Shot Learning

Following the partitioning strategy of , we split AwA\mathsf{AwA} dataset into 30/10/10 classes and CUB\mathsf{CUB} dataset into 100/50/50 classes for labeled training/ unlabeled training/ test data. We adopt att\mathit{att} attributes as a textual description of each class. For zero-shot learning, not only the labels of images are unknown in the unlabeled training and test set, but classes are also disjoint across labeled training/ unlabeled training/ test splits.

To verify how unlabeled training data could benefit the learning of ReViSE, we provide four variants: ReViSEa, ReViSEb, ReViSEc, and ReViSE. ReViSEa is when we only consider supervised objective. That is, α=0\alpha=0 in eq. (11). ReViSEb is when we further take unsupervised objective in labeled training data into account; that is, only Lreconstruct\mathcal{L}_{reconstruct} and LMMD\mathcal{L}_{MMD} are considered in Lunsupervised\mathcal{L}_{unsupervised} (see eq. (12)) for labeled training data. Next, for ReViSEc, we consider unlabeled training data in Lunsupervised\mathcal{L}_{unsupervised} without unsupervised-data adaptation technique (setting β=0\beta=0). Last, ReViSE denotes our complete training architecture.

For completeness, we also consider the technique of unsupervised-data adaptation inference (see section 3.4) for DeViSE and CMT . In other words, we also evaluate how DeViSE and CMT benefit from the unlabeled training data. We adopt the same procedure as in eq. (9) and report results as DeViSE* and CMT*, respectively.

Table 2 and 3 list the results for our zero-shot recognition and retrieval experiments. We first observe that NOT all the methods benefit from using unlabeled training data during training. For example, in AwA\mathsf{AwA} dataset for test images Vte\mathbf{V_{te}}, there is a 2.7%2.7\% retrieval deterioration from DeViSE to DeViSE* and a 3.1%3.1\% recognition deterioration from CMT to CMT*. On the other hand, our proposed method enjoys 2.6%2.6\% recognition improvement and 0.4%0.4\% retrieval improvement from ReViSEb to ReViSEc. This shows that the learning method of our proposed architecture can actually benefit from unlabeled training data Vut\mathbf{V_{ut}} and Tut\mathbf{T_{ut}}.

Next, we examine different variants in our proposed architecture. Comparing the average results from ReViSEa to ReViSEb, we observe 3.4%3.4\% recognition improvement and 3.9%3.9\% retrieval improvement. This indicates that taking unsupervised objectives Lreconstruct\mathcal{L}_{reconstruct} and LMMD\mathcal{L}_{MMD} into account results in learning better feature representations and thus yields a better recognition/ retrieval performance. Moreover, when unsupervised-data adaptation technique is introduced, we enjoy 5.0%5.0\% average recognition improvement and 5.8%5.8\% average retrieval improvement from ReViSEc to ReViSE. It is worth noting that the significant performance improvement for unlabeled training images Vut\mathbf{V_{ut}} further verifies that our unsupervised-data adaptation technique leads to a more accurate prediction on Vut\mathbf{V_{ut}}.

3 Transductive Zero-Shot Learning

In this subsection, we extend our experiments to a transductive setting, where test data are available during training. Therefore, the test data can now be regarded as the unlabeled training data (Vtr=Vut\mathbf{V_{tr}}=\mathbf{V_{ut}} and Ttr=Tut\mathbf{T_{tr}}=\mathbf{T_{ut}}). To perform the experiments, as in Table 1, we split AwA\mathsf{AwA} dataset into 40/10 disjoint classes and CUB\mathsf{CUB} dataset into 150/50 disjoint classes for labeled training/ test data.

In order to evaluate different components in ReViSE, we further provide two variants: ReViSE† and ReViSE††. ReViSE† is when we consider no distributional matching between the codes across modalities (β=0\beta=0). ReViSE†† is when we further consider no contractive loss in our visual auto-encoder (β=γ=0\beta=\gamma=0). Similar to previous subsection, we also consider DeViSE*, CMT*, and ReViSEc to evaluate the effect of our unsupervised-data adaptation inference.

Zero-Shot Recognition: Table 9 reports top-1 classification accuracy. Observe that ReViSE clearly outperforms other state-of-the-art methods by a large margin. On average, we have at least 17%17\% gain compared to the methods without using unsupervised objective and 7.5%7.5\% gain compared to DeViSE* and CMT*. Note that all the methods work better on human annotated attributes (att) than on unsupervised attributes (w2v and glo) in CUB\mathsf{CUB} dataset. One possible reason is that for visually and semantically similar classes in a fine-grained dataset (CUB\mathsf{CUB}), attributes obtained in an unsupervised way (glo word vectors) cannot fully differentiate between them. Nonetheless, for the more general concept dataset AwA\mathsf{AwA}, using either supervised or unsupervised textual attributes, the performance does not differ by that much. For instance, our method achieves comparable performance using att, w2v, and glo (93.4%93.4\%, 93.5%93.5\%, and 92.2%92.2\% top-1 classification accuracy) on AwA\mathsf{AwA} dataset.

The recognition performance for DeViSE* and CMT* (60.6%60.6\% and 60.5%60.5\% on average) compared to DeViSE and CMT (49.3%49.3\% and 50.5%50.5\% on average) further verifies that using unsupervised-data adaptation inference technique does benefit transductive zero-shot recognition. Furthermore, all of the variants of ReViSE using unsupervised-data adaptation inference (ReViSE††, ReViSE†, and ReViSE itself) have noticeable improvement over DeViSE* and CMT*. This shows that the proposed model succeeds in leveraging unsupervised information in test data for constructing more effective cross-modal embeddings.

Next, we evaluate the effects of different components designed in our architecture. First of all, we compare the results between ReViSE† (set β=0\beta=0) and ReViSE. The performance gain (66.8%66.8\% to 68.1%68.1\% on average) indicates that minimizing MMD distance between visual and textual codes enables our model to learn more robust visual-semantic embeddings. In other words, we can better associate cross-modal information when we match the distributions across visual and textual domains (please refer to Supplementary for the study of MMD distance). Second, we observe that, without contractive loss, performance slightly drop from 66.8%66.8\% (ReViSE†) to 65.8%65.8\% (ReViSE††). This is not surprising since the contractive auto-encoder aims at learning less varying features/codes with similar visual input, and therefore we can expect to learn more robust visual codes. Finally, similar to the observations found in comparing DeViSE/CMT to DeViSE*/CMT*, the unsupervised-data adaptation inference in ReViSE substantially improves the average top-1 classification accuracy from 53.6%53.6\% (ReViSEc) to 68.1%68.1\% (ReViSE). Please see Supplementary material for more detailed comparisons to the following non-deep-embeddings methods: SOC , ConSE , SSE , SJE , ESZSL , JLSE , LatEm , Sync , MTE , TMV , and SMS .

Zero-Shot Retrieval: In Table 5, we report zero-shot retrieval results by measuring the retrieval performance by mean average precision (mAP). On average, methods that leverage unsupervised information yield better performance compared to the methods using no unsupervised objective. However, in few cases, the performance drops when we take unsupervised information into account. For example, on CUB\mathsf{CUB} dataset, DeViSE* performs unfavorably compared to DeViSE when w2v and glo word embeddings are used as textual attributes.

Overall, our method does help improve zero-shot retrieval by at least 14.1%14.1\% compared to CMT*/DeViSE* and 21.5%21.5\% compared to CMT/DeViSE. It clearly demonstrates the effectiveness of leveraging unsupervised information for improving zero-shot retrieval (please see Supplementary for the plot of precision-recall curves).

In addition to quantitative results, we also provide qualitative results of ReViSE. Fig. 3 is the image retrieval experiments for classes Chestnut_sided_Warbler and White_eyed_Vireo. Given a class embedding, the nearest image neighbors are retrieved based on the cosine similarity between transformed visual and textual features. We consider two conditions: images from the same class and images from all test classes. In Chestnut_sided_Warbler, most of the images (71.7%71.7\%) are correctly classified, and we also observe that three nearest image neighbors are also in Chestnut_sided_Warbler. On the other hand, only 43.3%43.3\% images are correctly classified in White_eyed_Vireo, and two of the three nearest image neighbors are form wrong class Wilson_Warbler.

Availability of Unlabeled Test Images: We next evaluate the performance of our method w.r.t. to the availability of test images for unsupervised objective (see Fig. 4) on CUB\mathsf{CUB} dataset with att attributes. We alter the fraction pp of unlabeled test images used in the training stage from 0%0\% to 100%100\% by a step size of 10%10\%. That is, in eq. (12), only pp portion (randomly chosen) of test images contributes to Lunsupervised\mathcal{L}_{unsupervised}. Fig. 4 clearly indicates the performance increases when pp increases. That is, with more unsupervised information (test images) available, our model can better associate the supervised and unsupervised data. Another interesting observation is that with only 40%40\% test images available, ReViSE achieves favorable performance on both transductive zero-shot recognition and retrieval.

Expand the test-time search space: Note that most of the methods consider that, at test time, queries come from only test classes. For AwA\mathsf{AwA} dataset with att attributes, we expand the test-time search space to all training and test classes and perform transductive zero-shot recognition for DeViSE*, CMT*, and ReViSE. We discover severe performance drops from 90.7%90.7\%, 89.4%89.4\%, and 93.4%93.4\% to 47.4%47.4\%, 45.8%45.8\%, and 42.5%42.5\%. Similar results can also be observed in other non-deep-embeddings methods. Although challenging, it remains interesting to consider this generalized zero-/few-shot learning setting in our future work.

4 From Zero to Few-Shot Learning

In this subsection, we extend our experiments from transductive zero-shot to transductive few-shot learning. Compared to zero-shot learning, few-shot learning allows us to have a few labeled images in our test classes. Here, 3 images are randomly chosen to be labeled per test category. We use the same performance comparison metrics as in Sec. 4.2 to report the results.

Transductive Few-Shot Recognition and Retrieval: Tables 6 and 7 list the results of transductive few-shot recognition and retrieval tasks. Generally speaking, ReViSE achieves the best performance compared to its variants and other methods. Moreover, as expected, when we compare the results with transductive zero-shot recognition (Table 9) and retrieval (Table 5), every methods perform better when few (i.e., 33) labeled images are observed in the test classes. For example, for CUB\mathsf{CUB} dataset with w2v attributes, there is a 22.5%22.5\% recognition improvement for CMT* and a 32.3%32.3\% retrieval improvement for ReViSE.

We also observe that the performance gap between our proposed ReViSE and other methods is reduced compared to transductive zero-shot learning. For instance, in average retrieval performance, ReViSE has 15.5%15.5\% mAP improvement over DeViSE* under zero-shot experiments, while only 9.3%9.3\% improvement under few-shot experiments.

5 t-SNE Visualization

Next, we provide the t-SNE visualization on the output visual test scores fv(Vte)f_{v}(\mathbf{V_{te}}) for DeViSE*, CMT*, and ReViSE in Fig. 6. Clearly, ReViSE can better separate instances from different classes.

Conclusion

In this paper, we showed how we can augment a typical supervised formulation with unsupervised techniques for learning joint embeddings of visual and textual data. We empirically evaluate our proposed method on both general and fine-grained image classification datasets, with comparisons against the state-of-the-art methods in zero-shot and few-shot recognition and retrieval tasks, from inductive to transductive setting. In all the experiments, our method consistently outperforms other methods, substantially improving performance in some cases. We believe that this work sheds light on the advantages of combining supervised and unsupervised learning techniques, and makes a step towards learning more useful representations from multi-modal data.

References

Network Design

Fig. 7 provides an easy-to-understand design of ReViSE. In all of our experiments, GoogLeNet is pre-trained on ImageNet images. Without fine-tuning, we directly extract the top layer activations (1024-dim) as our input image features followed by a common log(1+v)log(1+v) pre-processing step. For the textual attributes, we pre-process them through a standard l2l_{2} normalization.

During the first 100 iterations of training, we set λ=0\lambda=0 so that no unsupervised-data adaptation is used while still updating I^i,c(ut)\hat{I}_{i,c}^{(ut)}. Note that I^i,c(ut)\hat{I}_{i,c}^{(ut)} are the inferred labels for unsupervised data, and not random at each iteration. Beginning with the 101th iteration, we set λ={0.1,1.0}\lambda=\{0.1,1.0\} (chosen by cross-validation), and the model typically converges within 2000 to 5000 iterations.

We implement ReViSE in TensorFlow . We use Adam for optimization with minibatches of size 10241024. We choose tanhtanh for all of our activation functions.

Parameters Choice

We have four parameters in our architecture: α,β,γ\alpha,\beta,\gamma, and κ\kappa. We fix α=1.0\alpha=1.0, γ=0.1\gamma=0.1, κ=32.0\kappa=32.0 for all the experiments. Then we set λ=0.0\lambda=0.0 (no unsupervised-data adaptation inference), and perform cross-validation on the splitting set as suggested by to determine β\beta from {0.1,1.0}\{0.1,1.0\}. Next, with chosen β\beta, we perform cross-validation to choose λ\lambda from {0.1,1.0}\{0.1,1.0\}. Table 8 lists the statistics of β\beta and λ\lambda.

Next, we study the power of unsupervised information. We now take CUB\mathsf{CUB} dataset with att attributes to test the advantage of using unsupervised information, which can be viewed as tuning the parameter α\alpha for the unsupervised objective in eq. (11). Originally, α\alpha was set to 1.01.0, which equally weights the contribution of supervised and unsupervised loss. We now alter α\alpha as follows: 0.10.1 to 1.01.0 by step size of 0.10.1 and 0.50.5 to 5.05.0 by step size of 0.50.5. The results are shown in Fig. 8. We observe that when α\alpha increases from 0.10.1 to 1.01.0, the performance increases; however, when α\alpha increase from 1.01.0 to 5.05.0, the performance stays relatively unchanged. Empirically, we find that ReViSE does not perform better when α>1.0\alpha>1.0, which is expected, since we should not view unsupervised information more important than supervised information.

Precision-Recall Curve

Fig. 10 is the precision-recall curve for zero-shot retrieval results on CUB\mathsf{CUB} dataset with att attributes.

MMD Distance

MMD distance in eq. (5) can be viewed as the distribution measurement between visual and textual code. For CUB\mathsf{CUB} dataset with att attributes under transductive zero-shot experiment, we calculate the MMD distance (on the test codes) in our method with (ReViSE) and without (ReViSE†) LMMD\mathcal{L}_{MMD}. The results of MMD distance w.r.t. the number of iterations are shown in Fig. 9. We clearly observe that the red curve (ReViSE) has consistently lower value than the blue curve (ReViSE†). Moreover, based on the previous results, ReViSE always performs better than ReViSE†. Hence aligning the distributions across visual and textual codes can better associate cross-modal information and thus lead to more robust visual-semantic embeddings.

Remarks on Contractive Loss

We find that adding contractive loss to textual auto-encoder doesn’t provide much benefit. One possible reason may be the limited number of textual features (200200 for CUB\mathsf{CUB}). On the other hand, the number of visual features is large (11,78611,786 for CUB\mathsf{CUB}).

Comparing with recent state-of-the-art methods

In our main paper, we focus on comparing with deep-embeddings methods. In Table 9, we compare other methods for inductive and transductive zero-shot learning. Note that SMSESZSL adopts ESZSL for its initialization.

References