Learning like a Child: Fast Novel Visual Concept Learning from Sentence Descriptions of Images
Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, Zhiheng Huang, Alan Yuille
Introduction
Recognizing, learning and using novel concepts is one of the most important cognitive functions of humans. When we were very young, we learned new concepts by observing the visual world and listening to the sentence descriptions of our parents. The process was slow at the beginning, but got much faster after we accumulated enough learned concepts . In particular, it is known that children can form quick and rough hypotheses about the meaning of new words in a sentence based on their knowledge of previous learned words , associate these words to the objects or their properties, and describe novel concepts using sentences with the new words . This phenomenon has been researched for over 30 years by the psychologists and linguists who study the process of word learning .
For the computer vision field, several methods are proposed to handle the problem of learning new categories of objects from a handful of examples. This task is important in practice because we sometimes do not have enough data for novel concepts and hence need to transfer knowledge from previously learned categories. Moreover, we do not want to retrain the whole model every time we add a few images with novel concepts, especially when the amount of data or model parameters is very big.
However, these previous methods concentrate on learning classifiers, or mappings, between single words (e.g. a novel object category) and images. We are unaware of any computer vision studies into the task of learning novel visual concepts from a few sentences and then using these concepts to describe new images – a task that children seem to do effortlessly. We call this the Novel Visual Concept learning from Sentences (NVCS) task (see Figure 2).
In this paper, we present a novel framework to address the NVCS task. We start with a model that has already been trained with a large amount of visual concepts. We propose a method that allows the model to enlarge its word dictionary to describe the novel concepts using a few examples and without extensive retraining. In particular, we do not need to retrain models from scratch on all of the data (all the previously learned concepts and the novel concepts). We propose three datasets for the NVCS task to validate our model, which are available on the project page.
Our method requires a base model for image captioning which will be adapted to perform the NVCS task. We choose the m-RNN model , which performs at the state of the art, as our base model. Note that we could use most of the current image captioning models as the base model in our method. But we make several changes to the model structure of m-RNN partly motivated by the desire to avoid overfitting, which is a particular danger for NVCS because we want to learn from a few new images. We note that these changes also improve performance on the original image captioning task, although this improvement is not the main focus of this paper. In particular, we introduce a transposed weight sharing (TWS) strategy (motivated by auto-encoders ) which reduces, by a factor of one half, the number of model parameters that need to be learned. This allows us to increase the dimension of the word-embedding and multimodal layers, without overfitting the data, yielding a richer word and multimodal dense representation. We train this image captioning model on a large image dataset with sentence descriptions. This is the base model which we adapt for the NVCS task.
Now we address the task of learning the new concepts from a small new set of data that contains these concepts. There are two main difficulties. Firstly, the weights for the previously learned concepts may be disturbed by the new concepts. Although this can be solved by fixing these weights. Secondly, learning the new concepts from positive examples can introduce bias. Intuitively, the model will assign a baseline probability for each word, which is roughly proportional to the frequency of the words in the sentences. When we train the model on new data, the baseline probabilities of the new words will be unreliably high. We propose a strategy that addresses this problem by fixing the baseline probability of the new words.
We construct three datasets to validate our method, which involves new concepts of man-made objects, animals, and activities. The first two datasets are derived from the MS-COCO dataset . The third new dataset is constructed by adding three uncommon concepts which do not occur in MS-COCO or other standard datasets. These concepts are: quidditch, t-rex and samisen (see section 5)The dataset is available at www.stat.ucla.edu/~junhua.mao/projects/child_learning.html. We are adding more novel concepts in this dataset. The latest version of the dataset contains 8 additional novel concepts: tai-ji, huangmei opera, kiss, rocket gun, tempura, waterfall, wedding dress, and windmill.. The experiments show that training our method on only a few examples of the new concepts gives us as good performance as retraining the entire model on all the examples.
Related Work
Deep neural network Recently there have been dramatic progress in deep neural networks for natural language and computer vision. For natural language, Recurrent Neural Networks (RNNs ) and Long-Short Term Memories (LSTMs ) achieve the state-of-the-art performance for many NLP tasks such as machine translation and speech recognition . For computer vision, deep Convolutional Neural Networks (CNN ) outperform previous methods by a large margin for the tasks of object classification and detection . The success of these methods for language and vision motivate their use for multimodal learning tasks (e.g. image captioning and sentence-image retrieval).
Multimodal learning of language and vision The methods of image-sentence retrieval , image description generation and visual question-answering have developed very fast in recent years. Very recent works of image captioning includes . Many of them (e.g. ) adopt an RNN-CNN framework that optimizes the log-likelihood of the caption given the image, and train the networks in an end-to-end way. An exception is , which incorporates visual detectors, language models, and multimodal similarity models in a high-performing pipeline. The evaluation metrics of the image captioning task is also discussed . All of these image captioning methods use a pre-specified and fixed word dictionary, and train their model on a large dataset. Our method can be directly applied to any captioning models that adopt an RNN-CNN framework, and our strategy to avoid overfitting is useful for most of the models in the novel visual concept learning task.
Zero-shot and one-shot learning For zero-shot learning, the task is to associate dense word vectors or attributes with image features . The dense word vectors in these papers are pre-trained from a large amount of text corpus and the word semantic representation is captured from co-occurrence with other words . developed this idea by only showing the novel words a few times. In addition, adopted auto-encoders with attribute representations to learn new class labels and proposed a method that scales to large datasets using label embeddings.
Another related task is one-shot learning task of new categories . They learn new objects from only a few examples. However, these work only consider words or attributes instead of sentences and so their learning target is different from that of the task in this paper.
The Image Captioning Model
We need an image captioning as the base model which will be adapted in the NVCS task. The base model is based on the m-RNN model . Its architecture is shown in Figure 2(a). We make two main modifications of the architecture to make it more suitable for the NVCS task which, as a side effect, also improves performance on the original image captioning task. Firstly and most importantly, we propose a transposed weight sharing strategy which significantly reduces the number of parameters in the model (see section 3.2). Secondly, we replace the recurrent layer in by a Long-Short Term Memory (LSTM) layer . LSTM is a recurrent neural network which is designed to solve the gradient explosion and vanishing problems. We briefly introduce the framework of the model in section 3.1 and describe the details of the transposed weight sharing strategy in section 3.2.
As shown in Figure 2(a), the input of our model for each word in a sentence is the index of the current word in the word dictionary as well as the image. We represent this index as a one-hot vector (a binary vector with only one non-zero element indicating the index). The output is the index of the next word. The model has three components: the language component, the vision component and the multimodal component. The language component contains two word embedding layers and a LSTM layer. It maps the index of the word in the dictionary into a semantic dense word embedding space and stores the word context information in the LSTM layer. The vision component contains a 16-layer deep convolutional neural network (CNN ) pre-trained on the ImageNet classification task . We remove the final SoftMax layer of the deep CNN and connect the top fully connected layer (a 4096 dimensional layer) to our model. The activation of this 4096 dimensional layer can be treated as image features that contain rich visual attributes for objects and scenes. The multimodal component contains a one-layer representation where the information from the language part and the vision part merge together. We build a SoftMax layer after the multimodal layer to predict the index of the next word. The weights are shared across the sub-models of the words in a sentence. As in the m-RNN model , we add a start sign and an end sign to each training sentence. In the testing stage for image captioning, we input the start sign into the model and pick the best words with maximum probabilities according to the SoftMax layer. We repeat the process until the model generates the end sign .
2 The Transposed Weight Sharing (TWS)
where is the activation of the multimodal layer and is the SoftMax non-linear function.
The Novel Concept Learning (NVCS) Task
Suppose we have trained a model based on a large amount of images and sentences. Then we meet with images of novel concepts whose sentence annotations contain words not in our dictionary, what should we do? It is time-consuming and unnecessary to re-train the whole model from scratch using all the data. In many cases, we cannot even access the original training data of the model. But fine-tuning the whole model using only the new data causes severe overfitting on the new concepts and decrease the performance of the model for the originally trained ones.
To solve these problems, we propose the following strategies that learn the new concepts with a few images without losing the accuracy on the original concepts.
2 Fixing the baseline probability
After that, we set every element in to be the average value of the elements in and fix when we train on the new images. We call this strategy Baseline Probability Fixation (BPF).
In the experiments, we adopt a stochastic gradient descent algorithm with an initial learning rate of 0.01 and use AdaDelta as the adaptive learning rate algorithm for both the base model and the novel concept model.
3 The Role of Language and Vision
In the novel concept learning (NVCS) task, the sentences serve as a weak labeling of the image. The language part of the model (the word embedding layers and the LSTM layer) hypothesizes the basic properties (e.g. the parts of speech) of the new words and whether the new words are closely related to the content of the image. It also hypothesizes which words in the original dictionary are semantically and syntactically close to the new words. For example, suppose the model meets a new image with the sentence description “A woman is playing with a cat”. Also suppose there are images in the original data containing sentence description such as “A man is playing with a dog”. Then although the model has not seen the word “cat” before, it will hypothesize that the word “cat” and “dog” are close to each other.
The vision part is pre-trained on the ImageNet classification task with 1.2 million images and 1,000 categories. It provides rich visual attributes of the objects and scenes that are useful not only for the 1,000 classification task itself, but also for other vision tasks .
Combining cues from both language and vision, our model can effectively learn the new concepts using only a few examples as demonstrated in the experiments.
Datasets
We use the annotations and images from the MS COCO to construct our Novel Concept (NC) learning datasets. The current release of COCO contains 82,783 training images and 40,504 validation images, with object instance annotations and 5 sentence descriptions for each image. To construct the NC dataset with a specific new concept (e.g. “cat”), we remove all images containing the object “cat” according to the object annotations. We also check whether there are some images left with sentences descriptions containing cat related words. The remaining images are treated as the Base Set where we will train, validate and test our base model. The removed images are used to construct the Novel Concept set (NC set), which is used to train, validate and test our model for the task of novel concept learning.
2 The Novel Visual Concepts Datasets
We construct three datasets involving five different novel visual concepts:
NewObj-Cat and NewObj-Motor The corresponding new concepts of these two datasets are “cat” and “motorcycle” respectively. The model need to learn all the related words that describe these concepts and their activities.
NC-3 datasetThe dataset is publicly available at www.stat.ucla.edu/~junhua.mao/projects/child_learning.html. We are actively expanding the dataset. The latest version contains 11 novel concepts. The two datasets mentioned above are all derived from the MS COCO dataset. To further verify the effectiveness of our method, we construct a new dataset contains three novel concepts: “quidditch” (a recently created sport derived from “Harry Potter”), “t-rex” (a dinosaur), and “samisen” (an instrument). It contains not only object concepts (e.g. t-rex and samisen), but also activity concepts (e.g. quidditch). We labeled 100 images for each concept with 5 sentence annotations for each image. To diversify the labeled sentences for different images in the same category, the annotators are instructed to label the images with different sentences by describing the details in each image. It leads to a different style of annotation from that of the MS COCO dataset. The average length of the sentences is also 26% longer than that of the MS COCO (13.5 v.s. 10.7). We construct this dataset for two reasons. Firstly, the three concepts are not included in the 1,000 categories of the ImageNet Classification task where we pre-trained the vision component of our model. Secondly, this dataset has richer and more diversified sentence descriptions compared to NewObj-Cat and NewObj-Motor. We denote this dataset as Novel Concept-3 dataset (NC-3). Some samples images and annotations are shown in Figure 5.
We randomly separate the above three datasets into training, testing and validation sets. The number of images for the three datasets are shown in Table 1. To investigate the possible overfitting issues on these datasets, in the testing stage, we randomly picked images from the testing set of the Base Set and treated them as a separate set of testing images. The number of added images is equal to the size of the original test set (e.g. 1000 images are picked for NewObj-Cat testing set). We denote the original new concept testing images as Novel Concept (NC) test set and the added base testing images as Base test set. A good novel visual concept learning method should perform better than the base model on NC test set and comparable on Base test set. The organization of NC datasets is illustrated in Figure 4.
Experiments
To evaluate the output sentence descriptions for novel visual concepts, we adopt two evaluation metrics that are widely used in recent image captioning work: BLEU scores (BLEU score for n-gram is denoted as as B-n in the paper) and METEOR .
Both BLEU scores and METEOR target on evaluating the overall quality of the generated sentences. In the NVCS task, however, we focus more on the accuracy for the new words than the previously learned words in the sentences. Therefore, to conduct a comprehensive evaluation, we also calculate the score for the words that describe the new concepts. E.g. for the cat dataset, there are 29 new words such as cat, cats, kitten, and pawing. The precision and recall for each new word in the dictionary () are calculated as follows:
A high with a low indicates that the model overfits the new data (We can always get if we output the new word every time) while a high with a low indicates underfiting. We use the as a balanced measurement between and . Best score is 1. Note that if either or . Compared to METEOR and BLEU, the score show the effectiveness of the model to learn new concepts more explicitly.
2 Effectiveness of TWS and BPF
We test our base model with the Transposed Weight Sharing (TWS) strategy in the original image captaining task on the MS COCO and compare to m-RNN , which does not use TWS. Our model performs better than m-RNN in this task as shown in Table 2. We choose the layer dimensions of our model so that the number of parameters matches that of . Models with different hyper-parameters, features or pipelines might lead to better performance, which is beyond the scope of this paper. E.g. further improve their results after the submission of this draft and achieve a B-4 score of 0.302, 0.309 and 0.308 respectively using, e.g., fine-tuned image features on COCO or consensus reranking , which are complementary with TWS.
We achieve 2.5% increase of performance in terms of using TWS (Deep-NVCS-BPF-TWS v.s. Deep-NVCS-BPF-noTWSWe tried two versions of the model without TWS: (\@slowromancapi@). the model with multimodal layer directly connected to softmax layer like , (\@slowromancapii@). the model with an additional intermediate layer like TWS but does not share the weights. In our experiments, (\@slowromancapi@) performs slightly better than (\@slowromancapii@) so we report the performance of (\@slowromancapi@) here.), and achieves 2.4% increase using BPF (Deep-NVCS-BPF-TWS v.s. Deep-NVCS-UnfixedBias). We use Deep-NVCS to represent Deep-NVCS-BPF-TWS in short for the rest of the paper.
3 Results on NewObj-Motor and NewObj-Cat
The results show that compared to the Model-base which is only trained on base set, the Deep-NVCS models perform much better on the novel concept test set while reaching comparable performance on the base test set. Deep-NVCS also performs better than the Model-word2vec model. The performance of our Deep-NVCS models is very close to that of the strong baseline Model-retrain but needs only less than 2% of the time. This demonstrates the effectiveness of our novel concept learning strategies. The model learns the new words for the novel concepts without disturbing the previous learned words.
The performance of Deep-NVCS is also comparable with, though slightly lower than Deep-NVCS-1:1Inc. Intuitively, if the image features can successfully capture the difference between the new concepts and the existing ones, it is sufficient to learn the new concept only from the new data. However, if the new concepts are very similar to some previously learned concepts, such as cat and dog, it is helpful to present the data of both novel and existing concepts to make it easier for the model to find the difference.
3.2 Using a few training samples
We also test our model under the one or few-shot scenarios. Specifically, we randomly sampled images from the training set of NewObj-Cat and NewObj-Motor, and trained our Deep-NVCS model only on these images ( ranges from 1 to 1000). We conduct the experiments 10 times and average the results to avoid the randomness of the sampling.
We show the performance of our model with different number of training images in Figure 6. We only show the results in terms of score, METEOR, B-3 and B-4 because of space limitation. The results of B-1 and B-2 and consistent with the shown metrics. The performance of the model trained with the full NC training set in the last section is indicated by the blue (Base test), red (NC test) or magenta (All test) dashed lines in Figure 6. These lines represent the experimental upper bounds of our model under the one or few-shot scenario. The performance of the Model-base is shown by a green dashed line. It serves as an experimental lower bound. We also show the results of Model-retrain for NC test with black dots in Figure 6 trained with 10 and 500 novel concepts images.
The results show that using about 10 to 50 training images, the model achieves comparable performance with the Deep-NVCS model trained on the full novel concept training set. In addition, using about 5 training images, we observe a nontrivial increase of performance compared to the base model. Our deep-NVCS also better handles the case for a few images and runs much faster than Model-retrain.
4 Results on NC-3
The NC-3 dataset has three main difficulties. Firstly, the concepts have very similar counterparts in the original image set, such as Samisen v.s. Guitar, Quidditch v.s. football. Secondly, the three concepts rarely appear in daily life. They are not included in the ImageNet 1,000 categories where we pre-trained our vision deep CNN. Thirdly, the way we describe the three novel concepts is somewhat different from that of the common objects included in the base set. The requirement to diversify the annotated sentences makes the difference of the style for the annotated sentences between NC-3 and MS COCO even larger. The effect of the difference in sentence style leads to decreased performance of the base model compared to that on the NewObj-Cat and NewObj-Motor dataset (see Model-base in Table 5 compared to that in Table 4 on NC test). Furthermore, it makes it harder for the model to hypothesize the meanings of new words from a few sentences.
Faced with these difficulties, our model still learns the semantic meaning of the new concepts quite well. The scores of the model shown in Table 5 indicate that the model successfully learns the new concepts with a high accuracy from only 50 examples.
It is interesting that Model-retrain performs very badly on this dataset. It does not output the word “quidditch” and “samisen” in the generated sentences. The BLEU scores and METEOR are also very low. This is not surprising since there are only a few training examples (i.e. 50) for these three novel concepts and so it is easy to be overwhelmed by other concepts from the original MS COCO dataset.
5 Qualitative Results
In Table 6, we show the five nearest neighbors of the new concepts using the activation of the word-embedding layer learned by our Deep-NVCS model. It shows that the learned novel word embedding vectors captures the semantic information from both language and vision. We also show some sample generated sentence descriptions of the base model and our Deep-NVCS model in Figure 7.
Conclusion
In this paper, we propose the Novel Visual Concept learning from Sentences (NVCS) task. In this task, methods need to learn novel concepts from sentence descriptions of a few images. We describe a method that allows us to train our model on a small number of images containing novel concepts. This performs comparably with the model retrained from scratch on all of the data if the number of novel concept images is large, and performs better when there are only a few training images of novel concepts available. We construct three novel concept datasets where we validate the effectiveness of our method. These datasets have been released to encourage future research in this area.
Acknowledgement
We thank the comments and suggestions of the anonymous reviewers, and help from Xiaochen Lian in the dataset collection process. We acknowledge support from NSF STC award CCF-1231216 and ARO 62250-CS.