X-Trans2Cap: Cross-Modal Knowledge Transfer using Transformer for 3D Dense Captioning
Zhihao Yuan, Xu Yan, Yinghong Liao, Yao Guo, Guanbin Li, Zhen Li, Shuguang Cui
Introduction
Hitherto, the computer vision community has witnessed significant progress in image captioning and dense captioning under the success of deep learning techniques. Unlike image captioning describing a 2D image with a single sentence, dense captioning (DC) better interprets “A picture is worth a thousand words”. That is to say, for DC task, each object in an image is first perceived, then is provided more customized and detailed descriptions according to its nature and context.
Most recently, 3D cross-modal learning in vision and language has gained an increasing amount of interest as well. Several datasets and downstream applications are proposed and investigated. Unlike 2D images with regular grids and dense pixels, 3D data represented by a set of points are unordered and scattered in the 3D space, impeding the direct extension of 2D-based methods to 3D scenarios. To perform dense captioning on 3D point clouds, proposes the first method, namely Scan2Cap, by directly combining 3D object detection with natural language generation. Specifically, Scan2Cap first employs a detection backbone to obtain object proposals, and then applies a relational graph and a context-aware attention captioning module to learn object relations and generate tokens. Besides, multi-view features extracted by the pre-trained E-Net are further projected onto the input point cloud to enhance final captioning. However, Scan2Cap still has several issues: 1) The object representations in Scan2Cap are defective since they are solely learned from sparse 3D point clouds, thus failing to provide strong texture and color information compared with the ones generated from 2D images. 2) It requires the extra 2D input in both training and inference phases, as shown in Figure 1 (a). However, the extra 2D information is usually computation-intensive and unavailable during inference. For instance, a model training with both 2D and 3D inputs cannot apply to LiDAR scenarios that only contains 3D point cloud.
To address the above issues, we explore how to ease the barrier of cross-modal learning on 2D and 3D data, and investigate how to effectively combine the merits of both modalities for 3D dense captioning in this paper. To this end, we first time present a flexible and novel cross-modal framework, namely X-Trans2Caphttps://github.com/CurryYuan/X-Trans2Cap, which transfers color and texture-aware information from 2D image into 3D object representation using Transformer . Concretely, all the instances in a given scene can be firstly extracted by 3D object detection. Subsequently, the 3D features of each instance and its 2D counterpart are processed by a teacher-student framework. Within this framework, the teacher network takes the multi-modal inputs, while the student one only leverages the 3D inputs. Considering different modalities for teacher and student streams, we innovatively design a Transformer-based knowledge transfer framework with more flexible input control and better representation. Moreover, to further enhance the knowledge transfer, a modified knowledge distillation operation with cross-modal fusion (CMF) module and cross-modal feature alignment objective is proposed for knowledge generalization. Owing to the end-to-end training scheme, the priors in the 2D modality can inherently improve the teacher network and the student as well, i.e., our model takes advantage of the color and texture aware 2D representation and reduces the extra computational cost. Therefore, in the inference phase, X-Trans2Cap can perform superior captioning performance with only 3D inputs, as shown in Figure 1 (b).
Sufficient experiments evaluated on the ScanRefer and Nr3D datasets have demonstrated the effectiveness of our proposed X-Trans2Cap. In specific, with the extra 2D priors and the novel framework design, X-Trans2Cap can effectively learn a better 3D object representation and boost the performance over the model without 2D priors, i.e., improving the CIDEr points on ScanRefer from 75.75 to 87.09. This result also exceeds the previous state-of-the-art Scan2Cap by about 21 CIDEr.
In summary, our main contributions are threefold:
We first time propose X-Trans2Cap, a simple but effective cross-modal knowledge transfer framework for 3D dense captioning, in which an enhanced 3D representation with 2D priors is achieved.
X-Trans2Cap leverages a modified knowledge distillation method through a novel cross-modal fusion module and feature alignment techniques merged in Transformer, eliminating extra computation burdens during inference while achieving superior knowledge transfer.
Our X-Trans2Cap gains significant performance boost on the ScanRefer (+21.0 CIDEr) and Nr3D (+16.7 CIDEr) datasets.
Related Work
A broad collection of methods have been proposed in the field of image captioning in the last few years . Recently, many methods focus on utilizing the attention mechanism to capture meaningful information in the image, e.g., over grid regions and detected objects . Furthermore, some works attempt to combine attention with graph neural networks or Transformer to boost performance.
For the dense captioning task, it needs to generate captions for all the detected objects. Johnson et al. is the pioneer in this challenging field. Along this line, considers the context outside the salient image regions and takes advantage of global image features. further introduces the object relations among detected regions. However, due to the limited views of a single image, the performance of image-based dense captioning methods is significantly degraded when directly transferred to 3D scenarios.
2 3D Vision and Language
Compared to image and language comprehension, 3D vision and language understanding is a relatively emerging research field. Existed works focus on using language to confine individual objects, e.g., detecting referred 3D objects or distinguishing objects according to language phrases . Recently, ScanRefer and ReferIt3D introduce a task of localizing objects within a 3D scene given the linguistic descriptions, namely 3D visual grounding. TGNN and InstanceRefer follow the above settings and exploit panoptic segmentation to reduce the number of proposals. 3D dense captioning is proposed very lately in Scan2Cap . It focuses on decomposing 3D scenes and describing the chromatic and spatial information of the objects. Very recently, combines the above task of 3D grounding and caption to mutually enhance the performance of two tasks. Though promising, it only takes point clouds as input to generate instance features. Compared with the well-organized 2D images containing stronger texture and color information, such representation inherently challenges the learning process.
3 Cross-modal Knowledge Transferring
Previous studies apply 2D images as the extra inputs to 3D tasks, e.g., 3D object detection , semantic segmentation and object tracking . However, they require extra 2D information in both the training and inference phases. Thus, it inevitably augments computational burdens during evaluation and severely limits the efficiency in real-world applications. The concept of knowledge distillation was first shown by Hinton et al. . Subsequent research enhanced distillation by matching intermediate representations in the networks along with outputs using different approaches. Zagoruyko et al. proposed to align attentional activation maps between networks. Srinivas and Fleuret improved it by applying Jacobian matching to networks. In recent years, cross-modal knowledge distillation extended knowledge distillation by applying it to transferring knowledge across different modalities. Very recently, there are works attempting to only utilize 2D images during training phase to address the above problems. Among them, the 2D-assisted pre-training , inflating 2D convolution kernels to 3D and joint training with mask attention are proposed. Unlike those, we adopt a well-designed teacher-student framework with cross-modal fusion for more efficacious knowledge transfer, and the experiment results also demonstrate that our approach is much better than previous knowledge transferring.
Method
Our -Trans2Cap is developed upon a teacher-student framework , which is widely exploited in the knowledge distillation research field. The detailed architecture of -Trans2Caps is presented in Figure 2. -Trans2Cap takes two types of features as input, i.e., pure 3D modal input for student and multi-modal input for teacher respectively. We first introduce the details of the above feature representation in Section. 3.1. Then we propose a baseline model for 3D dense captioning through Transformers in Section. 3.2, named TransCap. In Section 3.3, we illustrate how -Trans2Cap transfers the 2D priors to the 3D representations, in which a cross-modal fusion (CMF) module is proposed. The details of training objectives are presented in Section 3.4. Finally, by incorporating the above components in a whole architecture, we illustrate the data flow of -Trans2Cap in training and inference phases in Section 3.5.
As shown in Figure 2 (a), our framework takes object-level representation as input, and each object feature is denoted as a token. Given that there are objects in the 3D scene, in the remaining section, the objects set is represented as , in which and are depicted as the -th object and the attribute of the -th object, respectively. In each iteration, we randomly choose an object as the target object () to be described as in . The other objects, i.e., , are treated as the reference objects, and only provide the cues of locations or relations to the target object.
For the 3D modal input, each object is considered from the perspective of its 3D feature, semantic, size as well as relative position to the target object. Specifically, the object representation is computed as follows:
The first three elements in the positional encoding calculate the center offset between the target object and -th object, and the others denote their relative size. Two learnable projection matrices and in Eqn. (1) then transform the dimensions of and to . Finally, a transformation function generates the final object feature for the -th object.
Apart from 3D information in the multi-modal input, the corresponding 2D feature and 2D bounding box are introduced for the -th object as follows:
Concretely, for each object, its ground truth of the 3D bounding box is projected onto the original ScanNet videos to obtain the corresponding 2D boxes. In each training step, an image is randomly selected from the video sequences to generate an extra input. Features in the 2D box area are extracted by the Faster-RCNN detector pre-trained on the Visual Genome dataset, which are regarded as 2D features for the -th instance, i.e., . Finally, as shown in Eqn. (3), by applying linear and nonlinear transformations and , a -dimensional multi-modal representation is generated for the -th object.
2 Baseline Model: TransCap
The captioning decoder is conditioned on previously generated words and features from the encoder layers to generate the next token. Specifically, it integrates the features from different encoder layers and performs the cross-attention on the generated tokens.
3 Cross-Modal Fusion Module
The Cross-modal fusion (CMF) module enables cross-modal feature interaction between pure 3D and multi-modal feature representations. As shown in Figure 3, it is designed to construct interaction from the student network to the teacher at the same encoder level, thus building a bridge to fuse features between single and multiple modalities. Moreover, to further enhance the ability of the student network to learn the multi-modal representation, we exploit a random mask on the features from the teacher network. Owing to this framework, the strengths of multi-modal representation can be assimilated to reinforce the student network via an end-to-end training protocol. Specifically, we element-wise add the student features with the masked teacher features.
where the and denote features from -th encoder layer of student and teacher networks, respectively. The notation means element-wise addition. The mask indicator is initialized with and has the probability of change to . After that, we feed the fused features into the next encoder layer of the teacher network. It should be highlighted that, since our CMF module employs the single-directional connection from the student to the teacher, the teacher network can be discarded during inference, i.e., it introduces no extra computation for the student network. Moreover, various designs for the CMF module, including ablations, are shown in Section. 4.4.
4 Objective Function
Feature alignment loss. Following a standard practice in knowledge transfer, we use Huber loss (i.e., Smooth-L1 regression loss) to align decoder features between teacher and student networks.
Captioning loss. As in the previous work , we apply a conventional cross entropy loss function on the generated token probabilities in both teacher and student networks. Furthermore, to further boost the performance, we propose an enhanced version model -Trans2Cap (C) by applying the CIDEr-D score as reward. Following the previous work , we baseline the reward using the mean of the rewards rather than greedy decoding following previous methods .
Total objective loss. We combine all three loss terms linearly as our final objective loss function:
where , and are the weights for each individual loss. To guarantee the loss terms are roughly of the same magnitude, we fine-tune weights on the validation split, and set those to , , and empirically in the experiments.
5 Training and Inference Schemes
The black and red arrows in the Figure 2 (b) illustrate the information flow of the -Trans2Cap for training and inference. It is noteworthy that teacher and student networks are trained from scratch. In the training phase, both networks are exploited (see the black and red arrows in Figure 2 (b)), and CMF modules between corresponding encoder layers and feature alignment are conducted to enhance mutual representation. During the inference, if only the 3D modality exists, we only apply the student network (see the red arrow in Figure 2 (b)). However, if the auxiliary 2D information is available as well, the stronger teacher framework will be exploited. In our experiment, we demonstrate that our architecture can both enhance the performance of teacher and student networks with and without additional modality.
Experiment
We compare our method with Scan2Cap and 2D baselines proposed in their paper. Extending from , we further compare all methods on Nr3D dataset . More experiment results including subjective evaluation and ablations are in the supplementary material.
ScanRefer. The ScanRefer dataset annotates 800 3D indoor scenes in the ScanNet dataset with 51,583 language queries. It follows the official ScanNet splits and contains 36,665, 9,508, and 5,410 samples in train/val/test sets, respectively. Since the dataset is initially used in visual grounding and the labels of the test set are inaccessible, we follow the same setting as in to form the train and val sets for training and testing.
Nr3D. The Natural Reference in 3D (Nr3D) has the same train/val split as ScanRefer. It contains 41,503 queries annotated by Amazon Mechanical Turk (AMT) workers. Compared with ScanRefer dataset, Nr3D is more challenging since it does not contain the fixed or redundant sentence patterns, i.e., declarative sentences starting with “this is” or “that is”. We do not compare our method on its counterpart, the Spatial Reference in 3D (Sr3D) dataset, since it is totally generated by the machine templates.
2 Tasks and Metrics
Tasks. In our experiment, we follow and design two protocols to evaluate the generated caption:
Dense captioning with ground truth instances (Oracle DC): In this setting, the point cloud of each instance is given. Then one needs to generate faithful captions based on their attribute information and spatial relationships.
Dense captioning with 3D scans (Scan DC): This setting is more challenging. One needs to detect objects from the 3D scans first and then generate captions for each object according to the detection results.
Metrics. In Oracle DC, we directly apply CIDEr , BLEU-4 , METEOR and ROUGE averagely on all instances as metrics. For brevity, we simplify them as C, B-4, M and R, correspondingly.
In Scan DC, to jointly measure the quality of generated captions and detected bounding boxes, we evaluate them by combining above metrics with Intersection-over-Union (IoU) scores between predicted bounding boxes and GT bounding boxes. Specifically, we follow and define the combined metrics as IoU, where is set to 1 if the IoU score for the -th box exceeds , otherwise 0. We use to represent the above captioning metrics, e.g., CIDEr. is the number of detected object bounding boxes. We also use mean average precision (mAP) thresholded by IoU as the object detection metric.
3 3D Dense Captioning Results
Oracle dense captioning. The results of Oracle DC task are displayed in Table 1. In the upper part, we compare results without the extra 2D input for inference. Scan2Cap and Scan2Cap (Inst) denote the methods that exploit ground-truth (GT) boxes and GT instances as input, respectively. Merely using the baseline model TransCap, we improve the captioning result by a large margin compared to Scan2Cap (+9.96 and +7.24 CIDEr points on ScanRefer and Nr3D). Utilizing our cross-modal knowledge transfer training strategy further boost the performance on all the captioning metrics. Specifically, after using our teacher student framework, -Trans2Cap achieves 11.04 and 9.42 CIDEr improvement over TransCap on the ScanRefer and Nr3D datasets, where the performance on both datasets are about 20 CIDEr scores higher than those of Scan2Cap. The bottom part of Table 1 illustrates the result using extra 2D input in both the training and inference phases. Though using the extra 2D input for inference, the performance of Scan2Cap is still inferior to that of our propsed -Trans2Cap only exploiting 3D modal input, let alone using multi-modal. Besides, -Trans2Cap is better than TransCap when both using the extra 2D input, especially on Nr3D (85.38 vs 77.55 CIDEr), which illustrates that training with student network can even improve the result of teacher network. Furthermore, with CIDEr-D score optimization, i.e., -Trans2Cap (C) model, the performance of captioning can be further improved.
Scan dense captioning. In Table 2, we compare the result of Scan DC, which shows results without and with extra 2D input in the inference phase. The method for proposal generation is listed in the third column. Among these methods, 2D-3D Proj. and 3D-2D Proj. are two baseline methods proposed in . 2D-3D Proj. applies Mask R-CNN to generate 2D proposals in images, where the corresponding 2D bounding boxes and features are fed into the description generation module . On the contrary, 3D-2D Proj. exploits VoteNet to extract 3D proposals, which are projected back to 2D images. Then the projected 2D proposals are finally adopted to generate captions. As shown in Table 2, 2D-based methods obtain lowest captioning scores, which reveals that they cannot directly handle the 3D dense captioning task. Though Scan2Cap achieves better results than these 2D based methods, it is much inferior to -Trans2Cap without the assistance of appealing 2D priors and dedicated network structure. Surprisingly, we observe that the detection performance of -Trans2Cap is improved as well, though there is no extra 2D input fed into the detector during training and testing. It confirms that our -Trans2Cap is not only capable of faithful caption generation, but also acquires the knowledge mining capacity within multi-modalities for more complex applications, i.e., digging out the information embedded into language description for 3D visual detection. The results of Scan DC on Nr3D dataset are illustrated in the supplementary.
Visualization. Figure 4 presents the visualization results of -Trans2Cap, which demonstrates great improvements upon Scan2Cap for more faithful captions. Furthermore, we present the corresponding 2D counterparts within each 3D scene. Regarding 2D images, they can provide stronger texture and color information obviously when compared with sparse point cloud.
Comparison for knowledge transfer. To further verify the effectiveness of our proposed method upon common teach-student architecture and other cross-modal manners, we compare -Trans2Cap with typical approaches of knowledge transfer in Table 3. Among all the methods, Hinton et al. and Huang et al. are pure knowledge distillation designs, where the former is the pioneer for the research filed and the latter is newly proposed. As shown in the table, pure knowledge distillation manners cannot be directly adopted on the 3D DC scenario, and their improvement upon the baseline model is limited. Very recently, approaches and adopt cross-modal knowledge transfer technique in the 3D tasks. The core idea of is using extra 2D input to conduct 3D pre-training. We modify it by first training a TransCap with multi-modal input, then use its pre-trained parameters as initialization weights for pure 3D input training. For 2D semantic-assisted training (SAT) , it treats 2D features as additional tokens in (i.e., concatenated in sequence dimension) the same model, and then exploits an attention mask in Transformer layers. This mask only ignores the attention from 3D to the 2D. However, both methods cannot boost the performance as no mutual enhancement is introduced.
We also offer an offline distillation design, preparing a pre-trained teacher network before training the student network, called -Trans2Cap (pre-trained) in the bottom part of the Table 3. It can be noticed that using pre-trained teacher network results in a performance drop of 7.38 CIDEr, which may result from the distribution gap between multi-modality data. To the end, in the Table 3, -Trans2Cap significantly performs better, which illustrates the effectiveness of the teacher-student framework and cross-modal fusion (CMF) module.
4 Analysis and Ablation Studies
Does knowledge transfer help? As results shown in Table 1 and 2, when we adopt 2D prior during the training phase (-Trans2Cap), it can greatly improve the performance upon the baseline model (TransCap).
Does our proposed components help? To further verify the effectiveness of different components, we conduct the ablation studies in the Table 4. As shown in Table 4, model A is our baseline model (TransCap), and model B is our entire architecture of -Trans2Cap. The model C is the ablated architecture that discards the feature alignment loss . It can be found out that there is a great performance drop from 87.09 to 79.58 in terms of CIDEr metric. Fortunately, due to the advantage of CMF module, it still has 3.83 improvement upon the baseline model. Similarly, the performance drop is appearing (-5.25 CIDEr) when removing the CMF module (model D). This result demonstrates that both framework architecture and CMF module play important roles in -Trans2Cap.
How to design cross-modal fusion? We illustrate the results from different designs of CMF module in the bottom part of the Table 4. On the one hand, discarding the random mask hampers the caption results, as shown in model F. On the other hand, exploiting more complicated operations such as concatenation and attention mechanism cannot effectively improve the performance. There is only a slight improvement on the metrics of BLEU-4 and Rough for adopting attention. However, it will greatly increase the model complexity while make the CIDEr decline.
Conclusion
In this work, we propose an enhanced 3D dense captioning method via cross-modal knowledge transfer, named -Trans2Cap. By designing the network architecture and knowledge distillation method carefully, our -Trans2Cap outperforms previous methods by a large margin on multiple datasets with more faithful captions. We believe that our work can be applied to a wider range of 3D vision and language scenarios, and provide a solution to the comprehension of 3D scenes with severe texture details missing, i.e., leveraging the 2D priors and the cross-modal knowledge transfer to improve the performance.
Acknowledgment
This work was supported in part by NSFC-Youth 61902335, by Key Area R&D Program of Guangdong Province with grant No.2018B030338001, by the National Key R&D Program of China with grant No.2018YFB1800800, by Shenzhen Outstanding Talents Training Fund, by Guangdong Research Project No.2017ZT07X152, by Guangdong Regional Joint Fund-Key Projects 2019B1515120039, by the NSFC 61931024&81922046, by helixon biotechnology company Fund and CCF-Tencent Open Fund.
References
A Overview
In this supplementary material, we illustrate the implementation details, the efficiency of the model and the results of subjective evaluation in Section B, Section C and Section D, respectively. After that, we provide more experiments of Scan Dense Captioning (DC) on Nr3D dataset in Section E. Then we discuss the effectiveness of each attribute in instance representation in Section F.
B Implementation Details
In our experiment, we adopt the PointNet++ to generate 3D object features () in Oracle DC, and applies proposals’ features from VoteNet in the Scan DC. Furthermore, in the test with Oracle DC, we use ground truth category as while adopting the predicted results from detector in Scan DC task. We train the network for 30 epochs by using Adam optimizer with a batch size of 32. The probability of random mask in CMF module is set as 0.2 when achieving the best, and it does not greatly change the result. It should be noted that both teacher and student networks are trained from scratch. The learning rate is initialized as 0.0005 with the decay as 0.1 for every 10 epochs. Experiments are conducted on RTX2080Ti GPUs.
C Running Time Evaluation
We investigate the running time of our model in this section. Table 1 shows the number of parameters and inference time of per scan in Oracle DC setting. -Trans2Cap (3D) can speed up more than 20 compared with its baseline and -Trans2Cap using extra 2D modality.
D Subjective Evaluation
We conduct a subjective evaluation with three volunteers on randomly selected 100 descriptions generated by Scan2Cap and -Trans2Cap with Oracle DC setting on ScanRefer datasets. The subjective evaluation results are shown in Table 2. In practice, each volunteer is asked to manually check whether the descriptions correctly reflect two aspects of the object: object color attributes and spatial relations in local environment. As observed from Table 2, -Trans2Cap can generate more faithful captions regarding the attributes and spatial relationships.
E Scan Dense Captioning on Nr3D
In Table 3, we compare the results of Scan DC on Nr3D, including the results without and with extra 2D input in the inference phase. All methods exploit the same network, i.e., VoteNet, to generate proposals. 3D-2D Proj. projects proposals back to 2D images and captions in a 2D manner. However, it achieves the lowest captioning scores, which reveals that it cannot directly handle the 3D dense captioning task. Though Scan2Cap achieves better results than 3D-2D Proj., it also cannot generate faithful captioning results. Not surprisingly, -Trans2Cap obtains the highest score in all metrics. Specifically, it not only gains a +2.9 improvement in CIDEr@0.25 score upon baseline TransCap, but also achieves +5.5 boost over Scan2Cap. Finally, the experiment also confirms that our -Trans2Cap can improve 3D visual detection as well.
F Analysis and Ablation Studies
We further conduct an ablation study on different instance representation designs as shown in Table 4, where the upper part and lower part show the specific designs in teacher and student network, respectively.
Does object class help? From the results of model B in Table 4, it can be found out that there is a dramatic drop in metric of CIDEr, from 87.09 to 70.41, when we discard the object class . Thus, it shows that is an important attribute for instance representation. Note that the ablated model B is still +6 CIDEr higher than that of Scan2Cap.
Does 3D bounding box help? As shown in results of model C in Table 4, removing the 3D bounding box will not cause a large performance drop, only -4.1 CIDEr from 87.09 to 82.99. This result reflects that -Trans2Cap utilizes 3D object spatial coordinates to generate captions.
Does positional encoding help? The result of model D demonstrates a tremendous performance decrease in metric of CIDEr when positional encoding is not exploited, where the model can only obtain 33.52 in metric of CIDEr. Since our model only chooses one object as the target object and the remaining ones will be regarded as reference objects, positional encoding helps the model to identify the target one. Without its help, the network can hardly work.
Does 2D input help? The lower part of Table 4 describes the effectiveness of different attributes in the teacher network. There are three conclusions can be obtained: 1) Discarding 3D features in teacher network barely hampers the performance (model E). This is because the 3D features also exist in the input of the student network. 2) Utilizing the pre-trained network to extract 2D features is not necessary (model F). The result of model F shows that even if we only exploit the information of 2D bounding box, there is only an about -2 CIDEr drop for the caption results. 3) The 2D bounding box information seems to play a more important role compared with 2D features (see the model G). Without using , the model only obtains 83.85 CIDEr, and this result is even 0.4 lower than that of model F (without using ). Such results also emphasize the capability of -Trans2Cap in real-world applications, i.e., without pre-trained 2D network, only utilizing the 2D bounding box information can still greatly boost the captioning performance.