Auto-Encoding Knowledge Graph for Unsupervised Medical Report Generation
Fenglin Liu, Chenyu You, Xian Wu, Shen Ge, Sheng Wang, Xu Sun
Introduction
Medical images, such as radiology and pathology images, and their corresponding reports are widely used for clinical diagnosis and treatment . A medical report is usually a paragraph of multiple sentences which describes both the normal and abnormal findings in the medical image. In clinical practice, writing a report can be time-consuming and tedious for experienced radiologists, and error-prone for inexperienced radiologists . Therefore, given the large volume of medical images, automatically generating reports can improve current clinical practice in diagnostic radiology and assist radiologists in clinical decision-making . Specifically, it can relieve radiologists from such heavy workload and alert radiologists of the abnormalities to avoid misdiagnosis and missed diagnosis. Therefore, automatic medical report generation attracts remarkable attention in both artificial intelligence and clinical medicine.
Recently, inspired by the great success of neural machine translation , image captioning and medical imaging analysis , the data-driven deep neural models, particularly those based on the encoder-decoder frameworks , have achieved great success in advancing the state-of-the-art of medical report generation. However, these models are trained in a supervised learning manner and heavily rely on labeled paired image-report datasets , which are not easy to acquire in the real world. Specifically, the medical-related data can only be manually labeled by professional radiologists, and also involves privacy issues. Therefore, the medical report generation datasets are particularly labor-intensive and expensive to obtain. As a result, the scales of existing widely-used datasets for medical report generation models , i.e., MIMIC-CXR (0.22M samples) and IU X-ray (4K samples) , are relatively small compared to image recognition datasets, e.g., ImageNet (14M samples) , and image captioning datasets, e.g., Conceptual Captions (3.3M samples) . In addition, the MIMIC-CXR and IU X-Ray datasets only include Chest X-Ray images, for other types of medical images (MRI, Dermoscopy, Retinal, etc.) of other body parts (brain, skin, eye, etc.), the image-report pairs could be much less or even unavailable. Therefore, to relax the reliance on the paired data sets, making use of all available data, like independent image or report sets, is becoming increasingly important.
In this paper, we propose an unsupervised model Knowledge Graph Auto-Encoder (KGAE), which utilizes independent sets of images and reports in training (the image and report set are separate and have no overlap). KGAE consists of a pre-constructed knowledge graph, a knowledge-driven encoder and a knowledge-driven decoder. As shown in Figure 1, the knowledge graph works as the shared latent space of images and reports. The knowledge-driven encoder can take either the image or the report as queries and project them to corresponding coordinates and in the latent space. In this manner, since and share the same latent space, we can use positions in latent space to measure the relationship between images and reports which narrows the gap between visual and textual domains. In brief, to bridge the gap between vision and language domains without training on the pairs of images and reports, we adopt the knowledge graph to create a latent space and propose a knowledge-driven encoder, which includes a common mapping function to project images and reports to the same latent space. As a result, our encoder can extract the image and report knowledge representations, i.e., the knowledge related to the image and report, they (image, report knowledge) share the common latent space, which allows our model to bridge the gap between vision and language domains without the training on the pairs of image and report. Next, we introduce the knowledge-driven decoder to exploit and to generate the report. In the training stage, we estimate the parameters of the decoder by reconstructing the input report based on , i.e., auto-encoding pipeline; In the prediction stage, we directly input into the trained decoder to generate the report. In this way, our approach can produce desirable reports without any labeled image-report pairs.
Overall, the contributions of this paper are as follows:
In this paper, we make the first attempt to conduct unsupervised medical report generation where the image-report pairs are not available. To this end, we propose the Knowledge Graph Auto-Encoder (KGAE). By leveraging a pre-constructed knowledge graph, we introduce the knowledge-driven encoder and decoder which are trained with independent sets of images and reports. According to the experimental results, the unsupervised KGAE can even outperform several supervised approaches.
In addition to the unsupervised mode, KGAE can also be applied in a semi-supervised or supervised manner. Under the semi-supervised setting, by using only 60% of paired dataset, KGAE is able to achieve competitive results with current state-of-art models; Under the supervised setting, by training on fully paired datasets as in existing works, KGAE can set new state-of-the-art performances on the IU X-ray and MIMIC-CXR, respectively.
The analysis, both quantitative and qualitative, as well as a human evaluation conducted by professional radiologists, further proves the effectiveness of our approach.
Related Works
Medical report generation aims to generate a relatively long paragraph to describe a given medical image. It is similar to the image captioning task , which aims to generate a sentence to describe a given image. In image captioning, the encoder-decoder framework , where the encoder computes visual representations for the image and the decoder generates a target sentence based on the visual representations, has achieved great success . However, instead of only generating one single sentence, medical report generation aims to generate a long paragraph including multiple structured sentences that describe both the normal and abnormal parts . To this end, given the success of encoder-decoder framework on image captioning, most existing medical report generation models attempt to exploit the hierarchical LSTM (HLSTM) or Transformer to generate an accurate, long and coherent report. However, existing models require the paired image-report datasets, which are time-consuming and expensive to collect. In this paper, we propose the unsupervised model Knowledge Graph Auto-Encoder (KGAE) which doesn’t need paired images and reports.
Note that although the knowledge graph has been integrated in existing medical report generation models , these approaches are supervised and require paired images and reports. Thus, their objectives and motivations of using knowledge graph are different from our work. In detail, existing knowledge-graph based medical report generation methods aim to adopt the knowledge graph to boost the performance of supervised models. However, in our work, we aim to generate a medical report without using any coupled image-report training pairs, i.e., unsupervised medical report generation. A key challenge of unsupervised medical report generation is to bridge the gap between vision and language domains. To this end, we adopt the knowledge graph to create a latent space and propose a knowledge-driven encoder to project image and report to the same latent space.
Approach
We first formulate the conventional supervised medical report generation problems; Then, we describe the proposed Knowledge Graph Auto-Encoder for unsupervised medical report generation in detail.
Given a medical image , the goal is to generate a descriptive report . Most models normally include an image encoder and a report decoder, which can be formulated as:
2 Knowledge Graph Auto-Encoder
As shown in Figure 1, the proposed KGAE includes a knowledge graph, a knowledge-driven encoder and a knowledge-driven decoder, which will be described in detail in the following sections.
Knowledge-driven Decoder The decoder is designed to generate the reports based on the graph representations or . For clarity, we use to represent the and during the training and testing stages, respectively. In implementations, since medical report generation requires generating a long paragraph, we choose the (three-layer) Transformer as the basic module of our decoder and incorporate the proposed Knowledge-driven Attention (KA) to effectively model the long sequences.
Based on the above mechanism, the knowledge-driven attention in Eq. (5) is defined as:
where denotes the concatenation operation. It is worth noting that these operations are all differentiable, thus the bank can be learned in an end-to-end fashion.
In our subsequent analysis, we will show that the introduced knowledge memory mechanism indeed distills and preserves the desired medical knowledge, and thus boost the generation of reports.
3 Implementation Details
Unsupervised Training Details To train our KGAE in an unsupervised manner, instead of using the paired image-report dataset in the conventional supervised model, we only require an image set CheXpert , which includes 224,316 X-ray images, and a separate report corpus MIMIC-CXR + IU X-ray , which includes 222,758 + 2,770 = 225,528 reportsThere are no paired image-report samples between CheXpert and MIMIC-CXR+IU X-ray..
In detail, to train our knowledge-driven encoder and (see Eq. (4)), we feed the and into a common multi-label classification network trained with binary cross entropy loss for 14 common radiographic observations classificationAtelectasis, Cardiomegaly, Consolidation, Edema, Enlarged Cardiomediastinum, Fracture, Lung Lesion, Lung Opacity, No Finding, Pleural Effusion, Pleural Other, Pneumonia, Pneumothorax, Support Devices.. In this way, our encoder can extract the knowledge representations and of both image and report in a common latent space, effectively bridging the vision and the language domains. To train the knowledge-driven decoder, as well as the knowledge bank (see Eq. (5)), since there are no coupled image-report pairs, we propose to reconstruct the report based on the . Therefore, through Eq. (5), taking the input report as the ground truth report, we can train our approach by minimizing the cross-entropy loss:
In this way, we can train our decoder in the auto-encoding pipeline.
During testing, we first adopt the knowledge-driven encoder to extract the knowledge representations of the test image (see Eq. (4)). Then, we directly feed into the decoder to generate final report in the pipeline (see Eq. (5)). In this way, our approach can relax the reliance on the image-report pairs. In our following experiments, we validate the effectiveness of our approach, which even outperforms some supervised approaches.
Semi-Supervised and Supervised Training Details To further validate the effectiveness of our approach, we fine-tune the unsupervised KGAE using partial and full image-report pairs to acquire the KGAE-Semi(-Supervised) and KGAE-Supervised, respectively, where the former can evaluate the performance of our approach under limited labeled pairs for training and the latter can compare the performance of KGAE with state-of-the-art supervised approaches. In the (semi-)supervised setting, given the image-report pairs, i.e., -, we first incorporate the original visual information into the knowledge representation , and then train our KGAE by generating the ground truth report in the pipeline and minimizing the cross-entropy loss in Eq. (8). During testing, we also follow the unsupervised setting to generate the final report in the pipeline.
Experiments
We first introduce the datasets, metrics and detailed settings used for evaluation. Then, we present the evaluation of our approach under the unsupervised, semi-supervised and supervised training settings.
Datasets In this paper, we adopt the test sets of IU X-ray and MIMIC-CXR for evaluation. All protected health information (e.g., patient name and date of birth) was de-identified. In particular, the IU X-ray is a widely-used public benchmark dataset for medical report generation and contains 7,470 chest X-ray images associated with 3,955 fully de-identified medical reports. Each report is composed of impression, findings and indication sections, etc. . Following , our method also focuses on the findings section as it is the most important component of reports. Then, following , we randomly select 70%-10%-20% image-report pairs of dataset to form the training-validation-testing sets. The MIMIC-CXR includes 377,110 chest x-ray images associated with 227,835 reports. The dataset is officially split into 368,960 images (222,758 reports) for training, 2,991 images (1,808 reports) for validation and 5,159 images (3,269 reports) for testing. It is worth noting that we focus on the unsupervised medical report generation, where the image-report pairs are not available, thus the image-report training pairs of both the IU X-ray and MIMIC-CXR datasets are discarded and are not used in our unsupervised training stage. Only the training reports of MIMIC-CXR and IU X-ray, i.e., 222,758 + 2,770 = 225,528 reports, are used as the independent report corpus to train our unsupervised model. Only under the (semi-)supervised training setting, we will adopt the image-report training pairs to train our approach.
Metrics To fairly compare with existing models , we adopt the evaluation toolkit to calculate the widely-used natural language generation metrics, i.e., BLEU , METEOR and ROUGE-L , which measure the match between the generated reports and ground truth reports, but are not specialized for the abnormalities in the reports. Therefore, to measure the accuracy of descriptions for clinical abnormalities, we further report clinical efficacy metrics following the work of Chen et al. . The clinical efficacy metrics are calculated by comparing the generated reports with ground truth reports in 14 different categories related to thoracic diseases and support devices, producing the Precision, Recall and F1 scores.
2 Automatic Evaluation
We evaluate the performance of our approach under unsupervised, semi-supervised and supervised settings. The results are shown in Table 1, Table 2 and Figure 2. We select several supervised methods, including a recently state-of-the-art model R2Gen , for comparison. These models follow the encoder-decoder architecture, trained on the full pairs of images and reports.
Unsupervised Setting As shown in Table 1 and Table 2, our unsupervised model KGAE achieves competitive results with some supervised models in both IU X-ray and MIMIC-CXR datasets, and even outperforms several supervised models. Specifically, on the IU X-ray dataset, Table 1 shows that KGAE surpasses the NIC , AdaAtt and Att2in in terms of all metrics, and the Transformer in terms of BLEU-1,2,3. On the MIMIC-CXR dataset, Table 2 shows that KGAE outperforms the Trans. in terms of Precision, Recall and F1. The competitive results prove the effectiveness of our approach in addressing the unsupervised medical report generation, and thus can generate desirable medical reports without the training on the pairs of image and report.
Semi-Supervised Setting To further prove the effectiveness of our approach, we fine-tune the unsupervised KGAE using partial downstream paired image-report datasets (see Section 3.3), resulting in the KGAE-Semi. To this end, in Figure 2, we evaluate the performance of our approach on both IU X-ray and MIMIC-CXR datasets with respect to the increasing amount of paired data. For a fair comparison, we also re-train the state-of-the-art model R2Gen using the same amount of pairs. As we can see, our model outperforms the R2Gen under all ratios of paired dataset used for training. It is worth noting that the fewer the image-report pairs, the larger the margins, e.g., under the very limited pairs setting (20% of paired datasets), our approach significantly surpasses the R2Gen by 8.5% absolute BLEU-4 score on IU X-ray and 9.8% absolute F1 score on MIMIC-CXR. Intuitively, since our approach can relax the reliance on the paired datasets, we can make use of available unpaired image and report data as a solid bias for medical report generation task. Table 1 and Table 2 further prove the effectiveness of our approach, which achieves results competitive with current state-of-the-art models by using only 60% of paired dataset.
Supervised Setting We fine-tune the unsupervised KGAE using full image-report pairs, acquiring the KGAE-Supervised model (see Section 3.3). Table 1 and Table 2 show that KGAE-Supervised sets the new state-of-the-art results on the two datasets in all metrics. Moreover, in terms of the clinical efficacy metrics, our approach achieves 0.389 precision score, 0.362 recall score and 0.355 F1 score, outperforming the state-of-the-art model R2Gen . The superior clinical efficacy scores demonstrate the capability of our approach to produce higher quality descriptions for clinical abnormalities than existing models.
Overall Combining the results of unsupervised, semi-supervised and supervised settings, the proposed KGAE can relax the dependency on the paired datasets, and thus makes the medical report generation model use the available separate image and report data to boost the performance. The advantages under the scenarios with limited labeled pairs (i.e., semi-supervised setting) show that KGAE might be applied to other medical images (MRI, Dermoscopy, Retinal, etc.), where the coupled images and reports pairs could be much less or even unavailable.
3 Human Evaluation
We conduct human evaluations to verify the effectiveness of KGAE in clinical practice. Specifically, to assist radiologists in clinical decision-making and reduce their workload, it is important to generate accurate reports (faithfulness), i.e., the model does not generate normalities and abnormalities that does not exist according to doctors, with comprehensive abnormalities (comprehensiveness), the model does not leave out the abnormalities. Therefore, we randomly select 100 samples from the MIMIC-CXR and invite three professional clinicians to compare our approach and baselines independently. The clinicians are unaware of which model generates these reports. The results are shown in Table 3. As we can see, under the unsupervised setting, our approach achieves competitive results with the supervised model, outperforming the Trans with winning pick-up percentages. Under the (semi-)supervised setting, our method is better than state-of-the-art model R2Gen in all metrics, especially for the semi-supervised setting (20% of paired dataset), our method substantially surpasses the R2Gen, which is in accordance with the automatic evaluation, by and points in terms of the faithfulness and comprehensiveness metrics, respectively.
Analysis
In this section, we conduct several analysis to better understand our proposed approach.
In this section, to evaluate the knowledge graph sensitivity, we evaluate the performances using different knowledge graphs defined on IU X-Ray only, MIMIC-CXR only, both MIMIC-CXR and IU X-Ray. Table 4 shows the results of KGAE (0%) and KGAE-Supervised (100%) on the IU X-ray dataset. As we can see, our KGAE using different knowledge graphs can consistently outperform several existing supervised models, i.e., NIC, AdaAtt, Att2in, across all metrics (Table 1). Similarly, our KGAE-Supervised with different knowledge graphs can also consistently outperform existing state-of-the-art model, i.e., R2Gen (Table 1). The results prove the robustness of our proposed model to the pre-defined knowledge graph. Therefore, this work could provide a good basis or starting point for the research of unsupervised medical report generation in other clinical domains such as MRI and Dermoscopy.
2 Ablation Study
In Table 5, we conduct quantitative analysis to better understand our approach under both the unsupervised and supervised training settings. For different ablation settings, the KGAE-Supervised is acquired by further fine-tuning the KGAE using image-report pairs.
As we can see, under the unsupervised setting, both the introduced shared and knowledge bank can significantly boost the performance, which proves our arguments and verifies the effectiveness of our approach in performing the unsupervised medical report generation.
Under the supervised setting, as shown in settings (e,f), applying shared generates unchanged and impaired performance on the MIMIC-CXR and IU X-ray datasets, respectively. We speculate the reason is that the supervised model no longer requires the to bridge the vision and the language domains. Therefore, the performance is unchanged on the large dataset MIMIC-CXR. However, the increased parameters introduced by might bring overfitting or increase the difficulty in optimization on the small dataset IU X-ray, which somewhat hinders the performance. For the knowledge bank , we can find that the results of setting (h) outperforms (g) on the MIMIC-CXR dataset, but underperforms (g) on the IU X-ray dataset. We speculate the reason is that the large MIMIC-CXR dataset contains more knowledge than the small IU X-ray dataset, so a larger knowledge bank may learn more knowledge of MIMIC-CXR to boost the performance while introducing more noisy knowledge into the IU X-ray dataset to degrade the performance.
3 Qualitative Analysis
In Figure 3, we conduct the qualitative analysis to better understand our approach. As we can see, the visualization verifies the effectiveness of our knowledge-driven encoder in extracting the knowledge representations of both image and report. For the generated reports, our unsupervised KGAE generates a desirable report, which correctly describes “innumerable nodules are present” and “heart size is normal”. When removing the bank , the model tends to generate plausible general reports with no prominent abnormal narratives and some repeated reports, which shows that the knowledge memory mechanism can indeed distill and preserve the desired medical knowledge to boost the generation of reports. Under the semi-supervised setting, the R2Gen can not well handle the medical report generation task and generates some repeated sentences of normalities (Underlined text) and fails to depict some rare but important abnormalities, i.e., “nodules” and “scoliosis”, while our approach can generate fluent report supported by accurate abnormalities. Under the supervised setting, our approach can generate an accurate report showing significant alignment with the ground truth report. It further prove our arguments and the effectiveness of our proposed approach.
Conclusions and Discussions
In this paper, we propose the Knowledge Graph Auto-Encoder (KGAE). Without any image-report pairs, KGAE can extract the knowledge representations of both image and report from the knowledge graph to bridge the visual and textual domains, and generate desirable reports by being trained in the auto-encoding pipeline. The experiments verify the effectiveness of our approach, which even exceeds several supervised models. Moreover, by further fine-tuning KGAE using paired datasets, we achieve the state-of-the-art results on two public datasets with the best human preference.
In the future, 1) since we can relax the dependency on paired data, it can be interesting to apply the KGAE to other types of medical images of other body parts, where the image-report pairs could be much less or even unavailable, to assist radiologists in clinical decision-making and reduce their workload; 2) We can replace the explicit pre-defined knowledge graph with an implicit large matrix (e.g., knowledge bank in our decoder) to improve the generalization ability of our approach.
Societal Impacts: In this paper, we target the problem of medical report generation. Although the proposed model outperforms state-of-the-art approaches, it aims to assist the radiologists instead of replacing them. For experienced radiologists, given a large amount of medical images, our model can automatically generate medical reports, the radiologists only need to make revisions rather than write a new report from scratch. However, it is possible that some radiologists direct copy the generated report as the final report. Also for less experienced radiologists, they may not be able to correct the errors in machine-generated reports. In order to apply the proposed model in clinical practice, it is required to add process control to avoid unintended use.
Limitations: Although the proposed KGAE can work in an unsupervised manner, we still need the independent sets of medical images and medical reports which may still be difficult to collect for some types of medical images. In addition, our approach introduces the knowledge graph to bridge visual and textual domain, in the paper, we collect the frequent clinical findings as nodes and build the knowledge graph from the set of reports automatically. When applying to new domains, we need to collect a new set of clinical findings.