Federated Multi-View Learning for Private Medical Data Integration and Analysis

Sicong Che, Hao Peng, Lichao Sun, Yong Chen, Lifang He

Introduction

With the recent advance of technology, medical treatment has gradually become digitized, and a large and heterogeneous amount of medical data (e.g., EHR, claims, laboratory tests, imaging and genomics) has been accumulated in medical institutions. The availability of these data offers many potential benefits for healthcare. It not only facilitates sharing information in care-related activities, but also reduces medical errors and service time. Moreover, data sharing in the multi-institutional and large-scale collaboration context can effectively improve the generalizability of research, accelerate progress, enhance collaborations among institutions, and lead to new discoveries from data pooled from multiple sources (Lu et al., 2020). For these reasons, substantial efforts have been made in the past to improve the medical information systems enabling the collection and processing of huge amounts of data to work toward a multi-institutional collaborative setting. For example, District Health Information Software 2 (DHIS2) (Dehnavieh et al., 2019) is an open-source, web-based Health Management Information System (HMIS) widely used by countries for national-level aggregation of medical data. With the broad adoption of HMISs, security and privacy become critical (Singer et al., 1993), as medical data are highly sensitive to the patients.

In the last two decades, great efforts have been made to facilitate the availability and integrity of medical data while protecting confidentiality and privacy in medical information systems (Barrows Jr and Clayton, 1996; Barach and Small, 2000). Most of the research is centered around data de-identification (Sweeney, 2002; Wellner et al., 2007; El Emam, 2008; El Emam et al., 2009) and data anonymization (Deutsch and Papakonstantinou, 2005; Xiao and Tao, 2006; Machanavajjhala et al., 2007; Li et al., 2007), which removes the identifiable information from the published medical data to prevent an adversary from reasoning about the privacy of the patients. However, as pointed out in (Li et al., 2011), published medical data is not the only source that the adversaries can count on: with a large amount of information that people voluntarily share on the Web, sophisticated attacks that join disparate information pieces from multiple sources against medical data privacy become practical. As a result, many medical institutions are unwilling to share their data, as sharing may cause sensitive information to be leaked to researchers, other institutions, and unauthorized users. Even in multi-institutional collaborative research, they are expected to perform integrated analysis without leaving their medical data outside institutions. Therefore, security and privacy become an obstacle and challenge for data integration in medical research, especially when the data are required to be shared for secondary use.

However, machine learning requires a huge amount of data for better performance, and the current circumstances put deploying or developing medical AI applications in an extremely difficult situation. In order to address the data limitation and isolation issues, great progress has been made in the development of secure machine learning frameworks in recent years. A popular approach is the use of federated learning (FL) to support collaborative and distributed learning processes, which enables training machine learning models over remote devices or siloed data centers, such as mobile phones or hospitals, while keeping data localized (McMahan et al., 2017). For example, FL can be used to process EHR data distributed in multiple hospitals by sharing the local model instead of the patients’ data to prevent the raw data leakage (Boughorbel et al., 2019; Huang et al., 2019; Lee et al., 2018; Brisimi et al., 2018).

In the literature, many studies have explored the variants of FL schemes to support complicated tasks in real life. For example, Yang et al. (Yang et al., 2019) introduced a comprehensive federated learning framework that contains vertical federated learning (VFL) and horizontal federated learning (HFL) based on different local data availability. Specifically, VFL deals with the case that the datasets share the same sample ID space but their feature spaces are different, while HFL deals with the case that the datasets share the same feature space but are different in sample ID spaces. However, the previous works mainly focus on the difficulties in configuration of FL because of the data distributions but less consider the data complexity. In medical applications, many datasets are complex, heterogeneous and often collected from different instruments and/or measures, known as “multi-view” data (Xu et al., 2013). For example, electronic health records (EHRs) contain different types of patient-level variables, such as demographics, diagnoses, problem lists, medications, vital signs, and laboratory data. The mobile keyboard data are composed of various sensor data and keystroke records. It has been shown that multi-view learning can get better performance than the single view counterpart, especially when the strengths of one view complement the weaknesses of the other. Nevertheless, federated learning from multi-view data is still in its infancy, as most of the current solutions can only handle the single-view data. Moreover, existing works are only designed for a static setting without considering the temporal information of the data (e.g., EHRs).

Motivated by the aforementioned problems, in this paper we focus on multi-view learning tasks and associated problems of federated learning. Specifically, we first present a general Federated Multi-View Learning (FedMV) framework, which encompasses most of the existing multi-view fusion schemes and provides freedom to create their FL counterparts. Then, based on different types of local data availability, i.e., horizontal and vertical as shown in Fig. 1, we develop two multi-view federated algorithms: V-FedMV and H-FedMV. Both V-FedMV and H-FedMV approaches can perform the same tasks, but for different multi-view data distribution. In the V-FedMV approach, each client owns a single-view data, but they have common participants. In the H-FedMV approach, every client owns multi-view data equal to a subset of the overall data. Third, we investigate how multi-view sequential data can be arranged in the setting of sequential federated learning and present a sequential modelling strategy (S-FedMV) to consider the temporal information of the multi-view data. These methods together provide flexible and effective tools to support multi-institutional collaborations for multi-view data mining while solving the privacy and security challenges for data sharing in medical research.

Our contributions can be summarized as follows:

To the best of our knowledge, we are the first group that aims to propose a systematically solution for federated multi-view learning.

We have considered two types of distributed multi-view data and propose V-FedMV and H-FedMV to deal with these two cases separately.

This is the first work to consider multi-view sequential data in the federated learning setting, and S-FedMV method is developed in this regard.

Based on the experimental results, all three proposed methods (V-FedMV, H-FedMV, and S-FedMV) can make full use of multi-view data and are more effective than local training.

The rest of this paper is organized as follows. Section 2 describes the problem definition. Section 3 introduces the multi-view learning method and optimization. The V-FedMV and H-FedMV methods are presented in Section 4 and Section 5, respectively. Then in Section 6, we show how to adapt the model to deal with sequential data. In Section 7, we briefly discuss the privacy issue. The dataset, experiments and results are presented in Section 8. Section 9 discusses the related work, followed by the conclusion in Section 10.

PROBLEM DEFINITION

The key challenge to the adoption of distributed data is the security and privacy of highly sensitive data. For example, in many situations, medical data cannot leave the institutions, which limits the usefulness of these data to perform analytics in the data aggregation process. In order to solve this issue, federated learning has emerged as a promising technique for distributing machine learning (ML) model training, which performs local model training, and upload model parameters for global aggregation; thus, it enables collaborative model training while preserving each participant’s privacy, which is particularly beneficial to the medical field.

Based on the local data availability, we suppose that there are two types of federated multi-view learning: horizontal and vertical. Fig. 1 shows the example of vertical and horizontal multi-view scenarios. Our goal is to leverage all available data without sharing data between institutions for model learning in these respective scenarios, by distributing the model-training to the data-owners and aggregating their results. The problems are formally defined as follows:

Vertical Federated Multi-view Learning: It is supposed that each client shares the same ID space but has different single-view data. For example, several hospitals want to build a model to diagnose the bipolar affective disorder, they have some patients who have bipolar affective disorder in common because these patients have been to all these hospitals. However, each of them only has a single-view data collected from their own device. Our goal is to generate more accurate and robust classification results, by collectively training multiple single-view models through federated learning, so as to exploit vertical-level view information across multiple decentralized sites holding local private data samples without exchanging.

Horizontal Federated Multi-view Learning: It is supposed that each client owns multi-view data but shares different ID spaces, and each client’s data can be seen as a subset of the overall data. For example, each hospital has multiple devices and can obtain multiple views. Besides, the types of multiple views across different hospitals are the same. However, these hospitals have few or no patients in common. Our goal is to incorporate all samples in the model training to make the result more promising, by collaboratively training multiple multi-view models with the benefit of the federated learning framework.

MULTI-VIEW LEARNING

To deal with the federated multi-view learning problems in the vertical and horizontal situations, we start with a basic multi-view learning method to fuse multi-view data in this section, and then show how the formulas are modified to accommodate vertical and horizontal cases in the following two sections. Our approach is inspired by (Feng and Yu, 2020), which extended the idea of multi-view learning to study the multi-class VFL problems involving multiple parties.

Our aim is to get Wk\mathbf{W}_{k} from (1), k=1,⋯ ,Kk=1,\cdots,K, so when new testing data Xtest={X1test,⋯ ,XKtest}\mathbf{X}^{test}=\{\mathbf{X}_{1}^{test},\cdots,\mathbf{X}_{K}^{test}\} comes, we can use Wk\mathbf{W}_{k} to transform the data matrix Xktest\mathbf{X}_{k}^{test} to a label matrix. The objective function in the testing phase can be written as follows:

Finally, Ztest\mathbf{Z}^{test} will serve as a label matrix to facilitate the federated multi-view learning.

2. Optimization

We illustrate how to solve (1) in the training phase and how to solve (2) in the testing phase, respectively. Since the objective functions are non-convex and potentially non-smooth, we iteratively update the parameters one by one while fixing the other parameters (Feng and Yu, 2020). The overall algorithm is summarized in Algorithm 1.

In training phase, all parameters are iteratively updated according to the following three steps until (1) converges or reaches a maximum number of iterations.

Step 1: update Wk\mathbf{W}_{k}. When Zk\mathbf{Z}_{k} and Z\mathbf{Z} are fixed, the optimization problem of minimizing (1) over Wk\mathbf{W}_{k} can be written as

Following (Hou et al., 2013), (3) can be written as

Then, when Ak\mathbf{A}_{k} is fixed, Wk\mathbf{W}_{k} can be updated by:

So when Zk\mathbf{Z}_{k} and Z\mathbf{Z} are fixed, we can repeat the following operation to get Wk\mathbf{W}_{k} until the value of (3) converges: fix Wk\mathbf{W}_{k} to update Ak\mathbf{A}_{k} through (5) and then fix Ak\mathbf{A}_{k} to update Wk\mathbf{W}_{k} through (6).

Step 2: update Zk\mathbf{Z}_{k}. When Z\mathbf{Z} and Wk\mathbf{W}_{k} are fixed, (1) becomes:

Then Zk\mathbf{Z}_{k} can be updated directly by:

Step 3: update Z\mathbf{Z}. When Zk\mathbf{Z}_{k} and Wk\mathbf{W}_{k} are fixed, (1) becomes:

Then Z\mathbf{Z} can be updated directly by:

2.2. Testing Phase

Firstly, initialize Zktest\mathbf{Z}^{test}_{k} as Zktest=XktestWk\mathbf{Z}^{test}_{k}=\mathbf{X}^{test}_{k}\mathbf{W}_{k}. When Zktest\mathbf{Z}^{test}_{k} is fixed, (2) can be written as:

Then Ztest\mathbf{Z}^{test} can be updated directly by:

When Ztest\mathbf{Z}^{test} is fixed, Then Zktest\mathbf{Z}^{test}_{k} can be updated directly by:

VERTICAL FEDERATED MULTI-VIEW LEARNING

In this section, we elaborate on how proper design of the above multi-view learning mechanism can yield guaranteed vertical federated multi-view learning (V-FedMV) gains which consequently not only allows to connect multi-view relations over a global server for remote collaboration, but also to set up an environment across multiple decentralized data to distribute the model-training to the data-owners, without sharing data between institutions.

Fig. 2 provides an overview of the V-FedMV framework for vertical accumulation of different views of data for multi-view classification. The overall algorithm is summarized in Algorithm 2. Briefly, the main procedures of V-FedMV can be described as follows: We first establish a server with multiple distributed clients, then the computations involving the original sensitive data will be executed on clients, while some other computations related to insensitive data of all clients will be executed on the server. For example, (5), (6), (8) and (13) can be computed on each client LkL_{k}, while (10) and (12) can be computed on the server.

HORIZONTAL FEDERATED MULTI-VIEW LEARNING

In this section, we illustrate how multi-view learning mechanism can be used for horizontal federated multi-view learning (H-FedMV), which consequently allows to leverage all available multi-view samples from a collection of decentralized local data over a central global server for remote collaboration, but without delivering any raw data to the global server. Fig. 3 provides an overview of the H-FedMV framework for horizontal accumulation of different samples of data for multi-view classification.

The overall algorithm of H-FedMV is summarized in Algorithm 3. Briefly, the main procedures can be described as follows: firstly, all clients apply the algorithm of the multi-view learning on their own data and devices, then each of them can get Wkl\mathbf{W}_{k}^{l}, which are sent to the global server to compute the weighted average value. After the computation on server, it sends the output weighted average transformation matrix Wk\mathbf{W}_{k} back to the local client LlL_{l} for distributed model training. This process will be repeated for several times until convergence is reached or some stopping criteria are met. At last, the Wk\mathbf{W}_{k} of all KK views are obtained on the server and then server will send them to clients, which can be used to predict testing data of each client.

FEDERATED MULTI-VIEW SEQUENTIAL LEARNING

In medical practice, sequential or longitudinal data is very common, such as clinical research and epidemiological studies, However, V-FedMV and H-FedMV cannot be directly applied to sequential data for considering the temporal information of the multi-view data. We further propose a federated multi-view sequential learning method to adapt our model to deal with multi-view sequential data in a federated environment.

In the vertical setting, each client contains a single-view data, and can be independently executed (such as using GRU) on their own data to get the feature embedding matrix from sequential data and then apply the proposed V-FedMV method, which will not cause expose the raw data in the whole process. However, in the horizontal setting, each client owns multi-view data partially and they cannot process their own data independently, because the trained GRU models for each view are inconsistent locally. Hence, how to do federated multi-view learning in the horizontal setting when the raw data is sequential is a problem. Here we develop a sequential version of H-FedMV to illustrate how multi-view sequential data can be arranged by using federated learning, named S-FedMV.

Fig. 4 provides an overview of the S-FedMV framework for horizontal accumulation of different samples of multi-view sequential data for feature extraction. Briefly, the idea is based on the Federated Averaging (FedAvg), following a server-client setup with two repeated stages: (i) the clients train their models locally on their data, and (ii) the server collects and aggregates the models to obtain a global model by weighted averaging. For each client, we assign the aggregation weight as the ratio of data samples on each client to the total number of samples, namely Nl/NN_{l}/N. It is appropriate to process the data of each view separately in the case of distributed horizontal multi-view data, because there exist inconsistency among the models of different views. FedAvg is flexible to the model and the optimizer used for training, here we use the bidirectional GRU as the model and RMSProp as the optimizer in the experiments. For simplicity, in every round we choose all clients to participate in.

The overall algorithm is summarized in Algorithm 4, and the notations used in the proposed model are as below: As before, we use {Ll}\{L_{l}\}, (l=1,...,M)(l=1,...,M) to denote MM clients. Let {Dkl}\{D^{l}_{k}\} denote the sequential data on ll-th client in the kk-th view, wtk\mathbf{w}_{t}^{k} and wtk,l\mathbf{w}_{t}^{k,l} denote the parameters of kk-th view model on server and on the client in the tt round, and l(w;b)\mathscr{l}(\mathbf{w};b) denote the loss of the model with parameters w\mathbf{w} and bb, where bb is the bias.

Discussion

In V-FedMV, both Wk\mathbf{W}_{k} and Zk\mathbf{Z}_{k} are updated on local devices while Z\mathbf{Z} is updated on the server. ζk\zeta_{k} and Zk\mathbf{Z}_{k} are sent to the server from local devices when the server updates Z\mathbf{Z}, which would not leak the raw data because Xk\mathbf{X}_{k} is preserved at local devices throughout the process. Besides, in some cases, the server may not be set up by these clients, but by another specialized organization that would not provide the raw data but only do the computation, and in this case, although Y\mathbf{Y} is also needed when server updates Z\mathbf{Z}, it won’t cause leakage because Y\mathbf{Y} is only a one-hot label matrix which would not be useful when don’t know extra information.

In H-FedMV and S-FedMV, each client sends the parameters of its own model to the server in each round, and the server averages it and returns the result. Each client then continues to update locally with the parameters returned by the server. In the whole process, the server only contacts the parameters, not the raw data, and the local devices can only contact the updated parameters from the server, not the raw data from other clients, so this ensures privacy.

Experimental Evaluation

In this section, we evaluate the performance of the proposed V-FedMV, H-FedMV, and S-FedMV on a new multi-view sequential keystroke data which is collected from the BiAffecthttps://www.biaffect.com study.

BiAffect1, the first study on mood and cognition using mobile typing kinematics, provides the multi-view sequential data for our experiments. BiAffect invited 40 participants to use the customized smartphones in their daily life. These phones are equipped with a custom keyboard that would collect the data of keypress duration, typing behaviors, accelerometer value, and others. In the experiment, we use three types of metadata: alphanumeric characters, special characters, and accelerometer values, which can be seen as three views: Alphanumeric characters include the duration time of keypress, time consumed since the last key was pressed, and the distance between the current key and the previous key along two axes; Special characters are auto-correct, backspace, space, suggestion, switching-keyboard, and other special characters; Accelerometer values are collected by the sensor of the smartphone. The device records accelerometer values every 60ms, regardless of typing speed.

The diagnosis of patients with bipolar disorder was obtained by weekly assessment of participants through Hamilton Depression Rating Scale (HDRS) (Williams, 1988) and Young Mania Rating Scale (YMRS) (Young et al., 1978), which is a very reliable assessment standard for bipolar disorder. There are 7 participants with bipolar I disorder, 5 participants with bipolar II disorder, and 8 participants with no diagnosis per DSM-IV TR criteria (Kessler et al., 2005).

2. Experimental Setup

In the experiment, we investigate a session-based mood prediction problem same as prior works in the BiAffect project (Cao et al., 2017), which utilizes all three views to predict a user’s mood score. However, our goal is to train a binary mood classification. To get the binary labels, after consulting with the professional experts, we label the sessions with the HDRS score between 0 and 7 (inclusive) as negative samples and those with the HDRS score higher than score 7 as positive samples.

However, the original input data is not naturally fitted our proposed V-FedMV and H-FedMV, because the original keystroke data is time-sequential format, but these methods are designed for feature matrices. To address the data format problem, we applied a GRU (Cho et al., 2014) which is a simplified version of Long Short-Term Memory (LSTM) (Hochreiter and Schmidhuber, 1997) to each view respectively to preprocess the data and extract feature embedding matrix as the input of V-FedMV and H-FedMV methods.

In this project, we use Keras with Tensorflow as the backend to implement the code. We use RMSProp (Tieleman and Hinton, 2012) as the optimizer for GRU training. We retain sessions that any view contains the number of keypress between 10 and 100, and then we have 14,971 total samples. Besides, we set the batch size as 128, the epoch as 500, the learning rate from {0.001,0.005}\{0.001,0.005\}, and the dropout from {0.1,0.3}\{0.1,0.3\}. We use the validation dataset to select the optimal parameters and get the feature matrices which are the input of our V-FedMV and H-FedMV frameworks. Furthermore, the visualization of three views after data preprocessing is shown in Fig. 5a, demonstrating the original space of the input data of V-FedMV and H-FedMV. The visualizations of each view and Z\mathbf{Z} from V-FedMV and H-FedMV are shown in Fig. 5b and Fig. 5c, respectively. It is not hard to see after V-FedMV and H-FedMV, the data can be well classified.

In addition, we conduct various experiments to evaluate the performance of S-FedMV. The max communication round is 150 for alphanumeric and special views, 170 for accel view. We set the local epoch from {10,15}\{10,15\}, the dropout from {0.1,0.3}\{0.1,0.3\}. and the batch size as 128. The same as previous approaches, we select the optimal parameters through validation and finally get the best feature matrices.

After data preprocessing, we carry out experiments on V-FedMV and H-FedMV. In V-FedMV, we set up three clients, split the processed feature matrices according to the type of view, and make each client own one of three views. In H-FedMV, we set four clients, split the processed feature matrices, and make each client own the same number of samples with all three views. One more thing is that we also equally assigned positive and negative samples to each client for H-FedMV.

For more details of hyperparameter setting, V-FedMV uses βk\beta_{k} as 4, ζk\zeta_{k} and η\eta from 202^{0} to 252^{5}. In H-FedMV, we set the max communication rounds as 20, and the value of βkl\beta_{k}^{l} as 4, and ζkl\zeta_{k}^{l}, ηl\eta^{l} across different clients to be the same from 202^{0} to 252^{5}. For both V-FedMV and H-FedMV, the optimal parameters are also selected by validation. We repeat all the experiments ten times with four metrics, i.e., accuracy, precision, recall, and f1 score. In addition, we also calculate the average and standard deviation.

3. Vertical Federated Multi-View Learning

In this part, we evaluate V-FedMV with six other baselines. All approaches are summarized as follows:

V-FedMV: It is our proposed federated multi-view learning approach in the vertical setting, described in Algorithm 2. Note that, all three views are used in the experiment of V-FedMV.

Pairwise FL: These approaches are similar to V-FedMV that also applies the federated multi-view learning algorithm in vertical setting. The difference is that only two views are considered. In this case, we have three experiments with two views, including Pairwise FL w/o Special, Pairwise FL w/o Accel, Pairwise FL w/o Alphanum.

Single View w/o FL (Zhao et al., 2010): In this approach, we optimize the following objective function on each view:

where Y\mathbf{Y} in (20) refers to the one-hot label matrix. Similar to pairwise FL, we also have three single view w/o FL experiments, such as Alphanum w/o FL, Accel w/o FL, Special w/o FL.

Experimental results of V-FedMV and its baselines are shown in Table 1. It is not hard to see that V-FedMV shows the best performance with 88.76% accuracy and 88.93% f1 score. All methods have a low standard deviation, which means all approaches are very stable in the repeated experiments. Meanwhile, we also find the special view makes the most negligible contribution for this task based on the experimental results of Pairwise FL w/o Special and Special w/o FL. Especially, the Pairwise FL w/o Special approach can perform almost the same as V-FedMV. One of the main reasons is that the special view provides less information than the other two views. Besides, we can see that the alphanum view achieves the best accuracy (84.53%) and f1 score (84.59%) among all three views, and the accel view is between the others. While comparing all approaches together, we can easily conclude that the model’s performance increases with more views, demonstrating that the proposed method is effective for multi-view data. Finally, the proposed approach V-FedMV improves 4.23% accuracy, 3.31% precision, 5.28% recall, and 4.34 % f1 score than Alphanum w/o FL that is the best single view learning without FL.

4. Horizontal Federated Multi-View Learning

Here, we evaluate H-FedMV with seven other baselines. All approaches are summarized as follows:

H-FedMV: It is our proposed federated multi-view learning approach in the horizontal setting, described in Algorithm 3. H-FedMV also uses all three views like V-FedMV.

MV w/o FL: In this approach, each client applies the multi-view learning method on its local device. All client don’t share the parameters to the server for aggregation. Same as H-FedMV, MV w/o FL also uses all three views.

Pairwise FL: These approaches also apply the proposed multi-view federated learning approach in the horizontal setting, but with only two views. Thus, we have three experiments with two views, including Pairwise FL w/o Special, Pairwise FL w/o Accel, Pairwise FL w/o Alphanum.

Single View FL: In this approach, the data on clients are represented by a single view. All clients optimize Eq. (20) on their own devices firstly. Then each client sends its Wk\mathbf{W}_{k} to the server for aggregation. After the aggregation is done, the global Wk\mathbf{W}_{k} will be back to upload the local model for all clients. This process will continue until the maximum number of communication rounds is reached. Similar to pairwise FL, we also have three single view FL experiments: Single View FL - Alphanum, Single View FL - Accel, Single View FL - Special.

Table 2 shows the experimental results of H-FedMV and its baselines. It can be observed that H-FedMV outperforms the compared baselines with 88.98% accuracy and 89.17% f1 score. MV w/o FL trains each model locally and the average performance of all local models is worse than H-FedMV. One of the main reason is that federated learning enables H-FedMV to get more training sample information than MV w/o FL, resulting in a better performance. In addition, the significance of each view in model training is similar to the experimental results of H-FedMV and its baselines in last section, i.e., alphanum view ¿ accel view ¿ special view. Eventually, the proposed H-FedMV improves 1.06% accuracy, 0.8% precision, 1.27% recall, and 1.08% f1 score compared to MV w/o FL.

5. Federated Multi-View Sequential Learning

We further evaluate the S-FedMV approach with three other baselines.

S-FedMV: It is our proposed federated learning framework for multi-view sequential data, where each client interacts with the global model with H-FedMV.

Local Sequential H-FedMV: Each client runs feature representation learning locally with H-FedMV, which indicates no interactions of feature learning between any two clients.

Local Sequential LocalMV: There is no federated learning involved in this approach. Each client trains the multi-view sequential model by only using the local data.

Centralized Sequential H-FedMV: Centralized sequential approach has the access to all training data across all clients for feature leaning with H-FedMV. In general, it should achieve the best performance due to the best feature representation in federated learning.

Experimental results of different multi-view federated learning methods for sequential data are shown in Table 3. First, we can see that S-FedMV achieves outstanding performance with 88.90% accuracy and 89.22% f1 score. It demonstrates that S-FedMV is the best approach for distributed multi-view sequential data since its performance is close to centralized sequential H-FedMV. Note that centralized sequential H-FedMV is the up-bound performance of all approaches since the feature representation learning can access all training data. Then, we find the results of local sequential H-FedMV can confirm our statement that it is infeasible for each client to preprocess its data independently. It is because the trained models are inconsistent between different clients, and they obtained a bad representation learning of H-FedMV. We can also see that the proposed approach S-FedMV is better than local sequential LocalMV, which trains the local model on local sequential data only.

6. Hyperparameter Analysis

We use the hyperparameter ζ1,ζ2,ζ3\zeta_{1},\zeta_{2},\zeta_{3} to correspond with alphanumeric view, special view and accel view respectively, which controls the balance between three views. For instance, a large ζk\zeta_{k} indicates a higher impact of the kk-th view of model training. In this experimental setting, we test each hyperparameter ∈{20,21,22,23,24,25}\in\{2^{0},2^{1},2^{2},2^{3},2^{4},2^{5}\}. Note that while testing one hyperparameter, we fix the others as 232^{3}. We show the evaluation results with four metrics, i.e., accuracy, precision, recall, f1 scores shown in Fig. 6. We find that the hyperparameters affect the accuracy and f1 score consistently, which shows a similar trend line for the same hyperparameter evaluation. In summary, while ζ1=24,ζ2=22\zeta_{1}=2^{4},\zeta_{2}=2^{2} and ζ3=23\zeta_{3}=2^{3}, the model can achieve the best performance on accuracy and f1 score for both H-FedMV and V-FedMV approaches. While η=24\eta=2^{4}, H-FedMV is better than V-FedMV, but V-FedMV is better than H-FedMV once η=25\eta=2^{5}. Moreover, we also find the performance of the trained model is most sensitive to ζ1\zeta_{1} and ζ2\zeta_{2} and less sensitive to η\eta.

RELATED WORK

In this section, we review the related work, which can be placed into three main categories: multi-view learning, federated learning, and federated multi-view learning.

Multi-View Learning: Sun et al. (Sun, 2013) provided a survey for multi-view machine learning and pointed out that multi-view learning is related to the machine learning problem with the data represented by multiple distinct feature sets. Xu et al. analyzed different multi-view algorithms and indicated that it was consensus and complementary principles that ensure their promising performance (Xu et al., 2013). The aim of the consensus principle is to minimize the disagreement on multiple distinct views. A connection between the consensus of two hypotheses on two views respectively and their error rates was given by Dasgupta et al. (Dasgupta et al., 2002). The complementary principle means that multiple views can be complements for each other and can be exploited comprehensively to produce better learning performance for the reason that each view may contain some specific information that other views do not have. Wang and Zhou (Wang and Zhou, 2007) demonstrated that the performance of co-training algorithms was largely affected by the complementary information in distinct views. These two principles are very important for multi-view learning and should be taken into consideration when designing multi-view learning algorithms.

Xu et al. also categorized the classical approaches of combining multiple views into co-training style algorithms, multiple kernel learning algorithms, and subspace learning-based approaches (Xu et al., 2013).

Co-training style algorithm: Blum and Mitchell (Blum and Mitchell, 1998) proposed the original co-training algorithm to solve semi-supervised classification problems. Firstly, two classifiers are trained separately on each view, and then each classifier labels the unlabeled data which are then added to the training set of another classifier. Besides, Kumar and Daumé (Kumar and Daumé, 2011) extended the idea of co-training to an unsupervised setting. Since co-training algorithms usually separately train the base learners, it can be seen as a late combination method.

Multiple kernel learning algorithm: Combining different kernels is another way to integrate multiple views and can be regarded as an intermediate method for the reason that kernels are integrated just before or during the training phase. These methods include linear combination methods (Lanckriet et al., 2004; Joachims et al., 2001) and nonlinear combination methods (Cortes et al., 2009).

Subspace learning-based approach: In the subspace learning-based approach, there is an assumption that multiple views are generated from a latent subspace. This approach can be regarded as a prior combination of multiple views and the goal is to obtain the latent subspace. Canonical correlation analysis (CCA) (Hotelling, 1992) is a classical subspace learning-based approach that can be applied to the datasets that contain two views. Besides, it can be extended to cope with datasets represented by more than two views (Kettenring, 1971) and to kernel CCA (Fyfe and Lai, 2000).

Federated Learning: Federated Learning (FL) was proposed by McMahan et al. (McMahan et al., 2017). It is a collaborative machine learning paradigm for training models based on locally stored data from multiple organizations in a privacy-preserving way. Yang et al. gave a comprehensive survey for FL (Yang et al., 2019), which introduced horizontal federated learning, vertical federated learning, and federated transfer learning.

Horizontal Federated Learning: In horizontal federated learning, or sample-based federated learning, the data sets share the same feature space but have different sample ID space (Yang et al., 2019). Smith et al. proposed a novel framework for federated multi-task learning and considered high communication cost, stragglers, and fault tolerance in the federated environment for the first time (Smith et al., 2017). Bonawitz et al. designed a secure aggregation scheme that allows a server to securely compute users’ data from mobile devices (Bonawitz et al., 2017).

Vertical Federated Learning: In vertical federated learning, or feature-based federated learning, the data sets share the same sample ID space but have different feature space (Yang et al., 2019). Some algorithms and models have been proposed for vertical federated learning. (Gascón et al., 2016; Karr et al., 2009; Sanil et al., 2004) are about secure linear regression. Du et al. defined two types of secure 2-party multivariate statistical analysis problems: linear regression and classification problem. Secure methods also proposed to solve these two problems (Du et al., 2004). Liu et al. proposed asymmetrical vertical federated learning and showed the way to achieve the asymmetrical ID alignment (Liu et al., 2020b).

Federated Transfer Learning: In federated transfer learning, the data sets are different in both samples and feature space (Yang et al., 2019). Liu et al. proposed an end-to-end approach to the FTL problem and demonstrated it was comparable to transfer learning without privacy protection (Liu et al., 2020a). A federated transfer learning framework for wearable healthcare was proposed in (Chen et al., 2020).

Federated Multi-View Learning: Federated learning with multi-view data is currently less well explored in the literature. Adrian Flanagan et al. integrated multi-view matrix factorization with a federated learning framework for personalized recommendations and introduced a solution to the cold-start problem (Flanagan et al., 2020). Huang et al. proposed FL-MV-DSSM, which is a generic content-based federated multi-view framework for recommendation scenarios and can address the cold-start problem (Huang et al., 2020). Xu et al. (Xu et al., 2021) extend the DeepMood model (Cao et al., 2017) to a multi-view federated learning framework that suits the horizontal case. Feng et al. extended the idea of multi-view learning and proposed MMVFL that enables label sharing from its owner to other participants (Feng and Yu, 2020). MMVFL suits for vertical setting and it can deal with multi-participant and multi-class problems while the other existing VFL approaches can only handle two participants and binary classification problems. Kim et al. proposed a federated tensor factorization framework for horizontally partitioned data (Kim et al., 2017). To sum up, today’s federated multi-view learning is mainly applied to recommendation systems or tenor data, or only consider the horizontal or vertical situation, rather than consider the two together.

Conclusion

In this paper, we proposed a generic multi-view learning framework by using federated learning paradigm for privacy-preserving and secure sharing of medical data among institutions, which can well protect the patient privacy by keeping the data within their own confines and achieve multi-view data integration. Specifically, we investigated two types of multi-view learning (i.e., vertical and horizontal data integration) in the setting of federated learning based on different local data availability, and developed the vertical federated multi-view learning (V-FedMV) and horizontal federated multi-view learning (H-FedMV) algorithms to solve this problem. Moreover, we adapted our model to deal with multi-view sequential data in a federated environment and introduced a federated multi-view sequential learning (S-FedMV) method. Extensive experiments on real-world keyboard data demonstrated that our methods can make full use of multi-view data and obtain better classification results as compared to local training. Moreover, the result of S-FedMV is comparable with the result of centralized method that cannot protect the privacy of sensitive data and thus showing the effectiveness of our federated method.

Acknowledgments

This work is supported by NSF ONR N00014-18-1-2009. Hao Peng is supported by NSFC program (No. 62002007 and U20B2053), and S&T Program of Hebei through grant 20310101D.

References