On Deep Multi-View Representation Learning: Objectives and Optimization

Weiran Wang, Raman Arora, Karen Livescu, Jeff Bilmes

Introduction

In many applications, we have access to multiple “views” of data at training time while only one view is available at test time, or for a downstream task. The views can be multiple measurement modalities, such as simultaneously recorded audio + video (Kidron et al., 2005; Chaudhuri et al., 2009), audio + articulation (Arora and Livescu, 2013; Wang et al., 2015a), images + text (Hardoon et al., 2004; Socher and Li, 2010; Hodosh et al., 2013; Yan and Mikolajczyk, 2015), or parallel text in two languages (Vinokourov et al., 2003; Haghighi et al., 2008; Chandar et al., 2014; Faruqui and Dyer, 2014; Lu et al., 2015), but may also be different information extracted from the same source, such as words + context (Pennington et al., 2014) or document text + text of inbound hyperlinks (Bickel and Scheffer, 2004). The presence of multiple information sources presents an opportunity to learn better representations (features) by analyzing the views simultaneously. Typical approaches are based on learning a feature transformation of the “primary” view (the one available at test time) that captures useful information from the second view using a paired two-view training set. Under certain assumptions, theoretical results exist showing the advantages of multi-view techniques for downstream tasks (Kakade and Foster, 2007; Foster et al., 2009; Chaudhuri et al., 2009). Experimentally, prior work has shown the benefit of multi-view methods on tasks such as retrieval (Vinokourov et al., 2003; Hardoon et al., 2004; Socher and Li, 2010; Hodosh et al., 2013), clustering (Blaschko and Lampert, 2008; Chaudhuri et al., 2009), and classification/recognition (Dhillon et al., 2011; Arora and Livescu, 2013; Ngiam et al., 2011).

Recent work has introduced several approaches for multi-view representation learning based on deep neural networks (DNNs), using two main training criteria (objectives). One type of objective is based on deep autoencoders (Hinton and Salakhutdinov, 2006), where the objective is to learn a compact representation that best reconstructs the inputs. In the multi-view learning scenario, it is natural to use an encoder to extract the shared representation from the primary view, and use different decoders to reconstruct each view’s input features from the shared representation. This approach has been shown to be effective for speech and vision tasks (Ngiam et al., 2011).

The second main DNN-based multi-view approach is based on deep extensions of canonical correlation analysis (CCA, Hotelling, 1936), which learns features in two views that are maximally correlated. The CCA objective has been studied extensively and has a number of useful properties and interpretations (Borga, 2001; Bach and Jordan, 2002, 2005; Chechik et al., 2005), and the optimal linear projection mappings can be obtained by solving an eigenvalue system of a matrix whose dimensions equal the input dimensionalities. To overcome the limiting power of linear projections, a nonlinear extension–kernel canonical correlation analysis (KCCA)–has also been proposed (Lai and Fyfe, 2000; Akaho, 2001; Melzer et al., 2001; Bach and Jordan, 2002; Hardoon et al., 2004). CCA and KCCA have long been the workhorse for multi-view feature learning and dimensionality reduction (Vinokourov et al., 2003; Kakade and Foster, 2007; Socher and Li, 2010; Dhillon et al., 2011). Several alternative nonlinear CCA-like approaches based on neural networks have also been proposed (Lai and Fyfe, 1999; Hsieh, 2000), but the full DNN extension of CCA, termed deep CCA (DCCA, Andrew et al., 2013) has been developed only recently. Compared to kernel methods, DNNs are more scalable to large amounts of training data and also have the potential advantage that the non-linear mapping can be learned, unlike with a fixed kernel approach.

The contributions of this paper are as follows.We present a head-to-head comparison of several DNN-based approaches, along with linear and kernel CCA, in the unsupervised multi-view feature learning setting where the second view is not available at test time. We compare approaches based on prior work, as well as developing and comparing new variants. Empirically, we find that CCA-based approaches tend to outperform unconstrained reconstruction-based approaches. One of the new methods we propose, a DNN-based model combining CCA and autoencoder-based terms, is the consistent winner across several tasks. Finally, we study a stochastic optimization approach for deep CCA, both empirically and theoretically. To facilitate future work, we have released our implementations and a new benchmark dataset of simulated two-view data based on MNIST.Data and implementations can be found at http://ttic.uchicago.edu/~wwang5/dccae.html.

An early version of this work appeared in (Wang et al., 2015b). This paper expands on that work with expanded discussion of related work (Sections 3.2 and 3.3), additional experiments, and theoretical analysis. In particular we include additional experiments on word similarity tasks with multilingual word embedding learning (Section 4.3); extensive empirical comparisons between batch and stochastic optimization methods for DCCA; and comparisons between DCCA and two popular low-rank approximate KCCA methods, demonstrating their computational trade-offs (Section 4.4). We also analyze the stochastic optimization approach for DCCA theoretically in Appendix A.

DNN-based multi-view feature learning

In this section, we discuss several existing and new multi-view learning approaches based on deep feed-forward neural networks, including their objective functions and optimization procedures. Schematic diagrams summarizing the methods are given in Fig. 1.

1 Split autoencoders (SplitAE)

The intuition for this model is that the shared representation can be extracted from a single view, and can be used to reconstruct all views.The authors also propose a bimodal deep autoencoder combining DNN transformed features from both views; this model is more natural for the multimodal fusion setting, where both views are available at test time, but can also be used in our multi-view setting (see Ngiam et al., 2011). Empirically, however, Ngiam et al. (2011) report that SplitAE tends to work better in the multi-view setting than bimodal autoencoders. The autoencoder loss is the empirical expectation of the loss incurred at each training sample, and thus stochastic gradient descent (SGD) can be used to optimize the objective efficiently with the gradients estimated on small minibatches of samples.

2 Deep canonical correlation analysis (DCCA)

Andrew et al. (2013) propose a DNN extension of CCA termed deep CCA (DCCA; see Fig. 1 (b)). In DCCA, two DNNs f\mathbf{f} and g\mathbf{g} are used to extract nonlinear features for each view and the canonical correlation between the extracted features f(X)\mathbf{f}(\mathbf{X}) and g(Y)\mathbf{g}(\mathbf{Y}) is maximized:

where U=[u1,…,uL]\mathbf{U}=[\mathbf{u}_{1},\dots,\mathbf{u}_{L}] and V=[v1,…,vL]\mathbf{V}=[\mathbf{v}_{1},\dots,\mathbf{v}_{L}] are the CCA directions that project the DNN outputs and (rx,ry)>0(r_{x},r_{y})>0 are regularization parameters added to the diagonal of the sample auto-covariance matrices (Bie and Moor, 2003; Hardoon et al., 2004). In DCCA, U⊤f(⋅)\mathbf{U}^{\top}\mathbf{f}(\cdot) is the final projection mapping used for downstream tasks.In principle there is no need for the final linear projection; we could define DCCA such that the correlation objective and constraints are imposed on the final nonlinear layer of the two DNNs. The final linear projection is equivalent to constraining the final DNN layer to be linear. While this is in principle not needed, it is crucial for practical algorithmic implementations such as ours, and matches the original formulation of DCCA (Andrew et al., 2013). One intuition for CCA-based objectives is that, while it may be difficult to accurately reconstruct one view from the other view, it may be easier, and perhaps sufficient, to learn a predictor of a function (or subspace) of the second view. In addition, it should be helpful for the learned dimensions within each view to be uncorrelated so that they provide complementary information.

The DCCA objective couples all training samples through the whitening constraints (i.e., it is a fully batch objective), so stochastic gradient descent (SGD) cannot be applied in a standard way. However, it is observed that DCCA can still be optimized effectively as long as the gradient is estimated using a sufficiently large minibatch (Wang et al., 2015a, with the gradient formulas given as in Andrew et al., 2013). Intuitively, this approach works because a large minibatch of samples contains sufficient information for estimating the covariance matrices. We provide empirical analysis about the optimization of DCCA (in comparison with KCCA) in Section 4.4 and its theoretical analysis is given in Appendix A.

3 Deep canonically correlated autoencoders (DCCAE)

Inspired by both CCA and reconstruction-based objectives, we propose a new model that consists of two autoencoders and optimizes the combination of canonical correlation between the learned “bottleneck” representations and the reconstruction errors of the autoencoders. In other words, we optimize the following objective

where λ>0\lambda>0 is a trade-off parameter. Alternatively, this approach can be seen as adding an autoencoder regularization term to DCCA. We call this approach deep canonically correlated autoencoders (DCCAE). Fig. 1 (c) shows a schematic representation of the approach.

We apply stochastic optimization to the DCCAE objective. Notice obtaining good stochastic estimates of the gradient for the correlation and autoencoder terms may involve different minibatch sizes. We explore the minibatch sizes for each term separately (by training DCCA and an autoencoder) and select them using a validation set. The stochastic gradient is then the sum of the gradient for the DCCA term (usually estimated using large minibatches) and the gradient for the autoencoder term (usually estimated using small minibatches).

Interpretations

CCA maximizes the mutual information between the projected views for certain distributions (Borga, 2001), while training an autoencoder to minimize reconstruction error amounts to maximizing a lower bound on the mutual information between inputs and learned features (Vincent et al., 2010). The DCCAE objective offers a trade-off between the information captured in the (input, feature) mapping within each view on the one hand, and the information in the (feature, feature) relationship across views.

4 Correlated autoencoders (CorrAE)

In the next approach, we remove the uncorrelatedness constraints from the DCCAE objective, leaving only the sum of scalar correlations between pairs of learned dimensions and the reconstruction error term. This approach is intended to test how important the original CCA constraints are. We call this model correlated autoencoders (CorrAE), also represented by Fig. 1 (c). Its objective can be written as

where λ>0\lambda>0 is a trade-off parameter. It is clear that the constraint set in (3) is a relaxed version of that of (2). Later we demonstrate that this difference results in a large performance gap. We apply the same optimization strategy of DCCAE to CorrAE.

CorrAE is similar to the model of Chandar et al. (2014, 2015), who try to learn vectorial word representations using parallel corpora from two languages. They use a DNN in each view (language) to predict a bag-of-words representation of the input sentences, or that of the paired sentences from the other view, while encouraging the learned bottleneck layer representations to be highly correlated.

5 Minimum-distance autoencoders (DistAE)

The CCA objective can be seen as minimizing the distance between the learned projections of the two views, while satisfying the whitening constraints for the projections (Hardoon et al., 2004). The constraints complicate the optimization of CCA-based objectives, as pointed out above. This observation motivates us to consider additional objectives that decompose into sums over training examples, while maintaining the intuition of the CCA objective as a reconstruction error between two mappings. Here we consider two variants that we refer to as minimum-distance autoencoders (DistAE).

The first variant DistAE-1 optimizes the following objective:

which is a weighted combination of reconstruction errors of two autoencoders and the average discrepancy between the projected sample pairs. The denominator of the discrepancy term is used to keep the optimization from improving the objective by simply scaling down the projections (although they can never become identically zero due to the reconstruction terms). This objective is unconstrained and is the empirical average of the loss incurred at each training sample, so normal SGD applies using small (or any size) minibatches.

The second variant DistAE-2 optimizes a somewhat different objective:

Related work

Here we focus on related work on multi-view feature learning using neural networks and the kernel extension of CCA.

There have been several approaches to multi-view representation learning using neural networks with an objective similar to that of CCA. Under the assumption that the two views share a common cause (e.g., depth is the common cause for adjacent patches of images), Becker and Hinton (1992) propose to maximize a sample-based estimate of mutual information between outputs of neural networks for the two views. In this work, the 1-dimensional output of each network can be considered to be “internal teaching signals” for the other. It is less straightforward to extend their sample-based estimator of mutual information to higher dimensions, while the CCA objective is always closely related to maximal mutual information between the views (under the joint multivariate Gaussian distributions of the inputs, see Borga, 2001).

Lai and Fyfe (1999) propose to optimize the correlation (rather than canonical correlation) between the outputs of networks for each view, subject to scale constraints on each output dimension. Instead of directly solving this constrained formulation, the authors apply Lagrangian relaxation and solve the resulting unconstrained objective using SGD. Note, however, that their objective is different from that of CCA, as there are no constraints that the learned dimensions within each view be uncorrelated. Hsieh (2000) proposes a neural network-based model involving three modules: one module for extracting a pair of maximally correlated one-dimensional features for the two views; and a second and third module for reconstructing the original inputs of the two views from the learned features. In this model, the feature dimensions can be learned one after another, each learned using as input the reconstruction residual from previous dimensions. This approach is intuitively similar to CorrAE and DCCA, but the three modules are each trained separately, so there is no unified objective.

Kim et al. (2012) propose an algorithm that first uses deep belief networks and the autoencoder objective to extract features for two languages independently, and then applies linear CCA to the learned features (activations at the bottleneck layer of the autoencoders) to learn the final representation. In this two-step approach, the DNN weight parameters are not updated to optimize the CCA objective.

There has also been work on multi-view feature learning using deep Boltzmann machines (Srivastava and Salakhutdinov, 2014; Sohn et al., 2014). The models in this work stack several layers of restricted Boltzmann machines (RBM) to represent each view, with an additional top layer that provides the joint representation. These are probabilistic graphical models, for which the maximum likelihood objective is intractable and the training procedures are more complex. Although probabilistic models have some advantages (e.g., dealing with missing values and generating samples in a natural way), DNN-based models have the advantages of a tractable objective and efficient training.

2 Kernel CCA

Therefore we can conveniently work with the kernel (Gram) matrices instead of possibly infinite dimensional RKHS space and optimize directly over the coefficients. Following a derivation similar to that of CCA, one can show that the optimal solution (A∗,B∗)(\mathbf{A}^{*},\mathbf{B}^{*}) satisfies

where Σ\boldsymbol{\Sigma} is a diagonal matrix containing the leading correlation coefficients. Thus the optimal projection can be obtained by solving an eigenvalue problem of size N×NN\times N.

We can make a few observations on the KCCA method. First, the non-zero regularization parameters (rx,ry)>0(r_{x},r_{y})>0 are needed to avoid trivial solutions and correlations. Second, in KCCA, the mappings are not optimized over except that the kernel parameters are usually cross-validated. Third, exact KCCA is computationally challenging for large data sets as it would require performing an eigendecomposition of an N×NN\times N matrix which is expensive both in memory (storing the kernel matrices) and time (solving the N×NN\times N eigenvalue systems naively costs O(N3)\mathcal{O}(N^{3})).

Both techniques produce rank-MM approximations of the kernel matrices with computational complexity O(M3+M2N)\mathcal{O}(M^{3}+M^{2}N); but random Fourier features are data independent and more efficient to generate while Nyström tends to work better. Other approximation techniques such as incomplete Cholesky decomposition (Bach and Jordan, 2002), partial Gram-Schmidt (Hardoon et al., 2004), incremental SVD (Arora and Livescu, 2012) have also been proposed and applied to KCCA. However, for very large training sets, such as the ones in some of our tasks below, it remains difficult and costly to approximate KCCA well. Although recently iterative algorithms have been introduced for very large CCA problems (Lu and Foster, 2014), they are aimed at sparse matrices and do not have a natural out-of-sample extension.

3 Other related models

CCA is related to metric learning in a broad sense. In metric learning, the task is to learn a metric in the input space (or equivalently a projection mapping) such that the learned distances (or equivalently Euclidean distances in the projected space) between “similar” samples are small while distances between “dissimilar” samples are large. Metric learning often uses side information in the form of pairs of similar/dissimilar samples, which may be known a priori or derived from class labels (Xing et al., 2003; Shental et al., 2002; Bar-Hillel et al., 2005; Tsang and Kwok, 2003; Schultz and Joachims, 2004; Goldberger et al., 2005; Hoi et al., 2006; Globerson and Roweis, 2006; Davis et al., 2007; Weinberger and Saul, 2009).

Note that we can equivalently write the CCA objective as (by replacing max⁡\max with min⁡−\min- and adding 1/21/2 times the left-hand side of the whitening constraints)

In view of the above formulation, pairs of co-occurring two-view samples are mapped into similar locations in the projection, which are thus considered “similar” in CCA, and the whitening constraints set the scales of the projections so that the dataset does not collapse into a constant and so that the projection dimensions are uncorrelated. Unlike the typical metric learning setting, the two-view data in CCA may come from different domains/modalities and thus each view has its own projection mapping. Also, CCA uses no information regarding “dissimilar” pairs of two-view data. In this sense the CCA setting is more similar to that of Shental et al. (2002) and Bar-Hillel et al. (2005), which use only side information regarding groups of similar samples (“chunklets”) for single-view data.

For multi-view data, Globerson et al. (2007) propose an algorithm for learning Euclidean embeddings by defining a joint or conditional distribution of the views based on Euclidean distance in the embedding space and maximizing the data likelihood. The objective they minimize is the weighted mean of squared distances between embeddings of co-occurring pairs with a regularization term depending on the partition function of the defined distribution. This model differs from CCA in the global constraints/regularization used.

Recently, there has been increasing interest in learning (multi-view) representations using contrastive losses which aim to enforce that distances between dissimilar pairs are larger than distances between similar pairs by some margin (Hermann and Blunsom, 2014; Huang et al., 2013). In case there is no ground-truth information regarding similar/dissimilar pairs (as in our setting), random sampling is typically used to generate negative pairs. The sampling of negative pairs can be harmful in cases where the probability of mistakenly obtaining similar pairs from the sampling procedure is relatively high. In future work it would be interesting to compare contrastive losses with the CCA objective when only similar pairs are given, and to consider the effects of such sampling in a variety of data distributions.

Finally, CCA has a connection with the information bottleneck method (Tishby et al., 1999). Indeed, in the case of Gaussian variables, the information bottleneck method finds the same subspace as CCA (Chechik et al., 2005).

Experiments

We first demonstrate the proposed algorithms and related work on several multi-view feature learning tasks (Sections 4.1–4.3). In our setting, the second view is not available during test time, so we try to learn a feature transformation of the first/primary view that captures useful information from the second view using a paired two-view training set. Then, in Section 4.4, we explore the stochastic optimization procedure for the DCCA objective (1).

We focus on several downstream tasks including noisy digit image classification, speech recognition, and word pair semantic similarity. On these tasks, we compare the following methods in the multi-view learning setting:

DNN-based models, including SplitAE, CorrAE, DCCA, DCCAE, and DistAE.

Linear CCA (CCA), corresponding to DCCA with only a linear network with no hidden layers for both views.

Kernel CCA approximations. Exact KCCA is computationally infeasible for our tasks since they are large; we instead implement two kernel approximation techniques, using Gaussian RBF kernels. The first implementation, denoted FKCCA, uses random Fourier features (Lopez-Paz et al., 2014) and the second implementation, denoted NKCCA, uses the Nyström approximation (Williams and Seeger, 2001). As described in Section 3.2, in both FKCCA/NKCCA, we transform the original inputs to an MM-dimensional feature space where the inner products between samples approximate the kernel similarities (Yang et al., 2012). We apply linear CCA to the transformed inputs to obtain the approximate KCCA solution.

In this task, we generate two-view data using the MNIST dataset (LeCun et al., 1998), which consists of 28×2828\times 28 grayscale digit images, with 60K60K/10K10K images for training/testing. We generate a more challenging version of the dataset as follows (see Fig. 2 for examples). We first rescale the pixel values to (bydividingtheoriginalvaluesin(by dividing the original values in by 255255). We then randomly rotate the images at angles uniformly sampled from [−π/4,π/4][-\pi/4,\pi/4] and the resulting images are used as view 1 inputs. For each view 1 image, we randomly select an image of the same identity (0-9) from the original dataset, add independent random noise uniformly sampled from toeachpixel,andtruncatethepixelfinalvaluestoto each pixel, and truncate the pixel final values to to obtain the corresponding view 2 sample. The original training set is further split into training/tuning sets of size 50K50K/10K10K.

where sis_{i} is the ground truth label of sample xi\mathbf{x}_{i}, rir_{i} is the cluster label of xi\mathbf{x}_{i}, and map(ri)\text{map}(r_{i}) is an optimal permutation mapping between cluster labels and ground truth labels obtained by solving a linear assignment problem using the Hungarian algorithm (Munkres, 1957). The NMI considers the probability distribution over the ground truth label set CC and cluster label set C′C^{\prime} jointly, and is defined by the following set of equations

where p(ci)p(c_{i}) is interpreted as the probability of a sample having label cic_{i} and p(ci,cj′)p(c_{i},c^{\prime}_{j}) the probability of a sample having label cic_{i} while being assigned to cluster cj′c^{\prime}_{j} (all of which can be computed by counting the samples in the joint partition of CC and C′C^{\prime}). Larger values of these criteria (with an upper bound of 1) indicate better agreement between the clustering and ground-truth labeling.

Each algorithm has hyperparameters that are selected using the tuning set. The final dimensionality LL is selected from {5,10,20,30,50}\{5,10,20,30,50\}. For CCA, the regularization parameters rxr_{x} and ryr_{y} are selected via grid search. For KCCAs, we fix both rxr_{x} and ryr_{y} at a small positive value of 10−410^{-4} (as suggested by Lopez-Paz et al. (2014), FKCCA is robust to rxr_{x}, ryr_{y}), and do grid search for the Gaussian kernel width over {2,3,4,5,6,8,10,15,20}\{2,3,4,5,6,8,10,15,20\} for view 1 and {2.5,5,7.5,10,15,20,30}\{2.5,5,7.5,10,15,20,30\} for view 2 at rank M=5,000M=5,000, and then test with M=20,000M=20,000. For DNN-based models, the feature mappings (f,g)(\mathbf{f},\mathbf{g}) are implemented by networks of 33 hidden layers, each of 10241024 sigmoid units, and a linear output layer of LL units; reconstruction mappings (p,q)(\mathbf{p},\mathbf{q}) are implemented by networks of 33 hidden layers, each of 10241024 sigmoid units, and an output layer of 784784 sigmoid units. We fix rx=ry=10−4r_{x}=r_{y}=10^{-4} for DCCA and DCCAE. For SplitAE/CorrAE/DCCAE/DistAE we select the trade-off parameter λ\lambda via grid search over {0.001, 0.01, 0.1, 1, 10}, allowing a trade-off between the correlation term (with value in [0,L][0,L]) and the reconstruction term (varying roughly in the range $asthereconstructionerrorforthesecondviewisalwayslarge).Thenetworksas the reconstruction error for the second view is always large). The networks(\mathbf{f},\mathbf{p})arepre−trainedinalayerwisemannerusingrestrictedBoltzmannmachines(HintonandSalakhutdinov,2006)andsimilarlyforare pre-trained in a layerwise manner using restricted Boltzmann machines (Hinton and Salakhutdinov, 2006) and similarly for(\mathbf{g},\mathbf{q})$ with inputs from the corresponding view.

For DNN-based models, we use SGD for optimization with minibatch size, learning rate and momentum tuned on the tuning set (more on this in Section 4.4). A small weight decay parameter of 10−410^{-4} is used for all layers. We monitor the objective on the tuning set for early stopping. For each algorithm, we select the model with the best AC on the tuning set, and report its results on the test set. The AC and NMI results (in percent) for each algorithm are given in Table 1. As a baseline, we also cluster the original 784784-dimensional view 1 images.

Next, we also measure the quality of the projections via classification experiments. If the learned features are clustered well into classes, then one might expect that a simple linear classifier can achieve high accuracy on these projections. We train one-versus-one linear SVMs (Chang and Lin, 2011) on the projected training set (now using the ground truth labels), and test on the projected test set, while using the projected tuning set for selecting the SVM hyperparameter (the penalty parameter for hinge loss). Test error rates on the optimal embedding of each algorithm (with highest AC) are provided in Table 1 (last column). These error rates agree with the clustering results. Multi-view feature learning makes classification much easier on this task: Instead of using a heavily nonlinear classifier on the original inputs, a very simple linear classifier that can be trained efficiently on low-dimensional projections already achieves high accuracy.

All of the multi-view feature learning algorithms achieve some improvement over the baseline. The nonlinear CCA algorithms all perform similarly, and significantly better than SplitAE, CorrAE, and DistAE. We also qualitatively investigate the features by embedding the projected features in 2D using tt-SNE (van der Maaten and Hinton, 2008); the resulting visualizations are given in Figures 3 and 4. Overall, the visual class separation qualitatively agrees with the relative clustering and classification performance in Table 1.

In the embedding of input images (Figure 3 (a)), samples of each digit form an approximately one-dimensional, stripe-shaped manifold, and the degree of freedom along each manifold corresponds roughly to the variation in rotation angle (see Figure 3 (a’)). This degree of freedom does not change the identity of the image, which is common to both views. Projections by SplitAE/CorrAE/DistAE do achieve somewhat better separation for some classes, but the unwanted rotation variation is still prominent in the embeddings. On the other hand, without using any label information and with only paired noisy images, the nonlinear CCA algorithms manage to map digits of the same identity to similar locations while suppressing the rotational variation and separating images of different identities. Linear CCA also approximates the same behavior, but fails to separate the classes, presumably because the input variations are too complex to be captured by only linear mappings. Overall, DCCAE gives the cleanest embedding, with different digits pushed far apart and good separation achieved.

The different behavior of CCA-based methods from SplitAE/CorrAE/DistAE suggests two things. First, when the inputs are noisy, reconstructing the input faithfully may lead to unwanted degrees of freedom in the features (DCCAE tends to select a relatively small trade-off parameter λ=10−3\lambda=10^{-3} or 10−210^{-2}), further supporting that it is not necessary to fully minimize reconstruction error. We show the clustering accuracy of DCCAE at L=10L=10 for different λ\lambda values in Figure 5. Second, the hard CCA constraints, which enforce uncorrelatedness between different feature dimensions, appear essential to the success of CCA-based methods; these constraints are the difference between DCCAE and CorrAE. However, the constraints without the multi-view objective seem to be insufficient. To see this, we also visualize a 1010-dimensional locally linear embedding (LLE, Roweis and Saul, 2000) of the test images in Fig. 3 (b). LLE satisfies the same uncorrelatedness constraints as in CCA-based methods, but without access to the second view, it does not separate the classes as nicely.

2 Acoustic-articulatory data for speech recognition

We next experiment with the Wisconsin X-Ray Microbeam (XRMB) corpus (Westbury, 1994) of simultaneously recorded speech and articulatory measurements from 47 American English speakers. Multi-view feature learning via CCA/KCCA/DCCA has previously been shown to improve phonetic recognition performance when tested on audio alone (Arora and Livescu, 2013; Wang et al., 2015a).

We follow the setup of Arora and Livescu (2013) and use the learned features for speaker-independent phonetic recognition. Similarly to Arora and Livescu (2013), the inputs to multi-view feature learning are acoustic features (39D features consisting of mel frequency cepstral coefficients (MFCCs) and their first and second derivatives) and articulatory features (horizontal/vertical displacement of 8 pellets attached to different parts of the vocal tract) concatenated over a 7-frame window around each frame, giving 273D acoustic inputs and 112D articulatory inputs for each view.

We split the XRMB speakers into disjoint sets of 35/8/2/2 speakers for feature learning/recognizer training/tuning/testing. The 35 speakers for feature learning are fixed; the remaining 12 are used in a 6-fold experiment (recognizer training on 8 speakers, tuning on 2 speakers, and testing on the remaining 2 speakers). Each speaker has roughly 50K50K frames, giving 1.43M multi-view training frames; this is a much larger training set than those used in previous work on this data set. We remove the per-speaker mean and variance of the articulatory measurements for each training speaker. All of the learned feature types are used in a tandem approach (Hermansky et al., 2000), i.e., they are appended to the original 39D features and used in a standard hidden Markov model (HMM)-based recognizer with Gaussian mixture model observation distributions. The baseline is the recognition performance using the original MFCC features. The recognizer has one 3-state left-to-right HMM per phone, using the same language model as in Arora and Livescu (2013).

For each fold, we select the best hyperparameters based on recognition accuracy on the tuning speakers, and use the corresponding learned model for the test speakers. As before, models based on neural networks are trained via SGD with the optimization parameters tuned by grid search. Here we do not use pre-training for weight initialization. A small weight decay parameter of 5×10−45\times 10^{-4} is used for all layers. For each algorithm, the feature dimensionality LL is tuned over {30,50,70}\{30,50,70\}. For DNN-based models, we use hidden layers of 15001500 ReLUs. For DCCA, we vary the network depths (up to 3 nonlinear hidden layers) of each view. In the best DCCA architecture, f\mathbf{f} has 33 ReLU layers of 15001500 units followed by a linear output layer while g\mathbf{g} has only a linear output layer.

For CorrAE/DistAE/DCCAE, we use the same architecture of DCCA for the encoders, and we set the decoders to have symmetric architectures to the encoders. For this domain, we find that the best choice of architecture for the encoders/decoders for View 2 is linear while for View 1 it is typically three layers deep. For SplitAE, the encoder f\mathbf{f} is similarly deep and the View 1 decoder p\mathbf{p} has the symmetric architecture, while its View 2 decoder q\mathbf{q} was set to linear to match the best choice for the other methods. We fix (rx,ry)(r_{x},r_{y}) to small values as before. The trade-off parameter λ\lambda is tuned for each algorithm by grid search.

For FKCCA, we find it important to use a large number of random features MM to get a competitive result, consistent with the findings of Huang et al. (2014) when using random Fourier features for speech data. We tune kernel widths at M=5, ⁣000M=5,\!000 with FKCCA, and test FKCCA with M=30, ⁣000M=30,\!000 (the largest MM we could afford to obtain an exact SVD solution on a workstation with 32G main memory); We are not able to obtain results for NKCCA with M=30, ⁣000M=30,\!000 in 48 hours with our implementation, so we report its test performance at M=20, ⁣000M=20,\!000 with the optimal FKCCA hyperparameters. Notice that FKCCA has about 14.614.6 million parameters (random Gaussian samples + projection matrices from random Fourier features to the LL-dimensional KCCA features, which is more than the number of weight parameters in the largest DCCA model, so it is slower than DCCA for testing (the cost of computing test features is linear in the number of parameters for both KCCA and DNNs).

Phone error rates (PERs) obtained by different feature learning algorithms are given in Table 2. We see the same pattern as on MNIST: Nonlinear CCA-based algorithms outperform SplitAE/CorrAE/DistAE. Since the recognizer now is a nonlinear mapping (HMM), the performance of the linear CCA features is highly competitive. Again, DCCAE tends to select a relatively small λ\lambda, indicating that the canonical correlation term is more important.

3 Multilingual data for word embeddings

In this task, we learn a vectorial representation of English words from pairs of English-German word embeddings for improved semantic similarity. We follow the setup of Faruqui and Dyer (2014) and use as inputs 640640-dimensional monolingual word vectors trained via latent semantic analysis on the WMT 2011 monolingual news corpora and use the same 36K36K English-German word pairs for multi-view learning. The learned mappings are applied to the original English word embeddings (180K180K words) and the projections are used for evaluation.

We evaluate learned features on two groups of tasks. The first group consists of the four word similarity tasks from Faruqui and Dyer (2014): WS-353 and the two splits WS-SIM and WS-REL, RG-65, MC-30, and MTurk-287. The second group of tasks uses the adjective-noun (AN) and verb-object (VN) subsets from the bigram similarity dataset of Mitchell and Lapata (2010), and tuning and test splits (of size 649/1972) for each subset (we exclude the noun-noun subset as we find that the NN human annotations often reflect “topical” rather than “functional” similarity). We simply add the projections of the two words in each bigram to obtain an LL-dimensional representation of the bigram, as done in prior work (Blacoe and Lapata, 2012). We compute the cosine similarity between the two vectors of each bigram pair, order the pairs by similarity, and report the Spearman’s correlation (ρ\rho) between the model’s ranking and human rankings.

We tune the feature dimensionality LL over {128,384}\{128,384\}; other hyperparameters are tuned as in previous experiments. DNN-based models use ReLU hidden layers of width 1, ⁣2801,\!280. A small weight decay parameter of 10−410^{-4} is used for all layers. We use two ReLU hidden layers for encoders (f\mathbf{f} and g\mathbf{g}), and try both linear and nonlinear networks with two hidden layers for decoders (p\mathbf{p} and q\mathbf{q}). FKCCA/NKCCA are tested with M=20, ⁣000M=20,\!000 using kernel widths tuned at M=4, ⁣000M=4,\!000. We fix rx=ry=10−4r_{x}=r_{y}=10^{-4} for nonlinear CCAs and tune them over {10−6,10−4,10−2,1,102}\{10^{-6},10^{-4},10^{-2},1,10^{2}\} for CCA.

For word similarity tasks, we simply report the highest Spearman’s correlation obtained by each algorithm. For bigram similarity tasks, we select for each algorithm the model with the highest Spearman’s correlation on the 649 tuning bigram pairs, and we report its performance on the 1972 test pairs. The results are given in Table 3. Unlike MNIST and XRMB, it is important for the features to reconstruct the input monolingual word embeddings well, as can be seen from the superior performance of SplitAE over FKCCA/NKCCA/DCCA. This implies there is useful information in the original inputs that is not correlated across views. However, DCCAE still performs the best on the AN task, in this case using a relatively large λ=0.1\lambda=0.1.

4 Empirical analysis of DCCA optimization

We now explore the issue of stochastic optimization for DCCA as discussed in Section 2.2. We use the same XRMB dataset as in the acoustic-articulatory experiment of Section 4.2. In this experiment, we select utterances from a single speaker ‘JW11’ and divide them into training/tuning/test splits of roughly 30K30K/11K11K/9K9K pairs of acoustic and articulatory frames.Our split of the data is the same as the one used by Andrew et al. (2013). We note that Lopez-Paz et al. (2014) used the same speaker in their experiments, but they randomly shuffled all 50K50K frames before creating the splits. We suspect that DCCA (as well as their own algorithm) were under-tuned in their experiments. We ran experiments on a randomly shuffled dataset with careful tuning of kernel widths for FKCCA/NKCCA (using rank M=6000M=6000) and obtained canonical correlations of 99.2/105.6/107.6 for FKCCA/NKCCA/DCCA, which are better than those reported in Lopez-Paz et al. (2014). Since the sole purpose of this experiment is to study the optimization of nonlinear CCA algorithms rather than the usefulness of the features, we do not carry out any down-stream tasks or careful model selection for that purpose (e.g., a search for feature dimensionality LL).

We now consider the effect of our stochastic optimization procedure for DCCA, denoted STO below, and demonstrate the importance of minibatch size. We use a 3-layer architecture where the acoustic and articulatory networks have two hidden layers of 1800 and 1200 rectified linear units (ReLUs) respectively, and the output dimensionality (and therefore the maximum possible total canonical correlation over dimensions) is L=112L=112. We use a small weight decay γ=10−4\gamma=10^{-4}, and do grid search for several hyperparameters: rx,ry∈{10−4, 10−2, 1, 102}r_{x},r_{y}\in\{10^{-4},\ 10^{-2},\ 1,\ 10^{2}\}, constant learning rate in {10−4, 10−3, 10−2, 10−1}\{10^{-4},\ 10^{-3},\ 10^{-2},\ 10^{-1}\}, fixed momentum in {0, 0.5, 0.9, 0.95, 0.99}\{0,\ 0.5,\ 0.9,\ 0.95,\ 0.99\}, and minibatch size in {100, 200, 300, 400, 500, 750, 1000}\{100,\ 200,\ 300,\ 400,\ 500,\ 750,\ 1000\}. After learning the projection mappings on the training set, we apply them to the tuning/test set to obtain projections, and measure the canonical correlation between views. Figure 6 shows the learning curves on the tuning set for different minibatch sizes, each using the optimal values for the other hyperparameters. It is clear that for small minibatches (100100, 200200), the objective quickly plateaus at a low value, whereas for large enough minibatch size, there is always a steep increase at the beginning, which is a known advantage of stochastic first-order algorithms (Bottou and Bousquet, 2008), and a wide range of learning rate/momentum give very similar results. The reason for such behavior is that the stochastic estimate of the DCCA objective becomes more accurate as minibatch size nn increases. We provide theoretical analysis of the error between the true objective and its stochastic estimate in Appendix A.

Recently, Wang et al. (2015c) have proposed a nonlinear orthogonal iterations (NOI) algorithm for the DCCA objective which extends the alternating least squares procedure (Golub and Zha, 1995; Lu and Foster, 2014) and its stochastic version (Ma et al., 2015) for CCA. Each iteration of the NOI algorithm adaptively estimates the covariance matrices of the projections of each view, whitens the projections of a minibatch using the estimated covariance matrices, and takes a gradient step over DNN weight parameters of the nonlinear least squares problems of regressing each view’s input against the whitened projection of the other view for the minibatch. The advantage of NOI is that it performs well with smaller minibatch sizes and thus reduces memory consumption. Wang et al. (2015c) have shown that NOI can achieve the same objective value as STO using smaller minibatches. For the problems considered in this paper, however, each epoch of NOI takes a longer time than that of STO, due to the whitening operations at each iteration (involving eigenvalue decomposition of L×LL\times L covariance matrices). The smaller the minibatch size we use in NOI, the more frequently we run such operations. We report the learning curve of NOI with minibatch size n=100n=100 while all other hyper-parameters are tuned similarly to STO. NOI can eventually reach the same objective as STO given more training time. Ultimately, the choice of optimization algorithm will depend on the task-specific data and time/memory constraints.

Finally, we also train the same model using batch training with L-BFGSWe use the L-BFGS implementation of Schmidt (2012), which includes a good line-search procedure. with the same random initial weight parameters and tune (rx,ry)(r_{x},r_{y}) on the same grid. While L-BFGS does well on the training set, its performance on tuning/test is usually worse than that of stochastic optimization with reasonable hyperparameters.

We also train FKCCA and NKCCA on this dataset. We tune (rx,ry)(r_{x},r_{y}) on the same grid as for DCCA, tune the kernel width for each view, and vary the approximation rank MM from 10001000 to 60006000 for each method. We plot the best total canonical correlation achieved by each algorithm on the tuning set as a function of MM in Figure 7. Clearly both algorithms require relatively large MM to perform well on this task, but the return in further increasing MM is diminishing. NKCCA achieves better results than FKCCA, although forming the low-rank approximation also becomes more costly as MM increases. We plot the total canonical correlation achieved by each algorithm at different running times with various hyperparameters in Figure 8.

Finally, we select for each algorithm the best model on the tuning set and give the corresponding total canonical correlation on the test set in Table 4.

Conclusion

We have explored several approaches in the space of DNN-based multi-view representation learning. We have found that on several tasks, CCA-based models outperform autoencoder-based models (SplitAE) and models based on between-view squared distance (DistAE) or correlation (CorrAE) instead of canonical correlation. The best overall performer is a new DCCA extension, deep canonically correlated autoencoders (DCCAE). We have studied these objectives in the context of DNNs, but we expect that the same trends should apply also to other network architectures such as convolutional (LeCun et al., 1998) and recurrent (Elman, 1990; Hochreiter and Schmidhuber, 1997) networks, and this is one direction for future work.

In light of the empirical results, it is interesting to consider again the main features of each type of objective and corresponding constraints. Autoencoder-based approaches are based on the idea that the learned features should be able to accurately reconstruct the inputs (in the case of multi-view learning, the inputs in both views). The CCA objective, on the other hand, focuses on how well each view’s representation predicts the other’s, ignoring the ability to reconstruct each view. CCA is expected to perform well for clustering and classification when the two views are uncorrelated given the class label (Chaudhuri et al., 2009). The noisy MNIST dataset used here simulates exactly this scenario, and indeed this is the task where deep CCA outperforms other objectives by the largest margins. Even in the other tasks, however, there usually seems to be only a small advantage to being able to reconstruct the inputs faithfully.

The constraints in the various methods also have an important effect. The performance difference between DCCA and CorrAE demonstrates that uncorrelatedness between learned dimensions is important. On the other hand, the stronger DCCA constraint may still not be sufficiently strong; an even better constraint may be to require the learned dimensions to be independent (or approximately so), and this is an interesting avenue for future work.

We believe the applicability of the DCCA objective goes beyond the unsupervised feature learning setting. For example, it can be used as a data-dependent regularizer in supervised or semi-supervised learning settings where we have some labeled data as well as multi-view observations. The usefulness of CCA in such settings has previously been analyzed theoretically (Kakade and Foster, 2007; McWilliams et al., 2013) and has begun to be explored experimentally (Arora and Livescu, 2014).

A Analysis of stochastic optimization for DCCA

We denote by Θ(n)\boldsymbol{\Theta}^{(n)} the “partial” objective based on the nn samples, i.e.

And recall that our true objective function is

where Σxx=1n∑i=1Nfifi⊤+rxI{\boldsymbol{\Sigma}}_{xx}=\frac{1}{n}\sum_{i=1}^{N}\mathbf{f}_{i}\mathbf{f}_{i}^{\top}+r_{x}\mathbf{I}, Σyy=1n∑i=1Ngigi⊤+ryI{\boldsymbol{\Sigma}}_{yy}=\frac{1}{n}\sum_{i=1}^{N}\mathbf{g}_{i}\mathbf{g}_{i}^{\top}+r_{y}\mathbf{I}, and Σxy=1n∑i=1Nfigi⊤{\boldsymbol{\Sigma}}_{xy}=\frac{1}{n}\sum_{i=1}^{N}\mathbf{f}_{i}\mathbf{g}_{i}^{\top} are computed over the entire training set.

A1: The samples used for estimating Σxx{\boldsymbol{\Sigma}}_{xx}, Σyy{\boldsymbol{\Sigma}}_{yy} and Σxy{\boldsymbol{\Sigma}}_{xy} are chosen independently from each other, each by sampling examples (for Σxx{\boldsymbol{\Sigma}}_{xx} and Σyy{\boldsymbol{\Sigma}}_{yy}) or example pairs (for Σxy{\boldsymbol{\Sigma}}_{xy}) uniformly at random with replacement from the training set. In practice, we could randomly pick nn samples of f(x)\mathbf{f}(\mathbf{x}) for computing Σ^xx\hat{\boldsymbol{\Sigma}}_{xx}, another nn samples of g(y)\mathbf{g}(\mathbf{y}) for computing Σ^yy\hat{\boldsymbol{\Sigma}}_{yy}, and nn pairs of samples of (f(x),g(y))(\mathbf{f}(\mathbf{x}),\mathbf{g}(\mathbf{y})) for computing Σ^xy\hat{\boldsymbol{\Sigma}}_{xy}, and the computational cost of this modified procedure is twice the cost of the original naive implementation. This simple modification allows the expectation of matrix multiplications to factorize.

A2: Additionally, we assume an upper bound on the magnitude of the neural network outputs, namely, max⁡(∥fi∥2,∥gi∥2)≤B\max(\left\lVert\mathbf{f}_{i}\right\rVert^{2},\left\lVert\mathbf{g}_{i}\right\rVert^{2})\leq B, ∀i\forall i. This holds when nonlinear activations with a bounded range are used; e.g., with logistic sigmoid or hyperbolic tangent activations, the upper bound BB can be set to the output dimensionality (because the squared activation is bounded by 11 for each output unit), or when the inputs themselves are bounded and the functions f(x)\mathbf{f}(\mathbf{x}) and g(y)\mathbf{g}(\mathbf{y}) are Lipschitz.

Our analysis is based on the following formulation of the linear CCA solution (Borga, 2001). Notice that the optimal (U,V)(\mathbf{U},\mathbf{V}) for the CCA objective (1) satisfies

where Σ{\boldsymbol{\Sigma}} contains the top LL singular values of T{\mathbf{T}} on the diagonal. This is easy to verify as Σxx1/2U{\boldsymbol{\Sigma}}_{xx}^{1/2}\mathbf{U} and Σyy1/2V{\boldsymbol{\Sigma}}_{yy}^{1/2}\mathbf{V} contain the top left/right singular vectors of T{\mathbf{T}} (see, e.g., Section 2 of Andrew et al., 2013).

Multiplying both sides of (20) by \left[\begin{array}[]{cc}{\boldsymbol{\Sigma}}_{xx}^{-1/2}&\mathbf{0}\\ \mathbf{0}&{\boldsymbol{\Sigma}}_{yy}^{-1/2}\end{array}\right] gives

which implies \left[\begin{array}[]{c}\mathbf{U}\\ \mathbf{V}\end{array}\right] correspond to the top LL eigenvectors of the matrix A−1B{\mathbf{A}}^{-1}{\mathbf{B}}, where

(It is straightforward to check that eigenvalues of A−1B{\mathbf{A}}^{-1}{\mathbf{B}} are ±σ1(T),±σ2(T),… .\pm\sigma_{1}({\mathbf{T}}),\pm\sigma_{2}({\mathbf{T}}),\dots.)

When using minibatches, the estimate of the objective is of course computed based on the randomly chosen subset of samples, and is the sum of the top eigenvalues of the matrix A^−1B^\hat{\mathbf{A}}^{-1}\hat{\mathbf{B}}, where

We now try to bound the expected error between A^−1B^\hat{\mathbf{A}}^{-1}\hat{\mathbf{B}} and A−1B{\mathbf{A}}^{-1}{\mathbf{B}} measured in spectral norm. An important tool in our analysis is the matrix Bernstein inequality below.

Let A1,…,An\mathbf{A}_{1},\dots,\mathbf{A}_{n} be independent random matrices with common dimension d1×d2d_{1}\times d_{2}. Assume that each matrix has bounded deviation from its mean:

Form the sum B=∑k=1nAk\mathbf{B}=\sum_{k=1}^{n}\mathbf{A}_{k}, and introduce a variance parameter

The following result and its proof are similar to that of Lopez-Paz et al. (2014, Theorem. 4).

Assume that A1 and A2 hold for our stochastic estimate of the true DCCA objective, and assume that the eigenvalues of autocovariance matrices Σ^xx\hat{\boldsymbol{\Sigma}}_{xx} and Σ^yy\hat{\boldsymbol{\Sigma}}_{yy} are lower bounded by γx>0\gamma_{x}>0 and γy>0\gamma_{y}>0 respectively, i.e.,

Then we have the following bound for the expected error:

where the expectation is taken over random selection of training samples as described in A1, and

Proof Since A^−1B^−A−1B\hat{\mathbf{A}}^{-1}\hat{\mathbf{B}}-{\mathbf{A}}^{-1}{\mathbf{B}} is block diagonal, its spectral norm is bounded by the maximum of the spectral norms of the individual blocks:

We now focus on analyzing the first term in the above maximum as the other term follows analogously. Define the individual error terms

Thus our goal is to bound ∥Σ^xx−1Σ^xy−Σxx−1Σxy∥=∥E∥\left\lVert\hat{\boldsymbol{\Sigma}}_{xx}^{-1}\hat{\boldsymbol{\Sigma}}_{xy}-{\boldsymbol{\Sigma}}_{xx}^{-1}{\boldsymbol{\Sigma}}_{xy}\right\rVert=\left\lVert\mathbf{E}\right\rVert.

Notice that the matrices Σxx{\boldsymbol{\Sigma}}_{xx} and Σxy{\boldsymbol{\Sigma}}_{xy} are not random. And due to assumption A1, the samples used for estimating Σ^xx\hat{\boldsymbol{\Sigma}}_{xx} and Σ^xy\hat{\boldsymbol{\Sigma}}_{xy} are selected independently, so the expectation of Σ^xx−1figi⊤\hat{\boldsymbol{\Sigma}}_{xx}^{-1}\mathbf{f}_{i}\mathbf{g}_{i}^{\top} factorizes. Therefore we have

and the deviation of the individual error matrices from their expectation is

where we have used the triangle inequality in the first inequality, and Jensen’s inequality for the second and third inequality (since norms are convex functions). To apply the Matrix Bernstein inequality, we still need to bound the variance which is defined as

where we have used the fact that the Zi\mathbf{Z}_{i}’s are independent and have mean zero in the second equality.

Let us consider an individual term in the summand of the second term:

where Jensen’s inequality is used in the second inequality. A similar argument shows that

An invocation of the triangle inequality on the definition of σ2\sigma^{2} along with the above two bounds gives

We may now appeal to the matrix Bernstein inequality on the dx×dyd_{x}\times d_{y} matrices {Zi}i=1n\{\mathbf{Z}_{i}\}_{i=1}^{n} to obtain the bound

Using the same matrix Bernstein technique (now applied to ∑i=1nZi′\sum_{i=1}^{n}\mathbf{Z}_{i}^{\prime} with Zi′=1n(fifi⊤−Σxx)\mathbf{Z}_{i}^{\prime}=\frac{1}{n}\left(\mathbf{f}_{i}\mathbf{f}_{i}^{\top}-{\boldsymbol{\Sigma}}_{xx}\right)), we have

We are now ready to bound the quantity of interest as

Assuming that dx=dy=dd_{x}=d_{y}=d in our algorithm and γ=min⁡(γx,γy)\gamma=\min(\gamma_{x},\gamma_{y}), then we have the dependency of the error in spectral norm as

As expected, the error decreases as we use large minibatch size nn. Also, we observe the error decreases as (γx,γy)(\gamma_{x},\gamma_{y}) increase. Notice that from the definition of Σ^xx\hat{\boldsymbol{\Sigma}}_{xx} and Σ^yy\hat{\boldsymbol{\Sigma}}_{yy} we are guaranteed that γx≥rx\gamma_{x}\geq r_{x} and γy≥ry\gamma_{y}\geq r_{y}. This means we can increase the regularization constants to improve the estimation error, and this is to be expected as larger (rx,ry)(r_{x},r_{y}) render the auto-covariance matrices less relevant for estimating T\mathbf{T}. In practice, as long as one uses a minibatch size n>dn>d, the auto-covariance matrices are typically non-singular and (rx,ry)(r_{x},r_{y}) are underestimate of (γx,γy)(\gamma_{x},\gamma_{y}).

Acknowledgment

The authors would like to thank Louis Goldstein for providing phonetic alignments for the data used in the recognition experiment; Manaal Faruqui, Chris Dyer, Ang Lu, Mohit Bansal, and Kevin Gimpel for sharing resources for the multi-lingual embedding experiments; Nati Srebro for input on the stochastic optimization of DCCA; and Geoff Hinton for helpful discussion on the CCA objective.

References