Learning phase transitions by confusion

Evert P. L. van Nieuwenburg, Ye-Hua Liu, Sebastian D. Huber

I Introduction

Machine learning as a tool for analyzing data is becoming more and more prevalent in an increasing number of fields Jordan and Mitchell 2015. This is due to a combination of availability of large amounts of data, and the advances in hardware and computational power (most notably through the use of GPUs).

Two typical methods of machine learning can be distinguished, namely the unsupervised and supervised methods. In the former, the algorithm receives no input other than the data and is asked, e.g., to extract features or to cluster the samples. Such an unsupervised approach was applied to the physical problem of identifying phase transitions and order parameters, from images of classical configurations of Ising models Wang 2016. In the supervised learning methods, the data has to be supplemented by a labelling. Typical problems in this direction involve the classification of many samples, where each sample is assigned a class-label. The machine is trained to recognize samples and predict their associated label, demonstrating that it has learned by doing so via correctly predicting samples it has not encountered before. This approach, too, has been demonstrated on Ising models Carrasquilla and Melko 2016. Such approaches have also recently provided promising prospects for strongly correlated fermions Ch’ng et al. 2016; Li et al. 2016 and the fermion sign problem Broecker et al. 2016. Last, we mention a slightly orthogonal approach based on reinforcement learning, where the wavefunction was represented as a particular type of artificial neural network Carleo and Troyer 2016. A variational approach then allows for finding groundstates and performing time-evolution, and outperformed other state-of-the-art numerical methods for the two-dimensional Heisenberg spin model. Some topological states have been shown to have efficient and exact representations as an artificial neural network Deng et al. 2016.

Conversely, concepts from physics have also found their way into the field of machine learning. Examples of this are e.g. the relations between neural networks and statistical Ising models and renormalisation flow Mehta and Schwab 2014, the use of tensor network techniques to train them Stoudenmire and Schwab 2016, and indeed the very concept of phase transitions themselves Saitta and Sebag 2010.

Motivated by previous studies, we apply machine-learning techniques to the detection of phase transitions. In contrast to the previous works, however, we focus on a combination of supervised and unsupervised techniques. In most cases namely, it is exactly the labelling that one would like to find out (i.e. classification of phases). That implies that a labelling is not known beforehand, and hence supervised techniques are not directly applicable. In this Letter we demonstrate that it is possible to find the correct labels, by purposefully mislabelling the data and evaluating the performance of the machine learner. We will base our method on neural networks, which are capable of fitting arbitrary non-linear functions Haykin 1998. Indeed, if a linear feature extraction method worked, there would have been no need to explicitly find labels in the first place.

For quantum phase transitions, one tries to learn the quantum-mechanical wavefunction ∣ψ⟩|\psi\rangle, which contains exponentially many coefficients with increasing system size. As has been noted before Carrasquilla and Melko 2016, a similar problem exists in the field of machine learning: the number of samples in a dataset has to increase exponentially with the number of features one is trying to extract. To prevent having to deal with exponentially large wavefunctions, we pre-process the data in the form of the entanglement spectrum (ES) Li and Haldane 2008, which has been shown to contain important information about ∣ψ⟩|\psi\rangle Laflorencie 2015, although care has to be taken when interpreting phase transitions Chandran et al. 2014.

To justify the use of the ES, we note that recently the quantum entanglement has taken up a major role in the characterization of many-body quantum systems Amico et al. 2008; Laflorencie 2015. In particular, the ES has been used as an important tool in, for example, fingerprinting topological order Thomale et al. 2010a; Qi et al. 2012; Turner et al. 2011, tensor network properties Cirac et al. 2011; Schuch et al. 2013, quantum critical points, symmetry breaking phases Calabrese and Lefevre 2008; Alba et al. 2012, and even many-body localization Yang et al. 2015; Geraedts et al. 2016. Very recently, an experimental protocol for measuring the ES has been proposed Pichler et al. 2016. On the level of the ES, the information of phases is not clearly identifiable as in the classical images, which we will show in the following sections. However, patterns in the ES suggest that learning and generalization is still possible.

In the following, we will first consider the Kitaev chain as a demonstration of our method. The Kitaev chain serves as an excellent example since analytical results are available, and the ES shows a clear distinction between the two phases of the model Kitaev 2001. We demonstrate the generalizing power of the neural network by blanking out the training data around the transition, and show that it can still predict the transition accurately. We then purposefully mislabel the data, thereby confusing the network, and introduce the characteristic shape of the networks’ performance function. Next, this confusion method is demonstrated on a classical two-dimensional Ising model to show that it is not specific to the Kitaev model. Last, we apply the proposed method to the nontrivial task of finding the many-body-localization (MBL) transition from many disordered samples of a one-dimensional spin chain. We finish with a discussion and outlook.

II Results

We demonstrate the various machine-learning methods on the model of the Kitaev chain:

where t>0t>0 controls the hopping and the pairing of spinless fermions alike and μ\mu is a chemical potential. The groundstate of this model has a quantum phase transition from a topologically trivial (∣μ∣>2t|\mu|>2t) to a nontrivial state (∣μ∣<2t|\mu|<2t) as the chemical potential μ\mu is tuned across μ=±2t\mu=\pm 2t.

We use the ES to compress the quantum mechanical wavefunction. The ES is defined as follows. The whole system is first divided into two subsets AA and BB, after which the reduced density matrix of subset AA is calculated by partially tracing out the degrees of freedom in BB, i.e. ρA=TrB∣ψ⟩⟨ψ∣\rho_{A}=\text{Tr}_{B}|\psi\rangle\langle\psi|. Denoting the eigenvalues of ρA\rho_{A} as λi\lambda_{i}, the entanglement spectrum is then defined as the set of numbers −ln⁡λi-\ln\lambda_{i}. It is important to remark that various types of bipartition of the whole system into subsets AA and BB exist, such as dividing the bulk in extensive disconnected parts Hsieh and Fu 2014, divisions in momentum space Thomale et al. 2010b or indeed even random partitioning Vijay and Fu 2015. In this work, we use the usual spatial bipartition into left and right halves of the whole system.

As shown in Fig. 1a, the entanglement spectrum of the Kitaev chain is clearly distinguishable in the two phases, especially since the nontrivial phase has a degeneracy structure as do all symmetry protected topological phases Turner et al. 2011. This feature is clear also for human eyes, and a machine-learning routine seems to be an overkill. We use this model for demonstration purposes and in the following, we will apply the introduced methodology to more complex models. The data for machine learning is chosen to be the largest 10 eigenvalues λi\lambda_{i}, for L=20L=20 with an equal partitioning LA=LB=10L_{A}=L_{B}=10, and for various values of −4t≤μ≤0-4t\leq\mu\leq 0.

First we perform unsupervised learning, using an established method for feature extraction. The entanglement spectra are interpreted as points in a 1010-dimensional space, and we use principal component analysis (PCA) Pearson 1901 to extract mutually orthogonal axes along which most of the variance of the data can be observed. PCA amounts to a linear transformation Y=XWY=XW, where XX is an N×10N\times 10 matrix containing the entanglement spectra as rows (N=104N=10^{4} is the number of samples).

II.2 Supervised learning

Although it is as a whole unnecessary, since the PCA analysis already manages to extract the phases and the transition point, we still train a feedforward neural network (NN) on the 1010-dimensional inputs. This is for demonstration only, since we have in mind the application to models in which PCA is insufficient.

We train the network with 80 hidden sigmoid neurons in a single hidden layer, and 2 output neurons. The first/second output neuron predicts the (not necessarily normalized) probability for the data to be in trivial/nontrivial phase, and the predicted phase is the phase with the larger probability. We use stochastic gradient descent and L2 regularization to try and minimize a cross-entropy cost function (for details, see e.g. Ref. Nielsen 2015). The network easily learns to distinguish the spectra when trained on a subset of the data points and asked to predict the others.

Arguably the most important objective of machine-learning in general is that of generalization. After all, learning is demonstrated by being able to perform well on examples that have not been encountered before. By having trained the network only on a subset of the data, and having it correctly predict others, it has already demonstrated learning.

As another display of the generalizing power of the network, we blank out the data in a width ww around μ=−2t\mu=-2t and ask the network to interpolate and find the transition point. Figure 1c shows that the network has no difficulties doing so even for w=2tw=2t. We were able to go up to widths w=3tw=3t before training became unreliable.

II.3 Confusion scheme

The PCA as an unsupervised learning technique may be applied without perfectly known information of the system, but it is a linear analysis and is hence incapable of extracting non-linear relationships among the data. On the other hand, a NN is capably of fitting any non-linear function Haykin 1998, but a training phase with correctly labelled input-output pairs is needed. In the following, we propose a scheme combining the two methods which we refer to as a confusion scheme. This scheme is the main result of this work.

Suppose all data lies in the parameter range (a,b)(a,b), and we know there exists a critical point a<c<ba<c<b such that the data could be classified into two groups. However, we do not know the value of cc. We propose a critical point c′c^{\prime}, and train a network that we call Nc′\mathcal{N}_{c^{\prime}} by labelling all data with parameters smaller than c′c^{\prime} with label 00 and the others with label 11. Next, we evaluate the performance of Nc′\mathcal{N}_{c^{\prime}} on the entire data set and refer to its total performance, with respect to the proposed critical point c′c^{\prime}, as P(c′)P(c^{\prime}). We will show that the function P(c′)P(c^{\prime}) has a universal W-shape, with the middle peak at the correct critical point cc. Applying this to the Kitaev model, we can see from Fig. 1d that for −4t<μ<0-4t<\mu<0, the prediction performance from confusion scheme has a W-shape with the middle peak at μ=−2t\mu=-2t.

The W-shape can be understood as follows. We assume that the data has two different structures in the regimes below cc and above cc, and that the NN is able to find and distinguish them. We refer to these different structures as features. When we set c′=ac^{\prime}=a, the NN chooses to assign label 11 to both features and thus correctly predicts 100100% of the data. A similar analysis applies to c′=bc^{\prime}=b, except that every data point is assigned the label 00. When c′=cc^{\prime}=c is the correct labelling, the NN will choose to assign the right label to both sides of the critical point and again performs perfectly. When a<c′<ca<c^{\prime}<c, in the training phase the NN sees data with the same feature in the ranges from aa to c′c^{\prime} and from c′c^{\prime} to cc, but having different labels (hence the confusion). In this case it will choose to learn the label of the majority data, and the performance will be

Similar analysis applies to c<c′<bc<c^{\prime}<b. This gives the typical W-shape seen in Fig. 1d. Notice that if the point cc is not exactly centered between aa and bb, the W-shape will be slightly distorted. Its middle peak always corresponds to the correct labelling, but the depth of the minima will differ between the left and right.

We test the confusion scheme on the thermal phase transition in the two-dimensional classical Ising model Onsager 1944, which has been studied by both supervised learning Carrasquilla and Melko 2016 and unsupervised learning Wang 2016 methods. Here we train a NN (with L2L^{2} neurons in the input and hidden layers, and 2 neurons in the output layer) on the L×LL\times L classical configurations sampled from Monte Carlo simulation. As shown in Fig. 2, the W-shape again predicts the right transition temperature. Note the confusion scheme works better when the underlying feature in the data is shaper, i.e. for the larger system size L=20L=20.

To confirm that the confusion scheme indeed extracts nontrivial features from the input data. We have checked the performance curve from confusion scheme, when the NN is trained on unstructured random data. We use a fictive parameter as a tuning parameter, but have completely unstructured (random) data as a function of it. Hence, the network will not find structure in the data, and a correct labelling does not exist. The middle peak of the characteristic W-shape disappears, turning it into a V-shape.

We notice that the choice of the learning rate (α\alpha) and regularization (l2l_{2}) is essential for a successful training. The use of regularization is expected to reduce overfitting and make the network less sensitive to small variations of the data, hence forcing it to rather learn its structure Nielsen 2015. However, the confusion scheme depends solely on the ability of finding the majority label for the underlying structure in the data. In this sense, overfitting is not necessarily bad. Indeed we have observed that training with a negative l2l_{2} may lead to an equally good performance. We speculate that this is because a negative l2l_{2} tries to quickly increase the weights, making it harder for the network to change its opinion about data samples in later stages. If the initial training data is uniformly sampled, meaning the majority data is indeed represented by a majority, the network will rapidly adjust its weights to this majority. Lastly, we mention that trainings are performed in epochs. In each epoch all training data is passed once in batches of size NbN_{b} in a random order. The training is stopped when a clear W-shape is formed.

II.4 Random-field Heisenberg chain

We will now test our proposed scheme on an example where we the exact location of the transition point is now known Nandkishore and Huse 2015. We study a case of interest in recent literature, namely that of many-body localization. We consider the following model:

where S\bm{S} denote spin-1/21/2 operators. The local fields hiαh^{\alpha}_{i} are drawn from a uniform box distribution with zero mean and width hmaxαh^{\alpha}_{\textrm{max}}. We set hmaxx=hmaxz=hmaxh^{x}_{\textrm{max}}=h^{z}_{\textrm{max}}=h_{\textrm{max}} and hmaxy=0h^{y}_{\textrm{max}}=0. The disorder allows us to generate many samples at a fixed set of model parameters, in analogy to the different configurations for a fixed temperature in the classical spin systems Carrasquilla and Melko 2016; Wang 2016.

The model in equation (3) has a transition between thermalizing and non-thermalizing (i.e. many-body localized) behavior, driven by the disorder strength hmaxh_{\textrm{max}}. In particular, when varying hmaxh_{\textrm{max}}, both the energy level statistics as well as the statistics of the entanglement spectra change their nature Geraedts et al. 2016. For the case of the energy levels, the gaps (level spacings) follow either a Wigner-Dyson distribution for the thermalizing phase, or a Poisson distribution for the localized phase; while for the entanglement spectrum, the Wigner-Dyson distribution is replaced by a semi-Poisson distribution. Note that the change of ES can already be seen from the statistics in a single eigenstate Geraedts et al. 2016.

We numerically obtain the entanglement spectrum for the groundstate of the model in equation (3), for disorder strengths between hmax=Jh_{\textrm{max}}=J and hmax=5Jh_{\textrm{max}}=5J. The transition was shown to happen around hmax≈3Jh_{\textrm{max}}\approx 3J Geraedts et al. 2016, but we stress that our method does not rely on this knowledge. We would simply have started from a larger width of points, and then systematically narrow it down to the current range. At each value of hmaxh_{\textrm{max}} we generate 10510^{5} disorder realizations for system size L=12L=12 and calculate the entanglement spectrum for LA=LB=6L_{A}=L_{B}=6. These 26=642^{6}=64 levels are used as the input to the NN.

First, we try to use an unsupervised PCA to cluster the data. This analysis shows that the first two principal components are dominant, with the other components being of order 10−410^{-4} or less. However, a scatterplot of the data when projected onto the first two principal components (shown in Fig. 3a) does not reveal a clear clustering of the spectra.

We therefore turn to train a shallow feedforward network on the entanglement spectra in order to use the confusion scheme. Here we use a network with 64 input neurons, 100 hidden neurons and 2 output neurons. The results are shown in Fig. 3b. Also in this case, the characteristic W-shape is obtained and we detect the transition at hc≈3Jh_{c}\approx 3J. In addition to the previous cases, we also consider explicitly the performance of the network Nhc′\mathcal{N}_{h^{\prime}_{c}} at hc′h^{\prime}_{c}. We do this to confirm that the labelling with hc′h^{\prime}_{c} at 3J3J is indeed correct. We expect namely that the training of the network is most robust against changes in its parameters for the correct labelling. In other words, we may also look for the hc′h^{\prime}_{c} at which the training is most independent of chosen conditions. As shown in Fig. 3c, this point is also at hc′≈3Jh^{\prime}_{c}\approx 3J.

III Discussion

In this work we have explored the uses of machine learning in identifying (quantum) phase transitions from the overall performance of neural networks. As input data for training we proposed to use the entanglement spectrum of a quantum state, since it provides an excellent way of compressing the otherwise exponentially large wavefunction. We have shown that by confusing the neural network, by purposefully mislabelling the data, we are able to identify the phase transition between two phases of matter without prior knowledge of them. We demonstrated this for three different scenario’s. We expect our method to be useful also in situations where it is not clear whether two phases are identical, or whether intervening phases exist in between other phases. If the underlying data in such situations has a (hidden) pattern that distinguishes them, our machine learning approach may well be able to find it. Our method is not limited to applications in condensed matter physics and phase transitions, but might prove useful in any field in which large amounts of data can be classified using a priori unknown labels.

An interesting direction for future studies is the relaxation of the assumption that there are only two phases to be distinguished. If there are multiple phase transitions present in the data, the characteristic W-shape will be modified, and its new shape (i.e. the number of peaks) will signal the correct number of different labels. Additionally, it may be possible to formulate this method in a self-consistent way, with an adaptive labelling and having the algorithm determine the correct labels by itself.

Acknowledgements

E.v.N and S.H. gratefully acknowledge financial support from the Swiss National Science Foundation (SNSF). Y.-H.L. is supported by ERC Advanced Grant SIMCOFE. E.v.N. acknowledges fruitful discussions with Maciej Koch-Janusz on extending the confusion scheme to the case with multiple phases. E.v.N. and Y.-H.L. acknowledge helpful discussions with Giuseppe Carleo, Juan Osorio and Lei Wang. E.v.N and S.H. thank Andreas Krause for useful discussion on machine learning.

References