Machine learning phases of matter

Juan Carrasquilla, Roger G. Melko

Appendix A Details of the toy model

The analytical model encodes the low- and high-temperature phases of the Ising model through their magnetization. The hidden layer contains 3 perceptrons (a neuron with a Heaviside step nonlinearity); the first two perceptrons activate when the input states are mostly polarized, while the third one activate if the states are polarized up or unpolarized. Notice that the third neuron can also be choosen to activate if the states are polarized down or unpolarized. The resulting outcomes are recombined in the output layer and produce the desired classification of the state. The hidden layer is parametrized through a weight matrix and bias vector given by

where 0<ϵ<10<\epsilon<1 is the only free parameter of the model. The arguments of the three hidden layer neurons, in terms of the weight matrix, bias vector, and a particular Ising configuration x=(σ1σ2,...,σN)Tx=\left(\sigma_{1}\sigma_{2},...,\sigma_{N}\right)^{\text{T}}, are given by

where m(x)=1N∑i=1Nσim(x)=\frac{1}{N}\sum\limits_{i=1}^{N}\sigma_{i} is the magnetization of the Ising configuration. In Figure 5(A) we display the components of the Wx+bWx+b vector as a function of the magnetization of the Ising state m(x)m(x). The first and second neuron activate when the state is predominantly polarized, i.e., when m(x)>ϵm(x)>\epsilon or m(x)<−ϵm(x)<-\epsilon. The third neuron activates if the state has a magnetization m(x)>−ϵm(x)>-\epsilon, which means that, in the limit where 0<ϵ≪10<\epsilon\ll 1, it activates when the state is either polarized or unpolarized. The parameter ϵ\epsilon is thus a threshold value of the magnetization that helps deciding whether the state is considered polarized or not.

The output layer is parametrized through a weight matrix and bias vector given by

where these arbitrary choices ensure that the ordered, low-TT output neuron OLow-T=1O_{\text{Low-T}}=1 is active when either the spins polarize mostly ↑\uparrow or ↓\downarrow. On the other hand, when the ↑ ∥0\uparrow\,\parallel 0 neuron is active but the ↑\uparrow is not, then the high-temperature output neuron OHigh-T=1O_{\text{High-T}}=1, symbolizing a high-temperature state.

To illustrate what the effects of the training on the parameters WW and bb are, we consider the numerical training of a fully-connected neural network with only 3 neurons using the same setup and training/test data used in our ferromagnetic model (Figure 1) for the L=30L=30 system. In Figure 5(B) we display the argument W0x+b0W_{0}x+b_{0} of the input layer at training iteration t=0t=0 for configurations xx in the test set as a function of the magnetization of the configurations m(x)m(x). The weights have been randomly initialized at t=0t=0 from a normal distribution with zero mean and unit standard deviation. As the training proceeds, the parameters adjust such that the components of the vector Wtx+btW_{t}x+b_{t} approximately become linear functions of the magnetization m(x)m(x), as shown in Figure 5(C), in agreement with the assumptions of our toy model. These results clearly support our claim about the neural network’s ability encode and learn the magnetization in the hidden layer.

Appendix B Visualizing the action of a neural network on the Ising ferromagnet

A strategy to gain intuition for how these neural networks operate is to produce a low-dimensional visualization of data used in the training. We consider the t-distributed stochastic neighbor embedding (t-SNE) technique t-SNE where high-dimensional data is embedded in two or three dimensions so that data points close to each other in the original space are also positioned close to each other in the embedded low-dimensional space. Figure 6 displays a t-SNE visualization of the Ising configurations used the training of our ferromagnetic model. The two low-temperature blue regions correspond to the two ordered states with spins polarized either up or down. The high-temperature red region identifies the paramagnetic state. The resulting neural networks, which are functions defined over the high-dimensional state space, become trained so that the low-temperature output neuron takes a high value in the cool region (and vice versa), crossing over to a low value as the system is warmed through the orange hyperplane. This allows the classification of a state in terms of the neuron values.

Appendix C Details of the convolutional neural network of the Ising lattice gauge theory

The exact architecture of the convolutional neural network (CNN) Goodfellow-et-al-2016-Book , schematically described in Figure 4, is as follows. The input layer is a two-dimensional Ising spin configuration with N=16×16×2N=16\times 16\times 2 spins, where σi=±1\sigma_{i}=\pm 1. The first hidden layer convolves 64 2×22\times 2 filters on each of the two sublattices of the model with a unit stride, no padding, with periodic boundary conditions, followed by rectified linear unit (ReLu). The final hidden layer is a fully-connected layer with 64 ReLu units, while the output is a softmax layer with two outputs (correponding to T=0T=0 and T=∞T=\infty states). To prevent overfitting, we apply a dropout regularization in the fully-connected layer dropout . Our model has been implemented using TensorFlow tensorflow2015-whitepaper .

Since our CCN correctly classifies T=0T=0 and T=∞T=\infty states with 100%100\% accuracy, we would like to scrutinize the origin of its discriminiative power by asking whether it discerns the states by the presence (or absence) of the local Hamiltonian constraints or the extended closed-loop structure. Our strategy consists in the construction of new test sets with modified low temperature states, as detailed below. First, we consider transformations that do not destroy the topological order Castelnovo2007 of the T=0T=0 state but change the local constraints of the original Ising lattice gauge theory. As shown in the configurations in Figure 7(A) and (B), we consider transformations where a spin is flipped every m=2m=2 (A) (m=8m=8 (B)) plaquettes. The positions of the flipped spins are marked with red crosses. After optimizing the CNN using the original training set, the neural network classifies most of the transformed T=0T=0 states as high-temperature ones, resulting in an overall test accuracy of 50%50\% and 55%55\% for m=2m=2 and m=8m=8, respectively. This reveals that the neural network relies on the presence of satisfied local constraints of the original Ising lattice gauge theory, and not on the topological order of the state, in deciding whether a state is considered low or high temperature. Second, we consider a new test set where the T=0T=0 states retain most local constraints but disrupt non-local features like the extended closed-loop structure. We consider dividing the original states into 4 pieces as shown in Figure 7(C) and then reshuffling the 4 pieces among different states, subsequently stitching them to form new “low” temperature configurations. The new configurations will contain defects along the dashed lines in Figure 7(C), thus disrupting the extended closed-loop picture, but preserving the local constraints everywhere else in the configuration. We find that the trained CNN recognizes such states as ground states with high confidence, suggesting that the CNN does not use the extended closed-loop structure and indicating that local constraints are the only features that the CNN relies on for classification of the ground state.

In view of the conclusion above, we now present a toy model that uses a streamlined version of our original CNN constructed to explicitly detect satisfied energetic local constraints. The convolutional layer contains 16 2×\times2 filters per sublattice with unit stride in both directions and periodic boundary conditions. The convolutional layer is fully connected to two perceptron neurons in the output layer, as described below.

A schematic representation of the toy CNN is presented in Figure 8. The values of the filters WyxsfW_{yxsf} are presented in Table 1, where xx and yy represent the spatial indices of the convolution, and ss and ff label the sublattice and the filter, respectively. The purpose of the filters is to individually process each plaquette in the spin configuration and determine whether its energetic constraints are satisfied or not. The Ising gauge theory contains 1616 different spin configurations per plaquette, of which 8 satisfy the energetic constraints of the Hamiltonian. The first group of 16 filters WyxsfW_{yxsf} (f=1f=1 thorugh f=8f=8, left blue column in Table 1) detect satisfied plaquettes, while the remaining 16 (f=9f=9 through f=16f=16, right red column in Table 1) detect unsatisfied plaquettes. The bias of the convolutional layer is a 16-dimensional vector given by bc=−(2+δ)(1⋯1)Tb_{c}=-(2+\delta)\left(1\cdots 1\right)^{\text{T}}, where 0<δ≪10<\delta\ll 1 is a small parameter. The outcome of the convolutional layer consists of 16 two-dimensional arrays of size L×LL\times L processed through perceptrons. Here the total number of spins in the configuration is N=2×L×LN=2\times L\times L, where LL is the linear size of the system. The outcome of the convolutional layer is reshaped into a 16×L216\times L^{2}-dimensional vector such that the first 8L28L^{2} entries correspond to the outcome of the first group of filters (f=1f=1 thorugh f=8f=8) while the remaining 8L28L^{2} correspond to the last group of filters (f=9f=9 through f=16f=16). The output layer, which is fully connected to the the convolutional layer, contains two perceptron neurons denoted by O0O_{0} and O∞O_{\infty} for zero- and high-temperature states, respectively. It is parametrized through a weight matrix and a bias vector given by

These choices ensure that whenever an unsatisfied plaquette is encountered by the convolutional layer, the zero-temperature neuron is O0=0O_{0}=0 and the high-temperature O∞=1O_{\infty}=1, while only if all energetic constraints are satisfied O0=1O_{0}=1 and O∞=0O_{\infty}=0, thus allowing the classification of the states. When used on our test sets, the model performs the classification task with a 100%100\% accuracy, which means that all the high temperature states in the test set contain least one unsatisfied plaquette. Note that the classification error for this task is expected to be exponentially small in the volume of the system, since at infinte temperature the ground states appear with exponentially small probability. Having distilled the model’s basic ingredients, we proceed to train an analogue model numerically starting from random weights and biases WyxsfW_{yxsf}, WoW_{\text{o}}, bcb_{\text{c}}, and bob_{\text{o}}. Further, we replace the perceptron nonlinearities by ReLu units and a softmax output layer to enable a reliable numerical training. After the training, the model performs the classification task with a 100%100\% accuracy on the test sets, as expected.

As a consequence of the classification scheme provided by the analytical toy model, we observe that the values of the zero-temperature neuron O0O_{0} behave exactly like the amplitudes of one of the ground states of the toric code written in the σz\sigma_{z} basis Kitaev20032 . The ground state described by O0O_{0} is a linear combination of all 4 ground states with well defined parity on the torus. More precisely, such a state can be written as ∣ψtoric⟩=∑σz1,...,σzNO0(σz1...σzN)∣σz1...σzN⟩|\psi_{\text{toric}}\rangle=\sum_{\sigma_{z1},...,\sigma_{zN}}O_{0}(\sigma_{z1}...\sigma_{zN})|\sigma_{z1}...\sigma_{zN}\rangle, where the spin configurations σzi=±1\sigma_{zi}=\pm 1, and O0(σz1...σzN)O_{0}(\sigma_{z1}...\sigma_{zN}) corresponds to the value of O0O_{0} after a feed-forward pass of the neural network for a given a input configuration σz 1,...,σz N\sigma_{z\,1},...,\sigma_{z\,N}. Our model bears resemblance with the construction of the ground state of the toric code in terms of projected entangled pair states in that local tensors project out states containing plaquettes with odd parity PhysRevLett.109.260401 . These observations suggest that convolutional neural networks have the potential to represent ground states with topological order.

References