Solving the Quantum Many-Body Problem with Artificial Neural Networks
Giuseppe Carleo, Matthias Troyer
Neural-Network Quantum States —
Consider a quantum system with discrete-valued degrees of freedom , which may be spins, bosonic occupation numbers, or similar. The many-body wave function is a mapping of the dimensional set to (exponentially many) complex numbers which fully specify the amplitude and the phase of the quantum state. The point of view we take here is to interpret the wave function as a computational black box which, given an input many-body configuration , returns a phase and an amplitude according to . Our goal is to approximate this computational black box with a neural network, trained to best represent . Different possible choices for the artificial neural-network architectures have been proposed to solve specific tasks, and the best architecture to describe a many-body quantum system may vary from one case to another. For the sake of concreteness, in the following we specialize our discussion to restricted Boltzmann machines (RBM) architectures, and apply them to describe spin quantum systems. In this case, RBM artificial networks are constituted by one visible layer of nodes, corresponding to the physical spin variables in a chosen basis (say for example ) , and a single hidden layer of auxiliary spin variables () (see Fig. 1). This description corresponds to a variational expression for the quantum states which reads:
where is a set of hidden spin variables, and the weights fully specify the response of the network to a given input state . Since this architecture features no intra-layer interactions, the hidden variables can be explicitly traced out, and the wave function reads where . The network weights are, in general, to be taken complex-valued in order to provide a complete description of both the amplitude and the wave-function’s phase.
The mathematical foundations for the ability of NQS to describe intricate many-body wave functions are the numerously established representability theorems kolmogorov1961onthe ; hornik1991approximation ; leroux2008representational , which guarantee the existence of network approximates of high-dimensional functions, provided a sufficient level of smoothness and regularity is met in the function to be approximated. Since in most physically relevant situations the many-body wave function reasonably satisfies these requirements, we can expect the NQS form to be of broad applicability. One of the practical advantages of this representation is that its quality can, in principle, be systematically improved upon increasing the number of hidden variables. The number (or equivalently the density ) then plays a role analogous to the bond dimension for the MPS. Notice however that the correlations induced by the hidden units are intrinsically non local in space and are therefore well suited to describe quantum systems in arbitrary dimension. Another convenient point of the NQS representation is that it can be formulated in a symmetry-conserving fashion. For example, lattice translation symmetry can be used to reduce the number of variational parameters of the NQS ansatz, in the same spirit of shift-invariant RBM’s sohn2012learning ; norouzi2009stacksof . Specifically, for integer hidden variable density , the weight matrix takes the form of feature filters , for . These filters have a total of variational elements in lieu of the elements of the asymmetric case (see Supp. Mat. for further details).
Given a general expression for the quantum many-body state, we are now left with the task of solving the many-body problem upon machine learning of the network parameters . In the most interesting applications the exact many-body state is unknown, and it is typically found upon solution either of the static Schrödinger equation , either of the time-dependent one , for a given Hamiltonian . In the absence of samples drawn according to the exact wave function, supervised learning of is therefore not a viable option. Instead, in the following we derive a consistent reinforcement learning approach, in which either the ground-state wave function or the time-dependent one are learned on the basis of feedback from variational principles.
Ground State —
To demonstrate the accuracy of the NQS in the description of complex many-body quantum states, we first focus on the goal of finding the best neural-network representation of the unknown ground state of a given Hamiltonian . In this context, reinforcement learning is realized through minimization of the expectation value of the energy with respect to the network weights . In the stochastic setting, this is achieved with an iterative scheme. At each iteration , a Monte Carlo sampling of is realized, for a given set of parameters . At the same time, stochastic estimates of the energy gradient are obtained. These are then used to propose a next set of weights with an improved gradient-descent optimization sorella2007weakbinding . The overall computational cost of this approach is comparable to that of standard ground-state Quantum Monte Carlo simulations (see Supp. Material).
To validate our scheme, we consider the problem of finding the ground state of two prototypical spin models, the transverse-field Ising (TFI) model and the anti-ferromagnetic Heisenberg (AFH) model. Their Hamiltonians are
respectively, where are Pauli matrices.
In the following, we consider the case of both one and two dimensional lattices with periodic boundary conditions (PBC). In Fig. 2 we show the optimal network structure of the ground states of the two spin models for a hidden variables density and with imposed translational symmetries. We find that each filter learns specific correlation features emerging in the ground state wave function. For example, in the 2D case it can be seen (Fig. 2, rightmost panels) how the neural network learns patterns corresponding to anti-ferromagnetic correlations. The general behavior of the NQS is completely analogous to what observed in convolutional neural networks, where different layers learn specific structures of the input data.
In Fig. 3 we show the accuracy of the NQS states, quantified by the relative error on the ground-state energy , for several values of and model parameters. In the left panel, we compare the variational NQS energies with the exact result obtained by fermionization of the TFI model, on a one-dimensional chain with PBC. The most striking result is that NQS achieve a controllable and arbitrary accuracy which is compatible with a power-law behavior in . The hardest to learn ground-state is at the quantum critical point , where nonetheless a remarkable accuracy of one part per million can be easily achieved with a relatively modest density of hidden units. The same remarkable accuracy is obtained for the more complex one-dimensional AFH model (central panel). In this case we observe as well a systematic drop in the ground-state energy error, which for a small attains the same very high precision obtained for the TFI model at the critical point. Our results are compared with the accuracy obtained with the spin-Jastrow ansatz (dashed line in the central panel), which we improve by several orders of magnitude. It is also interesting to compare the value of with the MPS bond dimension , needed to reach the same level of accuracy. For example, on the AFH model with PBC, we find that with a standard DMRG implementation dolfi2014matrixproduct we need to reach the accuracy we have at . This points towards a more compact representation of the many-body state in the NQS case, which features about orders of magnitude less variational parameters than the corresponding MPS ansatz.
We next study the AFH model on a two-dimensional square lattice, comparing in the right panel of Fig. 3 to QMC results sandvik1997finitesize . As expected from entanglement considerations, the 2D case proves harder for the NQS. Nonetheless, we always find a systematic improvement of the variational energy upon increasing , qualitatively similar to the 1D case. The increased difficulty of the problem is reflected in a slower convergence. We still obtain results at the level of existing state-of-the-art methods or better. In particular, with a relatively small hidden unit density we already obtain results at the same level than the best known variational ansatz to-date for finite clusters (the EPS of Ref. mezzacapo2009groundstate and the PEPS states of Ref. lubasch2014algorithms ). Further increasing then leads to a sizable improvement and consequently yields the best variational results so-far-reported for this 2D model on finite lattices.
Unitary Dynamics —
NQS are not limited to ground-state problems but can be extended to the time-dependent Schrödinger equation. For this purpose we define complex-valued and time-dependent network weights which at each time are trained to best reproduce the quantum dynamics, in the sense of the Dirac-Frenkel time-dependent variational principle dirac1930noteon ; frenkel1934wavemechanics . In this context, the variational residuals
are the objective functions to be minimized as a function of the time derivatives of the weights (see Supp. Mat.) In the stochastic framework, this is achieved by a time-dependent VMC method carleo2012localization ; carleo2014lightcone , which samples at each time and provides the best stochastic estimate of the that minimize , with a computational cost . Once the time derivatives determined, these can be conveniently used to obtain the full time evolution after time-integration.
To demonstrate the effectiveness of the NQS in the dynamical context, we consider the unitary dynamics induced by quantum quenches in the coupling constants of our spin models. In the TFI model we induce a non-trivial quantum dynamics by means of an instantaneous change in the transverse field: the system is initially prepared in the ground-state of the TFI model for some transverse field, , and then let evolve under the action of the TFI Hamiltonian with a transverse field . We compare our results with the analytical solution obtained from fermionization of the TFI model for a one-dimensional chain with PBC. In the left panel of Fig. 4 the exact results for the time-dependent transverse spin polarization are compared to NQS with . In the AFH model, we study instead quantum quenches in the longitudinal coupling and monitor the time evolution of the nearest-neighbors correlations. Our results for the time evolution (and with ) are compared with the numerically-exact MPS dynamics white2004realtime ; vidal2004efficient ; daley2004timedependent for a system with open boundaries (see Fig. 4, right panel).
The high accuracy obtained also for the unitary dynamics further confirms that neural network-based approaches can be fruitfully used to solve the quantum many-body problem not only for ground-state properties but also to model the evolution induced by a complex set of excited quantum states. It is all in all remarkable that a purely stochastic approach can solve with arbitrary degree of accuracy a class of problems which have been traditionally inaccessible to QMC methods for the past years. The flexibility of the NQS representation indeed allows for an effective solution of the infamous phase problem plaguing the totality of existing exact stochastic schemes based on Feynman’s path integrals.
Outlook —
Variational quantum states based on artificial neural networks can be used to efficiently capture the complexity of entangled many-body systems both in one a two dimensions. Despite the simplicity of the restricted Boltzmann machines used here, very accurate results for both ground-state and dynamical properties of prototypical spin models can be readily obtained. Potentially many novel research lines can be envisaged in the near future. For example, the inclusion of the most recent advances in machine learning, like deep network architectures, might be further beneficial to increase the expressive power of the NQS. Furthermore, the extension of our approach to treat quantum systems other than interacting spins is, in principle, straightforward. In this respect, applications to answer the most challenging questions concerning interacting fermions in two-dimensions can already be anticipated. Finally, at variance with Tensor Network States, the NQS feature intrinsically non-local correlations which can lead to substantially more compact representations of many-body quantum states. A formal analysis of the NQS entanglement properties might therefore bring about substantially new concepts in quantum information theory.
References
Appendix A Stochastic Optimization For The Ground State
In the first part of our Paper we have considered the goal of finding the best representation of the ground state of a given quantum Hamiltonian . The expectation value over our variational states is a functional of the network weights . In order to obtain an optimal solution for which , several optimization approaches can be used. Here, we have found convenient to adopt the Stochastic Reconfiguration (SR) method of Sorella et al. sorella2007weakbinding , which can be interpreted as an effective imaginary-time evolution in the variational subspace. Introducing the variational derivatives with respect to the -th network parameter,
the SR updates at the th iteration are of the form
where we have introduced the (positive-definite) covariance matrix
and a scaling parameter . Since the covariance matrix can be non-invertible, denotes its Moore-Penrose pseudo-inverse. Alternatively, an explicit regularization can be applied, of the form . In our work we have preferred the latter regularization, with a decaying parameter and typically take , and .
Initially the network weights are set to some small random numbers and then optimized with the procedure outlined above. In Fig. 5 we show the typical behavior of the optimization algorithm, which systematically approaches the exact energy upon increasing the hidden units density .
Appendix B Time-Dependent Variational Monte Carlo
In the second part of our Paper we have considered the problem of solving the many-body Schrödinger equation with a variational ansatz of the NQS form. This task can be efficiently accomplished by means of the Time-Dependent Variational Monte Carlo (t-VMC) method of Carleo et al.
are a functional of the variational parameters derivatives, , and can be interpreted as the quantum distance between the exactly-evolved state and the variationally evolved one. Since in general we work with unnormalized quantum states, the correct Hilbert-space distance is given by the Fubini-Study metrics, given by
where the correlation matrix and the forces are defined analogously to the previous section. In this case the diagonal regularization, in general, cannot be applied, and strictly denotes the Moore-Penrose pseudo-inverse.
The outlined procedure is globally stable as also already proven for other wave functions in past works using the t-VMC approach. In Fig. 6 we show the typical behavior of the time-evolved physical properties of interest, which systematically approach the exact results when increasing .
Appendix C Efficient Stochastic Sampling
We complete the supplementary information giving an explicit expression for the variational derivatives previously introduced and of the overall computational cost of the stochastic sampling. We start rewriting the NQS in the form
In our stochastic procedure, we generate a Markov chain of many-body configurations sampling the square modulus of the wave function for a given set of variational parameters. This task can be achieved through a simple Metropolis-Hastings algorithm metropolis1953equation , in which at each step of the Markov chain a random spin is flipped and the new configuration accepted according to the probability
In order to efficiently compute these acceptances, as well as the variational derivatives, it is useful to keep in memory look-up tables for the effective angles and update them when a new configuration is accepted. These are updated according to
when the spin has been flipped. The overall cost of a Monte Carlo sweep (i.e. of single-spin flip moves) is therefore Notice that the computation of the variational derivatives comes at the same computational cost as well as the computation of the local energies after a Monte Carlo sweep.
Appendix D Iterative Solver
The most time-consuming part of both the SR optimization and of the t-VMC method is the solution of the linear systems (6 and 11) in the presence of a large number of variational parameters . Explicitly forming the correlation matrix , via stochastic sampling, has a dominant quadratic cost in the number of variational parameters, , where denotes the number of Monte Carlo sweeps. However, this cost can be significantly reduced by means of iterative solvers which never form the covariance matrix explicitly. In particular, we adopt the MINRES-QLP method of Choi and Saunders choi2014algorithm , which implements a modified conjugate-gradient iteration based on Lanczos tridiagonalization. This method iteratively computes the pseudo-inverse within numerical precision. The backbone of iterative solvers is, in general, the application of the matrix to be inverted to a given (test) vector. This can be efficiently implemented due to the product structure of the covariance matrix, and determines a dominant complexity of operations for the sparse solver. For example, in the most challenging case when translational symmetry is absent, we have , and the dominant computational cost for solving (6 and 11) is in line with the complexity of the previously described Monte Carlo sampling.
Appendix E Implementing Symmetries
Very often, physical Hamiltonians exhibit intrinsic symmetries which must be satisfied also by their ground- and dynamically-evolved quantum states. These symmetries can be conveniently used to reduce the number of variational parameters in the NQS.
where the network weights have now a different dimension with respect to the standard NQS. In particular, and are vectors in the feature space with and the connectivity matrix contains elements. Notice that this expression corresponds effectively to a standard NQS with hidden variables. Tracing out explicitly the hidden variables, we obtain
In the specific case of site translation invariance, we have that the symmetry group has an orbit of elements. For a given feature , the matrix can be seen as a filter acting on the translated copies of a given spin configuration. In other words, each feature has a pool of associated hidden variables that act with the same filter on the symmetry-transformed images of the spins.