Performance of Bayesian linear regression in a model with mismatch

Jean Barbier, Wei-Kuo Chen, Dmitry Panchenko, Manuel Sáenz

Introduction and set-up

Linear regression with random covariates in the high-dimensional regime, where both the number of data points and their dimension scale proportionally, is a paradigmatic problem for modern inference. It has been studied mostly from two complementary perspectives: (i)(i) In Bayesian statistics where, due to important simplifications that occur in this setting, analyses have focused on the case of the estimator obtained as the expectation of the “true” posterior distribution of the regression coefficients . This is usually referred to as the Bayesian-optimal estimator. But also on sparse Bayesian regression, see and the section below on related works. (ii)(ii) Another large body of literature has studied instead the performance of optimization procedures in the context of robust statistics and M-estimation, where an estimator is obtained as the mode of a log-concave posterior measure or, equivalently, as the minimizer of a convex cost function .

Our work focuses on Bayesian inference outside the optimal setting, for which much less is known. The lack of perfect knowledge of the model generating the data forces one to make certain inaccurate prior assumptions on the distribution of the regression coefficients as well as on the form of the model. This is, in a sense, more realistic and of practical interest. One natural setting for this, is to consider Bayesian linear regression with some sources of mismatch. Due to the lack of certain simplifying symmetries inherent to the optimal setting of inference, known as the “Nishimori identities” in physics , even a small mismatch can result in several theoretical challenges.

In the present work, we aim to study a mismatched linear regression model which is as follows. Let x∗=(xi∗)i⩽Nx^{*}=(x_{i}^{*})_{i\leqslant N} be an unknown ground-truth of regression coefficients. Assume that

for some deterministic γ⩾0\gamma\geqslant 0 and fixed CC independent of NN. Let α>0\alpha>0 be a fixed parameter and let M=⌊αN⌋M=\lfloor\alpha N\rfloorTo be more precise, we might take M=αN+o(N)M=\alpha N+o(N). But tracking this correction makes no difference.; ∥ ⋅ ∥\|\,\cdot\,\| is the usual L2L_{2} norm. Consider a matrix G=(gki)k⩽M,i⩽NG=(g_{ki})_{k\leqslant M,i\leqslant N} with i.i.d. standard normal entries. The columns in GG are understood as random covariate vectors, while the rows correspond to independent samples. The responses y=(yk)k⩽My=(y_{k})_{k\leqslant M} are obtained according to the linear model

for Sk(x∗):=(Gˉx∗)k,S_{k}(x^{*}):=(\bar{G}x^{*})_{k}, where Gˉ:=N−1/2G\bar{G}:=N^{-1/2}G and the variables z=(zk)k⩽M∼i.i.d.N(0,Δ∗)z=(z_{k})_{k\leqslant M}\stackrel{{\scriptstyle\rm{i.i.d.}}}{{\sim}}\mathcal{N}(0,\Delta_{*}) for some Δ∗>0\Delta_{*}>0 which gives the strength of the noise; N(m,σ2)\mathcal{N}(m,\sigma^{2}) is the normal law of mean mm and variance σ2\sigma^{2}. Here we will consider the regression task of recovering x∗x^{*} from the data D:=(y,Gˉ)\mathcal{D}:=(y,\bar{G}) with two sources of mismatch:

The distribution of x∗x^{*} is assumed to be unknown. Thus, we take the prior to be that of a gaussian vector with independent coordinates of mean and variance 1/(2κ)>01/(2\kappa)>0.

Under these sources of mismatch, the Bayesian posterior distribution is given by

2 Main results and contributions

Our main contributions are establishing the asymptotics of the log-normalizing constant and the performance of estimator (4). Informally, we prove that, whenever the four-dimensional system (17)–(20) presented below admits a unique solution, the log-normalizing constant converges to some non-random quantity F(q,ρ)F(q,\rho), that is

and the mean-square error converges to some non-random quantity qq, that is

Here (q,ρ)(q,\rho) is the unique critical point of the real-valued function FF with domain over the set {(q,ρ):0⩽q⩽ρ<∞}\{(q,\rho):0\leqslant q\leqslant\rho<\infty\} defined in (21) below.

There are two major findings that follow from our results. First of all, somewhat surprisingly (5) implies that the performance of the Bayesian estimator with gaussian prior is insensitive to the statistical properties of the coefficients x∗x^{*} other than its empirical second moment. The second major finding regards our main running example for our results which will play a special role given its tractability in the analysis and its importance in statistics. This example will correspond to the model where the function u(s)u(s) is of the form −s2/(2Δ)-s^{2}/(2\Delta) for some Δ>0\Delta>0. Notice that this model is still mismatched since Δ\Delta is not necessarily equal to the Δ∗\Delta_{*} appearing in the model generating the data. When we convert the posterior distribution into a Gibbs measure parametrized by the inverse temperature β\beta we discover that, for quadratic uu, the mean-square error is independent of the inverse temperature β\beta in the high-dimensional limit. This means that when the performance is evaluated using a square loss, Bayesian regression with gaussian prior and ridge regression have the same performance. In other words, sampling and optimizing are equivalent in this setting. One may argue that the closedness between the mean and the mode of the posterior distribution with gaussian prior is mainly due to the isotropy of the covariates. This intuition makes perfect sense in the classical regime, where a lot of data is accessible M≫NM\gg N when maximum likelihood estimation is optimal . But in the high-dimensional regime where both MM and NN are large and comparable, this feature becomes much less evident. To the best of our knowledge, this fact was not rigorously obtained prior to our work. Exploring the generality of this equivalence between optimization and sampling procedures is an interesting open direction for future research. See for related connections between M-estimation and Bayesian inference from statistical mechanics heuristics.

The approach in this paper relies primarily on a connection between our regression task and a generalized version of the Shcherbina-Tirozzi (ST) mean-field spin glass model . We study the limits of the corresponding log-partition function and some critical quantities called the overlaps. By utilizing the so-called smart path method, analogous results were intensively explored by Talagrand assuming that uu grows at most linearly. Whereas the machinery therein relies on quite heavy analysis, the advantage of our approach is methodological. In addition to being able to relax the growth condition for uu from linear to quadratic, our analysis is also rather simple and self-contained. It is based on methods coming from the mathematical physics of mean-field spin glasses: mainly a rigorous version of the cavity method also called the leave-one-out method in statistics (see Sections 5 and 6 for details). We expect for the simplicity of our approach to allow us to extend the analysis to generalized linear models in mismatched settings, for which results similar to the present ones exist only in the Bayesian-optimal setting of inference . This is left for future work.

3 Related works

In recent years, many works on mismatched regression have appeared: on non-rigorous statistical mechanics techniques such as the replica method like, e.g., in ; on the analysis of approximate message-passing algorithms, like in ; and on Gordon’s convex min-max theorem . All these approaches strongly rely on the fact that the estimator is obtained as a unique minimizer of a convex function and therefore do not generalize easily to mismatched Bayesian settings. The leave-one-out method was also used in but, like in the aforementioned references, only for the analysis of (non-Bayesian) M-estimators.

Complementary to our setup is a recent work on the mismatched Bayesian linear regression that derives conditions under which the “naive mean-field approximation” to the log-partition function is valid, and provides an infinite-dimensional variational formula for it (see for a related setting). Last but not least, Bayesian linear regression has also been considered in some different high-dimensional scaling regimes than the one presented here: in settings with sparse ground-truth vectors of regression coefficients with a possibly vanishing fraction of non-zero entries. Reference studies in detail the properties of the posterior distribution with sparsity inducing priors in this case, see also for related prior works on sparse linear regression. It has recently been understood that in the very sparse regime, and with the number of sampled responses much smaller than the covariates’ dimension, intriguing “all-or-nothing” phase transitions occur where the recovery error jumps from almost zero to its maximum value at a certain threshold . For a recent review on the Bayesian approach to high-dimensional inference, see .

4 Organization

The paper is organized as follows. In Section 2, we explain the link between the high-dimensional regression problem presented here and the ST spin glass model. In Section 3, we first state two results on the main quantities of interest in spin glass models, namely, the free energy (i.e., log-partition function) and overlaps (that will later be related to the mean-square error) in the ST model. We then state our main results on the regression task with a particular emphasis on the classical example of the Bayesian regression with quadratic loss and gaussian prior. Sections 4–6 are devoted to establishing results concerning the ST model. More explicitly, Section 4 establishes a number of concentration properties for the overlaps and free energy. Section 5 shows how to obtain certain fixed-point equations for the “order parameters” of the model based on the concentration results in Section 4 by utilizing the cavity method and simple gaussian integration by parts. Section 6 proves a closed-form expression for the limiting log-partition function of the posterior in a straightforward manner thanks to the convergence results for the order parameters. Finally, our results on the regression task are established in Section 7.

Connection to the Shcherbina-Tirozzi model

and the corresponding log-partition function, also called free energy, by

In the terminology of the large deviation theory , it is the scaled cumulant log-moment generating function associated with the energy per variable HN(σ)/NH_{N}(\sigma)/N. When zk=0z_{k}=0 and uu grows at most linearly, the Hamiltonian HNH_{N} is known as the original ST model extensively studied in . In the setting of the present paper, we allow uu to grow quadratically (see Assumption 1 below for details).

Two key quantities for our analysis will be the overlaps between replicas defined according to

Additionally, we also define a pair of “conjugate overlaps” in terms of the following two auxiliary quantities

As one shall see, the vector (R1,1,R1,2,Q1,1,Q1,2)(R_{1,1},R_{1,2},Q_{1,1},Q_{1,2}) of overlaps order parameters is the “correct” one, in the sense that all the statistical analysis can be asymptotically expressed solely in terms of the limits of these quantities and the model at the “macroscopic level” is completely determined by them.

In view of our regression task, the posterior measure (3) can be thought of as a Gibbs measure of the spin glass model with Hamiltonian

As in the ST model, we define the free energy by

In particular, from (4), we readily see that the Bayesian estimator

The connection between the ST and our models relies on a number of change of variables. First of all, by letting w:=x−x∗w:=x-x^{*}, we may write

To simplify this integral, take an orthogonal matrix O∈O(N)O\in\mathcal{O}(N) such that Ox∗=N−1/2∥x∗∥1NOx^{*}=N^{-1/2}{\|x^{*}\|}\bm{1}_{N}. Note that conditionally on x∗x^{*},

which implies that GO⊺∼GGO^{\intercal}\sim G conditionally on x∗x^{*}. Note also that ⟨x∗,O⊺a⟩=N−1/2∥x∗∥⟨1N,a⟩\langle x^{*},O^{\intercal}a\rangle=N^{-1/2}\|x^{*}\|\langle\bm{1}_{N},a\rangle and ∥O⊺a∥=∥a∥.\|O^{\intercal}a\|=\|a\|. From these and using the change of variable σ:=Ow\sigma:=Ow,

which implies, from the rotational invariance of GG mentioned above and the symmetry of zk,z_{k}, that

for γN:=∥x∗∥2/N.\gamma_{N}:=\|x^{*}\|^{2}/N. If now we randomize h=hN=2κγNh=h_{N}=2\kappa\sqrt{\gamma_{N}} in the Hamiltonian (6) of the ST model independently of all other randomness, then we readily get

Consequently, to understand our regression problem, we may study the limiting free energy and overlaps corresponding to the generalized ST model first. This characterization is part of our main results and is stated in the next section.

Main results

The first part of our main results consists of an expression for the limiting free energy and the convergence of the overlaps in the generalized ST model defined in (6). Throughout the rest of this paper, the following assumption on the function uu is in force.

The function u(s)u(s) is concave, non-positive, and there exists a constant d>0d>0 such that

The concavity of uu is conceptually an important assumption, which ensures that our posterior distribution is a log-concave measure and, as a result, the overlaps are concentrated under the Gibbs measure GNG_{N}. While the other assumptions are purely technical, overall they weaken the settings considered in . In particular, in this reference uu is allowed to grow at most linearly and thus its results do not apply to our main example, namely, a quadratic uu.

for 0⩽q<ρ0\leqslant q<\rho and 0⩽rˉ⩽r0\leqslant\bar{r}\leqslant r, where z~,ξ\widetilde{z},\xi are independent random variables with normal law N(0,1)\mathcal{N}(0,1). If we let θ=z~Δ∗+q+ξρ−q\theta=\widetilde{z}\sqrt{\Delta_{*}+q}+\xi\sqrt{\rho-q}, basic gaussian integration by parts with respect to z~\widetilde{z} and ξ\xi implies that the critical points of the function Fˉ\bar{F} must satisfy the following system of equations:

Observe that if (q,ρ,r,rˉ)(q,\rho,r,\bar{r}) satisfies (17)–(20), then Fˉ(q,ρ,r,rˉ)\bar{F}(q,\rho,r,\bar{r}) and F(q,ρ)F(q,\rho) match each other as can be seen by rewriting (17) and (18) as

and plugging these into the second line in (16) after writing rq−rˉρ=−r(ρ−q)+(r−rˉ)ρrq-\bar{r}\rho=-r(\rho-q)+(r-\bar{r})\rho. Here, if we express r,rˉr,\bar{r} in terms of q,ρq,\rho in (19)–(20), the pair (q,ρ)(q,\rho) is also a critical point of FF. Our first main result states that the limiting free energy in the ST model is equal to F(q,ρ)F(q,\rho) as long as (17)–(20) have a unique solution.

The next result establishes the convergence of the overlaps (9) and (11).

From these and with the help of the identities (14) and (15), we can now compute the limiting free energy and the mean-square error associated to our regression task.

Assume the convergence (1) for the norm of the ground-truth coefficients towards some deterministic γ⩾0\sqrt{\gamma}\geqslant 0. For κ>0\kappa>0, h=2κγh=2\kappa\sqrt{\gamma}, and Δ∗⩾0\Delta_{*}\geqslant 0, if (17)–(20) have a unique solution (q,ρ,r,rˉ)(q,\rho,r,\bar{r}), then

Let u(s)=−s2/(2Δ),u(s)=-{s^{2}}/(2\Delta), where Δ>0\Delta>0 is not necessarily equal to the true gaussian noise variance Δ∗.\Delta_{*}. Obviously uu satisfies Assumption 1. More importantly, it can be checked that (17)–(20) have a unique solution given by

To see this, note that (19) and (20) can be written as

Take a=ρ−qa=\rho-q and a′=r−rˉa^{\prime}=r-\bar{r}. From the difference between (17) and (18) and that of (19) and (20), we arrive at

from which we can eliminate a′a^{\prime} to get a single equation for aa,

It is easy to see that the only non-negative solution to this equation is a=ca=c and thus, a′=α/(Δ+c).a^{\prime}=\alpha/(\Delta+c). Finally, using these values of aa and a′a^{\prime} and the second equation of (25) leads to

Note that qq is positive since (26) implies that aa′<1aa^{\prime}<1 and aa′<αaa^{\prime}<\alpha, which ensures that (Δ+c)2/c2=(aa′)2<α(\Delta+c)^{2}/c^{2}=(aa^{\prime})^{2}<\alpha. From this qq, (ρ,r,rˉ)(\rho,r,\bar{r}) can then be uniquely determined through the second to the fourth equations in (24).

Surprisingly, if we parametrize the posterior distribution by an “inverse temperature” hyperparameter β>0\beta>0, namely, the original Gibbs measure G~N\widetilde{G}_{N} in (12) becomes

meaning that Δ\Delta and κ\kappa are replaced by Δ/β\Delta/\beta and βκ\beta\kappa in the original model, then the resulting constant qβq_{\beta}, i.e., the mean-square error, remains the same as the original qq for any inverse temperature β>0\beta>0. This can be easily seen from the above equations (for quadratic uu). At the same time, formula (27) matches the one for ridge regression at zero temperature (see, e.g., equation (26) in ). This non-trivial fact is due to the rotational invariance of the model following from the isotropy of the covariates. But as mentioned already in the introduction, this is not a priori evident in the high-dimensional regime we consider here, where MM and NN are comparable. It would be interesting to explore in future research for which settings posterior mean and mode (i.e., sampling and optimizing) yield the same reconstruction performance.

Concentration

The last assumption (G-3) will be used in an equivalent form,

Also, the assumptions (29), (30) may look redundant given (31), but we state them for convenience.

We will later specialize these assumptions for the following two specific choices:

For n=1n=1: ϕ(s)=s2\phi(s)=s^{2} and ψ(s)=u′′(s)+u′(s)2\psi(s)=u^{\prime\prime}(s)+u^{\prime}(s)^{2},

For n=2n=2: ϕ(s1,s2)=s1s2\phi(s_{1},s_{2})=s_{1}s_{2} and ψ(s1,s2)=u′(s1)u′(s2)\psi(s_{1},s_{2})=u^{\prime}(s_{1})u^{\prime}(s_{2}).

One can easily see that Assumption 1 on uu implies (G-1)–(G-3) in both cases, by possibly increasing the value of dd.

for functions ϕ\phi and ψ\psi verifying assumptions (G-1)–(G-3). Define the (non-averaged) free energy associated to XN,λ,ηX_{N,\lambda,\eta} by

Note that assumptions (28), (29), and (32) as well as ∣λ∣,∣η∣⩽min⁡(η0,κ/(2d))|\lambda|,|\eta|\leqslant\min(\eta_{0},\kappa/(2d)) ensure that this quantity is a.s. finite in GG and z=(zk)k⩽Mz=(z_{k})_{k\leqslant M} Denote by ⟨ ⋅ ⟩λ,η\langle\,\cdot\,\rangle_{\lambda,\eta} the expectation with respect to the Gibbs measure proportional to exp⁡XN,λ,η(σ⃗)\exp X_{N,\lambda,\eta}(\vec{\sigma}):

Under the above assumptions, there exists some K>0K>0 such that

Moreover, for any k⩾1k\geqslant 1, there exists some Ck>0C_{k}>0 such that

In the rest of this section, we will establish Proposition 4.1. To fix our notation, for any real M×NM\times N-matrix AA, define the Frobenius and operator norms, respectively, as

Recall ∥⋅∥\|\cdot\| applied to a vector is the L2L_{2} norm and note that ∥A∥op⩽∥A∥\|A\|_{op}\leqslant\|A\|.

Let zˉ:=z/N\bar{z}:=z/\sqrt{N} and recall also that Gˉ:=G/N\bar{G}:=G/\sqrt{N}.

There exist ε>0\varepsilon>0 and K⩾1K\geqslant 1 such that if t,∣λ∣,∣η∣⩽εt,|\lambda|,|\eta|\leqslant\varepsilon then

By (37), if t+d∣λ∣⩽κ/2t+d|\lambda|\leqslant\kappa/2, then

which, together with (32) for ∣η∣⩽η0|\eta|\leqslant\eta_{0} implies that,

Combining this with (37) and (38), we obtain that

which together with the fact that ln⁡x⩽x+1\ln x\leqslant x+1 for all x>0x>0 implies that

Our assertion follows by putting this and (39) together.

We let I(A)I(A) be the indicator function for event AA.

There exist constants ε>0\varepsilon>0 and K,K′>0K,K^{\prime}>0 such that, for any ∣λ∣,∣η∣⩽ε|\lambda|,|\eta|\leqslant\varepsilon,

Proof. Recall the constants ε\varepsilon and KK from the statement of Lemma 4.2. If t⩽εt\leqslant\varepsilon, then for any a>0a>0,

Our proof is completed by taking a=2t−1KN(∥zˉ∥2+∥Gˉ∥op2+1+h2).a=\sqrt{{2t^{-1}KN(\|\bar{z}\|^{2}+\|\bar{G}\|_{op}^{2}+1+h^{2})}}.

Let k⩾1.k\geqslant 1. There exist 0<ε<10<\varepsilon<1 and K⩾1K\geqslant 1 such that, for any ∣λ∣,∣η∣⩽ε|\lambda|,|\eta|\leqslant\varepsilon,

For any non-negative random variable XX, by Lemma 3.1.8 in ,

Using this inequality with X=t∥σ⃗∥2X=t\|\vec{\sigma}\|^{2}, we have

2 Convexity and pseudo-Lipschitz property of the overlaps

If ∣λ∣+∥Gˉ∥op2∣η∣⩽κ/(dn)|\lambda|+\|\bar{G}\|_{op}^{2}|\eta|\leqslant\kappa/(dn) then the following function is convex:

Proof. For simplicity, for σ⃗=(σ1,…,σn)\vec{\sigma}=(\sigma^{1},\ldots,\sigma^{n}), we denote σ(i)=(σi1,…,σin).\sigma(i)=(\sigma_{i}^{1},\ldots,\sigma_{i}^{n}). From (29),

where in the third line we used the Minkowski inequality. Again, by (29),

we can again use the Minkowski inequality to obtain

For any k⩾1k\geqslant 1 there exists some Ck>0C_{k}>0 such that

Proof. Let us denote by s⃗:=σ⃗/N.\vec{s}:=\vec{\sigma}/\sqrt{N}. Note that from Lemma 4.6,

where we used the assumption (30). Using Lemma 4.4 with λ=η=0\lambda=\eta=0, we see that there exists some CC such that

3 Concentration of the free energy

For clarity, here we will now denote FN(λ,η)F_{N}(\lambda,\eta) by F(G,z)F(G,z). Let us start with the following result.

There exist constants ε>0\varepsilon>0 and K>0K>0 such that for all ∣λ∣,∣η∣⩽ε|\lambda|,|\eta|\leqslant\varepsilon,

Using that ∣u′(s)∣⩽d(1+∣s∣)|u^{\prime}(s)|\leqslant d(1+|s|) by (28) (below the constant KK amy change from place to place),

where the last inequality used (41). Note that from (31), there exists a constant d′d^{\prime} such that

Using this inequality, we can argue similarly that

In similar manner, we readily compute that

for some δ>0\delta>0, which implies Lemma 4.8.

There exist absolute constants ε>0\varepsilon>0 and K>0K>0 such that, for all ∣λ∣,∣η∣⩽ε|\lambda|,|\eta|\leqslant\varepsilon,

Proof. For simplicity of notation, denote

If we define yˉ:=y/N\bar{y}:=y/\sqrt{N}, then the above Lemma proves that

Let L>0L>0 be a constant to be chosen later and Lh:=L(1+h2)L_{h}:=L(1+h^{2}). Since, for r(y),r(y′)⩽Lhr(y),r(y^{\prime})\leqslant L_{h},

and we have equality for y′=yy^{\prime}=y, if we define

then Φ(y)=f(y)\Phi(y)=f(y) for all yy such that r(y)⩽Lhr(y)\leqslant L_{h}. From this definition, it can be seen from the reverse triangle inequality that

By (44) and the Gaussian-Poincaré concentration inequality,

Next, using (45) and the Cauchy-Schwarz inequality, we control

4 Proof of Proposition 4.1

We will only consider the case of the overlap R(σ⃗)R(\vec{\sigma}), since the proof for Q(σ⃗)Q(\vec{\sigma}) is almost identical. In this section we will denote by KK a constant that depends only on the parameters of the model, but can vary from one occurrence to the next. Let us consider the event

By (43), we can and will choose K⩾1K\geqslant 1 so that the probability of the event ENE_{N} is at least 1−4e−N1-4e^{-N}. From Lemma 4.7,

Let FN′F_{N}^{\prime} be the λ\lambda-derivative of FNF_{N}. In the remainder we will control the same expectation on the event ENE_{N}, and we will use the convexity of FN(λ,η)F_{N}(\lambda,\eta) in λ\lambda, which implies that

Take ε>0\varepsilon>0 such that, for ∣λ∣,∣η∣⩽ε|\lambda|,|\eta|\leqslant\varepsilon, the inequalities in Lemmas 4.3, 4.4 and 4.9 hold. By Lemma 4.9, the expectation of the square of the term between the parentheses above can be bounded by K/(Nλ2)K/(N\lambda^{2}).

In order to control the first term FN′(λ,0)−FN′(−λ,0)F_{N}^{\prime}(\lambda,0)-F_{N}^{\prime}(-\lambda,0) above, our goal will be to control

For simplicity of notation we denote s⃗:=σ⃗/N\vec{s}:=\vec{\sigma}/\sqrt{N}. Since, for ∣λ∣,∣η∣⩽ε|\lambda|,|\eta|\leqslant\varepsilon, the inequalities in Lemmas 4.3 and 4.4 hold, on the event ENE_{N}, for some L⩾1L\geqslant 1,

where we also denoted s⃗′:=σ⃗′/N\vec{s}^{\prime}:=\vec{\sigma}^{\prime}/\sqrt{N}.

We now proceed as in the proof of Lemma 4.9. Take LL as in (49) and DD as in (50), denote r(σ⃗):=∥s⃗∥r(\vec{\sigma}):=\|\vec{s}\| and let

Then, from (50), ρ(σ⃗)=R(σ⃗)\rho(\vec{\sigma})=R(\vec{\sigma}) for all r(σ⃗)⩽Lr(\vec{\sigma})\leqslant L,

Using (52) and the Cauchy-Schwarz and Minkowski inequalities,

where in the last line we used the first equation in (49). Since R(0)=ϕ(0)R(0)=\phi(0), by (50)

Therefore, the second equation in (49) implies that

where the last inequality used (51). Together with (53), this implies that, on the event ENE_{N},

assuming as above that ∣λ∣,∣η∣⩽ε|\lambda|,|\eta|\leqslant\varepsilon and ∣λ∣+K∣η∣⩽κ/(dn)|\lambda|+K|\eta|\leqslant\kappa/(dn). In particular,

and \bigl{|}F_{N}^{\prime}(\lambda,0)-F_{N}^{\prime}(-\lambda,0)\bigr{|}\leqslant K\lambda. Plugging this into (48) and, as we mentioned above, using Lemma 4.9 for the second term,

by optimizing over λ\lambda. Using (54) on the event ENE_{N} and then combining with (47) gives

This proves the first inequality in (36). The proof of the second inequality is identical.

Fixed point equations

This section establishes the proof of Theorem 3.3. The key ingredient here is to show that equations (17)–(20) are satisfied by the subsequential limits of the overlaps, based on a cavity/leave-one-out argument and gaussian integration by parts.

For every k⩾1k\geqslant 1, there exists some Ck>0C_{k}>0 such that, uniformly on NN,

Proof. Here we will denote u(S1(σ)+z1)u(S_{1}(\sigma)+z_{1}) by u1⩽0u_{1}\leqslant 0. Because ∣u1∣k|u_{1}|^{k} is an increasing function of ∣u1∣|u_{1}| and exp⁡(−∣u1∣)\exp(-|u_{1}|) is decreasing, the Harris inequality implies ⟨∣u1∣kexp⁡(−∣u1∣)⟩c⩽⟨∣u1∣k⟩c⟨exp⁡(−∣u1∣)⟩c\langle|u_{1}|^{k}\exp(-|u_{1}|)\rangle_{c}\leqslant\langle|u_{1}|^{k}\rangle_{c}\langle\exp(-|u_{1}|)\rangle_{c}. As a consequence we have that

Furthermore, because the standard gaussian vector g1:=(g11,…,g1N)g_{1}:=(g_{11},\dots,g_{1N}) is independent of σ\sigma under the cavity measure, by the rotational invariance of g1g_{1} we have that ⟨g1,σ⟩∼g∥σ∥\langle g_{1},\sigma\rangle\sim g\|\sigma\| where gg is an independent standard gaussian variable, so under the cavity measure S1(σ)∼g∥σ∥/N=gR1,1S_{1}(\sigma)\sim g\|\sigma\|/\sqrt{N}=g\sqrt{R_{1,1}}. Thus, by Assumption 1

Finally, by the independence of R1,1R_{1,1} from gg and z1z_{1} and Proposition 4.1, the right-hand side of this last equation is easily seen to be smaller than some constant that does not depend on NN.

There exist constants K,K′>0K,K^{\prime}>0 such that

Proof. As in the previous lemma, we will denote u(S1(σ)+z1)u(S_{1}(\sigma)+z_{1}) by u1u_{1}. Let

for t∈t\in, and denote by ⟨ ⋅ ⟩t\langle\,\cdot\,\rangle_{t} the expectation with respect to the associated Gibbs measure. We then clearly have that

for some constants K,K′>0K,K^{\prime}>0 that do not depend on t∈t\in. Here, the bound for the first factor follows from the Brascamp-Lieb inequality, while the bound on the second term is a generalization of Lemma 5.1. The proof for R1,2R_{1,2} is similar.

2 Consistency equations for the overlaps

Throughout this subsection, we will repeatedly use the concentration results for the generalized overlaps QQ and RR from Section 4 that are applicable to the overlaps Q1,1Q_{1,1} using ψ(s)=u′′(s)+u′(s)2\psi(s)=u^{\prime\prime}(s)+u^{\prime}(s)^{2}, Q1,2Q_{1,2} using ψ(s1,s2)=u′(s1)u′(s2)\psi(s_{1},s_{2})=u^{\prime}(s_{1})u^{\prime}(s_{2}), R1,1R_{1,1} using ϕ(s)=s2\phi(s)=s^{2}, and R1,2R_{1,2} using ϕ(s1,s2)=s1s2\phi(s_{1},s_{2})=s_{1}s_{2}. Denote

By Proposition 4.1, there exists some compact set KK such that for all N⩾1N\geqslant 1, (ρN,qN,rˉN,rN)(\rho_{N},q_{N},\bar{r}_{N},r_{N}) belongs to KK. Thus, there exists a subsequence (Nm)m⩾1(N_{m})_{m\geqslant 1} of system sizes along which

for some (ρ,q,rˉ,r)(\rho,q,\bar{r},r) in KK. For the rest of this section, we work with this convergent subsequence and we aim to show that (ρ,q,rˉ,r)(\rho,q,\bar{r},r) satisfies (17)–(20) in Propositions 5.4 and 5.5. To begin with, we establish

By Proposition 4.1 and Lemma 5.2, Σm→L2Σ\Sigma_{m}\xrightarrow{L^{2}}\Sigma as m→+∞m\to+\infty where

The vector (ρ,q,rˉ,r)(\rho,q,\bar{r},r) satisfies (19) and (20).

By equation (55), the overlap definition (11), the spin symmetry and the geometric series (recall u(s)⩽0u(s)\leqslant 0) we have that

Note that by Assumption 1 the first term is the mean of a continuous and bounded function of a1,…,aKa_{1},\dots,a_{K}. Then by Lemma 5.3 we have that

The second term can be bounded in the following way

where in the first line we used Assumption 1, in the second Cauchy-Schwarz and Jensen’s inequalities, and in the last one that by Lemma 5.1 there exists a constant C>0C>0 that bounds the first factor in the square root. Now, because (1−exp⁡u(a1))2K(1-\exp{u(a_{1})})^{{2K}} is a continuous and bounded function of a1a_{1}, by Lemma 5.3 we have that

By taking the limit m→+∞m\to+\infty in equation (58) and using the convergence of the overlap Q1,1Q_{1,1} along the subsequence (Nm)m⩾1(N_{m})_{m\geqslant 1}, this implies that

Consider the two terms in fkf_{k} separately. Taking the limit K→+∞K\to+\infty, by the monotone convergence theorem,

and since ∣u′′(a1)∣⩽d|u^{\prime\prime}(a_{1})|\leqslant d by Assumption 1, by the dominated convergence theorem, we also have

To see that the second term goes to , note that 1−exp⁡u(θ)<11-\exp{u(\theta)}<1. The limit then follows by the dominated convergence theorem. This proves rˉ=Φˉ(q,ρ)\bar{r}=\bar{\Phi}(q,\rho), i.e., equation (20).

The proof that r=Φ(q,ρ)r=\Phi(q,\rho) (equation (19)) follows in a similar way. Let K⩾1K\geqslant 1. For each k⩽Kk\leqslant K define

By Assumption 1, all these are continuous and bounded functions of a1,…,a2ka_{1},\dots,a_{2k}. As before, by the cavity formula and the geometric series we have that

Finally, the conclusion follows in a similar manner as for the first equation: Lemma 5.3 is used for the sum while the second term can be seen to go to when K→+∞K\to+\infty.

To establish the second pair of fixed point equations, a feasible approach would be to perform a cavity argument over the variables indexed by i=1,…,Ni=1,\ldots,N analogous to the proof of Proposition 5.4 (this “cavity in NN” is in general more involved than the “cavity in MM”). Nevertheless, by the virtue of concentration of the overlaps, we show that they can also be obtained through an elementary derivation via gaussian integration by parts.

The vector (ρ,q,rˉ,r)(\rho,q,\bar{r},r) satisfies (17) and (18).

Proof. To see this, note that by the term −κ(σ11)2-\kappa(\sigma_{1}^{1})^{2} in the Hamiltonian (6) we can use gaussian integration by parts with respect to σ11\sigma_{1}^{1} and the gaussian variables (gk1)k⩽M(g_{k1})_{k\leqslant M}. In particular

and, denoting the modified Hamiltonian H^N(σ):=∑k⩽Mu(Sk(σ)+zk)−h∑i⩽Nσi\widehat{H}_{N}(\sigma):=\sum_{k\leqslant M}u(S_{k}(\sigma)+z_{k})-h\sum_{i\leqslant N}\sigma_{i},

These identities imply (recall the definitions (10))

The concentrations of R1,1R_{1,1}, R1,2R_{1,2}, Q1,1Q_{1,1}, and Q1,2Q_{1,2} (see Proposition 4.1) then implies that

Sending m→+∞m\to+\infty, these lead to q=Ψ(r,rˉ)q=\Psi(r,\bar{r}) and ρ=Ψˉ(r,rˉ)\rho=\bar{\Psi}(r,\bar{r}) (equations (17) and (18)) and complete our proof.

Propositions 5.4 and 5.5 together conclude that (ρ,q,rˉ,r)(\rho,q,\bar{r},r) satisfies (17)–(20). We are now in position to prove Theorem 3.3 for the ST spin glass model.

3 Proof of Theorem 3.3.

Consider any subsequence of system sizes. As explained above Lemma 5.3, there exists some further subsequence (Nm)n⩾1(N_{m})_{n\geqslant 1} such that

for some (ρ,q,rˉ,r)(\rho,q,\bar{r},r) in a compact set KK. By Propositions 5.4 and 5.5 and the continuity of the functions Φˉ\bar{\Phi}, Φ\Phi, Ψˉ\bar{\Psi}, and Ψ\Psi, we see that (ρ,q,rˉ,r)(\rho,q,\bar{r},r) is a solution of (17)–(20). Since this system of equations has only one unique solution as ensured by the given assumption, it follows that any convergent subsequence will share the same limit and this implies the assertion.

Proof of Theorem 3.2

In this section, we establish the convergence of FNF_{N} using a basic interpolation in this section. Let α,Δ∗,Δ,κ>0\alpha,\Delta_{*},\Delta,\kappa>0 be fixed. For t∈t\in, consider an auxiliary free energy

For any t0∈(0,1],t_{0}\in(0,1], there exist constants K>0K>0 and δ>0\delta>0 such that for all t∈[t0−δ,t0+δ]∩(0,1]t\in[t_{0}-\delta,t_{0}+\delta]\cap(0,1] and N⩾1N\geqslant 1 we have

Sending t↓0t\downarrow 0 in (17)–(20) also yields that lim⁡t↓0(ρt,qt,rˉt,rt)=(ρ0,q0,rˉ0,r0)\lim_{t\downarrow 0}(\rho_{t},q_{t},\bar{r}_{t},r_{t})=(\rho_{0},q_{0},\bar{r}_{0},r_{0}). Hence, ρt,qt,rˉt,rt\rho_{t},q_{t},\bar{r}_{t},r_{t} are continuous on $$.

where θt:=z~Δ∗+qt+ξρt−qt\theta_{t}:=\widetilde{z}\sqrt{\Delta_{*}+q_{t}}+\xi\sqrt{\rho_{t}-q_{t}} with i.i.d. standard gaussian z~,ξ\widetilde{z},\xi. Because the fixed point equations satisfy the critical point conditions

and ρt,qt,rˉt,rt\rho_{t},q_{t},\bar{r}_{t},r_{t} are locally Lipschitz on (0,1),(0,1), we conclude that df/dt=∂f/∂tdf/dt=\partial f/\partial t. To compute this partial derivative, let θt′:=z~Δ∗+qt+ξ′ρt−qt\theta_{t}^{\prime}:=\widetilde{z}\sqrt{\Delta_{*}+q_{t}}+\xi^{\prime}\sqrt{\rho_{t}-q_{t}} for ξ′\xi^{\prime} an independent copy of ξ\xi. Using gaussian integration by parts with respect to z~\widetilde{z} and ξ\xi yields

Proof of Theorem 3.4

The free energy and overlap for the ST model with a random hNh_{N} and fixed hh are asymptotically the same:

For 0⩽t⩽1,0\leqslant t\leqslant 1, consider the interpolated free energy,

Let HN,t(σ)H_{N,t}(\sigma) be the Hamiltonian defined by the above exponent. The associated Gibbs mean is denoted by ⟨ ⋅ ⟩t\langle\,\cdot\,\rangle_{t} and is defined by

From these, the mean value theorem and the Hölder inequality, there exists t∈t\in such that

The same proof as of Lemma 4.4 yields that for any k⩾1k\geqslant 1 there exists some K>0K>0 such that for all t∈t\in and N⩾1,N\geqslant 1,

Since the expectation of every term in the bracket is bounded by a universal constant, we obtain (62) by using (7) and taking k=1.k=1.

where KK is a constant independent of w,N,t.w,N,t. Therefore, letting w=(1,…,1)w=(1,\ldots,1) implies

we can apply (68) with the choices w=σ2w=\sigma^{2} and w=⟨σ1⟩tw=\langle\sigma^{1}\rangle_{t} to get

Putting these together and applying (67) with k=2k=2, we obtain that for all 0⩽t⩽10\leqslant t\leqslant 1 and N⩾1,N\geqslant 1,

for some constant K′K^{\prime} independent of t,N.t,N. Plugging (69) and (70) into (7) validates (63). ∎

Now we proceed to establish (22) and (23). From Theorem 3.3 and (15),

As for (22), from (LABEL:add:eq-5.1), we can write

and OO is an orthonormal matrix satisfying Ox∗=N−1/2∥x∗∥1N.Ox^{*}=N^{-1/2}{\|x^{*}\|}\bm{1}_{N}. Write

Here the first term on the right-hand side vanishes in L2L^{2} norm by the given assumption. For the second term, observe that conditionally on x∗x^{*}, FN(x∗)F_{N}(x^{*}) is the same as FN(hN)F_{N}(h_{N}) in distribution. Lemma 4.9 (with n=1n=1 and λ=η=0\lambda=\eta=0) allows us to write that with probability one,

where KK is independent of N.N. Hence, we can bound the second term by

We bound the third term in (71) by using the Jensen inequality,

W.-K.C. was partly supported by NSF DMS-17-52184. D.P. was partially supported by NSERC and Simons Fellowship.

References