The norm of polynomials in large random and deterministic matrices
C. Male
Introduction and statement of result
For a Hermitian matrix , let denote its empirical eigenvalue distribution, namely
where is the Dirac mass in and are the eigenvalues of . The empirical eigenvalue distribution of large dimensional random matrices has been studied with much interest for a long time. One pioneering result is Wigner’s theorem , from 1958. Let be an Wigner matrix. Then the theorem states that, under appropriate assumptions, the -th moment of converges in expectation to the -th moment of the semicircular law as goes to infinity for any integer . This result has been generalized in many directions, notably by Arnold for the almost sure convergence of the moments. The convergence of the empirical eigenvalue distribution for covariance matrices was first shown by Marc̆enko and Pastur in 1967, and has been generalized in the late 1970’s and the early 1980’s by many people, including Grenander and Silverstein , Wachter , Jonsson , Yin and Krishnaiah , Bai, Yin and Krishnaiah and Yin . In 1991, Voiculescu discovered a connection between large random matrices and free probability theory. He showed the so-called asymptotic freeness theorem, which has been generalized for instance in , which implies the almost sure weak convergence of the empirical eigenvalue distribution for Hermitian matrices of the form
is a fixed polynomial in non commutative indeterminates,
is a family of independent matrices of the normalized Gaussian Unitary Ensemble (GUE),
are matrices with appropriate assumptions (see Theorem 1.3 below).
Let be a family of independent, normalized GUE matrices and be a family of matrices, possibly random but independent of . Assume that for every Hermitian matrix of the form
where is a polynomial in non commutative indeterminates, we have with probability one that:
Convergence of the empirical eigenvalue distribution: there exists a compactly supported measure on the real line such that the empirical eigenvalue distribution of converges weakly to as goes to infinity.
Convergence of the spectrum: for any , almost surely there exists such that for all ,
Then almost surely the convergences of the empirical eigenvalue distribution and of the spectrum also hold for all Hermitian matrices , where is a polynomial in non commutative indeterminates.
Theorem 1.1 is a straightforward consequence of Theorem 1.6 below, where the language of free probability is used. Moreover, Theorem 1.6 specifies Theorem 1.1 by giving a description of the limit of the empirical eigenvalue distribution. For readers convenience, we recall some definitions (see and for details).
The elements of are called non commutative random variables. We will always assume that is a trace, i.e. that it satisfies for every . The trace is said to be faithful when it satisfies only if .
The non commutative law of a family of non commutative random variables is defined as the linear functional P\mapsto\tau\big{[}P(\mathbf{a},\mathbf{a}^{*})\ \big{]}, defined on the set of polynomials in non commutative indeterminates. The convergence in law is the pointwise convergence relative to this functional.
as soon as and \tau\big{[}P_{k}(\mathbf{a}_{i_{k}},\mathbf{a}_{i_{k}}^{*})\ \big{]}=0 for .
with the semicircle distribution.
Recall first the statement of Voiculescu’s asymptotic freeness theorem.
Let be a family of independent, normalized GUE matrices and be a family of matrices, possibly random but independent of . Let be a free semicircular system in a ∗-probability space and in be a family of non commutative random variables free from . Assume the following.
where denotes the normalized trace of matrices.
Boundedness of the spectrum: Almost surely, for one has
where denotes the operator norm.
In Haagerup and Thorbjørnsen strengthened the connection between random matrices and free probability. Limits of random matrices have now to be seen in more elaborated structure, called -probability space, which is endowed with a norm.
A -probability space consists of a ∗-probability space and a norm such that is a -algebra.
By the Gelfand-Naimark-Segal construction, one can always realize as a norm-closed -subalgebra of the algebra of bounded operators on a Hilbert space. Hence we can use functional calculus on . Moreover, if is a faithful trace, then the norm is uniquely determined by the following formula (see [28, Proposition 3.17]):
Let be independent, normalized GUE matrices and let be a free semicircular system in a -probability space with a faithful trace. Then almost surely, one has: for all polynomials in non commutative indeterminates, one has
This article is mainly devoted to the following theorem which is a generalization of Theorem 1.5 in the setting of Theorem 1.3.
Let be a family of independent, normalized GUE matrices and be a family of matrices, possibly random but independent of . Let and be a family of non commutative random variables in a -probability space with a faithful trace, such that is a free semicircular system free from . Assume the following. Strong convergence of : Almost surely, for all polynomials in non commutative indeterminates, one has
Then, almost surely, for all polynomials in non commutative indeterminates, one has
The convergence of the normalized traces stated in (1.13) is the content of Voiculescu’s asymptotic freeness theorem and is recalled in order to give a coherent and complete statement. Theorem 1.1 is easily deduced from Theorem 1.6 by applying Hamburger’s theorem for the convergence of the measure and functional calculus for the convergence of the spectrum. Organization of the paper: In Section 2 we give applications of Theorem 1.6 which are proved in Section 9. Sections 3 to 8 are dedicated to the proof of Theorem 1.6. Acknowledgments: The author would like to thank Alice Guionnet for dedicating much time for many discussions to the subjects of this paper and, along with Manjunath Krishnapur and Ofer Zeitouni, for the communication of Lemma 8.2. He is very much obliged to Dimitri Shlyakhtenko for his contribution to this paper. He would like to thank Benoit Collins for pointing out an error in a previous version of Corollary 2.1 and giving the idea to fix it. He also likes to thank Mikael de la Salle for useful discussions.
Applications
The first and the simpler matrix model that may be investigated to play the role of matrices in Theorem 1.6 consists of deterministic diagonal matrices with real entries and prescribed asymptotic spectral measure.
Let be a family of independent, normalized GUE matrices and let be deterministic real diagonal matrices, such that for any ,
the empirical spectral distribution of converges weakly to a compactly supported probability measure ,
the diagonal entries of are non decreasing:
for all , there exists such that for all , for all ,
Let in . We set \mathbf{D}_{N}^{v}=\big{(}D_{1}^{(N)}(v_{1}),\ldots,D_{q}^{(N)}(v_{q})\big{)}, where for any , one has
Let and \mathbf{d}^{v}=\big{(}d_{1}(v),\ldots,d_{q}(v)\big{)} be non commutative random variables in a -probability space with a faithful trace, such that
is a free semicircular system, free from ,
The variables commute, are selfadjoint and for all polynomials in indeterminates, one has
Then, with probability one, for all polynomials in non commutative indeterminates, one has
for any in except in a countable set.
Remark that the non commutative random variables can be realized as classical random variables, being -distributed for . The dependence between the random variables is trivial since Formula (2.1) exhibits a deterministic coupling. The convergence of the normalized trace (2.2) actually holds for any . In general, the convergence (2.3) of the norm can fail: the family of matrices where
gives a counterexample (consider their difference). Furthermore, let mention that it is clear that we always can take one of the to be zero.
2 Non-white Wishart matrices
Theorem 1.6 may be used to deduce the same result for some Wishart matrices as for the GUE matrices. Let be integers. Let be a family of independent positive definite Hermitian random matrices such that for the matrix is of size . Let be the family of matrices defined by: for each , , where is a matrix whose entries are random variables,
Let be a family of random matrices, independent of and . Assume that the families of matrices satisfy separately the assumptions of Theorem 1.6. Then, almost surely, for all polynomials in non commutative indeterminates, one has
where is given by Formula (1.9) with a faithful trace for which the non commutative random variables and are free.
In , motivated by applications in statistics and wireless communications, the authors study the global limiting behavior of the spectrum of the following matrix, referred as separable covariance matrix:
where is a random matrix, is a nonnegative definite square root of the nonnegative definite Hermitian matrix and is a diagonal matrix with nonnegative diagonal entries. It is shown in that, for large enough, almost surely the eigenvalues of belong in a small neighborhood of the limiting distribution under the following assumptions:
with .
The entries of are independent, identically distributed, standardized complex and with a finite fourth moment.
The empirical eigenvalue distribution (respectively ) of (respectively ) converges weakly to a compactly supported probability measure (respectively ) and the operator norms of and are uniformly bounded.
Now consider the following situation, where Corollary 2.2 may be applied
, for fixed positive integers and ,
the entries of are independent, identically distributed, standardized complex Gaussian,
the empirical eigenvalue distribution of (respectively ) converges weakly to a compactly supported probability measure,
for large enough, the eigenvalues of (respectively ) belong in a small neighborhood of its limiting distribution.
Then we obtain by Corollary 2.2 that for large enough, almost surely the eigenvalues of belong in a small neighborhood of the limiting distribution. The advantage of our version is the replacement of assumption 4 by assumption 4’. Replacing assumptions 1’ and 2’ by assumptions 1 and 2 could be an interesting question.
3 Block matrices
It will be shown as a consequence of Theorem 1.6 that the convergence of norms (1.14) also holds for block matrices.
4 Channel matrices
We give a potential application of Theorem 1.6 in the context of communication, where rectangular block random matrices are sometimes investigated for the study of wireless Multiple-input Multiple-Output (MIMO) systems . In the case of Intersymbol-Interference, the channel matrix reflects the channel effect during a transmission and is of the form
the family of matrices and the family of matrices satisfy separately the assumptions of Theorem 1.6,
the families of matrices , and are independent.
The strategy of proof
Let and be as in Theorem 1.6. We start with some remarks in order to simplify the proof.
We can suppose that the matrices of are Hermitian. Indeed for any , one has Re Im , where
It is sufficient to prove that for any polynomial the convergence of the norm in (1.14) holds almost surely (instead of almost surely the convergence holds for all polynomials). Indeed we can switch the words ”for all polynomials with rational coefficients“ and ”almost surely“ and both the left and the right hand side in (1.14) are continuous in .
where the trace is completely defined by:
is a free semicircular system,
is the limit in law of ,
Since is faithful on the ∗-algebra spanned by and , we can always assume that is a faithful trace on . Moreover, the matrices are uniformly bounded in operator norm. If we define in by Formula (1.9), then is finite for every . Hence, we can assume that is a -probability space endowed with the norm . Haagerup and Thorbjørnsen describe in a method to show that for all non commutative polynomials , almost surely one has
We present in this section this method with some modification to fit our situation. First, it is easy to see the following.
For all non commutative polynomials , almost surely one has
In a -algebra , one has , . Hence, without loss of generality, we can suppose that is non negative Hermitian and is selfadjoint. Let denote the empirical spectral distribution of :
It remains to show that the limsup is smaller than the right hand side in (3.3). The method is carried out in many steps.
the families , , , are free,
for any polynomials in non commutative indeterminates .
The intermediate object is therefore well defined as an element of . We use a theorem about norm convergence, due to D. Shlyakhtenko and stated in Appendix A, to relate the spectrum of with the spectrum of .
Step 2. An intermediate inclusion of spectrum: for all there exists such that for all , one has
for a constant and with denoting the operator norm.
Organization of the proof We tackle the different points of the proof described above in the following order:
Proof of Step 5. The asymptotic subordination property for random matrices is stated in Theorem 5.1 in a more general situation. The matrices can be random, independent of , satisfying a Poincaré inequality, without assumption on their asymptotic properties. This result is based on the Schwinger-Dyson equation and on the Poincaré inequality satisfied by the law of .
Proof of Estimate (3.9). The estimate will follow easily from the two previous items.
Proof of Step 3. The method is quite standard once Steps 2 and 4 are established. We use a version due to which is based on the use of local concentration inequalities.
Proof of Step 4: the subordination property for matrix-valued non commutative random variables
This lemma will be used throughout this paper. See [19, Lemma 3.1] for a proof.
The functional is well defined by Lemma 4.1 and satifies
The following proposition states the fundamental property of the amalgamated -transform, namely the subordination property, which is the keystone of our proof of Theorem 1.6.
Suppose that the families and are free. Then one has
Linearity property: There is a such that, in the domain , one has
Subordination property: There is such that, for every in , one has
Semicircular case: If is a free semicircular system, then we get
The linearity property has been shown by Voiculescu in and the -transform of has been computed by Lehner in . We deduce easily the subordination property since by Equation (4.3): there exists such that for all ,
Then there exists a such that, with instead of in the previous equality,
We compose by to obtain the result. ∎
Let and be as in Proposition 4.2, with a free semicircular system.
where .
Proof of Step 5: the asymptotic subordination property for random matrices
We denote by the functional
The proof of Theorem 5.1 is carried out in two steps.
In Section 5.1 we state a mean Schwinger-Dyson equation for random Stieltjes transforms (Proposition 5.2).
In Section 5.2 we deduce from Proposition 5.2 a Schwinger-Dyson equation for mean Stieltjes transforms (Proposition 5.3).
Theorem 5.1 is a direct consequence of Proposition 5.3 as it is shown in Section 5.3.
This induces an analogue formula for independent matrices of the GUE, called the Schwinger-Dyson equation, where the Hermitian symmetry of the matrices plays a key role. For instance, if is a monomial in non commutative indeterminates, one has for ,
the sum over all decompositions for and monomials being viewed as the partial derivative. This formula has an analogue for analytical maps instead of polynomials. The case of the function is investigated in details in [19, Formula (3.9)], our proof is obtained by minor modifications.
We take the partial trace in Equation (5.4) to obtain:
Re-injecting this expression in the left hand side of Equation (5.5), one gets Equation (5.3):
2 Schwinger-Dyson equation for mean Stieltjes transforms
is controlled in operator norm by the following estimate:
with c={k^{9/2}\sigma}\sum_{j=1}^{p}\|a_{j}\|^{2}\Big{(}\sum_{j=1}^{p}\left\|a_{j}\right\|+\sum_{j=1}^{q}\|b_{j}\|\Big{)}^{2}.
The Poincar inequality is compatible with tensor product and then such a formula is still valid when is a function of the matrices and with . We will often deal with matrices of size . Since the integer is fixed, we can use intensively the equivalence of norms, the constants appearing will not modify the order of convergence. For any integer , we denote the Euclidean norm of a matrix by
Recall that if are matrices we have the following inequalities
To estimate the operator norm of we use the domination by the infinity norm (5.15) in order to split the contributions due to and due to : we get
where we have denoted the matrices
Remark that by (5.16), for ,
Then by Cauchy-Schwarz inequality we get:
One is reduced to the study of variances of random variables. To use the Poincar inequality, we write for ,
The functions and their partial derivatives are bounded (see [19, Lemma 4.6] with minor modifications), so that, since the law of satisfies a Poincar inequality with constant , one has
We define the set of families of Hermitian matrices, with , , of unit Euclidean norm in . Then we have
For all in , for all selfadjoint matrices , :
The Cauchy-Schwarz inequality for (i.e. for ) gives
Using (5.16) to split Euclidean norms into the product of an operator norm and an Euclidean norm, we get:
Remark that, since , the norm of the matrices and is bounded by one. Then we have the following:
We then obtain as desired, by (5.18), (5.19) and (5.20):
where c={k^{9/2}\sigma}\sum_{j=1}^{p}\|a_{j}\|^{2}\Big{(}\sum_{j=1}^{p}\left\|a_{j}\right\|+\sum_{j=1}^{q}\|b_{j}\|\Big{)}^{2}.
3 Proof of Theorem 5.1
where \Theta_{N}(\Lambda)=\Theta_{N}\Big{(}\Lambda,\Lambda-\mathcal{R}_{s}\big{(}G_{S_{N}+T_{N}}(\Lambda)\ \big{)}\ \Big{)} is analytic in complex variables. Recall that by (4.7), we have \big{\|}\big{(}\Lambda-\mathcal{R}_{s}\big{(}G_{S_{N}+T_{N}}(\Lambda)\ \big{)}\ \big{)}^{-1}\big{\|}\leqslant\|(\Lambda)^{-1}\|, which gives (when replacing in (5.13) by ) the expected estimate of .
Proof of Estimate (3.9)
Let be as in Section 3. We assume that \big{(}\mathbf{x},\mathbf{y},(\mathbf{Y}_{N})_{N\geqslant 1}\big{)} are realized in a same -probability space with faithful trace, where
the families , , , are free,
for any polynomials in non commutative indeterminates .
On the other hand, since the matrices of are deterministic, we can apply Theorem 5.1 with
Then for , there exists such that for all and for any in , one has
Then by Proposition 4.3 with , one has
where denotes now the constant c={k^{9/2}}\sum_{j=1}^{p}\|a_{j}\|\Big{(}\sum_{j=1}^{p}\left\|a_{j}\right\|+\sum_{j=1}^{q}\|b_{j}\|\Big{)}^{2}\Big{(}\varepsilon^{-2}+2\sum_{j=1}^{p}\|a_{j}\|^{2}\Big{)}.
Proof of Step 2: An intermediate inclusion of spectrum
For a review on the theory of -algebras, we refer the readers to and . Notably, Appendix A of the second reference contains facts about ultrafilters and ultraproducts that are used in this section. Let \big{(}\mathbf{x},\mathbf{y},(\mathbf{Y}_{N})_{N\geqslant 1}\big{)} be as in Section 3. We assume that these non commutative random variables are realized in the same -probability space with faithful trace, where
the families , , , are free,
for any polynomials in non commutative indeterminates .
A consequence of Voiculescu’s theorem and of Shlyakhtenko’s Theorem A.1 in Appendix A is that for all polynomials in non commutative indeterminates,
Let and be unital -algebra. Let be a morphism of unital ∗-algebra. Then is contractive.
It is easy to see that for any in , the spectrum of is included in the spectrum of (since invertible implies that is also invertible). Hence we get that for all in
For the existence we consider the norm given by the spectral radius. The uniqueness follows from Lemma 7.1. ∎
Let be an integer. For all , let , respectively , be self-adjoint non commutative random variables in a - probability space , respectively . Assume that the traces and are faithful (hence the notation for the norms) and that for any polynomial in non commutative indeterminates,
The algebra is a -algebra whose norm is given by: for all in , equivalence class of
for all ultrafilter . Then the convergence holds when goes to infinity. ∎
The convergence extends to continuous function on the real line and then, with an appropriate choice of test functions, Step 2 follows. ∎
Proof of Step 3: from Stieltjes transforms to spectra
By Step 2, for all , there exists such that for all , one has
With (8.1) and (8.2) established, it is easy to show with minor modifications of [1, Lemma 5.5.5] the following result.
To get an almost sure control of , we use the fact that the entries of the matrices satisfy a concentration inequality.
With as in Lemma 8.1, there exists such that, almost surely
Then, for any smooth function , one has
where denotes the supremum of the considered function on the set of Hermitian matrices. Hence we get that . We fix a smooth function, non negative, compactly supported and vanishing on a neighborhood of the spectrum of . By the Tchebychev inequality
where we have used Lemma 8.1 ( also vanishes in a neighborhood of the spectrum of ). Moreover, since and are equals in and ,
Now, by (8.5) applied to : for all
By (8.10), (8.11), Lemma 8.1 and the Borel-Cantelli lemma, is almost surely of order at most. ∎
For every , there exists such that for
Proof of Corollaries 2.1, 2.2 and 2.4
converges weakly to . This sequence is tight, since there exists a such that for all , for all , one has . Hence it is sufficient to show the following: for all real numbers , for all , there exists such that
\mu_{j_{0}}\big{(}\ ]a_{j_{0}},a_{j_{0}}+\eta]\ \big{)}<\varepsilon/2.
for all , the real numbers and are points of continuity for .
By (9.3) with , there exists such that for all and , one has
But . Then we have
The are non decreasing, so we get
On the other hand, by (9.3) with and , there exists such that, for all , one has
But , so that
The are non decreasing, then we get
By (9.4) and (9.5) we obtain: for all
and then (9.2) is satisfied. So the convergence (9.1) holds when is zero. The convergence of traces, case in : To deduce the general case we shall need the following lemmas.
In particular, we have the convergence of the quantile of order :
Denote . Let be such that and and points of continuity for . Then, one has
Then, the being non decreasing, for any there exists such that for any , one has
Since is a point of continuity for , we get that . We chose . Then, we get . Hence, there exists such that, for any , one has i_{N}\geqslant\big{(}F(w-\eta)+\varepsilon\big{)}N and so, by (9.6): for any , there exists such that for all , one has . Hence, we get for all ,
and hence, letting go to zero, we obtain the expected result. ∎
Let denotes the cumulative distribution function of and its generalized inverse. We set , , and . Then, the empirical eigenvalue distribution of converges weakly the probability measure proportional to
We only show the lemma for , the general case can be deduce by adapting the reasoning. We then use, for conciseness, the symbols and instead of and respectively. If is not continuous in (i.e. if ) and , then for any in , the map is continuous in and . By Lemma 9.1, we get that
Hence, for any continuous function , we get
If is continuous in , we take in the following. We can always find , arbitrary small, such that is a point of continuity for . Remark that we then have
Moreover, we can always find in ]0,F^{-1}\big{(}F(w)+\beta\big{)}-w[, arbitrary small, such that is a point of continuity for and . Then, by (9.9), we get that, for large enough
Hence, for any continuous function , we get that for large enough
Letting go to zero, we get the result. ∎
Let in . We now show that, for any polynomial , one has
At the possible price of relabeling the matrices, we assume and set
For any , we decompose the matrices into
where for any , the matrix is . We set for any , the family . For any , we denote by the cumulative distribution function of the measure obtained in Lemma 9.2 with replaced by . Then, for any polynomial , one as
By Lemma 9.2 and by the case , we deduce that
with the convention . The merge of the different measures gives as expected
It is sufficient then to show that, for any , there exists such that for all , one has
and hence: for all , there exist and such that for all , for all
Suppose that (9.13) is not true: there exist and an increasing sequence of positive integer such that for all , there exists such that
By compactness, one can always assume that converges to in $j\{1,\ldots,q\}j_{0}u_{0}+v_{j}F^{-1}_{j}\lambda_{i_{k}+\lfloor v_{j}N_{k}\rfloor}^{(N_{k})}(j)F^{-1}_{j}(u_{0}+v_{j})$. Recall that
Then we have, for large enough and for all in $\big{|}\lambda_{i_{k}+\lfloor v_{j_{0}}N_{k}\rfloor}^{(N_{k})}(j_{0})-F_{j_{0}}^{-1}(u+v_{j_{0}})\big{|}>\eta$ i.e.
which is in contradiction with the fact that for large enough the eigenvalues of belong to a small neighborhood of the support of .
2 Proof of Corollary 2.2: Wishart matrices
Recall that by definition of the Wishart matrix model for
Let be a polynomial in non commutative indeterminates:
where the convergence holds almost surely since each term of the sum converges almost surely. ∎
For all polynomials in non commutative indeterminates, almost surely
, ,
,
, .
Define for the non commutative polynomial deduced by the formula
are free selfadjoint non commutative random variables,
is the limit in law of ,
For any polynomial in non commutative indeterminates
where the limits are almost sure. In particular we obtain that, for all polynomials in non commutative indeterminates, one has
Together with (9.43), this gives the expected result.
3 Proof of Corollary 2.4: Rectangular band matrices
We only give a sketch of the proof. Details are obtained by minor modification of the proofs of Corollaries 2.2 and 2.3. Let be as in Corollary 2.4:
We start with the following observation: the operator norm of is the square root of the operator norm of , which is a square block matrix. Its blocks consist of sums of matrices of the form , . By minor modifications of the proof of Corollary 2.2, we get the almost sure convergence of the normalized trace and of the norm for any polynomial in the matrices as goes to the infinity. By Proposition 7.3, we get that the convergences hold for square block matrices and in particular for any polynomial in . Hence the result follows by functional calculus.
Appendix A A theorem about norm convergence, by D. ShlyakhtenkoResearch supported by NSF grant DMS-0900776
Lemma Let be a -algebra with a faithful trace , and consider to be the universal -algebra generated by and elements satisfying for all . Moreover, consider the linear functional determined on by:
whenever and at least one of and is nonzero.
Consider the -Hilbert bimodule with the inner product
and the left and right actions given by
Let be the extended Cuntz-Pimsner algebra associated to (see ), i.e. the universal -algebra generated by and operators satisfying the relations
It follows from the results of that if we denote by the free product , then:
If is a finite tensor, write
From this we see that (by the universal property of ) there exists a -homomorphism , so that . Thus all we need to prove is that is injective. But by [30, Prop. 3.3], it follows that is isomorphic to the Toeplitz algebra (since in this case obviously ) acting on the Fock space . If we denote by the canonical conditional expectation from onto and consider the state , then the resulting Hilbert space is the closure of in the (faithful) norm ; from this we see that the GNS representation of associated to the state on is faithful. Since is exactly this GNS representation, it follows that is injective. ∎
Then is a -algebra.
Let now , , be self-adjoint random variables and assume that , are such that for any non-commutative polynomial ,
Let , be a family of free creation operators, free from each other and from . In other words, they satisfy:
and we denote by and the respective states on and (). We denote by and the respective states on and ().
We shall denote by the sequence . Then by assumption, we have that the map taking to extends to a state-preserving isomorphism from into with range .
We shall also denote by the constant sequence . Then for any element of represented by the sequence we have:
which (since the and operator norms coincide on multiples of identity) is equal to . It follows from the universality property that
Consider now a non-commutative -polynomial . Then
Since the left hand side does not depend on , we have proved:
Let , , be self-adjoint random variables and assume that , are such that for any non-commutative polynomial ,