One- versus multi-component regular variation and extremes of Markov trees

Johan Segers

Introduction

Imagine a random vector X=(X1,…,Xd)X=(X_{1},\ldots,X_{d}) of nonnegative variables. One of the components, say XiX_{i}, is known to have exceeded a large threshold. How does this information affect the conditional distribution of the whole vector XX? There could be a causal link from XiX_{i} to the other variables XjX_{j}, perhaps via a network of dependence relations, so that tampering with XiX_{i} would affect the whole system. Another possibility is that a large value of XiX_{i} is merely the result of a large value of some other variable XjX_{j}. The latter event, however, could have consequences for still other variables XkX_{k}.

Depending on which one of the dd components is known to have been exceptionally large, the conditional distribution of XX is likely to be different. Still, if high values of two variables XiX_{i} and XjX_{j} are not unlikely to arrive together, the conditional distribution of XX given that XiX_{i} is large must be connected to the one given that XjX_{j} is large.

In this paper, these questions are studied for general random vectors using the language of regular variation. The answers are worked out for the particular case that XX is a Markov tree. A large value at a particular node is found to spread through the tree via independent increments along the edges. The joint limit distribution is the one of a vector of coupled geometric random walks. The couplings occur through the common edges of different paths starting at the same root node.

Graphical models, of which Markov trees are a special case, bring structure and sparsity to the web of dependence relations between many random variables . Extreme value theory for such models is a fairly recent subject. In , a metric that takes the distance along a river into account underlies a spatial model for extremes of river networks. Recursive max-linear models on directed acyclic graphs are proposed in and put to work in . In , the density of a multivariate Pareto distribution is factorized through a version of the Hammersley–Clifford theorem. Such factorizations are also the theme in , where they form the basis of new inference methods for extremes of graphical models, including the identification of the graphical structure itself. Multivariate Hüsler–Reiss extreme-value copulas based on Gaussian Markov trees and higher-order truncated vines are introduced in , who propose composite likelihood methods based on bivariate margins to estimate the parameters.

Multivariate Pareto distributions arise as weak limits of normalized random vectors conditionally on the event that at least one component exceeds a high threshold. Although such conditioning events are covered by Theorem 3.9 below, the focus of this paper is rather on the case where the exceedance is known to have occurred at a specific variable. The message hinted at in the title is that both points of view are mathematically equivalent, but that, at least for Markov trees, the one-component limit is particularly elegant, as will be explained next.

For a Markov chain, it was discovered in that, conditionally on the event that the series is large at some time instant, the conditional distribution of the future of the system is that of a random walk, a process called tail chain in . For light-tailed marginal distributions, this random walk is additive, and for heavy-tailed margins it is geometric, i.e., multiplicative, which is the convention used in this paper.

A Markov tree can be viewed as a coupled collection of Markov chains with common stretches. Take for instance the four-variate Markov tree in Figure 1. The nodes of the tree are {1,2,3,4}\{1,2,3,4\} and the three pairs of neighbours are {1,2}\{1,2\}, {2,3}\{2,3\} and {2,4}\{2,4\}. The vector (X1,X2,X3)(X_{1},X_{2},X_{3}) is a Markov chain, and so is (X1,X2,X4)(X_{1},X_{2},X_{4}). These two chains are coupled via the common pair (X1,X2)(X_{1},X_{2}). Conditionally on X2X_{2}, the variables X1X_{1}, X3X_{3} and X4X_{4} are independent, since any path that connects two of the three nodes 1, 3 and 4 passes through node 2. This conditional independence property together with the distributions of the three pairs (X1,X2)(X_{1},X_{2}), (X2,X3)(X_{2},X_{3}) and (X2,X4)(X_{2},X_{4}) determines the joint distribution of (X1,X2,X3,X4)(X_{1},X_{2},X_{3},X_{4}).

For the moment, assume that the four variables have the same, regularly varying tail function. The set-up involving regular variation will be further motivated in Section 1.2. The effect (not necessarily causal) on X2X_{2} of a large value at X1X_{1} is via a multiplicative increment M1,2M_{1,2} whose distribution is equal to the weak limit of X2/X1X_{2}/X_{1} conditionally on X1=tX_{1}=t as t→∞t\to\infty. The existence of this limit is an assumption on the Markov kernel induced by the distribution of the pair (X1,X2)(X_{1},X_{2}). Similarly, a large value at X2X_{2} affects X3X_{3} and X4X_{4} via the increments M2,3M_{2,3} and M2,4M_{2,4}, respectively. The effect of X1X_{1} on X3X_{3} is then through the composite increment M1,2M2,3M_{1,2}M_{2,3}, whereas on X4X_{4} it is through M1,2M2,4M_{1,2}M_{2,4}. The conditional independence property ensures that the increments M1,2M_{1,2}, M2,3M_{2,3} and M2,4M_{2,4} are mutually independent. The common edge (1,2)(1,2) on the paths from node 11 to node 33 and from node 11 to node 44 induces dependence between the two tail chains (M1,2,M1,2M2,3)(M_{1,2},M_{1,2}M_{2,3}) and (M1,2,M1,2M2,4)(M_{1,2},M_{1,2}M_{2,4}) via the common increment M1,2M_{1,2}. In this paper, the random vector

is called the tail tree induced by XX with root at node u=1u=1.

The tail tree represents a network of stochastic dependence relations that are not necessarily causal. Suppose the Markov tree in Figure 1 represents water levels at four locations on a river network. If water flows from left to right, node 2 represents a point where the stream branches into two channels, as occurs for instance in a river delta. If water flows from right to left, however, node 2 represents the junction of two branches coming from nodes 3 and 4 into a larger stream flowing towards node 1. In the first case, the tail tree describes how a high water level at the upstream node 1 may cause high water levels at various locations in the delta further downstream. In the second, case however, it is nodes 3 and 4 that are situated upstream, and the tail tree models the sources of a high water volume at the downstream site 1. Still other set-ups are possible, such as for instance node 3 being upstream and nodes 1 and 4 being downstream: high water levels at nodes 1 and 4 are then related through a common cause at node 2, which can itself perhaps be traced back to node 3.

Whatever the causal relationships within XX, it may make sense to change the conditioning variable. In Figure 1, for instance, suppose it is known that a large value has occurred at node 3 rather than at node 1. Tracing the paths from node 3 to the three other nodes yields the tail tree with root at node u=3u=3:

The tail trees in (1) and (2) have a similar structure. The two edges on the path between the root nodes 1 and 3 have changed direction, however. The edge from node 2 to node 4 is common to both tail trees.

For each pair {a,b}\{a,b\} of neighbouring nodes, the choice of the root node uu determines which of the two increments appears in the tail tree: Ma,bM_{a,b} from XaX_{a} to XbX_{b} or Mb,aM_{b,a} from XbX_{b} to XaX_{a}. The distributions of Ma,bM_{a,b} and Mb,aM_{b,a} are connected by an expression that involves the marginal distributions of XaX_{a} and XbX_{b}. For stationary and reversible Markov chains, this relation underlies a sufficiency property discovered in . For tail chains of not necessarily reversible Markov chains, it was described in and for tail processes of regularly varying stationary time series in via the time change formula. This formula can be understood most easily through the connection between the tail process and the tail measure , and this is also the way in which the root change formula in Corollary 3.2 below will be derived, but then without the assumption of stationarity and for general random vectors, not necessarily Markov trees.

2 Regular variation

For multivariate distributions, regular variation can be described via multivariate cumulative distribution functions as well, but an approach via convergence of Borel measures is more versatile. Let the state space be \SS=[0,∞)d\SS=[0,\infty)^{d}. Generalizations to star-shaped metric spaces or abstract cones as in are left for further work. Let I⊂{1,…,d}I\subset\{1,\ldots,d\} denote the non-empty set of indices ii of variables of which the conditioning event Xi>tX_{i}>t is of possible interest. The marginal distributions of XiX_{i} for i∈Ii\in I are assumed to be regularly varying and the ratios of their tail functions are assumed to converge to positive constants. This set-up is a bit more general than the one of identical margins and comes at little technical or notational cost.

It is instructive to formulate statements in terms of weak convergence of distributions. For a high threshold tt tending to infinity and for a component i∈Ii\in I, consider the asymptotic distribution of the rescaled random vector X/tX/t given that Xi>tX_{i}>t. Decompose X/tX/t as (Xi/t,X/Xi)(X_{i}/t,X/X_{i}). Here, Xi/tX_{i}/t represents the overall level of XX with respect to tt whereas X/XiX/X_{i} represents a self-normalized version of XX. Convergence in distribution of (Xi/t,X/Xi)(X_{i}/t,X/X_{i}) given Xi>tX_{i}>t as t→∞t\to\infty is a special case of what is called one-component regular variation in , explored already in for the bivariate case but allowing for affine normalizations. The random variable Xi/tX_{i}/t is asymptotically Pa⁡(α)\operatorname{Pa}(\alpha) distributed and independent of X/XiX/X_{i}, whose weak limit, denoted by Θi=(Θi,j)j=1d\Theta_{i}=(\Theta_{i,j})_{j=1}^{d}, captures extremal dependence within XX given that XiX_{i} is large. Letting the index ii run through II produces multiple such one-component regular variation statements, which, together, are equivalent to what can be called multi-component regular variation. The limit distributions Θi\Theta_{i} that arise for various indices ii must be mutually consistent, and the tail measure mentioned at the end of the previous paragraph embraces them all at once.

In Section 3, the focus is on tying together multiple one-component regular variation limits. The theory is worked out for general random vectors, not necessarily Markov trees. A number of results in that section have already been formulated in the literature in one way or another, in slightly different settings. Some of the equivalence relations in Theorem 3.1, for instance, resemble those in [17, Theorem 1.4] and [34, Proposition 3.1]. The model consistency property between limit measures in Theorem 3.1(ii) is formulated in [6, Section 2] for the bivariate case. The root change formula in Corollary 3.4 extends the time change formula for regularly varying stationary time series stemming from and studied extensively in . Multivariate Pareto distributions as in Theorem 3.9 are foreshadowed in [29, Section 6.3] and appear in when ρ(x)=max⁡(x1,…,xd)\rho(x)=\max(x_{1},\ldots,x_{d}) and in for more general functionals ρ\rho. These are just a few connections, and the above list is by no means intended to be complete.

The set-up involving regular variation is intended to serve two purposes. First, to model tail dependence within a vector of random variables which have been transformed to the same, heavy-tailed distribution, such as the unit-Fréchet distribution, as is common in multivariate extreme value theory. Second, to model the joint distribution of a vector of regularly varying random variables, not necessarily identically distributed, but with equivalent tails, such as returns on financial portfolios composed of the same basket of underlying assets. The latter framework is more general than the former and comes at little additional notational cost.

3 Outline

For a Markov tree XX, convergence as t→∞t\to\infty of the conditional distribution of X/XuX/X_{u} given that Xu=tX_{u}=t is proved in Section 2. The main assumption is that, for edges e=(a,b)e=(a,b) directed away from the root uu, the conditional distribution of Xb/XaX_{b}/X_{a} given Xa=tX_{a}=t converges as t→∞t\to\infty. No regular variation is needed yet.

The tail trees pertaining to different roots uu can be linked up thanks to the theory of one- and multi-component regular variation developed in Section 3. The results do not rely on the Markov property and cover quite general random vectors XX on [0,∞)d[0,\infty)^{d}, as is illustrated briefly for max-linear models. An interesting special case of these are the recursive max-linear structural equation models introduced in , featuring a causal structure induced by a directed acyclic graph. Most of the proofs of this section are deferred to the Appendix.

When combined, the results in Section 2 and 3 serve to uncover the regular variation properties of Markov trees in Section 4. The common special case that the joint distribution of the Markov tree is absolutely continuous with respect to Lebesgue measure is the subject of Section 5. The theory then simplifies considerably and the limit distribution with respect to a single root uu is already sufficient to reconstruct the limit distributions with respect to all other possible roots uˉ\bar{u}.

In Sections 4 and 5, the distributions of the increments of the tail trees are calculated in case the pair distributions are max-stable, not necessarily absolutely continuous. For the Hüsler–Reiss distribution max-stable distribution, the tail tree is multivariate log-normal, constructed from partial sums of independent normal random variables along the edges of the tree.

The spectral tail tree of a Markov tree

A (finite) graph is a pair (V,E)(V,E) where VV is a non-empty finite set of vertices or nodes and where E⊂V×VE\subset V\times V is a set of edges. Self-loops are excluded, i.e., (u,u)∉E(u,u)\not\in E for all u∈Vu\in V. To avoid trivialities, VV is assumed to have at least two elements. Two nodes are neighbours if they are joined by an edge. A graph is undirected if (a,b)∈E(a,b)\in E implies (b,a)∈E(b,a)\in E. A path from a node uu to a node vv is a collection {e1,…,en}⊂E\{e_{1},\ldots,e_{n}\}\subset E of edges such that ek=(uk−1,uk)e_{k}=(u_{k-1},u_{k}) for all k=1,…,nk=1,\ldots,n, for n+1n+1 distinct nodes u0,u1,…,un∈Vu_{0},u_{1},\ldots,u_{n}\in V such that u0=uu_{0}=u and un=vu_{n}=v. An undirected tree T=(V,E)\mathcal{T}=(V,E) is an undirected graph such that for any pair of distinct nodes uu and vv, there exists a unique path from uu to vv, and this path is then denoted by [u⇝v][{u}\rightsquigarrow{v}].

Let T=(V,E)\mathcal{T}=(V,E) be an undirected tree and let X=(Xv)v∈VX=(X_{v})_{v\in V} be a random vector indexed by the nodes of the tree. The pair (X,T)(X,\mathcal{T}) is a Markov tree if it satisfies the global Markov property : whenever A,B,SA,B,S are disjoint, non-empty subsets of VV such that SS separates AA and BB (i.e., any path between a node a∈Aa\in A and a node b∈Bb\in B passes through some node in SS), the conditional independence relation

holds, where XWX_{W} denotes the random vector (Xv)v∈W(X_{v})_{v\in W} for W⊂VW\subset V.

For an undirected tree T=(V,E)\mathcal{T}=(V,E) and a node u∈Vu\in V, let Tu=(V,Eu)\mathcal{T}_{u}=(V,E_{u}) denote the directed, rooted tree that consists of directing the edges in EE outward starting from uu. Formally, EuE_{u} is the subset of EE that is obtained by choosing for every pair of edges (a,b)(a,b) and (b,a)(b,a) in EE the one such that the first node separates the second one from uu. If (a,b)∈Eu(a,b)\in E_{u}, then aa is the (necessarily unique) parent of bb in Tu\mathcal{T}_{u} whereas bb is a child of aa in Tu\mathcal{T}_{u}.

Let (X,T)(X,\mathcal{T}) be a nonnegative Markov tree, where T=(V,E)\mathcal{T}=(V,E) is an undirected tree.

There exists u∈Vu\in V with the following two properties.

For every directed edge e=(a,b)∈Eue=(a,b)\in E_{u}, there exists a version of the conditional distribution of XbX_{b} given XaX_{a} and a probability measure μe\mu_{e} on [0,∞)[0,\infty) such that

For edges e=(a,b)∈Eue=(a,b)\in E_{u} such that a≠ua\neq u and such that there exists an edge eˉ∈[u⇝a]\bar{e}\in[{u}\rightsquigarrow{a}] for which μeˉ({0})>0\mu_{\bar{e}}(\{0\})>0, we have

Assumption 2(ii) is similar to [26, equation (3.4)] and prevents non-extreme values to cause extreme ones. A similar assumption is [33, equation (2.4)], where it is illustrated [33, Example 7.5] what can go wrong without it.

Let (X,T)(X,\mathcal{T}) be a nonnegative Markov tree on T=(V,E)\mathcal{T}=(V,E). Assume Condition 2. Let (Me:e∈Eu)(M_{e}:e\in E_{u}) be a vector of independent random variables such that the law of MeM_{e} is μe\mu_{e} for all e∈Eue\in E_{u}. Then

The random vector (Θu,v)v∈V(\Theta_{u,v})_{v\in V} is called the tail tree of the Markov tree (Xv)v∈V(X_{v})_{v\in V}, adapting terminology for Markov chains in . In Figure 2, the tail tree is illustrated for a tree with seven nodes. For subvectors (Θu,w)w∈W(\Theta_{u,w})_{w\in W} where all nodes in WW lie on the same path starting at uu, the structure of the tail tree is that of a geometric random walk; take for instance u=1u=1 and W={1,4,5,7}W=\{1,4,5,7\} in Figure 2. The tail tree couples several geometric random walks together through the common edges in the underlying paths: in the same figure, consider for instance the vectors indexed by {1,4,5,7}\{1,4,5,7\} and by {1,4,6}\{1,4,6\}, respectively, which share the initial edge (1,4)(1,4).

Put d=∣V∣−1⩾1d=\lvert{V}\rvert-1\geqslant 1. The proof is by induction on dd.

If VV has only two elements, i.e., d=1d=1, then Condition 2(i) already confirms the convergence stated in (6) and (7). Therefore, we can henceforth assume that VV has at least three elements, i.e., d⩾2d\geqslant 2. Identify VV with {0,1,…,d}\{0,1,\ldots,d\} in such a way that the root is u=0u=0 and such that if (a,b)∈Eu(a,b)\in E_{u} then a<ba<b. Since X0/X0=1=Θ0,0X_{0}/X_{0}=1=\Theta_{0,0}, we do not need to consider the components X0X_{0} and Θ0,0\Theta_{0,0} in (6).

the joint distribution of Θ0,1:(d−1)=(Θ0,1,…,Θ0,d−1)\Theta_{0,1:(d-1)}=(\Theta_{0,1},\ldots,\Theta_{0,d-1}) being given by (7).

Recall that k∈{0,1,…,d−1}k\in\{0,1,\ldots,d-1\} denotes the parent node of dd. We need to distinguish between two cases: k=0k=0 is the root or k∈{1,…,d−1}k\in\{1,\ldots,d-1\} is a non-root vertex. The case k=0k=0 is similar to but easier than the case k∈{1,…,d−1}k\in\{1,\ldots,d-1\} and is left to the reader. We assume henceforth that k∈{1,…,d−1}k\in\{1,\ldots,d-1\}.

We will show that the expression (9) converges to zero as x0→∞x_{0}\to\infty (Step 3). Moreover, we will find a bound for the limit superior of (10) as x→∞x\to\infty. The bound will depend on δ\delta but will converge to zero as δ↓0\delta\downarrow 0 (Step 4). Together, these properties of (9) and (10) are sufficient to prove the theorem (Step 5).

Step 3: The term (9). — The vertex kk is the parent of dd in T0\mathcal{T}_{0}, and therefore it separates dd from the other vertices. By the conditional independence property (3),

To explain our notation: the integral is over x1:(d−1)=(x1,…,xd−1)x_{1:(d-1)}=(x_{1},\ldots,x_{d-1}) and is with respect to the conditional distribution of X1:(d−1)=(X1,…,Xd−1)X_{1:(d-1)}=(X_{1},\ldots,X_{d-1}) given that X0=tX_{0}=t. The integrand involves the conditional expectation of a function of XdX_{d} given that Xk=xkX_{k}=x_{k}.

We change variables and integrate with respect to the conditional distribution of X1:(d−1)/tX_{1:(d-1)}/t given that X0=tX_{0}=t: we get

By Assumption 2(i), we have X_{d}/x_{k}\mid X_{k}=x_{k}\raisebox{-0.5pt}{\,\scriptsize\stackrel{{\scriptstyle\raisebox{-0.5pt}{\mbox{\scriptsizedd}}}}{{\longrightarrow}}}\,M_{k,d} as xk→∞x_{k}\to\infty. Define

Recall that ff is bounded and (Lipschitz) continuous. By the extended continuous mapping theorem [37, Theorem 18.11], we have, for all vectors y1:(d−1)y_{1:(d-1)} such that yk≠δy_{k}\neq\delta and for all functions y1:(d−1)( ⋅ )y_{1:(d-1)}(\,\cdot\,) such that y1:(d−1)(t)→y1:(d−1)y_{1:(d-1)}(t)\to y_{1:(d-1)} as t→∞t\to\infty, the limit relation

Moreover, \mathcal{L}(X_{1:(d-1)}/t\mid X_{0}=t)\raisebox{-0.5pt}{\,\scriptsize\stackrel{{\scriptstyle\raisebox{-0.5pt}{\mbox{\scriptsizedd}}}}{{\longrightarrow}}}\,\mathcal{L}(\Theta_{1:(d-1)}) as t→∞t\to\infty by the induction hypothesis. By the same extended continuous mapping theorem, the integral (11) converges to

Recall that (Me)e∈Eu(M_{e})_{e\in E_{u}} is a vector of independent random variables such that the law of MeM_{e} is μe\mu_{e} for e∈Eue\in E_{u}. By construction, Mk,dM_{k,d} and Θ1:(d−1)\Theta_{1:(d-1)} are then independent too: each component Θ0,j\Theta_{0,j} of Θ1:(d−1)\Theta_{1:(d-1)} is a product of random variables Ma,bM_{a,b} with a,b∈{0,…,d−1}a,b\in\{0,\ldots,d-1\} and thus e=(a,b)≠(k,d)e=(a,b)\neq(k,d). The above integral may therefore be simplified to

since Θd=ΘkMk,d\Theta_{d}=\Theta_{k}M_{k,d}. It follows that the limit of (9) as t→∞t\to\infty is equal to zero.

Step 4.b.i: The term (12). — The term (12) is bounded by

The node kk separates the nodes and dd. By the global Markov property, the expectation on the right-hand side of (15) is therefore equal to

Let η∈(0,2/L)\eta\in(0,2/L). The conditional expectation in the integrand in (16) satisfies

Therefore, the integral in (16) is bounded by

By Assumption 2(ii), we can first take the limit superior as t→∞t\to\infty and then the limit superior as δ↓0\delta\downarrow 0 to find that

Since η\eta can be chosen arbitrarily close to zero, we find that the double limit superior above is equal to zero.

Step 4.b.ii: The term (13). — By the induction hypothesis, the term (13) converges to zero as t→∞t\to\infty.

Step 4.b.iii: The term (14). — Since Θd=ΘkMk,d\Theta_{d}=\Theta_{k}M_{k,d}, the term (14) is bounded by

By the dominated convergence theorem, the expectation on the right-hand side converges to zero as δ↓0\delta\downarrow 0.

This completes the proof of the induction step and thus of the theorem.

In the setting of Theorem 2.1, also \mathcal{L}(X/X_{u}\mid X_{u}>t)\raisebox{-0.5pt}{\,\scriptsize\stackrel{{\scriptstyle\raisebox{-0.5pt}{\mbox{\scriptsizedd}}}}{{\longrightarrow}}}\,\Theta_{u} as t→∞t\to\infty.

Given ε>0\varepsilon>0, Theorem 2.1 allows us to find t(ε)t(\varepsilon) sufficiently high such that the absolute value inside the integral is bounded by ε\varepsilon for all s⩾t(ε)s\geqslant t(\varepsilon). But then the left-hand side in the previous display is bounded by ε\varepsilon too, for all t⩾t(ε)t\geqslant t(\varepsilon). Since ε>0\varepsilon>0 was arbitrary, the stated convergence in distribution follows.

One- versus multi-component regular variation

Let X=(X1,…,Xd)X=(X_{1},\ldots,X_{d}) be a random vector of nonnegative variables. Upon an obvious change in notation, Corollary 2.3 concerned weak convergence of L(X/Xi∣Xi>t)\mathcal{L}(X/X_{i}\mid X_{i}>t) as t→∞t\to\infty for some i∈{1,…,d}i\in\{1,\ldots,d\}. This convergence plus regular variation of the marginal distribution of XiX_{i} is a special case of what is called one-component regular variation in . The weak limit, Θi=(Θi,j)j=1d\Theta_{i}=(\Theta_{i,j})_{j=1}^{d}, depends on the choice of ii. There may be good reasons to consider these limits for several indices ii. Let I⊂{1,…,d}I\subset\{1,\ldots,d\} be the set of all indices ii for which such a limit Θi\Theta_{i} exists. How are these random vectors Θi\Theta_{i} related?

In this section, several such one-component statements are combined into a single one which could be called multi-component regular variation. If I={1,…,d}I=\{1,\ldots,d\}, this is just ordinary multivariate regular variation. As discussed already in Section 1.2, the connections between the limits Θi\Theta_{i} generalize the time change formula for stationary regularly varying time series and can be deduced from their connections to a limiting tail measure.

For every i∈Ii\in I we have \mathcal{L}(X/X_{i}\mid X_{i}>t)\raisebox{-0.5pt}{\,\scriptsize\stackrel{{\scriptstyle\raisebox{-0.5pt}{\mbox{\scriptsizedd}}}}{{\longrightarrow}}}\,\mathcal{L}(\Theta_{i}) as t→∞t\to\infty for some random vector Θi=(Θi,j)j=1d\Theta_{i}=(\Theta_{i,j})_{j=1}^{d} on \SS\SS.

For every i∈Ii\in I we have \mathcal{L}(X_{i}/t,X/X_{i}\mid X_{i}>t)\raisebox{-0.5pt}{\,\scriptsize\stackrel{{\scriptstyle\raisebox{-0.5pt}{\mbox{\scriptsizedd}}}}{{\longrightarrow}}}\,\operatorname{Pa}(\alpha)\otimes\mathcal{L}(\Theta_{i}) as t→∞t\to\infty for some random vector Θi\Theta_{i} on \SS\SS.

For every i∈Ii\in I we have \mathcal{L}(X/t\mid X_{i}>t)\raisebox{-0.5pt}{\,\scriptsize\stackrel{{\scriptstyle\raisebox{-0.5pt}{\mbox{\scriptsizedd}}}}{{\longrightarrow}}}\,\mathcal{L}(Y_{i}) as t→∞t\to\infty for some random vector Yi=(Yi,j)j=1dY_{i}=(Y_{i,j})_{j=1}^{d} on \SS\SS.

In that case, the limiting objects are connected in the following ways: for all i∈Ii\in I,

YiY_{i} is equal in distribution to Yi,iΘiY_{i,i}\Theta_{i}, where L(Yi,i)=Pa⁡(α)\mathcal{L}(Y_{i,i})=\operatorname{Pa}(\alpha) and Yi,iY_{i,i} and Θi\Theta_{i} are independent;

νi\nu_{i} is equal to the restriction of ν\nu to \SS0,i\SS_{0,i};

for every Borel measurable f:\SS0,I→[0,∞]f:\SS_{0,I}\to[0,\infty], we have

The proof of Theorem 3.1, together with the proofs of the other theorems in this section, is given in Appendix A.

Apart from the characterizations (a)–(e) in Theorem 3.1, other equivalent ones are possible, for instance, involving sequences rather than functions, with a scaling function inside the probability rather than outside, or with respect to radial and ‘angular’ coordinates (ρ(X)/t,X/ρ(X))(\rho(X)/t,X/\rho(X)) for some appropriate functional ρ\rho. See for instance [29, Theorem 6.1] and [25, Theorem 3.1]. The tail measures ν\nu and νi\nu_{i} are homogeneous with index −α-\alpha [25, Theorem 3.1] and, upon a coordinate transformation, can be written as product measures. Since the focus here is on the weak limits Θi\Theta_{i}, these properties are not further elaborated upon. Statement (e) in Theorem 3.1 implies that the vector XI=(Xi)i∈IX_{I}=(X_{i})_{i\in I} is multivariate regularly varying with limit measure νI( ⋅ )=ν({x∈[0,∞)d:xI∈ ⋅ })\nu_{I}(\,\cdot\,)=\nu(\{x\in[0,\infty)^{d}:x_{I}\in\,\cdot\,\}) on [0,∞)I∖{0}[0,\infty)^{I}\setminus\{0\}, which in turn implies, among other things, that it is in the domain of attraction of a multivariate max-stable distribution with Fréchet margins and exponent measure νI\nu_{I}; see for instance .

A noteworthy special case of (17) is when ff is the indicator function of the orthant {x∈\SS0,I:∀j∈J, xj>yj}\{x\in\SS_{0,I}:\forall j\in J,\,x_{j}>y_{j}\}, where J⊂{1,…,d}J\subset\{1,\ldots,d\} has a non-empty intersection with II and where yj>0y_{j}>0 for all j∈Jj\in J. If i∈I∩Ji\in I\cap J, then f(x)\mathds1{xi>0}=f(x)f(x)\mathds{1}\{x_{i}>0\}=f(x), and thus

A remarkable consequence is that the right-hand side does not depend on the choice of i∈I∩Ji\in I\cap J. This invariance property is a special case of a more general mutual consistency property of the limit distributions L(Θi)\mathcal{L}(\Theta_{i}) for i∈Ii\in I that is formulated in Corollary 3.2 below.

If the equivalent conditions of Theorem 3.1 are fulfilled, then, for Borel measurable f:\SS0,I→[0,∞]f:\SS_{0,I}\to[0,\infty] and for i,j∈Ii,j\in I, we have

If the equivalent conditions of Theorem 3.1 are fulfilled, then, for all Borel measurable g:\SS0,I→[0,∞)g:\SS_{0,I}\to[0,\infty) and for all i,j∈Ii,j\in I, we have

If the limit measure ν\nu does not assign any mass to the coordinate hyperplane {x:xi=0}\{x:x_{i}=0\}, the indicator in (17) is redundant and ν\nu can be expressed entirely in terms of L(Θi)\mathcal{L}(\Theta_{i}). Moreover, whether this occurs or not can be read off from the α\alpha-th moments of the components of Θi\Theta_{i}.

If ν({x:xi=0})=0\nu(\{x:x_{i}=0\})=0 for some i∈Ii\in I, then, for all Borel measurable f:\SS0,I→[0,∞]f:\SS_{0,I}\to[0,\infty],

Moreover, all tail trees Θj\Theta_{j} for j∈Ij\in I are determined by Θi\Theta_{i} via (21).

Let ff be the indicator function of the set {x:∃j∈J, xj>yj}\{x:\exists j\in J,\,x_{j}>y_{j}\}, where J⊂{1,…,d}J\subset\{1,\ldots,d\} is non-empty and where yj>0y_{j}>0 for all j∈Jj\in J. If ν({x:xi=0})=0\nu(\{x:x_{i}=0\})=0, then, by (23),

In contrast to equation (18), equation (24) is true only when ν({x:xi=0})=0\nu(\{x:x_{i}=0\})=0, a prerequisite for which Corollary 3.6 gives a necessary and sufficient condition.

In Theorem 3.1, a sufficient condition for (a)–(e) to hold is that there exists a non-empty set K⊂IK\subset I with the following two properties:

For every i∈Ki\in K we have \mathcal{L}(X/X_{i}\mid X_{i}>t)\raisebox{-0.5pt}{\,\scriptsize\stackrel{{\scriptstyle\raisebox{-0.5pt}{\mbox{\scriptsizedd}}}}{{\longrightarrow}}}\,\mathcal{L}(\Theta_{i}) as t→∞t\to\infty for some random vector Θi\Theta_{i} on \SS\SS.

In that case, also \mathcal{L}(X/X_{j}\mid X_{j}>t)\raisebox{-0.5pt}{\,\scriptsize\stackrel{{\scriptstyle\raisebox{-0.5pt}{\mbox{\scriptsizedd}}}}{{\longrightarrow}}}\,\mathcal{L}(\Theta_{j}) as t→∞t\to\infty for j∈I∖Kj\in I\setminus K, where the law of Θj\Theta_{j} is given in terms of the one of Θi\Theta_{i} with i=i(j)i=i(j) via (21).

The focus so far has been on weak limits of conditional distributions involving a high-threshold exceedance by a specific component. In the spirit of the multivariate peaks-over-thresholds methodology , the following result covers, among other possibilities, the case where the conditioning event involves a high-threshold exceedance in at least one of a number of components.

Let X=(X1,…,Xd)X=(X_{1},\ldots,X_{d}) follow the max-linear model

where ai,r∈[0,∞)a_{i,r}\in[0,\infty) are scalars such that max⁡rai,r>0\max_{r}a_{i,r}>0 for all ii and where Z1,…,ZsZ_{1},\ldots,Z_{s} are independent and identically distributed nonnegative random variables whose common distribution function FF has a regularly varying tail function F‾=1−F\overline{F}=1-F with index −α<0-\alpha<0. The marginal tails satisfy F‾i(t)/F‾(t)→∑rai,rα=:ci\overline{F}_{i}(t)/\overline{F}(t)\to\sum_{r}a_{i,r}^{\alpha}=:c_{i} as t→∞t\to\infty. If XiX_{i} exceeds a large threshold t→∞t\to\infty, the probability that this was due to ZrZ_{r} is proportional to ai,rαa_{i,r}^{\alpha}, and then the other factors ZrˉZ_{\bar{r}} for rˉ≠r\bar{r}\neq r are of smaller order than ZrZ_{r}. It follows that (a) in Theorem 3.1 holds where the law of Θi\Theta_{i} is discrete with at most ss atoms and is given by

with ϵx( ⋅ )\epsilon_{x}(\,\cdot\,) denoting a unit point mass at xx. From (26), we find

Recursive max-linear models on directed acyclic graphs were introduced in . Borrowing some of their notation, consider a directed acyclic graph D=(V,E)\mathcal{D}=(V,E) with nodes V={1,…,d}V=\{1,\ldots,d\} and edges E={(k,i):i∈V,k∈pa⁡(i)}E=\{(k,i):i\in V,k\in\operatorname{pa}(i)\}, where pa⁡(i)⊂V\operatorname{pa}(i)\subset V denotes the possibly empty set of parents of ii. Consider a random vector X=(X1,…,Xd)X=(X_{1},\ldots,X_{d}) given by the structural equation model

where the random variables Z1,…,ZdZ_{1},\ldots,Z_{d} are as in Example 3.10 with s=ds=d and where all coefficients γki\gamma_{ki} and γii\gamma_{ii} are (strictly) positive; the maximum over the empty set is zero by convention. Then by [14, Theorem 2.2], the random vector XX admits the max-linear representation

with coefficients bjib_{ji}, for i,j∈{1,…,d}i,j\in\{1,\ldots,d\}, defined as follows: bii=γiib_{ii}=\gamma_{ii} and bji=0b_{ji}=0 if j∈V∖(an⁡(i)∪{i})j\in V\setminus(\operatorname{an}(i)\cup\{i\}), while

where PjiP_{ji} is the collection of paths p={e1,…,en}p=\{e_{1},\ldots,e_{n}\} from jj to ii in D\mathcal{D}; recall the definition of a path in the beginning of Section 2. The representation in (28) is of the form (25) with s=ds=d and with ai,r=bria_{i,r}=b_{ri} for i,r∈{1,…,d}i,r\in\{1,\ldots,d\}. It follows that ai,r=0a_{i,r}=0 unless r=ir=i or r∈an⁡(i)r\in\operatorname{an}(i).

In the special case that the directed acyclic graph is also a directed, rooted tree, every node ii has either exactly one parent or is equal to the root node, say uu. In that case, the collection of paths PjiP_{ji} between j∈an⁡(i)j\in\operatorname{an}(i) and i∈{1,…,d}∖{u}i\in\{1,\ldots,d\}\setminus\{u\} is a singleton, p=[j⇝i]p=[{j}\rightsquigarrow{i}], and the formula for bjib_{ji} in (28) simplifies to bji=γjj∏e∈[j⇝i]γeb_{ji}=\gamma_{jj}\prod_{e\in[{j}\rightsquigarrow{i}]}\gamma_{e}. Furthermore, the tail tree Θu\Theta_{u} in (26) starting from the root node uu simplifies to the degenerate distribution at the point θu=(θu,1,…,θu,d)\theta_{u}=(\theta_{u,1},\ldots,\theta_{u,d}) with coordinates θu,j=∏e∈[u⇝j]γe∈(0,∞)\theta_{u,j}=\prod_{e\in[{u}\rightsquigarrow{j}]}\gamma_{e}\in(0,\infty) for j∈{1,…,d}j\in\{1,\ldots,d\}. This is of the form in (7) with degenerate increments Me=γeM_{e}=\gamma_{e} for all e∈Ee\in E.

Regularly varying Markov trees

As in Section 2, let (X,T)(X,\mathcal{T}) be a nonnegative Markov tree on the undirected tree T=(V,E)\mathcal{T}=(V,E). The general theory in Section 3 sheds light on the relation between two tail trees emanating at different roots. For two different nodes uu and uˉ\bar{u} in VV, the sets of directed edges EuE_{u} and EuˉE_{\bar{u}} are the same except for the edges connecting nodes on the path between uu and uˉ\bar{u}, which are directed in opposite ways in the two edge sets: For every (a,b)∈[u⇝uˉ]=Eu∖Euˉ(a,b)\in[{u}\rightsquigarrow{\bar{u}}]=E_{u}\setminus E_{\bar{u}}, we have (b,a)∈[uˉ⇝u]=Euˉ∖Eu(b,a)\in[{\bar{u}}\rightsquigarrow{u}]=E_{\bar{u}}\setminus E_{u} and the other way around.

Condition 2 was formulated relative to a single root u∈Vu\in V. The next condition covers all nodes u∈Uu\in U in a non-empty subset UU of VV as possible roots. For such UU, let EU=⋃u∈UEuE_{U}=\bigcup_{u\in U}E_{u} denote the set of directed edges that appear in at least one of the directed trees EuE_{u}.

There exists a non-empty U⊂VU\subset V with the following two properties:

For every e=(a,b)∈EUe=(a,b)\in E_{U}, there exists a version of the conditional distribution of XbX_{b} given XaX_{a} and a probability measure μe\mu_{e} on [0,∞)[0,\infty) such that (4) holds.

For every edge e=(a,b)∈EUe=(a,b)\in E_{U} for which there exists u∈Uu\in U such that e∈Eue\in E_{u} and an edge eˉ∈[u⇝a]\bar{e}\in[{u}\rightsquigarrow{a}] such that μeˉ({0})>0\mu_{\bar{e}}(\{0\})>0, we have (5).

If u,v∈Uu,v\in U in Condition 4 then every node w∈Vw\in V that is on the path between uu and vv can be added to UU and Condition 4 remains true. Indeed, for such u,v,wu,v,w, we have Ew⊂Eu∪EvE_{w}\subset E_{u}\cup E_{v}, which takes care of (i), and [w⇝a]⊂[u⇝a]∪[v⇝a][{w}\rightsquigarrow{a}]\subset[{u}\rightsquigarrow{a}]\cup[{v}\rightsquigarrow{a}] for every node a∈Va\in V, which takes care of (ii). The author is grateful to an anonymous reviewer for having pointed this out.

Condition 4 and Corollary 2.3 imply that assumption (a) in Theorem 3.1 is satisfied for I=UI=U and with Θu\Theta_{u} the tail tree in (7), for every u∈Uu\in U. All equivalence relations and other properties are then as stated in Theorem 3.1.

In Corollary 4.1, if a,b∈Va,b\in V are neighbours in EE and if they both belong to UU, then the distributions of Ma,bM_{a,b} and Mb,aM_{b,a} mutually determine each other by

for all Borel measurable g:(0,∞)→[0,∞]g:(0,\infty)\to[0,\infty].

For different roots u,uˉ∈Uu,\bar{u}\in U, the tail trees Θu\Theta_{u} and Θuˉ\Theta_{\bar{u}} have the same multiplicative structure. The differences between their distributions lie in the starting nodes of the paths and in the distributions of the multiplicative increments for edges on the paths [u⇝uˉ][{u}\rightsquigarrow{\bar{u}}] and [uˉ⇝u][{\bar{u}}\rightsquigarrow{u}], since these edges change direction. For such edges of which the nodes belong to UU as well, the increment distributions are related by (29). See Figure 3 for an illustration.

Given the tree structure, the distribution of a Markov tree XX on T=(V,E)\mathcal{T}=(V,E) is entirely determined by the bivariate distributions (Xa,Xb)(X_{a},X_{b}) for e=(a,b)∈Ee=(a,b)\in E. Markov chains of which all pairs (Xi,Xi+1)(X_{i},X_{i+1}) are max-stable were proposed in [5, Section 4.6] and . When extended to trees, this construction method provides models meeting Condition 4.

Let the distribution of the random pair (X,Y)(X,Y) on (0,∞)2(0,\infty)^{2} be bivariate max-stable with cumulative distribution function

where A:→[1/2,1]A:\to[1/2,1] is a Pickands dependence function, that is, a convex function such that max⁡(w,1−w)⩽A(w)⩽1\max(w,1-w)\leqslant A(w)\leqslant 1 for all w∈w\in; see and the references therein. Both marginal distributions are unit-Fréchet, F(z,∞)=F(∞,z)=exp⁡(−1/z)F(z,\infty)=F(\infty,z)=\exp(-1/z) for z∈(0,∞)z\in(0,\infty). In particular, the marginal tail functions are regularly varying at infinity with index −α=−1-\alpha=-1.

Let A′A^{\prime} be the left-hand derivative of AA, which exists everywhere on (0,1](0,1], takes values between −1-1 and 11, and is non-decreasing and continuous from the left; define A′(0)A^{\prime}(0) as the right-hand limit. Since AA is convex, it is absolutely continuous, and the set of points in (0,1)(0,1) where it is not continuously differentiable is at most countable. For x,y∈(0,∞)x,y\in(0,\infty) such that AA is differentiable at w=x/(x+y)w=x/(x+y), we have

It follows that \mathcal{L}(Y/x\mid X=x)\raisebox{-0.5pt}{\,\scriptsize\stackrel{{\scriptstyle\raisebox{-0.5pt}{\mbox{\scriptsizedd}}}}{{\longrightarrow}}}\,M as x→∞x\to\infty, where

Absolutely continuous case

If the joint distribution of the Markov tree X=(Xv)v∈VX=(X_{v})_{v\in V} on T=(V,E)\mathcal{T}=(V,E) is absolutely continuous with respect to the Lebesgue measure on [0,∞)V[0,\infty)^{V}, the formulations of the conditions and results simplify considerably. Let ff denote the joint probability density function of XX and let fvf_{v}, for v∈Vv\in V, denote the marginal density of XvX_{v}.

By the Hammersley–Clifford theorem [23, Theorem 3.9], XX is a Markov tree as soon as the joint density factorizes as

The second product is over all unordered pairs of neighbours and fa,bf_{a,b} denotes the bivariate density function of (Xa,Xb)(X_{a},X_{b}).

For t∈(0,∞)t\in(0,\infty) such that fa(t)∈(0,∞)f_{a}(t)\in(0,\infty), the density of L(Xb/t∣Xa=t)\mathcal{L}(X_{b}/t\mid X_{a}=t) is tfa,b(t,ty)/fa(t)tf_{a,b}(t,ty)/f_{a}(t) for y∈(0,∞)y\in(0,\infty). The following condition replaces Condition 4.

For every e=(a,b)∈Ee=(a,b)\in E, there exists a probability density function qa,bq_{a,b} on (0,∞)(0,\infty) such that

Let the random vector XX on [0,∞)V[0,\infty)^{V} be a Markov tree on the undirected tree T=(V,E)\mathcal{T}=(V,E) with joint density function ff. Assume there exists a positive function gg, regularly varying at infinity with index −α−1<−1-\alpha-1<-1, such that fv(t)/g(t)→cv∈(0,∞)f_{v}(t)/g(t)\to c_{v}\in(0,\infty) as t→∞t\to\infty for every v∈(0,∞)v\in(0,\infty). If Condition 5 holds, then the conditions of Corollary 4.1 are satisfied with U=VU=V, the same constants cuc_{u}, and auxiliary function b(t)=α/{t g(t)}b(t)=\alpha/\{t\,g(t)\}. For all pairs of neighbours a,b∈Va,b\in V, the density of Ma,bM_{a,b} is qa,bq_{a,b} and for almost every y∈(0,∞)y\in(0,\infty), we have

for every u∈Vu\in V and for every Borel measurable f:[0,∞)V∖{0}→[0,∞]f:[0,\infty)^{V}\setminus\{0\}\to[0,\infty], with Θu\Theta_{u} the tail tree in (7). Moreover, all tail trees are connected through (21).

The function fvf_{v} is regularly varying at infinity with index −α−1-\alpha-1 too. By Karamata’s theorem [3, Proposition 1.5.10], we have tfv(t)/F‾v(t)→αtf_{v}(t)/\overline{F}_{v}(t)\to\alpha and thus F‾v(t)/{tg(t)}→cv/α\overline{F}_{v}(t)/\{tg(t)\}\to c_{v}/\alpha as t→∞t\to\infty.

Condition 4 with U=VU=V follows from Condition 5 and Scheffé’s theorem. Part (ii) of Condition 4 is void, since μe({0})=0\mu_{e}(\{0\})=0 for every e∈Ee\in E.

If a,b∈Va,b\in V are neighbours, we can apply (29) to g(y)=\mathds1(0,z)(y)g(y)=\mathds{1}_{(0,z)}(y), where z∈(0,∞)z\in(0,\infty), to find

Since this is true for every z∈(0,∞)z\in(0,\infty), we must have cb qb,a(y)=ca y−α−2 qa,b(y−1)c_{b}\,q_{b,a}(y)=c_{a}\,y^{-\alpha-2}\,q_{a,b}(y^{-1}) for almost every y∈(0,∞)y\in(0,\infty), whence (32).

Finally, M0\mathcal{M}_{0}-convergence to ν\nu with the stated expression follows from Theorem 3.1 and Corollary 3.6.

In Example 4.5, assume that AA is twice continuously differentiable on (0,1)(0,1) and that A′(0)=−1A^{\prime}(0)=-1 and A′(1)=1A^{\prime}(1)=1. The distribution of (X,Y)(X,Y) is then absolutely continuous and the conditional density of Y/xY/x given that X=xX=x converges as x→∞x\to\infty to the function

An interesting example in this respect is the bivariate Hüsler–Reiss distribution with Pickands dependence function

If all neighbouring pairs (Xa,Xb)(X_{a},X_{b}) for (a,b)∈E(a,b)\in E of the Markov tree follow such Hüsler–Reiss max-stable distributions, the joint distribution of the tail tree is multivariate log-normal, since log⁡Θu,v=∑e∈[u⇝v]log⁡Me\log\Theta_{u,v}=\sum_{e\in[{u}\rightsquigarrow{v}]}\log M_{e} for all u,v∈Vu,v\in V, where the random variables log⁡Me\log M_{e} are independent and normally distributed with expectation −2λe2-2\lambda_{e}^{2} and variance 4λe24\lambda_{e}^{2}, with dependence parameter λe∈(0,∞)\lambda_{e}\in(0,\infty) for all e∈Ee\in E.

Appendix A Proofs for Section 3

(b) implies (c) and (i). — Since (X/t)=(Xi/t)(X/Xi)(X/t)=(X_{i}/t)(X/X_{i}), statement (b) and the continuous mapping theorem [37, Theorem 2.3] imply that L(X/t∣Xi>t)\mathcal{L}(X/t\mid X_{i}>t) converges weakly to Yi=ZΘiY_{i}=Z\Theta_{i}, where ZZ is a Pa⁡(α)\operatorname{Pa}(\alpha) random variable independent of Θi\Theta_{i}. Since Θi,i=1\Theta_{i,i}=1 almost surely, we have Yi,i=ZY_{i,i}=Z.

(c) implies (a). — Since X/Xi=(X/t)/(Xi/t)X/X_{i}=(X/t)/(X_{i}/t), statement (c) and the continuous mapping theorem imply statement (a) with Θi=Yi/Yi,i\Theta_{i}=Y_{i}/Y_{i,i}.

(b) implies (d). — Define a Borel measure νi\nu_{i} on \SS0,i\SS_{0,i} by

By linearity of the integral and by monotone convergence, we find that

for every nonnegative Borel measurable function ff on \SS0,i\SS_{0,i}. The same expression is then true for real-valued Borel measurable functions ff on \SS0,i\SS_{0,i} for which at least one of the two integrals with ff replaced by ∣f∣\lvert{f}\rvert is finite. This includes bounded, Borel measurable functions that vanish on a set of the form {x∈\SS0,i:xi⩽ε}\{x\in\SS_{0,i}:x_{i}\leqslant\varepsilon\} for some ε>0\varepsilon>0.

Let f∈C0,if\in\mathcal{C}_{0,i} and let ε>0\varepsilon>0 be such that f(x)=0f(x)=0 as soon as xi⩽εx_{i}\leqslant\varepsilon. By (b), we have

where ZZ is a Pa⁡(α)\operatorname{Pa}(\alpha) random variable, independent of Θi\Theta_{i}. The limit is equal to

since f(zΘi)=0f(z\Theta_{i})=0 almost surely whenever z⩽εz\leqslant\varepsilon, as Θi,i=1\Theta_{i,i}=1 almost surely.

(d) implies (e). — For every z>0z>0, we have

For ε>0\varepsilon>0, let hε:[0,∞)→h_{\varepsilon}:[0,\infty)\to be the piece-wise linear function

Put ℏε=1−hε\hbar_{\varepsilon}=1-h_{\varepsilon}. Write I={i1,…,ik}I=\{i_{1},\ldots,i_{k}\}. Then

Each function fif_{i} belongs to C0,I\mathcal{C}_{0,I} too but has moreover the property that fi(x)=0f_{i}(x)=0 as soon as xi⩽ε/2x_{i}\leqslant\varepsilon/2. The restriction of fif_{i} to \SS0,i\SS_{0,i} thus belongs to C0,i\mathcal{C}_{0,i}. By (d),

The existence of a limit has thus been shown, and convergence in M0,I\mathcal{M}_{0,I} to some measure ν\nu as stated in (e) follows.

(e) implies (d), (ii), (iii) and (iv). — A function ff in C0,i\mathcal{C}_{0,i} can be extended to a function in C0,I\mathcal{C}_{0,I} denoted by the same symbol by putting f(x)=0f(x)=0 for x∈\SS0,I∖\SS0,ix\in\SS_{0,I}\setminus\SS_{0,i}. Hence, (e) implies (d), with νi\nu_{i} as described in (ii).

Statement (iii) follows from (ii) and the description of the law of YiY_{i} in terms of νi\nu_{i} in the proof above of the implication that (d) implies (c).

Similarly, (iv) follows from (ii), equation (33), and Fubini’s theorem.

It is sufficient to show statement (a) in Theorem 3.1. By property (i) in Theorem 3.8, the weak convergence in Theorem 3.1(a) already holds for all i∈Ki\in K, and we need to show that it also holds for all j∈I∖Kj\in I\setminus K. Choose j∈I∖Kj\in I\setminus K and let i=i(j)∈Ki=i(j)\in K be as in property (ii) of Theorem 3.8

We will show that L(X/Xj∣Xj>t)\mathcal{L}(X/X_{j}\mid X_{j}>t) converges weakly as t→∞t\to\infty to Θj\Theta_{j} whose law is defined in (21). Let G⊂\SSG\subset\SS be open and let δ>0\delta>0. We have

By Theorem 3.1 applied to KK, we have \mathcal{L}(X_{i}/s,X/X_{i}\mid X_{i}>s)\raisebox{-0.5pt}{\,\scriptsize\stackrel{{\scriptstyle\raisebox{-0.5pt}{\mbox{\scriptsizedd}}}}{{\longrightarrow}}}\,\operatorname{Pa}(\alpha)\otimes\mathcal{L}(\Theta_{i}) as s→∞s\to\infty. Let ZZ be a Pa⁡(α)\operatorname{Pa}(\alpha) random variable, independent of Θi\Theta_{i}. By the Portmanteau lemma for weak convergence, we have

The Portmanteau theorem for weak convergence implies the stated weak convergence of L(X/t∣ρ(X)>t)\mathcal{L}(X/t\mid\rho(X)>t) as t→∞t\to\infty. This proves statement (c) in Theorem 3.1 for the enlarged random vector (X,ρ(X))(X,\rho(X)).

The author is grateful to two anonymous reviewers whose suggestions have led to various improvements throughout the text. The author also wishes to thank Stefka Asenova and Gildas Mazo for inspiring discussions.

References