One- versus multi-component regular variation and extremes of Markov trees
Johan Segers
Introduction
Imagine a random vector of nonnegative variables. One of the components, say , is known to have exceeded a large threshold. How does this information affect the conditional distribution of the whole vector ? There could be a causal link from to the other variables , perhaps via a network of dependence relations, so that tampering with would affect the whole system. Another possibility is that a large value of is merely the result of a large value of some other variable . The latter event, however, could have consequences for still other variables .
Depending on which one of the components is known to have been exceptionally large, the conditional distribution of is likely to be different. Still, if high values of two variables and are not unlikely to arrive together, the conditional distribution of given that is large must be connected to the one given that is large.
In this paper, these questions are studied for general random vectors using the language of regular variation. The answers are worked out for the particular case that is a Markov tree. A large value at a particular node is found to spread through the tree via independent increments along the edges. The joint limit distribution is the one of a vector of coupled geometric random walks. The couplings occur through the common edges of different paths starting at the same root node.
Graphical models, of which Markov trees are a special case, bring structure and sparsity to the web of dependence relations between many random variables . Extreme value theory for such models is a fairly recent subject. In , a metric that takes the distance along a river into account underlies a spatial model for extremes of river networks. Recursive max-linear models on directed acyclic graphs are proposed in and put to work in . In , the density of a multivariate Pareto distribution is factorized through a version of the Hammersley–Clifford theorem. Such factorizations are also the theme in , where they form the basis of new inference methods for extremes of graphical models, including the identification of the graphical structure itself. Multivariate Hüsler–Reiss extreme-value copulas based on Gaussian Markov trees and higher-order truncated vines are introduced in , who propose composite likelihood methods based on bivariate margins to estimate the parameters.
Multivariate Pareto distributions arise as weak limits of normalized random vectors conditionally on the event that at least one component exceeds a high threshold. Although such conditioning events are covered by Theorem 3.9 below, the focus of this paper is rather on the case where the exceedance is known to have occurred at a specific variable. The message hinted at in the title is that both points of view are mathematically equivalent, but that, at least for Markov trees, the one-component limit is particularly elegant, as will be explained next.
For a Markov chain, it was discovered in that, conditionally on the event that the series is large at some time instant, the conditional distribution of the future of the system is that of a random walk, a process called tail chain in . For light-tailed marginal distributions, this random walk is additive, and for heavy-tailed margins it is geometric, i.e., multiplicative, which is the convention used in this paper.
A Markov tree can be viewed as a coupled collection of Markov chains with common stretches. Take for instance the four-variate Markov tree in Figure 1. The nodes of the tree are and the three pairs of neighbours are , and . The vector is a Markov chain, and so is . These two chains are coupled via the common pair . Conditionally on , the variables , and are independent, since any path that connects two of the three nodes 1, 3 and 4 passes through node 2. This conditional independence property together with the distributions of the three pairs , and determines the joint distribution of .
For the moment, assume that the four variables have the same, regularly varying tail function. The set-up involving regular variation will be further motivated in Section 1.2. The effect (not necessarily causal) on of a large value at is via a multiplicative increment whose distribution is equal to the weak limit of conditionally on as . The existence of this limit is an assumption on the Markov kernel induced by the distribution of the pair . Similarly, a large value at affects and via the increments and , respectively. The effect of on is then through the composite increment , whereas on it is through . The conditional independence property ensures that the increments , and are mutually independent. The common edge on the paths from node to node and from node to node induces dependence between the two tail chains and via the common increment . In this paper, the random vector
is called the tail tree induced by with root at node .
The tail tree represents a network of stochastic dependence relations that are not necessarily causal. Suppose the Markov tree in Figure 1 represents water levels at four locations on a river network. If water flows from left to right, node 2 represents a point where the stream branches into two channels, as occurs for instance in a river delta. If water flows from right to left, however, node 2 represents the junction of two branches coming from nodes 3 and 4 into a larger stream flowing towards node 1. In the first case, the tail tree describes how a high water level at the upstream node 1 may cause high water levels at various locations in the delta further downstream. In the second, case however, it is nodes 3 and 4 that are situated upstream, and the tail tree models the sources of a high water volume at the downstream site 1. Still other set-ups are possible, such as for instance node 3 being upstream and nodes 1 and 4 being downstream: high water levels at nodes 1 and 4 are then related through a common cause at node 2, which can itself perhaps be traced back to node 3.
Whatever the causal relationships within , it may make sense to change the conditioning variable. In Figure 1, for instance, suppose it is known that a large value has occurred at node 3 rather than at node 1. Tracing the paths from node 3 to the three other nodes yields the tail tree with root at node :
The tail trees in (1) and (2) have a similar structure. The two edges on the path between the root nodes 1 and 3 have changed direction, however. The edge from node 2 to node 4 is common to both tail trees.
For each pair of neighbouring nodes, the choice of the root node determines which of the two increments appears in the tail tree: from to or from to . The distributions of and are connected by an expression that involves the marginal distributions of and . For stationary and reversible Markov chains, this relation underlies a sufficiency property discovered in . For tail chains of not necessarily reversible Markov chains, it was described in and for tail processes of regularly varying stationary time series in via the time change formula. This formula can be understood most easily through the connection between the tail process and the tail measure , and this is also the way in which the root change formula in Corollary 3.2 below will be derived, but then without the assumption of stationarity and for general random vectors, not necessarily Markov trees.
2 Regular variation
For multivariate distributions, regular variation can be described via multivariate cumulative distribution functions as well, but an approach via convergence of Borel measures is more versatile. Let the state space be . Generalizations to star-shaped metric spaces or abstract cones as in are left for further work. Let denote the non-empty set of indices of variables of which the conditioning event is of possible interest. The marginal distributions of for are assumed to be regularly varying and the ratios of their tail functions are assumed to converge to positive constants. This set-up is a bit more general than the one of identical margins and comes at little technical or notational cost.
It is instructive to formulate statements in terms of weak convergence of distributions. For a high threshold tending to infinity and for a component , consider the asymptotic distribution of the rescaled random vector given that . Decompose as . Here, represents the overall level of with respect to whereas represents a self-normalized version of . Convergence in distribution of given as is a special case of what is called one-component regular variation in , explored already in for the bivariate case but allowing for affine normalizations. The random variable is asymptotically distributed and independent of , whose weak limit, denoted by , captures extremal dependence within given that is large. Letting the index run through produces multiple such one-component regular variation statements, which, together, are equivalent to what can be called multi-component regular variation. The limit distributions that arise for various indices must be mutually consistent, and the tail measure mentioned at the end of the previous paragraph embraces them all at once.
In Section 3, the focus is on tying together multiple one-component regular variation limits. The theory is worked out for general random vectors, not necessarily Markov trees. A number of results in that section have already been formulated in the literature in one way or another, in slightly different settings. Some of the equivalence relations in Theorem 3.1, for instance, resemble those in [17, Theorem 1.4] and [34, Proposition 3.1]. The model consistency property between limit measures in Theorem 3.1(ii) is formulated in [6, Section 2] for the bivariate case. The root change formula in Corollary 3.4 extends the time change formula for regularly varying stationary time series stemming from and studied extensively in . Multivariate Pareto distributions as in Theorem 3.9 are foreshadowed in [29, Section 6.3] and appear in when and in for more general functionals . These are just a few connections, and the above list is by no means intended to be complete.
The set-up involving regular variation is intended to serve two purposes. First, to model tail dependence within a vector of random variables which have been transformed to the same, heavy-tailed distribution, such as the unit-Fréchet distribution, as is common in multivariate extreme value theory. Second, to model the joint distribution of a vector of regularly varying random variables, not necessarily identically distributed, but with equivalent tails, such as returns on financial portfolios composed of the same basket of underlying assets. The latter framework is more general than the former and comes at little additional notational cost.
3 Outline
For a Markov tree , convergence as of the conditional distribution of given that is proved in Section 2. The main assumption is that, for edges directed away from the root , the conditional distribution of given converges as . No regular variation is needed yet.
The tail trees pertaining to different roots can be linked up thanks to the theory of one- and multi-component regular variation developed in Section 3. The results do not rely on the Markov property and cover quite general random vectors on , as is illustrated briefly for max-linear models. An interesting special case of these are the recursive max-linear structural equation models introduced in , featuring a causal structure induced by a directed acyclic graph. Most of the proofs of this section are deferred to the Appendix.
When combined, the results in Section 2 and 3 serve to uncover the regular variation properties of Markov trees in Section 4. The common special case that the joint distribution of the Markov tree is absolutely continuous with respect to Lebesgue measure is the subject of Section 5. The theory then simplifies considerably and the limit distribution with respect to a single root is already sufficient to reconstruct the limit distributions with respect to all other possible roots .
In Sections 4 and 5, the distributions of the increments of the tail trees are calculated in case the pair distributions are max-stable, not necessarily absolutely continuous. For the Hüsler–Reiss distribution max-stable distribution, the tail tree is multivariate log-normal, constructed from partial sums of independent normal random variables along the edges of the tree.
The spectral tail tree of a Markov tree
A (finite) graph is a pair where is a non-empty finite set of vertices or nodes and where is a set of edges. Self-loops are excluded, i.e., for all . To avoid trivialities, is assumed to have at least two elements. Two nodes are neighbours if they are joined by an edge. A graph is undirected if implies . A path from a node to a node is a collection of edges such that for all , for distinct nodes such that and . An undirected tree is an undirected graph such that for any pair of distinct nodes and , there exists a unique path from to , and this path is then denoted by .
Let be an undirected tree and let be a random vector indexed by the nodes of the tree. The pair is a Markov tree if it satisfies the global Markov property : whenever are disjoint, non-empty subsets of such that separates and (i.e., any path between a node and a node passes through some node in ), the conditional independence relation
holds, where denotes the random vector for .
For an undirected tree and a node , let denote the directed, rooted tree that consists of directing the edges in outward starting from . Formally, is the subset of that is obtained by choosing for every pair of edges and in the one such that the first node separates the second one from . If , then is the (necessarily unique) parent of in whereas is a child of in .
Let be a nonnegative Markov tree, where is an undirected tree.
There exists with the following two properties.
For every directed edge , there exists a version of the conditional distribution of given and a probability measure on such that
For edges such that and such that there exists an edge for which , we have
Assumption 2(ii) is similar to [26, equation (3.4)] and prevents non-extreme values to cause extreme ones. A similar assumption is [33, equation (2.4)], where it is illustrated [33, Example 7.5] what can go wrong without it.
Let be a nonnegative Markov tree on . Assume Condition 2. Let be a vector of independent random variables such that the law of is for all . Then
The random vector is called the tail tree of the Markov tree , adapting terminology for Markov chains in . In Figure 2, the tail tree is illustrated for a tree with seven nodes. For subvectors where all nodes in lie on the same path starting at , the structure of the tail tree is that of a geometric random walk; take for instance and in Figure 2. The tail tree couples several geometric random walks together through the common edges in the underlying paths: in the same figure, consider for instance the vectors indexed by and by , respectively, which share the initial edge .
Put . The proof is by induction on .
If has only two elements, i.e., , then Condition 2(i) already confirms the convergence stated in (6) and (7). Therefore, we can henceforth assume that has at least three elements, i.e., . Identify with in such a way that the root is and such that if then . Since , we do not need to consider the components and in (6).
the joint distribution of being given by (7).
Recall that denotes the parent node of . We need to distinguish between two cases: is the root or is a non-root vertex. The case is similar to but easier than the case and is left to the reader. We assume henceforth that .
We will show that the expression (9) converges to zero as (Step 3). Moreover, we will find a bound for the limit superior of (10) as . The bound will depend on but will converge to zero as (Step 4). Together, these properties of (9) and (10) are sufficient to prove the theorem (Step 5).
Step 3: The term (9). — The vertex is the parent of in , and therefore it separates from the other vertices. By the conditional independence property (3),
To explain our notation: the integral is over and is with respect to the conditional distribution of given that . The integrand involves the conditional expectation of a function of given that .
We change variables and integrate with respect to the conditional distribution of given that : we get
By Assumption 2(i), we have X_{d}/x_{k}\mid X_{k}=x_{k}\raisebox{-0.5pt}{\,\scriptsize\stackrel{{\scriptstyle\raisebox{-0.5pt}{\mbox{\scriptsize}}}}{{\longrightarrow}}}\,M_{k,d} as . Define
Recall that is bounded and (Lipschitz) continuous. By the extended continuous mapping theorem [37, Theorem 18.11], we have, for all vectors such that and for all functions such that as , the limit relation
Moreover, \mathcal{L}(X_{1:(d-1)}/t\mid X_{0}=t)\raisebox{-0.5pt}{\,\scriptsize\stackrel{{\scriptstyle\raisebox{-0.5pt}{\mbox{\scriptsize}}}}{{\longrightarrow}}}\,\mathcal{L}(\Theta_{1:(d-1)}) as by the induction hypothesis. By the same extended continuous mapping theorem, the integral (11) converges to
Recall that is a vector of independent random variables such that the law of is for . By construction, and are then independent too: each component of is a product of random variables with and thus . The above integral may therefore be simplified to
since . It follows that the limit of (9) as is equal to zero.
Step 4.b.i: The term (12). — The term (12) is bounded by
The node separates the nodes and . By the global Markov property, the expectation on the right-hand side of (15) is therefore equal to
Let . The conditional expectation in the integrand in (16) satisfies
Therefore, the integral in (16) is bounded by
By Assumption 2(ii), we can first take the limit superior as and then the limit superior as to find that
Since can be chosen arbitrarily close to zero, we find that the double limit superior above is equal to zero.
Step 4.b.ii: The term (13). — By the induction hypothesis, the term (13) converges to zero as .
Step 4.b.iii: The term (14). — Since , the term (14) is bounded by
By the dominated convergence theorem, the expectation on the right-hand side converges to zero as .
This completes the proof of the induction step and thus of the theorem.
In the setting of Theorem 2.1, also \mathcal{L}(X/X_{u}\mid X_{u}>t)\raisebox{-0.5pt}{\,\scriptsize\stackrel{{\scriptstyle\raisebox{-0.5pt}{\mbox{\scriptsize}}}}{{\longrightarrow}}}\,\Theta_{u} as .
Given , Theorem 2.1 allows us to find sufficiently high such that the absolute value inside the integral is bounded by for all . But then the left-hand side in the previous display is bounded by too, for all . Since was arbitrary, the stated convergence in distribution follows.
One- versus multi-component regular variation
Let be a random vector of nonnegative variables. Upon an obvious change in notation, Corollary 2.3 concerned weak convergence of as for some . This convergence plus regular variation of the marginal distribution of is a special case of what is called one-component regular variation in . The weak limit, , depends on the choice of . There may be good reasons to consider these limits for several indices . Let be the set of all indices for which such a limit exists. How are these random vectors related?
In this section, several such one-component statements are combined into a single one which could be called multi-component regular variation. If , this is just ordinary multivariate regular variation. As discussed already in Section 1.2, the connections between the limits generalize the time change formula for stationary regularly varying time series and can be deduced from their connections to a limiting tail measure.
For every we have \mathcal{L}(X/X_{i}\mid X_{i}>t)\raisebox{-0.5pt}{\,\scriptsize\stackrel{{\scriptstyle\raisebox{-0.5pt}{\mbox{\scriptsize}}}}{{\longrightarrow}}}\,\mathcal{L}(\Theta_{i}) as for some random vector on .
For every we have \mathcal{L}(X_{i}/t,X/X_{i}\mid X_{i}>t)\raisebox{-0.5pt}{\,\scriptsize\stackrel{{\scriptstyle\raisebox{-0.5pt}{\mbox{\scriptsize}}}}{{\longrightarrow}}}\,\operatorname{Pa}(\alpha)\otimes\mathcal{L}(\Theta_{i}) as for some random vector on .
For every we have \mathcal{L}(X/t\mid X_{i}>t)\raisebox{-0.5pt}{\,\scriptsize\stackrel{{\scriptstyle\raisebox{-0.5pt}{\mbox{\scriptsize}}}}{{\longrightarrow}}}\,\mathcal{L}(Y_{i}) as for some random vector on .
In that case, the limiting objects are connected in the following ways: for all ,
is equal in distribution to , where and and are independent;
is equal to the restriction of to ;
for every Borel measurable , we have
The proof of Theorem 3.1, together with the proofs of the other theorems in this section, is given in Appendix A.
Apart from the characterizations (a)–(e) in Theorem 3.1, other equivalent ones are possible, for instance, involving sequences rather than functions, with a scaling function inside the probability rather than outside, or with respect to radial and ‘angular’ coordinates for some appropriate functional . See for instance [29, Theorem 6.1] and [25, Theorem 3.1]. The tail measures and are homogeneous with index [25, Theorem 3.1] and, upon a coordinate transformation, can be written as product measures. Since the focus here is on the weak limits , these properties are not further elaborated upon. Statement (e) in Theorem 3.1 implies that the vector is multivariate regularly varying with limit measure on , which in turn implies, among other things, that it is in the domain of attraction of a multivariate max-stable distribution with Fréchet margins and exponent measure ; see for instance .
A noteworthy special case of (17) is when is the indicator function of the orthant , where has a non-empty intersection with and where for all . If , then , and thus
A remarkable consequence is that the right-hand side does not depend on the choice of . This invariance property is a special case of a more general mutual consistency property of the limit distributions for that is formulated in Corollary 3.2 below.
If the equivalent conditions of Theorem 3.1 are fulfilled, then, for Borel measurable and for , we have
If the equivalent conditions of Theorem 3.1 are fulfilled, then, for all Borel measurable and for all , we have
If the limit measure does not assign any mass to the coordinate hyperplane , the indicator in (17) is redundant and can be expressed entirely in terms of . Moreover, whether this occurs or not can be read off from the -th moments of the components of .
If for some , then, for all Borel measurable ,
Moreover, all tail trees for are determined by via (21).
Let be the indicator function of the set , where is non-empty and where for all . If , then, by (23),
In contrast to equation (18), equation (24) is true only when , a prerequisite for which Corollary 3.6 gives a necessary and sufficient condition.
In Theorem 3.1, a sufficient condition for (a)–(e) to hold is that there exists a non-empty set with the following two properties:
For every we have \mathcal{L}(X/X_{i}\mid X_{i}>t)\raisebox{-0.5pt}{\,\scriptsize\stackrel{{\scriptstyle\raisebox{-0.5pt}{\mbox{\scriptsize}}}}{{\longrightarrow}}}\,\mathcal{L}(\Theta_{i}) as for some random vector on .
In that case, also \mathcal{L}(X/X_{j}\mid X_{j}>t)\raisebox{-0.5pt}{\,\scriptsize\stackrel{{\scriptstyle\raisebox{-0.5pt}{\mbox{\scriptsize}}}}{{\longrightarrow}}}\,\mathcal{L}(\Theta_{j}) as for , where the law of is given in terms of the one of with via (21).
The focus so far has been on weak limits of conditional distributions involving a high-threshold exceedance by a specific component. In the spirit of the multivariate peaks-over-thresholds methodology , the following result covers, among other possibilities, the case where the conditioning event involves a high-threshold exceedance in at least one of a number of components.
Let follow the max-linear model
where are scalars such that for all and where are independent and identically distributed nonnegative random variables whose common distribution function has a regularly varying tail function with index . The marginal tails satisfy as . If exceeds a large threshold , the probability that this was due to is proportional to , and then the other factors for are of smaller order than . It follows that (a) in Theorem 3.1 holds where the law of is discrete with at most atoms and is given by
with denoting a unit point mass at . From (26), we find
Recursive max-linear models on directed acyclic graphs were introduced in . Borrowing some of their notation, consider a directed acyclic graph with nodes and edges , where denotes the possibly empty set of parents of . Consider a random vector given by the structural equation model
where the random variables are as in Example 3.10 with and where all coefficients and are (strictly) positive; the maximum over the empty set is zero by convention. Then by [14, Theorem 2.2], the random vector admits the max-linear representation
with coefficients , for , defined as follows: and if , while
where is the collection of paths from to in ; recall the definition of a path in the beginning of Section 2. The representation in (28) is of the form (25) with and with for . It follows that unless or .
In the special case that the directed acyclic graph is also a directed, rooted tree, every node has either exactly one parent or is equal to the root node, say . In that case, the collection of paths between and is a singleton, , and the formula for in (28) simplifies to . Furthermore, the tail tree in (26) starting from the root node simplifies to the degenerate distribution at the point with coordinates for . This is of the form in (7) with degenerate increments for all .
Regularly varying Markov trees
As in Section 2, let be a nonnegative Markov tree on the undirected tree . The general theory in Section 3 sheds light on the relation between two tail trees emanating at different roots. For two different nodes and in , the sets of directed edges and are the same except for the edges connecting nodes on the path between and , which are directed in opposite ways in the two edge sets: For every , we have and the other way around.
Condition 2 was formulated relative to a single root . The next condition covers all nodes in a non-empty subset of as possible roots. For such , let denote the set of directed edges that appear in at least one of the directed trees .
There exists a non-empty with the following two properties:
For every , there exists a version of the conditional distribution of given and a probability measure on such that (4) holds.
For every edge for which there exists such that and an edge such that , we have (5).
If in Condition 4 then every node that is on the path between and can be added to and Condition 4 remains true. Indeed, for such , we have , which takes care of (i), and for every node , which takes care of (ii). The author is grateful to an anonymous reviewer for having pointed this out.
Condition 4 and Corollary 2.3 imply that assumption (a) in Theorem 3.1 is satisfied for and with the tail tree in (7), for every . All equivalence relations and other properties are then as stated in Theorem 3.1.
In Corollary 4.1, if are neighbours in and if they both belong to , then the distributions of and mutually determine each other by
for all Borel measurable .
For different roots , the tail trees and have the same multiplicative structure. The differences between their distributions lie in the starting nodes of the paths and in the distributions of the multiplicative increments for edges on the paths and , since these edges change direction. For such edges of which the nodes belong to as well, the increment distributions are related by (29). See Figure 3 for an illustration.
Given the tree structure, the distribution of a Markov tree on is entirely determined by the bivariate distributions for . Markov chains of which all pairs are max-stable were proposed in [5, Section 4.6] and . When extended to trees, this construction method provides models meeting Condition 4.
Let the distribution of the random pair on be bivariate max-stable with cumulative distribution function
where is a Pickands dependence function, that is, a convex function such that for all ; see and the references therein. Both marginal distributions are unit-Fréchet, for . In particular, the marginal tail functions are regularly varying at infinity with index .
Let be the left-hand derivative of , which exists everywhere on , takes values between and , and is non-decreasing and continuous from the left; define as the right-hand limit. Since is convex, it is absolutely continuous, and the set of points in where it is not continuously differentiable is at most countable. For such that is differentiable at , we have
It follows that \mathcal{L}(Y/x\mid X=x)\raisebox{-0.5pt}{\,\scriptsize\stackrel{{\scriptstyle\raisebox{-0.5pt}{\mbox{\scriptsize}}}}{{\longrightarrow}}}\,M as , where
Absolutely continuous case
If the joint distribution of the Markov tree on is absolutely continuous with respect to the Lebesgue measure on , the formulations of the conditions and results simplify considerably. Let denote the joint probability density function of and let , for , denote the marginal density of .
By the Hammersley–Clifford theorem [23, Theorem 3.9], is a Markov tree as soon as the joint density factorizes as
The second product is over all unordered pairs of neighbours and denotes the bivariate density function of .
For such that , the density of is for . The following condition replaces Condition 4.
For every , there exists a probability density function on such that
Let the random vector on be a Markov tree on the undirected tree with joint density function . Assume there exists a positive function , regularly varying at infinity with index , such that as for every . If Condition 5 holds, then the conditions of Corollary 4.1 are satisfied with , the same constants , and auxiliary function . For all pairs of neighbours , the density of is and for almost every , we have
for every and for every Borel measurable , with the tail tree in (7). Moreover, all tail trees are connected through (21).
The function is regularly varying at infinity with index too. By Karamata’s theorem [3, Proposition 1.5.10], we have and thus as .
Condition 4 with follows from Condition 5 and Scheffé’s theorem. Part (ii) of Condition 4 is void, since for every .
If are neighbours, we can apply (29) to , where , to find
Since this is true for every , we must have for almost every , whence (32).
Finally, -convergence to with the stated expression follows from Theorem 3.1 and Corollary 3.6.
In Example 4.5, assume that is twice continuously differentiable on and that and . The distribution of is then absolutely continuous and the conditional density of given that converges as to the function
An interesting example in this respect is the bivariate Hüsler–Reiss distribution with Pickands dependence function
If all neighbouring pairs for of the Markov tree follow such Hüsler–Reiss max-stable distributions, the joint distribution of the tail tree is multivariate log-normal, since for all , where the random variables are independent and normally distributed with expectation and variance , with dependence parameter for all .
Appendix A Proofs for Section 3
(b) implies (c) and (i). — Since , statement (b) and the continuous mapping theorem [37, Theorem 2.3] imply that converges weakly to , where is a random variable independent of . Since almost surely, we have .
(c) implies (a). — Since , statement (c) and the continuous mapping theorem imply statement (a) with .
(b) implies (d). — Define a Borel measure on by
By linearity of the integral and by monotone convergence, we find that
for every nonnegative Borel measurable function on . The same expression is then true for real-valued Borel measurable functions on for which at least one of the two integrals with replaced by is finite. This includes bounded, Borel measurable functions that vanish on a set of the form for some .
Let and let be such that as soon as . By (b), we have
where is a random variable, independent of . The limit is equal to
since almost surely whenever , as almost surely.
(d) implies (e). — For every , we have
For , let be the piece-wise linear function
Put . Write . Then
Each function belongs to too but has moreover the property that as soon as . The restriction of to thus belongs to . By (d),
The existence of a limit has thus been shown, and convergence in to some measure as stated in (e) follows.
(e) implies (d), (ii), (iii) and (iv). — A function in can be extended to a function in denoted by the same symbol by putting for . Hence, (e) implies (d), with as described in (ii).
Statement (iii) follows from (ii) and the description of the law of in terms of in the proof above of the implication that (d) implies (c).
Similarly, (iv) follows from (ii), equation (33), and Fubini’s theorem.
It is sufficient to show statement (a) in Theorem 3.1. By property (i) in Theorem 3.8, the weak convergence in Theorem 3.1(a) already holds for all , and we need to show that it also holds for all . Choose and let be as in property (ii) of Theorem 3.8
We will show that converges weakly as to whose law is defined in (21). Let be open and let . We have
By Theorem 3.1 applied to , we have \mathcal{L}(X_{i}/s,X/X_{i}\mid X_{i}>s)\raisebox{-0.5pt}{\,\scriptsize\stackrel{{\scriptstyle\raisebox{-0.5pt}{\mbox{\scriptsize}}}}{{\longrightarrow}}}\,\operatorname{Pa}(\alpha)\otimes\mathcal{L}(\Theta_{i}) as . Let be a random variable, independent of . By the Portmanteau lemma for weak convergence, we have
The Portmanteau theorem for weak convergence implies the stated weak convergence of as . This proves statement (c) in Theorem 3.1 for the enlarged random vector .
The author is grateful to two anonymous reviewers whose suggestions have led to various improvements throughout the text. The author also wishes to thank Stefka Asenova and Gildas Mazo for inspiring discussions.