Global universal approximation of functional input maps on weighted spaces
Christa Cuchiero, Philipp Schmocker, Josef Teichmann
Introduction
We introduce a generalization of neural networks to infinite dimensional spaces which we call functional input neural networks (FNN). These neural networks can be applied as supervised learning tools in machine learning (see ), when the input and possibly also the output spaces are infinite dimensional.
In order to show the global universal approximation property of FNNs between infinite dimensional spaces, we use the structure of the weighted spaces and rely on a weighted version of the classical Stone-Weierstrass theorem (see e.g. ) proved in Section 3. Our formulation of this weighted Stone-Weierstrass theorem is inspired by Leopoldo Nachbin’s article , even though our setting differs in several important respects (definition of function spaces, weighted topologies, criteria for density, etc.). The crucial ingredient for global approximation is a weight function that controls the functions to be approximated outside of large compact sets and in turn allows to prove density of point separating moderate growth algebras (introduced in Definition 3.4) among all continuous functions which can be dominated by this weight function. Another Stone-Weierstrass theorem was proved in and for the space of continuous bounded functions over a non-compact space using the so-called strict topology (see also ), which can be seen as the projective limit of the weighted spaces introduced in this paper.
Let us remark that there is of course an extensive literature on neural networks with infinite dimensional input and output spaces. Early instances can be found in for learning a non-linear functional with continuous functions as input. For more recent results we refer to for the universal approximation on non-Euclidean spaces using feature maps and to for the approximation on Fréchet spaces. Moreover, echo-state network architectures were considered in and so-called metric hypertransformers in , which are both shown to be universal for adapted maps between suitably defined discrete-time path spaces. In addition, for learning the solution of a partial differential equation, we refer to the works on physics-informed neural networks in , neural operators in , DeepONets in , and the references therein.
The crucial novelty of the current work are the global UATs for (generalizations of) continuous functions beyond compacts. While classical UATs on compacts ensure the existence of an approximation over a fixed compact set, our results yield a global approximation. Hence, the model can be retrained with a new training set that might be not contained in an a priori fixed compact set. These non-compact UATs are highly relevant in areas like stochastic analysis or mathematical finance, where the model space is given by a set of paths, which is generically non-compact. An important class of functional input maps in mathematical finance are so-called non-anticipative functionals introduced in . We translate our general UAT to this setup and illustrate in Section 7 the numerical performance of FNNs when approximating the running maximum of a standard Brownian motion.
Apart from neural networks, there are many other families which can serve as universal approximators on function spaces. One well-known example is the signature of a path which actually serves as linear regression basis. It plays a central role in rough path theory (see e.g., ) and lately also in econometrics and mathematical finance see e.g., and the references therein). Similarly as for neural networks, UATs for linear functions of the signature have been proved on compact sets of paths (see e.g. for continuous semimartingales and for càdlàg paths). Note however that in the context of (semi-)martingales the Wiener-Ito chaos decomposition with iterated stochastic integrals instead of the signature (see for Brownian motion and for processes with stationary increments) can be seen as global UAT for the -norm with respect to the Wiener measure (see also the -version of for diffusion processes and for càdlàg semimartingales).
By choosing an appropriate weight function and by applying the weighted Stone-Weierstrass theorem we here obtain a global universal approximation result for linear functions of the signature of continuous geometric rough paths (see Section 5). Let us remark that in also a global approximation result is obtained, however not with the “true” signature but with a bounded normalization, as there the strict topology from on continuous bounded functions (see [23, Defintion 8]) is used. The disadvantage of this normalization is that the tractability properties of the signature, like computing its expected value analytically, are lost. In Section 6 we also introduce the viewpoint of Gaussian process regression in this setting and show that the reproducing kernel Hilbert space of the “true” signature kernels are Cameron-Martin spaces of certain Gaussian processes. This paves the way towards uncertainty quantification for signature kernel regression.
In the following subsection we provide three specific examples of FNNs that can be used to approximate -Hölder continuous functions, non-anticipative paths functionals as well as monetary risk measures.
In order to give an overview of the applicability of our results, we consider three particular examples. The first one is within functional data analysis, which is a branch of statistics where every sample corresponds to a continuous function (see ). Hence, we learn a continuous map with some growth conditions, where is a compact metric space, denotes the space of -Hölder continuous real-valued functions on , and is a Banach space. Then, our results show that the map can be approximated by learning a FNN of the form
These three examples provide a first glance over the broad range of applications for FNNs. Note that the input space can violate the vector space structure, as illustrated by the second example, and can be the dual of another Banach space, as in the third example.
The remainder of the article is structured as follows. In the next subsection we introduce relevant notation used throughout the article. Section 2 is dedicated to notions related to weighted spaces and the generalization of continuous functions, called -functions, defined thereon. In Section 3, we prove the weighted Stone-Weierstrass theorem in this setting, which we then use in Section 4 to lift the global universal approximation property of classical neural networks to FNNs. In Section 5, we present a further application of the weighted Stone-Weierstrass theorem to prove a global UAT for linear functions of path signatures. Section 6 then introduces a Gaussian process regression point of view and identifies the reproducing kernel Hilbert space of the signature kernel with the Cameron-Martin space of Gaussian processes taking values in -functions. Finally, we provide some numerical examples in Section 7.
2 Notation
Furthermore, for a Hausdorff topological space and a Banach space , we denote by the vector space of continuous maps , whereas denotes the vector subspace of bounded maps. If is additionally compact, every map is bounded. In this case, the supremum norm turns into a Banach space.
Moreover, if is a metric space and is a Banach space, a subset is called equicontinuous if for every there exists some such that for every and with it holds that . In addition, a subset is called pointwise bounded (resp. pointwise compact) if for every the set is bounded (resp. compact) in .
One of our most important example for a weighted space , as introduced in Section 2 below, are Hölder spaces defined as follows. For some , a compact metric space with designated origin , and a dual space equipped with the weak--topology, we denote by the space of -Hölder continuous functions with finite norm
where denotes the -Hölder seminorm of defined as
Unless otherwise specified, we endow with the -Hölder norm , which turns into a Banach space (see [47, Theorem 5.25 (ii)] and [102, Proposition 2.3(b)]). For , we can relate to the notion of globally Lipschitz continuous functions considered in . Moreover, we denote by the closed vector subspace of -Hölder continuous functions preserving the origin, i.e. .
The second important example are spaces of finite -variation. For some , , and a dual space equipped with the weak--topology, we denote by the space of continuous paths with finite norm
Hereby, denotes the -variation of defined as
We shall also need to intersect spaces of finite -variation with Hölder spaces. For some , with , and a dual space equipped with the weak--topology, we define the intersection which consists of -Hölder continuous paths with finite -variation. Unless otherwise specified, we endow with the norm
For the weighted space setting it will be essential to equip the Hölder spaces and the spaces also with weaker topologies than the norm topology, for example, with the uniform topology induced by or a weak--topology, which is precisely defined in Appendix A.
Weighted spaces and functions thereon
For the universal approximation results as well as the weighted Stone-Weierstrass theorems on infinite dimensional spaces, we use a weighted space as input space and a Banach space as output space. On the weighted input space, one can in turn introduce a weighted function space, which was under (slightly) different conditions also studied in .
Our weighted setting is in particular inspired by , where weighted spaces have been used for Kolmogorov equations, splitting schemes of (stochastic) partial differential equations, and generalized Feller processes. To make the current article self-contained, we shall recall all necessary definitions and provide several examples.
In order to define a weighted space we shall assume throughout that is a completely regular Hausdorff topological space.
A function is called an admissible weight function (on ) if every pre-image is compact with respect to , for all .
The definition of a weighted space has the following important consequences.
If is a separable locally convex topological vector space and is convex, then is continuous on if and only if is locally compact, see [36, Remark 2.2].
Note that by Baire’s category theorem can only be a Banach space if it is finite-dimensional.
One of our main examples for a weighted space (see Example 2.3 (ii) below) are spaces which are the dual of some Banach space, equipped with a weak--topology. Again, by Baire’s category theorem, these spaces are metrizable if and only if they are finite dimensional. Moreover, in infinite dimensions completeness with respect to the weak--topology also fails, since the completion would consist of all linear functions and not only of continuous ones.
In the following, we present some examples of weighted spaces , where the compactness of the pre-images needs to be verified with a suitable criterion (e.g. the Banach-Alaoglu theorem or the Arzelà-Ascoli theorem).
The examples of admissible weight functions that we present in the following are all of the form for a norm on and a continuous increasing function . In view of the so-called moderate growth condition defined below (see Definition 3.4), we consider in particular the continuous increasing function , for some and .
If is a dual space equipped with the weak--topology, i.e. if there exists a Banach space and an isometric isomorphism , we choose the weight function . Then, the pre-image is by the Banach-Alaoglu theorem compact in the weak--topology, for all . This shows that is a weighted space.
For , a compact metric space , and a dual space equipped with the weak--topology, let be the space of -Hölder continuous functions introduced in Section 1.2. Then, is by Theorem A.4 again a dual space, which means that can be equipped with a weak--topology. In this case, is by the Banach-Alaoglu theorem admissible, which shows that is a weighted space. Similarly, if is additionally pointed, we can consider functions in preserving the origin (see Section 1.2), where the same reasoning applies (see Remark A.7 (ii)).
Opposed to (iii), let be the space of -Hölder continuous functions , but now equipped with the supremum norm . Then, we choose and observe that the pre-image is equicontinuous and pointwise bounded, thus by the Banach-Alaoglu theorem pointwise compact with respect to the weak--topology of . Hence, is by the Arzelà-Ascoli theorem in [59, Theorem 7.17] compact with respect to , for all , which shows that is a weighted space. Note that instead of the uniform topology we could also use any -topology for due to the compact embedding (see Theorem A.3).
For and , let be the space of stopped -Hölder continuous paths (as introduced in Section 1.1), i.e.
For , with , and a dual space equipped with the weak--topology, let be the space of -Hölder continuous paths with finite -variation (see Section 1.2). Then, is by Theorem A.9 again a dual space and can be equipped with a weak--topology. Hence, is by the Banach-Alaoglu theorem admissible, which shows that is a weighted space. Similarly, we can consider paths in preserving the origin (see Section 1.2), where the same reasoning applies (see Remark A.12).
Opposed to (vi), let be the space of -Hölder continuous paths with finite -variation, but now equipped with the supremum norm . Then, we choose and use the compact embedding in Theorem A.8 to conclude that is compact with respect to , for all . Hence, is a weighted space. Note that instead of the uniform topology we could also use any -topology for due to the compact embedding (see Theorem A.8).
The above examples give rise to the following observations:
Note that in all the infinite dimensional examples (ii) to (ix) the idea is always to use a weaker topology than the norm topology which renders the closed unit ball compact. In the above cases this is either achieved with the weak--topology as in Example (ii), (iii), (vi), (viii), and (x), or via the following compact embeddings:
for as in Example (iv) and (v),
for with and with as in Example (vii), where , , and .
-embeddings of Sobolev–Slobodeckij spaces (see e.g. ).
Consider the setting of Example 2.3 (iii). Then, as shown in Theorem A.4, the weak--topology and the -topologies for (including the uniform topology with ) coincide on any -bounded and sequentially weak--closed subset of , but are globally clearly different. The same reasoning applies to Example 2.3 (vi), see Theorem A.9.
For a given weighted space with admissible weight function and a Banach space , we need to introduce an appropriate weighted function space that can be used for the subsequent global approximation theorems. Besides the vector space of bounded and continuous maps , we define the vector space
of maps , whose growth is controlled by the growth of the weight function . We equip with the weighted norm given by
for . Then, the continuous embedding holds true.
Note that if is compact, the weight function is admissible. However, on general spaces, grows on the compact pre-images , which means that the elements of are typically unbounded, but their growth is controlled by the growth of the weight function .
For simplicity, we always assume that the output space is a Banach space . However, all the following results (including the weighted vector-valued Stone-Weierstrass theorem in Theorem 3.8) still hold true with a locally convex topological vector space as output space. In this case, the topology on and is generated by families of seminorms instead of and , respectively (see ).
The following holds true for a function :
If then , for all , and
then . In particular, for every satisfying (2.3).
For (i), let and fix some . Then, there exists by definition of some such that . Hence, by choosing , it follows for every that
Since was chosen arbitrarily, we obtain (2.2). Moreover, with from above, we observe for every that
This shows that is continuous as uniform limit of continuous functions, for all . On the other hand, (ii) follows from [40, Theorem 2.7]. ∎
In addition, by following [40, Theorem 2.8], there exists for every with some such that , for all , which shows the analogy to functions vanishing at infinity on locally compact spaces.
Weighted Stone-Weierstrass theorems
In this section, we generalize the Stone-Weierstrass theorems to this weighted setting with non-compact input space. For this purpose, we shortly recall the classical Stone-Weierstrass theorems formulated for compact domains, which we will need for our results.
Subsequently, the Weierstrass approximation result in Theorem 3.1 was generalized by Marshall Harvey Stone in to more general spaces by using the notion of a subalgebra. Hereby, a vector subspace is called a subalgebra (of ) if is closed under multiplication, i.e. for every it holds that .
Let be a compact Hausdorff topological space and assume that is a subalgebra. Then, is dense in if and only if is point separating and vanishes nowhere.
Next, we state the vector-valued Stone-Weierstrass of R. Creighton Buck in . For a given subalgebra , a vector subspace is called an -submodule if , for all and , where . For further details on vector-valued approximation, we refer to the textbook .
Let be a compact Hausdorff space, let be a subalgebra and let be an -submodule. Then, the following are equivalent.
is point separating and vanishes nowhere, and is dense in , for all .
Later, the vector-valued Stone-Weierstrass was extended by Erret Bishop in to a measure theoretic version. In addition, Silvio Machado derived in a quantitative version, which relies on Zorn’s lemma, see also the proofs of and .
2 Weighted real-valued Stone-Weierstrass theorem
Note that Nachbin’s setting in differs in several respects, in particular in the definition of weighted topologies. It allows for multiple upper semicontinuous weights , e.g. , collected in a so-called Nachbin family and can therefore represent the compact-open topology, the strict topology, and other topologies.
For a given weight function , a subalgebra is called point separating of -moderate growth if there exists a point separating vector subspace such that , for all .
A subalgebra is point separating of -moderate growth under any given admissible weight function if the subalgebra is point separating and consists only of bounded maps. This is related to the so-called “bounded approximation problem” of Nachbin, see [79, Theorem 1].
Let be a subalgebra that is point separating of -moderate growth and vanishes nowhere. Then, is dense in .
It suffices to show the approximation of a map by an element since is by definition dense in . We first assume that consists only of bounded maps, where the condition being point separating of -moderate growth reduces to being point separating. Let , fix some , and define the constants as well as and . Observe that the restriction is a point separating subalgebra of , which vanishes nowhere. Hence, by using the classical real-valued Stone-Weierstrass in Theorem 3.2, there exists some such that
Hence, by combining (3.1) with (3.2), we conclude that
Since is a subalgebra, i.e. , and as well as were chosen arbitrarily, it follows that is dense in .
For the general case of a point separating subalgebra of -moderate growth, with point separating vector subspace such that , for all , we first show that the maps and , with , belong to the -closure of . Indeed, let and fix some . Then, by applying Lemma 2.7 (i), there exists some such that
Since was chosen arbitrarily, it follows that belongs to the -closure of , which holds analogously true for . Thus, the subalgebra
of is contained in the -closure of . Hence, by applying the previous step to the point separating subalgebra which vanishes nowhere and consists of bounded maps, we conclude that is dense in . However, since is contained in the -closure of , it follows that is also dense in . ∎
3 Weighted vector-valued Stone-Weierstrass theorem
In this section, we generalize the weighted Stone-Weierstrass to the vector-valued case, where is a Banach space. For this purpose, we introduce for every the tilted weight function defined by , for .
For every , the weight function is admissible.
We fix some . Clearly the set lies in a compact set of the form for some constant . A priori it is, however, not clear that is compact. It is of course sufficient to show that it is closed, since it lies in the compact set . By Lemma 2.7 (i) we know that \|w\|_{Y}\big{|}_{K} is continuous, whence \left(\|w\|_{Y}\big{|}_{K}\right)^{-1}([a,b]) is closed. Therefore
Let be a subalgebra that vanishes nowhere and is point separating of -moderate growth, for all , where is an -submodule such that is dense in , for all . Then, is dense in .
Let be the vector subspace of bounded maps. Then, we first show that the -closure of is a -submodule. For this purpose, we fix some and assume that belongs to the -closure of . Moreover, we fix some and define . Then, there exists some with , which implies that
Now, we observe that the tilted weight function is by Lemma 3.7 admissible. Moreover, for every , we have and as is an -submodule. Then, Lemma 2.7 (i) implies that
and that , for all . Thus, by using Lemma 2.7 (ii), we have , which shows that is also a subalgebra of as was chosen arbitrarily. Hence, by applying the weighted real-valued Stone-Weierstrass in Theorem 3.6 on , there exists some with , which implies that
By combining (3.4) with (3.5), we conclude that
Since is an -submodule, i.e. , and was chosen arbitrarily, it follows that belongs to the -closure of . By using that and were also chosen arbitrarily, the -closure of is a -submodule.
Finally, it suffices to prove the approximation of some by an element since is by definition dense in . Let , fix some , define as well as and choose . We observe that the restricted subalgebra of and the restricted -submodule of satisfy the assumptions of the classical vector-valued Stone-Weierstrass in Theorem 3.3 (i). Hence, there exists some such that
Now, for , it follows from (3.6) that . Moreover, we define by , for , which implies that , for all . Then, is bounded with . Hence, (3.6) and show that
Since the -closure of is by the previous step a -submodule, i.e. belongs to the -closure of , and as well as were chosen arbitrarily, it follows that is dense in . ∎
Theorem 3.8 can be generalized to the setting with a locally convex topological vector space as output space (see and Remark 2.6).
In Theorem 3.8, is point separating of -moderate growth for all , if either consists only of bounded functions (see Remark 3.5) or if is of -moderate growth and is a subset of . In both cases, the multiplication becomes jointly continuous.
The vector-valued Stone-Weierstrass in Theorem 3.8 is a version of Nachbin’s weighted approximation result in [79, Theorem 6]. Moreover, João B. Prolla extended in the measure theoretic version of Erret Bishop in to weighted spaces, see also .
Universal approximation on weighted spaces
First, we introduce the infinite dimensional analogue of affine maps, which are applied in classical neural networks on the hidden layer.
A subset is called an additive family (on ) if
is closed under addition, i.e. for every it holds that ,
is point separating, i.e. for distinct there is with .
We now give some examples of additive families. If the weighted space is a vector space equipped with a norm, then we can use the dual space of its completion.
Hence, Lemma 2.7 (ii) shows that and thus .
If is a Banach space in Lemma 4.9 such that is admissible, then is by Remark 2.2 (ii) finite dimensional. Hence, for an infinite dimensional weighted space , we consider in fact an incomplete normed vector space , where the dual of the -completion is used for the additive family.
Lemma 4.9 shows that a FNN over a weighted normed vector space can be chosen of the form
On the other hand, if the weighted space is a dual space equipped with the weak--topology as in Example 2.3 (ii), we obtain the following result.
Hence, Lemma 2.7 (ii) shows that and thus .
Lemma 4.11 shows that a FNN over a weighted dual space with isometric isomorphism and equipped with the weak--topology is given by
Let us give some examples of functional input neural networks, where we revisit among others the examples from Section 1.1. Hereby, is a continuous increasing function satisfying .
For the space of -Hölder continuous functions as in Example 2.3 (iv) with weight function , an additive family is given by
For the space of signed Radon measures on a weighted space as in Example 2.3 (x) with weight function , an additive family is given by
This can be used to learn a map with a (probability) measure as input.
These results give an overview of the broad variety of functional input neural networks, for which we show the universal approximation property in the following subsection.
2 Global universal approximation
Universal approximation results were first proved for neural networks between Euclidean spaces in and in using functional analytical arguments such as the Hahn-Banach theorem and Fourier theory. Later, further universal approximation results were found for different types of activation functions, such as in , while the results of e.g. , and are concerned with proving quantitative approximation rates under more restrictive assumptions on the involved functions.
Now, for the function , we define the corresponding FNN . Then, we conclude from (4.2) that
Since was chosen arbitrarily, it follows that the map belongs to the -closure of , which holds analogously true for the map . Hence, we conclude by the triangle inequality that linear combinations of these maps belong to the -closure of , which shows that the entire -submodule is contained in the -closure of .
Finally, we apply the weighted vector-valued Stone-Weierstrass in Theorem 3.8 to show that is dense in . For this purpose, we first see that vanishes nowhere as it contains the non-zero constant map . Moreover, we observe that is point separating and consists only of bounded maps, which shows that is point separating of -moderate growth, for all (see Remark 3.5). In addition, by using that is dense in , it follows that is dense in , for all . Hence, by applying the weighted vector-valued Stone-Weierstrass in Theorem 3.8, we conclude that is dense in . However, since is by the previous step contained in the -closure of , it follows that is also dense in . ∎
The global universal approximation result in Theorem 4.13 can be generalized to locally convex topological vector space as output spaces (see for a study of these infinite dimensional vector spaces and also Remark 2.6 and Remark 3.9).
3 Global universal approximation for non-anticipative functionals
The space of stopped paths was originally introduced in , and to generalize Föllmer’s pathwise Ito calculus to non-anticipative functionals. By introducing a special type of FNNs, so-called non-anticipative functional input neural networks, we now present the universal approximation result for this kind of functionals.
We consider again the space of stopped -Hölder continuous paths given by
Hence, we can apply the global universal approximation result below at least for continuous non-anticipative functionals, whose growth is controlled by .
Motivated by Example 4.12 (ii), we introduce a specific type of functional input neural network defined on , which preserve the non-anticipative behaviour.
We now apply the global universal approximation result in Theorem 4.13 to show that every can be approximated by some non-anticipative FNN.
Hence, the conclusion follows from Theorem 4.13. ∎
In Section 7, we provide some numerical examples, which illustrate how the theoretical approximation result in Corollary 4.17 can be applied to learn a non-anticipative functional by a non-anticipative FNNs. In particular for applications in finance, non-anticipative FNNs can be used within the framework of Stochastic Portfolio Theory (SPT) introduced by R. Fernholz in to learn optimal path-dependent portfolios, see .
Universal approximation on weighted spaces for linear functions of the signature
In this section, we give another application of the weighted real-valued Stone-Weierstrass Theorem 3.6, which is related to the theory of rough paths introduced by Terry Lyons in (we also refer to the textbooks and ). More precisely, we will show that a path space functional can be approximated globally by using linear functions of the paths’ signature, a notion which goes back to the work of Chen . This is in contrast to the usual results where only compact sets of paths are considered. Different versions of the latter, e.g. for finite variation paths or for continuous functions depending on the whole signature (instead of the rough path), are available in the literature (see for instance [69, Theorem 3.1], [62, Theorem 1], [72, Proposition 4.5] and [28, Section 3]). We refer to for the situation involving not just the approximation of a functional up to a fixed final value , but a uniform approximation on the whole time interval .
In general, one needs two key properties for universal approximation results for linear functions of the signature. First, that the signature at the terminal time determines uniquely the path (up to so-called tree-like equivalences, see and ), which ensures point separation. Second, that every polynomial of the signature can be realized as a linear function via the so-called shuffle product, which yields the algebra property of linear functions of the signature. While the classical Stone-Weierstrass then only guarantees the approximation on compact subsets of paths (e.g. with respect to a certain -variation metric), there have already been some attempts to go beyond this rather restrictive setting. As already mentioned in the introduction, used the strict topology originating from to formulate a global approximation result, however not with the “true” signature but with a bounded normalization implying that the tractability properties of the signature (like computing expected signature analytically as e.g. done in for generic classes of stochastic processes) are lost.
In contrast, our weighted space setup allows to work with the “true” signature and thus to preserve its tractability properties, while we can still provide a global approximation result. Hereby, we consider the path space of -Hölder continuous paths as weighted space, which we either equip with an -Hölder topology or the weak--topology as considered in Appendix A.
where , with and being determined via .
2 Global universal approximation of α𝛼\alpha-Hölder rough paths
The following lemma relates and and can be deduced from the estimates in [47, Proposition 8.15].
To ensure point separation (without tree-like equivalences), we shall always add an additional component to representing time. More precisely, we define the subspace
To get a weighted space, we choose the weight function defined by
Hence, the weight function is on both spaces admissible and it follows from Remark A.7 (iv) that the weighted function space is the same for each underlying space and , with .
Now, we present the global universal approximation theorem for linear functions of signature, which can approximate any path space functional in . Even though the proof looks involved, the most important ingredients can be summarized as follows: certain linear functions of the signature are given by the integrals
Since is the same for each underlying space and , with (see Remark A.7 (iv)), and the metrics and are equivalent (see Lemma 5.3), we choose in the following. Then, the result follows from the weighted real-valued Stone-Weierstrass in Theorem 3.6 applied to the set
Therefore, we need to prove that is a subspace of , and a point separating subalgebra of -moderate growth, which vanishes nowhere, where
is a possible candidate for the point separating vector subspace of -moderate growth.
is continuous on with respect to . This together with the continuity of the evaluation map
is continuous on with respect to . Furthermore, since linear functions are continuous, it follows that the map
is continuous on with respect to . Since was chosen arbitrarily, this shows that , for all . Moreover, by using the inequality
Since the exponential function dominates any polynomial, we conclude that
Hence, it follows from Lemma 2.7 (ii) that , which shows that .
where the last equality follows from the fact that the exponent tends to as is at most and . Hence, Lemma 2.7 (ii) shows that , which holds true for any .
Now, to show that is point separating, let be distinct. By contradiction, let us assume that for every and it holds that
for all . Thus, we conclude for every and that
By [23, Theorem 7], Theorem 5.4 thus also implies that the signature is characteristic for the measure space introduced in Example 2.3 (x). Indeed, the map
Laws of stochastic processes on path space are thus characterized by their expected signature if the following (super)-exponential moment condition is satisfied.
Let us remark that an alternative sufficient condition stating when the law of a group-valued random variable is characterized by its expected signature is given in [22, Corollary 6.6].
Alternatively to linear functions of the signature we can consider
which is again by the Stone-Weierstrass Theorem 3.6 dense in . We therefore have universality of the corresponding feature map and
3 Global universal approximation of p𝑝p-variation rough paths
In this section, we provide the universal approximation result for weakly geometric -variation rough paths (see [47, Definition 9.15 (i)]). However, in order to get a weighted space, we need to consider the subspace of weakly geometric -variation rough paths which are also Hölder continuous (see Example 2.3 (vi) and (vii) as well as Remark 2.4).
To get a weighted space, we choose similarly as above the weight function defined by
for and some and . Then, by the arguments as for but now using Remark A.12, we conclude that is admissible on both spaces and , with and . Moreover, does not depend on the choice of the underlying topology.
Now, we present our global universal approximation theorem for linear functions of the signature in .
Corollary 5.7 shows that every path space functional in can be learned with linear functions of the signature. Moreover, by following Remark 5.5, we observe that the signature is characteristic for the dual of .
Gaussian process regression with applications to signature kernels
As an important counterpart to the so far established density results, which make linear regressions feasible, we introduce a Gaussian process perspective, or, equivalently, the perspective of reproducing kernel Hilbert spaces on regression. Even though it can be done for general input spaces and general Banach spaces of functions thereon, we specialize here to weighted spaces as input spaces and -spaces as spaces of functions thereon. An important application is given through signature kernel regression on path spaces. The corresponding (general) signature kernels have already been considered in and are of the form
for appropriate choices of . We provide here a novel Gaussian process perspective in terms of -spaces, which together with the above results allows to treat approximation with the true signature kernels (without a normalization procedure as considered in ) rigorously. Consequently, uncertainty quantification results and a theory of regularization for regressions on signature components can be obtained.
Let be a general weighted space. A reproducing kernel Hilbert space on is a continuously embedded subspace of . Its kernel is uniquely defined through for all , i.e. the functional representation in of the Dirac measure located at via Riesz representation. Whence the kernel is symmetric, positive semi-definite and for . If, on the other hand, we are given a symmetric, positive semi-definite function such that for , and such that the span of with pre-scalar product is continuously embedded in , then the span’s closure is a reproducing kernel Hilbert space on with kernel .
We call a kernel universal if the corresponding reproducing kernel Hilbert space on is dense in . Since the signature feature map is universal (see Remark 5.5), it follows from [23, Proposition 29 and Proposition 43] that the (general) signature kernels as of form (6.1) are universal to with and as in Sections 5.2 and 5.3.
Let be a -valued centered Gaussian process and assume to be separable. Notice that separability is never a restriction if we are given a countable set of functions, like signature components, whose span (built with rational coefficients) is dense in , which, in other words, is the typical situation for linear regression.
The Cameron-Martin theorem then states that if and only if the law of is absolutely continuous to the law of . The well-known Radon-Nikodym derviative is given by \exp\big{(}(f\bullet Z)-\frac{1}{2}\|f\|_{H_{Z}}\big{)}.
In this context the meaning of the Gaussian process also corresponds to the generation of a prior distribution of functions, which are added to the searched function in question.
With these preparations we can now formulate the main theorem of this section, which has immediate applications for signature regression:
Given a kernel on and a countable system of functions in such as function on and with separable. Then the random series , for an i.i.d. sequence of standard Gaussian random variable , converges in the -norm to a Gaussian process with values in and with kernel . In particular its Cameron-Martin space coincides with the reproducing kernel Hilbert space generated by .
Under the condition the Ito-Nisio theorem, see e.g. [68, Theorem 2.4], guarantees that , for an i.i.d. sequence of standard Gaussian random variable , converges to a random variable , which is again a Gaussian random variable. Its kernel can be easily calculated by pointwise covariances
and coincides with by assumption. The pointwise covariances, however, determine and in turn the Cameron-Martin space as stated above. ∎
With this theorem uncertainty quantification for regression with signature kernels is now feasible, since for reasonable choices of the main condition of the theorem is satisfied. We shall show this in Remark 6.5 in the setting of Section 5.3.
Take together with
where are real numbers. More precisely, for satisfying , we require for that
where is some constant depending on and
holds. Using , Hölder’s inequality with , and with , we can estimate
where the second last inequality follows from the estimate for . The root criterion then implies that the expression is finite. Indeed, noticing that for , we have as , it follows that the first term is finite due to the assumption that for and some constant . Similarly the second term is finite as well.
Notice also that one could replace by a function of an absolutely converging sum of the form
with positive such that for , , and .
Under the above conditions on and the kernel
is thus the covariance of a Gaussian process with values in .
Choosing identical for all of the same length and defining
for bounded variation paths , we then get by the same arguments as in [18, Proposition 2.8] the following integral equation
where denotes the sequence of shifted by , i.e.
Note that for both and Condition 1 of which is required in [18, Proposition 2.8] is satisfied due to (6.2). In particular, satisfies
If and are differentiable, then the original signature kernel satisfies the so-called Goursat PDE
Numerical examples
In this section, we illustrate in two particular examplesThe experiments are implemented in Python using tensorflow (for FNN) and iisignature (for signatures) on a Lenovo ThinkPad X13 Gen2a with AMD Ryzen 7 PRO 5850U processor and Radeon Graphics (1901 Mhz, 8 Cores, 16 Logical Processors), see https://github.com/psc25/GlobalUAT. how path space functionals can be learned using non-anticipative functional input neural networks (see Section 4.3) or a linear function of the signature (see Section 5).
As input data we generate sample paths of a one-dimensional Brownian motion , for , with , which are discretized over equidistant time points . Since the sample paths of Brownian motion are a.s. continuous and -Hölder continuous, for all , we consider the weighted space of stopped -Hölder continuous paths in Example 2.3 (v), with .
We split up the data into for training and for testing, and apply stochastic gradient descent with the Adam algorithm (see ) over epochs with learning rate and batchsize to minimize the mean squared error (MSE)
For the FNNs, we consider with neurons (where are both the ReLU function and the classical neural networks have one hidden layer of neurons), while for signature, we choose .
Appendix A Predual of Banach spaces
In the following, we apply the result of to characterize a Banach space as dual space, which formally generalizes the Dixmier-Ng theorem (see and ).
Let be a Banach space, and let be a set of continuous linear functionals such that
Then, is a dual space and a predual of is the -closure of .
Now, we apply Theorem A.1 to show the existence of a predual for the Banach spaces presented in Example 2.3, i.e. Hölder spaces and spaces of finite -variation.
For every , the embedding is compact.
Fix some and define the constant . Then, by using the triangle inequality of , we conclude for every that
which shows that the embedding is continuous.
where . This shows that the embedding is compact. ∎
For every , the Banach space is a dual space. Moreover, on every -bounded and sequentially weak--closed subset of , the weak--topology and every -topology induced by , , are equivalent.
Let be a predual of . We want to apply Theorem A.1 with
On the other hand, the closed unit ball is by Theorem A.3 relatively compact in . Since the weak topology on induced by is weaker than the topology induced by , it follows that is compact with respect to this weak topology. Hence, by using Theorem A.1, the -closure of , denoted by , is a predual of .
Now, we denote by and the subspace topology on of the weak--topology and of the -topology induced by , respectively. Since is first countable, the previous argument implies that the identity is continuous, which shows that . Moreover, by using that is compact and that is Hausdorff, it follows from [78, Theorem 26.6] that the inverse is also continuous, which implies . This shows that and completes the proof. ∎
The point evaluations in (A.2) are also used in and to construct the Arens-Eells space and the Lipschitz-free space , respectively, which are both used as preduals of the space of globally Lipschitz continuous functions.
We now give some examples of -bounded and sequentially weak--closed subsets .
For every and sequentially weak--closed subset , the set
is -bounded and sequentially weak--closed. In particular, by choosing , the closed -Hölder ball is -bounded and sequentially weak--closed.
Choosing now the weight function , with a continuous increasing function , we consider the weighted space either equipped with the weak--topology as in Example 2.3 (iii) or with the -topology as in Example 2.3 (iv). Then, by using that these topologies coincide on the closed -Hölder balls (see Lemma A.5), we conclude that the weighted function space is the same for both choices of topologies on .
For any , let and denote the space equipped with the weak--topology and the -topology induced by , respectively. Then, .
On the other hand, by combining Lemma 2.7 (i)+(ii) for , we have
Now, the pre-image is for every a closed -Hölder ball (thus -bounded and sequentially weak--closed by Lemma A.5), which implies by Theorem A.4 that the weak--topology and the -topology induced by coincide on . Hence, by combining (A.3) and (A.4), it follows that if and only if . ∎
Let us summarize the consequences of this section for the closed vector subspace of -Hölder continuous functions preserving the origin.
Theorem A.3 implies the compact embedding , .
For every and sequentially weak--closed subset , the set
is -bounded and sequentially weak--closed (see Lemma A.5).
By Proposition A.6, the weighted function space with admissible weight function of the form does not depend on the choice of the topology on .
In this section, we show for some and a dual space equipped with the weak--topology that the -spaces introduced in Section 1.2 are also dual spaces, where with . For this purpose, we first consider the embeddings for with , where and .
For every with and with , the embedding is compact.
Fix some with and with . Then, [47, Proposition 5.3] implies for every that
which shows that is continuous.
where . This shows that is compact. ∎
For every satisfying , the Banach space is a dual space. Moreover, on every sequentially weak--closed and -bounded subset of , the weak--topology and every -topology induced by , with , are equivalent.
Let be a predual of . We want to apply Theorem A.1 with
Then, we follow the proof of Theorem A.4 and use Theorem A.1 to conclude that the -closure of is a predual of . Moreover, by using the same arguments as in the proof of Theorem A.4, the weak--topology and every -topology induced by , with and , coincide on each sequentially weak--closed and -bounded subset of . ∎
For every and sequentially weak--closed subset , the set
is -bounded and sequentially weak--closed. In particular, by choosing , the closed ball is -bounded and sequentially weak--closed.
The proof follow along the lines of the proof for Lemma A.5. ∎
For the weight function , with continuous increasing , we now consider the weighted space either equipped with the weak--topology as in Example 2.3 (vi) or with the -topology as in Example 2.3 (vii). Then, by using that these topologies coincide on the closed balls (see Lemma A.10), we conclude that the weighted function space is the same for both choices of topologies on .
For any with , let and denote the space equipped with the weak--topology and the -topology induced by , respectively. Then, .
The proof follows along the lines of the proof for Proposition A.6. ∎
For the closed vector subspace of -Hölder continuous functions with finite -variation preserving the origin, we obtain the analogous results as in Remark A.7.