Global universal approximation of functional input maps on weighted spaces

Christa Cuchiero, Philipp Schmocker, Josef Teichmann

Introduction

We introduce a generalization of neural networks to infinite dimensional spaces which we call functional input neural networks (FNN). These neural networks can be applied as supervised learning tools in machine learning (see ), when the input and possibly also the output spaces are infinite dimensional.

In order to show the global universal approximation property of FNNs between infinite dimensional spaces, we use the structure of the weighted spaces and rely on a weighted version of the classical Stone-Weierstrass theorem (see e.g. ) proved in Section 3. Our formulation of this weighted Stone-Weierstrass theorem is inspired by Leopoldo Nachbin’s article , even though our setting differs in several important respects (definition of function spaces, weighted topologies, criteria for density, etc.). The crucial ingredient for global approximation is a weight function that controls the functions to be approximated outside of large compact sets and in turn allows to prove density of point separating moderate growth algebras (introduced in Definition 3.4) among all continuous functions which can be dominated by this weight function. Another Stone-Weierstrass theorem was proved in and for the space of continuous bounded functions over a non-compact space using the so-called strict topology (see also ), which can be seen as the projective limit of the weighted spaces introduced in this paper.

Let us remark that there is of course an extensive literature on neural networks with infinite dimensional input and output spaces. Early instances can be found in for learning a non-linear functional with continuous functions as input. For more recent results we refer to for the universal approximation on non-Euclidean spaces using feature maps and to for the approximation on Fréchet spaces. Moreover, echo-state network architectures were considered in and so-called metric hypertransformers in , which are both shown to be universal for adapted maps between suitably defined discrete-time path spaces. In addition, for learning the solution of a partial differential equation, we refer to the works on physics-informed neural networks in , neural operators in , DeepONets in , and the references therein.

The crucial novelty of the current work are the global UATs for (generalizations of) continuous functions beyond compacts. While classical UATs on compacts ensure the existence of an approximation over a fixed compact set, our results yield a global approximation. Hence, the model can be retrained with a new training set that might be not contained in an a priori fixed compact set. These non-compact UATs are highly relevant in areas like stochastic analysis or mathematical finance, where the model space is given by a set of paths, which is generically non-compact. An important class of functional input maps in mathematical finance are so-called non-anticipative functionals introduced in . We translate our general UAT to this setup and illustrate in Section 7 the numerical performance of FNNs when approximating the running maximum of a standard Brownian motion.

Apart from neural networks, there are many other families which can serve as universal approximators on function spaces. One well-known example is the signature of a path which actually serves as linear regression basis. It plays a central role in rough path theory (see e.g., ) and lately also in econometrics and mathematical finance see e.g., and the references therein). Similarly as for neural networks, UATs for linear functions of the signature have been proved on compact sets of paths (see e.g. for continuous semimartingales and for càdlàg paths). Note however that in the context of (semi-)martingales the Wiener-Ito chaos decomposition with iterated stochastic integrals instead of the signature (see for Brownian motion and for processes with stationary increments) can be seen as global UAT for the L2L^{2}-norm with respect to the Wiener measure (see also the LpL^{p}-version of for diffusion processes and for càdlàg semimartingales).

By choosing an appropriate weight function and by applying the weighted Stone-Weierstrass theorem we here obtain a global universal approximation result for linear functions of the signature of continuous geometric rough paths (see Section 5). Let us remark that in also a global approximation result is obtained, however not with the “true” signature but with a bounded normalization, as there the strict topology from on continuous bounded functions (see [23, Defintion 8]) is used. The disadvantage of this normalization is that the tractability properties of the signature, like computing its expected value analytically, are lost. In Section 6 we also introduce the viewpoint of Gaussian process regression in this setting and show that the reproducing kernel Hilbert space of the “true” signature kernels are Cameron-Martin spaces of certain Gaussian processes. This paves the way towards uncertainty quantification for signature kernel regression.

In the following subsection we provide three specific examples of FNNs that can be used to approximate α\alpha-Hölder continuous functions, non-anticipative paths functionals as well as monetary risk measures.

In order to give an overview of the applicability of our results, we consider three particular examples. The first one is within functional data analysis, which is a branch of statistics where every sample corresponds to a continuous function (see ). Hence, we learn a continuous map f:Cα(S)→Yf:C^{\alpha}(S)\rightarrow Y with some growth conditions, where (S,dS)(S,d_{S}) is a compact metric space, Cα(S)C^{\alpha}(S) denotes the space of α\alpha-Hölder continuous real-valued functions on (S,dS)(S,d_{S}), and (Y,∥⋅∥Y)(Y,\|\cdot\|_{Y}) is a Banach space. Then, our results show that the map f:Cα(S)→Yf:C^{\alpha}(S)\rightarrow Y can be approximated by learning a FNN φ:Cα(S)→Y\varphi:C^{\alpha}(S)\rightarrow Y of the form

These three examples provide a first glance over the broad range of applications for FNNs. Note that the input space can violate the vector space structure, as illustrated by the second example, and can be the dual of another Banach space, as in the third example.

The remainder of the article is structured as follows. In the next subsection we introduce relevant notation used throughout the article. Section 2 is dedicated to notions related to weighted spaces and the generalization of continuous functions, called Bψ\mathcal{B}_{\psi}-functions, defined thereon. In Section 3, we prove the weighted Stone-Weierstrass theorem in this setting, which we then use in Section 4 to lift the global universal approximation property of classical neural networks to FNNs. In Section 5, we present a further application of the weighted Stone-Weierstrass theorem to prove a global UAT for linear functions of path signatures. Section 6 then introduces a Gaussian process regression point of view and identifies the reproducing kernel Hilbert space of the signature kernel with the Cameron-Martin space of Gaussian processes taking values in Bψ\mathcal{B}_{\psi}-functions. Finally, we provide some numerical examples in Section 7.

2 Notation

Furthermore, for a Hausdorff topological space (X,τX)(X,\tau_{X}) and a Banach space (Y,∥⋅∥Y)(Y,\|\cdot\|_{Y}), we denote by C0(X;Y)C^{0}(X;Y) the vector space of continuous maps f:X→Yf:X\rightarrow Y, whereas Cb0(X;Y)⊆C0(X;Y)C^{0}_{b}(X;Y)\subseteq C^{0}(X;Y) denotes the vector subspace of bounded maps. If (X,τX)(X,\tau_{X}) is additionally compact, every map f∈C0(X;Y)f\in C^{0}(X;Y) is bounded. In this case, the supremum norm ∥f∥C0(X;Y):=sup⁡x∈X∥f(x)∥Y\|f\|_{C^{0}(X;Y)}:=\sup_{x\in X}\|f(x)\|_{Y} turns C0(X;Y)C^{0}(X;Y) into a Banach space.

Moreover, if (S,dS)(S,d_{S}) is a metric space and (Z,∥⋅∥Z)(Z,\|\cdot\|_{Z}) is a Banach space, a subset A⊆C0(S;Z)A\subseteq C^{0}(S;Z) is called equicontinuous if for every ε>0\varepsilon>0 there exists some δ>0\delta>0 such that for every f∈Af\in A and s,t∈Ss,t\in S with dS(s,t)<δd_{S}(s,t)<\delta it holds that ∥f(s)−f(t)∥Z<ε\|f(s)-f(t)\|_{Z}<\varepsilon. In addition, a subset A⊆C0(S;Z)A\subseteq C^{0}(S;Z) is called pointwise bounded (resp. pointwise compact) if for every s∈Ss\in S the set {f(s):f∈A}\{f(s):f\in A\} is bounded (resp. compact) in ZZ.

One of our most important example for a weighted space XX, as introduced in Section 2 below, are Hölder spaces defined as follows. For some α∈(0,1]\alpha\in(0,1], a compact metric space (S,dS)(S,d_{S}) with designated origin 0∈S0\in S, and a dual space ZZ equipped with the weak-∗*-topology, we denote by Cα(S;Z)C^{\alpha}(S;Z) the space of α\alpha-Hölder continuous functions x:S→Zx:S\rightarrow Z with finite norm

where ∥⋅∥α\|\cdot\|_{\alpha} denotes the α\alpha-Hölder seminorm of x:S→Zx:S\rightarrow Z defined as

Unless otherwise specified, we endow Cα(S;Z)C^{\alpha}(S;Z) with the α\alpha-Hölder norm ∥⋅∥Cα(S;Z)\|\cdot\|_{C^{\alpha}(S;Z)}, which turns Cα(S;Z)C^{\alpha}(S;Z) into a Banach space (see [47, Theorem 5.25 (ii)] and [102, Proposition 2.3(b)]). For α=1\alpha=1, we can relate Cα(S;Z)C^{\alpha}(S;Z) to the notion of globally Lipschitz continuous functions x:S→Zx:S\rightarrow Z considered in . Moreover, we denote by C0α(S;Z)⊆Cα(S;Z)C^{\alpha}_{0}(S;Z)\subseteq C^{\alpha}(S;Z) the closed vector subspace of α\alpha-Hölder continuous functions x∈Cα(S;Z)x\in C^{\alpha}(S;Z) preserving the origin, i.e. x(0)=0∈Zx(0)=0\in Z.

The second important example are spaces of finite pp-variation. For some T>0T>0, p≥1p\geq 1, and a dual space ZZ equipped with the weak-∗*-topology, we denote by Cp−var([0,T];Z)C^{p-var}([0,T];Z) the space of continuous paths x:[0,T]→Zx:[0,T]\rightarrow Z with finite norm

Hereby, ∥x∥p−var\|x\|_{p-var} denotes the pp-variation of x:[0,T]→Zx:[0,T]\rightarrow Z defined as

We shall also need to intersect spaces of finite pp-variation with Hölder spaces. For some T>0T>0, (p,α)∈[1,∞)×(0,1)(p,\alpha)\in[1,\infty)\times(0,1) with pα<1p\alpha<1, and a dual space ZZ equipped with the weak-∗*-topology, we define the intersection Cp−var,α([0,T];Z):=Cp−var([0,T];Z)∩Cα([0,T];Z)C^{p-var,\alpha}([0,T];Z):=C^{p-var}([0,T];Z)\cap C^{\alpha}([0,T];Z) which consists of α\alpha-Hölder continuous paths x:[0,T]→Zx:[0,T]\rightarrow Z with finite pp-variation. Unless otherwise specified, we endow Cp−var,α([0,T];Z)C^{p-var,\alpha}([0,T];Z) with the norm

For the weighted space setting it will be essential to equip the Hölder spaces Cα(S;Z)C^{\alpha}(S;Z) and the spaces Cp−var,α(S;Z)C^{p-var,\alpha}(S;Z) also with weaker topologies than the norm topology, for example, with the uniform topology induced by ∥⋅∥C0(S;Z)\|\cdot\|_{C^{0}(S;Z)} or a weak-∗*-topology, which is precisely defined in Appendix A.

Weighted spaces and functions thereon

For the universal approximation results as well as the weighted Stone-Weierstrass theorems on infinite dimensional spaces, we use a weighted space as input space and a Banach space as output space. On the weighted input space, one can in turn introduce a weighted function space, which was under (slightly) different conditions also studied in .

Our weighted setting is in particular inspired by , where weighted spaces have been used for Kolmogorov equations, splitting schemes of (stochastic) partial differential equations, and generalized Feller processes. To make the current article self-contained, we shall recall all necessary definitions and provide several examples.

In order to define a weighted space we shall assume throughout that (X,τX)(X,\tau_{X}) is a completely regular Hausdorff topological space.

A function ψ:X→(0,∞)\psi:X\rightarrow(0,\infty) is called an admissible weight function (on (X,τX)(X,\tau_{X})) if every pre-image KR:=ψ−1((0,R])={x∈X:ψ(x)≤R}K_{R}:=\psi^{-1}((0,R])=\left\{x\in X:\psi(x)\leq R\right\} is compact with respect to τX\tau_{X}, for all R>0R>0.

The definition of a weighted space has the following important consequences.

If (X,τX)(X,\tau_{X}) is a separable locally convex topological vector space and ψ:X→(0,∞)\psi:X\rightarrow(0,\infty) is convex, then ψ:X→(0,∞)\psi:X\rightarrow(0,\infty) is continuous on E⊆XE\subseteq X if and only if EE is locally compact, see [36, Remark 2.2].

Note that by Baire’s category theorem (X,τX)(X,\tau_{X}) can only be a Banach space if it is finite-dimensional.

One of our main examples for a weighted space (see Example 2.3 (ii) below) are spaces XX which are the dual of some Banach space, equipped with a weak-∗*-topology. Again, by Baire’s category theorem, these spaces are metrizable if and only if they are finite dimensional. Moreover, in infinite dimensions completeness with respect to the weak-∗*-topology also fails, since the completion would consist of all linear functions and not only of continuous ones.

In the following, we present some examples of weighted spaces (X,ψ)(X,\psi), where the compactness of the pre-images KR=ψ−1((0,R])K_{R}=\psi^{-1}((0,R]) needs to be verified with a suitable criterion (e.g. the Banach-Alaoglu theorem or the Arzelà-Ascoli theorem).

The examples of admissible weight functions ψ:X→(0,∞)\psi:X\rightarrow(0,\infty) that we present in the following are all of the form ψ(x)=η(∥x∥)\psi(x)=\eta\left(\|x\|\right) for a norm ∥⋅∥\|\cdot\| on XX and a continuous increasing function η:[0,∞)→(0,∞)\eta:[0,\infty)\rightarrow(0,\infty). In view of the so-called moderate growth condition defined below (see Definition 3.4), we consider in particular the continuous increasing function η(r)=exp⁡(βrγ)\eta(r)=\exp\left(\beta r^{\gamma}\right), for some β>0\beta>0 and γ>1\gamma>1.

If (X,∥⋅∥X)(X,\|\cdot\|_{X}) is a dual space equipped with the weak-∗*-topology, i.e. if there exists a Banach space (E,∥⋅∥E)(E,\|\cdot\|_{E}) and an isometric isomorphism Φ:X→E∗\Phi:X\rightarrow E^{*}, we choose the weight function ψ(x)=η(∥x∥X)\psi(x)=\eta\left(\|x\|_{X}\right). Then, the pre-image KR=ψ−1((0,R])K_{R}=\psi^{-1}((0,R]) is by the Banach-Alaoglu theorem compact in the weak-∗*-topology, for all R>0R>0. This shows that (X,ψ)(X,\psi) is a weighted space.

For α>0\alpha>0, a compact metric space (S,dS)(S,d_{S}), and a dual space ZZ equipped with the weak-∗*-topology, let X=Cα(S;Z)X=C^{\alpha}(S;Z) be the space of α\alpha-Hölder continuous functions x:S→Zx:S\rightarrow Z introduced in Section 1.2. Then, Cα(S;Z)C^{\alpha}(S;Z) is by Theorem A.4 again a dual space, which means that Cα(S;Z)C^{\alpha}(S;Z) can be equipped with a weak-∗*-topology. In this case, ψ(x)=η(∥x∥Cα(S;Z))\psi(x)=\eta\left(\|x\|_{C^{\alpha}(S;Z)}\right) is by the Banach-Alaoglu theorem admissible, which shows that (Cα(S;Z),ψ)(C^{\alpha}(S;Z),\psi) is a weighted space. Similarly, if (S,dS)(S,d_{S}) is additionally pointed, we can consider functions in C0α(S;Z)C^{\alpha}_{0}(S;Z) preserving the origin (see Section 1.2), where the same reasoning applies (see Remark A.7 (ii)).

Opposed to (iii), let X=Cα(S;Z)X=C^{\alpha}(S;Z) be the space of α\alpha-Hölder continuous functions x:S→Zx:S\rightarrow Z, but now equipped with the supremum norm ∥⋅∥C0(S;Z)\|\cdot\|_{C^{0}(S;Z)}. Then, we choose ψ(x)=η(∥x∥Cα(S;Z))\psi(x)=\eta\left(\|x\|_{C^{\alpha}(S;Z)}\right) and observe that the pre-image KR=ψ−1((0,R])K_{R}=\psi^{-1}((0,R]) is equicontinuous and pointwise bounded, thus by the Banach-Alaoglu theorem pointwise compact with respect to the weak-∗*-topology of ZZ. Hence, KR=ψ−1((0,R])K_{R}=\psi^{-1}((0,R]) is by the Arzelà-Ascoli theorem in [59, Theorem 7.17] compact with respect to ∥⋅∥C0(S;Z)\|\cdot\|_{C^{0}(S;Z)}, for all R>0R>0, which shows that (Cα(S;Z),ψ)(C^{\alpha}(S;Z),\psi) is a weighted space. Note that instead of the uniform topology we could also use any Cα′C^{\alpha^{\prime}}-topology for 0<α′<α0<\alpha^{\prime}<\alpha due to the compact embedding Cα(S;Z)↪Cα′(S;Z)C^{\alpha}(S;Z)\hookrightarrow C^{\alpha^{\prime}}(S;Z) (see Theorem A.3).

For α>0\alpha>0 and T>0T>0, let X=ΛTαX=\Lambda^{\alpha}_{T} be the space of stopped α\alpha-Hölder continuous paths (as introduced in Section 1.1), i.e.

For T>0T>0, (p,α)∈[1,∞)×(0,1)(p,\alpha)\in[1,\infty)\times(0,1) with pα<1p\alpha<1, and a dual space ZZ equipped with the weak-∗*-topology, let X=Cp−var,α([0,T];Z)X=C^{p-var,\alpha}([0,T];Z) be the space of α\alpha-Hölder continuous paths x:[0,T]→Zx:[0,T]\rightarrow Z with finite pp-variation (see Section 1.2). Then, Cp−var,α([0,T];Z)C^{p-var,\alpha}([0,T];Z) is by Theorem A.9 again a dual space and can be equipped with a weak-∗*-topology. Hence, ψ(x)=η(∥x∥Cp−var,α([0,T];Z))\psi(x)=\eta\left(\|x\|_{C^{p-var,\alpha}([0,T];Z)}\right) is by the Banach-Alaoglu theorem admissible, which shows that (Cp−var,α([0,T];Z),ψ)(C^{p-var,\alpha}([0,T];Z),\psi) is a weighted space. Similarly, we can consider paths in C0p−var,α([0,T];Z)C^{p-var,\alpha}_{0}([0,T];Z) preserving the origin (see Section 1.2), where the same reasoning applies (see Remark A.12).

Opposed to (vi), let X=Cp−var,α([0,T];Z)X=C^{p-var,\alpha}([0,T];Z) be the space of α\alpha-Hölder continuous paths x:[0,T]→Zx:[0,T]\rightarrow Z with finite pp-variation, but now equipped with the supremum norm ∥⋅∥C0([0,T];Z)\|\cdot\|_{C^{0}([0,T];Z)}. Then, we choose ψ(x)=η(∥x∥Cp−var,α([0,T];Z))\psi(x)=\eta\left(\|x\|_{C^{p-var,\alpha}([0,T];Z)}\right) and use the compact embedding Cp−var,α([0,T];Z)↪C0([0,T];Z)C^{p-var,\alpha}([0,T];Z)\hookrightarrow C^{0}([0,T];Z) in Theorem A.8 to conclude that KR=ψ−1((0,R])K_{R}=\psi^{-1}((0,R]) is compact with respect to ∥⋅∥C0([0,T];Z)\|\cdot\|_{C^{0}([0,T];Z)}, for all R>0R>0. Hence, (Cp−var,α([0,T];Z),ψ)(C^{p-var,\alpha}([0,T];Z),\psi) is a weighted space. Note that instead of the uniform topology we could also use any Cp′−var,α′C^{p^{\prime}-var,\alpha^{\prime}}-topology for (p′,α′)∈(p,∞]×[0,α)(p^{\prime},\alpha^{\prime})\in(p,\infty]\times[0,\alpha) due to the compact embedding Cp−var,α([0,T];Z)↪Cp′−var,α′([0,T];Z)C^{p-var,\alpha}([0,T];Z)\hookrightarrow C^{p^{\prime}-var,\alpha^{\prime}}([0,T];Z) (see Theorem A.8).

The above examples give rise to the following observations:

Note that in all the infinite dimensional examples (ii) to (ix) the idea is always to use a weaker topology than the norm topology which renders the closed unit ball compact. In the above cases this is either achieved with the weak-∗*-topology as in Example (ii), (iii), (vi), (viii), and (x), or via the following compact embeddings:

Cα(S;Z)↪Cα′(S;Z)C^{\alpha}(S;Z)\hookrightarrow C^{\alpha^{\prime}}(S;Z) for 0≤α′<α0\leq\alpha^{\prime}<\alpha as in Example (iv) and (v),

Cp−var,α([0,T];Z)↪Cp′−var,α′([0,T];Z)C^{p-var,\alpha}([0,T];Z)\hookrightarrow C^{p^{\prime}-var,\alpha^{\prime}}([0,T];Z) for (p,α)∈[1,∞)×(0,1)(p,\alpha)\in[1,\infty)\times(0,1) with pα<1p\alpha<1 and (p′,α′)∈(p,∞]×[0,α)(p^{\prime},\alpha^{\prime})\in(p,\infty]\times[0,\alpha) with p′α′<1p^{\prime}\alpha^{\prime}<1 as in Example (vii), where C∞−var,0([0,T];Z):=C0([0,T];Z)C^{\infty-var,0}([0,T];Z):=C^{0}([0,T];Z), Cp′−var,0([0,T];Z):=Cp′−var([0,T];Z)C^{p^{\prime}-var,0}([0,T];Z):=C^{p^{\prime}-var}([0,T];Z), and C∞−var,α′([0,T];Z):=Cα′([0,T];Z)C^{\infty-var,\alpha^{\prime}}([0,T];Z):=C^{\alpha^{\prime}}([0,T];Z).

LpL^{p}-embeddings of Sobolev–Slobodeckij spaces (see e.g. ).

Consider the setting of Example 2.3 (iii). Then, as shown in Theorem A.4, the weak-∗*-topology and the Cα′(S;Z)C^{\alpha^{\prime}}(S;Z)-topologies for 0≤α′<α0\leq\alpha^{\prime}<\alpha (including the uniform topology with α′=0\alpha^{\prime}=0) coincide on any ∥⋅∥Cα(S;Z)\|\cdot\|_{C^{\alpha}(S;Z)}-bounded and sequentially weak-∗*-closed subset of Cα(S;Z)C^{\alpha}(S;Z), but are globally clearly different. The same reasoning applies to Example 2.3 (vi), see Theorem A.9.

For a given weighted space (X,ψ)(X,\psi) with admissible weight function ψ:X→(0,∞)\psi:X\rightarrow(0,\infty) and a Banach space (Y,∥⋅∥)(Y,\|\cdot\|), we need to introduce an appropriate weighted function space that can be used for the subsequent global approximation theorems. Besides the vector space of bounded and continuous maps Cb0(X;Y)C^{0}_{b}(X;Y), we define the vector space

of maps f:X→Yf:X\rightarrow Y, whose growth is controlled by the growth of the weight function ψ:X→(0,∞)\psi:X\rightarrow(0,\infty). We equip Bψ(X;Y)B_{\psi}(X;Y) with the weighted norm ∥⋅∥Bψ(X;Y)\|\cdot\|_{\mathcal{B}_{\psi}(X;Y)} given by

for f∈Bψ(X;Y)f\in B_{\psi}(X;Y). Then, the continuous embedding Cb0(X;Y)↪Bψ(X;Y)C^{0}_{b}(X;Y)\hookrightarrow B_{\psi}(X;Y) holds true.

Note that if XX is compact, the weight function ψ(x)=1\psi(x)=1 is admissible. However, on general spaces, ψ:X→(0,∞)\psi:X\rightarrow(0,\infty) grows on the compact pre-images KR:=ψ−1((0,R])K_{R}:=\psi^{-1}((0,R]), which means that the elements of Bψ(X;Y)\mathcal{B}_{\psi}(X;Y) are typically unbounded, but their growth is controlled by the growth of the weight function ψ:X→(0,∞)\psi:X\rightarrow(0,\infty).

For simplicity, we always assume that the output space is a Banach space (Y,∥⋅∥Y)(Y,\|\cdot\|_{Y}). However, all the following results (including the weighted vector-valued Stone-Weierstrass theorem in Theorem 3.8) still hold true with a locally convex topological vector space (Y,τY)(Y,\tau_{Y}) as output space. In this case, the topology on YY and Bψ(X;Y)\mathcal{B}_{\psi}(X;Y) is generated by families of seminorms instead of ∥⋅∥Y\|\cdot\|_{Y} and ∥⋅∥Bψ(X;Y)\|\cdot\|_{\mathcal{B}_{\psi}(X;Y)}, respectively (see ).

The following holds true for a function f:X→Yf:X\rightarrow Y:

If f∈Bψ(X;Y)f\in\mathcal{B}_{\psi}(X;Y) then f∣KR∈C0(KR;Y)f|_{K_{R}}\in C^{0}(K_{R};Y), for all R>0R>0, and

then f∈Bψ(X)f\in\mathcal{B}_{\psi}(X). In particular, f∈Bψ(X)f\in\mathcal{B}_{\psi}(X) for every f∈C0(X)f\in C^{0}(X) satisfying (2.3).

For (i), let f∈Bψ(X;Y)f\in\mathcal{B}_{\psi}(X;Y) and fix some ε>0\varepsilon>0. Then, there exists by definition of Bψ(X;Y)\mathcal{B}_{\psi}(X;Y) some g∈Cb0(X;Y)g\in C^{0}_{b}(X;Y) such that ∥f−g∥Bψ(X;Y)<ε/2\|f-g\|_{\mathcal{B}_{\psi}(X;Y)}<\varepsilon/2. Hence, by choosing R≥2ε−1sup⁡x∈X∥g(x)∥YR\geq 2\varepsilon^{-1}\sup_{x\in X}\|g(x)\|_{Y}, it follows for every x∈X∖KRx\in X\setminus K_{R} that

Since ε>0\varepsilon>0 was chosen arbitrarily, we obtain (2.2). Moreover, with g∈Cb0(X;Y)g\in C^{0}_{b}(X;Y) from above, we observe for every R>0R>0 that

This shows that f∣KR∈C0(KR;Y)f|_{K_{R}}\in C^{0}(K_{R};Y) is continuous as uniform limit of continuous functions, for all R>0R>0. On the other hand, (ii) follows from [40, Theorem 2.7]. ∎

In addition, by following [40, Theorem 2.8], there exists for every f∈Bψ(X)f\in\mathcal{B}_{\psi}(X) with sup⁡x∈Xf(x)>0\sup_{x\in X}f(x)>0 some x0∈Xx_{0}\in X such that f(x)/ψ(x)≤f(x0)/ψ(x0)f(x)/\psi(x)\leq f(x_{0})/\psi(x_{0}), for all x∈Xx\in X, which shows the analogy to functions vanishing at infinity on locally compact spaces.

Weighted Stone-Weierstrass theorems

In this section, we generalize the Stone-Weierstrass theorems to this weighted setting with non-compact input space. For this purpose, we shortly recall the classical Stone-Weierstrass theorems formulated for compact domains, which we will need for our results.

Subsequently, the Weierstrass approximation result in Theorem 3.1 was generalized by Marshall Harvey Stone in to more general spaces by using the notion of a subalgebra. Hereby, a vector subspace A⊆C0(X)\mathcal{A}\subseteq C^{0}(X) is called a subalgebra (of C0(X)C^{0}(X)) if A\mathcal{A} is closed under multiplication, i.e. for every a1,a2∈Aa_{1},a_{2}\in\mathcal{A} it holds that a1⋅a2∈Aa_{1}\cdot a_{2}\in\mathcal{A}.

Let (X,τX)(X,\tau_{X}) be a compact Hausdorff topological space and assume that A⊆C0(X)\mathcal{A}\subseteq C^{0}(X) is a subalgebra. Then, A\mathcal{A} is dense in C0(X)C^{0}(X) if and only if A\mathcal{A} is point separating and vanishes nowhere.

Next, we state the vector-valued Stone-Weierstrass of R. Creighton Buck in . For a given subalgebra A⊆C0(X)\mathcal{A}\subseteq C^{0}(X), a vector subspace W⊆C0(X;Y)\mathcal{W}\subseteq C^{0}(X;Y) is called an A\mathcal{A}-submodule if a⋅w∈Wa\cdot w\in\mathcal{W}, for all a∈Aa\in\mathcal{A} and w∈Ww\in\mathcal{W}, where x↦(a⋅w)(x):=a(x)w(x)x\mapsto(a\cdot w)(x):=a(x)w(x). For further details on vector-valued approximation, we refer to the textbook .

Let (X,τX)(X,\tau_{X}) be a compact Hausdorff space, let A⊆C0(X)\mathcal{A}\subseteq C^{0}(X) be a subalgebra and let W⊆C0(X;Y)\mathcal{W}\subseteq C^{0}(X;Y) be an A\mathcal{A}-submodule. Then, the following are equivalent.

A\mathcal{A} is point separating and vanishes nowhere, and W(x):={w(x):w∈W}\mathcal{W}(x):=\{w(x):w\in\mathcal{W}\} is dense in YY, for all x∈Xx\in X.

Later, the vector-valued Stone-Weierstrass was extended by Erret Bishop in to a measure theoretic version. In addition, Silvio Machado derived in a quantitative version, which relies on Zorn’s lemma, see also the proofs of and .

2 Weighted real-valued Stone-Weierstrass theorem

Note that Nachbin’s setting in differs in several respects, in particular in the definition of weighted topologies. It allows for multiple upper semicontinuous weights v:X→[0,∞)v:X\rightarrow[0,\infty), e.g. v(x)=ψ(x)−1v(x)=\psi(x)^{-1}, collected in a so-called Nachbin family and can therefore represent the compact-open topology, the strict topology, and other topologies.

For a given weight function ψ:X→(0,∞)\psi:X\rightarrow(0,\infty), a subalgebra A⊆Bψ(X)\mathcal{A}\subseteq\mathcal{B}_{\psi}(X) is called point separating of ψ\psi-moderate growth if there exists a point separating vector subspace A~⊆A\widetilde{\mathcal{A}}\subseteq\mathcal{A} such that x↦exp⁡(∣a~(x)∣)∈Bψ(X)x\mapsto\exp\left(\left|\widetilde{a}(x)\right|\right)\in\mathcal{B}_{\psi}(X), for all a~∈A~\widetilde{a}\in\widetilde{\mathcal{A}}.

A subalgebra A⊆Bψ(X)\mathcal{A}\subseteq\mathcal{B}_{\psi}(X) is point separating of ψ\psi-moderate growth under any given admissible weight function ψ:X→(0,∞)\psi:X\rightarrow(0,\infty) if the subalgebra A\mathcal{A} is point separating and consists only of bounded maps. This is related to the so-called “bounded approximation problem” of Nachbin, see [79, Theorem 1].

Let A⊆Bψ(X)\mathcal{A}\subseteq\mathcal{B}_{\psi}(X) be a subalgebra that is point separating of ψ\psi-moderate growth and vanishes nowhere. Then, A\mathcal{A} is dense in Bψ(X)\mathcal{B}_{\psi}(X).

It suffices to show the approximation of a map f∈Cb0(X)f\in C^{0}_{b}(X) by an element a∈Aa\in\mathcal{A} since Cb0(X)C^{0}_{b}(X) is by definition dense in Bψ(X)\mathcal{B}_{\psi}(X). We first assume that A\mathcal{A} consists only of bounded maps, where the condition A\mathcal{A} being point separating of ψ\psi-moderate growth reduces to A\mathcal{A} being point separating. Let f∈Cb0(X)f\in C^{0}_{b}(X), fix some ε>0\varepsilon>0, and define the constants M=(inf⁡x∈Xψ(x))−1>0M=\left(\inf_{x\in X}\psi(x)\right)^{-1}>0 as well as b=sup⁡x∈X∣f(x)∣+ε/(4M)b=\sup_{x\in X}|f(x)|+\varepsilon/(4M) and R≥4b/εR\geq 4b/\varepsilon. Observe that the restriction A∣KR\mathcal{A}|_{K_{R}} is a point separating subalgebra of C0(KR)C^{0}(K_{R}), which vanishes nowhere. Hence, by using the classical real-valued Stone-Weierstrass in Theorem 3.2, there exists some a∈Aa\in\mathcal{A} such that

Hence, by combining (3.1) with (3.2), we conclude that

Since A\mathcal{A} is a subalgebra, i.e. p∘a∈Ap\circ a\in\mathcal{A}, and ε>0\varepsilon>0 as well as f∈Cb0(X)f\in C^{0}_{b}(X) were chosen arbitrarily, it follows that A⊆Cb0(X)\mathcal{A}\subseteq C^{0}_{b}(X) is dense in Bψ(X)\mathcal{B}_{\psi}(X).

For the general case of a point separating subalgebra A⊆Bψ(X)\mathcal{A}\subseteq\mathcal{B}_{\psi}(X) of ψ\psi-moderate growth, with point separating vector subspace A~⊆A\widetilde{\mathcal{A}}\subseteq\mathcal{A} such that x↦exp⁡(∣a~(x)∣)∈Bψ(X)x\mapsto\exp\left(\left|\widetilde{a}(x)\right|\right)\in\mathcal{B}_{\psi}(X), for all a~∈A~\widetilde{a}\in\widetilde{\mathcal{A}}, we first show that the maps x↦cos⁡(a~(x))x\mapsto\cos\left(\widetilde{a}(x)\right) and x↦sin⁡(a~(x))x\mapsto\sin\left(\widetilde{a}(x)\right), with a~∈A~\widetilde{a}\in\widetilde{\mathcal{A}}, belong to the ∥⋅∥Bψ(X)\|\cdot\|_{\mathcal{B}_{\psi}(X)}-closure of A\mathcal{A}. Indeed, let a~∈A~\widetilde{a}\in\widetilde{\mathcal{A}} and fix some ε>0\varepsilon>0. Then, by applying Lemma 2.7 (i), there exists some R>4/εR>4/\varepsilon such that

Since ε>0\varepsilon>0 was chosen arbitrarily, it follows that x↦cos⁡(a~(x))x\mapsto\cos\left(\widetilde{a}(x)\right) belongs to the ∥⋅∥Bψ(X)\|\cdot\|_{\mathcal{B}_{\psi}(X)}-closure of A\mathcal{A}, which holds analogously true for x↦sin⁡(a~(x))x\mapsto\sin\left(\widetilde{a}(x)\right). Thus, the subalgebra

of Bψ(X)\mathcal{B}_{\psi}(X) is contained in the ∥⋅∥Bψ(X)\|\cdot\|_{\mathcal{B}_{\psi}(X)}-closure of A\mathcal{A}. Hence, by applying the previous step to the point separating subalgebra Atrig⊆Bψ(X)\mathcal{A}_{\text{trig}}\subseteq\mathcal{B}_{\psi}(X) which vanishes nowhere and consists of bounded maps, we conclude that Atrig\mathcal{A}_{\text{trig}} is dense in Bψ(X)\mathcal{B}_{\psi}(X). However, since Atrig\mathcal{A}_{\text{trig}} is contained in the ∥⋅∥Bψ(X)\|\cdot\|_{\mathcal{B}_{\psi}(X)}-closure of A\mathcal{A}, it follows that A\mathcal{A} is also dense in Bψ(X)\mathcal{B}_{\psi}(X). ∎

3 Weighted vector-valued Stone-Weierstrass theorem

In this section, we generalize the weighted Stone-Weierstrass to the vector-valued case, where (Y,∥⋅∥Y)(Y,\|\cdot\|_{Y}) is a Banach space. For this purpose, we introduce for every w∈Bψ(X;Y)w\in\mathcal{B}_{\psi}(X;Y) the tilted weight function ψw:X→(0,∞)\psi_{w}:X\rightarrow(0,\infty) defined by ψw(x)=ψ(x)1+∥w(x)∥Y\psi_{w}(x)=\frac{\psi(x)}{1+\|w(x)\|_{Y}}, for x∈Xx\in X.

For every w∈Bψ(X;Y)w\in\mathcal{B}_{\psi}(X;Y), the weight function ψw:X→(0,∞)\psi_{w}:X\rightarrow(0,\infty) is admissible.

We fix some R>0R>0. Clearly the set Kw,R:={x∈X:ψw(x)≤R}K_{w,R}:=\left\{x\in X:\psi_{w}(x)\leq R\right\} lies in a compact set of the form KC:={x∈X:ψ(x)≤C}K_{C}:=\left\{x\in X:\psi(x)\leq C\right\} for some constant C>0C>0. A priori it is, however, not clear that Kw,RK_{w,R} is compact. It is of course sufficient to show that it is closed, since it lies in the compact set KCK_{C}. By Lemma 2.7 (i) we know that \|w\|_{Y}\big{|}_{K} is continuous, whence \left(\|w\|_{Y}\big{|}_{K}\right)^{-1}([a,b]) is closed. Therefore

Let A⊆Bψ(X)\mathcal{A}\subseteq\mathcal{B}_{\psi}(X) be a subalgebra that vanishes nowhere and is point separating of ψw\psi_{w}-moderate growth, for all w∈Ww\in\mathcal{W}, where W⊆Bψ(X;Y)\mathcal{W}\subseteq\mathcal{B}_{\psi}(X;Y) is an A\mathcal{A}-submodule such that W(x):={w(x):w∈W}\mathcal{W}(x):=\{w(x):w\in\mathcal{W}\} is dense in YY, for all x∈Xx\in X. Then, W\mathcal{W} is dense in Bψ(X;Y)\mathcal{B}_{\psi}(X;Y).

Let Bb(X)⊆Bψ(X)\mathcal{B}_{b}(X)\subseteq\mathcal{B}_{\psi}(X) be the vector subspace of bounded maps. Then, we first show that the ∥⋅∥Bψ(X;Y)\|\cdot\|_{\mathcal{B}_{\psi}(X;Y)}-closure of W\mathcal{W} is a Bb(X)\mathcal{B}_{b}(X)-submodule. For this purpose, we fix some g∈Bb(X)g\in\mathcal{B}_{b}(X) and assume that w‾:X→Y\overline{w}:X\rightarrow Y belongs to the ∥⋅∥Bψ(X;Y)\|\cdot\|_{\mathcal{B}_{\psi}(X;Y)}-closure of W\mathcal{W}. Moreover, we fix some ε>0\varepsilon>0 and define c=sup⁡x∈X∣g(x)∣<∞c=\sup_{x\in X}|g(x)|<\infty. Then, there exists some w∈Ww\in\mathcal{W} with ∥w‾−w∥Bψ(X;Y)<ε/(2c)\|\overline{w}-w\|_{\mathcal{B}_{\psi}(X;Y)}<\varepsilon/(2c), which implies that

Now, we observe that the tilted weight function ψw:X→(0,∞)\psi_{w}:X\rightarrow(0,\infty) is by Lemma 3.7 admissible. Moreover, for every a∈Aa\in\mathcal{A}, we have a∈Bψ(X)a\in\mathcal{B}_{\psi}(X) and a⋅w∈W⊆Bψ(X;Y)a\cdot w\in\mathcal{W}\subseteq\mathcal{B}_{\psi}(X;Y) as W⊆Bψ(X;Y)\mathcal{W}\subseteq\mathcal{B}_{\psi}(X;Y) is an A\mathcal{A}-submodule. Then, Lemma 2.7 (i) implies that

and that a∣KR∈C0(KR)a|_{K_{R}}\in C^{0}(K_{R}), for all R>0R>0. Thus, by using Lemma 2.7 (ii), we have a∈Bψw(X)a\in\mathcal{B}_{\psi_{w}}(X), which shows that A\mathcal{A} is also a subalgebra of Bψw(X)\mathcal{B}_{\psi_{w}}(X) as a∈Aa\in\mathcal{A} was chosen arbitrarily. Hence, by applying the weighted real-valued Stone-Weierstrass in Theorem 3.6 on Bψw(X)\mathcal{B}_{\psi_{w}}(X), there exists some a∈Aa\in\mathcal{A} with ∥g−a∥Bψw(X)<ε/2\|g-a\|_{\mathcal{B}_{\psi_{w}}(X)}<\varepsilon/2, which implies that

By combining (3.4) with (3.5), we conclude that

Since W\mathcal{W} is an A\mathcal{A}-submodule, i.e. a⋅w∈Wa\cdot w\in\mathcal{W}, and ε>0\varepsilon>0 was chosen arbitrarily, it follows that g⋅w‾:X→Yg\cdot\overline{w}:X\rightarrow Y belongs to the ∥⋅∥Bψ(X;Y)\|\cdot\|_{\mathcal{B}_{\psi}(X;Y)}-closure of W\mathcal{W}. By using that g∈Bb(X)g\in\mathcal{B}_{b}(X) and w∈Ww\in\mathcal{W} were also chosen arbitrarily, the ∥⋅∥Bψ(X;Y)\|\cdot\|_{\mathcal{B}_{\psi}(X;Y)}-closure of W\mathcal{W} is a Bb(X)\mathcal{B}_{b}(X)-submodule.

Finally, it suffices to prove the approximation of some f∈Cb0(X;Y)f\in C^{0}_{b}(X;Y) by an element w∈Ww\in\mathcal{W} since Cb0(X;Y)C^{0}_{b}(X;Y) is by definition dense in Bψ(X;Y)\mathcal{B}_{\psi}(X;Y). Let f∈Cb0(X;Y)f\in C^{0}_{b}(X;Y), fix some ε>0\varepsilon>0, define b=sup⁡x∈X∥f(x)∥Yb=\sup_{x\in X}\|f(x)\|_{Y} as well as M=max⁡(1,(inf⁡x∈Xψ(x))−1)M=\max\left(1,\left(\inf_{x\in X}\psi(x)\right)^{-1}\right) and choose R>max⁡(4(b+1)/ε,1)R>\max(4(b+1)/\varepsilon,1). We observe that the restricted subalgebra A∣KR\mathcal{A}|_{K_{R}} of C0(KR)C^{0}(K_{R}) and the restricted A∣KR\mathcal{A}|_{K_{R}}-submodule W∣KR\mathcal{W}|_{K_{R}} of C0(KR;Y)C^{0}(K_{R};Y) satisfy the assumptions of the classical vector-valued Stone-Weierstrass in Theorem 3.3 (i). Hence, there exists some w∈Ww\in\mathcal{W} such that

Now, for c=sup⁡x∈KR∥w(x)∥Yc=\sup_{x\in K_{R}}\|w(x)\|_{Y}, it follows from (3.6) that c≤ε/(4M)+bc\leq\varepsilon/(4M)+b. Moreover, we define g∈Bb(X)g\in\mathcal{B}_{b}(X) by g(x)=min⁡(1+c,1+∥w(x)∥Y)/(1+∥w(x)∥Y)g(x)=\min(1+c,1+\|w(x)\|_{Y})/(1+\|w(x)\|_{Y}), for x∈Xx\in X, which implies that g(x)=1g(x)=1, for all x∈KRx\in K_{R}. Then, g⋅w:X→Yg\cdot w:X\rightarrow Y is bounded with sup⁡x∈X∥g(x)w(x)∥Y≤1+c≤1+b+ε/(4M)\sup_{x\in X}\|g(x)w(x)\|_{Y}\leq 1+c\leq 1+b+\varepsilon/(4M). Hence, (3.6) and M,R≥1M,R\geq 1 show that

Since the ∥⋅∥Bψ(X;Y)\|\cdot\|_{\mathcal{B}_{\psi}(X;Y)}-closure of W\mathcal{W} is by the previous step a Bb(X)\mathcal{B}_{b}(X)-submodule, i.e. g⋅w:X→Yg\cdot w:X\rightarrow Y belongs to the ∥⋅∥Bψ(X;Y)\|\cdot\|_{\mathcal{B}_{\psi}(X;Y)}-closure of W\mathcal{W}, and ε>0\varepsilon>0 as well as f∈Cb0(X;Y)f\in C^{0}_{b}(X;Y) were chosen arbitrarily, it follows that W\mathcal{W} is dense in Bψ(X;Y)\mathcal{B}_{\psi}(X;Y). ∎

Theorem 3.8 can be generalized to the setting with a locally convex topological vector space (Y,τY)(Y,\tau_{Y}) as output space (see and Remark 2.6).

In Theorem 3.8, A⊂Bψ(X)\mathcal{A}\subset\mathcal{B}_{\psi}(X) is point separating of ψw\psi_{w}-moderate growth for all w∈Ww\in\mathcal{W}, if either A⊂Bψ(X)\mathcal{A}\subset\mathcal{B}_{\psi}(X) consists only of bounded functions (see Remark 3.5) or if A⊂Bψ(X)\mathcal{A}\subset\mathcal{B}_{\sqrt{\psi}}(X) is of ψ\sqrt{\psi}-moderate growth and W\mathcal{W} is a subset of Bψ(X;Y)\mathcal{B}_{\sqrt{\psi}}(X;Y). In both cases, the multiplication A×W∋(a,w)↦a⋅w∈Bψ(X;Y)\mathcal{A}\times\mathcal{W}\ni(a,w)\mapsto a\cdot w\in\mathcal{B}_{\psi}(X;Y) becomes jointly continuous.

The vector-valued Stone-Weierstrass in Theorem 3.8 is a version of Nachbin’s weighted approximation result in [79, Theorem 6]. Moreover, João B. Prolla extended in the measure theoretic version of Erret Bishop in to weighted spaces, see also .

Universal approximation on weighted spaces

First, we introduce the infinite dimensional analogue of affine maps, which are applied in classical neural networks on the hidden layer.

A subset H⊆Bψ(X)\mathcal{H}\subseteq\mathcal{B}_{\psi}(X) is called an additive family (on XX) if

H\mathcal{H} is closed under addition, i.e. for every h1,h2∈Hh_{1},h_{2}\in\mathcal{H} it holds that h1+h2∈Hh_{1}+h_{2}\in\mathcal{H},

H\mathcal{H} is point separating, i.e. for distinct x1,x2∈Xx_{1},x_{2}\in X there is h∈Hh\in\mathcal{H} with h(x1)≠h(x2)h(x_{1})\neq h(x_{2}).

We now give some examples of additive families. If the weighted space (X,ψ)(X,\psi) is a vector space equipped with a norm, then we can use the dual space of its completion.

Hence, Lemma 2.7 (ii) shows that h∈Bψ(X)h\in\mathcal{B}_{\psi}(X) and thus H⊆Bψ(X)\mathcal{H}\subseteq\mathcal{B}_{\psi}(X).

If (X,∥⋅∥0)(X,\|\cdot\|_{0}) is a Banach space in Lemma 4.9 such that ψ:X→(0,∞)\psi:X\rightarrow(0,\infty) is admissible, then XX is by Remark 2.2 (ii) finite dimensional. Hence, for an infinite dimensional weighted space (X,ψ)(X,\psi), we consider in fact an incomplete normed vector space (X,∥⋅∥0)(X,\|\cdot\|_{0}), where the dual of the ∥⋅∥0\|\cdot\|_{0}-completion is used for the additive family.

Lemma 4.9 shows that a FNN φ:X→Y\varphi:X\rightarrow Y over a weighted normed vector space (X,ψ)(X,\psi) can be chosen of the form

On the other hand, if the weighted space (X,ψ)(X,\psi) is a dual space equipped with the weak-∗*-topology as in Example 2.3 (ii), we obtain the following result.

Hence, Lemma 2.7 (ii) shows that h∈Bψ(X)h\in\mathcal{B}_{\psi}(X) and thus H⊆Bψ(X)\mathcal{H}\subseteq\mathcal{B}_{\psi}(X).

Lemma 4.11 shows that a FNN φ:X→Y\varphi:X\rightarrow Y over a weighted dual space (X,ψ)(X,\psi) with isometric isomorphism Φ:X→E∗\Phi:X\rightarrow E^{*} and equipped with the weak-∗*-topology is given by

Let us give some examples of functional input neural networks, where we revisit among others the examples from Section 1.1. Hereby, η:[0,∞)→(0,∞)\eta:[0,\infty)\rightarrow(0,\infty) is a continuous increasing function satisfying lim⁡r→∞rη(r)=0\lim_{r\rightarrow\infty}\frac{r}{\eta(r)}=0.

For the space of α\alpha-Hölder continuous functions X=Cα(S)X=C^{\alpha}(S) as in Example 2.3 (iv) with weight function ψ(x)=η(∥x∥Cα(S))\psi(x)=\eta\left(\|x\|_{C^{\alpha}(S)}\right), an additive family is given by

For the space of signed Radon measures X=MψΩ(Ω)X=\mathcal{M}_{\psi_{\Omega}}(\Omega) on a weighted space (Ω,ψΩ)(\Omega,\psi_{\Omega}) as in Example 2.3 (x) with weight function ψ(x)=η(∥x∥MψΩ(Ω))\psi(x)=\eta\left(\|x\|_{\mathcal{M}_{\psi_{\Omega}}(\Omega)}\right), an additive family is given by

This can be used to learn a map with a (probability) measure as input.

These results give an overview of the broad variety of functional input neural networks, for which we show the universal approximation property in the following subsection.

2 Global universal approximation

Universal approximation results were first proved for neural networks between Euclidean spaces in and in using functional analytical arguments such as the Hahn-Banach theorem and Fourier theory. Later, further universal approximation results were found for different types of activation functions, such as in , while the results of e.g. , and are concerned with proving quantitative approximation rates under more restrictive assumptions on the involved functions.

Now, for the function x↦w(x):=cos⁡(h(x))y0∈Wx\mapsto w(x):=\cos(h(x))y_{0}\in\mathcal{W}, we define the corresponding FNN x↦φ(x):=y0ρ(h(x))∈NNX,YH,ρ,Lx\mapsto\varphi(x):=y_{0}\rho(h(x))\in\mathcal{NN}^{\mathcal{H},\rho,\mathcal{L}}_{X,Y}. Then, we conclude from (4.2) that

Since ε>0\varepsilon>0 was chosen arbitrarily, it follows that the map x↦w(x)=cos⁡(h(x))y0∈Wx\mapsto w(x)=\cos(h(x))y_{0}\in\mathcal{W} belongs to the ∥⋅∥Bψ(X;Y)\|\cdot\|_{\mathcal{B}_{\psi}(X;Y)}-closure of NNX,YH,ρ,L\mathcal{NN}^{\mathcal{H},\rho,\mathcal{L}}_{X,Y}, which holds analogously true for the map x↦sin⁡(h(x))y0∈Wx\mapsto\sin(h(x))y_{0}\in\mathcal{W}. Hence, we conclude by the triangle inequality that linear combinations of these maps belong to the ∥⋅∥Bψ(X;Y)\|\cdot\|_{\mathcal{B}_{\psi}(X;Y)}-closure of NNX,YH,ρ,L\mathcal{NN}^{\mathcal{H},\rho,\mathcal{L}}_{X,Y}, which shows that the entire A\mathcal{A}-submodule W\mathcal{W} is contained in the ∥⋅∥Bψ(X;Y)\|\cdot\|_{\mathcal{B}_{\psi}(X;Y)}-closure of NNX,YH,ρ,L\mathcal{NN}^{\mathcal{H},\rho,\mathcal{L}}_{X,Y}.

Finally, we apply the weighted vector-valued Stone-Weierstrass in Theorem 3.8 to show that NNX,YH,ρ,L\mathcal{NN}^{\mathcal{H},\rho,\mathcal{L}}_{X,Y} is dense in Bψ(X;Y)\mathcal{B}_{\psi}(X;Y). For this purpose, we first see that A⊆Bψ(X)\mathcal{A}\subseteq\mathcal{B}_{\psi}(X) vanishes nowhere as it contains the non-zero constant map x↦a(x):=cos⁡(0)=1∈Ax\mapsto a(x):=\cos(0)=1\in\mathcal{A}. Moreover, we observe that A⊆Bψ(X)\mathcal{A}\subseteq\mathcal{B}_{\psi}(X) is point separating and consists only of bounded maps, which shows that A⊆Bψ(X)\mathcal{A}\subseteq\mathcal{B}_{\psi}(X) is point separating of ψw\psi_{w}-moderate growth, for all w∈Ww\in\mathcal{W} (see Remark 3.5). In addition, by using that L\mathcal{L} is dense in YY, it follows that W(x)={w(x):w∈W}=L\mathcal{W}(x)=\left\{w(x):w\in\mathcal{W}\right\}=\mathcal{L} is dense in YY, for all x∈Xx\in X. Hence, by applying the weighted vector-valued Stone-Weierstrass in Theorem 3.8, we conclude that W\mathcal{W} is dense in Bψ(X;Y)\mathcal{B}_{\psi}(X;Y). However, since W\mathcal{W} is by the previous step contained in the ∥⋅∥Bψ(X;Y)\|\cdot\|_{\mathcal{B}_{\psi}(X;Y)}-closure of NNX,YH,ρ,L\mathcal{NN}^{\mathcal{H},\rho,\mathcal{L}}_{X,Y}, it follows that NNX,YH,ρ,L\mathcal{NN}^{\mathcal{H},\rho,\mathcal{L}}_{X,Y} is also dense in Bψ(X;Y)\mathcal{B}_{\psi}(X;Y). ∎

The global universal approximation result in Theorem 4.13 can be generalized to locally convex topological vector space as output spaces (see for a study of these infinite dimensional vector spaces and also Remark 2.6 and Remark 3.9).

3 Global universal approximation for non-anticipative functionals

The space of stopped paths was originally introduced in , and to generalize Föllmer’s pathwise Ito calculus to non-anticipative functionals. By introducing a special type of FNNs, so-called non-anticipative functional input neural networks, we now present the universal approximation result for this kind of functionals.

We consider again the space of stopped α\alpha-Hölder continuous paths given by

Hence, we can apply the global universal approximation result below at least for continuous non-anticipative functionals, whose growth is controlled by ψ:ΛTα→(0,∞)\psi:\Lambda^{\alpha}_{T}\rightarrow(0,\infty).

Motivated by Example 4.12 (ii), we introduce a specific type of functional input neural network defined on ΛTα\Lambda^{\alpha}_{T}, which preserve the non-anticipative behaviour.

We now apply the global universal approximation result in Theorem 4.13 to show that every f∈Bψ(ΛTα)f\in\mathcal{B}_{\psi}(\Lambda^{\alpha}_{T}) can be approximated by some non-anticipative FNN.

Hence, the conclusion follows from Theorem 4.13. ∎

In Section 7, we provide some numerical examples, which illustrate how the theoretical approximation result in Corollary 4.17 can be applied to learn a non-anticipative functional by a non-anticipative FNNs. In particular for applications in finance, non-anticipative FNNs can be used within the framework of Stochastic Portfolio Theory (SPT) introduced by R. Fernholz in to learn optimal path-dependent portfolios, see .

Universal approximation on weighted spaces for linear functions of the signature

In this section, we give another application of the weighted real-valued Stone-Weierstrass Theorem 3.6, which is related to the theory of rough paths introduced by Terry Lyons in (we also refer to the textbooks and ). More precisely, we will show that a path space functional can be approximated globally by using linear functions of the paths’ signature, a notion which goes back to the work of Chen . This is in contrast to the usual results where only compact sets of paths are considered. Different versions of the latter, e.g. for finite variation paths or for continuous functions depending on the whole signature (instead of the rough path), are available in the literature (see for instance [69, Theorem 3.1], [62, Theorem 1], [72, Proposition 4.5] and [28, Section 3]). We refer to for the situation involving not just the approximation of a functional up to a fixed final value T>0T>0, but a uniform approximation on the whole time interval [0,T][0,T].

In general, one needs two key properties for universal approximation results for linear functions of the signature. First, that the signature at the terminal time determines uniquely the path (up to so-called tree-like equivalences, see and ), which ensures point separation. Second, that every polynomial of the signature can be realized as a linear function via the so-called shuffle product, which yields the algebra property of linear functions of the signature. While the classical Stone-Weierstrass then only guarantees the approximation on compact subsets of paths (e.g. with respect to a certain pp-variation metric), there have already been some attempts to go beyond this rather restrictive setting. As already mentioned in the introduction, used the strict topology originating from to formulate a global approximation result, however not with the “true” signature but with a bounded normalization implying that the tractability properties of the signature (like computing expected signature analytically as e.g. done in for generic classes of stochastic processes) are lost.

In contrast, our weighted space setup allows to work with the “true” signature and thus to preserve its tractability properties, while we can still provide a global approximation result. Hereby, we consider the path space of α\alpha-Hölder continuous paths as weighted space, which we either equip with an 0≤α′<α0\leq\alpha^{\prime}<\alpha-Hölder topology or the weak-∗*-topology as considered in Appendix A.

where eI\shuffleeJ:=∑k=1KeIke_{I}\shuffle e_{J}:=\sum_{k=1}^{K}e_{I_{k}}, with KK and (Ik)k=1,…,K(I_{k})_{k=1,\ldots,K} being determined via I\shuffleJ=∑k=1KIkI\shuffle J=\sum_{k=1}^{K}I_{k}.

2 Global universal approximation of α𝛼\alpha-Hölder rough paths

The following lemma relates dcc,∞d_{cc,\infty} and dcc,0d_{cc,0} and can be deduced from the estimates in [47, Proposition 8.15].

To ensure point separation (without tree-like equivalences), we shall always add an additional component to XX representing time. More precisely, we define the subspace

To get a weighted space, we choose the weight function ψ:C^d,Tα→(0,∞)\psi:\widehat{C}^{\alpha}_{d,T}\rightarrow(0,\infty) defined by

Hence, the weight function ψ:C^d,Tα→(0,∞)\psi:\widehat{C}^{\alpha}_{d,T}\rightarrow(0,\infty) is on both spaces admissible and it follows from Remark A.7 (iv) that the weighted function space Bψ(C^d,Tα)\mathcal{B}_{\psi}(\widehat{C}^{\alpha}_{d,T}) is the same for each underlying space (C^d,Tα,τw∗)(\widehat{C}^{\alpha}_{d,T},\tau_{w^{*}}) and (C^d,Tα,dcc,α′)(\widehat{C}^{\alpha}_{d,T},d_{cc,\alpha^{\prime}}), with 0≤α′<α0\leq\alpha^{\prime}<\alpha.

Now, we present the global universal approximation theorem for linear functions of signature, which can approximate any path space functional in Bψ(C^d,Tα)\mathcal{B}_{\psi}(\widehat{C}^{\alpha}_{d,T}). Even though the proof looks involved, the most important ingredients can be summarized as follows: certain linear functions of the signature are given by the integrals

Since Bψ(C^d,Tα)\mathcal{B}_{\psi}(\widehat{C}^{\alpha}_{d,T}) is the same for each underlying space (C^d,Tα,τw∗)(\widehat{C}^{\alpha}_{d,T},\tau_{w^{*}}) and (C^d,Tα,dcc,α′)(\widehat{C}^{\alpha}_{d,T},d_{cc,\alpha^{\prime}}), with 0≤α′<α0\leq\alpha^{\prime}<\alpha (see Remark A.7 (iv)), and the metrics dcc,0d_{cc,0} and dcc,∞d_{cc,\infty} are equivalent (see Lemma 5.3), we choose (C^d,Tα,dcc,∞)(\widehat{C}^{\alpha}_{d,T},d_{cc,\infty}) in the following. Then, the result follows from the weighted real-valued Stone-Weierstrass in Theorem 3.6 applied to the set

Therefore, we need to prove that A\mathcal{A} is a subspace of Bψ(C^d,Tα)\mathcal{B}_{\psi}(\widehat{C}^{\alpha}_{d,T}), and a point separating subalgebra of ψ\psi-moderate growth, which vanishes nowhere, where

is a possible candidate for the point separating vector subspace of ψ\psi-moderate growth.

is continuous on KRK_{R} with respect to dcc,∞d_{cc,\infty}. This together with the continuity of the evaluation map

is continuous on KRK_{R} with respect to dcc,∞d_{cc,\infty}. Furthermore, since linear functions are continuous, it follows that the map

is continuous on KRK_{R} with respect to dcc,∞d_{cc,\infty}. Since R>0R>0 was chosen arbitrarily, this shows that a∣KR∈C0(KR)a|_{K_{R}}\in C^{0}(K_{R}), for all R>0R>0. Moreover, by using the inequality

Since the exponential function dominates any polynomial, we conclude that

Hence, it follows from Lemma 2.7 (ii) that a∈Bψ(C^d,Tα)a\in\mathcal{B}_{\psi}(\widehat{C}^{\alpha}_{d,T}), which shows that A⊆Bψ(C^d,Tα)\mathcal{A}\subseteq\mathcal{B}_{\psi}(\widehat{C}^{\alpha}_{d,T}).

where the last equality follows from the fact that the exponent tends to −∞-\infty as NN is at most ⌊1α⌋\lfloor\frac{1}{\alpha}\rfloor and γ>⌊1α⌋\gamma>\lfloor\frac{1}{\alpha}\rfloor. Hence, Lemma 2.7 (ii) shows that exp⁡(∣a~(⋅)∣)∈Bψ(C^d,Tα)\exp(|\widetilde{a}(\cdot)|)\in\mathcal{B}_{\psi}(\widehat{C}^{\alpha}_{d,T}), which holds true for any a~∈A~\widetilde{a}\in\widetilde{\mathcal{A}}.

Now, to show that A~\widetilde{\mathcal{A}} is point separating, let Y^,Z^∈C^d,Tα\widehat{\mathbf{Y}},\widehat{\mathbf{Z}}\in\widehat{C}^{\alpha}_{d,T} be distinct. By contradiction, let us assume that for every k,N∈{1,…,⌊1α⌋}k,N\in\{1,\ldots,\lfloor\frac{1}{\alpha}\rfloor\} and I∈{0,…,d}NI\in\{0,\ldots,d\}^{N} it holds that

for all X^∈C^d,Tα\widehat{\mathbf{X}}\in\widehat{C}^{\alpha}_{d,T}. Thus, we conclude for every k,N∈{1,…,⌊1α⌋}k,N\in\{1,\ldots,\lfloor\frac{1}{\alpha}\rfloor\} and I∈{0,…,d}NI\in\{0,\ldots,d\}^{N} that

By [23, Theorem 7], Theorem 5.4 thus also implies that the signature is characteristic for the measure space Mψ(C^d,Tα)≅Bψ(C^d,Tα)∗\mathcal{M}_{\psi}(\widehat{C}^{\alpha}_{d,T})\cong\mathcal{B}_{\psi}(\widehat{C}^{\alpha}_{d,T})^{*} introduced in Example 2.3 (x). Indeed, the map

Laws of stochastic processes X^\widehat{\mathbf{X}} on path space are thus characterized by their expected signature if the following (super)-exponential moment condition is satisfied.

Let us remark that an alternative sufficient condition stating when the law of a group-valued random variable is characterized by its expected signature is given in [22, Corollary 6.6].

Alternatively to linear functions of the signature we can consider

which is again by the Stone-Weierstrass Theorem 3.6 dense in B1(C^d,Tα)=Cb0(C^d,Tα)\mathcal{B}_{1}(\widehat{C}^{\alpha}_{d,T})=C_{b}^{0}(\widehat{C}^{\alpha}_{d,T}). We therefore have universality of the corresponding feature map and

3 Global universal approximation of p𝑝p-variation rough paths

In this section, we provide the universal approximation result for weakly geometric pp-variation rough paths (see [47, Definition 9.15 (i)]). However, in order to get a weighted space, we need to consider the subspace of weakly geometric pp-variation rough paths which are also Hölder continuous (see Example 2.3 (vi) and (vii) as well as Remark 2.4).

To get a weighted space, we choose similarly as above the weight function ψ:C^d,Tp−var,α→(0,∞)\psi:\widehat{C}^{p-var,\alpha}_{d,T}\rightarrow(0,\infty) defined by

for X^∈C^d,Tp−var,α\widehat{\mathbf{X}}\in\widehat{C}^{p-var,\alpha}_{d,T} and some β>0\beta>0 and γ>⌊1α⌋\gamma>\lfloor\frac{1}{\alpha}\rfloor. Then, by the arguments as for C^d,Tα\widehat{C}^{\alpha}_{d,T} but now using Remark A.12, we conclude that ψ:C^d,Tp−var,α→(0,∞)\psi:\widehat{C}^{p-var,\alpha}_{d,T}\rightarrow(0,\infty) is admissible on both spaces (C^d,Tp−var,α,τw∗)(\widehat{C}^{p-var,\alpha}_{d,T},\tau_{w^{*}}) and (C^d,Tp−var,α,dcc,p′−var,α′)(\widehat{C}^{p-var,\alpha}_{d,T},d_{cc,p^{\prime}-var,\alpha^{\prime}}), with (p′,α′)∈(p,∞]×[0,α)(p^{\prime},\alpha^{\prime})\in(p,\infty]\times[0,\alpha) and p′α′<1p^{\prime}\alpha^{\prime}<1. Moreover, Bψ(C^d,Tp−var,α)\mathcal{B}_{\psi}(\widehat{C}^{p-var,\alpha}_{d,T}) does not depend on the choice of the underlying topology.

Now, we present our global universal approximation theorem for linear functions of the signature in Bψ(C^d,Tp−var,α)\mathcal{B}_{\psi}(\widehat{C}^{p-var,\alpha}_{d,T}).

Corollary 5.7 shows that every path space functional in Bψ(C^d,Tp−var,α)\mathcal{B}_{\psi}(\widehat{C}^{p-var,\alpha}_{d,T}) can be learned with linear functions of the signature. Moreover, by following Remark 5.5, we observe that the signature is characteristic for the dual of Bψ(C^d,Tp−var,α)\mathcal{B}_{\psi}(\widehat{C}^{p-var,\alpha}_{d,T}).

Gaussian process regression with applications to signature kernels

As an important counterpart to the so far established density results, which make linear regressions feasible, we introduce a Gaussian process perspective, or, equivalently, the perspective of reproducing kernel Hilbert spaces on regression. Even though it can be done for general input spaces and general Banach spaces of functions thereon, we specialize here to weighted spaces as input spaces and Bψ\mathcal{B}_{\psi}-spaces as spaces of functions thereon. An important application is given through signature kernel regression on path spaces. The corresponding (general) signature kernels have already been considered in and are of the form

for appropriate choices of aIa_{I}. We provide here a novel Gaussian process perspective in terms of Bψ\mathcal{B}_{\psi}-spaces, which together with the above results allows to treat approximation with the true signature kernels (without a normalization procedure as considered in ) rigorously. Consequently, uncertainty quantification results and a theory of regularization for regressions on signature components can be obtained.

Let (X,ψ)(X,\psi) be a general weighted space. A reproducing kernel Hilbert space HH on (X,ψ)(X,\psi) is a continuously embedded subspace of H⊂Bψ(X)H\subset\mathcal{B}_{\psi}(X). Its kernel kk is uniquely defined through ∫Xf(y)δx(dy)=f(x)=⟨k(x,⋅),f⟩\int_{X}f(y)\delta_{x}(dy)=f(x)=\langle k(x,\cdot),f\rangle for all f∈Hf\in H, i.e. the functional representation in HH of the Dirac measure δx\delta_{x} located at xx via Riesz representation. Whence the kernel is symmetric, positive semi-definite and k(x,.)∈Bψ(X)k(x,.)\in\mathcal{B}_{\psi}(X) for x∈Xx\in X. If, on the other hand, we are given a symmetric, positive semi-definite function kk such that k(x,.)∈Bψ(X)k(x,.)\in\mathcal{B}_{\psi}(X) for x∈Xx\in X, and such that the span of k(x,⋅)k(x,\cdot) with pre-scalar product ⟨k(x1,⋅),k(x2,⋅)⟩:=k(x1,x2)\langle k(x_{1},\cdot),k(x_{2},\cdot)\rangle:=k(x_{1},x_{2}) is continuously embedded in Bψ(X)\mathcal{B}_{\psi}(X), then the span’s closure is a reproducing kernel Hilbert space on (X,ψ)(X,\psi) with kernel kk.

We call a kernel universal Bψ(X)\mathcal{B}_{\psi}(X) if the corresponding reproducing kernel Hilbert space HH on (X,ψ)(X,\psi) is dense in Bψ(X)\mathcal{B}_{\psi}(X). Since the signature feature map is universal (see Remark 5.5), it follows from [23, Proposition 29 and Proposition 43] that the (general) signature kernels as of form (6.1) are universal to Bψ(X)\mathcal{B}_{\psi}(X) with XX and ψ\psi as in Sections 5.2 and 5.3.

Let ZZ be a Bψ(X)\mathcal{B}_{\psi}(X)-valued centered Gaussian process and assume Bψ(X)\mathcal{B}_{\psi}(X) to be separable. Notice that separability is never a restriction if we are given a countable set of functions, like signature components, whose span (built with rational coefficients) is dense in Bψ(X)\mathcal{B}_{\psi}(X), which, in other words, is the typical situation for linear regression.

The Cameron-Martin theorem then states that f∈HZf\in H_{Z} if and only if the law of f+Zf+Z is absolutely continuous to the law of ZZ. The well-known Radon-Nikodym derviative is given by \exp\big{(}(f\bullet Z)-\frac{1}{2}\|f\|_{H_{Z}}\big{)}.

In this context the meaning of the Gaussian process ZZ also corresponds to the generation of a prior distribution of functions, which are added to the searched function in question.

With these preparations we can now formulate the main theorem of this section, which has immediate applications for signature regression:

Given a kernel kk on (X,ψ)(X,\psi) and a countable system of functions (hi)(h_{i}) in Bψ(X)\mathcal{B}_{\psi}(X) such k=∑ihi⊗hik=\sum_{i}h_{i}\otimes h_{i} as function on (X,ψ)(X,\psi) and ∑i∥hi∥Bψ(X)<∞\sum_{i}\|h_{i}\|_{\mathcal{B}_{\psi}(X)}<\infty with Bψ(X)\mathcal{B}_{\psi}(X) separable. Then the random series ∑iZihi\sum_{i}Z_{i}h_{i}, for an i.i.d. sequence of standard Gaussian random variable (Zi)(Z_{i}), converges in the Bψ(X)\mathcal{B}_{\psi}(X)-norm to a Gaussian process ZZ with values in Bψ(X)\mathcal{B}_{\psi}(X) and with kernel kk. In particular its Cameron-Martin space coincides with the reproducing kernel Hilbert space generated by kk.

Under the condition ∑i∥hi∥Bψ(X)<∞\sum_{i}\|h_{i}\|_{\mathcal{B}_{\psi}(X)}<\infty the Ito-Nisio theorem, see e.g. [68, Theorem 2.4], guarantees that ∑iZihi\sum_{i}Z_{i}h_{i}, for an i.i.d. sequence of standard Gaussian random variable (Zi)(Z_{i}), converges to a random variable ZZ, which is again a Gaussian random variable. Its kernel can be easily calculated by pointwise covariances

and coincides with kk by assumption. The pointwise covariances, however, determine CZC_{Z} and in turn the Cameron-Martin space as stated above. ∎

With this theorem uncertainty quantification for regression with signature kernels is now feasible, since for reasonable choices of ψ\psi the main condition of the theorem is satisfied. We shall show this in Remark 6.5 in the setting of Section 5.3.

Take X=C^d,Tp−var,αX=\widehat{C}^{p-var,\alpha}_{d,T} together with

where aIa_{I} are real numbers. More precisely, for 0<η<10<\eta<1 satisfying γ(1−η)>1\gamma(1-\eta)>1, we require for aIa_{I} that

where CpC_{p} is some constant depending on pp and

holds. Using ak:=max⁡∣I∣=k∣aI∣a_{k}:=\max_{|I|=k}|a_{I}|, Hölder’s inequality with 1/r+1/γ=11/r+1/\gamma=1, and 0<η<10<\eta<1 with γ(1−η)>1\gamma(1-\eta)>1, we can estimate

where the second last inequality follows from the estimate exp⁡(x)≥xkk!\exp(x)\geq\frac{x^{k}}{k!} for x≥0x\geq 0. The root criterion then implies that the expression is finite. Indeed, noticing that for ϵ>0\epsilon>0, we have (k!)ϵ/k→∞(k!)^{\epsilon/k}\to\infty as k→∞k\to\infty, it follows that the first term is finite due to the assumption that ak=max⁡∣I∣=k∣aI∣≤M(k!)δa_{k}=\max_{|I|=k}|a_{I}|\leq M(k!)^{\delta} for δ<η\delta<\eta and some constant M>0M>0. Similarly the second term is finite as well.

Notice also that one could replace ψ\psi by a function of an absolutely converging sum of the form

with positive bkb_{k} such that bk≥Mϵkk!κb_{k}\geq\frac{M\epsilon^{k}}{k!^{\kappa}} for k≥0k\geq 0, M>0M>0, ϵ>0\epsilon>0 and κ<γ(1−η)−1\kappa<\gamma(1-\eta)-1.

Under the above conditions on aIa_{I} and ψ\psi the kernel

is thus the covariance of a Gaussian process with values in Bψ(X)\mathcal{B}_{\psi}(X).

Choosing aI=a∣I∣a_{I}=a_{|I|} identical for all II of the same length and defining

for bounded variation paths X^,Y^\widehat{\mathbf{X}},\widehat{\mathbf{Y}}, we then get by the same arguments as in [18, Proposition 2.8] the following integral equation

where (a2)+1(a^{2})_{+1} denotes the sequence of a2a^{2} shifted by 11, i.e.

Note that for both a2a^{2} and a+12a^{2}_{+1} Condition 1 of which is required in [18, Proposition 2.8] is satisfied due to (6.2). In particular, k(X^,Y^)k(\widehat{\mathbf{X}},\widehat{\mathbf{Y}}) satisfies

If a∣I∣≡1a_{|I|}\equiv 1 and X^,Y^\widehat{\mathbf{X}},\widehat{\mathbf{Y}} are differentiable, then the original signature kernel satisfies the so-called Goursat PDE

Numerical examples

In this section, we illustrate in two particular examplesThe experiments are implemented in Python using tensorflow (for FNN) and iisignature (for signatures) on a Lenovo ThinkPad X13 Gen2a with AMD Ryzen 7 PRO 5850U processor and Radeon Graphics (1901 Mhz, 8 Cores, 16 Logical Processors), see https://github.com/psc25/GlobalUAT. how path space functionals can be learned using non-anticipative functional input neural networks (see Section 4.3) or a linear function of the signature (see Section 5).

As input data we generate M=50000M=50000 sample paths of a one-dimensional Brownian motion x(m)=(x(m)(t))t∈[0,T]x^{(m)}=(x^{(m)}(t))_{t\in[0,T]}, for m=1,…,Mm=1,\ldots,M, with T=1T=1, which are discretized over K=100K=100 equidistant time points (tk)k=1,…,K(t_{k})_{k=1,\ldots,K}. Since the sample paths of Brownian motion are a.s. continuous and (1/2−ε)(1/2-\varepsilon)-Hölder continuous, for all 0<ε<1/20<\varepsilon<1/2, we consider the weighted space of stopped α\alpha-Hölder continuous paths ΛTα\Lambda^{\alpha}_{T} in Example 2.3 (v), with α<1/2\alpha<1/2.

We split up the data into 80%80\% for training and 20%20\% for testing, and apply stochastic gradient descent with the Adam algorithm (see ) over 20002000 epochs with learning rate 10−510^{-5} and batchsize 500500 to minimize the mean squared error (MSE)

For the FNNs, we consider φ∈FNΛTαρ,ρ~\varphi\in\mathcal{FN}^{\rho,\widetilde{\rho}}_{\Lambda^{\alpha}_{T}} with NFNN=30N_{FNN}=30 neurons (where ρ(z)=ρ~(z)=max⁡(z,0)\rho(z)=\widetilde{\rho}(z)=\max(z,0) are both the ReLU function and the classical neural networks (ϕn,1)n=1,…,NFNN(\phi_{n,1})_{n=1,\ldots,N_{FNN}} have one hidden layer of N1=20N_{1}=20 neurons), while for signature, we choose NSig=7N_{Sig}=7.

Appendix A Predual of Banach spaces

In the following, we apply the result of to characterize a Banach space (X,∥⋅∥X)(X,\|\cdot\|_{X}) as dual space, which formally generalizes the Dixmier-Ng theorem (see and ).

Let (X,∥⋅∥X)(X,\|\cdot\|_{X}) be a Banach space, and let E0⊆X∗E_{0}\subseteq X^{*} be a set of continuous linear functionals such that

Then, XX is a dual space and a predual of XX is the ∥⋅∥X∗\|\cdot\|_{X^{*}}-closure of span⁡(E0)\operatorname*{span}(E_{0}).

Now, we apply Theorem A.1 to show the existence of a predual for the Banach spaces presented in Example 2.3, i.e. Hölder spaces and spaces of finite pp-variation.

For every 0≤α′<α0\leq\alpha^{\prime}<\alpha, the embedding Cα(S;Z)↪Cα′(S;Z)C^{\alpha}(S;Z)\hookrightarrow C^{\alpha^{\prime}}(S;Z) is compact.

Fix some 0≤α′<α0\leq\alpha^{\prime}<\alpha and define the constant C0:=sup⁡s∈SdS(s,0)<∞C_{0}:=\sup_{s\in S}d_{S}(s,0)<\infty. Then, by using the triangle inequality of ∥⋅∥Z\|\cdot\|_{Z}, we conclude for every x∈Cα(S;Z)x\in C^{\alpha}(S;Z) that

which shows that the embedding Cα(S;Z)↪Cα′(S;Z)C^{\alpha}(S;Z)\hookrightarrow C^{\alpha^{\prime}}(S;Z) is continuous.

where ∥x−xnk∥0:=sup⁡s,t∈S∥(x−xnk)(s)−(x−xnk)(t)∥Z≤2∥x−xnk∥C0(S;Z)\|x-x_{n_{k}}\|_{0}:=\sup_{s,t\in S}\|(x-x_{n_{k}})(s)-(x-x_{n_{k}})(t)\|_{Z}\leq 2\|x-x_{n_{k}}\|_{C^{0}(S;Z)}. This shows that the embedding Cα(S;Z)↪Cα′(S;Z)C^{\alpha}(S;Z)\hookrightarrow C^{\alpha^{\prime}}(S;Z) is compact. ∎

For every α>0\alpha>0, the Banach space (Cα(S;Z),∥⋅∥Cα(S;Z))(C^{\alpha}(S;Z),\|\cdot\|_{C^{\alpha}(S;Z)}) is a dual space. Moreover, on every ∥⋅∥Cα(S;Z)\|\cdot\|_{C^{\alpha}(S;Z)}-bounded and sequentially weak-∗*-closed subset of Cα(S;Z)C^{\alpha}(S;Z), the weak-∗*-topology and every Cα′C^{\alpha^{\prime}}-topology induced by ∥⋅∥Cα′(S;Z)\|\cdot\|_{C^{\alpha^{\prime}}(S;Z)}, 0≤α′<α0\leq\alpha^{\prime}<\alpha, are equivalent.

Let (V,∥⋅∥V)(V,\|\cdot\|_{V}) be a predual of (Z,∥⋅∥Z)(Z,\|\cdot\|_{Z}). We want to apply Theorem A.1 with

On the other hand, the closed unit ball B1(0)‾:={x∈Cα(S;Z):∥x∥Cα(S;Z)≤1}\overline{B_{1}(0)}:=\left\{x\in C^{\alpha}(S;Z):\|x\|_{C^{\alpha}(S;Z)}\leq 1\right\} is by Theorem A.3 relatively compact in C0(S;Z)C^{0}(S;Z). Since the weak topology on Cα(S;Z)C^{\alpha}(S;Z) induced by E0⊆Cα(S;Z)∗E_{0}\subseteq C^{\alpha}(S;Z)^{*} is weaker than the topology induced by ∥⋅∥C0(S;Z)\|\cdot\|_{C^{0}(S;Z)}, it follows that B1(0)‾\overline{B_{1}(0)} is compact with respect to this weak topology. Hence, by using Theorem A.1, the ∥⋅∥Cα(S;Z)∗\|\cdot\|_{C^{\alpha}(S;Z)^{*}}-closure of span⁡(E0)\operatorname*{span}(E_{0}), denoted by EE, is a predual of Cα(S;Z)C^{\alpha}(S;Z).

Now, we denote by τw∗B\tau^{B}_{w^{*}} and τα′B\tau^{B}_{\alpha^{\prime}} the subspace topology on B⊆Cα(S;Z)B\subseteq C^{\alpha}(S;Z) of the weak-∗*-topology and of the Cα′C^{\alpha^{\prime}}-topology induced by ∥⋅∥Cα′(S;Z)\|\cdot\|_{C^{\alpha^{\prime}}(S;Z)}, respectively. Since (B,τα′B)(B,\tau^{B}_{\alpha^{\prime}}) is first countable, the previous argument implies that the identity id⁡:(B,τα′B)→(B,τw∗B)\operatorname{id}:(B,\tau^{B}_{\alpha^{\prime}})\rightarrow(B,\tau^{B}_{w^{*}}) is continuous, which shows that τw∗B⊆τα′B\tau^{B}_{w^{*}}\subseteq\tau^{B}_{\alpha^{\prime}}. Moreover, by using that (B,τα′B)(B,\tau^{B}_{\alpha^{\prime}}) is compact and that (B,τw∗B)(B,\tau^{B}_{w^{*}}) is Hausdorff, it follows from [78, Theorem 26.6] that the inverse id⁡−1:(B,τw∗B)→(B,τα′B)\operatorname{id}^{-1}:(B,\tau^{B}_{w^{*}})\rightarrow(B,\tau^{B}_{\alpha^{\prime}}) is also continuous, which implies τα′B⊆τw∗B\tau^{B}_{\alpha^{\prime}}\subseteq\tau^{B}_{w^{*}}. This shows that τw∗B=τα′B\tau^{B}_{w^{*}}=\tau^{B}_{\alpha^{\prime}} and completes the proof. ∎

The point evaluations in (A.2) are also used in and to construct the Arens-Eells space \AE(S)\AE(S) and the Lipschitz-free space F(S)\mathcal{F}(S), respectively, which are both used as preduals of the space of globally Lipschitz continuous functions.

We now give some examples of ∥⋅∥Cα(S;Z)\|\cdot\|_{C^{\alpha}(S;Z)}-bounded and sequentially weak-∗*-closed subsets B⊆Cα(S;Z)B\subseteq C^{\alpha}(S;Z).

For every R>0R>0 and sequentially weak-∗*-closed subset G⊆ZG\subseteq Z, the set

is ∥⋅∥Cα(S;Z)\|\cdot\|_{C^{\alpha}(S;Z)}-bounded and sequentially weak-∗*-closed. In particular, by choosing G=ZG=Z, the closed α\alpha-Hölder ball BRZ(0)‾\overline{B^{Z}_{R}(0)} is ∥⋅∥Cα(S;Z)\|\cdot\|_{C^{\alpha}(S;Z)}-bounded and sequentially weak-∗*-closed.

Choosing now the weight function ψ(x):=η(∥x∥Cα(S;Z))\psi(x):=\eta\left(\|x\|_{C^{\alpha}(S;Z)}\right), with a continuous increasing function η:[0,∞)→(0,∞)\eta:[0,\infty)\rightarrow(0,\infty), we consider the weighted space (Cα(S;Z),ψ)(C^{\alpha}(S;Z),\psi) either equipped with the weak-∗*-topology as in Example 2.3 (iii) or with the Cα′C^{\alpha^{\prime}}-topology as in Example 2.3 (iv). Then, by using that these topologies coincide on the closed α\alpha-Hölder balls KR=ψ−1((0,R])K_{R}=\psi^{-1}((0,R]) (see Lemma A.5), we conclude that the weighted function space Bψ(Cα(S;Z))\mathcal{B}_{\psi}(C^{\alpha}(S;Z)) is the same for both choices of topologies on Cα(S;Z)C^{\alpha}(S;Z).

For any 0≤α′<α0\leq\alpha^{\prime}<\alpha, let Cτ∗α(S;Z)C^{\alpha}_{\tau^{*}}(S;Z) and Cα′α(S;Z)C^{\alpha}_{\alpha^{\prime}}(S;Z) denote the space Cα(S;Z)C^{\alpha}(S;Z) equipped with the weak-∗*-topology and the Cα′C^{\alpha^{\prime}}-topology induced by ∥⋅∥Cα′(S;Z)\|\cdot\|_{C^{\alpha^{\prime}}(S;Z)}, respectively. Then, Bψ(Cτ∗α(S;Z))=Bψ(Cα′α(S;Z))\mathcal{B}_{\psi}(C^{\alpha}_{\tau^{*}}(S;Z))=\mathcal{B}_{\psi}(C^{\alpha}_{\alpha^{\prime}}(S;Z)).

On the other hand, by combining Lemma 2.7 (i)+(ii) for Bψ(Cα′α(S;Z))\mathcal{B}_{\psi}(C^{\alpha}_{\alpha^{\prime}}(S;Z)), we have

Now, the pre-image KR=ψ−1((0,R])={x∈Cα(S;Z):∥x∥Cα(S;Z)≤η−1(R)}K_{R}=\psi^{-1}((0,R])=\left\{x\in C^{\alpha}(S;Z):\|x\|_{C^{\alpha}(S;Z)}\leq\eta^{-1}(R)\right\} is for every R>0R>0 a closed α\alpha-Hölder ball (thus ∥⋅∥Cα(S;Z)\|\cdot\|_{C^{\alpha}(S;Z)}-bounded and sequentially weak-∗*-closed by Lemma A.5), which implies by Theorem A.4 that the weak-∗*-topology and the Cα′C^{\alpha^{\prime}}-topology induced by ∥⋅∥Cα′α(S;Z)\|\cdot\|_{C^{\alpha}_{\alpha^{\prime}}(S;Z)} coincide on KRK_{R}. Hence, by combining (A.3) and (A.4), it follows that f∈Bψ(Cτ∗α(S;Z))f\in\mathcal{B}_{\psi}(C^{\alpha}_{\tau^{*}}(S;Z)) if and only if f∈Bψ(Cα′α(S;Z))f\in\mathcal{B}_{\psi}(C^{\alpha}_{\alpha^{\prime}}(S;Z)). ∎

Let us summarize the consequences of this section for the closed vector subspace C0α(S;Z)⊆Cα(S;Z)C^{\alpha}_{0}(S;Z)\subseteq C^{\alpha}(S;Z) of α\alpha-Hölder continuous functions preserving the origin.

Theorem A.3 implies the compact embedding C0α(S;Z)↪C0α′(S;Z)C^{\alpha}_{0}(S;Z)\hookrightarrow C^{\alpha^{\prime}}_{0}(S;Z), 0≤α′<α0\leq\alpha^{\prime}<\alpha.

For every R>0R>0 and sequentially weak-∗*-closed subset G⊆ZG\subseteq Z, the set

is ∥⋅∥Cα(S;Z)\|\cdot\|_{C^{\alpha}(S;Z)}-bounded and sequentially weak-∗*-closed (see Lemma A.5).

By Proposition A.6, the weighted function space Bψ(C0α(S;Z))\mathcal{B}_{\psi}(C^{\alpha}_{0}(S;Z)) with admissible weight function of the form ψ(x):=η(∥x∥Cα(S;Z))\psi(x):=\eta\left(\|x\|_{C^{\alpha}(S;Z)}\right) does not depend on the choice of the topology on C0α(S;Z)C^{\alpha}_{0}(S;Z).

In this section, we show for some T>0T>0 and a dual space ZZ equipped with the weak-∗*-topology that the Cp−var,α([0,T];Z)C^{p-var,\alpha}([0,T];Z)-spaces introduced in Section 1.2 are also dual spaces, where (p,α)∈[1,∞)×(0,1)(p,\alpha)\in[1,\infty)\times(0,1) with pα<1p\alpha<1. For this purpose, we first consider the embeddings Cp−var,α([0,T];Z)↪Cp′−var,α′([0,T];Z)C^{p-var,\alpha}([0,T];Z)\hookrightarrow C^{p^{\prime}-var,\alpha^{\prime}}([0,T];Z) for (p′,α′)∈(p,∞]×[0,α)(p^{\prime},\alpha^{\prime})\in(p,\infty]\times[0,\alpha) with p′α′<1p^{\prime}\alpha^{\prime}<1, where C∞−var,α′([0,T];Z):=Cα′([0,T];Z)C^{\infty-var,\alpha^{\prime}}([0,T];Z):=C^{\alpha^{\prime}}([0,T];Z) and Cp′−var,0([0,T];Z):=Cp′−var([0,T];Z)C^{p^{\prime}-var,0}([0,T];Z):=C^{p^{\prime}-var}([0,T];Z).

For every (p,α)∈[1,∞)×(0,1)(p,\alpha)\in[1,\infty)\times(0,1) with pα<1p\alpha<1 and (p′,α′)∈(p,∞]×[0,α)(p^{\prime},\alpha^{\prime})\in(p,\infty]\times[0,\alpha) with p′α′<1p^{\prime}\alpha^{\prime}<1, the embedding Cp−var,α([0,T];Z)↪Cp′−var,α′([0,T];Z)C^{p-var,\alpha}([0,T];Z)\hookrightarrow C^{p^{\prime}-var,\alpha^{\prime}}([0,T];Z) is compact.

Fix some (p,α)∈[1,∞)×(0,1)(p,\alpha)\in[1,\infty)\times(0,1) with pα<1p\alpha<1 and (p′,α′)∈(p,∞]×[0,α)(p^{\prime},\alpha^{\prime})\in(p,\infty]\times[0,\alpha) with p′α′<1p^{\prime}\alpha^{\prime}<1. Then, [47, Proposition 5.3] implies for every x∈Cp−var,α([0,T];Z)x\in C^{p-var,\alpha}([0,T];Z) that

which shows that Cp−var,α([0,T];Z)↪Cp′−var,α′([0,T];Z)C^{p-var,\alpha}([0,T];Z)\hookrightarrow C^{p^{\prime}-var,\alpha^{\prime}}([0,T];Z) is continuous.

where ∥x−xnk∥0:=sup⁡s,t∈S∥(x−xnk)(s)−(x−xnk)(t)∥Z≤2∥x−xnk∥C0(S;Z)\|x-x_{n_{k}}\|_{0}:=\sup_{s,t\in S}\|(x-x_{n_{k}})(s)-(x-x_{n_{k}})(t)\|_{Z}\leq 2\|x-x_{n_{k}}\|_{C^{0}(S;Z)}. This shows that Cp−var,α([0,T];Z)↪Cp′−var,α′([0,T];Z)C^{p-var,\alpha}([0,T];Z)\hookrightarrow C^{p^{\prime}-var,\alpha^{\prime}}([0,T];Z) is compact. ∎

For every (p,α)∈[1,∞)×(0,1)(p,\alpha)\in[1,\infty)\times(0,1) satisfying pα<1p\alpha<1, the Banach space (Cp−var,α([0,T];Z),∥⋅∥Cp−var,α([0,T];Z))(C^{p-var,\alpha}([0,T];Z),\|\cdot\|_{C^{p-var,\alpha}([0,T];Z)}) is a dual space. Moreover, on every sequentially weak-∗*-closed and ∥⋅∥Cp−var,α([0,T];Z)\|\cdot\|_{C^{p-var,\alpha}([0,T];Z)}-bounded subset of Cp−var,α([0,T];Z)C^{p-var,\alpha}([0,T];Z), the weak-∗*-topology and every Cp′−var,α′C^{p^{\prime}-var,\alpha^{\prime}}-topology induced by ∥⋅∥Cp′−var,α′([0,T];Z)\|\cdot\|_{C^{p^{\prime}-var,\alpha^{\prime}}([0,T];Z)}, (p′,α′)∈(p,∞]×[0,α)(p^{\prime},\alpha^{\prime})\in(p,\infty]\times[0,\alpha) with p′α′<1p^{\prime}\alpha^{\prime}<1, are equivalent.

Let (V,∥⋅∥V)(V,\|\cdot\|_{V}) be a predual of (Z,∥⋅∥Z)(Z,\|\cdot\|_{Z}). We want to apply Theorem A.1 with

Then, we follow the proof of Theorem A.4 and use Theorem A.1 to conclude that the ∥⋅∥Cp−var,α([0,T];Z)∗\|\cdot\|_{C^{p-var,\alpha}([0,T];Z)^{*}}-closure of span⁡(E0)\operatorname*{span}(E_{0}) is a predual of Cp−var,α([0,T];Z)C^{p-var,\alpha}([0,T];Z). Moreover, by using the same arguments as in the proof of Theorem A.4, the weak-∗*-topology and every Cp′−var,α′C^{p^{\prime}-var,\alpha^{\prime}}-topology induced by ∥⋅∥Cp′−var,α′([0,T];Z)\|\cdot\|_{C^{p^{\prime}-var,\alpha^{\prime}}([0,T];Z)}, with (p′,α′)∈(p,∞]×[0,α)(p^{\prime},\alpha^{\prime})\in(p,\infty]\times[0,\alpha) and p′α′<1p^{\prime}\alpha^{\prime}<1, coincide on each sequentially weak-∗*-closed and ∥⋅∥Cp−var,α([0,T];Z)\|\cdot\|_{C^{p-var,\alpha}([0,T];Z)}-bounded subset of Cp−var,α([0,T];Z)C^{p-var,\alpha}([0,T];Z). ∎

For every R>0R>0 and sequentially weak-∗*-closed subset G⊆ZG\subseteq Z, the set

is ∥⋅∥Cp−var,α([0,T];Z)\|\cdot\|_{C^{p-var,\alpha}([0,T];Z)}-bounded and sequentially weak-∗*-closed. In particular, by choosing G=ZG=Z, the closed ball BRZ(0)‾\overline{B^{Z}_{R}(0)} is ∥⋅∥Cα(S;Z)\|\cdot\|_{C^{\alpha}(S;Z)}-bounded and sequentially weak-∗*-closed.

The proof follow along the lines of the proof for Lemma A.5. ∎

For the weight function ψ(x):=η(∥x∥Cp−var,α([0,T];Z))\psi(x):=\eta\left(\|x\|_{C^{p-var,\alpha}([0,T];Z)}\right), with continuous increasing η:[0,∞)→(0,∞)\eta:[0,\infty)\rightarrow(0,\infty), we now consider the weighted space (Cp−var,α([0,T];Z),ψ)(C^{p-var,\alpha}([0,T];Z),\psi) either equipped with the weak-∗*-topology as in Example 2.3 (vi) or with the Cp′−var,α′C^{p^{\prime}-var,\alpha^{\prime}}-topology as in Example 2.3 (vii). Then, by using that these topologies coincide on the closed balls KR=ψ−1((0,R])K_{R}=\psi^{-1}((0,R]) (see Lemma A.10), we conclude that the weighted function space Bψ(Cp−var,α([0,T];Z))\mathcal{B}_{\psi}(C^{p-var,\alpha}([0,T];Z)) is the same for both choices of topologies on Cp−var,α([0,T];Z)C^{p-var,\alpha}([0,T];Z).

For any (p′,α′)∈(p,∞]×[0,α)(p^{\prime},\alpha^{\prime})\in(p,\infty]\times[0,\alpha) with p′α′<1p^{\prime}\alpha^{\prime}<1, let Cτ∗p−var,α([0,T];Z)C^{p-var,\alpha}_{\tau^{*}}([0,T];Z) and Cp′−var,α′p−var,α([0,T];Z)C^{p-var,\alpha}_{p^{\prime}-var,\alpha^{\prime}}([0,T];Z) denote the space Cp−var,α([0,T];Z)C^{p-var,\alpha}([0,T];Z) equipped with the weak-∗*-topology and the Cp′−var,α′C^{p^{\prime}-var,\alpha^{\prime}}-topology induced by ∥⋅∥Cp′−var,α′([0,T];Z)\|\cdot\|_{C^{p^{\prime}-var,\alpha^{\prime}}([0,T];Z)}, respectively. Then, Bψ(Cτ∗p−var,α([0,T];Z))=Bψ(Cp′−var,α′p−var,α([0,T];Z))\mathcal{B}_{\psi}(C^{p-var,\alpha}_{\tau^{*}}([0,T];Z))=\mathcal{B}_{\psi}(C^{p-var,\alpha}_{p^{\prime}-var,\alpha^{\prime}}([0,T];Z)).

The proof follows along the lines of the proof for Proposition A.6. ∎

For the closed vector subspace C0p−var,α([0,T];Z)⊆Cp−var,α([0,T];Z)C^{p-var,\alpha}_{0}([0,T];Z)\subseteq C^{p-var,\alpha}([0,T];Z) of α\alpha-Hölder continuous functions with finite pp-variation preserving the origin, we obtain the analogous results as in Remark A.7.

Appendix B Proof of Proposition 4.4

References