Signature Methods in Stochastic Portfolio Theory

Christa Cuchiero, Janka Möller

Introduction

Portfolio theory and portfolio optimization are a core topic in mathematical finance and have been the subject of vivid research for decades. The most prominent example is of course the mean-variance optimization problem as formulated by Harry Markowitz, the founder of Modern Portfolio Theory (see ). His framework takes into account risk-preferences of investors while offering a tractable form of optimization, namely being convex and quadratic. Although the choice of constant portfolio weights is arguably oversimplified, the elegance of the Markowitz model can not be denied. Ever since, researchers have been striving for more and more realistic models with fewer limiting assumptions while hoping to preserve tractability.

In the spirit of coming up with more realistic assumptions, for instance without the need of claiming a specific form of the hardly measurable drift, Robert Fernholz developed Stochastic Portfolio Theory (SPT) (see and also ). A key feature of SPT, again in the realm of more realistic modeling assumptions, is the rather relaxed no-arbitrage condition. Indeed, the price process is only assumed to be a (continuous) semimartingale satisfying No Unbounded Profit with Bounded Risk and not necessarily the stronger condition of No free lunch with vanishing risk (see ). It is also central that the portfolios’ performance is measured with respect to a benchmark portfolio, which is usually chosen to be the market portfolio, i.e. large indices like the S&P500. At least since the introduction of very liquid ETFs replicating market indices the out-performance of the market is indeed a major challenge for investors. In this context Fernholz introduced functionally generated portfolios and derived the so-called master formula describing the relative wealth process of such portfolios and allowing for the detection for relative arbitrages with respect to the market portfolio. These portfolios are constructed via the log-gradient of a function of the market weights. While appreciating the beauty of the framework and also its robustness in view of certain optimal worst case long-run growth rates (see ), one may point out that the log-gradient form as well as the neglect of past information is somewhat restrictive. To reduce these limitations various generalizations of the original framework have been developed, e.g., towards replacing the market weights by market-to-book ratios , considering trading strategies generated by Lyapunov functions , adding extra information in the form of a finite variation process and incorporating past information in a path-dependent setting . In this context it is worth mentioning that the latter two papers do not work in a stochastic setting but rather in a path-wise one, using functional Itô-calculus as in based on Föllmer integration. An alternative path-wise formulation allowing for more general strategies (beyond gradient type) and building on the theory of rough path has been studied in .

In this paper, we consider for simplicity the stochastic setting of continuous semimartingales (even though a path-wise formulation in the spirit of would also be possible), take inspiration from the functionally generated portfolios and generalize them to what we call path-functional portfolios. These are portfolios constructed via an auxiliary portfolio τ\tau and non-anticipative path-functionals with a general semimartingale as input. In doing so, we get rid of any gradient-type form of the functionals and allow for a general benchmark portfolio τ\tau and past information of the market or other exogenous relevant factors in form of a general continuous semimartingale. The choice of continuous semimartingales is here only for the ease of exposition, but could be replaced by any other (predictable stochastic) process.

One of our main interest concerns tractable portfolio optimization in the framework of SPT using such path-functional portfolios. To do so, we focus on linear path-functional portfolios, which are constructed via a linear combination of possibly path-dependent feature maps and constant optimization parameters.

This allows us to incorporate modern machine learning techniques for time-series data, in particular signature methods (see e.g. and the references therein), but also other tools, like random neural networks . In the current paper, our focus lies on signature methods, which play an important role in rough path theory (see e.g. ) and whose appeal stems from the (global) universal approximation theorem. It states that linear functions on the signature can approximate continuous (with respect to certain variation distances) path-functionals arbitrarily well on compact sets of paths or even globally when using the setup of weighted spaces (see ). This result then motivates the notion of signature portfolios, which are linear path-functional portfolios with feature maps being elements of the signature. Indeed, with this choice we can approximate generic path-functional portfolios including functionally generated portfolios and the growth-optimal portfolio in a large class of non-Markovian markets, arbitrary well.

Despite this versatility, signature portfolios and linear path-functional portfolios in general lead to remarkably tractable optimization tasks. More concretely, we show that maximizing the (expected) logarithmic wealth as well as the mean-variance optimization of portfolio returns lead to convex quadratic optimization problems. We would like to highlight that the Markowitz portfolio itself is included in the class of signature portfolios and moreover, that any chosen benchmark portfolio τ\tau can always be attained by signature portfolios constructed via τ\tau.

In addition to our theoretical results, we optimize signature portfolios in numerical experiments using simulated and real market data. By means of simulated data generated from the Black-Scholes model, volatility stabilized models and so-called signature market models, we illustrate that the trained signature portfolios (of small degree) are remarkably close to the theoretical growth-optimal portfolios.

In the real market situation we additionally incorporate transaction costs in our optimization by proposing a regularization under which the optimization problems remain convex and quadratic. To deal with high-dimensional market indices such as the NASDAQ and S&P500, we use dimension reduction techniques leading to what we call randomized signature portfolios and JL-signature portfolios, where the former rely on randomized signatures (see and also for a connection to neural signature kernels), while the latter are based on an explicit Johnson-Lindenstrauss projection. Note that we do not reduce the dimension of the underlying process, but only of its signature. To make the computation of the JL-signature portfolios feasible, we present a memory efficient algorithm to do so. We train the portfolios in both a log-relative wealth and mean-variance optimization in the NASDAQ, the SMI and S&P500 markets. We use randomized- and JL-signature portfolios in the high-dimensional markets (NASDAQ and S&P500) and signature portfolios for the SMI. In the NASDAQ market the portfolios are computed from the ranked market weights, while in the other two cases the name-based market weights are used.

We incorporate proportional transaction costs of 1% and 5% respectively in the name-based markets. In almost all configurations, our trained portfolios outperform the market portfolios during the out-of-sample period, even under 5% of proportional transaction costs. It is worth noting that in the setting with transaction costs we still re-balance our portfolios daily whereas the market portfolio does not pay any transaction costs.

To provide more context to our findings, we highlight two papers which are particularly related to our research, namely and . In , the authors study portfolio optimization in SPT over a family of rank-based functionally generated portfolios parameterized by an exponentially concave function. Their objective is to maximize the relative logarithmic growth rate, which in their parameterization leads to a convex optimization problem. In their empirical study they use (ranked) market data of the 100 largest US stocks, which they manage to out-perform during the out-of-sample period with their trained portfolios. As they do not include transaction costs in the learning procedure (which however could be done by adopting e.g. similar transaction cost treatments as ours), the out-performance does no longer work with 0.45% of proportional transaction costs. This could also be related to the fact that the considered portfolios are long only functionally generated and thus a rather small class. Indeed, the class of portfolios we consider is more general in several respects: we allow for short-selling, the inclusion of information from the past and from exogenous signals as well as for a general benchmark portfolio τ\tau in the construction of the portfolios. At the same time the portfolios of can be approximated arbitrarily well using signature portfolios, which is a result of the universal approximation theorem.

Very recently, the work of appeared, where the authors study mean-variance optimization of the wealth process using signature methods. In contrast to our multiplicative setting, they consider an additive approach where the trading strategies correspond to numbers of shares, which are then directly parameterized as linear functions on the signature. This means in particular that the strategies are only self-financing if a bank-account is included, while in our case the strategies are always self-financing without a bank-account, since we use portfolio weights. Even though the formulation of , where the wealth process itself is the quantity of interest, is the standard one in the literature of the continuous-time mean-variance portfolio selection problem (see e.g. ), it differs from the original approach of Markowitz , where mean-variance optimization of the returns was actually the objective. The latter is advantageous when optimizing (in a model-free setup) along the observed trajectory since the returns are more likely to be stationary than the assets themselves. For these reasons we use returns and a multiplicative approach, which allows in particular to include the classical Markowitz portfolio (as a very special case) in our setup (which is not possible in ). Another difference is that we consider general linear path-functionals and not only signature portfolios, while still obtaining a convex quadratic optimization problem. Additionally we include transaction costs and dimension reduction techniques to be able to deal with 500-dimensional price processes, whereas in the numerical implementations of only two-dimensional price processes are considered. Let us also point out that in the signal is augmented with the two-dimensional lead-lag process to express Itô-integrals as linear functions on the corresponding signature. We usually only augment with time, which – under certain conditions on the quadratic variation – still allows to get linear expressions for the Itô-integrals (see Remark 5.12). As the computation of the signature becomes expensive for higher dimensions, this might be of relevance.

The remainder of the article is structured as follows. In Section 2 we recall the notion of the signature of continuous semimartingales, provide two versions of the universal approximation theorem, state the Johnson-Lindenstrauss Lemma and specify the financial market setting as well as classes of important portfolios. Section 3 is dedicated to the introduction of (linear) path-functional portfolios with signature portfolios as special case, while Section 4 addresses their approximation properties, in particular in view of approximating the growth optimal portfolio. In Section 5 the convex quadratic optimization tasks for the mean-variance and (expected) log-(relative) wealth are discussed and Section 6 explains the treatment of proportional transaction costs. Section 7 concludes the paper with a presentation of the numerical results. Some proofs are gathered in the Appendix.

Preliminaries

We define the truncated tensor algebra at level NN by

Let us now state a few important and useful properties about the signature.

For two multiindices II, JJ we define the shuffle product \shuffle\shuffle as

Then, for any two multiindices II, JJ and any continuous semimartingale XX, it holds that

This follows directly from the properties of the Stratonovich integral. ∎

Let XX be a continuous semimartingale with at least one strictly monotone component. Then the signature uniquely determines XX up to vertical translations of the trajectories.

This follows from the fact, that the signature determines the underlying semimartingale uniquely up to tree-like equivalences, see . A more direct proof, in the case where that component is the time, can be found e.g. in . ∎

We now define a class of functions which plays a major role in this paper, namely linear functions on the signature:

2 Universal Approximation Theorem for Linear Functions on the Signature

Linear functions on the signature are dense in a set of certain continuous path-functionals defined on compact sets of paths. We call this the universal approximation property and state it rigorously in Theorem 2.11. To this end we need the following preparations which have been formulated in a similar manner in .

We define the space of lifted stopped paths via

We then equip ΛTN(φ,x){\Lambda}_{T}^{N}(\varphi,x) with the metric dΛd_{\Lambda} with its form being based on

Based on we formulate the notion of non-anticipative path-functionals in our setting:

The following continuity result for the signature is crucial for the universal approximation property.

For all t∈[0,T]t\in[0,T] and N≥2N\geq 2 there exists a unique bijection

with inverse Π≤2\Pi_{\leq 2} and SNS^{N} is continuous on bounded sets. We call SNS^{N} Lyons’ lift. Moreover, for any continuous semimartingale XX we have for almost all ω∈Ω\omega\in\Omega that

The first part of the theorem follows from [31, Theorem 8.10, Theorem 9.5 and Corollary 9.11] and the second statement follows from [31, Proposition 17.1 and Exercise 17.2]. ∎

Relying on two lemmas proved in Appendix A, we can now show the universal approximation theorem for linear functions on the signature. We here state a version that holds uniformly in time. For this we also provide a rigorous proof in Appendix A which is – up to our knowledge – not available in the literature. For similar assertions we however refer to . In the setting of càdlàg rough paths an analogous result has been proved in , however with different topologies.

To show the applicability of this result let us give some examples of non-anticipative path-functionals.

Then the components of the solution to this SDE (Yti)t∈[0,T](Y^{i}_{t})_{t\in[0,T]}, can be seen as a map

with x^=Π1(x^)\hat{x}=\Pi_{1}(\hat{\mathbf{x}}), y^=Π1(y^)\hat{y}=\Pi_{1}(\hat{\mathbf{y}}), where we replaced the max⁡k∈{1,…,N}\max_{k\in\{1,\dots,N\}} by k=1k=1 and sup⁡D∈D\sup_{D\in\mathcal{D}} by D=[0,t∨s]D=[0,t\vee s] in the definition of d2−var;[0,t∨s]d_{2-var;[0,t\vee s]} to obtain the inequality. The last equality follows from the fact that all paths in ΛT2(φ,x)\Lambda_{T}^{2}(\varphi,x) have the same initial value. Due to the continuity of hh, for every ϵ>0\epsilon>0 there exists a δ>0\delta>0 such that for all t1,t2∈[0,T]t_{1},t_{2}\in[0,T] and all paths X(ω1),Y(ω2)X(\omega_{1}),Y(\omega_{2}) satisfying ∣X^t1(ω1)−Y^t2(ω2)∣<δ\left|\hat{X}_{t_{1}}(\omega_{1})-\hat{Y}_{t_{2}}(\omega_{2})\right|<\delta we have

3 Global Universal Approximation Theorem for Linear Functions on the Signature

In this subsection, we state an alternative universal approximation theorem which is due to . We will refer to Theorem 2.18 also as global universal approximation theorem because we are no longer limited to compact sets, however this comes at the price of approximating with respect to weighted norms instead of the uniform norm. The following is based on the statements and the results of adapted to the setting considered in this paper. However, we do formulate everything in the path-functional setting here, leading to a approximation uniformly in time which has not been done in .

Let us start by introducing the following metrics:

In the following we shall denote by GtN,α(φ,x):=(GtN(φ,x),dα,[0,t])\mathcal{G}_{t}^{N,\alpha}(\varphi,x):=(\mathcal{G}_{t}^{N}(\varphi,x),d_{\alpha,[0,t]}) the space of lifted paths equipped with the homogeneous α\alpha-Hölder metric for some α<12\alpha<\frac{1}{2}. Moreover, we write ΛTN,α(φ,x):=(ΛTN(φ,x),dΛα)\Lambda^{N,\alpha}_{T}(\varphi,x):=(\Lambda^{N}_{T}(\varphi,x),d_{\Lambda^{\alpha}}) for the corresponding space of lifted stopped paths equipped with

We shall assume without loss of generality that the set Ωˉ\bar{\Omega} in the definition of GtN(φ,x)\mathcal{G}_{t}^{N}(\varphi,x) is such that the homogeneous α\alpha-Hölder norm is finite for all lifted paths and all α<12\alpha<\frac{1}{2}.

Due to the fact that we are now considering different topologies on the space of lifted (stopped) paths, the continuity statements of Lemma A.2, Corollary 2.10 and Lemma A.3 need to be proved again. In fact, they do hold analogously:

The proof that the map λ\lambda defined in Lemma A.2 is again continuous in the α\alpha-Hölder topology works very similar to the one of Lemma A.2, where one ultimately chooses δ>0\delta>0 such that

The continuity of Lyons’ lift holds by [31, Theorem 9.5, Corollary 9.11].

The proof of Lemma A.3 holds analogously in the α\alpha-Hölder topology.

In order to formulate the global universal approximation theorem we need the notion of a weighted space and certain weighted function spaces being generalizations of continuous functions.

Let (X,τX)(X,\tau_{X}) be a completely regular Hausdorff space and consider a function ψ:X→(0,∞)\psi:X\rightarrow(0,\infty). We call the pair (X,ψ)(X,\psi) a weighted space if it holds that KR:=ψ−1((0,R])={x∈X∣ψ(x)≤R}K_{R}:=\psi^{-1}((0,R])=\{x\in X\lvert\psi(x)\leq R\} is τX\tau_{X}-compact for all R>0R>0. Moreover, for a weighted space (X,ψ)(X,\psi), we define

and equip it with ∥f∥Bψ(X):=sup⁡x∈X∣f(x)∣ψ(x)\|f\|_{\mathcal{B}_{\psi}(X)}:=\sup_{x\in X}\frac{|f(x)|}{\psi(x)}. We denote by Bψ(X)\mathcal{B}_{\psi}(X) the closure of the space of bounded continuous functions Cb(X)C_{b}(X) under ∥f∥Bψ(X)\|f\|_{\mathcal{B}_{\psi}(X)}. We refer to Bψ(X)\mathcal{B}_{\psi}(X) as weighted function space.

The global universal approximation theorem via linear functions on the signature hinges on an application of the weighted Stone-Weierstrass theorem [22, Theorem 3.6] applied to Bψ(ΛT2(φ,x))\mathcal{B}_{\psi}(\Lambda_{T}^{2}(\varphi,x)) for the specific weight functionNote that in contrast to [22, Theorem 5.4] we can choose ξ>1\xi>1 (and do not need to require ξ>2\xi>2), as we are working in a semimartingale setting. We refer to the proof of Theorem 2.18 for a more detailed explanation.

Let α<12\alpha<\frac{1}{2} and consider the weight function ψ\psi given in (2.2).

Then ΛT2,α′(φ,x)\Lambda_{T}^{2,\alpha^{\prime}}(\varphi,x), i.e. ΛT2(φ,x)\Lambda_{T}^{2}(\varphi,x) equipped with dΛα′d_{\Lambda^{\alpha^{\prime}}} for 0≤α′<α0\leq\alpha^{\prime}<\alpha, is a weighted space.

We are finally ready to formulate the global universal approximation theorem for non-anticipative path-functionals. Its proof is also be given in Appendix A.

Let 0≤α′<α<120\leq\alpha^{\prime}<\alpha<\frac{1}{2} and consider the weight function ψ\psi given in (2.2). Then, for every non-anticipative path-functional f∈Bψ(ΛT2,α′(φ,x))f\in\mathcal{B}_{\psi}\left(\Lambda^{2,\alpha^{\prime}}_{T}(\varphi,x)\right) (i.e. ΛT2(φ,x)\Lambda^{2}_{T}(\varphi,x) is equipped with dΛα′d_{\Lambda^{\alpha^{\prime}}}) and for every ϵ>0\epsilon>0 there exists a linear function on the signature LL such that for almost all ω∈Ω\omega\in\Omega

4 The Johnson-Lindenstrauss Lemma

The Johnson-Lindenstrauss Lemma is an important mathematical result, which is interesting for machine learning and data science, as it can be seen as a dimension reduction technique. There are many variants of the Johnson-Lindenstrauss lemma, we will here use a version stated in [55, Theorem 3.1]:

In this paper, we will use Theorem 2.19 to obtain a random projection of the truncated signature. We will choose the components AijA_{ij} to be i.i.d. with Aij∼N(0,1k)A_{ij}\sim\mathcal{N}(0,\frac{1}{k}).

5 The Market and Important Portfolio Specifications

We call a vector πt=(πt1,...,πtd)\pi_{t}=(\pi^{1}_{t},...,\pi^{d}_{t}) of predictable processes fulfilling ∑i=1dπti≡1\sum_{i=1}^{d}\pi_{t}^{i}\equiv 1 and being integrable with respect to RR, with Rti:=∫0tdSuiSuiR_{t}^{i}:=\int_{0}^{t}\frac{dS_{u}^{i}}{S_{u}^{i}}, an (RR-integrable) portfolio.

The component πti\pi^{i}_{t} denotes the proportion of wealth invested in stock ii at time tt, hence every portfolio is self-financing. Consider investing with a portfolio π\pi over a time-horizon [t0,t1][t_{0},t_{1}] for 0≤t0<t1≤T0\leq t_{0}<t_{1}\leq T. We denote its wealth process portfolio π\pi with initial wealth ww by (Wtπ,w)t∈[t0,t1](W^{\pi,w}_{t})_{t\in[t_{0},t_{1}]}. Without loss of generality we will assume from now on that the initial investment is always one unit of currency, i.e. that w=1w=1 and therefore denote Wtπ:=Wtπ,1W^{\pi}_{t}:=W^{\pi,1}_{t}. The wealth process of a portfolio fulfils the stochastic differential equation

A special portfolio is the market portfolio whose portfolio weights are given by the relative capitalizations μt\mu_{t}, whence we also refer to μt\mu_{t} as the market weights. The wealth of the market portfolio is given by

We define the relative wealth process of a portfolio π\pi to be

By using Itô’s-formula, it follows directly from (2.3) that

Furthermore, by applying Itô’s formula again, we obtain

where [X,Y]t[X,Y]_{t} denotes the co-variation process of continuous semimartingales X,YX,Y.

We now turn to some portfolios which are of special interest, those are the numeraire, log-optimal and growth-optimal portfolio.

We call a portfolio ρ\rho the numeraire portfolio if WπWρ\frac{W^{\pi}}{W^{\rho}} is a supermartingale for any other portfolio π\pi.

In a semimartingale market the NUPBR condition (see e.g. ) holds if and only if the numeraire portfolio ρ\rho exists and WTρ<∞W^{\rho}_{T}<\infty almost surely.

See [43, Theorem 4.12], where in our case the predictable closed convex constraints C\mathfrak{C} are

Note that [43, Theorem 4.12] considers the setting with bank-account, however our constraint set C\mathfrak{C} prohibits us from investing into the bank account, i.e. π0≡0\pi^{0}\equiv 0 (in the notation of ). ∎

To connect the numeraire portfolio with the log-optimal one introduced in Definition 2.27 below, we recall Proposition 4.19 from .

has no solution or has infinitely many solutions, if the NUPBR condition does not hold.

Consider the log-utility optimization problem

If it is well-posed, we call the corresponding optimal portfolio log-optimal portfolio.

The following lemma can also be found in .

Under the NUPBR condition, the log-utility optimization problem admits a solution, if it is finite, in which case the numeraire portfolio is the log-optimal portfolio.

Finally, let us introduce the notion of the growth-optimal portfolio.

where ata_{t} is a dd-dimensional vector, Σt\Sigma_{t} a d×md\times m-matrix for m≥dm\geq d and BB a mm-dimensional Brownian motion. Assume ata_{t}, BtB_{t} are predictable processes and satisfy

and ΣtΣtT\Sigma_{t}\Sigma_{t}^{\mathsf{T}} is almost surely invertible for all t∈[0,T]t\in[0,T]. The log-wealth-process of a portfolio π\pi in this market model is then given by

We call gtπ=πtTat−12πtTΣtΣtTπtg^{\pi}_{t}=\pi_{t}^{\mathsf{T}}a_{t}-\frac{1}{2}\pi_{t}^{\mathsf{T}}\Sigma_{t}\Sigma_{t}^{\mathsf{T}}\pi_{t} the portfolios growth-rate and the portfolio with maximal growth-rate, if it exists, the growth-optimal portfolio.

Consider a market model given in Definition 2.29 and let us define the (instantaneous) market price of risk

where 1=(1,...,1)T\mathbf{1}=(1,...,1)^{\mathsf{T}}.

Regarding the existence of the growth-optimal and log-optimal portfolios, see , also for the general case of continuous semimartingale markets. The specific form of the growth-optimal portfolio is due to . ∎

Consider the setting of a market as in Definition 2.29. Then the growth-optimal portfolio exists if and only if the numeraire portfolio exists in which case they are the sameAssuming that their generated wealth is almost surely finite this holds if and only if NUPBR is satisfied (see Theorem 2.25).. Moreover, if the log-utility maximization problem is finite, then they also coincide with the log-optimal portfolio.

See and . The last assertion then follows from Lemma 2.28 ∎

In SPT, there is another class of portfolios of special interest, those are functionally generated portfolios, see .

The function GG is called the portfolio generating function and if it is concave, π\pi is a long-only portfolio.

Before we conclude this section, we introduce the ranked market-weights which are of particular interest in SPT, due to the remarkable stability of the capital distribution curves, see for example .

where we break ties by allocating a lower rank to smaller labels. To be precise, if Xti=XtjX_{t}^{i}=X_{t}^{j} with i≥ji\geq j, we set Xt(ri)=XtiX_{t}^{(r_{i})}=X_{t}^{i} and Xt(rj)=XtjX_{t}^{(r_{j})}=X_{t}^{j} where ri≥rjr_{i}\geq r_{j}.

Accordingly, we denote by μ\boldsymbol{\mu} the ranked market weights and by S\mathbf{S} the ranked capitalization process.

(Linear) Path-Functional Portfolios

Before we present our application of signature methods in the context of SPT and portfolio optimization, we generalize the concept of functionally generated portfolios by introducing so-called path-functional portfolios.

We introduce path-functional portfolios which are generated by a family of non-anticipating path functionals {fi}1≤i≤d\{f^{i}\}_{1\leq i\leq d} depending on time and the path of XX and which are constructed via an auxiliary portfolio τ\tau, where τ\tau is required to be uniformly bounded. Moreover, we assume that the non-anticipating path functionals {fi}1≤i≤d\{f^{i}\}_{1\leq i\leq d} fulfil the necessary integrability conditions such that the following processes π\pi are portfolios in the sense of Definition 2.22. We consider two types of such portfolios:

πti(τ,X)=τti(fi(t,X[0,t])+1−∑j=1dτtifj(t,X[0,t]))\pi_{t}^{i}(\tau,{X})=\tau_{t}^{i}\left(f^{i}\left(t,X_{[0,t]}\right)+1-\sum_{j=1}^{d}\tau_{t}^{i}f^{j}\left(t,X_{[0,t]}\right)\right)

πti(τ,X)=fi(t,X[0,t])+τti(1−∑j=1dfj(t,X[0,t]))\pi_{t}^{i}(\tau,{X})=f^{i}\left(t,X_{[0,t]}\right)+\tau_{t}^{i}\left(1-\sum_{j=1}^{d}f^{j}\left(t,X_{[0,t]}\right)\right)

In the above definition, we consider a general continuous semimartingale and a general auxiliary portfolio τ\tau in the construction of the path-functional portfolios. In the spirit of SPT, one would choose τ=μ\tau=\mu the market portfolio and X=μX=\mu, the process of market weights. One can then recover the classical functionally generated portfolios from the path-functional portfolios of type II by setting X=μX=\mu and choosing fi(t,μ[0,t])=fi(μt)=Dilog⁡(G(μt))f^{i}(t,\mu_{[0,t]})=f^{i}(\mu_{t})=D_{i}\log(G(\mu_{t})).

A simple example of a path-functional portfolio of type IIII would be any portfolio generated by a deep (recurrent) neural network, see for example . Denote by NNi(t,X[0,t])\mathcal{NN}^{i}(t,X_{[0,t]}) the ii-th output of a deep (recurrent) neural network. If

the last layer of the neural network is for example a softmax function and the portfolio weights are defined by πt(NN),i=NNi(t,X[0,t])\pi^{(\mathcal{NN}),i}_{t}=\mathcal{NN}^{i}(t,X_{[0,t]}), or,

the portfolio weights are constructed as πt(NN),i=NNi(t,X[0,t])∑j=1dNNj(t,X[0,t])\pi^{(\mathcal{NN}),i}_{t}=\frac{\mathcal{NN}^{i}(t,X_{[0,t]})}{\sum_{j=1}^{d}\mathcal{NN}^{j}(t,X_{[0,t]})},

then, the portfolios π(NN)\pi^{(\mathcal{NN})} are in both cases path-functional portfolios of type IIII but not of type II. Note that the choice of τ\tau is irrelevant because ∑j=1πt(NN),j=1\sum_{j=1}\pi^{(\mathcal{NN}),j}_{t}=1.

Note that for a fixed auxiliary portfolio τ\tau with weight-processes which are continuous semimartingales and for a continuous semimartingale (Xt)t∈[0,T](X_{t})_{t\in[0,T]} any path-functional portfolio of type II with portfolio controlling functions f(I),i:X[0,t]↦f(I),i(X[0,t])f^{(I),i}:X_{[0,t]}\mapsto f^{(I),i}(X_{[0,t]}) is a path-functional portfolio of type IIII with portfolio controlling function f(II),i:(τ,X)[0,t]↦τtif(I),i(X[0,t])f^{(II),i}:(\tau,X)_{[0,t]}\mapsto\tau_{t}^{i}f^{(I),i}(X_{[0,t]}). This also holds vice versa, if the components of τ\tau are all non-zero. Indeed, any path-functional portfolio of type IIII with portfolio controlling function f(II),i:X[0,t]↦f(II),i(X[0,t])f^{(II),i}:X_{[0,t]}\mapsto f^{(II),i}(X_{[0,t]}) is a path-functional portfolio of type II with portfolio controlling function f(I),i:(τ,X)[0,t]↦f(II),i(X[0,t])τtif^{(I),i}:(\tau,X)_{[0,t]}\mapsto\frac{f^{(II),i}(X_{[0,t]})}{\tau_{t}^{i}}. However, note that this conversion always involves adding the process of the auxiliary portfolio τ\tau to the input. Not only does this enlarge the dimensions of the inputs, but also is τ\tau then required to be a continuous semimartingale. Hence, this conversion is not always possible.

We now introduce a special class of path-functional portfolios which are particularly useful for optimizing path functional portfolios, as we will make more explicit in Section 5. This class is the one of linear path-functional portfolios.

For a path-functional portfolio of type II or IIII, we call it a linear path-functional portfolio of the corresponding type, if the portfolio controlling functions are of the form

Considering linear path-functional portfolios gives us a first idea on how to do portfolio optimization in that context. Namely, optimizing the parameters {lνi}1≤i≤d, ν∈V\{l_{\nu}^{i}\}_{{1\leq i\leq d,\,\nu\in\mathcal{V}}} for a given collection of feature maps. Note, that the auxiliary portfolio τ\tau can always be attained in this way by

setting lνi=lνjl_{\nu}^{i}=l_{\nu}^{j} for all i,j∈{1,…,d}i,j\in\{1,\dots,d\}, ν∈V\nu\in\mathcal{V} for portfolios of type II

setting lνi=0l_{\nu}^{i}=0 for all i∈{1,…,d}i\in\{1,\dots,d\}, ν∈V\nu\in\mathcal{V} for portfolios of type IIII.

Hence, if one wants to learn a linear path-functional portfolio which outperforms a given benchmark portfolio πˉ\bar{\pi}, one should use the benchmark portfolio as the auxiliary portfolio, i.e. set τ=πˉ\tau=\bar{\pi}, since then the benchmark portfolio is included in the family of portfolios one optimizes over. In the context of stochastic portfolio theory, we therefore often choose τ=μ\tau=\mu, because we aim to outperform the market portfolio.

Among all linear path-functional portfolios, we are in particularly interested in those, whose feature maps are (randomized) elements of the signature, resulting in (JL- or randomized-)signature portfolios. We make this more precise in the following definition:

Signature Portfolios: A signature portfolio of degree NN is a linear path-functional portfolio with V={I ∣ I=(i1,...,im)∈{1,…,n}m for 0≤m≤N}\mathcal{V}=\{I\,\mid\,I=(i_{1},...,i_{m})\in\{1,\dots,n\}^{m}\textrm{ for }0\leq m\leq N\} (the set of multiindices up to length NN) and with feature maps

In other words, signature portfolios are path-functional portfolios, where the portfolio controlling functions are linear functions on the signature.

JL-Signature Portfolios: A JL-signature portfolio of dimension (P,N)(P,N) is a linear path-functional portfolio with V={1,…,P}\mathcal{V}=\{1,\dots,P\}, where PP is the dimension of the projection, and with feature maps

where ApA^{p} is the pp-th column of the Johnson-Lindenstrauss projection introduced in Theorem 2.19. Hence, JL-signature portfolios are path-functional portfolios, where the portfolio controlling functions are linear functions on the Johnson-Lindenstrauss projected signature.

Randomized-Signature Portfolios: Let S\mathcal{S} denote the solution to

Hence, randomized-signature portfolios are path-functional portfolios, where the portfolio controlling functions are linear combinations of the elements of the randomized signature.

We have highlighted signature, JL- and randomized-signature portfolios as linear path-functional portfolios of special interest in this paper. However, we would like to mention some other examples of linear path-functional portfolios whose portfolio controlling functions are

increments of the signature. That is, for a fixed time span Θ\Theta, we observe a rolling window of length Θ<T\Theta<T and forget the information before that, i.e.

We will refer to such portfolios as signature portfolios with rolling windows. We shall not put particular emphasis on these portfolios in this paper, since our approximation results in Section 4 are tailored to signature portfolios without rolling windows and since portfolios with rolling windows are numerically less tractable, as we outline in Section 7.

random neural networks, which are neural networks where the parameters of the hidden layers are not trained but randomly sampled and only the (linear) read-out layer is trained, see . Training the read-out layer, exactly amounts to training the parameters {lνi}1≤i≤d,ν∈V\{l_{\nu}^{i}\}_{1\leq i\leq d,\nu\in\mathcal{V}} of path-functional portfolios. Such neural networks may have an infinite-dimensional input space, corresponding in our context to path spaces, as e.g. considered in or .

given by reservoir computers, see for instance . A special class of which are Echo State Networks or Quantum Reservoir Computers . Note, that there is an interesting connection between the Johnson-Lindenstrauss projection of signatures, randomized signature and reservoir computing, which is worked out in .

given by a constant, which leads to the case of constant portfolio weights when using a constant auxiliary portfolio τ\tau. In this case we recover the Markowitz portfolio optimization as a very special case in Subsection 5.1. Moreover, since the first element of the signature is constant, this class of portfolios is also included in the class of signature portfolios.

Approximation Properties of Signature Portfolios

This section is dedicated to prove universal approximation properties of signature portfolios. Let us recall that for a compact subset K⊂GT2(φ,x)K\subset\mathcal{G}_{T}^{2}(\varphi,x) the set

We present the proof for the case of portfolios of type II, however the proof is analogous for the case of type IIII.

Let us start by showing the first statement of the theorem. By the universal approximation result (Theorem 2.11), we know that for each {fi}1≤i≤d\{f^{i}\}_{1\leq i\leq d} there exits a linear function on the signature of X^\hat{X} that approximates fif^{i} arbitrarily well, on compact sets. Those linear functions can be chosen as the portfolio controlling functions fi,\*f^{i,\*} of a signature portfolio π∗\pi^{*}. Hence, there exists for each ϵ′>0\epsilon^{\prime}>0 a signature portfolio π∗\pi^{*} of type II such that for all i∈{1,...,d}i\in\{1,...,d\} it holds that

holds almost surely, where MM is the bound of τ\tau. The result follows by choosing ϵ′=ϵM(d+1)\epsilon^{\prime}=\frac{\epsilon}{M(d+1)}.

The proof of the second approximation statement follows by similar arguments. Recall Theorem 2.18 and use the same reasoning as above to translate the approximation result from the portfolio controlling functions to the portfolio weights itself, i.e. we obtain that there exists for each ϵ′>0\epsilon^{\prime}>0 a signature portfolio π∗\pi^{*} of type II such that for all i∈{1,...,d}i\in\{1,...,d\} it holds that

Note, that by Corollary 2.12, Theorem 4.1 also ensures that signature portfolios approximate classical functionally generated portfolios arbitrarily well, on compact sets. In particular, this includes the functionally generated portfolios as introduced by Fernholz and those considered for functional portfolio optimization in .

Having obtained a universal approximation theorem of path-functional portfolios, we will now apply it to obtain universal approximation results of the growth-optimal portfolio by signature portfolios in several markets.

In this subsection, we present a large class of markets where the growth-optimal portfolio can be regarded as a path-functional portfolio and show that it can therefore be approximated by a signature portfolio.

The following statements are to be understood in the market setting outlined in Definition 2.29 and we therefore assume that the corresponding necessary conditions are satisfied, in particular those ensuring the existence of the growth-optimal portfolio (Lemma 2.30). Moreover, recall also the definition of the compact set KΛ\mathcal{K}_{\Lambda} given in (4.1).

Consider the following class of markets with dd stocks:

where π(g)\pi^{(g)} is the growth-optimal portfolio of the respective market.

Note that the semimartingale XX in (4.4) could be (a functions of) SS, so that we are actually dealing with path-dependent SDEs. In the particular case where aa and Σ\Sigma are entire functions of the signature as introduced in we are then in the tractable setup of signature SDEs.

Recalling Lemma 2.30, it is straightforward that for any auxiliary portfolio τ\tau, the growth-optimal portfolio π(g)\pi^{(g)} is a path-functional portfolio of type IIII with portfolio controlling functions

Likewise, π(g)\pi^{(g)} is a path-functional portfolio of type II for any deterministic auxiliary portfolio and portfolio controlling functions

By Lemma B.1 and the form of π(g),i\pi^{(g),i} Theorem 4.1 is applicable for π(g),i\pi^{(g),i} (viewed as a path-functional portfolio of type II resp. type IIII). Hence, the statements follow. ∎

If X=μX=\mu or X=SX=S, the growth optimal is also a path-functional portfolio of type II with auxiliary portfolio τ=μ\tau=\mu.

We here give two prominent examples of markets, to which Theorem 4.3 can be applied.

and α≥0\alpha\geq 0. Volatility stabilized markets are of great interest in SPT because they reflect the observation in real markets that smaller stocks tend to have greater volatility than larger stocks, see for example . See also for further properties of volatility stabilized markets.

2 A Class of Non-Markovian Markets where the Growth-Optimal Portfolio is a Signature Portfolio

Let us now turn to a class of (possibly) non-Markovian markets, where the growth-optimal portfolio can not only be approximated by a signature portfolio but is a signature portfolio.

Let XX be a continuous semimartingale. For a market of dd stocks, consider the class of Sig-market models

The integrability and measurability assumptions required in Definition 2.29 and Lemma 2.30 follow from the integrability and measurability of elements of the signature. Applying Lemma 2.30, it follows that the growth-optimal weights are linear functions on the signature of X^\hat{X}. This is for any auxiliary portfolio a signature portfolio of type IIII, or a signature portfolio of type II for any constant and deterministic auxiliary portfolio τ\tau. ∎

As addressed in Remark 4.4, the existence of solutions is non-trivial for Sig-markets if X=h(S)X=h(S), i.e. if the semimartingale of which we construct the signature is a function of the price process itself. In particular, if hh is a real analytic function, then we deal with signature-SDEs in the spirit of . The study of existence of solutions to such equations goes beyond the subject of this paper, however, we give a simple example below, where the existence is guaranteed. To this end, let us denote by

Let X=log⁡StX=\log{S_{t}}. If αI(i)=0\alpha_{I}^{(i)}=0 for all I∉I1I\notin\mathcal{I}_{1}, then

The NUPBR condition in the above Sig-markets holds. This follows from the existence of the growth-optimal portfolio by Lemma 2.30, Lemma 2.31 and Theorem 2.25.

Optimization Tasks for Linear Path-Functional Portfolios

We now want to study some tasks for portfolio optimization and their form for linear path-functional portfolios. To formulate them in the most general way, let us introduce what we call a universe of stocks.

Consider a market of dd stocks with capitalization process S=(S1,...,Sd)S=(S^{1},...,S^{d}). We call U⊆{1,...,d}\mathcal{U}\subseteq\{1,...,d\} a universe of stocks with capitalization process SU=(Su)u∈US^{\mathcal{U}}=(S^{u})_{u\in\mathcal{U}}. Moreover, we define the universe weights (resp. universe portfolio) μU\mu^{\mathcal{U}} to be

Moreover, denote by (WtU)t∈[0,T](W_{t}^{\mathcal{U}})_{t\in[0,T]} the wealth process of universe U\mathcal{U}, which is given by

This notion of a universe of stocks is useful if one does not necessarily want to invest in all the stocks in the market but just in a subset of stocks. Of course, this can always be achieved by fixing certain weights of a portfolio to be zero, however using our notion of a universe of stocks is particularly useful, if one wants to compare the wealth process of a portfolio to the wealth process of the universe. Let us make this precise:

Take a universe U⊆{1,...,d}\mathcal{U}\subseteq\{1,...,d\} and consider a portfolio π\pi where πi≡0\pi^{i}\equiv 0 for i∉Ui\not\in\mathcal{U}. Then, the relative wealth process of π\pi with respect to μU\mu^{\mathcal{U}} is given by

Using the definition of π\pi and WUW^{\mathcal{U}}, the statement follows directly from (2.3) and applying Itô’s formula. ∎

When we consider path-functional portfolios investing in a universe U\mathcal{U}, it is convenient to construct them via the universe portfolio μU\mu^{\mathcal{U}}. Note that for path-functional portfolios π(I)\pi^{(I)} of type II it directly follows that π(I),i(μU,⋅)=0\pi^{(I),i}(\mu^{\mathcal{U}},\cdot)=0 for i∉Ui\not\in\mathcal{U}. However, for path-functional portfolios π(II)\pi^{(II)} of type IIII, we have to require additionally that π(II),i(μU,⋅)=0\pi^{(II),i}(\mu^{\mathcal{U}},\cdot)=0 for i∉Ui\not\in\mathcal{U} by setting fi≡0f^{i}\equiv 0 for i∉Ui\not\in\mathcal{U}.

Recall that our very basic assumptions on the financial market were that the stocks’ capitalizations are positive continuous semimartingales. Considering a market MM which fulfils these assumptions, we know that there exists another market which also fulfils these assumptions and for which the stocks’ capitalizations are the same as the ranked capitalizations of the market MM. This follows simply from the fact that the ranked capitalizations are again positive continuous semimartingales. Therefore, the optimization problems we present in the following hold also for investments in the ranked market.

It is sometimes useful to write optimization tasks in a vectorized form. To this end, let us introduce what we call a labelling function.

For a given linear path-functional portfolio π\pi investing in a universe of stocks U\mathcal{U}, let V\mathcal{V} be the set of features, ∣V∣|\mathcal{V}| the number of features and ∣U∣|\mathcal{U}| the number of stocks in a universe. A bijective function L\mathscr{L} of the form

Before studying two concrete optimization tasks for linear path-functional portfolios, let us make the following statement about a more general class of optimization tasks, which turn out to be quadratic optimization problems for linear path-functional portfolios:

Recall the form of linear path-functional portfolios of type II and IIII respectively. Therefore, it is clear that the parameters {lνi}\{l_{\nu}^{i}\} appear at most in quadratic terms. Moreover, since they are constant in time and deterministic, they can be pulled outside of the integrals.

Let us now, present two prominent optimization tasks which fall into the above class, namely mean-variance and log-wealth optimization. We assume that the following optimization problems are well-posed and that the expectations are finite.

In this optimization problem, we consider a strategy, where we invest at time tt with portfolio weights πt\pi_{t} of a linear path-functional portfolio into a universe U\mathcal{U} and follow a buy-and-hold strategy over the time-span Δ\Delta. The (relative) return of this strategy over the time-horizon [t,t+Δ][t,t+\Delta] is given by

and respectively the relative return (relative to the universe portfolio μU\mu^{\mathcal{U}})

In this setting, we can state the following corollary about the mean-variance optimization problem:

mean-variance optimization of returns we have

mean-variance optimization of relative return we have

We proof the above statement for the case of RW,π\mathcal{R}^{W,\pi} and for π\pi of type II. The proof of the other cases is completely analogous.

Note that the rightmost term of (5.5) do not contribute in the optimization problem, as they are independent of l\mathbf{l}. And,

Note that the expression of the optimization problem for the relative return simplifies considerably if we use a linear path-functional portfolio of type II and choose the universe portfolio as the auxiliary portfolio, i.e. τ=μU\tau=\mu^{\mathcal{U}}. To be precise, we obtain

for lL(i,ν)=lνi\mathbf{l}_{\mathscr{L}(i,\nu)}=l_{\nu}^{i} and YL(i,ν)V=ϕν(t,X[0,t])(μt+ΔU,i−μtU,i)\mathbf{Y}^{V}_{\mathscr{L}(i,\nu)}=\phi^{\nu}(t,X_{[0,t]})(\mu^{\mathcal{U},i}_{t+\Delta}-\mu^{\mathcal{U},i}_{t}), since Rt,t+ΔV,μU=0\mathcal{R}^{V,\mu^{\mathcal{U}}}_{t,t+\Delta}=0.

If Var(YV)\textrm{Var}(\mathbf{Y}^{V}) or Var(YW)\textrm{Var}(\mathbf{Y}^{W}) respectively are almost surely invertible, that is the optimization problem is strictly convex and hence admits a unique set of optimal parameters, then the optimal vector l∗\mathbf{l}^{*} can be found via first-order conditions, i.e.

and similarly for i<ni<n, taking just the history starting from into account. Assume now further that we have an invariant measure for the distribution of the rolling path segments of length nn, which we denote by ρ\rho. Then under this invariant measure we have for each i≥ni\geq n

and accordingly for i<ni<n assuming compatible invariant measures for the paths of shorter lengths. Under these assumptions we then obtain the following ergodicity result

This can be deduced from Birkhoff’s ergodic theorem for discrete time Markov processes (see e.g. [Theorem 2.2, Section 2.1.4] and also [21, Section 3] for the standard Markov situation with n=1n=1). Indeed, this becomes applicable by identifying the states of the Markov process as the past path segments of length nn, i.e. a state ziΔz_{i\Delta} is given by ziΔ=(μ(i−n)Δ,…,μiΔ)z_{i\Delta}=(\mu_{(i-n)\Delta},\ldots,\mu_{i\Delta}). The transition probabilities from ziΔz_{i\Delta} to z(i+1)Δz_{(i+1)\Delta} must then be chosen such that they assign probability to elements where there exists some j∈{1,…,n−1}j\in\{1,\ldots,n-1\} with z(i+1)Δj≠ziΔj+1z^{j}_{(i+1)\Delta}\neq z^{j+1}_{i\Delta}.

Clearly the above reasoning also holds true for Var[ViΔπ/V(i−1)Δπ]=Var[R(i−1)Δ,iΔV,π].\textrm{Var}\left[V^{\pi}_{i\Delta}/V^{\pi}_{(i-1)\Delta}\right]=\textrm{Var}[R_{(i-1)\Delta,i\Delta}^{V,\pi}]. Therefore, under such ergodicity assumptions all the expected values in Corollary 5.7 can be approximated by time averages, which will be important in our numerical implementations as outlined in Section 7.2.

2 Optimizing the (Expected) Log-(Relative)-Wealth

Similarly as the mean-variance portfolio optimization problem also the log-optimal portfolio over linear path functional portfolios can be found by solving a convex quadratic optimization task. We state it here in form of the relative log-utility optimization problem, i.e. the goal is to solve

where VtπV_{t}^{\pi} denotes again the relative wealth with respect to the universe U\mathcal{U}. Note however that under similar ergodicity conditions as in Remark 5.10 the use of time averages

Consider a universe U⊆{1,...,d}\mathcal{U}\subseteq\{1,...,d\} and a linear path functional portfolio π(μU,X)\pi(\mu^{\mathcal{U}},X). Let t0≥0t_{0}\geq 0 be the time at which we start investing. Denote by L\mathscr{L} an arbitrary but fixed labelling function. Then, for the (relative) log-utility optimization problem we have

where lL(i,ν)=lνi\mathbf{l}_{\mathscr{L}(i,\nu)}=l_{\nu}^{i} for i∈Ui\in\mathcal{U}, ν∈V\nu\in\mathcal{V} and

Moreover, it is a convex quadratic optimization problem.

Let us highlight how the expression of c\mathbf{c}, Q\mathbf{Q} simplify in the case of a signature portfolio of type II, t0=0t_{0}=0, U={1,…,d}\mathcal{U}=\{1,\dots,d\}, X=μX=\mu and μ\mu is such that it holds for some N0>0N_{0}>0

More generally, consider a signature portfolio π(μ,μ˘)\pi(\mu,\breve{\mu}) of type II where μ˘t:=(t,μt,[μ,μ]t)\breve{\mu}_{t}:=(t,\mu_{t},[\mu,\mu]_{t}) for all t∈[0,T]t\in[0,T] in a vectorized form, i.e. fix a labelling function κ(⋅,⋅)\kappa(\cdot,\cdot) such that μ˘κ(i,j)=[μi,μj]\breve{\mu}^{\kappa(i,j)}=[\mu^{i},\mu^{j}] for 1≤i,j≤d1\leq i,j\leq d and set t0=0t_{0}=0, U={1,…,d}\mathcal{U}=\{1,\dots,d\}. Then the elements of c\mathbf{c} and Q\mathbf{Q} are linear functions on the signature of μ˘\breve{\mu}, namely

Note that although we start trading at time t0≥0t_{0}\geq 0 we feed to the feature maps ϕν\phi^{\nu} the semimartingale XX starting at time . The interpretation of this is that we start “observing” the market at time and but only start trading at time t0≥0t_{0}\geq 0. This time-difference between observing and starting to invest can be important in non-Markovian settings.

Before we prove Theorem 5.11, let us state the following lemma.

For a path-functional portfolio π(μU,X)\pi(\mu^{\mathcal{U}},{X}) it holds that

We refer the reader to Appendix C for a proof of Lemma 5.14.

To show the equivalence in (5.7), we use Lemma 5.14 and make use of the fact that we are using linear path-functional portfolios.

from which the equivalence in (5.7) follows directly.

To show that the optimization problem given in (5.7) is convex, we show that the matrix Q(t)\mathbf{Q}(t) is positive semidefinite. Consider the stochastic process Yt=∑j=1d∫t0tγsjfj(s,X[0,s])dμsU,jY_{t}=\sum_{j=1}^{d}\int_{t_{0}}^{t}\gamma_{s}^{j}f^{j}(s,X_{[0,s]})d\mu^{\mathcal{U},j}_{s} whose quadratic variation [Y,Y]t[Y,Y]_{t} is given by

If the log-utility maximization problem is finite, it holds that

This simply follows from the following equivalences

Similar to Remark 5.9, the optimal vector l∗\mathbf{l}^{*} is given by

3 Constraints and Regularization

Given an optimization task involving a linear path-functional portfolio which is convex and quadratic, we can make the following statements about adding constraints or regularization:

which is convex quadratic, it remains convex quadratic under the following modifications.

Bounds Constraints: ∣lνi∣2≤bL(i,ν)|l_{\nu}^{i}|^{2}\leq b_{\mathscr{L}(i,\nu)} for bL(i,ν)≥0b_{\mathscr{L}(i,\nu)}\geq 0 for all i∈Ui\in\mathcal{U}, ν∈V\nu\in\mathcal{V}.

we can add an L2L^{2}-regularization term of the form γlTl\gamma\mathbf{l}^{T}\mathbf{l} such that the optimization problem becomes

Moreover, the above modifications can be combined and the optimization task remains convex.

∣lνi∣≤bL(i,ν)|l_{\nu}^{i}|\leq b_{\mathscr{L}(i,\nu)} for bL(i,ν)≥0b_{\mathscr{L}(i,\nu)}\geq 0 corresponds to the constraint

where δL(i,ν),L(i,ν)\delta_{\mathscr{L}(i,\nu),\mathscr{L}(i,\nu)} is the matrix

which is obviously positive semi-definite.

Transaction Costs for Portfolios with Short-Selling and Without Bank Account

When trading in the presence of (proportional) transaction costs, we need to consider re-balancing at discrete times. Let tt be an arbitrary but fixed time at which we re-balance. At time t−t^{-} (just before re-balancing) our portfolio has weights πt−\pi_{t^{-}} and a wealth of Wt−πW^{\pi}_{t^{-}}. We want to re-balance to some fixed target weights πt\pi_{t}. This situation has been studied in for the case of long-only weights. However, we allow for short-selling in our portfolios which makes the problem significantly more involved, as we illustrate below.

As the transaction costs are not paid on the portfolio weights, but on the dollar amounts invested in a stock, we define ψt−i\psi^{i}_{t^{-}}, ψti\psi^{i}_{t} to be the dollar amounts invested in stock ii directly before and after re-balancing. Note, that of course it holds

where cc is the percentage of transaction costs to be paid. The difficulty lies in solving for ψti\psi^{i}_{t}, since the amount which we can invest in stock ii depends on the amount of transaction costs to be paid and vice-versa. This difficulty arises from the fact, that we do not have a bank-account. Since we know the portfolio weights at each time, we need to solve for WtπW^{\pi}_{t}. We can use the self-financing identity which requires

Moreover, it is convenient to define α(t):=WtπWt−π\alpha(t):=\frac{W^{\pi}_{t}}{W^{\pi}_{t^{-}}} and rewrite (6.1) as

Hence, we reduced the problem to solving for αt\alpha_{t}. Since we consider an arbitrary but fixed re-balancing time tt, let us set α:=αt\alpha:=\alpha_{t} in the following. We can make the following statement about solutions for α\alpha:

Let α∗\alpha^{*} be a solution to the equation

and denote by L(α)=1−αL(\alpha)=1-\alpha and R(α)=c∑i=1d∣απti−πt−i∣R(\alpha)=c\sum_{i=1}^{d}|\alpha\pi^{i}_{t}-\pi^{i}_{t^{-}}|. We can make the following statements about α∗\alpha^{*}:

If ∑i=1d∣πt−i∣<1c\sum_{i=1}^{d}|\pi^{i}_{t^{-}}|<\frac{1}{c} a unique solution α∗∈\alpha^{*}\in exists.

If ∑i=1d∣πt−i∣≥1c\sum_{i=1}^{d}|\pi^{i}_{t^{-}}|\geq\frac{1}{c} there may not exist a solution, or at least none for α∈\alpha\in. If a solution exists, α∗≤1\alpha^{*}\leq 1 and it may not be unique.

If we have long-only portfolios (and if c<1c<1 i.e. less than 100% transaction costs) only the first case of Proposition 6.1 is relevant. Hence, in that case one always has a unique solution α∗∈\alpha^{*}\in. However, when we allow for short-selling, the second case of Proposition 6.1 becomes important.

We give a proof of Proposition 6.1 in Appendix D and move directly to the interpretation of situation in the second case of Proposition 6.1.

We want to give some interpretation of the three possible outcomes of the second case of Proposition 6.1.

no solution: The trade is infeasible. In order to pay the transaction costs, we have to reduce the dollar amounts we invest. However, reducing the dollar amounts ψt\psi_{t} leads to more transaction costs (because the difference to the previous investment amount grows, i.e. ∣ψt−ψt−∣|\psi_{t}-\psi_{t^{-}}| increases). If the transaction costs grow faster than the money we gain from reducing the dollar amount ψt\psi_{t}, the trade is infeasible.

no solution in $:Again,ifweneedtopaytransactioncostswithoutabank−accountwehavetofree−upthemoneybysellingstocks.Unlikecasea)thismightbefeasible,buttheamountoftransactioncoststobepaidismorethanthetotalwealthoftheportfolio.Thisleadstoasolutionwithnegative: Again, if we need to pay transaction costs without a bank-account we have to free-up the money by selling stocks. Unlike case a) this might be feasible, but the amount of transaction costs to be paid is more than the total wealth of the portfolio. This leads to a solution with negative\alpha$.

no unique solution in $:Here,weareinthecasewherethetransactioncostsdonotexceedourtotalwealth.Sincereducingthedollaramountleadstoanincreaseintransactioncosts,theremightbeseveralstrategieswithrespecttodollaramountsinvestedleadingtothedesiredportfolioweightsaftertransactioncosts.Forexample,onecoulddoamoreexpensivetrade(smaller: Here, we are in the case where the transaction costs do not exceed our total wealth. Since reducing the dollar amount leads to an increase in transaction costs, there might be several strategies with respect to dollar amounts invested leading to the desired portfolio weights after transaction costs. For example, one could do a more expensive trade (smaller\psi_{t})andpaymoretransactioncostsoralessexpensivetrade(larger) and pay more transaction costs or a less expensive trade (larger\psi_{t}$) where one pays less transaction costs.

In practice, we treat the above cases as follows. We consider cases a) and b) to lead to ruin, because both are caused by strategies which are infeasible given our wealth. If one of those cases occurs, we set our wealth to zero and terminate the investment. In case c) we choose the solution with highest α\alpha. This is what any reasonable investor would do: choose the trade with the least transaction costs.

Numerical Results

In this section we present our numerical results using simulated and real market data. The code corresponding to this section is available at https://github.com/janka-moeller/Sig-SPT

where Q^(t)=1M∑m=1MQ(t)(ωm)\hat{\mathbf{Q}}(t)=\frac{1}{M}\sum_{m=1}^{M}\mathbf{Q}(t)(\omega_{m}) and c^(t)=1M∑m=1Mc(t)(ωm)\hat{\mathbf{c}}(t)=\frac{1}{M}\sum_{m=1}^{M}\mathbf{c}(t)(\omega_{m}).

Moreover, we will compare the performance of the learned portfolio with the growth-optimal portfolio of the respective markets. This is justified by Lemma 2.31, which states that the growth-optimal portfolio (which is also the numeraire portfolio) solves the relative log-utility optimization problem.

For each market, we simulate MtrainM_{train} trajectories over the time-horizon $ofagivenmarketmodelandconsiderinvestmentsoverthewholetime−horizon,i.e.of a given market model and consider investments over the whole time-horizon, i.e.t_{0}=0,,t=1.WetrainasignatureportfoliousingtheMonte−Carlotypeoptimizationandtheformulasfor. We train a signature portfolio using the Monte-Carlo type optimization and the formulas for\mathbf{Q}(t)(\omega)andthevectorand the vector\mathbf{c}(t)(\omega)providedinTheorem5.11,whereweusesignatureportfoliosprovided in Theorem 5.11, where we use signature portfolios\pi(\mu,\hat{\mu}),i.e.choose, i.e. chooseX=\mutoconstructtheportfolios.Note,thatcalculatingthematrixto construct the portfolios. Note, that calculating the matrix\mathbf{Q}(t)(\omega)andvectorand vectorc(t)(\omega)foreachtraining−samplecanbeparallelizedandcomputedoffline.Havingobtainedtherespectivefor each training-sample can be parallelized and computed offline. Having obtained the respective\hat{\mathbf{Q}}(t)andand\hat{c}(t), we use the gurobipy.model.optimize()for more information, see https://www.gurobi.com/documentation/9.1/refman/py_python_api_overview.html#sec:Python. method to solve the convex quadratic optimization problem. We want to emphasize again, that we only need to solve the convex quadratic optimization task once and we never have to update\hat{\mathbf{Q}}(t)andand\hat{c}(t)$, which is very beneficial for the computational tractability of the optimization task.

After the optimization, we compare the weights and performance of the trained signature portfolio and the theoretical growth-optimal portfolio on MtestM_{test} out-of-sample trajectories.

More concretely, we consider the following three market models and signature portfolios:

We choose again d=3d=3, so that BB is a three-dimensional Brownian motion. We take α=10\alpha=10 and train a signature portfolio π(μ,μ^)\pi(\mu,\hat{\mu}) of type II of degree three. We simulate 10’000 time-steps.

We use Mtrain=100′000M_{train}=100^{\prime}000 in-sample trajectories for training and Mtest=100′000M_{test}=100^{\prime}000 out-of-sample trajectories for evaluation, i.e. computing the average logarithmic relative wealth of the respective portfolios on the test-samples. Furthermore, we used the bound constraints ∣lIi∣≤10′000|l_{I}^{i}|\leq 10^{\prime}000 for all II and for all 1≤i≤d1\leq i\leq d for each signature portfolio.

In a volatility stabilized market, signature portfolios π(μ,μ^)\pi(\mu,\hat{\mu}) of type II, with μ^=(t,μ)\hat{\mu}=(t,\mu) fall into the first case of Remark 5.12, where

Note that the growth-optimal portfolios of the Black-Scholes and volatility stabilized markets are also functionally generated portfolios in the classical sense. Indeed, in the Black-Scholes case the portfolio generating function is

with portfolio weights πt(BS),i=ci\pi^{(BS),i}_{t}=c_{i}, where

In the volatility stabilized market case the generating function of the growth optimal portfolio is given by

Therefore, these two examples are not only examples of signature portfolios approximating the respective growth-optimal portfolios but also examples for signature portfolios approximating functionally generated portfolios.

Figure 1 shows the trained signature portfolios (right) and the growth-optimal portfolios (left) evaluated at one out-of-sample trajectory in each market respectively (rows). As expected from the theoretical results, the signature weights are very similar to the theoretical growth-optimal weights in all markets. We want to emphasize, that the growth-optimal weights were never shown to the signature portfolio during training, but the signature portfolio was trained to maximize the expected logarithmic relative wealth.

Furthermore, we want to highlight, that although the growth-optimal portfolio in the Black-Scholes market has constant weights, this approximation task is far from trivial for a signature portfolio of type II. Because signature portfolios of type II approximate the controlling functions of the growth-optimal portfolio, the approximation task was actually

And likewise, in the volatility stabilized market, the approximation task was

We quantify the performance by comparing the average logarithmic relative wealth over the MtestM_{test} test samples of the growth-optimal and signature portfolios, where in both cases we compute the average logarithmic relative wealth numerically on the test samples. This explains why the signature portfolio sometimes even leads to higher values. We present these results in Table 1.

2 Real Market Data

Before we present the performance of signature portfolios, as well as, JL- and randomized-signature portfolios in real markets, we want to give some details on the optimization and investment procedure.

When working with real market data, we only have one realization available. Therefore, as already discussed in Remark 5.10 and at the beginning of Section 5.2 we replace the expectation by a time-average. More precisely, consider [0,Δ,2Δ,…,NΔ=t][0,\Delta,2\Delta,\dots,N\Delta=t] to be a equally spaced grid of [0,t][0,t] with distance Δ\Delta. Then, for s∈[0,t−Δ]s\in[0,t-\Delta] and analogously to Remark 5.10 we consider the following approximations:

Recall that sufficient ergodicity/stationarity conditions guaranteeing that these approximations are justified have been discussed in Remark 5.10.

Note however that we cannot expect that stationarity of the relative returns and log-relative returns is satisfied for every portfolio in our optimization class. Consider for example signature portfolios of type II with τ=(1d,…1d)T\tau=(\frac{1}{d},\dots\frac{1}{d})^{\mathsf{T}}, X^t=(t,Xt)\hat{X}_{t}=(t,X_{t}) and U={1,…,d}\mathcal{U}=\{1,\dots,d\}. It is easy to see that the stationarity assumption can not hold for all portfolios in this class, just by considering the following portfolio π1\pi^{1} where the only parameters that are non-zero are chosen to be l(1)dl^{d}_{(1)}. Then

However, there are of course classes of linear path-functional portfolios where the stationarity assumptions holds for all portfolios in the class. Indeed, the simplest example is the class of Markowitz-type portfolios, i.e. with τ\tau being constant and constant portfolio maps. A far-reaching generalization thereof are signature portfolios with with rolling windows, as defined in Remark 3.8 and for which stationarity holds under the conditions of Remark 5.10. Computing these portfolios is slightly more expensive, since one needs to compute the increments of the signature. Therefore we did not consider them in our implementations.

Finally, let us emphasize that the optimization using time-averages does make sense in practice, even if the stationarity of returns may not hold for every portfolio. For example consider Δ=1 day\Delta=1\textrm{ day}, then the mean-variance optimization using time-averages can be seen as looking for a portfolio with high average daily returns (in time) but without the daily returns varying too much over time and likewise for the log-relative wealth optimization. Moreover, we argue that the relative returns and log-relative returns of our optimized portfolios do exhibit some ”stability” in time, as almost all the learnt portfolios perform very well in the out-of-sample period, as we will present in the following. Note, that all of the above optimization problems remain convex quadratic optimization problems if we replace expectations by time-averages.

The main hyperparameters of our optimization procedure are:

t0t_{0}: this is the time at which we start to invest, with respect to the time when we start calculating the signature. The interpretation of this is that we can observe the market for some time before we start to invest, which is particularly relevant in the non-Markovian setting. In the following we choose t0=100t_{0}=100, hence, we always start calculating the signature 100 days before the respective investment period starts.

γ\gamma: this is the parameter of the L2L2-regularization described in Proposition 5.17. For the case of maximizing the (expected) log-relative wealth, we choose our γ\gamma during cross-validation. We do not train this regularization parameter in the mean-variance optimization but fix it a priori to a small value, as the minimization of the variance itself can already be regarded as regularizing.

For performing the cross-validation, we first split our data into a in-sample-period of TinsT_{ins} days, a cross-validation of TcvT_{cv} (consecutive to the in-sample period) and a testing period of TtestT_{test} days, starting at the end of the cross-validation period. Once we found the optimal hyper-parameters during cross-validation, we learn the parameters of our portfolios again on the TinsT_{ins} days prior to the testing period (using the obtained hyper-parameters). We then evaluate the performance of the learned portfolios during the testing-period.

In practice, we re-balance our portfolio-weights once a day since our market data is also of daily frequency. In this setting, we would like to incorporate transaction costs during optimization.

As explained in detail in Section 6, including transaction costs is not trivial in our setting. In particular, since we have shown that (6.1) does not necessarily have a (unique) solution, we cannot aim for any closed-form. Therefore, incorporating the exact amount of transaction-costs to be paid during the optimization would render the optimization task infeasible. However, we propose the following penalization which preserves the form of a convex quadratic optimization problem. Moreover, our empirical results in Subsection 7.2.3 verify that this penalization is effective in all cases and indeed useful for learning portfolios which perform well even under 5% of transaction costs.

Given a convex quadratic optimization problem of the form

adding the penalization for linear path-functional portfolios π\pi

preserves the form of a convex quadratic optimization problem. Here, β\beta is a hyperparameter that has to be chosen appropriately.

Clearly the penalization is quadratic in π\pi and therefore quadratic in the parameters {lνi}i∈U,ν∈V\{l_{\nu}^{i}\}_{i\in\mathcal{U},\nu\in\mathcal{V}}. Moreover, the penalization is positive for all {lνi}i∈U,ν∈V\{l_{\nu}^{i}\}_{i\in\mathcal{U},\nu\in\mathcal{V}} and hence the convexity is preserved as well. ∎

The motivation for this penalization is, first of all, that the universe portfolio μU\mu^{\mathcal{U}} is not punished, which should be the case because this portfolio has no transaction costs at all. Moreover, the penalization punishes changes in the weights which exceed changes in the market weights. Note that those are exactly the changes that lead to transaction costs.

The downside of this penalization is, that for a given level of transaction costs, we do not know how to choose β\beta. To find an appropriate β\beta, we propose the procedure described in Algorithm 1. That is, we solve the quadratic optimization problem for a given β\beta and calculate the in-sample performance under the true transaction costs for the found optimal portfolio. We then minimize over β\beta in order to find the one which best penalizes for transaction costs over the in-sample period. Note that choosing the initial value β0\beta_{0} can be delicate. Namely, if in a neighbourhood around β0\beta_{0} all optimal strategies lead to ruin under transaction costs, one may not move away from the initial value. Therefore, we test if β0\beta_{0} leads to ruin under transaction cost and, if so, increase it, in Algorithm 1. Note that the condition in Line 8 must be false for some β0\beta_{0} because the universe portfolio itself is included in the set of portfolios we optimize over.

In order to calculate the JL-signature, the true truncated signature needs to be calculated before applying the JL-projection. However, calculating the true signature can be computationally too heavy in large markets. We therefore propose the following memory-efficient procedure (Algorithm 2), which is based on the observation that for each component of the signature at level ll, at most ll components of the underlying path appear in the integrals. Hence, one can instead compute the signature of combinations of ll components of the path and apply the projection ”batch-wise”. It is important to note that by doing so, some words are computed multiple times, for example

can arise from the combination (X1,X2,X3)(X^{1},X^{2},X^{3}) and from (X1,X4,X5)(X^{1},X^{4},X^{5}) and from many more. Therefore, one needs to be careful which words to keep and our proposed algorithm takes care of that. Moreover, it is important to store the random projection matrix, once a realization is computed.

We parameterize the time-component of the time augmentation in the following way: Let X^t=(φT(t),Xt)\hat{X}_{t}=(\varphi_{T}(t),X_{t}). For a given trading horizon ThorT_{hor}, we consider the parametrization φThor(t)=tThor\varphi_{T_{hor}}(t)=\frac{t}{T_{hor}}. Concretely, for the in-sample period of 2000 days, we set Thor=2000T_{hor}=2000 and for the out-of-sample period we set Thor=750T_{hor}=750. This time-augmentation therefore contains information about the amount of time that is left (or has passed) in the current trading period. This is to compensate for the different training and testing periods. Another way to deal with this would be to use signature portfolios with rolling windows as defined in Remark 3.8.

Recall that we re-balance our portfolios once a day and follow a buy-and-hold strategy in between. We evaluate the out-of-sample performance of portfolios by their log-relative wealth

2.2 Rank-based approach: NASDAQ

We consider the ranked NASDAQ market and choose as a universe the stocks with ranks 1,…,1001,\dots,100. We obtained the data from the CRSP database.The raw/processed data required to reproduce our findings cannot be shared at this time due to legal reasons. We train JL-signature portfolios of dimension (50,3)(50,3) of type II and randomized signature portfolios of dimension 50 of type II investing in the universe of the 100 largest stocks and use the universe portfolio as the auxiliary portfolio. The portfolio is daily re-balanced. We train such portfolios in two ways, once in the log-wealth optimization and once in the mean-variance optimization, where in both optimization problems, we measure the performance of the portfolios by the log-relative wealth and the log-wealth achieved in the out-of-sample period.

We choose t0=100t_{0}=100 and take as an in-sample period 20002000 trading days and as an out-of-sample period the following 750750 trading days. For the mean-variance optimization task we set the L2L^{2}-regularization parameter γMV=10−6\gamma_{MV}=10^{-6} and for the log-wealth optimization task we found γLV=4.849⋅10−3\gamma_{LV}=4.849\cdot 10^{-3} for JL-signature portfolios and γLV=1⋅10−2\gamma_{LV}=1\cdot 10^{-2} for randomize-signature portfolios during the cross-validation. The grid-search for γLV\gamma_{LV} was performed over 100 equally-distant points in [10−6,10−2][10^{-6},10^{-2}].

We report the out-of-sample performance of the trained portfolios in Table 2 in terms of log-(relative)-wealth with respect to the universe portfolio. All but one of the trained portfolios outperform the universe portfolio.

In Figure 2(a) and Figure 2(b) we display the wealth processes of the trained portfolio, the one of the universe portfolio and as a benchmark also of the equally-weighted portfolio. The wealth processes of the mean-variance portfolio are the more volatile, the higher the risk-factor is, as one would expect. Moreover, the wealth-processes of the randomized signature portfolios are more tamed than their JL-counterparts.

We present the average values (in time) of the portfolio weights for each rank in Figure 3. The trained portfolios mainly take over- and under-weighted positions in the largest stocks. While the positions are extremer, the higher the risk-factor. Although we do not enforce any long-only constraints or regularization, we do not observe any extreme short-selling positions.

2.3 Name-based approach: SMI and S&P500

In the name-based setting we tackle the mean-variance and log-relative wealth optimization problems under transaction costs. We do this in two markets; the Swiss Market Index (SMI) and S&P500 Index. For the SMI, we consider the universe of stocks which survived between 2000-2022 and for the S&P500 those that survived between 2001-2022. For both markets, we obtained the data from Reuters Datastream.The raw/processed data required to reproduce our findings cannot be shared at this time due to legal reasons. We compare the performance of the trained portfolios with the universe portfolio. Note that the universe portfolio benefits from the same survivor ship bias as the signature portfolio and hence we do not have an unfair advantage. Again, we choose t0=100t_{0}=100, an in-sample period of 2000 days and an out-of-sample period of 750 days. When we include the regularization cost, we choose β0=0.5\beta_{0}=0.5 and enforce a lower bound β≥10−8\beta\geq 10^{-8}. We want to emphasize, that even with transaction costs our portfolios are re-balanced daily.

We train signature portfolios π(μU,μ^U)\pi(\mu^{\mathcal{U}},\hat{\mu}^{\mathcal{U}}) of type II degree two. We set γMV=10−8\gamma_{MV}=10^{-8} and found γLV=2.02⋅10−3\gamma_{LV}=2.02\cdot 10^{-3} during cross-validation, where the grid-search was performed over 100 equally-distant points in [10−8,10−2][10^{-8},10^{-2}]. We present the out-of-sample performance in terms of log-relative wealth in Table 3(a). The first three columns show the performance of portfolios with and without transaction costs, where we added no regularization for transaction costs. The next four columns show the performance with and without transactions costs, but with a regularization for transaction costs at the respective level. The corresponding regularization parameters β\beta which we found during the in-sample training are shown in Table 4. We note that the lower the risk-factor λ\lambda the less regularization for transaction costs is needed. However, adding such a regularization, the signature portfolios trained under the mean-variance optimization always outperform the universe portfolio under transaction costs during the out-of-sample period. This is remarkable, since the universe portfolio does not pay any transaction costs. However, the portfolio trained under log-relative wealth optimization does not out-perform the universe portfolio under transaction costs. Nevertheless, the regularization for transaction cost proved effective, since all portfolios performed better with the regularization than without, under the respective level of transaction costs.

We show the wealth processes of the signature portfolios for the SMI universe in Figure 4. It is obvious that the higher λ\lambda is, the more volatile the portfolios are, which lead to some signature portfolios under-performing the universe portfolio under transaction costs during some parts of the out-of-sample period. The signature portfolio with λ=0.05\lambda=0.05, however, did rather well over the entire period even under transaction costs, which we highlight in Figures 4(d) & 4(e).

We trained JL-signature portfolios π(μU,μ^U)\pi(\mu^{\mathcal{U}},\hat{\mu}^{\mathcal{U}}) of dimension (30,2) and type II as well as randomized-signature portfolios π(μU,μ^U)\pi(\mu^{\mathcal{U}},\hat{\mu}^{\mathcal{U}}) of dimension 30 and type II in the S&P500 market. For the mean-variance optimization, we set γMV=10−6\gamma_{MV}=10^{-6} and for obtained γLV=1.02⋅10−4\gamma_{LV}=1.02\cdot 10^{-4} for both the JL-signature portfolio and the randomized signature portfolio during cross-validation over an equally-spaced grid of 100 points in [10−6,10−2][10^{-6},10^{-2}]. We report the out-of-sample results of the optimized portfolios in Table 3(b) & 3(c) for the JL- and randomized-signature portfolios respectively. However, it is remarkable that many portfolios also out-performed the universe portfolio without any regularization for transaction costs. Indeed, in many cases the optimization yielded β=10−8\beta=10^{-8}, as we report in Table 4. In the cases where such a regularization was needed, the performance with transaction costs was always improved with the regularization. Apart from the log-wealth optimized portfolio under 5% of transaction costs, the learned portfolio all out-performed the universe portfolio under transaction costs. We show the corresponding wealth processes in Figure 6.

References

Appendix A Proofs of Section 2

Since the space of lifted stopped paths as introduced in Definition 2.7 contains paths defined on different time-intervals, we need the following definition.

We denote by (⋅)⊙(\cdot)_{\odot} the projection to the value at the final time, i.e. for each t∈[0,T]t\in[0,T]

For this projection on the final value we have the following continuity result.

where ∥⋅∥2−var;[0,T]\|\cdot\|_{2-var;[0,T]} is the norm induced by d2−var;[0,T]d_{2-var;[0,T]}. Let now ϵ>0\epsilon>0 and choose δ>0\delta>0 such that

Note that this is possible because ∥y^∥2−var;[s−δ,s+δ]→0\|\hat{\mathbf{y}}\|_{2-var;[s-\delta,s+\delta]}\to 0 as δ→0\delta\to 0.

Therefore, for all (t,x^[0,T])∈[0,T]×GTN(t,\hat{\mathbf{x}}_{[0,T]})\in[0,T]\times\mathcal{G}_{T}^{N} (equipped with the product topology) satisfying

The following lemma can be proved by means of Corollary 2.10.

is continuous on bounded sets of Gt2(φ,x)\mathcal{G}_{t}^{2}(\varphi,x) for each t∈[0,T]t\in[0,T]. Thus, for every ϵ>0\epsilon>0 there exists some δ1>0\delta_{1}>0 such that for all x^[0,t]\hat{\mathbf{x}}_{[0,t]} satisfying

is continuous for each fixed y^\hat{\mathbf{y}}. Hence, there exists some δ2>0\delta_{2}>0 such that for all tt with ∣t−s∣<δ2|t-s|<\delta_{2}

Choose now δ=min⁡(δ1,δ2)\delta=\min(\delta_{1},\delta_{2}) and let x^[0,t]\hat{\mathbf{x}}_{[0,t]} (in the bounded set) be such that

We are now prepared to prove Theorem 2.11.

Consider the function λ\lambda as introduced in Lemma A.2 and note that for any compact subset K⊂GT2(φ,x)K\subset\mathcal{G}_{T}^{2}(\varphi,x) the set

the set of path functionals induced by linear functions on the signature and by

which is equivalent to the claim of Theorem 2.11. ∎

The proof of the global UAT stated in Theorem 2.18 is based on the fact that ΛT2,α′(φ,x)\Lambda_{T}^{2,\alpha^{\prime}}(\varphi,x) with 0≤α′<α<120\leq\alpha^{\prime}<\alpha<\frac{1}{2} is a weighted space. This is proved below.

We have to prove that KR=ψ−1((0,R])K_{R}=\psi^{-1}((0,R]) is compact for any R>0R>0. To this end, note that the set

equipped with the dα′,[0,T]d_{\alpha^{\prime},[0,T]}-topology is for any R>0R>0 a ∥⋅∥α,[0,T]\|\cdot\|_{\alpha,[0,T]}-bounded subset of GTN(φ,x)\mathcal{G}^{N}_{T}(\varphi,x) and hence compact for 0≤α′<α0\leq\alpha^{\prime}<\alpha due to the compact embedding of α\alpha-Hölder spaces into α′\alpha^{\prime}-Hölder spaces, see [22, Theorem A.3]. Moreover, KR=λ([0,T]×KR′)K_{R}=\lambda([0,T]\times K_{R}^{\prime}) and therefore compact for any R>0R>0 due to Tychonoff’s theorem and the continuity of λ\lambda. ∎

Using Lemma 2.17 the proof is analogous to the proof of [22, Theorem 5.4]. The only difference here is that we can choose ξ>1\xi>1 and do not need to require ξ>2\xi>2 in the weight function ψ\psi as it would be the case in the setting of [22, Theorem 5.4]. The reason for this is that we consider here semimartingales which almost surely satisfy the ‘RIE’ property introduced in (see also [3, Proposition 4.1]). Hence, we can construct the Stratonovic lifts pathwise [58, Lemma 4.18] which then almost surely coincide with the stochastic lifts. Due to this pathwise construction, the first level of the paths uniquely determines the second order lifts and therefore the set

suffices as a point-separating sub-algebra and is of ψ\psi-moderate growth, even for ξ>1\xi>1. ∎

Appendix B Proofs of Section 4

The following lemma is needed in the proof of Theorem 4.3.

By assumption the determinant (ad−bc)(ad-bc) is non-zero, which implies continuity of the coefficients of M−1M^{-1}.

Assume the statement holds for an arbitrary but fixed n≥2n\geq 2.

Appendix C Proofs of Section 5

Note that the terms in (C.1) and (C.2) are both zero due to ∑i∈UμtU,i≡1\sum_{i\in\mathcal{U}}\mu_{t}^{\mathcal{U},i}\equiv 1.

Appendix D Proofs of Section 6

This case always holds for long-only portfolios and is discussed in . The argument in the long-only case is that L(1)<R(1)L(1)<R(1) and L(0)>R(0)L(0)>R(0), while both sides are obviously continuous and monotone. While the former statement still holds in the case with short-selling, the monotonicity is no longer true. However, one can show by a tedious case-by-case study that where R′(α)R^{\prime}(\alpha) exists it is piece-wise constant and increasing. Therefore a unique solution for α∗∈\alpha^{*}\in is still assured in this case.

Here the above argument does not hold any more. This is because now L(0)≤R(0)L(0)\leq R(0), while it still holds that L(1)<R(1)L(1)<R(1). Let us give three examples, one for each of the cases a) no solution b) no solution in ,c)nouniquesolutionin, c) no unique solution in.

no solution: consider a market of two stocks and c=0.05c=0.05, i.e. 5% transaction costs. For the portfolio weights

there does not exist a solution for α\alpha.

no solution in $:againfor: again forc=0.05$ and a market of two stocks. For the portfolio weights

no unique solution in $:consider: considerc=0.05$ and a market of three stocks with portfolio weights

then the solutions are α=0.3333\alpha=0.3333 and α=0.9535\alpha=0.9535 (numbers are rounded).