Robust pricing and hedging via neural SDEs

Patryk Gierjatowicz, Marc Sabate-Vidales, David Šiška, Lukasz Szpruch, Žan Žurič

Introduction

Model uncertainty is an essential part of mathematical modelling but is particularly acute in mathematical finance and economics where one cannot base models on well established physical laws. Until recently, these models were mostly conceived in a three step fashion: 1) gathering statistical properties of the underlying time-series or the so called stylized facts; 2) handcrafting a parsimonious model, which would best capture the desired market characteristics without adding any needless complexity and 3) calibration and validation of the handcrafted model. Indeed, model complexity was undesirable, amongst other reasons, for increasing the computational effort required to perform in particular calibration but also pricing and risk calculations. With greater uptake of machine learning methods and greater computational power more complex models can now be used. This is due to the fact that arguably the most complicated and computationally expensive step of calibration has been addressed. Indeed, in the seminal paper [Hernandez, 2016] used neural networks to learn the calibration map from market data directly to model parameters. Subsequently, many papers followed [Liu et al., 2019, Ruf and Wang, 2019, Ruf and Wang, 2019, Benth et al., 2020, Gambara and Teichmann, 2020, Sardroudi, 2019, Horvath et al., 2019, Bayer et al., 2019, Bayer and Stemper, 2018, Vidales et al., 2018]. However, these approaches focused on the calibration of fixed parametric model, but did not address perhaps even the more important issue which is model selection and model uncertainty.

The approach taken in this paper is fundamentally different. We let the data dictate the model, while still keeping a strong prior on the model form. This is achieved by using SDEs for the model dynamics but instead of choosing a fixed parametrization for the model SDEs we allow the drift and diffusion to be given by an overparametrized neural networks. We will refer to these as Neural SDEs. These are shown to not only provide a systematic framework for model selection, but also, quite remarkably, to produce robust estimates on the derivative prices. Here, the calibration and model selection are done simultaneously. In this sense, model selection is data-driven. Since the neural SDE model is overparametrised, there is a large pool of possible models and the training algorithm selects a model. Unlike in handcrafted models, individual parameters do not carry any meaning. This makes it hard to argue why one model is better than another. Hence the ability to efficiently compute interval estimators, which algorithms in this paper provide, is critical.

In parallel to this work, a similar approach to modelling was taken in [Cuchiero et al., 2020], where the authors considered local stochastic volatility models with the leverage function approximated with a neural network. Their model can be seen as an example of a Neural SDEs.

2. Neural SDEs

Observe that σS\sigma^{S} and σV\sigma^{V} encode arbitrary correlation structures between the traded assets and the non-tradable components. Moreover, we immediately see that (e−rtSt)t∈[0,T](e^{-rt}S_{t})_{t\in[0,T]} is a (local) martingale and thus the model is free of arbitrage.

As we shall see in Sections 5.1 and 5.2 there are many Neural SDE models that can be calibrated well to market data and that produce significantly different prices for derivatives that were not part of the calibration data. In practice these would be illiquid derivatives where we require model to obtain prices. Therefore, we compute price intervals for illiquid derivatives within the class of calibrated neural SDEs models. To be more precise we compute

We solve the above constraint optimisation problem by penalisation. See [Eckstein and Kupper, 2019] for related ideas.

3. Key conclusions and methodological contributions of this paper

The results in this paper presented below lead to the following conclusions.

Neural SDEs provide a systematic framework for model selection and produce robust estimates on the derivative prices. The calibration and model selection are done simultaneously and the thus the model selection is data-driven.

With neural SDEs, the modelling choices one makes are: networks architectures, structure of neural SDE (e.g. traded and non-traded assets), training methods and data. For classical handcrafted models the choice of the algorithm for calibrating parameters has not been considered as part of modelling choice, but for machine learning this is one of the key components. See Section 5, where we show how the change in initialisation of stochastic gradient method used for training leads to different prices of illiquid options, thus providing one way of obtaining price bounds. Furthermore even for basic local volalitly model that is unique for continuum of strikes and maturities, produces ranges of prices of illiquid derivatives when calibrated to finite data sets.

The above optimisation problem is not convex. Nonetheless, empirical experiments in Sections 5.1-5.2 demonstrate that the stochastic gradient decent methods used to minimise the loss functional DD converges to the set of parameters for which calibrated error is of order 10−510^{-5} to 10−410^{-4} for the square loss function. Theoretical framework for analysing such algorithms is being developed in [Šiška and Szpruch, 2020].

By augmenting classical risk models with modern machine learning approaches we are able to benefit from expressibility of neural networks while staying within realm of classical models, well understood by traders, risk managers and regulators. This mitigates, to some extent, the concerns that regulators have around use of black-box solutions to manage financial risk. Finally while our focus here is on SDE type models, the devised framework naturally extends to time-series type models.

The main methodological contributions of this work are as follows.

By leveraging martingale representation theorem we developed an efficient Monte Carlo based methods that simultaneously learns the model and the corresponding hedging strategy.

The calibration problem does not fit into classical framework of stochastic gradient algorithms, as the mini-batches of the gradient of the cost function are biased. We provide analysis of the bias and show how the inclusion of hedging strategies in training mitigates this bias.

We devise a novel, memory efficient randomised training procedure. The algorithm allows us to keep memory requirements constant, independently of the number of neural networks in neural SDEs. This is critical to efficiently calibrate to path dependent contingent claims. We provide theoretical analysis of our method in Section 4 and numerical experiment supporting the claims in Section 5.

The paper is organized as follows. In Section 2 we outline the exact optimization problem, introduce a deep neural network control variate (or hedging strategy), address the process of calibration to single/multiple option maturities and state the exact algorithms. In Section 3 we analyse the bias in Algorithms 1 and 2. In Section 4 we show that the novel, memory-efficient, drop-out-like training procedure for path-dependent derivatives does not introduce bias in the new estimator. Finally, the performance of Neural Local Volatility and Local Stochastic Volatility models is presented in Section 5. Some of the proofs and more detailed results from numerical experiments are relegated to the Appendix. The code used is available at github.com/msabvid/robust_nsde.

Robust pricing and hedging

Find model parameters θ∗\theta^{*} such that model prices match market prices:

Find model parameters θl,∗\theta^{l,*} and θu,∗\theta^{u,*} which provide robust arbitrage-free price bounds for an illiquid derivative, subject to available market data:

In the following we construct Φcv\Phi^{cv} using hedging strategy. Similar approach has recently been developed in [Vidales et al., 2018] in the context of pricing and hedging with deep networks.

Martingale representation theorem (see for example Th. 14.5.1 in [Cohen and Elliott, 2015]) provides a general methodology for finding Monte Carlo estimators with the above stated properties (2.3).

In a similar manner one can derive Ψcv\Psi^{cv} for the payoff of the illiquid derivative for which we seek the robust price bounds. Then (2.2) can be restated as

The learning problem (2.5) is better than (2.2) from the point of view of algorithmic implementation, as it will enjoy lower Monte Carlo variance and hence will require simulation of fewer paths of the Neural SDE in each step of the stochastic gradient algorithm. Furthermore, when using (2.5) we learn a (possibly abstract) hedging strategy for trading in the underlying asset to replicate the derivative payoff. Since the market may be incomplete this abstract hedging strategy may not be usable in practice. More precisely, since the process XθX^{\theta} will contain tradable as well as non-tradable assets, the control variate for the latter has to be adapted by either performing a projection or deriving a strategy for the corresponding tradable instrument.

To deduce a real hedging strategy recall that Xθ=(Sθ,Vθ)X^{\theta}=(S^{\theta},V^{\theta}) with SθS^{\theta} being the tradable assets and VθV^{\theta} the non-tradable components. Decompose the abstract hedging strategy as h=(hS,hV)\mathfrak{h}=(\mathfrak{h}^{S},\mathfrak{h}^{V}). Let Sˉtθ:=e−rtStθ\bar{S}^{\theta}_{t}:=e^{-rt}S^{\theta}_{t} and note that due to (1.2) we have

If we can solve hS=e−rthˉtSσS(t,Xtθ,θ)\mathfrak{h}^{S}=e^{-rt}\bar{\mathfrak{h}}^{S}_{t}\sigma^{S}(t,X^{\theta}_{t},\theta) for hˉtS\bar{\mathfrak{h}}^{S}_{t} then this is a real hedging strategy.

Therefore, an alternative approach to (2.4), possibly yielding a better hedge, but worse variance reduction would be to consider finding

for some other neural network hˉ\bar{\mathfrak{h}}. This is the version we present in Algorithms 1 and 2.

2. Time discretization

In order to implement the (2.5) we define partition π\pi of [0,T][0,T] as π:={t0,t1,…,tNsteps=T}\pi:=\{t_{0},t_{1},\ldots,t_{N_{\text{steps}}}=T\}. We first approximate the stochastic integral in (2.4) with the appropriate Riemann sum. Depending on the choice of the neural network architecture approximating σ\sigma in the Neural SDE (1.1) we may have σ\sigma which grows super-linearly as a function of xx. In such a case the moments of the classical Euler scheme are known blow up in the finite time, see [Hutzenthaler et al., 2011], even if moments of the solution to the SDE are finite. In order to avoid blow ups of moments of the simulated paths during training we apply tamed Euler method, see [Hutzenthaler et al., 2012, Szpruch and Zhāng, 2018]. The tamed Euler scheme is given by

with Δtk=tk+1−tk\Delta t_{k}=t_{k+1}-t_{k} and ΔWtk+1=Wtk+1−Wtk\Delta W_{t_{k+1}}=W_{t_{k+1}}-W_{t_{k}}.

3. Algorithms

We now present the algorithm to calibrate the Neural SDE (1.1) to market prices of derivatives (Algorithm 1) and the algorithm to find robust price bounds for an illiquid derivative (Algorithm 2). Note that during training we aim to calibrate the SDE (1.1), and at the same time, adapt the abstract hedging strategy to minimise the variance (2.4). Therefore, we alternate two optimisations:

During each epoch, we optimise the parameters ξ\xi of the parametrisation of the hedging strategy, while the parameters θ\theta of the Neural SDE are fixed.

4. Algorithm for multiple maturities

Algorithms 1 and 2 calibrate the SDE (5.3) to one set of derivatives. If the derivatives for which we have liquid market prices can be grouped by maturity, as is the case e.g. for call / put prices, we can use a more efficient algorithm to achieve the calibration.

This follows the natural approach used e.g. in [Cuchiero et al., 2020], and in [Vidales et al., 2018] in the context of learning PDEs, where the networks for b(t,Xtθ,θ)b(t,X_{t}^{\theta},\theta) and σ(t,Xtθ,θ)\sigma(t,X_{t}^{\theta},\theta) are split into different networks, one per maturity. Let θ=(θ1,…,θNm)\theta=(\theta_{1},\ldots,\theta_{N_{m}}), where NmN_{m} is the number of maturities. Let

with each bib^{i} and σi\sigma^{i} a feed forward neural network.

Regarding the SDE parametrisation, we fit feed-forward neural networks (see Appendix C) to the diffusion of the SDE of the price process under the risk-neutral measure. In the particular case, where we calibrate the Neural SDE to market data, without imposing any bounds on the resulting exotic option prices, one can then do an incremental learning as follows,

Consider the first maturity TiT_{i} with i=1i=1.

Calibrate the SDE using Algorithm 1 to the vanilla prices in maturity TiT_{i}.

Freeze the parameters of σi\sigma_{i}, set i:=i+1i:=i+1, and go back to previous step.

The above algorithm is memory efficient, as it only needs to backpropagate through that last maturity in each gradient descent step.

Analysis of the stochastic approximation algorithm for the calibration problem

Nevertheless we know that the classical gradient algorithm, with the learning rates (ηk)k=1∞(\eta_{k})_{k=1}^{\infty}, ηk>0\eta_{k}>0 for all kk, applied to this optimisation problem is given by

2. Stochastic algorithm for the calibration problem

Recall that our overall objective in calibration is to minimize some J=J(θ)J=J(\theta) given by

We differentiate J=J(θ)J=J(\theta) and work with the pathwise representation of this derivative (using language from [Glasserman, 2013]). For that we impose the following assumption.

We refer reader to [Glasserman, 2013, chapter 7] for exact conditions when exchanging integration and differentiation is possible. We also remark that for the payoffs for which the Assumption 3.1 does not hold, one can use likelihood method and more generally Malliavin weights approach for computing greeks, [Fournié et al., 1999]. We don’t pursue this here for simplicity.

If we had one network for each time step, leading to some resnet-like-network architecture for the time discretization then it may be more efficient to use a backward equation representation in the training. This representation can be derived using similar analysis as in [Jabir et al., 2019] (see also [Šiška and Szpruch, 2020]).

Since the summation plays effectively no role in further analysis we will assume, without loss of generality, that M=1M=1 and work with the objective

In Appendix A we provide a result on the bias for a general loss function. For the square loss function the bias is given below.

This is an immediate consequence of Theorem A.1. ∎

Hence, we see that by reducing the variance of the first term we are also reducing the bias of the gradient. This justifies superiority of learning task (2.5) over (2.2).

Analysis of the randomised training

While the idea of calibrating to one maturity at the time described in Section 2.4 works well if our aim is only to calibrate to vanilla options, it cannot be directly applied to learn robust bounds for path dependent derivatives, see (2.2). This is because the payoff of path dependent derivatives, in general, is not an affine function of maturity. On the other hand training all neural networks at every maturity all at once, makes every step of the gradient algorithm used for training computationally heavy.

In what follows we introduce a randomisation of the gradient so that at each step of the gradient algorithm the derivatives with respect to the network parameters are computed at only one maturity at the time while keeping parameters at all other maturities unchanged. This is similar to the popular dropout method, see [Srivastava et al., 2014], that is known to help with overfitting when training deep neural networks but for us the main aim is computational efficiency. Recall how we split the networks for drift and diffusion

where bU,σUb^{U},\sigma^{U} are simply neural networks sampled from the random index UU.

Less stringent assumption on derivatives of bb and σ\sigma are possible, but we do not want to overburden the present article with technical details.

It is well known, e.g [Krylov, 1999, Kunita, 1997], that

Now using Fubini-type Theorem for Conditional Expectation, [Hammersley et al., 2019, Lemma A5], we have

See also [Cuchiero et al., 2020] for the same observation. The gradient of hh is given by

To implement the above algorithm one simply needs to generate two independent sets of samples. Furthermore,

is an unbiased estimator of (∂θhN)(θ)(\partial_{\theta}h^{N})(\theta).

Testing neural SDE calibrations

All algorithms were implemented using PyTorch, see [Paszke et al., 2017] and [Paszke et al., 2019]. The code used is available at github.com/msabvid/robust_nsde. Our target data (European option prices for various strikes and maturities) is described in Appendix B. We assume that there is one traded asset S=(St)t∈[0,T]S=(S_{t})_{t\in[0,T]}. We calibrate to European option prices

for maturities of 2,4,…,122,4,\ldots,12 months and typically 2121 uniformly spaced strikes between in [0.8,1.2][0.8,1.2]. As an example of an illiquid derivative for which we wish to find robust bounds we take the lookback option

In this section we consider a Local Volatility (LV) Neural SDE model. It has been shown by [Dupire et al., 1994] (see also [Gyöngy, 1986]) that if the market data would consist of a continuum of call / put prices for all strikes and maturities then there is a unique function σ\sigma such that with the price process

the model prices and market prices match exactly. In practice only some calls / put prices are liquid in the market and so to apply [Dupire et al., 1994] one has to interpolate, in an arbitrage free way, the missing data. The choice of interpolation method is a further modelling choice on top of the one already made by postulating that the risky asset evolution is governed by (5.1).

We will use a Neural SDE instead of directly interpolating the missing data. Let our LV Neural SDE model be given by

2. Local stochastic volatility neural SDE model

In this section we consider a Local Stochastic Volatility (LSV) Neural SDE model. See, for example, [Tian et al., 2015]. As in the Local Volatility Neural SDE model (5.2), we have the risky asset price price process (St)t∈[0,T](S_{t})_{t\in[0,T]}, where the drift is equal to the risk-free bond rate rr. However, the volatility function in the LSV Neural SDE model now depends on tt, StS_{t} and a stochastic process (Vt)t∈[0,T](V_{t})_{t\in[0,T]}. Here (Vt)t∈[0,T](V_{t})_{t\in[0,T]} is not a traded asset. The model is then given by

3. Deep learning setting for the LV and LSV neural SDE models

In the SDE (5.2) the function σ\sigma and the SDE (5.3) the functions σS\sigma^{S}, bVb^{V} and σV\sigma^{V} are parametrised by one feed-forward neural network per maturity (see Section 2.4 and Appendix C) with 44 hidden layers with 5050 neurons in each layer. The non-linear activation function used in each of the hidden layers is the linear rectifier relu. In addition, in σS\sigma^{S} and σV\sigma^{V} after the output layer we apply the non-linear rectifier softplus(x)=log⁡(1+exp⁡(x))(x)=\log(1+\exp(x)) to ensure a positive output.

The parameterisation of the hedging strategy for the vanilla option prices is also a feed-forward linear network with 3 hidden layers, 20 neurons per hidden layer and relu activation functions. However, in order to get one hedging strategy per vanilla option considered in the market data, the output of h(tk,stk,θi,ξKj)\mathfrak{h}(t_{k},s_{t_{k},\theta}^{i},\xi_{K_{j}}) has as many neurons as strikes and maturities.

Finally, the parameterisation of the hedging strategy for the exotic options price is also a feed-forward network with 3 hidden layers, 20 neurons per hidden layer and relu activation functions.

The Neural SDEs (5.2) and (5.3) were discretized using the tamed Euler scheme (2.7) with Nsteps=8×12N_{\text{steps}}=8\times 12 uniform time steps for T=1T=1 year (i.e. 1616 for every 22 months). The number of Monte Carlo trajectories in each stochastic gradient descent iteration was N=4×104N=4\times 10^{4} and the abstract hedging strategy was used as a control variate.

4. Conclusions from calibrating for LV neural SDE

It is possible to obtain high accuracy of calibration with MSE of about 10−910^{-9} for 66 month maturity, about 10−810^{-8} for 12 month maturity when the only target is to fit market data. If we are minimizing / maximizing the illiquid derivative price at the same time then the MSE increases somewhat so that it is about 10−810^{-8} for both 66 and 1212 month maturities. See Figure 5.1. The calibration has been performed using K=21K=21 strikes.

The calibration is accurate not only in MSE on prices but also on individual implied volatility curves, see Figure 5.2 and others in Appendix D.

As we increase the number of strikes per maturity the range of possible values for the illiquid derivative narrows. See Figure F.1 and Tables 1, 2, 3 and 4. The conjecture is that as the number of strikes (and maturities) would increase to infinity we would recover the unique σ\sigma given by the Dupire formula that fits the continuum of European option prices.

With limited amount of market data (which is closer to practical applications) even the LV Neural SDE produces noticeable ranges for prices of illiquid derivatives, see again Figure 5.1.

In Appendix F we provide more details on how different random seeds, different constrained optimization algorithms and different number of strikes used in the market data input affect the illiquid derivative price.

In Appendix D we present market price and implied volatility fit for constrained and unconstrained calibrations. High level of accuracy in all calibrations is achieved due to the hedging neural network incorporated into model training.

5. Conclusions from calibrating for LSV neural SDE

We note that our methods achieve high calibration accuracy to the market data (measured by MSE) with consistent bounds on the exotic option prices. See Figure 5.3.

The calibration is accurate not only in MSE on prices but also on individual implied volatility curves, see Figure 5.4 and others in Appendix E.

The LSV Neural SDE produces noticeable ranges for prices of illiquid derivatives, see again Figure 5.3.

6. Hedging strategy evaluation

We calculate the error of the portfolio hedging strategy of the lookback option at maturity T=6T=6 months, given by the empirical variance

The histogram in Figure 5.5 is calculated on N=400 000N=400\,000 different paths and provides the values of ss,

Finally, we study the effect of the control variate parametrisation on the learning speed in Algorithm 1. Figure 5.6 displays the evolution of the Root Mean Squared Error of two runs of calibration to market vanilla option prices for two-months maturity: the blue line using Algorithm 1 with simultaneous learning of the hedging strategy, and the orange line without the hedging strategy. We recall that from Section 2.4, the Monte Carlo estimator ∂θhN(θ)\partial_{\theta}h^{N}(\theta) is a biased estimator of ∂θh(θ)\partial_{\theta}h(\theta) An upper bound of the bias is given by Corollary 3.2, that shows that by reducing the variance of Monte Carlo estimator of the option price then the bias of ∂θhN(θ)\partial_{\theta}h^{N}(\theta) is also reduced, yielding better convergence behaviour of the stochastic approximation algorithm. This can be observed in Figure 5.6.

Acknowledgements

This work was supported by the Alan Turing Institute under EPSRC grant no. EP/N510129/1. We thank Antoine Jacquier (Imperial) for fruitful discussions on the topic of the paper.

Declarations of Interest

The authors report no conflicts of interest. The authors alone are responsible for the content and writing of the paper.

References

Appendix A Bound on bias in gradient descent

We complete the analysis from Section 3.2 for a general loss function here.

Let Assumption 3.1 hold. Consider the family of neural SDEs (1.1). We have

Appendix B Data used in calibration

We used Heston model to generate prices of calls and puts. The model is

It is well know that for this model a semi-analytic formula can be used to calculate option prices, see [Heston, 1997] but also [Albrecher et al., 2007]. The choice of parameters below was used to generate target model calibration prices.

Options with bi-monthly maturities up to one year with varying range of strikes were used as market data for Neural SDE calibration. The call / put option prices were obtained from the Heston model using Monte Carlo simulation with 10710^{7} Brownian trajectories. We use bimonthly maturities up to one year for considered calibrations. Varying range of strikes is used among different calibrations. See Figure B.1 for the resulting “market” data.

Appendix C Feed-forward neural networks

Appendix D LV neural SDEs calibration accuracy

Figures 5.2, D.2, D.4 present implied volatility fit of local volatility neural SDE model (5.2) calibrated to: market vanilla data only; market vanilla data with lower bound constraint on lookback option payoff; market vanilla data with upper bound constraint on lookback option payoff respectively. Figures D.1, D.3, D.5 present target option price fit of local volatility neural SDE model (5.2) calibrated to: market vanilla data only; market vanilla data with lower bound constraint on lookback option payoff; market vanilla data with upper bound constraint on lookback option payoff respectively. High level of accuracy in all calibrations is achieved due to the hedging neural network incorporated into model training.

Appendix E LSV neural SDEs calibration accuracy

Figures E.1, 5.4 and E.2 provide the Vanilla call option price and the implied volatility curve for the calibrated models. In each plot, the blue line corresponds to the target data (generated using Heston model), and each orange line corresponds to one run of the Neural SDE calibration. We note again in this plots how the absolute error of the calibration to the vanilla prices is consistently of O(10−4)\mathcal{O}(10^{-4}).

Appendix F Exotic price in LV neural SDEs

Below we see how different random seeds, constrained optimization algorithms and number of strikes used in the market data input affect the illiquid derivative price in the Local Volatility Neural SDE model.