Deep Hedging: Learning to Simulate Equity Option Markets
Magnus Wiese, Lianjun Bai, Ben Wood, Hans Buehler
Introduction
There is growing interest in applying reinforcement learning techniques to the problem of managing a portfolio of derivatives . This involves buying and selling not only the relevant underlying assets, but also the available exchange-traded options. In order to train an option trading model, we therefore require time-series data that includes option prices.
Unfortunately, the amount of useful real-life data available is limited; if we take a sampling interval of one day, ten years of option prices translates into only a few thousand samples. This motivates the need for a realistic simulator: we can generate much larger volumes of data for training and evaluation while preserving the key distributional properties of the real samples, thus helping to avoid overfitting.
In this article, we build and test a collection of simulators based on neural networks (NNs). We begin by transforming the option prices to an equivalent representation in the form of discrete local volatilities (DLVs) with less complicated no-arbitrage constraints; we then formalize the role of the simulator as a mapping from a state and input noise to a new set of (transformed) option prices. Next, we define a series of benchmark scores based on the key distributional features of the transformed prices. We construct a set of generative models, varying the network architecture, training method, and state compression scheme. Finally, we evaluate these models against our proposed benchmark scores, and compare with a classical baseline.
Financial time series simulation
Recently, generative adversarial networks (GANs) have been successfully used to create realistic synthetic time series for asset prices . Zhang et al. and Zhou et al. reported on using the objective function of GANs to predict spot prices. Koshiyama et al. trained a conditional GAN to generate spot log-returns and provided a study of using generated paths to fine-tune trading strategies, and showed that the autocorrelation function (ACF) and partial autocorrelation function (PACF) could be generated accurately by using a two-layer perceptron. Takahashi et al. also reported on being able generate various stylised facts found in the historical series, but did not provide a detailed description of the methodology used. An unconditional approach to the generation of spot price log-returns, using Temporal Convolutional Networks (TCNs) , was first presented by Wiese et al. ; they also reproduce the relevant stylised facts, including volatility clustering and the leverage effect.
In contrast, generative models for time series of option prices are much less common: Cont performs a principal component analysis (PCA) on implied volatility data; Wissel provides a scheme to build a risk-neutral market model, focusing on ensuring the martingale property rather than realism. As far as we are aware, neural networks have not previously been applied to option market generation.
Our work extends the conditional modelling framework of to the multivariate setting by using GANs and other calibration techniques.
Option prices
Our aim in this article is to simulate the prices of standard “vanilla” equity index options, of the type commonly traded on exchanges in large volumesEurex reports an average of more than one million option contracts traded daily on EURO STOXX 50 in August 2019, corresponding to almost € 650MM in premiums paid.. An option is characterized by: the underlying index; the type (call or put); the strike, ; and the maturity, . The maturity is the expiry date of the option; on that date, the option holder receives an amount \max\bigl{(}0,s(I_{T}-K)\bigr{)}, where is the prevailing level of the underlying index, and for a call and for a put.
At any given time, not all strike/maturity/type combinations are tradable; market makers quote bid and/or offer prices for only the most relevant combinations, which broadly means those with strike closest to the current index level, and maturity closest to today. We will work with a representative grid of strikes and maturities.
Market participants commonly express option prices in terms of implied volatilities. The implied volatility is the number one must use in the standard Black-Scholes formula to obtain the option price. The other inputs to the formula are the discount factor and the index forward price. All three inputs to the price are stochastic, but the short-term price dynamics are primarily driven by changes in the implied volatility; in this paper, we focus on this contribution.
Option prices are subject to strict ordering constraints because of no-arbitrage considerations. For example, since the call option payoff is a non-increasing function of strike, the option price must also be non-increasing; violation of this rule would constitute an arbitrage opportunityStrictly, there is an arbitrage opportunity when the bid price of the higher-strike call option exceeds the offer price of the lower-strike option., i.e. the possibility of making a profit while guaranteeing no loss. True arbitrage opportunities are rare, fleeting, and small; an option price simulator is more useful if it does not generate arbitrageable market states.
For this reason, it is not convenient to work with option prices directly; the ordering constraints are too complicated. The equivalent constraints on implied volatilities are even more awkward. Instead, we transform the prices to DLVs , for which the absence of arbitrage corresponds to a simple requirement of positivity. In the absence of arbitrage, this mapping is bijective; an grid of option prices is converted into an grid of DLVs.
In this article we will focus on one of the largest option markets, namely options on the EURO STOXX 50 index. An overview of the distributional and dependence properties of DLVs is provided in Appendix B. The implied volatilities of the EURO STOXX 50 for the relevant sets of strikes and maturities are displayed in Figure 1.
Problem formulation
We assume a set of equispaced strikes
for which we obtain the -dimensional process of DLVs
where is -dimensional Gaussian noise and the state a function of the processes history, i.e. . An example of is a projection onto the most current component, i.e. , the last realisations, or an exponentially weighted moving average.
The objective is to approximate the mapping which ideally allows us to generate more data from a given state . In this paper our approach is to represent this mapping through a deep neural network
Models
The first (naive) approach to model log-DLVs is to approximate (2) directly and generate all DLVs yielding the generated time-series
The compressed version is motivated from the observation that log-DLV levels have high cross-correlations indicating that log-DLVs live on a lower-dimensional manifold. Figure 3 illustrates the cross-correlation matrix of log-DLVs and DLV log-returns for the set of relative strikes and maturities . Both cross-correlation matrices show high cross-correlations between groups of different maturities as well as within a specific maturity.
Figure 3 illustrates the cumulative sum of variance explained as a function of the number of principle components for the grid . In section 8 we report our results for principle components which explain of the variance yielding a good trade-off between dimensionality reduction and preserving information.
Optimization
The next step is to obtain a close approximation of . Various training procedures exist ranging from conventional parametric methods such as quasi maximum likelihood estimation (qMLE) to novel non-parametric ones such as GANs and the mean maximum discrepancy (MMD) . In this paper, we explore the performance of qMLE, GANs and Wasserstein GANs (WGAN-GP) . Following, we provide a brief explaination of the qMLE and GANs in the context of our explicit conditional model (3).
A challenge in qMLE is the correct specification of and the intractability of likelihood functions. GANs try to address this issue by introducing a min-max two-player game between the generator and the discriminator . The discriminator aims to discriminate between real samples from the data distribution and synthetic ones generated by the generator. The objective function proposed by the original paper from Goodfellow adapted to our setting is of the form
A drawback of GANs is that they are notoriously hard to train which lead to the introduction of various regularization techniques to stabilize training . In our numerical results, we grid search over spectral normalization imposed on the discriminator and generator and gradient penalities proposed by Mescheder . Furthermore, we also train our generative model by using WGAN-GP proposed by and report on these results separately in section 8.
Evaluating the generated paths
In GANs the objective function cannot be used to evaluate the performance of the generator. Equally, using the likelihood of a qMLE-trained model can give a distorted image. To measure the goodness and performance of a generative model we define and introduce various metrics and scores. These scores allow us to capture whether the generator is able to generate dynamics that are similar to those found in the historical DLV series such as highly cross-correlated log-DLVs and DLV log-returns, bimodal distributions or persistence in the autocorrelation.
During training we intentionally evaluate log-DLVs instead of implied volatilities due two reasons. First, a close approximation of the generating mechanism of DLVs yields a close approximation of the generating mechanism of implied volatilities. Second, transforming DLVs to implied volatilities is costly. Once a close approximation of is obtained we transform the generated DLVs to implied volatilities and compute metrics and scores for the generated implied volatilities and report on those.
We begin by introducing a distributional metric and distributional scores, then define dependence scores and at last two scores that take into account the cross-correlation structure.
2 Distributional scores
In financial applications higher order moments such as the skewness and kurtosis are of interest as they determine the propensity to generate extremal values. We therefore define the skewness score
3 Dependence scores
Since DLVs are persistent we adopted ACF score proposed in . It is defined by taking the Euclidean norm of the difference of the historical and the mean generated autocorrelation function
Similarly, we define the ACF score for the log-return process
In section 8 we report the score for the first lags and the score for the first lag.
4 Cross-correlation scores
To capture whether the generator generates cross-correlated log-DLVs and DLV log-returns we introduce two more scores. The cross-correlation score of log-DLVs is defined by taking the Euclidean norm of the cross-correlation matrix of log-DLVs where denote the cross-correlation matrix of the historical and generated respectively. Likewise, the cross-correlation score of DLV log-returns is defined where denote the cross-correlation matrix of the historical and generated DLV log-returns respectively.
Numerical results
In this section we evaluate the performance of qMLE-, GAN- and WGAN-GP-trained models for the compressed and explicit version.
We use call option prices of the EURO STOXX 50 from 2011-01-03 to 2019-08-30 and consider the set of relative strikes and maturities
For those option prices we compute the path of DLVs which we use for training (see Figure B.1). The total length of the time series is .
2 Benchmarks
For comparison we apply the vector autoregressive model and GAN-trained TCNs to the same data. is a standard model for multivariate time series and assumes that is an affine function of the past observations and some Gaussian noise :
TCNs model log-DLVs unconditionally. Here, the generated process takes the form
3 Training and evaluation time
We split the the historical series into a training and validation set through random sampling. The training set holds 85% of the data and is used to calibrate the parameters of the model. During training we compute the scores for generated paths of length . GAN- and WGAN-GP-trained models are evaluated every 100 generator gradient updates. qMLE-trained models are trained through early stopping and are only evaluated once after the criterion has been reached. The VAR model is also only evaluated once after the parameters are obtained through regression.
4 Results
subsection 8.4 summarises the best scores obtained for each of the generative models for a fixed parameter vector. For all except the score the explicit GAN-trained model achieves to generate the best paths.
The model that performs worst in terms of distributional properties is the VAR(2) - PCA(5) model. There the the fit of density and kurtosis is widely off. Notably, the scores fit least for the VAR(2) and qMLE-trained models, independent whether they were trained on all or the compressed log-DLVs. Overall, we can also conclude from subsection 8.4 that GAN- and WGAN-GP-calibrated models give the best fit. For TCNs we do not report on a compressed version as no good approximation could be obtained.
Although GAN-trained TCNs give a fairly good approximation in terms of distributional and cross-correlation scores the score is far off. This makes the generated paths look very noisy. From this observation we concluded that TCNs have difficulties generating time series with high persistence from a pure i.i.d. noise process.
5 Explicit GAN-trained model
We presented numerical results for a wide range of models and optimization algorithms. Now, we will take a look at the properties of the explicit GAN-trained model since it performed best across most benchmark scores.
As can be seen in Figure 5 the epdfs of the historical and generated log-implied volatilities match closely for most maturities and relative strikes. Arguably, the fit of the bimodal distribution for long-dated ( days) out of the money implied volatilies could be better. Taking a look at the historical and generated kurtosis in Figure 5 we can conclude that for most implied volatilities the approximation is accurate.
Figure 7 and Figure 7 illustrate the generated and historical cross-correlation matrices for log-implied volatilities and implied volatility log-returns. In both cases the historical is approximated accurately confirming the scores in subsection 8.4.
Last viewing the dependence properties Figure A.2 shows that for short lags the approximation of the ACF is very good. However, for longer lags the ACF of the historical decays slower than the generated. Figure A.2 displays the ACF of implied volatility log-returns and shows that the explicit GAN-trained model is able to approximate short dependencies quiet well. However, to model longer lags a larger history is necessary.
Figure A.3 and Figure A.4 display two synthetic paths generated by the explicit GAN-trained model. Notably, the model is able to generate long-lasting periods of low volatility and periods of stress and high volatility phases. When comparing visually the synthetic paths to the historical (see Figure 1) it is difficult to discriminate them from being synthetic.
Conclusion
In this paper, we demonstrated that the generating mechanism of implied volatilities can be closely approximated by employing adversarial training techniques. To measure the proximity of our synthetic paths to the historical we introduced a variety of scores that capture different features of implied volatilities. In section 8 we developed a benchmark and compared the performance of GANs against a wide range of models, training algorithms and explored the effects of compressing DLVs by using PCA. There we concluded that adversarial training outperforms conventional approaches such as the VAR model and qMLE training.
Finally, our work shows for the first time that network-based models can be successfully applied to the context of generative modelling of multivariate financial time series; opening new avenues for future research and applications.
References
Appendix A Explicit GAN-trained model
A.2 Generated paths
Appendix B Discrete local volatilities
DLVs define an arbitrage-free surface that is globally closest to the surface of implied volatilities. We work with DLVs instead of implied volatilities in order to satisfy arbitrage constraints. This eliminates the issue of generating option prices that contain butterfly, calendar or spread arbitrage. Having synthetic options prices that contain arbitrage is highly undesirable. An algorithm trained on these prices could potentially learn to exploit these synthetic arbitrage opportunities which do not occure in reality; yielding an unworldly algorithm that performs well on synthetic but not real data.
In this paper, we consider the option prices of the EURO STOXX 50 from 2011-01-03 to 2019-08-30 and for the sets of relative strikes and maturities
The historical time series of DLVs is obtained by transforming the EURO STOXX 50 implied volatilities and displayed in Figure B.1.
Figure B.3 shows the empirical densities of log-DLVs. Each row represents a specific maturity and columns the moneyness. Similiar to implied volatilities (see Figure 5) long-dated () out of the money (OTM) () DLVs are characterised by bimodal distributions. The large buckets in the epdfs of the short-dated OTM DLVs () represent the floor which is defined at .
The ACFs of log-DLVs and DLV log-returns are displayed in Figure B.5 and Figure B.5. Visually one can detect in Figure B.5 that the ACFs of short-dated in the money DLVs decay faster than long-dated OTM ones.
Appendix C Disclaimer
Opinions and estimates constitute our judgement as of the date of this Material, are for informational purposes only and are subject to change without notice. It is not a research report and is not intended as such. Past performance is not indicative of future results. This Material is not the product of J.P. Morgan’s Research Department and therefore, has not been prepared in accordance with legal requirements to promote the independence of research, including but not limited to, the prohibition on the dealing ahead of the dissemination of investment research. This Material is not intended as research, a recommendation, advice, offer or solicitation for the purchase or sale of any financial product or service, or to be used in any way for evaluating the merits of participating in any transaction. Please consult your own advisors regarding legal, tax, accounting or any other aspects including suitability implications for your particular circumstances. J.P. Morgan disclaims any responsibility or liability whatsoever for the quality, accuracy or completeness of the information herein, and for any reliance on, or use of this material in any way. Important disclosures at: www.jpmorgan.com/disclosures.