Modeling Continuous Stochastic Processes with Dynamic Normalizing Flows
Ruizhi Deng, Bo Chang, Marcus A. Brubaker, Greg Mori, Andreas Lehrmann
Introduction
Expressive models for sequential data form the statistical basis for downstream tasks in a wide range of domains, including computer vision, robotics, and finance. Recent advances in deep generative architectures, especially the concept of reversibility, have led to tremendous progress in this area and created a new perspective on many of the long-standing limitations that are typical in traditional approaches based on structured decompositions (e.g., state-space models).
We argue that the power of a time series model depends on its properties in the following areas: (1 – Resolution) Common time series models are discrete with respect to time. As a result, they make the implicit assumption of a uniformly spaced temporal grid, which precludes their application from asynchronous tasks with a separate arrival process. (2 – Structural assumptions) The expressiveness of a temporal model is determined by the dependencies and shapes of its variables. In particular, the topological structure should be rich enough to capture the dynamics of the underlying process but sparse enough to allow for robust learning and efficient inference. (3 – Generation) A good time series model must be able to generate unbiased samples from the true underlying process in an efficient way. (4 – Inference) Given a trained model, it should support standard inference tasks, such as interpolation, forecasting, and likelihood calculation.
Recently, deep generative modeling has enabled vastly increased flexibility while keeping generation and inference tractable, owing to novel techniques like amortized variational inference , reversible generative models , and networks based on differential equations .
In this work, we approach the modeling of continuous and irregular time series with a reversible generative model for stochastic processes. Our approach builds upon ideas from normalizing flows; however, instead of a static base distribution, we transform a dynamic base process into an observable one. In particular, we introduce the continuous-time flow process (CTFP), a novel type of generative model
that decodes the base continuous Wiener process into a complex observable process using a dynamic instance of normalizing flows. The resulting observable process is thus continuous in time. In addition to the appealing properties of static normalizing flows (e.g., efficient sampling and exact likelihood), this also enables a series of inference tasks that are typically unattainable in time series models with complex dynamics, such as interpolation and extrapolation at arbitrary timestamps. Furthermore, to overcome the simple covariance structure of the Wiener process, we augment the reversible mapping with latent variables and optimize this latent CTFP variant using variational optimization. Our approach is illustrated in Figure 1.
Contributions. In summary, we propose the continuous-time flow process (CTFP), a novel generative model for continuous stochastic processes. It has the following appealing properties: (1) it induces flexible and consistent joint distributions on arbitrary and irregular time grids, with easy-to-compute density and an efficient sampling procedure; (2) the stochastic process generated by CTFP is guaranteed to have continuous sample paths, making it a natural fit for data with continuously-changing dynamics; (3) CTFP can perform interpolation and extrapolation conditioned on given observations. We validate our model and its latent variant on various stochastic processes and real-world datasets and show superior performance to state-of-the-art methods, including the variational recurrent neural network (VRNN) and latent ordinary differential equation (latent ODE) .
Related Work
The following sections discuss the relevant literature on statistical models for sequential data and put it in context with the proposed approach.
Among the most popular traditional time series models are latent variable models following the state-space equations , including the well-known variants with discrete and linear state-space . In the non-linear case, exact inference is typically intractable and we need to resort to approximate techniques . Our CTFP can be viewed as a form of a continuous-time extended Kalman filter where the nonlinear observation process is noiseless and invertible and the temporal dynamics are a Wiener process. The final result, however, is more expressive than a Wiener process but retains some of its appealing properties like closed-form likelihood, interpolation, and extrapolation. Tree-based variants of non-linear Markov models have been proposed in . An augmentation with switching states increases the expressiveness of state-space models but introduces additional challenges for learning and inference . Marginalization over an expansion of the state-space equations in terms of non-linear basis functions extends classical Gaussian processes to Gaussian process dynamical models .
Variational Sequence Models.
Following its success on image data, many works extended the variational autoencoder (VAE) to sequential data . While RNN-based variational sequence models can model distributions over irregular timestamps, those timestamps have to be discrete and thus the models lack the notion of continuity. As a result, they are not suitable for modeling sequential data that have continuous underlying dynamics. Furthermore, it is not straightforward to perform interpolation at arbitrary timestamps using those models.
Latent ODEs use an ODE-RNN as encoder and propagate a latent variable along a time interval using a neural ODE. This formulation ensures that the latent trajectory is continuous in time. However, decoding of the latent variables to observations is done at each time step independently. As a result, there is no guarantee that sample paths are continuous, which causes problems similar to the ones observed in variational sequence models. Neural stochastic differential equations (neural SDEs) replace the deterministic latent trajectory of a latent ODE with a latent stochastic process but do also not generate continuous sample paths.
Recently, Qin et al. proposed the recurrent neural process model. However, members of the neural process family only model the conditional distribution of data given observations and are not generic generative models.
Normalizing Flows in Time Series Modeling.
Multiple recent works apply reversible generative models to sequential data and show promise at capturing complex distributions. Mehrasa et al. and Shchur et al. use normalizing flows to model the distribution of inter-arrival time between events in temporal point processes. Kumar et al. generate video frames using conditional normalizing flows. However, these models only use normalizing flows to model probability distributions in real space. In contrast, our model extends the domain of normalizing flows from distributions in real space to continuous-time stochastic processes.
Background
Our model is built upon the study of stochastic processes and recent advances in normalizing flow research. The following sections introduce the necessary background in these areas.
A stochastic process can be defined as a collection of random variables that are indexed by time. An example of a continuous stochastic process is the Wiener process. The -dimensional Wiener process can be characterized by the following properties: (1) ; (2) for , and is independent of past values of for all . The joint density of can be written as the product of the conditional densities: for .
The conditional distribution of , for , is multivariate Gaussian; its conditional density is
where is a -dimensional identity matrix. This equation also provides a way to sample from . Furthermore, given and , the conditional distribution of for is also Gaussian:
This is known as the Brownian bridge. An important property of the Wiener process is that the sample paths are continuous in time with probability one. This property allows our models to generate continuous sample paths and perform interpolation and extrapolation tasks.
2 Normalizing Flows
where we denote the inverse of by and is the Jacobian matrix of . Sampling from can be done by first drawing a sample from the simple distribution , and then apply the bijection .
Chen et al. , Grathwohl et al. proposed the continuous normalizing flow, which uses the neural ordinary differential equation (neural ODE) to model a flexible bijective mapping. Given sampled from the base distribution , it is mapped to based on the mapping defined by the ODE: . The change in log-density is computed by the instantaneous change of variables formula :
One potential disadvantage of the neural ODE model is that it preserves the topology of the input space, and there are classes of functions that cannot be represented by neural ODEs. Dupont et al. proposed the augmented neural ODE (ANODE) model to address this limitation. Note that the original formulation of ANODE is not a generative model and it does not support the computation of likelihoods or sampling from the target distribution . In this work, we formulate a modified version of ANODE that can be used as a (conditional) generative model.
Model
We define our proposed continuous-time flow process (CTFP) in Section 4.1. In Section 4.2, a generative variant of ANODE is presented as a component to implement CTFP. Since the proposed stochastic process is continuous in time, it enables interpolation and extrapolation at arbitrary time points, as described in Section 4.3. Finally, richer covariance structures are enabled by the latent CTFP model presented in Section 4.4.
Let denote a sequence of irregularly spaced time series data. We assume the time series to be an (incomplete) realization of a continuous stochastic process . In other words, this stochastic process induces a joint distribution of . Our goal is to model such that the log-likelihood of the observations
is maximized. We define the continuous-time flow process (CTFP) such that
The log-likelihood in Equation 5 can be rewritten using the change of variables formula. Let , then
where , , and is defined in Section 3.1. Figure 2(a) shows an example of the likelihood calculation. Sampling from a CTFP is straightforward: given the timestamps , we first sample a realization of the Wiener process , then map them to . Figure 2(b) illustrates this procedure.
The normalizing flow transforms a simple base distribution induced by on an arbitrary time grid into a more complex shape in the observation space. It is worth noting that given a continuous realization of , as long as is implemented as a continuous mapping, the resulting trajectory is also continuous.
2 Generative ANODE
Note that the index represents the independent variable in the initial value problem and should not be confused with , the timestamp of the observation.
Using Equation 4, the log-likelihood can be calculated as follows:
where is obtained by solving the ODE in Equation 8 backwards from to , and the trace of the Jacobian can be estimated by Hutchinson’s trace estimator .
3 Interpolation and Extrapolation with CTFP
Time-indexed normalizing flows and the Brownian bridge allow us to define conditional distributions on arbitrary timestamps. They also permit the CTFP model to perform interpolation and extrapolation given partial observations, which are of great importance in time series modeling.
Interpolation means that we can model the conditional distribution for all and . This can be done by mapping the values , and to , and , respectively. After that, Equation 2 can be applied to obtain the conditional density of . Finally, we have
Extrapolation can be done in a similar fashion using Equation 1. This allows the model to predict continuous trajectories into the future, given past observations. Figure 2(c) shows a visualization of interpolation and extrapolation using CTFP.
4 Latent Continuous-Time Flow Process
The generative model in Equation 6 is augmented to which induces the conditional distribution . Similar to the initial value problem in Equation 8, we define , where
Depending on the sample of the latent variable , the CTFP model has different gradient fields and thus different output distributions.
where is the number of samples from the approximate posterior distribution.
Experiments
In this section, we apply our models on synthetic data generated from common continuous-time stochastic processes and complex real-world datasets. The proposed CTFP and latent CTFP models are compared against two baseline models: latent ODEs and variational RNNs (VRNNs) . The latent ODE model with the ODE-RNN encoder is designed specifically to model time series data with irregular observation times. VRNN is a popular variational filtering model that demonstrated superior performance on structured sequential data.
For VRNNs, we append the time gap between two observations as an additional input to the neural network. Both latent CTFP and latent ODE models use ODE-RNN as the inference network; GRU is used as the RNN cell in latent CTFP, latent ODE, and VRNN models. All three latent variable models have the same latent dimension and GRU hidden state dimension. Please see the supplementary materials for details about our experimental setup and model implementations.
Results. The results are presented in Table 1. We report the exact negative log-likelihood (NLL) per observation for CTFP. For latent ODE, latent CTFP, and VRNN, we report the (upper bound of) NLL estimated by the IWAE bound in Equation 13, using samples of latent variables. We also show the NLL of the test set computed with the ground truth density function.
The results on the test set sampled from the GBM indicate that the CTFP model can recover the true data generation process as the NLL estimated by CTFP is close to the ground truth. In contrast, latent ODE and VRNN models fail to recover the true data distribution. On the M-OU dataset, the latent CTFP models show better performance than the other models. Moreover, latent CTFP outperforms CTFP by 0.016 nats, indicating its ability to leverage the latent variables.
Figure 3 provides a qualitative comparison between CTFP and latent ODE trained on the GBM data, both on the generation task (upper panels) and the interpolation task (lower panels). The results in the upper panels show that CTFP can generate continuous sample paths and accurately estimate the marginal mean and quantiles. In contrast, the sample paths generated by latent ODE are more volatile and discontinuous due to its lack of continuity. For the interpolation task, the results of CTFP are consistent with the ground truth in terms of both point estimation and uncertainty estimation. For latent ODE on the interpolation task, Figure 3(b) shows that the latent variables from the variational posterior shift the density to the region where the observations lie. However, although latent ODE is capable of performing interpolation, there is no guarantee that the (reconstructed) sample paths pass through the observed points (triangular marks in Figure 3(b)), as discussed in Section 2. In addition to these difficulties with the interpolation task, the qualitative comparison between samples further highlights the importance of our models’ continuity when generating samples of continuous dynamics.
2 Real-World Datasets
We also evaluate our models on real-world datasets with continuous and complex dynamics. The following three datasets are considered: Mujoco-Hopper consists of 10,000 sequences that are simulated by a “Hopper” model from the DeepMind Control Suite in a MuJoCo environment . PTB Diagnostic Database (PTBDB) consists of excerpts of ambulatory electrocardiography (ECG) recordings. Each sequence is one-dimensional and the sampling frequency of the recordings is 125 Hz. Beijing Air-Quality Dataset (BAQD) is a dataset consisting of multi-year recordings of weather and air quality data across different locations in Beijing. The variables in consideration are temperature, pressure, and wind speed, and the values are recorded once per hour. We segment the data into sequences, each covering the recordings of a whole week. Please refer to the supplementary materials for additional details about data preprocessing.
Similar to our synthetic data experiment settings, we compare the CTFP and latent CTFP models against latent ODE and VRNN. It is worth noting that the latent ODE model in the original work uses a fixed output variance and is evaluated using mean squared error (MSE); we adapt the model to our tasks with a predicted output variance (see supplementary materials). We further study the effect of using RealNVP as the invertible mapping . This experiment can be regarded as an ablation study and results are presented in the supplementary materials as well.
Results. The results are shown in Table 2. We report the exact negative log-likelihood (NLL) per observation for CTFP and the (upper bound of) NLL estimated by the IWAE bound, using samples of latent variables, for latent ODE, latent CTFP, and VRNN. For each setting, the mean and standard deviation of five evaluation runs are reported. The latent CTFP models trained and evaluated with Hutchinson’s trace estimator on multi-dimensional data could lead to a biased IWAE estimation and unstable training due to the variance of Hutchinson’s trace estimator and the log-sum-exp operation. Therefore, we also report results trained and estimated with the exact Jacobian trace, denoted by latent CTFP-ET. The evaluation results show that the latent CTFP model outperforms VRNN and latent ODE models on real-world datasets, indicating that CTFP is better at modeling irregular time series data with continuous dynamics. Table 2 also suggests that the latent CTFP model consistently outperforms the CTFP model, demonstrating that with the latent variables, the latent CTFP model is more expressive and able to capture the data distribution better. We defer additional experimental results with different observation processes but similar conclusions to the supplementary materials.
Conclusion
In summary, we propose the continuous-time flow process (CTFP), a reversible generative model for stochastic processes, and its latent variant. It maps a simple continuous-time stochastic process, i.e., the Wiener process, into a more complicated process in the observable space. As a result, many desirable mathematical properties of the Wiener process are retained, including the efficient sampling of continuous paths, likelihood evaluation on arbitrary timestamps, and inter-/extrapolation given observed data. Our experimental results demonstrate the superior performance of the proposed models on various datasets.
Broader Impact Statement
Time series models could be applied to a wide range of applications, including natural language processing, recommendation systems, traffic prediction, medical data analysis, forecasting, and others. Our research improves over the existing models on a particular type of data: irregular time series data.
There are opportunities for applications using the proposed models for beneficial purposes, such as weather forecasting, pedestrian behavior prediction for self-driving cars, and missing healthcare data interpolation or prediction. We encourage practitioners to understand the impacts of using CTFP in particular real-world scenarios.
One potential risk is that the capability of interpolation and extrapolation can be used in malicious ways. An adversary might be able to use the proposed model to infer private information given partial observations, which leads to privacy concerns. We would encourage further research to address this risk using tools like differential privacy.
Funding Transparency Statement
This work was conducted at Borealis AI and partly supported by Mitacs through the Mitacs Accelerate program.
References
Appendix A Finite-Dimensional Distribution of CTFP
Let be defined as Equations 8 and 9 in Section 4.2. The mapping from to defined by is measurable and therefore induces a pushforward measure .
Given a finite subset , the finite-dimensional distribution of is the same as the distribution of , where is a -dimensional random variable with finite-dimensional distribution of .
Appendix B Experiment Setup and Model Architecture Details
We describe the details on synthetic dataset generation, real-world dataset pre-processing, model architecture as well as training and evaluation settings in this section.
B.2 Real-World Dataset Details
As mentioned in Section 5.2 of the paper, we compare our models against the baselines on three datasets: Mujoco-Hopper, Beijing Air-Quality dataset (BAQD), and PTB Diagnostic Database(PTBDB). The three datasets can be downloaded using the following links:
http://www.cs.toronto.edu/~rtqichen/datasets/HopperPhysics/training.pt
https://www.kaggle.com/shayanfazeli/heartbeat/download
https://archive.ics.uci.edu/ml/datasets/Beijing+Multi-Site+Air-Quality+Data
We pad all sequences into the same length for each dataset. The sequence length of the Mujoco-Hopper dataset is 200 and the sequence length of BAQD is 168. The maximum sequence length in the PTBDB dataset is 650. We rescale the indices of sequences to real numbers in the interval of and take the rescaled values as observation timestamps for all datasets. To make the sequences asynchronous or irregularly-sampled, we sample observation timestamps from a homogeneous Poisson process with an intensity of 2 that is independent of the data. For each sampled timestamp, the value of the closest observation is taken as its corresponding value. The timestamps of all sampled sequences are shifted by a value of 0.2 since deterministically for the Wiener process and there’s no variance for the CTFP model’s prediction at .
B.3 Model Architecture Details
To ensure a fair comparison, we use the same values for hyper-parameters including the latent variable and hidden state dimensions across all models. Likewise, we keep the underlying architectures as similar as possible and use the same experimental protocol across all models.
For CTFP and Latent CTFP, we use a one-block augmented neural ODE module that maps the base process to the observation process. For the augmented neural ODE model, we use an MLP model consisting of 4 hidden layers of size 32–64–64–32 for the model in Equation 8 and Equation 12. In practice, the implementation of in the two equations is optional and its representation power can be fully incorporated into . This architecture is used for both synthetic and real-world datasets. For the latent CTFP and latent ODE models appearing in Section 5, we use the ODE-RNN model as the recognition network. For synthetic datasets, the ODE-RNN model consists of a one-layer GRU cell with a hidden dimension of 20 (the rec-dims parameter in its original implementation) and a one-block neural ODE module that has a single hidden layer of size 100, and it outputs a 10-dimensional latent variable. The same architecture is used by both latent ODE and latent CTFP models. For real-world datasets, the ODE-RNN architecture uses a hidden state of dimension 20 in the GRU cell and an MLP with a 128-dimensional hidden layer in the neural ODE module. The ODE-RNN model produces a 64-dimensional latent variable. For the generation network of the latent ODE (V2) model, we use an ODE function with one hidden layer of size 100 for synthetic datasets and 128 for real-world datasets. The decoder network has 4 hidden layers of size 32–64–64–32; it maps a latent trajectory to outputs of Gaussian distributions at different time steps.
The VRNN model is implemented using a GRU network. The hidden state of the VRNN models is 20-dimensional for synthetic and real-world datasets. The dimension of the latent variable is 64 for real-word datasets and 10 for synthetic datasets. We use an MLP of 4 hidden layers of size 32–64–64–32 for the decoder network, an MLP with one hidden layer that has the same dimension as the hidden state for the prior proposal network, and an MLP with two hidden layers for the posterior proposal network. For synthetic data sampled from Geometric Brownian Motion, we apply an exponential function to the samples of all models. Therefore the distribution precited by latent ODE and VRNN at each timestamp is a log-normal distribution.
B.4 Training and Evaluation Settings
For synthetic data, we train all models using the IWAE bound with 3 samples and a flat learning rate of for all models. We also consider models trained with or without the aggressive training scheme proposed by He et al. for latent ODE and latent CTFP. We choose the best-performing model among the ones trained with or without the aggressive scheme based IWAE bound, estimated with 25 samples on the validation set for evaluation. The batch size is 100 for CTFP models and 25 for all the other models. For experiments on real-world datasets, we did a hyper-parameter search on learning rates over two values of and , and whether using the aggressive training schemes for latent CTFP and latent ODE models. We report the evaluation results of the best-performing model based on IWAE bound estimated with 125 samples.
Appendix C Ablation Study Results
We provide additional experiment results on real-world datasets using different intensity value s of 1 and 5 to sample observation processes in Table 1 below.
C.2 I.I.D. Gaussian as Base Process
In this experiment, we replace the base Wiener process with I.I.D Gaussian random variables and keep the other components of the models unchanged. This model and its latent variant are named CTFP-IID-Gaussian and latent CTFP-IID-Gaussian. As a result, the trajectories sampled from CTFP-IID-Gaussian are not continuous and we use this experiment to study the continuous property of models and its impact on modeling irregular time series data with continuous dynamics. The results are presented in Table 4 and Table 5.
The results show that CTFP consistently outperforms CTFP-IID-Gaussian, and latent CTFP outperforms latent CTFP-IID-Gaussian. The results corroborate our hypothesis that the superior performance of CTFP models can be partially attributed to the continuous property of the model. Moreover, latent CTFP-IID-Gaussian shows similar but slightly better performance than latent ODE models. The results comply with our hypothesis as the models are very similar and both models have no notion of continuity in the decoder. We believe the performance gain of latent CTFP-IID-Gaussian comes from the use of (dynamic) normalizing flow which is more flexible than Gaussian distributions used by latent ODE.
C.3 CTFP-RealNVP
In this experiment, we replace the continuous normalizing flow in CTFP model with another popular choice of normalizing flow model, RealNVP . This is variant of CTFP is named CTFP-RealNVP and its latent version is called latent CTFP-RealNVP. Note that the trajectories sampled from CTFP-RealNVP model are still continuous. We evaluate CTFP-RealNVP and latent CTFP-RealNVP models on datasets with high dimensional data, Mujoco-Hopper, and BAQD. The results are shown in Table 6.
The table indicates that CTFP-RealNVP outperforms CTFP. However, when incorporating the latent variable, the latent CTFP-RealNVP performs significantly worse than latent CTFP. The worse performance might be because RealNVP cannot make full use of the information in the latent variable due to its structural constraints as we discussed in Section 4.2.
Appendix D Additional Details for Latent ODE Models on Mujoco-Hooper Data
The original latent ODE paper focuses on point estimation and uses the mean squared error as the performance metric . When applied to our problem setting and evaluated using the log-likelihood, the model performs unsatisfactorily. In Table 7, the first row shows the negative log-likelihood on the Mujoco-Hopper dataset. The inferior NLL of the original latent ODE is potentially caused by the use a fixed output variance of , which magnifies even a small reconstruction error.
To mitigate this issue, we propose two modified versions of the latent ODE model. For the first version (V1), given a pretrained (original) latent ODE model, we do a logarithmic scale search for the output variance and find the value that gives the best performance on the validation set. The second version (V2) uses an MLP to predict the output mean and variance. Both modified versions have much better performance than the original model, as shown in Table 7, rows 2–3. It also shows that the second version of the latent ODE model (V2) outperforms the first one (V1) on the Mujoco-Hopper dataset. Therefore, we use the second version (V2) for all the experiments in the main text.
Appendix E Qualitative Sample for VRNN Model
We sample trajectories from the VRNN model trained on Geometric Brownian Motion (GBM) by running the model on a dense time grid and show the trajectories in Figure 4. We compare the trajectories sampled from the model with trajectories sampled from GBM. As we can see, the sampled trajectories from VRNN are not continuous in time.
We also use VRNN to estimate the marginal density of for each and show the results in Figure 4. It is not straightforward to use VRNN model for marginal density estimation. For each timestamp , we get the marginal density of by running VRNN on a time grid with two timestamps, 0 and : at the first step, the input to VRNN model is and we can get prior distributions of the latent variable . Note that a sampled trajectory from GBM is always 1 when . Conditioned on the sampled latent codes and , VRNN proposes at the second step. We average the conditional density over 125 samples of and to estimate the marginal density.
The marginal density estimated using a time grid with two timestamps is not consistent with the trajectories sampled on a different dense time grid. The results indicate that the choice of time grid has a great impact on the distribution modeled by VRNN and the distributions modeled by VRNN on different time grids can be inconsistent. In contrast, our proposed CTFP models do not have such problems.