A Generalised Signature Method for Multivariate Time Series Feature Extraction

James Morrill, Adeline Fermanian, Patrick Kidger, Terry Lyons

Introduction

One approach is to construct models that directly accept some of these issues; for example recurrent neural networks handle correlated inputs with varying lengths. A second option is to use feature extraction techniques, which normalise the data so that other techniques may then be applied. Methods such as the shapelet transform (Ye & Keogh 2009; Grabocka et al. 2014; Kidger et al. 2020), Gaussian process adapters (Li & Marlin 2016; Futoma et al. 2017; Moor et al. 2020), and in particular the signature method (Levin et al. 2013), all fit into this category.

The approach taken by the signature method, coming from rough path theory (Lyons et al. 2007; Friz & Victoir 2010), is to interpret a multivariate time series as a discretisation of an underlying continuous path. The signature transform, also known as the path signature or signature, can then be applied, which produces a vector of real-valued features that are known to characterise the path.

Benefits of the signature method include: a high degree of flexibility, making it possible to customise the method to specific datasets; strong theoretical guarantees; an interpretable feature set; ease of handling irregularly sampled and/or partially observed data; and it being well-defined for some highly irregular processes such as ARMA, Gaussian processes or even Brownian motion. Also, signature features do not need to learned, which can make them particularly effective on (but not limited to) low sample datasets.

The flexibility of the signature method has made it possible to be tailored to specific applications and achieve state-of-the-art performance in wide range of problem domains, such as handwriting recognition (Wilson-Nunn et al. 2018; Yang et al. 2016b), action recognition (Yang et al. 2016a; Yang et al. 2017), and medical time series prediction tasks (Morrill et al. 2019; Morrill et al. 2020b). However, this flexibility comes at the cost of additional complexity in the model search space.

To the best of our knowledge, no comprehensive studies exist that collate and combine the most common method variations found in the literature and assemble them under a common mathematical framework. Additionally, no baseline signature model has ever been tested against other time series classification baselines. Our goal will be to address both of these issues, alongside the development of an open source implementation, so as to make the methods more accessible to a wider audience.

We introduce a generalised signature method that contains the many existing variations as special cases. In doing so we are able to understand their conceptual groupings into what we term augmentations, windows, transforms and rescalings. This involves a comprehensive review of the existing variations across the literature. By understanding their commonality, we are then able to combine different variations, and propose new options that fit into this framework.

We go on to examine which choices within this framework are most important to success by performing an extensive empirical study across 26 datasets. To the best of our knowledge this is the first study of this type.

In doing so, we are then able to produce a canonical signature pipeline. This represents a domain agnostic starting point that may then be adapted for the task at hand. We show that the performance of this canonical pipeline is comparable to current state-of-the-art classifiers for multivariate time series classification, including deep recurrent and convolutional neural networks. This has led to the implementation of this generalised approach in the open source [redacted] package.

Context

We begin with a few mathematical definitions necessary throughout the article.

We consider a dataset to be a collection of such samples. Note that the time stamps t\mathbf{t} for each sample may be different, and the sample lengths nn can vary. That is, we accept varying length and irregular sampling without modification. We are now in a position to define the signature of a time series.

where for any (i1,…,ik)∈{1,…,d}k(i_{1},\dots,i_{k})\in\{1,\dots,d\}^{k},

While this definition may seem somewhat technical, there are several intuitions that can be made with regard to the signature features. We present a geometric interpretation of the first two levels of the signature and log-signature in Figure 1. The depth-1 terms, S(x)(i)S(\mathbf{x})^{(i)}, equate to the displacement of the path over the interval in the iith coordinate, denoted by ΔXi\Delta X^{i} in Figure 1. The depth-2 terms, S(x)(i,j)S(\mathbf{x})^{(i,j)}, have interpretations in areas generated over the interval.

From a statistical point of view, the signature can be thought of as the equivalent of a moment-generating function for time series. Let ZZ be a random variable, then the moment-generating function of ZZ is the function

and, if well-defined, it characterizes the distribution of ZZ. Assume that XX is a random time series (that is a stochastic process), its signature now has the same properties as a moment-generating function: the powers of ZZ are replaced by integrals of products of coordinates and Chevyrev & Lyons 2016 show that the expected signature characterizes the law of XX.

Moreover, we have the following two properties that make the signature a good feature set in a machine-learning context—precise statements may be found in Bonnier et al. 2019.

Universal nonlinearity

Linear functionals on the signature are dense in the set of functions on x\mathbf{x}. Suppose we wish to learn the function ff that maps data x\mathbf{x} to labels yy, the universal nonlinearity property states that, under some assumptions, for any ε>0\varepsilon>0, there exists a linear function LL such that

Note that contrary to Fourier or wavelet basis, signatures provide a natural basis for functions of the time series rather than for the time series itself—equation 1 concerns f(x)f(\mathbf{x}) and not x\mathbf{x}. In the context of time series classification this shift of perspective is particularly well-suited since the object of interest is not the time series itself but its link to a label.

Logsignature transform

The signature contains some redundant information: for example we can see in the left panel of Figure 1 that the sum of the blue and orange areas is equal to the product of displacements ΔX1ΔX2\Delta X^{1}\Delta X^{2}:

The logsignature transform is essentially the signature with these redundancies removed. For example, the logsignature encodes the blue and orange areas from the left panel with the orange signed area in the right panel. However, the logsignature does not have a universal nonlinearity property such as equation 1. We refer the reader to Morrill et al. 2020a or Liao et al. 2019 for a precise definition of the logsignature.

A pedagogical introduction to the background theory of signatures is Lyons et al. 2007, whilst a comprehensive textbook is Friz & Victoir 2010. For introductions to the signature method, we recommend Bonnier et al. 2019 and Chevyrev & Kormilitzin 2016.

2 Related work

The signature transform has been used in a wide range of applications in machine learning predictive tasks. For example, as mentioned in the introduction, the signature has been used as a feature extraction layer in classifiers for both Arabic (Wilson-Nunn et al. 2018) and Chinese (Yang et al. 2016b) handwriting recognition. Similarly, it was successfully used in human action recognition by Li et al. 2017; Yang et al. 2017; Liao et al. 2019 and in the medical domain as part of the top performing model at the Physionet 2019 challenge for prediction of sepsis (Reyna et al. 2019; Morrill et al. 2019; Morrill et al. 2020b). Other applications involve finance (Lyons et al. 2014; Perez Arribas 2018), mental health (Kormilitzin et al. 2017; Arribas et al. 2018), and emotion recognition (Wang et al. 2019; Wang et al. 2020).

In almost all these applications, the method has been utilised in different ways. Many authors consider transformations of the input time series before application of the signature (Levin et al. 2013; Flint et al. 2016; Lyons & Oberhauser 2017; Yang et al. 2017; Liao et al. 2019; Kidger & Lyons 2020; Wu et al. 2020a). People have also explored different windows over which the signature transform should be taken, so as to extract information over different scales (Yang et al. 2017; Bonnier et al. 2019). Additionally a choice must be made between the signature and logsignature transforms, as must choices for the scaling of the terms in the signature (Chevyrev & Kormilitzin 2016; Lai et al. 2017).

The differences between some of these choice have been shown by Fermanian 2019 to significantly impact the performance of the methodology. However this study used a small collection of datasets and considered only some of the most common variations that exist in the literature. There is therefore a need for a comprehensive study and unification of all these different choices.

The generalized signature method

In this section we collate the modifications to the signature transform that have been proposed in signature literature to date. We will show that each can be categorised into one of the following groups:

Augmentations These describe the transformation of a time series into one or more new series, in order to return different information in the signature features and deal with dimensionality issues.

Windows Splitting the time series over different subsequences (or windows), so that signatures may be applied locally.

Transform The choice between the signature or the logsignature transform.

Rescaling Ways of normalising the terms in the signature.

We then go on to show that these groupings can themselves be synergised into a single mathematical framework that we term the generalised signature method. For clarity, we will begin by discussing each of these individually, and then afterwards show how they may be combined.

Remove the signature invariance to translation and/or reparametrization.

Lower the dimension dd of the time series, so that higher orders of the signature are reachable—recall that the depth-NN signature is of size O(dN)\mathcal{O}(d^{N}).

Preprocess the time series prior to the signature map so that information is more easily extracted.

There are many pre-signature operations which have been proposed in the literature, and which we categorise as augmentations. We refer the reader to Appendix A for full details of the many such operations proposed in the literature, but will focus on several important examples here.

This transformation, which basically consists in adding the timestamps as an extra coordinate, has two key properties: it guarantees the uniqueness of the signature (Hambly & Lyons 2010) and it adds information about the parametrization of the time series.

which simply adds a zero at the beginning of the time series—note that this zero could also be put at the end. This transformation makes the signature sensitive to translations of the time series. The invisibility-reset transformation (Yang et al. 2017; Wu et al. 2020b) also adds translation sensitivity, but does so by increasing the dimension.

In the second group of augmentations for dimensionality reduction, we consider random projections (Lyons & Oberhauser 2017), which consist in applying multiple random linear maps to the time series, or coordinate projections, which project along (multiple subsets of) the coordinate axes.

In the third group, the lead-lag augmentation (Chevyrev & Kormilitzin 2016; Flint et al. 2016; Yang et al. 2017) captures the quadratic variation by transforming the time series to

2 Windows

The second step is to choose a windowing operation. Much like the window functions used with a short time Fourier transform, this localises the signature computation to extract information over particular time intervals.

The expanding window produces time series of increasing length, and is analogous to the history processes of stochastic analysis whereas the sliding window produces time series of fixed length but shifted in time.

3 The signature and logsignature transforms

Central to the signature methodology is of course the signature transform itself. Two choices must be made; whether to use the signature or logsignature transform, and what depth to calculate the transform to—that is, what depth NN in Definition 2 to use. Choosing a logsignature lowers the feature vector dimension at the cost of loosing linear approximation properties. There is no consensus on which one should be favored for a machine learning task.

4 Rescaling

The depth-kk term in the signature is of size O(\nicefrac1k!)\mathcal{O}(\nicefrac{{1}}{{k!}}). Typically, rescaling these terms to O(1)\mathcal{O}(1) will aid in subsequent learning procedures. To this end, we can apply pre-signature scaling whereby we scale the path before signature computation, or post-signature where we scale the signature terms themselves. Specifics on how this is done in practice are given in Appendix B.

5 Putting the pieces together

over all i∈{1,…,p}i\in\{1,\dots,p\}, j∈{1,…,w}j\in\{1,\dots,w\}. We refer to the procedure of computing x↦(zi,j)\mathbf{x}\mapsto(\mathbf{z}_{i,j}) as the generalised signature method.

This final procedure is a little involved, but is simply a combination of different elementary operations used to impact the final feature set. The overall procedure now offers a degree of flexibility and generality which has, to our knowledge, never been achieved for signature methods.

The collection of features (zi,j)(\mathbf{z}_{i,j}) may then be fed into any later machine learning algorithm, which will depend on the application. In general, the zi,j\mathbf{z}_{i,j} will be stacked together and considered as a vector. However, if one wants to use a sequential algorithm such as a recurrent network, it is possible to turn the features zi,j\mathbf{z}_{i,j} into a sequence by choosing a sliding or expanding window. Indeed, these windows induce an ordering in the features: the terms zi,1\mathbf{z}_{i,1} will correspond to the first values of x\mathbf{x}, the terms zi,2\mathbf{z}_{i,2} to the following values, and so on.

Empirical study

We perform a first-of-its-kind empirical study across 26 datasets to determine the most important aspects of this framework.

The datasets used are the Human Activities and Postural Transitions dataset provided by Reyes-Ortiz et al. 2016, the Speech Commands dataset provided by Warden 2018, and 24 datasets from the UEA time series classification archive, provided by Bagnall et al. 2018. A few datasets from the UEA archive were excluded due to their high number of channels resulting in too large a computational burden.

Baseline

We begin by defining a single baseline procedure, representing a simple and straightforward collection of choices for the generalised signature method. This baseline is to take the augmentation ϕ\phi as appending time as defined by equation 2, WW as the global window defined by equation 5, have the transform be a signature transform of depth 33, and to use pre-signature scaling of the path. This means that the input features are the collection

Individual variations

With respect to this baseline procedure, we then consider, in turn, the groups described in Section 3. These were augmentations, windows, transform, and rescaling. For each group we modify the baseline by implementing each option in the group one-by-one. Each such variation defines a particular form of the generalised signature method as in equation 6. Example variations are to switch to using a logsignature transform of depth 5, or to use a sliding window instead of a global window. We discuss the precise variations below.

Models

On top of every variation, we then consider four different models: logistic regression, random forest, Gated Recurrent Unit (GRU) (Cho et al. 2014), and a residual Convolutional Neural Network (CNN) (He et al. 2015). We test nearly every combination of dataset, variation of the generalised signature method, and model. Different datasets and variations produce different numbers of features zi,j\mathbf{z}_{i,j}, so to reduce the computational burden we omit those cases for which the number of features is greater than 10510^{5}. Of the 9984 total combinations of dataset, variation, and model, this leaves out 1415 combinations. See Appendix C.2 for a break down of the omitted combinations by different cases.

Analysis

We define the performance of a variation on a dataset as the best performance across the four models considered, to reflect the fact that different models are better suited for different problems. We then follow the methodology of Demšar 2006; Benavoli et al. 2016; Ruiz et al. 2020 to compare the variations across the multiple datasets. We first perform a Friedman test to reject the null hypothesis that all methods are equivalent. If it is rejected, we perform pairwise Wilcoxon signed-rank tests to form cliques of not-significant methods, and use critical difference plots to visualize the performance of each signature method.

A critical difference plot shows the different variations ordered by their average rank: for example, in Figure 2, the best variation is “Time + Basepoint” with an average rank of 2.5. Then, a thick line indicates that the Wilcoxon test between variations inside the clique is not rejected at significance threshold of 5%, subject to Bonferroni’s multiple testing correction. In Figure 2 there are two groups of significantly different variations: one with “Basepoint” and “None” and one with all other variations.

We refer the reader to Appendix C for further details on the methodology, such as precise architectural choices, learning rates, and so on.

2 Results

Due to the large number of variations and datasets considered, we present only the critical difference plots in the main paper. See Appendix D for all the tables of the underlying numerical values.

We split the augmentations into two categories. The first category consists of those augmentations which remove the signature’s invariance to translation (basepoint augmentation, invisibility-reset augmentation) or reparameterisation (time augmentation). We see in Figure 2 that augmenting with time, and either basepoint or invisibility-reset, are both typically important. This is expected; in general a problem need not be invariant to either translation or reparametersiation.

The second category consists of those augmentations which either seek to reduce dimensionality or introduce additional information. We see in Figure 3 that most augmentations actually do not help matters, except for lead-lag which usually represents a good choice. We posit that the best augmentation is likely to be dataset dependent, so we break this down by dataset characteristics in Table 6.

Here we indeed see that there is generally a better choice than doing nothing at all, but that this better choice is dependent on some characteristic of the dataset. For example, learnt projections and multi-headed stream preserving transformations do substantially better on EEG datasets, while lead-lag is better for human action and motion recognition.

Windows

We consider the possibility of global, sliding, expanding, and dyadic windows. The results are shown in Figure 4. We see that the dyadic and expanding windows are significantly better than sliding and global windows. The poor performance of sliding windows is a little surprising, but tallies with the observations of Fermanian 2019. This is an important finding, as global and sliding windows tend to be commonly used with signature methods.

Signature versus logsignature transforms

We consider the signature and logsignature transforms with depths ranging from 1 to 6. As higher depths always produce more information, we define the performance of the (log)signature transform as the best performance across all depths. With this metric, the signature transform is significantly better than the logsignature transform, with a p-value of 0.01 for the Wilcoxon signed-rank test.

The key results

To conclude, these results show that invariance-removing transformations such as time and basepoint augmentations should a priori be used, that the lead-lag performs well but not significantly better than no additional augmentation, and that the hierarchical dyadic window performs significantly better than the sliding and global ones. The poor performance of deep learning approaches for augmentation is also notable and an additional motivation for this work: although slightly technical, the augmentations tailored to the signature transform are a significant addition in a machine learning pipeline and cannot be easily replaced by neural networks.

3 Further results

See Appendix D for further results, in particular on the running times, the different types of rescaling, augmentations broken down by dataset characteristics, an additional study on signature depth, and the precise numerical results for each individual test considered here.

The canonical signature pipeline

In this section we define the canonical signature pipeline. Using the results from Section 4 we evaluate the top performing options over all the datasets so as to provide a domain-agnostic starting point for any dataset, from which other variations can be easily explored. We show that this pipeline shows competitive performance against traditional benchmarks and even against deep neural networks.

In a nutshell, the pipeline consists in applying the basepoint and time augmentations, a hierarchical dyadic window and a signature transform, which can be written as a particular case of equation 6 as follows. Let WW be a hierarchical dyadic window of depth qq, ϕt\phi_{\mathbf{t}} and ϕb\phi^{b} be the time and basepoint augmentations, then the canonical signature pipeline may be written as

We give a graphical depiction of this in Figure 5. Signature and window depths (NN, qq) must be optimised for the problem (typically via cross-validation). We note that this canonical method may be adapted to the problem at hand in two ways: if the problem is known to be parametrization invariant, as is the case for example for characters recognition, then the time augmentation should not be applied. Moreover, if the problem is translation-invariant, then the basepoint augmentation is not applied. We emphasise that this pipeline does not represent a best option for every application, but is meant to represent a compromise between broad applicability, ease of implementation, computational cost, and good performance.

2 Performance

We validate the performance of the pipeline against the 26 datasets in the multivariate UEA archive This is not to be confused with the UCR archive which is a collection of 128 univariate datasets.. To our knowledge, the most recent benchmarks for the UEA archive are the results from Ruiz et al. 2020. We compare their results to the canonical signature pipeline with a random forest classifier—see Appendix C.3 for more details.

The benchmarks include variants on classical Dynamic Time Wrapping (DTWI, DTWD and DTWA); an ensemble of univariate classifiers, HIVE COTE (Bagnall et al. 2020), known to be highly perfromant in the univariate case; a random shapelet forest (Karlsson et al. 2016), denoted gRSF, and a bag of words based algorithm, MUSE (Schäfer & Leser 2017); two deep learning methods, TapNet (Wang et al. 2017) and MLCN (Karim et al. 2019). The MLCN architecture combines long short term memory layers (LSTM) and convolutional layers while TapNet combines 3 blocks: random projections on the different dimensions, convolutional layers and a final attention block to compare candidate time series representations.

Figure 6 shows the critical difference plot of this comparison. The signature pipeline is in the first clique, that is the group of classifiers that achieve the best accuracy while not being significantly different from one another. The two algorithms with a better rank than the signature pipeline are MUSE and HIVE-COTE. It is worth noting that MUSE is very memory intensive—Ruiz et al. 2020 report that it could not finish on 5 of the 26 UEA datasets on a computer with 500GB of memory—whilst HIVE-COTE is an ensemble of several sub-classifiers, and thus has very high training and inference costs. On the other hand, all experiments for the canonical signature pipeline were completed with no memory errors on a computer with less memory, and are significantly faster to run than HIVE-COTE—see Appendix D.

The canonical signature pipeline is meant to be a sensible starting point from which the user can propose additional variations following the structure defined in equation 6, but as a standalone classifier this pipeline performs comparably to state-of-the-art classifiers, on the UEA data, whilst being less computationally demanding.

Conclusion

We introduce a generalised signature method as a framework to capture recently proposed variations on the signature method. We go on to perform a first-of-its-kind extensive empirical investigation as to which elements of this framework are most important for performance in a domain-independent setting. In particular, we highlight the performance of hierarchical dyadic windows and signature-tailored augmentations such as lead-lag, time and basepoint. As a result, we are able to present a canonical signature pipeline that represents a best-practices domain-agnostic starting point, which shows competitive performance against state-of-the-art classifiers for multivariate time series classification.

References

Appendix A Augmentations

We give below the precise definition of the different augmentations considered in the study, which are summarized in Table 2. These augmentations were not typically introduced using such language, so this serves as a reference for how the existing literature may be interpreted through the generalised signature method.

We recall the definition of the time augmentation:

It ensures uniqueness of the signature transformation and removes the parametrization invariance (Levin et al. 2013).

Invisibility-reset augmentation

First introduced by Yang et al. 2017, the invisibility-reset augmentation consists in adding a coordinate to the sequence x\mathbf{x} that is constant equal to 1 but drops to 0 at the last time step, i.e.,

This augmentation adds information on the initial position of the path, which is otherwise not included in the signature as it is a translation-invariant map.

Basepoint augmentation

Introduced by Kidger & Lyons 2020, the basepoint augmentation has the same goal as the invisibility-reset augmentation: removing the translation-invariant property of the signature. It simply adds the point 0 at the beginning of the sequence:

The main difference compared to the invisibility-reset augmentation is that the signature of x\mathbf{x} is contained in the signature of the invisibility-reset augmented path, whereas it is not in the signature of the basepoint augmented path. The price paid is that the invisibility-reset augmentation introduces redundancy into the signature, and is more computationally expensive due to the additional channel. (Recall that the signature method scales as O(dN)\mathcal{O}(d^{N}), where dd is the input channels and NN is the depth of the (log)signature.)

Lead-lag augmentation

The lead-lag augmentation, introduced by Chevyrev & Kormilitzin 2016 and Flint et al. 2016 has been used in several applications (see for example Lyons et al. 2014; Kormilitzin et al. 2016; Yang et al. 2017). It adds lagged copies of the path as new coordinates. This then explicitly captures the quadratic variation of the underlying process (Flint et al. 2016). As many different lags as desired may be added. If there is a single lag of a single timestep, then this corresponds to

Coordinate projections

Then we define the singleton coordinate projection as

whilst considering all possible pairs of coordinates yields the augmentation

and all possible triples yields the augmentation

The decision to always include a time dimension is a somewhat arbitrary one, and it may alternatively be excluded if desired. (This is done so as to make sense of singleton coordinate projections; otherwise the result is a collection of univariate time series, for which the signature extracts only the increment due to the tree-like equivalence property.)

Random projections

Learnt projections

Rather than taking random projections, Liao et al. 2019 learn it from the data. This takes exactly the same form as the random projections, except that the AiA_{i} are learnt.

Stream-preserving neural network

Bonnier et al. 2019 introduce arbitrary learnt sequence-to-sequences maps prior to the signature transform, and refer to such maps, when parameterised as neural networks, as stream-preserving neural networks. For example these may be standard convolutional or recurrent architectures. In general this may be any learnt transformation

Multi-headed stream-preserving neural network

A straightforward extension of stream-preserving neural networks is to use multiple such networks, so as to avoid a potential bottleneck through the single signature map that it is eventually used in. Letting ϕ1,…,ϕp\phi^{1},\ldots,\phi^{p} be pp different stream-preserving neural networks, then this gives an augmentation

Appendix B Rescaling

The signature transform can be written as a sequence of tensors, indexed by k∈{1,…,N}k\in\{1,\ldots,N\}. The kk-th term is of size O(\nicefrac1k!)\mathcal{O}(\nicefrac{{1}}{{k!}}), as it is computed by an integral over a kk-dimensional simplex. It is typical that rescaling these terms to be O(1)\mathcal{O}(1) will aid subsequent learning procedures.

One option is to simply multiply the kk-th term by k!k!, which we call post-signature scaling.

Appendix C Implementation details

All the code for this project is available at [redacted for anonymity].

Libraries

The machine learning framework used was PyTorch (Paszke et al. 2017) version 1.3.1. Signatures and logsignatures were computed using the Signatory library (Kidger & Lyons 2020) version 1.1.6. Scikit-learn (Pedregosa et al. 2011) version 0.22.1 was used for the logistic regression and random forest models. The experiments were tracked using the Sacred framework (Greff et al. 2017) version 0.8.1.

Normalisation

Every dataset was normalised so that each channel has mean zero and unit variance.

Architectures

Two different GRU models were used on every dataset; a ‘small’ one with 32 hidden channels and 2 layers, and a ‘large’ one with 256 hidden channels and 3 layers.

Likewise, two different Residual CNN models were considered. The ‘small’ one used 6 blocks, each composed of batch normalisation, ReLU activation, convolution with 32 filters and kernel size 4, batch normalisation, ReLU activation, and a final convolution with 32 filters and kernel size 4, so that there are also 32 channels along the ‘residual path’. A final two-hidden-layer neural network with 256 neurons was placed on the output. The ‘large’ is similar, except that it used 128 filters in both the blocks and the residual path, had 8 blocks, used a kernel size of 8, and the final neural network had 1024 neurons.

The logistic regression was performed three times with different amounts of L2L^{2} regularisation, with scaling hyperparameters of 0.01, 0.2 and 1; for every experiment the regularization hyperparameter achieving the best accuracy on the test set was used.

The random forest used the default Scikit-learn implementation with a maximum depth of 6 and 100 trees.

Optimiser

The GRU and CNN were optimised using Adam (Kingma & Ba 2015). The learning rate was 0.01 for the GRU, and 0.001 for the residual CNN. The small models were trained for a maximum of 500 epochs; the large models were trained for a maximum of 1000 epochs. The learning rate was decreased by a factor of 10 if validation loss did not improve over a plateau of 10 epochs. Early stopping was used if the validation loss did not improve for 30 epochs. After training the parameters were always rolled back to those that demonstrated the best validation loss over the course of training. The batch size used varied by dataset; in each it was taken to be the power of two that meant that the number of batches per epoch was closest to 40.

Computing infrastructure

Experiments were run on an Amazon AWS G3 Instance (g3.16xlarge) equipped with 4 Tesla M60s, parallelized using GNUParallel (Tange 2011).

C.2 Analysis of variations of the signature method

The UEA archive comes with a pre-defined train-test split, which we respect. We take an 80%/20% train/validation split in the training data, stratified by class label. For the Human Activities and Postural Transitions dataset, we take a 60%/15%/25% train/validation/test split from the whole dataset. For the Speech Commands dataset, we take a 68%/17%/15% train/validation/test split from the whole dataset. (These somewhat odd choices corresponding to taking either 25% or 15% of the dataset as test, and then splitting the remaining 80%/20% between train and validation.) These train/validation splits are only used for the training of the GRU and CNN classifiers.

Combinations

In total we tested 8569 different combinations.

The variations tested are divided into groups. The first group consists of the sensitivity-adding augmentations, namely time, basepoint and invisibility-reset. Relative to the baseline model, we test every possible combination of these. (Including using none of them.)

The second group consists of those other augmentations, namely the lead-lag, singleton coordinate projection, pair coordinate projection, triplet coordinate projection, random projections, learnt projections, and multi-headed stream preserving neural networks, and finally also the case of no additional augmentation.

For the random projections, we consider four possibilities, with e∈{3,6}e\in\{3,6\} and p∈{2,5}p\in\{2,5\}, all relative to the baseline model.

For no additional augmentation, lead-lag, coordinate projections, learnt projections, and the multi-headed stream preserving neural networks, we compose them with the time, time+basepoint and time+invisibility-reset augmentations (the clear best three from the first group), all relative to the baseline model.

For the learnt projections, we consider four different possibilities corresponding to e∈{3,6}e\in\{3,6\} and p∈{2,5}p\in\{2,5\}; together with the time/time+basepoint/time+invisibility-reset cases this yields a total of twelve possibilities.

For the multi-headed stream-preserving neural networks, we again consider four different possibilities corresponding to e∈{3,6}e\in\{3,6\} and p∈{2,5}p\in\{2,5\}, for a total of twelve possible augmentation strategies. In each the neural network operates elementwise, so as to map one sequence to another, and is given by a feedforward neural network of three hidden layers separated by ReLU activation functions. When e=3e=3 the hidden layers have 16 neurons each, and when e=6e=6 they have 32 neurons each.

For both the learnt projections and multi-headed stream-preserving neural networks, training these requires backpropagating through the model, so these were only considered for the GRU and residual CNN model. (The logistic regression model would in principle be possible as well, except that we ended up implementing this through Scikit-learn rather than PyTorch.)

We note that there are a great many possible ways of doing stream preserving neural networks, of which these are a small fraction. Their relatively weak performance here may likely be improved upon with greater tuning on an individual task, or the selection of better final models than were considered here.

The third group consisted of the different windows. Recall that the baseline model used a global window; we then consider varying this to two possible sliding windows, two possible expanding windows, and three possible dyadic windows. The two possible sliding/expanding windows are chosen so that either 5 or 20 windows are applied across the full length of the dataset. The three possible dyadic windows are depths 2, 3, 4. Thus in total there are 8 possible window combinations we consider.

The fourth group consists of rescaling options, namely no rescaling, pre-signature rescaling, and post-signature rescaling.

Omissions

For the empirical study on the variations on the signature method, we excluded those UEA datasets with a dimension dd over 60, so as to reduce the computational cost. This results removes 6 of the 30 datasets from the study, namely DuckDuckGeese, FaceDetection, Heartbeat, InsectWingbeat, MotorImagery, and PEMS-SF. These were nonetheless used in the demonstration of performance of the canonical signature method in Figure 6. Furthermore those combinations of dataset/variation/model which produced more than 10510^{5} signature features were omitted, to keep the computation managable. See Table 3.

C.3 The canonical signature pipeline

For each dataset, we implement the following steps. First, the sequences are augmented with time and basepoint augmentations. Then, we consider every combination of signature depth in {1,2,3,4,5,6}\{1,2,3,4,5,6\} and hierarchical dyadic window depth in {2,3,4}\{2,3,4\}. For each of these choices, we perform a randomized grid search on a random forest classifier to optimize its number of trees and maximal depth parameters. We test 20 combinations randomly sampled from the following grids:

Note that a maximal depth set to ‘None’ means that the trees are expanded until all leaves contain exactly one sample. Finally, we choose the combination of signature and hierarchical dyadic window depths which maximise the out-of-bag score.

Appendix D Additional results

To get a sense of the cost of each augmentation or window, we present the run times of each augmentation/model combination, and each window/model combination. (The times for varying between signature and logsignature, and between different rescalings, are largely insignificant.) See Table 4.

The run times are averaged over every UEA dataset. As the datasets are of very different sizes this thus represents quite a crude statistic, and in particular produces very large variances, so these are most meaningful simply with respect to each other.

Sensitivity-inducing augmentations broken down by dataset type

Table 5 shows the average rank of each of the first group of augmentations (that add sensitivity to certain kinds of perturbation) by dataset type, where the types are taken from Bagnall et al. 2018. (This may be regarded as a companion to Table 6.)

It is interesting to note that for EEG data, it seems better not to consider the time augmentation, whereas it is the case for other applications. In particular the combination of time and basepoint augmentations achieve the best ranks for human action and motion recognition (HAR and MOTION in Table 5). Recognizing an action may not be translation-invariant nor invariant by time reparametrization.

Other augmentations broken down by dataset characteristichs

Table 6presents the average ranks of the other augmentations borken down by some characteristics of the datasets.

Here we see that there is generally a better choice than doing nothing at all, but that this better choice. For example on long or high-dimensional datasets, coordinate projections often perform well, whilst multi-headed stream preserving transformations do substantially better on EEG datasets. Lead-lag remains a strong choice in many cases.

Depth study on the signature transform

In the main text we focused on the difference between the signature and logsignature transforms, and stated that larger depths must be chosen by a bias-variance tradeoff. Here we consider varying the depth together with the choice of signature or logsignature, and taking the best transform for each depth. See Figure 7. We see that larger depths do indeed generally correspond to increased performance, up to a point. The optimal depth will depend on the complexity of the task, as the number of features increases exponentially with the depth.

Rescaling critical difference diagram

In Figure 8, we see that pre-signature rescaling performs significantly worse than the other two options and that no significant difference between post-rescaling and no rescaling is found.

D.2 Complete results

We present in Tables 7, 8, 9, 10, 11 and 12 the performance of the different signature variations on each dataset. The tables were obtained by maximizing the test accuracy of the signature method over the different classifiers considered. Recall that some values are omitted due to the large number of signature features that would be obtained.

D.3 Canonical signature method

In Table 13 we give the full results for our canonical signature method on all UEA datasets, together with the results of Ruiz et al. 2020 used in Figure 6.

Finally, we give in Table 14 the hyperparameters that were selected for each dataset in the signature pipeline model.