Classification Accuracy Score for Conditional Generative Models

Suman Ravuri, Oriol Vinyals

Introduction

Evaluating generative models of high-dimensional data remains an open problem. Despite a number of subtleties in generative model assessment , in a quest to improve generative models of images, researchers, and particularly those who have focused on Generative Adversarial Networks , have identified desirable properties such as “sample quality” and “diversity” and proposed automatic metrics to measure these desiderata. As a result, recent years have witnessed a rapid improvement in the quality of deep generative models. While ultimately the utility of these models is their performance in downstream tasks, the focus on these metrics has led to models whose samplers now generate nearly photorealistic images . For one model in particular, BigGAN-deep , results on standard GAN metrics such as Inception Score (IS) and Frechet Inception Distance (FID) approach those of the data distribution. The results on FID, which purports to be the Wasserstein-2 metric in a perceptual feature space, in particular suggest that BigGANs are capturing the data distribution.

A similar, though less heralded, improvement has occurred for models whose objectives are (bounds of) likelihood, with the result that many of these models now also produce photorealistic samples. Examples include: Subscale Pixel Networks , unconditional autoregressive models of 128×\times128 ImageNet that achieve state-of-the-art test set log-likelihoods; Hierarchical Autoregressive Models (HAMs) , class-conditional autoregressive models of 128×\times128 and 256×\times256 ImageNet; and the recently introduced Vector-Quantized Variational Autoencoder-2 (VQ-VAE-2) , a variational autoencoder that uses vector quantization and an autoregressive prior to produce high-quality samples. Notably, these models measure diversity using test set likelihood and assess sample quality through visual inspection, eschewing the metrics typically used in GAN research.

As these models increasingly seem “to learn the distribution” according to these metrics, it is natural to consider their use in downstream tasks. Such a view certainly has a precedent: improved test set likelihoods in language models, unconditional models of text, also improve performance in tasks such as speech recognition . While a generative model need not learn the data distribution to perform well on a downstream task, poor performance on such tasks allows us to diagnose specific problems with both our generative models and the task-agnostic metrics we use to evaluate them. To that end, we use a general framework (first posed in and further studied in ) in which we use conditional generative models to perform approximate inference and measure the quality of that inference. The idea is simple: for any generative model of the form pθ(x∣y)p_{\theta}(x|y), we learn an inference network p^(y∣x)\hat{p}(y|x) using only samples from the conditional generative model and measure the performance of the inference network on a downstream task. We then compare performance to that of an inference network trained on real data.

We apply this framework to conditional image models where yy is the image label, xx is the image, and task is image classification. (N.B. this approach has been used for evaluating smaller scale GANs ). The performance measure we use, Top-1 and Top-5 accuracy, denote a Classification Accuracy Score (CAS). The gap in performance between networks trained on real and synthetic data allows us to understand specific deficiencies in the generative model. Although a simple metric, CAS reveals some surprising results:

When using a state-of-the-art GAN (BigGAN-deep) and an off-the-shelf ResNet-50 classifier as the inference network, we found that Top-1 and Top-5 accuracies decrease by 27.9% and 41.6%, respectively, compared to using real data.

Conditional generative models based on likelihood, such as VQ-VAE-2 and HAM, perform well compared to BigGAN-deep, despite achieving relatively poor Inception Scores and Frechet Inception Distances. Since these models produce visually appealing samples, the result suggests that IS and FID are poor measures of non-GAN model performance.

CAS automatically surfaces particular classes for which BigGAN-deep and VQ-VAE-2 fail to capture the data distribution and were previously unknown in the literature. Figure 1 shows four such classes for BigGAN-deep.

We find that neither IS, nor FID, nor combinations thereof are predictive of CAS. As generative models may soon be deployed in downstream tasks, these results suggest that we should create metrics that better measure task performance.

We calculate a Naive Augmentation Score (NAS), a variant of CAS where the image classifier is trained on both real and synthetic images, to demonstrate that classification performance improves in limited circumstances. Augmenting the ImageNet training set with low-diversity BigGAN-deep images improves Top-5 accuracy by 0.2%, while augmenting the dataset with any other synthetic images degrades classification performance.

In Section 2 we provide a few definitions, desiderata of metrics, and shortcomings of the most popular metrics in relation to different research directions for generative modeling. Section 3 defines CAS. Finally, Section 4 provides a large-scale study of current state-of-the-art generative models using FID, IS, and CAS on both the ImageNet and CIFAR-10 datasets.

Metrics for Generative Models

Much of the difficulty in evaluating any generative model is not knowing the task for which the model will be used. Understanding how the model will be deployed, however, has important implications on its desired properties. For example, consider the seemingly similar tasks of automatic speech recognition and speech synthesis. While both tasks may share the same generative model of speech—such as a hidden Markov Model pθ(o,l)p_{\theta}(o,l) with the observed and latent variables being the waveform oo and word sequence ll, respectively—the implications of model misspecification are vastly different. In speech recognition, the model should be able to infer words for all possible speech waveforms, even if the waveforms themselves are degraded. In speech synthesis, however, the model should produce the most realistic-sounding samples, even if it cannot produce all possible speech waveforms. In particular, for automatic speech recognition, we care about pθ(l∣o)p_{\theta}(l|o), while for speech synthesis, we care about o∼pθ(o∣l)o\sim p_{\theta}(o|l).

In absence of a known downstream task, we assess to what extent the model distribution pθ(x)p_{\theta}(x) matches the data distribution pdata(x)p_{data}(x), a less specific and often more difficult goal. Two consequences of the trivial observation that pθ(x)=pdata(x)p_{\theta}(x)=p_{data}(x) are: 1) each sample x∼pθ(x)x\sim p_{\theta}(x) “comes” from the data distribution (i.e., it is a “plausible” sample from the data distribution), and 2) that all possible examples from the data distribution are represented by the model. Different metrics that evaluate the degree of model mismatch weigh these criteria differently.

Furthermore, we expect our metrics to be relatively fast to calculate. This last desideratum often depends on the model class. The most popular seem to be:

(Inexact) Likelihood models using variational inference (e.g., VAE )

Likelihood using autoregressive models (e.g., PixelCNN )

Likelihood models based on bijections (e.g., GLOW , rNVP )

(Possibly inexact) likelihood using energy-based models (e.g., RBM )

For the first four of these classes, the likelihood objective provides us scaled estimates of the KL-divergence between the data and model. Furthermore, test set likelihood is also an implicit measure of diversity. The likelihood, however, is a fairly poor measure of sample quality and often scores out-of-domain data more highly than in-domain data .

Reliance on IS and FID in particular has led to improvement in GAN models but has certain deficiencies. IS does not penalize a lack of intra-class diversity, and certain out-of-distribution samples produce Inception Scores three times higher than that of the data . FID, on the other hand, suffers from a high degree of bias . Moreover, the pool3 feature layer may not even correlate well with human judgment of sample quality . In this work, we also find that non-GAN models have rather poor Inception Scores and Frechet Inception Distances, even though the samples are visually appealing.

Rather than creating ad-hoc heuristics aimed at broadly measuring sample quality and diversity, we instead evaluate generative models by assessing their performance on a downstream task. This is akin to measuring a generative model of speech by evaluating it on automatic speech recognition. Since models considered here are implicit or do not admit exact likelihoods, exact inference is difficult. To circumvent this issue, we train an inference network on samples from the model. If the generative model is indeed capturing the data distribution, then we could replace the original distribution with a model-generated one, perform any downstream task, and obtain the same result. In this work, we study perhaps the simplest downstream task: image classification.

This idea is not necessarily new: for GAN evaluation, it has been independently discovered at least four times. first introduced the metric (denoted “adversarial accuracy”) to measure their proposed Layer-Recursive GAN and connected image classification to approximate inference. more systematically studied this idea of approximate inference to measure the boundary distortion induced by GANs. They did this by training separate per-label unconditional generative models, and then trained classifiers on synthetic data to understand how the boundary shifted and to measure the sample diversity of GANs. Predating , used “Train on Synthetic, Test on Real” to measure a recurrent conditional GAN for medical data. trained on synthetic data, tested on real (denoted “GAN-train”) as an approximate recall metric for GANs. They also trained on real data and tested on synthetic (denoted “GAN-test”) as an approximate precision test. Unlike previous work, they tested on larger datasets such as 128×\times128 ImageNet, but with smaller scale models such as SNGAN .

The metrics mentioned above are by no means the only ones, and researchers have proposed methods to evaluate other properties of generative models. constructs approximate manifolds from data and samples, and applies the method to GAN samples to determine whether mode collapse occurred. attempts to determine the support size of GANs by using a Birthday Paradox test, though the procedure requires a human to identify two nearly-identical samples. Maximum Mean Discrepancy is a two-sample test that has many nice theoretical properties but seems to be less used because the choice of kernels do not necessarily coincide with human judgment. Procedurally similar to our method, proposes a “reverse LM score”, which trains a language model on GAN data and tests on a real held-out set. measures the quality of generative models of text by training a sentiment analysis classifier. Finally, classifies real data using a student network mimicking a teacher network pretrained on real data but distilled on GAN data.

Our work most closely mirrors , but differs in a some key respects. First, since we view image classification as approximate inference, we are able to describe its limitations in Section 3, and verify the approximation in Section 4.5. Second, while in performance on GAN-train correlates with improved IS and FID, we focus more on large-scale and non-GAN models, such as VQ-VAE-2 and HAMs, where FID and IS are not indicative of classification performance. Third, by polling the inference network, we can identify classes for which the model failed to capture the data distribution. Finally, we open-source the metric for ImageNet for ease of evaluating large-scale generative models.

Classification Accuracy Score

At the heart of CAS lies a very simple idea: if the model captures the data distribution, performance on any downstream task should be similar whether using the original or model data. To make this intuition more precise, suppose that data comes from a distribution p(x,y)p(x,y), the task is to infer yy from xx, and we suffer a loss L(y,y^)L(y,\hat{y}) for predicting y^\hat{y} when the true label is yy. The risk associated with a classifier y^=f(x)\hat{y}=f(x) is:

As we only have samples from p(x,y)p(x,y), we measure the empirical risk 1NL(yi,f(xi))\frac{1}{N}L(y_{i},f(x_{i})). From the right hand side of Equation 1, of the set of predictions Y\mathcal{Y}, the optimal one y^\hat{y} minimizes the expected posterior loss:

Assuming we know the label distribution p(y)p(y), a generative modeling approach to this problem is to model the conditional distribution pθ(x∣y)p_{\theta}(x|y), and infer labels using Bayes rule: pθ(y∣x)=pθ(x∣y)p(y)pθ(x)p_{\theta}(y|x)=\frac{p_{\theta}(x|y)p(y)}{p_{\theta}(x)}. If pθ(y∣x)=p(y∣x)p_{\theta}(y|x)=p(y|x), then we can make predictions that minimize the risk for any loss function. If the risk is not minimized, however, then we can conclude that distributions are not matched, and we can interrogate pθ(y∣x)p_{\theta}(y|x) to better understand how our generative models failed.

In the case of conditional generative models of images, yy is the class label for image xx, and the model of p^(y∣x)\hat{p}(y|x) is an image classifier. We use ResNets in this work. The loss functions LL we explore are the standard ones for image classification. One is 0-1, which yields Top-1 accuracy, and the other is 0-1 in the Top-5, which yields Top-5 accuracy.It is more correct to state that the losses yield errors, but we present results as accuracies instead as they are standard in computer vision literature. Procedurally, we train a classifier on synthetic data, and evaluate the performance of the classifier on real data. We call the accuracy the Classification Accuracy Score (CAS).

Note that a CAS close to that for the data does not imply that the generative model accurately modeled the data distribution. This may happen for a few reasons. First, pθ(y∣x)=p(y∣x)p_{\theta}(y|x)=p(y|x) for any generative model that satisfies pθ(x∣y)pθ(x)=p(x∣y)p(x)\frac{p_{\theta}(x|y)}{p_{\theta}(x)}=\frac{p(x|y)}{p(x)} for all x,y∼p(x,y)x,y\sim p(x,y). One example is a generative model that samples from the true distribution with probability pp, and from a noise distribution with a support disjoint from the true distribution with probability 1−p1-p. In this case, our inference model is good but the underlying generative model is poor.

Second, since the losses considered here are not proper scoring rules , one could obtain reasonable CAS from suboptimal inference networks. For example, suppose that p(y∣x)=1.0p(y|x)=1.0 for the correct class while p^(y∣x)=0.51\hat{p}(y|x)=0.51 for the correct class due to poor synthetic data. CAS for both is 100%. Using a proper scoring rule, such as Brier Score, eliminates this issue, but experimentally we found limited practical benefit from using one.

Finally, a generative model that memorizes the training set will achieve the same CAS as the original data.N.B. IS and FID also suffer the same failure mode. In general, however, we hope that generative models produce samples disjoint from the set on which they are trained. If the samples are sufficiently different, we can train a classifier on both the original data and model data and expect improved accuracy. We denote the performance of classifiers trained on this “naive augmentation” Naive Augmentation Score (NAS). Our CAS results, however, indicate that the current models still significantly underfit, rendering the conclusions less compelling. For completeness, we include results on augmentation in Section 4.4.

Despite these theoretical issues, we find that generative models have Classification Accuracy Scores lower than the original data, indicating that they fail to capture the data distribution.

Computationally, training classifiers is significantly more demanding than calculating FID or IS over 50,000 samples. We believe, however, that now is the right time for such a metric due to a few key advances in training classifiers: 1) the training of ImageNet classifiers has been reduced to minutes , 2) with cloud services, the variance due to implementation details of such a metric is largely mitigated, and 3) the price and time cost of training classifiers on cloud services is reasonable and will only improve over time. Moreover, many class-conditional generative models are computationally expensive to train, and thus, even a relatively expensive metric such as CAS comprises a small percentage of the training budget.

We open-source our metric on Google Cloud for others to use. The instructions are given in Appendix B. At the time of writing, one can compute the metric in 10 hours for roughly 15,orin45minutesforroughly15, or in 45 minutes for roughly85 using TPUs. Moreover, depending on affiliation, one may be able to access TPUs for free using the Tensorflow Research Cloud (TFRC) (https://www.tensorflow.org/tfrc/).

Experiments

Our experiments are simple: on ImageNet, we use three generative models—BigGAN-deep at 256×\times256 and 128×\times128 resolutions, HAM with masked self-prediction auxiliary decoder at 128×\times128 resolution, and VQ-VAE-2 at 256×\times256 resolution—to replace the ImageNet training set with a model-generated one, train an image classifier, and evaluate performance on the ImageNet validation set. To calculate CAS, we replace the ImageNet training set with one sampled from the model, and each example from the original training set is replaced with a model sample from the same class.

In addition, we compare CAS to two traditional GAN metrics: IS and FID, as these metrics are the current gold standard for GAN comparison and are being used to compare non-GAN to GAN models. Both rely on a feature space from a classifier trained on ImageNet, suggesting that if metrics are useful at predicting performance on a downstream task, it would indeed be this one.

Further details about the experiment can be found in Appendix A.1.

Table 1 shows the performance of classifiers trained on model-generated datasets compared to those on the real dataset for 256×\times256 and 128×\times128, respectively. At 256×\times256 resolution, BigGAN-deep achieves a CAS Top-5 of 65.92%, suggesting that BigGANs are learning nontrivial distributions. Perhaps surprisingly, VQ-VAE-2, though performing quite poorly compared to BigGAN-deep on both FID and IS, obtains a CAS Top-5 accuracy of 77.59%. Both models, however, lag the original 256×\times256 dataset, which achieves a CAS Top-5 Accuracy of 91.47%.

We find nearly identical results for the 128×\times128 models. BigGAN-deep achieves CAS Top-5 and Top-1 similar to the 256×\times256 model (note that IS and FID results for 128×\times128 and 256×\times256 BigGAN-deep are vastly different). HAMs, similar to VQ-VAE-2, perform poorly on FID and Inception Score but outperform BigGAN-deep on CAS. Moreover, both models underperform relative to the original 128×\times128 dataset.

2 Uncovering Model Deficiencies

To better understand what accounts for the gap between generative model and dataset CAS, we broke down the performance by class (Figure 2). As shown in the left pane, nearly every class of BigGAN-deep suffers a drop in performance compared to the original dataset, though six classes—partridge, red fox, jaguar/panther, squirrel monkey, African elephant, and strawberry—show marginal improvement over the original dataset. The left pane of Figure 3 shows the two best and two worst performing categories, as measured by the difference in classification performance. Notably, for the two worst performing categories and two others—balloon, paddlewheel, pencil sharpener, and spatula—classification accuracy was 0% on the validation set.

The per-class breakdown of VQ-VAE-2 (middle pane of Figure 2) shows that this model also underperforms the real data for most classes (only 31 classes perform better than the original data), though the gap is not as large as for BigGAN-deep. Furthermore, VQ-VAE-2 has better generalization performance in 87.6% of classes compared to BigGAN-deep, and suffers 0% classification accuracy for no classes. The middle pane of Figure 3 shows the two best and two worst performing categories.

The right panes of Figures 2 and 3 show the per-class breakdown and top and bottom two classes, respectively, for HAMs. Results broadly mirror those of VQ-VAE-2.

3 A Note on FID and a Second Note on IS

We note that IS and FID have very little correlation with CAS, suggesting that alternative metrics are needed if we intend to deploy our models on downstream tasks. As a controlled experiment, we calculate CAS, IS, and FID for BigGAN-deep models with input noise distributions truncated at different values (known as the “truncation trick”). As noted in , lower truncation values seem to improve sample quality at the expense of diversity. For CAS, the correlation coefficient between Top-1 Accuracy and FID is 0.16, and IS -0.86, the latter result incorrectly suggesting that improved IS is highly correlated with poorer performance. Moreover, the best-performing truncation values (1.5 and 2.0) have rather poor ISs and FIDs. That these poor IS/FID also seem to indicate poor performance on this metric is no surprise; that other models, with well-performing ISs and FIDs yield poor performance on CAS suggests that alternative metrics are needed. One can easily diagnose the issue with IS: as noted in , IS does not account for intra-class diversity, and a training set with little intra-class diversity may make the classifier fail to generalize to a more diverse test set. FID should better account for this lack of diversity at least grossly, as the metric, calculated as FID(Px,Py)=∥μx−μy∥2+tr(Σx+Σy−2(ΣxΣy)1/2)FID(P_{x},P_{y})=\|\mu_{x}-\mu_{y}\|^{2}+tr(\Sigma_{x}+\Sigma_{y}-2(\Sigma_{x}\Sigma_{y})^{1/2}), compares the covariance matrices of the data and model distribution. By comparison, CAS offers a finer measure of model performance, as it provides us a per-class metric to identify which classes have better or worse performance. While in theory one could calculate a per-class FID, FID is known to suffer from high bias for a low number of samples, suggesting that per-class estimates would be unreliable. proposed KID, an unbiased alternative to FID, but this metric suffers from variance too large to be reliable when using the number of per-class samples in the ImageNet training set (roughly 1,000 per class), much less when using the 50 in the validation set.

Perhaps a larger issue is that IS and FID heavily penalize non-GAN models, suggesting that these heuristics are not suitable for inter-model-class comparisons. A particularly egregious failure is that IS and FID aggressively penalize certain types of samples that look nearly identical to the original dataset. For example, we computed CAS, IS, and FID on the ImageNet training set at 256×\times256 resolution and on reconstructions from VQ-VAE-2. As shown in Figure 4, the reconstructions look nearly identical to the original data. As noted in Table 2, however, IS decreases by 128 points and FID increases by 3.5×\times. The drop in performance is so great that BigGAN-deep at 1.00 truncation achieves better IS and FID than nearly-identical reconstructions. CAS Top-1 and Top-5 for the reconstructions, however, drops by 2.2% and 4.4%, respectively, relative to the original dataset. CAS for BigGAN-deep model at 1.00 truncation, on the other hand, drops by 31.1% and 46.5% relative.

4 Naive Augmentation Score

To calculate NAS, we add to the original ImageNet training set 25%, 50%, or 100% more data from each of our models. The original ImageNet training set achieves a Top-5 accuracy of 92.97%.

Although the CAS results for BigGAN-deep, and to a lesser extent VQ-VAE-2, suggest that augmenting the original training set with model samples will not result in improved classification performance, we wanted to study whether the relative ordering on the CAS experiments would hold for the NAS ones. Figure 5 illustrates the performance of the classifiers as we increase the amount of synthetic training data. Perhaps somewhat surprisingly, BigGAN-deep models that sample from lower truncation values, and have lower sample diversity, are able to perform better for data augmentation compared to those models that performed well on CAS. In fact, for some of the lowest truncation values, we found a modest improvement in classification performance: roughly 0.2%0.2\%. Moreover, VQ-VAE-2 underperforms relative to BigGAN-deep models. Of course, the caveat is that the former model does not yet have a mechanism to trade off sample quality from sample diversity.

The results on augmentation highlight different desiderata for samples that are added to the dataset rather than replaced. Clearly, the samples added should be sufficiently different from the data to allow the classifier to better generalize, yet poorer sample quality may lead to poorer generalization compared to the original dataset. This may be the reason why extending the dataset with samples generated from a lower truncation value —which are higher-quality, but less diverse—perform better on NAS than CAS. Furthermore, this may also explain why IS, FID, and CAS are not predictive of NAS.

5 Model Comparison on CIFAR-10

Finally, we also compare CAS for different model classes on CIFAR-10. We compare four models: BigGAN, cGAN with Projection Discriminator , PixelCNN , and PixelIQN . We train a ResNet-56 following the training procedure of . More details can be found in Appendix A.2. Similar to the ImageNet experiments, we find that both GANs produce samples with a certain degree of generalization. GANs also significantly outperform PixelCNN on this benchmark. Furthermore, since PixelCNN is an exact likelihood model, we can compare classification performance with exact inference using Bayes rule to that with approximate inference using a classifier. Perhaps surprisingly, CAS for PixelCNN is modestly better than classification accuracy using exact inference, though both results are similar. Finally, PixelIQN has similar performance to the newer GANs.

Conclusion

Good metrics have long been an important, and perhaps underappreciated, component in driving improvements in models. It may be particularly important now, as generative models have reached a maturity that they may be deployed in downstream tasks. We proposed one, Classification Accuracy Score, for conditional generative models of images and found the metric practically useful in uncovering model deficiencies. Furthermore, we find that GAN models of ImageNet, despite high sample quality, tend to underperform models based on likelihood. Finally, we find that IS and FID unfairly penalize non-GAN models.

An open question in this work is understanding to what extent these models generalize beyond the training set. While current results suggest that even state-of-the-art models currently underfit, recent progress indicates that underfitting may be a temporary issue. Measuring generalization will then be of primary importance, especially if models are deployed on downstream tasks.

Acknowledgments

We would like to thank Ali Razavi, Aaron van den Oord, Andy Brock, Jean-Baptiste Alayrac, Jeffrey De Fauw, Sander Dieleman, Jeff Donahue, Karen Simonyan, Takeru Miyato, and Georg Ostrovski for discussion and help with models. Furthermore, we would like to thank those who contacted us, pointing us to prior work.

References

Appendix A Experimental Setup

We use a ResNet-50 classifier with single-crop evaluation for our models on ImageNet. The classifier is trained for 90 epochs using TensorFlow’s momentum optimizer with a learning rate schedule linearly increasing from 0.0 to 0.4 for the first 5 epochs and decreasing by a factor of 10 at epochs 30, 60, and 80. It mirrors the 8,192 batch setup of with gradual warmup. These were trained on 128 TPU chips, and training completed in roughly 45 minutes. We also compared results to those trained on 8 TPU chips with batch size 1,024 (and completed in roughly 10 hours), and found that Top-1 and Top-5 accuracy were within 0.4%.

For BigGAN-deep models, since the truncation trick – which resamples dimensions that are outside the mean of the distribution – seems to trade off quality for diversity, we perform experiments for a sweep of truncation parameters: 0.2, 0.42, 0.5, 0.8, 1.0, 1.5, and 2.0.Dimensions of the noise vector zz whose value are greater outside the range of −2τ-2\tau to 2τ2\tau (τ\tau is the truncation parameter) are resampled. Lower values of τ\tau lead to less diverse datasets.

A.2 CIFAR-10

We use a ResNet-56 classifier for our models on CIFAR-10, using 45,000 samples for the training set and 5,000 for validation. We train for 182 epochs, starting at learning rate 1.0 and decaying by a factor of 10 at epochs 91 and 136. We use batch size 128. Note that this setup mirrors the one in .

Appendix B Instructions for Operating on Google Cloud

(these broadly follow https://cloud.google.com/tpu/docs/tutorials/resnet, with some changes for the metric)

For first-time users: 1. Create a project using: https://console.cloud.google.com/cloud-resource-manager

2. Enable Billing: https://cloud.google.com/billing/docs/how-to/modify-project

3. Create a storage bucket (used for storing data and models): https://console.cloud.google.com/storage/browser. Make sure this is in the zone us-central, as this has the cheapest pricing.

1. Launch google cloud shell (https://cloud.google.com/shell/)

(where is a,b for paying customers, and f for those in the TFRC program)

3. You will now be in a virtual machine. Run tmux to keep a persistent ssh.

1. Launch google cloud shell (https://cloud.google.com/shell/)

(where is a,b for paying customers, and f for those in the TFRC program)

3. You will now be in a virtual machine. Run tmux to keep a persistent ssh.

13. follow steps of 10-hour training, except for steps 6 and 9.