On Aliased Resizing and Surprising Subtleties in GAN Evaluation
Gaurav Parmar, Richard Zhang, Jun-Yan Zhu
Introduction
With the proliferation of generative modeling techniques, such as Generative Adversarial Networks (GANs) , accurately discerning which methods are performing better has become a critical aspect of the field. For visual data, metrics such as Inception Score (IS) , Kernel Inception Distance (KID) , and the ubiquitously-used Fréchet Inception Distance (FID) have become standard practice for developing and adopting models. Under the hood, these methods evaluate the discrepancy between generated and natural images, in a deep feature space, to capture relevant features of the two distributions. After all, at its core, generative modeling involves learning and mimicking high-order, complex statistics of visual data.
However, we find that low-level, seemingly innocuous operations, can induce surprisingly large discrepancies in high-level statistics. For example, consider Figure 1. Given the same input image, different image processing libraries produce drastically different results. Specifically, the implementations using OpenCV, TensorFlow and PyTorch libraries with default flags, contain severe aliasing artifacts. Similarly, the simple act of saving images in a JPEG operation with the default parameters, either when building the training dataset or collection of generated images, adds quantization and low-level statistical differences to the underlying data. The low-level statistical differences induced by these differences cause meaningful variations when used for evaluation protocols. As the Fréchet Inception Distance (FID) metric is the most ubiquitous , it is the focus of our experiments. We offer a standard benchmark, clean-fid (github.com/GaParmar/clean-fid), and concrete suggestions on resizing and quantization procedures to enable clean comparisons in future evaluation protocols.
First, we investigate the implications of image resizing. When downsampling, signal processing techniques recommend “prefiltering” the input, to prevent high-frequency elements from aliasing into the output. When the downsampling factor is larger, the prefilter kernel should be correspondingly stretched. However, as shown in Figure 2, the resizing function used by the FID implementations in TensorFlow and PyTorch do not prefilter the image, resulting in aliasing artifacts shown in Figure 1. Resizing can occur in two locations – during data preprocessing (training with lower resolution) or at evaluation time (resizing to 299 resolution to compute the FID metric). In both cases, inconsistent resizing functions induce variations downstream. If used for data preprocessing, the training data distribution itself is changed. When used for the evaluation metric, small variations in resizing can cause changes in subsequent feature extraction. We quantify the effects of these inconsistencies and offer standard recommendations. Specifically, we propose to use a stronger bicubic filter ; more importantly, we propose to adjust prefiltering width based on the resizing factors, as guided by signal processing principles.
Secondly, we investigate the implication of image compression. While the JPEG protocol is a lossy compression scheme, designed to preserve perceptual similarity to the original , it can perturb an image enough to corrupt downstream feature extraction. This affects performance drastically and can create mismatches when comparing methods. Perhaps more surprisingly, when training images are saved with JPEG compression, modern GANs are unable to fully mimic the induced artifacts, and large FID improvements can actually be artificially achieved by tweaking the JPEG compression ratios when storing the generated images. We quantify the surprising effects of this compression operation, and again offer a concrete, standardized protocol to avoid inconsistencies and hindrances to proper evaluation.
In conclusion, we characterize the surprising importance of low-level image processing steps, resizing and quantization, when training and evaluating generative models, such as GANs. We focus our experiments on the widely adopted FID metric, and show additional results on the KID metric as well as IS and Perceptual Path Length (PPL) metrics (in the supplement). Importantly, any metric, present or future, that derives statistics from images undergoing these processing steps, will be affected by these factors. More details and results can be found on our website.
Related Work
A wide range of image and video synthesis applications have been enabled, as a result of tremendous progress in deep generative models such as GANs , VAEs , autoregressive models , flow-based models , and energy-based models . It is often relatively easier to evaluate individual model’s performance on downstream computer vision and graphics tasks, as they have a clear target for a given input. However, evaluating unconditional generative models remains an open problem. It is still an important goal, as most generative models are not tailored to any downstream task.
Evaluating generative models.
The community has introduced many evaluation protocols. One idea is to conduct user studies on cloud-sourcing platforms for either assessing the samples’ image quality or identifying duplicate images . Due to the subtle differences in user study protocols (e.g., UI design, fees, date/time), it is not easy to replicate results across different papers. Large-scale user studies can also be expensive, prohibiting its usage when evaluating hundreds of model variants and checkpoints during the development stage. Several methods propose evaluating generative models from a self-supervised feature learning perspective, by repurposing the learned discriminators or accompanying encoders for a downstream classification task. However, the representation power of the discriminator or encoder does not directly reflect the generators’ sample quality and diversity. In addition, not every generative model is trained with a discriminator or encoder.
To overcome the previous issues, an area of focus is developing automatic metrics that directly assess the samples of generative models. Various metrics been proposed, criticized, and modified. Commonly-used ones include log-likelihood , density estimate with Parzen window , Inception Score , Perceptual Path Length , Fréchet Inception Distance (FID) , Classification Accuracy Score and its early variants , Classifier Two-sample Tests , precision and recall , Kernel Inception Distance (KID) , among others. Each metric has associated pros and cons and none are perfect.
Among them, Fréchet Inception Distance (FID) has become the most widely-used metrics, as it can model intra-class diversity better than Inception Score. FID is also easy and fast to compute without training additional classifiers , and has been shown to be consistent with human perception . As a result, it has been used in recent GANs papers as well as large-scale evaluation study , despite facing criticism about the fact that FID is a biased estimator and sensitive to the number of samples used in the evaluation . Our goal here is not to study which one is a better metric. Instead, we focus our study on the popular FID metric and how subtle details and aliased image resizing functions can affect the final scores. Note that the resizing and quantization we study in are applicable to any evaluation metric that contains such operations.
Antialiasing and robustness.
The study of resampling signals is central in signal processing , image processing , and computer graphics . In particular, when downsampling a signal, one must consider the Nyquist sampling criterion and antialias to prevent high-frequency information from aliasing into the output. Without proper antialiasing, in the worst case, an adversary can embed a completely different image in the original, resulting in a “scaling attack” . In convolutional network design, antialiasing has taken form in average pooling and Gaussian filtering . While it was replaced by operations such as max-pooling, based on empirical performance , recent works have demonstrated that antialiasing can be compatible and improve performance in convolutional networks , transformers , NeRFs , and GANs . Despite these advances, generative methods continue to be detectable , and discriminative networks continue to be sensitive to small perturbations, such as shifts and JPEG compression . Achieving robustness to such perturbations remains an open problem , and the preprocessing steps, such as image resizing, used before feature extraction remain consequential. We study the effect of such steps in a generative modeling pipeline and propose a standardization following signal processing principles, in order to facilitate easy and fair comparisons.
Preliminaries
In this section, we discuss several low-level image processing steps using different popular libraries. We find that many of these details can have a large effect on the FID score being computed. Figure 3 details the step-by-step process for both dataset preparation and model evaluations.
The Fréchet Inception Distance (FID) score aims to measure the gap between two data distributions , such as between a training set and samples from a generator.
Evaluating a generator with FID.
After the images are appropriately resized, and the features are extracted, the mean (, ) and covariance matrix (, ) of the corresponding set of features and are used to compute the Fréchet distance shown in the equation below.
The Tr operation calculates the trace of the matrix.The different choices for the resizing functions () and quantization functions () adds potential sources of inconsistencies in generative modeling pipelines.
2 Image Resizing
Depending on the dataset and training size, the resizing operations (, ) in Figure 3 can either be downsampling or upsampling. Downsampling is the primary focus of this investigation, as it involves throwing away information. Methods for downsampling is a common study in the fields of signal and image processing .
The most naive approach is to simply subsample (taking every N element if performing downsampling by an integer factor N), sometimes referred to as nearest. This corresponds to filtering the input image with Kronecker delta function, as only a single value is drawn. Such an approach leads to aliasing, as high-frequency elements of the input alias to the output.
A central principle in image processing, signal processing, graphics, and vision is to blur or “prefilter” before subsampling, as a means of removing high-frequency information (thus preventing its misrepresentation downstream). For linear filters, this corresponds to a “depth-wise convolution”, using deep learning parlance . We explain two important ways in which prefiltering implementations can vary.
Filter size adaptation to downsampling factor.
First, according to signal processing principles, the size of the filter should be adjusted, in accordance with the downsampling factor. Widening the low-pass filter in the spatial domain corresponds to reducing its bandwidth and filtering more aggressively in frequency space. As a larger downsampling factor means a lower bandwidth can be represented on the output signal, widening the filter accordingly is necessary to prevent aliasing. However, in many common implementations, this is not implemented (or is not used by default); instead, a filter of fixed, non-adaptive size is used.
Choice of filters.
Secondly, there is a choice of different convolutional filters. The idealized low-pass filter is a sinc, requiring infinite support. As such, approximate filters with different subtle tradeoffs in runtime and behavior are used instead. The box, also known as area filter, corresponds to a rectangular filter, computing the average of values within a neighborhood. The bilinear filter is a triangular filter, bicubic is a stronger cubic function, and the lanczos filter is an enveloped sinc. All perform a weighted average and have stronger antialiasing, closer to the idealized sinc. See Appendix 6.1 for additional details about the different interpolation filters.
Practical implications of implementation variations.
We investigate the inconsistencies that can arise, when these two factors are varied, and show a toy example in Figure 1 in downsampling a circle. While the choice of filter is largely constant across libraries (lanczos, bicubic, bilinear are shown in each column), the choice of whether the filter adapts to the downsampling factor is not. While the PIL library adapts the filter (top row), other libraries do not by default, leading to aliased results. In particular, FID implementations of TensorFlow-FID and PyTorch-FID, use bilinear downsampling implementations that exhibit aliasing, and thus are the focus of our study.
An implication of aliasing is a suboptimal representation of the original image. In Figure 4, we show the result of downsampling and upsampling an image, and comparing it to the original with PSNR (averaged over 300 FFHQ images). The methods with non-adaptive filters achieve a worse reconstruction than a method that adapts the filter. This effect is more significantly accentuated with larger downsampling factors, where high-frequency aliasing dominates when using non-adaptive filters. Figure 5 shows how the Inception features are affected by aliased resizing functions for various datasets.
Recommendation.
Above, we have established that the implementations of FID are inconsistent and aliased. Ideally, the community can (a) use a consistent pipeline to facilitate fair comparisons across papers, and (b) follows signal processing principles and antialiases, in order to best represent the underlying data it is trying to characterize. We propose to use an adaptive filter (and thus produce consistently antialiased results). Second, we propose to use a bicubic, instead of bilinear filter, which offers stronger reconstruction. While such an implementation is currently found in PIL, future implementations that are computationally equivalent would be of use).
3 Quantization and Image Compression
Image compression.
Saving the image as a raw matrix of values is data-intensive. However, an image contains redundant information that can be exploited. For example, the PNG format compresses an image losslessly. To further save storage, images are commonly saved using the JPEG codec. While JPEG is a lossy compression technique, it aims to make changes that the human visual system is less sensitive to, namely reducing information in higher frequencies and chroma (color) components . JPEG converts an image into a YCbCr space, subsamples the chroma components, divides images into 88 blocks, computes the Discrete Cosine Transform (DCT), and performs quantization. The quantization step facilitates a trade-off between the fidelity of the original image and the amount of the storage saved. In the PIL implementation , this is done using a “quality” option (0-100), which linearly scales the quantization tables (which controls which frequencies are quantized to what granularity). Note that setting the quality flag to 100 is not a lossless operation. Even when the quantization tables are not scaled, the DCT coefficients are quantized to integer values and the chroma components are subsampled.
Image compression changes deep network activations.
In Figure 6, we show a real image sampled from the FFHQ dataset at a resolution of 256, saved with lossless PNG and lossy JPEG (quality flags set to 100, 90, and 75). Despite being perceptually indistinguishable (with high PSNR values of ), the FID scores increase. The PIL default of 75 results in a high score (21), for example. Note that this FID score is far higher than the score from a powerful generative model, StyleGAN2 (around 3). Also, variations across recent methods are typically within FID on FFHQ. We further investigate the implications of using JPEG compression in various parts of the pipeline in experiments below.
Experiments
In Section 3, we outlined the various image processing steps involved in generative modeling pipelines and evaluation. In this section, we introduce sources of variation at these steps and empirically quantify their impacts. As depicted in Figure 3, the variations in the FID score arises from three distinct steps: resizing in the FID evaluation step (, ), resizing in the data preprocessing step (), and quantizing of images (, ). We investigate each of these steps in Section 4.1, Section 4.2, and Section 4.3 respectively.
Here we investigate the effects of different resizing methods (, ) used in the FID calculation step.
We start with two sets of full-resolution face images - from the FFHQ dataset, and from a pre-trained StyleGAN2 generator. Each of the sets of images is resized from 1024299 using different methods. In Table 1 (left), we compare the set of real images resized with the antialiased resizing operation (PIL bicubic) to the same set of real images, resized using other aliased functions that use a fixed width prefiltering kernel. As we compare the same set of images, we anticipate all FID and KID scores to be close to 0 and the PSNR values to be very high. However, as shown in Figures 1, 2, and 5, only a subset of the commonly used resizing operators adjust the filter width and antialias the images. These differences in resizing operations cause drastic changes in the Inception-V3 activation maps.
Filters that adapt their size and antialias are more consistent, even with different filter types – PIL-bilinear has FID 0.64 as compared to PIL-bicubic. On the other hand, implementations that ignore the downsampling factor (PyTorch and TensorFlow) show much larger deviation (FID 4.3), with scores nearing naive nearest (FID 7.4), that does not filter at all. This indicates that whether the filter adapts to the downsampling filter can change the modeled data distribution by non-trivial amounts.
Variation induced by resizing functions on generated images.
After studying the effects on real images, we evaluate how different resizing function choices affect the FID score when used in a full generative modeling pipeline. Here, we evaluate a pretrained StyleGAN2 generator trained on FFHQ (1024), MetFaces (1024), and AFHQ (512) dataset images, and calculate FID with 50,000 images. In Table 1 (right), we consider the asymmetric case, where features for the real images and generated images use different resizing functions. This case arises when features for real images are pre-computed and shared by one group of authors, while generated features may be calculated on the fly with a different library. Here, we observe that using the same resizing function as the reference dataset (PIL-bicubic) achieves the lowest performance. Using a different resize function, such as PIL-bilinear increases the score to 4. Using an aliased function increases the score drastically to 7, close to naive subsampling ().
Next, in Table 2, we show a comparison when the same resizing function is used for the real dataset images and the StyleGAN2 generated images. Interestingly, we observe that the aliased resizing functions result in lower FID scores across multiple commonly used datasets - FFHQ (1024), MetFaces (1024), and AFHQ (512). This indicates that using the antialiased function as preprocessing makes the downstream FID calculation more sensitive at measuring the discrepancies between distributions.
2 Variation due to Dataset Resizing
Previously, we considered the scenario when the dataset was not downsampled. However, as discussed in Section 1 and illustrated in Figure 3, dataset downsampling is needed when training a model on a low-resolution version of the original dataset (e.g., for FFHQ or for ImageNet). Before, the target distribution was fixed, and differences were purely introduced during post-hoc metric evaluation. Now, the situation is much more intricate. Different resizing choices will result in different training distributions entirely.
In Table 3, we train three different StyleGAN2 (config-e) models, following the official PyTorch implementationhttps://github.com/NVlabs/stylegan2-ada for 25k iterations. We resize FFHQ to using Naive Nearest, PIL–bicubic, PyTorch–bilinear, and TensorFlow–bilinear. We use the same PIL–bicubic function (, ) for FID evaluation; note that here, it is upsampling (). Qualitatively, using an aliased downsampling function produces a training distribution with visual artifacts for the generative model to mimic, likely different than the natural visual data we wish to model. Quantitatively, interestingly, we observe that that the aliased pre-processing results in lower FID values. As the antialiased function better preserves signal in the original images, we hypothesize that retaining more information from the original input actually produces a more difficult distribution to model.
3 Variation due to Quantization/Compression
In Figure 7, we test the effect of quantization applied to real FFHQ images at different resolutions on the FID (left) and KID (right) metrics. For each resolution, the real dataset images are correspondingly downsampled using PIL–bicubic, and the scores are computed between the resized uncompressed PNG images and the resized JPEG-compressed images. Figure 7 shows that the effect of the JPEG compression on both metrics. The effect is more pronounced for lower resolutions, where the artifacts remain after the subsequent resampling step.
JPEG on training images.
In both comparisons above, each method was compared with the FFHQ dataset images, which were collected as uncompressed PNG files. Any additional compression only monotonically increases the FID score (Figure 8 right). This is expected, as information is being removed from the generator.
However, this does not apply to other datasets which were collected as JPEG images. To study this effect, we train a StyleGAN2 model on the LSUN outdoor Church dataset , which saved as JPEG-75 images during data collection. In Figure 8 (left), we plot the FID of the trained generator as a function of JPEG compression. Surprisingly, we observe that the FID score for the StyleGAN2 model actually improves when slight JPEG compression is added. This indicates that interestingly, though the model is able to capture complex variations in the dataset, it is unable to fully model the low-level statistics induced by JPEG compression. Interestingly, the best FID score (3.48) is obtained when the generated images are compressed with JPEG quality 87 (not the full 75), indicating the model is able to replicate some of the artifacts, but not all. The FID score for the generated images stores as PNG files is 4.00. Furthermore, this indicates that the metric is sensitive to low-level statistics, and a large gain in the metric could be achieved simply through manual post-processing. Following these observations, we recommend that researchers curate and store training images as PNG formats for the future image synthesis datasets.
4 Consequences in model selection
In this section, we show that using an aliased, as opposed to antialiased implementation can result in different conclusions, both when comparing across different methods and when choosing a “best” model checkpoint. In particular, in Figure 9 (left) we evaluate the different intermediate checkpoints when training an image-to-image translation model on the horse2zebra dataset. In Figure 9 (right) we evaluate the StyleGAN2 models with different data augmentation trained to generate FFHQ images in a few shot setting (2000 training images). Note that using an aliased resizing implementation for computing the FID metric and choosing the best model can lead to a different best model getting selected.
Recommendations
We have shown surprisingly large sensitivities to seemingly inconsequential implementation details when evaluating generative models. The resize operation and the image quantization/compression are especially impactful. Based on our observations, we discuss some best practices when training and evaluating a generative model. We recommend using implementations that adapt the filter size to the downsampling factor, following signal processing principles, at each of the resizing steps (, , and ) involved. There are many details one needs to keep track of when computing the FID score. Any inconsistency in the steps leads to results that are no longer comparable to other methods. To facilitate an easy comparison, avoid inconsistent comparisons, and encourage the usage of critical operations that are correctly implemented, we provide an easy-to-use library, clean-fid, at github.com/GaParmar/clean-fid and pre-computed statistics of Inception features for commonly used datasets.
Acknowledgments. We thank Jaakko Lehtinen and Assaf Shocher for bringing attention to this issue and for helpful discussion. We thank Sheng-Yu Wang, Nupur Kumari, Kangle Deng, and Andrew Liu for useful discussions. We thank William S. Peebles, Shengyu Zhao, and Taesung Park for proofreading our manuscript. We are grateful for the support of Adobe, Naver Corporation, and Sony Corporation.
References
Appendix
Figure 10 shows the image downsampling procedure. When the resizing ratio is an integer, downsampling can be implemented as a discrete convolution with an interpolation kernel, followed by subsampling. As discussed in Section 3.2, the kernel needs to be widened according to the resizing ratio to prevent aliasing in the resized image. All of the commonly used interpolation filters are separable, meaning the two-dimensional interpolation over an image can be decomposed into one-dimensional interpolations along each dimension. Here, represent the spatial coordinates for the two-dimensional case and represents the spatial coordinates for the one-dimension case.
Next we describe each of the different interpolation functions in one-dimension.
The simplest form of image interpolation is the nearest neighbor interpolation which only considers the value of the neighboring point. This is equivalent to interpolating with the function shown below.
Bilinear Interpolation.
The bilinear image interpolation corresponds to interpolating using the triangle filter defined below.
Lanczos Interpolation.
The Lanczos image interpolation is the normalized sinc functions windowed by the Lanczos window .
Bicubic Interpolation.
The bicubic interpolation uses the interpolation kernel .
The common choices for the free parameter are .
Filter scaling.
As shown in Figure 2 and discussed in Section 3.2, whether to adapt the kernel width to the downsampling factor has a large qualitative and quantitative effect on the downsampled image. The continuous filter is sampled at a set of discrete locations and yield a discrete filter and normalized to sum to . The difference between adaptive and non-adaptive filters arise at which locations are sampled.
For an adaptive filter, is sampled at for even downsampling factors and for odd factors. The filter width widens with larger downsampling factor .
For a non-adaptive filter, for even factors and for odd factors. Notice the sampling locations do not scale as a function of downsampling factor .
From here, one can observe why a non-adaptive filter behaves similarly to nearest. For even factors, plugging in the sampling locations yields a 2-tap filter . For odd factors, yields delta function for all filters. In contrast, for an adaptive filter, a bilinear downsample yields a 4-tap filter, yields an 8-tap filter, etc.
2 JPEG Compression.
In Sections 3.3 and 4.3 in the main paper, we discuss the compression of images and the effects on evaluation metrics such as FID and KID. Next, we detail the JPEG compression protocol in Figure 11, and outline the three steps that result in a loss of information. Motivated by the observation that the human vision is less sensitive to color components, the first lossy step is the subsampling of color channels Cr, Cb after the color space transformation. Next, the image channels are divided into smaller blocks and the Discrete Cosine Transformation (DCT) is computed. The DCT coefficients are subsequently divided by the quantization table to suppress the higher frequencies and rounded to integers. The quantization table is determined by the user specified ”quality” option (0-100) and controls the tradeoff between the storage space and image information retained. When the quality option is set to 100, the color subsampling and integer rounding are the primary sources of information loss.
3 Library Implementation Details.
The library implementations used for the comparisons are detailed below.
Pillow Image Library (PIL) v8.0.1 : We use the standard Image.resize function; the library provides consistently antialiased results across filters.
OpenCV v4.5.5 : We use the standard cv2.resize function.
TensorFlow (TF) v2.0 : For the comparisons in this section we use the flags used by the original TensorFlow implementation of FID. The TensorFlow library has changed substantially through the versions. In this work we use the new TensorFlow version 2.0. Note that the newer version of the library has an optional flag antialias. However this option is set to False by default and not used in the current FID implementations.
PyTorch v1.9 : We use the differentiable function F.interpolate on data tensors.A separate function, torchvision.transforms.Resize, is a wrapper around the PIL library and is often used in the data pre-processing step. This resizing method has been used by popular PyTorch implementations of FID .
MXNet v1.8 : The resizing method provided in the MXNet framework is a wrapper around the OpenCV implementation.
Keras v2.6.0 : The library is built on top of the TensorFlow framework and shares the implementation for resizing images.
4 Additional resizing example
In Figure 1 in the main paper, we showed an example resizing a sparse circle. We observe that when the bicubic, lanczos, and bilinear filters do not adjust their filter widths to the downsampling factor, aliasing patterns occur. This occurs in several libraries, including the settings used in PyTorch and TensorFlow for FID calculation.
Here, in Figure 12, we show an image with varying frequency content, in order to further illustrate the behavior of different downsampling filters and implementations. The input is of size and is downsampled by to resolution . The input image is of concentric circles, with low frequency in the middle and increasing frequency towards the outside.
When the image is heavily downsampled, the high frequencies on the outside cannot be represented by a low resolution. As seen in the bottom left of Figure 12, naive subsampling results in heavy aliasing, with a grid of additional circles being hallucinated in the output. A well-filtered downsampling result would instead retain the circle in the middle, while filtering out the high-frequency content into gray. This is observed in implementations where the filter is adjusted based on the downsampling factor – namely the PIL implementations of bicubic, lanczos, and bilinear and Tensorflow with antialias flag set as True. As before, using a fixed-width filter, as in the other rows, results in heavy aliasing.
In addition, we also show the area filter. Here, we observe a mixed results. Because implementations of the area filter do adjust to the downsampling factor across all libraries, the aliasing is not as apparent as in naive subsampling, or the fixed-width implementations of bicubic, lanczos, and bilinear. However, as described in L417 in the main paper, this particular filter corresponds to a box, or rectangular filter, which does not have strong antialiasing properties as the other filters. As a result, there are significantly more artifacts (additional hallucinated concentric circles) compared to the stronger filters (bicubic, lanczos, and bilinear) which adjust the filter widths.
In conclusion, this shows that in practical implementations, the variations in whether the filter width and the actual filter type both have an effect on the aliasing artifacts on the output.