NIMA: Neural Image Assessment
Hossein Talebi, Peyman Milanfar
I Introduction
Quantification of image quality and aesthetics have been a long-standing problem in image processing and computer vision. While technical quality assessment deals with measuring low-level degradations such as noise, blur, compression artifacts, etc., aesthetic assessment quantifies semantic level characteristics associated with emotions and beauty in images. In general, image quality assessment can be categorized into full-reference and no-reference approaches. While availability of a reference image is assumed in the former (metrics such as PSNR, SSIM, etc.), typically blind (no-reference) approaches rely on a statistical model of distortions to predict image quality. The main goal of both categories is to predict a quality score that correlates well with human perception. Yet, the subjective nature of image quality remains the fundamental issue. Recently, more complex models such as deep convolutional neural networks (CNNs) have been used to address this problem. Emergence of labeled data from human ratings has encouraged these efforts . In a typical deep CNN approach, weights are initialized by training on classification related datasets (e.g. ImageNet ), and then fine tuned on annotated data for perceptual quality assessment tasks.
Machine learning has shown promising success in predicting technical quality of images . Kang et. al. show that extracting high level features using CNNs can result in state-of-the-art blind quality assessment performance. It appears that replacing hand-crafted features with an end-to-end feature learning system is the main advantage of using CNNs for pixel-level quality assessment tasks . The proposed method in is a shallow network with one convolutional layer and two fully-connected layers, and input patches are of size . Bosse et al. use a deep CNN with 12 layers to improve on image quality predictions of . Given the small input size ( patch), both methods require score aggregation across the whole image. Bianco et al. in propose a deep quality predictor based on AlexNet . Multiple CNN features are extracted from image crops of size , and then regressed to the human scores.
Success of CNNs on object recognition tasks has significantly benefited the research on aesthetic assessment. This seems natural, as semantic level qualities are directly related to image content. Recent CNN-based methods show a significant performance improvement compared to earlier works based on hand-crafted features . Murray et al. is the benchmark on aesthetic assessment. They introduce the AVA dataset and propose a technique to use manually designed features for style classification. Later, Lu et al. show that deep CNNs are well suited to the aesthetic assessment task. Their double-column CNN consists of four convolutional and two fully-connected layers, and its inputs are the resized image and cropped windows of size . Predictions from these global and local image views are aggregated to an overall score by a fully-connected layer. Similar to Murray et al. , in images are also categorized to low and high aesthetics based on mean human ratings. A regression loss and an AlexNet inspired architecture is used in to predict the mean scores. In a similar approach to , Bin et al. fine-tune a VGG network to learn the human ratings of the AVA dataset. They use a regression framework to predict the histogram of ratings. A recent method by Zheng et al. retrains AlexNet and ResNet CNNs to predict quality of photos. More recently, uses an adaptive spatial pooling to allow for feeding multiple scales of the input image with fixed size aspect ratios to their CNN. This work presents a multi-net (each network a pre-trained VGG) approach which extracts features at multiple scales, and uses a scene aware aggregation layer to combine predictions of the sub-networks. Similarly, Ma et al. propose a layout-aware framework in which a saliency map is used to select patches with highest impact on predicted aesthetic score. Overall, none of these methods reported correlation of their predictions with respect to ground truth ratings. Recently, Kong et al. in proposed a method to aesthetically rank photos by training on AVA with a rank-based loss function. They trained an AlexNet-based CNN to learn the difference of the aesthetic scores from two input images, and as a result, indirectly optimize for rank correlation. To the best of our knowledge, is the only work that performed a correlation evaluation against AVA ratings.
I-B Our Contributions
In this work, we introduce a novel approach to predict both technical and aesthetic qualities of images. We show that models with the same CNN architecture, trained on different datasets, lead to state-of-the-art performance for both tasks. Since we aim for predictions with higher correlation with human ratings, instead of classifying images to low/high score or regressing to the mean score, the distribution of ratings are predicted as a histogram. To this end, we use the squared EMD (earth mover’s distance) loss proposed in , which shows a performance boost in classification with ordered classes. Our experiments show that this approach also leads to more accurate prediction of the mean score. Also, as shown in aesthetic assessment case , non-conventionality of images is directly related to score standard deviations. Our proposed paradigm allows for predicting this metric as well.
It has recently been shown that perceptual quality predictors can be used as learning loss to train image enhancement models . Similarly, image quality predictors can be used to adjust parameters of enhancement techniques . In this work we use our quality assessment technique to effectively tune parameters of image denoising and tone enhancement operators to produce perceptually superior results.
This paper begins with reviewing three widely used datasets for quality assessment. Then, our proposed method is explained in more detail. Finally, performance of this work is quantified and compared to the existing methods.
I-C A Large-Scale Database for Aesthetic Visual Analysis (AVA) [1]
The AVA dataset contains about 255,000 images, rated based on aesthetic qualities by amateur photographersAVA images are obtained from www.dpchallenge.com, which is an on-line community for amateur photographers.. Each photo is scored by an average of 200 people in response to photography contests. Each image is associated to a single challenge theme, with nearly 900 different contests in the AVA. The image ratings range from 1 to 10, with 10 being the highest aesthetic score associated to an image. Histograms of AVA ratings are shown in Fig. 1. As can be seen, mean ratings are concentrated around the overall mean score (5.5). Also, ratings of roughly half of the photos in AVA dataset have a standard deviation greater than 1.4. As pointed out in , presumably images with high score variance tend to be subject to interpretation, whereas images with low score variance seem to represent conventional styles or subject matter. A few examples with ratings associated with different levels of aesthetic quality and unconventionality are illustrated in Fig. 2. It seems that aesthetic quality of a photograph can be represented by the mean score, and unconventionality of it closely correlates to the score deviation. Given the distribution of AVA scores, typically, training a model on AVA data results in predictions with small deviations around the overall mean (5.5).
It is worth mentioning that the joint histogram in Fig. 1 shows higher deviations for very low/high ratings (compared to the overall mean 5.5, and mean standard deviation 1.43). In other words, divergence of opinion is more consistent in AVA images with extreme aesthetic qualities. As discussed in , distribution of ratings with mean value between 2 and 8 can be closely approximated by Gaussian functions, and highly skewed ratings can be modeled by Gamma distributions.
I-D Tampere Image Database 2013 (TID2013) [2]
TID2013 is curated for evaluation of full-reference perceptual image quality. It contains 3000 images, from 25 reference (clean) images (Kodak images ), 24 types of distortions with 5 levels for each distortion. This leads to 120 distorted images for each reference image; including different types of distortions such as compression artifacts, noise, blur and color artifacts.
Human ratings of TID2013 images are collected through a forced choice experiment, where observers select a better image between two distorted choices. Set up of the experiment allows raters to view the reference image while making a decision. In each experiment, every distorted image is used in 9 random pairwise comparisons. The selected image gets one point, and other image gets zero points. At the end of the experiment, sum of the points is used as the quality score associated with an image (this leads to scores ranging from 0 to 9). To obtain the overall mean scores, total of 985 experiments are carried out.
Mean and standard deviation of TID2013 ratings are shown in Fig. 3. As can be seen in Fig. 3(c), the mean and score deviation values are weakly correlated. A few images from TID2013 are illustrated in Fig. 4 and Fig. 5. All five levels of JPEG compression artifacts and the respective ratings are illustrated in Fig. 4. Evidently higher distortion level leads to lower mean scoreThis is a quite consistent trend for most of the other distortions too (namely noise, blur and color distortions). However, in case of the contrast change (Fig. 5), this trend is not obvious. This is due to the order of contrast compression/stretching from level 1 to level 5). Effect of contrast compression/stretching distortion on the human ratings is demonstrated in Fig. 5. Interestingly, stretch of contrast (Fig. 5(c) and Fig. 5(e)) leads to relatively higher perceptual quality.
I-E LIVE In the Wild Image Quality Challenge Database [26]
LIVE dataset contains 1162 photos captured by mobile devices. Each image is rated by an average of 175 unique subjects. Mean and standard deviation of LIVE ratings are shown in Fig. 6. As can be seen in the joint histogram, images that are rated near overall mean score show higher standard deviation. A few images from LIVE dataset are illustrated in Fig. 7. It is worth noting that in this paper, LIVE scores are scaled to .
Unlike AVA, which includes distribution of ratings for each image, TID2013 and LIVE only provide mean and standard deviation of the opinion scores. Since our proposed method requires training on score probabilities, the score distributions are approximated through maximum entropy optimization .
The rest of the paper is organized as follows. In Section II, a detailed explanation of the proposed method is described. Next, in SectionIII, applications of our algorithm in ranking photos and image enhancement are exemplified. We also provide details of our implementation. Finally, this paper is concluded in SectionIV.
II Proposed Method
Our proposed quality and aesthetic predictor stands on image classifier architectures. More explicitly, we explore a few different classifier architectures such as VGG16 , Inception-v2 , and MobileNet for image quality assessment task. VGG16 consists of 13 convolutional and 3 fully-connected layers. Small convolution filters of size are used in the deep VGG16 architecture . Inception-v2 is based on the Inception module which allows for parallel use of convolution and pooling operations. Also, in the Inception architecture, traditional fully-connected layers are replaced by average pooling, which leads to a significant reduction in number of parameters. MobileNet is an efficient deep CNN, mainly designed for mobile vision applications. In this architecture, dense convolutional filters are replaced by separable filters. This simplification results in smaller and faster CNN models.
We replaced the last layer of the baseline CNN with a fully-connected layer with 10 neurons followed by soft-max activations (shown in Fig. 8). Baseline CNN weights are initialized by training on the ImageNet dataset , and then an end-to-end training on quality assessment is performed. In this paper, we discuss performance of the proposed model with various baseline CNNs.
In training, input images are rescaled to , and then a crop of size is randomly extracted. This lessens potential over-fitting issues, especially when training on relatively small datasets (e.g. TID2013). It is worth noting that we also tried training with random crops without rescaling. However, results were not compelling. This is due to the inevitable change in image composition. Another random data augmentation in our training process is horizontal flipping of the image crops.
Our goal is to predict the distribution of ratings for a given image. Ground truth distribution of human ratings of a given image can be expressed as an empirical probability mass function with , where denotes the th score bucket, and denotes the total number of score buckets. In both AVA and TID2013 datasets , in AVA, and , and in TID and . Since , represents the probability of a quality score falling in the th bucket. Given the distribution of ratings as p, mean quality score is defined as , and standard deviation of the score is computed as . As discussed in the previous section, one can qualitatively compare images by mean and standard deviation of scores.
Each example in the dataset consists of an image and its ground truth (user) ratings p. Our objective is to find the probability mass function that is an accurate estimate of p. Next, our training loss function is discussed.
Soft-max cross-entropy is widely used as training loss in classification tasks. This loss can be represented as (where denotes estimated probability of th score bucket) to maximize predicted probability of the correct labels. However, in the case of ordered-classes (e.g. aesthetic and quality estimation), cross-entropy loss lacks the inter-class relationships between score buckets. One might argue that ordered-classes can be represented by a real number, and consequently, can be learned through a regression framework. Yet, it has been shown that for ordered classes, the classification frameworks can outperform regression models . Hou et al. show that training on datasets with intrinsic ordering between classes can benefit from EMD-based losses. These loss functions penalize mis-classifications according to class distances.
For image quality ratings, classes are inherently ordered as , and norm distance between classes is defined as , where . EMD is defined as the minimum cost to move the mass of one distribution to another. Given the ground truth and estimated probability mass functions p and , with ordered classes of distance , the normalized Earth Mover’s Distance can be expressed as :
where is the cumulative distribution function as . It is worth noting that this closed-form solution requires both distributions to have equal mass as . As shown in Fig. 8, our predicted quality probabilities are fed to a soft-max function to guarantee that . Similar to , in our training framework, is set as 2 to penalize the Euclidean distance between the CDFs. allows easier optimization when working with gradient descent.
III Experimental Results
We train two separate models for aesthetics and technical quality assessment on AVA, TID2013, and LIVE. For each case, we split each dataset into train and test sets, such that 20% of the data is used for testing. In this section, performance of the proposed models on the test sets are discussed and compared to the existing methods. Then, applications of the proposed technique in photo ranking and image enhancement are explored. Before moving forward, details of our implementation are explained.
The CNNs presented in this paper are implemented using TensorFlow . The baseline CNN weights are initialized by training on ImageNet , and the last fully-connected layer is randomly initialized. The weight and bias momentums are set to 0.9, and a dropout rate of 0.75 is applied on the last layer of the baseline network. The learning rate of the baseline CNN layers and the last fully-connected layers are set as and , respectively. We observed that setting a low learning rate on baseline CNN layers results in easier and faster optimization when using stochastic gradient descent. Also, after every 10 epochs of training, an exponential decay with decay factor 0.95 is applied to all learning rates.
Accuracy, correlation and EMD values of our evaluations on the aesthetic assessment model on AVA are presented in Table I. Most methods in Table I are designed to perform binary classification on the aesthetic scores, and as a result, only accuracy evaluations of two-class quality categorization are reported. In this binary classification, predicted mean scores are compared to as cut-off score. Images with predicted scores above the cut-off score are categorized as high quality. In two-class aesthetic categorization task, results from , and NIMA(Inception-v2) show the highest accuracy. Also, in terms of rank correlation, NIMA(VGG16) and NIMA(Inception-v2) outperform . NIMA is much cheaper: applies multiple VGG16 nets on image patches to generate a single quality score, whereas computational complexity of NIMA(Inception-v2) is roughly one pass of Inception-v2 (see Table V).
Our technical quality assessment model on TID2013 is compared to other existing methods in Table II. While most of these methods regress to the mean opinion score, our proposed technique predicts the distribution of ratings, as well as mean opinion score. Correlation between ground truth and results of NIMA(VGG16) are close to the state-of-the-art results in and . It is worth highlighting that Bianco et al. feed multiple image crops to a deep CNN, whereas our method takes only the rescaled image.
The predicted distributions of AVA scores are presented in Fig. 9. We used NIMA(Inception-v2) model to predict the ground truth scores from our AVA test set. As can be seen, distribution of the ground truth mean scores is closely predicted by NIMA. However, predicting distribution of the ground truth standard deviations is a more challenging task. As we discussed previously, unconventionality of subject matter or style has a direct impact on score standard deviations.
III-B Cross Dataset Evaluation
As a cross validation test, performance of our trained models are measured on other datasets. These results are presented in Table III and Table IV. We test NIMA(Inception-v2) model trained on AVA, TID2013 and LIVE across all three test sets. As can be seen, on average, training on AVA dataset shows the best performance. For instance, training on AVA and testing on LIVE results in and linear and rank correlations, respectively. However, training on LIVE and testing on AVA leads to and linear and rank correlation coefficients. We believe this observation shows that NIMA models trained on AVA can generalize to other test examples more effectively, whereas training on TID2013 results in poor performance on LIVE and AVA test sets. It is worth mentioning that AVA dataset contains roughly 250 times more examples (in comparison to the LIVE dataset), which allows training NIMA models without any significant overfitting.
III-C Photo Ranking
Predicted mean scores can be used to rank photos, aesthetically. Some test photos from AVA dataset are ranked in Fig. 10 and Fig. 11. Predicted NIMA scores and ground truth AVA scores are shown below each image. Results in Fig. 10 suggest that in addition to image content, other factors such as tone, contrast and composition of photos are important aesthetic qualities. Also, as shown in Fig. 11, besides image semantics, framing and color palette are key qualities in these photos. These aesthetic attributes are closely predicted by our trained models on AVA.
Predicted mean scores are used to qualitatively rank photos in Fig. 12. These images are part of our TID2013 test set, which contain various types and levels of distortions. Comparing ground truth and predicted scores indicates that our trained model on TID2013 accurately ranks the test images.
III-D Image Enhancement
Quality and aesthetic scores can be used to perceptually tune image enhancement operators. In other words, maximizing NIMA score as a prior can increase the likelihood of enhancing perceptual quality of an image. Typically, parameters of enhancement operators such as image denoising and contrast enhancement are selected by extensive experiments under various photographic conditions. Perceptual tuning could be quite expensive and time consuming, especially when human opinion is required. In this section, our proposed models are used to tune a tone enhancement method , and an image denoiser . A more detailed treatment is presented in .
The multi-layer Laplacian technique enhances local and global contrast of images. Parameters of this method control the amount of detail, shadow, and brightness of an image. Fig. 13 shows a few examples of the multi-layer Laplacian with different sets of parameters. We observed that the predicted aesthetic ratings from training on the AVA dataset can be improved by contrast adjustments. Consequently, our model is able to guide the multi-layer Laplacian filter to find aesthetically near-optimal settings of its parameters. Examples of this type of image editing are represented in Fig. 14, where a combination of detail, shadow and brightness change is applied on each image. In each example, 6 levels of detail boost, 11 levels of shadow change, and 11 levels of brightness change account for a total of 726 variations. The aesthetic assessment model tends to prefer high contrast images with boosted details. This is consistent with the ground truth results from AVA illustrated in Fig. 10.
Turbo denoising is a technique which uses the domain transform as its core filter. Performance of Turbo denoising depends on spatial and range smoothing parameters, and consequently, proper tuning of these parameters can effectively boost performance of the denoiser. We observed that varying the spatial smoothing parameter makes the most significant perceptual difference, and as a result, we use our quality assessment model trained on TID2013 dataset to tune this denoiser. Application of our no-reference quality metric as a prior in image denoising is similar to the work of Zhu et al. . Our results are shown in Fig. 15. Additive white Gaussian noise with standard deviation 30 is added to the clean image, and Turbo denoising with various spatial parameters is used to denoise the noisy image. To reduce the score deviation, 50 random crops are extracted from denoised image. These scores are averaged to obtain the plots illustrated in Fig. 15. As can be seen, although the same amount of noise is added to each image, maximum quality scores correspond to different denoising parameters in each example. For relatively smooth images such as (a) and (g), optimal spatial parameter of Turbo denoising is higher (which implies stronger smoothing) than the textured image in (j). This is probably due to the relatively high signal-to-noise ratio of (j). In other words, the quality assessment model tends to respect textures and avoid over-smoothing of details. Effect of the denoising parameter can be visually inspected in Fig. 16. While the denoised result in Fig. 16 (a) is under-smoothed, (c), (e) and (f) show undesirable over-smoothing effects. The predicted quality scores validate this perceptual observation.
III-E Computational Costs
Computational complexity of NIMA models are compared in Table V. Our inference TensorFlow implementation is tested on an Intel Xeon CPU @ 3.5 GHz with 32 GB memory and 12 cores, and NVIDIA Quadro K620 GPU. Timings of one pass of NIMA models on an image of size are reported in Table V. Evidently, NIMA(MobileNet) is significantly lighter and faster than other models. This comes at the expense of a slight performance drop (shown in Table I and Table II).
IV Conclusion
In this work we introduced a CNN-based image assessment method, which can be trained on both aesthetic and pixel-level quality datasets. Our models effectively predict the distribution of quality ratings, rather than just the mean scores. This leads to a more accurate quality prediction with higher correlation to the ground truth ratings. We trained two models for high level aesthetics and low level technical qualities, and utilized them to steer parameters of a few image enhancement operators. Our experiments suggest that these models are capable of guiding denoising and tone enhancement to produce perceptually superior results.
As part of our future work, we will exploit the trained models on other image enhancement applications. Our current experimental setup requires the enhancement operator to be evaluated multiple times. This limits real-time application of the proposed method. One might argue that in case of an enhancement operator with well-defined derivatives, using NIMA as the loss function is a more efficient approach.
Acknowledgment
We would like to thank Dr. Pascal Getreuer for valuable discussions and helpful advice on approximation of score distributions.