Fine-grained Recognition with Learnable Semantic Data Augmentation
Yifan Pu, Yizeng Han, Yulin Wang, Junlan Feng, Chao Deng, Gao Huang
I Introduction
Fine-grained image recognition aims to distinguish objects with subtle differences in visual appearance within the same general category, e.g., different species of animals , different models of aircraft , different kinds of retail products . The key challenge therefore lies in comprehending fine-grained visual differences that sufficiently discriminate between objects that are highly similar in overall appearance but differ in subtle traits . In recent years, deep learning has emerged as a powerful tool to learn discriminative image representations and has achieved great success in the field of fine-grained visual recognition .
As deep neural networks dominate the field of visual object recognition , data augmentation techniques further boost the generalization ability of neural networks in the generic image classification scenario. Popular data augmentation techniques, e.g., Mixup , CutMix , RandAugment, and Random Erasing , have become a standard recipe in training modern convolution networks and vision transformers . However, in the scenario of fine-grained image recognition, these data augmentation techniques are rarely applied because of the discriminative region loss problem. Specifically, as illustrated in Figure 1, the random editing behavior of image-level data augmentation approaches have a risk of destroying discriminative regions, which is of great significance in performing fine-grained recognition. For example, random drop based data augmentations, such as Cutout , Random Erasing , have the potential to drop the discriminative regions of fine-grained objects. Geometric transformation, which is contained in AutoAugment and RandAugment , is also likely to cause a loss to discriminative visual cues. Besides, mix-based techniques (e.g. Mixup , CutMix ) would even result in a noisy label problem by replacing discriminative regions of the current image with those of another class. Figure 2 (a) further shows such limitations of the image-level augmentation techniques in the feature space: the deep feature of an image-level augmented image would probably distribute on the classification boundary or even intrude into the feature space of another category.
To cope with the aforementioned problem, we propose to augment the training samples at the feature level (Figure 2 (b)) rather than the image level. By translating data samples in the deep feature space along their corresponding meaningful semantic directions, the feature-level data augmentation would produce diversified augmented image features, which correspond to semantic meaningful images in the pixel space. In this way, the implicit data augmentation method alleviates the discriminative region loss problem caused by the random editing manner of image-level data augmentation techniques.
The performance of implicit data augmentation techniques heavily relies on the quality of the semantic directions. The existing feature-level data augmentation method (implicit semantic data augmentation, ISDA ) verifies that, in the generic image classification problem, a global set of semantic directions shared by all classes is inferior to maintaining a number of sets of semantic directions for each category. Therefore, ISDA takes the class-conditional covariance matrices of deep features as the candidate semantic directions of each category and estimates the covariance matrices statistically in an online manner. However, in the fine-grained scenario, the limitations of the augmentation approach in are two-fold: 1) the online estimation strategy is sub-optimal due to the limited amount of training data in each class; 2) the class-conditional semantic directions are not suitable in the fine-grained recognition for its large intra-class variation and small inter-class variation . Specifically, the meaningful semantic directions of the image samples within the same sub-category vary because of the large intra-class variance. For example, in Figure 3 the landed birds have some meaningful semantic directions that flying birds do not have, and vice versa. As a result, augmenting all samples within one sub-category along the same set of semantic directions is improper. Intuitively, diversifying different training samples along their corresponding semantic directions is preferable.
In this paper, we propose a learnable semantic data augmentation method for the fine-grained image recognition problem. The meaningful semantic directions is automatically learned in a sample-wise manner based on a covariance matrix prediction network (CovNet) rather than estimated class-conditionally with statistical method . The CovNet takes in the deep features of each training sample and predicts their meaningful semantic directions. The covariance matrix prediction network and the classification network are jointly trained in a meta-learning manner. The meta-learning framework optimizes the two networks in an alternate way with different objectives, which solves the degeneration problem when optimizing them jointly (see the theoretical and empirical analysis in Sec. III-B and Sec. IV-D, respectively). Compared with the online estimation approach , our proposed sample-wise prediction method can effectively produce appropriate semantic directions for each training sample, and therefore boost the network performance in the fine-grained image classification task.
We evaluate our method on four popular fine-grained image recognition benchmarks (i.e. CUB-200-2011 , Stanford Cars , FGVC Aircrarts and NABirds ). Experimental results show that our approach effectively enhances the intra-class compactness of learned features and significantly improves the performance of mainstream classification networks (e.g. ResNet , DenseNet , EfficientNet , RegNet and ViT ) on various fine-grained classification datasets. Combined with a recent proposed fine-grained recogntion method (P2P-Net ), the proposed method achieves state-of-the-art performance on CUB-200-2011.
II Related work
Fine-grained recognition. Fine-grained image recognition aims to discriminate numerous visually similar subordinate categories that belong to the same basic category . Recognizing fine-grained categories is difficult due to the challenges of discriminative region localization and fine-grained feature learning. Broadly, existing fine-grained recognition approaches can be divided into two main paradigms: localization methods and feature encoding methods. The former usually creates models that capture the discriminative semantic parts of fine-grained objects and then construct a mid-level representation corresponding to these parts for the final classification. Common methods could be divided as employ detection or segmentation techniques , utilize deep filters and leverage attention mechanisms . The feature encoding methods aim to learn a unified, yet discriminative, image representation for modeling subtle differences between fine-grained categories. Common practices include performing high-order feature interactions and designing novel loss functions . However, because image-level data augmentation techniques tend to be harmful to discriminative subtle regions of fine-grained images, the designing of data augmentation strategies for fine-grained vision tasks is rarely explored.
Data Augmentation. In recent years, data augmentation techniques have been widely used in training deep neural networks . Basic data augmentation, such as rotation, translation, cropping, flipping , are commonly used to increase the diversity of training samples. Beyond these, Cutout , Mixup and CutMix are manually design with domain knowledge. Recently, inspired by the neural architecture search algorithms, some works attempt to automate learning data augmentation policies, such as AutoAugment and RandAugment . Although these data augmentation methods are commonly used and some of them have even become a routine in training deep neural networks in the generic classification problem, they are rarely adopted in the fine-grained scenario because random crop, mix and deform operations in the image level would easily destroy the discriminative information of fine-grained objects. Inspired by a recently proposed technique , recent studies have achieved significant success by applying feature-level data augmentation to augment minority classes in the long-tailed recognition problem (MetaSAug ) and enhancing classifier adaptability in domain adaptation through the generation of source features aligned with target semantics (TSA ). These methods demonstrate the ability to achieve state-of-the-art performance in their problems. In contrast, this paper focus on the longstanding fine-grained recognition problem and proposes an important extension of ISDA , aiming to address the discriminative region loss problem through the augmentation of training data at the feature level.
Meta-learning. The field of meta-learning has seen a dramatic rise in interest in recent years . Contrary to the conventional deep learning approach which solves the optimization problem with a fixed learning algorithm, meta-learning aims to improve the learning algorithm itself. Under the meta-learning framework, a machine learning model could gain experience over multiple learning episodes, and uses this experience to improve its future learning performance. Meta-learning has proven useful both in multi-task scenarios where task-agnostic knowledge is extracted from a family of tasks and used to improve the learning of new tasks from that family , and in single-task scenarios where a single problem is solved repeatedly and improved over multiple episodes . Successful applications have been demonstrated in areas spanning few-shot image recognition , unsupervised learning , data-efficient and self-directed reinforcement learning, hyperparameter optimization , domain generalizable person ReID, and neural architecture search . In this paper, we design a single-task meta-learning algorithm, which learns the meaningful semantic directions for each training sample during the classification model training process.
III Method
In this section, we first introduce the preliminaries of our method, implicit semantic data augmentation (ISDA) and its online estimation algorithm for the covariance matrices. Then our meta-learning-based framework will be presented. We further present the convergence proof of our proposed meta-learning algorithm. For better readability, we list the notations used in this paper in Table I.
Most conventional data augmentation methods make modifications directly on training images. In contrast, ISDA performs data augmentation at the feature level, i.e., translating image features along meaningful semantic directions. Such directions are determined based on the covariance matrices of deep features. Specifically, for a -class classification problem, ISDA in statistically estimates the class-wise covariance matrices in an online manner at each training iteration. For the -th sample with ground truth , ISDA randomly samples transformation directions from the Gaussian distribution to augment the deep feature , where is the learned feature to be fed to the last fully-connected layer in a deep network, and is a hyperparameter controlling the augmentation strength. One should sample a large number () of directions in to get the sufficiently augmented features . The modified cross-entropy loss on these augmented data can be written as
where and are the trainable parameters of the last fully connected layer, is the number of sampled directions. Take a step further, if infinite directions are sampled, ISDA derives the upper bound of the expected cross-entropy loss on all augmented features:
III-B Covariance Matrix Prediction Network
The performance of the aforementioned ISDA heavily relies on the estimated class-wise covariance matrices, which directly affect the quality of semantic directions. In the fine-grained visual recognition scenario, the limited amount of training data, along with its large intra-class variance and small inter-class variance characteristic, poses great challenges on the covariance estimation. The unsatisfying covariance matrices can further cause limited improvement in model performance. To this end, we propose to automatically learn sample-wise semantic directions based on a covariance matrix prediction network instead of statistically estimating the class-wise covariance matrices as in the existing method .
Our covariance matrix prediction network (CovNet) is established as a multilayer perception (MLP), denoted as parameterized by . The CovNet take the deep feature as input, and predicts its sample-wise semantic directions: As a result, the ISDA loss function with our predicted covariance matrices can be rewritten as
where is the deep feature extracted by , and refers to the parameters of the classification head (fully-connected layer). In practice, due to GPU memory limitation, the CovNet only predicts the diagonal elements of the covariance matrices and we set all other elements as zero following . We use the Sigmoid activation in the last layer of to ensure the produced covariance matrices are positive definite.
Therefore, the naive joint training strategy would trivially encourage the CovNet to produce a zero-valued matrix , and limited improvement for the classification accuracy will be obtained (see the results in Section IV-D).
III-C The Meta-learning Method
We propose a meta-learning-based approach to deal with the degenerate solution problem illustrated in Eq. (4). By setting a meta-learning objective and optimizing with meta-gradient , the CovNet could learn to produce appropriate covariance matrices by mining the meta knowledge from metadata. In this subsection, we first formulate the forward pass and the training objectives of the classification network and the CovNet , respectively. Then the detailed optimization pipeline of the two networks is presented.
where is the number of training data. The optimization target of the classification network is minimizing the training loss under the learned covariance matrices
It can be observed from Eq. (6) that the optimized is a function of the parameter of the CovNet .
The network parameter of the CovNet is learned in a meta-learning approach. We use the meta data from the meta dataset to optimize the CovNet. The meta loss is the cross-entropy loss between the network prediction and corresponding ground truth
Since the CovNet is optimized by backpropagating the meta-gradient through the gradient operator in the meta-objective and the meta-objective is established with cross-entropy loss rather than ISDA loss, the degeneration problem (illustrated in Eq. (4)) when training an in one gradient step with the same optimization objective will not happen.
III-C2 The optimization pipeline
Starting from a initial state of the -th iteration, the classification network take a pseudo update step under the current CovNet parameter using training data
The updated parameter by gradient descent is
where is the learning rate of the classification network at the current time step . Note that the pseudo update process only update the classification network. The gradients would not pass backward to the CovNet. See Figure 4 (grey dashed lines) for the gradient flow of this pseudo update.
The meta update process updates the CovNet using the meta data. By minimizing the meta-objective over the metadata
we could get the updated CovNet parameter
where is the current learning rate of . The meta knowledge contained in the metadata helps the CovNet learn the appropriate semantic directions for each samples in the deep feature space. The red dashed lines in Figure 4 illustrate the gradient flow for updating the parameter of CovNet .
The real update process updates the classification network based on the prediction of the updated CovNet . The updated CovNet helps to predict the semantic directions of each training sample and facilitate the optimization procedure of adopting feature-level data augmentation
Finally, we could get the updated classification network parameter by taking a optimization step
III-D Convergence of The Meta-learning Algorithm
The proposed meta-learning-based algorithm involves a bi-level optimization procedure. We proved that our method converges to some critical points for both the training loss and the meta loss under some mild conditions. Specifically, we proved that the expectation of the meta loss gradient would be smaller than an infinitely small quantity in finite step, and the expectation of the gradient of the training loss will converge to zero. The theorems is demonstrated as the following. The complete proof is present in the Appendix Fine-grained Recognition with Learnable Semantic Data Augmentation.
1) the proposed algorithm can always achieve
III-E Accelerating the meta-learning framework
The cost of the meta update process is relatively high because updating requires computing second-order gradients (see Eq. (12)). In order to make the training procedure more efficient, we adopt an approximate method to accelerate the meta update process by freezing part of the classification network. In the pseudo update process, we freeze the first several blocks and only the late blocks will have gradients. Consequently, during the meta update process, the meta gradient will be computed solely from the gradient of a subset of the pseudo network. In this way, the algorithm significantly improves the training efficiency without sacrificing model performance. The detailed analysis is presented in Sec. IV-D4.
IV Experiments
In section, we empirically evaluate our semantic data augmentation method on different fine-grained visual recognition datasets. We first introduce the detailed experiment setup in Section IV-A, including datasets and training configurations. Then the main results of our method with various backbone architectures on different datasets are presented in Section IV-B. We also conduct comparison experiments with competing approaches in Section IV-C. Finally, the ablation study (Section IV-D), the experiment results on the general recognition task (Section VII), and the visualization results (Section IV-F) further validate the effectiveness of the proposed method.
Datasets. We evaluate our proposed methods on four widely used fine-grained benchmarks, i.e, CUB-200-2011 , FGVC Aircraft , Stanford Cars and NABirds . The CUB-200-2011 is the most wildly-used dataset for fine-grained visual categorization tasks, which contains 11,788 photographs of 200 subcategories belonging to birds, 5,994 for training and 5,794 for testing. The FGVC Aircraft dataset contains 10,200 images of aircraft, with 100 images for each of 102 different aircraft model variants, most of which are airplanes. The Stanford Cars dataset contains 16,185 images of 196 classes of cars and is split into 8,144 training images and 8,041 testing images. The NABirds dataset is a collection of 48,000 annotated photographs of the 400 species of birds that are commonly observed in North America. In our experiments, we only use the category label as supervision, although some datasets have additional available annotations.
For all the fine-grained datasets, we first resize the image to 600 600 pixels and crop it into 448 448 resolution (random cropping for training and center crop for testing), with a random horizontal flipping operation following behind.
Implementation Details For all the classification network structures, we load the pretrained network from the torchvision library, except that the ViT pretrained weight is taken following TransFG . We use stochastic gradient descent (SGD) optimizer to train the classification network , with a momentum of 0.9 and a weight decay of 0.0. The learning rate of the classification network is initialized as 0.03 with a batch size of 64, decaying with a cosine shape. All the models are trained for 100 epochs except that the combination experiment with P2P-Net in Sec. IV-C follows the setting in .
The covariance prediction network is established as a multilayer perceptron with one hidden layer. The width of the hidden layer is set as a quarter of the final feature dimension. We also provide the ablation study on the structure of the CovNet in Sec. IV-D5. Due to GPU memory limitation on high resolution images, following the practice of the original implicit data augmentation method , we approximate the covariance matrices by their diagonals, i.e., the variance of each dimension of the features. We also adopt the SGD optimizer with a learning rate of 0.001 to optimize . The CovNet is only used for training, and no extra computation cost will be brought during the inference procedure.
IV-B Effectiveness with different datasets and architectures
We compare our proposed feature-level data augmentation method with counterparts that only use basic data augmentation (i.e. random horizontal flipping) on ResNet-50. The experiment results on the above-mentioned fine-grained datasets are shown in Table II. From the results, we find that our method significantly improves the generalization ability of classification networks in various fine-grained scenarios, including birds , aircraft and cars . To be specific, on the most popular fine-grained dataset CUB-200-2011, the proposed method achieves 2.2% improvement on the Top-1 Accuracy compared to the baseline counterpart. Our method also earns more than 1.2% Top-1 accuracy improvement on other popular fine-grained datasets. Compared with the semantic data augmentation with class-wise estimated semantic directions , our sample-wise prediction approach shows its superiority in improving network performance.
We also apply our method on various popular classification network architectures, i.e, DenseNet , EfficientNet , MobileNetV2 , RegNet , Vision Transfromer (ViT) . The results in Table III show that our method could improve the performance of fine-grained classification accuracy among various neural network architectures. Specifically, our method could improve the performance of networks designed for server (ResNet, DenseNet and EfficientNet) for more than 1.0% Top-1 Accuracy. For mobile devices intended networks, our method could also get remarkable improvement (1.7% accuracy improvement for MobileNetV2, 1.2% for RegNetX-400MF). Our method is also effective on recent proposed vision transformer architectures (0.4% accuracy improvement for ViT-B_16). The results also show that, in the fine-grained scenario, sample-wise predicted semantic directions method is more effective than the class-wise estimated counterparts among various neural network architectures.
IV-C Comparisons with Competing Methods
Comparison and compatibility with state-of-the-arts. We combine our proposed sample-wise predicted semantic data augmentation method with a recent proposed fine-grained recognition technique P2P-Net . In addition to the supervision from image label, the P2P-Net further utilizes the localized discriminative parts to promote discrimination of image representations and learns a graph matching for part alignment in order to alleviate the variation of object pose. The P2P-Net achieves state-of-the-art performance on the CUB-200-2011 dataset using the ResNet as base model.
In the P2P-Net , there are totally classifiers, where the first classifiers are attached to intermediate feature maps of the base model, the classifier takes the final classification feature as input, and the last one make prediction depending on the combination of all the features mentioned previous. For simplicity, we only combine our method with the cross-entropy loss of the final classifier and keep other settings the same as P2P-Net . We set the augmentation strength as 5.0 and construct the CovNet as a MLP with one hidden layer, whose size is a half of the deep feature size. The meta update process in Eq. (12) only occurs every ten iterations for training efficiency consideration.
Experimental results in Table. IV show that combined with our proposed method, the P2P-Net achieves a 91.0% Top-1 Accuracy on CUB-200-2011, which is 0.8 higher than the original P2P-Net. This result verifies the effectiveness when combining our method with other competing methods.
Superiority over online estimated covariance. We compare our meta-learning-based sample-wise covariance matrix prediction method with the class-wise online estimation technique proposed in ISDA . When combining with P2P-Net , ours method shows a superior performance comparing with the online estimation technique in ISDA (shown in Table. IV). We further conduct experiments on ResNets with different depths on CUB-200-2011, and the results are shown in Figure 6. We could observe from the results that although ISDA in could improve the generalization ability over the baseline method, our proposed strategy could further surpass the it by a large margin. This phenomenon verifies the superiority of our sample-wise covariance matrix prediction over ISDA with online estimated covariance in the fine-grained scenario.
Comparison with image-level data augmentation. In Table V, we compare the performance of our method and some popular image-level data augmentation methods . It can be found that the proposed method is more effective than image-level counterparts in the fine-grained scenario.
IV-D Ablation Studies
We conduct ablation studies on our method to analyze how its variants affect the fine-grained visual classification result. We first show how optimizing the covariance matrix prediction network and the classification network simultaneously (mentioned in Section III-B) would result in a zero-valued output. Then, we ablate the strength of augmentation , the growth scheduling of and the structure of the proposed CovNet. Finally, the effect of accelerating the meta update process by freezing part of the pseudo network is presented.
We train the and the with the same target as described in Section III-B and tuning the learning rate of as a proportion of the learning rate of . The experiments are conducted on CUB-200-2011 dataset with a ResNet-50 classification network. In Figure 7(a), we show the result from two perspectives: the mean of the CovNet output at the end of each epoch and the corresponding final accuracy. The mean of the CovNet output will converge to zero quickly if the learning rate is large and have a tendency to zero under a small learning rate, which verifies our claim in Section III-B. Although the final classification accuracy is slightly higher than that of the baseline training method, it is inferior to ISDA and remarkably worse than that of our proposed sample-wise predicted semantic data augmentation method.
IV-D2 Influence of the strength of augmentation λ𝜆\lambda
In our work, the strength of augmentation is linear increases along with the training epoch , where is the augmentation strength of the current epoch, is the index of the current epoch, is the total training epoch and is a hyperparameter to control the overall augmentation strength. We ablate the strength of semantic data augmentation by conducting experiments on CUB-200-2011 with three different models (i.e., ResNet-50, DenseNet-161 and MobileNetV2). As shown in Fig 7(b), as increase from zero, the accuracy gradually increase and reach its peak around . As a result, in our experiments we simply set as 10.0 expect for the P2P-Net. Despite the performance of the proposed method is affected by the augmentation strength , our method always outperforms the baseline ( in Figure 7(b)).
IV-D3 Influence of growth scheduling of λ𝜆\lambda
Considering the simple linear growth strategy of may not be optimal, we ablate the growth scheduling of by tuning the hyperparameter in the scheduling function . The scheduling of different is illustrated in Figure 11 (a) and Figure 11 (b) shows the corresponding results demonstrated on a ResNet-50 model with CUB-200-2011. The results show that the simple linear increasing schedule is superior to convex or concave growing counterparts. As a result, we choose the linear increasing schedule for all the experiments.
IV-D4 Impact of the frozen blocks
The meta-learning framework introduces extra training cost because it has two extra update steps. The pseudo update process takes about the same time as the final real update process, while the meta update process takes more time because it needs to compute second order gradient. In a 4 Nvidia V100 GPU server, the vanilla meta-learning algorithm take 6.5 hour to train a ResNet-50 model in CUB-200-2011 for 100 epoch, while the baseline method takes 1.7 hour. We freeze the ResNet-50 network stem and the first residual blocks and record the corresponding accuracy and training time on the CUB-200-2011 dataset. The results in Figure 12 show that the network performance is similar to the counterpart without freezing when , while the total training time is monotonically decreasing. As a result, we freeze the first ten residual blocks of ResNet-50 for all the fine-grained recognition experiments to get a good accuracy-efficiency trade-off. In this way, the training cost in CUB-200-2011 is reduced from 6.5 hour to 4.5 hour. It is worth mention that our method do not add any extra cost in inference. For other neural network structures, including P2P-Net , we do not freeze any part of the network structure because finding the proper number of frozen blocks need extra experiments.
IV-D5 Impact of the CovNet Structure
We ablate the structure of the covariance matrix prediction network by varying its depth (the number of hidden layers) and width (the number of neurons of the hidden layer). Experimental results on CUB-200-2011 dataset with ResNet-50, which is shown in Table VI, demonstrate that although different CovNet structures could affect the results, the whole method is still effective. In practice we set the depth as one and the width of the hidden layer as a quarter of the final feature dimension for all the model architectures except otherwise mentioned.
IV-E Evaluation on Generic Recognition Benchmark
As the proposed sample-wise feature-level augmentation approach does not rely on additional fine-grained annotations (such as bounding boxes, part annotations, and hierarchical labels), it can be easily adapted to general image classification. We evaluate our proposed method on ImageNet with a ResNet-50 model. For a fair comparison, we keep the same training configuration as ISDA , in which the model is trained from scratch for 120 epochs with a momentum of 0.9 and a weight decay of 0.0001. The augmentation strength is set as 7.5, which is also the same as that in ISDA . The CovNet has the same architecture configuration as in the CUB-200-2011 and is optimized every 100 iterations with a learning rate of 0.0005. The results in Table VII show that, on the generic image recognition problem, our sample-wise semantic data augmentation method is also effective over the class-wise estimated semantic directions approach .
IV-F Visualization Results
Augmented samples in the pixel space. To demonstrate that our method is able to generate meaningful semantically augmented samples, we present the Top-5 nearest neighbors of the augmented features in image space on four fine-grained datasets (CUB-200-2011, FGVC Aircraft, Stanford Cars and NABirds). As shown in Figure 8, our feature-level data augmentation strategy is able to adjust the semantics of training sample, such as visual angels, background, pose of the birds, painting of the aircraft, color of the cars.
Quality of learned feature. We compare the learned feature quality between training with basic data augmentation (baseline) and training with our method. We select the meta-classes, which contains more than five sub-classes, on CUB-200-2011. We extract the deep features using the pre-trained network learned with the baseline method and our feature-level data augmentation method, respectively. These high-dimensional deep features are downscaled using t-SNE , and results are illustrated in Figure 9. We can find that our method effectively enhances the intra-class compactnesss and the inter-class separability of the learned features.
Comparison with image-level data augmentation. Finally, we visualize the feature of different data augmentation methods, including Random Erasing , RandAugment , and our method. For fair comparison we use the same feature extractor pre-trained with basic augmentation. For image-level augmentations (Random Erasing and RandAugment), we extract the feature of original images and that of the augmented images and reduce its dimension by t-SNE. For our feature-level data augmentation method, we extract the feature of original images, diversify them with a pre-trained CovNet, then downscale both of them using t-SNE. The result in Figure 10 (a) (b) reveals that a portion of the image-level augmented samples would loss its discriminative region and thus are clustered together to form a new clustering center. This phenomenon verifies the discriminative region loss problem we proposed in Figure 1. Furthermore, our feature-level data augmentation method always produce appropriate semantic directions and help the training samples to be translated into reasonable locations. Our method provide an ingenious solution to avoid the discriminative region loss problem induced by image-level data augmentation methods on fine-grained images.
V Conclusion
In this paper, we propose a meta-learning based implicit data augmentation method for fine-grained image recognition. Our approach aims to cope with the discriminative region loss problem in the fine-grained scenario, which is induced by the random editing behavior of image-level data augmentation techniques. We diversify the training samples in the feature space rather than the image space to alleviate this problem. The sample-wise meaningful semantic direction is predicted by the covariance prediction network, which is joint optimized with the classification network in a meta-learning manner. Experiment results over multiple fine-grained benchmarks and neural network structures show the effectiveness of our proposed method on the fine-grained recognition problem.
Acknowledgement. This work is supported in part by the National Key R&D Program of China under Grant 2021ZD0140407, the National Natural Science Foundation of China under Grants 62022048 and 62276150, Guoqiang Institute of Tsinghua University and Beijing Academy of Artificial Intelligence.