InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions
Wenhai Wang, Jifeng Dai, Zhe Chen, Zhenhang Huang, Zhiqi Li, Xizhou Zhu, Xiaowei Hu, Tong Lu, Lewei Lu, Hongsheng Li, Xiaogang Wang, Yu Qiao
Introduction
With the remarkable success of transformers in large-scale language models , vision transformers (ViTs) have also swept the computer vision field and are becoming the primary choice for the research and practice of large-scale vision foundation models. Some pioneers have made attempts to extend ViTs to very large models with over a billion parameters, beating convolutional neural networks (CNNs) and significantly pushing the performance bound for a wide range of computer vision tasks, including basic classification, detection, and segmentation. While these results suggest that CNNs are inferior to ViTs in the era of massive parameters and data, we argue that CNN-based foundation models can also achieve comparable or even better performance than ViTs when equipped with similar operator-/architecture-level designs, scaling-up parameters, and massive data.
To bridge the gap between CNNs and ViTs, we first summarize their differences from two aspects: (1) From the operator level , the multi-head self-attention (MHSA) of ViTs has long-range dependencies and adaptive spatial aggregation (see Fig. 1(a)). Benefiting from the flexible MHSA, ViTs can learn more powerful and robust representations than CNNs from massive data. (2) From the architecture view , besides MHSA, ViTs contain a series of advanced components that are not included in standard CNNs, such as Layer Normalization (LN) , feed-forward network (FFN) , GELU , etc. Although recent works have made meaningful attempts to introduce long-range dependencies into CNNs by using dense convolutions with very large kernels (e.g., 3131) as shown in Fig. 1 (c), there is still a considerable gap with the state-of-the-art large-scale ViTs in terms of performance and model scale.
In this work, we concentrate on designing a CNN-based foundation model that can efficiently extend to large-scale parameters and data. Specifically, we start with a flexible convolution variant—deformable convolution (DCN) . By combining it with a series of tailored block-level and architecture-level designs similar to transformers, we design a brand-new convolutional backbone network, termed InternImage. As shown in Fig. 1, different from recently improved CNNs with very large kernels such as 3131 , the core operator of InternImage is a dynamic sparse convolution with a common window size of 33, (1) whose sampling offsets are flexible to dynamically learn appropriate receptive fields (can be long- or short-range) from given data; (2) the sampling offsets and modulation scalars are adaptively adjusted according to the input data, which can achieve adaptive spatial aggregation like ViTs, reducing the over-inductive bias of regular convolutions; and (3) the convolution window is a common 33, avoiding the optimization problems and expensive costs caused by large dense kernels .
With the aforementioned designs, the proposed InternImage can efficiently scale to large parameter sizes and learn stronger representations from large-scale training data, achieving comparable or even better performance to large-scale ViTs on a wide range of vision tasks. In summary, our main contributions are as follows:
(1) We present a new large-scale CNN-based foundation model—InternImage. To our best knowledge, it is the first CNN that effectively scales to over 1 billion parameters and 400 million training images and achieves comparable or even better performance than state-of-the-art ViTs, showing that convolutional models are also a worth-exploring direction for large-scale model research.
(2) We successfully scale CNNs to large-scale settings by introducing long-range dependencies and adaptive spatial aggregation using an improved 33 DCN operator, and explore the tailored basic block, stacking rules, and scaling strategies centered on the operator. These designs make effective use of the operator, enabling our models to obtain the gains from large-scale parameters and data.
(3) We evaluate the proposed model on representative vision tasks including image classification, object detection, instance and semantic segmentation, and compared it with state-of-the-art CNNs and large-scale ViTs by scaling the model size ranging from 30 million to 1 billion, the data ranging from 1 million to 400 million. Specifically, our model with different parameter sizes can consistently outperform prior arts on ImageNet . InternImage-B achieves 84.9% top-1 accuracy trained only on the ImageNet-1K dataset, outperforming CNN-based counterparts by at least 1.1 points. With large-scale parameters (i.e., 1 billion) and training data (i.e., 427 million), the top-1 accuracy of InternImage-H is further boosted to 89.6%, which is close to well-engineering ViTs and hybrid-ViTs . In addition, on COCO , a challenging downstream benchmark, our best model InternImage-H achieves state-of-the-art 65.4% box mAP with 2.18 billion parameters, 2.3 points higher than SwinV2-G (65.4 vs. 63.1) with 27% fewer parameters as shown in Fig. 2.
Related Work
Vision foundation models. Convolutional neural networks (CNNs) became the mainstream for visual recognition after the large-scale dataset and computation resources were available. Straining from AlexNet , lots of deeper and more effective neural network architectures have been proposed, such as VGG , GoogleNet , ResNet , ResNeXt , EfficientNet , etc. In addition to the architectural design, more sophisticated convolution operations such as depth-wise convolution and deformable convolution are formulated. By considering the advanced designs of transformers, modern CNNs showed promising performance on the vision tasks by discovering better components in macro/micro designs and introducing improved convolutions with long-range dependencies or dynamic weights .
In recent years, a new line of vision foundation models focuses on transformer-based architecture. ViT is the most representative model, which achieves great success in vision tasks thanks to global receptive fields and dynamic spatial aggregation. However, global attention in ViT suffers from expensive computational/memory complexity, especially on large feature maps, which limits its application in downstream tasks. To address this problem, PVT and Linformer performed global attention on the downsampled key and value maps, DAT employed deformable attention to sparsely sample information from value maps, while HaloNet and Swin transformer developed local attention mechanisms and used haloing and shift operations to transfer information among adjacent local regions.
Large-scale models. Scaling up models is an important strategy to improve feature representation quality, which has been well-studied in the natural language processing (NLP) domain . Inspired by the success in the NLP field, Zhai et al. first extended ViT to 2 billion parameters. Liu et al. enlarged the hierarchical-structure Swin transformer to a deeper and wider model with 3 billion parameters. Some researchers developed large-scale hybrid ViTs by combining the advantages of ViTs and CNNs at different levels. Recently, BEiT-3 further explored stronger representations based on ViT with large-scale parameters using multimodal pre-training. These methods significantly raise the upper bound of basic vision tasks. However, research on CNN-based large-scale models has lagged behind transformer-based architectures in terms of the total number of parameters and performance. Although newly-proposed CNNs introduce long-range dependencies by using convolutions with very large kernels or recursive gated kernels, there is still a considerable gap with state-of-the-art ViTs. In this work, we aim to develop a CNN-based foundation model that can extend efficiently to a large scale comparable to ViT.
Proposed Method
To design a large-scale CNN-based foundation model, we start with a flexible convolution variant, namely deformable convolution v2 (DCNv2) and make some tune-ups based on it to better suit the requirements of large-scale foundation models. Then, we build the basic block by combining the tuned convolution operator with advanced block designs used in modern backbones . Finally, we explore the stacking and scaling principles of DCN-based blocks to build a large-scale convolutional model that can learn strong representations from massive data.
Convolution vs. MHSA. Previous works have extensively discussed the differences between CNNs and ViTs. Before deciding on the core operator of InternImage, we first summarize the main differences between regular convolution and MHSA.
(1) Long-range dependencies. Although it has long been recognized that models with large effective receptive fields (long-range dependencies) usually perform better on downstream vision tasks , the de-facto effective receptive field of CNNs stacked by 33 regular convolutions is relatively small. Even with very deep models, the CNN-based model still cannot acquire long-range dependencies like ViTs, which limits its performance.
(2) Adaptive spatial aggregation. Compared to MHSA whose weights are dynamically conditioned by the input, regular convolution is an operator with static weights and strong inductive biases such as 2D locality, neighborhood structure, translation equivalence, etc. With the highly-inductive properties, models composed by regular convolutions might converge faster and require less training data than ViTs, but it also restricts CNNs from learning more general and robust patterns from web-scale data.
Extending DCNv2 for Vision Foundation Models. In common practice, DCNv2 is usually used as an extension to regular convolutions, loading pre-trained weights of regular convolutions and fine-tuning for better performance, which is not exactly suitable for large-scale vision foundation models that need to be trained from scratch. In this work, to address this problem, we extend DCNv2 from aspects as follows:
(1) Sharing weights among convolutional neurons. Similar to regular convolution, different convolutional neuronsA 33 regular convolution has 9 linear projection neurons. in original DCNv2 have independent linear projection weights, and thus its parameter and memory complexity are linear with the total number of sampling points, which significantly limits the efficiency of the model, especially in large-scale models. To remedy this problem, we borrow the idea from the separable convolution and detach the original convolution weights into depth-wise and point-wise parts, where the depth-wise part is responsible by the original location-aware modulation scalar , and the point-wise part is the shared projection weights among sampling points.
(2) Introducing multi-group mechanism. The multi-group (head) design first appeared in group convolution , and it is widely used in MHSA of transformers and works with adaptive spatial aggregation to effectively learn richer information from different representation subspaces at different locations. Inspired by this, we split the spatial aggregation process into groups, each of which has individual sampling offsets and modulation scale , and thus different groups on a single convolution layer can have different spatial aggregation patterns, resulting in stronger features for downstream tasks.
(3) Normalizing modulation scalars along sampling points. The modulation scalars in the original DCNv2 are element-wise normalized by the sigmoid function. Therefore, each modulation scalar is in the range , and the sum of the modulation scalars of all sample points is not stable and varies from 0 to . This leads to unstable gradients in DCNv2 layers when training with large-scale parameters and data. To alleviate the instability issues, we change element-wise sigmoid normalization to softmax normalization along sample points. In this way, the sum of the modulation scalars is constrained to 1, which makes the training process of models at different scales more stable.
Combining the aforementioned modifications, the extended DCNv2, marked as DCNv3, can be formulated as Eqn. (2).
In general, DCNv3, as an extension of the DCN series, enjoys three merits as follows: (1) This operator made up for the deficiencies of regular convolution in terms of long-range dependencies and adaptive spatial aggregation; (2) Compared with attention-based operators such as common MHSA and closely-related deformable attention , this operator inherits the inductive bias of convolution, making our model more efficient with fewer training data and shorter training time; (3) This operator is based on sparse sampling, which is more computational and memory efficient than previous methods such as MHSA and re-parameterizing large kernel . In addition, due to the sparse sampling, DCNv3 only needs a 33 kernel to learn long-range dependencies, which is easier to be optimized and avoids extra auxiliary techniques such as re-parameterizing used in large kernels.
2 InternImage Model
Using DCNv3 as the core operator brings a new problem: how to build a model that can make effective use of the core operator? In this section, we first present the details of the basic block and other integral layers of our model, and then we construct a new CNN-based foundation model termed InternImage, by exploring a tailored stacking strategy for these basic blocks. Finally, we study scaling-up rules for the proposed model to obtain the gain from increasing parameters.
Basic block. Unlike the widely used bottlenecks in traditional CNNs , the design of our basic block is closer to ViTs, which is equipped with more advanced components including LN , feed-forward networks (FFN) , and GELU . This design is proved to be efficient in various vision tasks. The details of our basic block are illustrated in Fig. 3, where the core operator is DCNv3, and the sampling offsets and modulation scales are predicted by passing input feature through a separable convolution (a 33 depth-wise convolution followed by a linear projection). For other components, we use the post-normalization setting by default and follow the same design as that of the plain transformer .
Stem & downsampling layers. To obtain hierarchical feature maps, we use convolutional stem and downsampling layers to resize the feature maps to different scales. As shown in Fig. 3, the stem layer is placed before the first stage to reduce the input resolution by 4 times. It consists of two convolutions, two LN layers, and one GELU layer, where the kernel size of the two convolutions is 3, the stride is 2, the padding is 1, and the output channel of the first convolution is half of the second one. Similarly, the downsampling layer is made up of a 33 convolution with a stride of 2 and a padding of 1, followed by one LN layer. It sits between the two stages and is used to downsample the input feature map by 2 times.
Stacking rules. To clarify the block-stacking process, we first list the integral hyperparameters of the InternImage as follows: : the channel number of the -th stage; : the group number of the DCNv3 in the -th stage; : the number of basic blocks in the -th stage. Since our model has 4 stages, a variant is decided by 12 hyper-parameters, whose search space is too large to exhaustively enumerate and find the best variant. To reduce the search space, we summarize the design experiences of prior arts into 4 rules as shown in Fig. 3, where the first rule makes the channel numbers of the last three stages determined by the channel number of the first stage, and the second rule lets the group number correspond to the channel number of stages. For the number of stacked blocks in different stages, we simplify the stacking pattern to “AABA”, which means the block number of stage 1, 2, and 4 are the same, and are not greater than that of the stage 3 as illustrated in the last two rules. With these rules, a InternImage variant can be defined by using only 4 hyper-parameters .
Let us choose a model with 30 million parameters as the origin and discretize to , to , and to . In this way, the original huge search space is reduced to 30, and we can find the best model from the 30 variants by training and evaluating them in ImageNet . In practice, we use the best hyper-parameter setting to define the origin model and scale it to different scales.
Scaling rules. Based on the optimal origin model under the aforementioned constraints, we further explore the parameter scaling rules inspired by . Specifically, we consider two scaling dimensions: depth (i.e., ) and width , and scale the two dimensions using , and a composite factor . The scaling rules can be written as: and , where , , and . Here, 1.99 is specific for InternImage and calculated by doubling the model width and keeping the depth constant. We experimentally find out that the best scaling setting is and , and then we base on it to construct InternImage variants with different parameter scales, namely InternImage-T/S/B/L/XL, whose complexity is similar to those of ConvNeXt . To further test the capability, we built a larger InternImage-H with 1 billion parameters, and to accommodate very large model widths, we also change the group dimension to 32. The configurations are summarized in Table 1.
Experiment
We analyze and compare InternImage with the leading CNNs and ViTs on representative vision tasks including image classification, object detection, instance and semantic segmentation. Besides the experiments in the main paper, due to space constraints, more experimental setups and ablation studies are presented in the supplementary material.
Settings. We evaluate the classification performance of InternImage on ImageNet . For fair comparisons, following common practices , InternImage-T/S/B are trained on ImageNet-1K (1.3 million) for 300 epochs, and InternImage-L/XL are first trained on ImageNet-22K (14.2 million) for 90 epochs and then fine-tuned on ImageNet-1K for 20 epochs. To further explore the capability of our model and match the large-scale private data used in previous methods , we adopt M3I Pre-training , a unified pre-training approach available for both unlabeled and weakly-labeled data, to pre-train InternImage-H on a 427 million joint dataset of public Laion-400M , YFCC-15M , and CC12M for 30 epochs, and then we fine-tune the model on ImageNet-1K for 20 epochs.
Results. Table 2 shows the classification results of models with different scales. With similar parameters and computational costs, our models are comparable or even superior to the state-of-the-art transformer-based and CNN-based models. For example, InternImage-T achieves 83.5% top-1 accuracy, outperforming ConvNext-T with a clear margin of 1.4 points. InternImage-S/B keeps the leading position and InternImage-B surpasses the hybrid-ViT CoAtNet-2 by 0.8 points. When pre-trained on ImageNet-22K and the large-scale joint dataset, the top-1 accuracy of InternImage-XL and -H are boosted to 88.0% and 89.6%, respectively, which is better than previous CNNs also trained with large-scale data, and closes the gap with the state-of-the-art large-scale ViTs to about 1 point. This gap may be caused by the discrepancy between large-scale inaccessible private data and the aforementioned joint public data. These results show that our InternImage not only has good performance on the common parameter scale and the public training data, but also can effectively extend to large-scale parameters and data.
2 Object Detection
Settings. We verify the detection performance of our InternImage on the COCO benchmark , on top of two representative object detection frameworks: Mask R-CNN , and Cascade Mask R-CNN . We follow common practices to initialize the backbone with pre-trained classification weights, and train models use a 1 (12 epochs) or 3 (36 epochs) schedule by default.
Results. As shown in Table 3, when using Mask R-CNN for object detection, we find that under a comparable number of parameters, our models significantly surpass their counterparts. For example, with the training schedule, the box AP (APb) of InternImage-T is 4.5 points better than Swin-T (47.2 vs. 42.7), and 3.0 points higher than ConvNeXt-T (47.2 vs. 44.2). With the 3 multi-scale training schedule, more parameters, and more advanced Cascade Mask R-CNN , InternImage-XL achieves APb of 56.2, surpassing ConvNeXt-XL by 1.0 points (56.2 vs. 55.2). Similar results are also seen in instance segmentation experiments. With the 1 training schedule, InternImage-T yields 42.5 mask AP (i.e., APm), which outperforms Swin-T and ConvNeXt-T by 3.2 points (42.5 vs. 39.3) and 2.4 points (42.5 vs. 40.1), respectively. The best APm 48.8 is obtained by InternImage-XL with Cascade Mask R-CNN, which is at least 1.1 points higher than its counterparts.
To further push the performance bound of object detection, we follow the advanced setting used in leading methods to initialize the backbone with the weights pre-trained on ImageNet-22K or the large-scale joint dataset, and double its parameters via the composite techniques (see the model with 2 billion parameters in Fig. 2). Then, we fine-tune it along with the DINO detector on the Objects365 and COCO datasets one after another for 26 epochs and 12 epochs, respectively. As shown in Table 4, our method achieves the best results of 65.0 AP and 65.4 AP on COCO val2017 and test-dev. Compared to previous state-of-the-art models, we surpass FD-SwinV2-G by 1.2 points (65.4 vs. 64.2), with 27% fewer parameters and without complicated distillation processes, which shows the effectiveness of our models on the detection task.
3 Semantic Segmentation
Settings. To evaluate the semantic segmentation performance of InternImage, we initialize the backbone with pre-trained classification weights and train our models with UperNet on ADE20K for 160k iterations and compare fairly with previous CNN-based and transformer-based backbones. To further reach top performance, we arm InternImage-H with more advanced Mask2Former , and adopt the same training settings in .
Results. As shown in Table 5, when using UperNet for semantic segmentation, our InternImage consistently outperforms prior arts . For example, with almost the same parameter numbers and FLOPs, our InternImage-B reports 50.8 mIoU on the ADE20K val, which is outstanding from the strong counterparts such as ConvNeXt-B (50.8 vs. 49.1) and RepLKNet-31B (50.8 vs. 49.9). Furthermore, our InternImage-H yields 60.3 MS mIoU, which is better than SwinV2-G , while the parameter number is much smaller (1.12B vs. 3.00B).
It is worth noting that, when using Mask2Former and multi-scale testing, our InternImage-H achieves the best mIoU of 62.9, higher than the current best BEiT-3 on the ADE20K benchmark. These results demonstrate that the CNN-based foundation model can also enjoy the dividends of massive data and challenge the leading position of transformer-based models.
4 Ablation Study
Sharing weights among convolution neurons matters. Large-scale models are sensitive to parameters and memory cost of the core operator, due to hardware limitations. To address this problem, we share weights among convolution neurons of DCNv3. As shown in Fig. 4, we compare the parameters and memory cost of the models based on DCNv3 with shared or unshared weights. We see that the parameters and memory cost of models with unshared weights are much higher than the shared one, especially for the -H scale, the ratio of saved parameters and GPU memory is 42.0% and 84.2%, respectively. As shown in Table 6, we also examine that the two models at -T scale have similar top-1 accuracy on ImageNet (83.5 vs. 83.6) and AP on COCO (47.2 vs. 47.4), even the model without shared weights has 66.1% more parameters.
Multi-group spatial aggregation brings stronger features. We introduce aggregation groups to allow our model to learn information from different representation subspaces like transformers . As shown in Fig. 5, for the same query pixel, the offsets from different groups are concentrated in different regions, resulting in hierarchical semantic features. We also compare the performance of the model with and without multiple groups. As reported in Table 6, the model significantly drops 1.2 points on ImageNet and 3.4 points on COCO val2017. In addition, we also see that in the first two stages, the learned effective receptive field (ERF) is relatively small, and as the model goes deeper (i.e., stages 3 and 4), the ERF increases to be global. This phenomenon is different from ViTs whose ERF is usually global.
Conclusion & Limitations
We introduce InternImage, a new large-scale CNN-based foundation model that can provide strong representations for versatile vision tasks, such as image classification, object detection, and semantic segmentation. We tune the flexible DCNv2 operator to satisfy the requirement of foundation models, and develop a series of blocks, stacking and scaling rules centered on the core operator. Extensive experiments on object detection and semantic segmentation benchmarks verify that our InternImage can obtain comparable or better performance than well-designed large-scale vision transformers trained with massive data, showing that CNN is also a considerable choice for large-scale vision foundation model research. Nonetheless, latency remains an issue for DCN-based operators adapting to downstream tasks with high-speed requirements. Also, large-scale CNNs are still in their early stages of development, and we hope InternImage can serve as a good starting point.
Appendix A Detailed Training Settings
In this section, we present the detailed training recipes for image classification, object detection, and semantic segmentation.
ImageNet image classification. The training details of image classification on ImageNet are shown in Table 7, which are similar to common practices and with some tweaks. To further explore the capability of our model and match the large-scale private data used in previous methods , we adopt M3I Pre-training , a unified pre-training approach available for both unlabeled and weakly-labeled data, to pre-train InternImage-H on a 427 million joint dataset of public Laion-400M , YFCC-15M , and CC12M for 30 epochs, and then we fine-tune the model on ImageNet-1K for 20 epochs. For the more detailed pre-training settings of InternImage-H, please refer to M3I Pre-training .
COCO object detection. We verify the detection performance of our InternImage on the COCO benchmark , on top of Mask R-CNN and Cascade Mask R-CNN . For fair comparisons, we follow common practices to initialize the backbone with pre-trained classification weights, and train these models using a 1 (12 epochs) or 3 (36 epochs) schedule by default. For 1 schedule, the image is resized to have a shorter side of 800 pixels, while the longer side does not exceed 1,333 pixels. During testing, the shorter side of the input image is fixed to 800 pixels. For 3 schedule, the shorter side is resized to 480800 pixels, while the longer side does not exceed 1,333 pixels. All these detection models are trained with a batch size of 16 and optimized by AdamW with an initial learning rate of .
ADE20K semantic segmentation. We evaluate our InternImage models on the ADE20K dataset , and initialize them with the pre-trained classification weights. For the InternImage-T/S/B models, we optimize them using AdamW with an initial learning rate of 6, and 2 for InternImage-X/XL. The learning rate is decayed following the polynomial decay schedule with a power of 1.0. Following previous methods , the crop size is set to 512 for InternImage-T/S/B, and 640 for InternImage-L/XL. All segmentation models are trained using UperNet with a batch size of 16 for 160k iterations, and compared fairly with previous CNN-based and transformer-based backbones.
A.2 Settings for System-Level Comparison
COCO object detection. For system-level comparison with state-of-the-art large-scale detection models , we first initialize the InternImage-XL/H backbone with the weights pre-trained on ImageNet-22K or the 427M large-scale joint dataset, and double its parameters using the composite techniques . Then, we pre-train the model along with the DINO detector on the Objects365 for 26 epochs, with an initial learning rate of and a batch size of 256. The shorter size of input images is resized to 6001200 pixels during pre-training, and the learning rate drops by 10 times at epoch 22. Finally, we fine-tune these detectors on the COCO dataset for 12 epochs, where the batch size is 64, and the initial learning rate is , which drops by 10 times at the final epoch.
ADE20K semantic segmentation. To further reach leading segmentation performance, we first initialize our InternImage-H backbone with the pre-trained weights on the 427M large-scale joint dataset, and arm it with the state-of-the-art segmentation method Mask2Former . We follow the same training settings in , i.e. pre-training and fine-tuning the model on COCO-Stuff and ADE20K datasets both for 80k iterations, with a crop size of 896 and an initial learning rate of 1.
Appendix B Exploration of Hyper-parameters
As discussed in Section 3.2, our model is constructed in four stacking rules, and we further restrict the model parameters to 30M for the origin model. We discretize the stacking hyperparameters to , to , and to . And is determined by selecting the model size to approximately 30M. In this way, we obtained 30 models by combining the three hyper-parameters.
We adopt the training recipe listed in Table 7 to train our -T models unless otherwise stated. Fig. 6 shows the ImageNet-1K top-1 accuracy of these models under the same training settings, with darker green indicating higher accuracy, i.e., models with stronger representational capability. When equals 16, models are generally higher than that with of 32, and works best at 4, thanks to a reasonable stacking ratio. A large number of channels allows for more gain. Finally, through the above exploration experiments, we determine our basic stacking hyper-parameter to .
B.2 Model Scaling
In Section 3.2, we have shown the constraints on the depth scaling factor and the width scaling factor . Based on this condition and the -T model (30M), we display reasonable scaling possibilities for extending the -T model to -B models (100M). As illustrated in Table 8, the first two columns show the formulas for and . The penultimate column indicates model parameters, and the last column indicates the ImageNet-1K top-1 accuracy of these models after 300 training epochs.
It is worth noting that the model width needs to be divisible by . Therefore some adjustment is required in determining the specific scaling parameters. This results in a small fluctuation in the number of parameters, but this is acceptable. Our exploratory experiments prove that when is set at for the best performance. In addition, the other size models -S/L/XL/H also confirmed the effectiveness of our scaling rules.
B.3 Kernel Size
As mentioned in Section 3.1, we argue 33 dynamic sparse convolution is enough for the large receptive field. Here, we explore the role played by the number of convolutional neurons in the DCNv3 operator. Specifically, we replaced the kernel in the DCNv3 operator with the or kernel. They are all trained by the -T training recipes (see Table 7) and validated on the ImageNet-1K validation set. The results are shown in Table 9.
The results show that when enlarging the convolution kernel, the parameters and FLOPs are followed by the surge, while the accuracy is not significantly improved (83.5 v.s 83.6) or even decreased (83.5 v.s 82.8). These results show that when the number of convolutional neurons in a single layer increases, the model becomes more difficult to optimize. This phenomenon is also confirmed in RepLKNet , and it addresses this problem by re-parameterizing techniques, which might bring extra time and memory costs in the training phase. In this work, we avoid this problem by adopting the simple yet effective DCNv3 as InternImage’s core operator.
Fig. 7 shows the effective receptive fields (ERF) of ResNet-101 and InternImage-S. A wider distribution of bright areas indicates a larger ERF. We uniformly activate the input image at the dog’s eye, count the gradient map of each block, aggregate by channel, and map back to the input image. We see that the ERF of ResNet-101 without training is limited to a local area, while the fully trained ResNet-101 still has an ERF around the eye, and the gradient amplitude is lower, and the distribution is more sparse. Therefore, the area that ResNet-101 can effectively perceive is very limited. For the InternImage-S without training, its ERF is concentrated around the activation point. Since the offset is not learned, its ERF is also very small in the last two blocks. But after sufficient training, InternImage-L can effectively perceive the information of the entire image in the 3-rd and 4-th stages.
Appendix C Additional Downstream Tasks
iNaturalist 2018 is a read-word long-tailed dataset containing 8142 fine-graned species. The dataset comprises 437.5K training images and an imbalance factor of 500. For this experiment, we initialize our InternImage-H model with the pre-trained weights on the 427M large-scale joint dataset, and fine-tune it on the training set of iNaturalist 2018 for 100 epochs. We follow MetaFormer to adopt a resolution of 384384 for fine-tuning, with the utilization of meta information. Other training settings are the same as the recipe for fine-tuning InternImage-H on ImageNet-1K, as reported in Table 7. As a result, our method achieves the state-of-the-art accuracy of 92.6 (see Table 10) on the validation set of iNaturalist 2018, 3.9 points better than the previous best model MetaFormer .
Places205 is a dataset containing 2.5 million images of 205 scene categories, which are dedicated to the scene recognition task. The images in this dataset cover a wide range of indoor and outdoor scenes, such as offices, kitchens, forests, and beaches. We initialize our model with pre-trained weights on a large-scale joint dataset, consisting of 427 million images, and fine-tune it on the Places205 training set. Other training settings are the same as the recipe for fine-tuning InternImage-H on ImageNet-1K, as reported in Table 7. Our method achieves state-of-the-art accuracy of 71.7 (see Table 10) on the validation set of Places205, outperforming the previous best model MixMIM-L by 2.4 points.
Places365 is a dataset containing 1.8 million images of 365 scene categories, which are dedicated to the scene recognition task. The images in this dataset cover a wide range of indoor and outdoor scenes, such as airports, bedrooms, deserts, and waterfalls. The specific pre-training and fine-tuning strategies are the same as for Places205. Our method achieves state-of-the-art accuracy of 61.2 (see Table 10) on the validation set of Places365, outperforming the previous best model SWAG by 0.5 points. The Places365 dataset provides a more fine-grained classification task compared to Places205, allowing our model to learn more subtle differences between similar scenes.
C.2 Object Detection
LVIS v1.0 is a large-scale vocabulary dataset for object detection and instance segmentation tasks, which contains 1203 categories in 164k images. For this dataset, we initialize our InternImage-H with the Objects365 pre-trained weights, then fine-tune it on the training set of LVIS v1.0. Here, we report the box AP (i.e., APb) with multi-scale testing on the minival set and the val set, respectively. As shown in Table 10, our InternImage-H creates a new record of 65.8 APb on the LVIS minival, and 63.2 APb on the LVIS val, outperforming previous state-of-the-art methods by clear margins.
Pascal VOC contains 20 object classes, which has been widely used as a benchmark for object detection tasks. We adopt this dataset to further evaluate the detection performance of our model. Specifically, we employ the Objects365 pre-trained weights to initialize our InternImage-H, and fine-tune it on the trainval set of Pascal VOC 2007 and Pascal VOC 2012 following previous method . As shown in Table 10, on the Pascal VOC 2007 test set, our InternImage-H yields 94.0 AP50 with single-scale testing, which is 4.7 points better than previous best Cascade Eff-B7 NAS-FPN . On the Pascal VOC 2012 test set, our method achieves 97.2 mAP, 4.3 points higher than the best record on the official leaderboard .
OpenImages v6 is a dataset of about 9 million images with 16M bounding boxes for 600 object classes on 1.9 million images dedicated to the object detection task, which are very diverse and often embrace complex scenes with multiple objects (8.3 per image on average). For this dataset, we use the same settings as the previous two datasets. In addition, we follow to use the class-aware sampling during fine-tuning. As reported in Table 10, our InternImage-H yields 74.1 mAP, achieving 1.9 mAP improvement compared to the previous best results .
CrownHuman is a benchmark dataset to better evaluate detectors in crowd scenarios. The CrowdHuman dataset is large, rich-annotated and contains high diversity. CrowdHuman contains 15000, 4370 and 5000 images for training, validation, and testing, respectively. There are a total of 470K human instances from train and validation subsets and 23 persons per image, with various kinds of occlusions in the dataset. We used the same training setup as for the previous dataset. Our pre-trained model reached optimal performance in 3750 iterations, exceeding the previous best model Iter-Deformable-DETR by 3.1 AP.
BDD100K is a dataset of around 100K high-resolution images with diverse weather and lighting conditions, containing 10 object categories, including pedestrians, cars, buses, and bicycles, dedicated to the object detection task. The images in this dataset are captured from a moving vehicle, simulating real-world scenarios. For this experiment, we initialize our InternImage-H model with the pre-trained weights on the 427M joint dataset and fine-tune it on the BDD100K training set for 12 epochs. As reported in Table 10, our InternImage-H achieves 38.8 mAP on the validation set, which is the state-of-the-art performance, surpassing the previous best model by 3.2 mAP. Our method demonstrates superior performance in detecting objects in real-world driving scenarios, which can benefit autonomous driving and intelligent transportation systems.
C.3 Semantic Segmentation
COCO-Stuff includes the images from the COCO dataset for semantic segmentation, spanning over 171 categories. Specifically, COCO-Stuff-164K is the full set that contains all 164k images, while COCO-Stuff-10K is a subset of the -164K that splits into 9,000 and 1,000 images for training and testing. Here, we equip our InternImage-H with the advanced Mask2Former , and pre-train the model on the COCO-Stuff-164K for 80k iterations. Then we fine-tune it on the COCO-Stuff-10K for 40k iterations and report the multi-scale mIoU. The crop size is set to 512512 in this experiment. As shown in Table 10, our model achieves 59.6 MS mIoU on the test set, outperforming the previous best ViT-Adapter by 5.4 mIoU.
Pascal Context contains 59 semantic classes. It is divided into 4,996 images for training and 5,104 images for testing. For this dataset, we also employ Mask2Former with our InternImage-H, and follow the training settings in . Specifically, we first load the classification pre-trained weights to initialize the model, then fine-tune it on the training set of Pascal Context for 40k iterations. The crop size is set to 480480 in this experiment. As shown in Table 10, our method reports 70.3 MS mIoU on the test set, which is 2.1 points better than ViT-Adapter .
Cityscapes is a high-resolution dataset recorded in street scenes including 19 classes. In this experiment, we use Mask2Former as the segmentation framework. Following common practices , we first pre-train on Mapillary Vistas and then fine-tune on Cityscapes for 80k iterations, respectively. The crop size is set to 10241024 in this experiment. As shown in Table 10, our InternImage-H achieves 87.0 MS mIoU on the validation set, and 86.1 MS mIoU on the test set.
NYU Depth V2 comprises of 1449 RGB-D images, each with a size of 640480. These images are divided into 795 training and 654 testing images, each with annotations on 40 semantic categories. We adopt the same training settings as we used when fine-tuning on Pascal Context. As shown in Table 10, our method achieves a big jump to 68.1 MS mIoU on the validation set, which is 11.2 points better than CMX-B5 .
Appendix D Throughput Analysis
In this section, we benchmark the throughput of our InternImage with counterparts, including a variant equipped with DCNv2 , ConvNext , RepLKNet , and a vision transformer with deformable attention (DAT) . As shown in Table 11, compared to the variant with DCNv2 , our model enjoys better parameter-efficient and significantly faster inference speed under both 224224 and 800800 input resolutions. Compared to RepLKNet-B and DAT-B , our model has a throughput advantage at a high input resolution (i.e., 800800). This resolution is widely used in dense prediction tasks such as object detection. Compared with ConvNeXt , despite the throughput gap due to DCN-based operators, our model still has an accuracy advantage (84.9 vs. 83.8), and we are also looking for an efficient DCN to make our model more suitable for downstream tasks that require high efficiency.
Appendix E Robustness Evaluation on ImageNet
In this section, we evaluate the robustness of different models under different transformations (see Fig. 8). We consider translation, rotation, and scaling to evaluate. The models we choose for comparison include a convolutional model (ConvNeXt-T ), a local attention-based model (Swin-T ), a global attention-based model (PVTv2-B2 ), and our InternImage-T.
Translation invariance describes the capability of the model to retain the original output when the input image is translated. We evaluate the translation invariance in the classification task by dithering the image from 0 to 64 pixels. The invariance is measured by the probability that the model predicts the same label when the same input image is translated. The first row of Fig. 8 indicates our InternImagehas the translation invariance of the different methods. It is evident that the robustness of the four models to translation is shown as our method is the best, followed by convolution-based ConvNeXt, followed by global attention-based PVTv2, and the worst local attention-based Swin transformer.
E.2 Rotation Invariance
To evaluate the rotation invariance of the classification task, we rotate the image from to in steps of . In a similar way to translation invariance, the predicted consistency under different rotation angles is used to evaluate the rotational invariance. From the second row of Fig. 8, we found that the consistency performance of all models is comparable in the small angle phase. However, at large-angle rotation (i.e., ), our model is clearly superior to the other models.
E.3 Scaling Invariance
We evaluate the scaling invariance on object detection. The scaling factor of the input image varies from 0.25 to 3.0 in steps of 0.25. Detection consistency is defined as the invariance metric for the detection task. The predicted boxes on the scaled images are first converted back to the original resolution, and then the predicted boxes at the original resolution are used as the ground truth boxes to calculate the box mAP. As seen in the last row of Fig. 8, we can observe that all methods of our experiments are sensitive to down-scaling. And they show invariance comparable to the input at small resolutions. Our method performs better when scaling up the images. Both box consistency and bounding box mAP are better than the others.
E.4 How Hungry the Model is for Data Scale?
In order to verify the robustness of the model to the data scale. We uniformly sampled the ImageNet-1K data to obtain 1%, 10%, and 100% data, respectively. And we chose ResNet-50 , ConvNeXt-T , Swin-T , InternImage-T-dattn and our InternImage-T to conduct 300 rounds of training experiments on these data. The experimental settings are consistent with Table 7. The experimental results can be viewed in Table 12. We see that ResNet performs best on the 1% and 10% data (12.2% & 57.5%), benefiting from its inductive biases. But its upper limitation is low (80.4%) when the data is sufficient. Swin-T fails completely in 1% datasets and shows good performance only on the 100% dataset. The proposed InternImage-T has strong robustness not only on 1% and 10% data (5.9% and 56.0%) but also on full data (83.5%), which is consistently better than the InternImage-T variant with deformable attention (dattn) and ConvNeXt . These results indicate the robustness of our model with respect to the data scale.