Group DETR: Fast DETR Training with Group-Wise One-to-Many Assignment

Qiang Chen, Xiaokang Chen, Jian Wang, Shan Zhang, Kun Yao, Haocheng Feng, Junyu Han, Errui Ding, Gang Zeng, Jingdong Wang

Introduction

Detection Transformer (DETR) conducts end-to-end object detection without the need for many hand-crafted components, such as non-maximum suppression (NMS) and anchor generation . The architecture consists of a CNN and transformer encoder , and a transformer decoder that consists of self-attention, cross-attention and FFNs, followed by class and box prediction FFNs. During training, one-to-one assignment, where one ground-truth object is assigned to one single prediction, is applied for learning to only promote the predictions assigned to ground-truth objects, and demote the duplicate predictions.

This work explores the solutions to accelerate the DETR training process. Previous solutions contain two main lines. The one line is to modify cross-attention so that informative image regions are selected for effectively and efficiently collecting the information from image features. Example methods include sparse sampling, through deformable attention , and spatial modulations with modifying object queries . The other line is to stabilize one-to-one assignment during training, e.g., feeding ground-truth bounding boxes with noises into transformer decoder .

We are interested in the second line. Instead of focusing on stabilizing the assignment like DN-DETR , we study the assignment scheme for efficient DETR training from a new perspective: introducing more supervision. It has been proven that assigning one ground-truth object to multiple predictions, i.e., one-to-many assignment, is successful in traditional object detection methods, e.g., Faster R-CNN and FCOS with more anchors and pixels assigned to one ground-truth object. Unfortunately, naive one-to-many assignment does not work for DETR training. It remains a challenge to apply one-to-many assignment to DETR training.

We present a simple yet efficient DETR training approach that uses a group-wise way for one-to-many assignment, called Group DETR. Our approach is based on that end-to-end detection with successful removal of NMS post-processing for DETR comes from the joint effect of two components : decoder self-attention, which collects the information of other predictions, and one-to-one assignment, which expects to learn to score one prediction higher and other duplicate predictions lower for one ground-truth object.

Our approach adopts KK groups of object queries, and introduces group-wise one-to-many assignment. This assignment scheme conducts one-to-one assignment within each group of object queries, resulting in that one ground-truth object is assigned to multiple predictions. It is encouraged that the prediction assigned to the ground-truth object gets a high score, and other duplicate predictions from the same group of queries get low scores. In other words, the predictions make competition within each group. Thus, our approach uses separate self-attention, i.e., self-attention is done for each group separately, eliminating the influence of predictions from other groups and easing DETR training. Regarding inference, it is the same as DETR trained normally, and only needs a single group of object queries.

The resulting architecture is equivalent to DETR with a group of parallel decoders, illustrated in Figure 2 (a). During training, the parallel decoders boost each other through sharing decoder parameters and using different object queries. On the other hand, using more groups of object queries resembles data augmentation, and behaves as query augmentation. It introduces more supervision and improve the decoder training. In addition, it is empirically observed that the encoder training is also improved, presumably with the help of the improved decoder.

Group DETR is versatile and is applicable to various DETR variants. Extensive experiments demonstrate that our approach is effective in achieving fast training convergence, shown in Figure 1. Group DETR obtains consistent improvements on various DETR-based methods . For instance, Group DETR significantly improves Conditional DETR-C55 by 5.0{5.0} mAP with 1212-epoch training on COCO . The non-trivial improvements hold when we adopt longer training schedules (e.g., 3636 epochs and 5050 epochs). Furthermore, Group DETR outperforms baseline methods for multi-view 33D object detection and instance segmentation .

Background

DETR Architecture. DETR is composed of an encoder, a transformer decoder, and object class and box position predictors. The encoder takes an image I\mathbf{I} as input, and outputs the image feature X\mathbf{X},

The decoder is a sequence of multiple layers. Each layer includes: (i) self-attention over object queries, which performs interactions among queries for collecting the information about duplicate detection; (ii) cross-attention between queries and image features, which collects the information from image features that is useful for object detection; (iii) feed-forward network that processes the queries separately to benefit object detection.

DETR Training. The predictions during DETR training are in the set form, and have no correspondence to the ground-truth objects. DETR uses one-to-one assignment, i.e., one ground-truth object is assigned to one predictions and vice versa, through building a bipartite matching between the predictions and the ground-truth objects:

Here, σ(⋅)\sigma(\cdot) is the optimal permutation of NN indices, and [yˉ1 yˉ2,… yˉN]=Yˉ[\bar{\mathbf{y}}_{1}~{}\bar{\mathbf{y}}_{2},\dots~{}\bar{\mathbf{y}}_{N}]=\bar{\mathbf{Y}} correponds to ground truth. The loss is then formulated as below:

Optimization with one-to-one assignment aims to score the predictions for promoting one prediction for one ground-truth object, and demoting duplicate predictions. Such scoring needs the comparison of one prediction with other predictions, and the information of other predictions is provided from decoder self-attention over queries. The two designs, one-to-one assignment and self-attention over object queries, are critical for end-to-end detection without the need of the post-processing NMS.

One-to-many assignment for non-end-to-end detection. One-to-many assignment is successfully adopted for introducing more supervision to non-end-to-end detection training, such as Faster R-CNN , FCOS , and so on . One ground-truth object is assigned to multiple anchors or multiple pixels. During inference, a post-processing NMS is conducted for duplicate detection removal.

Group DETR

Naive one-to-many assignment. We start from a naive way for one-to-many assignment depcited in Figure 2 (c). We replace one-to-one assignment with one-to-many assignment: assign one ground-truth object to multiple predictions. It does not work and the performance is much low. The reason is that the model is trained to output multiple predictions for one ground-truth object, and lacks the scoring mechanism to promote one single prediction and demote duplicate predictions for one ground-truth object.

Group-wise one-to-many assignment. We adopt the multi-group object query mechanism: form the initial NN queries as the primary group and introduce more (K−1)(K-1) groups of NN queries, totally KK groups, {Q1,Q2,…,QK}\{\mathbf{Q}_{1},\mathbf{Q}_{2},\dots,\mathbf{Q}_{K}\}. Accordingly, we have KK groups of predictions, {Y1,Y2,…,YK}\{\mathbf{Y}_{1},\mathbf{Y}_{2},\dots,\mathbf{Y}_{K}\}. We perform one-to-one assignment for each group, and find a bipartite matching σk(⋅)\sigma_{k}(\cdot), between each group of predictions and the ground-truth objects (Yk,Yˉ)(\mathbf{Y}_{k},\bar{\mathbf{Y}}). This results in that only one prediction for one ground-truth object is expected to score higher, and duplicate predictions is expected to score lower within one group other than within all the groups.

Separate self-attention. One-to-one assignment in one group means that the prediction assigned to one ground-truth object is superior to other predictions within the same group. This implies that we only need to collect the information of the predictions only from the same group, rather than from all the groups. Thus we perform self-attention (abbreviated as SA⁡\operatorname{SA}) over queries for each group separately:

Training architecture. The resulting architecture for training is very simple: the encoder keeps the same, and the decoder contains KK separate parallel decoders as shown in Figure 2 (a):

Here, the parameters of the decoder and the predictor for the KK groups are shared. Decoder separation and parallelism are feasible in that there is no interaction among queries for the other two operations, cross-attention and FFN. Our approach is called Group Decoder. In model inference, the process is the same as DETR trained normally and only needs one group of queries without any architecture modification. The pseudo-code is shown in Algorithm 1.

Loss function. The loss is an aggregation of KK losses, each for one decoder. It is written as follows,

where σk(⋅)\sigma_{k}(\cdot) is the optimal permutation of NN indices for the kkth decoder.

2 Analysis

Explanation with parameter-shared models. We discuss Group DETR from the perspective of training multiple models with parameter sharing. Training with Group DETR can be regarded as simultaneously training KK DETR models, which share the parameters of the encoder, the decoder, and the predictor, and only differ in the initialization of object queries. This leads to the shared parameters receive more back-propagated gradients. Thus, these parameters are better trained and accordingly the training process converges faster.

As a side benefit, we observe that Group DETR makes the assignment more stable, as shown in Figure 5. We speculate that the stability is because the improved network leads to more reliable predictions, and thus the assignment quality is better.

Explanation with object query augmentation. The multi-group object query mechanism introduces additional (K−1)(K-1) group of queries, which can be regarded as an augmentation of the primary group of queries. This is empirically illustrated in Figure 3. The reference points predicting the same objects are spatially close, and thus the corresponding object queries are similar. This may suggest that the multi-group object query mechanism resembles data augmentation, and at each iteration, more automatically-learned augmented queries are included, which equivalently introduces more supervision for decoder training. The results in Figure 4 empirically suggest that different groups of augmented queries lead to similar results.

The point about more supervision is also observed from the comparison between Equation 6 (for training with Group DETR) and Equation 2 (for normal DETR training). Group DETR training includes KK pairs of image feature and object query group {(X,Q1),(X,Q2),…,(X,QK)}\{(\mathbf{X},\mathbf{Q}_{1}),(\mathbf{X},\mathbf{Q}_{2}),\dots,(\mathbf{X},\mathbf{Q}_{K})\}, and thus the loss contains more components as shown in Equation 7.

Encoder training improvement. The additional supervision introduces more box regression and classification supervision from more queries assigned to each ground-truth object. The gradients with more supervision are also back-propagated from the decoder to the encoder. It is presumable that the encoder also gets benefit, verified by the empirical results in Table 1.

Computation and memory complexity. Group DETR uses more decoders during training. It is expected that Group DETR will bring additional training computation costs (FLOPs) as well as training memory costs. But the parallel decoders can be implemented as a single decoder by replacing normal self-attention with parallel self-attention (depicted in Figure 6) and we can use an efficient attention implementation, FlashAttention . As a result, Group DETR only takes a small increase in training GPU memory and training time. For example, with Conditional DETR and DAB-DETR , the memory increases are just 1.21.2 G and 1.71.7 G (Figure 7). The training time is increased by 55 minutes per epoch (from 2323 minutes to 2828 minutes and from 2828 minutes to 3333 minutes, respectively).

We provide the results by increasing the training time for normal DETR training to see if Group DETR benefits simply from more training time. The results given in Table 2 show that normal training with more training time brings a little benefit and the performance is still much lower than Group DETR, implying that the performance gain from our approach is not from training time increase.

Connection to DN-DETR. DN-DETR aims to stabilize one-to-one assignment during DETR training. DN-DETR forms the additional queries by adding the noises to ground-truth objects, which can be regarded as a variant of our multi-group mechanism with clear differences. In DN-DETR , on the one hand, the number of queries within each additional group is the same as the number of ground-truth objects. Each one correspond to one ground-truth object, and there is no query corresponding to no-object. In contrast, our approach automatically learns a number of NN (e.g., 300300) object queries that correspond to both ground-truth objects and no-object.

On the other hand, DN-DETR performs self-attention over noised queries, mainly for collecting the information from predictions for other objects other than from duplicate predictions. Self-attention in Group DETR instead collects both duplicate predictions and predictions for other objects.

The above two comparisons imply that DN-DETR brings the major help for the box and classification prediction, through the introduction of more positive queries corresponding to ground-truth objects (like FCOS), and no direct help for duplicate prediction removal. Our approach introduces both positive queries and negative queries (no-object), also brings the help for duplicate prediction removal.

Figure 8 shows that the performance of Group DETR is better than DN-DETR. We further investigate if Group DETR still benefits from introducing more positive queries with noised queries. As shown in Figure 8, the performance gain over Group DETR is non-trivial, a 1.51.5 mAP. This implies that Group DETR and DN-DETR are complementary and their major roles are different, though they have some similarities.

Experiments

We demonstrate the effectiveness of Group DETR in various DETR variants, and its extension to 33D detection and instance segmentation . The training setting is almost the same as baseline models, for illustrating the effectiveness of our Group DETR. We adopt the same training settings and hyper-parameters as the baseline models, such as learning rate, optimizer, pre-trained model, initialization methods, and data augmentationsWe may adjust the batch size due to the limitation of the GPU memory size for both the baseline model and our approach so that the batch size is the same..

Setting. We study various representative DETR-based detectors, such as basic baselines (Conditional DETR , DAB-DETR , DN-DETR ) with dense attentions, and strong baselines (DAB-Deformable-DETR and DINO ) with deformable attentions. We report the results on two training schedules, training for 1212 epochs and training for more epochs (3636 or 5050). Unless specified, the models are trained with ResNet-50 as the backbone on the COCO train2017 and evaluated on the COCO val2017. More implementation details are provided in Appendix.

Results. We first report the results of training with 1212 epochs in Table 3. Group DETR brings consistent improvements over the baselines with dense attentions that already are superior to the original DETR . It boosts Conditional DETR (-DC55) by 5.0\mathbf{5.0} (4.8\mathbf{4.8}) mAP, improves DAB-DETR (-DC55) by 3.9\mathbf{3.9} (4.4\mathbf{4.4}) mAP, and brings a 2.0\mathbf{2.0} (2.6\mathbf{2.6}) mAP gain to DN-DETR (-DC55) .

Group DETR also works well on those strong baselines with deformable attentions that are equipped with two or more accelerating techniques. It gives a 1.5\mathbf{1.5} mAP improvements over DAB-Deformable-DETR . When applying to DINO , Group DETR also exceeds it by 0.70.7 mAP. The gain is non-trivial over such a stronger baseline, considering that DINO is a well-tuned modelIn fact, our approach is compatible with query denoising and two-stage. In Table 33, for example, DN-DETR utilizes query denoising, and our method improves it by 2.02.0. Similarly, DAB-D-DETR adopts a two-stage structure, and our method achieves a 1.51.5 improvement. based on DAB-Deformable-DETR that combines improved hyper-parameters, improved two-stage design, improved query denoising task, and other tricks.

Furthermore, we report the results with 5050 training epochs that is commonly adopted in many acceleration methods . Table 4 presents that Group DETR outperforms baseline models by large margins. For the stronger backbone, Swin-Large , our approach achieves 58.458.4 mAP (still a 0.40.4 mAP higher than its baseline DINO (58.058.0 mAP with Swin-Large)). This verifies the generalization ability of our Group DETR.

Last, we compare the training convergence curves of the baseline models and their Group DETR counterparts. The results, as shown in Figure 1, provide more evidence that Group DETR speeds DETR training convergence on various DETR variants.

System-level Results on COCO test-dev with ViT-Huge. We also have the system-level performance on COCO test-dev with ViT-Huge . We apply Group DETR to DINO and follow its training pipeline and settings: pretrain the encoder with a self-supervised method, then pretrain the whole model on Object365 , and last fine-tune the whole model on COCO . Our model is the first to achieve 64.5\mathbf{64.5} mAP on COCO test-dev, which is still superior to other methods with larger encoder and more pre-training data . The details and comparisons with other methods are provided in Appendix.

2 More Applications

Group DETR is applicable to DETR-style techniques to other vision problems. We report the results for two additional problems: multi-view 33D object detection and instance segmentation , to further demonstrate the effectiveness.

Multi-view 3D object detection. We report the results over PETR and PETR v22 on the nuScenes val dataset . Table 5 shows that Group DETR brings significant gains to PETR and PETR v22 with 2424 training epochs in terms of both the nuScenes Detection Score (NDS) and mAP scores.

We demonstrate the effectiveness of the representative method, Mask2Former . The results are given in Table 6. Group DETR achieves a 1.21.2 (0.30.3) mAPm gain with 1212 (5050) epochs.

3 Ablation Study

We conduct the ablation study by using Conditional DETR as the baseline. The CNN backbone is ResNet-50 , and the training epoch nubmer is 1212. The performances are evaluated on COCO val2017 . We mainly study the effects of the key design: group-wise one-to-many assignment, separate self-attention, and group number.

Table 7 shows how group-wise one-to-many (o2m) assignment and separate self-attention make contributions. In comparison to the baseline (a), group-wise o2m assignment improves the mAP score from 32.632.6 mAP to 34.834.8 mAP: with the gain 2.22.2. The separate self attention (Sep. SA) further gets a 2.82.8 mAP gain. In addition, we report naive one-to-many assignment. The results are very poor, which is reasonable in that there are duplicate predictions and there is a lack of scoring mechanisms for demoting them. The results suggest that both group-wise o2m assignment and separate self-attention are effective.

Group number. Figure 9 shows the influence of the number of groups KK in Group DETR. The detection performance improves when increasing the number of groups, and becomes stable when the group number reaches 1111. Thus, we adopt K=11K=11 by default in Group DETR in our experiments.

Related Works

There are two main lines for accelerating DETR training: modify cross-attention and stabilize one-to-one assignment. The two are complementary and can be combined to further boost the performance.

Modifying cross-attention. Cross-attention module aims to collect the information from the image features useful to classification and localization. Various methods are proposed to select the informative image regions more efficiently and effectively . For example, Deformable attention selects the highly informative positions dynamically according to the previous decoder embedding. Conditional DETR instead continues to use the normal global attention, and dynamically computes the spatial attention to softly select the informative regions. SMCA uses the Gaussian-like weight for spatial modulation.

Stabilizing one-to-one assignment. DETR relies on one-to-one assignment, where each ground-truth object is assigned to a single prediction through building a bipartite matching between the predictions and the ground-truth objects. DN-DETR finds the assignment process is unstable and attributes the slow convergence issue to the instabilities. Thus, DN-DETR introduces groups of noisy queries by adding noises to ground-truth objects, to stabilize the assignment, leading to faster convergence. DINO makes further improvement through contrastive denoising training to generate both positive and negative noise queries with different noise levels. Our approach studies the assignment mechanism instead for introducing more supervision.

One-to-many assignment is widely adopted in deep detectors , and has attracted a lot of interest . For example, Faster R-CNN and FCOS produce multiple positive anchors and pixels for each ground-truth object. In this paper, we investigate one-to-many assignment in a feasible manner for the end-to-end detector DETR.

Concurrent with our work, H\mathcal{H}-DETR also uses one-to-many assignment to speed up DETR training convergence. Our Group DETR and H\mathcal{H}-DETR are related, but different: (1) Group DETR introduces group-wise one-to-many assignment with separate self-attention with the same number of object queries in each group. H\mathcal{H}-DETR adopts hybrid assignments in two different groups: One group uses one-to-one assignment and another uses one-to-many assignment with more object queries. (2) All the decoders in Group DETR can be used for inference. But the additional decoder in H\mathcal{H}-DETR is not directly used and requires NMS for inference. (3) During training, our architecture introduces one parameter: the number of groups. In contrast, H\mathcal{H}-DETR introduces the number of additional queries and the number of additional positive queries.

DETA is another concurrent work with our Group DETR. DETA directly uses one-to-many assignment and brings NMS back to DETR frameworks. While our method provides group-wise one-to-many assignment and maintains end-to-end detection.

Conclusion

The key points in Group DETR include group-wise one-to-many assignment and parallel self-attention. The success stems from involving more groups of object queries as an addition to the primary group of object queries, and thus introducing more supervision. Group-wise assignment mechanism makes sure that the competition among predictions happens within each group separately, and separate self-attention eases the training, Thus, the NMS pose-processing is not necessary, and the inference process is kept the same as normally trained DETR and not dependent on the group design. Our approach is simple, easily implemented, and general.

Acknowledgements. This work is supported by the Sichuan Science and Technology Program (2023YFSY0008), National Natural Science Foundation of China (61632003, 61375022, 61403005), Grant SCITLAB-20017 of Intelligent Terminal Key Laboratory of SiChuan Province, Beijing Advanced Innovation Center for Intelligent Robots and Systems (2018IRS11), and PEK-SenseTime Joint Laboratory of Machine Vision.

Appendix

A More Details and Results

We perform the object detection and instance segmentation experiments on the COCO 20172017 dataset, which contains about 118118K training (train2017) images, 55K validation (val2017) images, and 2020K testing (test-dev) images. Following the common practice, we train our model on COCO train2017 and report the standard mean average precision (mAP) result (box mAP for object detection and mask mAP for instance segmentation) on the COCO val2017 dataset under different IoU thresholds (from 0.50.5 to 0.950.95) and object scales (small, medium, and large). We also report the result on COCO test-dev with a large foundation model (ViT-Huge ).

We perform multi-view 3D object detection experiments on the nuScenes dataset, which contains 10001000 driving sequences. There are 700700 for train set, 150150 for val set and 150150 for test set. We report the standard nuScenes Detection Score (NDS) and mean Average Precision (mAP) result on the nuScenes val set.

A.2 Implementation Details

Our Group DETR adopts multiple groups of object queries. Each group shares the same architectures and numbers of object queriesWhen applying Group DETR to DN-DETR and DINO , we add the corresponding query denoising task in each group to keep the same architecture with the original implementation.. It resembles data augmentation with automatically-learned object query augmentation and is also equivalent to simultaneously training parameter-sharing networks of the same architecture.

In one-stage DETR frameworks, including Conditional DETR , DAB-DETR , DN-DETR , and DAB-Deformable-DETR , we can easily implement Group DETR by adopting multiple groups of learnable object queries. While the situation is different in two-stage DETR frameworks, such as DINO . The initializations of object queries are dependent on the top-NN predicted boxes of the first stage. To make the object queries in multiple groups similar to each other, we construct multiple pairs of classification and regression prediction heads in the first stage, each pair of which provides initialization for the object queries in the corresponding group. As for model inference, we only need one pair of these prediction heads, the same as the original model.

A.3 More Results of DN-DETR

We conduct experiments with different numbers of denoising queries in DN-DETR . The results in Figure 10 suggest that increasing the number of denoising queries can not achieve further improvements and show unstable performances. The effects of denoising queries differ from the ones of Group DETR (Figure 8 in the main paper). We choose to use 100100 denoising queries in our experiments in Table 3 and Table 4 in the main paper by following the setting in the original paper . To make direct comparisons with DN-DETR , we report the best results across different numbers of denoising queries in Figure 10 (38.838.8 mAP).

A.4 Applying Group DETR to SAM-DETR series

We also apply Group DETR to another stream of work to accelerate DETR training, SAM-DETR and SAM-DETR++ . The results are given in Table 9. Improvements on SAM-DETR (gains: 3.13.1 mAP with 12e and 1.91.9 mAP with 50e) and SAM-DETR++ (gains: 2.22.2 mAP with 12e and 1.31.3 mAP with 50e) show that Group DETR is complementary to them as well.

B More Comparisons on COCO test-dev

To compare state-of-the-art results on COCO test-dev, we follow DINO to build our model with a large foundation model, ViT-Huge. We follow its training pipeline and settings: (i) pre-train and fine-tune the ViT-Huge on ImageNet-1K , (ii) pre-train the whole detector on Object365 for 2424 epochs with 6464 A100 GPUs, and (iii) finetune the detector on COCO for 2020 epochs with 3232 A100 GPUs. When pre-training the detector on Object365, we follow DINO to only leave the first 55k out of 8080k validation images as the validation set and add the other images to the training set. We also use other schemes when training the detector on Object365 and COCO, such as enlarging the image size to 1.5×1.5\times when finetuning and adopting test time augmentation. In addition, we apply the exponential moving average (EMA) technique , use CDN queries , and adopt 1111 groups with Group DETR during detector pre-training and fine-tuning. When fine-tuning the detector on COCO, we find that applying learning rate decay for the components of the detector gives a ∼\sim0.90.9 mAP gain on COCO. During testing, we adopt test time augmentation with various scales and their flipped counterparts and perform fusionAccording to our experiments, the fusion on the query features builds a robust feature across different scales and gives a ∼\sim0.80.8 mAP improvement. on the query features and the final predictions .

Results.

Table 8 shows the results. Our model is the first to achieve 64.5\mathbf{64.5} mAP on COCO test-dev. Only pre-training the ViT-Huge on ImageNet-1K , our model can outperform other methods with larger models (e.g., BEIT-3 and SwinV2-G ) and more pre-training data. Models such as EVA and InterImage-H , with larger foundation models (ViT-giant or InterImage-H ) and more data , give higher results (64.764.7 mAP and 65.465.4 mAP) than our model. We expect that our results will be further improved with more pre-training data and larger models.

References