Dense Distinct Query for End-to-End Object Detection

Shilong Zhang, Xinjiang Wang, Jiaqi Wang, Jiangmiao Pang, Chengqi Lyu, Wenwei Zhang, Ping Luo, Kai Chen

Introduction

Object detection is one of the most fundamental tasks in computer vision, which aims at answering what objects are in an image and where they are. To achieve the objective, the detector is expected to detect all objects and mark each object with only one bounding box.

Due to the complex spatial distribution and the vast shape variance of objects, detecting all objects is quite challenging. To solve the problem, traditional detectors first lay predefined dense grid queries1 to achieve a high recall. Convolutions with shared weights are then applied to quickly process dense queries in a sliding-window manner. At last, one ground truth bounding box is assigned to multiple similar candidate queries for optimization. However, the one-to-many assignment results in redundant predictions and thus requires extra duplicate-removal operations (e.g., non-maximum suppression) during inference, which causes misaligned inference with training and hinders the pipeline from being end-to-end (as shown in Fig. 1(a)).

This paradigm is broken by DETR , which assigns only one positive query to each ground truth bounding box (one-to-one assignment) to achieve end-to-end. This scheme requires heavy computation to refine queries and adopts self-attention to model interactions between queries to facilitate the optimization of one-to-one assignment, which unfortunately limits the number of queries. For example, DETR only initializes hundreds of learnable object queries. Therefore, compared to the densely distributed queries in conventional detectors, the sparse queries fall short in recall rate, as shown in Fig. 1(b).

Some recent works have also tried to integrate dense queries into one-to-one assignment . However, dense queries in end-to-end detectors face unique challenges. For example, our analysis shows that this paradigm would inevitably introduce many similar queries (potentially representing the same instance) and that it suffers difficult and inefficient optimization as similar queries are assigned opposite labels under one-to-one assignment.(Fig. 1.(c)).

Now that both sparse queries (low recall) and dense queries (optimization difficulty) under one-to-one assignment are sub-optimal, what are the expected queries in end-to-end object detection?

In this study, we demonstrate that the solution should be dense distinct queries (DDQ), meaning that the queries for object detection should be both densely distributed to detect all objects and also distinct from each other to facilitate the optimization of one-to-one label assignment. Guided by such a principle, we consistently improve the performance of various detector architectures, including FCN, R-CNN, and DETRs. For one-stage detectors composed of fully convolutional networks (FCN), we first propose a pyramid shuffle operation to replace heavy self-attentions to model the interaction between dense queries. Then, a distinct queries selection scheme ensures that the one-to-one assignment is only imposed on the selected distinct queries, preventing contradictory labels from being assigned to similar queries. Such an end-to-end one-stage detector is named DDQ FCN and achieves state-of-the-art performance in one-stage detectors. DDQ also naturally extends to DETR and R-CNN structures by first laying dense queries as in and then selecting distinct queries for later refining stages, which are respectively dubbed DDQ R-CNN and DDQ DETR.

We have conducted experiments on two datasets—MS-COCO and CrowdHuman . DDQ FCN and DDQ R-CNN obtain 41.5/44.6 AP, respectively, on the MS-COCO detection dataset with the standard 1x schedule. Compared to recent DETRs, DDQ DETR achieved 52.1 AP in just 12 epochs with the DETR-style augmentation . The strong performance demonstrates that DDQ overcomes the optimization difficulty in end-to-end detectors and converges as fast as traditional detectors with higher performance.

Object detection in crowded scenes such as CrowdHuman is another arena to testify to the effectiveness of DDQ. It is extremely cumbersome to tune the post-processing NMS in traditional detectors, as a low IoU threshold leads to missing overlapping objects, while a high threshold brings too many false positives. Recent end-to-end structures also struggle to distinguish between duplicated predictions and overlapping objects due to their difficult optimization. In this study, DDQ FCN/R-CNN/DETR achieve 92.7/93.5/93.8 AP and 98.2/98.6/98.7 recall on CrowdHuman , surpassing both traditional and end-to-end detectors by a large margin.

Related Work

Dense Queries with One-To-Many Assignment. One-stage detectors such as RetinaNet and FCOS use densely distributed queries for regression and classification.The same manner is also applied to the region proposal network (RPN) of multi-stage models . And one-to-many assignments are a common practice for these traditional detectors. Despite the fast development of one-to-many assignments from static label assignments (such as IOU-based and center-based ones ) to prediction-aware dynamic label assignments , these strategies are also long criticized for they pair each ground truth with multiple queries and thus require additional postprocessing to remove duplicate predictions at inference, which prevents the pipeline from being end-to-end. Sparse Queries with One-To-One Assignment. DETR designs a small set of learned positional embeddings that represent the position in an image to focus on. These queries are then optimized with one-to-one assignments, making an end-to-end pipeline. Sparse R-CNN reformulates queries in the traditional R-CNN framework as a bounding box and its corresponding embedding. Anchor DETR provides the correspondence between anchor points and query position. DAB-DETR explicitly learns a set of 4-D anchor boxes as queries. Though the formulation of queries varies, they share the same core idea of sparse queries and one-to-one assignments. Therefore, a low recall rate is an expected issue for these detectors. Dense Queries with One-To-One Assignment. Both DeFCN and OneNet try to integrate one-to-one assignment with dense queries. Despite their competitive performance compared to FCOS , there is still a clear performance gap with recent detectors with dynamic one-to-many assignment strategies . It is the optimization difficulty of similar queries under one-to-one assignments that accounts for the performance gap. Efficient DETR , and Two-Stage Deformable DETR can also be regarded as a multi-stage version of this paradigm. Although DINO , Group DETR , and H-DeformableDETR have introduced more positive samples to speed up convergence, the hindrance effect between similar queries and one-to-one assignments still remains unrevealed.

Analysis of Sparse and Dense Queries

Current end-to-end detectors use either dense or sparse queries, both of which are however problematic during training. Specifically, sparse queries suffer a low recall rate, and dense queries have issues in optimization. To illustrate this, we increase the number of queries from 10 to 7000 in Sparse R-CNN, and the performance is shown as the black line in Fig. 2. The performance first keeps rising as the number of queries increases to around 2000, implying that the sparse queries (∼\sim 300) in Sparse R-CNN are far from enough due to the low recall rate. On the other hand, the performance finally plateaus and even decreases as queries number further increases. This phenomenon can be explained by the difficulty in distinguishing similar queries in end-to-end detectors with one-to-one assignment, especially when queries become denser.

To understand how similar queries would hinder optimization, we provide a simplified example where we assume there exist two identical queries. In this case, the one-to-one assignment assigns a foreground label to one of them but a background label to another. Without loss of generality, we adopt binary cross-entropy loss for classification. Therefore, the loss from these two queries becomes L1=−log⁡(p1)−log⁡(1−p2)L_{1}=-\log(p_{1})-\log(1-p_{2}), where p1p_{1} and p2p_{2} are the probability scores of the positive and negative query, respectively, and satisfy p1=p2=pp_{1}=p_{2}=p as they are identical queries. In contrast, the loss value when only one of the duplicated queries exists is L0=−log⁡(p)L_{0}=-\log(p). The ratio α\alpha of the gradient between the duplicate and non-duplicate query is.

It is obvious that the gradient is scaled down (i.e., α<1\alpha<1) at 0<p<0.50<p<0.5 and may even cause negative training (i.e., α<0\alpha<0) at p>0.5p>0.5.

As shown in the toy example, duplicated queries reduce gradients and even cause negative training, which dramatically suppresses convergence. To avoid this issue, we impose a distinct queries selection operation before the one-to-one assignment process. The distinct queries selection strategy is realized by a simple class-agnostic NMS. The filtered distinct queries are thus easier to optimize, and such an operation improves the performance by a clear margin, as seen from the red curve in Fig. 2. More surprisingly, the performance margin consistently increases along with more queries. A similar trend is also observed for Deformable DETR, which can be found in the supplementary material.

In other words, once we make sure the selected queries are distinct, the performance of Sparse R-CNN can be improved consistently with the increasing number of distinct queries. However, using a large number of distinct queries causes a significant memory footprint. For example, Sparse R-CNN requires around 45G memory per GPU with 7000 queries. To leverage the advantage of dense distinct queries(DDQ) with a reasonable computation cost, we give practical designs for all popular detector architectures (FCN, R-CNN, and DETRs).

Method

Dense distinct queries (DDQ) is our principle for designing an object detector and can be integrated into different architectures. We first briefly describe the design of DDQ followed by detailed descriptions for the three architectures: FCN, R-CNN, and DETRs. The overall pipeline is sketched in Fig. 3.

Dense Queries. As shown in Fig. 2, the memory cost soars for dense queries. The main reason for this is the heavy calculation for each query. Instead of adopting learnable positional embedding in DETR, DDQ directly takes the feature point on each feature map as densely distributed initial queries. The number of queries in the feature pyramid can easily surpass 10000 given an input resolution of 800x800. To discriminate dense queries with reasonable computation cost, a light-weighted convolutional/linear network serves as the first stage and processes all queries in a sliding window manner.

Distinct Queries Now that the importance of query distinctness for optimization has been revealed in Sec. 2, we would discuss in this section why a class-agnostic non-maximum suppression (NMS) can be used to select distinct queries and how it differs from the traditional NMS as post-processing in traditional detectors. Since each query represents a potential instance in an image, and an instance can be uniquely represented by its location in an image , it comes naturally to detect similar queries using the class-agnostic overlapping ratio between the bounding boxes predicted by queries. More specifically, we apply a class-agnostic NMS to select distinct queries for the following one-to-one assignment. The loss is thus only computed on the selected distinct queries. It should be noted that such an operation is adopted in both training and inference, instead of only in inference as an extra post-processing in traditional detectors. Therefore, such a pipeline still abides by the definition of end-to-end detectors. Compared to the training-unaware NMS in traditional detectors, it is designed to relieve the burden of one-to-one assignment during training, and can thus be set with an aggressive IoU threshold (0.7 in DDQ FCN and DDQ R-CNN, 0.8 in DDQ DETR), which is robust even on CrowdHuman dataset . Such crowd scenes can not be properly handled by NMS as post-processing in traditional detectors. We validate this in Table. 4.

Loss Components (1). Main Loss for Dense Distinct Queries. We simply apply the bipartite matching algorithm in DETR with the same cost weight in the one-to-one assignment. No extra prior (such as center priors in ) is adopted for a fair comparison with DETRs. After discriminating positive and negative samples, DDQ FCN adopts GIoU loss and QFocal loss with weights 2 and 1. For DDQ R-CNN and DDQ DETR, we just follow the implementation of Sparse R-CNN and DINO .

(2). Auxiliary Loss for Dense Queries. Despite the more efficient optimization in DDQ due to the removal of similar queries, it also results in numerous ”leaf” queries through which no gradients are back-propagated. Therefore, we design an auxiliary head and an auxiliary loss to further harness the potential of the filtered queries following the design in DeFCN . The auxiliary head is mostly identical to the main head, except that it adopts a soft one-to-many assignment for dense queries to allow for dense gradients and more positive samples to speed up training. More details can be found in our supplementary material.

2 DDQ FCN

As shown in Fig. 3. (a), the DDQ principle is first applied to FCOS as an example of the FCN structure for object detection. It is found that dense queries are already available on the dense feature pyramid. However, as the dense queries are processed level by level with convolutional layers. The missing interaction across different levels poses a challenge for the optimization of one-to-one assignments.

Inspired by channel shuffle operation in ShuffleNet , we propose a pyramid shuffle to compensate for the interaction between queries in different levels where SS channels across adjacent levels are shuffled to form a new feature pyramid. Specifically, features at level ii exchange SS channels with those at level i−1i-1 and level i+1i+1 simultaneously. To account for the different spatial dimensions on the feature pyramid, a bilinear interpolation is adopted when exchanging features. We apply the pyramid shuffle operation on the last two and one convolution layer in the classification and regression branches, respectively. This approach stabilizes training and improves performance with negligible additional computation costs. In this work, we set SS to 64 which means each feature level exchanges information from 128 channels with other levels. (Comparison with other approaches to model the interaction among dense queries, ablation, and the analysis of pyramid shuffle can be found in our supplementary material.)

As for the distinct queries selection module in DDQ FCN, we first select the top 1000 predictions according to the classification score from each feature level and then apply a class-agnostic non-maximum suppression with a threshold of 0.7 to ensure both distinctness and generality across different datasets.

3 DDQ R-CNN

We combine DDQ FCN with two refine stages in Sparse R-CNN to construct the DDQ R-CNN. As shown in Fig. 3. (b), thanks to the fast processing of dense distinct queries in DDQ FCN, we select 300 most representative queries according to the classification score from the remaining distinct queries. Then we concatenate the feature in the distinct position of the last feature map of the classification branch and regression branch to construct the query embedding. The query embedding and the corresponding bounding box prediction will be passed to the refinement head of Sparse R-CNN. Different from Sparse R-CNN which requires 6 stages of iterative query refinement, DDQ R-CNN needs as few as 2 refinement stages. Actually, the long iteration stages in Sparse R-CNN mainly compensate for the drawbacks caused by the sparse and sometimes similar input queries. For one thing, sparse queries could not cover all instances at initialization and thus need long cascading stages to refine. For another, similar queries also require long refinements to distinguish from each other to output a one-hot prediction for each instance . In contrast, the dense distinct queries from DDQ R-CNN have addressed the above issues, and hence the number of iterative refinements can be significantly reduced. We also report the results when we change the number of queries and refinement heads of DDQ R-CNN in the supplementary material.

4 DDQ DETR

We construct DDQ DETR based on Deformable DETR* . As shown in Fig. 3. (c), We follow Two-Stage Deformable DETR to process dense queries. Instead of initializing the content part with transformed coordinates, we fuse the feature map embedding of distinct positions as the content part, which makes the initial queries more distinct. A class-agnostic NMS with a threshold of 0.8 is set to select distinct queries before each refining stage. To compare with recent DETRs, we keep the original 6 refining stages and select KK distinct queries for the refining stages. We also select the top 1.5KK queries directly according to classification scores as dense queries for the auxiliary head in the decoder. The parallel forward of dense queries and distinct queries follows the H-DeformableDETR and Group DETR. We set KK to 900, following DINO .

Experiments

In this section, we first introduce two standard benchmarks MS COCO and CrowdHuman . Then we introduce the setting of training and inference on both datasets. We also present three examples to show how end-to-end detectors with different architectures evolve to our DDQ step by step. At last, we compare DDQ with state-of-the-art conventional detectors and recent end-to-end detectors on MS COCO and CrowdHuman, which show that DDQ blends the advantages of two design paradigms. The latency of current popular models and DDQ is compared in our supplementary material.

MS COCO 2017 detection dataset is mainly used for comparison and ablation studies. It contains 118k training, 5k validation images, and 20k test images without annotations. There are on average 7 instances per image in this dataset. We report bounding box mean average precision (AP) as the performance metric, which is the mean average precision over multiple thresholds. If not specified, AP on the validation set is set as default.

Besides, we also report the performance on the CrowdHuman dataset , which has 15k training images and 4.4k validation images with around 23 heavily occluded instances per image. For evaluation, we use AP, mMR, and Recall as the metrics. mMR means the average log miss rate over false positives per image ranging in [10−2,100]\left[10^{-2},10^{0}\right] following the official report . A lower value of mMR means a better quality of high-scoring bounding boxes. All evaluation results are reported on the CrowdHuman validation subset.

2 Setting

COCO ResNet-50 is the default backbone in this study if not specified. Most models adopt the 1x(12 epochs) training protocol in MMDetection . AdamW optimizer is used. For DDQ FCN, we set the initial learning rate to 5×10−55\times 10^{-5} and weight decay to 0.1. For DDQ R-CNN, we used a learning rate of 10−410^{-4} and weight decay of 0.05. The learning rate for both two CNN-based detectors decayed with a ratio of 0.1 at epoch 9 and epoch 12. For DDQ DETR, we utilized a learning rate of 2×10−42\times 10^{-4} and a weight decay of 0.05, and the learning rate decayed with a ratio of 0.1 at epoch 12 only. To ensure a fair comparison with other studies, we classified our data augmentation into three types : Normal, Multi-Scale, and DETR. Normal augmentation rescaled images to a short side of 800 pixels, with only random flips applied. For Multi-Scale augmentation, we used the classic multi-scale training range (480–800). Finally, the DETR augmentation followed that of the study by Carion et al. .

CrowdHuman ResNet-50 is the default backbone. All conventional detectors and DDQ adopt the 3x (36 epochs) schedule with multi-scale training (480-800). All optimizer-related parameters are consistent with the setting on COCO. For end-to-end detectors with sparse queries, we follow the schedule(50 epochs) in Sparse R-CNN due to their slow convergence. The max detected instance number is changed to 500 for all conventional detectors following the . For a fair comparison, We increase the number of queries to 500 for DDQ FCN/R-CNN, Sparse R-CNN, and Deform DETR. For DDQ DETR, we just keep the same 900 queries as COCO.

3 Evolving to DDQ

In this section, we show how detectors of different architectures evolve to DDQ. We can validate the importance of both density and distinctness from such a progressive development.

From FCOS* to DDQ FCN In Table. 1, We start from an FCOS equipped with the bipartite matching algorithm in DETR and our main loss components mentioned in Sec. 3, which is denoted as FCOS*. We adopt the normal augmentation that is mentioned in 5.2 and train the model for 12 epochs. Due to the lack of cross-level interaction, its performance is quite unstable and fluctuates between 24.5 AP and 36.5 AP in a few successful experiments. We select the best result 36.5 as our baseline. After adding pyramid shuffle operations to interact with cross-level queries, the training becomes stable and gets 1.1 AP improvement with only 0.2 G flops and 0.2 ms latency increase. Adding a distinct queries selection operation boosts the performance from 37.6 AP to 40.6 AP with only 0.3 ms latency. Such a 3 AP improvement demonstrates that the distinctness of queries is vital for the one-to-one assignment. After adding an auxiliary loss for dense queries following DeFCN , we get DDQ FCN with a state-of-the-art performance of 41.5 AP. DQS on the strong baseline (equipped with pyramid shuffle and auxiliary loss) can be found in Table. 14. DQS still improves 2 AP (from 39.5 to 41.5).

From Sparse R-CNN to DDQ R-CNN Table. 2 shows a progressive development from Sparse R-CNN to DDQ R-CNN. Sparse R-CNN with 300 queries achieves 39.4 AP within 12 epochs using the normal augmentation that is mentioned in Sec.5.2. Increasing the number of queries to 7000 improves the performance to 40.6 AP, at the cost of a quite heavy detector. Applying a distinct queries selection at the beginning of each stage boosts the performance by 2.5 AP to 43.1 AP. At last, by replacing the first four refining stages with our DDQ FCN, which not only makes the structure more light-weighted but also allows even denser input queries, the performance further increases to 44.6 AP.

From Deformable DETR* to DDQ DETR Table 3 illustrates the progressive development from Deformable DETR* to DDQ DETR. Deformable DETR* achieves 45.4 AP with 900 queries within 12 epochs using the DETR augmentation mentioned in Sec.5.2. By employing a linear layer to process the dense queries on the feature pyramid and constructing content parts with feature embeddings, the performance increases to 48.5 AP. However, initializing the content part as Two-Stage Deformable DETR(TS D-DETR) with mapped coordinates only achieves 46.7 AP, which is due to the lack of distinctness in the coordinates compared to the feature embedding. Adding an auxiliary loss for the decoder improves performance to 50.0 AP. Furthermore, by adding DQS before each refining stage, the performance further increases to 50.7 AP. Finally, by adding the P2 feature and 100 CDN queries as in DINO , we achieve an impressive 52.1 AP, surpassing all detectors in the same setting. We show distinctness can be complementary to CDN in Table. 14 and analyze the reason in our supplementary material.

4 Comparison with Other Detectors

Results on CrowdHuman We select some recent representative studies for comparison with DDQ on crowded scenes. It is seen that traditional detectors struggle between a low recall rate and a high false positive rate. Although DW assignment is the recent state-of-the-art one-to-many assignment strategy and shows a clear increase in Recall compared to ATSS, it suffers from more serious false predictions and thus leads to a high mMR. The performance of such traditional detectors is limited by the post-processing NMS. In the supplementary material, we also show it can not be properly handled by adjusting the IoU threshold because it is training unaware.

End-to-end detectors can achieve a higher theoretical recall rate due to the removal of NMS as a post-process. However, a high recall is not guaranteed in Sparse R-CNN and Deformable DETR due to their sparse query design. Although DeFCN achieves a better performance than other end-to-end methods by adopting dense queries, it is still difficult for DeFCN to distinguish between crowded objects and duplicated predictions(optimization difficulty) which affects the mMR.

In contrast, DDQ surpasses these detectors on all metrics by a clear margin. For one thing, DDQ leads in Recall due to the dense queries that could cover most objects. For the other, DDQ also achieves the lowest mMR, as a merit of the distinctness among queries so that the detector can better differentiate false predictions.

Results on COCO We adopt heavier backbones and longer schedules to fairly compare with other detectors on COCO. As shown in Table. 5, we get all the results from the original study except those marked with *. We divide the results into two parts according to the augmentation. The first part adopts the augmentations in DETR and reports the results on COCO validation dataset. DDQ remains its advantage among end-to-end object detectors using different backbone structures. It is worth emphasizing that DDQ FCN without any refinement architecture can already surpass most end-to-end detectors. DDQ R-CNN surpasses these methods by a large margin with only two refining heads and without encoder architecture. The performance of DDQ R-CNN (R-50) can be further improved by adopting an encoder structure as in SEPC or DyHead . For example, It achieves an impressive 51.0 AP by adopting 6 blocks in DyHead as encoder structure, which is denoted as DDQ R-CNNwith_encoder(details about this model can be found in supplementary material). DDQ DETR outperforms recent DETR with a clear margin using R-50 as its backbone. When adopting a Swin-L backbone, it also surpasses the SOTA method DINO by 0.7 AP.

The second part adopts a multi-scale training (480-800) strategy for 24 epochs and reports the results on the COCO test-dev using ResNet-101, which is widely used by conventional detectors.

Ablation study

We analyze the recall of IoU threshold 0.5. As shown in Table. 6, we report the recall of the 5th stage input queries of Sparse R-CNN to make a fair comparison with the input queries of the refinement head in DDQ R-CNN. It can be seen that the Sparse R-CNN with 300 queries has a significantly lower recall(10.2 AR100) than that with 7000 queries. In DDQ R-CNN, the queries from the DDQ FCN achieve a comparable recall to 7000 queries but with much less latency.

2 DQS with Different IoU Threshold

In this section, we show the robustness of distinct queries selection(DQS) with different IoU thresholds. As shown in Table. 14, the performance of DDQ FCN/R-CNN is quite robust when the IoU threshold ranges from 0.6 to 0.8. The performance drops slightly when the threshold is lower than 0.6, which is due to the lower recall rate for overlapping objects. The performance also starts to degrade when the IoU threshold is larger than 0.8 due to its incapability to suppress similar queries that slow the optimization. DDQ DETR exhibits a similar trend, as observed in Table 3. We can find even though CDN training in has been adopted in DDQ DETR, distinctness still improves the performance. By the way, we also report the performance of ATSS at different post-processing NMS IoU thresholds and show its sensitivity to this hyperparameter, in contrast to the robust behavior of DDQ in which a class-agnostic NMS is adopted in both training and inference to filter out distinct queries.

Acknowledgement

This project is supported by the National Key R&D Program of China No.2022ZD0161600, No.2022ZD0161000 and the General Research Fund of HK No.17200622.

Conclusions

This paper reveals that both sparse and dense queries in end-to-end detection are problematic. We propose that the expected queries should be both dense and distinct. Such a paradigm significantly improves the performance of various detectors including FCN, R-CNN, and DETRs. This proves the paradigm blends advantages from the traditional detectors and the recent end-to-end detectors. We hope it can inspire researchers to consider the complementarity between traditional methods and end-to-end detectors.

References

Appendix A Analysis of DDQ in Deformable DETR

For fast verification of Distinct Queries Selection(DQS) in such a heavy model, we adopt the standard 1x setting on COCO and keep other hyperparameters (such as learning rate and weight decay) the same with that in Deformable DETR .

As shown in Fig.5, when we increase the number of queries in Deformable DETR, there is a similar trend as it in Sparse R-CNN , as shown in the main manuscript. When the number of queries naively increases without distinct queries selection, the performance increases at the beginning but decreases as the number of queries reaches ∼5000\sim 5000. It is due to the more difficult training with more similar queries out of the dense queries. By imposing a distinct queries selection pre-processing to filter out similar queries and keeping only distinct queries before each stage of iterative refinement, the performance is improved with a clear margin, and the performance margin consistently increases along with more queries.

Therefore, we believe Dense Distinct Queries (DDQ) is a principle of designing an object detector with a fast convergence based on recent end-to-end detectors.

Appendix B Auxiliary Loss for Dense Queries

We follow the TOOD to design our auxiliary loss. We select KK samples with the smallest cost of each ground truth as positive samples. PP means the index set of positive samples which correspond to the same ground truth. The classification score target of a sample ii in this set is

The GIoU loss of each sample is reweighted by the classification target. The classification loss and regression loss weight keep consistent with the main loss weight (1 and 2 respectively) for distinct queries. The performance is quite stable for DDQ FCN when KK ranges between 5 and 16. We adopt 8 and 4 for DDQ FCN and DDQ DETR respectively in this study. The auxiliary loss also works for refining heads in DDQ RCNN. Due to time issues, we will supplement relevant results in the future version.

Appendix C More Analysis and Ablation for Pyramid Shuffle

We provide an in-depth analysis of pyramid shuffle. Firstly, we compare it with 3D MAX Filter in DeFCN in Table. 9. Then we visualize the change of score maps in different levels after adding pyramid shuffle operations in Fig. 6. At last, Table. 11 gives the results under different shuffle channels.

Comparision with 3D MAX Filter Table. 9 shows the comparison between pyramid shuffle and 3D MAX Filtering in DeFCN. DeFCN believes there should be extra parameters and max pooling operation to facilitate the optimization under the one-to-one assignment. However, we argue that only the interaction of cross-level queries matters, and extra parameters or max operations are unnecessary. Our pyramid shuffle is more lightweight and with better performance. When we add an extra convolution to regression branches to fair compare with the 3D MAX Filter, we can surpass it by 0.8 AP.

Visualization of Adjacent Level Score Map Fig. 6 shows the scores map of adjacent levels. The left side of each subfigure(with blue background) shows score maps with pyramid shuffles, and the corresponding right side (with yellow background) means without pyramid shuffles. The top-left corner marks the feature level. The red circle represents the duplication predictions in the adjacent level. We can find pyramid shuffle effectively reduces the cross-level high score false positives.

Results under Different Shuffle Channels Table 11 gives the results under different shuffle channels; we can find that the number of channels can even be reduced to 16 when there is already cross-level distinct queries selection, making it more lightweight. When the number of shuffle channels is 128, which means no remaining channels for the current level, the performance will dramatically drop 1.5 AP because of missing information on the current level queries in the interaction.

We report the results of DDQ FCN with the different number of pyramid shuffle operations in the classification and regression branches. When no pyramid shuffle is adopted in the DDQ FCN, its performance is unstable and fluctuates between 40.8 AP and 41.1 AP. We report an average performance of 41.0 AP. Even though there has been a cross-level distinct queries selection operation, compared to adopting only 2 and 1 operations to two branches respectively, there is still a 0.5 AP drop.

Appendix D DDQ R-CNN with Encoder

The encoder can provide a more powerful feature representation for the decoder head which is explored in . We simply add 6 dynamic blocks in DyHead as our encoder. 500 queries and 3 refinement stages are adopted. It is observed in the main manuscript that there is about 3 AP improvement for DDQ R-CNN.

Appendix E Details of Improved DeformableDETR and Comparison with DINO

DINO adopt some techniques that significantly improve the Deformable DETR. We remove the CDN and mix query selection from DINO to form our baseline. DDQ is a concurrent work of DINO. The contrastive denoising training (CDN) is not intended to relieve the optimization difficulty of very similar queries among dense queries. In their implementation, the generated positive and negative samples in each pair are always significantly distinct from each other. Mix query selection increases the distinctness of queries by additionally initializing content embeddings, but the position embeddings are still created from top-k dense regression predictions which can be very similar and still hinder the optimization. We have shown our components can be combined with DINO.

Appendix F Number of Queries and Stages in DDQ R-CNN

We analyze the combination of different numbers of stages and queries for DDQ R-CNN. Table. 12 shows that the best number of stages is proportional to the number of queries. This is easy to understand. When the number of queries increases, the newly added queries are of low quality, and more stages are needed to refine these queries. It is worth emphasizing that we use 2 stages and 300 queries to trade off the performance and latency. When using 3 stages with the same number of queries, our method achieves an even higher performance of 45.1 AP on MS COCO.

Appendix G Latency Benchmark

As for the details of latency calibration, We compare the speed (forwarding + post-processing) of different methods with batch size 1 in Table. 8. All evaluations were performed on Tesla A100 GPU with Intel(R) Xeon(R) Gold 6348 CPU @ 2.60GHz. The Pytorch version is 1.9.0 with CudaToolkit 11.1 and Cudnn 8.0.5. An average of 200 iterations during model inference is adopted as the latency reported in this study.

Appendix H Impact of Query Construction in DDQ R-CNN

We try five ways to construct queries in DDQ R-CNN. As shown in Table. 13, None means all query embeddings are set to a zero tensor, and the refinement stages only get meaningful query bounding boxes. This attempt reduces the performance to 43.2 AP. Simply constructing queries from the FPN results in 1.0 AP degradation. Reg means only using the last feature map of the regression branch, which drops the performance by 0.6 AP. Constructing queries from the last feature map in the classification branch can be an alternative as it can get a comparable performance(only 0.3 AP drop).

Appendix I DQS with Different IoU Threshold in CrowdHuman

In this section, we show the robustness of distinct queries selection(DQS) with different IoU thresholds in CrowdHuman. We can find there is a clear performance bottleneck for traditional detector ATSS even with carefully adjusting the threshold of NMS.

Appendix J Social Impact

The potential social impact of this work inherits from object detection. Because human behaviors often cause crowded scenes, and DDQ achieves excellent performance in such scenes, it may be applied to some applications that violate human privacy, such as surveillance.