DN-DETR: Accelerate DETR Training by Introducing Query DeNoising

Feng Li, Hao Zhang, Shilong Liu, Jian Guo, Lionel M. Ni, Lei Zhang

Introduction

Object detection is a fundamental task in computer vision that aims to predict the bounding boxes and classes of objects in an image. While having made remarkable progress, classical detectors were mainly based on convolutional neural networks, until Carion et al. recently introduced Transformers into object detection and proposed DETR (DEtection TRansformer).

In contrast to previous detectors, DETR uses learnable queries to probe image features from the output of Transformer encoders and bipartite graph matching to perform set-based box prediction. Such a design effectively eliminates hand-designed anchors and non-maximum suppression (NMS) and makes object detection end-to-end optimizable. However, DETR suffers from prohibitively slow training convergence compared with previous detectors. To obtain a good performance, it usually takes 500500 epochs of training on the COCO detection dataset, in contrast to 1212 epochs used in the original Faster-RCNN training.

Much work has tried to identify the root cause and mitigate the slow convergence issue. Some of them address the problem by improving the model architecture. For example, Sun et al. attributed the slow convergence issue to the low efficiency of the cross-attention and proposed an encoder-only DETR. Dai et al. designed an RoI-based dynamic decoder to help the decoder focus on regions of interest. More recent works propose to associate each DETR query with a specific spatial position rather than multiple positions for more efficient feature probing . For instance, Conditional DETR decouples each query into a content part and a positional part, enforcing a query to have a clear correspondence with a specific spatial position. Deformable DETR and Anchor DETR directly treat 2D2D reference points as queries to perform cross-attention. DAB-DETR interprets queries as 44-D anchor boxes and learns to progressively improve them layer by layer.

Despite all the progress, few works pay attention to the bipartite graph matching part for more efficient training. In this study, we find that the slow convergence issue also results from the discrete bipartite graph matching component, which is unstable especially in the early stages of training due to the nature of stochastic optimization. As a consequence, for the same image, a query is often matched with different objects in different epochs, which makes optimization ambiguous and inconstant.

To address this problem, we propose a novel training method by introducing a query denoising task to help stabilize bipartite graph matching in the training process. Since previous works have shown effectiveness in interpreting queries as reference points or anchor boxes , which contain positional information, we follow their viewpoint and use 44D anchor boxes as queries. Our solution is to feed noised GT bounding boxes as noised queries together with learnable anchor queries into Transformer decoders. Both kinds of queries have the same input format of (x,y,w,h)(x,y,w,h) and can be fed into Transformer decoders simultaneously. For noised queries, we perform a denoising task to reconstruct their corresponding GT boxes. For other learnable anchor queries, we use the same training loss and bipartite matching as in the vanilla DETR. As the noised bounding boxes do not need to go through the bipartite graph matching component, the denoising task can be regarded as an easier auxiliary task, helping DETR alleviate the unstable discrete bipartite matching and learn bounding box prediction more quickly. Meanwhile, the denoising task also helps lower the optimization difficulty because the added random noise is usually small. To maximize the potential of this auxiliary task, we also regard each decoder query as a bounding box + a class label embedding so that we are able to conduct both box denoising and label denoising.

In summary, our method is a denoising training approach. Our loss function consists of two components. One is a reconstruction loss and the other is a Hungarian loss which is the same as in other DETR-like methods. Our method can be easily plugged into any existing DETR-like method. For convenience, we utilize DAB-DETR to evaluate our method since their decoder queries are explicitly formulated as 44D anchor boxes (x,y,w,h)(x,y,w,h). For DETR variants that only support 22D anchor points such as anchor DETR , we can do denoising on anchor points. For those that do not support anchors like the vanilla DETR , we can do linear transformation to map 44D anchor boxes to the same latent space as for other learnable queries.

To the best of our knowledge, this is the first work to introduce the denoising principle into detection models. We summarize our contribution as follows:

We design a novel training method to speed up DETR training. Experimental results show that our method not only accelerates training convergence but also leads to a remarkably better training result — achieving the best result among all detection algorithms in the 1212-epoch setting. Moreover, our method shows a remarkable improvement (+1.9+\textbf{1.9} AP) over our baseline DAB-DETR and can be easily integrated into other DETR-like methods.

We analyze the slow convergence of DETR from a novel viewpoint and give a deeper understanding of DETR training. We design a metric to evaluate the instability of bipartite matching and verify that our method can effectively lower the instability.

We conduct a series of ablation studies to analyze the effectiveness of different components of our model, such as noise, label embedding, and attention mask.

This paper is an extension of our previous paper that was accepted to CVPR’2022 as an oral presentation. Compared with its conference version, this paper brings some new contributions as follows.

We achieve better results and faster convergence by introducing deformable attention into our decoder layer.

We further demonstrate the effectiveness of denoising training by adding it to other DETR-like models without 4D anchor design, including Vanilla DETR without explicit anchors and Anchor DETR with only 2D anchors. We also show denoising training can improve segmentation models such as Mask2Former and Mask DINO.

We incorporate denoising training to the traditional CNN detector Faster R-CNN to show its generalization ability.

We provide more experimental results and analysis to get a better understanding of our method.

Related Work

Most modern object detection models are based on convolutional networks, which have achieved significant success in recent years. Classical CNN-based detectors can be divided into 22 categories, one-stage, and two-stage methods. Two-stage methods like HTC and Fast R-CNN first generate some region proposals and then decide whether each region contains an object and do bounding box regression to get a refined box. Ren et al. proposed an end-to-end method that utilizes a Region Proposal Network to predict anchor boxes. In contrast to two-stage methods, one-stage methods, including YOLO900 and YOLOv3 directly predict the offset of real boxes relative to anchor boxes.

Though these methods achieve top performance on many datasets, they are sensitive to the way how anchors are generated. In addition, they require some hand-crafted components like non-maximum suppression (NMS) and label assignment rules. Therefore, they suffer from these drawbacks and can not be end-to-end optimized.

2 DETR-based Detectors

Carion et al. proposed an end-to-end object detector based on Transformers named DETR (DEtection TRansformer) without using anchors. While DETR achieves comparable results with Faster-RCNN , its training suffers severely from the slow convergence problem — it needs 500500 epochs of training to obtain a good performance.

Many recent works have attempted to speed up the training process of DETR. Some find the cross attention of Transformer decoders in DETR inefficient and make improvements in different ways. For example, Dai et al. [Dai_2021_ICCV] designed a dynamic decoder that can focus on regions of interest in a coarse-to-fine manner and lower the learning difficulty. Sun et al. discarded the Transformer decoder and proposed an encoder-only DETR. Another series of works make improvements in decoder queries. Zhu et al. designed an attention module that only attends to some sampling points around a reference point. Meng et al. decoupled each decoder query into a content part and a position part and only utilized the content-to-content and position-to-position terms in the cross-attention formulation. Yao et al. utilized a Region Proposal Network (RPN) to propose top-KK anchor points. DAB-DETR uses 44-D box coordinates as queries and updates boxes layer by layer in a cascade manner.

Despite all the progress, none of them treats bipartite graph matching used in the Hungarian loss as the main reason for slow convergence. Sun et al. analyzed the impact of Hungarian loss by using a pre-trained DETR as a teacher to provide the GT label assignment for a student model and train the student model. They found that the label assignment only helps the convergence in the early stage of training but does not influence the final performance significantly. Therefore, they concluded that the Hungarian loss is not the main reason for the slow convergence. In this work, we give a different analysis with an effective solution that leads to a different conclusion.

We adopt DAB-DETR as the basic detection architecture to evaluate our training method, where the label embedding appended with an indicator is used to replace the decoder embedding part to support label denoising. The difference between our method and other methods is mainly in the training method. In addition to the Hungarian loss, we add a denoising loss as an easier auxiliary task that can accelerate training and boost performance significantly. Chen et al. augments their sequence with synthetic noise objects, but is totally different from our method. They set the targets of noise objects to the ”noise” class (not belonging to any ground-truth classes) so that they can delay the End-of-Sentence (EOS) token and improve the recall. In contrast to their method, we set the target of noised boxes to the original boxes, and the motivation is to bypass bipartite graph matching and directly learn to approximate ground-truth boxes.

We are pleased to see that many very recent detection models adopt our proposed denoising training to accelerate convergence for detection and segmentation models, such as DINO , Mask DINO , Group DETR , and SAM-DETR++ . DINO further develops our denoising training by feeding hard-negative samples and training the model to reject them. Therefore, the proposed Contrastive Denoising (CDN) further improves the performance. Mask DINO extends denoising to three image segmentation tasks (instance, panoptic, and semantic) by reconstructing masks from noised boxes. Group DETR and SAM-DETR+++ also adopt denoising training in their model to achieve better performance. These models demonstrate the effectiveness and generalization capabilities of our methods.

Why Denoising accelerates DETR training?

Hungarian matching is a popular algorithm in graph matching. Given a cost matrix, the algorithm outputs an optimal matching result. DETR is the first algorithm that adopts Hungarian matching in object detection to solve the matching problem between predicted objects and ground-truth objects. DETR turns ground-truth assignment into a dynamic process, which brings in an instability problem due to its discrete bipartite matching and the stochastic training process. There are works showing that Hungarian matching does not result in stable matching since blocking pairs exist. A small change in the cost matrix may cause an enormous change in the matching result, which will further lead to inconsistent optimization goals for decoder queries.

We view the training process of DETR-like models as two stages, learning “good anchors” and learning relative offsets. Decoder queries are responsible for learning anchors as shown in previous works and . The inconsistent update of anchors can make it difficult to learn relative offsets. Therefore, in our method, we leverage a denoising task as a training shortcut to make relative offset learning easier, as the denoising task bypasses bipartite matching. Since we interpret each decoder query as a 44-D anchor box, a noised query can be regarded as a “good anchor” which has a corresponding ground-truth box nearby. The denoising training thus has a clear optimization goal - to predict the original bounding box, which essentially avoids the ambiguity brought by Hungarian matching.

To quantitatively evaluate the instability of the bipartite matching result, we design a metric as follows. For a training image, we denote the predicted objects from Transformer decoders as Oi={O0i,O1i,...,ON−1i}\mathbf{O^{i}}=\left\{O_{0}^{i},O_{1}^{i},...,O_{N-1}^{i}\right\} in the ii-th epoch, where NN is the number of predicted objects, and the ground-truth objects as T={T0,T1,T2,...,TM−1}\mathbf{T}=\left\{T_{0},T_{1},T_{2},...,T_{M-1}\right\} where MM is the number of ground-truth objects. After bipartite matching, we compute an index vector Vi={V0i,V1i,...,VN−1i}\mathbf{V^{i}}=\left\{V^{i}_{0},V_{1}^{i},...,V_{N-1}^{i}\right\} to store the matching result of epoch ii as follows.

We define the instability of epoch ii for one training image as the difference between its ViV^{i} and Vi−1V^{i-1}, which is calculated as

Fig. 3 shows a comparison of ISIS between our DN-DETR (DeNoising DETR) and DAB-DETR. We conduct this evaluation on the COCO 2017 validation set , which has 7.367.36 objects per image on average. So the largest possible ISIS is 7.36×2=14.727.36\times 2=14.72. Fig. 3 clearly shows that our method effectively alleviates the instability of matching.

2 Make Query Search More Locally

We also show that DN-DETR can help detection by reducing the distance between anchors and the corresponding targets. DETR shows from the visualization that its positional queries have several operating modes, which makes a query search from a wide region for a predicted box. However, DN-DETR has much smaller mean distances between initial anchors (positional queries) and targets. As shown in Fig. 4(a), we compute the mean l1\textit{l}_{1} distance between initial anchors and the matched ground-truth boxes in the last decoder layer for DAB-DETR and our model.

As denoising training trains the model to reconstruct boxes from the noised ones that are close to the ground truth, the model will search more locally for prediction, which makes each query focus on regions nearby and prevents potential prediction conflicts between queries. Fig. 4(b) and (c) are some examples of anchors and targets in DAB-DETR and DN-DETR. Each arrow starts from an anchor and ends with its matched ground-truth box. We use color to reflect the length of the arrows. The shortened distances between anchors and targets make the training process easier and therefore converge faster.

DN-DETR

We base on the architecture of DAB-DETR to implement our training method. Similar to DAB-DETR, we explicitly formulate the decoder queries as box coordinates. The only difference between our architecture and theirs lies in the decoder embedding, which is specified as class label embedding to support label denoising. Our main contribution is the training method as shown in Fig. 6.

Similar to DETR, our architecture contains a Transformer encoder and a Transformer decoder. On the encoder side, the image features are extracted with a CNN backbone and then fed into the Transformer encoder with positional encodings to attain refined image features. On the decoder side, queries are fed into the decoder to search for objects through cross-attention.

We denote decoder queries as q={q0,q1,...,qN−1}\mathbf{q}=\left\{q_{0},q_{1},...,q_{N-1}\right\} and the output of the Transformer decoder as o={o0,o1,...,oN−1}\mathbf{o}=\left\{o_{0},o_{1},...,o_{N-1}\right\}. We also use FF and AA to denote the refined image features after the Transformer encoder, and the attention mask derived based on the denoising task design. We can formulate our method as follows.

where DD denotes the Transformer decoder.

There are two parts to decoder queries. One is the matching part. The inputs of this part are learnable anchors, which are treated in the same way as in DETR. That is, the matching part adopts bipartite graph matching and learns to approximate the ground-truth box-label pairs with matched decoder outputs. The other is the denoising part. The inputs of this part are noised ground-truth (GT) box-label pairs which are called GT objects in the rest of the paper. The outputs of the denoising part aim to reconstruct GT objects.

In the following, we abuse the notations to denote the denoising part as q={q0,q1,...,qK−1}\mathbf{q}=\left\{q_{0},q_{1},...,q_{K-1}\right\} and the matching part as Q={Q0,Q1,...,QL−1}\mathbf{Q}=\left\{Q_{0},Q_{1},...,Q_{L-1}\right\}. So the formulation of our method becomes

To increase the denoising efficiency, we propose to use multiple versions of noised GT objects in the denoising part. Furthermore, we utilize an attention mask to prevent information leakage from the denoising part to the matching part and among different noised versions of the same GT object.

2 Intro to DAB-DETR

Many recent works associate DETR queries with different positional information. DAB-DETR follows this analysis and explicitly formulates each query as 4D anchor coordinates. As shown in Fig. 5(a), a query is specified as a tuple (x,y,w,h)(x,y,w,h), where x,yx,y are the center coordinates and w,hw,h are the corresponding width and height of each box. In addition, the anchor coordinates are dynamically updated layer by layer. The output of each decoder layer contains a tuple (Δx,Δy,Δw,Δh)(\Delta x,\Delta y,\Delta w,\Delta h) and the anchor is updated to (x+Δx,y+Δy,w+Δw,h+Δh)(x+\Delta x,y+\Delta y,w+\Delta w,h+\Delta h).

Note that our proposed method is mainly a training method that can be integrated into any DETR-like model. To test on DAB-DETR, we only add minimal modifications: specifying the decoder embedding as label embedding, as shown in Fig. 5(b).

3 Denoising

For each image, we collect all GT objects and add random noises to both their bounding boxes and class labels. To maximize the utility of denoising learning, we use multiple noised versions for each GT object.

We consider adding noise to boxes in two ways: center shifting and box scaling. We define λ1\lambda_{1} and λ2\lambda_{2} as the noise scale of these 22 noises. 1) center shifting: we add a random noise (Δx,Δy)(\Delta x,\Delta y), to the box center and make sure that ∣Δx∣<λ1w2|\Delta x|<\frac{\lambda_{1}w}{2} and ∣Δy∣<λ1h2|\Delta y|<\frac{\lambda_{1}h}{2}, where λ1∈(0,1)\lambda_{1}\in(0,1) so that the center of the noised box will still lie inside the original bounding box. 2) box scaling: we set a hyper-parameter λ2∈(0,1)\lambda_{2}\in(0,1). The width and height of the box are randomly sampled in [(1−λ2)w,(1+λ2)w]\left[(1-\lambda_{2})w,(1+\lambda_{2})w\right] and [(1−λ2)h,(1+λ2)h]\left[(1-\lambda_{2})h,(1+\lambda_{2})h\right], respectively.

For label noising, we adopt label flipping, which means we randomly flip some GT labels to other labels. Label flipping forces the model to predict the GT labels according to the noised boxes to better capture the label-box relationship. We have a hyper-parameter γ\gamma to control the ratio of labels to flip. The reconstruction losses are l1l_{1} loss and GIOU loss for boxes and focal loss for class labels as in DAB-DETR. We use a function δ(⋅)\delta(\cdot) to denote the noised GT objects. Therefore, each query in the denoising part can be represented as qk=δ(tm)q_{k}=\delta(t_{m}) where tmt_{m} is mm-th GT object.

Notice that denoising is only considered in training, during inference the denoising part is removed, leaving only the matching part.

4 Attention Mask

Attention mask is a component of great importance in our model. Without an attention mask, the denoising training will compromise the performance instead of improving it as shown in Table V.

To introduce an attention mask, we need first to divide the noised GT objects into groups. Each group is a noised version of all GT objects. The denoising part becomes

where gp\mathbf{g_{p}} is defined as the pp-th denoising group. Each denoising group contains MM queries where MM is the number of GT objects in the image. So we have

The purpose of the attention mask is to prevent information leakage. There are two types of potential information leakage. One is that the matching part may see the noised GT objects and easily predict GT objects. The other is that one noised version of a GT object may see another version. Therefore, our attention mask is to make sure the matching part cannot see the denoising part and the denoising groups cannot see each other as shown in Fig. 6.

We use A=[aij]W×W\mathbf{A}=\left[\mathbf{a}_{ij}\right]_{W\times W} to denote the attention mask where W=P×M+NW=P\times M+N. PP and MM are the number of groups and GT objects. NN is the number of queries in the matching part. We let the first P×MP\times M rows and columns represent the denoising part and the latter represents the matching part. aij=1a_{ij}=1 means the ii-th query cannot see the jj-th query and aij=0a_{ij}=0 otherwise. We devise the attention mask as follows

Note that whether the denoising part can see the matching part or not will not influence the performance, since the queries of the matching part are learned queries that contain no information about the GT objects.

The extra computation introduced by multiple denoising groups is negligible—when 55 denoising groups are introduced, GFLOPs for training are only increased from 94.494.4 to 94.694.6 for DAB-DETR with a ResNet-50 backbone, and there is no computation overhead for testing.

5 Label Embedding

The decoder embedding is specified as label embedding in our model to support both box denoising and label denoising. Except for the 8080 classes in COCO 2017 , we also consider an unknown class embedding that is used in the matching part to be semantically consistent with the denoising part. We also append an indicator to label embedding. The indicator is 11 if a query belongs to the denoising part and otherwise.

6 Compatibility with Deformable Attention Design

DN-Deformable-DETR: To show the effectiveness of denoising training applied in other attention designs, we also integrate denoising training into Deformable DETR as DN-Deformable-DETR. We follow the same setting as Deformable DETR but specify its query into 44D boxes as in DAB-DETR to better use denoising training. Note that this is our original deformable model in the conference version, in which we only add deformable attention to Transformer encoders.

When comparing in the standard 50 epoch setting, to eliminate any misleading information that the performance improvement of DN-Deformable-DETR may result from the explicit query formulation of anchor boxes, we also implement a strong baseline DAB-Defromable-DETR for comparison. It formulates the queries of Deformable DETR as anchor boxes without using denoising training, while all the other settings are the same. DN-Deformable-DETR++: We further incorporate the deformable attention in our decoder and optimize our model to build DN-Deformable-DETR++, which converges much faster and improves the final results. We also follow DAB-Defromable-DETR to build a strong baseline DAB-Defromable-DETR++ to show our performance improvement in the ablations.

7 Introducing DN to Other DETR-like models with different anchor formulations

In the aforementioned sections, we build DN-DETR upon DAB-DETR with explicit 4D anchor box formulation. As shown in Fig. 6, denoising is only a training method and can be plugged into other detection models to accelerate training. In this section, we will extend denoising training to other DETR-like models.

We first demonstrate its effectiveness by adding it to Anchor DETR , which formulates positional queries as 2D anchor points. For DN-Anchor-DETR, though it can be easily modified to 4D anchors to achieve better results, we strictly follow Anchor DETR to add noise only to 2D anchors. A 2D anchor corresponds to the center point of a box. Hence we only use center shifting noise (described in Sec. 4.3). In this way, we plug in the denoising training task for anchor points without introducing other modifications.

7.2 Introducing DN to Vanilla DETR without Explicit Anchors

Vanilla DETR differs from DAB-DETR in that its positional queries are high dimensional vectors without explicit meanings. For DN-Vanilla-DETR, we can simply use linear box embedding to embed noised boxes into the same dimension as DETR queries. The content query part is the same as DAB-DETR, and we use label embedding to embed labels into content queries. After obtaining content and position queries, following Vanilla DETR, we can add the label embedding and box embedding together as DETR queries.

8 Introducing DN to Faster R-CNN for Traditional Detectors

Apart from accelerating DETR-like models, denoising training can also be used to accelerate traditional CNN detectors. We take Faster R-CNN as an example and add denoising training to it. The detection head of Faster R-CNN works in a similar way as the decoder of DETR-based models, where the major differences lie in 1) feature extraction: Faster R-CNN uses RoI pooling while DETR uses cross attention to extract features. and 2) label-assignment scheme: Faster R-CNN adopts a one-to-many label assignment (one GT object can be matched with multiple predicted objects), while DETR adopts a one-to-one label assignment (one GT object can only be matched with one predicted object). As the denoising part trains in parallel with the original matching part in detection models and is irrelevant to feature extraction schemes, denoising training can be easily applied to these traditional detectors.

Fundamentally, the idea of denoising training in DETR is to bypass the unstable label assignment and directly learn bounding box regression. Though Faster R-CNN does not have bipartite matching, it also has label assignment controlled by the IoU threshold. Therefore, denoising training can also serve as a shortcut to help learn bounding box regression without label assignment in traditional models. Therefore, we add noised boxes to the detection head of Faster R-CNN in parallel with the original boxes from the RPN. These noised boxes will directly regress the GT to improve training. Note that as Faster R-CNN does not have an initial content part, we only use box denoising training.

9 Introducing DN to Mask2Former for Segmentation Models

We also show the feasibility of adding denoising training to segmentation models such as Mask2Former . Mask2Former adopts a DETR-like architecture and proposes masked attention to extract features for segmentation tasks. More specifically, each decoder layer predicts segmentation masks, which are passed to the subsequent decoder layer as the attention mask to pool features. Therefore, following the idea of denoising training in detection models, we can add noise to the GT masks and feed them to the decoder as the attention mask. The training objective of these noised masks is to directly predict the GT mask, which bypasses the bipartite match and serves as a shortcut to directly learn mask refinement.

To verify the effectiveness of denoising training on masks, we build a simple baseline by adding simple shifting noise to the mask. Without changing the shape or size of the mask, we shift the whole GT mask on the x-axis and y-axis by a random value, which is the same as the center shifting noise as described in Sec. 4.3. This simple baseline already demonstrates the effectiveness of denoising training.

Experiment

Dataset: We show the effectiveness of DN-DETR on the challenging MS-COCO 2017 Detection task. MS-COCO is composed of 160K images with 80 categories. These images are divided into train2017 with 118K images, val2017 with 5K images, and test2017 with 41K images. In all our experiments, we train the models on train2017 and test on val2017. Following the common practice, we report the standard mean average precision (AP) result on the COCO validation dataset under different IoU thresholds and object scales. Implementation Details: We test the effectiveness of the denoising training on DAB-DETR, which is composed of a CNN backbone, multiple Transformer encoder layers, and decoder layers. We also show that denoising training can be plugged into other DETR-like models to boost performance. For example, our DN-Deformable-DETR is built upon Deformable DETR in a multi-scale setting.

We adopt several ResNet models pre-trained on ImageNet as our backbones and report our results on 44 ResNet settings: ResNet-50 (R50), ResNet-101 (R101), and their 16×-resolution extensions ResNet-50-DC5 (DC5-R50) and ResNet-101-DC5 (DC5-R101). For hyperparameters, we follow DAB-DETR to use a 66-layer Transformer encoder and a 66-layer Transformer decoder and 256256 as the hidden dimension. We add uniform noise on boxes and set the hyperparameters with respect to noise as λ1=0.4\lambda_{1}=0.4, λ2=0.4\lambda_{2}=0.4, and γ=0.2\gamma=0.2. For the learning rate scheduler, we use an initial learning rate (lr) 1×10−41\times 10^{-4} and drop lr at the 40-th epoch by multiplying 0.1 for the 50-epoch setting and at the 11-th epoch by multiplying 0.1 for the 12-epoch setting. We use the AdamW optimizer with weight decay of 1×10−41\times 10^{-4} and train our model on 8 Nvidia A100 GPUs. The batch size is 16. Unless otherwise specified, we use 5 denoising groups.

We conduct a series of experiments to demonstrate the performance improvement as shown in Table I, where we follow the basic settings in DAB-DETR without any bells and whistles in training. To compare with the state-of-the-art performance in the 12 epoch setting (the so-called 1×1\times setting in Detectron2) and the standard 50 epoch setting (most widely used in DETR-like models) in Table II and IV, we follow DAB-DETR to use 3 pattern embeddings as in Anchor DETR . All our comparisons with DAB-DETR and its variants are under exactly the same setting. DN-Deformable-DETR and DN-Deformable-DETR++: For DN-Deformable-DETR with only deformable encoder, we use 1010 denoising groups. For DN-Deformable-DETR++ with deformable attention in both encoder and decoder, we use 55 denoising groups. Note that we strictly follow Deformable DETR to use multi-scale (44 scale) features without FPN. Dynamic DETR adds FPN and more scales (55 scales) which can further boost the performance, but our performance still exceeds theirs. Faster R-CNN and Anchor DETR: We use 1010 and 55 denoising groups respectively. DINO: To test the effectiveness of denoising training in DINO, we only use our proposed DN without its proposed contrastive DN and keep all the other components in DINO. We use 55 denoising groups. Mask DINO: Mask DINO incorporates both box denoising and mask denoising. To show performance improvement over segmentation tasks, we keep the box denoising part and only remove the mask denoising to study its effectiveness. We use 55 denoising groups under this setting. Mask2Former: Mask2Former is only designed for segmentation tasks. Therefore, we only add mask denoising training in our experiments. We use 55 denoising groups under this setting.

Our proposed denoising training has been incorporated into many subsequent works and also implemented in detrex (https://github.com/IDEA-Research/detrex).

2 Denoising Training Improves Performance

To show the absolute performance improvement compared with DAB-DETR and other single-scale DETR models, we conduct a series of experiments using different backbones under the basic single-scale settings. The results are summarized in Table I.

The results show that we achieve the best results among single-scale models with all four commonly used backbones. For example, compared with our baseline DAB-DETR under exactly the same setting, we achieve +1.9 AP absolute improvement with ResNet-50. The table also shows that denoising training adds negligible parameters and computation.

3 1×1\times Setting

With denoising training, the detection task can be accelerated by a large margin. As shown in Table II, we compare our method with both a traditional detector and some DETR-like models, including DETR , Dynamic DETR , and Deformable DETR . Note that Dynamic DETR adopts a dynamic encoder, for a fair comparison, we also compare with its version without a dynamic encoder.

Under the same setting with the DC5-R50 backbone, DN-DETR can outperform DAB-DETR by +3.7 AP within 1212 epochs. Compared with other models, DN-Deformable-DETR achieves the best results in the 1212 epoch setting. It is worth noting that our DN-Deformable-DETR achieves 44.1 AP within 12 epochs with the ResNet-101 backbone, which surpasses Faster R-CNN ResNet-101 trained for 108108 epochs (9×9\times faster).

4 Extending DN to Other Detection and Segmentation Models

To further validate the effectiveness of denoising training, we extend this method to other detection and segmentation model, as shown in Table III. The experimental results indicate that denoising training is a universal training method to boost performance.

For example, we improve the DETR-like detection models significantly by 1.2−2.61.2-2.6 AP under the 12-epoch setting. The results also reveal that

Denoising training is compatible with other positional query formulations, for example, Vallina DETR with high dimensional vectors, Anchor DETR with 2D anchor points, and DAB-DETR with 4D anchor boxes.

Our method is only a training method and also compatible with other methods, for example, deformable attention , semantic-alignment , and query selection, etc.

5 Compared with State-of-Art Detectors

We also conduct experiments to compare our method with multi-scale models. The results is summarized in Table IV. Our proposed DN-Deformable-DETR achieves the best result 48.6 AP with the ResNet-50 backbone. To eliminate the performance improvement from formulating the queries of deformable DETR as anchor boxes, we further use a strong baseline DAB-Deformable-DETR without denoising training. The results show that we can still yield 1.71.7 AP absolute improvement. The performance improvement of DN-Deformable-DETR also indicates that denoising training can be integrated into other DETR-like models and improve their performance. Though it is not a fair comparison with Dynamic DETR as it includes a dynamic encoder and more scales (55 scales) with FPN, we still yield +1.4+1.4 AP improvement.

We also show the convergence curve in both single-scale and multi-scale settings in Fig. 7, where we drop the learning rate by 0.10.1 in multiple epochs in Fig. 7(b).

6 Ablation Study

We conduct a series of ablation studies with the ResNet-50 backbone trained for 50 epochs to verify the effectiveness of each component and report the results in Table V and Table VI. The results in Table V show that each component in denoising training contributes to performance improvement. Notably, without an attention mask to prevent information leakage, the performance degenerates significantly.

6.2 Effectiveness of using more denoising groups

We also analyze the influence of the number of denoising groups in our model, as shown in Table VI. The results indicate that adding more denoising groups improves performance, but the performance improvement becomes marginal as the number of denoising groups increases. Therefore, in our experiment, our default setting uses 5 denoising groups, but more denoising groups can further boost performance as well as faster convergence.

In Fig. 8, We explore the influence of noise scale. We run 2020 epochs with batch size 6464 and ResNet-50 backbone without learning rate drop. The results show that both center shifting and box scaling improve performance. But when the noise is too large, the performance drops.

6.3 Acceleration Analysis

We show how much our method can speed up training exactly in Table I. Our method achieves results comparable to the baseline with only half of the training epochs, resulting in 22x acceleration.

6.4 The training wall clock time and GFLOPs

We tested the training wall clock time and GFLOPs with 8 NVIDIA A100 GPUs as shown in Table II.

The total training time is calculated by multiplying the number of training epochs and the training time for each epoch. The training time per epoch is 51.151.1min and 57.757.7min for DAB-DETR-R50 and DN-DAB-DETR-R50, respectively. While denoising training introduces a minor training cost increase, it only needs about half the number of training epochs (2525 epochs) to achieve the same performance as DAB-DETR-R50. The practical training speedup is indeed remarkable.

7 Other tasks and future work

In addition to regular detection, our design of queries as anchor box + label makes the detection model capable of handling other tasks. For example, known object detection and known label detection. Note that the results shown in this section are just a preliminary exploration and not based on our well-trained model with the best hyper-parameters.

Known Object Detection: Assume we know a part of the objects in an image and want to predict the remaining objects. We want the known objects to help predict the unknown objects through co-occurrence relations. We did some preliminary exploration. We randomly divide the 8080 classes of MS COCO2017 into 22 parts, including known classes and unknown classes. We put objects of known classes in the denoising part and want the matching part to predict the objects of the unknown classes. We do not use an attention mask so that the matching part can get useful information from the denoising part. Our experimental results are shown in Table III. Compared with the evaluation without known boxes, the evaluation of the known object improves the performance, which indicates that co-occurrence helps the prediction of unknown boxes. Moreover, our DN-DETR trained with known objects exceeds DAB-DETR only trained on unknown classes when evaluating without known objects. This means the denoising of extra boxes from extra (known) classes also helps the performance of the unknown objects.

Known Label Detection: For each image, we assume we know all the class labels in the image without box information. Since our model has interpreted the query embedding into class label embedding, we can seamlessly utilize these known labels to detect the boxes of each class label. For each class cc in the image, we concatenate its label embedding with the indicator 11, which denotes a known label. We feed the concatenated vector into the decoder and let the decoder output all boxes of class cc. To compare with methods without known labels and detect all objects in an image, we concatenate outputs of all classes and evaluate the result as shown in Table IV. By finetuning with known labels, the detection performance can be improved in only one epoch. Within 1010 epochs of finetuning on pre-trained DN-DETR, the known label detection performance is improved to 46.646.6. This result demonstrates that given labels can significantly improve the detection performance.

7.2 Future Work

There are three potential future works to be mentioned here. One is zero-shot detection, and the other is progressive inference. Zero-shot or Open Set Detection: Since we have decoupled decoder queries as anchor boxes and class labels, pre-trained class label embeddings can be fed into the class label part of the queries. To enable zero-shot detection, one can take 8080 classes of MSCOCO as phrases and collect phrase embeddings from a pre-trained language model as the class label embedding. With the pre-trained label embedding, it is possible to train a given class detector that takes a class label embedding as input and detects objects of the given classes. In inference time, class label embeddings from unseen classes can be fed into the decoder to achieve zero-shot detection.

Progressive inference: Based on known object detection, a progressive inference method can be designed. For example, we can train a DN-DETR capable of doing known object detection. In inference time, we let the detector predict objects, and then, we can choose the objects with the highest score and treat them as known objects to do known object detection. For each step of prediction, we choose objects with the highest score and add them to the known box set. After repeating for many times, we get the final prediction.

Classification before detection: As shown in Table IV, given labels can significantly improve the detection performance. Therefore, one potential future work is to add a muli-label classification network to provide labels and feed them to DN-DETR, which may help improve detection performance.

Conclusion

In this paper, we have analyzed the reason for the slow convergence of DETR training lying in the unstable bipartite matching and proposed a novel denoising training method to address this problem. Based on this analysis, we proposed DN-DETR by integrating denoising training into DAB-DETR to test its effectiveness. DN-DETR specifies the decoder embedding as label embedding and introduces denoising training for both boxes and labels. We also added denoising training to Deformable DETR to show its generality. The results show that denoising training significantly accelerates convergence and improves performance, leading to the best results in the 1x (12 epochs) setting with both ResNet-50 and ResNet-101 as the backbone. This study shows that denoising training can be easily integrated into DETR-like models as a general training method with only a small training cost overhead and bring in a remarkable improvement in terms of both training convergence and detection performance. Limitations: In this work, the added noises are simply sampled from a uniform distribution. We have not explored more complex noising schemes and leave these for future work. Reconstructing noised data achieves great success in unsupervised learning and diffusion models. This work is an initial step to apply it to object detection. In the future, we will explore how to pre-train detectors on weakly labeled data with unsupervised learning techniques and explore applying other denoising training schemes in detection models.

References